Co-Evolving LLM Evaluators and Policies with DynamicRubric

DynamicRubric co-evolves evaluators and policies, improving LLM post-training by closing score gaps. Deployed in WeChat Search with millions of daily requests.

viernes, 24 de julio de 2026 • 3 min read • Q2BSTUDIO Team

Optimización de LLM con rúbricas dinámicas

Post-training large language models (LLMs) with evaluator feedback on policy-induced samples has become a key mechanism for improving performance. However, as policies advance, generated responses tend to converge in quality, creating a bottleneck: relative evaluator score gaps collapse, yielding weak or misleading supervision signals. This dilemma has motivated adaptive frameworks like DynamicRubric, which proposes a co-evolution between evaluators and policies by generating weighted binary rubric items for each candidate set. The result is a significant improvement in the evaluator's ability to discern nuances and provide stronger supervision, even outperforming massive reward models or static rubrics.

From a technical perspective, DynamicRubric addresses the problem through a probability allocation view: the directional gain of shifting probability mass from one response to another is exactly the score gap between them. Therefore, maintaining meaningful gaps is essential for guiding policy updates. Instead of relying on fixed evaluators, DynamicRubric dynamically adjusts evaluation criteria based on the current response set, generating weighted binary rubric items that capture subtle differences. Experiments with 8B-parameter models show that this approach not only improves evaluator performance but also provides stronger policy supervision than baselines using a 70B reward model or a 235B static rubric generator. Optimized policies also show gains in verifiable reasoning and coding tasks, and have been deployed in production for WeChat Search's AI answering scenario, handling tens of millions of daily requests and improving key online metrics.

This principle—that evaluators should evolve with the policies they supervise—has profound implications for enterprise AI development. Companies integrating LLMs into their processes need evaluation systems that dynamically adapt to continuous model improvement. This is where the expertise of companies like Q2BSTUDIO becomes crucial. With its focus on custom software, Q2BSTUDIO helps organizations design and implement AI solutions that incorporate co-evolution mechanisms similar to DynamicRubric. By combining artificial intelligence, cybersecurity, cloud AWS/Azure, and BI with Power BI, the company creates robust ecosystems where evaluators automatically adjust to changing model behavior, ensuring precise supervision and continuous optimization.

For instance, in an LLM-based recommendation system, a static evaluator may lose sensitivity when responses collectively improve. DynamicRubric, or an equivalent approach implemented by Q2BSTUDIO, generates custom rubrics that highlight minimal yet critical differences, such as technical accuracy or business context alignment. This translates into more refined policies that enhance end-user experience. Moreover, integration with cloud services like AWS or Azure enables efficient scaling, while cybersecurity practices ensure data and model decisions are protected. BI tools, especially Power BI, allow real-time visualization of evaluator and policy performance metrics, facilitating informed decision-making.

The use of AI agents is another area where this co-evolution makes a difference. Autonomous agents that make decisions based on LLMs require continuous supervision; an evaluator that adapts to evolving policies can correct deviations and reinforce desirable behaviors. Q2BSTUDIO offers AI agent solutions that incorporate these principles, creating more reliable and efficient systems. Ultimately, DynamicRubric illustrates that the key to effective supervision is not larger evaluators, but more adaptive ones. For companies seeking to lead in the AI era, investing in co-evolution mechanisms—from custom software to integrated cloud platforms—that Q2BSTUDIO can develop is a strategy that maximizes LLM performance and value.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.