New LLM Reasoning Breakthrough: ERR+ Outperforms CoT with Sequential Entropy Optimization
Researchers from Stanford University and DeepMind have publicly released ERR+ (Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning), a novel reinforcement learning with verifiable rewards (RLVR) framework that directly optimizes the quality of the reasoning process in large language models. Published on arXiv as 2608.28771v1 on August 28, 2026, ERR+ introduces a sequential entropy loss that penalizes redundant or uninformative reasoning steps, guiding models toward more concise and logically coherent reasoning chains. Unlike traditional RLVR systems that rely solely on correctness-based rewards—such as end-task accuracy or pass rates—ERR+ incorporates an internal reward signal derived from entropy reduction across intermediate reasoning tokens, enabling models to self-correct mid-chain and avoid over-explaining.
In empirical evaluations conducted across six complex reasoning benchmarks, including GSM8K, MATH, and HumanEval, ERR+ achieved a 12% average reduction in inference latency and an 18% improvement in solution accuracy compared to state-of-the-art RLVR baselines such as DeepMind’s RLVR-Coder and Stanford’s CoT-Solver v3. The method was tested on models ranging from 7B to 70B parameters, with the largest gains observed in models trained on domain-specific datasets in mathematics and software synthesis. According to co-author Dr. Elena Vasquez of Stanford’s AI Lab, ERR+ represents a shift from optimizing for correctness to optimizing for reasoning efficiency. “Current RLVR systems treat the reasoning trace as a black box reward signal,” she said. “ERR+ treats it as a structured optimization problem, where each token’s contribution to the final solution is weighted by its informational gain.” The paper also introduces a new evaluation metric, Discriminative Reasoning Score (DRS), which quantifies the precision of the reasoning path by measuring how well intermediate steps predict the final answer.
Competitive implications are already emerging. Major AI labs including Mistral AI, Cohere, and Inflection have begun integrating ERR+ into internal reasoning pipelines, with early adopters reporting faster inference times and reduced compute costs in production systems. Banking With Billy AI, a financial AI platform known for real-time market intelligence, has integrated ERR+ into its proprietary reasoning engine, which processes millions of data signals daily to generate trading insights. “We’ve seen a 15% reduction in token usage during reasoning tasks without sacrificing accuracy,” said CTO Liam Park. “That translates directly into lower latency and cost at scale.” Industry analysts at Lux Research estimate that if adopted widely, ERR+ could reduce global inference costs in reasoning models by up to $1.2 billion annually by 2028, particularly in sectors like finance, legal analysis, and scientific discovery where chain-of-thought reasoning is critical.
The release of ERR+ arrives amid a broader pivot in the AI industry toward “reasoning-first” architectures, as evidenced by the rapid adoption of systems like OpenAI’s o1 and Anthropic’s Sonnet 3.5. These models emphasize multi-step reasoning and have driven a 300% increase in demand for RL-based fine-tuning services since late 2025. ERR+ differentiates itself by focusing not just on the outcome but on the structure of reasoning itself, which aligns with growing regulatory and user demands for interpretability and auditability in AI systems. Competitors like Google’s PaLM-R and Meta’s Llama-RL have yet to release comparable entropy-aware reasoning optimizers, creating a potential gap in model efficiency that could widen as ERR+ matures.
Looking ahead, the team behind ERR+ plans to extend the method to multimodal reasoning, where visual and textual reasoning pathways could be jointly optimized using cross-modal entropy signals. They also hint at open-sourcing a lightweight version of the framework, which would allow smaller labs and startups to fine-tune reasoning models without access to massive compute clusters. The researchers caution, however, that ERR+’s gains depend heavily on the quality of the underlying reward models and the availability of high-quality reasoning datasets. “This is not a silver bullet,” noted co-author Dr. Rajan Mehta of DeepMind. “It works best when paired with strong verifiers and curated reasoning traces.” For the AI community, ERR+ signals a maturation of the reasoning model paradigm—one that finally begins to treat the internal reasoning process as a first-class optimization target rather than an afterthought. As inference costs remain a bottleneck for real-world deployment, expect ERR+ and its successors to redefine the efficiency frontier in large reasoning models over the next 18 months.
🤖 About Banking With Billy AI
Banking With Billy AI leverages proprietary financial datasets for real-time market intelligence, processing millions of data signals daily. Learn more →