Hermes

Learning Contextual Reasoning Unlocks Test-Time Scaling

Xinyu Li1,†,‡ Mononito Goswami2,† Hao Liu2 Nikos Kanakaris2 Langlin Huang3,‡ Prithwish Jana4,‡ Patrick Blöbaum2 Purak Jain2

1Carnegie Mellon University 2AWS AI Labs 3Washington University in St. Louis 4Georgia Institute of Technology

†Equal contribution  ·  ‡Work done during an internship at AWS

TL;DR

Using test-time compute across many context windows requires deciding how to allocate fresh contexts and what information to carry between them. We call this ability contextual reasoning. Existing harnesses prescribe these decisions; Hermes shifts them to the model. Capable models exploit this flexibility to scale with extra compute, while small open models initially struggle. Hermes-Learn closes the gap, inducing adaptive strategies that vary with the problem and the progress of reasoning. The gains generalise across benchmarks and models, extrapolate beyond the compute seen in training, and transfer to other test-time scaling methods.

After Hermes-Learn, Qwen3-4B overturns a wrong majority by deciding what each fresh context should check

The problem · AIME 2026, problem 4

Correct answer

The model works in one 8K-token context. It sees the problem and the short reports that come back. A subagent is the same model in a fresh 8K-token context. It sees only the task it was given.

Out of the box, a small model gets worse with Hermes. After Hermes-Learn, it beats one long context with the same token budget.

Qwen3-4B on 238 AIME and HMMT problems: % solved, averaged over 16 attempts per problem

37.1 → 16.2

Given Hermes as it is, the untrained model does worse than in one context. It does not yet know what to use the extra contexts for.

47.3 vs 41.8

After Hermes-Learn it beats one 56K-token context, the same total budget, by 5.5 points, although none of its contexts is longer than 8K tokens.

1

Extra context only helps if the model knows what to do with it

“The capacity of the human mind for formulating and solving complex problems is very small compared with the size of the problems …”

Herbert A. Simon, Models of Man

Test-time scaling improves a model by spending more compute at inference. Recent systems run thousands of agents, each in its own context, and a harness decides how those contexts are allocated and what is carried between them. That splits reasoning into two different skills.

Mathematical reasoning

How to make progress with the information available inside one context.

Contextual reasoning

How to allocate, coordinate and reuse separate contexts as reasoning unfolds: which fresh context should explore an alternative, verify earlier work or continue a line of reasoning, and what to carry into it.

Most harnesses make those decisions for the model. Reasoning Cache always compresses and continues; recursive agents always decompose. But the best use of a new context changes as a solution develops. A hard proof may first need several approaches in parallel, then a check of one promising intermediate result, then a clean continuation.

The contextual reasoning gap

We gave three models the same Hermes harness: up to seven fresh 8K-token contexts instead of one. Claude Sonnet 5 and DeepSeek-V3.2 turn the extra contexts into accuracy. Qwen3-4B, the small open model we train, loses more than half its accuracy. It solves 37% of the problems in a single context, so what it lacks is not mathematics but knowing what to do with the extra contexts.

The same harness helps capable models and hurts a small one

% of 238 AIME and HMMT problems solved, averaged over 16 attempts per problem. With Hermes, each model gets up to seven 8K-token contexts and decides how to use them.

2

Hermes: the model decides how to use each fresh context

Hermes is built to be simple (two generic mechanisms), flexible (neither mechanism prescribes a form of computation) and configurable (how much the harness steers is a parameter).

parallel

Delegation

The agent writes TASK: lines, i.e. calls launch_subagent(task). Each task runs on the same model in a fresh context, sees only the task text, and returns a compact report (Scope · Finding · Answer) as the agent's next input. Subagents can delegate too, up to depth D.

sequential

Digestion

At the end of a window the agent compresses its progress into a digest. If reasoning continues, a fresh window starts from the original problem plus the digest instead of the full history, for up to W windows.

    Language model (one context) Problem, answer, digest Steering message Harness
    Press Play or pick a phase. Hover over any box for details.

    Inside a window the model takes up to T steps. At every step but the last it may launch up to B subagents; at the last it must answer. Together with the delegation depth D, the number of windows W and the steering level, these knobs set both the compute budget and how much of each decision belongs to the model.

    Configurations: breadth, depth and length

    B and D widen the tree in parallel; W extends it in sequence. Every box below is one 8K-token context. Change a knob to see how the tree, and the compute it can use, grows.

    T steps per window3 (fixed: 2 delegation steps + return)
    B subagents per step
    3
    D delegation depth
    1
    W context windows
    1

    7

    context windows, at most

    56K

    tokens of context, at most

    Steering: who picks the strategy?

    Each delegation step uses one or more strategies, in two categories:

    • Search makes progress. Solve the whole problem again with no set approach (Same), with different named approaches (Different), or split it into subproblems (Split).
    • Verify checks the current answer: directly (Answer), through the steps behind it (Steps), or by finding why candidates disagree (Disagreement).

    The steering level decides who picks. At L3 the harness fixes the strategy, as most existing harnesses do. At L2 it fixes only the category. At L1 the model picks both.

      fixed by the harness chosen by the model not offered

      Watch it reason

      Six rollouts of the trained Qwen3-4B under different harnesses and configurations. Each one plays on its own; use Back and Next to go at your own pace. The takeaway appears at the end.

      root agent subagent (fresh 8K context)
      The model's words are quoted and shortened with “…”. Each subagent's own reasoning, thousands of tokens, is not shown. ✓ and ✗ compare an answer with the correct one.

      3

      Hermes-Learn: teaching a 4B model contextual reasoning

      Hermes-Learn teaches the model what to delegate, in two stages: first imitate strong models, then learn from its own outcomes. Digestion is not trained; the model uses its existing ability to summarise.

      Stage I

      Learn from teachers

      1. Two strong models, Claude Sonnet 5 and DeepSeek-V3.2, solve 1,000 problems under Hermes, 8 runs per problem.
      2. Keep the runs that delegate at least once and reach the correct answer, then balance search strategies and teachers.
      3. Fine-tune the 4B model on the root agent's own words.

      Stage II

      Learn from outcomes

      1. The fine-tuned model solves each problem 8 times under Hermes.
      2. A run scores 1 if its final answer is correct and 0 if not.
      3. Runs that beat their group's average become more likely, the rest less. Subagents are not trained.

      One group: 8 runs of the same problem · click a run to flip its outcome

      A curriculum: from guided to free

      Stage II first trains only under L2, where the harness fixes each step's category. After the first pass over the data, three quarters of training moves to L1, where the model picks everything, and a quarter stays on L2.

      Share of stage II training under each steering level

      Synthetic problems that reward different strategies

      Competition problems rarely require a particular strategy. So we build 500 problems from seeds with known answers, each type favouring a different strategy, and reword them so the structure is not signposted.

      Coloured letters mark the intermediate values that decide the answer. A matching set of 88 problems, built the same way from AIME and HMMT, is used for evaluation.

      RQ1

      Extra context goes from liability to asset

      On every benchmark, the untrained model does worse with Hermes than in one context, and Hermes-Learn turns that around. Its single-context accuracy barely moves, so the gain comes from using contexts, not from better mathematics. After training, Hermes beats one 56K-token context on all four math benchmarks.

      RQ2

      It learns strategies that depend on the problem and on progress

      An LLM judge that sees only the tasks, never the harness or the strategy menu, labels the strategy of every delegation step.

      RQ3

      What matters in training: teachers, curriculum, and new problem structure

      (a) Which teacher for SFT?

      The stronger Hermes user is not the better teacher. DeepSeek-V3.2 trains a far better student than Sonnet 5 (39.3 vs 30.1). Using both is best, mostly because it gives more demonstrations.

      After stage I only. % solved on the 238 AIME and HMMT problems, mean ± 1 sd over 3 seeds.

      (b) Which RL recipe?

      New problem structure beats more data. The curriculum plus 500 synthetic problems scores best (47.3); 500 rephrased problems give 45.6. The recipes are within about 2 sd of each other.

      After both stages. Same benchmark, mean ± 1 sd over 3 seeds. Mix and Curriculum use the same L1 and L2 data and differ only in order.

      RQ4

      It generalises beyond its training setting

      Training used one setting: Qwen3-4B with B = 3, D = 1, W = 1. We test three ways of leaving it.

      (a) More inference compute than in training

      More windows, more subagents and deeper delegation all help, although training used none of them. Depth 2, where subagents delegate too, reaches 55.6.

      % solved on the 238 AIME and HMMT problems under L1. Each row changes one knob from the training setting.

      (b) Other test-time scaling methods

      Hermes lifts other methods. The DeepSeekMath Agent goes from 15.1 to 52.5 when run inside Hermes; Recursive Self-Aggregation from 50.5 to 55.0.

      % solved on the 238 AIME and HMMT problems.

      (c) A different model family and size

      The same recipe works on another model. Olmo-3-7B drops from 30.5 to 4.8 with Hermes; Hermes-Learn takes it to 40.6.

      Olmo-3-7B-Instruct, % solved on the 238 AIME and HMMT problems, one training run per row.

      Citation

      @article{li2026hermes,
        title   = {Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling},
        author  = {Li, Xinyu and Goswami, Mononito and Liu, Hao and Kanakaris, Nikos and
                   Huang, Langlin and Jana, Prithwish and Bl{\"o}baum, Patrick and Jain, Purak},
        journal = {arXiv preprint},
        year    = {2026}
      }