From strategy-free scaffolds to specialist harnesses
Most agent stacks today shove everything into context: high-level “what’s the plan?”, low-level “which tool now?”, retry policies, and book-keeping all live as instructions and scratchpad around the model.
That makes every task an island. The agent repeatedly *re-derives* the same control decisions from scratch, even when the task stream is homogeneous.
Li et al.’s *Growing Harness* paper asks a different question: can you grow the harness—the code that orchestrates the model and tools—directly from task feedback, and then reuse it across tasks and model sizes? The idea is to:
- Fix model and tool interfaces.
- Start with a strategy-free scaffold that has no task-solving controller.
- Use failures, traces, and an optimizer to iteratively edit the harness code, not the prompts.
- Gate those edits so you don’t regress on past capabilities.
Over time, “control structure” moves from per-task prompts into a shared, low-cost program that executes around the model 1.
The result is a reusable specialist agent: same harness, smaller models still succeed, with far fewer calls and much lower inference cost on web tasks like BrowseComp-Plus and WebArena-Verified 1.
This article focuses on what this means for people building agent systems today: where the paradigm plugs into your stack, what’s actually being learned, and what trade-offs show up in the numbers that *are* documented.
---
What Growing Harness actually is
The key abstraction is:
Harness: executable code that orchestrates an LLM and tools to solve a class of tasks.
By contrast:
Strategy-free scaffold: a shell that exposes fixed interfaces to the model and tools but does not encode a task-solving controller 1.
Growing Harness is a failure-guided training paradigm that turns that scaffold into a working harness by repeatedly:
- Running tasks through the current harness.
- Capturing function-level execution traces.
- Using failures and traces to localize where the harness logic went wrong.
- Jointly repairing a *window of failures* via an optimizer.
- Applying a success-first held-out gate that rolls back any repair sequence that harms previous capabilities 1.
Crucially, all accepted edits “accumulate in one shared harness, allowing its control structure to emerge from task feedback” 1.
So instead of:
- “Prompt the LLM to decide what to do next” inside every context, you:
- Incrementally build a program that *encodes* those decisions once, and then run it for every future task.
---
Step-by-step: how the harness grows
The paper gives only a high-level description, but the control flow can be mapped to something like this, in conceptual terms:
1. Start from a strategy-free scaffold
You define:
- Model interface: how your harness can call an LLM (e.g., a function that takes messages and returns completions).
- Tool interfaces: how to invoke browsing, clicking, form-filling, etc., on benchmarks like BrowseComp-Plus and WebArena-Verified 1.
But you *do not* embed per-task strategy in code. The scaffold “exposes fixed model and tool interfaces but encodes no task-solving controller” 1.
In a typical stack, this would correspond to:
- The shell that wires up your environment.
- A run loop that can call tools and the LLM.
- No hand-written policies like “if page contains X, do Y”.
2. Execute tasks and record traces
For each task:
- The current harness runs, making calls into:
- The LLM (for “task-specific semantic reasoning”).
- Tools (for actions in the environment).
- A function-level execution trace is recorded 1:
- Which functions executed.
- Their arguments and results.
- Which model/tool calls they triggered.
This gives you a structured account of how control flowed through the harness code on that run.
3. Localize failures to a bounded code surface
When a task fails, the trace lets you localize the failure:
- The error is attributed to a small region of harness code—a bounded code surface in the authors’ terminology 1.
- That bounded region defines where edits are allowed for this “repair window”.
The important property for builders: failures don’t cause global, undirected edits; they’re scoped to specific functions or small modules where the harness logic is implicated.
4. Jointly repair a window of failures
Rather than patching each failed task in isolation, Growing Harness uses:
“an optimizer [that] repairs a window of failures jointly” 1.
Interpretation for practice:
- Collect a batch of failures whose traces point to overlapping code regions.
- Propose and evaluate code edits that satisfy *multiple* such tasks at once.
This joint repair step encourages more general harness logic, as changes that only fix one weird edge-case but break others won’t survive.
5. Guard with a success-first held-out gate
To avoid catastrophic regressions, Growing Harness maintains a success-first held-out gate which:
- Evaluates repair sequences on a held-out set of tasks where the harness previously succeeded.
- Rolls back any edit that “harm[s] prior capability” 1.
So the harness’s competence is monotone non-decreasing on that held-out set: it can expand to cover more tasks, but must preserve what it already solved, at least under this gate’s test battery.
This is a pattern production teams will recognize from online policy updates or A/B gating: don’t ship an edit unless you can verify you didn’t regress on known workloads.
6. Accumulate accepted edits in one shared harness
The loop repeats, and:
“Accepted edits accumulate in one shared harness, allowing its control structure to emerge from task feedback” 1.
You end up with:
- A *single* program that contains the agent’s control policies.
- Persistent, reusable logic that solves common subtasks without re-eliciting them from an LLM prompt.
- A harness that can now be paired with different deployment models (4B to 120B parameters in the experiments) without re-training it 1.
---
How this fits in an agent stack
Given a modern agent stack, where does Growing Harness live?
From the paper’s description, you can view your system as three layers:
- Environment + tools
- Browsers, click/scroll forms, APIs, structured data tools.
- In the paper’s experiments, this is instantiated via web environments like BrowseComp-Plus and WebArena-Verified 1.
- Harness (grown)
- Orchestrates tool use, decides which subgoals to pursue, when to stop, how to recover from errors.
- Encodes control structure in executable code.
- Calls the LLM only for task-specific semantic reasoning 1.
- LLM(s)
- 4B–120B parameter models in the paper’s deployments 1.
- Treated as black boxes accessed via a fixed interface defined in the scaffold.
Growing Harness affects layer (2) exclusively:
- You don’t change the environment.
- You don’t fine-tune the LLM.
- You *evolve* the harness code around it, using execution traces and task-level feedback.
This is appealing for production because it decouples:
- Model choice (which foundation model, what size).
- Harness competence (how good your agent is at controlling the environment).
The paper’s results explicitly showcase that decoupling: on WebArena-Verified, the same grown harness maintains 44.7–45.3% success across model scales, while a more conventional Tool-Calling agent drops to 6.7% success at 4B 1.
---
Why it matters: costs, robustness, and small models
The authors report three headline kinds of results.
1. Higher success with less context work
Across BrowseComp-Plus and WebArena-Verified, and across three deployment models between 4B and 120B parameters, Growing Harness:
- Achieves the highest mean success in five out of six benchmark–model combinations, and
- “Trails the best mean by 0.7 percentage points in the sixth” 1.
That’s with a harness that is learned, not hand-coded.
If you’re currently manually evolving your orchestration layer, this gives evidence that task feedback and automated repair can reach or exceed the performance of bespoke Tool-Calling agents on realistic web tasks.
2. Massive reductions in LLM calls and cost
The central economic argument:
“Relative to a Tool-Calling agent, [Growing Harness] reduces LLM calls by 76.0–91.8% and deployed-agent inference cost by 74.4–98.6%” 1.
Mechanistically, that’s what you’d expect when:
- Repeated control logic is compiled into code,
- Only task-specific semantic interpretations still hit the LLM.
For a production system, this shifts the optimization frontier:
- You can hit similar success rates with far fewer tokens.
- Your harness becomes an asset that amortizes over every task and every model deployment.
3. Small models stay useful when the harness is strong
On WebArena-Verified:
- Growing Harness maintains 44.7–45.3% success “across model scales” 1.
- A Tool-Calling baseline plummets to 6.7% success with the 4B model 1.
From a builder’s perspective, this is big:
- You can deploy much smaller, cheaper models in production while preserving a large fraction of capability, provided your harness is competent.
- Harness quality acts as a *force multiplier* on model size. A better harness compresses “agentic overhead” out of the model.
The paper also reports that ablations show:
“trace-local edits, joint repair, and gate-based rollback each improve final success” 1.
So all three mechanisms—localization, batch repair, gated rollback—matter for reaching the reported performance.
---
How to adapt the paradigm in your own stack
The paper doesn’t give implementation details, but the high-level recipe suggests some concrete design moves if you want to borrow the idea:
- Define a strategy-free scaffold
- Separate “model calls” from “control logic” as explicit functions.
- Avoid writing bespoke per-task if/else logic; keep the harness minimal at first.
- Add function-level tracing
- Instrument your harness so every run produces a structured trace of function entries/exits and tool/model calls.
- Ensure you can tie a failed task to specific functions.
- Localize repair surfaces
- When tasks fail, identify the minimal set of functions that could have caused the failure via their traces.
- Restrict edits to those regions to avoid destabilizing the entire harness.
- Batch failures into windows
- Instead of patching one task at a time, collect a “window” of failed tasks whose traces touch similar surfaces.
- Optimize edits to simultaneously resolve as many of those as possible (the paper’s “window of failures jointly” 1).
- Gate changes against a success-first suite
- Maintain a held-out suite of tasks where your harness currently works.
- Reject any code edits that cause regressions on this suite (the “success-first held-out gate” 1).
- Share the harness across models
- Treat the harness as model-agnostic.
- Evaluate it with both small and large models to check whether control logic is consistently reusable, as in the paper’s 4B–120B experiments 1.
Because the paper doesn’t specify the optimizer or the repair representation, those choices are open: you could imagine anything from search over code edits to LLM-assisted program synthesis, but that is not described in the source.
---
What to watch next
Based on the documented results, questions worth tracking (or experimenting with yourself) include:
- Transfer across task distributions: the paper evaluates on BrowseComp-Plus and WebArena-Verified; it doesn’t document how a harness grown on one domain transfers to another 1.
- Stability under distribution shift: the success-first gate protects a fixed held-out set; how the harness behaves when the web environment changes is not covered.
- Interaction with model training: Growing Harness leaves the LLM weights fixed; combining this with model fine-tuning is an open direction not discussed in the paper.
- Tool evolution: the strategy-free scaffold assumes fixed tool interfaces 1; how often you can safely change tools while reusing a grown harness isn’t documented.
For now, the main practical lesson is clear and supported by the data:
Moving recurring control out of per-task context and into persistent code can both improve success rates and cut LLM call volume by large margins, while keeping small models viable 1.
If you’re building web or API agents at scale, that’s an attractive target.
---
What is not documented
The sources do not establish:
- The exact programming language, runtime, or framework used to implement the harness.
- The specific optimizer or algorithm used for code repair (e.g., gradient-based, search-based, LLM-based).
- The precise definitions and sizes of “windows” of failures or the criteria for selecting them.
- Details of the function-level trace schema or data structures.
- Any hyperparameters, training curves, or timings for the harness growth process.
- Any qualitative examples of learned harness code or emergent control patterns.
- How well a harness trained on one benchmark transfers to different domains or tools.