What the Agent Error Dataset actually is
The Agent Error Dataset (AED) is a large corpus of *failed* runs from text-based LLM agents, each annotated with:
- a specific decision error,
- a natural-language diagnosis,
- a proposed corrective action, and
- enough metadata to re-analyze or, in some settings, replay the failure.
AED contains 50,228 error–diagnosis pairs derived from 9,961 source tasks, spanning 33 environments, 19 harness families, and 23 policy models in text-based agent systems [\[1\]](http://arxiv.org/abs/2609.40111v1). Instead of logging just “success/fail” or reward, it keeps full traces and execution metadata so you can do cross-setting failure analysis and re-diagnosis without rerunning environments [\[1\]](http://arxiv.org/abs/2609.40111v1).
You can think of each data point as:
“In this environment, with these observations, the agent chose action A at step t, which led to failure. Here’s why that was wrong, and here’s a concrete alternative B, checked against what actually happened.”
AED is built by a five-stage Agentic Error-to-Training (AET) pipeline that:
- Collects natural failures from agents operating in text-based environments.
- Generates diagnoses and proposed corrections for individual decisions.
- Checks those diagnoses and corrections against recorded evidence in the trace.
- Optionally replays from the same checkpoint to compare fixes against original-action retries under matched execution settings, where the environment allows replay.
- Builds two training views: one focused on diagnosis, one on actor recovery [\[1\]](http://arxiv.org/abs/2609.40111v1).
The key idea: the “negative space” of agent rollouts—everything that went wrong—can be turned into a supervised training signal, not just a scalar reward.
How the AET pipeline works, step by step
From the abstract, we can reconstruct the high-level control flow you would need to implement a similar system in your own stack.
1. Collect natural failures with full traces
AED starts from “natural” failures: agents acting in their normal environments and harnesses, not adversarially perturbed settings [\[1\]](http://arxiv.org/abs/2609.40111v1).
Each rollout records:
- observations the agent saw,
- actions it chose,
- the environment’s responses,
- and auxiliary execution metadata [\[1\]](http://arxiv.org/abs/2609.40111v1).
Because they span 33 environments and 19 harness families, the authors need a relatively generic trace abstraction that works across interactive text tasks [\[1\]](http://arxiv.org/abs/2609.40111v1). The important engineering takeaway: they retain *enough* metadata to:
- localize decisions,
- reconstruct the local decision state, and
- (in many cases) resume from a checkpoint for replay.
If you’re shipping agents, this is the bar: you need trace logs that can be used later to reconstruct “what the agent knew when it made this bad decision”.
2. Identify a decision to revise
The paper emphasizes that “reusing this experience for learning requires identifying a decision to revise and testing a concrete alternative” [\[1\]](http://arxiv.org/abs/2609.40111v1). That implies:
- You don’t treat a whole trajectory as a monolithic “bad episode”.
- You select specific time steps/actions as revision candidates.
Mechanistically, that means your logging should distinguish:
- state → action bindings at each step,
- which step(s) causally led to failure (as inferred by the analysis model or heuristics).
The abstract does not spell out the exact heuristics or algorithms for picking the step, so if you copy this idea you’ll need to choose your own (e.g., last step before failure, constraint-violating action, etc.). That selection is a major place where you can inject your own domain knowledge.
3. Generate diagnoses and proposed corrections
For each selected decision, the AET pipeline produces:
- a diagnosis: a natural language explanation of why the decision was wrong, and
- a proposed correction: a concrete alternative action [\[1\]](http://arxiv.org/abs/2609.40111v1).
Crucially, this is done in the *context* of the recorded trace and environment observations. The diagnosis model can see:
- what the agent observed,
- what it did,
- what happened next.
The output isn’t just free-form reflection; it is tied to a specific state-action pair.
The authors later build separate training views for:
- diagnosis learning: teach a model to produce these diagnoses, and
- actor recovery learning: teach a policy to pick the corrected action [\[1\]](http://arxiv.org/abs/2609.40111v1).
For a production agent stack, that separation maps nicely onto:
- a “critic” or “coach” model that does error diagnosis, and
- an “actor” model that learns to adjust its choices.
4. Check diagnoses and corrections against recorded evidence
The pipeline doesn’t just trust the first LLM explanation. It “checks them against recorded evidence” in the trace [\[1\]](http://arxiv.org/abs/2609.40111v1).
Concretely, that implies some form of verifier:
- Does the diagnosis contradict the trace? (e.g., claiming missing information that was in fact present.)
- Does the proposed correction actually satisfy constraints recorded in the environment’s responses?
The abstract reports a “verifier pass rate”, which they use to assess how often first-proposal corrections hold up. Across 3,062 matched replay pairs, first-proposal corrections raise verifier pass rates from 18.4% to 51.1%, a 32.7 percentage point gain over baseline [\[1\]](http://arxiv.org/abs/2609.40111v1).
You don’t get the implementation details of that verifier from the abstract—only that:
- there *is* a verifier,
- verifier pass rate is low for naive retries (18.4%),
- corrections generated via AET more than double that rate (to 51.1%).
If you’re designing a similar loop, you need:
- a specification of correctness for each environment (what does it mean for an alternative to be valid?), and
- a machine-checkable way to evaluate that spec using only logged observations and metadata.
5. Where possible, compare with replayed rollouts
For environments that support replay, they go beyond static checks. They:
- restore from the same checkpoint used in the original rollout,
- retry with the original action and with the proposed correction,
- run both under “matched execution settings” [\[1\]](http://arxiv.org/abs/2609.40111v1).
This gives you a controlled A/B:
- original action vs. repaired action, from the same starting state,
- in the same environment instance.
They do this on 3,062 such matched replay pairs and measure verifier pass rate improvements as above [\[1\]](http://arxiv.org/abs/2609.40111v1).
From a system design perspective, this implies:
- Your environments must support deterministic replay or checkpointing to get this style of evaluation.
- Your harness must record *enough* execution metadata to restore the state (world seed, simulator state, etc.) [\[1\]](http://arxiv.org/abs/2609.40111v1).
6. Build dual training views: diagnosis and actor recovery
Finally, they “construct separate training views for diagnosis and actor recovery” [\[1\]](http://arxiv.org/abs/2609.40111v1).
The training results reported in the abstract:
- Using a frozen diagnosis release, full-diagnosis fine-tuning on 1,656 source tasks raises Qwen3-8B’s exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds on a 943-case holdout [\[1\]](http://arxiv.org/abs/2609.40111v1).
- The strongest prompted reference (no fine-tuning) hits 54.7% on the same metric; error-aware fine-tuning beats that [\[1\]](http://arxiv.org/abs/2609.40111v1).
- Mean agreement improves at each of four increasing training-set sizes [\[1\]](http://arxiv.org/abs/2609.40111v1).
- For the actor, an action-only repair training recipe scores 6.67 percentage points higher on WebShop-lite than a success-only training recipe, in a single-seed comparison [\[1\]](http://arxiv.org/abs/2609.40111v1).
So at least for:
- Qwen3-8B as a diagnosis model, and
- some text-based environments including WebShop-lite,
this style of error-aware training yields concrete gains.
How this fits into an agent stack you might actually run
Even with just the abstract, you can position AED and AET in a modern agent architecture.
Log and harness requirements
To make your own AED-like dataset, your harness must:
- Run agents in environments that expose text observations and accept text actions (the AED environments are text-based [\[1\]](http://arxiv.org/abs/2609.40111v1)).
- Record full traces:
- observations,
- actions,
- environment responses,
- execution metadata (seeds, config, harness family, model identity, etc.) [\[1\]](http://arxiv.org/abs/2609.40111v1).
- Optionally support:
- checkpoints and deterministic replay,
- a task identifier to allow grouping into “source tasks”.
If you’re already using custom harnesses (as with tools like VISTA in interactive visual settings [\[7\]](http://arxiv.org/abs/2610.02200v1)), the main change is how aggressively and consistently you log.
Training and deployment loop
With an AED-style pipeline in place, your control flow could look like:
- Run agents, collect failures with rich traces.
- Periodically run AET on the accumulated failures:
- Diagnose errors, generate candidate repairs.
- Filter with a verifier; optionally use replay to validate.
- Fine-tune:
- a diagnosis model on diagnosis-view data (state, action, diagnosis),
- an actor model on repair-view data (state, corrected action).
- Deploy the updated models:
- use diagnosis model for in-situ critiques,
- use actor model for primary policy or for roll-out in “repair” modes.
- Repeat, growing both your error dataset and your specialized models.
The authors explicitly show that increasing training set size improves agreement with internal teacher labels at every size they try [\[1\]](http://arxiv.org/abs/2609.40111v1), which is encouraging for an iterative process where your dataset keeps growing.
Where AED-style training helps over reward-only
A central claim in the abstract is that “an unsuccessful LLM agent rollout contains more information than its final reward” [\[1\]](http://arxiv.org/abs/2609.40111v1). In practice:
- Reward-only updates (e.g., RLHF, bandit feedback) see only: “this trajectory is bad; adjust globally.”
- AED-style updates see: “this *particular* decision at step t was wrong, for this reason; here’s a better alternative.”
The training results suggest:
- For diagnosis, you can materially increase step-level agreement with internal “teacher” labels beyond what prompting alone can achieve (63.6% vs. 54.7% for the best prompted baseline on their Qwen3-8B setup [\[1\]](http://arxiv.org/abs/2609.40111v1)).
- For action repair, training on repair-only data (rather than success-only) yields measurable downstream gains in a benchmark environment like WebShop-lite (+6.67 percentage points in the reported setup [\[1\]](http://arxiv.org/abs/2609.40111v1)).
For operators of production agents, this is a concrete argument for:
- investing in *structured failure logging*,
- allocating some training budget to error-aware post-training, not just scaling the base model or reward modeling.
Design trade-offs and failure modes to watch
Based on what’s described in the abstract, there are several practical trade-offs and risks.
Coverage vs. quality of diagnoses
AED covers 23 policy models and 33 environments [\[1\]](http://arxiv.org/abs/2609.40111v1). That breadth is good for generalization, but likely yields heterogeneous:
- error types,
- environment conventions,
- action spaces.
The AET pipeline’s verifier pass rate numbers (18.4% → 51.1% [\[1\]](http://arxiv.org/abs/2609.40111v1)) show that first-proposal corrections are *still* wrong or unverifiable about half the time, even after diagnosis. For your own system, expect:
- noisy labels in both diagnoses and corrections,
- a need for either strong automatic filters or human-in-the-loop curation for safety-critical domains.
Replay constraints
Replay is only possible “where replay is supported”; they use 3,062 matched replay pairs for the verifier-pass rate study [\[1\]](http://arxiv.org/abs/2609.40111v1). That’s a subset of the 50,228 error–diagnosis pairs [\[1\]](http://arxiv.org/abs/2609.40111v1).
If your environment stack includes:
- non-deterministic APIs,
- external services without snapshotting,
- real-world robotics,
you may have to settle for static trace-based verification only, which will be weaker than replay-based evaluation.
Overfitting to logged environments
AED’s improvements are reported:
- on a 943-case holdout drawn from the same general setup used for training the diagnosis model [\[1\]](http://arxiv.org/abs/2609.40111v1),
- on WebShop-lite for the actor repair comparison [\[1\]](http://arxiv.org/abs/2609.40111v1).
The abstract does not report cross-environment generalization outside this cluster. If you’re using a similar pipeline, you should expect:
- strong in-distribution gains,
- unknown behavior out-of-distribution (new environments, tools, or action abstractions).
Mitigation: keep environment diversity high and avoid making your diagnosis/repair model overly specific to a narrow harness family.
Why this matters for agent builders now
Even from the abstract-level view, AED and AET signal a shift in how we treat agent logs:
- Instead of looking at rollouts only through reward, we can operationalize error analysis at scale.
- Instead of manual red-teaming, we can bootstrap a continuous, model-driven failure-diagnosis loop.
The concrete numbers—32.7 percentage point verifier pass-rate gain for first-proposal corrections, ~16.4 absolute points of step-wise agreement gain over the unfine-tuned Qwen3-8B baseline, and a 6.67-point WebShop-lite gain from repair-focused actor training [\[1\]](http://arxiv.org/abs/2609.40111v1)—show that this is not just a conceptual nicety.
If you’re already running:
- multi-agent coding environments (like Offrun coordinating several coding agents [\[3\]](https://offrun.dev/)),
- or your own harnesses for generic models as in VISTA [\[7\]](http://arxiv.org/abs/2610.02200v1),
then you likely have access to exactly the data AET needs—you’re just not structuring or exploiting it this aggressively yet.
The core actionable move:
- Start treating each failure as a *data point for two models*:
- the main actor,
- a critic/diagnostician.
- Build logging and harness support so that future AED- or AET-style pipelines are feasible.
What is not documented
The abstract and metadata for AED and AET do *not* document:
- The exact schemas of traces and execution metadata, or how to implement them.
- The identities of the 33 environments, 19 harness families, or 23 policy models beyond being text-based agent systems.
- How the specific decision to revise is selected in each rollout.
- The architecture, training objective, or hyperparameters of the diagnosis model and actor models beyond mentioning Qwen3-8B.
- The implementation details of the verifier, or how verifier pass rate is computed.
- The exact nature of the internal teacher labels used to score step-wise agreement.
- The training hardware, compute budget, or optimizer schedules.
- Licensing terms for AED or any released models or code.