What HumanoidToolBench actually is
HumanoidToolBench is a benchmark for evaluating humanoid robots on end-to-end tool use: from choosing an appropriate tool to manipulating it and, when necessary, walking to complete the task 1.
The authors define:
- 18 tasks
- Spanning three scenarios
- At three execution levels
- Under two tool-set modes
Together with that, they introduce ToolBook, a dataset of 3.1k demonstrations collected both in simulation and on a real Unitree G1 humanoid 1.
They evaluate:
- Seven policies in simulation
- Three policies on the real robot
and find “substantial gaps between selecting a suitable tool and completing the task” 1.
For you as an agent builder, the important bit is that this is not a pure perception benchmark or a pure motion benchmark. It’s explicitly structured to jointly stress:
- Tool selection (which object, given the goal?)
- Humanoid manipulation (how to use it with humanlike hands/arms?)
- Locomotion when needed (how to get to where the tool or task is?)
On top of that, they run focused probes on GR00T N1.7 and observe:
- Reduced tool selection accuracy on unseen tools
- Continued task execution under unrelated instructions 1
This is a red flag for anyone wiring LLM-style control into robots: the policy may keep “doing something” even when the instruction no longer matches the situation.
Code and data are available from the authors 1, so you can actually run your stack through this instead of guessing.
The structure: scenarios, execution levels, tool-set modes
The paper specifies the high-level design but doesn’t enumerate all details. What we do know:
- Three scenarios
These organize the 18 tasks into distinct settings for humanoid tool use 1.
- Three execution levels
These are meant to separate stages of the problem. At a minimum, they distinguish “just pick the right tool” from “actually complete the physical task,” because the results explicitly talk about a gap between tool selection and task completion 1.
- Two tool-set modes
These control what tools are available and how predictable the set is. The GR00T N1.7 probes specifically look at unseen tools versus known ones 1, so at least one mode includes tools outside the training distribution.
For your own system design, think of these three axes as knobs:
- Scenario → layout, task semantics, and environmental constraints.
- Execution level → which parts of your stack you’re actually testing (selection only vs. selection + control).
- Tool-set mode → how much you lean on memorization versus generalization over tool affordances.
Even without every task spelled out, this structure lets you treat HumanoidToolBench as a matrix of (environment × cognitive demand × generalization regime) and decide which slices your agent should actually be good at first.
ToolBook: why the demonstrations matter
ToolBook is a 3.1k-demonstration dataset, spanning both simulation and the real Unitree G1 1.
The important part is that:
- It is aligned with the benchmark tasks; it’s not arbitrary motion capture.
- It is captured across sim and real, so you have a concrete source for cross-domain transfer experiments 1.
For an agentic stack, ToolBook enables at least three use cases:
- Policy learning / cloning
Use it to pretrain a low-level controller (classic behavioral cloning or imitation) that knows how to wield tools in a humanoid body, leaving your high-level agent to mostly worry about what to do and when.
- Representation learning
You can learn joint embeddings over:
- Visual observations of tools and environments
- Proprioception and joint states
- Actions / trajectories
All of that is grounded in explicit tool-use behaviors.
- Sim-to-real diagnostics
Because demonstrations exist in both domains, you can benchmark how brittle your agent’s internal representations are when you swap sim for real data, or when using a policy initialized on sim-only demos.
The key point: instead of trying to learn “tool use” abstractly, you can ground your system on exactly the kind of humanoid manipulation + locomotion that the benchmark will later test.
What the early results tell you about failure modes
The authors run seven policies in simulation and three on the real robot 1.
Their main reported finding: there are “substantial gaps between selecting a suitable tool and completing the task” 1.
For an AI engineer, this says:
- Your tool-selection module can look great in isolation (picking the right object by name or appearance).
- The actual embodied execution—getting the robot to pick up, orient, move, and apply the tool while coordinating its body—remains the bottleneck.
The GR00T N1.7 probes add two more concrete failure patterns 1:
- Reduced selection accuracy on unseen tools
When faced with tools it hasn’t seen during training, GR00T N1.7 is worse at picking the right one. That suggests:
- Overreliance on memorized mappings (“this shape → hammer”) rather than affordance-level understanding.
- Weakness in generalizing from descriptions or visual cues to function.
- Continued task execution under unrelated instructions
Even when the instructions are no longer aligned with the task, the policy continues to execute. That implies:
- Insufficient grounding of language in the current state.
- A bias toward “always act” rather than “stop and re-check the goal.”
Both issues show up in language-based tool-use agents across domains; here they’re explicitly documented in a humanoid robotics setting.
Operationally, this motivates explicit architectural choices:
- Separation of selection and execution: don’t treat “pick tool + use it” as one monolithic LLM step.
- State-grounded re-evaluation: regularly re-check whether the current plan still matches the instruction and environment.
- Out-of-distribution tool handling: instrument tests that introduce novel tools and look specifically for selection and safety failures.
HumanoidToolBench gives you a concrete harness where these patterns actually show up, rather than only reasoning about them abstractly.
How this fits in your agent stack
Given what the benchmark measures, a practical humanoid tool-use stack that can engage with HumanoidToolBench will typically decompose along a few layers:
- Perception / state estimator
- Detect and localize tools and task-relevant objects.
- Estimate robot pose and contacts.
HumanoidToolBench itself does not specify perception details, but any policy you test will implicitly depend on some state representation.
- Tool selection module
- Given a goal and perceived tools, choose which tool to use.
- This might be:
- A classifier over tools,
- A retrieval model over a catalog,
- Or an LLM-based decision.
HumanoidToolBench directly evaluates this capability as a separate stage, since they can measure the gap between correct selection and actual completion 1.
- High-level planner / orchestrator
- Break the task into substeps: approach tool, grasp, reposition, act, verify outcome.
- Route control to locomotion vs manipulation primitives as needed.
The three execution levels in the benchmark are exactly about controlling how much of this pipeline you’re exercising 1.
- Low-level control policies
- Locomotion: walking, turning, balancing with changing loads.
- Manipulation: grasping and wielding tools.
This is where ToolBook’s 3.1k demonstrations are directly useful as supervision 1.
- Monitoring and safety layer
- Check that actions are consistent with the task spec.
- Abort on large deviations or ambiguous states.
GR00T N1.7’s tendency to continue executing under unrelated instructions shows you why you can’t assume “the policy will just stop” 1.
An end-to-end agent for HumanoidToolBench is then a controller that:
- Receives scenario/task-level instructions.
- Runs a tool selection pass (benchmarked explicitly).
- Generates a plan consistent with the execution level in question.
- Calls down to motion policies (possibly trained or fine-tuned on ToolBook).
- Monitors execution and re-queries selection or the planner when reality diverges.
HumanoidToolBench gives you the shared environment, tasks, and metrics to compare variations of this stack, rather than hand-wavy “robot demo” claims.
Why this benchmark matters now
From the paper’s perspective, the need is straightforward:
- Existing benchmarks “do not jointly evaluate” tool selection, manipulation, and locomotion for humanoids 1.
- But in practical deployments, these three dimensions are tightly coupled; your robot doesn’t get to assume a perfect gripper or a fixed base.
HumanoidToolBench plus ToolBook close at least three major gaps:
- Coupled evaluation of cognition and control
Instead of benchmarking:
- Tool recognition, or
- Grasping, or
- Walking
in isolation, you can evaluate agents whose *core challenge* is integrating these abilities in a way that respects tools and tasks.
- Shared embodied context across sim and real
The same 18 tasks and ToolBook dataset live in both simulation and on a real Unitree G1 1. That matters for:
- Testing learned policies before deployment.
- Measuring real-world degradation without reinventing tasks.
- Ground truth about generalization and instruction brittleness
The GR00T N1.7 probes are an early example of using this setup to expose:
- Generalization failures on unseen tools.
- Dangerous persistence when instructions become unrelated 1.
For builders working on agentic humanoids, this gives you a concrete target: “Our stack should at least close the documented selection–completion gap on HumanoidToolBench.”
How to start using HumanoidToolBench as an engineer
Given the information from the paper, a practical adoption plan looks like:
- Treat execution levels as stages in your CI
Use the three execution levels as:
- Level A: unit tests for tool-selection logic.
- Level B: integration tests for manipulation given a fixed tool.
- Level C: full-stack tests with locomotion + manipulation.
That way, when you refactor the high-level planner, you can see whether you broke selection, execution, or just the combined behavior.
- Leverage ToolBook for policy bootstrapping
Since ToolBook includes 3.1k demonstrations across sim and real on the Unitree G1 1:
- Pretrain manipulation and locomotion policies in simulation.
- Fine-tune or validate on real-robot demos.
- Reserve a slice of tasks for evaluation-only to avoid overfitting.
- Explicitly probe unseen-tool behavior
The documented drop in selection accuracy for unseen tools in GR00T N1.7 1 suggests adding your own probes:
- Hold out subsets of tools entirely during training.
- Evaluate selection accuracy and safety behaviors on those.
- Log when the agent insists on executing with obviously bad choices.
- Guard against instruction mismatch
Since policies can continue executing under unrelated instructions 1, build in:
- A separate module that monitors alignment between current state and task description.
- A “halt / query human” path triggered when divergence crosses a threshold.
- Periodic re-interpretation of the instruction in light of fresh observations, not just at the start.
- Benchmark multiple policy families
The authors already report seven policies in sim and three on real 1. To make your own results meaningful:
- Run diverse architectures: purely learned, rule-augmented, hierarchical.
- Compare where each fails: at selection, execution, or generalization.
- Keep your comparisons at the same execution level and tool-set mode to avoid apples-to-oranges interpretation.
The point is to make HumanoidToolBench something that runs every time you change a major component in your stack, not a one-off paper result.
What to watch next
HumanoidToolBench is an early step; the abstract hints at multiple directions:
- Extension of ToolBook
More tasks, more scenarios, more robots would expand what you can reasonably claim to handle.
- Better generalization metrics
Right now, we just know GR00T N1.7 struggles with unseen tools and instruction mismatch 1. Future work could standardize OOD splits and robustness scores for embodied tool use.
- Cross-benchmark synthesis
There are tool-use benchmarks in other modalities (e.g., CLIs, as in KaliBench 2), but HumanoidToolBench is specific to humanoid robots. It will be interesting to see whether techniques that help with textual tools transfer to physical tools and vice versa.
As more teams adopt it, you should expect to see:
- Reported curves for selection vs completion on the 18 tasks.
- Policies that close some of the selection–execution gap.
- Better diagnostics for the exact circumstances where humanoid agents decide to “keep going” on the wrong task.
For now, the benchmark gives you a reference point and shared language for talking about humanoid tool use with other builders.
What is not documented
From the available text, the following are not established:
- The specific identities or descriptions of the 18 tasks, three scenarios, three execution levels, or two tool-set modes.
- The architecture, training setup, or performance numbers of the seven simulated and three real-robot policies, beyond the existence of a gap between tool selection and task completion.
- Detailed quantitative metrics (success rates, error rates, or learning curves) for any policy, including GR00T N1.7, beyond the qualitative findings about unseen tools and unrelated-instruction execution.
- The exact structure, modalities, or annotation schema inside ToolBook beyond its size (3.1k demonstrations) and that they are collected in simulation and on a Unitree G1.
- Any implementation details of the released code and data (APIs, file formats, environment setup) beyond their availability.