DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

A slow brain proposes. A fast brain grounds, monitors, refuses and substitutes at 2 Hz. One execution contract bounds every command to a frozen VLA, analytic skills and recovery skills. Failure evidence becomes validated capability revisions.

Haoyuan Deng1 · Jiebin Liu1 · Tengxiao Zhang1 · Langning Yan1 · Hongye Cao2 · Ziwei Wang1†
1Nanyang Technological University  ·  2Nanjing University  ·  †corresponding author
Task"put the mug in the basket"
safety 50 Hz
controller 20 Hz
fast brain 2 Hz
slow brain, on demand
Schematic animation; the rates, percentages and step costs quoted in the captions are the paper's.

Task choose a situation

Slow brain Qwen3-VL-4B · on demand · 913 ms

idle

Fast brain 2 Hz · 0.046 ms

progress
0.00 stagnation
0 risk
low
idleno command running

Evidence store append-only

LIBERO-Pro, 800 new initial states after freezing
0%vs 17.5% frozen π0.5
+57.8 pp · 473 / 11 exclusive wins · same policy, 4B upper model
Same capability library, full dynamic execution vs nominal one-step replanning
0%vs 63.9%
+10.1 pp · 89 / 8 exclusive wins · p = 2.0e-18
Fast-brain decision vs slow-brain call (median under load)
0 msvs 913 ms
47,199 decisions · 3,090 planner calls · one evidence store
Abstract

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system component should be revised.

DynaHarness is a dynamic physical harness that couples semantic reasoning with physical governance through a shared execution contract and turns failure evidence into validated capability revisions. The slow brain proposes capabilities and symbolic arguments; the fast brain grounds and monitors commands, refuses unresolved actions, substitutes capabilities and requests replans when needed. The contract bounds each accepted command and records execution evidence across analytic skills, recovery skills and the frozen VLA. Failure attribution localizes faults in these records and directs targeted revisions; paired regression checks govern admission or rejection.

On LIBERO-Pro, DynaHarness achieves 75.2% on 800 newly sampled initial states, compared with 17.5% for the frozen policy. With the same capability library, full dynamic execution reaches 74.0% versus 63.9% under nominal one-step replanning.

1
Structured physical capabilities

One execution contract grounds, bounds and refuses commands across analytic skills, recovery skills and a frozen VLA, recording evidence for revision.

2
Fast-slow capability execution

On-demand semantic reasoning and 2 Hz physical governance coordinate monitoring, substitution, recovery and replanning beyond nominal one-step replanning.

3
Attribution-driven self-evolution

Execution records localize failures and direct targeted revisions; paired regression checks govern admission; frozen snapshots keep their gains on new states.

800new initial states
13attribution layers
5 / 1admitted / reverted
4 × 10real tasks × trials
DynaHarness at a glance
Figure 1 of the paper. (a) From existing robot-agent paradigms to DynaHarness. (b) LIBERO-Pro success against upper-model scale. (c) Execution time: fast brain 0.046 ms, slow brain 0.913 s, Sonnet 5 via Claude Code 6.67 s. Click to enlarge.
Why a fast brain

Physical events are shorter than a slow decision.

In turn off the stove, the frozen policy turned the knob off at step 44 and back on by step 55: success held for 0.55 s. A planner that observes the scene once every few seconds never sees it. DynaHarness reads the benchmark predicate at every control chunk, latches it, and lets the 2 Hz fast brain end the command.

Sampling every 20 steps, no latch
25%
20 replays of the stove request
Bare policy, no harness
80%
the policy's own success on this request
Latched, chunk-level read
95%
paired ablation over 800 episodes: latch −28 net successes, p = 6.2e-5

Three clocks, one authority

Control steps run at 20 Hz, fast-brain decisions at 2 Hz, plan steps on demand. The slow brain does not drive the robot; the fast brain holds execution authority because it runs at a rate at which physical events can be observed. A 50 Hz envelope, lease and budget check sits underneath both.

ContinueInterveneSwitch capabilityEscalate to slow brainEnd on latched verdict

Failures run out the budget

In the analyzed round, 254 of 258 recorded failures ended at 99% or more of their budget while successes ended at a median of 67%. Across 5,805 episodes a failure spends 173 steps after its last progress against 51 for a success, and 220 when the planner layer failed. Detecting unproductive execution and supplying an effective replacement are separate requirements: the harness does both, through refusal, substitution and recovery inside the episode.

Positioning

From policy adaptation and agentic orchestration to fast-slow physical self-evolution.

Hover or tap a paradigm. Existing harnesses place a planner over a pretrained policy and treat checking as a pipeline stage; self-improving systems start from a failed task. DynaHarness gives execution an authority at the physical rate and gives evolution a layer to charge.

PropertyPolicy learningPlanner + frozen skillsFailure → skill updateDynaHarness
Semantic reasoning above the policy–✓✓✓
Execution authority at the physical rate (2 Hz decision, latched verdict)–pipeline stage–✓
Refusal is a recorded value with a reason–––✓
Frozen VLA is one callable capability next to analytic and recovery skills–somesome✓
Improvement without weight updates–✓✓✓
Failure localized to one of N ordered layers before revision––task-level✓
Paired gate and broader regression check that can reject a change––varies✓

The matrix summarizes the paradigms as characterized in the paper's introduction and related work; individual systems differ.

Method

One harness, two brains, one contract.

Figure 2 of the paper, live: particles trace the context, advisory, command and evidence paths. Click any panel or box for its description and equations.

Physical execution contract

The fast brain within one episode

Scrub through recorded episodes, frozen policy alongside.

Each DynaHarness clip is re-rendered from the episode's own event store, tool call by tool call, and reproduces the paper's keyframes pixel for pixel. The frozen-policy clip is a fresh rollout of the same seed recorded for this page, since π0.5 sampling is not seeded; it failed at the full budget, as the paper's episode did. Press play or drag the slider; the evidence store follows.

Capability composition in the records

Successful episodes exist despite individual capability failures. turn_knob_object completes 21% of its invocations while its source cell improves from 0% to 95%. Refusing push_object in favor of carrying the plate turns a cell from 0% to 40%: the library supplies the alternative and execution governance decides when to select it.

The VLA is one capability, not the controller

Of 770 archived control episodes, 570 never invoke the VLA, with 96.7% success. Policy calls concentrate in difficult episodes after analytic execution struggles; in retained failures the VLA reaches 91.4% of steps in the hardest cells, where no cell exceeds 35%.

Attribution-driven self-evolution

Charge the failure to the layer that broke. Admit only what a paired gate cannot reject.

One failed task is compatible with quite different repairs: the target was not grounded, the dispatched capability does not work, the verifier read the wrong instant, or recovery restored nothing. Attribution names the earliest of 13 ordered layers whose check fails, and the revision is stated in physical quantities rather than keyed to a task.

Σ successes ↑total successes over the paired development cells do not fall
Σ harness-attributed failures ↓failures charged to a harness layer do not rise
policy wins intactevery cell the bare policy already wins stays won
zero contaminationno contaminated episode; then broader regression checks

A local improvement is not a global one: the rejected round also contained a genuine local gain (the hinged-door capability raised the microwave cell from 0% to 25%). Judged by the round total the two changes would have been kept or discarded together; the per-cell comparison kept one and reverted the other.

Results

Explore the results by cell, by scope and by mechanism.

Every chart has a table twin. Numbers are those of the paper: full-denominator benchmark success, 200 episodes per cell, benchmark predicate as the verdict.

Cell
Scope

Slow-brain replacement

Replacing Qwen3-VL-4B with Claude Sonnet 5 through a coding agent gives 80% against 75% over 40 episodes on four untuned cells, and 75% with a tuned prompt, at 5.9× the latency per call and 18× the prompt tokens. In matched PhyAgentOS runs, changing the upper model from Qwen3-VL-4B to GPT-4o-mini changes success by one point.

Learning the fast brain

The evidence stores hold 2.59 million decisions over 27,581 episodes with the state the fast brain saw, the capability mask, the action and its reason. A learned variant matches the deterministic rule at 593 of 800 on the paired online comparison and escalates less often (12.2 against 15.4 per episode).

Real-world experiments

The same harness on a UR7e, with cameras instead of simulator state.

Qwen 4B as the slow brain, SAM3 for segmentation, RealSense D435 and D405 cameras. Four tasks, ten trials each, object positions varied across trials.

Ring placement · external + wrist
Grasp the yellow ring, place it over the green post, release and move away. One untrimmed trial, 29.5 s: 94 fast-brain ticks, two analytic skills, 16 tool calls, no intervention. Two stagnation flags (a missed-grasp check at 6 s, a stalled transport at 19 s) cleared at the next verification; the run ended on its own verified completion. Plan from the task manifest, no slow-brain call in this trial.
Cup stacking · external view
Stack the pink cup on the blue cup, then the green cup on the pink cup. One untrimmed trial, 117.9 s: four slow-brain calls (1.5 s, then 0.34 s each), 447 fast-brain ticks, four analytic skills. The run stopped on its own verified completion, a two-frame acceptance of both nestings.
Bread into the toaster · external view
Grasp the bread, place it in the left toaster slot, then press the white lever down. One untrimmed trial, 94.7 s. Placement stage: two slow-brain calls, 133 fast-brain ticks; seven stagnation flags (three missed-grasp checks, four stalled transports) cleared without an intervention. The lever press ends on a measured goal contract.
Drawer · refused before motion
Ring into the red drawer: the preflight scene check refused the command 3.2 s in, before any arm motion, with the recorded reason handle: object moved from the taught scene. External and wrist views, 3.3 s.
Real robot setup and four tasks
Workspace and the four tasks. Click to enlarge.

The verdict comes from the cameras

  • Ring placement ends on a placement check from a fresh observation: the measured ring-to-post relation must hold before the command completes.
  • Cup stacking ends on a two-frame acceptance of both nestings from SAM3 segmentation; the run stops on verified completion.
  • Bread placement and the lever press each end on a measured goal contract read from fresh hardware state, not on a model's judgment.
  • Drawer placement starts with a scene check against the taught scene; a moved handle refuses the command before any motion, reason recorded.
Real-world execution sequences
Execution sequences for ring placement, cup stacking, bread placement and ring placement in a drawer: the initial scene and four stages each.
Where the difference lies

Diagnostic runs of the baselines, with the same 4B model where possible.

Frames and quotes come from the baselines' own records. A retry that acts as extra budget, a verifier that changes what is recorded rather than what happens, a tool that reads simulator poses, a decision every several seconds.

Every video, frame, quote and number in this section comes from the baselines' own run records, re-encoded for the web and otherwise unchanged. Published rows are taken from the cited papers; matched rows are our runs with the same Qwen3-VL-4B upper model and the same frozen π0.5.

Citation

BibTeX

Code, evaluation configurations and scripts, and the run records behind every aggregate and paired comparison: github.com/Denghaoyuan123/DynaHarness.