ROBOTNESS
FortgeschrittenarXiv

ChunkTrust: Ausführungshorizonte von Roboter-Policies anhand von Signalen des Action-Experts anpassen

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi WangTsinghua University, Beijing Academy of Artificial Intelligence (BAAI), Renmin University of China, Shenzhen Technology University, Hefei University of Technology, Jiangnan University, Chongqing University, The Chinese University of Hong Kong
In 30 Sekunden

Roboter-KI-Modelle planen immer einen kurzen Block künftiger Bewegungen, und wie viele davon der Roboter ausführt, bevor er wieder hinschaut, legen Entwickler bislang meist starr fest. Dieser Preprint der Tsinghua-Universität, der Beijing Academy of Artificial Intelligence (BAAI) und weiterer Partner ergänzt ein Zusatzmodul, das diese Zahl im laufenden Betrieb wählt und die Erfolgsquoten vorhandener Modelle von Physical Intelligence und Nvidia in der Simulation und auf einem realen Zweiarmroboter ohne Nachtraining erhöht.

Forschungsfrage

Untersucht wird, ob eine Policy, die Aktionsblöcke (Action Chunks) ausgibt, bei jeder Neuplanung ohne Nachtraining selbst entscheiden kann, wie viel des vorhergesagten Blocks sie gefahrlos ausführen darf und wann sie die Szene erneut beobachten sollte.

Problem

Aktuelle Robot-Foundation-Policies wie π0 und π0.5 sagen einen Block künftiger Aktionen voraus und führen einen festen Anteil davon aus, bevor sie neu planen. Die Autoren zeigen, dass der beste feste Anteil je nach Aufgabe, Aufgabenlänge und Testbedingung variiert und sich sogar innerhalb einer Episode ändert. Eine Bewegung durch freien Raum verträgt lange Ausführung ohne Rückkopplung, eine heikle Kontaktphase nicht. Ein starrer Horizont verschwendet daher entweder Rechenzeit durch zu häufiges Neuplanen oder lässt den Roboter blind weiterfahren, wenn sein Plan nicht mehr zur Realität passt. In 1.600 Episoden mit festem Horizont scheiterten Durchläufe mit zwei hohen Instabilitätssignalen zu 97,4 Prozent, bei zwei niedrigen Signalen zu 50,2 Prozent.

Bisheriger Ansatz

In der Praxis wird der Ausführungshorizont meist pro Aufgabe durch Ausprobieren festgelegt. Frühere adaptive Verfahren leiten ihn aus Attention-Mustern, aus der Entropie der vorhergesagten Aktionen oder aus der Übereinstimmung unterschiedlich langer Vorhersagen ab. Bidirectional Decoding (BID) wählt unter mehreren Stichproben nach Rückwärtskohärenz und Vorwärtskontrast, FASTER priorisiert beim Sampling die nahen Zeitschritte. Die Autoren kritisieren, dass diese Methoden einen Block vor allem nach seiner inneren Stimmigkeit beurteilen. Ein in sich schlüssiger Plan, der nicht zur bereits ausgeführten Bewegung passt, kann so durchrutschen. Zudem sammeln sie keine Evidenz über mehrere Neuplanungen hinweg und reagieren nicht auf Phasenwechsel der Aufgabe.

Neuer Ansatz

ChunkTrust besteht aus zwei Teilen. Der trainingsfreie Action-aware Horizon Selector (AHS) misst online für jede Kandidatenlänge zwei Signale. Das erste ist die spektrale Stabilität innerhalb des Blocks, also wie stark sich der hochfrequente Anteil der vorhergesagten Geschwindigkeiten über die Entrauschungsschritte des Flow- oder Diffusions-Aktionskopfs verschiebt. Das zweite ist die Kontinuität zwischen Blöcken, also ob die Geschwindigkeit dort glatt bleibt, wo bereits ausgeführte Aktionen und neuer Abschnitt aneinanderstoßen. Beide Werte werden gleich gewichtet und mit einer Zuverlässigkeitsschätzung je Horizontlänge multipliziert, die als Beta-Verteilung nach jeder Neuplanung aktualisiert wird. Ein Vergessensfaktor von 0,99 lässt Evidenz aus früheren Aufgabenphasen abklingen, benachbarte Horizontlängen teilen ihr Feedback. Der optionale Query-based Horizon Adapter (QHA) ist ein kleines gelerntes Modul, das Bild-Sprach-Kontext und latente Aktionsdarstellungen der eingefrorenen Basis-Policy liest und darauf trainiert ist, die Horizontpräferenz des AHS in einem Durchlauf vorherzusagen. Im Einsatz wird diese Vorab-Verteilung mit der Online-Evidenz kombiniert. Die Basis-Policy selbst bleibt unverändert.

Ergebnisse

Auf allen 50 Aufgaben von RoboTwin 2.0 hob AHS die mittlere Erfolgsquote einer Multitask-π0.5 von 56,70 auf 63,50 Prozent, ein Plus von 6,80 Prozentpunkten bei einem 95-Prozent-Bootstrap-Intervall von 3,75 bis 9,95 Punkten. Einfache Varianten stiegen von 65,40 auf 73,10 Prozent, schwere von 48,00 auf 53,90 Prozent. In einer Teilmenge von acht Aufgaben mit aufgabenspezifischem Training verbesserte sich π0 von 18,19 auf 21,38 Prozent und mit QHA auf 24,88 Prozent, π0.5 von 29,63 auf 36,25 und dann 39,06 Prozent. Das bereits starke Fast-WAM legte von 86,81 auf 88,13 Prozent zu, X-VLA blieb nahezu gleich (47,44 auf 47,75 Prozent). Bei nicht trainierten RoboTwin-Aufgaben brachte QHA für π0.5 weitere 2,75 Punkte gegenüber AHS allein (40,00 auf 42,75 Prozent). Auf den 24 Tischaufgaben von RoboCasa GR1 gewann π0.5 2,42 Punkte (40,08 auf 42,50 Prozent), Nvidias Isaac GR00T N1.5 im Zero-Shot-Einsatz 1,67 Punkte, GR00T N1.6 3,80 Punkte (47,61 auf 51,42 Prozent) und eine Qwen3-basierte GR00T-Variante 9,67 Punkte (47,83 auf 57,50 Prozent). Auf einem zweiarmigen AgileX COBOT Magic mit π0.5 stieg bei vier Haushaltsaufgaben mit je 15 Durchläufen (Handtuch falten, Brot umsetzen, Getränk umsetzen, Spielzeugente in eine Schublade legen) der gleich gewichtete Mittelwert, der Teilpunkte für erledigte Zwischenschritte vergibt, von 50,4 auf 57,5 Prozent. Die Zugewinne reichten von 1,7 Punkten beim Handtuch bis 13,3 Punkten bei der Schubladenaufgabe. Der Selektor kostet je Neuplanung 0,7 bis 3,6 Prozent zusätzlich, abhängig von der Zahl der bewerteten Kandidaten. Die gesamte Inferenzzeit pro Episode stieg im 50-Aufgaben-Lauf jedoch von 0,97 auf 1,55 Sekunden, also um 58 Prozent, weil häufiger neu geplant wird. Asynchron mit Real-Time Chunking (RTC) betrieben, sank die Wartezeit des AHS bei schweren Aufgaben von 7,893 auf 0,344 Sekunden pro Episode.

Grenzen

Es handelt sich um einen Preprint ohne Peer-Review. Die Autoren räumen ein, dass das Verfahren Zugriff auf den Zwischenverlauf der Entrauschung braucht und sich daher nicht an Black-Box-Schnittstellen oder an Policies anschließen lässt, die Aktionen in einem Schritt ausgeben. Die Gewinne seien je Aufgabe ungleich verteilt, und hinter den Mittelwerten verbergen sich Rückschritte einzelner Aufgaben, vor allem wenn der gelernte QHA-Prior mit Online-Evidenz kombiniert wird, etwa bei einer Übergabeaufgabe unter schweren Bedingungen. Geringe Kosten je Neuplanung bedeuten zudem keine gleiche Episodendauer. Aus Sicht von ROBOTNESS ruht der Praxisnachweis auf einer Plattform, einer Basis-Policy und vier Aufgaben mit je 15 Durchläufen, der Gewinn von sieben Punkten ist entsprechend unsicher. Bei stärkeren oder anders gebauten Policies schrumpft der Effekt fast auf null (Fast-WAM 1,31 Punkte, X-VLA 0,31 Punkte), was darauf hindeutet, dass das Verfahren vor allem Schwächen mittelstarker Modelle ausgleicht. Die um 58 Prozent längere Inferenzzeit pro Episode zählt bei knappen Edge-Rechenbudgets, und die absoluten Erfolgsquoten in der Simulation liegen bei den meisten Aufgaben weit unter industrietauglichem Niveau.

Bedeutung für die Branche

Die Arbeit zielt auf eine Stellgröße, die jedes Team beim Einsatz blockweise planender VLA-Policies bislang von Hand einstellt. Getestet wurden unter anderem π0 und π0.5 von Physical Intelligence sowie Nvidias Isaac GR00T N1.5 und N1.6. Weil AHS kein Nachtraining braucht und der Code veröffentlicht ist, können Systemintegratoren und Mittelständler, die auf offenen VLA-Checkpoints aufbauen, das Verfahren binnen Wochen testen, am ehesten in Anwendungen, in denen freie Bewegungen und kontaktreiche Schritte wechseln, etwa in der Kommissionierung oder Montagevorbereitung. Die dauerhaftere Lehre für Modellanbieter lautet, dass der Ausführungshorizont zur Laufzeit aus den Unsicherheitssignalen der Policy selbst bestimmt werden sollte. Anbieter geschlossener Modelle dürften in den nächsten 12 bis 24 Monaten eher eigene Logik dieser Art in ihre Inferenz-Stacks einbauen, als genau diese Methode zu übernehmen. Für Käufer von Roboterhardware ist der kurzfristige Effekt gering, denn bei den stärksten Policies lag der Gewinn bei rund einem Punkt.

Vollständige Studie

ChunkTrust: Adapting Execution Horizons for Robot Policies with Action-Expert Evidence

Fanding Huang, Jingyan Jiang, Shifeng Bao, Mingkang Pu, Shiwei Li, Jing Xu, Shijia Xu, Guanbo Huang, Chenghao Gu, Yuzhi Huang, Chenxin Li, Faisal Nadeem Khan, Huan Yang, Yan Wang, Cheng Chi, Zhi Wang

Veröffentlicht unter CC BY 4.0. Wiedergabe mit Namensnennung, Original unter arXiv:2609.39754 (PDF).

Abstract

Robot foundation policies predict action chunks, but how many actions to execute before replanning depends on the current task phase. We introduce , which treats the execution horizon as a latent variable inferred from action-expert evidence rather than a fixed hyperparameter. Its training-free Action-aware Horizon Selector (AHS) combines intra-chunk spectral stability of generation traces with inter-chunk continuity between executed history and predicted actions. An online Beta posterior with kernel forgetting tracks horizon preferences across replans. A lightweight Query-based Horizon Adapter (QHA) optionally learns a context-conditioned dense prior from complementary evidence, fused with current evidence and episode-local Beta memory while the base policy remains frozen. Across RoboTwin2.0 and RoboCasa GR1 Tabletop, AHS improves overall task-averaged success for each evaluated base-policy configuration, including gains of +6.80 percentage points on \pi_{0.5} over all 50 RoboTwin2.0 tasks and +9.67 percentage points on Qwen3GR00T in RoboCasa. AHS+QHA raises the gain over Base to +9.44 percentage points on the eight-task \pi_{0.5} evaluation. On four real-world household tasks, AHS improves the equal-task mean normalized process score from 50.4% to 57.5%. Ablations examine the contributions of both evidence terms, temporal memory, and the learned prior.

Figure 1: Fixed execution horizons are brittle across tasks and settings. Top: the reliable executable prefix is phase dependent in a real bread-manipulation rollout (qualitative trust illustration K_{\mathrm{exec}}=K_{t}). (a) A shared robot policy with fixed or adaptive execution horizons. (b) Fixed horizons K\in\{10,20,30,40,50\} versus adaptive execution on \pi_{0} RoboTwin2.0 evaluations, grouped by task length and environment setting. Error bars indicate \pm 1 standard error of the mean across task–setting combinations. Dashed lines show the corresponding adaptive references.

1 Introduction

Robot foundation policies, including Vision-Language-Action (VLA) models (Black et al., 2024; Intelligence et al., 2025) and World-Action Models (WAMs) (Yuan et al., 2026; Ye et al., 2026b), predict action chunks to amortize inference and maintain temporal coherence. The execution horizon can differ from the generated chunk length, reflecting a trade-off between longer execution that can preserve smooth progress but delays correction and shorter execution that increases feedback frequency but can disrupt coherent motion (Lu et al., 2026). In manipulation, this trade-off changes within an episode: free-space approach may tolerate longer open-loop execution, whereas contact or delicate transport demands earlier replanning. Fig. 1 illustrates this phase dependence in a real rollout and shows that the best fixed horizon varies across task lengths and evaluation settings. The resulting trust boundary problem is to determine how many actions from the current chunk the robot should execute before replanning.

Adaptive chunking methods derive execution horizons from attention structure, action entropy, or cross-horizon agreement (Wang et al., 2026b; Liang et al., 2026; Jing et al., 2026), and increasingly from the policy’s denoising trajectory (Feng et al., 2026; Chen et al., 2026c). Other approaches learn when to replan (Zhao et al., 2026; Xu et al., 2026b) or monitor execution to trigger correction (Pan et al., 2026). These approaches tackle when to replan from different perspectives, yet an internally consistent prediction can still be incompatible with the motion already executed. Meanwhile, evidence at individual replans can be noisy, while similar execution contexts recur within and across episodes. This raises a central question: How can a robot identify complementary evidence native to its action expert and internalize it as reusable knowledge for adaptive execution?

We introduce ChunkTrust, a framework that couples online evidence accumulation with context-conditioned horizon learning (Fig. 3). ChunkTrust evaluates candidate execution prefixes along two complementary dimensions. We assess spectral stability during action generation, drawing on frequency-domain analyses of diffusion and flow models (Si et al., 2024; Huang et al., 2026a). We also assess continuity with recently executed motion, reflecting the importance of cross-chunk consistency in robot control (Liu et al., 2025b; Black et al., 2025). Our retrospective analysis in Fig. 2 provides empirical support for this pairing. Failed episodes exhibit higher median spectral instability and boundary variation, with the highest failure rate observed when both risks are elevated.

ChunkTrust internalizes this evidence through online memory and a learned prior. The Action-aware Horizon Selector (AHS) accumulates evidence in an episode-local Beta state with forgetting, retaining useful horizon preferences while adapting to phase changes. The Query-based Horizon Adapter (QHA) learns context-conditioned horizon preferences from action-expert evidence, enabling their reuse beyond the current episode. At deployment, QHA supplies a dense prior that complements online AHS evidence, combining learned preferences with adaptation to the current rollout.

We evaluate the benefits of online horizon adaptation and learned horizon preferences across RoboTwin2.0, RoboCasa GR1 Tabletop, and four real-world household tasks. AHS improves aggregate success across all evaluated simulation base-policy configurations, including gains of 6.80 percentage points on the complete 50-task RoboTwin2.0 suite and 9.67 percentage points with Qwen3GR00T on RoboCasa. Adding QHA further improves aggregate success on the eight-task evaluation and benefits two tasks excluded from horizon-head training, supporting reuse of the learned preferences beyond the QHA training tasks. On real robots, AHS increases the equal-task mean normalized process score by 7.1 percentage points. Controlled comparisons and ablations examine the contributions of complementary evidence, temporal memory, and the learned prior, alongside alternative horizon selectors and inference costs.

Our contributions are threefold: (1) we formulate execution-horizon adaptation as inference over candidate prefixes, grounded in the action expert’s generation dynamics and compatibility with executed history. (2) we introduce AHS, which integrates dual evidence with a phase-aware Beta posterior, and QHA, which learns a context-conditioned horizon prior that complements online evidence without updating the base policy. (3) we evaluate across policy families, two simulation benchmarks, and real robots, with full task-level results and controlled evidence, memory, transfer, and cost analyses.

2 Related Work

Figure 2: Horizon-aware evidence and episode outcomes. Analysis of 1,600 RoboTwin2.0 episodes with \pi_{0} and fixed K=H=50. (a,b) Episode means of replan-level evidence: \bar{z}_{\mathrm{intra}}(k)=R_{z}^{-1}\sum_{t}z_{\mathrm{intra},t}(k), \bar{u}_{\mathrm{inter}}(k)=R_{u}^{-1}\sum_{t}u_{\mathrm{inter},t}(k). Sums and counts R_{z},R_{u} use valid replans for each metric and horizon. Lines/bands show medians/interquartile ranges across episodes. (c) Failure rates after median-splitting episode risks (within-episode 75th percentiles of 1-q_{\mathrm{intra}} and 1-q_{\mathrm{inter}}). “Inter only”/“Intra only” means only the named risk is high. Details: Appendix B.
Action generation and reactive execution.

Mobile ALOHA (Fu et al., 2024b) and Diffusion Policy (Chi et al., 2025) use multi-step action predictions for temporal coherence. Flow-based policies such as \pi_{0} (Black et al., 2024) and \pi_{0.5} (Intelligence et al., 2025), and world-action models such as Fast-WAM (Yuan et al., 2026), extend action generation to broader task distributions. FASTER (Lu et al., 2026) prioritizes near-term sampling through horizon-aware scheduling and streaming execution. ChainVLA (Huang et al., 2026b) conditions successive queries on task progress and the unexecuted action suffix. Execution monitors offer another route to reactivity. VLA-Corrector (Pan et al., 2026) uses a learned latent dynamics model to detect persistent execution drift and guide action correction. React When You Need To (Wu et al., 2026) triggers asynchronous inference in response to scene changes. ChunkTrust instead selects a prefix at each replan using evidence already available from the action expert and executed history. It neither modifies the generated actions nor monitors new observations during that prefix.

Evidence-based horizon selection.

BID (Liu et al., 2025b) selects among sampled chunks using backward coherence and forward contrast, targeting consistency across predictions. Horizon adaptation instead changes the executed prefix. Mixture of Horizons (MoH) uses cross-horizon consensus (Jing et al., 2026), AutoHorizon uses action self-attention as a predictive-limit proxy (Wang et al., 2026b), and Adaptive Action Chunking (AAC) uses action entropy (Liang et al., 2026). HiPolicy combines multi-frequency chunk generation with entropy-guided execution (Zhang et al., 2026b). More recent methods expand the available signals. Knowing When to Stop (Xu et al., 2026a) detects entropy plateaus in action-to-observation cross-attention, while DVAC (Feng et al., 2026) measures variation in clean-action estimates during denoising. GeoAAC (Chen et al., 2026c) constructs prefix-wise geometric profiles from a single denoising trajectory. PACE (Nie et al., 2026) instead identifies low-speed transition points directly in the predicted chunk. ChunkTrust pairs spectral variation during generation with speed variation after stitching a candidate prefix to executed history. DVAC uses rolling history to calibrate its variance threshold, whereas AHS maintains horizon-indexed Beta states that accumulate the paired evidence as soft feedback with forgetting.

Learned horizon selection.

DEHP (Zhao et al., 2026) and BCP (Xu et al., 2026b) train horizon or continuation heads through reinforcement learning with frozen base policies. EQRL (Wang et al., 2026a) jointly learns to select the latent input, denoising budget, and chunk length, while Spatial Attention (SA) (Park et al., 2026) learns to forecast observation sensitivity and uses it to allocate execution horizons. QHA instead learns a context-conditioned horizon prior from complementary action-expert evidence and combines it with current evidence and episode-local Beta memory at deployment. Appendix G discusses additional connections.

3 Horizon-Aware Evidence from the Action Expert

3.1 Preliminaries

At replan step t, a frozen policy conditions on visual observation o_{t}, proprioceptive state \mathbf{x}_{t}^{\mathrm{prop}}, and instruction \ell. Its backbone produces context tokens \mathbf{C}_{t}=f_{\theta}(o_{t},\mathbf{x}_{t}^{\mathrm{prop}},\ell), and the action expert predicts

\hat{\mathbf{A}}_{t}=[\hat{\mathbf{a}}_{t,1},\ldots,\hat{\mathbf{a}}_{t,H}]\in\mathbb{R}^{H\times d_{a}}.
(1)

The controller executes a prefix of length K_{t}\in\{1,\ldots,H\} before re-observation. We score candidates k\in\mathcal{K}\subseteq\{1,\ldots,H\} using internal generation stability and compatibility with executed history. The expected-round rule in Sec. 4 can select intermediate integer lengths rather than only grid points.

A single policy call with trace recording returns (\hat{\mathbf{A}}_{t},\mathcal{F}_{t})=\pi_{\theta}(o_{t},\mathbf{x}_{t}^{\mathrm{prop}},\ell). For flow-based action experts (Lipman et al., 2023), \mathcal{F}_{t} contains velocity predictions recorded during generation, without changing the actions. We write the trace and executed history as

\mathcal{F}_{t}=\{\mathbf{v}_{t,\tau}\in\mathbb{R}^{H\times d_{a}}\}_{\tau=0}^{T-1},\qquad\mathcal{H}_{t}=[\mathbf{a}^{\mathrm{exec}}_{n_{t}-N_{t}+1},\ldots,\mathbf{a}^{\mathrm{exec}}_{n_{t}}],
(2)

where \tau indexes sampling steps, n_{t} counts actions executed before replan t, and N_{t} is the available history length. The following evidence terms use \mathcal{F}_{t} and (\mathcal{H}_{t},\hat{\mathbf{A}}_{t}), respectively.

Figure 3: Overview of horizon adaptation. AHS normalizes raw evidence, maintains per-candidate Beta reliability, and forms selection distribution \mu_{t}. QHA learns a context-conditioned prior from dense evidence.

3.2 Intra-Chunk Spectral Stability

We measure variation in the velocity-prefix spectrum during generation. Specifically, we apply the Fourier transform along the action horizon at each denoising step and compare the resulting spectra across steps. For each candidate k, we pad its velocity prefix to length H and apply a one-dimensional Fourier transform along the action-horizon axis, separately for action dimensions j\in\{1,\ldots,d_{a}\} and nonnegative frequencies \omega\in\Omega:

\widehat{\mathbf{V}}^{(k)}_{t,\tau}(\omega,j)=\operatorname{FFT}_{h}\!\left(\bar{\mathbf{v}}^{(k)}_{t,\tau}[:,j]\right)_{\omega},\quad\text{where}\quad\bar{\mathbf{v}}^{(k)}_{t,\tau}=\operatorname{Pad}_{H}\!\left(\mathbf{v}_{t,\tau,1:k,:}\right)\in\mathbb{R}^{H\times d_{a}},
(3)

Let \Omega_{\mathrm{hi}}\subset\Omega contain frequencies above cutoff fraction c_{\mathrm{cut}} of the discrete frequency grid. The fraction of energy in this band, aggregated across action dimensions, is

r_{\mathrm{hi},t,\tau}(k)=\frac{\sum_{\omega\in\Omega_{\mathrm{hi}}}\sum_{j=1}^{d_{a}}\left|\widehat{\mathbf{V}}^{(k)}_{t,\tau}(\omega,j)\right|^{2}}{\sum_{\omega\in\Omega}\sum_{j=1}^{d_{a}}\left|\widehat{\mathbf{V}}^{(k)}_{t,\tau}(\omega,j)\right|^{2}}.
(4)

High-frequency energy reflects rapid variation along the future-action axis. To track changes during generation, we define an early baseline \bar{r}_{\mathrm{hi},t,0}(k)=\frac{1}{W_{\tau}}\sum_{\tau=0}^{W_{\tau}-1}r_{\mathrm{hi},t,\tau}(k) from the first W_{\tau} sampling steps. The raw intra-chunk instability is its root mean square (RMS) deviation over all denoising steps, including the baseline window:

z_{\mathrm{intra},t}(k)=\left[\frac{1}{T}\sum_{\tau=0}^{T-1}\left(r_{\mathrm{hi},t,\tau}(k)-\bar{r}_{\mathrm{hi},t,0}(k)\right)^{2}\right]^{1/2}.
(5)

The early window supplies a reference rather than being discarded from the RMS calculation. The score measures deviation from this reference, not simply the final chunk’s high-frequency energy. Because raw scales vary across tasks, policies, and action normalizations, we min-max normalize this score within \mathcal{K} and set q_{\mathrm{intra},t}(k)=\exp[-\alpha_{\mathrm{intra}}\tilde{z}_{\mathrm{intra},t}(k)], where \alpha_{\mathrm{intra}}>0 controls the penalty strength. Larger spectral deviations thus receive lower quality. Normalization makes this a relative comparison among candidate prefixes at the current replan. The resulting quality need not be comparable in absolute scale across unrelated episodes. Fig. 2(a) uses the episode mean \bar{z}_{\mathrm{intra}}(k) of this same replan-level score, with the exact aggregation specified in the caption.

3.3 Inter-Chunk Continuity

An internally stable prefix may still be incompatible with recent motion. For continuity window W_{h}, we concatenate the available history suffix and candidate future prefix:

\mathbf{S}_{t}^{(k)}=\left[\operatorname{Suffix}_{(W_{h}-k)_{+}}(\mathcal{H}_{t}),\hat{\mathbf{a}}_{t,1},\ldots,\hat{\mathbf{a}}_{t,k}\right].
(6)

Here (x)_{+}=\max(0,x). Let \delta_{i}^{(k)}=\lVert\mathbf{S}_{t,i+1}^{(k)}-\mathbf{S}_{t,i}^{(k)}\rVert_{2} denote first-difference speed along the stitched trajectory. The raw inter-chunk discontinuity is the speed coefficient of variation:

u_{\mathrm{inter},t}(k)=\operatorname{Std}\left(\{\delta_{i}^{(k)}\}_{i}\right)/\operatorname{Mean}\left(\{\delta_{i}^{(k)}\}_{i}\right).
(7)

This proxy penalizes irregular speed, including boundary jumps and stop-and-go motion. A smaller value indicates a more uniform stitched trajectory, whereas a larger value can reflect an abrupt correction despite a smooth candidate viewed in isolation. As above, we min-max normalize within \mathcal{K} and set q_{\mathrm{inter},t}(k)=\exp[-\beta_{\mathrm{inter}}\tilde{u}_{\mathrm{inter},t}(k)], with \beta_{\mathrm{inter}}>0. Fig. 2(b) reports the corresponding episode mean \bar{u}_{\mathrm{inter}}(k) over valid replans. Failed episodes have higher median raw scores for both evidence terms across the candidate horizons. The shaded bands show interquartile ranges.

3.4 Dual-Evidence Fusion

We combine internal stability and compatibility with recent motion into

q_{\mathrm{mix},t}(k)=(1-\lambda)q_{\mathrm{intra},t}(k)+\lambda q_{\mathrm{inter},t}(k),\qquad\lambda\in[0,1].
(8)

The default is \lambda=0.5, with \lambda=0 and \lambda=1 recovering the intra-only and inter-only ablations. This score supplies action-expert evidence for horizon inference, not a deterministic execution rule or a calibrated task-success probability.

Fig. 2(c) summarizes episode risks by the 75th percentiles of 1-q_{\mathrm{intra}} and 1-q_{\mathrm{inter}} over replans, then splits each risk at its median. Failure rises from 50.2% when both risks are low to 97.4% when both are high, with intermediate rates when only one is high. These associations motivate combining the evidence, but they do not establish that a low-risk prefix guarantees success. The diagnostic episodes use fixed K=H=50, so the comparison characterizes the association of evidence with outcomes without selecting trajectories based on AHS decisions. Full distributions and aggregation details are reported in Appendix B.

4 Action-aware Horizon Selection and Query-based Adaptation

AHS uses the candidate quality q_{\mathrm{mix},t}(k) from Sec. 3 to maintain an online reliability posterior. QHA optionally learns a context-conditioned dense horizon prior from the same evidence (Fig. 3). Neither updates the base action generator.

4.1 Action-aware Horizon Selector

Online reliability posterior.

Manipulation alternates between phases such as approach, contact, transport, and placement. A horizon that was useful in a stable phase may become undesirable after a contact change, motivating memory that can also forget. To retain information across noisy replans, AHS maintains a Beta state for each k\in\mathcal{K}, initialized by a_{0}(k)=b_{0}(k)=1. This is episode-local memory of evidence-derived reliability, not a supervised task-success model. Following Thompson sampling (Russo et al., 2018), we draw a sample and combine it with the current quality:

\hat{\xi}_{t}(k)\sim\operatorname{Beta}(a_{t}(k),b_{t}(k)),\qquad s_{t}(k)=\hat{\xi}_{t}(k)\,q_{\mathrm{mix},t}(k).
(9)

Here q_{\mathrm{mix},t} measures the current chunk, while sampled reliability reflects accumulated episode-local evidence. Their product combines both. Except for uniform exploration with probability \epsilon_{\mathrm{exp}}, temperature scaling and expected-round selection give:

\mu_{t}(k)\propto s_{t}(k)^{1/T_{\mathrm{sel}}},\quad\sum_{k\in\mathcal{K}}\mu_{t}(k)=1,\qquad K_{t}=\operatorname{clip}_{[1,H_{t}^{\mathrm{avail}}]}\operatorname{round}\!\Bigl[\sum_{k\in\mathcal{K}}k\,\mu_{t}(k)\Bigr].
(10)

Here H_{t}^{\mathrm{avail}} is the available chunk length. If all scores degenerate, \mu_{t} falls back to uniform over valid candidates. Expected-round combines candidate preferences; exploration samples one valid candidate uniformly. Clipping keeps the prefix within the available prediction. AHS executes \hat{\mathbf{a}}_{t,1:K_{t}}, appends the executed actions to \mathcal{H}, and replans from the next observation.

Posterior update with kernel forgetting.

After execution, the soft feedback is y_{t}=\sum_{k\in\mathcal{K}}\mu_{t}(k)q_{\mathrm{mix},t}(k), or the sampled candidate quality on exploration steps. A Gaussian kernel w_{t}(k)\propto\exp[-(k-K_{t})^{2}/(2\sigma_{K}^{2})], normalized over \mathcal{K}, shares feedback among nearby horizons. Exponential forgetting (Raj and Kalyani, 2017) updates the state:

\displaystyle a_{t+1}(k) \\ \displaystyle=\rho_{f}a_{t}(k)+\eta\,w_{t}(k)y_{t}, \\ \displaystyle b_{t+1}(k) \\ \displaystyle=\rho_{f}b_{t}(k)+\eta\,w_{t}(k)(1-y_{t}),
(11) (12)

where \rho_{f}, \eta, and \sigma_{K} control forgetting, update strength, and neighborhood sharing. Decaying old evidence lets the selector adapt as the task changes phase. The feedback comes from action-expert quality, not an observed task-success label. Neighboring candidates receive shared soft evidence rather than independent rollout outcomes. Kernel sharing couples nearby lengths; forgetting discounts earlier phases.

4.2 Training-Time Query-Based Horizon Adapter

Table 1: AHS performance across simulation benchmarks. Success rate (%). Each comparison uses the same policy checkpoint. \Delta=\mathrm{SR}_{\mathrm{AHS}}-\mathrm{SR}_{\mathrm{Base}}.
  • RoboTwin2.0 Easy and Hard
    Task Scope
    Keine Daten
    Train Recipe
    Keine Daten
    Base
    Keine Daten
    +AHS
    Keine Daten
    \Delta (pp)
    Keine Daten
  • \pi_{0.5}
    Task Scope
    50 tasks
    Train Recipe
    multitask post-training
    Base
    56.70
    +AHS
    63.50
    \Delta (pp)
    +6.80
  • \pi_{0}
    Task Scope
    8 tasks
    Train Recipe
    task-specific post-training
    Base
    18.19
    +AHS
    21.38
    \Delta (pp)
    +3.19
  • \pi_{0.5}
    Task Scope
    8 tasks
    Train Recipe
    task-specific post-training
    Base
    29.63
    +AHS
    36.25
    \Delta (pp)
    +6.63
  • Fast-WAM
    Task Scope
    8 tasks
    Train Recipe
    multitask post-training
    Base
    86.81
    +AHS
    88.13
    \Delta (pp)
    +1.31
  • RoboCasa GR1 Tabletop
    Task Scope
    Keine Daten
    Train Recipe
    Keine Daten
    Base
    Keine Daten
    +AHS
    Keine Daten
    \Delta (pp)
    Keine Daten
  • \pi_{0.5}
    Task Scope
    24 tasks
    Train Recipe
    multitask post-training
    Base
    40.08
    +AHS
    42.50
    \Delta (pp)
    +2.42
  • GR00T N1.5
    Task Scope
    24 tasks
    Train Recipe
    zero-shot
    Base
    43.25
    +AHS
    44.92
    \Delta (pp)
    +1.67
  • GR00T N1.6
    Task Scope
    24 tasks
    Train Recipe
    zero-shot
    Base
    47.61
    +AHS
    51.42
    \Delta (pp)
    +3.80
  • Qwen3GR00T
    Task Scope
    24 tasks
    Train Recipe
    multitask post-training
    Base
    47.83
    +AHS
    57.50
    \Delta (pp)
    +9.67

QHA predicts a dense distribution p_{\phi}(K_{t}=h\mid\mathbf{C}_{t},\mathbf{Z}_{t}^{A}) for h\in\{1,\ldots,H\} from VLM context tokens \mathbf{C}_{t} and action-latent tokens \mathbf{Z}_{t}^{A}. With a sparse candidate grid, AHS explicitly scores each candidate, while QHA predicts a preference for every action step in one forward pass. Learned horizon queries use bridge cross-attention over both token streams (Fig. 3). Architecture and teacher construction are specified in Appendix A.6. The teacher normalizes dense evidence without online Beta memory:

\mu_{t}^{\star}(h)\propto q_{\mathrm{mix},t}(h)^{1/T_{\mathrm{teach}}},\qquad\sum_{h=1}^{H}\mu_{t}^{\star}(h)=1,\qquad h=1,\ldots,H.
(13)

Only QHA is trained, minimizing \mathcal{L}_{\mathrm{QHA}}=\operatorname{KL}(\mu_{t}^{\star}\,\|\,p_{\phi}). QHA thus encodes horizon preferences across training episodes in its parameters, providing cross-episode memory that complements AHS’s episode-local state. At deployment, let \tilde{q}_{\mathrm{mix},t}(h) denote quality on the dense grid, interpolated from sparse evidence or scored directly on that grid. QHA supplies a prior, combined with dense Beta reliability samples as

s_{\mathrm{eff},t}(h)=\hat{\xi}_{t}(h)\,\tilde{q}_{\mathrm{mix},t}(h)\,p_{\phi}(K_{t}=h\mid\mathbf{C}_{t},\mathbf{Z}_{t}^{A})^{\gamma},\qquad h=1,\ldots,H,
(14)

where \gamma\geq 0 controls prior strength. Applying temperature scaling and the selection rule above to this dense score yields AHS+QHA, combining learned preferences with current episode evidence. The dense prior supplies a context-conditioned preference before the episode-local state has accumulated much evidence. It does not remove the need to record generation traces or evaluate AHS candidates in the hybrid setting, as those computations provide the online correction to the prior.

5 Experiments

Table 2: QHA augmentation and transfer on RoboTwin2.0. Success rate (%). Easy and Hard denote clean and randomized settings. A: Evaluation on the eight tasks used to train QHA. B: For \pi_{0.5}, QHA is trained on six tasks and evaluated on two held-out tasks. Results are averaged over Easy and Hard. Bold marks the best method for each policy and evaluation condition.
  • —
    A. QHA training-task evaluation
    \pi_{0} (Black et al., 2024)
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    \pi_{0.5} (Intelligence et al., 2025)
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    Keine Daten
  • Task
    A. QHA training-task evaluation
    Base
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    AHS
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    AHS+QHA
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    Base
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    AHS
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    AHS+QHA
    A. QHA training-task evaluation
    Keine Daten
  • —
    A. QHA training-task evaluation
    Easy
    A. QHA training-task evaluation
    Hard
    A. QHA training-task evaluation
    Easy
    A. QHA training-task evaluation
    Hard
    A. QHA training-task evaluation
    Easy
    A. QHA training-task evaluation
    Hard
    A. QHA training-task evaluation
    Easy
    A. QHA training-task evaluation
    Hard
    A. QHA training-task evaluation
    Easy
    A. QHA training-task evaluation
    Hard
    A. QHA training-task evaluation
    Easy
    A. QHA training-task evaluation
    Hard
  • Blocks Ranking RGB
    A. QHA training-task evaluation
    19.00
    A. QHA training-task evaluation
    0.00
    A. QHA training-task evaluation
    28.00
    A. QHA training-task evaluation
    1.00
    A. QHA training-task evaluation
    31.00
    A. QHA training-task evaluation
    5.00
    A. QHA training-task evaluation
    36.00
    A. QHA training-task evaluation
    20.00
    A. QHA training-task evaluation
    48.00
    A. QHA training-task evaluation
    29.00
    A. QHA training-task evaluation
    56.00
    A. QHA training-task evaluation
    33.00
  • Handover Block
    A. QHA training-task evaluation
    41.00
    A. QHA training-task evaluation
    10.00
    A. QHA training-task evaluation
    41.00
    A. QHA training-task evaluation
    5.00
    A. QHA training-task evaluation
    64.00
    A. QHA training-task evaluation
    9.00
    A. QHA training-task evaluation
    44.00
    A. QHA training-task evaluation
    14.00
    A. QHA training-task evaluation
    44.00
    A. QHA training-task evaluation
    14.00
    A. QHA training-task evaluation
    32.00
    A. QHA training-task evaluation
    10.00
  • Handover Mic
    A. QHA training-task evaluation
    100.00
    A. QHA training-task evaluation
    2.00
    A. QHA training-task evaluation
    100.00
    A. QHA training-task evaluation
    24.00
    A. QHA training-task evaluation
    100.00
    A. QHA training-task evaluation
    30.00
    A. QHA training-task evaluation
    98.00
    A. QHA training-task evaluation
    64.00
    A. QHA training-task evaluation
    100.00
    A. QHA training-task evaluation
    56.00
    A. QHA training-task evaluation
    98.00
    A. QHA training-task evaluation
    59.00
  • Hanging Mug
    A. QHA training-task evaluation
    17.00
    A. QHA training-task evaluation
    3.00
    A. QHA training-task evaluation
    14.00
    A. QHA training-task evaluation
    9.00
    A. QHA training-task evaluation
    23.00
    A. QHA training-task evaluation
    10.00
    A. QHA training-task evaluation
    14.00
    A. QHA training-task evaluation
    8.00
    A. QHA training-task evaluation
    13.00
    A. QHA training-task evaluation
    12.00
    A. QHA training-task evaluation
    13.00
    A. QHA training-task evaluation
    14.00
  • Place A2B Left
    A. QHA training-task evaluation
    24.00
    A. QHA training-task evaluation
    1.00
    A. QHA training-task evaluation
    34.00
    A. QHA training-task evaluation
    0.00
    A. QHA training-task evaluation
    39.00
    A. QHA training-task evaluation
    1.00
    A. QHA training-task evaluation
    41.00
    A. QHA training-task evaluation
    4.00
    A. QHA training-task evaluation
    47.00
    A. QHA training-task evaluation
    4.00
    A. QHA training-task evaluation
    47.00
    A. QHA training-task evaluation
    15.00
  • Place Bread Basket
    A. QHA training-task evaluation
    11.00
    A. QHA training-task evaluation
    8.00
    A. QHA training-task evaluation
    18.00
    A. QHA training-task evaluation
    16.00
    A. QHA training-task evaluation
    15.00
    A. QHA training-task evaluation
    10.00
    A. QHA training-task evaluation
    30.00
    A. QHA training-task evaluation
    17.00
    A. QHA training-task evaluation
    50.00
    A. QHA training-task evaluation
    29.00
    A. QHA training-task evaluation
    55.00
    A. QHA training-task evaluation
    41.00
  • Place Bread Skillet
    A. QHA training-task evaluation
    15.00
    A. QHA training-task evaluation
    2.00
    A. QHA training-task evaluation
    23.00
    A. QHA training-task evaluation
    5.00
    A. QHA training-task evaluation
    23.00
    A. QHA training-task evaluation
    2.00
    A. QHA training-task evaluation
    24.00
    A. QHA training-task evaluation
    10.00
    A. QHA training-task evaluation
    36.00
    A. QHA training-task evaluation
    19.00
    A. QHA training-task evaluation
    44.00
    A. QHA training-task evaluation
    19.00
  • Place Can Basket
    A. QHA training-task evaluation
    33.00
    A. QHA training-task evaluation
    5.00
    A. QHA training-task evaluation
    22.00
    A. QHA training-task evaluation
    2.00
    A. QHA training-task evaluation
    33.00
    A. QHA training-task evaluation
    3.00
    A. QHA training-task evaluation
    35.00
    A. QHA training-task evaluation
    15.00
    A. QHA training-task evaluation
    50.00
    A. QHA training-task evaluation
    29.00
    A. QHA training-task evaluation
    55.00
    A. QHA training-task evaluation
    34.00
  • Mean by setting
    A. QHA training-task evaluation
    32.50
    A. QHA training-task evaluation
    3.88
    A. QHA training-task evaluation
    35.00
    A. QHA training-task evaluation
    7.75
    A. QHA training-task evaluation
    41.00
    A. QHA training-task evaluation
    8.75
    A. QHA training-task evaluation
    40.25
    A. QHA training-task evaluation
    19.00
    A. QHA training-task evaluation
    48.50
    A. QHA training-task evaluation
    24.00
    A. QHA training-task evaluation
    50.00
    A. QHA training-task evaluation
    28.13
  • Overall
    A. QHA training-task evaluation
    18.19
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    21.38
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    24.88
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    29.63
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    36.25
    A. QHA training-task evaluation
    Keine Daten
    A. QHA training-task evaluation
    39.06
    A. QHA training-task evaluation
    Keine Daten

We evaluate on RoboTwin2.0 (Chen et al., 2025), RoboCasa GR1 Tabletop (Nasiriany et al., 2024), and real robots. Each Base–AHS comparison fixes the policy checkpoint. AHS changes execution without updating policy weights. Simulation tables report success rates (%). In Tabs. 1 and 2, tasks are weighted equally after averaging Easy and Hard within each RoboTwin2.0 task. Rounding follows aggregation. Easy and Hard denote clean and randomized settings. Appendix A specifies checkpoints, candidate horizons, and evaluation protocols.

5.1 AHS across Simulation Benchmarks

Tab. 1 tests AHS across benchmarks and training regimes. Its task scope and training recipe distinguish evaluation coverage from checkpoint provenance. Absolute scores across regimes are not matched-policy comparisons. The evaluation spans task-specific post-training, multitask post-training, and zero-shot deployment on the target benchmark. Gains in these regimes test whether execution adaptation remains useful across the evaluated backbones. They do not imply that AHS repairs every failed task or replaces policy training.

RoboTwin2.0.

The full-suite evaluation uses one multitask-post-trained \pi_{0.5} (Intelligence et al., 2025) checkpoint on 50 tasks, with 20 rollouts per task and setting. AHS improves success from 56.70% to 63.50% (+6.80 percentage points). The eight-task evaluation uses task-specific \pi_{0} (Black et al., 2024) and \pi_{0.5} checkpoints and multitask Fast-WAM (Yuan et al., 2026), with 100 rollouts per task and setting. All three improve in aggregate. These are distinct checkpoint and evaluation cohorts, not subset and full-suite results for one policy. Fast-WAM improves from 86.81% to 88.13%, showing that execution adaptation can still help a stronger baseline, although its aggregate gain is smaller than those of the task-specific policies. Appendix Tabs. 8 and 9 retain full per-task outcomes, including regressions.

RoboCasa GR1 Tabletop.

We evaluate all 24 tasks; the \pi_{0.5} control uses 50 rollouts per task. The multitask-post-trained \pi_{0.5} and Qwen3GR00T (Ye et al., 2026a; Community, 2026) improve from 40.08% to 42.50% and from 47.83% to 57.50%, respectively. Both zero-shot Isaac-GR00T baselines (Bjorck et al., 2025) also improve in aggregate. All four use \mathcal{K}=\{4,8,12,16\}. Appendix Tab. 10 provides the complete breakdown. Fixed-horizon, action-only, temporal-memory, and AAC (Liang et al., 2026) comparisons appear in Appendix C.5.

5.2 QHA Augmentation and Transfer

Augmentation on QHA training tasks.

Each policy family uses one QHA head trained on eight tasks with the base policies frozen (Tab. 2A). AHS+QHA improves overall success over AHS from 21.38% to 24.88% for \pi_{0} and from 36.25% to 39.06% for \pi_{0.5}. Both Easy and Hard means improve, but \pi_{0.5} regresses on Handover Block in both settings. This may reflect a mismatch between the shared horizon prior and the feedback timing required during bimanual object transfer.

Transfer to tasks held out from QHA training.

A separate \pi_{0.5} head trains on six tasks and tests on two held-out tasks (Tab. 2B). AHS+QHA improves the equal-task mean from 40.00% to 42.75%. Blocks Ranking RGB improves from 38.50% to 43.50%, compared with 41.50% to 42.00% on Place Bread Basket. Only the horizon head is task-held-out: base policies remain task-specific. Appendix C.1 gives the split, per-setting results, QHA-only control, and decision agreement.

5.3 Real-World Deployment

Figure 4: Real-world deployment. (a) Four bimanual household tasks: fold towels, move bread to a plate, move a drink to a basket, and put a duck toy into a drawer. (b) Mean normalized process scores for \pi_{0.5} and \pi_{0.5} + AHS over 15 rollouts per task and method. Error bars show \pm one standard error of the mean across rollouts.

We deploy \pi_{0.5} on the AgileX COBOT Magic ALOHA-style bimanual platform, comparing fixed K=25 with AHS over 15 rollouts on each of four household tasks (Fig. 4). Object positions vary across rollouts to test spatial generalization. Scores average predefined sub-steps and rollouts to measure partial progress, rather than binary success (Appendix A.1). AHS raises the equal-task mean from 50.4% to 57.5%, with the largest gain on Duck Toy to Drawer: 50.0% to 63.3%.

5.4 Ablation Studies

Table 3: Candidate-set scaling on \pi_{0.5}. Success and per-replan overhead over four RoboTwin2.0 tasks as |\mathcal{K}| increases. Fixed denotes \pi_{0.5} without horizon selection. \DeltaSuccess is in percentage points.
  • Fixed
    |\mathcal{K}|
    –
    Success (%)
    22.00
    \DeltaSuccess
    –
    Replan time (ms)
    95.83
    Overhead (%)
    –
    Replan counts
    4.22
    Ep. time (s)
    22.43
  • \mathcal{K}_{5}
    |\mathcal{K}|
    5
    Success (%)
    33.00
    \DeltaSuccess
    +11.00
    Replan time (ms)
    96.50
    Overhead (%)
    0.70
    Replan counts
    8.36
    Ep. time (s)
    23.17
  • \mathcal{K}_{10}
    |\mathcal{K}|
    10
    Success (%)
    32.50
    \DeltaSuccess
    +10.50
    Replan time (ms)
    96.84
    Overhead (%)
    1.05
    Replan counts
    9.34
    Ep. time (s)
    23.08
  • \mathcal{K}_{50}
    |\mathcal{K}|
    50
    Success (%)
    34.88
    \DeltaSuccess
    +12.88
    Replan time (ms)
    99.28
    Overhead (%)
    3.60
    Replan counts
    10.57
    Ep. time (s)
    21.39
Table 4: Component ablation (\pi_{0}). Four-task mean success. Top two: bold/underline.
  • Base \pi_{0} (default)
    Easy (%)
    20.75
    Hard (%)
    4.00
    Overall (%)
    12.38
  • Inter-chunk only
    Easy (%)
    23.50
    Hard (%)
    4.00
    Overall (%)
    13.75
  • Intra-chunk only
    Easy (%)
    23.75
    Hard (%)
    5.25
    Overall (%)
    14.50
  • Both evidence terms, no posterior
    Easy (%)
    25.50
    Hard (%)
    3.75
    Overall (%)
    14.63
  • Full AHS
    Easy (%)
    24.25
    Hard (%)
    5.75
    Overall (%)
    15.00
  • Full AHS+QHA
    Easy (%)
    27.50
    Hard (%)
    4.00
    Overall (%)
    15.75

All three studies use Place A2B Left, Place Bread Basket, Place Bread Skillet, and Place Can Basket, with 100 rollouts per task and condition. All three studies include Easy and Hard.

Candidate-set scaling.

On \pi_{0.5}, five candidates raise success from 22.0% with fixed K=50 to 33.0%, versus 34.9% with 50 candidates (Tab. 3). Added per-replan cost rises from 0.67 to 3.45 ms, or 0.70–3.60% of policy inference time. Sparse candidates thus capture most of the gain. The table’s episode times cover successes only, so they do not establish all-attempt cost. Appendix D.1 reports all-episode costs for additional cohorts. Appendix E gives hyperparameter sweeps.

Component ablation.

On \pi_{0}, combining both evidence terms without memory gives 14.6% overall but 3.75% on Hard (Tab. 4). Full AHS reaches 15.0% overall and the best Hard result, 5.75%, supporting temporal memory. QHA raises overall success to 15.8% while Hard falls to 4.00%. Appendix C.5 compares memory designs with shared evidence and candidates.

Latency and asynchronous execution.

We combine AHS with real-time chunking (RTC; Black et al., 2025) on task-specific \pi_{0.5} policies (Fig. 5). All predict H=50 actions. Fixed uses target budget K=40, and AHS scores candidates in \{10,20,30,40\} without QHA. At +200 ms, RTC reduces AHS waiting from 7.893 to 0.344 s/episode on Hard and from 6.310 to 0.344 s on Easy. Within RTC, AHS lowers mean observation age from 4.92 to 3.07 s on Hard and from 5.00 to 3.18 s on Easy, while SR rises from 22.00% to 26.75% and from 39.00% to 42.75%, respectively. This requires more policy calls and model computation (Appendix D.2). RTC does not uniformly improve SR over synchronous AHS: at +200 ms on Easy, SR is 42.75% versus 44.75%.

Figure 5: AHS with asynchronous execution on \pi_{0.5} Hard. (a) Mean waiting and observation age. Points run left to right as +0/+100/+200 ms additional delay. (b) Final SR. (c) SR(t) at +200 ms. All 400 outcomes per condition are included. Bars and shading are marginal/pointwise 95% paired-seed bootstrap intervals within four tasks. Time uses a controlled physical clock with a 142 ms base delay, not native deployment timing. Easy curves and full results: Appendix D.2.

6 Conclusion

ChunkTrust adapts robot-policy execution horizons using action-expert evidence. AHS combines spectral stability, inter-chunk continuity, and online temporal memory, while QHA learns a context-conditioned horizon prior from the same evidence without changing the base policy. Experiments support aggregate gains across two simulation benchmarks and real-world manipulation, while transfer, ablation, and cost analyses characterize where adaptation helps. The method requires accessible generation traces, and gains are not uniform across tasks. The evaluation also separates per-replan overhead from episode-level cost: a small selector cost does not imply identical total rollout time. Task breakdowns likewise expose regressions that aggregate gains can conceal, particularly when a learned prior is combined with online evidence. Future work includes multi-chunk horizon inference, cross-policy transfer of the learned prior, and broader real-world validation across hardware and domain shifts.

References

  1. Bai et al. (2025) S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631.
  2. Bi et al. (2025) H. Bi, H. Tan, S. Xie, Z. Wang, S. Huang, H. Liu, R. Zhao, Y. Feng, C. Xiang, Y. Rong, et al. Motus: a unified latent action world model. arXiv preprint arXiv:2512.13030.
  3. Bjorck et al. (2025) J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al. Gr00t n1: an open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734.
  4. Black et al. (2024) K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al. \pi_{0}: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164.
  5. Black et al. (2025) K. Black, M. Galliker, and S. Levine Real-time execution of action chunking flow policies. In Advances in Neural Information Processing Systems, Vol. 38, Main Conference, pp. 33383–33407. External Links: Document, Link
  6. Brohan et al. (2023) A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. S. Ryoo, G. Salazar, P. R. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. H. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich RT-1: Robotics Transformer for Real-World Control at Scale. In Proceedings of Robotics: Science and Systems, Daegu, Republic of Korea. External Links: Document
  7. Chen et al. (2026a) F. Chen, X. Wang, Y. Chen, B. Li, Y. He, Z. Zhang, and Y. Wu Dynamic Execution Commitment of Vision-Language-Action Models. arXiv preprint arXiv:2605.11567. External Links: Link
  8. Chen et al. (2025) T. Chen, Z. Chen, B. Chen, Z. Cai, Y. Liu, Q. Liang, Z. Li, X. Lin, Y. Ge, Z. Gu, et al. RoboTwin 2.0: a scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation. arXiv preprint arXiv:2506.18088.
  9. Chen et al. (2026b) W. Chen, K. Zhang, C. Lin, Z. Zhang, Y. She, Y. Liu, R. A. Yeh, S. Mou, and Y. Gu DREAM-Chunk: Reactive Action Chunking with Latent World Model. arXiv preprint arXiv:2606.18589. External Links: Link
  10. Chen et al. (2026c) X. Chen, S. Chen, Y. Ding, J. Liu, G. Wang, W. Ye, H. T. Shen, and Y. Bin GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies. arXiv preprint arXiv:2609.20776. External Links: Link
  11. Chi et al. (2025) C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song Diffusion policy: visuomotor policy learning via action diffusion. The International Journal of Robotics Research 44 (10-11), pp. 1684–1704.
  12. Community (2026) S. Community StarVLA: a lego-like codebase for vision-language-action model developing. arXiv preprint arXiv:2604.05014.
  13. Driess et al. (2023) D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence PaLM-e: an embodied multimodal language model. In Proceedings of the 40th International Conference on Machine Learning, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (Eds.), Proceedings of Machine Learning Research, Vol. 202, pp. 8469–8488. External Links: Link
  14. Feng et al. (2026) X. Feng, Y. Cheng, C. Shi, B. Han, Y. Yan, Y. Hong, Z. Tian, and L. Jiang Denoising Tells When to Replan: Denoising-Variance Adaptive Chunking for Flow-Based Robot Policies. arXiv preprint arXiv:2606.03847. External Links: Link
  15. Fu et al. (2024a) Z. Fu, T. Z. Zhao, and C. Finn Mobile aloha: learning bimanual mobile manipulation using low-cost whole-body teleoperation. In 8th Annual Conference on Robot Learning,
  16. Fu et al. (2024b) Z. Fu, T. Z. Zhao, and C. Finn Mobile ALOHA: learning bimanual mobile manipulation using low-cost whole-body teleoperation. In 8th Annual Conference on Robot Learning, External Links: Link
  17. Ghosh et al. (2024) D. Ghosh, H. R. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y. L. Tan, L. Y. Chen, Q. Vuong, T. Xiao, P. R. Sanketi, D. Sadigh, C. Finn, and S. Levine Octo: An Open-Source Generalist Robot Policy. In Proceedings of Robotics: Science and Systems, Delft, Netherlands. External Links: Document
  18. Guo et al. (2026) J. Guo, F. Chen, Z. Mao, W. L. H. Kenny, Z. Wu, Y. Li, Y. Cai, Y. Chen, Y. Ban, K. Chen, et al. Frequency-Aware Flow Matching for Continuous and Consistent Robotic Action Generation. In Advances in Neural Information Processing Systems,
  19. Guo and Guo (2026) T. Guo and M. Guo Action ControlNet: A Lightweight Delay-Aware Adapter for Smooth Asynchronous Control in Vision-Language-Action Models. arXiv preprint arXiv:2606.25985. External Links: Link
  20. He et al. (2026) Q. He, Z. Yang, W. Liang, C. Hao, N. Sebe, and J. Tian FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy. In Proceedings of the 43rd International Conference on Machine Learning,
  21. Hu et al. (2022) E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations, External Links: Link
  22. Huang et al. (2026a) G. Huang, J. Mao, F. Huang, F. Liu, X. Luo, Y. Liang, J. Lu, X. Wang, P. Liu, R. Fu, R. Huang, and S. Huang Exposure bias can alleviate itself via directional and frequency rectification in flow matching. In European Conference on Computer Vision (ECCV), External Links: Link
  23. Huang et al. (2026b) Y. Huang, W. Bu, Z. Xiong, J. Wu, F. Huang, J. Jiang, and Z. Wang ChainVLA: chaining vision-language-action queries through a unified execution state for long-horizon manipulation. arXiv preprint arXiv:2608.02326.
  24. Intelligence et al. (2025) P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al. \pi_{0.5}: A vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054.
  25. Jing et al. (2026) D. Jing, G. Wang, J. Liu, W. Tang, Z. Sun, Y. Yao, Z. Wei, Y. Liu, Z. Lu, and M. Ding Mixture of horizons in action chunking. In Proceedings of the 43rd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 306, pp. 54350–54370. External Links: Link
  26. Kim et al. (2024) M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn OpenVLA: an open-source vision-language-action model. In 8th Annual Conference on Robot Learning, External Links: Link
  27. Lazzati et al. (2026) F. Lazzati, K. Stachowicz, W. Chen, A. M. Metelli, A. Wagenmaker, and S. Levine Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?. arXiv preprint arXiv:2608.02547. External Links: Link
  28. Liang et al. (2026) Y. Liang, X. Wang, K. Wang, S. Wang, X. Peng, H. Chen, D. K. H. Chua, and P. Vadakkepat Adaptive action chunking at inference-time for vision-language-action models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 20802–20811.
  29. Lipman et al. (2023) Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations,
  30. Liu et al. (2025a) S. Liu, L. Wu, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu RDT-1b: a diffusion foundation model for bimanual manipulation. In The Thirteenth International Conference on Learning Representations, External Links: Link
  31. Liu et al. (2025b) Y. Liu, J. I. Hamid, A. Xie, Y. Lee, M. Du, and C. Finn Bidirectional decoding: improving action chunking via guided test-time sampling. In International Conference on Learning Representations, External Links: Link
  32. Liu et al. (2026) Y. Liu, H. Yu, J. Zhao, B. Li, D. Zhang, M. Li, W. Wu, Y. Hu, J. Xie, J. Guo, D. Wang, and Y. Gao Learning Native Continuation for Action Chunking Flow Policies. In Proceedings of Robotics: Science and Systems, Sydney, Australia. External Links: Document
  33. Lu et al. (2026) Y. Lu, Z. Liu, X. Fan, Z. Yang, J. Hou, J. Li, K. Ding, and H. Zhao FASTER: rethinking real-time flow VLAs. arXiv preprint arXiv:2603.19199.
  34. Nasiriany et al. (2024) S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y. Zhu RoboCasa: large-scale simulation of everyday tasks for generalist robots. In RSS 2024 Workshop: Data Generation for Robotics,
  35. Nie et al. (2026) J. Nie, J. Li, C. Liu, J. Lao, J. Zhang, T. Zhang, L. Lin, and S. Huang PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking. arXiv preprint arXiv:2606.00537. External Links: Link
  36. O’Neill et al. (2024) A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al. Open x-embodiment: robotic learning datasets and rt-x models: open x-embodiment collaboration 0. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pp. 6892–6903.
  37. Pan et al. (2026) Y. Pan, M. Pan, Q. Lu, J. Huang, M. Zhang, S. Huang, X. Li, J. Zhang, Y. Shen, X. Zhang, and W. Zhang VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon. In Advances in Neural Information Processing Systems (NeurIPS),
  38. Park et al. (2026) C. Park, J. Ha, J. Fu, and F. C. Park Spatial Attention: Adapting Execution Horizons for Diffusion Policies via Observation Sensitivity. arXiv preprint arXiv:2607.04739. External Links: Link
  39. Raj and Kalyani (2017) V. Raj and S. Kalyani Taming non-stationary bandits: a Bayesian approach. Note: arXiv preprint arXiv:1707.09727 External Links: 1707.09727, Link
  40. Russo et al. (2018) D. J. Russo, B. Van Roy, A. Kazerouni, I. Osband, and Z. Wen A tutorial on thompson sampling. Foundations and Trends in Machine Learning 11 (1), pp. 1–96.
  41. Sendai et al. (2025) K. Sendai, M. Alvarez, T. Matsushima, Y. Matsuo, and Y. Iwasawa Leave No Observation Behind: Real-time Correction for VLA Action Chunks. arXiv preprint arXiv:2509.23224. External Links: Link
  42. Si et al. (2024) C. Si, Z. Huang, Y. Jiang, and Z. Liu FreeU: free lunch in diffusion U-Net. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4733–4743. External Links: Link
  43. So et al. (2026) J. So, C. Lee, S. Lee, J. Ok, and E. Park Improving generative behavior cloning via self-guidance and adaptive chunking. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link
  44. Team et al. (2025) B. R. Team, M. Cao, H. Tan, Y. Ji, X. Chen, M. Lin, Z. Li, Z. Cao, P. Wang, E. Zhou, et al. Robobrain 2.0 technical report. arXiv preprint arXiv:2507.02029.
  45. Vaswani et al. (2017) A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin Attention is all you need. Advances in neural information processing systems 30.
  46. Wang et al. (2026a) G. Wang, X. Tan, X. Li, M. Luo, C. Yao, S. Yan, J. Yang, F. Feng, H. Cai, X. Wang, Z. Mai, Y. Zhao, Y. Han, and Z. Li Elastic Queries Reinforcement Learning: Self-Aware Policy Execution for VLA Models. arXiv preprint arXiv:2606.14375. External Links: Link
  47. Wang et al. (2026b) H. Wang, G. Zhang, Y. Yan, R. R. Kompella, and G. Liu VLA knows its limits. arXiv preprint arXiv:2602.21445.
  48. Wang et al. (2026c) H. Wang, G. Zhang, Y. Yan, Y. Shang, R. Kompella, and G. Liu Real-Time Robot Execution with Masked Action Chunking. In International Conference on Learning Representations, pp. 120161–120179.
  49. Wang et al. (2026d) R. Wang, Y. Zhang, C. Chen, J. Lin, Z. Wang, and X. Qi When to Trust Imagination: Adaptive Action Execution for World Action Models. arXiv preprint arXiv:2605.06222. External Links: Link
  50. Wang et al. (2026e) Z. Wang, Z. Lin, R. Li, Y. Zhang, X. Yang, S. Mi, and X. Wei Open-Loop Planning, Closed-Loop Verification: Speculative Verification for VLA. arXiv preprint arXiv:2604.02965. External Links: Link
  51. Wen et al. (2025) J. Wen, Y. Zhu, J. Li, Z. Tang, C. Shen, and F. Feng DexVLA: vision-language model with plug-in diffusion expert for general robot control. In 9th Annual Conference on Robot Learning,
  52. Wu et al. (2026) Y. Wu, H. Li, T. Hou, L. Chen, and A. Knoll React When You Need To: Event-Triggered Asynchronous Inference for VLA Policies. arXiv preprint arXiv:2609.22587. External Links: Link
  53. Xu et al. (2026a) R. Xu, X. Shan, S. Dai, Y. Wang, and J. Yu Knowing When to Stop: Adaptive Action Chunking via Internal Cross-Attention Dynamics in VLAs. arXiv preprint arXiv:2609.00908. External Links: Link
  54. Xu et al. (2026b) W. Xu, Z. Liu, L. Luo, Y. Liang, C. Yao, Q. Mei, J. Cao, X. Cao, X. Zhang, J. Yang, and B. Guo Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution. arXiv preprint arXiv:2608.03483. External Links: Link
  55. Yang et al. (2026) Z. Yang, Y. Shi, M. Yao, W. Xue, Y. Jueluo, and L. Liu ChunkFlow: Towards Continuity-Consistent Chunked Policy Learning. arXiv preprint arXiv:2607.12992. External Links: Link
  56. Ye et al. (2026a) J. Ye, N. Gao, S. Yang, J. Zheng, Z. Wang, Y. Chen, P. Chen, Y. Chen, S. Liu, and J. Jia StarVLA-\alpha: reducing complexity in vision-language-action systems. arXiv preprint arXiv:2604.11757.
  57. Ye et al. (2026b) S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y. L. Tan, C. Zhu, J. Xiang, A. N. Malik, K. Lee, W. Liang, N. R. Arachchige, J. Gu, Y. Xu, G. Wang, F. Hu, A. Narayan, J. Bjorck, J. Wang, G. Kim, D. Niu, R. Zheng, Y. Xie, J. Wu, Q. Wang, D. Xu, Y. Du, R. Julian, Y. Chebotar, S. Reed, J. Kautz, Y. Zhu, L. Fan, and J. Jang World action models are zero-shot policies. In ICLR 2026 the 2nd Workshop on World Models: Understanding, Modelling and Scaling, External Links: Link
  58. Yuan et al. (2026) T. Yuan, Z. Dong, Y. Liu, and H. Zhao Fast-wam: do world action models need test-time future imagination?. arXiv preprint arXiv:2603.16666.
  59. Zeng et al. (2026) M. Zeng, A. Agarwal, A. Bati, B. Lee, S. Ancha, and R. Tedrake Revisiting Open-Loop Execution in Robotics: Toward Reactive, Higher-Performing Policies. arXiv preprint arXiv:2608.15938. External Links: Link
  60. Zhan et al. (2026) D. Zhan, X. Xu, J. Li, and J. Tang SEAM: Smooth Execution of Action-Chunked Motion for Vision-Language-Action Policies. arXiv preprint arXiv:2607.04609. External Links: Link
  61. Zhang et al. (2026a) J. Zhang, R. Liu, Y. Zhang, and Y. Yang TraceFlow: Guiding Frozen Flow-Matching Robot Policies with Success and Failure Traces. arXiv preprint arXiv:2609.20646. External Links: Link
  62. Zhang et al. (2026b) J. Zhang, Z. Han, J. Wang, X. Wu, S. Lin, J. Li, H. Fan, R. Wu, D. Li, and H. Dong HiPolicy: hierarchical multi-frequency action chunking for policy learning. arXiv preprint arXiv:2604.06067.
  63. Zhao et al. (2026) Y. Zhao, M. Bogdanovic, A. Sohal, L. Tao, K. Darvish, A. Aspuru-Guzik, F. Shkurti, and A. Garg Dynamic Execution Horizon Prediction for Chunk-based Robot Policies. arXiv preprint arXiv:2606.11408. External Links: Link
  64. Zheng et al. (2026) J. Zheng, J. Li, Z. Wang, D. Liu, X. Kang, Y. Feng, Y. Zheng, J. Zou, Y. Chen, J. Zeng, T. Wang, Y. Zhang, J. Liu, and X. Zhan X-VLA: soft-prompted transformer as scalable cross-embodiment vision-language-action model. In The Fourteenth International Conference on Learning Representations, External Links: Link
  65. Zhou et al. (2026) Y. Zhou, R. Qiu, Y. Chen, J. Cui, and W. Zhi PATCH: Action-Chunk-Conditioned Latent Patch Innovation Monitoring for Robot Manipulation. arXiv preprint arXiv:2606.16690. External Links: Link
  66. Zitkovich et al. (2023) B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V. Vanhoucke, H. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. Sanketi, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y. Lu, S. Levine, L. Lee, T. E. Lee, I. Leal, Y. Kuang, D. Kalashnikov, R. Julian, N. J. Joshi, A. Irpan, B. Ichter, J. Hsu, A. Herzog, K. Hausman, K. Gopalakrishnan, C. Fu, P. Florence, C. Finn, K. A. Dubey, D. Driess, T. Ding, K. M. Choromanski, X. Chen, Y. Chebotar, J. Carbajal, N. Brown, A. Brohan, M. G. Arenas, and K. Han RT-2: vision-language-action models transfer web knowledge to robotic control. In Proceedings of The 7th Conference on Robot Learning, J. Tan, M. Toussaint, and K. Darvish (Eds.), Proceedings of Machine Learning Research, Vol. 229, pp. 2165–2183. External Links: Link

ChunkTrust: Adapting Execution Horizons for Robot Policies with Action-Expert Evidence Appendix

Appendix A Experimental Setup

We specify evaluation cohorts, checkpoints, scoring criteria, and implementation settings for the results in Sec. 5.

A.1 Real-world Experiments

Hardware Setup.

We conduct real-world experiments on an AgileX COBOT Magic platform configured as an ALOHA-style bimanual system (Fu et al., 2024a; Fu et al., 2024b), as shown in Fig. 6. The platform consists of four 6-DoF Piper arms, with two leader arms used for human teleoperation and two follower arms used for data collection and autonomous policy execution. The perception system includes three RealSense D435 cameras: one front-view camera and two wrist-mounted cameras, one on each follower arm.

Figure 6: Real-world robot platform. AgileX COBOT Magic configured as an ALOHA-style bimanual system for teleoperation, data collection, and autonomous policy execution.
Tasks and evaluation.

We collect 200 human-teleoperated demonstrations for each of four bimanual tasks and evaluate \pi_{0.5} with fixed K=25 or AHS over 15 rollouts per task. Data are recorded at 30 FPS, with task-relevant object positions varied across evaluation rollouts. Instructions are “Fold the towel with both arms,” “Move the bread to the plate with both arms,” “Move the drink to the basket with both arms,” and “Put the duck toy into the drawer with both arms.” These tasks cover approach, contact, transport, alignment, and placement (Fig. 4).

Process scores.

Each sub-step receives 0 for failure, 0.5 for recovered or imperfect completion, and 1 for smooth, accurate completion. We average equally over the task’s sub-steps and its rollouts, reporting the result as a percentage. Table 5 lists the task-specific criteria. Error bars in Fig. 4(b) show \pm s/\sqrt{15}, where s is the sample standard deviation of the 15 normalized rollout scores for each task and method. In the drawer task, the arm assignment depends on the layout: one arm opens/closes the drawer and the other manipulates the toy.

Table 5: Real-world scoring criteria. Scoring is independent of arm assignment.
  • Bread: grasp
    0
    Fails to grasp
    0.5
    Multiple attempts
    1
    Smooth first attempt
  • Bread: handover
    0
    Handover fails
    0.5
    Unstable/awkward receiving grasp
    1
    Stable, aligned grasp
  • Bread: place on plate
    0
    Not placed
    0.5
    Poor alignment or rough placement
    1
    Clean placement
  • Drink: push
    0
    Only tilts, no useful displacement
    0.5
    Acceptable position, tilt/misalignment
    1
    Good position, stable alignment
  • Drink: grasp
    0
    Fails to grasp
    0.5
    Multiple attempts
    1
    Smooth first attempt
  • Drink: place in basket
    0
    Fails or drops drink
    0.5
    Rough placement/collision
    1
    Clean placement
  • Towel: first fold, second fold
    0
    Not folded over
    0.5
    Folded, misaligned
    1
    Folded, well aligned
  • Duck: open drawer, grasp toy, place toy, close drawer
    0
    Sub-step fails
    0.5
    Multiple attempts
    1
    Smooth first attempt

A.2 Simulation Experiments

RoboTwin2.0 (Chen et al., 2025) contains 50 bimanual manipulation tasks with strong domain randomization. The complete-suite evaluation in Tab. 1 uses one multitask \pi_{0.5} checkpoint on all 50 tasks, with 20 rollouts per task, setting, and method (2,000 per method). The eight-task evaluation uses 100 rollouts per task, setting, and method (1,600 per method), on Blocks Ranking RGB, Handover Block, Handover Mic, Hanging Mug, Place A2B Left, Place Bread Basket, Place Bread Skillet, and Place Can Basket. Each task is evaluated under two settings: Easy (clean scenes) and Hard (randomized object poses, lighting, and distractor placement). Results report success rate, averaged equally across the tasks and settings in each cohort.

RoboCasa GR1 Tabletop (Nasiriany et al., 2024) provides 24 pick and-place tasks spanning everyday object categories and novel source–target combinations. We evaluate all 24 tasks, using 50 rollouts per task for the \pi_{0.5} control. Results for the other policies use their respective evaluation budgets and reported numerical precision. Task success is determined by the native RoboCasa success checker.

A.3 Base Policy Checkpoints

For the eight-task RoboTwin2.0 evaluation, checkpoints follow the protocol of each base policy:

  • \pi_{0} (Black et al., 2024): we train LoRA adapters (Hu et al., 2022) with the released RoboTwin2.0 fine-tuning protocol. For each evaluated task, one adapter is trained on that task’s clean50 demonstrations and then evaluated. The action chunk size is H=50, with a flow-matching action expert (Lipman et al., 2023) using 10 sampling steps.
  • \pi_{0.5} (Intelligence et al., 2025): we use the official RoboTwin2.0 fine-tuning recipe for full-parameter fine-tuning. For each evaluated task, one task-specific checkpoint is trained on clean50 demonstrations and then evaluated. The action chunk size is H=50, with a flow-matching action expert (Lipman et al., 2023).
  • X-VLA (Zheng et al., 2026): we evaluate the released X-VLA RoboTwin2.0 checkpoint. The action chunk size is H=30, with a flow-matching action expert.
  • Fast-WAM (Yuan et al., 2026): we evaluate the released Fast-WAM multitask RoboTwin2.0 checkpoint on our eight-task subset, rather than post-training a separate policy for each task. The action chunk size is H=32, with a flow-matching action expert.
Multitask \pi_{0.5} checkpoints.

The complete-suite controls initialize from the official pi05_base and use benchmark-specific full-parameter post-training. RoboTwin2.0 uses 50 clean demonstrations per task (2,500 trajectories and 549,787 frames). RoboCasa GR1 uses 24,000 demonstrations and 6,020,058 frames across 24 tasks. Both use task-uniform sampling, one global quantile normalizer per benchmark, global batch size 256, seed 42, and eight H100 GPUs for training, with the checkpoint fixed at step 30,000. RoboTwin uses H=50 and Base K=50, while RoboCasa uses H=16 and Base K=16. The latter maps the raw 44-dimensional GR1 interface to 29 effective absolute controls in the padded policy interface. These controls share a policy architecture, not weights across benchmarks. Within each benchmark, Base and AHS use identical policy weights. Their evaluation runs on one RTX 4090.

For the other RoboCasa policies, Isaac-GR00T N1.5 and Isaac-GR00T N1.6 (Bjorck et al., 2025) use released base checkpoints without additional benchmark post-training. Qwen3GR00T uses the released 24-task multitask-post-trained GR1 checkpoint from StarVLA (Ye et al., 2026a; Community, 2026). Thus, “zero-shot” in Tab. 1 describes benchmark adaptation, not an absence of robot pretraining.

A.4 AHS Hyperparameter Configurations

Candidate sets and continuity windows depend on the policy and benchmark:

  • RoboTwin2.0, \pi_{0} and \pi_{0.5}
    Candidate horizons \mathcal{K}
    \{10,20,30,40,50\}
    Continuity window W_{h}
    60
  • RoboTwin2.0, X-VLA and Fast-WAM
    Candidate horizons \mathcal{K}
    \{10,20,30\}
    Continuity window W_{h}
    40
  • RoboCasa GR1, all policies
    Candidate horizons \mathcal{K}
    \{4,8,12,16\}
    Continuity window W_{h}
    20

The scaling ablation compares \mathcal{K}_{5}=\{10,20,30,40,50\}, \mathcal{K}_{10}=\{5,10,\ldots,50\}, and \mathcal{K}_{50}=\{1,\ldots,50\}. Shared defaults are \alpha_{\mathrm{intra}}=\beta_{\mathrm{inter}}=4, \lambda=0.5, a_{0}=b_{0}=1, \rho_{f}=0.99, \eta=1, and \sigma_{K}=10. The eight-task study uses expected-round selection at T_{\mathrm{sel}}=0.8, with Thompson exploration probability \epsilon_{\mathrm{exp}}=0.05. The 50-task evaluation uses T_{\mathrm{sel}}=1.0, and the RoboCasa \pi_{0.5} control uses 0.8.

Spectral scoring zero-pads each velocity prefix to H before FFT, uses the first W_{\tau}=3 denoising steps as its reference, window RMS for z, and cutoff fraction c_{\mathrm{cut}}=0.25. Without sufficient executed history, AHS uses intra-chunk-only scoring (\lambda=0). Costs appear in Appendix D.1.

A.5 Evaluation Protocol

For every (base policy, task, setting) combination, we run a fixed number of rollouts with AHS enabled. Each rollout starts from the standard initial state distribution of the benchmark. The online AHS posterior is reset to its prior (a_{0}=b_{0}=1) at the beginning of each episode, so its rollout history is episode-local. When QHA is used, its learned weights are reused across episodes without online training. Success is determined by the benchmark’s native success checker. We report the raw success rate (percentage of successful rollouts) and equal-weight averages over the stated task and setting groups. Absolute differences are in percentage points. The \Delta (pp) columns subtract the corresponding success-rate percentages, rather than reporting relative percentage changes.

In the 50-task evaluation, 62 of the 100 task–setting conditions use exact same-seed and instruction pairing between Base and AHS. The other 38 use deterministic, method-specific fallback identities selected without outcome information after repeated scene-construction failures. The full-suite averages include both groups. Identical episode identities are not assumed for the fallback group.

For the main RoboTwin2.0 benchmarks, Fast-WAM inference uses one NVIDIA H100 (80 GB); \pi_{0}, \pi_{0.5}, and X-VLA use one RTX 4090 (24 GB). The RoboCasa \pi_{0.5} control also uses one RTX 4090. Hardware details for separate runtime profiles are discussed in Appendix D.1. Replanning frequency depends on the selected horizon and execution scheduler.

Figure 7: Full evidence distributions and boundary behavior. (a,b) Episode means of the operational scores, with violin marks indicating medians. Widths show within-group density, not relative sample counts. (c,d) Normalized speed profiles for candidate k=10 and k=50, with dashed vertical lines marking the history–prediction join. Curves show means and bands show P10–P90 across episode profiles, not confidence intervals or the IQR bands used in the main figure.

A.6 QHA Training Configurations

Training-task protocol (Tab. 2A).

We train one shared QHA per policy family (\pi_{0} or \pi_{0.5}) across all eight tasks, keeping the task-specific base generators frozen. Training uses their clean50 Aloha-AgileX LeRobot datasets, with task-balanced batches of 256 (32 per task). Inputs comprise three RGB streams, robot state, action windows, VLM context tokens \mathbf{C}_{t} with mask \mathbf{M}_{t}, and action-latent tokens \mathbf{Z}_{t}^{A} from the final chunk. The action-query window includes 64 history and 50 future steps. The frozen checkpoint generates the chunk \hat{\mathbf{A}}_{t}, evidence \mathcal{F}_{t}, and dense teacher \mu_{t}^{\star} online. The teacher follows Eq. 13 over h=1,\ldots,50 at T_{\mathrm{teach}}=1, using current evidence without episode-local Beta memory. The six-task held-out protocol is specified separately in Appendix C.1.

Architecture and optimization.

Context and action tokens are projected into a shared d=256 space with positional encodings. One bridge layer updates eight learned queries by cross-attending to masked context, then action tokens, followed by self-attention, each with residual connections (Vaswani et al., 2017). It uses eight attention heads, zero dropout, and an enabled fusion gate. Query pooling and an MLP produce 50 logits followed by a softmax. We minimize \operatorname{KL}(\mu_{t}^{\star}\|p_{\phi}) with weight 1 and numerical floor 10^{-6}. Only QHA parameters are updated and the base flow-matching loss is zero. AdamW runs for 10,000 steps with batch size 256, global-norm clipping 1, weight decay 10^{-10}, and EMA 0.99. The warmup-cosine schedule uses 1,000 warmup steps, peak learning rate 5\times 10^{-5}, and final rate 5\times 10^{-6}.

Deployment (Tab. 2A).

Each policy family uses its QHA checkpoint saved at step 10,000 after joint training on the eight tasks, with prior strength \gamma=1 in Eq. 14. Online evidence q_{\mathrm{mix},t}(k) is computed on \mathcal{K}=\{10,20,30,40,50\} and interpolated onto \{1,\ldots,H\} before Beta sampling. Both this sparse-evidence variant and the dense-evidence held-out variant maintain dense Beta states and apply temperature scaling, expected-round selection, and uniform exploration as in Sec. 4.1. Posterior feedback uses prior-weighted quality \tilde{q}_{\mathrm{mix},t}(h)p_{\phi}(h)^{\gamma}.

Held-out deployment (Tab. 2B).

This evaluation instead computes dense evidence over h=1,\ldots,50, with \gamma=1 and expected-round selection at T_{\mathrm{sel}}=1. Appendix C.1 specifies the six-task training split and checkpoints.

Figure 8: Joint-risk landscape on the same 1,600 episodes. Dashed lines mark the original median thresholds. Points denote episodes, with outcome colors as in Fig. 7. Color shows the smoothed local failure fraction. Lower opacity denotes lower smoothed occupancy. The color scale saturates below 0.25 and above 0.95. The map is descriptive. Group failure rates below are computed directly from episode outcomes, not read from the smoothed colors.

Appendix B Action-Expert Evidence Diagnostics

B.1 Evidence Distributions and Boundary Profiles

Data and aggregation.

We analyze the same 1,600 \pi_{0} RoboTwin2.0 episodes as Fig. 2: eight tasks, two settings, and 100 episodes per task–setting pair, with 291 successes and 1,309 failures. All episodes execute fixed K=H=50. The candidates k\in\{10,20,30,40,50\} are scored on these recorded traces. They are not five separate executed-horizon experiments.

For each candidate, we average finite scores over replans within an episode, then report medians and IQRs over episode means. The valid counts can differ between the two metrics because inter-chunk evidence requires executed history. The intra score follows Eq. 5, with prefixes zero-padded to H, a three-step baseline, and RMS over all ten denoising steps. Recomputed scores agree with the recorded values within 1.5\times 10^{-8}.

For boundary profiles, we compute first-difference speeds in the W_{h}=60 stitched window of Eq. 6, normalize by the window mean, and average within each episode before summarizing across episodes. This prevents longer episodes receiving extra weight.

B.2 Joint Risk and Episode Outcomes

At the executed K=50, each episode’s intra/inter risk is the 75th percentile of 1-q_{\mathrm{intra},t}(50) or 1-q_{\mathrm{inter},t}(50) over valid replans. Pooled median thresholds are 0.98168436 and 0.96278395, respectively. Values strictly below the threshold are low and ties are high, so group sizes can differ.

Table 6: Exact median-split group outcomes. Counts and failure rates reproduce Fig. 2(c).
  • Both low
    Intra risk
    Low
    Inter risk
    Low
    Episodes
    313
    Failures
    157
    Failure rate
    50.2%
  • Inter only
    Intra risk
    Low
    Inter risk
    High
    Episodes
    193
    Failures
    144
    Failure rate
    74.6%
  • Intra only
    Intra risk
    High
    Inter risk
    Low
    Episodes
    487
    Failures
    417
    Failure rate
    85.6%
  • Both high
    Intra risk
    High
    Inter risk
    High
    Episodes
    607
    Failures
    591
    Failure rate
    97.4%

The map uses a 23\times 23 grid with 7% range padding and separable kernel [1,4,6,4,1]/16. Success/failure counts are smoothed separately before division (stabilizer 10^{-12}). Opacity scales to the 90th occupancy percentile, hiding cells below 10^{-3}.

Scope of the evidence.

These are pooled, retrospective associations from one policy and a fixed execution horizon. They do not establish causality, calibrated failure prediction, within-task effects independent of task difficulty, or how long a particular prefix remains reliable at the current replan. The benefit of adaptive execution is evaluated separately in the policy comparisons and ablations, not inferred from this heatmap.

Appendix C Full Simulation Benchmark Results

C.1 QHA Transfer to Held-out Tasks

For Tab. 2B, QHA is trained on Handover Block, Handover Mic, Hanging Mug, Place A2B Left, Place Bread Skillet, and Place Can Basket. Blocks Ranking RGB and Place Bread Basket are held out. The QHA head uses step 10,000, and frozen task-specific \pi_{0.5} bases use step-20,000 clean50 checkpoints without quantile normalization. All selectors use \{1,\ldots,50\} and expected-round selection at temperature 1.0. AHS+QHA uses \gamma=1.0. This cohort differs from the eight-task study in Tab. 2A.

AHS, QHA-only, and AHS+QHA share 100 episode identities per task and setting (400 per method). The Base results in Tab. 2B come from a separate evaluation outside this paired cohort. The transfer comparison is AHS+QHA versus AHS. Table 7 retains all four conditions: fusion improves two and reduces success in two. The aggregate gain does not establish uniform improvement or transfer of the base policy.

Table 7: Complete held-out comparison. Success rate (%) for the \pi_{0.5} policy with QHA trained on six tasks and evaluated on two held-out tasks. Each task and setting uses 100 paired episodes. \Delta=\mathrm{SR}_{\mathrm{AHS+QHA}}-\mathrm{SR}_{\mathrm{AHS}}.
  • Blocks Ranking RGB
    Setting
    Easy
    AHS
    44.00
    QHA-only
    48.00
    AHS+QHA
    57.00
    \Delta (pp)
    +13.00
  • Blocks Ranking RGB
    Setting
    Hard
    AHS
    33.00
    QHA-only
    29.00
    AHS+QHA
    30.00
    \Delta (pp)
    -3.00
  • Place Bread Basket
    Setting
    Easy
    AHS
    52.00
    QHA-only
    50.00
    AHS+QHA
    49.00
    \Delta (pp)
    -3.00
  • Place Bread Basket
    Setting
    Hard
    AHS
    31.00
    QHA-only
    31.00
    AHS+QHA
    35.00
    \Delta (pp)
    +4.00
  • Overall
    Setting
    Both
    AHS
    40.00
    QHA-only
    39.50
    AHS+QHA
    42.75
    \Delta (pp)
    +2.75
Same-state decision disagreement.

In AHS+QHA traces, the QHA and AHS component choices differ at 77.02% of 12,573 replans, with a mean absolute gap of 1.83 action steps. Easy contributes 5,533 replans (77.17%, 1.84 steps) and Hard 7,040 (76.90%, 1.82 steps). This measures differences at the same states, without establishing complementarity or explaining success gains.

C.2 Complete 50-task RoboTwin2.0 Evaluation

Table 8: Complete 50-task RoboTwin2.0 results. Success rate (%) for the shared multitask \pi_{0.5} checkpoint, with 20 episodes per task, setting, and method. \Delta is AHS minus Base, averaged over Easy and Hard.
  • —
    Easy
    Base
    AHS
    Hard
    Base
    AHS
    \Delta (pp)
    Keine Daten
  • adjust bottle
    Easy
    100.00
    100.00
    Hard
    90.00
    95.00
    \Delta (pp)
    +2.50
  • beat block hammer
    Easy
    70.00
    80.00
    Hard
    15.00
    50.00
    \Delta (pp)
    +22.50
  • blocks ranking rgb
    Easy
    80.00
    95.00
    Hard
    40.00
    70.00
    \Delta (pp)
    +22.50
  • blocks ranking size
    Easy
    25.00
    45.00
    Hard
    25.00
    35.00
    \Delta (pp)
    +15.00
  • click alarmclock
    Easy
    60.00
    65.00
    Hard
    45.00
    45.00
    \Delta (pp)
    +2.50
  • click bell
    Easy
    35.00
    70.00
    Hard
    55.00
    55.00
    \Delta (pp)
    +17.50
  • dump bin bigbin
    Easy
    95.00
    95.00
    Hard
    80.00
    75.00
    \Delta (pp)
    -2.50
  • grab roller
    Easy
    100.00
    100.00
    Hard
    90.00
    90.00
    \Delta (pp)
    +0.00
  • handover block
    Easy
    65.00
    65.00
    Hard
    10.00
    35.00
    \Delta (pp)
    +12.50
  • handover mic
    Easy
    100.00
    100.00
    Hard
    15.00
    15.00
    \Delta (pp)
    +0.00
  • hanging mug
    Easy
    20.00
    30.00
    Hard
    5.00
    10.00
    \Delta (pp)
    +7.50
  • lift pot
    Easy
    65.00
    70.00
    Hard
    30.00
    35.00
    \Delta (pp)
    +5.00
  • move can pot
    Easy
    50.00
    75.00
    Hard
    10.00
    25.00
    \Delta (pp)
    +20.00
  • move pillbottle pad
    Easy
    55.00
    80.00
    Hard
    30.00
    30.00
    \Delta (pp)
    +12.50
  • move playingcard away
    Easy
    90.00
    80.00
    Hard
    75.00
    75.00
    \Delta (pp)
    -5.00
  • move stapler pad
    Easy
    20.00
    35.00
    Hard
    20.00
    5.00
    \Delta (pp)
    +0.00
  • open laptop
    Easy
    95.00
    90.00
    Hard
    85.00
    70.00
    \Delta (pp)
    -10.00
  • open microwave
    Easy
    85.00
    50.00
    Hard
    35.00
    25.00
    \Delta (pp)
    -22.50
  • pick diverse bottles
    Easy
    45.00
    55.00
    Hard
    35.00
    70.00
    \Delta (pp)
    +22.50
  • pick dual bottles
    Easy
    55.00
    70.00
    Hard
    75.00
    75.00
    \Delta (pp)
    +7.50
  • place a2b left
    Easy
    75.00
    70.00
    Hard
    55.00
    65.00
    \Delta (pp)
    +2.50
  • place a2b right
    Easy
    70.00
    75.00
    Hard
    50.00
    50.00
    \Delta (pp)
    +2.50
  • place bread basket
    Easy
    55.00
    75.00
    Hard
    55.00
    55.00
    \Delta (pp)
    +10.00
  • place bread skillet
    Easy
    35.00
    55.00
    Hard
    60.00
    55.00
    \Delta (pp)
    +7.50
  • place burger fries
    Easy
    95.00
    95.00
    Hard
    90.00
    100.00
    \Delta (pp)
    +5.00
  • place can basket
    Easy
    40.00
    85.00
    Hard
    0.00
    35.00
    \Delta (pp)
    +40.00
  • place cans plasticbox
    Easy
    40.00
    85.00
    Hard
    55.00
    80.00
    \Delta (pp)
    +35.00
  • place container plate
    Easy
    75.00
    90.00
    Hard
    80.00
    80.00
    \Delta (pp)
    +7.50
  • place dual shoes
    Easy
    75.00
    85.00
    Hard
    50.00
    60.00
    \Delta (pp)
    +10.00
  • place empty cup
    Easy
    100.00
    100.00
    Hard
    65.00
    70.00
    \Delta (pp)
    +2.50
  • place fan
    Easy
    75.00
    75.00
    Hard
    45.00
    40.00
    \Delta (pp)
    -2.50
  • place mouse pad
    Easy
    15.00
    45.00
    Hard
    25.00
    25.00
    \Delta (pp)
    +15.00
  • place object basket
    Easy
    50.00
    60.00
    Hard
    20.00
    35.00
    \Delta (pp)
    +12.50
  • place object scale
    Easy
    60.00
    85.00
    Hard
    45.00
    50.00
    \Delta (pp)
    +15.00
  • place object stand
    Easy
    85.00
    80.00
    Hard
    55.00
    90.00
    \Delta (pp)
    +15.00
  • place phone stand
    Easy
    65.00
    65.00
    Hard
    45.00
    35.00
    \Delta (pp)
    -5.00
  • place shoe
    Easy
    85.00
    90.00
    Hard
    75.00
    55.00
    \Delta (pp)
    -7.50
  • press stapler
    Easy
    100.00
    90.00
    Hard
    60.00
    70.00
    \Delta (pp)
    +0.00
  • put bottles dustbin
    Easy
    45.00
    65.00
    Hard
    50.00
    55.00
    \Delta (pp)
    +12.50
  • put object cabinet
    Easy
    25.00
    35.00
    Hard
    20.00
    15.00
    \Delta (pp)
    +2.50
  • rotate qrcode
    Easy
    80.00
    85.00
    Hard
    30.00
    30.00
    \Delta (pp)
    +2.50
  • scan object
    Easy
    20.00
    40.00
    Hard
    20.00
    50.00
    \Delta (pp)
    +25.00
  • shake bottle
    Easy
    100.00
    100.00
    Hard
    100.00
    100.00
    \Delta (pp)
    +0.00
  • shake bottle horizontally
    Easy
    100.00
    100.00
    Hard
    100.00
    100.00
    \Delta (pp)
    +0.00
  • stack blocks three
    Easy
    75.00
    55.00
    Hard
    35.00
    45.00
    \Delta (pp)
    -5.00
  • stack blocks two
    Easy
    85.00
    100.00
    Hard
    80.00
    90.00
    \Delta (pp)
    +12.50
  • stack bowls three
    Easy
    60.00
    65.00
    Hard
    35.00
    30.00
    \Delta (pp)
    +0.00
  • stack bowls two
    Easy
    100.00
    85.00
    Hard
    85.00
    90.00
    \Delta (pp)
    -5.00
  • stamp seal
    Easy
    35.00
    30.00
    Hard
    25.00
    30.00
    \Delta (pp)
    +0.00
  • turn switch
    Easy
    40.00
    40.00
    Hard
    25.00
    25.00
    \Delta (pp)
    +0.00
  • Mean by setting
    Easy
    65.40
    73.10
    Hard
    48.00
    53.90
    \Delta (pp)
    +6.80
  • Overall
    Easy
    Base: 56.70
    Keine Daten
    Hard
    AHS: 63.50
    Keine Daten
    \Delta (pp)
    +6.80

Table 8 expands the multitask \pi_{0.5} row of Tab. 1. Base and AHS succeed in 1,134 and 1,270 of 2,000 episodes, respectively, giving 56.70% and 63.50% overall. Easy success increases from 65.40% to 73.10%, and Hard from 48.00% to 53.90%. All 50 tasks are retained, including nine tasks with a negative change after averaging the two settings.

Statistical robustness and pairing scope.

The full-matrix gain is 6.80 percentage points, with a task-cluster 95% bootstrap interval of [3.75,\,9.95]. This interval resamples the 50 tasks, retaining both settings and methods within each task. Only 62 of 100 task–setting cells preserve exact seed-and-instruction pairing, while the remaining 38 cells use deterministic, outcome-blind, method-specific fallback identities after scene-construction failures. The 1,240 exact episode pairs yield a gain of 7.58 percentage points with a paired 95% interval of [5.08,\,10.08], resampling episodes within these cells. Both intervals use 10,000 bootstrap draws, but their statistical units and populations differ: the paired interval applies only to the exact-identity subset, not the full matrix.

C.3 Eight-task Policy-family Evaluation

The task-specific \pi_{0} and \pi_{0.5} matrices appear in Tab. 2A. Table 9 provides the corresponding Base–AHS breakdown for Fast-WAM and X-VLA (Zheng et al., 2026), evaluated on the same eight tasks with 100 rollouts per task and setting. Both use \mathcal{K}=\{10,20,30\}. Each Base–AHS comparison uses the same released policy checkpoint. Fast-WAM’s overall success increases from 86.81% to 88.13%. For X-VLA, AHS slightly improves overall success from 47.44% to 47.75%: Easy increases from 74.50% to 75.75%, while Hard decreases from 20.38% to 19.75%. Individual task regressions are retained for both policies.

Table 9: Fast-WAM and X-VLA per-task success rates on RoboTwin2.0. Success rate (%) over 100 rollouts per task and setting. Easy and Hard denote clean and randomized settings. Each Base–AHS comparison fixes the policy checkpoint. Yellow columns use AHS, and bold marks the better method within each policy and setting, including ties.
  • Task
    Fast-WAM (Yuan et al., 2026)
    Base
    Keine Daten
    AHS
    Keine Daten
    X-VLA (Zheng et al., 2026)
    Base
    Keine Daten
    AHS
    Keine Daten
  • —
    Fast-WAM (Yuan et al., 2026)
    Easy
    Hard
    Easy
    Hard
    X-VLA (Zheng et al., 2026)
    Easy
    Hard
    Easy
    Hard
  • Blocks Ranking RGB
    Fast-WAM (Yuan et al., 2026)
    99.00
    98.00
    100.00
    100.00
    X-VLA (Zheng et al., 2026)
    87.00
    34.00
    95.00
    38.00
  • Handover Block
    Fast-WAM (Yuan et al., 2026)
    94.00
    81.00
    91.00
    82.00
    X-VLA (Zheng et al., 2026)
    89.00
    2.00
    88.00
    1.00
  • Handover Mic
    Fast-WAM (Yuan et al., 2026)
    100.00
    99.00
    100.00
    100.00
    X-VLA (Zheng et al., 2026)
    97.00
    1.00
    100.00
    1.00
  • Hanging Mug
    Fast-WAM (Yuan et al., 2026)
    66.00
    64.00
    71.00
    69.00
    X-VLA (Zheng et al., 2026)
    32.00
    8.00
    35.00
    4.00
  • Place A2B Left
    Fast-WAM (Yuan et al., 2026)
    95.00
    95.00
    93.00
    93.00
    X-VLA (Zheng et al., 2026)
    31.00
    25.00
    40.00
    21.00
  • Place Bread Basket
    Fast-WAM (Yuan et al., 2026)
    91.00
    91.00
    93.00
    94.00
    X-VLA (Zheng et al., 2026)
    87.00
    41.00
    80.00
    42.00
  • Place Bread Skillet
    Fast-WAM (Yuan et al., 2026)
    91.00
    91.00
    94.00
    94.00
    X-VLA (Zheng et al., 2026)
    86.00
    22.00
    88.00
    23.00
  • Place Can Basket
    Fast-WAM (Yuan et al., 2026)
    71.00
    63.00
    69.00
    67.00
    X-VLA (Zheng et al., 2026)
    87.00
    30.00
    80.00
    28.00
  • Mean by setting
    Fast-WAM (Yuan et al., 2026)
    88.38
    85.25
    88.88
    87.38
    X-VLA (Zheng et al., 2026)
    74.50
    20.38
    75.75
    19.75
  • Overall
    Fast-WAM (Yuan et al., 2026)
    86.81
    Keine Daten
    88.13
    Keine Daten
    X-VLA (Zheng et al., 2026)
    47.44
    Keine Daten
    47.75
    Keine Daten

C.4 RoboCasa GR1 Tabletop

Table 10 reports the full 24-task RoboCasa GR1 Tabletop breakdown for \pi_{0.5} and the three GR00T-family policies summarized in Tab. 1. QwenFAST (discrete tokens) and QwenPI (flow-matching action expert) are additional StarVLA baselines (Ye et al., 2026a; Community, 2026), both using Qwen3VL (Bai et al., 2025). Neither has an AHS counterpart in this table. For \pi_{0.5}, the 24-task means are 40.08% for Base and 42.50% for AHS, matching Tab. 1.

Table 10: Full per-task success rates on RoboCasa GR1 Tabletop. Full 24-task breakdown, with 50 rollouts per task for \pi_{0.5}. Yellow columns use AHS and parentheses report percentage-point deltas over Base. Bold marks the best result in each row, including ties.
  • PnPBottleToCabinetClose
    QwenFAST +Qwen3VL
    38.0
    QwenPI +Qwen3VL
    26.0
    Isaac-GR00T N1.5 Base
    64.0
    Isaac-GR00T N1.5 AHS
    68.0 (+4.0)
    Isaac-GR00T N1.6 Base
    51.5
    Isaac-GR00T N1.6 AHS
    54.0 (+2.5)
    Qwen3GR00T +Qwen3VL Base
    46.0
    Qwen3GR00T +Qwen3VL AHS
    64.0 (+18.0)
    \boldsymbol{\pi_{0.5}} Base
    64.00
    \boldsymbol{\pi_{0.5}} AHS
    68.00 (+4.00)
  • PnPCanToDrawerClose
    QwenFAST +Qwen3VL
    44.0
    QwenPI +Qwen3VL
    62.0
    Isaac-GR00T N1.5 Base
    18.0
    Isaac-GR00T N1.5 AHS
    12.0 (-6.0)
    Isaac-GR00T N1.6 Base
    13.0
    Isaac-GR00T N1.6 AHS
    12.0 (-1.0)
    Qwen3GR00T +Qwen3VL Base
    80.0
    Qwen3GR00T +Qwen3VL AHS
    80.0 (+0.0)
    \boldsymbol{\pi_{0.5}} Base
    58.00
    \boldsymbol{\pi_{0.5}} AHS
    56.00 (-2.00)
  • PnPCupToDrawerClose
    QwenFAST +Qwen3VL
    56.0
    QwenPI +Qwen3VL
    42.0
    Isaac-GR00T N1.5 Base
    12.0
    Isaac-GR00T N1.5 AHS
    4.0 (-8.0)
    Isaac-GR00T N1.6 Base
    8.5
    Isaac-GR00T N1.6 AHS
    14.0 (+5.5)
    Qwen3GR00T +Qwen3VL Base
    54.0
    Qwen3GR00T +Qwen3VL AHS
    52.0 (-2.0)
    \boldsymbol{\pi_{0.5}} Base
    34.00
    \boldsymbol{\pi_{0.5}} AHS
    40.00 (+6.00)
  • PnPMilkToMicrowaveClose
    QwenFAST +Qwen3VL
    44.0
    QwenPI +Qwen3VL
    50.0
    Isaac-GR00T N1.5 Base
    38.0
    Isaac-GR00T N1.5 AHS
    34.0 (-4.0)
    Isaac-GR00T N1.6 Base
    14.0
    Isaac-GR00T N1.6 AHS
    20.0 (+6.0)
    Qwen3GR00T +Qwen3VL Base
    48.0
    Qwen3GR00T +Qwen3VL AHS
    42.0 (-6.0)
    \boldsymbol{\pi_{0.5}} Base
    40.00
    \boldsymbol{\pi_{0.5}} AHS
    44.00 (+4.00)
  • PnPPotatoToMicrowaveClose
    QwenFAST +Qwen3VL
    14.0
    QwenPI +Qwen3VL
    42.0
    Isaac-GR00T N1.5 Base
    54.0
    Isaac-GR00T N1.5 AHS
    36.0 (-18.0)
    Isaac-GR00T N1.6 Base
    41.5
    Isaac-GR00T N1.6 AHS
    50.0 (+8.5)
    Qwen3GR00T +Qwen3VL Base
    28.0
    Qwen3GR00T +Qwen3VL AHS
    28.0 (+0.0)
    \boldsymbol{\pi_{0.5}} Base
    22.00
    \boldsymbol{\pi_{0.5}} AHS
    30.00 (+8.00)
  • PnPWineToCabinetClose
    QwenFAST +Qwen3VL
    14.0
    QwenPI +Qwen3VL
    32.0
    Isaac-GR00T N1.5 Base
    16.0
    Isaac-GR00T N1.5 AHS
    20.0 (+4.0)
    Isaac-GR00T N1.6 Base
    16.5
    Isaac-GR00T N1.6 AHS
    24.0 (+7.5)
    Qwen3GR00T +Qwen3VL Base
    46.0
    Qwen3GR00T +Qwen3VL AHS
    52.0 (+6.0)
    \boldsymbol{\pi_{0.5}} Base
    52.00
    \boldsymbol{\pi_{0.5}} AHS
    56.00 (+4.00)
  • PnPNovelFromCuttingboardToBasket
    QwenFAST +Qwen3VL
    54.0
    QwenPI +Qwen3VL
    40.0
    Isaac-GR00T N1.5 Base
    50.0
    Isaac-GR00T N1.5 AHS
    52.0 (+2.0)
    Isaac-GR00T N1.6 Base
    58.0
    Isaac-GR00T N1.6 AHS
    54.0 (-4.0)
    Qwen3GR00T +Qwen3VL Base
    48.0
    Qwen3GR00T +Qwen3VL AHS
    70.0 (+22.0)
    \boldsymbol{\pi_{0.5}} Base
    34.00
    \boldsymbol{\pi_{0.5}} AHS
    34.00 (+0.00)
  • PnPNovelFromCuttingboardToCardboardbox
    QwenFAST +Qwen3VL
    42.0
    QwenPI +Qwen3VL
    46.0
    Isaac-GR00T N1.5 Base
    36.0
    Isaac-GR00T N1.5 AHS
    34.0 (-2.0)
    Isaac-GR00T N1.6 Base
    46.5
    Isaac-GR00T N1.6 AHS
    46.0 (-0.5)
    Qwen3GR00T +Qwen3VL Base
    40.0
    Qwen3GR00T +Qwen3VL AHS
    54.0 (+14.0)
    \boldsymbol{\pi_{0.5}} Base
    32.00
    \boldsymbol{\pi_{0.5}} AHS
    40.00 (+8.00)
  • PnPNovelFromCuttingboardToPan
    QwenFAST +Qwen3VL
    58.0
    QwenPI +Qwen3VL
    60.0
    Isaac-GR00T N1.5 Base
    68.0
    Isaac-GR00T N1.5 AHS
    64.0 (-4.0)
    Isaac-GR00T N1.6 Base
    68.5
    Isaac-GR00T N1.6 AHS
    80.0 (+11.5)
    Qwen3GR00T +Qwen3VL Base
    68.0
    Qwen3GR00T +Qwen3VL AHS
    80.0 (+12.0)
    \boldsymbol{\pi_{0.5}} Base
    54.00
    \boldsymbol{\pi_{0.5}} AHS
    58.00 (+4.00)
  • PnPNovelFromCuttingboardToPot
    QwenFAST +Qwen3VL
    58.0
    QwenPI +Qwen3VL
    40.0
    Isaac-GR00T N1.5 Base
    34.0
    Isaac-GR00T N1.5 AHS
    56.0 (+22.0)
    Isaac-GR00T N1.6 Base
    65.0
    Isaac-GR00T N1.6 AHS
    64.0 (-1.0)
    Qwen3GR00T +Qwen3VL Base
    52.0
    Qwen3GR00T +Qwen3VL AHS
    76.0 (+24.0)
    \boldsymbol{\pi_{0.5}} Base
    46.00
    \boldsymbol{\pi_{0.5}} AHS
    40.00 (-6.00)
  • PnPNovelFromCuttingboardToTieredbasket
    QwenFAST +Qwen3VL
    40.0
    QwenPI +Qwen3VL
    44.0
    Isaac-GR00T N1.5 Base
    46.0
    Isaac-GR00T N1.5 AHS
    32.0 (-14.0)
    Isaac-GR00T N1.6 Base
    46.5
    Isaac-GR00T N1.6 AHS
    54.0 (+7.5)
    Qwen3GR00T +Qwen3VL Base
    56.0
    Qwen3GR00T +Qwen3VL AHS
    44.0 (-12.0)
    \boldsymbol{\pi_{0.5}} Base
    22.00
    \boldsymbol{\pi_{0.5}} AHS
    28.00 (+6.00)
  • PnPNovelFromPlacematToBasket
    QwenFAST +Qwen3VL
    36.0
    QwenPI +Qwen3VL
    44.0
    Isaac-GR00T N1.5 Base
    50.0
    Isaac-GR00T N1.5 AHS
    46.0 (-4.0)
    Isaac-GR00T N1.6 Base
    58.5
    Isaac-GR00T N1.6 AHS
    48.0 (-10.5)
    Qwen3GR00T +Qwen3VL Base
    42.0
    Qwen3GR00T +Qwen3VL AHS
    54.0 (+12.0)
    \boldsymbol{\pi_{0.5}} Base
    42.00
    \boldsymbol{\pi_{0.5}} AHS
    46.00 (+4.00)
  • PnPNovelFromPlacematToBowl
    QwenFAST +Qwen3VL
    38.0
    QwenPI +Qwen3VL
    52.0
    Isaac-GR00T N1.5 Base
    50.0
    Isaac-GR00T N1.5 AHS
    62.0 (+12.0)
    Isaac-GR00T N1.6 Base
    57.5
    Isaac-GR00T N1.6 AHS
    60.0 (+2.5)
    Qwen3GR00T +Qwen3VL Base
    44.0
    Qwen3GR00T +Qwen3VL AHS
    66.0 (+22.0)
    \boldsymbol{\pi_{0.5}} Base
    34.00
    \boldsymbol{\pi_{0.5}} AHS
    32.00 (-2.00)
  • PnPNovelFromPlacematToPlate
    QwenFAST +Qwen3VL
    42.0
    QwenPI +Qwen3VL
    50.0
    Isaac-GR00T N1.5 Base
    62.0
    Isaac-GR00T N1.5 AHS
    66.0 (+4.0)
    Isaac-GR00T N1.6 Base
    63.0
    Isaac-GR00T N1.6 AHS
    82.0 (+19.0)
    Qwen3GR00T +Qwen3VL Base
    48.0
    Qwen3GR00T +Qwen3VL AHS
    72.0 (+24.0)
    \boldsymbol{\pi_{0.5}} Base
    46.00
    \boldsymbol{\pi_{0.5}} AHS
    48.00 (+2.00)
  • PnPNovelFromPlacematToTieredshelf
    QwenFAST +Qwen3VL
    18.0
    QwenPI +Qwen3VL
    28.0
    Isaac-GR00T N1.5 Base
    14.0
    Isaac-GR00T N1.5 AHS
    26.0 (+12.0)
    Isaac-GR00T N1.6 Base
    28.5
    Isaac-GR00T N1.6 AHS
    36.0 (+7.5)
    Qwen3GR00T +Qwen3VL Base
    18.0
    Qwen3GR00T +Qwen3VL AHS
    20.0 (+2.0)
    \boldsymbol{\pi_{0.5}} Base
    28.00
    \boldsymbol{\pi_{0.5}} AHS
    28.00 (+0.00)
  • PnPNovelFromPlateToBowl
    QwenFAST +Qwen3VL
    52.0
    QwenPI +Qwen3VL
    52.0
    Isaac-GR00T N1.5 Base
    58.0
    Isaac-GR00T N1.5 AHS
    58.0 (+0.0)
    Isaac-GR00T N1.6 Base
    57.0
    Isaac-GR00T N1.6 AHS
    58.0 (+1.0)
    Qwen3GR00T +Qwen3VL Base
    60.0
    Qwen3GR00T +Qwen3VL AHS
    60.0 (+0.0)
    \boldsymbol{\pi_{0.5}} Base
    40.00
    \boldsymbol{\pi_{0.5}} AHS
    44.00 (+4.00)
  • PnPNovelFromPlateToCardboardbox
    QwenFAST +Qwen3VL
    30.0
    QwenPI +Qwen3VL
    40.0
    Isaac-GR00T N1.5 Base
    40.0
    Isaac-GR00T N1.5 AHS
    48.0 (+8.0)
    Isaac-GR00T N1.6 Base
    43.5
    Isaac-GR00T N1.6 AHS
    58.0 (+14.5)
    Qwen3GR00T +Qwen3VL Base
    50.0
    Qwen3GR00T +Qwen3VL AHS
    54.0 (+4.0)
    \boldsymbol{\pi_{0.5}} Base
    28.00
    \boldsymbol{\pi_{0.5}} AHS
    30.00 (+2.00)
  • PnPNovelFromPlateToPan
    QwenFAST +Qwen3VL
    48.0
    QwenPI +Qwen3VL
    36.0
    Isaac-GR00T N1.5 Base
    44.0
    Isaac-GR00T N1.5 AHS
    48.0 (+4.0)
    Isaac-GR00T N1.6 Base
    51.0
    Isaac-GR00T N1.6 AHS
    68.0 (+17.0)
    Qwen3GR00T +Qwen3VL Base
    54.0
    Qwen3GR00T +Qwen3VL AHS
    54.0 (+0.0)
    \boldsymbol{\pi_{0.5}} Base
    32.00
    \boldsymbol{\pi_{0.5}} AHS
    36.00 (+4.00)
  • PnPNovelFromPlateToPlate
    QwenFAST +Qwen3VL
    50.0
    QwenPI +Qwen3VL
    48.0
    Isaac-GR00T N1.5 Base
    66.0
    Isaac-GR00T N1.5 AHS
    74.0 (+8.0)
    Isaac-GR00T N1.6 Base
    78.7
    Isaac-GR00T N1.6 AHS
    82.0 (+3.3)
    Qwen3GR00T +Qwen3VL Base
    70.0
    Qwen3GR00T +Qwen3VL AHS
    74.0 (+4.0)
    \boldsymbol{\pi_{0.5}} Base
    54.00
    \boldsymbol{\pi_{0.5}} AHS
    54.00 (+0.00)
  • PnPNovelFromTrayToCardboardbox
    QwenFAST +Qwen3VL
    28.0
    QwenPI +Qwen3VL
    34.0
    Isaac-GR00T N1.5 Base
    44.0
    Isaac-GR00T N1.5 AHS
    52.0 (+8.0)
    Isaac-GR00T N1.6 Base
    51.5
    Isaac-GR00T N1.6 AHS
    48.0 (-3.5)
    Qwen3GR00T +Qwen3VL Base
    38.0
    Qwen3GR00T +Qwen3VL AHS
    56.0 (+18.0)
    \boldsymbol{\pi_{0.5}} Base
    48.00
    \boldsymbol{\pi_{0.5}} AHS
    50.00 (+2.00)
  • PnPNovelFromTrayToPlate
    QwenFAST +Qwen3VL
    34.0
    QwenPI +Qwen3VL
    64.0
    Isaac-GR00T N1.5 Base
    50.0
    Isaac-GR00T N1.5 AHS
    60.0 (+10.0)
    Isaac-GR00T N1.6 Base
    71.0
    Isaac-GR00T N1.6 AHS
    68.0 (-3.0)
    Qwen3GR00T +Qwen3VL Base
    56.0
    Qwen3GR00T +Qwen3VL AHS
    62.0 (+6.0)
    \boldsymbol{\pi_{0.5}} Base
    40.00
    \boldsymbol{\pi_{0.5}} AHS
    40.00 (+0.00)
  • PnPNovelFromTrayToPot
    QwenFAST +Qwen3VL
    46.0
    QwenPI +Qwen3VL
    44.0
    Isaac-GR00T N1.5 Base
    46.0
    Isaac-GR00T N1.5 AHS
    50.0 (+4.0)
    Isaac-GR00T N1.6 Base
    64.5
    Isaac-GR00T N1.6 AHS
    64.0 (-0.5)
    Qwen3GR00T +Qwen3VL Base
    50.0
    Qwen3GR00T +Qwen3VL AHS
    66.0 (+16.0)
    \boldsymbol{\pi_{0.5}} Base
    54.00
    \boldsymbol{\pi_{0.5}} AHS
    54.00 (+0.00)
  • PnPNovelFromTrayToTieredbasket
    QwenFAST +Qwen3VL
    36.0
    QwenPI +Qwen3VL
    50.0
    Isaac-GR00T N1.5 Base
    44.0
    Isaac-GR00T N1.5 AHS
    38.0 (-6.0)
    Isaac-GR00T N1.6 Base
    57.0
    Isaac-GR00T N1.6 AHS
    56.0 (-1.0)
    Qwen3GR00T +Qwen3VL Base
    36.0
    Qwen3GR00T +Qwen3VL AHS
    56.0 (+20.0)
    \boldsymbol{\pi_{0.5}} Base
    34.00
    \boldsymbol{\pi_{0.5}} AHS
    36.00 (+2.00)
  • PnPNovelFromTrayToTieredshelf
    QwenFAST +Qwen3VL
    16.0
    QwenPI +Qwen3VL
    28.0
    Isaac-GR00T N1.5 Base
    34.0
    Isaac-GR00T N1.5 AHS
    38.0 (+4.0)
    Isaac-GR00T N1.6 Base
    31.5
    Isaac-GR00T N1.6 AHS
    34.0 (+2.5)
    Qwen3GR00T +Qwen3VL Base
    16.0
    Qwen3GR00T +Qwen3VL AHS
    44.0 (+28.0)
    \boldsymbol{\pi_{0.5}} Base
    24.00
    \boldsymbol{\pi_{0.5}} AHS
    28.00 (+4.00)
  • Average
    QwenFAST +Qwen3VL
    39.00
    QwenPI +Qwen3VL
    43.92
    Isaac-GR00T N1.5 Base
    43.25
    Isaac-GR00T N1.5 AHS
    44.92 (+1.67)
    Isaac-GR00T N1.6 Base
    47.61
    Isaac-GR00T N1.6 AHS
    51.42 (+3.80)
    Qwen3GR00T +Qwen3VL Base
    47.83
    Qwen3GR00T +Qwen3VL AHS
    57.50 (+9.67)
    \boldsymbol{\pi_{0.5}} Base
    40.08
    \boldsymbol{\pi_{0.5}} AHS
    42.50 (+2.42)

C.5 Horizon-selection Comparisons

Table 11 reports additional controls for horizon selection, separately from the cross-policy results in Tab. 1.

Fixed, action-only, and evidence-update selectors.

The \pi_{0} study uses Place A2B Left, Place Bread Basket, Place Bread Skillet, and Place Can Basket under Easy and Hard settings, with 16 paired episodes per task and setting (128 per selector). Instantaneous is a separate reference. This cohort differs from the 100-rollout studies in Tabs. 3 and 4. The global fixed horizon K=20 is selected retrospectively from previous evaluations, not from a held-out validation set. Jerk-min uses an action-only smoothness criterion on \mathcal{K}=\{10,20,30,40,50\}.

The evidence-update variants share this candidate grid and the same instantaneous evidence, using expected-round selection with temperature 0.8. Instantaneous uses q_{\mathrm{mix},t}(k), whereas EMA maintains m_{t}(k)=\rho_{\mathrm{EMA}}m_{t-1}(k)+(1-\rho_{\mathrm{EMA}})q_{\mathrm{mix},t}(k) with \rho_{\mathrm{EMA}}=0.99. Neither comparator uses Beta counts or kernel neighborhood sharing. Full Beta is the shared AHS reference. Results describe success and call-count trade-offs without exact compute matching. The paired 95% SR-difference interval between AHS and Jerk-min includes zero, so this compact study does not establish an SR advantage.

Table 11: Horizon-selection success rates and computational costs. RoboTwin2.0 uses \pi_{0} on four tasks under Easy and Hard (128 episodes per selector). Instantaneous is a separate reference without timing measurements. Inference and wall times are seconds per episode. RoboCasa references are non-paired. Bold/underline mark best/second-best SR among the displayed methods within each benchmark.
  • Global fixed (K=20)
    RoboTwin2.0 SR (%) \uparrow
    14.84
    RoboTwin2.0 Calls/ep
    25.96
    RoboTwin2.0 Policy infer
    2.75
    RoboTwin2.0 Total infer
    2.75
    RoboTwin2.0 Wall
    54.33
  • Jerk-min
    RoboTwin2.0 SR (%) \uparrow
    14.06
    RoboTwin2.0 Calls/ep
    21.14
    RoboTwin2.0 Policy infer
    2.28
    RoboTwin2.0 Total infer
    2.28
    RoboTwin2.0 Wall
    52.01
  • EMA q_{\mathrm{mix}}
    RoboTwin2.0 SR (%) \uparrow
    10.94
    RoboTwin2.0 Calls/ep
    19.72
    RoboTwin2.0 Policy infer
    2.05
    RoboTwin2.0 Total infer
    2.07
    RoboTwin2.0 Wall
    46.72
  • AHS (Full Beta)
    RoboTwin2.0 SR (%) \uparrow
    15.63
    RoboTwin2.0 Calls/ep
    20.71
    RoboTwin2.0 Policy infer
    2.26
    RoboTwin2.0 Total infer
    2.28
    RoboTwin2.0 Wall
    53.47
  • Instantaneous q_{\mathrm{mix}} (reference)
    RoboTwin2.0 SR (%) \uparrow
    13.28
    RoboTwin2.0 Calls/ep
    19.58
    RoboTwin2.0 Policy infer
    —
    RoboTwin2.0 Total infer
    —
    RoboTwin2.0 Wall
    —
AAC-core on RoboCasa.

We evaluate a joint-space adapter of AAC (Liang et al., 2026) on Qwen3GR00T over 24 tasks with 50 rollouts per task. Each replan draws 20 action-head samples under the same observation and instruction. A prefix-entropy elbow and a movement guard determine K\in\{2,\ldots,16\}, and the first sampled chunk supplies the executed actions. The adapter operates on the policy’s native 29-dimensional absolute joint targets. It is not an exact reproduction of the published Cartesian-action implementation. AAC-core obtains 52.58% success. Base/AHS values from separate evaluations in Tab. 1 provide non-paired context only. We do not report a paired difference or infer compute equivalence from policy call counts.

Appendix D Runtime and Latency

D.1 Policy-call and Runtime Costs

We group runtime measurements by timing scope. Means include all episodes, including failures. Policy inference excludes separately timed horizon selection, but total inference includes it. Episode wall time includes simulation and within-episode overhead, not physical robot execution time. These are descriptive profiles, not hardware-matched cross-policy rankings.

Compact-selector timings are included with their success rates in Tab. 11.

Table 12: Policy-inference and episode costs for \pi_{0.5}. RoboTwin2.0 means include failures. Panels use separate cohorts. Panel B also differs from the success-rate evaluation in Tab. 2B.
  • Method
    Calls/
    episode
    Policy inference
    (s/episode)
    Episode wall
    time (s/episode)
  • A. Full 50-task evaluation 2,000 episodes per method
    Calls/
    Keine Daten
    Policy inference
    Keine Daten
    Episode wall
    Keine Daten
  • Base
    Calls/
    8.00
    Policy inference
    0.97
    Episode wall
    38.19
  • AHS
    Calls/
    14.50
    Policy inference
    1.55
    Episode wall
    36.83
  • B. QHA timing evaluation 2 held-out tasks, 400 episodes per method
    Calls/
    Keine Daten
    Policy inference
    Keine Daten
    Episode wall
    Keine Daten
  • AHS
    Calls/
    33.77
    Policy inference
    3.29
    Episode wall
    68.60
  • QHA-only
    Calls/
    29.81
    Policy inference
    18.87
    Episode wall
    81.53
  • AHS+QHA
    Calls/
    31.43
    Policy inference
    19.88
    Episode wall
    83.32
Shared timing scope for \pi_{0.5}.

Table 12 reports policy-only inference time. Total inference time is unavailable for the 50-task study. QHA selection is outside the policy timer and is not timed separately. The AHS reference in Panel B has a total inference time of 3.49 s per episode. Both QHA variants make fewer calls but have higher policy-inference and wall times than this reference. These timings and the success rates in Tab. 2B come from different cohorts and cannot be combined to estimate success-normalized efficiency. These costs characterize the evaluated implementation, not an intrinsic QHA cost.

AAC-core sampling cost.

On RoboCasa GR1 Tabletop with Qwen3GR00T (24 tasks, 1,200 episodes), AAC-core averages 180.95 calls, 20.98 s of synchronized total inference, and 47.56 s of wall time per episode. Each call samples 20 action chunks, giving 4,342,740 chunks in total. No policy-only timing aggregate is available. Base/AHS references come from separate evaluations and are not paired with the AAC-core measurements. AAC-core, the 50-task study, the four-task selector comparison, and the QHA timing evaluation were run on an NVIDIA RTX 4090 GPU.

D.2 Latency and Asynchronous Execution

The policy prediction horizon is H=50. Fixed K=40 is a target execution budget. RTC can replace a chunk at the first legal waypoint boundary after readiness. The AHS candidates are \{10,20,30,40\}. This cohort uses no QHA and is separate from the K=50 candidate-scaling experiment.

Timing and aggregation.

For episode i, let a_{ij}=t^{\mathrm{start}}_{ij}-t^{\mathrm{obs}}_{ij} be the age of the observation used to generate executed action j. We compute \bar{a}_{i}=n_{i}^{-1}\sum_{j}a_{ij} and average \bar{a}_{i} equally over episodes within each task, then equally over the four tasks, separately for Easy and Hard. Failures are included.

Wait s/ep averages recorded hold time per episode. Wait % is 100\sum_{i}W_{i}/\sum_{i}T_{i}, where T_{i} includes both action execution and waiting. Calls/ep includes requests discarded at episode termination. Infer s/ep reports model computation time per episode, amortized across active batch requests and excluding selector computation.

Table 13: Latency results by task and in aggregate on RoboTwin2.0 Hard. 100 paired episodes per task/method/delay, including failures. Costs and mean observation age use +200 ms. Fixed uses target K=40. Bold/underline mark best/second-best SR, waiting, and age within each group, including displayed ties. Calls and inference time are unranked.
  • —
    SR (%) \uparrow
    +0 ms
    +100 ms
    +200 ms
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • All four tasks
    SR (%) \uparrow
    Keine Daten
    Keine Daten
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • Sync-Fixed
    SR (%) \uparrow
    19.50
    19.50
    19.25
    Wait (%) \downarrow
    3.77
    Wait s/ep \downarrow
    4.432
    Calls/ep
    12.89
    Infer s/ep
    5.04
    Mean age (s) \downarrow
    5.02
  • Sync-AHS
    SR (%) \uparrow
    26.25
    24.25
    23.75
    Wait (%) \downarrow
    6.80
    Wait s/ep \downarrow
    7.893
    Calls/ep
    22.95
    Infer s/ep
    8.59
    Mean age (s) \downarrow
    3.00
  • RTC-Fixed
    SR (%) \uparrow
    23.00
    21.50
    22.00
    Wait (%) \downarrow
    0.31
    Wait s/ep \downarrow
    0.344
    Calls/ep
    13.32
    Infer s/ep
    6.47
    Mean age (s) \downarrow
    4.92
  • RTC-AHS
    SR (%) \uparrow
    26.50
    25.75
    26.75
    Wait (%) \downarrow
    0.31
    Wait s/ep \downarrow
    0.344
    Calls/ep
    24.30
    Infer s/ep
    10.62
    Mean age (s) \downarrow
    3.07
  • Place A2B Left
    SR (%) \uparrow
    Keine Daten
    Keine Daten
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • Sync-Fixed
    SR (%) \uparrow
    15.00
    21.00
    16.00
    Wait (%) \downarrow
    3.37
    Wait s/ep \downarrow
    3.106
    Calls/ep
    9.03
    Infer s/ep
    3.29
    Mean age (s) \downarrow
    5.35
  • Sync-AHS
    SR (%) \uparrow
    20.00
    16.00
    20.00
    Wait (%) \downarrow
    5.92
    Wait s/ep \downarrow
    5.790
    Calls/ep
    16.83
    Infer s/ep
    6.21
    Mean age (s) \downarrow
    3.30
  • RTC-Fixed
    SR (%) \uparrow
    26.00
    22.00
    22.00
    Wait (%) \downarrow
    0.41
    Wait s/ep \downarrow
    0.344
    Calls/ep
    9.53
    Infer s/ep
    4.09
    Mean age (s) \downarrow
    5.15
  • RTC-AHS
    SR (%) \uparrow
    22.00
    21.00
    22.00
    Wait (%) \downarrow
    0.37
    Wait s/ep \downarrow
    0.344
    Calls/ep
    17.94
    Infer s/ep
    7.28
    Mean age (s) \downarrow
    3.37
  • Place Bread Basket
    SR (%) \uparrow
    Keine Daten
    Keine Daten
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • Sync-Fixed
    SR (%) \uparrow
    26.00
    25.00
    25.00
    Wait (%) \downarrow
    3.20
    Wait s/ep \downarrow
    5.246
    Calls/ep
    15.25
    Infer s/ep
    6.12
    Mean age (s) \downarrow
    5.77
  • Sync-AHS
    SR (%) \uparrow
    38.00
    34.00
    32.00
    Wait (%) \downarrow
    6.19
    Wait s/ep \downarrow
    9.409
    Calls/ep
    27.36
    Infer s/ep
    10.14
    Mean age (s) \downarrow
    3.20
  • RTC-Fixed
    SR (%) \uparrow
    28.00
    29.00
    29.00
    Wait (%) \downarrow
    0.22
    Wait s/ep \downarrow
    0.344
    Calls/ep
    15.61
    Infer s/ep
    7.80
    Mean age (s) \downarrow
    5.68
  • RTC-AHS
    SR (%) \uparrow
    37.00
    31.00
    37.00
    Wait (%) \downarrow
    0.24
    Wait s/ep \downarrow
    0.344
    Calls/ep
    28.68
    Infer s/ep
    12.60
    Mean age (s) \downarrow
    3.38
  • Place Bread Skillet
    SR (%) \uparrow
    Keine Daten
    Keine Daten
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • Sync-Fixed
    SR (%) \uparrow
    13.00
    9.00
    12.00
    Wait (%) \downarrow
    5.15
    Wait s/ep \downarrow
    4.097
    Calls/ep
    11.91
    Infer s/ep
    4.51
    Mean age (s) \downarrow
    3.89
  • Sync-AHS
    SR (%) \uparrow
    13.00
    13.00
    12.00
    Wait (%) \downarrow
    8.56
    Wait s/ep \downarrow
    6.994
    Calls/ep
    20.33
    Infer s/ep
    7.06
    Mean age (s) \downarrow
    2.58
  • RTC-Fixed
    SR (%) \uparrow
    13.00
    12.00
    10.00
    Wait (%) \downarrow
    0.45
    Wait s/ep \downarrow
    0.346
    Calls/ep
    12.35
    Infer s/ep
    5.59
    Mean age (s) \downarrow
    3.81
  • RTC-AHS
    SR (%) \uparrow
    15.00
    16.00
    13.00
    Wait (%) \downarrow
    0.45
    Wait s/ep \downarrow
    0.345
    Calls/ep
    21.71
    Infer s/ep
    8.43
    Mean age (s) \downarrow
    2.54
  • Place Can Basket
    SR (%) \uparrow
    Keine Daten
    Keine Daten
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • Sync-Fixed
    SR (%) \uparrow
    24.00
    23.00
    24.00
    Wait (%) \downarrow
    3.90
    Wait s/ep \downarrow
    5.280
    Calls/ep
    15.35
    Infer s/ep
    6.27
    Mean age (s) \downarrow
    5.08
  • Sync-AHS
    SR (%) \uparrow
    34.00
    34.00
    31.00
    Wait (%) \downarrow
    7.07
    Wait s/ep \downarrow
    9.381
    Calls/ep
    27.27
    Infer s/ep
    10.95
    Mean age (s) \downarrow
    2.93
  • RTC-Fixed
    SR (%) \uparrow
    25.00
    23.00
    27.00
    Wait (%) \downarrow
    0.27
    Wait s/ep \downarrow
    0.344
    Calls/ep
    15.78
    Infer s/ep
    8.39
    Mean age (s) \downarrow
    5.02
  • RTC-AHS
    SR (%) \uparrow
    32.00
    35.00
    35.00
    Wait (%) \downarrow
    0.27
    Wait s/ep \downarrow
    0.344
    Calls/ep
    28.86
    Infer s/ep
    14.19
    Mean age (s) \downarrow
    3.01
Success within a time budget.

For T_{i}^{\mathrm{success}}, set the recorded completion time for a successful episode and +\infty for a failed episode. Then

\mathrm{SR}(t)=\frac{100}{4}\sum_{q=1}^{4}\frac{1}{100}\sum_{i\in q}\mathbf{1}\{T_{i}^{\mathrm{success}}\leq t\}.

Here t denotes physical time in seconds. Computed separately for each setting, \mathrm{SR}(t) measures task completion under a time budget, with all attempted episodes retained in the denominator.

Easy and Hard outcomes.

Both settings use the same task checkpoints, execution horizon and AHS candidates; episodes are paired across methods and delays within each setting. At +200 ms, RTC reduces AHS waiting by 94.5% on Easy and 95.6% on Hard. Relative to RTC-Fixed, RTC-AHS achieves higher aggregate SR in all six setting/delay conditions, with observed gains of 3.50–8.25 percentage points. At +200 ms, the paired 95% interval for this gain is [-0.75,8.26] points on Easy and [0.75,8.75] on Hard. Relative to Sync-AHS, RTC-AHS changes final SR by -2.00 points on Easy and +3.00 points on Hard at +200 ms. Thus, reduced waiting does not imply uniformly higher success. Policy calls and model computation remain higher than for RTC-Fixed (Tabs. 13, 14, and 15).

Table 14: Latency results by task and in aggregate on RoboTwin2.0 Easy. 100 paired episodes per task/method/delay, including failures. Costs and mean observation age use +200 ms. Fixed uses target K=40. Bold/underline mark best/second-best SR, waiting, and age within each group, including displayed ties. Calls and inference time are unranked.
  • —
    SR (%) \uparrow
    +0 ms
    +100 ms
    +200 ms
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • All four tasks
    SR (%) \uparrow
    Keine Daten
    Keine Daten
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • Sync-Fixed
    SR (%) \uparrow
    38.00
    35.50
    37.00
    Wait (%) \downarrow
    3.90
    Wait s/ep \downarrow
    3.856
    Calls/ep
    11.21
    Infer s/ep
    4.38
    Mean age (s) \downarrow
    5.12
  • Sync-AHS
    SR (%) \uparrow
    46.00
    46.25
    44.75
    Wait (%) \downarrow
    6.63
    Wait s/ep \downarrow
    6.310
    Calls/ep
    18.35
    Infer s/ep
    6.90
    Mean age (s) \downarrow
    3.16
  • RTC-Fixed
    SR (%) \uparrow
    38.75
    40.25
    39.00
    Wait (%) \downarrow
    0.37
    Wait s/ep \downarrow
    0.344
    Calls/ep
    11.50
    Infer s/ep
    5.57
    Mean age (s) \downarrow
    5.00
  • RTC-AHS
    SR (%) \uparrow
    46.00
    48.50
    42.75
    Wait (%) \downarrow
    0.37
    Wait s/ep \downarrow
    0.344
    Calls/ep
    20.32
    Infer s/ep
    8.80
    Mean age (s) \downarrow
    3.18
  • Place A2B Left
    SR (%) \uparrow
    Keine Daten
    Keine Daten
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • Sync-Fixed
    SR (%) \uparrow
    49.00
    45.00
    50.00
    Wait (%) \downarrow
    3.37
    Wait s/ep \downarrow
    2.401
    Calls/ep
    6.98
    Infer s/ep
    2.48
    Mean age (s) \downarrow
    5.65
  • Sync-AHS
    SR (%) \uparrow
    54.00
    58.00
    56.00
    Wait (%) \downarrow
    5.86
    Wait s/ep \downarrow
    4.107
    Calls/ep
    11.94
    Infer s/ep
    4.25
    Mean age (s) \downarrow
    3.41
  • RTC-Fixed
    SR (%) \uparrow
    50.00
    51.00
    50.00
    Wait (%) \downarrow
    0.51
    Wait s/ep \downarrow
    0.344
    Calls/ep
    7.55
    Infer s/ep
    3.36
    Mean age (s) \downarrow
    5.44
  • RTC-AHS
    SR (%) \uparrow
    52.00
    57.00
    57.00
    Wait (%) \downarrow
    0.51
    Wait s/ep \downarrow
    0.344
    Calls/ep
    12.83
    Infer s/ep
    5.21
    Mean age (s) \downarrow
    3.51
  • Place Bread Basket
    SR (%) \uparrow
    Keine Daten
    Keine Daten
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • Sync-Fixed
    SR (%) \uparrow
    40.00
    35.00
    35.00
    Wait (%) \downarrow
    3.41
    Wait s/ep \downarrow
    4.548
    Calls/ep
    13.22
    Infer s/ep
    5.29
    Mean age (s) \downarrow
    5.51
  • Sync-AHS
    SR (%) \uparrow
    47.00
    47.00
    49.00
    Wait (%) \downarrow
    6.10
    Wait s/ep \downarrow
    7.317
    Calls/ep
    21.27
    Infer s/ep
    8.19
    Mean age (s) \downarrow
    3.28
  • RTC-Fixed
    SR (%) \uparrow
    41.00
    42.00
    41.00
    Wait (%) \downarrow
    0.28
    Wait s/ep \downarrow
    0.344
    Calls/ep
    13.19
    Infer s/ep
    6.56
    Mean age (s) \downarrow
    5.48
  • RTC-AHS
    SR (%) \uparrow
    46.00
    48.00
    41.00
    Wait (%) \downarrow
    0.29
    Wait s/ep \downarrow
    0.344
    Calls/ep
    24.49
    Infer s/ep
    10.41
    Mean age (s) \downarrow
    3.33
  • Place Bread Skillet
    SR (%) \uparrow
    Keine Daten
    Keine Daten
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • Sync-Fixed
    SR (%) \uparrow
    26.00
    26.00
    27.00
    Wait (%) \downarrow
    4.98
    Wait s/ep \downarrow
    3.643
    Calls/ep
    10.59
    Infer s/ep
    4.02
    Mean age (s) \downarrow
    4.25
  • Sync-AHS
    SR (%) \uparrow
    36.00
    33.00
    29.00
    Wait (%) \downarrow
    7.86
    Wait s/ep \downarrow
    5.853
    Calls/ep
    17.02
    Infer s/ep
    6.00
    Mean age (s) \downarrow
    2.81
  • RTC-Fixed
    SR (%) \uparrow
    25.00
    31.00
    21.00
    Wait (%) \downarrow
    0.48
    Wait s/ep \downarrow
    0.345
    Calls/ep
    11.41
    Infer s/ep
    5.18
    Mean age (s) \downarrow
    4.02
  • RTC-AHS
    SR (%) \uparrow
    39.00
    39.00
    34.00
    Wait (%) \downarrow
    0.51
    Wait s/ep \downarrow
    0.345
    Calls/ep
    17.39
    Infer s/ep
    7.00
    Mean age (s) \downarrow
    2.90
  • Place Can Basket
    SR (%) \uparrow
    Keine Daten
    Keine Daten
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • Sync-Fixed
    SR (%) \uparrow
    37.00
    36.00
    36.00
    Wait (%) \downarrow
    4.12
    Wait s/ep \downarrow
    4.832
    Calls/ep
    14.05
    Infer s/ep
    5.75
    Mean age (s) \downarrow
    5.09
  • Sync-AHS
    SR (%) \uparrow
    47.00
    47.00
    45.00
    Wait (%) \downarrow
    6.84
    Wait s/ep \downarrow
    7.964
    Calls/ep
    23.15
    Infer s/ep
    9.16
    Mean age (s) \downarrow
    3.12
  • RTC-Fixed
    SR (%) \uparrow
    39.00
    37.00
    44.00
    Wait (%) \downarrow
    0.33
    Wait s/ep \downarrow
    0.344
    Calls/ep
    13.83
    Infer s/ep
    7.17
    Mean age (s) \downarrow
    5.06
  • RTC-AHS
    SR (%) \uparrow
    47.00
    50.00
    39.00
    Wait (%) \downarrow
    0.30
    Wait s/ep \downarrow
    0.344
    Calls/ep
    26.57
    Infer s/ep
    12.60
    Mean age (s) \downarrow
    3.00
Figure 9: AHS with asynchronous execution on \pi_{0.5} Easy. The three panels mirror Fig. 5: waiting versus mean observation age at +0/+100/+200 ms, final SR at each delay, and SR(t) at +200 ms. All 400 outcomes per condition are included. Bars and shading show marginal/pointwise 95% paired-seed bootstrap intervals within four tasks.
Table 15: Aggregate costs at +0 and +100 ms on Easy and Hard. Each method/delay has 400 outcomes per setting, including failures. The +200 ms aggregates appear with the per-task results in Tabs. 13 and 14. Observation age averages within episodes, then equally across episodes and tasks. Within each setting and delay, best and second-best distinct waiting and age values are bold and underlined. Displayed ties share a mark. Calls/ep and Infer s/ep are descriptive quantities and are not ranked.
  • Easy
    Method
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • 0
    Method
    Sync-Fixed
    Wait (%) \downarrow
    1.67
    Wait s/ep \downarrow
    1.600
    Calls/ep
    11.11
    Infer s/ep
    4.32
    Mean age (s) \downarrow
    4.93
  • —
    Method
    Sync-AHS
    Wait (%) \downarrow
    2.87
    Wait s/ep \downarrow
    2.614
    Calls/ep
    18.15
    Infer s/ep
    6.78
    Mean age (s) \downarrow
    3.01
  • —
    Method
    RTC-Fixed
    Wait (%) \downarrow
    0.16
    Wait s/ep \downarrow
    0.147
    Calls/ep
    11.20
    Infer s/ep
    5.38
    Mean age (s) \downarrow
    5.00
  • —
    Method
    RTC-AHS
    Wait (%) \downarrow
    0.16
    Wait s/ep \downarrow
    0.147
    Calls/ep
    17.86
    Infer s/ep
    7.81
    Mean age (s) \downarrow
    3.24
  • 100
    Method
    Sync-Fixed
    Wait (%) \downarrow
    2.82
    Wait s/ep \downarrow
    2.755
    Calls/ep
    11.29
    Infer s/ep
    4.37
    Mean age (s) \downarrow
    4.98
  • —
    Method
    Sync-AHS
    Wait (%) \downarrow
    4.80
    Wait s/ep \downarrow
    4.384
    Calls/ep
    17.97
    Infer s/ep
    6.60
    Mean age (s) \downarrow
    3.08
  • —
    Method
    RTC-Fixed
    Wait (%) \downarrow
    0.26
    Wait s/ep \downarrow
    0.244
    Calls/ep
    11.42
    Infer s/ep
    5.45
    Mean age (s) \downarrow
    4.97
  • —
    Method
    RTC-AHS
    Wait (%) \downarrow
    0.28
    Wait s/ep \downarrow
    0.244
    Calls/ep
    18.73
    Infer s/ep
    8.09
    Mean age (s) \downarrow
    3.12
  • Hard
    Method
    Keine Daten
    Wait (%) \downarrow
    Keine Daten
    Wait s/ep \downarrow
    Keine Daten
    Calls/ep
    Keine Daten
    Infer s/ep
    Keine Daten
    Mean age (s) \downarrow
    Keine Daten
  • 0
    Method
    Sync-Fixed
    Wait (%) \downarrow
    1.60
    Wait s/ep \downarrow
    1.853
    Calls/ep
    12.87
    Infer s/ep
    5.14
    Mean age (s) \downarrow
    4.84
  • —
    Method
    Sync-AHS
    Wait (%) \downarrow
    2.93
    Wait s/ep \downarrow
    3.244
    Calls/ep
    22.53
    Infer s/ep
    8.37
    Mean age (s) \downarrow
    2.85
  • —
    Method
    RTC-Fixed
    Wait (%) \downarrow
    0.13
    Wait s/ep \downarrow
    0.146
    Calls/ep
    12.84
    Infer s/ep
    6.31
    Mean age (s) \downarrow
    4.99
  • —
    Method
    RTC-AHS
    Wait (%) \downarrow
    0.14
    Wait s/ep \downarrow
    0.147
    Calls/ep
    22.23
    Infer s/ep
    9.83
    Mean age (s) \downarrow
    3.01
  • 100
    Method
    Sync-Fixed
    Wait (%) \downarrow
    2.72
    Wait s/ep \downarrow
    3.168
    Calls/ep
    12.99
    Infer s/ep
    5.12
    Mean age (s) \downarrow
    4.89
  • —
    Method
    Sync-AHS
    Wait (%) \downarrow
    4.85
    Wait s/ep \downarrow
    5.523
    Calls/ep
    22.64
    Infer s/ep
    8.39
    Mean age (s) \downarrow
    2.93
  • —
    Method
    RTC-Fixed
    Wait (%) \downarrow
    0.22
    Wait s/ep \downarrow
    0.244
    Calls/ep
    13.34
    Infer s/ep
    6.64
    Mean age (s) \downarrow
    4.84
  • —
    Method
    RTC-AHS
    Wait (%) \downarrow
    0.22
    Wait s/ep \downarrow
    0.244
    Calls/ep
    24.05
    Infer s/ep
    10.49
    Mean age (s) \downarrow
    2.98

Appendix E AHS Hyperparameter Sensitivity

We evaluate \pi_{0.5} + AHS on eight RoboTwin2.0 tasks under Easy and Hard settings, with \mathcal{K}=\{10,20,30,40,50\} and expected-round selection at T_{\mathrm{sel}}=0.8. The two studies below use different episode budgets and pairing protocols. Neither changes our default configuration.

Inter-chunk continuity window.

We vary only W_{h}\in\{30,60,90\}, sharing 16 episode identities per task and setting (256 per configuration). The default W_{h}=60 is evaluated afresh as the paired reference in Fig. 10.

Figure 10: Inter-chunk continuity-window sensitivity. \pi_{0.5} + AHS on eight RoboTwin2.0 tasks, Easy and Hard, with 256 matched episodes per configuration. Lines connect tested points. Paired difference intervals appear in the text.

Compared with W_{h}=60, gains for 30 and 90 are +2.73 and +3.91 percentage points, with paired 95% confidence intervals [-2.34,\,7.81] and [-0.78,\,8.59]. We use 10,000 bootstrap draws of matched episodes within task–setting cells, weighted equally. Both intervals include zero, so neither a success-rate difference nor equivalence is established. We retain W_{h}=60.

Other evidence and posterior hyperparameters.

Each configuration in Tab. 16 uses 100 episodes per task and setting (1,600 total), unpaired across configurations. One parameter family changes at a time, with the two penalties varied jointly. Other settings follow Appendix A. These evaluations are separate from the continuity-window study. Overall SR spans 34.38–37.13%, characterizing sensitivity rather than validation-based selection or an established optimum.

Table 16: Sensitivity to other AHS hyperparameters. Success rates (%) on eight RoboTwin2.0 tasks, with 1,600 episodes per configuration. Bold/underline mark higher/lower SR within each parameter family. Defaults specify parameters, not additional evaluations.
  • Forgetting \rho_{f}
    Default
    0.99
    Tested
    0.95
    Easy \uparrow
    46.88
    Hard \uparrow
    26.75
    Overall \uparrow
    36.81
  • —
    Default
    Keine Daten
    Tested
    0.995
    Easy \uparrow
    46.50
    Hard \uparrow
    25.88
    Overall \uparrow
    36.19
  • Update strength \eta
    Default
    1
    Tested
    0.5
    Easy \uparrow
    45.38
    Hard \uparrow
    24.38
    Overall \uparrow
    34.88
  • —
    Default
    Keine Daten
    Tested
    2
    Easy \uparrow
    47.38
    Hard \uparrow
    26.88
    Overall \uparrow
    37.13
  • Kernel bandwidth \sigma_{K}
    Default
    10
    Tested
    5
    Easy \uparrow
    45.50
    Hard \uparrow
    26.13
    Overall \uparrow
    35.81
  • —
    Default
    Keine Daten
    Tested
    20
    Easy \uparrow
    44.25
    Hard \uparrow
    24.50
    Overall \uparrow
    34.38
  • Intra reference window W_{\tau}
    Default
    3
    Tested
    1
    Easy \uparrow
    47.38
    Hard \uparrow
    24.75
    Overall \uparrow
    36.06
  • —
    Default
    Keine Daten
    Tested
    5
    Easy \uparrow
    46.38
    Hard \uparrow
    23.88
    Overall \uparrow
    35.13
  • Joint \alpha_{\mathrm{intra}}=\beta_{\mathrm{inter}}
    Default
    4
    Tested
    2
    Easy \uparrow
    46.63
    Hard \uparrow
    23.63
    Overall \uparrow
    35.13
  • —
    Default
    Keine Daten
    Tested
    8
    Easy \uparrow
    46.25
    Hard \uparrow
    24.88
    Overall \uparrow
    35.56
Figure 11: \pi_{0} Handover Block case study. The selected execution horizon changes with task phase in a successful RoboTwin2.0 rollout. K_{\mathrm{exec}} denotes the selected K_{t}.

Appendix F More Case Studies

Handover Block.

Fig. 11 illustrates phase-dependent execution in a successful \pi_{0} rollout. AHS selects shorter horizons during grasping and transport phases that require closed-loop correction, and longer horizons once the motion stabilizes. This example illustrates the behavior of the online selector in Sec. 4.1. It is not a controlled comparison of posterior mechanisms.

Additional tasks and settings.

Fig. 12 shows four additional \pi_{0} RoboTwin2.0 rollouts across Easy and Hard settings. Fig. 13 adds \pi_{0.5}+AHS cases on four real-world tasks. These qualitative examples illustrate horizon changes across task phases, not additional scored trials.

Figure 12: Additional \pi_{0} case studies on RoboTwin2.0. Four successful rollouts show phase-dependent AHS horizons across Place Bread Basket and Blocks Ranking RGB under Easy and Hard settings.
Figure 13: Real-world \pi_{0.5}+AHS case studies. Highlighted observations are selected at large adjacent-replan changes in K_{t}. Curves show all recorded execution horizons without smoothing, through the end of each recording. Phase labels are manual visual annotations, not ground-truth boundaries or trial scores.

Appendix G Discussion and Limitations

Additional background.

RT-1, RT-2, and PaLM-E connect large-scale learning with robot control (Brohan et al., 2023; Zitkovich et al., 2023; Driess et al., 2023). Open X-Embodiment, Octo, and OpenVLA extend shared data and generalist initialization (O’Neill et al., 2024; Ghosh et al., 2024; Kim et al., 2024), while RoboBrain 2.0 broadens embodied reasoning interfaces (Team et al., 2025). RDT-1B and DexVLA extend diffusion-based action generation (Liu et al., 2025a; Wen et al., 2025). Other approaches couple actions to world models (Bi et al., 2025) or adapt chunking through self-guidance (So et al., 2026). ChunkTrust adapts chunk execution using action-expert evidence.

Additional horizon and verification methods.

Adaptive execution can draw on additional samples, learned sensitivity, or observations acquired during rollout. A3 (Chen et al., 2026a) uses group-sampled consensus and conditional re-decoding to verify a contiguous execution prefix. SA (Park et al., 2026) forecasts the sensitivity of the action distribution to observation changes and allocates shorter horizons to more sensitive phases. EQRL (Wang et al., 2026a) jointly learns the latent input, denoising budget, and chunk length through reinforcement learning. These methods differ in the information and computation used to choose a horizon. AHS scores prefixes of one generated chunk using its existing generation trace and executed history, with no conditional re-decoding or joint optimization of the generator’s inference schedule. Verification methods instead use fresh observations to assess a running plan. FFDC (Wang et al., 2026d) compares imagined futures with reality for world-action models, while DREAM-Chunk (Chen et al., 2026b) uses a latent world model to match candidate chunks’ predicted futures to observed execution. SV-VLA (Wang et al., 2026e) compares planned actions with a lightweight closed-loop reference, and PATCH (Zhou et al., 2026) accumulates localized visual residuals along an action-conditioned execution corridor to trigger intervention. These observation-driven mechanisms address disturbances that become visible after a chunk has been selected. ChunkTrust’s evidence instead informs how much of the current prediction to execute before observing again, so it does not provide the same within-prefix monitoring capability.

Continuity and frequency-aware action generation.

Cross-chunk consistency can be improved by changing how actions are generated or corrected. ChunkFlow (Yang et al., 2026) trains with seam and derivative-continuity losses and blends overlapping predictions at execution. SEAM (Zhan et al., 2026) steers denoising toward the previous chunk’s unexecuted tail, while Legato (Liu et al., 2026) learns continuation through action-conditioned initialization and modified flow dynamics. REMAC (Wang et al., 2026c) uses masked action conditioning for real-time execution, and ACNet (Guo and Guo, 2026) conditions a lightweight delay-aware adapter on executed motion. A2C2 (Sendai et al., 2025) instead applies a learned per-step correction using the latest observation and the base policy’s action. Frequency-aware approaches act on the representation or training objective. FAFM (Guo et al., 2026) generates continuous action trajectories in a frequency-domain representation, while FocalPolicy (He et al., 2026) combines proximal time-domain supervision with multi-chunk spectral regularization. These works motivate attention to temporal coherence, but they do not make the same intervention as ChunkTrust. Our spectral signal measures variation during generation, and our continuity signal evaluates a prefix stitched to executed history. Both are used to select an execution length, leaving the generated action values and base-policy weights unchanged. A frequency-domain training loss is therefore distinct from the generation-time diagnostic used here.

Experience reuse and the role of chunking.

TraceFlow (Zhang et al., 2026a) reuses successful and failed rollouts through a retrieval bank and outcome-conditioned guidance of a frozen flow-matching action expert. ChunkTrust reuses information in a different form and for a different decision. AHS retains evidence-derived horizon preferences within an episode, while QHA learns a context-conditioned horizon prior offline. Neither component retrieves rollout trajectories to steer action generation, and QHA does not update its weights during deployment. Recent analyses also caution against treating long open-loop execution as universally beneficial. Lazzati et al. (2026) study non-Markovian expressivity and implicit ensembling as explanations for the benefits of chunking. Zeng et al. (2026) show that the value of open-loop execution depends on demonstration non-Markovianity and policy context length. ChunkTrust addresses horizon selection for existing chunk policies, not a claim that longer open-loop execution is intrinsically preferable to reactive control.

Limitations.

AHS requires access to action chunks and generation-time velocity traces. It scores a candidate grid, although expected-round selection can return intermediate integer lengths. QHA learns a dense prior, but AHS+QHA still computes online evidence, interpolated or scored directly on the dense grid. Horizon decisions are made at replanning time, so the selector cannot directly detect a new disturbance that arises during the chosen prefix. Our evaluations cover multiple policies and two simulation benchmarks, plus four real-world bimanual tasks, rather than all embodiments, safety-critical tasks, or long-horizon mobile manipulation.

Broader impact.

Adaptive horizons may reduce unnecessary replanning while preserving reactivity near contact. An incorrect horizon can still commit a robot to unsafe motion outside the tested distribution. Deployment should retain workspace and speed limits, emergency stops, human supervision during evaluation, and task-specific validation.

Third-party resources.

We use the cited RoboTwin2.0 and RoboCasa benchmarks, policy checkpoints, and associated software for research evaluation. These third-party resources remain subject to their original licenses, terms of use, and attribution requirements.