ChunkTrust: 행동 전문가 신호로 로봇 정책의 실행 구간을 조절하는 방법
로봇 AI 모델은 앞으로의 동작을 한 묶음씩 예측하는데, 그중 몇 개를 실행한 뒤 다시 주변을 살필지는 대개 사람이 고정값으로 정해 왔다. 칭화대와 베이징즈위안인공지능연구원(BAAI) 등이 낸 이번 프리프린트는 이 값을 작업 중에 스스로 고르는 부가 모듈을 제안했고, 피지컬인텔리전스와 엔비디아의 기존 모델을 다시 학습시키지 않고도 시뮬레이션과 실제 양팔 로봇에서 성공률을 끌어올렸다.
동작 묶음(action chunk)을 출력하는 로봇 정책이 재계획 시점마다 재학습 없이, 예측한 묶음 가운데 어디까지를 믿고 실행한 뒤 다시 관측해야 하는지 스스로 판단할 수 있는지를 따졌다.
π0와 π0.5 같은 최신 로봇 파운데이션 정책은 미래 동작 묶음을 한꺼번에 예측한 뒤 정해진 길이만큼 실행하고 다시 계획한다. 연구진은 최적의 고정 길이가 과제와 과제 길이, 평가 조건에 따라 다르고 한 에피소드 안에서도 단계마다 달라진다는 점을 보였다. 빈 공간에서 팔을 뻗는 구간은 관측 없이 오래 움직여도 되지만 물체와 접촉하는 섬세한 구간은 그렇지 않다. 고정값은 재계획을 너무 자주 해 연산을 낭비하거나, 계획이 현실과 어긋난 뒤에도 로봇을 그대로 움직이게 만든다. 연구진이 모은 고정 구간 에피소드 1,600개에서 두 불안정 신호가 모두 높았던 경우의 실패율은 97.4%로, 둘 다 낮았던 경우의 50.2%를 크게 웃돌았다.
현장에서는 과제마다 실행 구간을 시행착오로 하나 정해 쓰는 경우가 많다. 기존 적응형 기법은 어텐션 구조나 예측 동작의 엔트로피, 길이가 다른 예측 간 일치도에서 구간을 끌어냈다. 양방향 디코딩(BID)은 여러 샘플 묶음 가운데 역방향 일관성과 순방향 대비로 하나를 고르고, FASTER는 가까운 시점의 샘플링을 우선한다. 논문은 이들 방법이 묶음 자체의 내부 일관성을 주로 따지기 때문에, 그 자체로는 매끄럽지만 이미 실행한 동작과 이어지지 않는 계획을 걸러내지 못한다고 지적했다. 재계획 사이에 근거를 쌓거나 과제 단계가 바뀔 때 판단을 조정하는 장치도 없다고 봤다.
ChunkTrust는 두 부분으로 구성된다. 학습이 필요 없는 행동 인지 구간 선택기(AHS)는 작업 중에 후보 구간 길이마다 두 신호를 잰다. 첫째는 묶음 내부의 스펙트럼 안정성으로, 플로우나 디퓨전 방식 행동 헤드가 노이즈를 걷어내는 단계를 거치는 동안 예측 속도의 고주파 성분이 얼마나 흔들리는지를 본다. 둘째는 묶음 사이의 연속성으로, 이미 실행한 동작과 새 구간이 이어지는 지점에서 속도가 매끄럽게 유지되는지를 측정한다. 두 점수를 같은 비중으로 섞은 뒤 구간 길이별 신뢰도를 베타 분포로 관리해 곱한다. 이 신뢰도는 재계획마다 갱신되고 0.99의 망각 계수로 이전 단계의 근거가 서서히 옅어지며, 인접한 구간 길이끼리 피드백을 나눈다. 선택 사항인 질의 기반 구간 어댑터(QHA)는 고정된 기반 정책의 시각 언어 문맥과 행동 잠재 표현을 읽어 AHS가 선호할 구간을 한 번의 연산으로 예측하도록 학습한 소형 모듈이다. 배치 단계에서는 이 사전 분포를 실시간 근거와 결합한다. 기반 정책 자체는 전혀 수정하지 않는다.
RoboTwin 2.0 50개 과제 전체에서 멀티태스크 π0.5의 평균 성공률은 AHS 적용으로 56.70%에서 63.50%로 6.80%포인트 올랐다. 95% 부트스트랩 구간은 3.75%포인트에서 9.95%포인트였다. 쉬운 조건은 65.40%에서 73.10%로, 어려운 조건은 48.00%에서 53.90%로 높아졌다. 과제별로 학습한 8개 과제 부분 집합에서 π0는 18.19%에서 AHS 적용 시 21.38%, QHA까지 더하면 24.88%가 됐고, π0.5는 29.63%에서 36.25%, 39.06%로 올랐다. 이미 성능이 높은 Fast-WAM은 86.81%에서 88.13%로, X-VLA는 47.44%에서 47.75%로 소폭 오르는 데 그쳤다. 학습에 쓰지 않은 과제에서 QHA는 π0.5 기준 AHS 단독 대비 2.75%포인트(40.00%에서 42.75%)를 더했다. RoboCasa GR1 탁상 24개 과제에서는 π0.5가 40.08%에서 42.50%로, 엔비디아 Isaac GR00T N1.5 제로샷이 1.67%포인트, GR00T N1.6이 47.61%에서 51.42%로, Qwen3 기반 GR00T 변형이 47.83%에서 57.50%로 개선됐다. π0.5를 탑재한 AgileX COBOT Magic 양팔 로봇으로 수건 개기, 빵 옮기기, 음료 옮기기, 오리 인형 서랍에 넣기 등 4개 가정 과제를 과제당 15회씩 시험한 결과, 세부 단계 완료에 부분 점수를 주는 평균 점수가 50.4%에서 57.5%로 올랐다. 과제별 개선 폭은 수건 개기 1.7%포인트에서 서랍 과제 13.3%포인트까지 벌어졌다. 선택기는 평가하는 후보 수에 따라 재계획당 0.7%에서 3.6%의 부담을 더하지만, 50개 과제 실험에서 에피소드당 총 추론 시간은 재계획이 잦아지면서 0.97초에서 1.55초로 58% 늘었다. AHS를 실시간 청킹(RTC)과 함께 비동기로 돌리자 어려운 과제의 대기 시간은 에피소드당 7.893초에서 0.344초로 줄었다.
이 연구는 동료 심사를 거치지 않은 프리프린트다. 저자들은 이 방법이 정책의 중간 디노이징 궤적에 접근해야 하므로 내부를 볼 수 없는 API형 모델이나 동작을 한 번에 출력하는 정책에는 붙일 수 없다고 인정했다. 개선 폭이 과제마다 고르지 않고 평균 뒤에 일부 과제의 성능 하락이 숨어 있으며, 특히 학습한 QHA 사전 분포를 실시간 근거와 결합할 때 이런 하락이 나타났다고 밝혔다. 어려운 조건의 블록 건네기 과제가 그 예다. 재계획당 선택 비용이 작다고 에피소드 전체 시간이 같다는 뜻은 아니라는 점도 짚었다. ROBOTNESS가 보기에 실제 로봇 근거는 한 기종, 한 기반 정책, 4개 과제와 과제당 15회 시행에 그쳐 7%포인트 개선의 불확실성이 크다. 성능이 높거나 구조가 다른 정책에서는 효과가 거의 사라져(Fast-WAM 1.31%포인트, X-VLA 0.31%포인트) 중간 수준 모델의 약점을 메우는 성격이 강하다. 에피소드당 추론 시간 58% 증가는 엣지 연산 예산이 빠듯한 제품에 부담이 되고, 시뮬레이션의 절대 성공률도 대부분 과제에서 상용 수준과 거리가 멀다.
이번 연구는 동작 묶음 방식 VLA 정책을 현장에 배치하는 팀이 손으로 맞춰 온 설정값을 겨냥했다. 시험 대상에는 피지컬인텔리전스의 π0와 π0.5, 엔비디아의 Isaac GR00T N1.5와 N1.6이 포함됐다. AHS는 재학습이 필요 없고 코드도 공개돼 있어, 공개 VLA 체크포인트를 쓰는 국내 로봇 통합 업체와 연구실도 몇 주 안에 시험해 볼 수 있다. 빈 공간 이동과 접촉 작업이 번갈아 나오는 주방 보조나 물류 피킹 같은 작업에서 효과가 가장 클 것으로 보인다. 모델 개발사에 남는 더 오래가는 교훈은 실행 구간을 정책 스스로의 불확실성 신호로 작업 중에 정해야 한다는 점이다. 폐쇄형 모델 업체는 이 방법을 그대로 쓰기보다 앞으로 12개월에서 24개월 사이에 비슷한 논리를 자체 추론 스택에 넣을 가능성이 높다. 가장 강한 정책에서는 개선 폭이 1%포인트 안팎이어서 로봇 하드웨어 구매자에게 당장의 영향은 크지 않다.
논문 전문
ChunkTrust: Adapting Execution Horizons for Robot Policies with Action-Expert Evidence
CC BY 4.0 라이선스로 공개된 논문입니다. 출처를 밝혀 전재하며, 원문은 arXiv:2609.39754(PDF)에서 볼 수 있습니다.
Abstract
Robot foundation policies predict action chunks, but how many actions to execute before replanning depends on the current task phase. We introduce , which treats the execution horizon as a latent variable inferred from action-expert evidence rather than a fixed hyperparameter. Its training-free Action-aware Horizon Selector (AHS) combines intra-chunk spectral stability of generation traces with inter-chunk continuity between executed history and predicted actions. An online Beta posterior with kernel forgetting tracks horizon preferences across replans. A lightweight Query-based Horizon Adapter (QHA) optionally learns a context-conditioned dense prior from complementary evidence, fused with current evidence and episode-local Beta memory while the base policy remains frozen. Across RoboTwin2.0 and RoboCasa GR1 Tabletop, AHS improves overall task-averaged success for each evaluated base-policy configuration, including gains of +6.80 percentage points on \pi_{0.5} over all 50 RoboTwin2.0 tasks and +9.67 percentage points on Qwen3GR00T in RoboCasa. AHS+QHA raises the gain over Base to +9.44 percentage points on the eight-task \pi_{0.5} evaluation. On four real-world household tasks, AHS improves the equal-task mean normalized process score from 50.4% to 57.5%. Ablations examine the contributions of both evidence terms, temporal memory, and the learned prior.
K_{\mathrm{exec}}=K_{t}). (a) A shared robot policy with fixed or adaptive execution horizons. (b) Fixed horizons K\in\{10,20,30,40,50\} versus adaptive execution on \pi_{0} RoboTwin2.0 evaluations, grouped by task length and environment setting. Error bars indicate \pm 1 standard error of the mean across task–setting combinations. Dashed lines show the corresponding adaptive references.1 Introduction
Robot foundation policies, including Vision-Language-Action (VLA) models (Black et al., 2024; Intelligence et al., 2025) and World-Action Models (WAMs) (Yuan et al., 2026; Ye et al., 2026b), predict action chunks to amortize inference and maintain temporal coherence. The execution horizon can differ from the generated chunk length, reflecting a trade-off between longer execution that can preserve smooth progress but delays correction and shorter execution that increases feedback frequency but can disrupt coherent motion (Lu et al., 2026). In manipulation, this trade-off changes within an episode: free-space approach may tolerate longer open-loop execution, whereas contact or delicate transport demands earlier replanning. Fig. 1 illustrates this phase dependence in a real rollout and shows that the best fixed horizon varies across task lengths and evaluation settings. The resulting trust boundary problem is to determine how many actions from the current chunk the robot should execute before replanning.
Adaptive chunking methods derive execution horizons from attention structure, action entropy, or cross-horizon agreement (Wang et al., 2026b; Liang et al., 2026; Jing et al., 2026), and increasingly from the policy’s denoising trajectory (Feng et al., 2026; Chen et al., 2026c). Other approaches learn when to replan (Zhao et al., 2026; Xu et al., 2026b) or monitor execution to trigger correction (Pan et al., 2026). These approaches tackle when to replan from different perspectives, yet an internally consistent prediction can still be incompatible with the motion already executed. Meanwhile, evidence at individual replans can be noisy, while similar execution contexts recur within and across episodes. This raises a central question: How can a robot identify complementary evidence native to its action expert and internalize it as reusable knowledge for adaptive execution?
We introduce ChunkTrust, a framework that couples online evidence accumulation with context-conditioned horizon learning (Fig. 3). ChunkTrust evaluates candidate execution prefixes along two complementary dimensions. We assess spectral stability during action generation, drawing on frequency-domain analyses of diffusion and flow models (Si et al., 2024; Huang et al., 2026a). We also assess continuity with recently executed motion, reflecting the importance of cross-chunk consistency in robot control (Liu et al., 2025b; Black et al., 2025). Our retrospective analysis in Fig. 2 provides empirical support for this pairing. Failed episodes exhibit higher median spectral instability and boundary variation, with the highest failure rate observed when both risks are elevated.
ChunkTrust internalizes this evidence through online memory and a learned prior. The Action-aware Horizon Selector (AHS) accumulates evidence in an episode-local Beta state with forgetting, retaining useful horizon preferences while adapting to phase changes. The Query-based Horizon Adapter (QHA) learns context-conditioned horizon preferences from action-expert evidence, enabling their reuse beyond the current episode. At deployment, QHA supplies a dense prior that complements online AHS evidence, combining learned preferences with adaptation to the current rollout.
We evaluate the benefits of online horizon adaptation and learned horizon preferences across RoboTwin2.0, RoboCasa GR1 Tabletop, and four real-world household tasks. AHS improves aggregate success across all evaluated simulation base-policy configurations, including gains of 6.80 percentage points on the complete 50-task RoboTwin2.0 suite and 9.67 percentage points with Qwen3GR00T on RoboCasa. Adding QHA further improves aggregate success on the eight-task evaluation and benefits two tasks excluded from horizon-head training, supporting reuse of the learned preferences beyond the QHA training tasks. On real robots, AHS increases the equal-task mean normalized process score by 7.1 percentage points. Controlled comparisons and ablations examine the contributions of complementary evidence, temporal memory, and the learned prior, alongside alternative horizon selectors and inference costs.
Our contributions are threefold: (1) we formulate execution-horizon adaptation as inference over candidate prefixes, grounded in the action expert’s generation dynamics and compatibility with executed history. (2) we introduce AHS, which integrates dual evidence with a phase-aware Beta posterior, and QHA, which learns a context-conditioned horizon prior that complements online evidence without updating the base policy. (3) we evaluate across policy families, two simulation benchmarks, and real robots, with full task-level results and controlled evidence, memory, transfer, and cost analyses.
2 Related Work
\pi_{0} and fixed K=H=50. (a,b) Episode means of replan-level evidence: \bar{z}_{\mathrm{intra}}(k)=R_{z}^{-1}\sum_{t}z_{\mathrm{intra},t}(k), \bar{u}_{\mathrm{inter}}(k)=R_{u}^{-1}\sum_{t}u_{\mathrm{inter},t}(k). Sums and counts R_{z},R_{u} use valid replans for each metric and horizon. Lines/bands show medians/interquartile ranges across episodes. (c) Failure rates after median-splitting episode risks (within-episode 75th percentiles of 1-q_{\mathrm{intra}} and 1-q_{\mathrm{inter}}). “Inter only”/“Intra only” means only the named risk is high. Details: Appendix B.Action generation and reactive execution.
Mobile ALOHA (Fu et al., 2024b) and Diffusion Policy (Chi et al., 2025) use multi-step action predictions for temporal coherence. Flow-based policies such as \pi_{0} (Black et al., 2024) and \pi_{0.5} (Intelligence et al., 2025), and world-action models such as Fast-WAM (Yuan et al., 2026), extend action generation to broader task distributions. FASTER (Lu et al., 2026) prioritizes near-term sampling through horizon-aware scheduling and streaming execution. ChainVLA (Huang et al., 2026b) conditions successive queries on task progress and the unexecuted action suffix. Execution monitors offer another route to reactivity. VLA-Corrector (Pan et al., 2026) uses a learned latent dynamics model to detect persistent execution drift and guide action correction. React When You Need To (Wu et al., 2026) triggers asynchronous inference in response to scene changes. ChunkTrust instead selects a prefix at each replan using evidence already available from the action expert and executed history. It neither modifies the generated actions nor monitors new observations during that prefix.
Evidence-based horizon selection.
BID (Liu et al., 2025b) selects among sampled chunks using backward coherence and forward contrast, targeting consistency across predictions. Horizon adaptation instead changes the executed prefix. Mixture of Horizons (MoH) uses cross-horizon consensus (Jing et al., 2026), AutoHorizon uses action self-attention as a predictive-limit proxy (Wang et al., 2026b), and Adaptive Action Chunking (AAC) uses action entropy (Liang et al., 2026). HiPolicy combines multi-frequency chunk generation with entropy-guided execution (Zhang et al., 2026b). More recent methods expand the available signals. Knowing When to Stop (Xu et al., 2026a) detects entropy plateaus in action-to-observation cross-attention, while DVAC (Feng et al., 2026) measures variation in clean-action estimates during denoising. GeoAAC (Chen et al., 2026c) constructs prefix-wise geometric profiles from a single denoising trajectory. PACE (Nie et al., 2026) instead identifies low-speed transition points directly in the predicted chunk. ChunkTrust pairs spectral variation during generation with speed variation after stitching a candidate prefix to executed history. DVAC uses rolling history to calibrate its variance threshold, whereas AHS maintains horizon-indexed Beta states that accumulate the paired evidence as soft feedback with forgetting.
Learned horizon selection.
DEHP (Zhao et al., 2026) and BCP (Xu et al., 2026b) train horizon or continuation heads through reinforcement learning with frozen base policies. EQRL (Wang et al., 2026a) jointly learns to select the latent input, denoising budget, and chunk length, while Spatial Attention (SA) (Park et al., 2026) learns to forecast observation sensitivity and uses it to allocate execution horizons. QHA instead learns a context-conditioned horizon prior from complementary action-expert evidence and combines it with current evidence and episode-local Beta memory at deployment. Appendix G discusses additional connections.
3 Horizon-Aware Evidence from the Action Expert
3.1 Preliminaries
At replan step t, a frozen policy conditions on visual observation o_{t}, proprioceptive state \mathbf{x}_{t}^{\mathrm{prop}}, and instruction \ell. Its backbone produces context tokens \mathbf{C}_{t}=f_{\theta}(o_{t},\mathbf{x}_{t}^{\mathrm{prop}},\ell), and the action expert predicts
\hat{\mathbf{A}}_{t}=[\hat{\mathbf{a}}_{t,1},\ldots,\hat{\mathbf{a}}_{t,H}]\in\mathbb{R}^{H\times d_{a}}.The controller executes a prefix of length K_{t}\in\{1,\ldots,H\} before re-observation. We score candidates k\in\mathcal{K}\subseteq\{1,\ldots,H\} using internal generation stability and compatibility with executed history. The expected-round rule in Sec. 4 can select intermediate integer lengths rather than only grid points.
A single policy call with trace recording returns (\hat{\mathbf{A}}_{t},\mathcal{F}_{t})=\pi_{\theta}(o_{t},\mathbf{x}_{t}^{\mathrm{prop}},\ell). For flow-based action experts (Lipman et al., 2023), \mathcal{F}_{t} contains velocity predictions recorded during generation, without changing the actions. We write the trace and executed history as
\mathcal{F}_{t}=\{\mathbf{v}_{t,\tau}\in\mathbb{R}^{H\times d_{a}}\}_{\tau=0}^{T-1},\qquad\mathcal{H}_{t}=[\mathbf{a}^{\mathrm{exec}}_{n_{t}-N_{t}+1},\ldots,\mathbf{a}^{\mathrm{exec}}_{n_{t}}],where \tau indexes sampling steps, n_{t} counts actions executed before replan t, and N_{t} is the available history length. The following evidence terms use \mathcal{F}_{t} and (\mathcal{H}_{t},\hat{\mathbf{A}}_{t}), respectively.
\mu_{t}. QHA learns a context-conditioned prior from dense evidence.3.2 Intra-Chunk Spectral Stability
We measure variation in the velocity-prefix spectrum during generation. Specifically, we apply the Fourier transform along the action horizon at each denoising step and compare the resulting spectra across steps. For each candidate k, we pad its velocity prefix to length H and apply a one-dimensional Fourier transform along the action-horizon axis, separately for action dimensions j\in\{1,\ldots,d_{a}\} and nonnegative frequencies \omega\in\Omega:
\widehat{\mathbf{V}}^{(k)}_{t,\tau}(\omega,j)=\operatorname{FFT}_{h}\!\left(\bar{\mathbf{v}}^{(k)}_{t,\tau}[:,j]\right)_{\omega},\quad\text{where}\quad\bar{\mathbf{v}}^{(k)}_{t,\tau}=\operatorname{Pad}_{H}\!\left(\mathbf{v}_{t,\tau,1:k,:}\right)\in\mathbb{R}^{H\times d_{a}},Let \Omega_{\mathrm{hi}}\subset\Omega contain frequencies above cutoff fraction c_{\mathrm{cut}} of the discrete frequency grid. The fraction of energy in this band, aggregated across action dimensions, is
r_{\mathrm{hi},t,\tau}(k)=\frac{\sum_{\omega\in\Omega_{\mathrm{hi}}}\sum_{j=1}^{d_{a}}\left|\widehat{\mathbf{V}}^{(k)}_{t,\tau}(\omega,j)\right|^{2}}{\sum_{\omega\in\Omega}\sum_{j=1}^{d_{a}}\left|\widehat{\mathbf{V}}^{(k)}_{t,\tau}(\omega,j)\right|^{2}}.High-frequency energy reflects rapid variation along the future-action axis. To track changes during generation, we define an early baseline \bar{r}_{\mathrm{hi},t,0}(k)=\frac{1}{W_{\tau}}\sum_{\tau=0}^{W_{\tau}-1}r_{\mathrm{hi},t,\tau}(k) from the first W_{\tau} sampling steps. The raw intra-chunk instability is its root mean square (RMS) deviation over all denoising steps, including the baseline window:
z_{\mathrm{intra},t}(k)=\left[\frac{1}{T}\sum_{\tau=0}^{T-1}\left(r_{\mathrm{hi},t,\tau}(k)-\bar{r}_{\mathrm{hi},t,0}(k)\right)^{2}\right]^{1/2}.The early window supplies a reference rather than being discarded from the RMS calculation. The score measures deviation from this reference, not simply the final chunk’s high-frequency energy. Because raw scales vary across tasks, policies, and action normalizations, we min-max normalize this score within \mathcal{K} and set q_{\mathrm{intra},t}(k)=\exp[-\alpha_{\mathrm{intra}}\tilde{z}_{\mathrm{intra},t}(k)], where \alpha_{\mathrm{intra}}>0 controls the penalty strength. Larger spectral deviations thus receive lower quality. Normalization makes this a relative comparison among candidate prefixes at the current replan. The resulting quality need not be comparable in absolute scale across unrelated episodes. Fig. 2(a) uses the episode mean \bar{z}_{\mathrm{intra}}(k) of this same replan-level score, with the exact aggregation specified in the caption.
3.3 Inter-Chunk Continuity
An internally stable prefix may still be incompatible with recent motion. For continuity window W_{h}, we concatenate the available history suffix and candidate future prefix:
\mathbf{S}_{t}^{(k)}=\left[\operatorname{Suffix}_{(W_{h}-k)_{+}}(\mathcal{H}_{t}),\hat{\mathbf{a}}_{t,1},\ldots,\hat{\mathbf{a}}_{t,k}\right].Here (x)_{+}=\max(0,x). Let \delta_{i}^{(k)}=\lVert\mathbf{S}_{t,i+1}^{(k)}-\mathbf{S}_{t,i}^{(k)}\rVert_{2} denote first-difference speed along the stitched trajectory. The raw inter-chunk discontinuity is the speed coefficient of variation:
u_{\mathrm{inter},t}(k)=\operatorname{Std}\left(\{\delta_{i}^{(k)}\}_{i}\right)/\operatorname{Mean}\left(\{\delta_{i}^{(k)}\}_{i}\right).This proxy penalizes irregular speed, including boundary jumps and stop-and-go motion. A smaller value indicates a more uniform stitched trajectory, whereas a larger value can reflect an abrupt correction despite a smooth candidate viewed in isolation. As above, we min-max normalize within \mathcal{K} and set q_{\mathrm{inter},t}(k)=\exp[-\beta_{\mathrm{inter}}\tilde{u}_{\mathrm{inter},t}(k)], with \beta_{\mathrm{inter}}>0. Fig. 2(b) reports the corresponding episode mean \bar{u}_{\mathrm{inter}}(k) over valid replans. Failed episodes have higher median raw scores for both evidence terms across the candidate horizons. The shaded bands show interquartile ranges.
3.4 Dual-Evidence Fusion
We combine internal stability and compatibility with recent motion into
q_{\mathrm{mix},t}(k)=(1-\lambda)q_{\mathrm{intra},t}(k)+\lambda q_{\mathrm{inter},t}(k),\qquad\lambda\in[0,1].The default is \lambda=0.5, with \lambda=0 and \lambda=1 recovering the intra-only and inter-only ablations. This score supplies action-expert evidence for horizon inference, not a deterministic execution rule or a calibrated task-success probability.
Fig. 2(c) summarizes episode risks by the 75th percentiles of 1-q_{\mathrm{intra}} and 1-q_{\mathrm{inter}} over replans, then splits each risk at its median. Failure rises from 50.2% when both risks are low to 97.4% when both are high, with intermediate rates when only one is high. These associations motivate combining the evidence, but they do not establish that a low-risk prefix guarantees success. The diagnostic episodes use fixed K=H=50, so the comparison characterizes the association of evidence with outcomes without selecting trajectories based on AHS decisions. Full distributions and aggregation details are reported in Appendix B.
4 Action-aware Horizon Selection and Query-based Adaptation
AHS uses the candidate quality q_{\mathrm{mix},t}(k) from Sec. 3 to maintain an online reliability posterior. QHA optionally learns a context-conditioned dense horizon prior from the same evidence (Fig. 3). Neither updates the base action generator.
4.1 Action-aware Horizon Selector
Online reliability posterior.
Manipulation alternates between phases such as approach, contact, transport, and placement. A horizon that was useful in a stable phase may become undesirable after a contact change, motivating memory that can also forget. To retain information across noisy replans, AHS maintains a Beta state for each k\in\mathcal{K}, initialized by a_{0}(k)=b_{0}(k)=1. This is episode-local memory of evidence-derived reliability, not a supervised task-success model. Following Thompson sampling (Russo et al., 2018), we draw a sample and combine it with the current quality:
\hat{\xi}_{t}(k)\sim\operatorname{Beta}(a_{t}(k),b_{t}(k)),\qquad s_{t}(k)=\hat{\xi}_{t}(k)\,q_{\mathrm{mix},t}(k).Here q_{\mathrm{mix},t} measures the current chunk, while sampled reliability reflects accumulated episode-local evidence. Their product combines both. Except for uniform exploration with probability \epsilon_{\mathrm{exp}}, temperature scaling and expected-round selection give:
\mu_{t}(k)\propto s_{t}(k)^{1/T_{\mathrm{sel}}},\quad\sum_{k\in\mathcal{K}}\mu_{t}(k)=1,\qquad K_{t}=\operatorname{clip}_{[1,H_{t}^{\mathrm{avail}}]}\operatorname{round}\!\Bigl[\sum_{k\in\mathcal{K}}k\,\mu_{t}(k)\Bigr].Here H_{t}^{\mathrm{avail}} is the available chunk length. If all scores degenerate, \mu_{t} falls back to uniform over valid candidates. Expected-round combines candidate preferences; exploration samples one valid candidate uniformly. Clipping keeps the prefix within the available prediction. AHS executes \hat{\mathbf{a}}_{t,1:K_{t}}, appends the executed actions to \mathcal{H}, and replans from the next observation.
Posterior update with kernel forgetting.
After execution, the soft feedback is y_{t}=\sum_{k\in\mathcal{K}}\mu_{t}(k)q_{\mathrm{mix},t}(k), or the sampled candidate quality on exploration steps. A Gaussian kernel w_{t}(k)\propto\exp[-(k-K_{t})^{2}/(2\sigma_{K}^{2})], normalized over \mathcal{K}, shares feedback among nearby horizons. Exponential forgetting (Raj and Kalyani, 2017) updates the state:
\displaystyle a_{t+1}(k) \\ \displaystyle=\rho_{f}a_{t}(k)+\eta\,w_{t}(k)y_{t}, \\ \displaystyle b_{t+1}(k) \\ \displaystyle=\rho_{f}b_{t}(k)+\eta\,w_{t}(k)(1-y_{t}),where \rho_{f}, \eta, and \sigma_{K} control forgetting, update strength, and neighborhood sharing. Decaying old evidence lets the selector adapt as the task changes phase. The feedback comes from action-expert quality, not an observed task-success label. Neighboring candidates receive shared soft evidence rather than independent rollout outcomes. Kernel sharing couples nearby lengths; forgetting discounts earlier phases.
4.2 Training-Time Query-Based Horizon Adapter
- RoboTwin2.0 Easy and Hard
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- \pi_{0.5}
- 50 tasks
- multitask post-training
- 56.70
- 63.50
- +6.80
- \pi_{0}
- 8 tasks
- task-specific post-training
- 18.19
- 21.38
- +3.19
- \pi_{0.5}
- 8 tasks
- task-specific post-training
- 29.63
- 36.25
- +6.63
- Fast-WAM
- 8 tasks
- multitask post-training
- 86.81
- 88.13
- +1.31
- RoboCasa GR1 Tabletop
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- \pi_{0.5}
- 24 tasks
- multitask post-training
- 40.08
- 42.50
- +2.42
- GR00T N1.5
- 24 tasks
- zero-shot
- 43.25
- 44.92
- +1.67
- GR00T N1.6
- 24 tasks
- zero-shot
- 47.61
- 51.42
- +3.80
- Qwen3GR00T
- 24 tasks
- multitask post-training
- 47.83
- 57.50
- +9.67
QHA predicts a dense distribution p_{\phi}(K_{t}=h\mid\mathbf{C}_{t},\mathbf{Z}_{t}^{A}) for h\in\{1,\ldots,H\} from VLM context tokens \mathbf{C}_{t} and action-latent tokens \mathbf{Z}_{t}^{A}. With a sparse candidate grid, AHS explicitly scores each candidate, while QHA predicts a preference for every action step in one forward pass. Learned horizon queries use bridge cross-attention over both token streams (Fig. 3). Architecture and teacher construction are specified in Appendix A.6. The teacher normalizes dense evidence without online Beta memory:
\mu_{t}^{\star}(h)\propto q_{\mathrm{mix},t}(h)^{1/T_{\mathrm{teach}}},\qquad\sum_{h=1}^{H}\mu_{t}^{\star}(h)=1,\qquad h=1,\ldots,H.Only QHA is trained, minimizing \mathcal{L}_{\mathrm{QHA}}=\operatorname{KL}(\mu_{t}^{\star}\,\|\,p_{\phi}). QHA thus encodes horizon preferences across training episodes in its parameters, providing cross-episode memory that complements AHS’s episode-local state. At deployment, let \tilde{q}_{\mathrm{mix},t}(h) denote quality on the dense grid, interpolated from sparse evidence or scored directly on that grid. QHA supplies a prior, combined with dense Beta reliability samples as
s_{\mathrm{eff},t}(h)=\hat{\xi}_{t}(h)\,\tilde{q}_{\mathrm{mix},t}(h)\,p_{\phi}(K_{t}=h\mid\mathbf{C}_{t},\mathbf{Z}_{t}^{A})^{\gamma},\qquad h=1,\ldots,H,where \gamma\geq 0 controls prior strength. Applying temperature scaling and the selection rule above to this dense score yields AHS+QHA, combining learned preferences with current episode evidence. The dense prior supplies a context-conditioned preference before the episode-local state has accumulated much evidence. It does not remove the need to record generation traces or evaluate AHS candidates in the hybrid setting, as those computations provide the online correction to the prior.
5 Experiments
- —
- \pi_{0} (Black et al., 2024)
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- \pi_{0.5} (Intelligence et al., 2025)
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- Task
- Base
- 자료 없음
- AHS
- 자료 없음
- AHS+QHA
- 자료 없음
- Base
- 자료 없음
- AHS
- 자료 없음
- AHS+QHA
- 자료 없음
- —
- Easy
- Hard
- Easy
- Hard
- Easy
- Hard
- Easy
- Hard
- Easy
- Hard
- Easy
- Hard
- Blocks Ranking RGB
- 19.00
- 0.00
- 28.00
- 1.00
- 31.00
- 5.00
- 36.00
- 20.00
- 48.00
- 29.00
- 56.00
- 33.00
- Handover Block
- 41.00
- 10.00
- 41.00
- 5.00
- 64.00
- 9.00
- 44.00
- 14.00
- 44.00
- 14.00
- 32.00
- 10.00
- Handover Mic
- 100.00
- 2.00
- 100.00
- 24.00
- 100.00
- 30.00
- 98.00
- 64.00
- 100.00
- 56.00
- 98.00
- 59.00
- Hanging Mug
- 17.00
- 3.00
- 14.00
- 9.00
- 23.00
- 10.00
- 14.00
- 8.00
- 13.00
- 12.00
- 13.00
- 14.00
- Place A2B Left
- 24.00
- 1.00
- 34.00
- 0.00
- 39.00
- 1.00
- 41.00
- 4.00
- 47.00
- 4.00
- 47.00
- 15.00
- Place Bread Basket
- 11.00
- 8.00
- 18.00
- 16.00
- 15.00
- 10.00
- 30.00
- 17.00
- 50.00
- 29.00
- 55.00
- 41.00
- Place Bread Skillet
- 15.00
- 2.00
- 23.00
- 5.00
- 23.00
- 2.00
- 24.00
- 10.00
- 36.00
- 19.00
- 44.00
- 19.00
- Place Can Basket
- 33.00
- 5.00
- 22.00
- 2.00
- 33.00
- 3.00
- 35.00
- 15.00
- 50.00
- 29.00
- 55.00
- 34.00
- Mean by setting
- 32.50
- 3.88
- 35.00
- 7.75
- 41.00
- 8.75
- 40.25
- 19.00
- 48.50
- 24.00
- 50.00
- 28.13
- Overall
- 18.19
- 자료 없음
- 21.38
- 자료 없음
- 24.88
- 자료 없음
- 29.63
- 자료 없음
- 36.25
- 자료 없음
- 39.06
- 자료 없음
We evaluate on RoboTwin2.0 (Chen et al., 2025), RoboCasa GR1 Tabletop (Nasiriany et al., 2024), and real robots. Each Base–AHS comparison fixes the policy checkpoint. AHS changes execution without updating policy weights. Simulation tables report success rates (%). In Tabs. 1 and 2, tasks are weighted equally after averaging Easy and Hard within each RoboTwin2.0 task. Rounding follows aggregation. Easy and Hard denote clean and randomized settings. Appendix A specifies checkpoints, candidate horizons, and evaluation protocols.
5.1 AHS across Simulation Benchmarks
Tab. 1 tests AHS across benchmarks and training regimes. Its task scope and training recipe distinguish evaluation coverage from checkpoint provenance. Absolute scores across regimes are not matched-policy comparisons. The evaluation spans task-specific post-training, multitask post-training, and zero-shot deployment on the target benchmark. Gains in these regimes test whether execution adaptation remains useful across the evaluated backbones. They do not imply that AHS repairs every failed task or replaces policy training.
RoboTwin2.0.
The full-suite evaluation uses one multitask-post-trained \pi_{0.5} (Intelligence et al., 2025) checkpoint on 50 tasks, with 20 rollouts per task and setting. AHS improves success from 56.70% to 63.50% (+6.80 percentage points). The eight-task evaluation uses task-specific \pi_{0} (Black et al., 2024) and \pi_{0.5} checkpoints and multitask Fast-WAM (Yuan et al., 2026), with 100 rollouts per task and setting. All three improve in aggregate. These are distinct checkpoint and evaluation cohorts, not subset and full-suite results for one policy. Fast-WAM improves from 86.81% to 88.13%, showing that execution adaptation can still help a stronger baseline, although its aggregate gain is smaller than those of the task-specific policies. Appendix Tabs. 8 and 9 retain full per-task outcomes, including regressions.
RoboCasa GR1 Tabletop.
We evaluate all 24 tasks; the \pi_{0.5} control uses 50 rollouts per task. The multitask-post-trained \pi_{0.5} and Qwen3GR00T (Ye et al., 2026a; Community, 2026) improve from 40.08% to 42.50% and from 47.83% to 57.50%, respectively. Both zero-shot Isaac-GR00T baselines (Bjorck et al., 2025) also improve in aggregate. All four use \mathcal{K}=\{4,8,12,16\}. Appendix Tab. 10 provides the complete breakdown. Fixed-horizon, action-only, temporal-memory, and AAC (Liang et al., 2026) comparisons appear in Appendix C.5.
5.2 QHA Augmentation and Transfer
Augmentation on QHA training tasks.
Each policy family uses one QHA head trained on eight tasks with the base policies frozen (Tab. 2A). AHS+QHA improves overall success over AHS from 21.38% to 24.88% for \pi_{0} and from 36.25% to 39.06% for \pi_{0.5}. Both Easy and Hard means improve, but \pi_{0.5} regresses on Handover Block in both settings. This may reflect a mismatch between the shared horizon prior and the feedback timing required during bimanual object transfer.
Transfer to tasks held out from QHA training.
A separate \pi_{0.5} head trains on six tasks and tests on two held-out tasks (Tab. 2B). AHS+QHA improves the equal-task mean from 40.00% to 42.75%. Blocks Ranking RGB improves from 38.50% to 43.50%, compared with 41.50% to 42.00% on Place Bread Basket. Only the horizon head is task-held-out: base policies remain task-specific. Appendix C.1 gives the split, per-setting results, QHA-only control, and decision agreement.
5.3 Real-World Deployment
\pi_{0.5} and \pi_{0.5} + AHS over 15 rollouts per task and method. Error bars show \pm one standard error of the mean across rollouts.We deploy \pi_{0.5} on the AgileX COBOT Magic ALOHA-style bimanual platform, comparing fixed K=25 with AHS over 15 rollouts on each of four household tasks (Fig. 4). Object positions vary across rollouts to test spatial generalization. Scores average predefined sub-steps and rollouts to measure partial progress, rather than binary success (Appendix A.1). AHS raises the equal-task mean from 50.4% to 57.5%, with the largest gain on Duck Toy to Drawer: 50.0% to 63.3%.
5.4 Ablation Studies
- Fixed
- –
- 22.00
- –
- 95.83
- –
- 4.22
- 22.43
- \mathcal{K}_{5}
- 5
- 33.00
- +11.00
- 96.50
- 0.70
- 8.36
- 23.17
- \mathcal{K}_{10}
- 10
- 32.50
- +10.50
- 96.84
- 1.05
- 9.34
- 23.08
- \mathcal{K}_{50}
- 50
- 34.88
- +12.88
- 99.28
- 3.60
- 10.57
- 21.39
- Base \pi_{0} (default)
- 20.75
- 4.00
- 12.38
- Inter-chunk only
- 23.50
- 4.00
- 13.75
- Intra-chunk only
- 23.75
- 5.25
- 14.50
- Both evidence terms, no posterior
- 25.50
- 3.75
- 14.63
- Full AHS
- 24.25
- 5.75
- 15.00
- Full AHS+QHA
- 27.50
- 4.00
- 15.75
All three studies use Place A2B Left, Place Bread Basket, Place Bread Skillet, and Place Can Basket, with 100 rollouts per task and condition. All three studies include Easy and Hard.
Candidate-set scaling.
On \pi_{0.5}, five candidates raise success from 22.0% with fixed K=50 to 33.0%, versus 34.9% with 50 candidates (Tab. 3). Added per-replan cost rises from 0.67 to 3.45 ms, or 0.70–3.60% of policy inference time. Sparse candidates thus capture most of the gain. The table’s episode times cover successes only, so they do not establish all-attempt cost. Appendix D.1 reports all-episode costs for additional cohorts. Appendix E gives hyperparameter sweeps.
Component ablation.
On \pi_{0}, combining both evidence terms without memory gives 14.6% overall but 3.75% on Hard (Tab. 4). Full AHS reaches 15.0% overall and the best Hard result, 5.75%, supporting temporal memory. QHA raises overall success to 15.8% while Hard falls to 4.00%. Appendix C.5 compares memory designs with shared evidence and candidates.
Latency and asynchronous execution.
We combine AHS with real-time chunking (RTC; Black et al., 2025) on task-specific \pi_{0.5} policies (Fig. 5). All predict H=50 actions. Fixed uses target budget K=40, and AHS scores candidates in \{10,20,30,40\} without QHA. At +200 ms, RTC reduces AHS waiting from 7.893 to 0.344 s/episode on Hard and from 6.310 to 0.344 s on Easy. Within RTC, AHS lowers mean observation age from 4.92 to 3.07 s on Hard and from 5.00 to 3.18 s on Easy, while SR rises from 22.00% to 26.75% and from 39.00% to 42.75%, respectively. This requires more policy calls and model computation (Appendix D.2). RTC does not uniformly improve SR over synchronous AHS: at +200 ms on Easy, SR is 42.75% versus 44.75%.
\pi_{0.5} Hard. (a) Mean waiting and observation age. Points run left to right as +0/+100/+200 ms additional delay. (b) Final SR. (c) SR(t) at +200 ms. All 400 outcomes per condition are included. Bars and shading are marginal/pointwise 95% paired-seed bootstrap intervals within four tasks. Time uses a controlled physical clock with a 142 ms base delay, not native deployment timing. Easy curves and full results: Appendix D.2.6 Conclusion
ChunkTrust adapts robot-policy execution horizons using action-expert evidence. AHS combines spectral stability, inter-chunk continuity, and online temporal memory, while QHA learns a context-conditioned horizon prior from the same evidence without changing the base policy. Experiments support aggregate gains across two simulation benchmarks and real-world manipulation, while transfer, ablation, and cost analyses characterize where adaptation helps. The method requires accessible generation traces, and gains are not uniform across tasks. The evaluation also separates per-replan overhead from episode-level cost: a small selector cost does not imply identical total rollout time. Task breakdowns likewise expose regressions that aggregate gains can conceal, particularly when a learned prior is combined with online evidence. Future work includes multi-chunk horizon inference, cross-policy transfer of the learned prior, and broader real-world validation across hardware and domain shifts.
References
ChunkTrust: Adapting Execution Horizons for Robot Policies with Action-Expert Evidence Appendix
Appendix A Experimental Setup
We specify evaluation cohorts, checkpoints, scoring criteria, and implementation settings for the results in Sec. 5.
A.1 Real-world Experiments
Hardware Setup.
We conduct real-world experiments on an AgileX COBOT Magic platform configured as an ALOHA-style bimanual system (Fu et al., 2024a; Fu et al., 2024b), as shown in Fig. 6. The platform consists of four 6-DoF Piper arms, with two leader arms used for human teleoperation and two follower arms used for data collection and autonomous policy execution. The perception system includes three RealSense D435 cameras: one front-view camera and two wrist-mounted cameras, one on each follower arm.
Tasks and evaluation.
We collect 200 human-teleoperated demonstrations for each of four bimanual tasks and evaluate \pi_{0.5} with fixed K=25 or AHS over 15 rollouts per task. Data are recorded at 30 FPS, with task-relevant object positions varied across evaluation rollouts. Instructions are “Fold the towel with both arms,” “Move the bread to the plate with both arms,” “Move the drink to the basket with both arms,” and “Put the duck toy into the drawer with both arms.” These tasks cover approach, contact, transport, alignment, and placement (Fig. 4).
Process scores.
Each sub-step receives 0 for failure, 0.5 for recovered or imperfect completion, and 1 for smooth, accurate completion. We average equally over the task’s sub-steps and its rollouts, reporting the result as a percentage. Table 5 lists the task-specific criteria. Error bars in Fig. 4(b) show \pm s/\sqrt{15}, where s is the sample standard deviation of the 15 normalized rollout scores for each task and method. In the drawer task, the arm assignment depends on the layout: one arm opens/closes the drawer and the other manipulates the toy.
- Bread: grasp
- Fails to grasp
- Multiple attempts
- Smooth first attempt
- Bread: handover
- Handover fails
- Unstable/awkward receiving grasp
- Stable, aligned grasp
- Bread: place on plate
- Not placed
- Poor alignment or rough placement
- Clean placement
- Drink: push
- Only tilts, no useful displacement
- Acceptable position, tilt/misalignment
- Good position, stable alignment
- Drink: grasp
- Fails to grasp
- Multiple attempts
- Smooth first attempt
- Drink: place in basket
- Fails or drops drink
- Rough placement/collision
- Clean placement
- Towel: first fold, second fold
- Not folded over
- Folded, misaligned
- Folded, well aligned
- Duck: open drawer, grasp toy, place toy, close drawer
- Sub-step fails
- Multiple attempts
- Smooth first attempt
A.2 Simulation Experiments
RoboTwin2.0 (Chen et al., 2025) contains 50 bimanual manipulation tasks with strong domain randomization. The complete-suite evaluation in Tab. 1 uses one multitask \pi_{0.5} checkpoint on all 50 tasks, with 20 rollouts per task, setting, and method (2,000 per method). The eight-task evaluation uses 100 rollouts per task, setting, and method (1,600 per method), on Blocks Ranking RGB, Handover Block, Handover Mic, Hanging Mug, Place A2B Left, Place Bread Basket, Place Bread Skillet, and Place Can Basket. Each task is evaluated under two settings: Easy (clean scenes) and Hard (randomized object poses, lighting, and distractor placement). Results report success rate, averaged equally across the tasks and settings in each cohort.
RoboCasa GR1 Tabletop (Nasiriany et al., 2024) provides 24 pick and-place tasks spanning everyday object categories and novel source–target combinations. We evaluate all 24 tasks, using 50 rollouts per task for the \pi_{0.5} control. Results for the other policies use their respective evaluation budgets and reported numerical precision. Task success is determined by the native RoboCasa success checker.
A.3 Base Policy Checkpoints
For the eight-task RoboTwin2.0 evaluation, checkpoints follow the protocol of each base policy:
\pi_{0}(Black et al., 2024): we train LoRA adapters (Hu et al., 2022) with the released RoboTwin2.0 fine-tuning protocol. For each evaluated task, one adapter is trained on that task’s clean50 demonstrations and then evaluated. The action chunk size isH=50, with a flow-matching action expert (Lipman et al., 2023) using 10 sampling steps.\pi_{0.5}(Intelligence et al., 2025): we use the official RoboTwin2.0 fine-tuning recipe for full-parameter fine-tuning. For each evaluated task, one task-specific checkpoint is trained on clean50 demonstrations and then evaluated. The action chunk size isH=50, with a flow-matching action expert (Lipman et al., 2023).- X-VLA (Zheng et al., 2026): we evaluate the released X-VLA RoboTwin2.0 checkpoint. The action chunk size is
H=30, with a flow-matching action expert. - Fast-WAM (Yuan et al., 2026): we evaluate the released Fast-WAM multitask RoboTwin2.0 checkpoint on our eight-task subset, rather than post-training a separate policy for each task. The action chunk size is
H=32, with a flow-matching action expert.
Multitask \pi_{0.5} checkpoints.
The complete-suite controls initialize from the official pi05_base and use benchmark-specific full-parameter post-training. RoboTwin2.0 uses 50 clean demonstrations per task (2,500 trajectories and 549,787 frames). RoboCasa GR1 uses 24,000 demonstrations and 6,020,058 frames across 24 tasks. Both use task-uniform sampling, one global quantile normalizer per benchmark, global batch size 256, seed 42, and eight H100 GPUs for training, with the checkpoint fixed at step 30,000. RoboTwin uses H=50 and Base K=50, while RoboCasa uses H=16 and Base K=16. The latter maps the raw 44-dimensional GR1 interface to 29 effective absolute controls in the padded policy interface. These controls share a policy architecture, not weights across benchmarks. Within each benchmark, Base and AHS use identical policy weights. Their evaluation runs on one RTX 4090.
For the other RoboCasa policies, Isaac-GR00T N1.5 and Isaac-GR00T N1.6 (Bjorck et al., 2025) use released base checkpoints without additional benchmark post-training. Qwen3GR00T uses the released 24-task multitask-post-trained GR1 checkpoint from StarVLA (Ye et al., 2026a; Community, 2026). Thus, “zero-shot” in Tab. 1 describes benchmark adaptation, not an absence of robot pretraining.
A.4 AHS Hyperparameter Configurations
Candidate sets and continuity windows depend on the policy and benchmark:
- RoboTwin2.0, \pi_{0} and \pi_{0.5}
- \{10,20,30,40,50\}
- 60
- RoboTwin2.0, X-VLA and Fast-WAM
- \{10,20,30\}
- 40
- RoboCasa GR1, all policies
- \{4,8,12,16\}
- 20
The scaling ablation compares \mathcal{K}_{5}=\{10,20,30,40,50\}, \mathcal{K}_{10}=\{5,10,\ldots,50\}, and \mathcal{K}_{50}=\{1,\ldots,50\}. Shared defaults are \alpha_{\mathrm{intra}}=\beta_{\mathrm{inter}}=4, \lambda=0.5, a_{0}=b_{0}=1, \rho_{f}=0.99, \eta=1, and \sigma_{K}=10. The eight-task study uses expected-round selection at T_{\mathrm{sel}}=0.8, with Thompson exploration probability \epsilon_{\mathrm{exp}}=0.05. The 50-task evaluation uses T_{\mathrm{sel}}=1.0, and the RoboCasa \pi_{0.5} control uses 0.8.
Spectral scoring zero-pads each velocity prefix to H before FFT, uses the first W_{\tau}=3 denoising steps as its reference, window RMS for z, and cutoff fraction c_{\mathrm{cut}}=0.25. Without sufficient executed history, AHS uses intra-chunk-only scoring (\lambda=0). Costs appear in Appendix D.1.
A.5 Evaluation Protocol
For every (base policy, task, setting) combination, we run a fixed number of rollouts with AHS enabled. Each rollout starts from the standard initial state distribution of the benchmark. The online AHS posterior is reset to its prior (a_{0}=b_{0}=1) at the beginning of each episode, so its rollout history is episode-local. When QHA is used, its learned weights are reused across episodes without online training. Success is determined by the benchmark’s native success checker. We report the raw success rate (percentage of successful rollouts) and equal-weight averages over the stated task and setting groups. Absolute differences are in percentage points. The \Delta (pp) columns subtract the corresponding success-rate percentages, rather than reporting relative percentage changes.
In the 50-task evaluation, 62 of the 100 task–setting conditions use exact same-seed and instruction pairing between Base and AHS. The other 38 use deterministic, method-specific fallback identities selected without outcome information after repeated scene-construction failures. The full-suite averages include both groups. Identical episode identities are not assumed for the fallback group.
For the main RoboTwin2.0 benchmarks, Fast-WAM inference uses one NVIDIA H100 (80 GB); \pi_{0}, \pi_{0.5}, and X-VLA use one RTX 4090 (24 GB). The RoboCasa \pi_{0.5} control also uses one RTX 4090. Hardware details for separate runtime profiles are discussed in Appendix D.1. Replanning frequency depends on the selected horizon and execution scheduler.
k=10 and k=50, with dashed vertical lines marking the history–prediction join. Curves show means and bands show P10–P90 across episode profiles, not confidence intervals or the IQR bands used in the main figure.A.6 QHA Training Configurations
Training-task protocol (Tab. 2A).
We train one shared QHA per policy family (\pi_{0} or \pi_{0.5}) across all eight tasks, keeping the task-specific base generators frozen. Training uses their clean50 Aloha-AgileX LeRobot datasets, with task-balanced batches of 256 (32 per task). Inputs comprise three RGB streams, robot state, action windows, VLM context tokens \mathbf{C}_{t} with mask \mathbf{M}_{t}, and action-latent tokens \mathbf{Z}_{t}^{A} from the final chunk. The action-query window includes 64 history and 50 future steps. The frozen checkpoint generates the chunk \hat{\mathbf{A}}_{t}, evidence \mathcal{F}_{t}, and dense teacher \mu_{t}^{\star} online. The teacher follows Eq. 13 over h=1,\ldots,50 at T_{\mathrm{teach}}=1, using current evidence without episode-local Beta memory. The six-task held-out protocol is specified separately in Appendix C.1.
Architecture and optimization.
Context and action tokens are projected into a shared d=256 space with positional encodings. One bridge layer updates eight learned queries by cross-attending to masked context, then action tokens, followed by self-attention, each with residual connections (Vaswani et al., 2017). It uses eight attention heads, zero dropout, and an enabled fusion gate. Query pooling and an MLP produce 50 logits followed by a softmax. We minimize \operatorname{KL}(\mu_{t}^{\star}\|p_{\phi}) with weight 1 and numerical floor 10^{-6}. Only QHA parameters are updated and the base flow-matching loss is zero. AdamW runs for 10,000 steps with batch size 256, global-norm clipping 1, weight decay 10^{-10}, and EMA 0.99. The warmup-cosine schedule uses 1,000 warmup steps, peak learning rate 5\times 10^{-5}, and final rate 5\times 10^{-6}.
Deployment (Tab. 2A).
Each policy family uses its QHA checkpoint saved at step 10,000 after joint training on the eight tasks, with prior strength \gamma=1 in Eq. 14. Online evidence q_{\mathrm{mix},t}(k) is computed on \mathcal{K}=\{10,20,30,40,50\} and interpolated onto \{1,\ldots,H\} before Beta sampling. Both this sparse-evidence variant and the dense-evidence held-out variant maintain dense Beta states and apply temperature scaling, expected-round selection, and uniform exploration as in Sec. 4.1. Posterior feedback uses prior-weighted quality \tilde{q}_{\mathrm{mix},t}(h)p_{\phi}(h)^{\gamma}.
Held-out deployment (Tab. 2B).
This evaluation instead computes dense evidence over h=1,\ldots,50, with \gamma=1 and expected-round selection at T_{\mathrm{sel}}=1. Appendix C.1 specifies the six-task training split and checkpoints.
Appendix B Action-Expert Evidence Diagnostics
B.1 Evidence Distributions and Boundary Profiles
Data and aggregation.
We analyze the same 1,600 \pi_{0} RoboTwin2.0 episodes as Fig. 2: eight tasks, two settings, and 100 episodes per task–setting pair, with 291 successes and 1,309 failures. All episodes execute fixed K=H=50. The candidates k\in\{10,20,30,40,50\} are scored on these recorded traces. They are not five separate executed-horizon experiments.
For each candidate, we average finite scores over replans within an episode, then report medians and IQRs over episode means. The valid counts can differ between the two metrics because inter-chunk evidence requires executed history. The intra score follows Eq. 5, with prefixes zero-padded to H, a three-step baseline, and RMS over all ten denoising steps. Recomputed scores agree with the recorded values within 1.5\times 10^{-8}.
For boundary profiles, we compute first-difference speeds in the W_{h}=60 stitched window of Eq. 6, normalize by the window mean, and average within each episode before summarizing across episodes. This prevents longer episodes receiving extra weight.
B.2 Joint Risk and Episode Outcomes
At the executed K=50, each episode’s intra/inter risk is the 75th percentile of 1-q_{\mathrm{intra},t}(50) or 1-q_{\mathrm{inter},t}(50) over valid replans. Pooled median thresholds are 0.98168436 and 0.96278395, respectively. Values strictly below the threshold are low and ties are high, so group sizes can differ.
- Both low
- Low
- Low
- 313
- 157
- 50.2%
- Inter only
- Low
- High
- 193
- 144
- 74.6%
- Intra only
- High
- Low
- 487
- 417
- 85.6%
- Both high
- High
- High
- 607
- 591
- 97.4%
The map uses a 23\times 23 grid with 7% range padding and separable kernel [1,4,6,4,1]/16. Success/failure counts are smoothed separately before division (stabilizer 10^{-12}). Opacity scales to the 90th occupancy percentile, hiding cells below 10^{-3}.
Scope of the evidence.
These are pooled, retrospective associations from one policy and a fixed execution horizon. They do not establish causality, calibrated failure prediction, within-task effects independent of task difficulty, or how long a particular prefix remains reliable at the current replan. The benefit of adaptive execution is evaluated separately in the policy comparisons and ablations, not inferred from this heatmap.
Appendix C Full Simulation Benchmark Results
C.1 QHA Transfer to Held-out Tasks
For Tab. 2B, QHA is trained on Handover Block, Handover Mic, Hanging Mug, Place A2B Left, Place Bread Skillet, and Place Can Basket. Blocks Ranking RGB and Place Bread Basket are held out. The QHA head uses step 10,000, and frozen task-specific \pi_{0.5} bases use step-20,000 clean50 checkpoints without quantile normalization. All selectors use \{1,\ldots,50\} and expected-round selection at temperature 1.0. AHS+QHA uses \gamma=1.0. This cohort differs from the eight-task study in Tab. 2A.
AHS, QHA-only, and AHS+QHA share 100 episode identities per task and setting (400 per method). The Base results in Tab. 2B come from a separate evaluation outside this paired cohort. The transfer comparison is AHS+QHA versus AHS. Table 7 retains all four conditions: fusion improves two and reduces success in two. The aggregate gain does not establish uniform improvement or transfer of the base policy.
- Blocks Ranking RGB
- Easy
- 44.00
- 48.00
- 57.00
- +13.00
- Blocks Ranking RGB
- Hard
- 33.00
- 29.00
- 30.00
- -3.00
- Place Bread Basket
- Easy
- 52.00
- 50.00
- 49.00
- -3.00
- Place Bread Basket
- Hard
- 31.00
- 31.00
- 35.00
- +4.00
- Overall
- Both
- 40.00
- 39.50
- 42.75
- +2.75
Same-state decision disagreement.
In AHS+QHA traces, the QHA and AHS component choices differ at 77.02% of 12,573 replans, with a mean absolute gap of 1.83 action steps. Easy contributes 5,533 replans (77.17%, 1.84 steps) and Hard 7,040 (76.90%, 1.82 steps). This measures differences at the same states, without establishing complementarity or explaining success gains.
C.2 Complete 50-task RoboTwin2.0 Evaluation
- —
- Base
- AHS
- Base
- AHS
- 자료 없음
- adjust bottle
- 100.00
- 100.00
- 90.00
- 95.00
- +2.50
- beat block hammer
- 70.00
- 80.00
- 15.00
- 50.00
- +22.50
- blocks ranking rgb
- 80.00
- 95.00
- 40.00
- 70.00
- +22.50
- blocks ranking size
- 25.00
- 45.00
- 25.00
- 35.00
- +15.00
- click alarmclock
- 60.00
- 65.00
- 45.00
- 45.00
- +2.50
- click bell
- 35.00
- 70.00
- 55.00
- 55.00
- +17.50
- dump bin bigbin
- 95.00
- 95.00
- 80.00
- 75.00
- -2.50
- grab roller
- 100.00
- 100.00
- 90.00
- 90.00
- +0.00
- handover block
- 65.00
- 65.00
- 10.00
- 35.00
- +12.50
- handover mic
- 100.00
- 100.00
- 15.00
- 15.00
- +0.00
- hanging mug
- 20.00
- 30.00
- 5.00
- 10.00
- +7.50
- lift pot
- 65.00
- 70.00
- 30.00
- 35.00
- +5.00
- move can pot
- 50.00
- 75.00
- 10.00
- 25.00
- +20.00
- move pillbottle pad
- 55.00
- 80.00
- 30.00
- 30.00
- +12.50
- move playingcard away
- 90.00
- 80.00
- 75.00
- 75.00
- -5.00
- move stapler pad
- 20.00
- 35.00
- 20.00
- 5.00
- +0.00
- open laptop
- 95.00
- 90.00
- 85.00
- 70.00
- -10.00
- open microwave
- 85.00
- 50.00
- 35.00
- 25.00
- -22.50
- pick diverse bottles
- 45.00
- 55.00
- 35.00
- 70.00
- +22.50
- pick dual bottles
- 55.00
- 70.00
- 75.00
- 75.00
- +7.50
- place a2b left
- 75.00
- 70.00
- 55.00
- 65.00
- +2.50
- place a2b right
- 70.00
- 75.00
- 50.00
- 50.00
- +2.50
- place bread basket
- 55.00
- 75.00
- 55.00
- 55.00
- +10.00
- place bread skillet
- 35.00
- 55.00
- 60.00
- 55.00
- +7.50
- place burger fries
- 95.00
- 95.00
- 90.00
- 100.00
- +5.00
- place can basket
- 40.00
- 85.00
- 0.00
- 35.00
- +40.00
- place cans plasticbox
- 40.00
- 85.00
- 55.00
- 80.00
- +35.00
- place container plate
- 75.00
- 90.00
- 80.00
- 80.00
- +7.50
- place dual shoes
- 75.00
- 85.00
- 50.00
- 60.00
- +10.00
- place empty cup
- 100.00
- 100.00
- 65.00
- 70.00
- +2.50
- place fan
- 75.00
- 75.00
- 45.00
- 40.00
- -2.50
- place mouse pad
- 15.00
- 45.00
- 25.00
- 25.00
- +15.00
- place object basket
- 50.00
- 60.00
- 20.00
- 35.00
- +12.50
- place object scale
- 60.00
- 85.00
- 45.00
- 50.00
- +15.00
- place object stand
- 85.00
- 80.00
- 55.00
- 90.00
- +15.00
- place phone stand
- 65.00
- 65.00
- 45.00
- 35.00
- -5.00
- place shoe
- 85.00
- 90.00
- 75.00
- 55.00
- -7.50
- press stapler
- 100.00
- 90.00
- 60.00
- 70.00
- +0.00
- put bottles dustbin
- 45.00
- 65.00
- 50.00
- 55.00
- +12.50
- put object cabinet
- 25.00
- 35.00
- 20.00
- 15.00
- +2.50
- rotate qrcode
- 80.00
- 85.00
- 30.00
- 30.00
- +2.50
- scan object
- 20.00
- 40.00
- 20.00
- 50.00
- +25.00
- shake bottle
- 100.00
- 100.00
- 100.00
- 100.00
- +0.00
- shake bottle horizontally
- 100.00
- 100.00
- 100.00
- 100.00
- +0.00
- stack blocks three
- 75.00
- 55.00
- 35.00
- 45.00
- -5.00
- stack blocks two
- 85.00
- 100.00
- 80.00
- 90.00
- +12.50
- stack bowls three
- 60.00
- 65.00
- 35.00
- 30.00
- +0.00
- stack bowls two
- 100.00
- 85.00
- 85.00
- 90.00
- -5.00
- stamp seal
- 35.00
- 30.00
- 25.00
- 30.00
- +0.00
- turn switch
- 40.00
- 40.00
- 25.00
- 25.00
- +0.00
- Mean by setting
- 65.40
- 73.10
- 48.00
- 53.90
- +6.80
- Overall
- Base: 56.70
- 자료 없음
- AHS: 63.50
- 자료 없음
- +6.80
Table 8 expands the multitask \pi_{0.5} row of Tab. 1. Base and AHS succeed in 1,134 and 1,270 of 2,000 episodes, respectively, giving 56.70% and 63.50% overall. Easy success increases from 65.40% to 73.10%, and Hard from 48.00% to 53.90%. All 50 tasks are retained, including nine tasks with a negative change after averaging the two settings.
Statistical robustness and pairing scope.
The full-matrix gain is 6.80 percentage points, with a task-cluster 95% bootstrap interval of [3.75,\,9.95]. This interval resamples the 50 tasks, retaining both settings and methods within each task. Only 62 of 100 task–setting cells preserve exact seed-and-instruction pairing, while the remaining 38 cells use deterministic, outcome-blind, method-specific fallback identities after scene-construction failures. The 1,240 exact episode pairs yield a gain of 7.58 percentage points with a paired 95% interval of [5.08,\,10.08], resampling episodes within these cells. Both intervals use 10,000 bootstrap draws, but their statistical units and populations differ: the paired interval applies only to the exact-identity subset, not the full matrix.
C.3 Eight-task Policy-family Evaluation
The task-specific \pi_{0} and \pi_{0.5} matrices appear in Tab. 2A. Table 9 provides the corresponding Base–AHS breakdown for Fast-WAM and X-VLA (Zheng et al., 2026), evaluated on the same eight tasks with 100 rollouts per task and setting. Both use \mathcal{K}=\{10,20,30\}. Each Base–AHS comparison uses the same released policy checkpoint. Fast-WAM’s overall success increases from 86.81% to 88.13%. For X-VLA, AHS slightly improves overall success from 47.44% to 47.75%: Easy increases from 74.50% to 75.75%, while Hard decreases from 20.38% to 19.75%. Individual task regressions are retained for both policies.
- Task
- Base
- 자료 없음
- AHS
- 자료 없음
- Base
- 자료 없음
- AHS
- 자료 없음
- —
- Easy
- Hard
- Easy
- Hard
- Easy
- Hard
- Easy
- Hard
- Blocks Ranking RGB
- 99.00
- 98.00
- 100.00
- 100.00
- 87.00
- 34.00
- 95.00
- 38.00
- Handover Block
- 94.00
- 81.00
- 91.00
- 82.00
- 89.00
- 2.00
- 88.00
- 1.00
- Handover Mic
- 100.00
- 99.00
- 100.00
- 100.00
- 97.00
- 1.00
- 100.00
- 1.00
- Hanging Mug
- 66.00
- 64.00
- 71.00
- 69.00
- 32.00
- 8.00
- 35.00
- 4.00
- Place A2B Left
- 95.00
- 95.00
- 93.00
- 93.00
- 31.00
- 25.00
- 40.00
- 21.00
- Place Bread Basket
- 91.00
- 91.00
- 93.00
- 94.00
- 87.00
- 41.00
- 80.00
- 42.00
- Place Bread Skillet
- 91.00
- 91.00
- 94.00
- 94.00
- 86.00
- 22.00
- 88.00
- 23.00
- Place Can Basket
- 71.00
- 63.00
- 69.00
- 67.00
- 87.00
- 30.00
- 80.00
- 28.00
- Mean by setting
- 88.38
- 85.25
- 88.88
- 87.38
- 74.50
- 20.38
- 75.75
- 19.75
- Overall
- 86.81
- 자료 없음
- 88.13
- 자료 없음
- 47.44
- 자료 없음
- 47.75
- 자료 없음
C.4 RoboCasa GR1 Tabletop
Table 10 reports the full 24-task RoboCasa GR1 Tabletop breakdown for \pi_{0.5} and the three GR00T-family policies summarized in Tab. 1. QwenFAST (discrete tokens) and QwenPI (flow-matching action expert) are additional StarVLA baselines (Ye et al., 2026a; Community, 2026), both using Qwen3VL (Bai et al., 2025). Neither has an AHS counterpart in this table. For \pi_{0.5}, the 24-task means are 40.08% for Base and 42.50% for AHS, matching Tab. 1.
- PnPBottleToCabinetClose
- 38.0
- 26.0
- 64.0
- 68.0 (+4.0)
- 51.5
- 54.0 (+2.5)
- 46.0
- 64.0 (+18.0)
- 64.00
- 68.00 (+4.00)
- PnPCanToDrawerClose
- 44.0
- 62.0
- 18.0
- 12.0 (-6.0)
- 13.0
- 12.0 (-1.0)
- 80.0
- 80.0 (+0.0)
- 58.00
- 56.00 (-2.00)
- PnPCupToDrawerClose
- 56.0
- 42.0
- 12.0
- 4.0 (-8.0)
- 8.5
- 14.0 (+5.5)
- 54.0
- 52.0 (-2.0)
- 34.00
- 40.00 (+6.00)
- PnPMilkToMicrowaveClose
- 44.0
- 50.0
- 38.0
- 34.0 (-4.0)
- 14.0
- 20.0 (+6.0)
- 48.0
- 42.0 (-6.0)
- 40.00
- 44.00 (+4.00)
- PnPPotatoToMicrowaveClose
- 14.0
- 42.0
- 54.0
- 36.0 (-18.0)
- 41.5
- 50.0 (+8.5)
- 28.0
- 28.0 (+0.0)
- 22.00
- 30.00 (+8.00)
- PnPWineToCabinetClose
- 14.0
- 32.0
- 16.0
- 20.0 (+4.0)
- 16.5
- 24.0 (+7.5)
- 46.0
- 52.0 (+6.0)
- 52.00
- 56.00 (+4.00)
- PnPNovelFromCuttingboardToBasket
- 54.0
- 40.0
- 50.0
- 52.0 (+2.0)
- 58.0
- 54.0 (-4.0)
- 48.0
- 70.0 (+22.0)
- 34.00
- 34.00 (+0.00)
- PnPNovelFromCuttingboardToCardboardbox
- 42.0
- 46.0
- 36.0
- 34.0 (-2.0)
- 46.5
- 46.0 (-0.5)
- 40.0
- 54.0 (+14.0)
- 32.00
- 40.00 (+8.00)
- PnPNovelFromCuttingboardToPan
- 58.0
- 60.0
- 68.0
- 64.0 (-4.0)
- 68.5
- 80.0 (+11.5)
- 68.0
- 80.0 (+12.0)
- 54.00
- 58.00 (+4.00)
- PnPNovelFromCuttingboardToPot
- 58.0
- 40.0
- 34.0
- 56.0 (+22.0)
- 65.0
- 64.0 (-1.0)
- 52.0
- 76.0 (+24.0)
- 46.00
- 40.00 (-6.00)
- PnPNovelFromCuttingboardToTieredbasket
- 40.0
- 44.0
- 46.0
- 32.0 (-14.0)
- 46.5
- 54.0 (+7.5)
- 56.0
- 44.0 (-12.0)
- 22.00
- 28.00 (+6.00)
- PnPNovelFromPlacematToBasket
- 36.0
- 44.0
- 50.0
- 46.0 (-4.0)
- 58.5
- 48.0 (-10.5)
- 42.0
- 54.0 (+12.0)
- 42.00
- 46.00 (+4.00)
- PnPNovelFromPlacematToBowl
- 38.0
- 52.0
- 50.0
- 62.0 (+12.0)
- 57.5
- 60.0 (+2.5)
- 44.0
- 66.0 (+22.0)
- 34.00
- 32.00 (-2.00)
- PnPNovelFromPlacematToPlate
- 42.0
- 50.0
- 62.0
- 66.0 (+4.0)
- 63.0
- 82.0 (+19.0)
- 48.0
- 72.0 (+24.0)
- 46.00
- 48.00 (+2.00)
- PnPNovelFromPlacematToTieredshelf
- 18.0
- 28.0
- 14.0
- 26.0 (+12.0)
- 28.5
- 36.0 (+7.5)
- 18.0
- 20.0 (+2.0)
- 28.00
- 28.00 (+0.00)
- PnPNovelFromPlateToBowl
- 52.0
- 52.0
- 58.0
- 58.0 (+0.0)
- 57.0
- 58.0 (+1.0)
- 60.0
- 60.0 (+0.0)
- 40.00
- 44.00 (+4.00)
- PnPNovelFromPlateToCardboardbox
- 30.0
- 40.0
- 40.0
- 48.0 (+8.0)
- 43.5
- 58.0 (+14.5)
- 50.0
- 54.0 (+4.0)
- 28.00
- 30.00 (+2.00)
- PnPNovelFromPlateToPan
- 48.0
- 36.0
- 44.0
- 48.0 (+4.0)
- 51.0
- 68.0 (+17.0)
- 54.0
- 54.0 (+0.0)
- 32.00
- 36.00 (+4.00)
- PnPNovelFromPlateToPlate
- 50.0
- 48.0
- 66.0
- 74.0 (+8.0)
- 78.7
- 82.0 (+3.3)
- 70.0
- 74.0 (+4.0)
- 54.00
- 54.00 (+0.00)
- PnPNovelFromTrayToCardboardbox
- 28.0
- 34.0
- 44.0
- 52.0 (+8.0)
- 51.5
- 48.0 (-3.5)
- 38.0
- 56.0 (+18.0)
- 48.00
- 50.00 (+2.00)
- PnPNovelFromTrayToPlate
- 34.0
- 64.0
- 50.0
- 60.0 (+10.0)
- 71.0
- 68.0 (-3.0)
- 56.0
- 62.0 (+6.0)
- 40.00
- 40.00 (+0.00)
- PnPNovelFromTrayToPot
- 46.0
- 44.0
- 46.0
- 50.0 (+4.0)
- 64.5
- 64.0 (-0.5)
- 50.0
- 66.0 (+16.0)
- 54.00
- 54.00 (+0.00)
- PnPNovelFromTrayToTieredbasket
- 36.0
- 50.0
- 44.0
- 38.0 (-6.0)
- 57.0
- 56.0 (-1.0)
- 36.0
- 56.0 (+20.0)
- 34.00
- 36.00 (+2.00)
- PnPNovelFromTrayToTieredshelf
- 16.0
- 28.0
- 34.0
- 38.0 (+4.0)
- 31.5
- 34.0 (+2.5)
- 16.0
- 44.0 (+28.0)
- 24.00
- 28.00 (+4.00)
- Average
- 39.00
- 43.92
- 43.25
- 44.92 (+1.67)
- 47.61
- 51.42 (+3.80)
- 47.83
- 57.50 (+9.67)
- 40.08
- 42.50 (+2.42)
C.5 Horizon-selection Comparisons
Table 11 reports additional controls for horizon selection, separately from the cross-policy results in Tab. 1.
Fixed, action-only, and evidence-update selectors.
The \pi_{0} study uses Place A2B Left, Place Bread Basket, Place Bread Skillet, and Place Can Basket under Easy and Hard settings, with 16 paired episodes per task and setting (128 per selector). Instantaneous is a separate reference. This cohort differs from the 100-rollout studies in Tabs. 3 and 4. The global fixed horizon K=20 is selected retrospectively from previous evaluations, not from a held-out validation set. Jerk-min uses an action-only smoothness criterion on \mathcal{K}=\{10,20,30,40,50\}.
The evidence-update variants share this candidate grid and the same instantaneous evidence, using expected-round selection with temperature 0.8. Instantaneous uses q_{\mathrm{mix},t}(k), whereas EMA maintains m_{t}(k)=\rho_{\mathrm{EMA}}m_{t-1}(k)+(1-\rho_{\mathrm{EMA}})q_{\mathrm{mix},t}(k) with \rho_{\mathrm{EMA}}=0.99. Neither comparator uses Beta counts or kernel neighborhood sharing. Full Beta is the shared AHS reference. Results describe success and call-count trade-offs without exact compute matching. The paired 95% SR-difference interval between AHS and Jerk-min includes zero, so this compact study does not establish an SR advantage.
- Global fixed (K=20)
- 14.84
- 25.96
- 2.75
- 2.75
- 54.33
- Jerk-min
- 14.06
- 21.14
- 2.28
- 2.28
- 52.01
- EMA q_{\mathrm{mix}}
- 10.94
- 19.72
- 2.05
- 2.07
- 46.72
- AHS (Full Beta)
- 15.63
- 20.71
- 2.26
- 2.28
- 53.47
- Instantaneous q_{\mathrm{mix}} (reference)
- 13.28
- 19.58
- —
- —
- —
AAC-core on RoboCasa.
We evaluate a joint-space adapter of AAC (Liang et al., 2026) on Qwen3GR00T over 24 tasks with 50 rollouts per task. Each replan draws 20 action-head samples under the same observation and instruction. A prefix-entropy elbow and a movement guard determine K\in\{2,\ldots,16\}, and the first sampled chunk supplies the executed actions. The adapter operates on the policy’s native 29-dimensional absolute joint targets. It is not an exact reproduction of the published Cartesian-action implementation. AAC-core obtains 52.58% success. Base/AHS values from separate evaluations in Tab. 1 provide non-paired context only. We do not report a paired difference or infer compute equivalence from policy call counts.
Appendix D Runtime and Latency
D.1 Policy-call and Runtime Costs
We group runtime measurements by timing scope. Means include all episodes, including failures. Policy inference excludes separately timed horizon selection, but total inference includes it. Episode wall time includes simulation and within-episode overhead, not physical robot execution time. These are descriptive profiles, not hardware-matched cross-policy rankings.
Compact-selector timings are included with their success rates in Tab. 11.
- Method
- episode
- (s/episode)
- time (s/episode)
- A. Full 50-task evaluation 2,000 episodes per method
- 자료 없음
- 자료 없음
- 자료 없음
- Base
- 8.00
- 0.97
- 38.19
- AHS
- 14.50
- 1.55
- 36.83
- B. QHA timing evaluation 2 held-out tasks, 400 episodes per method
- 자료 없음
- 자료 없음
- 자료 없음
- AHS
- 33.77
- 3.29
- 68.60
- QHA-only
- 29.81
- 18.87
- 81.53
- AHS+QHA
- 31.43
- 19.88
- 83.32
Shared timing scope for \pi_{0.5}.
Table 12 reports policy-only inference time. Total inference time is unavailable for the 50-task study. QHA selection is outside the policy timer and is not timed separately. The AHS reference in Panel B has a total inference time of 3.49 s per episode. Both QHA variants make fewer calls but have higher policy-inference and wall times than this reference. These timings and the success rates in Tab. 2B come from different cohorts and cannot be combined to estimate success-normalized efficiency. These costs characterize the evaluated implementation, not an intrinsic QHA cost.
AAC-core sampling cost.
On RoboCasa GR1 Tabletop with Qwen3GR00T (24 tasks, 1,200 episodes), AAC-core averages 180.95 calls, 20.98 s of synchronized total inference, and 47.56 s of wall time per episode. Each call samples 20 action chunks, giving 4,342,740 chunks in total. No policy-only timing aggregate is available. Base/AHS references come from separate evaluations and are not paired with the AAC-core measurements. AAC-core, the 50-task study, the four-task selector comparison, and the QHA timing evaluation were run on an NVIDIA RTX 4090 GPU.
D.2 Latency and Asynchronous Execution
The policy prediction horizon is H=50. Fixed K=40 is a target execution budget. RTC can replace a chunk at the first legal waypoint boundary after readiness. The AHS candidates are \{10,20,30,40\}. This cohort uses no QHA and is separate from the K=50 candidate-scaling experiment.
Timing and aggregation.
For episode i, let a_{ij}=t^{\mathrm{start}}_{ij}-t^{\mathrm{obs}}_{ij} be the age of the observation used to generate executed action j. We compute \bar{a}_{i}=n_{i}^{-1}\sum_{j}a_{ij} and average \bar{a}_{i} equally over episodes within each task, then equally over the four tasks, separately for Easy and Hard. Failures are included.
Wait s/ep averages recorded hold time per episode. Wait % is 100\sum_{i}W_{i}/\sum_{i}T_{i}, where T_{i} includes both action execution and waiting. Calls/ep includes requests discarded at episode termination. Infer s/ep reports model computation time per episode, amortized across active batch requests and excluding selector computation.
- —
- +0 ms
- +100 ms
- +200 ms
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- All four tasks
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- Sync-Fixed
- 19.50
- 19.50
- 19.25
- 3.77
- 4.432
- 12.89
- 5.04
- 5.02
- Sync-AHS
- 26.25
- 24.25
- 23.75
- 6.80
- 7.893
- 22.95
- 8.59
- 3.00
- RTC-Fixed
- 23.00
- 21.50
- 22.00
- 0.31
- 0.344
- 13.32
- 6.47
- 4.92
- RTC-AHS
- 26.50
- 25.75
- 26.75
- 0.31
- 0.344
- 24.30
- 10.62
- 3.07
- Place A2B Left
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- Sync-Fixed
- 15.00
- 21.00
- 16.00
- 3.37
- 3.106
- 9.03
- 3.29
- 5.35
- Sync-AHS
- 20.00
- 16.00
- 20.00
- 5.92
- 5.790
- 16.83
- 6.21
- 3.30
- RTC-Fixed
- 26.00
- 22.00
- 22.00
- 0.41
- 0.344
- 9.53
- 4.09
- 5.15
- RTC-AHS
- 22.00
- 21.00
- 22.00
- 0.37
- 0.344
- 17.94
- 7.28
- 3.37
- Place Bread Basket
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- Sync-Fixed
- 26.00
- 25.00
- 25.00
- 3.20
- 5.246
- 15.25
- 6.12
- 5.77
- Sync-AHS
- 38.00
- 34.00
- 32.00
- 6.19
- 9.409
- 27.36
- 10.14
- 3.20
- RTC-Fixed
- 28.00
- 29.00
- 29.00
- 0.22
- 0.344
- 15.61
- 7.80
- 5.68
- RTC-AHS
- 37.00
- 31.00
- 37.00
- 0.24
- 0.344
- 28.68
- 12.60
- 3.38
- Place Bread Skillet
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- Sync-Fixed
- 13.00
- 9.00
- 12.00
- 5.15
- 4.097
- 11.91
- 4.51
- 3.89
- Sync-AHS
- 13.00
- 13.00
- 12.00
- 8.56
- 6.994
- 20.33
- 7.06
- 2.58
- RTC-Fixed
- 13.00
- 12.00
- 10.00
- 0.45
- 0.346
- 12.35
- 5.59
- 3.81
- RTC-AHS
- 15.00
- 16.00
- 13.00
- 0.45
- 0.345
- 21.71
- 8.43
- 2.54
- Place Can Basket
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- Sync-Fixed
- 24.00
- 23.00
- 24.00
- 3.90
- 5.280
- 15.35
- 6.27
- 5.08
- Sync-AHS
- 34.00
- 34.00
- 31.00
- 7.07
- 9.381
- 27.27
- 10.95
- 2.93
- RTC-Fixed
- 25.00
- 23.00
- 27.00
- 0.27
- 0.344
- 15.78
- 8.39
- 5.02
- RTC-AHS
- 32.00
- 35.00
- 35.00
- 0.27
- 0.344
- 28.86
- 14.19
- 3.01
Success within a time budget.
For T_{i}^{\mathrm{success}}, set the recorded completion time for a successful episode and +\infty for a failed episode. Then
\mathrm{SR}(t)=\frac{100}{4}\sum_{q=1}^{4}\frac{1}{100}\sum_{i\in q}\mathbf{1}\{T_{i}^{\mathrm{success}}\leq t\}.Here t denotes physical time in seconds. Computed separately for each setting, \mathrm{SR}(t) measures task completion under a time budget, with all attempted episodes retained in the denominator.
Easy and Hard outcomes.
Both settings use the same task checkpoints, execution horizon and AHS candidates; episodes are paired across methods and delays within each setting. At +200 ms, RTC reduces AHS waiting by 94.5% on Easy and 95.6% on Hard. Relative to RTC-Fixed, RTC-AHS achieves higher aggregate SR in all six setting/delay conditions, with observed gains of 3.50–8.25 percentage points. At +200 ms, the paired 95% interval for this gain is [-0.75,8.26] points on Easy and [0.75,8.75] on Hard. Relative to Sync-AHS, RTC-AHS changes final SR by -2.00 points on Easy and +3.00 points on Hard at +200 ms. Thus, reduced waiting does not imply uniformly higher success. Policy calls and model computation remain higher than for RTC-Fixed (Tabs. 13, 14, and 15).
- —
- +0 ms
- +100 ms
- +200 ms
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- All four tasks
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- Sync-Fixed
- 38.00
- 35.50
- 37.00
- 3.90
- 3.856
- 11.21
- 4.38
- 5.12
- Sync-AHS
- 46.00
- 46.25
- 44.75
- 6.63
- 6.310
- 18.35
- 6.90
- 3.16
- RTC-Fixed
- 38.75
- 40.25
- 39.00
- 0.37
- 0.344
- 11.50
- 5.57
- 5.00
- RTC-AHS
- 46.00
- 48.50
- 42.75
- 0.37
- 0.344
- 20.32
- 8.80
- 3.18
- Place A2B Left
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- Sync-Fixed
- 49.00
- 45.00
- 50.00
- 3.37
- 2.401
- 6.98
- 2.48
- 5.65
- Sync-AHS
- 54.00
- 58.00
- 56.00
- 5.86
- 4.107
- 11.94
- 4.25
- 3.41
- RTC-Fixed
- 50.00
- 51.00
- 50.00
- 0.51
- 0.344
- 7.55
- 3.36
- 5.44
- RTC-AHS
- 52.00
- 57.00
- 57.00
- 0.51
- 0.344
- 12.83
- 5.21
- 3.51
- Place Bread Basket
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- Sync-Fixed
- 40.00
- 35.00
- 35.00
- 3.41
- 4.548
- 13.22
- 5.29
- 5.51
- Sync-AHS
- 47.00
- 47.00
- 49.00
- 6.10
- 7.317
- 21.27
- 8.19
- 3.28
- RTC-Fixed
- 41.00
- 42.00
- 41.00
- 0.28
- 0.344
- 13.19
- 6.56
- 5.48
- RTC-AHS
- 46.00
- 48.00
- 41.00
- 0.29
- 0.344
- 24.49
- 10.41
- 3.33
- Place Bread Skillet
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- Sync-Fixed
- 26.00
- 26.00
- 27.00
- 4.98
- 3.643
- 10.59
- 4.02
- 4.25
- Sync-AHS
- 36.00
- 33.00
- 29.00
- 7.86
- 5.853
- 17.02
- 6.00
- 2.81
- RTC-Fixed
- 25.00
- 31.00
- 21.00
- 0.48
- 0.345
- 11.41
- 5.18
- 4.02
- RTC-AHS
- 39.00
- 39.00
- 34.00
- 0.51
- 0.345
- 17.39
- 7.00
- 2.90
- Place Can Basket
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- Sync-Fixed
- 37.00
- 36.00
- 36.00
- 4.12
- 4.832
- 14.05
- 5.75
- 5.09
- Sync-AHS
- 47.00
- 47.00
- 45.00
- 6.84
- 7.964
- 23.15
- 9.16
- 3.12
- RTC-Fixed
- 39.00
- 37.00
- 44.00
- 0.33
- 0.344
- 13.83
- 7.17
- 5.06
- RTC-AHS
- 47.00
- 50.00
- 39.00
- 0.30
- 0.344
- 26.57
- 12.60
- 3.00
\pi_{0.5} Easy. The three panels mirror Fig. 5: waiting versus mean observation age at +0/+100/+200 ms, final SR at each delay, and SR(t) at +200 ms. All 400 outcomes per condition are included. Bars and shading show marginal/pointwise 95% paired-seed bootstrap intervals within four tasks.- Easy
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 0
- Sync-Fixed
- 1.67
- 1.600
- 11.11
- 4.32
- 4.93
- —
- Sync-AHS
- 2.87
- 2.614
- 18.15
- 6.78
- 3.01
- —
- RTC-Fixed
- 0.16
- 0.147
- 11.20
- 5.38
- 5.00
- —
- RTC-AHS
- 0.16
- 0.147
- 17.86
- 7.81
- 3.24
- 100
- Sync-Fixed
- 2.82
- 2.755
- 11.29
- 4.37
- 4.98
- —
- Sync-AHS
- 4.80
- 4.384
- 17.97
- 6.60
- 3.08
- —
- RTC-Fixed
- 0.26
- 0.244
- 11.42
- 5.45
- 4.97
- —
- RTC-AHS
- 0.28
- 0.244
- 18.73
- 8.09
- 3.12
- Hard
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 자료 없음
- 0
- Sync-Fixed
- 1.60
- 1.853
- 12.87
- 5.14
- 4.84
- —
- Sync-AHS
- 2.93
- 3.244
- 22.53
- 8.37
- 2.85
- —
- RTC-Fixed
- 0.13
- 0.146
- 12.84
- 6.31
- 4.99
- —
- RTC-AHS
- 0.14
- 0.147
- 22.23
- 9.83
- 3.01
- 100
- Sync-Fixed
- 2.72
- 3.168
- 12.99
- 5.12
- 4.89
- —
- Sync-AHS
- 4.85
- 5.523
- 22.64
- 8.39
- 2.93
- —
- RTC-Fixed
- 0.22
- 0.244
- 13.34
- 6.64
- 4.84
- —
- RTC-AHS
- 0.22
- 0.244
- 24.05
- 10.49
- 2.98
Appendix E AHS Hyperparameter Sensitivity
We evaluate \pi_{0.5} + AHS on eight RoboTwin2.0 tasks under Easy and Hard settings, with \mathcal{K}=\{10,20,30,40,50\} and expected-round selection at T_{\mathrm{sel}}=0.8. The two studies below use different episode budgets and pairing protocols. Neither changes our default configuration.
Inter-chunk continuity window.
We vary only W_{h}\in\{30,60,90\}, sharing 16 episode identities per task and setting (256 per configuration). The default W_{h}=60 is evaluated afresh as the paired reference in Fig. 10.
\pi_{0.5} + AHS on eight RoboTwin2.0 tasks, Easy and Hard, with 256 matched episodes per configuration. Lines connect tested points. Paired difference intervals appear in the text.Compared with W_{h}=60, gains for 30 and 90 are +2.73 and +3.91 percentage points, with paired 95% confidence intervals [-2.34,\,7.81] and [-0.78,\,8.59]. We use 10,000 bootstrap draws of matched episodes within task–setting cells, weighted equally. Both intervals include zero, so neither a success-rate difference nor equivalence is established. We retain W_{h}=60.
Other evidence and posterior hyperparameters.
Each configuration in Tab. 16 uses 100 episodes per task and setting (1,600 total), unpaired across configurations. One parameter family changes at a time, with the two penalties varied jointly. Other settings follow Appendix A. These evaluations are separate from the continuity-window study. Overall SR spans 34.38–37.13%, characterizing sensitivity rather than validation-based selection or an established optimum.
- Forgetting \rho_{f}
- 0.99
- 0.95
- 46.88
- 26.75
- 36.81
- —
- 자료 없음
- 0.995
- 46.50
- 25.88
- 36.19
- Update strength \eta
- 1
- 0.5
- 45.38
- 24.38
- 34.88
- —
- 자료 없음
- 2
- 47.38
- 26.88
- 37.13
- Kernel bandwidth \sigma_{K}
- 10
- 5
- 45.50
- 26.13
- 35.81
- —
- 자료 없음
- 20
- 44.25
- 24.50
- 34.38
- Intra reference window W_{\tau}
- 3
- 1
- 47.38
- 24.75
- 36.06
- —
- 자료 없음
- 5
- 46.38
- 23.88
- 35.13
- Joint \alpha_{\mathrm{intra}}=\beta_{\mathrm{inter}}
- 4
- 2
- 46.63
- 23.63
- 35.13
- —
- 자료 없음
- 8
- 46.25
- 24.88
- 35.56
\pi_{0} Handover Block case study. The selected execution horizon changes with task phase in a successful RoboTwin2.0 rollout. K_{\mathrm{exec}} denotes the selected K_{t}.Appendix F More Case Studies
Handover Block.
Fig. 11 illustrates phase-dependent execution in a successful \pi_{0} rollout. AHS selects shorter horizons during grasping and transport phases that require closed-loop correction, and longer horizons once the motion stabilizes. This example illustrates the behavior of the online selector in Sec. 4.1. It is not a controlled comparison of posterior mechanisms.
Additional tasks and settings.
Fig. 12 shows four additional \pi_{0} RoboTwin2.0 rollouts across Easy and Hard settings. Fig. 13 adds \pi_{0.5}+AHS cases on four real-world tasks. These qualitative examples illustrate horizon changes across task phases, not additional scored trials.
\pi_{0} case studies on RoboTwin2.0. Four successful rollouts show phase-dependent AHS horizons across Place Bread Basket and Blocks Ranking RGB under Easy and Hard settings.\pi_{0.5}+AHS case studies. Highlighted observations are selected at large adjacent-replan changes in K_{t}. Curves show all recorded execution horizons without smoothing, through the end of each recording. Phase labels are manual visual annotations, not ground-truth boundaries or trial scores.Appendix G Discussion and Limitations
Additional background.
RT-1, RT-2, and PaLM-E connect large-scale learning with robot control (Brohan et al., 2023; Zitkovich et al., 2023; Driess et al., 2023). Open X-Embodiment, Octo, and OpenVLA extend shared data and generalist initialization (O’Neill et al., 2024; Ghosh et al., 2024; Kim et al., 2024), while RoboBrain 2.0 broadens embodied reasoning interfaces (Team et al., 2025). RDT-1B and DexVLA extend diffusion-based action generation (Liu et al., 2025a; Wen et al., 2025). Other approaches couple actions to world models (Bi et al., 2025) or adapt chunking through self-guidance (So et al., 2026). ChunkTrust adapts chunk execution using action-expert evidence.
Additional horizon and verification methods.
Adaptive execution can draw on additional samples, learned sensitivity, or observations acquired during rollout. A3 (Chen et al., 2026a) uses group-sampled consensus and conditional re-decoding to verify a contiguous execution prefix. SA (Park et al., 2026) forecasts the sensitivity of the action distribution to observation changes and allocates shorter horizons to more sensitive phases. EQRL (Wang et al., 2026a) jointly learns the latent input, denoising budget, and chunk length through reinforcement learning. These methods differ in the information and computation used to choose a horizon. AHS scores prefixes of one generated chunk using its existing generation trace and executed history, with no conditional re-decoding or joint optimization of the generator’s inference schedule. Verification methods instead use fresh observations to assess a running plan. FFDC (Wang et al., 2026d) compares imagined futures with reality for world-action models, while DREAM-Chunk (Chen et al., 2026b) uses a latent world model to match candidate chunks’ predicted futures to observed execution. SV-VLA (Wang et al., 2026e) compares planned actions with a lightweight closed-loop reference, and PATCH (Zhou et al., 2026) accumulates localized visual residuals along an action-conditioned execution corridor to trigger intervention. These observation-driven mechanisms address disturbances that become visible after a chunk has been selected. ChunkTrust’s evidence instead informs how much of the current prediction to execute before observing again, so it does not provide the same within-prefix monitoring capability.
Continuity and frequency-aware action generation.
Cross-chunk consistency can be improved by changing how actions are generated or corrected. ChunkFlow (Yang et al., 2026) trains with seam and derivative-continuity losses and blends overlapping predictions at execution. SEAM (Zhan et al., 2026) steers denoising toward the previous chunk’s unexecuted tail, while Legato (Liu et al., 2026) learns continuation through action-conditioned initialization and modified flow dynamics. REMAC (Wang et al., 2026c) uses masked action conditioning for real-time execution, and ACNet (Guo and Guo, 2026) conditions a lightweight delay-aware adapter on executed motion. A2C2 (Sendai et al., 2025) instead applies a learned per-step correction using the latest observation and the base policy’s action. Frequency-aware approaches act on the representation or training objective. FAFM (Guo et al., 2026) generates continuous action trajectories in a frequency-domain representation, while FocalPolicy (He et al., 2026) combines proximal time-domain supervision with multi-chunk spectral regularization. These works motivate attention to temporal coherence, but they do not make the same intervention as ChunkTrust. Our spectral signal measures variation during generation, and our continuity signal evaluates a prefix stitched to executed history. Both are used to select an execution length, leaving the generated action values and base-policy weights unchanged. A frequency-domain training loss is therefore distinct from the generation-time diagnostic used here.
Experience reuse and the role of chunking.
TraceFlow (Zhang et al., 2026a) reuses successful and failed rollouts through a retrieval bank and outcome-conditioned guidance of a frozen flow-matching action expert. ChunkTrust reuses information in a different form and for a different decision. AHS retains evidence-derived horizon preferences within an episode, while QHA learns a context-conditioned horizon prior offline. Neither component retrieves rollout trajectories to steer action generation, and QHA does not update its weights during deployment. Recent analyses also caution against treating long open-loop execution as universally beneficial. Lazzati et al. (2026) study non-Markovian expressivity and implicit ensembling as explanations for the benefits of chunking. Zeng et al. (2026) show that the value of open-loop execution depends on demonstration non-Markovianity and policy context length. ChunkTrust addresses horizon selection for existing chunk policies, not a claim that longer open-loop execution is intrinsically preferable to reactive control.
Limitations.
AHS requires access to action chunks and generation-time velocity traces. It scores a candidate grid, although expected-round selection can return intermediate integer lengths. QHA learns a dense prior, but AHS+QHA still computes online evidence, interpolated or scored directly on the dense grid. Horizon decisions are made at replanning time, so the selector cannot directly detect a new disturbance that arises during the chosen prefix. Our evaluations cover multiple policies and two simulation benchmarks, plus four real-world bimanual tasks, rather than all embodiments, safety-critical tasks, or long-horizon mobile manipulation.
Broader impact.
Adaptive horizons may reduce unnecessary replanning while preserving reactivity near contact. An incorrect horizon can still commit a robot to unsafe motion outside the tested distribution. Deployment should retain workspace and speed limits, emergency stops, human supervision during evaluation, and task-specific validation.
Third-party resources.
We use the cited RoboTwin2.0 and RoboCasa benchmarks, policy checkpoints, and associated software for research evaluation. These third-party resources remain subject to their original licenses, terms of use, and attribution requirements.