ROBOTNESS
전문가arXiv

실험실 자동화 분말 계량을 위한 Tool-Policy Co-Design 프레임워크

Nikola Radulov, Xin Yang, Kevin S. Luck, Gabriella Pizzuto
30초 요약

연구진은 로봇 팔이 쓰는 분말 분배 도구의 형태와 강화학습 제어 정책을 함께 최적화하는 Tool-Policy Co-Design 프레임워크를 제안했다. BOHB 기반 외부 루프와 SAC 기반 내부 루프로 도구 깊이, 폭, 림 스파이크 구조를 탐색하고, 유사한 기하 구조의 정책을 재사용해 탐색 속도를 높였다. 실물 로봇 실험에서 공동 설계한 도구는 표준 도구 대비 실제 계량 오차를 45% 줄였고, 이 논문은 동료 심사를 거치지 않은 프리프린트다.

연구 질문

로봇이 분말을 계량할 때 도구 형태와 제어 정책을 동시에 최적화하면, 사람 손 기준의 고정 도구와 정책만 최적화하는 기존 방식보다 다양한 분말 유동성에서 계량 오차를 줄일 수 있는가?

문제

자율 분말 계량은 이질적인 입자 재료의 비선형 동역학 때문에 실험실 자동화의 병목이다. 기존 로봇 화학자는 사람 손의 민첩성을 기준으로 설계된 스푼이나 주걱을 그대로 쓰는데, 도구 기하가 행동 대비 질량 이동 함수를 결정해 제어 정책이 달성할 수 있는 정밀도가 고정된 형태에 제한된다.

기존 접근

최근 연구는 고정된 도구를 가정하고 심층 강화학습으로 제어 정책만 적응시키거나, SCU-Hand 같은 특수 엔드이펙터를 수작업으로 설계했다. 로봇 형태와 정책 공동 설계도 일부 시도됐지만, 도구 재료 조작에서는 정책을 고정하거나 초기 물체 자세만 바꾸는 등 도구와 정책을 함께 최적화하지 않았다.

새 접근

이 연구는 도구 형태와 정책을 이중 수준 최적화로 묶었다. 외부 루프는 BOHB가 도구 깊이, 폭, 14개 림 스파이크 높이의 이산 구성을 탐색하고, 내부 루프는 후보 도구마다 SAC 정책을 훈련한다. 기하 유사성 지표 S를 도입해 부피, 폭 대비 깊이 비율, 스파이크 둘레 위상을 비교하고, 유사한 기존 설계의 캐시 정책으로 웜스타트해 같은 계산 예산에서 약 28% 더 많은 형태를 평가한다. 보상은 단계별 오차 변화량으로 바꿔 최종 정밀도와 정렬하고, 유동성 범위를 균일 샘플링해 BOHB의 successive halving과 충돌하지 않게 했다.

결과

시뮬레이션 그리드 서치에서는 수동 설계 스파이크 구성이 저충실도와 고충실도 환경에서 각각 1.27 mg, 1.54 mg의 평균 최종 오차로 가장 낮았다. 실물 Franka Research 3 로봇에서 15 mg 목표, 재료당 10회 평가한 결과 BOHB가 찾은 Config b는 분포 내 재료 평균 오차 0.91 mg, 유사도 가속 BOHB가 찾은 Config B는 0.81 mg을 기록해 표준 도구의 1.93 mg보다 낮았다. 유사도 웜스타트는 같은 200,000 에피소드 예산에서 106개 대신 136개 고유 형태를 탐색했고, 시뮬레이션 대비 실물 오차 격차는 분포 내 재료 기준 고충실도 환경에서 평균 1.78 mg였다. 분포 밖 밀가루, 펙틴, 중탄산나트륨을 포함한 전체 실제 오차는 Config B가 2.23 mg로 표준 도구 4.09 mg 대비 약 45% 낮았다.

한계

저자들은 응집성 분말의 동역학을 시뮬레이션으로 정확히 재현하지 못한다고 밝혔다. 실제 평가는 단일 Franka Research 3 로봇과 3D 프린팅 도구, 재료당 10회 실행에 그쳤고, 초기 분말을 스쿠핑해 담는 단계는 제외하고 수동으로 파라미터를 맞췄다. 최적화에는 200,000 에피소드가 약 7.85일의 병렬 계산을 요구했으며, 시뮬레이션 최고 순위와 실물 최고 순위가 뒤바뀌는 경우도 있었다. OOD 응집성 재료로 갈수록 sim-to-real 오차가 커진다.

업계 영향

자율 실험실과 로봇 화학 플랫폼에서 분말 시료 준비를 자동화하는 기업이 이 방법을 적용할 수 있다. 특히 제약, 화학, 신소재 연구실의 밀리그램급 계량 공정에 이식 가능성이 있지만, 실제 도입은 1~3년 뒤로 보인다. 그 전에 응집성 분말을 다루는 고충실도 시뮬레이션, 스쿠핑부터 계량까지 전체 공정 자동화, 다양한 로봇과 저울 인터페이스 검증이 필요하다.

논문 전문

Tool-Policy Co-Design for Powder Weighing in Laboratory Automation

Nikola Radulov, Xin Yang, Kevin S. Luck, Gabriella Pizzuto

CC BY 4.0 라이선스로 공개된 논문입니다. 출처를 밝혀 전재하며, 원문은 arXiv:2609.39797(PDF)에서 볼 수 있습니다.

Abstract

Autonomous powder weighing is one of many bottlenecks in laboratory automation due to the complex, non-linear dynamics of heterogeneous materials. Robot chemists performing this task utilise standard tools shaped for the dexterity of human hands, whose fixed geometry sets the dynamics that the control policy needs to regulate. This work introduces a tool-policy co-design framework that concurrently optimises the morphology of a dispensing tool and its control policy for use by robots in chemistry laboratories, formulated as a bi-level optimisation that minimises dispensing error over a target distribution of powder flowabilities. The outer loop varies tool-design parameters such as tool depth, width and rim spike topology using Bayesian optimisation and hyperband, while an inner loop optimises a control policy for each candidate morphology. We also introduce a geometric similarity metric that warm-starts policy training from cached policies of structurally similar designs, exploring 28\% more configurations under the same compute budget. The proposed framework is evaluated on a robotic powder weighing task across seven materials with distinct physical dynamics in a flowability-informed robot-material simulation framework. Experimental results demonstrate that our co-designed tool morphology reduces real-world weighing errors by 45\% relative to a standard tool, including on previously unseen materials. These results demonstrate our method can adapt both the control policy and the physical tool to the dynamics of the target material, bringing a new paradigm for material manipulation to the field of laboratory automation.

I Introduction

Accelerating the discovery of materials is crucial for addressing global challenges and self-driving laboratories (SDLs) pursue this by integrating automated hardware with data-driven decision making [1]. To date, laboratory robots have been deployed predominantly for sample transportation, where the manipulated object is rigid. Sample preparation remains a fundamental challenge, as materials deform and flow under contact and their response is governed by bulk properties that differ across materials. Robot-material manipulation underpins solid-state chemistry [1], yet robots inherit tools such as spatulas [2, 3] shaped for use by the dexterous human hand rather than for a manipulator or the material. Recent approaches use deep reinforcement learning (RL) to adapt robot manipulators to these dynamics [2, 3], which optimise a control policy around a fixed tool. Tool geometry defines the action-to-mass transfer function, bounding the task precision achievable by any control policy using that tool. While co-design of morphology and control is increasingly explored, it remains largely unaddressed for tool-material manipulation which is fundamental to the success of robot-driven chemistry lab automation.

Fig. 1: Tool-policy co-design method for powder weighing. We use BOHB to propose tool morphologies \xi. For each candidate \xi, a control policy \pi_{\xi} is trained in a low-fidelity environment and evaluated in a high-fidelity digital twin. The highest-performing design and paired policy are transferred for real-world validation.

To address this, we present a co-design framework that concurrently optimises both the morphology of a dispensing tool and the robot’s reinforcement learning control policy, formulated as a bi-level optimisation problem that minimises dispensing error across a target distribution of powder flowabilities (Fig. 1). In an outer loop, Bayesian optimisation and hyperband (BOHB) explores a parameterised, low-dimensional morphological search space defining the depth, width, and discrete rim spike topology of an ellipsoidal tool. We select BOHB over standard Bayesian optimisation because successive halving evaluates candidates at low training budgets and promotes only top performers, saving compute on sub-optimal morphologies. In the inner loop, an RL control policy is trained for each candidate tool. To make the search tractable, we introduce a geometric similarity metric that warm-starts policy training from cached policies of structurally similar designs. We validate the framework in simulation and by zero-shot transfer to a physical robot.

In summary, the main contributions of this work are:

  • A bi-level co-design framework for milligram-scale powder weighing that jointly optimises a reinforcement learning control policy and the tool morphology across a distribution of material flowabilities, over a parameterised design space derived from standard tool geometries that combines continuous bowl scaling with discrete rim spike configurations.
  • A morphological similarity metric that accelerates the search by warm-starting policy training from structurally related candidates, without compromising the real-world performance of the discovered tool.
  • An empirical evaluation in simulation and on a real robotic manipulator, demonstrating that the co-designed tools outperform both a standard tool and grid-search baselines across a diverse range of powder behaviours.

II Related Work

II-A Robot Skill Learning for Laboratory Automation

Robotic scientists have transitioned from open-loop automation to adaptive modular systems, motivated by the need for generalisable approaches capable of handling unpredictable, heterogeneous samples [1]. While self-driving labs have advanced high-level planning and experimental decision making [1], low-level manipulation remains a fundamental bottleneck. Sample grinding [4], powder weighing [2] and scooping [5] are challenging due to the non-linear, unpredictable dynamics of materials. To address this, recent works leverage deep reinforcement learning through sim-to-real domain randomisation [2] or embed material properties directly into the learning process [3]. Prior work optimises a control policy around a fixed, human-centric tool, so task success is bounded by a geometry that was not designed for robot-material manipulation. We instead formulate the tool as a design variable of the robot skill itself.

II-B Robotic Tool Co-Design

Co-design jointly optimises a robot’s morphology and its control policy, on the premise that a fixed embodiment bounds control performance. Recent works search a parametrised design space, either through bi-level formulations pairing Bayesian optimisation with reinforcement learning, e.g. to design mobile manipulator mountings [6], or through generative models and cross-embodiment policies combined with grammar-based search to scale the synthesis of tool shapes [7]. Closest to our work, Li et al. [8] optimise tool geometry for contact-rich tasks over a distribution of task variations using differentiable simulation. However, they vary the initial object pose instead of its dynamics and hold the policy fixed, leaving joint optimisation under such variation unaddressed. Within laboratory automation, specialised end-effectors have been engineered e.g., the SCU-Hand uses a soft, reconfigurable conical sheet for adaptive scooping [9], later extended with a single-sheet valve for milligram-scale dispensing [10]. However, the morphologies of these tools are manually handcrafted. Our approach concurrently optimises the control policy alongside the tool’s physical design, tailoring the tool to the flow properties of the material.

III Methodology

Our method formulates robotic powder weighing as a joint optimisation problem with bi-level structure over tool morphology and control policy. The goal is to identify a tool design that minimises the expected dispensing error over a distribution of materials with diverse physical dynamics.

III-A Problem Formulation

We consider the task of autonomous powder weighing with a robotic manipulator that holds a custom tool at its end-effector and dispenses powder into a container on an analytical balance (Fig. 1). The balance provides weight feedback at each control step and an episode terminates after a fixed number T of steps. Following Radulov et al.’s method [3], we characterise a material by its static flowability. This is quantified by the angle of repose (AoR), where a larger AoR corresponds to a lower flowability. The flowability F of a material is sampled from a representative range F_{r}=[AoR_{min},AoR_{max}] across the task distribution.

Let the morphology of the tool be described by a parameter vector \xi\in\Xi, where \Xi denotes the design space defined in Section III-B. Let \pi_{\xi} be the robot control policy trained with that tool. As the geometry of the tool determines how much powder is displaced per action, \xi and \pi_{\xi} cannot be chosen independently, i.e., a morphology is only as good as the policy that can be learnt for it and vice versa. Furthermore, powder dispensing is difficult to mathematically model as the dynamics vary strongly with material properties and with the amount of powder remaining in the tool [2, 3].

To find the optimal combination (\xi^{*},\pi^{*}_{\xi^{*}}) we define the bi-level optimisation problem as:

\displaystyle\xi^{*}=\arg\min_{\xi}\mathbb{E}_{F_{r}}[\mathcal{L}(\xi,\pi^{*}_{\xi},F_{r})], \\ \displaystyle\text{s.t.~}\pi^{*}_{\xi}=\arg\max_{\pi}\mathbb{E}_{\begin{subarray}{c}\pi(a_{t}|s_{t})\\
F\sim F_{r}\\
p(s_{t+1}|s_{t},a_{t},\xi,F)\end{subarray}}\left[G\right]
(1) (2)

Here, we optimise the morphology \xi and policy \pi_{\xi} using a morphology objective \mathcal{L} and a policy learning objective G. The complete bi-level co-design framework is illustrated in Fig. 1.

III-B Outer Loop: Tool Morphology Parametrisation

We base the design space on a standard spoon-like geometry, where the bowl is modelled as an ellipsoid of height h, width w, and length l, sliced at the equator. Three groups of parameters, illustrated in Fig. 2, form \xi.

Fig. 2: Tool morphology parameters of the design space; l, h, w represent the length, depth and width of the spoon bowl. H_{k} represents the height of the k^{th} spike.
  1. Tool depth (h) governs the vertical extent of the spoon bowl. A shallower depth yields a flatter geometry, requiring less aggressive agitation and smaller pitch adjustments, which can be ideal for cohesive materials.
  2. Tool width (w) controls the lateral scale of the bowl. Increased widths broaden the surface area, enabling the capture of larger volumes per scoop.
  3. Tool rim spike heights (H_{k}) are used for finer control over material outflow. The outer optimisation loop can independently vary the height of each spike. As the spikes are modelled along the edge of the bowl, their spatial distribution is inherently coupled with the tool depth.

The bowl length l is constant, which keeps the tool centre point identical for all candidates. Should l be allowed to vary, the parameters of the low-level Cartesian position controller executing the shake and incline primitives would have to be retuned per candidate, in simulation and on the robot, to avoid collisions with the vial and analytical balance. Since the bowl volume varies across \Xi, the granular simulation is configured so that its fidelity is equivalent for every candidate.

III-C Outer Loop: Tool Design Generation

We search the morphological space \Xi with BOHB [11]. This method proposes candidates from a surrogate model fitted to past evaluations and allocates a budget b, defined as the number of training episodes granted to the inner loop. Budgets are assigned by successive halving, where a large pool of morphologies is trained on a small budget, the worst performers are discarded and the survivors are promoted to a longer budget. At each outer-loop iteration, BOHB samples a candidate tool design \xi\in\Xi. This configuration is passed to the robotic simulator, which applies the required scaling to instantiate the custom tool geometry before initiating the inner loop to train a corresponding control policy for b episodes. Due to the stochasticity inherent in policy optimisation, we evaluate n random seeds per configuration. We define the outer-loop objective function \mathcal{L}(\xi,\pi^{*}_{\xi},F_{r}) as the expected weighing error across a flowability range F_{r}.

\mathcal{L}(\xi,\pi^{*}_{\xi},F_{r})=\mathbb{E}_{F\in F_{r}}\left|w_{\mathrm{target}}-w_{\pi^{*}_{\xi},T}^{F}\right|
(3)

where w_{\mathrm{target}} represents the target weight to be dispensed, and w_{\pi^{*}_{\xi},T}^{F} represents the final dispensed weight for material F, using the candidate tool design \xi and its associated control policy \pi^{*}_{\xi}. In practice, to increase stability we sample over a distribution/set of optimised policies (i.e., multiple seeds), i.e., \mathcal{L}_{\text{avg}}=\frac{1}{n}\sum_{\pi^{*}_{\xi}\in\Pi^{*}_{\xi}}\mathcal{L}(\xi,\pi^{*}_{\xi},F_{r}).

III-D Inner Loop: Policy Optimisation

Given a morphological design \xi, the aim of the inner loop is to optimise a corresponding control policy \pi_{\xi} and provide a reliable estimate of its performance over the target flowability range F_{r}. We formulate the powder weighing task as a Markov decision process (MDP) [12] \mathcal{M}=\langle\mathcal{S},\mathcal{A},\mathcal{P},\mathcal{R},\gamma\rangle. We employ the same parameterised action space as prior works [2, 3], where at each step, the agent adjusts the tool pitch by a_{\mathrm{incline}}, and subsequently performs a shaking motion by retracting the tool a distance a_{\mathrm{shake}} before returning it to its home position. The observation o\in\mathcal{S} follows prior work: o=(w_{current},w_{target},\theta_{spoon}). We introduce two enhancements from Radulov et al.’s work [3].

III-D1 Reward

We redesign the reward to penalise the step-wise change in error rather than the absolute error at the current step:

\mathcal{R}_{t}=\frac{\Delta_{t-1}-\Delta_{t}}{w_{\mathrm{target}}}
(4)

where \Delta_{t}=|w_{\mathrm{target}}-w_{t}| is the absolute weight error at step t, with the initial condition \Delta_{0}=w_{\mathrm{target}}. This formulation telescopes such that the cumulative return over an episode becomes G=\sum_{t=1}^{T}\mathcal{R}_{t}=1-\frac{\Delta_{T}}{w_{\mathrm{target}}}, which depends only on the final dispensing error. Under the original step-wise absolute error reward, the inner loop implicitly optimised for speed as well as accuracy, encouraging the agent to close the mass gap rapidly and sometimes overshoot. Since our outer loop scores a morphology strictly on final precision \mathcal{L} (Equation 3), this reward corrects the mismatch between the inner loop’s objective and the outer loop’s cost function.

III-D2 Sampling of the Flowability Range

We replace the curriculum learning strategy with uniform random sampling of the flowability range F_{r}. A curriculum makes the training distribution a function of the budget b which conflicts with the successive halving of BOHB in our outer loop. At low budgets the agent would only be exposed to the easy end of F_{r}, so Equation (3) would rank morphologies on a narrow subset of materials, shifting the ranking as candidates are promoted. Uniform sampling makes the score at every budget an unbiased estimate of the same quantity, so that evaluations obtained at different fidelities remain comparable.

III-E Similarity-based Evaluation

Evaluating the outer loop objective function \mathcal{L}(\xi,\pi^{*}_{\xi},F_{r}) presents a computational bottleneck as it requires training the control policy \pi_{\xi} from scratch for every proposed tool morphology \xi. However, during the BOHB optimisation process, the algorithm frequently samples configurations that represent minor parametric perturbations of previously evaluated designs. As the physical dynamics of powder flow remain largely consistent across minor morphological adjustments, we leverage the fully trained policy of a structurally similar design to warm-start the training of the new control policy. To facilitate this accelerated search, we introduce a morphological similarity metric S(\xi_{1},\xi_{2})\in[0,1] that quantifies the resemblance between two geometries based on their total volume, geometry and spike topology.

For both the total volume V and the width-to-depth ratio R, we employ a min–max ratio to ensure scale-invariant similarity bounds within [0,1]. The total volume V of a given morphology is computed as the sum of the ellipsoidal half-bowl volume and the volume of the watertight elliptical cylinder formed by the rim spikes. The height of this cylinder is bounded by the shortest spike. The volume and ratio similarity scores are computed as:

S_{V}=\frac{\min(V_{1},V_{2})}{\max(V_{1},V_{2})},\quad S_{R}=\frac{\min(R_{1},R_{2})}{\max(R_{1},R_{2})}
(5)

With spike height already captured volumetrically, the final term measures perimeter topology, where we weight the squared differences between the first-order discrete derivatives of adjacent spikes by their arc ratio c_{i} and map the result into [0,1] with a negative exponential function:

S_{\mathrm{spikes}}=\exp\left(-\sum_{i=1}^{M}c_{i}(\delta_{1,i}-\delta_{2,i})^{2}\right)
(6)

The overall similarity score S(\xi_{1},\xi_{2}) is defined as a weighted linear combination of S_{V}, S_{R}, and S_{\mathrm{spikes}}:

S(\xi_{1},\xi_{2})=\alpha S_{V}+\beta S_{R}+\gamma S_{\mathrm{spikes}}
(7)

Here, \alpha+\beta+\gamma=1.

We incorporate the similarity metric into our optimisation method as detailed in Algorithm 1, where the similarity logic is implemented in lines 10 to 16. Throughout the optimisation, the system maintains an observation history \mathcal{H} to train the Bayesian optimisation surrogate model and a policy registry \mathcal{P} to cache learned control policies. The algorithm first calculates the maximum bracket index s_{\max}, which dictates the maximum number of successive halving stages. BOHB iterates through several brackets, starting from the most aggressive one with s_{\max} successive halvings. At the start of a stage, a set of N initial morphologies \mathcal{C} is sampled from a Bayesian surrogate model \mu(\xi|\mathcal{H}) conditioned on prior history. Within each successive halving stage k, the target budget b is computed. Before training a policy \pi_{\xi_{i}} for a morphology \xi_{i}, the algorithm computes the similarity score S(\xi_{i},\xi_{j}) (line 10) against all previously trained morphologies \xi_{j}\in\mathcal{H}. It filters for a candidate set \mathcal{S} of morphologies that meet a minimum similarity threshold \sigma and precisely match the target budget b. If a valid match is found, the policy \pi^{*}\in\mathcal{P} of the most similar morphology \xi^{*} is used to warm-start \pi_{\xi_{i}} and the execution budget b is reduced to \phi. Conversely, if \mathcal{S} is empty, the policy undergoes standard training for the full budget b. After training, the policy is evaluated to determine its task error e_{i} and both the registry \mathcal{P} and history \mathcal{H} are updated. Finally, the configuration pool \mathcal{C} is pruned, retaining only the top N_{k} performers to advance to the next iteration.

Algorithm 1 Tool-Policy Optimisation for Powder Weighing

IV Experimental Evaluation

In this section, we evaluated the proposed co-design framework in simulation and on a physical robot, with experiments designed to address the following research questions: (1) Can simulation predict real-world task performance and the relative ranking of tool morphologies? (2) Can BOHB co-design discover morphologies that outperform a standard (commercial) tool and grid-search baselines in the real world? (3) Does the morphological similarity threshold accelerate the search without compromising the tool performance?

IV-A Experimental Setup

IV-A1 Simulation Setup

Granular materials were simulated using the position-based dynamics (PBD) solver of NVIDIA Isaac Sim [13], which resolves constraints such as contacts at the position level prioritising computational efficiency over physical fidelity and making it suitable for reinforcement learning. We constructed two environments (Fig. 1) to balance accuracy against cost. The high-fidelity environment is a digital twin of our real-world setup and runs at a physics time step of dt=1/240 \text{\,}\mathrm{s}. The low-fidelity environment removes the analytical balance and the target receptacle, reducing collision calculations per step and allowing a coarser time step of dt=1/100\text{\,}\mathrm{s}. In both environments, a 3D model of the standard tool is attached to the robot’s end-effector via a fixed joint matching the grasp pose of the real-world system. Episodes are initialised with the powder already contained within the tool, bypassing the computationally expensive tool-picking and scooping phases. To simulate how a tool’s physical dimensions dictate its retained powder volume, we scale the baseline initial particle count bounds (N_{\text{low}},N_{\text{high}}) proportionally with the volumetric capacity of the tool morphology \xi. Specifically, the initial particle count scales with the tool’s width and its effective depth, accounting for both base bowl and the added volume from the spikes. Tool morphologies are instantiated by scaling the width and depth of the standard model. The M=14 rim spikes are modelled as independent rigid bodies at fixed positions relative to the tool centre. As dispensing is primarily directed through the tool tip, eight smaller spikes each spanning c_{i}=1/32 of the ellipsoid arc are positioned there, while the remaining six span c_{i}=1/8. Spike height is set by vertical scaling to one of four discrete levels [0,0.33,0.66,1]; if the morphological configuration \xi sets a spike’s height to zero, its corresponding rigid body is removed from the simulation. The depth and width scale factors are continuous over [0.7,1.5] relative to the standard tool, so \Xi combines two continuous scale factors with 4^{14} discrete spike configurations. Episode length is fixed to 10 steps. All experiments are distributed across four Ubuntu 22.04 workstations with NVIDIA RTX 4090/5090 GPUs; on the RTX 5090 machine, an episode averages 13.57\text{\,}\mathrm{s} in the low-fidelity environment and 24.95\text{\,}\mathrm{s} in the high-fidelity environment.

IV-A2 Materials

We defined the target material distribution over an AoR range F_{r}=[28^{\circ},41^{\circ}], as in prior works [3]. For our experiments, we sample materials at four fixed flowability points: 28^{\circ}, 32^{\circ}, 36^{\circ}, and 41^{\circ}. Highly cohesive materials (e.g., flour) are excluded from the range of simulated materials, as capturing their clumping dynamics would require alternative techniques that are prohibitively expensive for our framework.

IV-A3 Real World Setup

We use a Franka Research 3 (FR3) [14] manipulator equipped with a Robotiq 85F gripper, mounted parallel to the working surface for precise tool manoeuvring. A Sartorius Entris II precision analytical balance measures the dispensed mass, feeding the data directly to the control policy in real time. The tool is loaded using a parabolic scooping trajectory [5], without the vision-based volume estimation, as calibrating it across all tool–material combinations is prohibitively expensive. Instead, we manually tune the scooping parameters so the initial acquired mass is within the simulated training distribution. To transfer a morphology from simulation, we fabricated the custom tool using a Prusa XL 3D printer equipped with a 0.4\text{\,}\mathrm{m}\mathrm{m} nozzle, which yields a vertical print resolution of 0.1\text{\,}\mathrm{m}\mathrm{m} and a horizontal resolution of 0.4\text{\,}\mathrm{m}\mathrm{m}. Consequently, the morphological search space \Xi is discretised by these manufacturing constraints: the horizontal resolution sets the lower bound of 0.7 on the scale factors, below which the rim spikes fall under the minimum reproducible feature size, while the vertical resolution sets the spacing of the four spike height levels.

IV-B Sim-to-Real Transferability and Simulation Fidelity

To evaluate the capability of our simulation to predict real-world task performance, we perform a grid search across our morphological design space \Xi. We apply a step size of 0.4 for both the depth and width scaling factors, which evaluates the extremities and the midpoint of these parameter ranges. We constrain the grid search to three predefined spike configurations: (1) all spikes set to maximum height; (2) all spikes disabled (height set to zero); and (3) a manually-designed configuration where the six front-most spikes are disabled while the remaining eight are set to their maximum height (represented by the array \{0,0,0,1,1,1,1,1,1,1,1,0,0,0\}). To train the control policy \pi_{\xi}, we use Soft Actor-Critic (SAC) [15], a model-free, off-policy deep reinforcement learning algorithm. We adopt the neural network architecture and physics-informed modelling from prior work [3], which includes seven optimised data points per flowability level. Training terminated after 3000 episodes with n=3 random seeds per configuration. This is reduced from the 4000 episodes in prior works, as standard tool policies converge well before the limit and the tighter budget also favours morphologies that remain sample efficient.

For each configuration, we measured the empirical weighing error (Equation 3) across the sampled materials for a fixed target mass w_{\mathrm{target}}=$15\text{\,}\mathrm{m}\mathrm{g}$, reporting the mean and standard deviation across seeds. While policy training occurred exclusively within the computationally-efficient low-fidelity environment, the final weight error was evaluated in both environments (Fig. 3). In both environments, tools with all spikes enabled performed worst, yielding a mean error of 2.81\text{\,}\mathrm{m}\mathrm{g} in the low-fidelity environment and 4.69\text{\,}\mathrm{m}\mathrm{g} in the high-fidelity environment. The manually-designed spike configuration achieves the lowest error in both cases: 1.27\text{\,}\mathrm{m}\mathrm{g} and 1.54\text{\,}\mathrm{m}\mathrm{g} respectively. In general, configurations with increased internal volume, i.e., those in which both depth and width are scaled up, performed worse than their smaller counterparts. This is likely because the larger retained powder mass makes it harder for the policy to exert fine-grained control over the dispensing flow.

Fig. 3: The performance of each chosen morphology as the mean final test error and the respective standard deviation, averaged across 3 seeds. The temperature represents the mean final test error (\text{\,}\mathrm{m}\mathrm{g}). The top graph (3(a)) shows the raw task error when evaluated on the low-fidelity environment and the bottom graph (3(b)) shows the performance when evaluated on the high-fidelity test environment.

To evaluate the simulation’s predictive capability for tool optimisation, we selected nine morphologies for zero-shot transfer to the real setup, i.e., three from each spike configuration, encompassing the best- and worst-performing designs for the search. We evaluated the configurations on in-distribution (IID) materials and out-of-distribution (OOD) materials. We compare them against the standard tool, trained under identical conditions to those described in Section III-D. The results are presented in Table I. The most effective geometry is Config. 2, which performs best on both the IID set and the full material set. Both Config. 2 and Config. 4 achieve IID errors lower than those of the standard tool. Config. 2 yields the lowest error for IID powders, apart from salt, where Config. 4 performs best.

The sim-to-real prediction gap, defined as the absolute difference between simulated and physical error, averages 1.78\pm 1.94\text{\,}\mathrm{m}\mathrm{g} for the high-fidelity environment and 2.63\pm 2.83\text{\,}\mathrm{m}\mathrm{g} for the low-fidelity environment on IID materials. Over the entire test set (including OOD powders), this prediction gap widens to 4.18\pm 1.78 mg and 5.10\pm 2.34 mg respectively. This degradation is expected, as we explicitly exclude highly cohesive materials from simulation due to the prohibitive computational cost of modelling their complex dynamics, so the tool morphology cannot be optimised for clumping or compressible behaviour.

Fig. 4: Sim-to-real rank shifts across morphologies. Each spoon image shows its baseline simulation rank; blue arrows show shifts for IID powders.

We ranked configurations by high-fidelity simulation performance using an anchor-based scheme: configurations are sorted by mean error, and each shares the current anchor’s rank unless a one-tailed Student’s t-test at 95% confidence finds it significantly worse, where it becomes a new anchor. With n=3 seeds, configurations differing only slightly tend to share a rank. Fig. 4 illustrates the relative ranking shifts from simulation to the real world for IID materials. Only three of the tools change their standing: one simulation Rank 1 configuration drops to Rank 2, and two Rank 3 configurations are upgraded to Rank 2. The best and worst configurations by mean error retain their respective ranks across the sim-to-real gap, which supports the use of the simulation as a proxy for morphological optimisation.

TABLE I: Zero-shot transfer results on w_{target}=15\text{\,}\mathrm{m}\mathrm{g} of selected policy-tool pairs, evaluated using the best-performing policy in simulation for each morphology. For each powder, results report the mean absolute error and standard deviation over 10 real-world runs. Sodium bicarbonate, pectin and flour represent OOD materials, i.e., the policy was not trained on these specific powders or materials with similar flow. Data points marked with * represent cases where no material was dispensed.
  • Standard Tool
    \Block2-4Spoon Configuration Depth
    자료 없음
    \Block2-4Spoon Configuration Width
    자료 없음
    \Block2-4Spoon Configuration Spike Configuration
    자료 없음
    Powder Weighing Error (mg) \Block2-1Sand
    1.49\pm 1.36
    Powder Weighing Error (mg) \Block2-1Sugar
    2.05\pm 1.73
    Powder Weighing Error (mg) \Block2-1Salt
    1.66\pm 1.19
    Powder Weighing Error (mg) \Block2-1Semolina
    2.52\pm 3.66
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    4.45\pm 5.57
    Powder Weighing Error (mg) \Block2-1Pectin
    \mathbf{3.58\pm 3.64}
    Powder Weighing Error (mg) \Block2-1Flour
    12.91\pm 13.58
    \Block3-1In-distribution Error
    1.93\pm 2.16
    \Block3-1Overall Error
    4.09\pm 6.82
  • 1
    \Block2-4Spoon Configuration Depth
    1.1
    \Block2-4Spoon Configuration Width
    1.1
    \Block2-4Spoon Configuration Spike Configuration
    {1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1}
    Powder Weighing Error (mg) \Block2-1Sand
    2.12\pm 4.21
    Powder Weighing Error (mg) \Block2-1Sugar
    4.78\pm 5.02
    Powder Weighing Error (mg) \Block2-1Salt
    12.56\pm 2.98
    Powder Weighing Error (mg) \Block2-1Semolina
    *
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    14.48\pm 0.23
    Powder Weighing Error (mg) \Block2-1Pectin
    13.15\pm 3.96
    Powder Weighing Error (mg) \Block2-1Flour
    13.29\pm 4.06
    \Block3-1In-distribution Error
    8.58\pm 6.36
    \Block3-1Overall Error
    10.74\pm 5.79
  • 2
    \Block2-4Spoon Configuration Depth
    1.1
    \Block2-4Spoon Configuration Width
    0.7
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    \mathbf{0.67\pm 0.63}
    Powder Weighing Error (mg) \Block2-1Sugar
    \mathbf{0.53\pm 0.53}
    Powder Weighing Error (mg) \Block2-1Salt
    1.01\pm 0.54
    Powder Weighing Error (mg) \Block2-1Semolina
    \mathbf{2.47\pm 1.41}
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    3.62\pm 3.45
    Powder Weighing Error (mg) \Block2-1Pectin
    8.21\pm 4.04
    Powder Weighing Error (mg) \Block2-1Flour
    11.41\pm 8.72
    \Block3-1In-distribution Error
    \mathbf{1.17\pm 1.13}
    \Block3-1Overall Error
    \mathbf{4.04\pm 5.43}
  • 3
    \Block2-4Spoon Configuration Depth
    1.5
    \Block2-4Spoon Configuration Width
    1.5
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    4.38\pm 2.27
    Powder Weighing Error (mg) \Block2-1Sugar
    2.29\pm 1.27
    Powder Weighing Error (mg) \Block2-1Salt
    1.66\pm 1.49
    Powder Weighing Error (mg) \Block2-1Semolina
    14.48\pm 0.99
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    14.49\pm 0.97
    Powder Weighing Error (mg) \Block2-1Pectin
    *
    Powder Weighing Error (mg) \Block2-1Flour
    14.66\pm 0.34
    \Block3-1In-distribution Error
    5.70\pm 5.44
    \Block3-1Overall Error
    9.52\pm 6.06
  • 4
    \Block2-4Spoon Configuration Depth
    0.7
    \Block2-4Spoon Configuration Width
    1.1
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    0.81\pm 0.38
    Powder Weighing Error (mg) \Block2-1Sugar
    0.76\pm 0.46
    Powder Weighing Error (mg) \Block2-1Salt
    \mathbf{0.56\pm 0.40}
    Powder Weighing Error (mg) \Block2-1Semolina
    2.76\pm 2.57
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    5.03\pm 4.08
    Powder Weighing Error (mg) \Block2-1Pectin
    7.04\pm 4.18
    Powder Weighing Error (mg) \Block2-1Flour
    13.70\pm 1.4
    \Block3-1In-distribution Error
    1.22\pm 1.56
    \Block3-1Overall Error
    4.38\pm 5.05
  • 5
    \Block2-4Spoon Configuration Depth
    1.5
    \Block2-4Spoon Configuration Width
    1.1
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    0.90\pm 0.53
    Powder Weighing Error (mg) \Block2-1Sugar
    1.18\pm 0.46
    Powder Weighing Error (mg) \Block2-1Salt
    2.23\pm 1.20
    Powder Weighing Error (mg) \Block2-1Semolina
    4.73\pm 3.76
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    9.16\pm 8.89
    Powder Weighing Error (mg) \Block2-1Pectin
    9.78\pm 5.53
    Powder Weighing Error (mg) \Block2-1Flour
    15.71\pm 14.68
    \Block3-1In-distribution Error
    2.26\pm 2.45
    \Block3-1Overall Error
    6.24\pm 8.42
  • 6
    \Block2-4Spoon Configuration Depth
    1.5
    \Block2-4Spoon Configuration Width
    1.5
    \Block2-4Spoon Configuration Spike Configuration
    {1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1}
    Powder Weighing Error (mg) \Block2-1Sand
    14.26\pm 2.14
    Powder Weighing Error (mg) \Block2-1Sugar
    14.12\pm 1.87
    Powder Weighing Error (mg) \Block2-1Salt
    *
    Powder Weighing Error (mg) \Block2-1Semolina
    *
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    *
    Powder Weighing Error (mg) \Block2-1Pectin
    *
    Powder Weighing Error (mg) \Block2-1Flour
    *
    \Block3-1In-distribution Error
    14.59\pm 1.42
    \Block3-1Overall Error
    14.76\pm 1.09
  • 7
    \Block2-4Spoon Configuration Depth
    0.7
    \Block2-4Spoon Configuration Width
    0.7
    \Block2-4Spoon Configuration Spike Configuration
    {1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1}
    Powder Weighing Error (mg) \Block2-1Sand
    0.98\pm 0.94
    Powder Weighing Error (mg) \Block2-1Sugar
    1.31\pm 2.29
    Powder Weighing Error (mg) \Block2-1Salt
    0.71\pm 0.66
    Powder Weighing Error (mg) \Block2-1Semolina
    5.21\pm 4.01
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    \mathbf{1.56\pm 1.40}
    Powder Weighing Error (mg) \Block2-1Pectin
    10.37\pm 5.36
    Powder Weighing Error (mg) \Block2-1Flour
    10.88\pm 9.15
    \Block3-1In-distribution Error
    2.05\pm 2.94
    \Block3-1Overall Error
    4.43\pm 5.94
  • 8
    \Block2-4Spoon Configuration Depth
    1.1
    \Block2-4Spoon Configuration Width
    1.1
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    1.12\pm 1.71
    Powder Weighing Error (mg) \Block2-1Sugar
    0.98\pm 0.65
    Powder Weighing Error (mg) \Block2-1Salt
    2.98\pm 4.50
    Powder Weighing Error (mg) \Block2-1Semolina
    4.55\pm 0.87
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    3.14\pm 1.63
    Powder Weighing Error (mg) \Block2-1Pectin
    10.44\pm 4.99
    Powder Weighing Error (mg) \Block2-1Flour
    13.89\pm 9.19
    \Block3-1In-distribution Error
    2.40\pm 2.79
    \Block3-1Overall Error
    5.30\pm 6.25
  • 9
    \Block2-4Spoon Configuration Depth
    1.1
    \Block2-4Spoon Configuration Width
    1.1
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    0.76\pm 0.94
    Powder Weighing Error (mg) \Block2-1Sugar
    1.09\pm 1.29
    Powder Weighing Error (mg) \Block2-1Salt
    3.41\pm 2.78
    Powder Weighing Error (mg) \Block2-1Semolina
    3.46\pm 3.84
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    5.81\pm 4.43
    Powder Weighing Error (mg) \Block2-1Pectin
    7.83\pm 4.74
    Powder Weighing Error (mg) \Block2-1Flour
    \mathbf{10.65\pm 4.36}
    \Block3-1In-distribution Error
    2.18\pm 2.72
    \Block3-1Overall Error
    4.71\pm 4.75

IV-C Tool Optimisation via BOHB

We ran a BOHB-driven search, as described in Section III, over the morphological design space \Xi. We used the BOHB implementation from SMAC3 [16], with a random forest as the surrogate model and expected improvement (EI) as the acquisition function. We set the successive halving proportion to \eta=3, with a minimum budget b_{\mathrm{min}}=330 episodes and a maximum budget b_{\mathrm{max}}=3000 episodes. We provide the configurations evaluated in Section IV-B as warm-start points for the optimiser. When a configuration is promoted to a higher budget, training resumes from its checkpoint rather than re-initialising. All other hyperparameters follow Section IV-B. We evaluate n=2 random seeds per configuration at the budget b prescribed by successive halving. Accounting for each seed independently, the total 200,000-episode optimisation budget required approximately 7.85 days of compute across four parallel workstations.

Fig. 5: Top three morphologies discovered by the BOHB optimiser. The simulated models are shown in the top row, with their corresponding physical 3D-printed prototypes below. The table details the specific width, depth, and spike configurations for each design.
TABLE II: Simulation and real-world deployment comparison of the best 3 tool configurations found by BOHB, reported from the high-fidelity simulation environment and for IID powders, averaged over 10 real-world runs for w_{target}=15\text{\,}\mathrm{m}\mathrm{g}.
  • Config a
    Powder Weighing Error (mg) Sand
    0.80\pm 0.45
    Powder Weighing Error (mg) Sugar
    0.96\pm 0.89
    Powder Weighing Error (mg) Salt
    1.03\pm 0.89
    Powder Weighing Error (mg) Semolina
    2.58\pm 1.58
    \Block2-1Simulation Error
    0.97\pm 0.31
    \Block2-1Real-world Error
    1.31\pm 1.17
  • Config b
    Powder Weighing Error (mg) Sand
    \mathbf{0.55\pm 0.24}
    Powder Weighing Error (mg) Sugar
    \mathbf{0.71\pm 0.35}
    Powder Weighing Error (mg) Salt
    \mathbf{0.65\pm 0.73}
    Powder Weighing Error (mg) Semolina
    1.85\pm 1.43
    \Block2-1Simulation Error
    0.94\pm 0.09
    \Block2-1Real-world Error
    \mathbf{0.91\pm 0.93}
  • Config c
    Powder Weighing Error (mg) Sand
    0.76\pm 0.58
    Powder Weighing Error (mg) Sugar
    1.0\pm 0.9
    Powder Weighing Error (mg) Salt
    1.18\pm 1.44
    Powder Weighing Error (mg) Semolina
    \mathbf{1.52\pm 1.36}
    \Block2-1Simulation Error
    \mathbf{0.85\pm 0.17}
    \Block2-1Real-world Error
    1.10\pm 1.10

We selected the top three morphologies identified by the BOHB optimiser, ranked by their weighing error in the high-fidelity environment at the maximum budget b_{\mathrm{max}}. For these three designs, we train two additional random seeds, bringing the total to four per configuration. We report the simulated error as the average across four seeds and deployed the best-performing policy from simulation for each morphology to the physical robot. Fig. 5 illustrates the three designs in simulation alongside their 3D-printed counterparts, with their morphological configurations. Table II compares the simulated error against the average real-world error. We report the real-world error only for in-distribution materials, as the findings in Section IV-B demonstrated that zero-shot transfer for OOD, highly cohesive materials that are poorly modelled in our simulation framework is unreliable.

All three configurations are predicted by the simulator to outperform the best tool of Section IV-B. However only Config. b and Config. c achieve lower real-world errors. Config. b demonstrates an overall improvement of 0.26\text{\,}\mathrm{m}\mathrm{g} across the in-distribution powders, exhibiting better performance on sand, salt, and semolina, while Config. c achieves the lowest absolute error on semolina. Although the optimiser predicted Config. c to be the best design overall, Config. b outperformed it in physical trials and exceeded its own simulated prediction. The sim-to-real prediction gaps for these configurations nonetheless remain small (0.25\text{\,}\mathrm{m}\mathrm{g} for Config. c and 0.03\text{\,}\mathrm{m}\mathrm{g} for Config. b). We attribute this rank inversion to the reality gap between the simulator and physical setup, which includes unmodelled dynamics and morphological variations introduced by 3D printing.

IV-D Accelerated Co-design via Morphological Similarity

We evaluated the similarity-based strategy introduced in Section III-E, to determine whether it accelerates the optimisation process without degrading the real-world performance of the discovered tools. To compute the spike similarity score S_{\mathrm{spikes}}, we defined the spike arc ratios as c_{i}\in\{1/32,1/8\}. The overall similarity S(\xi_{1},\xi_{2}) was then computed using Equation 7, with the coefficients \alpha, \beta, and \gamma tuned to 0.5, 0.3, and 0.2, respectively. To examine whether the metric captures task-relevant structure, we aggregated the morphologies from the initial grid search (Section IV-B) and the maximum budget evaluations from the BOHB optimisation (Section IV-C), and projected them into two dimensions using kernel principal component analysis (KPCA) [17], with S(\xi_{1},\xi_{2}) supplied as a precomputed kernel instead of a distance in the raw parameter space. Fig. 6 illustrates the resulting distribution of the dimensionality-reduced morphologies, mapped alongside their corresponding simulated weighing errors. Morphologies that cluster in the embedding, which appear structurally similar upon visual inspection, exhibit comparable task performance, indicating that the proposed similarity is correlated with dispensing behaviour. This supports using this metric to warm-start and accelerate policy training.

Fig. 6: Each point is two-dimensional representation of a spoon configuration. Spoon designs that are located closer together in the 2D space have higher morphological similarity. The colour of the points represents their task error in simulation.

We repeated the search from Section IV-C using identical hyperparameters, with the similarity-based logic from Algorithm 1. The similarity threshold \sigma and warmup fraction \phi were set to 0.85 and 0.33, respectively. Figure 7 compares unique configurations explored within the fixed budget of 200,000 episodes by the BOHB (Section IV-C) and its similarity-augmented variant. Both methods exhibit similar initial exploration rates, but the similarity-augmented BOHB accelerates as the number of sampled configurations increase. Ultimately, the similarity approach explores 136 unique configurations, compared to 106 for the standard baseline. As the cached policy pool grows, the warm-start hit-rate increases, reducing the average training cost per configuration.

Following the evaluation protocol of Section IV-C, we select the top three configurations, train two additional random seeds each and deploy the best-performing policy on the physical robot. Table III reports the sim-to-real transfer with the predicted simulation errors. All three morphologies outperformed the grid-search tool of Section IV-B in simulated and physical experiments, with predicted errors comparable to the standard BOHB run. Notably, Config. B is morphologically identical to Config. b from Section IV-C. Upon sim-to-real transfer, Config. B now shows a minor improvement of 0.10\text{\,}\mathrm{m}\mathrm{g} weighing error, emerging as the best-performing physical morphology in this experiment as well. While this variance in physical performance is likely attributable to policy seeding, the convergence on the same geometry under both settings indicates that similarity-based warm-starting accelerates morphological search without degrading the quality of the final tool.

Fig. 7: BOHB performance with and without morphology similarity filter. The plot shows total unique sampled configurations at 50000, 125000 and 200000 total episodes used.
TABLE III: Simulation and real-world deployment comparison of the best 3 tool configurations found by BOHB with similarity, reported from the high-fidelity simulation environment and for IID powders, averaged over 10 real-world runs for w_{target}=15\text{\,}\mathrm{m}\mathrm{g}.
  • Config A
    Powder Weighing Error (mg) Sand
    0.65\pm 0.54
    Powder Weighing Error (mg) Sugar
    \mathbf{0.47\pm 0.24}
    Powder Weighing Error (mg) Salt
    0.70\pm 0.36
    Powder Weighing Error (mg) Semolina
    1.82\pm 1.25
    \Block2-1Simulation Error
    0.95\pm 0.05
    \Block2-1Real-world Error
    0.91\pm 0.87
  • Config B
    Powder Weighing Error (mg) Sand
    0.62\pm 0.55
    Powder Weighing Error (mg) Sugar
    0.56\pm 0.23
    Powder Weighing Error (mg) Salt
    \mathbf{0.41\pm 0.18}
    Powder Weighing Error (mg) Semolina
    \mathbf{1.66\pm 0.92}
    \Block2-1Simulation Error
    0.92\pm 0.08
    \Block2-1Real-world Error
    \mathbf{0.81\pm 0.73}
  • Config C
    Powder Weighing Error (mg) Sand
    \mathbf{0.59\pm 0.27}
    Powder Weighing Error (mg) Sugar
    0.73\pm 0.50
    Powder Weighing Error (mg) Salt
    0.91\pm 1.33
    Powder Weighing Error (mg) Semolina
    1.97\pm 0.91
    \Block2-1Simulation Error
    \mathbf{0.86\pm 0.22}
    \Block2-1Real-world Error
    1.05\pm 0.98

IV-E Discussion

Our framework identifies tools that outperform human-centric standard ones and manually-designed baselines. While our high-fidelity simulation predicts real-world performance for IID materials (28^{\circ} to 41^{\circ} AoR), we deployed our best spoon (Config. B) on OOD materials to evaluate generalisability (Table IV). As material dynamics drift from the training distribution, the control policy becomes less reliable as the sim-to-real gap widens. Despite this degradation on OOD powders, Config. B achieves lower overall errors than the baselines across the combined set. While our co-designed tool is specialised for its target distribution of powder dynamics, it remains a more effective all-purpose tool for robotic chemists than standard human tools.

TABLE IV: Zero-shot transfer results on OOD materials with their respective AoR, averaged over 10 runs for w_{target}=15\text{\,}\mathrm{m}\mathrm{g}.
  • Standard Tool
    IID Error
    1.93\pm 2.16
    Sodium Bicarbonate (43^{\circ})
    4.45\pm 5.57
    Pectin (46^{\circ})
    3.58\pm 3.64
    Flour (51^{\circ})
    12.91\pm 13.58
    Overall Real-world Error
    4.09\pm 6.82
  • Config. 2
    IID Error
    1.17\pm 1.13
    Sodium Bicarbonate (43^{\circ})
    3.62\pm 3.45
    Pectin (46^{\circ})
    8.21\pm 4.04
    Flour (51^{\circ})
    11.41\pm 8.72
    Overall Real-world Error
    4.04\pm 5.43
  • Config. B
    IID Error
    \mathbf{0.81\pm 0.73}
    Sodium Bicarbonate (43^{\circ})
    \mathbf{1.98\pm 1.53}
    Pectin (46^{\circ})
    \mathbf{2.71\pm 2.43}
    Flour (51^{\circ})
    \mathbf{7.69\pm 4.01}
    Overall Real-world Error
    \mathbf{2.23\pm 3.00}

V Conclusion

We presented a co-design framework for autonomous powder weighing that jointly optimises a tool’s morphology and its reinforcement learning control policy, using a geometric similarity metric to warm-start the search from structurally related candidates. Real-world experiments show the co-designed tools outperform both a standard tool and grid-search baselines across materials. The main limitation is that our simulation cannot accurately reproduce cohesive powders; future work will develop higher-fidelity granular simulation that remains tractable for search. Ultimately, extending joint morphology–control optimisation to other contact-rich laboratory tasks will improve robot–material manipulation in autonomous scientific discovery.

Acknowledgements

Gemini 3.1 Pro was used in support for code-generation and creating functions for plotting results; all code and results were reviewed by all authors.

References

  1. [1] G. Tom, S. P. Schmid, S. G. Baird, Y. Cao, K. Darvish, H. Hao, S. Lo, S. Pablo-Garcia, E. M. Rajaonson, M. Skreta, N. Yoshikawa, S. Corapi, G. D. Akkoc, F. Strieth-Kalthoff, M. Seifrid, and A. Aspuru-Guzik (2024) Self-driving laboratories for chemistry and materials science. Chemical Reviews 124 (16), pp. 9633–9732.
  2. [2] Y. Kadokawa, M. Hamaya, and K. Tanaka (2023) Learning robotic powder weighing from simulation for laboratory automation. In IEEE/RSJ IROS, Vol. . External Links: Document
  3. [3] N. Radulov, A. Wright, T. Little, A. I. Cooper, and G. Pizzuto (2026) FLIP: flowability-informed powder weighing. In IEEE ICRA,
  4. [4] Y. Nakajima, M. Hamaya, K. Tanaka, T. Hawai, F. von Drigalski, Y. Takeichi, Y. Ushiku, and K. Ono (2023) Robotic powder grinding with audio-visual feedback for laboratory automation in materials science. In IEEE/RSJ IROS, Vol. . External Links: Document
  5. [5] N. Radulov, T. Little, A. I. Cooper, and G. Pizzuto (2026) Vision-guided adaptive scooping for powder weighing in autonomous chemistry laboratories. Digital Discovery 5 (5), pp. 2120–2127. External Links: ISSN 2635-098X, Document
  6. [6] R. Schneider, D. Honerkamp, T. Welschehold, and A. Valada (2025) Task-driven co-design of mobile manipulators. IEEE Robotics and Automation Letters 10 (7), pp. 7158–7165. External Links: Document
  7. [7] C. Lin, H. Yuan, Y. Wang, X. Qiu, T. Wang, M. Guo, B. Wang, Y. Narang, D. Fox, and C. Gan (2025) RobotSmith: generative robotic tool design for acquisition of complex manipulation skills. External Links: 2506.14763
  8. [8] M. Li, R. Antonova, D. Sadigh, and J. Bohg (2023) Learning tool morphology for contact-rich manipulation tasks with differentiable simulation. In IEEE ICRA, Vol. , pp. 1859–1865. External Links: Document
  9. [9] T. Takahashi, C. C. Beltran-Hernandez, Y. Kuroda, K. Tanaka, M. Hamaya, and Y. Ushiku (2025) SCU-hand: soft conical universal robotic hand for scooping granular media from containers of various sizes. In IEEE ICRA, Vol. , pp. 5549–5555. External Links: Document
  10. [10] T. Takahashi, Y. Nakajima, C. C. Beltran-Hernandez, Y. Kuroda, K. Tanaka, M. Hamaya, K. Ono, and Y. Ushiku (2026) SCU-hand with integrated single-sheet valve: a funnel-shaped robotic hand for milligram-scale powder handling. In IEEE ICRA,
  11. [11] S. Falkner, A. Klein, and F. Hutter (2018) BOHB: robust and efficient hyperparameter optimization at scale. ICML.
  12. [12] R. S. Sutton and A. G. Barto (2018) Reinforcement learning: an introduction. A Bradford Book, Cambridge, MA, USA. External Links: ISBN 0262039249
  13. [13] Isaac Sim External Links: Link
  14. [14] S. Haddadin (2024) The franka emika robot: a standard platform in robotics research [survey]. IEEE Robotics & Automation Magazine 31 (4), pp. 136–148. External Links: Document
  15. [15] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine (2018) Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. In ICML,
  16. [16] M. Lindauer, K. Eggensperger, M. Feurer, A. Biedenkapp, D. Deng, C. Benjamins, T. Ruhkopf, R. Sass, and F. Hutter (2022) SMAC3: a versatile bayesian optimization package for hyperparameter optimization. Journal of Machine Learning Research 23 (54), pp. 1–9.
  17. [17] B. Schölkopf, A. Smola, and K. Müller (1998) Nonlinear component analysis as a kernel eigenvalue problem. Neural Computation 10 (5), pp. 1299–1319. External Links: Document