ROBOTNESS
ExpertenarXiv

Tool-Policy-Co-Design für präzises Pulverwiegen in der Laborautomatisierung

Nikola Radulov, Xin Yang, Kevin S. Luck, Gabriella Pizzuto
In 30 Sekunden

Das Preprint stellt ein Tool-Policy-Co-Design vor, das Werkzeuggeometrie und Reinforcement-Learning-Politik für die Pulverdosierung gemeinsam optimiert. In Simulation und auf einem Franka-Roboter erreicht das co-designte Werkzeug 45 Prozent weniger Wägefehler als ein Standardwerkzeug, auch bei ungesehenen Materialien. Eine neue Ähnlichkeitsmetrik warmstartet das Politiklernen aus ähnlichen Kandidaten und erhöht die Zahl explorierter Konfigurationen um 28 Prozent.

Forschungsfrage

Wie lassen sich Werkzeugmorphologie und Steuerungspolitik gemeinsam optimieren, um den Dosierfehler beim robotischen Pulverwiegen über ein Spektrum von Fließfähigkeiten zu minimieren?

Problem

Autonomes Pulverwiegen ist wegen nichtlinearer, heterogener Materialdynamik ein Engpass in der Laborautomatisierung. Übliche Werkzeuge sind für die menschliche Hand geformt; ihre feste Geometrie begrenzt die mit einer Roboterpolitik erreichbare Präzision.

Bisheriger Ansatz

Bisherige Ansätze trainieren Reinforcement-Learning-Politiken für feste, menschenzentrierte Werkzeuge oder setzen handgefertigte Spezialwerkzeuge wie die SCU-Hand ein. Einige Arbeiten optimieren die Werkzeuggeometrie bei fester Politik, aber die gemeinsame Optimierung von Morphologie und Politik unter variierenden Materialdynamiken blieb offen.

Neuer Ansatz

Das Verfahren formuliert eine Bi-level-Optimierung: Die äußere Schleife nutzt BOHB zur Suche über Tiefe, Breite und 14 Randspitzen-Höhen eines ellipsoiden Löffels. Die innere Schleife trainiert eine SAC-Politik mit umgestellter Belohnung, die die schrittweise Fehleränderung statt des absoluten Fehlers belohnt, und mit uniformem Sampling des Fließfähigkeitsbereichs statt Curriculum. Eine geometrische Ähnlichkeitsmetrik aus Volumen, Breite-zu-Tiefe-Verhältnis und Spitzen-Topologie startet das Training aus dem Cache ähnlicher Designs und reduziert so das Rechenbudget pro Kandidat.

Ergebnisse

Im Gittersuch-Baseline-Experiment erreicht die beste Konfiguration (Config. 2) real einen mittleren Fehler von 1,17 mg für in-distribution Pulver, das Standardwerkzeug 1,93 mg. Die BOHB-Suche findet Config. b mit 0,91 mg realem IID-Fehler und Config. c mit 1,10 mg. Über alle sieben Materialien inklusive OOD-Pulver erzielt Config. B 2,23 mg gegenüber 4,09 mg des Standardwerkzeugs, eine Reduktion um 45 Prozent. Die Ähnlichkeitsvariante exploriert 136 statt 106 Konfigurationen (+28 Prozent) und findet Config. B mit 0,81 mg realem IID-Fehler. Der mittlere Sim-to-Real-Gap im Hochfidelitätssimulator liegt bei 1,78 ± 1,94 mg für IID-Materialien; auf OOD-Pulver weitet er sich auf 4,18 ± 1,78 mg aus.

Grenzen

Die Simulation kann kohäsive Pulver, etwa Mehl mit Ruhewinkel über 41 Grad, nicht präzise abbilden, weshalb der Transfer auf OOD-Materialien unzuverlässig wird. Die reale Validierung umfasst nur einen Franka FR3, sieben Materialien und 10 Läufe pro Pulver. Der Suchraum ist durch die 3D-Druckauflösung diskretisiert. Im Paper nicht berichtet: Code-Verfügbarkeit, Peer-Review, Übertragung auf andere Roboter oder Waagen.

Bedeutung für die Branche

Selbstfahrende Labore in Chemie, Pharma und Materialforschung könnten co-designte Dosierwerkzeuge in den nächsten 1 bis 3 Jahren als Spezialkomponenten für Roboterstationen einsetzen, sofern die Simulation für kohäsive Stoffe verbessert und die Übertragbarkeit auf weitere Manipulatoren validiert wird. Werkzeughersteller könnten solche Geometrien als roboteroptimierte Alternativen zu menschenzentrierten Löffeln produzieren. Längerfristig ist eine Ausweitung auf andere kontaktintensive Laboraufgaben denkbar.

Vollständige Studie

Tool-Policy Co-Design for Powder Weighing in Laboratory Automation

Nikola Radulov, Xin Yang, Kevin S. Luck, Gabriella Pizzuto

Veröffentlicht unter CC BY 4.0. Wiedergabe mit Namensnennung, Original unter arXiv:2609.39797 (PDF).

Abstract

Autonomous powder weighing is one of many bottlenecks in laboratory automation due to the complex, non-linear dynamics of heterogeneous materials. Robot chemists performing this task utilise standard tools shaped for the dexterity of human hands, whose fixed geometry sets the dynamics that the control policy needs to regulate. This work introduces a tool-policy co-design framework that concurrently optimises the morphology of a dispensing tool and its control policy for use by robots in chemistry laboratories, formulated as a bi-level optimisation that minimises dispensing error over a target distribution of powder flowabilities. The outer loop varies tool-design parameters such as tool depth, width and rim spike topology using Bayesian optimisation and hyperband, while an inner loop optimises a control policy for each candidate morphology. We also introduce a geometric similarity metric that warm-starts policy training from cached policies of structurally similar designs, exploring 28\% more configurations under the same compute budget. The proposed framework is evaluated on a robotic powder weighing task across seven materials with distinct physical dynamics in a flowability-informed robot-material simulation framework. Experimental results demonstrate that our co-designed tool morphology reduces real-world weighing errors by 45\% relative to a standard tool, including on previously unseen materials. These results demonstrate our method can adapt both the control policy and the physical tool to the dynamics of the target material, bringing a new paradigm for material manipulation to the field of laboratory automation.

I Introduction

Accelerating the discovery of materials is crucial for addressing global challenges and self-driving laboratories (SDLs) pursue this by integrating automated hardware with data-driven decision making [1]. To date, laboratory robots have been deployed predominantly for sample transportation, where the manipulated object is rigid. Sample preparation remains a fundamental challenge, as materials deform and flow under contact and their response is governed by bulk properties that differ across materials. Robot-material manipulation underpins solid-state chemistry [1], yet robots inherit tools such as spatulas [2, 3] shaped for use by the dexterous human hand rather than for a manipulator or the material. Recent approaches use deep reinforcement learning (RL) to adapt robot manipulators to these dynamics [2, 3], which optimise a control policy around a fixed tool. Tool geometry defines the action-to-mass transfer function, bounding the task precision achievable by any control policy using that tool. While co-design of morphology and control is increasingly explored, it remains largely unaddressed for tool-material manipulation which is fundamental to the success of robot-driven chemistry lab automation.

Fig. 1: Tool-policy co-design method for powder weighing. We use BOHB to propose tool morphologies \xi. For each candidate \xi, a control policy \pi_{\xi} is trained in a low-fidelity environment and evaluated in a high-fidelity digital twin. The highest-performing design and paired policy are transferred for real-world validation.

To address this, we present a co-design framework that concurrently optimises both the morphology of a dispensing tool and the robot’s reinforcement learning control policy, formulated as a bi-level optimisation problem that minimises dispensing error across a target distribution of powder flowabilities (Fig. 1). In an outer loop, Bayesian optimisation and hyperband (BOHB) explores a parameterised, low-dimensional morphological search space defining the depth, width, and discrete rim spike topology of an ellipsoidal tool. We select BOHB over standard Bayesian optimisation because successive halving evaluates candidates at low training budgets and promotes only top performers, saving compute on sub-optimal morphologies. In the inner loop, an RL control policy is trained for each candidate tool. To make the search tractable, we introduce a geometric similarity metric that warm-starts policy training from cached policies of structurally similar designs. We validate the framework in simulation and by zero-shot transfer to a physical robot.

In summary, the main contributions of this work are:

  • A bi-level co-design framework for milligram-scale powder weighing that jointly optimises a reinforcement learning control policy and the tool morphology across a distribution of material flowabilities, over a parameterised design space derived from standard tool geometries that combines continuous bowl scaling with discrete rim spike configurations.
  • A morphological similarity metric that accelerates the search by warm-starting policy training from structurally related candidates, without compromising the real-world performance of the discovered tool.
  • An empirical evaluation in simulation and on a real robotic manipulator, demonstrating that the co-designed tools outperform both a standard tool and grid-search baselines across a diverse range of powder behaviours.

II Related Work

II-A Robot Skill Learning for Laboratory Automation

Robotic scientists have transitioned from open-loop automation to adaptive modular systems, motivated by the need for generalisable approaches capable of handling unpredictable, heterogeneous samples [1]. While self-driving labs have advanced high-level planning and experimental decision making [1], low-level manipulation remains a fundamental bottleneck. Sample grinding [4], powder weighing [2] and scooping [5] are challenging due to the non-linear, unpredictable dynamics of materials. To address this, recent works leverage deep reinforcement learning through sim-to-real domain randomisation [2] or embed material properties directly into the learning process [3]. Prior work optimises a control policy around a fixed, human-centric tool, so task success is bounded by a geometry that was not designed for robot-material manipulation. We instead formulate the tool as a design variable of the robot skill itself.

II-B Robotic Tool Co-Design

Co-design jointly optimises a robot’s morphology and its control policy, on the premise that a fixed embodiment bounds control performance. Recent works search a parametrised design space, either through bi-level formulations pairing Bayesian optimisation with reinforcement learning, e.g. to design mobile manipulator mountings [6], or through generative models and cross-embodiment policies combined with grammar-based search to scale the synthesis of tool shapes [7]. Closest to our work, Li et al. [8] optimise tool geometry for contact-rich tasks over a distribution of task variations using differentiable simulation. However, they vary the initial object pose instead of its dynamics and hold the policy fixed, leaving joint optimisation under such variation unaddressed. Within laboratory automation, specialised end-effectors have been engineered e.g., the SCU-Hand uses a soft, reconfigurable conical sheet for adaptive scooping [9], later extended with a single-sheet valve for milligram-scale dispensing [10]. However, the morphologies of these tools are manually handcrafted. Our approach concurrently optimises the control policy alongside the tool’s physical design, tailoring the tool to the flow properties of the material.

III Methodology

Our method formulates robotic powder weighing as a joint optimisation problem with bi-level structure over tool morphology and control policy. The goal is to identify a tool design that minimises the expected dispensing error over a distribution of materials with diverse physical dynamics.

III-A Problem Formulation

We consider the task of autonomous powder weighing with a robotic manipulator that holds a custom tool at its end-effector and dispenses powder into a container on an analytical balance (Fig. 1). The balance provides weight feedback at each control step and an episode terminates after a fixed number T of steps. Following Radulov et al.’s method [3], we characterise a material by its static flowability. This is quantified by the angle of repose (AoR), where a larger AoR corresponds to a lower flowability. The flowability F of a material is sampled from a representative range F_{r}=[AoR_{min},AoR_{max}] across the task distribution.

Let the morphology of the tool be described by a parameter vector \xi\in\Xi, where \Xi denotes the design space defined in Section III-B. Let \pi_{\xi} be the robot control policy trained with that tool. As the geometry of the tool determines how much powder is displaced per action, \xi and \pi_{\xi} cannot be chosen independently, i.e., a morphology is only as good as the policy that can be learnt for it and vice versa. Furthermore, powder dispensing is difficult to mathematically model as the dynamics vary strongly with material properties and with the amount of powder remaining in the tool [2, 3].

To find the optimal combination (\xi^{*},\pi^{*}_{\xi^{*}}) we define the bi-level optimisation problem as:

\displaystyle\xi^{*}=\arg\min_{\xi}\mathbb{E}_{F_{r}}[\mathcal{L}(\xi,\pi^{*}_{\xi},F_{r})], \\ \displaystyle\text{s.t.~}\pi^{*}_{\xi}=\arg\max_{\pi}\mathbb{E}_{\begin{subarray}{c}\pi(a_{t}|s_{t})\\
F\sim F_{r}\\
p(s_{t+1}|s_{t},a_{t},\xi,F)\end{subarray}}\left[G\right]
(1) (2)

Here, we optimise the morphology \xi and policy \pi_{\xi} using a morphology objective \mathcal{L} and a policy learning objective G. The complete bi-level co-design framework is illustrated in Fig. 1.

III-B Outer Loop: Tool Morphology Parametrisation

We base the design space on a standard spoon-like geometry, where the bowl is modelled as an ellipsoid of height h, width w, and length l, sliced at the equator. Three groups of parameters, illustrated in Fig. 2, form \xi.

Fig. 2: Tool morphology parameters of the design space; l, h, w represent the length, depth and width of the spoon bowl. H_{k} represents the height of the k^{th} spike.
  1. Tool depth (h) governs the vertical extent of the spoon bowl. A shallower depth yields a flatter geometry, requiring less aggressive agitation and smaller pitch adjustments, which can be ideal for cohesive materials.
  2. Tool width (w) controls the lateral scale of the bowl. Increased widths broaden the surface area, enabling the capture of larger volumes per scoop.
  3. Tool rim spike heights (H_{k}) are used for finer control over material outflow. The outer optimisation loop can independently vary the height of each spike. As the spikes are modelled along the edge of the bowl, their spatial distribution is inherently coupled with the tool depth.

The bowl length l is constant, which keeps the tool centre point identical for all candidates. Should l be allowed to vary, the parameters of the low-level Cartesian position controller executing the shake and incline primitives would have to be retuned per candidate, in simulation and on the robot, to avoid collisions with the vial and analytical balance. Since the bowl volume varies across \Xi, the granular simulation is configured so that its fidelity is equivalent for every candidate.

III-C Outer Loop: Tool Design Generation

We search the morphological space \Xi with BOHB [11]. This method proposes candidates from a surrogate model fitted to past evaluations and allocates a budget b, defined as the number of training episodes granted to the inner loop. Budgets are assigned by successive halving, where a large pool of morphologies is trained on a small budget, the worst performers are discarded and the survivors are promoted to a longer budget. At each outer-loop iteration, BOHB samples a candidate tool design \xi\in\Xi. This configuration is passed to the robotic simulator, which applies the required scaling to instantiate the custom tool geometry before initiating the inner loop to train a corresponding control policy for b episodes. Due to the stochasticity inherent in policy optimisation, we evaluate n random seeds per configuration. We define the outer-loop objective function \mathcal{L}(\xi,\pi^{*}_{\xi},F_{r}) as the expected weighing error across a flowability range F_{r}.

\mathcal{L}(\xi,\pi^{*}_{\xi},F_{r})=\mathbb{E}_{F\in F_{r}}\left|w_{\mathrm{target}}-w_{\pi^{*}_{\xi},T}^{F}\right|
(3)

where w_{\mathrm{target}} represents the target weight to be dispensed, and w_{\pi^{*}_{\xi},T}^{F} represents the final dispensed weight for material F, using the candidate tool design \xi and its associated control policy \pi^{*}_{\xi}. In practice, to increase stability we sample over a distribution/set of optimised policies (i.e., multiple seeds), i.e., \mathcal{L}_{\text{avg}}=\frac{1}{n}\sum_{\pi^{*}_{\xi}\in\Pi^{*}_{\xi}}\mathcal{L}(\xi,\pi^{*}_{\xi},F_{r}).

III-D Inner Loop: Policy Optimisation

Given a morphological design \xi, the aim of the inner loop is to optimise a corresponding control policy \pi_{\xi} and provide a reliable estimate of its performance over the target flowability range F_{r}. We formulate the powder weighing task as a Markov decision process (MDP) [12] \mathcal{M}=\langle\mathcal{S},\mathcal{A},\mathcal{P},\mathcal{R},\gamma\rangle. We employ the same parameterised action space as prior works [2, 3], where at each step, the agent adjusts the tool pitch by a_{\mathrm{incline}}, and subsequently performs a shaking motion by retracting the tool a distance a_{\mathrm{shake}} before returning it to its home position. The observation o\in\mathcal{S} follows prior work: o=(w_{current},w_{target},\theta_{spoon}). We introduce two enhancements from Radulov et al.’s work [3].

III-D1 Reward

We redesign the reward to penalise the step-wise change in error rather than the absolute error at the current step:

\mathcal{R}_{t}=\frac{\Delta_{t-1}-\Delta_{t}}{w_{\mathrm{target}}}
(4)

where \Delta_{t}=|w_{\mathrm{target}}-w_{t}| is the absolute weight error at step t, with the initial condition \Delta_{0}=w_{\mathrm{target}}. This formulation telescopes such that the cumulative return over an episode becomes G=\sum_{t=1}^{T}\mathcal{R}_{t}=1-\frac{\Delta_{T}}{w_{\mathrm{target}}}, which depends only on the final dispensing error. Under the original step-wise absolute error reward, the inner loop implicitly optimised for speed as well as accuracy, encouraging the agent to close the mass gap rapidly and sometimes overshoot. Since our outer loop scores a morphology strictly on final precision \mathcal{L} (Equation 3), this reward corrects the mismatch between the inner loop’s objective and the outer loop’s cost function.

III-D2 Sampling of the Flowability Range

We replace the curriculum learning strategy with uniform random sampling of the flowability range F_{r}. A curriculum makes the training distribution a function of the budget b which conflicts with the successive halving of BOHB in our outer loop. At low budgets the agent would only be exposed to the easy end of F_{r}, so Equation (3) would rank morphologies on a narrow subset of materials, shifting the ranking as candidates are promoted. Uniform sampling makes the score at every budget an unbiased estimate of the same quantity, so that evaluations obtained at different fidelities remain comparable.

III-E Similarity-based Evaluation

Evaluating the outer loop objective function \mathcal{L}(\xi,\pi^{*}_{\xi},F_{r}) presents a computational bottleneck as it requires training the control policy \pi_{\xi} from scratch for every proposed tool morphology \xi. However, during the BOHB optimisation process, the algorithm frequently samples configurations that represent minor parametric perturbations of previously evaluated designs. As the physical dynamics of powder flow remain largely consistent across minor morphological adjustments, we leverage the fully trained policy of a structurally similar design to warm-start the training of the new control policy. To facilitate this accelerated search, we introduce a morphological similarity metric S(\xi_{1},\xi_{2})\in[0,1] that quantifies the resemblance between two geometries based on their total volume, geometry and spike topology.

For both the total volume V and the width-to-depth ratio R, we employ a min–max ratio to ensure scale-invariant similarity bounds within [0,1]. The total volume V of a given morphology is computed as the sum of the ellipsoidal half-bowl volume and the volume of the watertight elliptical cylinder formed by the rim spikes. The height of this cylinder is bounded by the shortest spike. The volume and ratio similarity scores are computed as:

S_{V}=\frac{\min(V_{1},V_{2})}{\max(V_{1},V_{2})},\quad S_{R}=\frac{\min(R_{1},R_{2})}{\max(R_{1},R_{2})}
(5)

With spike height already captured volumetrically, the final term measures perimeter topology, where we weight the squared differences between the first-order discrete derivatives of adjacent spikes by their arc ratio c_{i} and map the result into [0,1] with a negative exponential function:

S_{\mathrm{spikes}}=\exp\left(-\sum_{i=1}^{M}c_{i}(\delta_{1,i}-\delta_{2,i})^{2}\right)
(6)

The overall similarity score S(\xi_{1},\xi_{2}) is defined as a weighted linear combination of S_{V}, S_{R}, and S_{\mathrm{spikes}}:

S(\xi_{1},\xi_{2})=\alpha S_{V}+\beta S_{R}+\gamma S_{\mathrm{spikes}}
(7)

Here, \alpha+\beta+\gamma=1.

We incorporate the similarity metric into our optimisation method as detailed in Algorithm 1, where the similarity logic is implemented in lines 10 to 16. Throughout the optimisation, the system maintains an observation history \mathcal{H} to train the Bayesian optimisation surrogate model and a policy registry \mathcal{P} to cache learned control policies. The algorithm first calculates the maximum bracket index s_{\max}, which dictates the maximum number of successive halving stages. BOHB iterates through several brackets, starting from the most aggressive one with s_{\max} successive halvings. At the start of a stage, a set of N initial morphologies \mathcal{C} is sampled from a Bayesian surrogate model \mu(\xi|\mathcal{H}) conditioned on prior history. Within each successive halving stage k, the target budget b is computed. Before training a policy \pi_{\xi_{i}} for a morphology \xi_{i}, the algorithm computes the similarity score S(\xi_{i},\xi_{j}) (line 10) against all previously trained morphologies \xi_{j}\in\mathcal{H}. It filters for a candidate set \mathcal{S} of morphologies that meet a minimum similarity threshold \sigma and precisely match the target budget b. If a valid match is found, the policy \pi^{*}\in\mathcal{P} of the most similar morphology \xi^{*} is used to warm-start \pi_{\xi_{i}} and the execution budget b is reduced to \phi. Conversely, if \mathcal{S} is empty, the policy undergoes standard training for the full budget b. After training, the policy is evaluated to determine its task error e_{i} and both the registry \mathcal{P} and history \mathcal{H} are updated. Finally, the configuration pool \mathcal{C} is pruned, retaining only the top N_{k} performers to advance to the next iteration.

Algorithm 1 Tool-Policy Optimisation for Powder Weighing

IV Experimental Evaluation

In this section, we evaluated the proposed co-design framework in simulation and on a physical robot, with experiments designed to address the following research questions: (1) Can simulation predict real-world task performance and the relative ranking of tool morphologies? (2) Can BOHB co-design discover morphologies that outperform a standard (commercial) tool and grid-search baselines in the real world? (3) Does the morphological similarity threshold accelerate the search without compromising the tool performance?

IV-A Experimental Setup

IV-A1 Simulation Setup

Granular materials were simulated using the position-based dynamics (PBD) solver of NVIDIA Isaac Sim [13], which resolves constraints such as contacts at the position level prioritising computational efficiency over physical fidelity and making it suitable for reinforcement learning. We constructed two environments (Fig. 1) to balance accuracy against cost. The high-fidelity environment is a digital twin of our real-world setup and runs at a physics time step of dt=1/240 \text{\,}\mathrm{s}. The low-fidelity environment removes the analytical balance and the target receptacle, reducing collision calculations per step and allowing a coarser time step of dt=1/100\text{\,}\mathrm{s}. In both environments, a 3D model of the standard tool is attached to the robot’s end-effector via a fixed joint matching the grasp pose of the real-world system. Episodes are initialised with the powder already contained within the tool, bypassing the computationally expensive tool-picking and scooping phases. To simulate how a tool’s physical dimensions dictate its retained powder volume, we scale the baseline initial particle count bounds (N_{\text{low}},N_{\text{high}}) proportionally with the volumetric capacity of the tool morphology \xi. Specifically, the initial particle count scales with the tool’s width and its effective depth, accounting for both base bowl and the added volume from the spikes. Tool morphologies are instantiated by scaling the width and depth of the standard model. The M=14 rim spikes are modelled as independent rigid bodies at fixed positions relative to the tool centre. As dispensing is primarily directed through the tool tip, eight smaller spikes each spanning c_{i}=1/32 of the ellipsoid arc are positioned there, while the remaining six span c_{i}=1/8. Spike height is set by vertical scaling to one of four discrete levels [0,0.33,0.66,1]; if the morphological configuration \xi sets a spike’s height to zero, its corresponding rigid body is removed from the simulation. The depth and width scale factors are continuous over [0.7,1.5] relative to the standard tool, so \Xi combines two continuous scale factors with 4^{14} discrete spike configurations. Episode length is fixed to 10 steps. All experiments are distributed across four Ubuntu 22.04 workstations with NVIDIA RTX 4090/5090 GPUs; on the RTX 5090 machine, an episode averages 13.57\text{\,}\mathrm{s} in the low-fidelity environment and 24.95\text{\,}\mathrm{s} in the high-fidelity environment.

IV-A2 Materials

We defined the target material distribution over an AoR range F_{r}=[28^{\circ},41^{\circ}], as in prior works [3]. For our experiments, we sample materials at four fixed flowability points: 28^{\circ}, 32^{\circ}, 36^{\circ}, and 41^{\circ}. Highly cohesive materials (e.g., flour) are excluded from the range of simulated materials, as capturing their clumping dynamics would require alternative techniques that are prohibitively expensive for our framework.

IV-A3 Real World Setup

We use a Franka Research 3 (FR3) [14] manipulator equipped with a Robotiq 85F gripper, mounted parallel to the working surface for precise tool manoeuvring. A Sartorius Entris II precision analytical balance measures the dispensed mass, feeding the data directly to the control policy in real time. The tool is loaded using a parabolic scooping trajectory [5], without the vision-based volume estimation, as calibrating it across all tool–material combinations is prohibitively expensive. Instead, we manually tune the scooping parameters so the initial acquired mass is within the simulated training distribution. To transfer a morphology from simulation, we fabricated the custom tool using a Prusa XL 3D printer equipped with a 0.4\text{\,}\mathrm{m}\mathrm{m} nozzle, which yields a vertical print resolution of 0.1\text{\,}\mathrm{m}\mathrm{m} and a horizontal resolution of 0.4\text{\,}\mathrm{m}\mathrm{m}. Consequently, the morphological search space \Xi is discretised by these manufacturing constraints: the horizontal resolution sets the lower bound of 0.7 on the scale factors, below which the rim spikes fall under the minimum reproducible feature size, while the vertical resolution sets the spacing of the four spike height levels.

IV-B Sim-to-Real Transferability and Simulation Fidelity

To evaluate the capability of our simulation to predict real-world task performance, we perform a grid search across our morphological design space \Xi. We apply a step size of 0.4 for both the depth and width scaling factors, which evaluates the extremities and the midpoint of these parameter ranges. We constrain the grid search to three predefined spike configurations: (1) all spikes set to maximum height; (2) all spikes disabled (height set to zero); and (3) a manually-designed configuration where the six front-most spikes are disabled while the remaining eight are set to their maximum height (represented by the array \{0,0,0,1,1,1,1,1,1,1,1,0,0,0\}). To train the control policy \pi_{\xi}, we use Soft Actor-Critic (SAC) [15], a model-free, off-policy deep reinforcement learning algorithm. We adopt the neural network architecture and physics-informed modelling from prior work [3], which includes seven optimised data points per flowability level. Training terminated after 3000 episodes with n=3 random seeds per configuration. This is reduced from the 4000 episodes in prior works, as standard tool policies converge well before the limit and the tighter budget also favours morphologies that remain sample efficient.

For each configuration, we measured the empirical weighing error (Equation 3) across the sampled materials for a fixed target mass w_{\mathrm{target}}=$15\text{\,}\mathrm{m}\mathrm{g}$, reporting the mean and standard deviation across seeds. While policy training occurred exclusively within the computationally-efficient low-fidelity environment, the final weight error was evaluated in both environments (Fig. 3). In both environments, tools with all spikes enabled performed worst, yielding a mean error of 2.81\text{\,}\mathrm{m}\mathrm{g} in the low-fidelity environment and 4.69\text{\,}\mathrm{m}\mathrm{g} in the high-fidelity environment. The manually-designed spike configuration achieves the lowest error in both cases: 1.27\text{\,}\mathrm{m}\mathrm{g} and 1.54\text{\,}\mathrm{m}\mathrm{g} respectively. In general, configurations with increased internal volume, i.e., those in which both depth and width are scaled up, performed worse than their smaller counterparts. This is likely because the larger retained powder mass makes it harder for the policy to exert fine-grained control over the dispensing flow.

Fig. 3: The performance of each chosen morphology as the mean final test error and the respective standard deviation, averaged across 3 seeds. The temperature represents the mean final test error (\text{\,}\mathrm{m}\mathrm{g}). The top graph (3(a)) shows the raw task error when evaluated on the low-fidelity environment and the bottom graph (3(b)) shows the performance when evaluated on the high-fidelity test environment.

To evaluate the simulation’s predictive capability for tool optimisation, we selected nine morphologies for zero-shot transfer to the real setup, i.e., three from each spike configuration, encompassing the best- and worst-performing designs for the search. We evaluated the configurations on in-distribution (IID) materials and out-of-distribution (OOD) materials. We compare them against the standard tool, trained under identical conditions to those described in Section III-D. The results are presented in Table I. The most effective geometry is Config. 2, which performs best on both the IID set and the full material set. Both Config. 2 and Config. 4 achieve IID errors lower than those of the standard tool. Config. 2 yields the lowest error for IID powders, apart from salt, where Config. 4 performs best.

The sim-to-real prediction gap, defined as the absolute difference between simulated and physical error, averages 1.78\pm 1.94\text{\,}\mathrm{m}\mathrm{g} for the high-fidelity environment and 2.63\pm 2.83\text{\,}\mathrm{m}\mathrm{g} for the low-fidelity environment on IID materials. Over the entire test set (including OOD powders), this prediction gap widens to 4.18\pm 1.78 mg and 5.10\pm 2.34 mg respectively. This degradation is expected, as we explicitly exclude highly cohesive materials from simulation due to the prohibitive computational cost of modelling their complex dynamics, so the tool morphology cannot be optimised for clumping or compressible behaviour.

Fig. 4: Sim-to-real rank shifts across morphologies. Each spoon image shows its baseline simulation rank; blue arrows show shifts for IID powders.

We ranked configurations by high-fidelity simulation performance using an anchor-based scheme: configurations are sorted by mean error, and each shares the current anchor’s rank unless a one-tailed Student’s t-test at 95% confidence finds it significantly worse, where it becomes a new anchor. With n=3 seeds, configurations differing only slightly tend to share a rank. Fig. 4 illustrates the relative ranking shifts from simulation to the real world for IID materials. Only three of the tools change their standing: one simulation Rank 1 configuration drops to Rank 2, and two Rank 3 configurations are upgraded to Rank 2. The best and worst configurations by mean error retain their respective ranks across the sim-to-real gap, which supports the use of the simulation as a proxy for morphological optimisation.

TABLE I: Zero-shot transfer results on w_{target}=15\text{\,}\mathrm{m}\mathrm{g} of selected policy-tool pairs, evaluated using the best-performing policy in simulation for each morphology. For each powder, results report the mean absolute error and standard deviation over 10 real-world runs. Sodium bicarbonate, pectin and flour represent OOD materials, i.e., the policy was not trained on these specific powders or materials with similar flow. Data points marked with * represent cases where no material was dispensed.
  • Standard Tool
    \Block2-4Spoon Configuration Depth
    Keine Daten
    \Block2-4Spoon Configuration Width
    Keine Daten
    \Block2-4Spoon Configuration Spike Configuration
    Keine Daten
    Powder Weighing Error (mg) \Block2-1Sand
    1.49\pm 1.36
    Powder Weighing Error (mg) \Block2-1Sugar
    2.05\pm 1.73
    Powder Weighing Error (mg) \Block2-1Salt
    1.66\pm 1.19
    Powder Weighing Error (mg) \Block2-1Semolina
    2.52\pm 3.66
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    4.45\pm 5.57
    Powder Weighing Error (mg) \Block2-1Pectin
    \mathbf{3.58\pm 3.64}
    Powder Weighing Error (mg) \Block2-1Flour
    12.91\pm 13.58
    \Block3-1In-distribution Error
    1.93\pm 2.16
    \Block3-1Overall Error
    4.09\pm 6.82
  • 1
    \Block2-4Spoon Configuration Depth
    1.1
    \Block2-4Spoon Configuration Width
    1.1
    \Block2-4Spoon Configuration Spike Configuration
    {1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1}
    Powder Weighing Error (mg) \Block2-1Sand
    2.12\pm 4.21
    Powder Weighing Error (mg) \Block2-1Sugar
    4.78\pm 5.02
    Powder Weighing Error (mg) \Block2-1Salt
    12.56\pm 2.98
    Powder Weighing Error (mg) \Block2-1Semolina
    *
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    14.48\pm 0.23
    Powder Weighing Error (mg) \Block2-1Pectin
    13.15\pm 3.96
    Powder Weighing Error (mg) \Block2-1Flour
    13.29\pm 4.06
    \Block3-1In-distribution Error
    8.58\pm 6.36
    \Block3-1Overall Error
    10.74\pm 5.79
  • 2
    \Block2-4Spoon Configuration Depth
    1.1
    \Block2-4Spoon Configuration Width
    0.7
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    \mathbf{0.67\pm 0.63}
    Powder Weighing Error (mg) \Block2-1Sugar
    \mathbf{0.53\pm 0.53}
    Powder Weighing Error (mg) \Block2-1Salt
    1.01\pm 0.54
    Powder Weighing Error (mg) \Block2-1Semolina
    \mathbf{2.47\pm 1.41}
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    3.62\pm 3.45
    Powder Weighing Error (mg) \Block2-1Pectin
    8.21\pm 4.04
    Powder Weighing Error (mg) \Block2-1Flour
    11.41\pm 8.72
    \Block3-1In-distribution Error
    \mathbf{1.17\pm 1.13}
    \Block3-1Overall Error
    \mathbf{4.04\pm 5.43}
  • 3
    \Block2-4Spoon Configuration Depth
    1.5
    \Block2-4Spoon Configuration Width
    1.5
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    4.38\pm 2.27
    Powder Weighing Error (mg) \Block2-1Sugar
    2.29\pm 1.27
    Powder Weighing Error (mg) \Block2-1Salt
    1.66\pm 1.49
    Powder Weighing Error (mg) \Block2-1Semolina
    14.48\pm 0.99
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    14.49\pm 0.97
    Powder Weighing Error (mg) \Block2-1Pectin
    *
    Powder Weighing Error (mg) \Block2-1Flour
    14.66\pm 0.34
    \Block3-1In-distribution Error
    5.70\pm 5.44
    \Block3-1Overall Error
    9.52\pm 6.06
  • 4
    \Block2-4Spoon Configuration Depth
    0.7
    \Block2-4Spoon Configuration Width
    1.1
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    0.81\pm 0.38
    Powder Weighing Error (mg) \Block2-1Sugar
    0.76\pm 0.46
    Powder Weighing Error (mg) \Block2-1Salt
    \mathbf{0.56\pm 0.40}
    Powder Weighing Error (mg) \Block2-1Semolina
    2.76\pm 2.57
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    5.03\pm 4.08
    Powder Weighing Error (mg) \Block2-1Pectin
    7.04\pm 4.18
    Powder Weighing Error (mg) \Block2-1Flour
    13.70\pm 1.4
    \Block3-1In-distribution Error
    1.22\pm 1.56
    \Block3-1Overall Error
    4.38\pm 5.05
  • 5
    \Block2-4Spoon Configuration Depth
    1.5
    \Block2-4Spoon Configuration Width
    1.1
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    0.90\pm 0.53
    Powder Weighing Error (mg) \Block2-1Sugar
    1.18\pm 0.46
    Powder Weighing Error (mg) \Block2-1Salt
    2.23\pm 1.20
    Powder Weighing Error (mg) \Block2-1Semolina
    4.73\pm 3.76
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    9.16\pm 8.89
    Powder Weighing Error (mg) \Block2-1Pectin
    9.78\pm 5.53
    Powder Weighing Error (mg) \Block2-1Flour
    15.71\pm 14.68
    \Block3-1In-distribution Error
    2.26\pm 2.45
    \Block3-1Overall Error
    6.24\pm 8.42
  • 6
    \Block2-4Spoon Configuration Depth
    1.5
    \Block2-4Spoon Configuration Width
    1.5
    \Block2-4Spoon Configuration Spike Configuration
    {1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1}
    Powder Weighing Error (mg) \Block2-1Sand
    14.26\pm 2.14
    Powder Weighing Error (mg) \Block2-1Sugar
    14.12\pm 1.87
    Powder Weighing Error (mg) \Block2-1Salt
    *
    Powder Weighing Error (mg) \Block2-1Semolina
    *
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    *
    Powder Weighing Error (mg) \Block2-1Pectin
    *
    Powder Weighing Error (mg) \Block2-1Flour
    *
    \Block3-1In-distribution Error
    14.59\pm 1.42
    \Block3-1Overall Error
    14.76\pm 1.09
  • 7
    \Block2-4Spoon Configuration Depth
    0.7
    \Block2-4Spoon Configuration Width
    0.7
    \Block2-4Spoon Configuration Spike Configuration
    {1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1}
    Powder Weighing Error (mg) \Block2-1Sand
    0.98\pm 0.94
    Powder Weighing Error (mg) \Block2-1Sugar
    1.31\pm 2.29
    Powder Weighing Error (mg) \Block2-1Salt
    0.71\pm 0.66
    Powder Weighing Error (mg) \Block2-1Semolina
    5.21\pm 4.01
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    \mathbf{1.56\pm 1.40}
    Powder Weighing Error (mg) \Block2-1Pectin
    10.37\pm 5.36
    Powder Weighing Error (mg) \Block2-1Flour
    10.88\pm 9.15
    \Block3-1In-distribution Error
    2.05\pm 2.94
    \Block3-1Overall Error
    4.43\pm 5.94
  • 8
    \Block2-4Spoon Configuration Depth
    1.1
    \Block2-4Spoon Configuration Width
    1.1
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    1.12\pm 1.71
    Powder Weighing Error (mg) \Block2-1Sugar
    0.98\pm 0.65
    Powder Weighing Error (mg) \Block2-1Salt
    2.98\pm 4.50
    Powder Weighing Error (mg) \Block2-1Semolina
    4.55\pm 0.87
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    3.14\pm 1.63
    Powder Weighing Error (mg) \Block2-1Pectin
    10.44\pm 4.99
    Powder Weighing Error (mg) \Block2-1Flour
    13.89\pm 9.19
    \Block3-1In-distribution Error
    2.40\pm 2.79
    \Block3-1Overall Error
    5.30\pm 6.25
  • 9
    \Block2-4Spoon Configuration Depth
    1.1
    \Block2-4Spoon Configuration Width
    1.1
    \Block2-4Spoon Configuration Spike Configuration
    {0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0}
    Powder Weighing Error (mg) \Block2-1Sand
    0.76\pm 0.94
    Powder Weighing Error (mg) \Block2-1Sugar
    1.09\pm 1.29
    Powder Weighing Error (mg) \Block2-1Salt
    3.41\pm 2.78
    Powder Weighing Error (mg) \Block2-1Semolina
    3.46\pm 3.84
    Powder Weighing Error (mg) \Block2-1Sodium Bicarbonate
    5.81\pm 4.43
    Powder Weighing Error (mg) \Block2-1Pectin
    7.83\pm 4.74
    Powder Weighing Error (mg) \Block2-1Flour
    \mathbf{10.65\pm 4.36}
    \Block3-1In-distribution Error
    2.18\pm 2.72
    \Block3-1Overall Error
    4.71\pm 4.75

IV-C Tool Optimisation via BOHB

We ran a BOHB-driven search, as described in Section III, over the morphological design space \Xi. We used the BOHB implementation from SMAC3 [16], with a random forest as the surrogate model and expected improvement (EI) as the acquisition function. We set the successive halving proportion to \eta=3, with a minimum budget b_{\mathrm{min}}=330 episodes and a maximum budget b_{\mathrm{max}}=3000 episodes. We provide the configurations evaluated in Section IV-B as warm-start points for the optimiser. When a configuration is promoted to a higher budget, training resumes from its checkpoint rather than re-initialising. All other hyperparameters follow Section IV-B. We evaluate n=2 random seeds per configuration at the budget b prescribed by successive halving. Accounting for each seed independently, the total 200,000-episode optimisation budget required approximately 7.85 days of compute across four parallel workstations.

Fig. 5: Top three morphologies discovered by the BOHB optimiser. The simulated models are shown in the top row, with their corresponding physical 3D-printed prototypes below. The table details the specific width, depth, and spike configurations for each design.
TABLE II: Simulation and real-world deployment comparison of the best 3 tool configurations found by BOHB, reported from the high-fidelity simulation environment and for IID powders, averaged over 10 real-world runs for w_{target}=15\text{\,}\mathrm{m}\mathrm{g}.
  • Config a
    Powder Weighing Error (mg) Sand
    0.80\pm 0.45
    Powder Weighing Error (mg) Sugar
    0.96\pm 0.89
    Powder Weighing Error (mg) Salt
    1.03\pm 0.89
    Powder Weighing Error (mg) Semolina
    2.58\pm 1.58
    \Block2-1Simulation Error
    0.97\pm 0.31
    \Block2-1Real-world Error
    1.31\pm 1.17
  • Config b
    Powder Weighing Error (mg) Sand
    \mathbf{0.55\pm 0.24}
    Powder Weighing Error (mg) Sugar
    \mathbf{0.71\pm 0.35}
    Powder Weighing Error (mg) Salt
    \mathbf{0.65\pm 0.73}
    Powder Weighing Error (mg) Semolina
    1.85\pm 1.43
    \Block2-1Simulation Error
    0.94\pm 0.09
    \Block2-1Real-world Error
    \mathbf{0.91\pm 0.93}
  • Config c
    Powder Weighing Error (mg) Sand
    0.76\pm 0.58
    Powder Weighing Error (mg) Sugar
    1.0\pm 0.9
    Powder Weighing Error (mg) Salt
    1.18\pm 1.44
    Powder Weighing Error (mg) Semolina
    \mathbf{1.52\pm 1.36}
    \Block2-1Simulation Error
    \mathbf{0.85\pm 0.17}
    \Block2-1Real-world Error
    1.10\pm 1.10

We selected the top three morphologies identified by the BOHB optimiser, ranked by their weighing error in the high-fidelity environment at the maximum budget b_{\mathrm{max}}. For these three designs, we train two additional random seeds, bringing the total to four per configuration. We report the simulated error as the average across four seeds and deployed the best-performing policy from simulation for each morphology to the physical robot. Fig. 5 illustrates the three designs in simulation alongside their 3D-printed counterparts, with their morphological configurations. Table II compares the simulated error against the average real-world error. We report the real-world error only for in-distribution materials, as the findings in Section IV-B demonstrated that zero-shot transfer for OOD, highly cohesive materials that are poorly modelled in our simulation framework is unreliable.

All three configurations are predicted by the simulator to outperform the best tool of Section IV-B. However only Config. b and Config. c achieve lower real-world errors. Config. b demonstrates an overall improvement of 0.26\text{\,}\mathrm{m}\mathrm{g} across the in-distribution powders, exhibiting better performance on sand, salt, and semolina, while Config. c achieves the lowest absolute error on semolina. Although the optimiser predicted Config. c to be the best design overall, Config. b outperformed it in physical trials and exceeded its own simulated prediction. The sim-to-real prediction gaps for these configurations nonetheless remain small (0.25\text{\,}\mathrm{m}\mathrm{g} for Config. c and 0.03\text{\,}\mathrm{m}\mathrm{g} for Config. b). We attribute this rank inversion to the reality gap between the simulator and physical setup, which includes unmodelled dynamics and morphological variations introduced by 3D printing.

IV-D Accelerated Co-design via Morphological Similarity

We evaluated the similarity-based strategy introduced in Section III-E, to determine whether it accelerates the optimisation process without degrading the real-world performance of the discovered tools. To compute the spike similarity score S_{\mathrm{spikes}}, we defined the spike arc ratios as c_{i}\in\{1/32,1/8\}. The overall similarity S(\xi_{1},\xi_{2}) was then computed using Equation 7, with the coefficients \alpha, \beta, and \gamma tuned to 0.5, 0.3, and 0.2, respectively. To examine whether the metric captures task-relevant structure, we aggregated the morphologies from the initial grid search (Section IV-B) and the maximum budget evaluations from the BOHB optimisation (Section IV-C), and projected them into two dimensions using kernel principal component analysis (KPCA) [17], with S(\xi_{1},\xi_{2}) supplied as a precomputed kernel instead of a distance in the raw parameter space. Fig. 6 illustrates the resulting distribution of the dimensionality-reduced morphologies, mapped alongside their corresponding simulated weighing errors. Morphologies that cluster in the embedding, which appear structurally similar upon visual inspection, exhibit comparable task performance, indicating that the proposed similarity is correlated with dispensing behaviour. This supports using this metric to warm-start and accelerate policy training.

Fig. 6: Each point is two-dimensional representation of a spoon configuration. Spoon designs that are located closer together in the 2D space have higher morphological similarity. The colour of the points represents their task error in simulation.

We repeated the search from Section IV-C using identical hyperparameters, with the similarity-based logic from Algorithm 1. The similarity threshold \sigma and warmup fraction \phi were set to 0.85 and 0.33, respectively. Figure 7 compares unique configurations explored within the fixed budget of 200,000 episodes by the BOHB (Section IV-C) and its similarity-augmented variant. Both methods exhibit similar initial exploration rates, but the similarity-augmented BOHB accelerates as the number of sampled configurations increase. Ultimately, the similarity approach explores 136 unique configurations, compared to 106 for the standard baseline. As the cached policy pool grows, the warm-start hit-rate increases, reducing the average training cost per configuration.

Following the evaluation protocol of Section IV-C, we select the top three configurations, train two additional random seeds each and deploy the best-performing policy on the physical robot. Table III reports the sim-to-real transfer with the predicted simulation errors. All three morphologies outperformed the grid-search tool of Section IV-B in simulated and physical experiments, with predicted errors comparable to the standard BOHB run. Notably, Config. B is morphologically identical to Config. b from Section IV-C. Upon sim-to-real transfer, Config. B now shows a minor improvement of 0.10\text{\,}\mathrm{m}\mathrm{g} weighing error, emerging as the best-performing physical morphology in this experiment as well. While this variance in physical performance is likely attributable to policy seeding, the convergence on the same geometry under both settings indicates that similarity-based warm-starting accelerates morphological search without degrading the quality of the final tool.

Fig. 7: BOHB performance with and without morphology similarity filter. The plot shows total unique sampled configurations at 50000, 125000 and 200000 total episodes used.
TABLE III: Simulation and real-world deployment comparison of the best 3 tool configurations found by BOHB with similarity, reported from the high-fidelity simulation environment and for IID powders, averaged over 10 real-world runs for w_{target}=15\text{\,}\mathrm{m}\mathrm{g}.
  • Config A
    Powder Weighing Error (mg) Sand
    0.65\pm 0.54
    Powder Weighing Error (mg) Sugar
    \mathbf{0.47\pm 0.24}
    Powder Weighing Error (mg) Salt
    0.70\pm 0.36
    Powder Weighing Error (mg) Semolina
    1.82\pm 1.25
    \Block2-1Simulation Error
    0.95\pm 0.05
    \Block2-1Real-world Error
    0.91\pm 0.87
  • Config B
    Powder Weighing Error (mg) Sand
    0.62\pm 0.55
    Powder Weighing Error (mg) Sugar
    0.56\pm 0.23
    Powder Weighing Error (mg) Salt
    \mathbf{0.41\pm 0.18}
    Powder Weighing Error (mg) Semolina
    \mathbf{1.66\pm 0.92}
    \Block2-1Simulation Error
    0.92\pm 0.08
    \Block2-1Real-world Error
    \mathbf{0.81\pm 0.73}
  • Config C
    Powder Weighing Error (mg) Sand
    \mathbf{0.59\pm 0.27}
    Powder Weighing Error (mg) Sugar
    0.73\pm 0.50
    Powder Weighing Error (mg) Salt
    0.91\pm 1.33
    Powder Weighing Error (mg) Semolina
    1.97\pm 0.91
    \Block2-1Simulation Error
    \mathbf{0.86\pm 0.22}
    \Block2-1Real-world Error
    1.05\pm 0.98

IV-E Discussion

Our framework identifies tools that outperform human-centric standard ones and manually-designed baselines. While our high-fidelity simulation predicts real-world performance for IID materials (28^{\circ} to 41^{\circ} AoR), we deployed our best spoon (Config. B) on OOD materials to evaluate generalisability (Table IV). As material dynamics drift from the training distribution, the control policy becomes less reliable as the sim-to-real gap widens. Despite this degradation on OOD powders, Config. B achieves lower overall errors than the baselines across the combined set. While our co-designed tool is specialised for its target distribution of powder dynamics, it remains a more effective all-purpose tool for robotic chemists than standard human tools.

TABLE IV: Zero-shot transfer results on OOD materials with their respective AoR, averaged over 10 runs for w_{target}=15\text{\,}\mathrm{m}\mathrm{g}.
  • Standard Tool
    IID Error
    1.93\pm 2.16
    Sodium Bicarbonate (43^{\circ})
    4.45\pm 5.57
    Pectin (46^{\circ})
    3.58\pm 3.64
    Flour (51^{\circ})
    12.91\pm 13.58
    Overall Real-world Error
    4.09\pm 6.82
  • Config. 2
    IID Error
    1.17\pm 1.13
    Sodium Bicarbonate (43^{\circ})
    3.62\pm 3.45
    Pectin (46^{\circ})
    8.21\pm 4.04
    Flour (51^{\circ})
    11.41\pm 8.72
    Overall Real-world Error
    4.04\pm 5.43
  • Config. B
    IID Error
    \mathbf{0.81\pm 0.73}
    Sodium Bicarbonate (43^{\circ})
    \mathbf{1.98\pm 1.53}
    Pectin (46^{\circ})
    \mathbf{2.71\pm 2.43}
    Flour (51^{\circ})
    \mathbf{7.69\pm 4.01}
    Overall Real-world Error
    \mathbf{2.23\pm 3.00}

V Conclusion

We presented a co-design framework for autonomous powder weighing that jointly optimises a tool’s morphology and its reinforcement learning control policy, using a geometric similarity metric to warm-start the search from structurally related candidates. Real-world experiments show the co-designed tools outperform both a standard tool and grid-search baselines across materials. The main limitation is that our simulation cannot accurately reproduce cohesive powders; future work will develop higher-fidelity granular simulation that remains tractable for search. Ultimately, extending joint morphology–control optimisation to other contact-rich laboratory tasks will improve robot–material manipulation in autonomous scientific discovery.

Acknowledgements

Gemini 3.1 Pro was used in support for code-generation and creating functions for plotting results; all code and results were reviewed by all authors.

References

  1. [1] G. Tom, S. P. Schmid, S. G. Baird, Y. Cao, K. Darvish, H. Hao, S. Lo, S. Pablo-Garcia, E. M. Rajaonson, M. Skreta, N. Yoshikawa, S. Corapi, G. D. Akkoc, F. Strieth-Kalthoff, M. Seifrid, and A. Aspuru-Guzik (2024) Self-driving laboratories for chemistry and materials science. Chemical Reviews 124 (16), pp. 9633–9732.
  2. [2] Y. Kadokawa, M. Hamaya, and K. Tanaka (2023) Learning robotic powder weighing from simulation for laboratory automation. In IEEE/RSJ IROS, Vol. . External Links: Document
  3. [3] N. Radulov, A. Wright, T. Little, A. I. Cooper, and G. Pizzuto (2026) FLIP: flowability-informed powder weighing. In IEEE ICRA,
  4. [4] Y. Nakajima, M. Hamaya, K. Tanaka, T. Hawai, F. von Drigalski, Y. Takeichi, Y. Ushiku, and K. Ono (2023) Robotic powder grinding with audio-visual feedback for laboratory automation in materials science. In IEEE/RSJ IROS, Vol. . External Links: Document
  5. [5] N. Radulov, T. Little, A. I. Cooper, and G. Pizzuto (2026) Vision-guided adaptive scooping for powder weighing in autonomous chemistry laboratories. Digital Discovery 5 (5), pp. 2120–2127. External Links: ISSN 2635-098X, Document
  6. [6] R. Schneider, D. Honerkamp, T. Welschehold, and A. Valada (2025) Task-driven co-design of mobile manipulators. IEEE Robotics and Automation Letters 10 (7), pp. 7158–7165. External Links: Document
  7. [7] C. Lin, H. Yuan, Y. Wang, X. Qiu, T. Wang, M. Guo, B. Wang, Y. Narang, D. Fox, and C. Gan (2025) RobotSmith: generative robotic tool design for acquisition of complex manipulation skills. External Links: 2506.14763
  8. [8] M. Li, R. Antonova, D. Sadigh, and J. Bohg (2023) Learning tool morphology for contact-rich manipulation tasks with differentiable simulation. In IEEE ICRA, Vol. , pp. 1859–1865. External Links: Document
  9. [9] T. Takahashi, C. C. Beltran-Hernandez, Y. Kuroda, K. Tanaka, M. Hamaya, and Y. Ushiku (2025) SCU-hand: soft conical universal robotic hand for scooping granular media from containers of various sizes. In IEEE ICRA, Vol. , pp. 5549–5555. External Links: Document
  10. [10] T. Takahashi, Y. Nakajima, C. C. Beltran-Hernandez, Y. Kuroda, K. Tanaka, M. Hamaya, K. Ono, and Y. Ushiku (2026) SCU-hand with integrated single-sheet valve: a funnel-shaped robotic hand for milligram-scale powder handling. In IEEE ICRA,
  11. [11] S. Falkner, A. Klein, and F. Hutter (2018) BOHB: robust and efficient hyperparameter optimization at scale. ICML.
  12. [12] R. S. Sutton and A. G. Barto (2018) Reinforcement learning: an introduction. A Bradford Book, Cambridge, MA, USA. External Links: ISBN 0262039249
  13. [13] Isaac Sim External Links: Link
  14. [14] S. Haddadin (2024) The franka emika robot: a standard platform in robotics research [survey]. IEEE Robotics & Automation Magazine 31 (4), pp. 136–148. External Links: Document
  15. [15] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine (2018) Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. In ICML,
  16. [16] M. Lindauer, K. Eggensperger, M. Feurer, A. Biedenkapp, D. Deng, C. Benjamins, T. Ruhkopf, R. Sass, and F. Hutter (2022) SMAC3: a versatile bayesian optimization package for hyperparameter optimization. Journal of Machine Learning Research 23 (54), pp. 1–9.
  17. [17] B. Schölkopf, A. Smola, and K. Müller (1998) Nonlinear component analysis as a kernel eigenvalue problem. Neural Computation 10 (5), pp. 1299–1319. External Links: Document