Membrane-Coupled Delta Array가 43.3mm 액추에이터 간격보다 작은 15mm 물체까지 조작
연구팀은 3자유도 델타 로봇 64대(8×8)의 끝단을 신축성 원단으로 연결해 192자유도 연속 표면을 만들었다. 이 표면은 중심 간격 43.3mm보다 작은 15mm 물체부터 90mm 물체까지 하나의 플랫폼에서 조작했고, 19개 델타 주변에 저차 DCT 계수로 명령하는 정책이 전체 배열 명령보다 배치 오차를 절반으로 줄였다. 실물 전이 실험에서는 적응 없이 76% 성공률을 보였다. 이 논문은 arXiv 프리프린트다.
액추에이터 중심 간격이 물체 크기의 하한으로 작용하는 분산 조작 시스템에서, 끝단을 막으로 연결한 연속 표면이 간격보다 작은 물체까지 조작할 수 있는가?
기존 분산 조작 시스템은 물체가 동시에 여러 액추에이터에 접촉해야 안정되므로 액추에이터 중심 간격이 조작 가능한 물체 크기의 하한이 된다. 간격보다 작은 물체는 액추에이터 사이로 빠져 접촉을 유지하지 못하고, 배열을 조밀하게 만들어도 하한을 낮출 뿐 제거하지 못한다.
기존 이산 액추에이터 배열은 대개 액추에이터 간격보다 훨씬 큰 물체를 대상으로 하며, 물체 크기가 간격에 가까워지면 이산 접촉 효과가 지배적이 되어 연속 벡터장 모델이 깨진다. 연속 표면 방식은 팽팽한 인터페이스와 느슨한 인터페이스로 나뉘지만 설계 시 재질과 장착으로 고정되고, 프리즘 구동만 쓰는 경우 면내 변형과 면외 변위를 독립적으로 명령할 수 없다.
64개 델타 로봇의 끝단을 80/20 나일론 스판덱스 원단(두께 0.4mm, 최대 변형률 70%)으로 묶어 192자유도 연속 표면을 구성했다. 조작을 표면 위 변위장으로 재구성해 준정적 형상 필드 12개(강체 모드와 막 변형, 굽힘 곡률)와 주기적 진행파를 사용한다. 셀과 링 단위 로컬 프리미티브를 A* 플래너로 연결해 간격보다 작은 물체와 큰 물체를 동시에 라우팅하고, 강화학습에서는 19개 델타 주변(R2)의 저차 이산 코사인 변환 계수 9개로 정책을 훈련했다. MuJoCo 시뮬레이션에서 막 스팬 제약과 보상, PPO, 관측 잡음과 캘리브레이션 편차를 넣었다.
연결 작업공간은 최대 변형률 0.7에서 인접 델타 자세 쌍의 54.4%가 허용되고, 최대 높이 차 59.1mm, 국소 기울기 0.94rad였다. 15mm 큐브 라우팅 30회에서 6개 고정 경로 25mm에서 185mm를 모두 목표 셀에 도달했고, 순속도는 7.6±2.9mm/s, 전이당 진행은 28.4±10.7mm였다. 강화학습 평가 392개 홀드아웃 에피소드에서 R2 19개 델타 xyz 정책은 성공률 95.1%, 중앙값 오차 3.5mm, 90분위 오차 8.4mm를 기록해 전체 배열 명령(93.5%, 6.5mm, 14.2mm)보다 중앙값과 90분위 오차가 약 절반이었다. R2 z 정책은 성공률 95.2%였고 R1 7개 델타는 83.9%에 그쳤다. 실물 전이에서는 30mm에서 90mm EGAD 물체 5개, 10개 시작 및 목표 쌍 총 50회에서 적응 없이 76% 성공률을 보였고 같은 조건 시뮬레이션은 100%였다. 하드웨어는 15mm에서 90mm 물체를 조작해 43.3mm 간격 양쪽을 아우르는 6배 범위를 보였다.
저자들은 실물에서 미세 배치 성능이 시뮬레이션보다 떨어지며, 이는 MuJoCo가 충돌을 볼록 껍질로 해석해 오목한 물체 특징에 끝단이 걸리는 현상이 훈련에 없었기 때문이라고 밝혔다. 원단은 비선형, 이방성이고 TPU 링크 컴플라이언스, 모터 백래시, 경계 하중 불균일이 이상 표면과의 편차를 만든다. 변형률 한계와 이축 변형 때문에 실제 작업공간은 모델보다 작을 수 있고, 끝단 부착부의 유한 면적 때문에 끝단 위 원단이 변형되지 않아 작은 물체가 끝단 쪽으로 몰릴 수 있다. 하드웨어 평가는 5개 물체 50회로 제한됐고 코드 공개 여부는 논문에 보고되지 않았다.
이 플랫폼은 소형 비정형 부품 정렬, 전자제품 조립, 물류 센터의 다품종 소량 이송 등에 적용 가능성이 있다. 다만 현재는 실험실 수준이고 원단 내구성, 접촉 상태 추정, 대형 배열의 비용, 고차 DCT 제어가 확보되어야 실용화할 수 있다. 산업 적용은 1년에서 3년 뒤에나 가능할 것으로 보인다.
논문 전문
Making Waves: A Membrane-Coupled Delta Array for Manipulating Objects Below the Actuator Spacing
CC BY 4.0 라이선스로 공개된 논문입니다. 출처를 밝혀 전재하며, 원문은 arXiv:2609.39652(PDF)에서 볼 수 있습니다.
Abstract
Distributed manipulator systems manipulate objects through the coordinated motion of many actuators. However, an object must be supported by several actuators at once, so the centre-to-centre actuator spacing imposes a hard lower bound on manipulable object size. We remove this bound by coupling the end-effectors of an 8\text{\times}8 array of three degrees-of-freedom delta robots with a stretchable fabric, turning 64 discrete contacts into a continuous surface capable of manipulating objects smaller than the actuator spacing. Viewing the array as a displacement field over that surface, we investigate local quasi-static and cyclic fields as manipulation primitives. These primitives can be applied globally across the array or locally confined around each tracked object to independently manipulate several objects in parallel. We then train a policy acting on low-order discrete cosine transform coefficients: at equal action dimension, commanding a nineteen-delta neighbourhood halves the placement error of commanding the whole array. The policy transfers to hardware without adaptation at 76\text{\,}\mathrm{\%} success. The platform manipulates objects from 15\text{\,}\mathrm{mm}90\text{\,}\mathrm{mm}, a six-fold range spanning both sides of the actuator spacing, 43.3\text{\,}\mathrm{mm}, on a single surface.
Index Terms:
Surface Manipulation, Multi-Robot Systems, Distributed Manipulator Systems, Soft Robot Applications
I Introduction
DMS manipulate objects not with a single end-effector but through the coordinated motion of many actuators arranged in an array. This motion is stabilised through many points of contact occurring simultaneously. Such systems have shown their utility manipulating a range of object morphologies, including those that may prove challenging for traditional gripper based approaches. Their high bandwidth offers the potential for parallel manipulation capabilities, and actuator redundancy gives them inherent robustness to failure[1].
However, a limitation shared by discrete actuator arrays is that the object must rest on several actuators simultaneously; at least three non-collinear supporting actuators [2]. It follows that the C2C (C2C) between actuators sets a lower bound on object size: anything smaller falls between the actuators. The range of manipulable objects is therefore bounded from below by the hardware; a denser array can only lower this floor, not remove it. Extending distributed manipulation to very small objects without a proportionally dense array requires a way to provide continuous contact, not only at discrete points.
In this paper, we remove the constraint imposed by the C2C by interconnecting the end-effectors of an actuator array with a stretchable material (Fig. 1). The end-effector positions control the shape of the continuous surface formed, so objects smaller than the C2C, can be manipulated by utilizing the local slope, curvature, and strain, here demonstrated down to 15\text{\,}\mathrm{mm} on an array with a 43.3\text{\,}\mathrm{mm} C2C. Larger objects can still be driven by the coordinated motion of the actuators beneath them, with the membrane mediating contact in-between. A single platform therefore spans both regimes. As the actuators now control a continuous surface, it becomes natural to reframe actuation as a field over the surface: quasi-static fields deform the surface locally, moving an object either under gravity or ejection from the strained material; cyclic fields impart velocity directly by producing travelling waves and gaits.
8\times 8 array of 3-DoF delta robots whose end-effectors are coupled by a stretchable fabric, turning 64 discrete contacts into a continuous surface. Objects both larger and smaller than the inter-actuator spacing are manipulated by shaping the surface.We demonstrate these ideas on an 8\times 8 array of 3- DoF (DoF) delta robots, 192 DoF in total, coupled together by an elastic membrane. Here, we make the following contributions:
- Characterisation of the connected workspace of the coupled array including the strain induced in the material as the deltas move through it.
- Quasi-static fields that use the first order in-plane and second order out-of-plane linearised surface kinematics to translate objects, and cyclic fields producing travelling waves and gaits for object manipulation.
- Closed-loop manipulation across object scales: objects smaller than the actuator C2C are routed via local deformation of the surface with dynamic re-planning and collision avoidance, orchestrated by a high-level planner. Larger objects are translated and rotated independently by localised cyclic fields.
- A trained policy acting on low-order coefficients of the discrete cosine transform of the delta positions. The model translates objects spanning a range from
30\text{\,}\mathrm{mm}90\text{\,}\mathrm{mm}in characteristic length and transfers from simulation to hardware without adaptation.
II Related Work
II-A Distributed Manipulation Systems
DMS, systems composed of a distributed array of actuators, have been researched for their object manipulation capabilities, including arrays of pistons [3, 4, 5], wheels [6], parallel mechanisms [7, 8], and cilia-like MEMS devices [9], differing in scale and in the DoF available in and out of plane. Most systems use a single DoF per element; however parallel mechanisms have been investigated for their utility in providing independent in-plane and out-of-plane motion [7, 8]. Despite the variety, most share a common architecture: a dense array manipulating objects substantially larger than the C2C between adjacent actuators. This is required for the continuous vector-field model to hold [1]; as object size approaches the C2C, discrete contact effects dominate [2], and the object can no longer maintain the multiple contacts needed for stable support.
II-B Continuous Surface-Based Manipulation
Rather than grasping objects, surface-based platforms manipulate objects by dynamically deforming a continuous surface on which the objects rest, using motions that leverage friction, gravity, and induced impulse to reposition one or multiple objects simultaneously [4, 10]. These systems enable non-prehensile manipulation, making them well suited to objects that are soft, fragile, or of irregular geometry. Surface-based systems are also not restricted by the typical DMS actuator C2C constraints, these platforms can transport smaller objects owing to continuous surface contact.
Existing surface-based manipulators can be categorised by the tension of the contact interface. In taut interfaces [10], the surface is shaped by elastic deformation of the material, so its geometry between contacts follows the actuator displacements. In slack interfaces [4, 8], the material hangs under gravity between sparse fixings, forming a catenary whose shape is set by the fixing points and the material’s weight; this conforms to cradle irregular objects, but the actuators can only directly control the fixings and must shape the span between them at a distance. In both cases, the choice between elastic deformation and slack is fixed during design by material and mounting. Where actuation is purely prismatic [10, 4], in-plane strain is coupled to out-of-plane displacement and cannot be commanded independently; Dacre et al. [8] provide three DoF per element via Canfield joints, but use an inextensible sheet. Our system combines three DoF per node with a stretchable membrane held at neutral strain, so local strain becomes a bidirectional command: a region can be relaxed into the slack regime to cradle objects smaller than the C2C, or stretched elastically to shape and drive them.
II-C Learned Control
ArrayBot [3] learns policies for a rigid array of pistons by acting on low-order DCT (DCT) modes. This keeps the action space tractable. We adopt the same spectral action space on a compliant, membrane-coupled surface, where the smoothness of the commanded field is also a physical constraint of the material.
III Materials and Methods
We use a system consisting of 64 delta robots, arranged in a hexagonal 8\times 8 array. The C2C between each delta’s base is 43.3\text{\,}\mathrm{mm} [7].
2\times 2 unit. The array is made up of 16 of these units. Kinematic parameters (Section IV-A) given in red.The delta mechanisms are a combination of three parallelogram linkage structures made from TPU (TPU) and PETG (PETG) (Fig. 2(a)), 3D-printed using a dual-extrusion printer (Bambu Lab H2D). TPU is used for its high compliance, allowing the delta to bend along living hinges printed in the structure. Internal PETG beams reinforce the parallel linkages against bending.
Each delta is controlled by three linear actuators (Actuonix L12-100-50-12-P), with an effective 100\text{\,}\mathrm{mm} stroke length and inbuilt potentiometers for positional feedback.
The array comprises 16 modular 2 × 2 units each with four deltas (12 actuators)[7](Fig. 2(b)). Each unit has its own microcontroller (Adafruit Feather M0), three motor drivers (DC Motor Featherwing), ADC (ADC), attached to a custom PCB.
A fabric membrane is attached to the deltas via a quick-release twist-lock mechanism, designed for ease of access and to allow the fabric to be replaced as needed. The inter connective material is an 80/20 Nylon Spandex blend with a thickness 0.4\text{\,}\mathrm{mm}, density 200\text{\,}\mathrm{g}\text{\,}\mathrm{m} and a manufacturer supplied maximum achievable strain of 70%. This forms a continuous surface with a 263\times325\text{\,}\mathrm{mm} footprint.
For object tracking, we utilise an array of four cameras (MOKOSE, Logitech) positioned to view the surface from different vantage points. The cameras capture in 1080p at 30 FPS. Points of interest are tracked by an intensity-weighted image moments tracker built from the OpenCV library [11] and localised with AprilTags.
IV Workspace Analysis
IV-A Single-module workspace
Each delta is driven by three parallel prismatic actuators [7], so the end effector stays parallel to the base and each link acts as a scalar distance constraint. Actuator i sits at azimuth \theta_{i}\in\{\frac{\pi}{2},\frac{7\pi}{6},\frac{11\pi}{6}\} about the base centroid, with an in-plane direction \hat{\mathbf{u}}_{i}; base and platform circumradii are r_{b}=$12.35\text{\,}\mathrm{mm}$ and r_{p}=$6.93\text{\,}\mathrm{mm}$, link length \ell=$55\text{\,}\mathrm{mm}$, and the tip sits h=$24\text{\,}\mathrm{mm}$ above the platform (Fig. 2). Each lateral section of the workspace \mathcal{W} is thus the intersection of three discs of radius \ell centred at (r_{b}-r_{p})\hat{\mathbf{u}}_{i}, truncated in z by the 100\text{\,}\mathrm{mm} actuator stroke. \mathcal{W} is therefore C_{3v} symmetric, with mirror planes aligned with the actuator azimuths.
IV-B Connected workspace
In the membrane-connected array, each delta can strain the membrane according to its pose relative to its neighbours. The tips sit on a regular triangular lattice of pitch d=$43.3\text{\,}\mathrm{mm}$, the C2C. This gives each interior delta six neighbours. The membrane is mounted flat at its unstretched length, \varepsilon=0; any residual pretension is negligible against commanded strains. Modelling the span between adjacent tips as a straight line and neglecting out-of-plane deflection and shear, two neighbours separated by lattice vector \mathbf{a}, \lVert\mathbf{a}\rVert=d, with tip displacements from neutral \mathbf{x}_{a},\mathbf{x}_{b}\in\mathcal{W} and relative displacement \bm{\delta}=\mathbf{x}_{b}-\mathbf{x}_{a}, span s=\lVert\mathbf{a}+\bm{\delta}\rVert and carry engineering strain \varepsilon=(s-d)/d. The connected workspace is the set of pose pairs with \varepsilon\leq\varepsilon_{\max}, equivalently s\leq s_{\max}, where s_{\max}=(1+\varepsilon_{\max})\,d.
For a workspace discretised into N points, enumerating all 6N^{2} neighbour pairs for an interior delta is unnecessary. The span depends on a pair only through \bm{\delta}, so what is needed is the multiplicity of each displacement, which on a grid with occupancy indicator A is the autocorrelation
C(\bm{\delta})=\sum_{\mathbf{x}}A(\mathbf{x})A(\mathbf{x}+\bm{\delta})=\mathcal{F}^{-1}\!\left\{\lvert\mathcal{F}\{A\}\rvert^{2}\right\}.By the convolution theorem [12], this calculation costs O(M\log M) in grid cells rather than O(N^{2}) in pairs. The six neighbour directions fall into two C_{3} orbits,\frac{\pi}{6},\frac{5\pi}{6},\frac{3\pi}{2} and \frac{\pi}{2},\frac{7\pi}{6},\frac{11\pi}{6}, which the C_{3v} symmetry of W alone does not relate. Reversing an edge, however, negates \bm{\delta} without changing the span, so C(-\bm{\delta})=C(\bm{\delta}) for any W; inversion relates the two orbits and all six neighbours share one span-length distribution. The distribution is obtained by binning \lVert\mathbf{a}+\bm{\delta}\rVert against weights C(\bm{\delta}) for a single neighbour. At the \epsilon_{\max}=0.7 strain limit of the fabric used here, 54.4% of pose pairs remain, as shown in Fig. 4.
The model assumes uniaxial loading along each edge. Within it, adjacent tips reach a maximum height differential of 59.1\text{\,}\mathrm{mm}, a local gradient of 0.94\text{\,}\mathrm{rad}.
However, pairwise admissibility does not compose: for a delta and its six neighbours, 21 DoF and twelve spans, uniform Monte Carlo over 4\times 10^{7} configurations finds 0.42\text{\,}\mathrm{\%} of joint space admissible at \varepsilon_{\max}=0.7, against 0.544^{12}=$0.07\text{\,}\mathrm{\%}$ if spans were independent; spans sharing a delta are positively correlated. Admissible configurations therefore form a small, coherent subset of joint space, which is why the array is commanded through fields (section V) and low-order modes (section VI).
z=$133\text{\,}\mathrm{mm}$, coloured by the maximum membrane strain between one delta and its six neighbours, all held at their neutral poses (black triangles).4.9\times 10^{11} pose pairs of an adjacent delta pair, from the workspace autocorrelation at 1\text{\,}\mathrm{mm} resolution. (a) Histogram in 10\text{\,}\mathrm{\%} bands. (b) Admissible fraction of the joint workspace under \varepsilon\leq\varepsilon_{\max}, with the 70\text{\,}\mathrm{\%} limit of the membrane used marked.V Surface Fields and Local Primitives
The membrane surface is set by the tip positions of the deltas, so we command it through a spatially sampled displacement field \mathbf{u}(\mathbf{p}_{i},t)=(\mathbf{v},w): the in-plane and out-of-plane displacement of delta i from neutral, as a function of its lattice position \mathbf{p}_{i} and a global clock t. Because all deltas share a common orientation, a field specified over the surface is directly a valid per-delta command.
To describe a field, we denote \hat{\mathbf{d}} as a freely specified in-plane unit direction, used to orient the field. For fields with a centre \mathbf{p}_{c} we write \mathbf{r}=\mathbf{p}_{i}-\mathbf{p}_{c}=(X,Y).
The same primitives serve both open and closed-loop control, differing only in how these parameters are set. Under open-loop control, the parameters are fixed in advance and the field is applied globally across the whole array. Under closed-loop control, the parameters can be updated at each time step from the tracked object pose, obtained by triangulation with the four-camera rig and transformed to the array frame via static AprilTags. For example \mathbf{p}_{c} could follow or lead the object, \hat{\mathbf{d}} may be used to direct toward its goal. Applied to only the subset of deltas surrounding an object, a field can be applied locally. In such a way, several objects can be manipulated independently and concurrently on a single surface.
V-A Field Primitives
Quasi-static fields. We define twelve quasi-static shape fields spanning the linearised kinematics of the surface to first order in-plane and second order out-of-plane (Table I): six rigid-body modes (three translations, two tilts, one in-plane rotation) and six deformation modes, comprising three membrane strains (dilation and two shears) and three bending curvatures (a parabola and two saddles).
By controlling the membrane’s curvature and slope, these fields can be used for manipulation: tilt drives sliding or rolling under gravity, the parabola collects or disperses rolling objects about \mathbf{p}_{c}, and the in-plane strains reposition and reorient objects resting on the surface.
- Out-of-plane (\mathbf{v}=\mathbf{0}): w=
- 자료 없음
- 자료 없음
- Heave
- h
- 1
- Tilt
- gX\;\mid\;gY
- 2
- Parabola
- \tfrac{1}{2}\kappa\,(X^{2}+Y^{2})
- 1
- Saddle
- \tfrac{1}{2}\kappa\,(X^{2}-Y^{2})\;\mid\;\kappa XY
- 2
- In-plane (w=0): \mathbf{v}=
- 자료 없음
- 자료 없음
- Translation
- c\,(1,0)\;\mid\;c\,(0,1)
- 2
- Dilation
- \beta\,(X,Y)
- 1
- Rotation
- \Omega\,(-Y,X)
- 1
- Shear
- \gamma\,(X,-Y)\;\mid\;\gamma\,(Y,X)
- 2
- Total
- 자료 없음
- 12
Cyclic fields. A cyclic field drives each tip around a closed stroke \sigma, returning to its initial pose each cycle. Transport arises from the asymmetry between the loaded and unloaded portions of the stroke. These are demonstrated on delta arrays by Patil et al. as gaits [7].
Three parameter sets specify cyclic fields. The stroke, here common to all deltas, fixing per-cycle displacement and duty ratio. The direction field \hat{\mathbf{d}}(\mathbf{r}) orients each stroke and is drawn from the same in-plane basis as Table I - uniform, radial, tangential, and shear - so the transport modes inherit the completeness of the quasi-static set. The phase field \varphi(\mathbf{r}) assigns a temporal offset, with delta i executing \mathbf{\sigma}(\omega t+\varphi(\mathbf{p}_{i})); this determines how travelling waves are sequenced across the surface.
The membrane acts as a physical reconstruction filter: the deltas sample the wave at lattice sites and the membrane interpolates between them, presenting a continuous surface rather than a sequence of discrete steps. Objects smaller than the C2C are therefore transported without loss of contact; we return to the implications of reconstruction in Section VII.
For a travelling wave, \phi(\mathbf{r})=\mathbf{k}\cdot\mathbf{r}\bmod 2\pi with wave vector \mathbf{k}, heading \theta=\operatorname{atan2}(k_{y},k_{x}) and wavelength \lambda=2\pi/\|\mathbf{k}\|. Independent actuation of each delta admits a continuously valued \phi(\mathbf{r}), so the wave may take any heading; a gait built from N discrete phases [7] is exact only at headings where the phase advance between neighbours is a multiple of 2\pi/N.
R_{1}. Purple marks the unit cell of the reciprocal lattice and \lambda_{\min} in the principal directions. Inter-cell primitives - edge (P_{0}, red), node (P_{1}, blue) and deflect (P_{2}, green) - transfer an object between cells, C_{n}, triangular regions of fabric between adjacent deltas.The C2C, d, still bounds the reproducible wavelength. Since the wave is rendered only where deltas exist, a \mathbf{k} advancing the phase by 2\pi between neighbours moves them in unison and is indistinguishable from infinite wavelength; shortening the wave further aliases to a longer wave at a different heading rather than finer structure. Over all headings the admissible \mathbf{k} therefore lie within the unit cell of the reciprocal lattice [13, 14], a hexagon (Fig. 5) whose boundary corresponds to a \pi advance along a lattice direction, adjacent deltas in antiphase; its boundary distance sets a heading-dependent minimum wavelength, \lambda_{\min}:
\lambda_{\min}^{\theta}=\sqrt{3}\,d\max_{i\in\{0,1,2\}}\bigl|\cos\bigl(\theta-\tfrac{i\pi}{3}\bigr)\bigr|\in\left[\tfrac{3}{2}d,\ \sqrt{3}\,d\right].The three DoF decouple where objects travel from how support is sequenced. A vertically actuated element encloses no orbit and can only drag an object along the propagating wave, fixing transport to the phase gradient. A closed stroke instead drives the object along \hat{\mathbf{d}}(\mathbf{r}) during contact, leaving \mathbf{k} free to serve support alone. Fig. 6 shows how this can be utilised to efficiently gait an object in a direction in which \lambda_{\min}^{\theta} would be too long to provide stable support, if \mathbf{k} and \hat{\mathbf{d}}(\mathbf{r}) were forced to be aligned.
\diameter=$70\text{\,}\mathrm{mm}$,\frac{3d}{2}<\diameter<\sqrt{3}d) across the array using an elliptical gait with \hat{\mathbf{d}}=[1,0], averaged (bold line) over five trials (faint). Red: wave vector \mathbf{k} aligned with \hat{\mathbf{d}}; blue: \mathbf{k} perpendicular, showing a greater translational velocity. Each uses the shortest wavelength available at that heading, \lambda_{\min}^{\theta}.V-B Cell and Ring Local Primitives
In their neutral positions, the tips form an equilateral triangular lattice of C2C d (Fig. 5). Treating the tips as nodes of this lattice, the triangular regions of fabric between three mutually adjacent tips are cells, whose local shape is set primarily by these three deltas. An object small enough to sit within a cell is supported by the membrane alone and is manipulated by shaping that cell. A larger object spans several cells and is supported by the membrane and the tips beneath it. A local surface enclosed by the neighbourhood of deltas centred around the tip nearest the object, which we define as a ring. Each case has its own family of primitives, and a planner routes objects through the corresponding graph: cell to cell or ring to ring.
Ring primitives. Let R_{n}(i) denote the ring-n patch about tip i: every tip within n lattice hops of i, centre included, so |R_{1}|=7 and |R_{2}|=19, spanning 2dn and truncated at the array boundary. A ring primitive applies one of the out-of-plane quasi-static modes of Table I or a cyclic field to the patch: tilt translates an object toward a neighbouring node under gravity (Fig. 7a), and the parabola collects or disperses objects about the centre. Cyclic fields confined to R_{n} translate or rotate objects too large to be moved by tilt alone (Fig. 7b).
R_{1} tilt primitive. (b) A local cyclic wave rotates and centres the assembly. (c) Final desired pose.Cell primitives. Ring primitives use only the out-of-plane kinematics of the surface. Objects smaller than the C2C sit within a cell, where the bounding deltas and their neighbours can shape the fabric using all three DoF. We define three cell primitives that transfer an object from a source cell C_{0} to a neighbour (Fig. 5, Table II): edge, P_{0}, crossing a shared edge into C_{1}; node, P_{1}, crossing over a tip into the opposite cell C_{2}; and deflect, turning across a tip into C_{3}. Each is a set of commanded offsets for the deltas around the source cell, tuned once and applied by symmetry to every cell and heading. Deflection has two mirror-image variants, and measuring \hat{\mathbf{c}} toward \Delta_{0} lets one tuning serve both. These primitives require the object to fit within a cell; the inscribed circle has diameter d/\sqrt{3}\approx 25 mm, but in practice, transfer remained reliable up to approximately 40 mm, as some out of plane overhang beyond the cell boundary was not detrimental. Utilizing the 3-DoF of each actuator allows for slack to be created to capture objects, or elastic strain to be created to eject objects. Each primitive is a single target pose. As the deltas move to it, the strained fabric ejects the object from the source cell and it descends the resulting gradient into the target; the pose is held until tracking confirms traversal.
- Delta
- \hat{a}
- \hat{c}
- \hat{z}
- Delta
- \hat{a}
- \hat{c}
- \hat{z}
- Delta
- \hat{a}
- \hat{c}
- \hat{z}
- \Delta_{3}
- -5
- 0
- +15
- \Delta_{1}
- -3
- +10
- +25
- \Delta_{1}
- -5
- 0
- +15
- \Delta_{1}
- 0
- +15
- -10
- \Delta_{3}
- -3
- -10
- +25
- \Delta_{3}
- -10
- -17.3
- -10
- \Delta_{0}
- 0
- -15
- -10
- \Delta_{0}
- 0
- 0
- -20
- \Delta_{0}
- 0
- +20
- -10
- \Delta_{2}
- -10
- 0
- -30
- \Delta_{4}
- -10
- -10
- -25
- \Delta_{5}
- +10
- -17.3
- -20
- —
- 자료 없음
- 자료 없음
- 자료 없음
- \Delta_{6}
- -10
- +10
- -25
- \Delta_{6}
- -10
- 0
- -30
V-C Routing and Multi-Object Deconfliction
For either family of local primitive, an A* planner routes objects from any ring to any ring, or any cell to any cell. The path is verified at every control step against the tracked pose, so a disturbed object is re-planned rather than lost; after three consecutive failed transfers the current edge is penalised and the object re-routed around the obstruction. We evaluated this using a single 15\text{\,}\mathrm{mm} cube over 30 runs on six fixed routes of 25-$185\text{\,}\mathrm{mm}$ spanning the array. Every run reached its goal cell, at a net speed of 7.6\pm 2.9 mm/s and 28.4\pm 10.7 mm of progress per transfer. 12.3\% stalled, leaving the cube in its source cell, but retries succeeded often enough that re-routing was triggered on only 1.5\% of transfers.
15\text{\,}\mathrm{mm} cubes using cell primitives. Objects route dynamically to colour-matched targets (⌀= 50\text{\,}\mathrm{mm}), traversing the array simultaneously without conflict.Addressability on a coupled surface is not free: a locally issued pose perturbs the fabric beyond the cell it targets, and objects in adjacent cells share bounding deltas by construction. Each object therefore raises the cost of the cells and nodes within its area of influence, so that concurrent routes stay outside one another’s influence and objects traverse the array simultaneously without conflict (Fig. 8).
VI Reinforcement Learning
The array has 192 DoF, making it hard to intuit a control policy. Hand-designed primitives make the array tractable by restricting it to configurations whose behaviour is predictable. However, different geometries and object scale need their own primitives and none improves with experience. We instead learn a policy that generalises across scale, requiring only a reward defined on the object state rather than a forward model of the surface.
Explored directly, a 192-dimensional action space produces large local deformations that destabilise the soft-body solver. We therefore act on low-order DCT modes of the tip field, following Xue et al. [3]. The action is the 3\times 3 low-order coefficients of a 2D cosine basis fitted to the bounding box of the commanded patch and evaluated at each tip, nine per commanded channel. Discarding high-frequency spatial components keeps the commanded surface smooth. We also investigate the effect of acting on just a subsection of the array, the ring neighbourhoods R_{n} (Section V-B) against controlling the array as a whole.
VI-A Training Environment
We train a policy using PPO (PPO) in MuJoCo simulator to transport a rigid object to within 20\text{\,}\mathrm{mm} of a planar goal by shaping an elastic membrane over the 64-tip array. Goals are drawn uniformly over the tip grid’s bounding box, inset by 20 % of its span per edge so every goal lies within the array. Objects start at rest at a uniformly sampled position over the same inset region. Episodes run for 100 steps (20\text{\,}\mathrm{s}) and terminate early when the object leaves the array. The goal height is pinned to the object’s resting height and the vertical axis is unscored. Each policy is trained for 1\times 10^{6} environment steps.
The membrane is a MuJoCo flex element at 21.65\text{\,}\mathrm{mm} inter-node distance, attached to actuated tips modelled as servos with first-order lag. Commanded tip positions pass through a closed-form span constraint: the commanded displacement is scaled by \rho\in[0,1] until no neighbouring pair exceeds \eta\,s_{\max}=66.25 mm, where s_{\max} is the admissible span of the connected workspace (Section IV-B) and \eta=0.9 a safety factor, so every output pose lies within the connected workspace. Physics is integrated at 2\text{\,}\mathrm{kHz} and the policy queried at 5\text{\,}\mathrm{Hz}, within the update rate of the physical system.
Objects are drawn from the EGAD dataset [15], each scaled to a square bounding box of edge S_{e} sampled log-uniformly over [30,150] mm at fixed density. The observation comprises object and goal position and their difference, the third-order DCT of the tip state, the projection deficit 1-\rho from the last command, and S_{e}. Each entry is standardised by a running mean and variance estimated online and frozen at evaluation. The reward is:
r_{t}=10\,(d_{t-1}-d_{t})+0.085\,\mathds{1}_{\mathrm{suc}}-0.01\,(1-\rho_{t})-10\,\mathds{1}_{\mathrm{lost}},with d_{t} the in-plane error [m], \mathds{1}_{\mathrm{suc}} indicating d_{t}<20 mm and \mathds{1}_{\mathrm{lost}} the object leaving the array.
Three disturbances are injected in training. Object position carries zero-mean Gaussian noise of 2\text{\,}\mathrm{mm} per axis, redrawn each step, modelling a camera estimate. Each actuator carries a per-episode calibration bias (2\text{\,}\mathrm{mm} in x,y; 0.5\text{\,}\mathrm{mm} in z) plus 0.2\text{\,}\mathrm{mm} per-step jitter, applied to the executed command but not the stored target, so the policy acts on a command the actuators do not execute exactly. Observed size carries 3 % per-episode relative error. Training is asymmetric: the critic sees the clean state and drawn biases, the actor only the corrupted stream.
VI-B Model Comparison
- R_{2}, 19 deltas
- xyz
- 95.1
- 3.5
- 8.4
- R_{2}, 19 deltas
- z
- 95.2
- 4.0
- 10.2
- R_{1}, 7 deltas
- xyz
- 83.9
- 7.1
- 61.1
- Whole array, 64 deltas
- xyz
- 93.5
- 6.5
- 14.2
Table III compares action-space extent and commanded channels over 392 held-out episodes (49 shapes, eight each). At equal action dimension, an R_{2} patch of 19 deltas matched the whole array on success rate and roughly halved its median and 90th-percentile error, indicating that transport is governed by deformation local to the object. Restricting to R_{1} held the median error but degraded the tail sharply, indicating a catastrophic failure on a subset of objects, possibly consistent with a patch too small to steer objects spanning several C2C. Commanding z alone on R_{2} gave comparable success to xyz with a larger median error, suggesting that vertical shaping dominates long-range transport while lateral tip motion refines placement near the goal.
VI-C Sim-to-Real Transfer
30\text{\,}\mathrm{mm}90\text{\,}\mathrm{mm} in bounding box edge length, S_{e} .We deploy the R_{2} xyz policy on the hardware array without fine-tuning or adaptation. Object pose is estimated by an overhead camera array and passed to the policy in place of the simulator state; all other observation channels are computed as in training. The policy is queried at 5\text{\,}\mathrm{Hz}, matching the simulated rate.
We evaluate 5 EGAD objects, Fig. 9, spanning 30\text{\,}\mathrm{mm} to 90\text{\,}\mathrm{mm}, each over the same 10 start-goal pairs drawn to cover the array, giving 50 trials. Episodes terminate on 100 time steps (20\text{\,}\mathrm{s}) as in training. The identical (object, start, goal) triples are replayed in simulation, so the transfer can be compared, Table IV.
The R_{2} xyz policy transferred to hardware without adaptation, reaching 76\text{\,}\mathrm{\%} success against 100\text{\,}\mathrm{\%} in simulation. Degradation was seen in fine placement capabilities, as reflected in the increased median final distance. Qualitatively, failures were observed to be dominated by tips catching on concave object features. This can be seen as a direct consequence of MuJoCo resolving collisions against the convex hull; a failure mode absent from training and a true sim-to-real gap. Fine placement depends most heavily on contact dynamics at the tip, which are also where the idealised model departs furthest from the real surface.
- A1, 30mm
- 80
- 100
- 18.8
- 3.5
- 14.7
- 0.8
- B2, 45mm
- 90
- 100
- 11.4
- 3.6
- 4.6
- 0.4
- C3, 60mm
- 70
- 100
- 25.3
- 3.6
- 18.3
- 0.3
- D4, 75mm
- 80
- 100
- 16.4
- 3.3
- 6.3
- 0.3
- E5, 90mm
- 60
- 100
- 24.9
- 3.3
- 15.6
- 0.6
- All
- 76
- 100
- 19.4
- 3.46
- 11.9
- 0.48
VII Discussion
Scale and Geometry. Manipulation strategy is dictated by object scale relative to the C2C. With S_{e} the characteristic object dimension, three regimes arise. Sub-cell objects (S_{e}<d/\sqrt{3}, the inscribed diameter of a cell) rest on the membrane alone within a single cell and are translated by deforming that cell only. Cell-scale objects (d/\sqrt{3}\leq S_{e}<1.5d) are supported by the tips bounding one cell and move between cells by travelling deformation. Multi-cell objects (S_{e}\geq 1.5d) bridge several cells and are supported at enough points for their pose to be commanded directly, permitting gait-like transport and rotation about the surface normal - unavailable at smaller scales, where a single contact region offers no moment arm. Across these regimes the hardware manipulated objects that span 15\text{\,}\mathrm{mm}90\text{\,}\mathrm{mm}, a six-fold range, and the learned policy covered 30\text{\,}\mathrm{mm}150\text{\,}\mathrm{mm} in simulation.
Object geometry also determines which primitive is effective, most sharply seen seen between objects that roll and those that slide. The present observation space does not expose any geometric or contact information, so a general policy cannot condition on it; contact estimation, from vision or sensing at the tips, is the natural route to supplying it.
Sampling and surface reconstruction. The array renders a continuous surface through N discrete contacts, so representable geometry is bounded by delta number and density, in analogy with multidimensional sampling [13]. The analogy is imperfect in two respects. First, the samples are mobile. Lateral delta motion makes their positions reconfigurable, so local density can be raised where finer detail is needed at the cost of density elsewhere. Second, the membrane is the reconstruction filter, and it is neither passive nor ideal. For independent deltas the map from joint space to rendered surface would be many-to-one, owing to their lateral travel. Under tension, the membrane instead assumes the shape minimising its stored elastic energy while passing through all delta tips, so joint space is in this sense fully expressed in the rendered surface and strain. The same energy minimisation makes the surface non-local: equilibrium is reached globally, so the geometry within a cell is not determined solely by its three bounding deltas; the fabric may bow out of their plane when neighbouring deltas impose strain across the cell. Intra-cell geometry is consequently not independently commandable. Complexity further arises when the membrane is not taut; slack regions admit wrinkling, which is multi-stable and hysteretic.
Resolution. The DCT order bounds the resolution of any commanded deformation: higher orders resolve finer detail. The transform is fitted to the bounding box of the ring, so a smaller neighbourhood resolves finer detail at a given order. The minimum admissible wavelength depends on heading (Equation 2), so a component is free of aliasing in every heading only if its wavelength exceeds \max\lambda_{\min}^{\theta}=\sqrt{3}\,d=$75\text{\,}\mathrm{mm}$. All policies use only the first three coefficients of the DCT so that the basis is matched across neighbourhood sizes and Table III compares neighbourhoods at equal action dimension. However, at third order, the whole array is only able to resolve waves up to a minimum of 263\text{\,}\mathrm{mm}, the x axis span of the array. Sufficient to transport objects but not to shape the surface locally near the goal, consistent with its near-equal success rate and doubled median error in Table III. Higher order coefficients may allow for finer control and a lower median error, though they come at the cost of increased dimensionality. This trade off is a promising consideration for future work.
Departures from the ideal surface. The nylon–spandex (80/20) fabric departs from the ideal membrane. Its loading curve is non-linear and its weave anisotropic. This along with manufacturing tolerances, compliance in the TPU links, and motor backlash compound this deviation. We must also consider the array boundary loads the fabric unevenly, so nominally identical configurations strain differently at the edge than at the centre. The effect of these factors combine and the result is visible deviation, such as residual ripple in a commanded plane. The associated variation in local contact geometry is absent from the idealised model and is the most likely source of the fine-placement gap observed between real and simulation in Section VI-C.
The strain limit of the fabric restricts the joint workspace (Fig. 4). For a single delta it binds only in poses with large height gradient or lateral travel, such as passing a neighbouring delta (Fig. 3); both draw on the same strain budget, so each extreme is reachable alone but not in combination. Under the biaxial strain present in the array the fabric cannot contract laterally, so the true workspace is likely smaller than modelled. Even under these limitations, the shared workspace was sufficient for every task demonstrated here. Fabrics with higher strain limits and lower anisotropy remain a promising direction for future iterations.
The ideal model also treats each attachment as a point, whereas the fabric has bonded fabric locks with finite area (Fig. 2), which must be large enough for reliable adhesion. Being fixed, over each tip the fabric cannot deform, leaving patches of uncontrolled surface gradient whose extent scales with tip size. These patches matter most for objects comparable in size to the tips. A taut, unloaded membrane is harmonic and so has no interior minima; the local low points of the surface therefore lie at the tips, and small objects are funnelled towards regions from which they cannot be driven out. Domed caps prevent the tips from acting as stable equilibria, but the out-of-plane geometry they introduce can deflect objects during manipulation.
VIII Conclusion
We have presented a manipulable surface with 192 DoF created from coupling the end-effectors of an 8\times 8 delta array with a stretchable membrane. The continuous surface removes the imposed floor of the inter-actuator spacing on object size seen by other discrete actuator arrays: objects too small to bridge adjacent tips are supported by the membrane and manipulated through its local slope, curvature, and strain, while larger objects remain driven by the tips beneath them.
Treating actuation as a field proved productive at every level of control. Quasi-static and cyclic fields gave open-loop transport at arbitrary heading; localised primitives routed several sub-pitch objects in parallel under closed-loop control; and a policy commanding low-order discrete cosine transform modes of a 19-delta neighbourhood halved the median error compared to commanding the whole array, transferring from simulation to hardware without adaptation. In each case, the effective command set proved smaller and more local than the actuator count suggests. Fine placement on hardware remains an outstanding gap and is the natural target for a contact-aware observation space.