POINT IT, STRIKE IT↗
DEFORMX 2.0 / ROBOT LEARNING

Point It,
Strike It.

Direction-Conditioned Dynamic Manipulation of Deformable Linear Objects

One Model: Any rope. Any point. Any direction.

Yi Yang1,2,3,*Xiang Fei1,*Lehong Wang1,*Zilin Dai4,†Ruogu Li1,†Jiting Cai1,†Liyao Chang1Xinyi Yang1Henry Kou1Ruijie Fu1Lu Li1Howie Choset1
1 Carnegie Mellon University2 School of Ocean and Civil Engineering, Shanghai Jiao Tong University3 Zhiyuan College, Shanghai Jiao Tong University4 Harvard University

* Equal contribution   ·   † Equal contribution

REAL-WORLD HIGHLIGHT

Balloon Boom

ORIGINAL SPEED → 8× SLOW MOTION
ONE POLICY · THREE ROPESSIM + REAL ROBOT
Control where the rope tip arrives—and its direction.

Download GIF ↓

87%

Real-world position success

Within 5 cm · three-rope mean
79%

Real-world position + direction

Within 10 cm and 10° · three-rope mean
8

Calibration swings per rope

Calibrate once, then strike new goals

Balloon & cup striking.

Real-robot strikes with a green braided rope at different target positions.

Balloon striking

Random target positions, measured by motion capture.

Cup striking

Targeted cup strikes with the same green braided rope.

Pick a target. Watch it strike.

The yellow boundary marks the reachable region. Keep the pink target inside it: drag the target or type its coordinates, then press Run — the policy plans the swing live in simulation. Or browse ten recorded swings. Drag to orbit, scroll to zoom.

Browse

Simulate. Learn. Adapt.

Generate strike data, learn a base policy, and adapt to a real rope with eight calibration swings.

SIMULATION ENGINE

DeformX 2.0

Stable GPU rope dynamics power large-scale motion search and policy evaluation.

01 / 03 · TRACE

One policy. Three ropes.

25 targets × 3 swings per rope. One calibration, then a single strike per goal.

Success within 5 cm of the target.

Base policyRECAP (ours)
Mean success across three ropes72% → 87%

Base: candidate selection in the nominal simulator. RECAP: fitted simulator with action correction. Values from Table II. 4D success additionally requires at least 1 m/s velocity along the commanded direction.

Inside the researchMethod diagrams, ablations & evaluation details

Hitting the point is only
part of the task.

A rope tip can reach the same point from very different directions. We learn to control both: a 3D target position and a 1D arrival angle.

Our framework generates striking motions in simulation, learns multiple ways to reach a goal, and adapts to a real rope through eight calibration swings. Once calibrated, the robot executes a single strike for each new goal without repeated real-world attempts.

THE 4D GOAL3D position + arrival angle

The arrival angle is measured after projecting the tip velocity onto the tangent plane of the arm-centered sphere; it is not a free 3D orientation. Goals are evaluated within the robot’s reachable workspace.

Simulate. Learn. Adapt.

From a physics simulator to a direction-conditioned strike on real hardware.

01

DeformX 2.0

Generate at scale

A GPU-accelerated Cosserat rod solver with cross-flow aerodynamic drag supports parallel rope simulation and motion search.

Over 20,000× batched throughput¹
02

TRACE + flow matching

Learn reliable swings

Starting from one manually tuned swing verified on hardware, TRACE warm-starts each new search from the stored tip path closest to the goal—not simply the closest previously solved target. A conditional flow-matching policy learns distinct swings, with the goal injected into every residual block.

98.3% data-generation coverage
03

RECAP

Adapt to a real rope

Eight swings identify rope and rig parameters. A simulation-trained correction policy proposes actions, and the fitted simulator selects the best candidate.

One calibration, many new goals
System overview: TRACE data generation and conditional flow matching, followed by RECAP calibration, action correction and simulator verification.
System overview and network architecture. Training is performed in simulation; real-world data is used for per-rope calibration.

¹ Reported throughput relative to DeformX, with 8,192 parallel environments on an RTX 4090. See Table I in the paper for the benchmark setup.

Eight swings to adapt.
One strike per new goal.

RECAP fits 11 rope and rig parameters to motion-captured rope-tip trajectories. It then generates eight corrected actions and compares them with the nominal action in the fitted simulator.

The robot executes the best candidate. The correction policy is trained in simulation; calibration is performed once per rope, and execution has no online feedback.

Hardware means across three ropes · Table II
Method3D ≤ 5 cm4D ≤ 5 cm, 10°4D ≤ 10 cm, 10°4D mean miss
Base72%28%50%10.6 cm
Learned identification baseline74%28%51%10.2 cm
RECAP87%52%79%5.6 cm

The learned identification baseline follows the paper’s Wiggle&Go-style parameter estimator with candidate selection and no action correction. All 4D success rates also require ≥ 1 m/s along the commanded direction.

Rope A workspace success and distance maps for Base, Baseline and RECAP
Rope A: RECAP improves success across the evaluated workspace. Dots are measured targets; shaded regions interpolate those measurements. The 3D and 4D panels use different distance thresholds, as labeled.

Grow reliable motion patterns
into a training dataset.

A successful tip trajectory contains useful starting points for many nearby goals.

TRACE searches from the closest stored tip path, penalizes bending and abrupt tip motion, refines neighboring solutions, and relabels successful trajectory segments with their achieved goals. This creates consistent goal–action pairs without a human demonstration corpus.

Data generation and downstream policy · Table IV
GeneratorCoverageOne-shotBest-of-64
CEM with example49.5%58.5%68.5%
TRACE without example96.8%75.5%83.9%
TRACE98.3%73.5%92.1%

Coverage measures solved generation targets. The example-free variant has higher one-shot accuracy; full TRACE provides higher coverage and stronger accuracy after candidate selection.

Learning more than one
way to hit a goal.

Distinct swings can reach the same target. A generative policy preserves these alternatives instead of averaging them into an invalid motion.

Goal and flow-time conditioning enter every residual block through zero-initialized modulation. At inference, 64 candidates are sampled with 50 Euler steps and ranked in simulation.

Conditional flow-matching ablation comparing samples and flow fields across motion patterns
Figure 5: conditioning in every block concentrates samples around distinct motion patterns. Regression may average incompatible swings.
Policy evaluation on 5,700 held-out goals
MethodOne-shotBest-of-64
Ours73.5%92.1%
Naive flow matching67.9%88.9%
Regression21.4%—
Nearest neighbor33.6%—

Success within 5 cm and 10°. One-shot evaluates one policy sample; best-of-64 selects a candidate using the simulator. Table III in the paper.

The simulation engine
behind the strikes.

A stable GPU Cosserat rod solver and quadratic cross-flow drag make large-scale search, policy training and candidate verification practical.

The reported >20,000× improvement compares batch throughput at 8,192 environments on an RTX 4090 with the original single-thread CPU DeformX. It does not mean a single environment runs 20,000× faster.

Same 1 m, 20-segment rope · Table I
EngineEnvironmentsBatch update rateSim / wall time
DeformX · CPU, 1 thread112.9 Hz0.22×
DeformX 2.0 · RTX 40901136 Hz2.3×
DeformX 2.0 · RTX 40908,19234.5 Hz4,710×

What the current system covers.

The demonstrated task is single-swing rope-tip striking within the tested workspace. It does not yet address sustained contact, obstacle interaction or multi-stage manipulation. Execution remains open-loop after calibration, and transfer may degrade outside the dynamics represented during training.

Point It, Strike It.

Position and arrival direction, learned together.