Tactile Behavior Priors for Dexterous Reorientation
TacBPM distills multi-scale sphere specialists into a reusable tactile-proprioceptive latent controller, then freezes the tactile prior so downstream policies learn compact residual latent actions instead of raw joint commands while finetuning the decoder for new geometry and achieving successful sim-to-real transfer.
TacBPM learns a contact-aware behavior prior model for dexterous reorientation. The task is difficult because policies must explore from low-level 22-DoF hand actions while maintaining contact under shape, scale, axis-command, and hardware shifts. TacBPM first trains sphere reorientation specialists across scale factors 0.3 to 1.0, corresponding to roughly 2.4-8 cm grasp diameters and nearly covering the hand's reachable aperture workspace, then distills them into a single variational latent controller for downstream reuse.
The downstream policy predicts only a residual latent code around the prior mean. This keeps exploration close to stable contact behaviors while still allowing task-specific adaptation to anisotropic objects, multi-object settings, and commanded rotation axes, leading to successful sim-to-real transfer on both real hand rotation and arm-hand hammer rollouts.
Eight sphere-scale teachers span 2.4-8 cm grasp diameters, nearly covering the hand's reachable aperture workspace while exposing different tactile force patterns and rolling or regrasping regimes.
A variational encoder, learned prior, and decoder are conditioned on joint history, previous targets, contact magnitudes, 3D contact locations, and scale cues.
Downstream policies keep the tactile prior fixed, learn compact residual latent actions, and finetune the decoder for geometry adaptation.
The table below condenses the main quantitative claims from the paper into reviewer-facing takeaways. It highlights where the tactile-conditioned prior improves transfer, rather than listing every object-specific column on the project page.
| Question | Setting | Key number | Takeaway |
|---|---|---|---|
| Generalist coverage | Seen sphere scales 0.3-1.0 | 95.62% success | One distilled prior preserves the multi-scale teacher family. |
| Extrapolation | Held-out scales and unseen objects | 60.38% success vs. 40.05% w/o tactile | Tactile conditioning improves transfer beyond the sphere teachers. |
| Downstream transfer | Complex-object AnyPose-to-AnyPose | 70.00% multi-object success with main finetuned decoder | Residual latent control remains stable when several object geometries are trained jointly. |
| Axis-conditioned reuse | Continuous commanded rotation | 342.39 deg on block_03 z-axis | The learned substrate can be reused for rotation rewards, not only target-pose matching. |
| Arm-hand generalization | Diverse Grasp-to-AnyPose tools | 96.73% avg success vs. 49.24% raw-action PPO | The same prior-centered paradigm scales to coupled grasp, lift, transport, and goal-pose reaching. |
| Sim-to-real transfer | Real hand and arm-hand hardware | Successful real-robot deployment | TacBPM preserves contact-stable behavior under hardware noise, calibration error, and object inertia. |
The gallery is organized as a reviewer-facing stress test of the same learned interface: TacBPM first distills reusable tactile contact behavior, then reuses it for harder object geometries, signed-axis commands, successful sim-to-real transfer, and finally a larger arm-hand embodiment that must acquire, lift, transport, and reorient a tool.
TacBPM starts from multi-scale sphere teachers whose 2.4-8 cm grasp diameters nearly span the hand's reachable aperture workspace. This controlled scale sweep is meant to cover the main contact regimes the hand can physically realize before the prior is reused on harder shapes.
The distilled TacBPM generalist handles sphere scales from small fingertip grasps to near-maximum hand aperture, illustrating reusable contact behavior before downstream finetuning.
With the tactile prior fixed, downstream policies only learn compact residual latent actions for complex shapes and multi-object reuse while the decoder is finetuned for geometry adaptation. The paper also reports a frozen-decoder ablation to diagnose strict prior reuse.
These rollouts use the TacBPM prior-centered residual latent interface while a downstream policy adapts to different object geometries. The prior remains fixed and the decoder is finetuned for the downstream geometry.
Residual latent policy transfers the sphere-trained contact prior to a block-shaped downstream object.
A harder block variant tests whether the latent prior remains useful under changed contact affordances.
The downstream controller reuses tactile behavior on an elongated object with different rolling dynamics.
The policy adjusts contact on a sharp-edged object rather than relying on sphere-specific rolling.
The same latent interface adapts to an irregular object while preserving stable contact behavior.
This downstream rollout tests TacBPM on a broader, container-like object whose contact geometry differs from the sphere teachers.
A downstream policy rotates five different objects in the same setting, testing whether TacBPM's contact prior can support joint object coverage.
Continuous commanded-axis videos are grouped by domain: simulation first, followed by real-robot rollouts demonstrating successful sim-to-real transfer. The policy responds to six signed rotation commands and performs continuous transitions as the requested rotation axis changes.
These simulation rollouts follow the automatic recording schedule used by the axis-conditioned downstream evaluation: each object receives +x/-x, +y/-y, and +z/-z commands. The same residual-latent controller changes direction without resetting the grasp, illustrating continuous contact through axis switches.
These real-robot rollouts evaluate whether the downstream policy can condition on a commanded rotation axis while TacBPM supplies stable tactile contact behavior through the hand prior, showing successful sim-to-real transfer without real-robot policy finetuning.
Matched single-axis comparisons and representative failures are grouped together because they answer the same reviewer question: when contact becomes hard on hardware, does the tactile prior preserve usable rolling behavior while ablations drift, stall, drop, or leave the training distribution?
The following videos compare TacBPM, No tactile, and From scratch on isolated real-robot axis commands. The comparison highlights that tactile-conditioned residual latent control maintains usable rolling contact, while the baselines are more likely to show slow progress, drops, or out-of-distribution postures.
Tactile feedback keeps the multiface object in controllable rolling contact.
Without contact feedback, contact drift quickly becomes an object drop.
Raw-action learning produces weaker and slower progress under the same command.
The tactile prior recovers usable rolling contact on a faceted object.
The object becomes stuck, trapped between fingers instead of rotating.
The raw-action policy keeps partial contact but makes slow progress under the +y command.
TacBPM completes the recorded segment while maintaining stable contact.
The policy keeps partial contact but makes slow, side-biased progress.
The raw-action policy keeps partial contact but makes slow progress instead of stable commanded rotation.
TacBPM keeps contact and rotation stable on the small ball.
The no-tactile policy starts rotating but shows obvious off-axis, slow motion; without stopping the program or the ball getting stuck between fingers, it drifts OOD and stops moving.
The from-scratch policy leaves the contact regime and drops the ball.
Representative failures show how real contact errors accumulate into drops, stuck grasps, slow rotations, out-of-distribution hand states, and command-switching failures.
Failure summary. Across the recorded No tactile and From scratch trials, failures mostly appear as drops after contact drift, stuck configurations, slow or side-biased rotations, and out-of-distribution hand postures. A rare TacBPM failure occurs under frequent command switching when an axis change arrives during unstable contact.
Edge contact drift accumulates until the block is dropped.
The hand keeps contact but cannot escape a stuck grasp.
The policy leaves the contact regime and drops the ball.
Facet contacts push the hand into an out-of-distribution posture.
52 s: -x to +z switch during unstable contact, causing a drop. Occurs after nearly one minute of frequent command switches.
Grasp-to-AnyPose instantiates the same prior-centered pipeline on the North right arm and hand. Instead of starting from an already supported in-hand object, the robot must grasp a tool-like object, lift and transport it, and then move it toward sampled goal poses while preserving useful fingertip contact. In this setting, TacBPM reuses a rubber-hammer teacher prior with fingertip contacts and a frozen decoder, testing whether the latent interface remains useful when arm motion, gravity loading, object geometry, and hand reorientation are coupled.
This rollout stress-tests whether the arm-hand policy can maintain contact-stable grasp, lift, transport, and reorientation over a long sequence rather than a single goal. The visible counter marks 50 consecutive target goals; repeated success is important because small contact errors would otherwise accumulate into drops, OOD hand states, or failed pose reaching.
The four simulation clips show category-level generalization from the rubber-hammer teacher family to new tool geometries. Each rollout requires the policy to acquire the object, maintain contact through lift and transport, and reduce both position and orientation error toward a random goal pose.
A compact hammer tests whether the prior can preserve grasp stability while adapting to a smaller handle and head geometry.
The elongated brush changes the load distribution and contact affordances relative to the rubber-hammer teachers.
The marker geometry stresses grasp acquisition and goal-pose reaching on a thinner object with different support contacts.
The heavier mallet-like shape tests transport and reorientation under a larger head and stronger moment arm.
The real-robot clips are not meant as isolated demos of a scripted grasp; they expose the same contact-maintenance challenge under hardware effects such as compliance, calibration error, tool inertia, and changing wrist posture. The 2x2 layout highlights successful sim-to-real transfer across several rollouts rather than a single favorable take.
Initial tool acquisition: the arm places the hand while the fingers stabilize an asymmetric hammer instead of a compact object.
Lift under changing load: contact must remain usable as gravity and the hammer's moment arm perturb the grasp.
Post-lift pose adjustment: the hand reorients the hammer after support is removed, where small slips can quickly compound.
Repeated hardware validation: the same prior-centered interface supports another grasp-lift-reorientation rollout.
@article{anonymous2026tacbpm,
author = {Anonymous Authors},
title = {TacBPM: A Tactile-conditioned Behavior Prior Model for Dexterous Reorientation},
journal = {Manuscript in preparation},
year = {2026},
}