TacBPM: A Tactile-conditioned Behavior Prior Model for Dexterous Reorientation

Anonymous Submission

Anonymous Authors
Anonymous Institution(s)

Tactile Behavior Priors for Dexterous Reorientation

TacBPM distills multi-scale sphere specialists into a reusable tactile-proprioceptive latent controller, then freezes the tactile prior so downstream policies learn compact residual latent actions instead of raw joint commands while finetuning the decoder for new geometry and achieving successful sim-to-real transfer.

Overview timeline: 0-2.5 s TacBPM motivation, 2.5-6.5 s multi-scale tactile prior, 6.5-12.5 s residual latent downstream reuse, 12.5-17.2 s multi-object rotation, 17.2-29.2 s continuous axis control in simulation, 29.2-37.2 s successful sim-to-real axis control, 37.2-42.2 s single-axis multiface comparisons, 42.2-46.2 s challenging +z tennis comparisons, 46.2-54.2 s arm-hand consecutive-goal transfer, 54.2-59.2 s arm-hand held-out tool transfer, 59.2-64.2 s successful real arm-hand transfer, 64.2-66.7 s takeaway. For compact presentation, some clips are played back faster than real time.

TL;DR

TacBPM learns a contact-aware behavior prior model for dexterous reorientation. The task is difficult because policies must explore from low-level 22-DoF hand actions while maintaining contact under shape, scale, axis-command, and hardware shifts. TacBPM first trains sphere reorientation specialists across scale factors 0.3 to 1.0, corresponding to roughly 2.4-8 cm grasp diameters and nearly covering the hand's reachable aperture workspace, then distills them into a single variational latent controller for downstream reuse.

The downstream policy predicts only a residual latent code around the prior mean. This keeps exploration close to stable contact behaviors while still allowing task-specific adaptation to anisotropic objects, multi-object settings, and commanded rotation axes, leading to successful sim-to-real transfer on both real hand rotation and arm-hand hammer rollouts.

Stage 1 Specialist behavior coverage

Eight sphere-scale teachers span 2.4-8 cm grasp diameters, nearly covering the hand's reachable aperture workspace while exposing different tactile force patterns and rolling or regrasping regimes.

Stage 2 Tactile latent controller

A variational encoder, learned prior, and decoder are conditioned on joint history, previous targets, contact magnitudes, 3D contact locations, and scale cues.

Stage 3 Residual latent transfer

Downstream policies keep the tactile prior fixed, learn compact residual latent actions, and finetune the decoder for geometry adaptation.

Representative Results

The table below condenses the main quantitative claims from the paper into reviewer-facing takeaways. It highlights where the tactile-conditioned prior improves transfer, rather than listing every object-specific column on the project page.

Question Setting Key number Takeaway
Generalist coverage Seen sphere scales 0.3-1.0 95.62% success One distilled prior preserves the multi-scale teacher family.
Extrapolation Held-out scales and unseen objects 60.38% success vs. 40.05% w/o tactile Tactile conditioning improves transfer beyond the sphere teachers.
Downstream transfer Complex-object AnyPose-to-AnyPose 70.00% multi-object success with main finetuned decoder Residual latent control remains stable when several object geometries are trained jointly.
Axis-conditioned reuse Continuous commanded rotation 342.39 deg on block_03 z-axis The learned substrate can be reused for rotation rewards, not only target-pose matching.
Arm-hand generalization Diverse Grasp-to-AnyPose tools 96.73% avg success vs. 49.24% raw-action PPO The same prior-centered paradigm scales to coupled grasp, lift, transport, and goal-pose reaching.
Sim-to-real transfer Real hand and arm-hand hardware Successful real-robot deployment TacBPM preserves contact-stable behavior under hardware noise, calibration error, and object inertia.

BibTeX

@article{anonymous2026tacbpm,
  author  = {Anonymous Authors},
  title   = {TacBPM: A Tactile-conditioned Behavior Prior Model for Dexterous Reorientation},
  journal = {Manuscript in preparation},
  year    = {2026},
}