Temporal Visuo-Tactile Learning
for Dexterous Grasp Stability

LASR Lab, TU Dresden, Germany

Watch on YouTube

Overview

Vision, proprioception, and four Digit 360 tactile streams feed a temporal stability predictor. A sufficiently high stability score leads to lifting; otherwise the robot releases and regrasps.

Stable grasping with a multi-fingered hand depends on how contacts form and evolve as the fingers close around an object. We investigate how high-resolution, dynamic tactile sensing supports grasp stability prediction and model-guided grasping. Using a dataset of 10,000 grasp trials across 200 objects, we train temporal multimodal models to predict post-lift stability from pre-lift observations. On the robot, the learned predictor acts as an online stability gate: the robot lifts when confidence is high and otherwise releases its grip and attempts a regrasp.

Hardware Setup

Hardware setup: an xArm7 with a four-fingered Tilburg Hand, a Digit 360 sensor on each fingertip, and an external RealSense D435i camera on a stationary arm.

Our platform uses a 7-DoF xArm7 with a four-fingered, 16-DoF Tilburg Hand. Each fingertip carries a Digit 360 sensor that records tactile images, audio, inertial measurements, and pressure. A RealSense D435i RGB-D camera on a stationary arm provides an external view of the workspace. Visual observations, hand and arm proprioception, and fingertip tactile streams are recorded continuously throughout each grasp trial.

Temporal Multimodal Learning

Modality-specific video and sequence encoders feed tactile attention across modalities, time, and fingers, followed by cross-modal attention to predict grasp success.

The model observes a pre-lift window anchored at grasp initiation. VideoMAEv2 backbones encode external RGB and tactile image sequences, while Transformer branches encode proprioception and other tactile streams. Tactile features are integrated across modalities, time, and fingers, then fused with vision and proprioception to predict the probability of a stable grasp.

During deployment, the robot lifts only when the predicted probability exceeds a fixed threshold. Otherwise, it opens the hand, adjusts the end-effector position, and samples another grasp. This uses learned contact cues without explicit contact or force modeling.

Data Collection

Data collection proceeds through object localization, randomized reaching and grasping, lifting, automatic stability labeling, and resetting, with continuous multimodal recording.

We designed an automated data collection procedure comprising object localization, randomized reaching and grasping, lifting, stability labeling, and resetting. A trial is labeled successful when the object is lifted and held for three seconds without dropping.

One automated trial: localize, reach and grasp, lift, label, reset.
The 200 everyday objects used for data collection, including toys, tools, bottles, and household items.
The 200 objects used to collect 10,000 grasp trials.

The dataset of 10,000 grasp trials covers 200 objects with diverse shapes, sizes, and materials. The outcomes are nearly balanced: 47.1% stable and 52.9% unstable trials. Each trial records external RGB frames, proprioception, tactile images, audio, IMU, pressure, and commanded actions asynchronously throughout the full grasp attempt, together with phase start times. These recordings can be resampled over any interval into synchronized multimodal sequences; the pre-lift window used in this paper is one example. This flexibility supports broader uses, including contact dynamics modeling and multimodal representation learning.

Results

Effect of Sensing Modalities and Backbone Design

Prediction accuracy across sensing modalities and backbone designs. Temporal touch-enabled models perform best, with vision, proprioception, and touch achieving 83.55 percent mean accuracy.
Mean prediction accuracy across sensing modalities and backbones.

We evaluate stability prediction using 5-fold object-disjoint cross-validation, holding out all trials from each validation object. Touch-enabled models achieve the highest mean accuracies in the VideoMAEv2-based modality comparison. Removing touch from the vision-proprioception-touch model reduces accuracy by 3.99 percentage points, from 83.55% to 79.56%, supporting the value of local contact cues. For vision-proprioception-touch inputs, the static single-frame baseline achieves 79.79% accuracy, compared with 81.42% for ResNet-Transformer and 83.55% for VideoMAEv2, suggesting the value of expressive temporal modeling.

The following controlled input ablations examine spatial contact detail and temporal tactile information while keeping the VideoMAEv2-based architecture and training protocol fixed. Each condition is trained separately, with the same input modification applied during training and evaluation.

Effect of Tactile Spatial Resolution

Mean grasp stability prediction accuracy rises steeply from 79.40 percent at 1 by 1 tactile resolution to 82.51 percent at 28 by 28, then improves more gradually, reaching 83.55 percent at both 112 by 112 and 224 by 224. The non-tactile baseline achieves 79.56 percent.
Mean accuracy across tactile-image resolutions. The dashed line shows the non-tactile baseline.

We area-downsample tactile frames to lower resolutions and bilinearly restore them to 224 × 224 before encoding, keeping all other inputs and training settings fixed. At 1 × 1, only mean RGB values remain, yielding 79.40% accuracy, close to the non-tactile baseline. Accuracy improves sharply up to 28 × 28, then more gradually, reaching 83.55% at both 112 × 112 and 224 × 224, supporting the value of spatial contact detail.

Effect of Temporal Observations

Temporal input ablations in single-modality models. Repeating the final observation lowers vision-only accuracy from 79.66 to 72.47 percent and touch-only from 82.64 to 77.53 percent, while proprioception-only stays at about 80 percent.
Temporal input ablations in single-modality models.

Repeating a stream's final observation throughout the window preserves the final grasp state while removing its temporal variation. Applied to single-modality models, it lowers vision-only accuracy from 79.66% to 72.47% and touch-only from 82.64% to 77.53%, while proprioception-only is nearly unchanged at about 80%.

Temporal input ablations in the vision-proprioception-touch model: full sequences achieve 83.55 percent accuracy; repeating the final tactile observation lowers accuracy to 81.48 percent; jointly permuting vision and tactile-camera frames gives 81.93 percent.
Temporal input ablations in the vision-proprioception-touch model. V: vision; P: proprioception; T: touch.

In the vision-proprioception-touch model, repeating vision or proprioception has little effect, whereas repeating touch lowers accuracy by 2.07 percentage points, from 83.55% to 81.48%.

Applying the same random permutation to the external RGB and four tactile-camera clips preserves all frames and their alignment across image streams, but disrupts their temporal order. Accuracy falls to 81.93%, a 1.62-point reduction. These results support the value of tactile dynamics and temporal structure beyond the final contact state.

Real-Robot Grasping Evaluation

Executed-lift success for direct lifting, a random gate, a non-tactile gate, and a tactile gate on 20 unseen objects. The tactile gate achieves 82 percent overall success.
Success among executed lifts on 20 unseen objects, with 10 lifts per object and policy.

We deploy the learned predictors on the real robot as online stability gates that decide whether to lift or regrasp. On 20 unseen objects, the tactile stability gate achieves 82.0% success among executed lifts (164/200), improving on the non-tactile gate by 10.5 percentage points.

Baseball
Light bulb
Matcha can

The videos above show example grasping sequences leading to successful lifts with the tactile stability gate. Qualitatively, contact cues from multiple fingers appeared to help identify partial or off-center grasps, particularly on visually ambiguous objects such as the light bulb.

The paper (coming soon on arXiv) provides more detailed experimental results and additional analyses not shown here, including the effects of dataset size, tactile submodalities, and finger-specific inputs, as well as the relationship between fingertip contacts and grasp success.

Summary

  • We collect a dataset of 10,000 multimodal dexterous grasp trials across 200 everyday objects using four Digit 360 fingertips. Vision, proprioception, touch, and commanded actions are continuously recorded throughout each grasp, with phase annotations and stability labels supporting reuse beyond grasp stability prediction.
  • We systematically study what drives grasp stability prediction from pre-lift sequences of vision, proprioception, and touch. Modality and encoding comparisons, together with controlled analyses of tactile spatial resolution and temporal observations, show that high-resolution, dynamic touch provides a particularly strong stability signal.
  • We deploy the learned predictor on the real robot to guide online decisions to lift or regrasp. On 20 unseen objects, the tactile stability gate achieves 82.0% success among executed lifts, exceeding the non-tactile gate by 10.5 percentage points.

BibTeX

@unpublished{nakahara2026temporal,
  title   = {Temporal Visuo-Tactile Learning for Dexterous Grasp Stability},
  author  = {Nakahara, Ken and Buvailik, Aleksei and
             Kotov, Prokhor and Calandra, Roberto},
  note    = {Manuscript; arXiv preprint forthcoming},
  year    = {2026}
}

Acknowledgments

This work was partly supported by the project GeniusRobot (01IS24083) funded by the Federal Ministry of Research, Technology and Space (BMFTR), by BMFTR in the DAAD project 57616814 (SECAI, School of Embedded Composite AI), by BMFTR under the Robotics Institute Germany (RIG), and by the German Research Foundation (DFG) under the Cluster of Excellence CARE: Climate-Neutral And Resource-Efficient Construction (EXC 3115, project number 533767731).

We thank the ZIH at TU Dresden for providing computing resources.