Opinion

Visual-Tactile Fusion Method Estimates Grasping Force on Flexible Objects

Archive editionAmara DialloSep 15, 2024· 11,341 views

A transformer-based approach fuses visual and tactile data to estimate safe grasping forces for deformable objects, improving accuracy by 10.19%.

In context

In 2024, robotic manipulation of flexible objects such as paper cups, toys, and produce remained challenging due to deformation and uncertain physical properties. Traditional force-control methods often relied on rigid assumptions, while emerging learning-based approaches struggled with generalization. This research addressed the need for accurate, safe grasping force estimation in dynamic operations.

What was reported

Researchers proposed a visual-tactile fusion method called MultiSense Local-Enhanced Transformer (MSLET) to estimate grasping forces on flexible objects. The model learns low-dimensional features from each sensor modality, infers physical characteristics like friction and center of mass, and integrates these to predict grasp outcomes and optimal force.

Two key modules were introduced: a Feature-to-Patch module that extracts shallow features from visual and tactile images to capture edge information, and a Local-Enhanced module that applies depth-wise separable convolution to enhance local feature processing, improving spatial correlation between adjacent tokens.

In comparative experiments, MSLET improved grasping accuracy by 10.19% over state-of-the-art models while maintaining operational efficiency, demonstrating its effectiveness in estimating appropriate forces for flexible objects.

Why it mattered

This work advanced robotic handling of deformable objects by combining complementary sensory modalities, offering a pathway to more reliable and damage-free automation in industries such as food processing and logistics, where gentle yet secure grasping is critical.

Source: 《机器人》期刊 (robot.sia.cn) · Published 2024-09-15 · “一种视/触觉融合的柔性物体抓取力估计方法”