ICRA 2027 submission/ Project page

Learning Visual Contact Representations
from Triboelectric Signals
for Tactile-Free Robot Manipulation

Anonymous authors

Learning from touch during training.
Acting with vision, language, and robot state at deployment.

MOBILE ALOHA · RECORDED TRIAL

From contact to action

Touch teaches.
Vision acts.

A robot probes an artificial banana, then places it in the gray box according to the sorting instruction.

Task instruction

“Probe the fruit. If it is artificial, put it in the gray box.”

Deployment inputsWrist RGB · Instruction · Robot state

Selected successful trial · 1× playback · No live TENG input

+9.43pp

Multiple-choice QA accuracy

Paired vs. No TENG, SFT
66.25%

Robot sorting macro success

Four settings · 20 trials per setting
0

Tactile inputs at deployment

TENG supervision is used in training

01 / Interactive hardware

Look closer.
Explore the contact interface.

Turn the gripper in your hands. Separate the illustrated components and inspect the fingertip sensing layers.

Interactive model ready to load

PHOTO-GUIDED ILLUSTRATION

Drag to rotate · Scroll or pinch to zoom · Arrow keys to orbit

Loading interactive hardware…

A structural illustration, not dimensionally validated CAD. Added sensors, electronics, and wiring are approximated from photographs; exploded spacing is enlarged for visibility. Base geometry: AgileX Pika (MIT). Viewer: model-viewer and Draco (Apache 2.0).

View the hardware figure and recording specifications

Data collection

A handheld interface
for recording contact.

A modified Pika gripper records wrist RGB and two TENG channels during object interactions. Temporal annotations connect observations to contact transitions, with gripper measurements providing additional context.

500 HzTENG sampling
2 channelsPaired fingertip signals
Modified Pika gripper with fingertip sensing and wrist vision for collecting paired contact recordings.

Can recorded touch teach a visual policy?

Robot manipulation depends on understanding how objects respond to contact. We investigate whether triboelectric nanogenerator (TENG) recordings provide useful supervision beyond questions and answers about those interactions.

A modified handheld gripper records paired wrist RGB and two-channel TENG signals. Visual question answering and signal alignment update shared EO-1 adapters, followed by visual QA reinforcement learning and action learning from separate Mobile ALOHA demonstrations. The signal branch is used only during representation training; deployment receives wrist RGB, instructions, and robot state.

Framework overview: paired handheld visual and tactile data, EO-1 post-training, and tactile-free robot sorting.
From handheld interaction recordings to visual QA and robot sorting. Paired TENG supervision enters during training.

02 / Method

A contact-informed route
from perception to action.

One retained set of EO-1 adapters connects visual QA, RGB–TENG alignment, and robot action learning.

Three-stage architecture: shared-adapter QA SFT and TENG InfoNCE alignment, visual QA GRPO, and flow-matching action fine-tuning.
The TENG encoder and alignment projections are discarded after SFT. The adapted EO-1 model is retained for the next stages.
01 · REPRESENTATION LEARNING

SFT + TENG alignment

Eight RGB frames and a paired two-second TENG window are mapped to 128-dimensional embeddings. Visual QA and symmetric InfoNCE update shared LoRA adapters through separate forward passes.

02 · ANSWER REFINEMENT

Visual QA GRPO

Groups of sampled answers receive verifiable correctness rewards for multiple-choice and binary questions. This stage uses RGB and questions, with no signal branch.

03 · POLICY LEARNING

Robot action fine-tuning

Separate Mobile ALOHA demonstrations train a flow-matching action head and the retained adapters. The policy predicts action chunks and replans from updated observations.

Train with paired contact experience.Deploy with wrist RGB + instruction + robot state.

03 / VTLA-QA

Physical interactions.
Visually grounded questions.

Paired recordings connect object attributes, contact events, and instruction-conditioned decisions.

50Physical objects
~3,000Recorded episodes
100,000Expanded QA instances
12Question families

Approximately 20,000 source questions are expanded through paraphrasing. All variants from the same source episode remain in one partition; expanded questions are not additional physical recordings.

Cup interaction with wrist RGB frames, Ch0 and Ch1 TENG traces, contact annotations, and example questions about width, contact state, and timing.
A recorded cup interaction and representative questions. Each question is tied to its designated visual observation; the contact-state example uses a pre-contact window.
T1 · Object attributesT2 · Observability & uncertaintyT3 · Interaction understandingT4 · Property comparisonT5 · Outcome predictionT6 · Action recommendationT7 · Temporal reasoningT8 · Visual disambiguationT9 · Sorting decisionsT10 · Anomaly & riskT11 · Counterfactual reasoningT12 · Description & transfer
Dataset composition and evaluation split

79,998 training, 10,004 validation, and 9,998 test QA instances. Across all partitions: 50,662 multiple-choice, 29,975 binary open-answer, and 19,363 free-form open-answer instances. All three formats supply SFT targets; only multiple-choice and binary items enter GRPO and the primary accuracy metrics.

45,376 questions use pre-contact observations and 54,624 use interaction observations. The split is by source episode, not a claim that all test objects are unseen.

Dataset statistics from the manuscript: question families and formats, evidence timing, and question lengths.

Research film

The complete pipeline, in motion.

Hardware, data collection, annotation, training, and a recorded robot demonstration. Illustrative hardware rendering includes AgileX Pika assets under the MIT license.

04 / Evaluation

Does tactile supervision
carry over to vision?

Comparisons separate the contribution of paired signals from QA training and later reinforcement learning.

Macro success rate (%) · Higher is better · Four equally weighted sorting settings

+11.25 percentage points over the No TENG route. All policies undergo action training. The paired and no-signal EO-1 conditions share the action-training recipe; external VLA baselines differ in pretraining and adaptation.

Successful placements / trials
Model / routeFruitBaseballCupsSpongesMacro SR
π₀.₅10 / 2010 / 209 / 208 / 2046.25%
OpenVLA-OFT11 / 2010 / 2010 / 207 / 2047.50%
Base EO-19 / 209 / 208 / 206 / 2040.00%
SFT + GRPO · No TENG13 / 2012 / 209 / 2010 / 2055.00%
SFT + GRPO · Paired TENG15 / 2013 / 2013 / 2012 / 2066.25%
Evaluation scope and limitations

Values are reproduced from Tables II and III of the linked manuscript. Trained QA conditions use one seed (29). QA accuracy counts correctly parsed answers; invalid outputs count as errors. These are ordinary accuracies, not balanced accuracies.

Each robot policy receives 20 trials per setting, 10 per class. Success requires placement in the correct box within 30 seconds. Macro SR averages the four settings. Fruit trials reuse five held-out instances. Success combines destination choice and execution, and class-level effects vary. These experiments do not establish broad task generalization or identify the specific visual contact cues the model uses.