SFT + TENG alignment
Eight RGB frames and a paired two-second TENG window are mapped to 128-dimensional embeddings. Visual QA and symmetric InfoNCE update shared LoRA adapters through separate forward passes.
ICRA 2027 submission/ Project page
Learning from touch during training.
Acting with vision, language, and robot state at deployment.
From contact to action
A robot probes an artificial banana, then places it in the gray box according to the sorting instruction.
“Probe the fruit. If it is artificial, put it in the gray box.”
Selected successful trial · 1× playback · No live TENG input
Multiple-choice QA accuracy
Paired vs. No TENG, SFTRobot sorting macro success
Four settings · 20 trials per settingTactile inputs at deployment
TENG supervision is used in training01 / Interactive hardware
Turn the gripper in your hands. Separate the illustrated components and inspect the fingertip sensing layers.
Drag to rotate · Scroll or pinch to zoom · Arrow keys to orbit
Loading interactive hardware…
A structural illustration, not dimensionally validated CAD. Added sensors, electronics, and wiring are approximated from photographs; exploded spacing is enlarged for visibility. Base geometry: AgileX Pika (MIT). Viewer: model-viewer and Draco (Apache 2.0).
Data collection
A modified Pika gripper records wrist RGB and two TENG channels during object interactions. Temporal annotations connect observations to contact transitions, with gripper measurements providing additional context.

Robot manipulation depends on understanding how objects respond to contact. We investigate whether triboelectric nanogenerator (TENG) recordings provide useful supervision beyond questions and answers about those interactions.
A modified handheld gripper records paired wrist RGB and two-channel TENG signals. Visual question answering and signal alignment update shared EO-1 adapters, followed by visual QA reinforcement learning and action learning from separate Mobile ALOHA demonstrations. The signal branch is used only during representation training; deployment receives wrist RGB, instructions, and robot state.

02 / Method
One retained set of EO-1 adapters connects visual QA, RGB–TENG alignment, and robot action learning.

Eight RGB frames and a paired two-second TENG window are mapped to 128-dimensional embeddings. Visual QA and symmetric InfoNCE update shared LoRA adapters through separate forward passes.
Groups of sampled answers receive verifiable correctness rewards for multiple-choice and binary questions. This stage uses RGB and questions, with no signal branch.
Separate Mobile ALOHA demonstrations train a flow-matching action head and the retained adapters. The policy predicts action chunks and replans from updated observations.
03 / VTLA-QA
Paired recordings connect object attributes, contact events, and instruction-conditioned decisions.
Approximately 20,000 source questions are expanded through paraphrasing. All variants from the same source episode remain in one partition; expanded questions are not additional physical recordings.

79,998 training, 10,004 validation, and 9,998 test QA instances. Across all partitions: 50,662 multiple-choice, 29,975 binary open-answer, and 19,363 free-form open-answer instances. All three formats supply SFT targets; only multiple-choice and binary items enter GRPO and the primary accuracy metrics.
45,376 questions use pre-contact observations and 54,624 use interaction observations. The split is by source episode, not a claim that all test objects are unseen.

Research film
Hardware, data collection, annotation, training, and a recorded robot demonstration. Illustrative hardware rendering includes AgileX Pika assets under the MIT license.
04 / Evaluation
Comparisons separate the contribution of paired signals from QA training and later reinforcement learning.
Macro success rate (%) · Higher is better · Four equally weighted sorting settings
+11.25 percentage points over the No TENG route. All policies undergo action training. The paired and no-signal EO-1 conditions share the action-training recipe; external VLA baselines differ in pretraining and adaptation.
| Model / route | Fruit | Baseball | Cups | Sponges | Macro SR |
|---|---|---|---|---|---|
| π₀.₅ | 10 / 20 | 10 / 20 | 9 / 20 | 8 / 20 | 46.25% |
| OpenVLA-OFT | 11 / 20 | 10 / 20 | 10 / 20 | 7 / 20 | 47.50% |
| Base EO-1 | 9 / 20 | 9 / 20 | 8 / 20 | 6 / 20 | 40.00% |
| SFT + GRPO · No TENG | 13 / 20 | 12 / 20 | 9 / 20 | 10 / 20 | 55.00% |
| SFT + GRPO · Paired TENG | 15 / 20 | 13 / 20 | 13 / 20 | 12 / 20 | 66.25% |
Multiple-choice QA accuracy (%) · Same SFT recipe · 4,999 test questions
+9.43 percentage points with paired supervision. The larger difference occurs in interaction windows: 52.70% to 69.91%. Shuffled signals perform below the QA-only baseline, supporting the value of recorded correspondence under this training setup.
| Alignment | MC overall | MC pre-contact | MC interaction | Binary overall |
|---|---|---|---|---|
| No TENG | 60.10 | 66.50 | 52.70 | 55.50 |
| Shuffled TENG | 54.80 | 65.50 | 42.40 | 46.80 |
| Paired TENG | 69.53 | 69.21 | 69.91 | 72.56 |
Multiple-choice QA accuracy (%) · Held-out QA partition · 4,999 test questions
GRPO adds 4.38 percentage points after paired-TENG SFT. These are comparisons of complete training routes. Alignment is applied only during SFT; the same tactile-free answer interface is used for evaluation.
| Route | TENG alignment | MC overall | Binary overall |
|---|---|---|---|
| Base (zero-shot) | No TENG | 40.93 | 41.59 |
| GRPO-only | No TENG | 56.41 | 52.96 |
| SFT | No TENG | 60.10 | 55.50 |
| SFT | Shuffled TENG | 54.80 | 46.80 |
| SFT | Paired TENG | 69.53 | 72.56 |
| SFT + GRPO | No TENG | 64.70 | 57.60 |
| SFT + GRPO | Paired TENG | 73.91 | 74.48 |
Values are reproduced from Tables II and III of the linked manuscript. Trained QA conditions use one seed (29). QA accuracy counts correctly parsed answers; invalid outputs count as errors. These are ordinary accuracies, not balanced accuracies.
Each robot policy receives 20 trials per setting, 10 per class. Success requires placement in the correct box within 30 seconds. Macro SR averages the four settings. Fruit trials reuse five held-out instances. Success combines destination choice and execution, and class-level effects vary. These experiments do not establish broad task generalization or identify the specific visual contact cues the model uses.