Robot Learning · Vision-Language-Action
Tingting Du
B.S. student in Computer Science and Mathematics at the University of Wisconsin–Madison. I build robot learning systems that connect language, perception, and physical action — currently on trimanual manipulation in the RT² Lab.
- RT² Lab, UW–Madison
- Trimanual Manipulation
- Madison, WI
Demos
A look at what the systems I work on actually do — one robot, one model.
Trimanual Manipulation — RT² Lab, UW–Madison
A single recorded episode, played back across the four synchronized camera streams the policy sees: one scene overview plus a wrist view per arm. Three coordinating arms rearrange coloured blocks across pegboards — a setting where the hard part is not the individual grasp but keeping three end-effectors, three viewpoints, and one shared workspace consistent over a long horizon.
ROCKET — Spatially-Aware Vision-Language-Action Models
Click to enlargeVLA models read pixels well but reason about space poorly. ROCKET aligns the VLA's residual stream layer by layer with a frozen 3D foundation model through a shared projector, using Matryoshka-style activation so depth cues at different ranges land in different parts of the representation. Spatial grounding is distilled into the policy during training, so no depth sensor is needed at inference.
About
I am an undergraduate at the University of Wisconsin–Madison, studying Computer Science and Mathematics. Before Madison I was a visiting student in Computer Science at UC Berkeley, and I began my studies in Linguistics at Ningbo University.
I currently work with Prof. Mike Hagenow in the Robot Teaching and Teaming (RT²) Lab at UW–Madison on trimanual manipulation. Previously I worked with Prof. Ang Li in the CASE Lab at the University of Maryland on Vision-Language-Action models, with Prof. Meng Jiang at the University of Notre Dame on student modeling and question generation, and with Prof. Alane Suhr at Berkeley AI Research on situated language understanding.
Research Focus
Vision-Language-Action Models
Giving policies a sense of 3D space and scale, and understanding what the data and benchmarks behind them actually measure.
Multi-Arm Manipulation
Coordinating several arms and viewpoints on long-horizon, contact-rich tasks, and learning those behaviours from human demonstration.
Grounded Language & Reasoning
How people and models use language in situated, collaborative settings — and how memory and simulation improve agent reasoning.
Updates
- 2026.08 Joined the RT² Lab at UW–Madison with Prof. Mike Hagenow, working on trimanual manipulation.New
- 2026.04 Our Vision-Language-Action survey was accepted to TMLR.
- 2026.02 ROCKET released on arXiv.
- 2025.12 Collaborative situated game paper released on arXiv.
- 2025.01 QG-SMS accepted to ACL 2025.
- 2024.06 Completed a research internship at Berkeley AI Research.
Publications
- arXiv 2026 ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models
arXiv:2602.17951
- TMLR 2026 Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
Transactions on Machine Learning Research
- ACL 2025 QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics
- arXiv 2025 Characterizing Language Use in a Collaborative Situated Game
arXiv:2512.03381
- ICML 2025 Workshop Agent KB: A Hierarchical Memory Framework for Cross-Domain Agentic Problem Solving
Workshop on Collaborative and Federated Agentic Workflows
Education
- Jan 2025 – May 2027 University of Wisconsin–Madison B.S. in Computer Science and Mathematics (expected)
- 2023 – 2024 University of California, Berkeley Visiting Student in Computer Science
- 2021 – 2023 Ningbo University Undergraduate studies in Linguistics
