Robot Learning · Vision-Language-Action

Tingting Du

B.S. student in Computer Science and Mathematics at the University of Wisconsin–Madison. I build robot learning systems that connect language, perception, and physical action — currently on trimanual manipulation in the RT² Lab.

  • RT² Lab, UW–Madison
  • Trimanual Manipulation
  • Madison, WI

Demos

A look at what the systems I work on actually do — one robot, one model.

01

Trimanual Manipulation — RT² Lab, UW–Madison

Ongoing, with Prof. Mike Hagenow · Robot Teaching & Teaming Lab

Scene Wrist · 1 Wrist · 2 Wrist · 3
Episode 0011

A single recorded episode, played back across the four synchronized camera streams the policy sees: one scene overview plus a wrist view per arm. Three coordinating arms rearrange coloured blocks across pegboards — a setting where the hard part is not the individual grasp but keeping three end-effectors, three viewpoints, and one shared workspace consistent over a long horizon.

02

ROCKET — Spatially-Aware Vision-Language-Action Models

arXiv 2026 · with the CASE Lab, University of Maryland

ROCKET architecture: a VLA model's residual stream is aligned layer-by-layer with a frozen 3D foundation model through a shared projector, with Matryoshka-style activation, producing end-effector deltas. Click to enlarge

VLA models read pixels well but reason about space poorly. ROCKET aligns the VLA's residual stream layer by layer with a frozen 3D foundation model through a shared projector, using Matryoshka-style activation so depth cues at different ranges land in different parts of the representation. Spatial grounding is distilled into the policy during training, so no depth sensor is needed at inference.

About

I am an undergraduate at the University of Wisconsin–Madison, studying Computer Science and Mathematics. Before Madison I was a visiting student in Computer Science at UC Berkeley, and I began my studies in Linguistics at Ningbo University.

I currently work with Prof. Mike Hagenow in the Robot Teaching and Teaming (RT²) Lab at UW–Madison on trimanual manipulation. Previously I worked with Prof. Ang Li in the CASE Lab at the University of Maryland on Vision-Language-Action models, with Prof. Meng Jiang at the University of Notre Dame on student modeling and question generation, and with Prof. Alane Suhr at Berkeley AI Research on situated language understanding.

Research Focus

Vision-Language-Action Models

Giving policies a sense of 3D space and scale, and understanding what the data and benchmarks behind them actually measure.

Multi-Arm Manipulation

Coordinating several arms and viewpoints on long-horizon, contact-rich tasks, and learning those behaviours from human demonstration.

Grounded Language & Reasoning

How people and models use language in situated, collaborative settings — and how memory and simulation improve agent reasoning.

Updates

  • 2026.08 Joined the RT² Lab at UW–Madison with Prof. Mike Hagenow, working on trimanual manipulation.New
  • 2026.04 Our Vision-Language-Action survey was accepted to TMLR.
  • 2026.02 ROCKET released on arXiv.
  • 2025.12 Collaborative situated game paper released on arXiv.
  • 2025.01 QG-SMS accepted to ACL 2025.
  • 2024.06 Completed a research internship at Berkeley AI Research.

Publications

  • arXiv 2026 ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models

    G. Sun, T. Du, K. Feng, C. Luo, X. Ding, Z. Shen, Z. Wang, Y. He, A. Li

    arXiv:2602.17951

  • TMLR 2026 Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines

    Z. Wang, B. Wang, H. Zhang, T. Du, T. Chen, G. Sun, Y. He, Z. Shen, W. Ye, A. Li

    Transactions on Machine Learning Research

  • ACL 2025 QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation

    B. Nguyen, T. Du, M. Yu, L. Angrave, M. Jiang

    Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics

  • arXiv 2025 Characterizing Language Use in a Collaborative Situated Game

    N. Tomlin, N. Zhou, E. Fleisig, L. Chen, T. Wright, L. Vinh, L.X. Ma, S. Eisape, T. Du, T. Zhang, A. Koller, A. Suhr

    arXiv:2512.03381

  • ICML 2025 Workshop Agent KB: A Hierarchical Memory Framework for Cross-Domain Agentic Problem Solving

    X. Tang, T. Qin, T. Peng, Z. Zhou, D. Shao, T. Du, X. Wei, H. Zhu, G. Zhang, et al.

    Workshop on Collaborative and Federated Agentic Workflows

Education

  • Jan 2025 – May 2027 University of Wisconsin–Madison B.S. in Computer Science and Mathematics (expected)
  • 2023 – 2024 University of California, Berkeley Visiting Student in Computer Science
  • 2021 – 2023 Ningbo University Undergraduate studies in Linguistics