arXiv 2026
DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects
Tianshan Zhang, Yijia Duan, Yanjun Li, Zeyu Zhang, Hao Tang
B.S. Candidate in Computer Science & Materials Science
张天山
I work on generative models of the world — video generation and world models — and on bringing them to robots that perceive, reason, and act in the physical world.
I am interested in generative models that learn how the world looks and changes: video generation, and world models that predict what happens next under an action. A model that can imagine the future is a natural substrate for planning, simulation, and learning without a real robot in the loop.
On the robotics side I work on vision-language-action models and humanoid robots: policies that ground language and vision in physical action, from dexterous hand-object interaction to whole-body control.
arXiv 2026
Tianshan Zhang, Yijia Duan, Yanjun Li, Zeyu Zhang, Hao Tang

Vision-Language-Action Models, Generative Models, Robotic Manipulation
2025 – Present
Mathematical Reasoning, LLM Inference
2025 – 2026

Generative Models
2024 – 2025
Classical mechanics can be organized around force, action, or energy. The Hamiltonian—energy on phase space—turns out to be the object that survives into quantum theory as the generator of time evolution on Hilbert space, while the action reappears as the phase weighting every path in the path integral. Following these two threads, together with the Poisson-bracket-to-commutator correspondence and the symmetry–conservation link, leads through quantization, the harmonic oscillator, statistical mechanics, quantum field theory, renormalization, and effective field theory.
Vision-Language-Action models are transforming robot learning from a collection of task-specific controllers into general-purpose embodied agents. Rather than treating VLA as a single model architecture, we examine it as a sequence of responses to increasingly difficult questions: how robots understand tasks, how semantic knowledge becomes continuous action, how policies operate over long horizons, and how they improve beyond static demonstrations.
A mathematical walkthrough of sequence modeling, tracing the evolution from recurrent neural networks and LSTM gating mechanisms to self-attention, FlashAttention, KV-cache optimization, RoPE, state space models, Mamba, RWKV, and hybrid long-context architectures.