Vision–Language–Action
Unified perception, reasoning, planning, and control for complex real-world environments.
Foundation models for embodied intelligence
I am Kun Zhan, Head of Foundation Models and Autonomous Driving at Li Auto. I build vehicle-scale AI systems that connect perception, language, decision-making, and action—and carry them from frontier research into production.
Beijing / San Jose Li Auto Autonomous Driving · Foundation Models · Embodied AI
My work focuses on unifying the capabilities a physical agent needs: understanding three-dimensional scenes, reasoning about intent and risk, planning under uncertainty, and executing safely in real time.
Build physical-world AGI, starting with autonomous driving and expanding toward robots and intelligent spaces.
Unified perception, reasoning, planning, and control for complex real-world environments.
Simulation, generative scene models, closed-loop evaluation, and learning from physical feedback.
On-vehicle inference, model–chip co-design, data engines, and reliable fleet-scale deployment.
Selected moments across leadership, production systems, model releases, and research.
Leading Li Auto's unified foundation-model and autonomous-driving agenda across Mach VLA, Mach Mind, world models, reinforcement learning, infrastructure, and deployment.
Presented Li Auto's software and embodied-intelligence roadmap, connecting intelligent vehicles with a broader physical-world AI stack.
Contributed to ReconDreamer, StreetCrafter, and DrivingSphere—advancing reconstruction, controllable scene generation, and closed-loop 4D simulation for autonomous driving.
Released DriveVLM and contributed to Street Gaussians and TOD3Cap—bridging multimodal reasoning, dynamic urban reconstruction, and 3D scene understanding.
Helped evolve the driving stack from Highway NoA and City NoA through end-to-end, VLM-assisted, and VLA-based architectures running across production vehicles.
Led L4 prediction and pre-decision algorithms for robo-taxi pilots and production-oriented autonomous-driving systems.
Head of Foundation Models & Autonomous Driving
Site Manager, U.S. R&D Center
Algorithm Lead, L4 Prediction & Planning
Beihang University · M.S. in Navigation, Guidance and Control
Research focused on object recognition and tracking.
University of Science and Technology Beijing · B.Eng. in Automation
Selected work across VLM/VLA systems, world models, 3D reconstruction, planning, and simulation.
VLM · 2024
Combining multimodal reasoning with driving perception, planning, and interaction.
3DGS · ECCV 2024
Modeling dynamic urban scenes with Gaussian splatting.
World model · CVPR 2025
Crafting world models for driving-scene reconstruction via online restoration.
Video diffusion · CVPR 2025
Street-view synthesis with controllable video diffusion models.
4D simulation · CVPR 2025
Building a high-fidelity 4D world for closed-loop simulation.
Short notes for model releases, talks, research, and milestones—without the overhead of a full blog.
For research conversations, speaking, advisory work, or collaboration, reach out directly.