Vision–Language–Action
Unified perception, reasoning, planning, and control for complex real-world environments.
Foundation models for embodied intelligence
I am Kun Zhan, Head of Foundation Models and Autonomous Driving at Li Auto. I build vehicle-scale AI systems that connect perception, language, decision-making, and action—and carry them from frontier research into production.
Beijing / San Jose Li Auto Autonomous Driving · Foundation Models · Embodied AI
Overall citation metrics from Google Scholar ↗ · Updated . Publication list and per-paper citations: snapshot.
My work focuses on unifying the capabilities a physical agent needs: understanding three-dimensional scenes, reasoning about intent and risk, planning under uncertainty, and executing safely in real time.
Build physical-world AGI, starting with autonomous driving and expanding toward robots and intelligent spaces.
Unified perception, reasoning, planning, and control for complex real-world environments.
Simulation, generative scene models, closed-loop evaluation, and learning from physical feedback.
On-vehicle inference, model–chip co-design, data engines, and reliable fleet-scale deployment.
Eight principles on technology, organizations, and lasting value.
Ambitious in goals. Deliberate in choices. Accountable for outcomes.
In the age of AI, I want to do more than keep pace with technology. I want to push the boundaries of intelligence, create value in the real world, and help exceptional people do difficult things together over the long term.
These are not things I claim to have achieved. They are standards I hold myself to in research, in leading teams, and in making trade-offs.
Open each principle to read the full reflection.
A vision I believe in must answer three questions: Why do this? Where should the best resources go? Which opportunities are worth passing up?
Fine words are not enough. When short-term gains conflict with long-term goals, choices are what count. What I truly believe should be visible in my decisions.
I value fair returns, but I do not aim to capture every benefit. I do not need to take every reward or own every part of the work.
Before securing my own share, I care more about whether cooperation and sharing can improve our chances of achieving something significant. Restraint does not mean a lack of ambition; it means directing ambition toward what matters more.
Something worth doing is not necessarily ours to do. The fact that everyone else is doing it is not a reason for us to follow.
I care less about appearing to cover everything than about truly solving the key problems. Intermediate wins are worth pursuing, but they must not replace the ultimate goal. Saying no to good opportunities outside our focus is also a responsibility.
My responsibility is not to make every judgment for everyone. It is to make our shared direction clear and build trust and collaboration.
I want to create an environment where excellent people have the confidence to exercise judgment, the willingness to collaborate, and the opportunity to grow. Being able to keep moving forward together through setbacks matters more than a momentarily impressive roster.
What we have promised deserves serious follow-through. Directions that have not yet been proven also deserve a chance to be tested.
I do not accept using “research” to evade delivery, or filling every hour of everyone’s schedule to create a sense of managerial security. Well-defined work needs efficient execution; uncertain exploration needs time, resources, and patience.
I care both about the limits of intelligence and whether it can actually be used within real constraints on compute, cost, and latency.
Enabling the same resources to support stronger models, and stronger models to serve more people, is a core problem worth investing in. Algorithmic breakthroughs and engineering efficiency belong together: efficiency is not only about savings, but about making previously impossible things feasible.
I care not only about what a model can do this time, but whether it can learn from interaction and feedback to do better next time.
I see continual learning as a core direction for long-term investment. This applies to models, to myself, and to the team: we should not simply repeat work, but let every experience improve the next judgment.
AI should be more than the product of our research and development. It should become part of our ability to do that work.
I want each generation of models to help us find problems, test ideas, and improve systems faster. Progress should be measured not only by how capable this version is, but also by whether it brings the next breakthrough closer.
Deepen intelligence. Make things happen. Keep progress going.
Selected moments across leadership, production systems, model releases, and research.
Leading Li Auto's unified foundation-model and autonomous-driving agenda across Mach VLA, Mach Mind, world models, reinforcement learning, infrastructure, and deployment.
Presented Li Auto's software and embodied-intelligence roadmap, connecting intelligent vehicles with a broader physical-world AI stack.
Contributed to ReconDreamer, StreetCrafter, and DrivingSphere—advancing reconstruction, controllable scene generation, and closed-loop 4D simulation for autonomous driving.
Released DriveVLM and contributed to Street Gaussians and TOD3Cap—bridging multimodal reasoning, dynamic urban reconstruction, and 3D scene understanding.
Helped evolve the driving stack from Highway NoA and City NoA through end-to-end, VLM-assisted, and VLA-based architectures running across production vehicles.
Led L4 prediction and pre-decision algorithms for robo-taxi pilots and production-oriented autonomous-driving systems.
Head of Foundation Models & Autonomous Driving
Site Manager, U.S. R&D Center
Algorithm Lead, L4 Prediction & Planning
Beihang University · M.S. in Navigation, Guidance and Control
Research focused on object recognition and tracking.
University of Science and Technology Beijing · B.Eng. in Automation
Selected work across VLM/VLA systems, world models, 3D reconstruction, planning, and simulation.
VLM · 2024
Combining multimodal reasoning with driving perception, planning, and interaction.
3DGS · ECCV 2024
Modeling dynamic urban scenes with Gaussian splatting.
World model · CVPR 2025
Crafting world models for driving-scene reconstruction via online restoration.
Video diffusion · CVPR 2025
Street-view synthesis with controllable video diffusion models.
4D simulation · CVPR 2025
Building a high-fidelity 4D world for closed-loop simulation.
Short notes for model releases, talks, research, and milestones—without the overhead of a full blog.
For research conversations, speaking, advisory work, or collaboration, reach out directly.