詹锟 | Kun Zhan

理想汽车基座模型与自动驾驶负责人

portrait.jpeg

Email: zk_1028@aliyun.com

微信: KevinZhan1990

北京,中国

面向具身智能的基座模型

构建从感知、语言到行动的整车级 AI 技术栈

我负责理想汽车基座模型与自动驾驶业务,打造理想同学 VLA 与 Mind GPT 模型家族,并将其作为可靠的物理世界智能系统部署到量产车上。

Google Scholar zk_1028@aliyun.com KevinZhan1990 北京,中国

关于我

我是 詹锟,现任理想汽车基座模型与自动驾驶负责人。我的工作聚焦于将机器智能与语言智能统一为量产级具身系统:以 理想同学 VLA 实现三维感知、决策与行动,以 Mind GPT-Pro / Mind GPT-Edge 实现云端与端侧的 Agent 推理。

自 2021 年加入理想汽车以来,我一直负责自动驾驶核心架构的演进,推动系统从高速 NoA、城市 NoA 发展至端到端、VLM 辅助和 VLA 架构。目前,我的工作覆盖模型研究、数据引擎、强化学习、世界模型、芯片与模型协同设计、车端推理,以及汽车规模的 OTA 部署。

此前,我于北京航空航天大学获得自动化专业硕士学位,并在百度 Apollo 负责行为预测工作,为 L4 自动驾驶项目打造大规模运动预测与规划系统。

我的使命是打造物理世界 AGI,以自动驾驶为起点,逐步拓展至机器人、智能空间以及更广阔的真实世界具身智能。

核心亮点

贯穿完整具身智能技术栈的团队领导、模型发布与量产落地。

统一 AI 业务领导

负责理想汽车基座模型与自动驾驶路线图,覆盖理想同学 VLA、Mind GPT、3D ViT、世界模型、强化学习、模型基础设施与车端部署。

模型产品发布

推动理想同学 VLA 与 Mind GPT-Pro / Mind GPT-Edge 模型家族发布,为车辆、座舱和未来具身产品连接语言智能与机器智能。

全栈量产交付

通过模型、芯片、操作系统与域控制器的一体化,将研究系统转化为用户可用的产品能力,包括围绕理想自研 M100 AI 芯片和 Livis 车型构建的部署方案。

近期里程碑

近期公开发布与演讲,呈现我当前工作的方向。

Livis 发布会主题演讲

作为理想汽车软件与具身智能发布会主讲人,发布全新理想同学模型技术栈,并阐述从智能汽车走向具身智能产品的路径。

观看发布会回放

理想同学 VLA

发布理想汽车基于 VLA 的下一代自动驾驶架构,通过扩展模仿学习、强化学习、模型规模与算力,让真实道路驾驶更安全、更高效。

阅读发布报道

Mind GPT

发布面向云端与端侧 Agent 智能的 Mind GPT-Pro 和 Mind GPT-Edge 模型家族,实现原生车辆控制、多模态交互与持续在线的本地感知。

阅读模型报道

研究方向

构成我当前研究与工程工作的主要方向。

物理世界基座模型VLA、世界模型、3D ViT、强化学习、实时决策与行动
自动驾驶端到端驾驶、安全、规划与闭环部署
Agent LLM/VLM 系统云端与边缘 Agent、工具使用、多模态交互、人车协同
AI 系统与芯片模型与芯片协同设计、编译器与运行时、车端推理、低成本部署
仿真与数据引擎车队级数据飞轮、合成数据、闭环评测与训练基础设施
机器人与具身智能从车辆泛化至空间智能、人形机器人、操作与导航

工作经历

塑造我对应用型 AI 系统理解的核心项目与岗位。

理想汽车

2021 年 4 月 - 至今

北京 / 圣何塞
基座模型与自动驾驶负责人
  • 负责理想汽车基座模型与自动驾驶团队,统一理想同学 VLA、Mind GPT、3D ViT、世界模型、强化学习、数据基础设施与车端部署。
  • 发布理想同学 VLA 与 Mind GPT-Pro / Mind GPT-Edge 具身智能模型系列,并将其量产集成至 Livis 车型和理想汽车全栈 AI 平台。
  • 围绕自研 AI 芯片、车辆操作系统和域控制器,推动芯片与模型协同设计及量产部署。
  • 从高速 NoA 和城市 NoA 开始,构建理想汽车自动驾驶技术栈,并演进至在数十万辆车上运行的端到端、VLM 辅助与 VLA 架构。
  • 指导由感知、规划、基座模型、仿真、数据与部署等方向组成的 100+ 人团队。
美国研发中心负责人
  • 从零建立理想汽车海外研究中心,负责当地战略、预算与人才招聘。
  • 通过跨境项目评审与路线图对齐,连接硅谷创新与北京执行。

百度 Apollo

2016 年 4 月 - 2021 年 3 月

北京,中国
L4 预测与规划算法负责人
  • 负责 RoboTaxi 试点项目的 L4 预测与前决策算法,提升复杂城市场景中的运动预测能力。
  • 交付规划控制模块和深度学习车端组件,支持自动驾驶车队在北京和广州运行。

学术成果

基于 Google Scholar 的论文与引用快照

论文数 63
总引用 2068
h-index 18
i10-index 29

Top 10 引用论文

按 Google Scholar 引用量排序,更新时间见卡片上方。

Google Scholar
Drivevlm: The convergence of autonomous driving and large vision-language models
arXiv preprint arXiv:2402.12289 2024 引用 714

Drivevlm: The convergence of autonomous driving and large vision-language models

X Tian, J Gu, B Li, Y Liu, Y Wang, Z Zhao, K Zhan, P Jia, X Lang, H Zhao

Street gaussians: Modeling dynamic urban scenes with gaussian splatting
European Conference on Computer Vision, 156-173 2024 引用 488

Street gaussians: Modeling dynamic urban scenes with gaussian splatting

Y Yan, H Lin, C Zhou, W Wang, H Sun, K Zhan, X Lang, X Zhou, S Peng

Recondreamer: Crafting world models for driving scene reconstruction via online restoration
Proceedings of the Computer Vision and Pattern Recognition Conference, 1559-1569 2025 引用 106

Recondreamer: Crafting world models for driving scene reconstruction via online restoration

C Ni, G Zhao, X Wang, Z Zhu, W Qin, G Huang, C Liu, Y Chen, Y Wang, ...

Planagent: A multi-modal large language agent for closed-loop vehicle motion planning
IEEE Transactions on Cognitive and Developmental Systems 2026 引用 76

Planagent: A multi-modal large language agent for closed-loop vehicle motion planning

Y Zheng, Z Xing, Q Zhang, B Jin, P Li, Y Zheng, Z Xia, Y Chen, D Zhao

Unleashing generalization of end-to-end autonomous driving with controllable long video generation
arXiv preprint arXiv:2406.01349 2024 引用 56

Unleashing generalization of end-to-end autonomous driving with controllable long video generation

E Ma, L Zhou, T Tang, Z Zhang, D Han, J Jiang, K Zhan, P Jia, X Lang, ...

Streetcrafter: Street view synthesis with controllable video diffusion models
Proceedings of the Computer Vision and Pattern Recognition Conference, 822-832 2025 引用 52

Streetcrafter: Street view synthesis with controllable video diffusion models

Y Yan, Z Xu, H Lin, H Jin, H Guo, Y Wang, K Zhan, X Lang, H Bao, X Zhou, ...

Tod3cap: Towards 3d dense captioning in outdoor scenes
European Conference on Computer Vision, 367-384 2024 引用 47

Tod3cap: Towards 3d dense captioning in outdoor scenes

B Jin, Y Zheng, P Li, W Li, Y Zheng, S Hu, X Liu, J Zhu, Z Yan, H Sun, ...

Drivingsphere: Building a high-fidelity 4d world for closed-loop simulation
Proceedings of the Computer Vision and Pattern Recognition Conference, 27531 … 2025 引用 40

Drivingsphere: Building a high-fidelity 4d world for closed-loop simulation

T Yan, D Wu, W Han, J Jiang, X Zhou, K Zhan, C Xu, J Shen

The better you learn, the smarter you prune: Towards efficient vision-language-action models via differentiable token pruning
arXiv preprint arXiv:2509.12594 2025 引用 33

The better you learn, the smarter you prune: Towards efficient vision-language-action models via differentiable token pruning

T Jiang, X Jiang, Y Ma, X Wen, B Li, K Zhan, P Jia, Y Liu, S Sun, X Lang

Finetuning generative trajectory model with reinforcement learning from human feedback
arXiv e-prints, arXiv: 2503.10434 2025 引用 33

Finetuning generative trajectory model with reinforcement learning from human feedback

D Li, J Ren, Y Wang, X Wen, P Li, L Xu, K Zhan, Z Xia, P Jia, X Lang, N Xu, ...

专利、学术服务与社区

量产技术栈之外的研究服务与技术转化。

专利

已授权或公开专利 20 项,其中中国 18 项、美国 2 项,覆盖感知、规划和高精地图流程。

审稿服务

担任 CVPR、ICCV、ECCV、NeurIPS、AAAI、IROS,以及 TPAMI、T-ITS、T-IV 等期刊的审稿人。

社区

Livis 发布会主题演讲嘉宾、CVPR 2023 Autonomous Driving Workshop 组织者,并持续分享 VLA 在量产环境中的部署实践。