Innovator Coffee EP-38 World Models: The Missing Layer Between AI and the Physical WorldInnovator Coffee

Innovator Coffee EP-38 World Models: The Missing Layer Between AI and the Physical World

76分钟 ·
播放数8
·
评论数0

Welcome to the Innovator Coffee, a podcast that bridges the gap between people and the world of AI and innovation. Follow us to explore the top AI products, ecosystem insights, and the emerging trends.


Guest Bios

Dr. ChongKai Gao

Dr. ChongKai Gao is a Visiting Ph.D. Student at Stanford University in Dr. Feifei Li’s lab, where he conducts research on robotic manipulation, visual planning, and world models. His work focuses on enabling robots to reason about future physical states, perform long-horizon task planning, and bridge AI perception with real-world decision making.


Dr. JunFan Zhu

Dr. JunFan Zhu is the organizer of the San Francisco Robotics & World Model Reading Club, bringing together researchers from leading AI labs to discuss frontier advances in robotics, embodied AI, and world models. His interests span world model architectures, evaluation, tactile intelligence, VLA systems, and the future roadmap of physical AI.


Episode Description

Large Language Models (LLMs) have transformed how AI understands language. But what happens when AI must understand and interact with the physical world?


In this episode of Innovator Coffee, Stanford researcher Dr. ChongKai Gao and robotics researcher Dr. JunFan Zhu explain World Models, one of the fastest-growing areas in AI, Robotics, and Embodied AI. We explore how world models differ from LLMs, their relationship with Vision-Language-Action (VLA) models and Spatial Intelligence, and why companies like NVIDIA, Meta, and Google DeepMind are investing heavily in this field.


We also discuss the biggest challenges in Physical AI, from data and simulation to evaluation, and what it will take for robotics to reach its own "ChatGPT moment."


Timestamps

00:0009:20 | What Is a World Model?

  • World Models vs. Large Language Models
  • Why predicting the future matters more than predicting the next token

09:2121:40 | Mapping the World Model Landscape

  • Five major technical routes
  • JEPA, Dreamer, VLA, diffusion models, and hybrid architectures

21:4029:45 | Spatial Intelligence vs. Decision Making

  • Fei-Fei Li's Spatial Intelligence
  • Why robotics may need action before perfect 3D reconstruction

29:4538:30 | Visual Planning for Robot Manipulation

  • Why robots need to "imagine" before acting
  • Visual planning versus language planning

38:3046:00 | Research Bottlenecks and Missing Pieces

  • Why deployment is much harder than demos
  • Tactile sensing, evaluation, calibration, and the "missing layer" between intelligence and capability

46:0055:00 | Where Will World Models Create Real Business Value?

  • Games, autonomous driving, warehouse automation, and robotics
  • Which applications may commercialize first?

55:0001:10:30 | Robotics' ChatGPT Moment

  • Why robotics doesn't have a scaling law yet
  • Data flywheels, simulation, evaluation, and the future of embodied AI

01:10:30 – End | Betting on the Future of Physical Intelligence

  • Why leading researchers are investing in world models
  • The biggest open questions and what founders, researchers, and investors should watch next


Tom Kong

*Stanford EE alumni,

*Founder@ Stanford AGI Adventist Community (10K+ members so far from top VC, Engineers, startups from Silicon Valley )

*AI Lecturer, a serial entrepreneur in media and data. Advisor @ techtimes.com and heyboss.ai

*AI deployment for 8 years, with NLP and recent LLMs (RAG, Agent, Diffusion)


Wickey Wang

*IT Security Compliance Leader & University Faculty

*VC advisor and Angel Investor with cybersecurity and AI focus

*GAI Security book co-author

Co-founderQuestions, Suggestions, Feedback and Comments? You can find us in LinkedIn:

www.linkedin.com

www.linkedin.com