BYD's AI Team Revealed for the First Time, Harbin Institute of Technology's Robotics Genes Shine, Large Model Achieves SOTA on Debut
BYD's AI team has unveiled its first multimodal foundational model, HyWorldVLA, achieving state-of-the-art (SOTA) results on the NAVSIM v1 autonomous driving benchmark. This marks BYD's public entry into advanced physical AI research, moving beyond traditional ADAS.
Intelligence analysis by Gemini 2.5 Flash
The article highlights BYD's previously understated AI capabilities, revealing a dedicated team focused on foundational physical AI models, particularly for autonomous driving. Their HyWorldVLA model combines vision-language-action (VLA) and hybrid world modeling to enable AI to understand and predict real-world driving scenarios efficiently, challenging the perception of BYD's intell…
Imagine a super-smart car brain that doesn't just see the road but can also guess what will happen next, like if a ball will roll out or another car will turn. BYD built a special AI brain called HyWorldVLA that learns from videos and words to understand the world better and drive safely and smoothly, almost like it has a tiny crystal ball inside. This helps their cars drive smarter, even in tricky weather.
Analysis
BYD's AI Ambition Takes Center Stage
For years, BYD's prowess in intelligent vehicles was largely associated with its robust supply chain integration, large-scale deployment, and hardware capabilities, often leaving its "code-level self-developed" AI technology and academic contributions in the shadows. The recent unveiling of HyWorldVLA, a multimodal foundational model from BYD Automotive New Technology Research Institute, marks a significant departure from this perception. This paper, achieving state-of-the-art (SOTA) results on the NAVSIM v1 autonomous driving benchmark, signals BYD's serious commitment to foundational physical AI research, a domain previously thought to be dominated by tech giants like Tesla and Waymo.
The HyWorldVLA model specifically targets the cutting-edge Vision-Language-Action (VLA) and World Model paradigms in autonomous driving. This approach aims to move AI beyond mere road recognition towards a deeper "understanding of the world," enabling it to predict future scenarios and make informed driving decisions. This public disclosure not only showcases BYD's advanced research capabilities but also challenges the narrative that its intelligence strategy was primarily focused on integration rather than deep, proprietary AI development.
HyWorldVLA: A Hybrid Approach to World Models
The core innovation of HyWorldVLA lies in its hybrid world modeling approach, designed to resolve the inherent tension between understanding real-world details and achieving efficient inference in autonomous driving. Traditional methods often struggle, with pixel-level models offering detail but high computational cost, and latent world models providing efficiency but potentially losing crucial real-world nuances. BYD's solution integrates both, leveraging a three-step training process.
Initially, a video compressor transforms continuous video frames into compact latent features, guided by textual information to enhance reconstruction quality. This is followed by a pre-training phase where the model learns to predict future frames and latent features simultaneously, ensuring the latent space retains sufficient information. Finally, a joint fine-tuning stage focuses solely on outputting latent features, which are then used by an action generation module to produce driving trajectories. This phased approach, validated by ablation studies, demonstrates that the hybrid model's superior performance, particularly in challenging conditions like rain and fog, stems from its ability to "feed" the latent space with pixel-level supervision during pre-training, then switch to efficient latent-space inference during fine-tuning, thereby gaining robustness against visual noise.
The Harbin Institute of Technology Robotics Pedigree
The independent development of HyWorldVLA by BYD Automotive New Technology Research Institute, without external academic co-authorship, is noteworthy. A deeper look into the team behind this and other recent BYD AI papers, such as EMoE-Planner (published in 2025 according to the article), reveals a distinct "robotics AI faction" within the company. Key figures like Liulong Ma, a corresponding author on multiple BYD AI papers, and Hongbiao Zhu, the first author of EMoE-Planner, share strong academic backgrounds from the Harbin Institute of Technology (HIT) and Carnegie Mellon University (CMU).
HIT, particularly its State Key Laboratory of Robotics and System, is renowned for its robotics research. This strong academic lineage in robotics, rather than traditional automotive electronics or ADAS engineering, suggests a strategic recruitment of top-tier computer science and robotics talent. This organizational philosophy aligns closely with Tesla's integrated "robotics + autonomous driving" model, where the underlying principles of physical AI for robots are directly applied to autonomous vehicles. This foundational approach, rooted in first-principles thinking for physical AI, positions BYD to compete at the forefront of the evolving autonomous driving landscape, where the focus is shifting from mere feature stacking to comprehensive physical AI capabilities.
Key points
- BYD's AI team, primarily from its Automotive New Technology Research Institute, has publicly revealed its first multimodal foundational model, HyWorldVLA.
- HyWorldVLA achieved SOTA results on the NAVSIM v1 benchmark for autonomous driving, focusing on Vision-Language-Action (VLA) and hybrid world modeling.
- The model aims to enable AI to "understand the world" and predict future scenarios, combining the strengths of pixel-level and latent world models for efficiency and detail.
- The core AI team shows a strong "robotics AI faction" background, with key members having ties to Harbin Institute of Technology and Carnegie Mellon University.
- This move signifies BYD's shift towards deep, self-reliant AI research, aligning with a "robotics + autonomous driving" integrated approach similar to Tesla.
BYD's demonstrated capability in foundational physical AI research could significantly accelerate its autonomous driving development, reduce reliance on external suppliers, and enhance its competitive edge in the global smart electric vehicle market. This internal expertise could also foster innovation across its broader robotics initiatives, aligning with a more integrated AI strategy.
While achieving SOTA in benchmarks is promising, translating simulation results to real-world mass production faces significant engineering challenges, including extensive real-world data collection, vehicle-side computing constraints, and robust handling of diverse, long-tail scenarios. There's no guarantee this research will quickly lead to a market-leading product or overcome these complex deployment hurdles.


