On August 27, XPeng Group held an in-person AI sharing event and the second-generation VLA new version experience day with the theme "TIME - Time". The second-generation VLA large model has undergone its first major upgrade since its launch, with the new XOS 6.3.0 version making its global debut on the XPeng G9L. The core of this model upgrade is enabling AI to truly understand time, driving the physical AI foundation model from static 3D space understanding to dynamic 4D spatiotemporal understanding at a fundamental level.
XPeng has been continuously verifying the Scaling Law in the field of autonomous driving. This time, the edge-side model parameter count was increased by 3.5 times, creating a significant gap with industry common small models. With ultra-long sequence context memory capabilities, the AI driver can now "see accurately, think effectively, and react quickly". At the same time, the new version introduces the vehicle's master agent, achieving the integration of VLA and VLM in the cockpit, and bringing some Level 4 robotaxi experiences down to mass-produced vehicles.

The real world faced by physical AI is not a series of static images, but a continuous process. For the model to understand the world like humans, it must first understand "time". Although traditional models can recognize vehicles and pedestrians, vehicles need to know what might happen in the past, present, and future. To address this, the second-generation VLA has introduced the Infini-VLA long-time sequence architecture, which can remember the world for the previous 30 seconds, giving driving decisions a longer context. Continuously occurring road events are no longer fragmented into individual frames.
To keep up with the constantly changing real-world road conditions, the second-generation VLA has simultaneously improved model reasoning efficiency, adopting Streaming Inference (streaming autoregressive inference), moving from discrete inference to continuous inference, and increasing end-to-end response speed by 300%. The model can perform parallel processing of "seeing, thinking, and acting", reacting as skillfully and agilely as experienced drivers when facing sudden situations such as sudden braking of the car ahead, pedestrians turning back, or surrounding vehicles forcing a lane change.
In addition, XPeng's X-Foresight prediction world model has been deployed on vehicles for the first time, capable of predicting events within the next 6 seconds and simulating the possible behaviors of surrounding traffic participants. The new model also introduces the MoT hybrid architecture, dynamically allocating model capabilities for different tasks, reducing task interference between different scenarios such as urban areas, campuses, and parking. With larger parameters, advanced trajectory prediction, ultra-long effective time sequences, and faster decision-making responses, multi-dimensional comprehensive safety capabilities have improved by 20 times.
In terms of the vehicle cockpit, XPeng has launched the Master Agent for the first time, reconfiguring the vehicle brain using a robotic approach to achieve the integration of VLA and VLM in the cockpit. It not only understands natural language and ambiguous semantics but can also automatically break down user intent into tasks, scheduling vertical agents such as intelligent driving, chassis, cockpit, and body control to execute collaboratively. The new version also brings features such as "voice-guided parking" and "voice-controlled nearby parking", achieving a complete closed-loop from "voice commands" to "autonomous execution". The L4-level robotaxi technology is accelerating its deployment to mass-produced vehicles.
As a global embodied intelligence company, XPeng is committed to unifying the technical foundation to connect cars and robots, pushing physical AI into a stage of large-scale replication. The Robotaxi equipped with the second-generation VLA has recently obtained Guangzhou's qualification for remote testing of intelligent connected vehicles, allowing testing on relevant roads without a safety driver. In terms of the robot body, the XPeng general-purpose humanoid robot IRON, equipped with three Turing AI chips, has a computing power of 2250 TOPS and has achieved edge-side deployment, capable of independently completing complex tasks. XPeng's robot business has also recently completed a $900 million first-round equity financing, with a post-investment valuation exceeding $6.3 billion.
To achieve technological inclusivity, XPeng has deployed the second-generation VLA base model capabilities on platforms with lower computing power through learning-based Token compression and distillation training. The distilled Turing VLA2.0Lite will be first launched on the XPeng G9L Max version in September. At the same time, XPeng is actively advancing the global deployment of the second-generation VLA. Recently, it completed localized acceptance testing in Germany, aiming to obtain regulatory approval in Europe first half of next year and deliver to overseas users.
Behind this series of advancements lies XPeng's long-term built complete AI Infra system. Relying on a million-vehicle fleet and ten-billion data assets, XPeng has established a data flywheel, with single-train data throughput reaching 100 million clips, ten times higher than six months ago. By continuously accumulating AI Infra, XPeng is expanding the evolution of model capabilities beyond single automotive products to different embodiments such as robots, continuously extending the boundaries of physical AI interacting with the real world.