At the JD.com Global Tech Explorer Conference (JDD) on September 9th, the JD Explore Research Institute unveiled two new developments in the JoyAI-Echo series: the upgraded JoyAI-Echo1.5 long video generation model and the newly released and open-sourced interactive audio-visual world model EchoWM. The former surpasses the 10-minute threshold for continuous audio-visual generation, while the latter integrates native audio-visual generation with user control, taking AI content from "watchable" to "enterable and exploratory."
In terms of long video, JoyAI-Echo1.5 uses the audio-visual Memory Bank architecture. It can firmly maintain character images, voices, and story lines during multi-camera, long-duration generation, supporting continuous production of audio-visual content lasting more than 10 minutes. By training on long videos, the model's inference speed is increased by up to 7.5 times, requiring only 2 H200s to generate 24 frames per second at 480P resolution, achieving real-time 1:1 generation; the Director Agent can also complete the entire process of story development, shot planning, generation review, and final assembly through natural language.

In the area of world models, there are usually two approaches: one generates visuals continuously based on user actions but lacks natural sounds; the other can jointly generate audio and video but users find it hard to keep making changes over time. EchoWM combines these two capabilities, generating native audio-visual content including 720P video, ambient sounds, music, and voice, allowing sound to change synchronously with scene, camera, and subject movements. Through a unified "camera intention," the model supports first-person navigation, third-person camera and subject collaboration control, as well as multi-round continuous exploration. In the WBench Navigation evaluation, EchoWM and the four-step causal streaming version EchoWM-Flash took the top two positions, with EchoWM securing the overall first place on a list covering 158 multi-round navigation cases.
On the foundation level, JD.com has concentrated on showcasing the JoyAI basic model matrix, covering image and video generation, multimodal understanding, real-time interaction, world modeling, intelligent voice, and embodied operations, pushing JoyAI from content generation towards perception, simulation, and interaction in dynamic environments. At the JDD event, the new audio-video foundation model JoyAI-Video was also demonstrated, focusing on enhancing commercial video generation, physical law representation, first-person perspective, and multi-shot continuous storytelling, targeting e-commerce advertising, short drama creation, and embodied intelligence training, with a planned official release in October. In addition to the model, JD.com also exhibited EgoLive real-world data, the JoyAI-Sim simulation conversion toolchain, and large model infrastructure, building a support system from real-world data, simulation conversion, to model development.





