Jingdong has announced the open-source of its self-developed real-time streaming video editing model JoyAI-Video-Edit. Users can modify characters and scenes while watching videos, transforming video creation from "editing after having materials" to real-time interactive editing. Jingdong's official stated that this model comprehensively outperforms current industry representative models in real-time streaming video editing, reaching world-class standards.

image.png

Real-time video editing first needs to solve the speed issue. Videos are made up of continuous frames, and if the model's processing speed cannot keep up with the playback speed, lag and delay will occur. JoyAI-Video-Edit is based on the fully self-developed JoyAI-Video model, achieving a model inference speed of 30 frames per second at 720P resolution. It can maintain high image clarity while keeping up with the playback rhythm of common films and videos.

Video length is also a barrier for streaming editing. Previously, streaming editing models could only handle seconds or minute-level segments, but JoyAI-Video-Edit supports stable streaming editing of any duration, allowing continuous processing as the video plays. Users don't need to wait for the entire video to be generated before they can make real-time changes to scenes, replace objects, adjust character appearances, or change the visual style, achieving "create while watching."

Comprehensive authoritative evaluation, outstanding advantages in four types of editing tasks

To evaluate the model's capabilities, the research team compared JoyAI-Video-Edit with international representative streaming video editing models such as SANA Streaming, LiveEdit, and Xmax-X2.0. The results showed that JoyAI-Video-Edit comprehensively outperformed others in evaluation data covering typical video editing scenarios.

In the industry's authoritative OpenVE-Bench general evaluation, this model surpassed all current streaming editing methods, especially excelling in global style, local replacement, local deletion, and subtitle editing tasks, significantly outperforming industry peers. In specific application scenarios, creators can let characters continuously change clothes in the video, transform ordinary streets into animated worlds during live streams, or try different furniture, wall colors, and lighting effects in home design.

Not just video creation, but also producing training data for robots

Notably, JoyAI-Video-Edit also provides a new technical path for large-scale synthesis of embodied intelligence data. Robot training requires a large amount of video about object grasping, moving, and operation, but real machine data collection is costly, and dangerous or rare scenarios are difficult to collect repeatedly.

With controllable editing capabilities over scenes, objects, and visual content, JoyAI-Video-Edit can convert manually operated videos into mechanical hand or arm operation materials. It can retain the position of objects, spatial relationships, and action trajectories while replacing scenes, objects, and robot forms, expanding a single video into more training samples. From video creation to robot training data synthesis, the open-sourcing of this model may open up new possibilities across multiple fields.