DeepSeek has officially open-sourced the first experimental multimodal model of its V4 series - DeepSeek-V4-Flash-Vision-Exp. This model is deeply integrated with a vision module based on the existing V4-Flash underlying architecture, enabling it to have comprehensive deep image understanding and analysis capabilities on top of its strong text processing capabilities.

In terms of core technical parameters, the experimental model has a total parameter scale of 30.5 billion. By combining the vision module with the text architecture, it can not only accurately interpret various types of image content but also independently complete complex agent (intelligent entity) tasks by seamlessly integrating external tools. Official evaluation data shows that the model performs at a comparable level to the previous V4-Flash in pure text agent tasks, while demonstrating strong competitiveness in multimodal comprehensive tests. Its actual multimodal agent capabilities have shown outstanding performance in multiple industry tests, approaching the top level of the industry.

To further promote the development of the open-source ecosystem and the popularization of technology, DeepSeek has chosen the open and friendly MIT license for this experimental multimodal model. This move allows global developers, research institutions, and enterprises to conduct secondary development, commercial integration, and technological innovation more freely under compliance conditions, which is expected to trigger a new wave of multimodal application deployment in the open-source community.