August 21st news: DeepSeek has carried out intensive iteration on its large model toolchain, officially upgrading DeepSeek Harness to version 0.1.1-rc.1. The biggest highlight of this update is the addition of a brand-new visual model experience version - V4Flash Vision-Exp, built upon the existing DeepSeek V4Flash and V4Pro models.

The v0.1.0-rc.8 version released the day before already enabled DeepSeek Harness to access multimodal core capabilities. By supporting native image requests, key commands such as /goal and /plan now support mixed input of text and images. This means users can now send work screenshots and specific requirements at the same time, allowing AI agents (Agent) to directly "work by looking at images," greatly simplifying the development and operation process.
Although DeepSeek previously provided a good image recognition mode on the web, integrating it deeply and applying it to the development toolchain marks the completion of the last puzzle in multimodal processing. For developers, they can simply upgrade by specifying an npm command to experience this new visual model capability that combines speed and accuracy.




