Benchmarks? It is sceptical.
I appreciate WebBrain publishing the DeepSeek-V4-Flash-0731 vision overlay and component hashes. However, the current public artifacts appear to establish packaging rather than validated multimodal capability.
VISION_ADAPTER_MANIFEST.json currently reports gpu_validated_for_this_0731_package=false, and the model card states that a fresh full-GPU loader/startup and image-generation smoke test has not yet been rerun for this exact assembled checkpoint. The repository also does not provide training-step/data/loss logs, checkpoint-resume evidence, or paired visual evaluations such as ScreenSpot vision-vs-blind/shuffled, TextVQA, DocVQA, or OCRBench.
In my version MoonViT to DeepSeek V4 Flash 0731 Vision projector, I didn't see the vision benchmark works well: https://huggingface.co/cyjin-yl/DeepSeek-V4-Flash-0731-Vision and https://github.com/cyjin-yl/moonvit-deepseek-v4-glue
Could you clarify whether this exact 0731 checkpoint has been independently loaded and tested with real images? If so, publishing the command, hardware/software environment, and raw sample outputs would make the result much easier to verify. Otherwise, it may be more accurate to label it as an experimental, unvalidated vision overlay rather than a working DeepSeek vision model.