DeepSeek V4 Flash 0731: A Big Step Up in Agent Capabilities
On July 31, DeepSeek announced that the DeepSeek-V4-Flash API is now in public beta. The model name and calling convention are unchanged — you still set the model to deepseek-v4-flash.
The 0731 release keeps the same architecture and size as the previous Preview; it was only re-post-trained. The capability gains, however, are substantial. In the official agent benchmarks, Terminal Bench 2.1 jumped from 56.9 to 82.7, and Toolathlon verified from 51.8 to 70.3. Other scores: Cybergym 76.7, NL2Repo 54.2, DeepSWE 54.4, Agent Last Exam 25.2, DSBench-FullStack 68.7.
Against GPT-5.6 Terra: Terminal Bench 82.7 vs 78.4, Toolathlon 70.3 vs 53.1, DeepSWE 54.4 vs 69.6. Mixed results, but a 284B-parameter MoE model (13B active) trading blows with flagship-tier models says a lot on its own.
The HN thread (670 points) was enthusiastic. One commenter noted that V4-Flash-0731 runs on a single B300, which opens up local inference; others compared it to GLM-5.2 and Claude Opus 4.8, saying it is approaching that tier.
Pricing is unchanged: $0.14/M input tokens (cache miss), $0.28/M output, and $0.0028/M for cache-hit input. Context is 1M tokens with a 384K max output. It natively supports the Responses API, with specific adaptation for Codex.
DeepSeek says the official V4-Pro release will follow soon. This update only affects the Flash API; the V4-Pro API and App/Web models are unchanged.
MiniMax H3: Video Models Go Omni-Modal
On the same day, MiniMax launched H3, which it calls a general-purpose omni-modal generation model. H3 jointly understands text, image, video, and audio inputs, and generates video with native stereo audio at up to 2K resolution and 15 seconds in length.
Instead of the task-split approach of Hailuo 01/02, H3 unifies text-to-image, text-to-video, text-to-audio, and various reference and editing tasks into a single model. The official blog shows an example: "Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3" — H3 handles the multimodal context and generates the video. Accurate text and brand presentation, along with V2V motion transfer, are highlighted strengths.
On pricing, MiniMax says H3's per-second cost at 2K is less than a third of mainstream models, and less than half of mainstream 720p pricing at 768p. Technically, H3 rebuilt its tokenizer (H3-VAE) for a 4x gain in effective sequence length, and uses "In-Context Regeneration" — the base model regenerates its own low-resolution output to 2K using the original multimodal context, instead of a separate super-resolution module.
The bigger news: MiniMax plans to open-source H3's weights within days, subject to applicable laws. Video generation has been dominated by closed models, so if the weights actually ship, the ecosystem could get a lot livelier.
Quick Take
Two releases, one day, different directions. DeepSeek is pushing agent capability deeper while keeping prices aggressive; MiniMax is pushing video generation toward omni-modality and open weights. Both are Chinese vendors closing the capability gap while keeping the pressure on price and openness.




