DeepSeek-V4-Flash-0731: 284 billion parameters, 1M context and an MIT licence
A mixture-of-experts model with 284B total parameters and only 13B active per token, a one-million-token window, and weights published with no access gating. The agentic gains are large: 82.7 vs 72.1 on Terminal Bench 2.1 and 54.4 vs 12.8 on DeepSWE.