MiMo V2 Flash
Xiaomi · text, image, video, audio → text
MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion parameters and 15 billion activated parameters. It is designed for high-speed inference and proxy workflows, adopting a novel hybrid attention architecture and multi-token prediction (MTP), significantly reducing inference costs while achieving state-of-the-art performance.
Input$0.1918 /M
Output$0.5754 /M
Cache read$0.0384 /M
MiMo V2 Flash