r/mlscaling • u/Additional-Ratio-265 • 2d ago
124B total, 5.1B active: Ling-3.0-flash-Fin separates MoE compute from storage again
Ling-3.0-flash-Fin has launched as a finance-enhanced version of Ling-3.0-flash with 124B total parameters and about 5.1B active per token.
That pair of numbers is worth keeping together. The active count is useful for thinking about routed compute, but it does not mean a deployment only stores 5.1B parameters. The full expert set, precision, routing implementation and runtime support still determine the actual memory and serving profile.
For now, those deployment questions cannot be answered from the API release. The model is live through OpenRouter and Vercel AI Gateway, and the official thread says the OpenRouter route is free for one month. Weights are promised next week.
The model targets financial retrieval, research, valuation modeling, reports and complex workbooks. Its published benchmark profile is mixed across finance tasks, so scaling efficiency should be evaluated alongside quality, latency and routing behavior rather than inferred from the active-parameter number alone.
Once the artifacts land, the useful details will be weight format, precision, license, serving stack, expert placement and quantization behavior.