Overview: Agnes 3.0 Flash와 DeepSeek V4.1 Flash가 1M 컨텍스트 및 초고속 추론 능력으로 오픈AI 생태계를 확장합니다.
1. Agnes 3.0 Flash by Sapiens AI
Released on September 11, 2026, Agnes 3.0 Flash delivers high-throughput reasoning with an output speed of 234.7 tokens per second and an extensive 1.0M token context window. Designed for cost-sensitive enterprise agentic workflows, it costs only $0.05 per 1M input tokens and $0.15 per 1M output tokens ($0.03 blended rate with prompt caching).
2. DeepSeek V4.1 Flash: Open Weights MoE Frontier
Announced on September 10, 2026 under the permissive MIT License, DeepSeek V4.1 Flash features a 552-billion parameter sparse Mixture-of-Experts architecture activating 16 billion parameters per token. Generating at 214.4 tokens per second with a TTFT of 1.21s, it achieves an Intelligence Index of 40 and full vision multimodal capabilities.
3. Leaderboard Integration
Both models are now live on LLMPodium across all 16 locales, with complete side-by-side benchmark comparisons, pricing calculators, and speed metrics.