Provider model hub
Fireworks AI
AI models.
Browse every Fireworks AI model in the public Swarm catalog, with dedicated pages for capabilities, context, thinking controls, and catalog identity.
- Models
- 23
- Provider ID
- fireworks
- Catalog total
- 688
- Snapshot
- 2026-08-28
All Fireworks AI models
23 crawlable model pages
DeepSeek-V4-Flash-0731
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.
- Context
- 1M tokens
- Thinking
- off · high · xhigh
DeepSeek-V4-Pro
DeepSeek-V4-Pro is a flagship open-source Mixture-of-Experts model designed for frontier reasoning, advanced coding, and long-context intelligence at scale (up to 1M tokens). It introduces a hybrid attention architecture that dramatically improves long-context efficiency while reducing KV and compute overhead, along with stability and training enhancements for deep multi-step reasoning. It represents a top-tier open-source system for complex agentic workflows, high-precision reasoning, and demanding production workloads.
- Context
- 1M tokens
- Thinking
- off · high · xhigh
DeepSeek-V4-Pro-0813
DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached.
- Context
- 1M tokens
- Thinking
- off · high · xhigh
GLM 5.2
GLM-5.2 introduces a robust 1M-token context and advanced, multi-effort coding capabilities to significantly enhance performance on long-horizon tasks. Its new IndexShare architecture and improved MTP layer simultaneously boost efficiency by reducing per-token FLOPs and increasing speculative decoding lengths.
- Context
- 1.05M tokens
- Thinking
- off · high · xhigh
GLM 5.2 Fast
GLM-5.2 introduces a robust 1M-token context and advanced, multi-effort coding capabilities to significantly enhance performance on long-horizon tasks. Its new IndexShare architecture and improved MTP layer simultaneously boost efficiency by reducing per-token FLOPs and increasing speculative decoding lengths.
- Context
- 1.05M tokens
- Thinking
- off · high · xhigh
OpenAI gpt-oss-120b
Welcome to the gpt-oss series, OpenAI's open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. gpt-oss-120b is used for production, general purpose, high reasoning use-cases that fits into a single H100 GPU.
- Context
- 131.1K tokens
- Thinking
- low · medium · high
OpenAI gpt-oss-20b
Welcome to the gpt-oss series, OpenAI's open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. gpt-oss-20b is used for lower latency, and local or specialized use-cases.
- Context
- 131K tokens
- Thinking
- low · medium · high
Inkling
Inkling - Thinking Machines multimodal (audio+vision) MoE. Inkling is the first open-weights model released by Thinking Machines Lab. It is a 975B Mixture-of-Experts with 41B active parameters, trained natively across text, image, and audio. Built as a broad generalist foundation for fine-tuning, with controllable thinking effort.
- Context
- 1.05M tokens
- Thinking
- Not supported
Kimi K2.6
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
- Context
- 262.1K tokens
- Thinking
- Not supported
Kimi K2.6 Fast
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
- Context
- 262.1K tokens
- Thinking
- Not supported
Kimi K2.7 Code
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.
- Context
- 262.1K tokens
- Thinking
- Not supported
Kimi K2.7 Code Fast
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.
- Context
- 262.1K tokens
- Thinking
- Not supported
Kimi K3
Kimi K3 is Kimi’s most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, with comparable performance to leading close-source models. It is available on both Fast and Priority serverless tiers, as well as with US-only serverless endpoints for workloads in regulated industries. All Fireworks inference comes with zero data retention enabled by default. → Use Priority for max reliability during congestion; priced at +25% from standard rates. → Use Fast for max speed or latency-sensitive workloads; priced at +50% from standard. To use Fast, switch model ID to the Fast variant: accounts/fireworks/routers/kimi-k3-fast → Use the US-only endpoint for necessary workloads; priced at +10% from standard rates. To do so, switch to model ID: accounts/fireworks/routers/kimi-k3-us
- Context
- 1.05M tokens
- Thinking
- off · low · high · max
Kimi K3 Fast
Kimi K3 is Kimi’s most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, with comparable performance to leading close-source models. It is available on both Fast and Priority serverless tiers, as well as with US-only serverless endpoints for workloads in regulated industries. All Fireworks inference comes with zero data retention enabled by default. → Use Priority for max reliability during congestion; priced at +25% from standard rates. → Use Fast for max speed or latency-sensitive workloads; priced at +50% from standard. To use Fast, switch model ID to the Fast variant: accounts/fireworks/routers/kimi-k3-fast → Use the US-only endpoint for necessary workloads; priced at +10% from standard rates. To do so, switch to model ID: accounts/fireworks/routers/kimi-k3-us
- Context
- 1.05M tokens
- Thinking
- off · low · high · max
MiniMax M2.7
Mixture-of-Experts language model. M2.7 is capable of building complex agent harnesses and completing highly elaborate productivity tasks, leveraging Agent Teams, complex Skills, and dynamic tool search.
- Context
- 196.6K tokens
- Thinking
- low · medium · high
Minimax M3
MiniMax-M3 is a native multimodal model with 512K context running ~428B parameters and ~23B activated parameters. It brings native multimodality. enabling deeper semantic fusion across text, image, and video. M3 also introduces MiniMax Sparse Attention (MSA) to improve long context efficiency, achieving frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork.
- Context
- 512K tokens
- Thinking
- Not supported
Muse Glimmer 30B
Muse Glimmer 30B is a dense causal language model distilled from Muse Spark and purpose-built for autonomous agentic work. It combines multi-step reasoning, reliable schema-based tool calling, and failure recovery with multimodal understanding via a ~1.8B ViT-G/14 perception encoder, supporting interleaved text and image input, a 131K+ context window, and selectable reasoning strength (low through xhigh). Trained on data from over 100 languages, Muse Glimmer performs strongly for its size class on agentic benchmarks including MCP Atlas, DeepSearch QA, Gaia2 and SWE-Bench Pro, and is released under Apache 2.0.
- Context
- 131.1K tokens
- Thinking
- Not supported
NVIDIA Nemotron 3 Ultra NVFP4
Nemotron-3-Ultra-550B-A55B-NVFP4 is a frontier-scale large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for the most demanding workloads, including complex multi-step agents, long-context analysis, and high-accuracy reasoning over code, math, and science. The model employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. Like the Super model, the Ultra model incorporates Multi-Token Prediction (MTP) layers for faster text generation and improved quality, and it is trained using an NVFP4 pre-training recipe to maximize compute efficiency. The model has 55B active parameters and 550B parameters in total.
- Context
- 262.1K tokens
- Thinking
- Not supported
Nemotron Lightning 3.5 30B A3B
Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.
- Context
- 262.1K tokens
- Thinking
- Not supported
Qwen3 Embedding 8B
The Qwen3 Embedding 8B model is the latest proprietary model of the Qwen family, specifically designed for text embedding tasks. This model inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills building upon the dense foundational models of the Qwen3 series. The model represents significant advancements in multiple text embedding tasks including text retrieval, code retrieval, text classification, text clustering.
- Context
- 41K tokens
- Thinking
- Not supported
Qwen3 Reranker 8B
significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining
- Context
- 41K tokens
- Thinking
- Not supported
Qwen3.7 Plus
Qwen 3.7 Plus is Alibaba's latest flagship closed model, available exclusively through Fireworks outside of Alibaba's own infrastructure. For dedicated instances or fine-tuning, please contact the Fireworks team.
- Context
- 262.1K tokens
- Thinking
- off · low · medium · high
Qwen3.8-2.4T-A95B
Qwen3.8-2.4T-A95B is Alibaba's most capable Qwen model to date, a 2.4T-parameter sparse MoE with ~95B active parameters. It is built for autonomous, long-horizon work: multi-day coding runs, research-paper reproduction and self-improvement.
- Context
- 262.1K tokens
- Thinking
- off · low · medium · high
Public catalog snapshot. A listed model does not guarantee real-time provider or account availability.