Provider model hub

Fireworks AI
AI models.

Browse every Fireworks AI model in the public Swarm catalog, with dedicated pages for capabilities, context, thinking controls, and catalog identity.

Models
23
Provider ID
fireworks
Catalog total
688
Snapshot
2026-08-28

All Fireworks AI models

23 crawlable model pages

Deepseek · Available

DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.

Context
1M tokens
Thinking
off · high · xhigh
Deepseek · Available

DeepSeek-V4-Pro

DeepSeek-V4-Pro is a flagship open-source Mixture-of-Experts model designed for frontier reasoning, advanced coding, and long-context intelligence at scale (up to 1M tokens). It introduces a hybrid attention architecture that dramatically improves long-context efficiency while reducing KV and compute overhead, along with stability and training enhancements for deep multi-step reasoning. It represents a top-tier open-source system for complex agentic workflows, high-precision reasoning, and demanding production workloads.

Context
1M tokens
Thinking
off · high · xhigh
Deepseek · Available

DeepSeek-V4-Pro-0813

DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached.

Context
1M tokens
Thinking
off · high · xhigh
Glm · Available

GLM 5.2

GLM-5.2 introduces a robust 1M-token context and advanced, multi-effort coding capabilities to significantly enhance performance on long-horizon tasks. Its new IndexShare architecture and improved MTP layer simultaneously boost efficiency by reducing per-token FLOPs and increasing speculative decoding lengths.

Context
1.05M tokens
Thinking
off · high · xhigh
Glm · Available

GLM 5.2 Fast

GLM-5.2 introduces a robust 1M-token context and advanced, multi-effort coding capabilities to significantly enhance performance on long-horizon tasks. Its new IndexShare architecture and improved MTP layer simultaneously boost efficiency by reducing per-token FLOPs and increasing speculative decoding lengths.

Context
1.05M tokens
Thinking
off · high · xhigh
Gpt Oss · Available

OpenAI gpt-oss-120b

Welcome to the gpt-oss series, OpenAI's open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. gpt-oss-120b is used for production, general purpose, high reasoning use-cases that fits into a single H100 GPU.

Context
131.1K tokens
Thinking
low · medium · high
Gpt Oss · Available

OpenAI gpt-oss-20b

Welcome to the gpt-oss series, OpenAI's open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. gpt-oss-20b is used for lower latency, and local or specialized use-cases.

Context
131K tokens
Thinking
low · medium · high
Fireworks AI · Available

Inkling

Inkling - Thinking Machines multimodal (audio+vision) MoE. Inkling is the first open-weights model released by Thinking Machines Lab. It is a 975B Mixture-of-Experts with 41B active parameters, trained natively across text, image, and audio. Built as a broad generalist foundation for fine-tuning, with controllable thinking effort.

Context
1.05M tokens
Thinking
Not supported
Kimi · Available

Kimi K2.6

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

Context
262.1K tokens
Thinking
Not supported
Kimi · Available

Kimi K2.6 Fast

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

Context
262.1K tokens
Thinking
Not supported
Kimi · Available

Kimi K2.7 Code

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

Context
262.1K tokens
Thinking
Not supported
Kimi · Available

Kimi K2.7 Code Fast

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

Context
262.1K tokens
Thinking
Not supported
Kimi · Available

Kimi K3

Kimi K3 is Kimi’s most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, with comparable performance to leading close-source models. It is available on both Fast and Priority serverless tiers, as well as with US-only serverless endpoints for workloads in regulated industries. All Fireworks inference comes with zero data retention enabled by default. → Use Priority for max reliability during congestion; priced at +25% from standard rates. → Use Fast for max speed or latency-sensitive workloads; priced at +50% from standard. To use Fast, switch model ID to the Fast variant: accounts/fireworks/routers/kimi-k3-fast → Use the US-only endpoint for necessary workloads; priced at +10% from standard rates. To do so, switch to model ID: accounts/fireworks/routers/kimi-k3-us

Context
1.05M tokens
Thinking
off · low · high · max
Kimi · Available

Kimi K3 Fast

Kimi K3 is Kimi’s most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, with comparable performance to leading close-source models. It is available on both Fast and Priority serverless tiers, as well as with US-only serverless endpoints for workloads in regulated industries. All Fireworks inference comes with zero data retention enabled by default. → Use Priority for max reliability during congestion; priced at +25% from standard rates. → Use Fast for max speed or latency-sensitive workloads; priced at +50% from standard. To use Fast, switch model ID to the Fast variant: accounts/fireworks/routers/kimi-k3-fast → Use the US-only endpoint for necessary workloads; priced at +10% from standard rates. To do so, switch to model ID: accounts/fireworks/routers/kimi-k3-us

Context
1.05M tokens
Thinking
off · low · high · max
Minimax · Available

MiniMax M2.7

Mixture-of-Experts language model. M2.7 is capable of building complex agent harnesses and completing highly elaborate productivity tasks, leveraging Agent Teams, complex Skills, and dynamic tool search.

Context
196.6K tokens
Thinking
low · medium · high
Minimax · Available

Minimax M3

MiniMax-M3 is a native multimodal model with 512K context running ~428B parameters and ~23B activated parameters. It brings native multimodality. enabling deeper semantic fusion across text, image, and video. M3 also introduces MiniMax Sparse Attention (MSA) to improve long context efficiency, achieving frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork.

Context
512K tokens
Thinking
Not supported
Fireworks AI · Available

Muse Glimmer 30B

Muse Glimmer 30B is a dense causal language model distilled from Muse Spark and purpose-built for autonomous agentic work. It combines multi-step reasoning, reliable schema-based tool calling, and failure recovery with multimodal understanding via a ~1.8B ViT-G/14 perception encoder, supporting interleaved text and image input, a 131K+ context window, and selectable reasoning strength (low through xhigh). Trained on data from over 100 languages, Muse Glimmer performs strongly for its size class on agentic benchmarks including MCP Atlas, DeepSearch QA, Gaia2 and SWE-Bench Pro, and is released under Apache 2.0.

Context
131.1K tokens
Thinking
Not supported
Nemotron · Available

NVIDIA Nemotron 3 Ultra NVFP4

Nemotron-3-Ultra-550B-A55B-NVFP4 is a frontier-scale large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for the most demanding workloads, including complex multi-step agents, long-context analysis, and high-accuracy reasoning over code, math, and science. The model employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. Like the Super model, the Ultra model incorporates Multi-Token Prediction (MTP) layers for faster text generation and improved quality, and it is trained using an NVFP4 pre-training recipe to maximize compute efficiency. The model has 55B active parameters and 550B parameters in total.

Context
262.1K tokens
Thinking
Not supported
Nemotron · Available

Nemotron Lightning 3.5 30B A3B

Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.

Context
262.1K tokens
Thinking
Not supported
Qwen · Available

Qwen3 Embedding 8B

The Qwen3 Embedding 8B model is the latest proprietary model of the Qwen family, specifically designed for text embedding tasks. This model inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills building upon the dense foundational models of the Qwen3 series. The model represents significant advancements in multiple text embedding tasks including text retrieval, code retrieval, text classification, text clustering.

Context
41K tokens
Thinking
Not supported
Qwen · Available

Qwen3 Reranker 8B

significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining

Context
41K tokens
Thinking
Not supported
Qwen · Available

Qwen3.7 Plus

Qwen 3.7 Plus is Alibaba's latest flagship closed model, available exclusively through Fireworks outside of Alibaba's own infrastructure. For dedicated instances or fine-tuning, please contact the Fireworks team.

Context
262.1K tokens
Thinking
off · low · medium · high
Qwen · Available

Qwen3.8-2.4T-A95B

Qwen3.8-2.4T-A95B is Alibaba's most capable Qwen model to date, a 2.4T-parameter sparse MoE with ~95B active parameters. It is built for autonomous, long-horizon work: multi-day coding runs, research-paper reproduction and self-improvement.

Context
262.1K tokens
Thinking
off · low · medium · high
← Compare all model providers

Public catalog snapshot. A listed model does not guarantee real-time provider or account availability.