Fireworks AI model

Kimi K3 Fast

Kimi K3 is Kimi’s most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, with comparable performance to leading close-source models. It is available on both Fast and Priority serverless tiers, as well as with US-only serverless endpoints for workloads in regulated industries. All Fireworks inference comes with zero data retention enabled by default. → Use Priority for max reliability during congestion; priced at +25% from standard rates. → Use Fast for max speed or latency-sensitive workloads; priced at +50% from standard. To use Fast, switch model ID to the Fast variant: accounts/fireworks/routers/kimi-k3-fast → Use the US-only endpoint for necessary workloads; priced at +10% from standard rates. To do so, switch to model ID: accounts/fireworks/routers/kimi-k3-us

Provider
Fireworks AI
Availability
Available
Model ID
kimi-k3-fast
Snapshot
2026-08-28

Kimi K3 Fast model overview

Kimi K3 Fast is listed in the public Swarm catalog through Fireworks AI. This page summarizes the verified catalog fields used by Swarm without claiming real-time access for every provider account.

Provider
Fireworks AI
Family
Kimi
Context window
1.05M tokens
Thinking support
off · low · high · max
Availability status
Available
Catalog ID
fireworks/kimi-k3-fast
  • Image input
  • Text input
  • Text output
  • Reasoning
  • Tools
  • Fine-tuning
← All Fireworks AI models

Public catalog snapshot. Catalog availability is not a real-time provider status or an access guarantee.