Fireworks AI model
Kimi K3
Kimi K3 is Kimi’s most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, with comparable performance to leading close-source models. It is available on both Fast and Priority serverless tiers, as well as with US-only serverless endpoints for workloads in regulated industries. All Fireworks inference comes with zero data retention enabled by default. → Use Priority for max reliability during congestion; priced at +25% from standard rates. → Use Fast for max speed or latency-sensitive workloads; priced at +50% from standard. To use Fast, switch model ID to the Fast variant: accounts/fireworks/routers/kimi-k3-fast → Use the US-only endpoint for necessary workloads; priced at +10% from standard rates. To do so, switch to model ID: accounts/fireworks/routers/kimi-k3-us
- Provider
- Fireworks AI
- Availability
- Available
- Model ID
- kimi-k3
- Snapshot
- 2026-08-28
Kimi K3 model overview
Kimi K3 is listed in the public Swarm catalog through Fireworks AI. This page summarizes the verified catalog fields used by Swarm without claiming real-time access for every provider account.
- Provider
- Fireworks AI
- Family
- Kimi
- Context window
- 1.05M tokens
- Thinking support
- off · low · high · max
- Availability status
- Available
- Catalog ID
- fireworks/kimi-k3
Public catalog snapshot. Catalog availability is not a real-time provider status or an access guarantee.