About Fireworks AI
Fireworks AI is a drop-in replacement for closed-model APIs that routes to the best open or closed model for each task. It serves developers and ML teams who want to own their AI stack and reduce spending while preserving performance.

Key Features
Engineers juggle multiple models and want to avoid vendor lock while controlling costs. Nexus lets you route tasks to the best open or closed model and swap on the fly to balance cost and accuracy:
Model Routing Engine
Routes each task to the best open or closed model, with dynamic re-routing as needs change.
Drop-in API Replacement
Works as a direct substitute for existing closed-model APIs, so you can switch without rewriting client code.
Cost Reduction
Aims for a 50–75% cut in AI coding spend by steering queries to cheaper models when suitable.
Specialized Intelligence via SoTA Training & Inference
Transforms open models into your tailored capabilities through state-of-the-art training and inference.
Summary
Best for ML/AI engineers, platform teams embedding AI in apps, and data science groups aiming to optimize model usage.
Pricing
View pricingBase model (Less than 4B parameters)
$0.10 / per 1M tokens
Base model (4B - 16B parameters)
$0.20 / per 1M tokens
Base model (More than 16B parameters)
$0.90 / per 1M tokens
MoE model (0B - 56B parameters)
$0.50 / per 1M tokens
MoE model (56.1B - 176B parameters)
$1.20 / per 1M tokens
DeepSeek V3 family
$0.56 input, $1.68 output / per 1M tokens
DeepSeek R1 0528
$1.35 input, $5.4 output / per 1M tokens
GLM-4.5, GLM-4.6
$0.55 input, $2.19 output / per 1M tokens
Meta Llama 3.1 405B
$3.00 / per 1M tokens
Meta Llama 4 Maverick (Basic)
$0.22 input, $0.88 output / per 1M tokens
Meta Llama 4 Scout (Basic)
$0.15 input, $0.60 output / per 1M tokens
Qwen3 235B Family
$0.22 input, $0.88 output / per 1M tokens
Qwen3 30B, Qwen Coder Flash
$0.15 input, $0.60 output / per 1M tokens
Kimi K2 Instruct, Kimi K2 Thinking
$0.60 input, $2.50 output / per 1M tokens
Qwen3 Coder 480B
$0.45 input, $1.80 output / per 1M tokens
OpenAI gpt-oss-120b
$0.15 input, $0.60 output / per 1M tokens
OpenAI gpt-oss-20b
$0.07 input, $0.30 output / per 1M tokens
Whisper-v3-large
$0.0015 / per audio minute
Whisper-v3-large-turbo
$0.0009 / per audio minute
Streaming ASR v1
$0.0032 / per audio minute
Streaming ASR v2
$0.0035 / per audio minute
All Non-Flux Models
$0.0002 / per image
FLUX.1 [dev]
N/A / per image
FLUX.1 [schnell]
N/A / per image
FLUX.1 Kontext Pro
$0.04 / per image
FLUX.1 Kontext Max
$0.08 / per image
up to 150M
$0.008 / per 1M input tokens
150M - 350M
$0.016 / per 1M input tokens
Qwen3 8B
$0.1 / per 1M input tokens
Supervised Fine Tuning (Models up to 16B parameters)
$0.50 / per 1M training tokens
Supervised Fine Tuning (Models 16.1B - 80B)
$3.00 / per 1M training tokens
Supervised Fine Tuning (Models 80B - 300B)
$6.00 / per 1M training tokens
Supervised Fine Tuning (Models >300B)
$10.00 / per 1M training tokens
On-Demand - A100 80 GB GPU
$2.90 / per hour
On-Demand - H100 80 GB GPU
$4.00 / per hour
On-Demand - H200 141 GB GPU
$6.00 / per hour
On-Demand - B200 180 GB GPU
$9.00 / per hour