HelloStackHelloStack
Fireworks AI logo

Own Your Specialized Intelligence

About Fireworks AI

Fireworks AI is a drop-in replacement for closed-model APIs that routes to the best open or closed model for each task. It serves developers and ML teams who want to own their AI stack and reduce spending while preserving performance.

Fireworks AI screenshot

Key Features

Engineers juggle multiple models and want to avoid vendor lock while controlling costs. Nexus lets you route tasks to the best open or closed model and swap on the fly to balance cost and accuracy:

Model Routing Engine

Routes each task to the best open or closed model, with dynamic re-routing as needs change.

Drop-in API Replacement

Works as a direct substitute for existing closed-model APIs, so you can switch without rewriting client code.

Cost Reduction

Aims for a 50–75% cut in AI coding spend by steering queries to cheaper models when suitable.

Specialized Intelligence via SoTA Training & Inference

Transforms open models into your tailored capabilities through state-of-the-art training and inference.

Summary

Best for ML/AI engineers, platform teams embedding AI in apps, and data science groups aiming to optimize model usage.

Base model (Less than 4B parameters)

$0.10 / per 1M tokens

Base model (4B - 16B parameters)

$0.20 / per 1M tokens

Base model (More than 16B parameters)

$0.90 / per 1M tokens

MoE model (0B - 56B parameters)

$0.50 / per 1M tokens

MoE model (56.1B - 176B parameters)

$1.20 / per 1M tokens

DeepSeek V3 family

$0.56 input, $1.68 output / per 1M tokens

DeepSeek R1 0528

$1.35 input, $5.4 output / per 1M tokens

GLM-4.5, GLM-4.6

$0.55 input, $2.19 output / per 1M tokens

Meta Llama 3.1 405B

$3.00 / per 1M tokens

Meta Llama 4 Maverick (Basic)

$0.22 input, $0.88 output / per 1M tokens

Meta Llama 4 Scout (Basic)

$0.15 input, $0.60 output / per 1M tokens

Qwen3 235B Family

$0.22 input, $0.88 output / per 1M tokens

Qwen3 30B, Qwen Coder Flash

$0.15 input, $0.60 output / per 1M tokens

Kimi K2 Instruct, Kimi K2 Thinking

$0.60 input, $2.50 output / per 1M tokens

Qwen3 Coder 480B

$0.45 input, $1.80 output / per 1M tokens

OpenAI gpt-oss-120b

$0.15 input, $0.60 output / per 1M tokens

OpenAI gpt-oss-20b

$0.07 input, $0.30 output / per 1M tokens

Whisper-v3-large

$0.0015 / per audio minute

Whisper-v3-large-turbo

$0.0009 / per audio minute

Streaming ASR v1

$0.0032 / per audio minute

Streaming ASR v2

$0.0035 / per audio minute

All Non-Flux Models

$0.0002 / per image

FLUX.1 [dev]

N/A / per image

FLUX.1 [schnell]

N/A / per image

FLUX.1 Kontext Pro

$0.04 / per image

FLUX.1 Kontext Max

$0.08 / per image

up to 150M

$0.008 / per 1M input tokens

150M - 350M

$0.016 / per 1M input tokens

Qwen3 8B

$0.1 / per 1M input tokens

Supervised Fine Tuning (Models up to 16B parameters)

$0.50 / per 1M training tokens

Supervised Fine Tuning (Models 16.1B - 80B)

$3.00 / per 1M training tokens

Supervised Fine Tuning (Models 80B - 300B)

$6.00 / per 1M training tokens

Supervised Fine Tuning (Models >300B)

$10.00 / per 1M training tokens

On-Demand - A100 80 GB GPU

$2.90 / per hour

On-Demand - H100 80 GB GPU

$4.00 / per hour

On-Demand - H200 141 GB GPU

$6.00 / per hour

On-Demand - B200 180 GB GPU

$9.00 / per hour

More in Inference

See all →