Skip to main content

Overview

Fireworks serves open-weight models — Llama, Mixtral, Qwen, DeepSeek and Yi — behind one fast inference endpoint, so you can pick an open model without hosting it yourself.

Available Models

DeepSeek V3

1 credit • Latest DeepSeek model. Improved performance and capabilities.
  • 64K context window
  • Good reasoning
  • Speed: Fast • Cost: Low
  • Model id: deepseek-v3

DeepSeek R1

1 credit • Specialized reasoning model. Enhanced analytical capabilities.
  • 64K context window
  • Excellent reasoning (reasoning model)
  • Speed: Medium • Cost: Low
  • Model id: deepseek-r1

Llama V3 (8B)

1 credit • Instruction-tuned 8B model. Efficient for general tasks.
  • 8K context window
  • Good reasoning
  • Speed: Very_fast • Cost: Very_low
  • Model id: llama-v3-8b-instruct

Llama V3 (70B)

2 credits • Large instruction-tuned model. Powerful for complex tasks.
  • 8K context window
  • Excellent reasoning
  • Speed: Medium • Cost: Medium
  • Model id: llama-v3-70b-instruct

Llama V3.1 (8B)

1 credit • Updated 8B instruction model. Improved performance and efficiency.
  • 128K context window
  • Good reasoning
  • Speed: Very_fast • Cost: Very_low
  • Model id: llama-v3p1-8b-instruct

Llama V3.1 (70B)

1 credit • Updated 70B instruction model. Enhanced capabilities and performance.
  • 128K context window
  • Excellent reasoning
  • Speed: Fast • Cost: Low
  • Model id: llama-v3p1-70b-instruct

Llama V3.1 (405B)

2 credits • Massive 405B parameter model. Exceptional performance on complex tasks.
  • 128K context window
  • Exceptional reasoning
  • Speed: Slow • Cost: High
  • Model id: llama-v3p1-405b-instruct

Llama V3.3 (70B)

2 credits • Latest 70B model version. Improved instruction following and reasoning.
  • 128K context window
  • Excellent reasoning
  • Speed: Fast • Cost: Medium
  • Model id: llama-v3p3-70b-instruct

Yi Large

2 credits • Large Yi model. Strong performance across diverse applications.
  • 32K context window
  • Excellent reasoning
  • Speed: Medium • Cost: Medium
  • Model id: yi-large

Mixtral (8x7B)

1 credit • Instruction-tuned Mixtral model. Efficient mixture of experts architecture.
  • 32K context window
  • Good reasoning
  • Speed: Fast • Cost: Low
  • Model id: mixtral-8x7b-instruct

Mixtral (8x22B)

2 credits • Larger instruction-tuned Mixtral. Enhanced capabilities with larger capacity.
  • 64K context window
  • Excellent reasoning
  • Speed: Medium • Cost: Medium
  • Model id: mixtral-8x22b-instruct

Qwen 2.5 (72B)

1 credit • Large Qwen model. Advanced capabilities for complex tasks.
  • 128K context window
  • Excellent reasoning
  • Speed: Medium • Cost: Medium
  • Model id: qwen2p5-72b-instruct

Setup

Using BoostGPT-Hosted API Keys

1

Select a model

In your BoostGPT dashboard, choose any Fireworks AI model when creating or configuring an agent.
2

Start chatting

Credits are deducted per step at the rate shown on each model above. Nothing else to configure.

Using Your Own API Key

Bring your own key to bill Fireworks AI directly instead of spending BoostGPT credits. See Bring Your Own Keys.

Next Steps

Model Comparison

Compare every model across all providers

Provider Overview

See all supported providers