Skip to main content

Overview

Reasoning models use explicit step-by-step thinking to solve complex problems. Unlike standard models that generate immediate responses, reasoning models “think” through problems methodically, making them ideal for analytical tasks.

What Are Reasoning Models?

Reasoning models employ a different approach:
  1. Explicit Thinking - Show their thought process
  2. Multi-Step Analysis - Break problems into steps
  3. Self-Correction - Refine answers progressively
  4. Higher Token Usage - Require more tokens (2000+ minimum)
  5. Slower Response - Take longer but more accurate
Think of reasoning models as “showing their work” like in math class - they explain how they reached the answer, not just what the answer is.

Available Reasoning Models

A reasoning model spends tokens thinking before it answers, so it needs a higher max_reply_tokens than a standard model — each one below declares its own minimum. Credits are charged per step.

OpenAI

GPT-5.6 Sol

6 credits • The frontier GPT-5.6 model, built for complex professional work requiring the highest reasoning capability.
  • 1.05M context window
  • Needs at least 3000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: gpt-5.6-sol

GPT-5.6 Terra

3 credits • Balanced GPT-5.6 tier offering strong reasoning and multimodal input at moderate cost.
  • 1.05M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Medium
  • Vision: images supported
  • Model id: gpt-5.6-terra

GPT-5.6 Luna

1 credit • The most cost-efficient GPT-5.6 tier, designed for high-volume, cost-sensitive workloads.
  • 1.05M context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Low
  • Vision: images supported
  • Model id: gpt-5.6-luna

GPT-5.5

6 credits • Previous-generation frontier GPT model with strong reasoning and a 1M+ context window.
  • 1.05M context window
  • Needs at least 3000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: gpt-5.5

GPT-5.4

6 credits • Previous-generation flagship GPT model with strong reasoning, creativity, and contextual understanding.
  • 1M context window
  • Needs at least 3000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: gpt-5.4

GPT-5.2

5 credits • The model for coding and agentic tasks across industries.
  • 400K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: gpt-5.2

GPT-5.1

4 credits • An enhanced version of GPT-5 with improved reasoning, creativity, and contextual understanding for the most demanding applications.
  • 200K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: gpt-5.1

GPT-5

5 credits • The flagship GPT-5 model offering state-of-the-art performance in reasoning, creativity, and contextual understanding for the most demanding applications.
  • 200K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: gpt-5

GPT-5.4 Mini

3 credits • The mid-size GPT-5.4 tier: strong reasoning with a 400K context at a fraction of the flagship cost.
  • 400K context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Medium
  • Vision: images supported
  • Model id: gpt-5.4-mini

GPT-5.4 Nano

2 credits • The smallest GPT-5.4 tier, for high-volume work where latency and cost matter more than depth.
  • 400K context window
  • Needs at least 2000 completion tokens
  • Speed: Very fast • Cost: Low
  • Vision: images supported
  • Model id: gpt-5.4-nano

GPT-5 Mini

3 credits • A faster and more affordable GPT-5 variant optimized for everyday use while maintaining strong reasoning and contextual understanding.
  • 200K context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Medium
  • Vision: images supported
  • Model id: gpt-5-mini

GPT-5 Nano

2 credits • Ultra-lightweight GPT-5 optimized for speed and scale. Perfect for simple tasks, prototyping, and high-volume applications.
  • 128K context window
  • Needs at least 2000 completion tokens
  • Speed: Very_fast • Cost: Low
  • Vision: images supported
  • Model id: gpt-5-nano

Google

Gemini 3.7 Flash

2 credits • The newest Gemini Flash model, with a 1M-token context window and native multimodal input.
  • 1.05M context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Medium
  • Vision: images supported
  • Model id: gemini-3.7-flash

Gemini 3.6 Flash

2 credits • Balances speed with intelligence for strong performance on agentic and multimodal tasks.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Medium
  • Vision: images supported
  • Model id: gemini-3.6-flash

Gemini 3.5 Flash

3 credits • Most intelligent Gemini model for sustained frontier performance on agentic and coding tasks.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Medium
  • Vision: images supported
  • Model id: gemini-3.5-flash

Gemini 3.5 Flash Lite

1 credit • The fastest, most cost-effective Gemini 3.5 model for high-throughput execution.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Very_fast • Cost: Very_low
  • Vision: images supported
  • Model id: gemini-3.5-flash-lite

Gemini 3.1 Flash Lite

1 credit • Frontier-class performance at a fraction of the cost, optimized for low latency and high-volume tasks.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Very_fast • Cost: Very_low
  • Vision: images supported
  • Model id: gemini-3.1-flash-lite

Gemini 3.1 Pro

3 credits • Reasoning-first Gemini model for complex agentic workflows and coding, with adaptive thinking.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: gemini-3.1-pro-preview

Gemini 3 Flash

1 credit • The latest Gemini 3 Flash model in preview, optimized for speed and efficiency while delivering strong reasoning and language capabilities.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Medium
  • Vision: images supported
  • Model id: gemini-3-flash-preview

Gemini 2.5 Flash

1 credit • A fast and capable model in the Gemini 2.5 series, balancing speed and reasoning performance for production use.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Medium
  • Vision: images supported
  • Model id: gemini-2.5-flash

Gemini 2.5 Flash Lite

1 credit • A lightweight, cost-efficient variant of Gemini 2.5 Flash, optimized for low latency and basic tasks.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Very_fast • Cost: Very_low
  • Vision: images supported
  • Model id: gemini-2.5-flash-lite

Anthropic

Claude Fable 5

10 credits • Anthropic’s most capable model, for the most demanding reasoning and long-running agentic work.
  • 1M context window
  • Needs at least 3000 completion tokens
  • Speed: Slow • Cost: High
  • Vision: images supported
  • Model id: claude-fable-5

Claude Opus 5

5 credits • Anthropic’s model for complex agentic coding and enterprise work, with a 1M-token context window.
  • 1M context window
  • Needs at least 3000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: claude-opus-5

Claude Opus 4.8

5 credits • The most capable Claude model, built for complex agentic coding and enterprise work.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: claude-opus-4-8

Claude Sonnet 5

3 credits • The best combination of speed and intelligence, with near-Opus quality on coding and agentic work.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Medium
  • Vision: images supported
  • Model id: claude-sonnet-5

Claude Opus 4.7

5 credits • Highly autonomous Claude Opus model, strong on long-horizon agentic work, vision, and memory.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: claude-opus-4-7

Claude Opus 4.6

5 credits • Claude Opus model with superior reasoning, creativity, tool use, and contextual understanding.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: claude-opus-4-6

Claude Sonnet 4.6

3 credits • Fast Claude Sonnet model with strong reasoning, creativity, tool use, and contextual understanding.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Medium
  • Vision: images supported
  • Model id: claude-sonnet-4-6

xAI

Grok 4.6

1 credit • The newest Grok model, with a 500K context window and strong multimodal reasoning.
  • 500K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Medium
  • Vision: images supported
  • Model id: grok-4.6

Grok 4.5

1 credit • The flagship Grok model, with native video input and strong multimodal reasoning.
  • 500K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Medium
  • Vision: images supported
  • Model id: grok-4.5

Grok 4.3

1 credit • A versatile 1M-context model offering strong language and reasoning capabilities at low cost.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Low
  • Vision: images supported
  • Model id: grok-4.3

Grok 4.20 Reasoning

1 credit • Grok 4.20 in reasoning mode, with enhanced analytical and problem-solving capabilities.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Low
  • Vision: images supported
  • Model id: grok-4.20-0309-reasoning

Zhipu AI

GLM-5.3

1 credit • Zhipu’s flagship for complex software engineering and long-horizon agent work, with a 1M-token context window.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Low
  • Model id: glm-5.3

GLM-5.3 Flash

1 credit • The fast, low-cost GLM-5.3 tier, keeping the 1M-token context window.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Low
  • Model id: glm-5.3-flash

GLM-5.2

1 credit • The previous GLM flagship, with a 1M-token context window and strong long-horizon performance.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Low
  • Model id: glm-5.2

GLM-4.7

1 credit • Newer GLM-4 generation model with adaptive reasoning for complex agentic and coding tasks.
  • 200K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Low
  • Model id: glm-4.7

GLM-5.1

1 credit • Zhipu’s frontier GLM-5 model, a large MoE for the most demanding reasoning and agentic work.
  • 200K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Low
  • Model id: glm-5.1

MiniMax

MiniMax M2.7 Highspeed

1 credit • The accelerated M2.7 tier — roughly 100 tokens/sec against the standard 60.
  • 205K context window
  • Needs at least 2000 completion tokens
  • Speed: Very fast • Cost: Low
  • Model id: MiniMax-M2.7-highspeed

MiniMax M2.5 Highspeed

1 credit • The accelerated M2.5 tier — roughly 100 tokens/sec against the standard 60.
  • 205K context window
  • Needs at least 2000 completion tokens
  • Speed: Very fast • Cost: Low
  • Model id: MiniMax-M2.5-highspeed

MiniMax M3

1 credit • MiniMax’s flagship model with a 1M-token context window and native multimodal (image) input.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Low
  • Vision: images supported
  • Model id: MiniMax-M3

MiniMax M2.7

1 credit • Agentic MiniMax M2 model tuned for coding and reasoning, with a 205K context window.
  • 205K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Low
  • Model id: MiniMax-M2.7

MiniMax M2.5

1 credit • Cost-efficient MiniMax M2 tier for general-purpose tasks, with a 205K context window.
  • 205K context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Very_low
  • Model id: MiniMax-M2.5

DeepSeek

DeepSeek V4 Flash

1 credit • Fast, low-cost general-purpose model with a 1M context window. Good for everyday conversations and high-volume tasks.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Low
  • Model id: deepseek-v4-flash

DeepSeek V4 Flash Vision

1 credit • Experimental vision variant of V4 Flash. Same 1M context and pricing; images are billed as input tokens.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: Low
  • Vision: images supported
  • Model id: deepseek-v4-flash-vision-exp

DeepSeek V4 Pro

1 credit • Higher-performance tier for complex reasoning, with enhanced logical and analytical capabilities.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Low
  • Model id: deepseek-v4-pro

Moonshot AI

Kimi K3

4 credits • Moonshot AI’s flagship multimodal reasoning model with a one-million-token context window.
  • 1M context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: High
  • Vision: images supported
  • Model id: kimi-k3

Kimi K2.7 Code

3 credits • A reasoning-first Kimi model optimized for agentic software engineering and complex coding work.
  • 256K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Medium
  • Model id: kimi-k2.7-code

Kimi K2.7 Code Highspeed

4 credits • The high-throughput Kimi K2.7 Code variant for latency-sensitive coding and agent workflows.
  • 256K context window
  • Needs at least 2000 completion tokens
  • Speed: Fast • Cost: High
  • Model id: kimi-k2.7-code-highspeed

Kimi K2.6

3 credits • A multimodal Kimi model for general reasoning, tool use, and long-context tasks.
  • 256K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Medium
  • Vision: images supported
  • Model id: kimi-k2.6

Kimi K2.5

2 credits • A capable multimodal Kimi model for reasoning over text, images, and video inputs.
  • 256K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Medium
  • Vision: images supported
  • Model id: kimi-k2.5

Fireworks AI

DeepSeek R1

1 credit • Specialized reasoning model. Enhanced analytical capabilities.
  • 64K context window
  • Needs at least 2000 completion tokens
  • Speed: Medium • Cost: Low
  • Model id: deepseek-r1

Reasoning Model Comparison

When to Use Reasoning Models

Perfect For

Solving complex math, physics, or engineering problems that require step-by-step work.
Analyzing code to find bugs, understand logic, and suggest improvements.
Business analysis, market research, competitive analysis.Best models: O1, Claude 3 Opus Extended
Research paper analysis, experimental design, data interpretation.Best models: O1, Gemini 2.0 Flash Thinking
Solving riddles, logic games, complex reasoning challenges.Best models: O1 Mini, DeepSeek R1 (budget option)

Not Ideal For

Reasoning models are overkill for these tasks. Use standard models instead:
  • Simple conversations
  • Quick factual questions
  • Creative writing (use GPT-5 or Claude Opus instead)
  • High-volume simple queries (too slow and expensive)
  • Real-time chat applications (too slow)

Usage Examples

OpenAI O1 - Deep Analysis

DeepSeek R1 - Budget Reasoning

Gemini 2.0 Flash Thinking - Large Context

Override Reasoning Mode

All reasoning models support these modes:

Best Practices

Set Appropriate Token Limits

Handle Longer Wait Times

Cost Optimization

Cache Common Analyses

Performance Comparison

Speed vs Accuracy Tradeoff

Best Value: DeepSeek R1 (2 credits) or Gemini 2.0 Flash Thinking (3 credits) offer the best balance of cost and capability.

Troubleshooting

Expected: Reasoning models take longerSolutions:
  • Use O1 Mini instead of O1
  • Use DeepSeek R1 for faster reasoning
  • Try Groq’s reasoning models for ultra-fast
  • Add progress indicators for users
Cause: Reasoning models use more tokens and creditsSolutions:
  • Use only for complex tasks
  • Try DeepSeek R1 (2 credits)
  • Cache results for common queries
  • Use standard models for simple tasks
Note: Some models show thinking, others don’tDetails:
  • O1 series: Shows detailed thinking
  • DeepSeek R1: Shows reasoning steps
  • Claude Extended: Implicit thinking
  • Gemini Thinking: Shows analysis process
Cause: Insufficient max_reply_tokensSolutions:
  • Set minimum 2000 tokens
  • Use 3000+ for O1
  • Use 2500+ for Claude Extended
  • Check model-specific requirements

Real-World Use Cases

Code Review Assistant

Business Strategy Advisor

Math Tutor

Next Steps

Model Comparison

Compare all models

Provider Overview

See all providers

OpenAI O-Series

Learn about O1 models

DeepSeek R1

Budget reasoning option