Model catalog

All models available through RouterLab.

Explore available models, compare API formats, context windows and pricing. Use model identifiers directly in your integrations.

45

available models

Active public catalog

7

catalog categories

Categories derived from the catalog

3

API formats

Routes published in the catalog

6

embedding models

Also visible in this catalog

Quick access

Common starting points

Dedicated embeddings view

45 of 45 models

DeepSeek V4 FlashHigh availability
deepseek-v4-flash
DeepSeekOpen sourceOpenAI-compatibleClaude MessagesToolsReasoningStreaming

Model origin

DeepSeek

Context

1M

Format

OpenAI-compatible

Max output

384K

Input / 1M

$0.09

Output / 1M

$0.18

Cache read / 1M

$0.018

Cache write / 1M

$0.09

View model
DeepSeek V4 Flash 0731High availability
deepseek-v4-flash-0731
DeepSeekOpen sourceOpenAI-compatibleClaude MessagesToolsReasoningStreaming

Model origin

DeepSeek

Context

1M

Format

OpenAI-compatible

Max output

384K

Input / 1M

$0.09

Output / 1M

$0.18

Cache read / 1M

$0.018

Cache write / 1M

$0.09

View model
DeepSeek V4 ProHigh availability
deepseek-v4-pro
DeepSeekOpen sourceOpenAI-compatibleClaude MessagesToolsReasoningStreaming

Model origin

DeepSeek

Context

1M

Format

OpenAI-compatible

Max output

384K

Input / 1M

$1.25

Output / 1M

$2.55

Cache read / 1M

$0.10

Cache write / 1M

$0.10

View model
DeepSeek V4 Pro 0813High availability
deepseek-v4-pro-0813
DeepSeekOpen sourceOpenAI-compatibleToolsReasoningStreaming

Model origin

DeepSeek

Context

1M

Format

OpenAI-compatible

Max output

384K

Input / 1M

$1.25

Output / 1M

$2.55

Cache read / 1M

$0.10

Cache write / 1M

$0.10

View model
DeepSeek V4.1 FlashHigh availability
deepseek-v4.1-flash
DeepSeekOpen sourceOpenAI-compatibleVisionToolsReasoning

Model origin

DeepSeek

Context

1M

Format

OpenAI-compatible

Max output

384K

Input / 1M

$0.25

Output / 1M

$1.15

Cache read / 1M

$0.006

Cache write / 1M

$0.03

View model
GLM 5.2High availability
glm-5.2
GLMOpen sourceOpenAI-compatibleClaude MessagesToolsReasoningStreaming

Model origin

GLM

Context

1.0M

Format

OpenAI-compatible

Max output

131.1K

Input / 1M

$0.95

Output / 1M

$3.00

Cache read / 1M

$0.18

Cache write / 1M

$0.95

View model
GLM 5.3High availability
glm-5.3
GLMOpen sourceOpenAI-compatibleToolsReasoningStreaming

Model origin

GLM

Context

1.3M

Format

OpenAI-compatible

Max output

262.1K

Input / 1M

$1.40

Output / 1M

$4.40

Cache read / 1M

$0.26

Cache write / 1M

$1.40

View model
GLM 5.3 FlashHigh availability
glm-5.3-flash
GLMOpen sourceOpenAI-compatibleVisionToolsReasoning

Model origin

GLM

Context

1.0M

Format

OpenAI-compatible

Max output

131.1K

Input / 1M

$0.15

Output / 1M

$0.50

Cache read / 1M

$0.03

Cache write / 1M

$0.15

View model
Hy4 PreviewHigh availability
hy4-preview
OpenAIOpen sourceOpenAI-compatibleToolsReasoningStreaming

Model origin

OpenAI

Context

960K

Format

OpenAI-compatible

Max output

64K

Input / 1M

$0.834

Output / 1M

$2.501

Cache read / 1M

$0.042

Cache write / 1M

Same as input

View model
Kimi K2.7 CodeHigh availability
kimi-k2.7-code
KimiOpen sourceOpenAI-compatibleClaude MessagesVisionToolsReasoning

Model origin

Kimi

Context

262.1K

Format

OpenAI-compatible

Max output

32.8K

Input / 1M

$0.70

Output / 1M

$3.45

Cache read / 1M

$0.15

Cache write / 1M

$0.70

View model
Kimi K3High availability
kimi-k3
KimiOpen sourceOpenAI-compatibleVisionToolsReasoning

Model origin

Kimi

Context

1.0M

Format

OpenAI-compatible

Max output

131.1K

Input / 1M

$3.00

Output / 1M

$15.00

Cache read / 1M

$0.30

Cache write / 1M

$3.00

View model
MiniMax M2.7High availability
minimax-m2.7
MiniMaxOpen sourceOpenAI-compatibleClaude MessagesToolsReasoningStreaming

Model origin

MiniMax

Context

204.8K

Format

OpenAI-compatible

Max output

204.8K

Input / 1M

$0.29

Output / 1M

$1.19

Cache read / 1M

$0.29

Cache write / 1M

$0.29

View model
MiniMax M3High availability
minimax-m3
MiniMaxOpen sourceOpenAI-compatibleClaude MessagesToolsReasoningStreaming

Model origin

MiniMax

Context

204.8K

Format

OpenAI-compatible

Max output

204.8K

Input / 1M

$0.28

Output / 1M

$1.15

Cache read / 1M

$0.06

Cache write / 1M

$0.28

View model
Qwen3.8 MaxHigh availability
qwen3.8-max
QwenOpen sourceOpenAI-compatibleVisionToolsReasoning

Model origin

Qwen

Context

991.8K

Format

OpenAI-compatible

Max output

131.1K

Input / 1M

$2.00

Output / 1M

$6.00

Cache read / 1M

$0.25

Cache write / 1M

$2.50

View model
aws-claude-haiku-4-5
AWS ClaudeClaudeClaude MessagesVisionToolsReasoning

Model origin

AWS Claude

Context

200K

Format

Claude Messages

Max output

64K

Input / 1M

$0.60

Output / 1M

$3.00

Cache read / 1M

$0.06

Cache write / 1M

$0.75

View model
aws-claude-opus-5
AWS ClaudeClaudeClaude MessagesVisionToolsReasoning

Model origin

AWS Claude

Context

200K

Format

Claude Messages

Max output

64K

Input / 1M

$3.00

Output / 1M

$15.00

Cache read / 1M

$0.30

Cache write / 1M

$3.75

View model
aws-claude-sonnet-5
AWS ClaudeClaudeClaude MessagesVisionToolsReasoning

Model origin

AWS Claude

Context

200K

Format

Claude Messages

Max output

64K

Input / 1M

$1.20

Output / 1M

$6.00

Cache read / 1M

$0.12

Cache write / 1M

$1.50

View model
BAAI/bge-m3
Open sourceEmbeddingsEmbeddings

Model origin

Open source

Context

8.2K

Format

Embeddings

Max output

Input / 1M

$0.01

Output / 1M

$0.00

Cache read / 1M

Not applicable

Cache write / 1M

Not applicable

View model
claude-fable-5
ClaudeClaudeClaude MessagesVisionToolsReasoning

Model origin

Claude

Context

1M

Format

Claude Messages

Max output

128K

Input / 1M

$10.00

Output / 1M

$50.00

Cache read / 1M

$1.00

Cache write / 1M

$12.50

View model
claude-haiku-4-5
ClaudeClaudeClaude MessagesVisionToolsReasoning

Model origin

Claude

Context

200K

Format

Claude Messages

Max output

64K

Input / 1M

$1.00

Output / 1M

$5.00

Cache read / 1M

$0.10

Cache write / 1M

$1.25

View model
claude-opus-5
ClaudeClaudeClaude MessagesVisionToolsReasoning

Model origin

Claude

Context

1M

Format

Claude Messages

Max output

128K

Input / 1M

$5.00

Output / 1M

$25.00

Cache read / 1M

$0.50

Cache write / 1M

$6.25

View model
claude-sonnet-5
ClaudeClaudeClaude MessagesVisionToolsReasoning

Model origin

Claude

Context

1M

Format

Claude Messages

Max output

128K

Input / 1M

$2.00

Output / 1M

$10.00

Cache read / 1M

$0.20

Cache write / 1M

$2.50

View model
gemini-3.1-pro-preview
GoogleGoogleOpenAI-compatibleVisionToolsReasoning

Model origin

Google

Context

1.0M

Format

OpenAI-compatible

Max output

65.5K

Input / 1M

$4.00

Output / 1M

$18.00

Cache read / 1M

$0.40

Cache write / 1M

$0.40

View model
gemini-3.5-flash
GoogleGoogleOpenAI-compatibleVisionToolsReasoning

Model origin

Google

Context

1.0M

Format

OpenAI-compatible

Max output

65.5K

Input / 1M

$1.50

Output / 1M

$9.00

Cache read / 1M

$0.15

Cache write / 1M

$0.15

View model
gemini-3.6-flash
GoogleGoogleOpenAI-compatibleVisionToolsReasoning

Model origin

Google

Context

1.0M

Format

OpenAI-compatible

Max output

65.5K

Input / 1M

$0.75

Output / 1M

$3.75

Cache read / 1M

$0.075

Cache write / 1M

$0.075

View model
gemini-3.7-flash
GoogleGoogleOpenAI-compatibleVisionToolsReasoning

Model origin

Google

Context

1.0M

Format

OpenAI-compatible

Max output

65.5K

Input / 1M

$0.75

Output / 1M

$3.75

Cache read / 1M

$0.075

Cache write / 1M

$0.075

View model
gemini-3.8-flash
GoogleGoogleOpenAI-compatibleVisionToolsReasoning

Model origin

Google

Context

1.0M

Format

OpenAI-compatible

Max output

65.5K

Input / 1M

$0.75

Output / 1M

$3.75

Cache read / 1M

$0.075

Cache write / 1M

$0.075

View model
gpt-5.6-luna
OpenAIOpenAIOpenAI-compatibleClaude MessagesVisionToolsReasoning

Model origin

OpenAI

Context

922K

Format

OpenAI-compatible

Max output

128K

Input / 1M

$0.20

Output / 1M

$1.20

Cache read / 1M

$0.02

Cache write / 1M

$0.25

View model
gpt-5.6-sol
OpenAIOpenAIOpenAI-compatibleClaude MessagesVisionToolsReasoning

Model origin

OpenAI

Context

922K

Format

OpenAI-compatible

Max output

128K

Input / 1M

$5.00

Output / 1M

$30.00

Cache read / 1M

$0.50

Cache write / 1M

$6.25

View model
gpt-5.6-terra
OpenAIOpenAIOpenAI-compatibleClaude MessagesVisionToolsReasoning

Model origin

OpenAI

Context

922K

Format

OpenAI-compatible

Max output

128K

Input / 1M

$2.00

Output / 1M

$12.00

Cache read / 1M

$0.20

Cache write / 1M

$2.50

View model
gpt-6-astra
OpenAIOpenAIOpenAI-compatibleVisionToolsReasoning

Model origin

OpenAI

Context

922K

Format

OpenAI-compatible

Max output

128K

Input / 1M

$10.00

Output / 1M

$50.00

Cache read / 1M

$1.00

Cache write / 1M

$12.50

View model
grok-4.6
xAI (SpaceX)xAIOpenAI-compatibleVisionToolsReasoning

Model origin

xAI (SpaceX)

Context

500K

Format

OpenAI-compatible

Max output

Input / 1M

$2.00

Output / 1M

$6.00

Cache read / 1M

$0.50

Cache write / 1M

Same as input

View model
Qwen/Qwen3-Embedding-0.6B
QwenEmbeddingsEmbeddings

Model origin

Qwen

Context

32.8K

Format

Embeddings

Max output

Input / 1M

$0.012

Output / 1M

$0.00

Cache read / 1M

Not applicable

Cache write / 1M

Not applicable

View model
Qwen/Qwen3-Embedding-4B
QwenEmbeddingsEmbeddings

Model origin

Qwen

Context

32.8K

Format

Embeddings

Max output

Input / 1M

$0.02

Output / 1M

$0.00

Cache read / 1M

Not applicable

Cache write / 1M

Not applicable

View model
Qwen3-VL-Embedding-2B
QwenEmbeddingsEmbeddingsVision

Model origin

Qwen

Context

32.8K

Format

Embeddings

Max output

Input / 1M

$0.09

Output / 1M

$0.00

Cache read / 1M

Not applicable

Cache write / 1M

Not applicable

View model
text-embedding-3-large
OpenAIEmbeddingsEmbeddings

Model origin

OpenAI

Context

8.2K

Format

Embeddings

Max output

Input / 1M

$0.13

Output / 1M

$0.00

Cache read / 1M

Not applicable

Cache write / 1M

Not applicable

View model
z-image-turbo
OpenAIOpen sourceOpenAI-compatible

Model origin

OpenAI

Context

Format

OpenAI-compatible

Max output

Input / 1M

$0.00

Output / 1M

$0.00

Cache read / 1M

Same as input

Cache write / 1M

Same as input

View model
zembed-1
Open sourceEmbeddingsEmbeddings

Model origin

Open source

Context

32.8K

Format

Embeddings

Max output

Input / 1M

$0.05

Output / 1M

$0.00

Cache read / 1M

Not applicable

Cache write / 1M

Not applicable

View model
deepseek-v4-flash-0731-trial
DeepSeekTrialOpenAI-compatibleToolsReasoningStreaming

Model origin

DeepSeek

Context

1M

Format

OpenAI-compatible

Max output

384K

Input / 1M

$0.01

Output / 1M

$0.01

Cache read / 1M

$0.001

Cache write / 1M

$0.001

View model
DeepSeek V4 Pro 0813Trial variants
deepseek-v4-pro-0813-trial
DeepSeekTrialOpenAI-compatibleToolsReasoningStreaming

Model origin

DeepSeek

Context

1M

Format

OpenAI-compatible

Max output

384K

Input / 1M

$0.01

Output / 1M

$0.01

Cache read / 1M

$0.001

Cache write / 1M

$0.001

View model
Gemini 3.7 FlashTrial variants
gemini-3.7-flash-trial
GoogleTrialOpenAI-compatibleVisionToolsReasoning

Model origin

Google

Context

1.0M

Format

OpenAI-compatible

Max output

65.5K

Input / 1M

$0.01

Output / 1M

$0.01

Cache read / 1M

$0.001

Cache write / 1M

$0.001

View model
GLM 5.3 FlashTrial variants
glm-5.3-flash-trial
GLMTrialOpenAI-compatibleVisionToolsReasoning

Model origin

GLM

Context

1.0M

Format

OpenAI-compatible

Max output

131.1K

Input / 1M

$0.01

Output / 1M

$0.01

Cache read / 1M

$0.001

Cache write / 1M

$0.001

View model
Grok 4.6Trial variants
grok-4.6-trial
xAI (SpaceX)TrialOpenAI-compatibleVisionToolsReasoning

Model origin

xAI (SpaceX)

Context

500K

Format

OpenAI-compatible

Max output

Input / 1M

$0.01

Output / 1M

$0.01

Cache read / 1M

$0.001

Cache write / 1M

$0.001

View model
MiniMax M3Trial variants
MiniMax-M3-trial
MiniMaxTrialOpenAI-compatibleToolsReasoningStreaming

Model origin

MiniMax

Context

204.8K

Format

OpenAI-compatible

Max output

204.8K

Input / 1M

$0.01

Output / 1M

$0.01

Cache read / 1M

$0.001

Cache write / 1M

$0.001

View model
Qwen3.8 MaxTrial variants
qwen3.8-max-trial
QwenTrialOpenAI-compatibleVisionToolsReasoning

Model origin

Qwen

Context

991.8K

Format

OpenAI-compatible

Max output

131.1K

Input / 1M

$0.01

Output / 1M

$0.01

Cache read / 1M

$0.001

Cache write / 1M

$0.001

View model

Need a dedicated embeddings view?

Compare dimensions, context and cost on the specialized page.

Open embeddings view

Next step

Need to understand how a route works?

See how to integrate a model into your application in minutes.