No models found
Try a different search term, or broaden your search by removing filters.
claude-fable-5
Claude Fable 5 is Anthropic's most capable widely released model, built for the most demanding reasoning and long-horizon agentic work. Adaptive thinking is always on, and the model supports a 1M token context window with up to 128k output tokens per request.
- Third-party
- Context: 1M tokens
- Maximum output: 128K tokens
- Pricing listed
claude-fable-5.1
Claude Fable 5.1 is Anthropic's next model in the Fable family, with improvements in agentic coding, long-running agentic workflows, knowledge work, front-end and visual code generation, and finance and analysis tasks. It supports adaptive thinking and a 1M token context window.
- Third-party
- Context: 1M tokens
- Maximum output: 128K tokens
- Pricing listed
claude-haiku-4.5
Claude Haiku 4.5 delivers similar levels of coding performance at one-third the cost and more than twice the speed of larger models.
- Third-party
- Zero data retention
- Context: 200K tokens
- Maximum output: 8.2K tokens
- Pricing listed
claude-opus-4.5
Claude Opus 4.5 brings further reasoning, coding, and agentic improvements over Opus 4.1, with stronger tool use and tighter instruction following.
- Third-party
- Zero data retention
- Context: 200K tokens
- Maximum output: 32K tokens
- Pricing listed
claude-opus-4.6
Claude Opus 4.6 is Anthropic's flagship language model built for complex, multi-step work in coding, financial analysis, and legal reasoning. It uses extended thinking to work through complex problems carefully and features a one million token context window.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 128K tokens
- Pricing listed
claude-opus-4.7
Claude Opus 4.7 is Anthropic's most capable generally available model, with a step-change improvement in agentic coding over Claude Opus 4.6. It uses adaptive thinking to calibrate reasoning per task and supports a one million token context window at standard pricing.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 128K tokens
- Pricing listed
claude-opus-4.8
Claude Opus 4.8 is Anthropic's most capable generally available model, with a step-change improvement in agentic coding over Claude Opus 4.7. It uses adaptive thinking to calibrate reasoning per task and supports a one million token context window at standard pricing.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 128K tokens
- Pricing listed
claude-opus-5
Claude Opus 5 is Anthropic's model for complex agentic coding and enterprise work, delivering intelligence close to Claude Fable 5 at half the price. It uses adaptive thinking to calibrate reasoning per task and supports a one million token context window at standard pricing. Unlike Fable 5, Opus 5 has no data retention requirements for general access.
- Third-party
- Context: 1M tokens
- Maximum output: 128K tokens
- Pricing listed
claude-sonnet-4.5
Claude Sonnet 4.5 is the best coding model to date, with significant improvements across the entire development lifecycle.
- Third-party
- Zero data retention
- Context: 200K tokens
- Maximum output: 8.2K tokens
- Pricing listed
claude-sonnet-4.6
Claude Sonnet 4.6 is Anthropic's latest balanced model offering strong coding, reasoning, and agentic capabilities with improved instruction following.
- Third-party
- Zero data retention
- Context: 200K tokens
- Maximum output: 128K tokens
- Pricing listed
claude-sonnet-5
Claude Sonnet 5 is Anthropic's most agentic Sonnet model yet, built for coding, tool use, reasoning, and long-horizon professional work at lower cost than Opus-class models.
- Third-party
- Context: 1M tokens
- Maximum output: 128K tokens
- Pricing listed
deepseek-v4-pro
DeepSeek V4 Pro is a high-capability reasoning model from DeepSeek, served via Fireworks infrastructure for production-grade inference.
- Third-party
- Context: 131.1K tokens
- Pricing listed
gemini-2.5-flash
Google's fast multimodal Gemini 2.5 model with strong reasoning and a 1M token context window.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 8.2K tokens
- Pricing listed
gemini-2.5-flash-lite
Google's lightest and most cost-efficient Gemini 2.5 model for high-throughput tasks.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 8.2K tokens
- Pricing listed
gemini-2.5-pro
Google's most capable Gemini 2.5 model with strong reasoning, thinking support, and a 1M token context window.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 65.5K tokens
- Pricing listed
gemini-3-flash
Gemini 3 Flash is Google's fast multimodal model with frontier intelligence, superior search, and grounding capabilities.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 8.2K tokens
- Pricing listed
gemini-3.1-flash-lite
Google's lightest and most cost-efficient Gemini model for high-throughput tasks.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 8.2K tokens
- Pricing listed
gemini-3.1-pro
Google's most intelligent Gemini model with improved reasoning, a medium thinking level, and a 1M token context window.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 65.5K tokens
- Pricing listed
gemini-3.5-flash
Gemini 3.5 Flash is Google's fast multimodal model with frontier intelligence, superior search, and grounding capabilities.
- Third-party
- Context: 1M tokens
- Maximum output: 65.5K tokens
- Pricing listed
gemini-3.5-flash-lite
Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing.
- Third-party
- Context: 1M tokens
- Maximum output: 65.5K tokens
- Pricing listed
gemini-3.6-flash
Gemini 3.6 Flash provides sustained frontier-level intelligence optimized for real-world tasks at a higher speed and lower cost, excelling at code generation, agentic execution, and spatial reasoning.
- Third-party
- Context: 1M tokens
- Maximum output: 65.5K tokens
- Pricing listed
gemini-3.7-flash
Gemini 3.7 Flash is a highly capable, natively multimodal reasoning model optimized for agentic workflows and real-world tasks.
- Third-party
- Context: 1M tokens
- Maximum output: 65.5K tokens
- Pricing listed
gemini-3.8-flash
Our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
- Third-party
- Context: 1M tokens
- Maximum output: 65.5K tokens
- Pricing listed
nano-banana-2-lite
Google's fastest Gemini image generation model for rapid image creation and iteration.
- Third-party
- Zero data retention
- Context: 65.5K tokens
- Pricing listed
m2.7
MiniMax's M2.7 language model with multilingual capabilities.
- Third-party
- Zero data retention
- Context: 128K tokens
- Maximum output: 4.1K tokens
- Pricing listed
m3
MiniMax's M3 language model with frontier coding and agentic capabilities, a 1M token context window, and multilingual support.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 4.1K tokens
- Pricing listed
kimi-k3
Kimi K3 is Moonshot's flagship 2.8 trillion-parameter model, built on Kimi Delta Attention (a hybrid linear attention mechanism) with Attention Residuals. It offers native visual understanding, always-on reasoning, and a 1M-token context window for long-horizon coding, knowledge work, and deep reasoning tasks.
- Third-party
- Context: 1M tokens
- Maximum output: 1M tokens
- Pricing listed
gpt-4.1
OpenAI's flagship GPT model for complex tasks with a million-token context window.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 32.8K tokens
- Pricing listed
gpt-4.1-mini
Fast, affordable version of GPT-4.1 with a million-token context window.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 32.8K tokens
- Pricing listed
gpt-4.1-nano
GPT-4.1 Nano is OpenAI’s smallest and cheapest GPT-4.1 variant, optimized for high-throughput, low-latency tasks.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 32.8K tokens
- Pricing listed
gpt-4o
GPT-4o is OpenAI’s multimodal flagship, accepting text and images and responding quickly across a wide range of tasks.
- Third-party
- Zero data retention
- Context: 128K tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-4o-mini
GPT-4o Mini is the lightweight, low-cost variant of GPT-4o, well suited to high-volume tasks with multimodal inputs.
- Third-party
- Zero data retention
- Context: 128K tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-5
OpenAI's model excelling at coding, writing, and reasoning.
- Third-party
- Zero data retention
- Context: 128K tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-5-mini
GPT-5 Mini is the lightweight, low-cost variant of GPT-5, well suited to high-volume coding and reasoning tasks.
- Third-party
- Zero data retention
- Context: 128K tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-5-nano
GPT-5 Nano is OpenAI’s smallest GPT-5 variant, optimized for low latency and cheap, high-throughput tasks.
- Third-party
- Zero data retention
- Context: 128K tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-5.1
GPT-5.1 is OpenAI’s incremental improvement over GPT-5, with stronger coding, reasoning, and writing.
- Third-party
- Zero data retention
- Context: 128K tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-5.4
GPT-5.4 is OpenAI's flagship model with strong coding, reasoning, and multimodal capabilities.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-5.4-mini
GPT-5.4 mini is a smaller, faster, and more cost-efficient version of GPT-5.4 for lightweight tasks.
- Third-party
- Zero data retention
- Context: 128K tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-5.4-nano
GPT-5.4 nano is OpenAI's smallest and fastest model, optimized for edge and low-latency use cases.
- Third-party
- Zero data retention
- Context: 128K tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-5.4-pro
GPT-5.4 pro uses OpenAI's Responses API with built-in tools, improved reasoning, and stateful context management.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-5.5
GPT-5.5 is OpenAI's flagship model with strong coding, reasoning, and multimodal capabilities.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-5.5-pro
GPT-5.5 pro uses OpenAI's Responses API with built-in tools, improved reasoning, and stateful context management.
- Third-party
- Zero data retention
- Context: 1M tokens
- Maximum output: 16.4K tokens
- Pricing listed
gpt-5.6-luna
GPT-5.6 Luna is an OpenAI GPT-5.6 model optimized for cost-sensitive workloads, using the Responses API for efficient text generation.
- Third-party
- Context: 1.1M tokens
- Maximum output: 128K tokens
- Pricing listed
gpt-5.6-sol
GPT-5.6 Sol is OpenAI's frontier GPT-5.6 model for complex professional work, using the Responses API for reasoning and stateful context management.
- Third-party
- Context: 1.1M tokens
- Maximum output: 128K tokens
- Pricing listed
gpt-5.6-terra
GPT-5.6 Terra is an OpenAI GPT-5.6 model that balances intelligence and cost, using the Responses API for reasoning and stateful context management.
- Third-party
- Context: 1.1M tokens
- Maximum output: 128K tokens
- Pricing listed
gpt-6-astra
GPT-6 Astra is OpenAI's most capable model, built for complex reasoning, coding, computer use, research, and document creation.
- Third-party
- Context: 1.1M tokens
- Maximum output: 128K tokens
- Pricing listed
o3
o3 is OpenAI’s general-purpose reasoning model, balancing strong analytical performance with reasonable latency and cost.
- Third-party
- Zero data retention
- Context: 200K tokens
- Maximum output: 100K tokens
- Pricing listed
o3-mini
o3-mini is the lightweight, low-cost reasoning variant of o3, well suited to quick analytical tasks at scale.
- Third-party
- Zero data retention
- Context: 200K tokens
- Maximum output: 100K tokens
- Pricing listed
o4-mini
OpenAI's fast, lightweight reasoning model optimized for multi-step problem solving at lower cost.
- Third-party
- Zero data retention
- Context: 200K tokens
- Maximum output: 100K tokens
- Pricing listed
grok-4.20-0309-non-reasoning
xAI's Grok 4.20 non-reasoning model. Skips the thinking trace for fast, single-pass responses while keeping the same training as the reasoning variant.
- Third-party
- Zero data retention
- Context: 2M tokens
- Pricing listed
grok-4.20-0309-reasoning
xAI's Grok 4.20 reasoning model. Uses extended thinking to work through complex problems, returning a reasoning trace alongside the final answer.
- Third-party
- Zero data retention
- Context: 2M tokens
- Pricing listed
grok-4.20-multi-agent-0309
xAI's Grok 4.20 multi-agent model with a 2M-token context window. Multiple agents collaborate in parallel to perform deep research tasks, with function calling, structured outputs, and reasoning capabilities.
- Third-party
- Zero data retention
- Context: 2M tokens
- Pricing listed
grok-4.5
xAI's Grok 4.5, a frontier model built for coding, agentic tasks, and knowledge work. Accepts text and image inputs, and supports function calling, structured outputs, and configurable reasoning effort (low, medium, high).
- Third-party
- Zero data retention
- Context: 500K tokens
- Pricing listed
grok-4.6
xAI's Grok 4.6, a flagship reasoning model for coding, agentic tasks, and visual work. Accepts text and image inputs, and supports function calling and structured outputs.
- Third-party
- Zero data retention
- Context: 500K tokens
- Pricing listed
bge-base-en-v1.5
BAAI general embedding (Base) model that transforms any given text into a 768-dimensional vector
- Cloudflare-hosted
- Batch
- Context: 153.6K tokens
- Pricing listed
bge-m3
Multi-Functionality, Multi-Linguality, and Multi-Granularity embeddings model.
- Cloudflare-hosted
- Context: 60K tokens
- Pricing listed
deepseek-r1-distill-qwen-32b
DeepSeek-R1-Distill-Qwen-32B is a model distilled from DeepSeek-R1 based on Qwen2.5. It outperforms OpenAI-o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.
- Cloudflare-hosted
- Reasoning
- Context: 80K tokens
- Pricing listed
deepseek-v4-flash-0731
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Context: 1M tokens
- Pricing listed
deepseek-v4-pro-0813
DeepSeek V4 Pro is a high-capability reasoning model from DeepSeek with a one million token context window, built for long-horizon agentic workflows and complex, multi-step problem-solving
- Cloudflare-hosted
- Function calling
- Reasoning
- Context: 1M tokens
- Pricing listed
gemma-4-26b-a4b-it
Gemma 4 is Google's most intelligent family of open models, built from Gemini 3 research to maximize intelligence-per-parameter.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Vision
- Context: 256K tokens
- Pricing listed
gemma-sea-lion-v4-27b-it
SEA-LION stands for Southeast Asian Languages In One Network, which is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.
- Cloudflare-hosted
- Context: 128K tokens
- Pricing listed
glm-4.7-flash
GLM-4.7-Flash is a fast and efficient multilingual text generation model with a 131,072 token context window. Optimized for dialogue, instruction-following, and multi-turn tool calling across 100+ languages.
- Cloudflare-hosted
- Function calling
- Reasoning
- Context: 131.1K tokens
- Pricing listed
glm-5.2
Z.ai's flagship agentic coding model
- Cloudflare-hosted
- Function calling
- Reasoning
- Context: 262.1K tokens
- Pricing listed
glm-5.3
GLM-5.3 is Z.ai's flagship agentic coding model, pairing a 1M-token context window with reasoning, function calling, and structured outputs to power multi-step, tool-driven development workflows.
- Cloudflare-hosted
- Function calling
- Reasoning
- Context: 1.3M tokens
- Pricing listed
gpt-oss-120b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases – gpt-oss-120b is for production, general purpose, high reasoning use-cases.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Context: 128K tokens
- Pricing listed
gpt-oss-20b
OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases – gpt-oss-20b is for lower latency, and local or specialized use-cases.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Context: 128K tokens
- Pricing listed
kimi-k2.6
Kimi K2.6 is a frontier-scale open-source 1T parameter model with a 262.1k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Vision
- Context: 262.1K tokens
- Pricing listed
kimi-k2.7-code
Kimi K2.7 is a frontier-scale open-source 1T parameter model with a 262.1k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads.
- Cloudflare-hosted
- Function calling
- Reasoning
- Vision
- Context: 262.1K tokens
- Pricing listed
llama-3.1-8b-instruct-fp8
Llama 3.1 8B quantized to FP8 precision
- Cloudflare-hosted
- Context: 32K tokens
- Pricing listed
llama-3.2-11b-vision-instruct
The Llama 3.2-Vision instruction-tuned models are optimized for visual recognition, image reasoning, captioning, and answering general questions about an image.
- Cloudflare-hosted
- LoRA
- Vision
- Context: 128K tokens
- Pricing listed
llama-3.2-1b-instruct
The Llama 3.2 instruction-tuned text only models are optimized for multilingual dialogue use cases, including agentic retrieval and summarization tasks.
- Cloudflare-hosted
- Context: 60K tokens
- Pricing listed
llama-3.2-3b-instruct
The Llama 3.2 instruction-tuned text only models are optimized for multilingual dialogue use cases, including agentic retrieval and summarization tasks.
- Cloudflare-hosted
- LoRA
- Context: 80K tokens
- Pricing listed
llama-3.3-70b-instruct-fp8-fast
Llama 3.3 70B quantized to fp8 precision, optimized to be faster.
- Cloudflare-hosted
- Batch
- Function calling
- Context: 24K tokens
- Pricing listed
melotts
MeloTTS is a high-quality multi-lingual text-to-speech library by MyShell.ai.
- Cloudflare-hosted
- Pricing listed
nemotron-3-120b-a12b
NVIDIA Nemotron 3 Super is a hybrid MoE model with leading accuracy for multi-agent applications and specialized agentic AI systems.
- Cloudflare-hosted
- Function calling
- Reasoning
- Context: 256K tokens
- Pricing listed
qwen3-embedding-0.6b
The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks.
- Cloudflare-hosted
- Context: 8.2K tokens
- Pricing listed
qwen3.8-27b
Qwen 3.8 27B is a 27-billion-parameter instruction-tuned language model from Alibaba's Qwen family, designed for vision, efficient general-purpose text generation and agentic workloads.
- Cloudflare-hosted
- Batch
- Function calling
- Reasoning
- Vision
- Context: 262.1K tokens
- Pricing listed