Commercial multimodal understanding for text, images, and visual prompts.
Context 1M · Max output 128K45 models. One account. Clear usage pricing.
Browse proprietary commercial models from providers in North America and Europe, compare their current public API rates, and choose the right endpoint for each workload.
Compare every currently listed model and its public rate.
Prices are shown in USD using the billing unit published by each model provider. Search by model or provider, or narrow the directory by workload type.
Catalog pricing snapshot · 2026-09-18Vision
Visual understanding and multimodal reasoning
Commercial multimodal understanding for text, images, and visual prompts.
Context 1M · Max output 128KCommercial multimodal understanding for text, images, and visual prompts.
Context 1M · Max output 128KAnthropic's fastest current Claude model for interactive, high-throughput, and cost-sensitive workloads.
Context 200K · Max output 64KCommercial multimodal understanding for text, images, and visual prompts.
Context 1M · Max output 64KCommercial multimodal understanding for text, images, and visual prompts.
Context 1.05M · Max output 64KGoogle's stable, low-latency multimodal model for high-volume automation, document parsing, and tool use.
Context 1.05M · Max output 64KGoogle's preview multimodal reasoning model for demanding software engineering and agent workflows.
Context 1.05M · Max output 64KMeta's multimodal reasoning model for long-horizon agentic and coding work, with text, image, video, audio, and PDF input.
Context 1.05M · Max output 128KOpenAI's flagship reasoning model with text and image input, billed at double the input rate above 272K input tokens.
Context 1.05M · Max output 128KCommercial multimodal understanding for text, images, and visual prompts.
Context 1.05M · Max output 128KCommercial multimodal understanding for text, images, and visual prompts.
Context 1.05M · Max output 128KThe GPT-5.6 reasoning stack for cost-sensitive, high-volume text and image workloads.
Context 1.05M · Max output 128KOpenAI's model for long-running agentic software engineering in Codex-style environments.
Context 400K · Max output 128KCommercial multimodal understanding for text, images, and visual prompts.
Context 500K · Max output 450KImage
Text-led image generation
Commercial image generation and editing through a dedicated media endpoint.
Google's fast image generation and editing model, billed per output image at 1K and 2K resolutions.
Context 128K · Max output 32KGoogle's flagship image generation and editing model with 2K and 4K output, multi-image fusion, and web-search grounding.
Context 64K · Max output 32KCommercial image generation and editing through a dedicated media endpoint.
OpenAI's fast image generation and editing model for high-volume production work.
OpenAI's most capable image generation and editing model, tuned for detailed instructions and reference-guided edits.
Commercial image generation and editing through a dedicated media endpoint.
xAI's commercial image generation and editing model with 1K and 2K output and multi-image references.
Video
Text- and image-led video generation
Commercial video generation from text or image inputs.
Google's stable video model for fast generation and conversational editing from text, images, or short video inputs.
Context 1.05MCommercial video generation from text or image inputs.
Commercial video generation from text or image inputs.
Commercial video generation from text or image inputs.
xAI's commercial model for text-to-video, image-to-video, and reference-guided generation up to 1080p.
Audio
Speech recognition, transcription, and synthesis
Commercial speech recognition, transcription, and text-to-speech generation.
Commercial speech recognition, transcription, and text-to-speech generation.
Google's stable speech-to-text model for accurate, low-latency transcription across multilingual audio.
Google's preview audio-to-audio model for low-latency, voice-first applications with multimodal input and tool use.
Context 128K · Max output 64KCommercial speech recognition, transcription, and text-to-speech generation.
OpenAI's high-accuracy speech-to-text model for completed files, streamed files, and committed Realtime turns.
OpenAI's realtime reasoning model for interactive voice agents, image input, and live tool use.
Context 128K · Max output 32KOpenAI's generally available audio model for Chat Completions workflows that accept and return audio.
Context 128K · Max output 16KCommercial speech recognition, transcription, and text-to-speech generation.
Context 2KxAI's real-time speech-to-speech service for bidirectional voice agents with reasoning and tool use.
xAI's streamed and batch text-to-speech model with expressive multilingual voices and configurable delivery.
Embedding
Semantic representation for search and retrieval
Hosted vector embeddings for semantic search and retrieval.
Context 128KHosted vector embeddings for semantic search and retrieval.
Context 8KOpenAI's highest-capability text embedding model for English and multilingual retrieval.
Context 8KOpenAI's lower-cost third-generation embedding model for high-volume search and classification.
Context 8KReranker
Relevance scoring for retrieval pipelines
Hosted relevance scoring for improving retrieved result ordering.
Context 32KFund your account and make the first API call.
Buy tokens, create an API key, and call the selected model by its model ID.