AI Provider Guides
Complete setup guides for all supported AI providers.
🆓 Free Tier Providers
Start with zero cost using these free-tier options:
Hugging Face
100,000+ open-source models
- ✅ Free inference API
- 🌍 Largest model collection
- 🔓 Fully open source
- 📊 Models by task: chat, classification, NER, summarization
Google AI Studio
Gemini models with generous free tier
- ✅ 1,500 requests/day free
- ⚡ Fast Gemini 2.0 Flash
- 🎯 15 requests/minute
- 💰 Pay-as-you-go option
🤖 Direct AI Providers
Access leading AI models directly from their creators:
OpenAI
GPT-5.4, GPT-5, GPT-4o, and o-series reasoning models
- 🧠 GPT-5.4 and GPT-5 series flagships with up to 400K context
- 👁️ GPT-4o multimodal (vision) and o3 / o3-pro / o4-mini reasoning models
- 🔧 Full tool/function calling and embeddings support
- 🔑 Auth: API Key (
OPENAI_API_KEY)
Anthropic
Claude models with API key or OAuth authentication
- 🧠 Claude 4.5 Opus/Sonnet/Haiku, Claude 4.0 Opus/Sonnet
- 🔐 API key or OAuth (Pro/Max subscription)
- 💭 Extended thinking for deep reasoning
- 📄 200K context window, multimodal support
🏢 Enterprise Providers
Production-grade providers for enterprise deployments:
Azure OpenAI
Enterprise AI with Microsoft Azure
- 🔒 SOC2, HIPAA, ISO 27001 compliant
- 🌍 Multi-region deployment (30+ regions)
- 🛡️ Private endpoints with VNet
- 💼 Enterprise SLAs
Google Vertex AI
Google Cloud ML platform
- ☁️ GCP integration
- 🔐 IAM, VPC, service accounts
- 🌏 Global deployment
- 🎯 Gemini, PaLM, Codey models
AWS Bedrock
Serverless AI on AWS
- 📦 13 foundation models (Claude, Llama, Mistral)
- 🔐 IAM, VPC integration
- 🌍 Multi-region (us-east-1, eu-west-1, ap-southeast-1)
- 💰 Pay-per-use pricing
AWS SageMaker
Custom model endpoints on AWS SageMaker infrastructure
- 🎯 Deploy fine-tuned, Hugging Face, or JumpStart models
- 🔐 IAM, VPC, PrivateLink, KMS encryption
- ⚠️
generate()only — streaming is not implemented - 💰 Full control over instance types and autoscaling
🌍 Compliance-Focused
Providers with specific compliance certifications:
Mistral AI
European AI with GDPR compliance
- 🇪🇺 EU data residency
- ✅ GDPR compliant by default
- 🔓 Open source models
- 💰 Cost-effective
🧑💻 Hosted Inference Providers
Access frontier models via hosted cloud inference APIs:
DeepSeek
deepseek-chat (V3) and deepseek-reasoner (R1)
- 🧠 deepseek-chat — high-quality general chat at low cost
- 💭 deepseek-reasoner — R1 chain-of-thought reasoning model
- 🔑 API key from platform.deepseek.com
- 🔄 Aliases:
ds
NVIDIA NIM
400+ models via NVIDIA's hosted and self-hosted inference platform
- 🚀 Llama 3.3 70B Instruct (default), Mistral, Nemotron, and 400+ catalog models
- 🔧 NIM-specific extras: top_k, min_p, repetition_penalty, reasoning_budget
- 🔑 API key from build.nvidia.com
- 🖥️ Also supports self-hosted NIM endpoints via
NVIDIA_NIM_BASE_URL - 🔄 Aliases:
nim,nvidia
xAI Grok
Grok 3 / 3 Mini / 2 / 2 Vision via api.x.ai
- 🧠 Grok 3 — flagship reasoning + coding
- ⚡ Grok 3 Mini — faster + cheaper
- 👁️ Grok 2 Vision — multimodal text + images
- 🔑 API key from console.x.ai
- 🔄 Aliases:
grok
Groq
Sub-100ms inference via LPU acceleration
- ⚡ <100ms TTFT — fastest hosted inference available
- 🦙 Llama 3.3 70B Versatile (default), Llama 3.1 8B Instant, Mixtral, Gemma 2
- 👁️ Llama 3.2 vision-preview variants for multimodal
- 🔑 API key from console.groq.com/keys
Cerebras
Wafer-scale inference at ~3000 tokens/s
- 🚀 Fastest generation speed of any hosted provider (WSE hardware)
- 🤖 GPT-OSS 120B (default), Gemma 4 31B — roster live-verified 2026-08-27
- 💳 Free $5 credit requires a saved payment method
- 🔑 API key from cloud.cerebras.ai
SambaNova
RDU-accelerated open-weight flagships
- 🧠 Llama 3.3 70B (default), GPT-OSS 120B, DeepSeek V3.x, MiniMax, Gemma 4
- 👁️ Vision on gemma-4-31B-it (image+video) and MiniMax-M3
- 💳 No free allowance — credits required before first call
- 🔑 API key from cloud.sambanova.ai/apis
Together AI
Hosted open-model gateway
- 📚 Llama 3.3 / 3.1 (8B–405B), Mixtral, Qwen 2.5, DeepSeek R1/V3, WizardLM
- ⚡ Turbo variants for low latency
- 🔑 API key from api.together.xyz/settings/api-keys
- 🔄 Aliases:
together
Fireworks AI
Fast open-model serving
- 🔥 Llama v3.1 70B/405B, Mixtral 8x22B, Qwen 2.5 Coder, DeepSeek V3
- 👁️ Phi-3-Vision and Llama 3.2 vision variants
- 🔑 API key from fireworks.ai/account/api-keys
Perplexity
Sonar models with built-in web grounding
- 🌐 sonar / sonar-pro / sonar-reasoning / sonar-deep-research
- 📚 Built-in web search + citations
- 🔑 API key from perplexity.ai/settings/api
- 🔄 Aliases:
pplx
Cloudflare Workers AI
Edge-served open models
- 🌍 Lowest cost tier — bills per "neuron" not token
- 🦙 Llama 3.3 70B FP8, Llama 3.1, Mistral, Qwen, Gemma
- 🔑 Token from dash.cloudflare.com/profile/api-tokens (Workers AI Read+Write)
- ⚠️ Requires both
CLOUDFLARE_API_KEYANDCLOUDFLARE_ACCOUNT_ID - 🔄 Aliases:
workers-ai,cf-ai
Cohere
Command R / R+ chat + Embed v3 / Rerank v3 (RAG-essential)
- 💬 Command R+ flagship + Command R + Command R7B
- 🔍 Embed v3 (English / multilingual) + Rerank v3 — top-tier RAG
- 🔑 API key from dashboard.cohere.com/api-keys
Baseten
Hosted open-model inference
- 🤖 Default model:
zai-org/GLM-5.3-Flash - 🔑
BASETEN_API_KEY— no dedicated setup guide yet
GMI Cloud
Hosted open-model inference
- 🤖 Default model:
MiniMaxAI/MiniMax-M3 - 🔑
GMICLOUD_API_KEY— no dedicated setup guide yet
Inception Labs
Diffusion LLMs
- 🤖 Default model:
mercury-2 - 🔑
INCEPTION_LABS_API_KEY— no dedicated setup guide yet
io.net Intelligence
Decentralized GPU inference
- 🤖 Default model:
meta-llama/Llama-3.3-70B-Instruct - 🔑
IO_INTELLIGENCE_API_KEY— no dedicated setup guide yet
Mancer
Hosted open-model inference
- 🤖 Default model:
deepseek-v4-flash - 🔑
MANCER_API_KEY— no dedicated setup guide yet
Upstage
Solar models
- 🤖 Default model:
solar-pro4 - 🔑
UPSTAGE_API_KEY— no dedicated setup guide yet
API Route
OpenAI-compatible passthrough
- 🤖 Default model:
claude-sonnet-4-6 - 🔑
API_ROUTE_API_KEY— no dedicated setup guide yet
Replicate
Multi-modal gateway — LLM + image + video + avatar + music in one auth
- 🎯 One
REPLICATE_API_TOKENfor 5 modalities - 📚 Llama 3.1 70B/405B, Mistral, Mixtral
- 🎨 FLUX 1.1 Pro, SDXL, Stable Diffusion 3.5
- 🎬 Wan-Alpha video, MuseTalk avatar, MusicGen music
- 🔑 Token from replicate.com/account/api-tokens
🔍 Embedding-Only Providers
Specialised embedding providers for RAG / retrieval pipelines (no chat):
Voyage AI
Top-tier RAG embeddings
- 📊 voyage-4-large newest flagship; voyage-3.5 default; voyage-code-3 / voyage-code-4 for code
- 🌍 voyage-multilingual-2 + domain-tuned (finance, law)
- 🔑 API key from dash.voyageai.com/api-keys
Jina AI
Embeddings + reranking
- 📊 jina-embeddings-v3 multilingual flagship
- 🔄 jina-reranker-v2 for retrieval reranking
- 🔍 jina-colbert-v2 late-interaction retrieval
- 🔑 API key from jina.ai
🎨 Direct Image Generation
Specialised image-gen providers (in addition to Vertex Imagen / OpenAI DALL-E / Anthropic / Bedrock):
Stability AI
Stable Image Ultra/Core + SD 3.5 family
- 🎨 Stable Image Ultra (flagship), Core (fast), SD 3.5 Large/Large-Turbo/Medium
- 🖼️ PNG output, aspect-ratio + negative-prompt + seed support
- 🔑 API key from platform.stability.ai/account/keys
- 🔄 Aliases:
stability-ai,sd
Ideogram
Strong typography + design-focused image generation
- 📝 V3 default; V2/V2-Turbo/V1 also supported
- 🎨 magic_prompt + style + aspect_ratio controls
- 🔑 API key from developer.ideogram.ai
Recraft
Vector / illustration-focused image generation
- 🎨 recraftv3 (raster), recraftv3-svg (vector), recraftv2
- 📐 OpenAI-compat shape + style + size controls
- 🔑 API token from recraft.ai/api
💻 Local Providers
Run models entirely on your own hardware — no API key or internet required for inference:
Ollama
Run open-source models locally with full privacy
- 🖥️ 100% local inference — no data leaves your machine
- 🦙 70+ models: Llama, Mistral, Qwen, DeepSeek, Gemma, Phi, CodeLlama
- 🌐 Native Ollama API and OpenAI-compatible mode
- 🆓 No API key required
LM Studio
Run any supported model locally with a GUI app
- 🖥️ Download and run models via the LM Studio desktop application
- 🔍 Auto-discovers the loaded model from
/v1/models(no model name required) - 🌐 OpenAI-compatible API at
http://localhost:1234/v1by default - 🆓 No API key needed for local use (key optional for reverse-proxy setups)
- 🔄 Aliases:
lmstudio,lms
llama.cpp
High-performance local inference via llama-server
- ⚡ Run GGUF models with llama-server at
http://localhost:8080/v1by default - 🔍 Auto-discovers the loaded model from
/v1/models - 🛠️ Tool support requires
--jinjaflag when starting llama-server - 🆓 No API key needed for local use (key optional for reverse-proxy setups)
- 🔄 Aliases:
llama.cpp
🔌 Aggregators & Proxies
Access multiple providers through unified interfaces:
OpenRouter
300+ models from 60+ providers
- 🌐 Single API across many AI providers (Anthropic, OpenAI, Google, Meta, etc.)
- ⚡ Automatic failover and routing
- 💰 Competitive pricing with cost optimization
- 🎯 Zero lock-in - switch models instantly
- 📊 Usage tracking dashboard
- 🆓 Free models available
OpenAI Compatible
OpenRouter, vLLM, LocalAI, and more
- 🌐 100+ models through OpenRouter
- 💻 Local deployment with vLLM
- 🔓 Self-hosted with LocalAI
- 🔄 Drop-in OpenAI replacement
LiteLLM
100+ providers through proxy
- 🔄 Unified API for 100+ providers
- 📊 Load balancing and fallbacks
- 💰 Cost tracking
- 🎯 Model routing
🧠 Decision-Only Providers
The two providers that serve decide rather than generate/stream. Each
returns typed, calibrated judgments and emits no text, so neither appears in
generation fallback chains or the health sweep.
TypeSafe (Jev)
Typed, calibrated judgments instead of text
- 🎯
boolean/choice/scoreanswers, each with a calibrated confidence - ⚡ Latency flat in question count — 1 question ~393 ms, 400 questions ~465 ms
- 💰
$0.00002 per decision ($0.042/M input, output billed at zero) - 🔌 Two transports: TypeSafe direct, or the Vercel AI Gateway
- 🛡️ Fails open — with no key configured, every consumer behaves exactly as before
- 🔑 API key from console.typesafe.ai/keys
- 🔄 Aliases:
jev,typesafe-ai
Laya
Open-weights decision provider — the same typed boolean/choice/score answers as Jev, from Convai Innovations' Apache-2.0 checkpoints, at a Laya server or LiteLLM proxy route you configure
- 🧭 Serves
decideonly; built-in features use it when its key and base URL are set and neitherTYPESAFE_API_KEYnorAI_GATEWAY_API_KEYis - 📏 Refuses more than ~768 tokens of state (320 on
english) before any network call, since its encoders read only 1,024 (512) - 🔌 No built-in endpoint:
LAYA_BASE_URL(orcredentials.laya.baseURL) names a Laya server or a LiteLLM pass-through route to one - 🔑
LAYA_API_KEYis the key that endpoint accepts; on LiteLLM, the route must be in the key's Allowed Routes
🧩 Additional Catalog Providers
Every provider below is a Tier-2 catalog entry. Provider-specific
behavioural quirks may apply, so check the provider's catalog file. The whole
integration is one JSON file under src/lib/providers/catalog/. Each page is
generated from that file, which is also what the CI onboarding gate reads.
API Route
Claude Sonnet 4.6
- 🤖 8 models; default
claude-sonnet-4-6(1M context) - 🛠️ Native tool calling + structured output together
- 💳 Free tier available
- ✅ Roster verified 2026-09-17 (authenticated GET /v1/models)
- 🔑 API key from api-route.com
- 🔄 Aliases:
apiroute
Baseten
GLM 5.3 Flash
- 🤖 16 models; default
zai-org/GLM-5.3-Flash(1M context) - 🛠️ Native tool calling + structured output together
- 💳 Free tier available
- ✅ Roster verified 2026-09-03 (authenticated GET /v1/models)
- 🔑 API key from app.baseten.co
Friendli
zai-org/GLM-5.3
- 🤖 7 models; default
zai-org/GLM-5.3(1M context) - 🛠️ Native tool calling — not combined with structured output in one request (HTTP 422 if both are set)
- 💳 Free tier available (rate limits are tight — pace requests 20s+ apart)
- ✅ Roster verified 2026-09-06 (authenticated GET /serverless/v1/models)
- 🔑 API key from suite.friendli.ai
GMI Cloud
MiniMaxAI/MiniMax-M3
- 🤖 1 model; default
MiniMaxAI/MiniMax-M3(1M context) - 🛠️ Native tool calling + structured output together
- 💳 Free tier available
- ✅ Roster verified 2026-09-03 (authenticated GET /v1/models)
- 🔑 API key from console.gmicloud.ai
- 🔄 Aliases:
gmi-cloud
Inception Labs
Mercury 2, Inception's enterprise diffusion LLM (dLLM)
- 🤖 1 model; default
mercury-2(125K context) - 🛠️ Native tool calling + structured output together
- 💳 Free tier available
- ✅ Roster verified 2026-09-03 (authenticated GET /v1/models)
- 🔑 API key from platform.inceptionlabs.ai/dashboard/api-keys
- 🔄 Aliases:
inception,mercury
io.net Intelligence
Meta: Llama 3.3 70B Instruct
- 🤖 34 models; default
meta-llama/Llama-3.3-70B-Instruct(125K context) - 🛠️ Native tool calling + structured output together
- 💳 Free tier available
- ✅ Roster verified 2026-09-03 (authenticated GET /v1/models)
- 🔑 API key from ai.io.net
- 🔄 Aliases:
io-net
Mancer
DeepSeek V4 Flash
- 🤖 10 models; default
deepseek-v4-flash(1M context) - ⚠️ No tool calling — text generation and structured output only
- 💳 Free tier available
- ✅ Roster verified 2026-09-03 (authenticated GET /oai/v1/models; full response retained as evidence/mancer-roster-authenticated.json in the campaign scratchpad and every catalog price/limit machine-checked against it (Mancer re-prices — gpt-oss-120b input moved 0.024 → 0.022 within the day))
- 🔑 API key from mancer.tech/dashboard
- 🔄 Aliases:
mancer-tech
Morph
Morph v3 Large
- 🤖 2 models; default
morph-v3-large(262K context) - ⚠️ No tool calling — text generation and structured output only