AI Provider Guides
Complete setup guides for all supported AI providers.
🆓 Free Tier Providers
Start with zero cost using these free-tier options:
Hugging Face
100,000+ open-source models
- ✅ Free inference API
- 🌍 Largest model collection
- 🔓 Fully open source
- 📊 Models by task: chat, classification, NER, summarization
Google AI Studio
Gemini models with generous free tier
- ✅ 1,500 requests/day free
- ⚡ Fast Gemini 2.0 Flash
- 🎯 15 requests/minute
- 💰 Pay-as-you-go option
🤖 Direct AI Providers
Access leading AI models directly from their creators:
Anthropic
Claude models with API key or OAuth authentication
- 🧠 Claude 4.5 Opus/Sonnet/Haiku, Claude 4.0 Opus/Sonnet
- 🔐 API key or OAuth (Pro/Max subscription)
- 💭 Extended thinking for deep reasoning
- 📄 200K context window, multimodal support
🏢 Enterprise Providers
Production-grade providers for enterprise deployments:
Azure OpenAI
Enterprise AI with Microsoft Azure
- 🔒 SOC2, HIPAA, ISO 27001 compliant
- 🌍 Multi-region deployment (30+ regions)
- 🛡️ Private endpoints with VNet
- 💼 Enterprise SLAs
Google Vertex AI
Google Cloud ML platform
- ☁️ GCP integration
- 🔐 IAM, VPC, service accounts
- 🌏 Global deployment
- 🎯 Gemini, PaLM, Codey models
AWS Bedrock
Serverless AI on AWS
- 📦 13 foundation models (Claude, Llama, Mistral)
- 🔐 IAM, VPC integration
- 🌍 Multi-region (us-east-1, eu-west-1, ap-southeast-1)
- 💰 Pay-per-use pricing
🌍 Compliance-Focused
Providers with specific compliance certifications:
Mistral AI
European AI with GDPR compliance
- 🇪🇺 EU data residency
- ✅ GDPR compliant by default
- 🔓 Open source models
- 💰 Cost-effective
🧑💻 Hosted Inference Providers
Access frontier models via hosted cloud inference APIs:
DeepSeek
deepseek-chat (V3) and deepseek-reasoner (R1)
- 🧠 deepseek-chat — high-quality general chat at low cost
- 💭 deepseek-reasoner — R1 chain-of-thought reasoning model
- 🔑 API key from platform.deepseek.com
- 🔄 Aliases:
ds
NVIDIA NIM
400+ models via NVIDIA's hosted and self-hosted inference platform
- 🚀 Llama 3.3 70B Instruct (default), Mistral, Nemotron, and 400+ catalog models
- 🔧 NIM-specific extras: top_k, min_p, repetition_penalty, reasoning_budget
- 🔑 API key from build.nvidia.com
- 🖥️ Also supports self-hosted NIM endpoints via
NVIDIA_NIM_BASE_URL - 🔄 Aliases:
nim,nvidia
xAI Grok
Grok 3 / 3 Mini / 2 / 2 Vision via api.x.ai
- 🧠 Grok 3 — flagship reasoning + coding
- ⚡ Grok 3 Mini — faster + cheaper
- 👁️ Grok 2 Vision — multimodal text + images
- 🔑 API key from console.x.ai
- 🔄 Aliases:
grok
Groq
Sub-100ms inference via LPU acceleration
- ⚡ <100ms TTFT — fastest hosted inference available
- 🦙 Llama 3.3 70B Versatile (default), Llama 3.1 8B Instant, Mixtral, Gemma 2
- 👁️ Llama 3.2 vision-preview variants for multimodal
- 🔑 API key from console.groq.com/keys
Together AI
Hosted open-model gateway
- 📚 Llama 3.3 / 3.1 (8B–405B), Mixtral, Qwen 2.5, DeepSeek R1/V3, WizardLM
- ⚡ Turbo variants for low latency
- 🔑 API key from api.together.xyz/settings/api-keys
- 🔄 Aliases:
together
Fireworks AI
Fast open-model serving
- 🔥 Llama v3.1 70B/405B, Mixtral 8x22B, Qwen 2.5 Coder, DeepSeek V3
- 👁️ Phi-3-Vision and Llama 3.2 vision variants
- 🔑 API key from fireworks.ai/account/api-keys
Perplexity
Sonar models with built-in web grounding
- 🌐 sonar / sonar-pro / sonar-reasoning / sonar-deep-research
- 📚 Built-in web search + citations
- 🔑 API key from perplexity.ai/settings/api
- 🔄 Aliases:
pplx
Cloudflare Workers AI
Edge-served open models
- 🌍 Lowest cost tier — bills per "neuron" not token
- 🦙 Llama 3.3 70B FP8, Llama 3.1, Mistral, Qwen, Gemma
- 🔑 Token from dash.cloudflare.com/profile/api-tokens (Workers AI Read+Write)
- ⚠️ Requires both
CLOUDFLARE_API_KEYANDCLOUDFLARE_ACCOUNT_ID - 🔄 Aliases:
workers-ai,cf-ai
Cohere
Command R / R+ chat + Embed v3 / Rerank v3 (RAG-essential)
- 💬 Command R+ flagship + Command R + Command R7B
- 🔍 Embed v3 (English / multilingual) + Rerank v3 — top-tier RAG
- 🔑 API key from dashboard.cohere.com/api-keys
Replicate
Multi-modal gateway — LLM + image + video + avatar + music in one auth
- 🎯 One
REPLICATE_API_TOKENfor 5 modalities - 📚 Llama 3.1 70B/405B, Mistral, Mixtral
- 🎨 FLUX 1.1 Pro, SDXL, Stable Diffusion 3.5
- 🎬 Wan-Alpha video, MuseTalk avatar, MusicGen music
- 🔑 Token from replicate.com/account/api-tokens
🔍 Embedding-Only Providers
Specialised embedding providers for RAG / retrieval pipelines (no chat):
Voyage AI
Top-tier RAG embeddings
- 📊 voyage-3-large flagship; voyage-3.5 default; voyage-code-3 for code
- 🌍 voyage-multilingual-2 + domain-tuned (finance, law)
- 🔑 API key from dash.voyageai.com/api-keys
Jina AI
Embeddings + reranking
- 📊 jina-embeddings-v3 multilingual flagship
- 🔄 jina-reranker-v2 for retrieval reranking
- 🔍 jina-colbert-v2 late-interaction retrieval
- 🔑 API key from jina.ai
🎨 Direct Image Generation
Specialised image-gen providers (in addition to Vertex Imagen / OpenAI DALL-E / Anthropic / Bedrock):
Stability AI
Stable Image Ultra/Core + SD 3.5 family
- 🎨 Stable Image Ultra (flagship), Core (fast), SD 3.5 Large/Large-Turbo/Medium
- 🖼️ PNG output, aspect-ratio + negative-prompt + seed support
- 🔑 API key from platform.stability.ai/account/keys
- 🔄 Aliases:
stability-ai,sd
Ideogram
Strong typography + design-focused image generation
- 📝 V3 default; V2/V2-Turbo/V1 also supported
- 🎨 magic_prompt + style + aspect_ratio controls
- 🔑 API key from developer.ideogram.ai
Recraft
Vector / illustration-focused image generation
- 🎨 recraftv3 (raster), recraftv3-svg (vector), recraftv2
- 📐 OpenAI-compat shape + style + size controls
- 🔑 API token from recraft.ai/api
💻 Local Providers
Run models entirely on your own hardware — no API key or internet required for inference:
LM Studio
Run any supported model locally with a GUI app
- 🖥️ Download and run models via the LM Studio desktop application
- 🔍 Auto-discovers the loaded model from
/v1/models(no model name required) - 🌐 OpenAI-compatible API at
http://localhost:1234/v1by default - 🆓 No API key needed for local use (key optional for reverse-proxy setups)
- 🔄 Aliases:
lmstudio,lms
llama.cpp
High-performance local inference via llama-server
- ⚡ Run GGUF models with llama-server at
http://localhost:8080/v1by default - 🔍 Auto-discovers the loaded model from
/v1/models - 🛠️ Tool support requires
--jinjaflag when starting llama-server - 🆓 No API key needed for local use (key optional for reverse-proxy setups)
- 🔄 Aliases:
llama.cpp
🔌 Aggregators & Proxies
Access multiple providers through unified interfaces:
OpenRouter
300+ models from 60+ providers
- 🌐 Single API for all major providers (Anthropic, OpenAI, Google, Meta, etc.)
- ⚡ Automatic failover and routing
- 💰 Competitive pricing with cost optimization
- 🎯 Zero lock-in - switch models instantly
- 📊 Usage tracking dashboard