AI21 Labs Provider Guide
AI21 Labs is a Tier-2 catalog provider: NeuroLink reaches it through its
shared chat-completions client, and AI21 documents a chat endpoint at
POST /chat/completions with Bearer authentication, a choices array in the
response and server-sent-event streaming ending in data: [DONE]
(Chat request,
Chat response). No
behavioural quirks are declared, so its integration is one JSON file
(src/lib/providers/catalog/ai21.json) rather than hand-written code. That file
is the source of truth for this page.
Verification status: this entry is docs-verified only, not yet live-verified. AI21's
GET /studio/v1/modelsanswers200with an emptydataarray without a key, so the model ids come from AI21's public docs (Jamba models page, retrieved 2026-09-29) and the roster has not been checked. No account was created and no API key was used to build this entry, and no POST request was sent.evidence.liveMatrixisnulluntil someone runs the live capability matrix with a real key (see Verification status below).
Key Facts
- Provider id:
ai21 - Protocol: chat completions at
POST https://api.ai21.com/studio/v1/chat/completionswithmodelandmessages, Bearer authentication (Chat request, Authentication) - Base URL:
https://api.ai21.com/studio/v1 - Default model:
jamba-large - Models in catalog: 6 — the six ids named in the API Versioning list of the Jamba models page
- Context: the models page lists Max Tokens
256Kfor Jamba Large and Jamba Mini; the catalog stores the number 256,000 as a placeholder for that256Klabel - Max output: the Chat request page states "For Jamba models, the maximum
allowed value is 4096 tokens." for
max_tokens - Streaming: supported — the
streamparameter is documented as "Stream results one token at a time using server-sent events." (Chat request); the Chat response page says "The final message will bedata: [DONE]." - Tool calling:
true— the Function Calling page says "Function calling is available for all Jamba models via the Chat Completions API." - Tools while streaming:
false— thestreamparameter text on the Chat request page says "Must be False if using tools." - Structured output:
true, based on JSON mode — the Chat request page documentsresponse_format: "Setting it to{ "type": "json_object" }activates JSON mode, ensuring the generated message adheres to valid JSON structure." NeuroLink's chat-completions client builds ajson_schemaresponse_formatwhen a schema is supplied; that form was not sent to AI21 because no POST request was sent. - Structured output + tools together:
false— no combined tools + schema request was sent - Embeddings: not declared (
false) - Thinking: not declared (
false) - Rate limits: the Rate Limits page lists Jamba Large and Jamba Mini at 10 RPS and 200 RPM on the AI21 Platform
- Billing:
free-tier— see Billing below - Key format: none declared (
apiKeyFormatisnull)
Quick Start
1. Get an API key
- Visit: https://docs.ai21.com/docs/create-api-key and follow its steps: sign in to AI21 Studio (the page says "You can sign in using your email address, Google account, GitHub account, or SSO."), click Settings in the bottom-left corner of the sidebar, open the API Keys tab, then click Create new key
- Copy the key when it is shown; the page says "you won’t be able to view it again"
- Free trial: https://docs.ai21.com/docs/usage-cost says "New accounts are given a $10 credit good for three months on the AI21 platform." and, once the trial is exceeded or expires, "you must provide valid billing information in your account to continue using the AI21 platform". https://www.ai21.com/pricing shows "$10 credits for 7 days." and "No credit card needed" under Free Trial
- Set
AI21_API_KEYin your .env file
2. Configure
export AI21_API_KEY=your-api-key
export AI21_MODEL=jamba-large # optional — overrides the default model
export AI21_BASE_URL=https://api.ai21.com/studio/v1 # optional — proxy or self-hosted gateway
3. Use it
import { NeuroLink } from "@juspay/neurolink";
const neurolink = new NeuroLink();
const result = await neurolink.generate({
input: { text: "Explain context windows in one paragraph." },
provider: "ai21",
model: "jamba-large",
});
console.log(result.content);
# CLI
npx @juspay/neurolink generate "Hello" --provider ai21
Per-request credentials are passed like this:
await neurolink.generate({
input: { text: "Hello" },
provider: "ai21",
credentials: { ai21: { apiKey: process.env.AI21_API_KEY } },
});
Billing
AI21's docs and website state different free-trial durations:
- Pricing (docs): "New accounts are given a $10 credit good for three months on the AI21 platform." After the trial is exceeded or expires, "you must provide valid billing information in your account to continue using the AI21 platform."
- Pricing page (website): "Start with a free trial, no credit card required, then pay based on usage." Under Free Trial it lists "$10 credits for 7 days." and "No credit card needed".
The entry records free-tier. The
Account page says "You will
need to provide billing information when your introductory usage credits
expire."
Models
| Model | Context | Max output | Vision | $/M in · out | Notes |
|---|---|---|---|---|---|
jamba-large ⭐ | 256K | 4,096 | no | $2 / $8 | NeuroLink default. The models page: "Our most powerful and advanced model, designed to handle complex tasks at enterprise scale with superior performance." |
jamba-mini | 256K | 4,096 | no | $0.2 / $0.4 | NeuroLink fallback. The pricing page: "Efficient & lightweight model for a wide range of tasks". |
jamba-large-1.7-2025-07 | 256K | 4,096 | no | — | Dated snapshot of Jamba Large (Version 1.7, Snapshot 2025-07); jamba-large currently points to it. |
jamba-mini-2-2026-01 | 256K | 4,096 | no | — | Dated snapshot of Jamba Mini (Version 2, Snapshot 2026-01); jamba-mini currently points to it. |
jamba-large-1.7 | 256K | 4,096 | no | — | Version name; points to jamba-large-1.7-2025-07. |
jamba-mini-2 | 256K | 4,096 | no | — | Version name; points to jamba-mini-2-2026-01. |
Prices for jamba-large and jamba-mini come from
AI21's pricing page ("Jamba Large" at $2 / 1M
input tokens and $8 / 1M output tokens; "Jamba Mini" at $0.2 / 1M input tokens
and $0.4 / 1M output tokens), matched to the ids through the Model Details table
on the Jamba models page
(retrieved 2026-09-29). That page also lists the API Versioning mapping used in
the notes above and gives Input Modality: Text, so the catalog records
vision: false for these models. The same page states "We advise using dated versions of the
Jamba API to avoid disruptions from model updates and breaking changes."
models.defaultContextWindow and the per-model contextWindow values (256,000)
are a numeric placeholder for the 256K label printed on the
Jamba models page; no
exact token count is taken from a vendor page. models.defaultMaxOutputTokens
(4,096) is the max_tokens limit stated on the
Chat request page.
Fallback order when the default is unavailable:
jamba-mini → jamba-large-1.7-2025-07 → jamba-mini-2-2026-01. The runtime
fallback model name is set explicitly to jamba-mini.
Verification status
Tier-2 onboarding requires evidence before a provider is accepted, and
pnpm run verify:provider-onboarding gates it in CI. This is what the catalog
currently records for AI21 Labs — docs-verified only, not live-verified:
| Probe | Result |
|---|---|
| Roster | unauthenticated GET /studio/v1/models, HTTP 200, body {"data":[]}, 2026-09-29 — no key used; model ids come from the Jamba models page, not from the API |
| Chat endpoint | unauthenticated GET /studio/v1/chat/completions, HTTP 405, body {"detail":"Method Not Allowed"}, header allow: POST, 2026-09-29 |
| Billing | usage-cost and pricing pages, 2026-09-29 — the two pages state different trial terms, quoted under Billing |
| Auth-failure shape | unauthenticated GET /studio/v1/library/files (a documented endpoint), HTTP 400, body {"detail":"Bad request: bad or missing authentication."}, 2026-09-29. The Chat response page lists 401, 403, 422, 429, 500 and 503. errorRules match status 401, status 429, status 503 and the text bad or missing authentication |
| Tools / structured output | function calling is documented on the Function Calling page and response_format with { "type": "json_object" } on the Chat request page. Neither was exercised live. No POST request was sent, so stream_options on streaming calls, tool_choice on calls with tools and a json_schema response_format, which NeuroLink's chat-completions client can add, have not been exercised against AI21, and no combined tools + schema request was sent — structuredOutputWithTools stays false |
| Live capability sweep | not run. evidence.liveMatrix is null. |
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| HTTP 401 | The Chat response page documents 401 as "Unauthorized (Incorrect API key provided/Invalid Authentication)" | The Authentication page says to use your API key as the Bearer token; NeuroLink reads it from AI21_API_KEY |
| HTTP 429 | The Chat response page documents 429 as "Too Many Requests (You are sending requests too quickly.)" | The Rate Limits page lists 10 RPS and 200 RPM for Jamba Large and Jamba Mini and says to contact [email protected] if you require a higher limit |
| HTTP 503 | The Chat response page documents 503 as "Service Unavailable (The engine is currently overloaded, please try again later)" | The same text says "please try again later" |
| Trial credit exceeded or expired | The usage-cost page: "you must provide valid billing information in your account to continue using the AI21 platform" | Add billing information in your account; the page points to Account > Billing & Plans |
See also
- Provider setup overview
- All providers
- Tier-2 onboarding — how this provider's JSON becomes a working integration
- Provider feature compatibility