Skip to main content

AI21 Labs Provider Guide

AI21 Labs is a Tier-2 catalog provider: NeuroLink reaches it through its shared chat-completions client, and AI21 documents a chat endpoint at POST /chat/completions with Bearer authentication, a choices array in the response and server-sent-event streaming ending in data: [DONE] (Chat request, Chat response). No behavioural quirks are declared, so its integration is one JSON file (src/lib/providers/catalog/ai21.json) rather than hand-written code. That file is the source of truth for this page.

Verification status: this entry is docs-verified only, not yet live-verified. AI21's GET /studio/v1/models answers 200 with an empty data array without a key, so the model ids come from AI21's public docs (Jamba models page, retrieved 2026-09-29) and the roster has not been checked. No account was created and no API key was used to build this entry, and no POST request was sent. evidence.liveMatrix is null until someone runs the live capability matrix with a real key (see Verification status below).


Key Facts​

  • Provider id: ai21
  • Protocol: chat completions at POST https://api.ai21.com/studio/v1/chat/completions with model and messages, Bearer authentication (Chat request, Authentication)
  • Base URL: https://api.ai21.com/studio/v1
  • Default model: jamba-large
  • Models in catalog: 6 — the six ids named in the API Versioning list of the Jamba models page
  • Context: the models page lists Max Tokens 256K for Jamba Large and Jamba Mini; the catalog stores the number 256,000 as a placeholder for that 256K label
  • Max output: the Chat request page states "For Jamba models, the maximum allowed value is 4096 tokens." for max_tokens
  • Streaming: supported — the stream parameter is documented as "Stream results one token at a time using server-sent events." (Chat request); the Chat response page says "The final message will be data: [DONE]."
  • Tool calling: true — the Function Calling page says "Function calling is available for all Jamba models via the Chat Completions API."
  • Tools while streaming: false — the stream parameter text on the Chat request page says "Must be False if using tools."
  • Structured output: true, based on JSON mode — the Chat request page documents response_format: "Setting it to { "type": "json_object" } activates JSON mode, ensuring the generated message adheres to valid JSON structure." NeuroLink's chat-completions client builds a json_schema response_format when a schema is supplied; that form was not sent to AI21 because no POST request was sent.
  • Structured output + tools together: false — no combined tools + schema request was sent
  • Embeddings: not declared (false)
  • Thinking: not declared (false)
  • Rate limits: the Rate Limits page lists Jamba Large and Jamba Mini at 10 RPS and 200 RPM on the AI21 Platform
  • Billing: free-tier — see Billing below
  • Key format: none declared (apiKeyFormat is null)

Quick Start​

1. Get an API key​

  1. Visit: https://docs.ai21.com/docs/create-api-key and follow its steps: sign in to AI21 Studio (the page says "You can sign in using your email address, Google account, GitHub account, or SSO."), click Settings in the bottom-left corner of the sidebar, open the API Keys tab, then click Create new key
  2. Copy the key when it is shown; the page says "you won’t be able to view it again"
  3. Free trial: https://docs.ai21.com/docs/usage-cost says "New accounts are given a $10 credit good for three months on the AI21 platform." and, once the trial is exceeded or expires, "you must provide valid billing information in your account to continue using the AI21 platform". https://www.ai21.com/pricing shows "$10 credits for 7 days." and "No credit card needed" under Free Trial
  4. Set AI21_API_KEY in your .env file

2. Configure​

export AI21_API_KEY=your-api-key
export AI21_MODEL=jamba-large # optional — overrides the default model
export AI21_BASE_URL=https://api.ai21.com/studio/v1 # optional — proxy or self-hosted gateway

3. Use it​

import { NeuroLink } from "@juspay/neurolink";

const neurolink = new NeuroLink();

const result = await neurolink.generate({
input: { text: "Explain context windows in one paragraph." },
provider: "ai21",
model: "jamba-large",
});

console.log(result.content);
# CLI
npx @juspay/neurolink generate "Hello" --provider ai21

Per-request credentials are passed like this:

await neurolink.generate({
input: { text: "Hello" },
provider: "ai21",
credentials: { ai21: { apiKey: process.env.AI21_API_KEY } },
});

Billing​

AI21's docs and website state different free-trial durations:

  • Pricing (docs): "New accounts are given a $10 credit good for three months on the AI21 platform." After the trial is exceeded or expires, "you must provide valid billing information in your account to continue using the AI21 platform."
  • Pricing page (website): "Start with a free trial, no credit card required, then pay based on usage." Under Free Trial it lists "$10 credits for 7 days." and "No credit card needed".

The entry records free-tier. The Account page says "You will need to provide billing information when your introductory usage credits expire."


Models​

ModelContextMax outputVision$/M in · outNotes
jamba-large ⭐256K4,096no$2 / $8NeuroLink default. The models page: "Our most powerful and advanced model, designed to handle complex tasks at enterprise scale with superior performance."
jamba-mini256K4,096no$0.2 / $0.4NeuroLink fallback. The pricing page: "Efficient & lightweight model for a wide range of tasks".
jamba-large-1.7-2025-07256K4,096no—Dated snapshot of Jamba Large (Version 1.7, Snapshot 2025-07); jamba-large currently points to it.
jamba-mini-2-2026-01256K4,096no—Dated snapshot of Jamba Mini (Version 2, Snapshot 2026-01); jamba-mini currently points to it.
jamba-large-1.7256K4,096no—Version name; points to jamba-large-1.7-2025-07.
jamba-mini-2256K4,096no—Version name; points to jamba-mini-2-2026-01.

Prices for jamba-large and jamba-mini come from AI21's pricing page ("Jamba Large" at $2 / 1M input tokens and $8 / 1M output tokens; "Jamba Mini" at $0.2 / 1M input tokens and $0.4 / 1M output tokens), matched to the ids through the Model Details table on the Jamba models page (retrieved 2026-09-29). That page also lists the API Versioning mapping used in the notes above and gives Input Modality: Text, so the catalog records vision: false for these models. The same page states "We advise using dated versions of the Jamba API to avoid disruptions from model updates and breaking changes."

models.defaultContextWindow and the per-model contextWindow values (256,000) are a numeric placeholder for the 256K label printed on the Jamba models page; no exact token count is taken from a vendor page. models.defaultMaxOutputTokens (4,096) is the max_tokens limit stated on the Chat request page.

Fallback order when the default is unavailable: jamba-mini → jamba-large-1.7-2025-07 → jamba-mini-2-2026-01. The runtime fallback model name is set explicitly to jamba-mini.


Verification status​

Tier-2 onboarding requires evidence before a provider is accepted, and pnpm run verify:provider-onboarding gates it in CI. This is what the catalog currently records for AI21 Labs — docs-verified only, not live-verified:

ProbeResult
Rosterunauthenticated GET /studio/v1/models, HTTP 200, body {"data":[]}, 2026-09-29 — no key used; model ids come from the Jamba models page, not from the API
Chat endpointunauthenticated GET /studio/v1/chat/completions, HTTP 405, body {"detail":"Method Not Allowed"}, header allow: POST, 2026-09-29
Billingusage-cost and pricing pages, 2026-09-29 — the two pages state different trial terms, quoted under Billing
Auth-failure shapeunauthenticated GET /studio/v1/library/files (a documented endpoint), HTTP 400, body {"detail":"Bad request: bad or missing authentication."}, 2026-09-29. The Chat response page lists 401, 403, 422, 429, 500 and 503. errorRules match status 401, status 429, status 503 and the text bad or missing authentication
Tools / structured outputfunction calling is documented on the Function Calling page and response_format with { "type": "json_object" } on the Chat request page. Neither was exercised live. No POST request was sent, so stream_options on streaming calls, tool_choice on calls with tools and a json_schema response_format, which NeuroLink's chat-completions client can add, have not been exercised against AI21, and no combined tools + schema request was sent — structuredOutputWithTools stays false
Live capability sweepnot run. evidence.liveMatrix is null.

Troubleshooting​

SymptomCauseFix
HTTP 401The Chat response page documents 401 as "Unauthorized (Incorrect API key provided/Invalid Authentication)"The Authentication page says to use your API key as the Bearer token; NeuroLink reads it from AI21_API_KEY
HTTP 429The Chat response page documents 429 as "Too Many Requests (You are sending requests too quickly.)"The Rate Limits page lists 10 RPS and 200 RPM for Jamba Large and Jamba Mini and says to contact [email protected] if you require a higher limit
HTTP 503The Chat response page documents 503 as "Service Unavailable (The engine is currently overloaded, please try again later)"The same text says "please try again later"
Trial credit exceeded or expiredThe usage-cost page: "you must provide valid billing information in your account to continue using the AI21 platform"Add billing information in your account; the page points to Account > Billing & Plans

See also​