Skip to main content

W&B Inference Provider Guide

W&B Inference (the vendor's docs call it Serverless Inference) is a Tier-2 catalog provider: OpenAI-wire-compatible with no behavioural quirks, so its entire integration is one JSON file (src/lib/providers/catalog/wandb-inference.json) rather than hand-written code. That file is the source of truth for everything on this page.

Verification status: this entry is docs-verified only, not yet live-verified. The vendor's GET /v1/models needs a key (an unauthenticated GET answers HTTP 401), so the model ids come from the vendor's public models page and the roster has not been checked. No account was created and no API key was used to build this entry. evidence.liveMatrix is null until someone runs the live capability matrix with a real key (see Verification status below).


Key Facts​

  • Provider id: wandb-inference (alias wandb)
  • Protocol: OpenAI-compatible (/chat/completions) — the chat completions reference says the endpoint "follows the OpenAI format for sending messages and receiving responses"
  • Base URL: https://api.inference.wandb.ai/v1
  • Default model: openai/gpt-oss-120b
  • Models in catalog: 12 (taken from the 20 rows of the models page's "Generally available models" table)
  • Streaming: supported — the streaming page says "All hosted models support streaming output."
  • Tool calling: model-dependent — the tool-calling page documents function calling ("Serverless Inference only supports calling functions") with an example on openai/gpt-oss-20b; the models page describes ibm-granite/granite-4.2-8b as "capable of enhanced tool calling" and google/gemma-4-26B-A4B-it as having "function calling for agentic workflows"
  • Tools while streaming: not declared (false)
  • Structured output: supported — the structured-output page documents json_schema as the response_format type, and the JSON-mode page documents json_object. The vendor's 422 article says "Some parameters (such as frequency_penalty, logprobs, or response_format) are not supported by all models."
  • Structured output + tools together: not declared (false) — no combined request was sent; the requests made were unauthenticated GETs
  • Embeddings: not declared — the API reference lists two methods, Chat Completions and List Models
  • Thinking: not declared (false). The vendor's reasoning page says reasoning information appears in the reasoning field of responses and that chat_template_kwargs.enable_thinking turns it on or off for models that allow toggling; this entry does not set that flag
  • Billing: free-tier — see Billing below
  • Key format: none declared

Quick Start​

1. Get an API key​

  1. Visit: https://app.wandb.ai/login?signup=true and sign up for a W&B account (prerequisites)
  2. Create an API key: click your user profile icon, then User Settings, then Create new API key, and copy it immediately because W&B shows the full key once (prerequisites); the API reference also names https://wandb.ai/settings
  3. Project: the prerequisites page lists a W&B project among the items you must have: "Create a project in your W&B account to track usage." The chat completions reference lists "Your W&B team and project: [YOUR-TEAM]/[YOUR-PROJECT] (optional)."
  4. Credits: "Serverless Inference credits come with Free, Pro, and Academic plans for a limited time." Free accounts must activate pay-as-you-go inference on the Billing tab or upgrade to a paid plan when credits run out; the default cap for a Free account is $100/month (usage limits)
  5. Set WANDB_INFERENCE_API_KEY in your .env file

2. Configure​

export WANDB_INFERENCE_API_KEY=your-api-key
export WANDB_INFERENCE_MODEL=openai/gpt-oss-120b # optional — overrides the default model
export WANDB_INFERENCE_BASE_URL=https://api.inference.wandb.ai/v1 # optional — proxy or self-hosted gateway

The chat completions reference shows an optional team and project for usage tracking (project="[YOUR-TEAM]/[YOUR-PROJECT]" on the Python OpenAI client, or an OpenAI-Project header). The API reference says: "If you don’t specify these, W&B uses your default entity and the project name inference." The catalog schema has no field for extra request headers, so this entry does not send one.

3. Use it​

import { NeuroLink } from "@juspay/neurolink";

const neurolink = new NeuroLink();

const result = await neurolink.generate({
input: { text: "Explain context windows in one paragraph." },
provider: "wandb-inference",
model: "openai/gpt-oss-120b",
});

console.log(result.content);
# CLI
npx @juspay/neurolink generate "Hello" --provider wandb-inference

Per-request credentials use the same shape as for the other catalog providers:

await neurolink.generate({
input: { text: "Hello" },
provider: "wandb-inference",
credentials: {
wandbInference: { apiKey: process.env.WANDB_INFERENCE_API_KEY },
},
});

Billing​

The usage limits page says "Serverless Inference credits come with Free, Pro, and Academic plans for a limited time. Enterprise availability may vary." When credits run out, Free accounts "must activate pay-as-you-go inference on the Billing tab, or upgrade to a paid plan to continue using Serverless Inference", and the page adds "W&B requires prepayment for paid Inference access." Its account-tier table lists a default cap of $100/month for Free, $6,000/month for Pro and $700,000/year for Enterprise. Per-token prices are on the pricing page, which links from the docs as https://wandb.ai/site/pricing/inference. The entry records free-tier.


Models​

ModelContextVision$/M in · outNotes
openai/gpt-oss-120b ⭐131,072no$0.03 / $0.17NeuroLink default. "Efficient Mixture-of-Experts model designed for high-reasoning, agentic and general-purpose use cases." Reasoning page: "Always on".
deepseek-ai/DeepSeek-V4-Flash-0731262,144no$0.13 / $0.28 (cache hit $0.07)"an MoE model great for coding, reasoning, and agentic workloads."
meta-llama/Llama-3.3-70B-Instruct128,000no$0.71 / $0.71"Multilingual model excelling in conversational tasks, detailed instruction-following, and coding."
openai/gpt-oss-20b131,072no$0.03 / $0.13"Lower latency Mixture-of-Experts model trained on OpenAI's Harmony response format with reasoning capabilities." Example model on the tool-calling, structured-output and JSON-mode pages.
meta-llama/Llama-3.1-8B-Instruct131,072no$0.22 / $0.22"Efficient conversational model optimized for responsive multilingual chatbot interactions." Model in the quick-start example on the Serverless Inference page.
zai-org/GLM-5.21,048,576no$0.76 / $2.42 (cache hit $0.14)"a Mixture-of-Experts language model featuring 40 billion activated parameters and a total of 744 billion parameters."
deepseek-ai/DeepSeek-V4-Pro-08131,048,576no$1.31 / $3.96 (cache hit $0.044)"a 1.6T-parameter MoE model excelling at advanced reasoning, coding, and complex agentic workloads."
google/gemma-4-26B-A4B-it262,144yes$0.10 / $0.30 (cache hit $0.05)"a multimodal MoE model with LoRA support and function calling for agentic workflows." models.visionModel.
Qwen/Qwen3.8-27B262,144yes$0.40 / $3.00 (cache hit $0.15)"a dense multimodal model suited for coding, research, vision, and long-running agent tasks."
moonshotai/Kimi-K2.6262,144yes$0.65 / $3.41 (cache hit $0.15)"a multimodal Mixture-of-Experts language model featuring 32 billion activated parameters and a total of 1 trillion parameters."
deepseek-ai/DeepSeek-V4.1-Flash1,048,576yes$0.20 / $0.65 (cache hit $0.03)"a multimodal MoE model for coding, reasoning, and agentic workloads with long contexts."
ibm-granite/granite-4.2-8b131,072no$0.10 / $0.15 (cache hit $0.05)"an instruct model capable of enhanced tool calling, instruction following, and chat capabilities."

Context comes from the contextWindow field of the vendor's public catalog API, GET https://trace.wandb.ai/inference/catalog/models (called unauthenticated on 2026-09-29; the vendor's Service API overview gives https://trace.wandb.ai as the base URL, and the endpoint page says "This API is available without authentication."). The models page shows the same figures in rounded form, for example "262k". Prices come from the pricing page, shown per 1 million tokens; its rows are named as on the models page, and a - in its cache hit column means no cache-hit price is recorded here. Notes quote the "Description" column of the models page, retrieved 2026-09-29.

models.defaultContextWindow (128,000) and models.defaultMaxOutputTokens (8,192) are placeholders: the vendor does not publish a general default or a maximum output figure, and no per-model maxOutputTokens is set.

Fallback order when the default is unavailable: deepseek-ai/DeepSeek-V4-Flash-0731 → meta-llama/Llama-3.3-70B-Instruct → openai/gpt-oss-20b → meta-llama/Llama-3.1-8B-Instruct → zai-org/GLM-5.2 → deepseek-ai/DeepSeek-V4-Pro-0813 → google/gemma-4-26B-A4B-it → Qwen/Qwen3.8-27B → moonshotai/Kimi-K2.6 → deepseek-ai/DeepSeek-V4.1-Flash → ibm-granite/granite-4.2-8b.

Vision: vision: true marks the rows whose "Type" column on the models page reads "Text, Vision"; rows that read "Text" are vision: false. No image request was sent to confirm either.

Lifecycle: the lifecycle page defines "Generally available" as "The model is fully supported and recommended for use." and says requests to retired models "fail and return an HTTP 404 status code." The 12 models in this catalog are rows of the models page's "Generally available models" table.


Verification status​

Tier-2 onboarding requires evidence before a provider is accepted, and pnpm run verify:provider-onboarding gates it in CI. This is what the catalog currently records for W&B Inference — docs-verified only, not live-verified:

ProbeResult
Rosterunauthenticated GET https://api.inference.wandb.ai/v1/models answered HTTP 401 on 2026-09-29, so the roster has not been checked; model ids are taken from the models page (retrieved 2026-09-29)
Public catalog cross-checkunauthenticated GET https://trace.wandb.ai/inference/catalog/models answered HTTP 200 with 47 models on 2026-09-29; the 12 catalog ids appear with lifecycleStage general-availability, and their contextWindow and price fields agree with the pricing page
Billingthe usage limits page states credits "come with Free, Pro, and Academic plans for a limited time", 2026-09-29
Auth-failure shapeunauthenticated GET /v1/models and unauthenticated GET /v1/chat/completions both answered HTTP 401 with {"error":{"code":"invalid_api_key","message":"Missing bearer authentication in header","type":"invalid_request_error"}} on 2026-09-29. The errorRules for 402, 403, 404 and 429 follow the vendor's support articles; those responses were not observed
Tools / structured outputdocumented on the tool-calling, structured-output and JSON-mode pages; neither was exercised live, and no combined tools+schema request was sent — structuredOutputWithTools stays false
Live capability sweepnot run. evidence.liveMatrix is null. Before treating this provider as production-ready, run npx tsx test/continuous-test-suite-provider-matrix.ts --provider=wandb-inference with a real key and record the result.

Do not treat this entry as equivalent to a live-verified Tier-2 provider (e.g. FriendliAI, Novita AI) until that live matrix has been run and evidence.liveMatrix is filled in.


Troubleshooting​

Each row restates a vendor support article.

SymptomCause (vendor)Fix (vendor)
HTTP 401 "Authentication failed"An API key that is "invalid, expired, or revoked", or an incorrect W&B project entity name or project name (401 article)Regenerate the key from your W&B settings; check the project entity and name
HTTP 402 "You exceeded your current quota"No remaining credits, or the monthly spending cap was reached (402 article)Check plan and billing details, get more credits, or increase limits
HTTP 403 "Country, region, or territory not supported"Access from an unsupported location, determined from the IP address of the request (403 article)Check the geographic restrictions
HTTP 403 "The inference gateway is not enabled for your organization"The organization has not enabled the inference gateway (403 article)Ask your organization's W&B administrator to enable it
HTTP 404A model id that does not match an available model, or a model that was deprecated or removed (404 article)Verify the id against the models page; ids are case-sensitive
HTTP 400 or 422A parameter the model does not support, an out-of-range value, or a malformed messages payload (422 article)Read the response body before changing the request
HTTP 429 "Concurrency limit reached for requests"Too many concurrent requests (429 article)Reduce concurrent requests; use exponential backoff when retrying
HTTP 503 "The engine is currently overloaded"High traffic (503 article)Retry after a short delay with exponential backoff

See also​