Skip to main content

Nebius Token Factory Provider Guide

Nebius Token Factory is a Tier-2 catalog provider: OpenAI-wire-compatible, with no catalog quirks set, so its entire integration is one JSON file (src/lib/providers/catalog/nebius.json) rather than hand-written code. That file is the source of truth for everything on this page.

Verification status: this entry is docs-verified, not yet live-verified. The vendor's GET /v1/models needs a key (an unauthenticated request answers HTTP 401), so the model ids come from the vendor's public model catalog and the roster has not been checked against the API. No account was created and no API key was used to build it. evidence.liveMatrix is null until someone runs the live capability matrix with a real key (see Verification status below).


Key Facts​

  • Provider id: nebius
  • Protocol: OpenAI-compatible (/chat/completions) — the API introduction says "Nebius Token Factory offers an OpenAI-compatible API for inference and fine-tuning." (docs)
  • Base URL: https://api.tokenfactory.nebius.com/v1
  • Default model: Qwen/Qwen3-235B-A22B-Instruct-2507
  • Models in catalog: 7, listed in the Models table below
  • Streaming: supported — the chat completion reference documents stream, with data-only server-sent events and a stream "terminated by a data: [DONE] message" (reference)
  • Tool calling: model-dependent — the Function calling & Tools page documents tools and tool_choice, and the public model catalog lists function_calling in use_cases for 6 of the 7 models in this entry (not for openbmb/MiniCPM-V-4_5)
  • Tools while streaming: declared (true) — the reference's streaming delta schema has a tool_calls field described as "The tool calls generated by the model, such as function calls."
  • Structured output: declared (true) — the Structured output & JSON page documents response_format with json_schema and json_object, and says "Use JSON mode tag in on a model card to find a model with structured output supported." The public catalog carries the JSON mode tag on 6 of the 7 models here: the default, the four fallbacks and openbmb/MiniCPM-V-4_5
  • Structured output + tools together: not declared (false) — no combined probe was possible without credentials
  • Vision: openbmb/MiniCPM-V-4_5 is vision: true (catalog type image2text with image in use_cases). The Vision capabilities example passes an image as image_url, either "A URL to the image" or "A base64 encoded image directly in the request."
  • Embeddings: not declared on this catalog entry. The public model catalog lists an embedding model type, but NeuroLink's generic ConfiguredOpenAICompatProvider (which Tier-2 catalog entries use) has no native embed()/embedMany() — BaseProvider's default throws — so this flag tracks that, not the vendor's own API surface
  • Thinking: not declared (false) — the reference lists a reasoning_effort request field (none, minimal, low, medium, high, xhigh, max) and reasoning_content / reasoning response fields; no request using them was sent
  • Billing: free-with-card — the billing page says "Setting up a billing account requires a bank card." and "Upon first sign-up, you receive $1 in trial credit, valid for 30 days."
  • Key format: none declared

Quick Start​

1. Get an API key​

  1. Visit: https://tokenfactory.nebius.com/ and create an account; the quickstart (https://docs.tokenfactory.nebius.com/quickstart) says: "Log in using your Google or GitHub account"
  2. Create an API key (https://docs.tokenfactory.nebius.com/api-reference/introduction#authentication): go to the API keys section, click Create API key, enter the key name, click Create, and save the displayed key — "Save the displayed API key. You cannot open it later in Nebius Token Factory."
  3. Billing (https://docs.tokenfactory.nebius.com/other-capabilities/billing-new): "Billing setup is mandatory—you cannot complete onboarding without it." "Setting up a billing account requires a bank card." "Upon first sign-up, you receive $1 in trial credit, valid for 30 days."
  4. Set NEBIUS_API_KEY in your .env file (the vendor's docs read NEBIUS_API_KEY)

2. Configure​

export NEBIUS_API_KEY=your-api-key
export NEBIUS_MODEL=Qwen/Qwen3-235B-A22B-Instruct-2507 # optional — overrides the default model
export NEBIUS_BASE_URL=https://api.tokenfactory.nebius.com/v1 # optional — proxy or self-hosted gateway

3. Use it​

import { NeuroLink } from "@juspay/neurolink";

const neurolink = new NeuroLink();

const result = await neurolink.generate({
input: { text: "Explain context windows in one paragraph." },
provider: "nebius",
model: "Qwen/Qwen3-235B-A22B-Instruct-2507",
});

console.log(result.content);
# CLI
npx @juspay/neurolink generate "Hello" --provider nebius

Per-request credentials:

await neurolink.generate({
input: { text: "Hello" },
provider: "nebius",
credentials: { nebius: { apiKey: process.env.NEBIUS_API_KEY } },
});

Models​

ModelContext (max_model_len)VisionPrice per M tokens, in / outNotes
Qwen/Qwen3-235B-A22B-Instruct-2507 ⭐262,144no0.2 / 0.6NeuroLink default; JSON mode tag; function_calling in use_cases.
deepseek-ai/DeepSeek-V4-Flash-07311,024,000no0.14 / 0.28Fallback; JSON mode tag; function_calling in use_cases.
openai/gpt-oss-120b131,072no0.15 / 0.6Fallback; JSON mode tag; function_calling in use_cases.
Qwen/Qwen3-30B-A3B-Instruct-2507262,144no0.1 / 0.3Fallback; JSON mode tag; function_calling in use_cases.
NousResearch/Hermes-4-405B131,072no1 / 3Fallback; JSON mode tag; function_calling in use_cases.
nvidia/Nemotron-3_5-Lightning1,048,576no0.06 / 0.24function_calling in use_cases.
openbmb/MiniCPM-V-4_532,000yes0.658 / 1.11JSON mode tag; catalog type image2text; models.visionModel.

Model ids, context lengths and prices come from the vendor's public model catalog, retrieved 2026-09-29: /api/public/models_info (its Markdown view is at /model-catalog.md, which says "The JSON endpoint is the authoritative machine-readable source"). Context is each model's max_model_len; the catalog also gives a rounded context_window_k. Prices are the catalog's input_price_per_million_tokens and output_price_per_million_tokens values, copied as published. The catalog's status: active is recorded here as production.

The default and the four fallbacks are the models in this entry whose catalog record carries both the JSON mode tag and function_calling in use_cases; the other two follow in topModels.

Vendor descriptions (the public model catalog's description field, quoted):

  • Qwen/Qwen3-235B-A22B-Instruct-2507 — "Balanced Qwen3 flagship tuned for strong general reasoning, chat quality, and tool use."
  • deepseek-ai/DeepSeek-V4-Flash-0731 — "DeepSeek V4 Flash 0731 is a 1M-context reasoning model designed for coding and agentic workloads."
  • openai/gpt-oss-120b — "Open-weight agentic model with configurable reasoning, full CoT visibility, strong tool use, and fine-tuning support."
  • Qwen/Qwen3-30B-A3B-Instruct-2507 — "Versatile 30B instruct model optimized for high-quality chat, reasoning, and coding."
  • NousResearch/Hermes-4-405B — "Hybrid-reasoning model trained on verified CoT traces for strong math, coding, and step-by-step reliability."
  • nvidia/Nemotron-3_5-Lightning — "NVIDIA's 30B-parameter hybrid MoE model with 3B active parameters per token, designed for efficient agentic reasoning, tool use, coding, and long-context workflows."
  • openbmb/MiniCPM-V-4_5 — "MiniCPM-V-4.5 – Compact multimodal model for image, multi-image, high-FPS/long-video, OCR/PDF understanding, with switchable fast/deep thinking."

models.defaultMaxOutputTokens (8,192) is the value the chat completion reference gives for an omitted max_tokens: "If omitted or set to null, defaults to 8192 tokens." models.defaultContextWindow (32,000) is a placeholder the vendor does not publish.

The inference overview describes two model flavors, Base and Fast, and says: "To use the Fast flavor, append -fast to the model name in the API." The ids in this entry are the ones in the public model catalog's flavors[].model_id. Code samples in the docs name other ids, for example deepseek-ai/DeepSeek-R1-0528 on the quickstart and meta-llama/Meta-Llama-3.1-70B-Instruct on the API introduction.

Nebius records serverless model removals in its June 2026 and August 2026 deprecation notices, both of which say "Token Factory does not automatically reroute requests from deprecated models." The August notice lists deepseek-ai/DeepSeek-V4-Flash-0731 and nvidia/Nemotron-3_5-Lightning in its "Recommended Replacement" column.

Fallback order when the default is unavailable: deepseek-ai/DeepSeek-V4-Flash-0731 → openai/gpt-oss-120b → Qwen/Qwen3-30B-A3B-Instruct-2507 → NousResearch/Hermes-4-405B. The runtime fallback model name the loader derives (fallbacks[1]) is openai/gpt-oss-120b.


Verification status​

Tier-2 onboarding requires evidence before a provider is accepted, and pnpm run verify:provider-onboarding gates it in CI. This is what the catalog currently records for Nebius Token Factory — docs-verified, not live-verified:

ProbeResult
Rosterunauthenticated GET /v1/models answered HTTP 401 with {"detail":"Couldn't authenticate. Reason: token is not present"}, 2026-09-29 — no key used. Model ids are taken from the public model catalog (/api/public/models_info), so the roster is not verified against the API
Chat endpointunauthenticated GET /v1/chat/completions answered HTTP 405 with allow: POST and {"detail":"Method Not Allowed"}, 2026-09-29; no POST was sent
Billingthe billing page states a bank card is required to set up a billing account and a $1 trial credit valid for 30 days on first sign-up, 2026-09-29
Error shapesstatuses and example bodies from the chat completion reference: 401 Couldn't authenticate. Reason: token is not present, 402 Payment Required: You have exhausted your budget. Please add funds to continue using the API., 404 The model `unknown-model` does not exist., 429 (response description "Rate limit exceeded"). errorRules match the status codes 401, 402, 404 and 429; the 429 rule points to the Retry-After header on the rate-limits page
Tools / structured outputdocumented on the function-calling and structured-output pages; neither was exercised live, and no combined tools+schema request was sent — structuredOutputWithTools stays false. The chat completion reference's response_format description reads Only {'type': 'json_object'} or {'type': 'text' } is supported., while its ResponseFormat schema lists text, json_object and json_schema
Live capability sweepnot run. evidence.liveMatrix is null. Before treating this provider as production-ready, run npx tsx test/continuous-test-suite-provider-matrix.ts --provider=nebius with a real key and record the result.

Do not treat this entry as equivalent to a live-verified Tier-2 provider (e.g. FriendliAI, Novita AI) until that live matrix has been run and evidence.liveMatrix is filled in.


Troubleshooting​

Each row uses a status and body the vendor's chat completion reference shows, or a statement from a vendor page named in the row.

SymptomCauseFix
HTTP 401The reference's 401 example body is Couldn't authenticate. Reason: token is not presentThe API introduction says to include the key in the Authorization: Bearer header; NeuroLink sends NEBIUS_API_KEY that way
HTTP 402The reference's 402 example body is Payment Required: You have exhausted your budget. Please add funds to continue using the API.Add funds — see "Topping up your balance" on the billing page
HTTP 404The reference's 404 example body is The model `unknown-model` does not exist.Take the id from the public model catalog
HTTP 429The rate-limits page says "If you exceed the active limit you will receive HTTP 429."The same page describes the Retry-After header as "The time in seconds to wait before making another request if the rate limit is exceeded."
Structured output ignored with tools attachedstructuredOutputWithTools is false on this entry — untested combinationNeuroLink omits response_format automatically whenever tools are present, before sending

See also​