Credential-free provider acceptance gate
Suite: test/continuous-test-suite-acceptance-gate.ts
Run: pnpm run test:acceptance-gate
Script: test:acceptance-gate in package.json
Mock server: test/helpers/acceptanceGateServer.ts
Why this exists
The prior full provider capability matrix (test/continuous-test-suite-provider-matrix.ts,
1,037 cells) was audited and found to be able to report green without proving
anything, in at least 10 ways: tool cells that passed with zero tool calls
made, no disableInternalFallback and no provider/model identity check on
any cell, structured-output cells that asserted a value's type rather than
its contents, and a harness that could exit 0 on zero coverage.
This gate replaces that surface area with a small, falsifiable one: at most 9 cells per (provider, pinned model), fail-fast, cheapest first, run entirely against a local mock vendor HTTP server — no real vendor calls, no real API keys, ever. Every provider is pointed at the mock through its own documented base-URL environment override, with a fake key, so the SDK and the CLI both round-trip through the exact same server (that's what makes cell 9, CLI parity, meaningful instead of a re-derivation).
The 9 cells
Each cell asserts an exact value the mock server returns, never a shape
or a "non-empty" check. A row stops (remaining cells report SKIP, not a
second FAIL) the first time a cell genuinely fails — this is the
"fail-fast" contract; a capability-gated skip (e.g. a provider that doesn't
declare tools) never counts as a failure and never poisons the rest of the
row.
- Identity pin — every call carries
{ provider, model, disableInternalFallback: true }; the result'sproviderandmodelfields are asserted against what was requested. - Exact-output generate —
generate()'scontentis asserted equal to a specific constant (GATE_EXACT_VALUE), not merely non-empty. - Drained stream, identity asserted — the stream is fully drained and
the accumulated content is asserted equal to a constant. The
StreamResult'sprovider/modelare then asserted — and for the OpenAI-compatible protocol rows, against what the mock server actually served (${model}::mock-server-resolved), not what was requested. See "The stream-identity defect" below — this cell is the reason that defect was caught and fixed. - Tool-nonce proof — the test tool's
execute()returns a freshrandomUUID()nonce that exists nowhere else. The final answer is asserted to containACCEPTANCE_GATE_TOOL_CONFIRMED:<nonce>, which the mock server only emits once it sees that exact nonce echoed back in a tool-result message — i.e. the assertion can only pass if a real tool call round-tripped through the model. - Structured-exact —
generate({ schema })'s parsedstructuredDatais asserted equal (viaJSON.stringify) to an exact object, andjsonTruncatedis asserted nottrue— truncation is surfaced, never silently swallowed. - Thinking proof — only for providers whose catalog/descriptor declares
thinking: true(currentlyanthropicand thedeepseekcatalog row). Drivesstream()with an explicitthinkingConfig: { enabled: true, budgetTokens: 2048 }(the SDK path has no barethinkingLevelconvenience — that folding only exists on the CLI's option parser) and asserts the accumulatedreasoningdeltas and the post-thinking answer both equal exact constants. - Embeddings proof — only for providers whose row declares
embeddings: true. Goes throughProviderFactory.createProvider()directly (NeuroLinkhas noembed()method) and asserts the returned vector equals an exact constant array. - Budget ceiling enforced — a dedicated, low-ceiling (
1) mock server instance proves the ceiling mechanism itself is real: a first call succeeds, a second call against the same server must fail, and the server's own request counter is asserted to have observed both. This runs against its own server, never the shared one every other cell uses, so it can't be tripped by unrelated traffic. - CLI parity — the same identity + exact-output assertions as cells 1-2,
but driven through the built CLI
(
node dist/cli/index.js generate ... --format json), parsing its JSON stdout, against the same mock server. Not a separate re-derivation of the assertions — literally the same constants.
Known, documented gap in cell 9
The CLI's generate/stream commands have no flag that reaches
disableInternalFallback — commandFactory.ts's processOptions()
whitelist omits it (only the interactive REPL's separate options schema
supports it). Cell 9 therefore runs without disableInternalFallback. This
is an intentional, documented gap, not an oversight: fixing it is a CLI
option-surface change out of scope for this gate.
What keeps the gap harmless is that the suite holds no real credential to
fall back to. test/helpers/credentialFreeEnv.ts is its first import: it
points DOTENV_CONFIG_PATH at /dev/null, so neither the SDK's nor the
harness's .env load reads a developer's .env, it deletes every
credential-named variable the shell exported, and it points HOME at an
empty temp directory. The last one matters as much as the first two: the
Anthropic provider prefers a stored OAuth token
(~/.neurolink/anthropic-credentials.json) over ANTHROPIC_API_KEY and
refreshes it against Anthropic's live token endpoint, so a gate that only
cleaned the environment could still send a real refresh token to a real
host. The suite checks the environment and the home directory before it
runs a row, and the CLI child inherits that stripped environment plus the
current row's fake key.
The stream-identity defect (cell 3)
Confirmed and fixed in src.
StreamResult.model (the public value cell 3 asserts against) was reporting
the requested model, not the model the server actually served. Internally,
neurolink.ts already captured the resolved model — the bug was that this
resolved value never reached the public StreamResult.
- OpenAI-compatible protocol path (
src/lib/providers/openaiChatCompletionsBase.ts,src/lib/types/openaiCompatible.ts) — the substantive fix, and the one cell 3 actually exercises: addedStreamLoopArgs.onModelObserved?: (model: string) => void, fired once per step that echoes amodelfield on the wire (SSEchunk.model, captured byparseSSEStreamintoOpenAICompatSSEResult.model).executeStream's publicresult.modelwas a plain data property fixed at the pre-call resolved/requestedmodelId; it is now a live getter (observedServerModel ?? modelId) so a caller reading.modelafter draining.streamsees what the server actually served — the mock's::mock-server-resolvedsuffix on theSTREAMcell is what makes this difference observable and the fix provable. It has to be a getter, not a reassignment after the fact:preserveLiveStreamAccessorsspreadsresultinto a new object at each provider boundary before the background loop produces its first chunk, which would freeze a plain field at its construction-time value. - Anthropic path (
src/lib/providers/anthropic/client.ts) — a narrower, related fix:StreamResult.modelwas reporting the raw constructor argumentthis.modelName, which isundefinedwhenever a caller relies on the default model rather than passing one explicitly. It now reportsmodelId, the value already resolved asthis.modelName || getDefaultAnthropicModel()and actually put on the wire. Note this is a different bug from the OpenAI-compatible one: the Anthropic Messages API always echoes the request'smodelback verbatim in its envelope, so there is no "server resolved a different model than requested" case on this protocol for cell 3 to catch — the mock'santhropicrow therefore asserts the request-echoed value, not a server-observed one. This fix is real and independently correct (it fixes default-model callers), but it is not the same class of defect cell 3's OpenAI-compatible assertion proves, and cell 3 does not exercise it for the coveredanthropicrow specifically, since that row always passes an explicitmodel.
The mock server deliberately makes this defect observable rather than
hiding it: on the OpenAI-compatible protocol, the STREAM cell's response
appends ::mock-server-resolved to the requested model in both the content
and finish SSE chunks, simulating a gateway/router rewriting an alias to a
concrete served model id. Cell 3 asserts the server-observed value for
protocol: "openai" rows and the request-echoed value for protocol: "anthropic" rows (per the Messages API's own envelope semantics).
Additional src defects found and fixed by cells 4 and 8
Building this gate surfaced three more real defects, all in provider-side tool-capability allowlists and one test-sizing constant — none in the mock itself. Each is a one-line, strictly capability-adding fix.
ollamacell 4 (tool-nonce proof) —OllamaProvider.supportsTools()checks the model name againstmodelConfiguration.ts'screateOllamaConfig().modelBehavior.toolCapableModelsallowlist. That list (OLLAMA_TOOL_CAPABLE_MODELS) included"llama3.1"but not"llama3.2"— the gate's own pinned default Ollama model (providerMatrix.ts'sollama.defaultModel, which declarestools: true).supportsTools()returnedfalsefor the provider's own default model, which silently dropped the user-supplied tool from both the wiretoolsfield and the execution lookup (executeNativeGenerate()'stoolsRecordis built empty whensupportsTools()is false), producing a"Tool not found"error instead of a real tool round-trip. Fixed insrc/lib/core/modelConfiguration.tsby adding"llama3.2"to the defaulttoolCapableModelsarray.openroutercell 4 (tool-nonce proof) —OpenRouterProvider.supportsTools()has two paths: a dynamically cached capability set (populated bycacheModelCapabilities(), which fetches/modelsand readssupported_parameters) and a hardcoded fallback pattern list used until that cache is populated.cacheModelCapabilities()is never invoked automatically anywhere in the realgenerate()/stream()code paths, so the fallback list is the only path ever exercised in practice. That list included"meta-llama/llama-3.2"and"meta-llama/llama-3.3"but not"meta-llama/llama-3.1"— the family of the gate's pinned default OpenRouter model,"meta-llama/llama-3.1-8b-instruct"(chosen, per an existing code comment, specifically because it supports every capability the matrix exercises). Same failure mode as the ollama defect above. Fixed insrc/lib/providers/openRouter/client.tsby adding"meta-llama/llama-3.1"toknownToolCapablePatterns.litellmcell 8 (budget ceiling) — test sizing, not asrcdefect — a singlenl.generate()call againstlitellmmakes 3 wire requests against this mock, not 1:ensureModelLimits()performs its ownGET /model/infodiscovery request before the real chat-completions call, and NeuroLink constructs two separateLiteLLMProviderinstances per logical call, so that discovery request fires twice. Against a real, slower backend the second firing is normally absorbed by the module-level in-flight-promise dedup cache insrc/lib/providers/litellm/client.ts; against this fast local mock the first discovery call can complete (and 404, since the mock has no/model/infohandler) before the secondLiteLLMProviderconstruction reuses the cache entry, so the dedup misses and both land on the wire. This is a test-infrastructure timing artifact specific to a fast local mock, not a genuine product defect — the underlying two-instances-per-call construction pattern is sanctioned architecture and is normally invisible against real network latency. Fixed in the test suite only: thelitellmrow'srequestsPerCall(seeGateRow.requestsPerCallintest/continuous-test-suite-acceptance-gate.ts) is3, so cell 8 sizes its dedicated ceiling server correctly and doesn't trip mid-way through the row's own first call.
Mock server
test/helpers/acceptanceGateServer.ts is a real node:http server (not
fetch interception) that speaks the two wire protocols the covered providers
actually use:
- OpenAI-compatible chat completions —
/chat/completions(including SSE streaming, tool calls,response_format, reasoning deltas),/models,/embeddings. - Anthropic Messages API —
/messages, for the oneANTHROPIC_BASE_URLrow.
Every provider is redirected at it through its own documented base-URL
environment override (<ID>_BASE_URL for catalog providers, each
hand-written provider's own env var — see "Coverage" below) with a fake API
key (gate-fake-key-not-a-real-credential). The server enforces a
request-count ceiling per instance and returns HTTP 400 past it (cell 8).
Every gate marker, exact value, and mock behavior the suite asserts against
is exported from that file (GATE_MARKERS, GATE_EXACT_VALUE,
GATE_STREAM_VALUE, GATE_STRUCTURED_VALUE, GATE_TOOL_NAME,
GATE_TOOL_CONFIRM_PREFIX, GATE_THINKING_FULL, GATE_THINKING_ANSWER,
GATE_EMBEDDING_VECTOR, GATE_SERVER_MODEL_SUFFIX, ...) so the suite and
the mock can never silently drift apart.
Coverage
Provider rows are derived the same way test/helpers/providerMatrix.ts
already does: hand-written rows are listed explicitly, catalog rows are
derived from the built catalog (dist/providers/catalog/index.generated.js
and dist/providers/catalog/loader.js), matching every other suite that
consumes providerMatrix.ts.
31 providers covered — 9 hand-written OpenAI-compatible providers
(openai, openai-compatible, openrouter, ollama, litellm,
nvidia-nim, lm-studio, llamacpp, cohere), azure (OpenAI-compatible
protocol via its own deployment URL), anthropic (Anthropic Messages
protocol), and 20 catalog providers whose base URL is a plain env override
(every catalog provider except cloudflare, which needs a
computedBaseURL/account-id template the mock can't satisfy — see below).
13 providers not covered, each for a structural reason (no silent caps —
the suite itself throws at import time if a PROVIDERS entry is in neither
list):
| Provider | Reason |
|---|---|
vertex | Signed Google Cloud client (service account / ADC), not a plain bearer-token HTTP call; no base-URL override, credentials can't be faked with a static string. |
google-ai | Speaks Google's Gemini wire protocol — a third protocol family this gate's mock does not implement (only OpenAI-compatible and Anthropic Messages are). |
bedrock | AWS SigV4-signed requests via the AWS SDK — same class of problem as vertex. |
sagemaker | AWS SigV4-signed requests via the AWS SDK — same class of problem as vertex. |
cloudflare | Catalog entry declares a baseURLTemplate with a hardcoded api.cloudflare.com host; only the {accountId} path segment is overridable via CLOUDFLARE_ACCOUNT_ID, not the host. |
voyage | Embeddings-only (text: false) — no chat surface for cells 1-3, and its embeddings wire shape is a custom protocol not implemented here. |
jina | Embeddings/reranking-only (text: false) — same reasoning as voyage. |
stability | Image-generation-only (text: false) — no chat surface, image generation out of scope. |
ideogram | Image-generation-only (text: false) — same as stability. |
recraft | Image-generation-only (text: false) — same as stability. |
replicate | Predictions API (poll-based, streaming: false) — a custom protocol distinct from both implemented protocol families. |
typesafe | Decide-only (SystemOneDecisionProvider) — every generation capability is false; no text surface for any cell. |
laya | Decide-only, like typesafe — no text surface for any cell. |
Cells 4 (tools), 5 (structured output), 6 (thinking) and 7 (embeddings) run
only where the covered provider's row/catalog descriptor declares that
capability; otherwise the cell reports SKIP with a reason, never a
silent pass or a fabricated failure.
Running it
pnpm run build # dist/ must be fresh — the suite calls
# assertDistFresh({ entrypoints: ["dist/cli/index.js"] })
# and refuses to run against a stale build, since
# cell 9 drives the built CLI directly.
pnpm run test:acceptance-gate
No environment variables or credentials are required, and none are used:
real ones are stripped at startup (see "Known, documented gap in cell 9"),
and every credential the suite sets is a fake string, scoped per-row via
withEnv() and restored after each row.