Skip to main content

Claude Proxy Configuration Reference

This document is the authoritative reference for every configurable aspect of the NeuroLink Claude proxy. It covers CLI flags, the YAML config file schema, environment variables, auto-configured Claude Code settings, and all file locations.


1. CLI Flags​

Start the Claude multi-account proxy server.

FlagAliasTypeDefaultDescription
--port-pnumber55669Port to listen on.
--share-portnumber--port + 1Gate-only listener port for peer sharing (see below).
--host-Hstring127.0.0.1Host/IP to bind to. Use 0.0.0.0 to listen on all interfaces.
--strategy-sstringfill-firstAccount selection strategy. Choices: fill-first, round-robin.
--health-intervalnumber30Health check interval in seconds.
--quiet-qbooleanfalseSuppress non-essential output (banner, status messages).
--debug-dbooleanfalseEnable debug output (stack traces on errors, verbose logging).
--config-cstring~/.neurolink/proxy-config.yamlPath to proxy config file (YAML or JSON).
--env-filestringPath to .env file for provider API keys (overrides cwd .env).
--passthroughbooleanfalseTransparent forwarding: no retry, rotation, or polyfill.

Examples:

# Start with defaults (port 55669, fill-first strategy)
neurolink proxy start

# Custom port and explicit round-robin strategy
neurolink proxy start -p 8080 -s round-robin

# Start with 60-second health checks, debug output
neurolink proxy start --health-interval 60 --debug

# Use a custom config file
neurolink proxy start --config /path/to/my-proxy.yaml

The share listener​

proxy start runs a second, gate-only listener whenever this node has at least one active share grant. It serves the same routes on a different port and refuses every request that carries no valid share token; the main port keeps serving the operator's own untokened client exactly as before.

Which port a request arrived on is decided by the accepting socket, so nothing a client sends can move it across. That is the reason for a second port rather than an origin check: cloudflared and every reverse proxy connect from 127.0.0.1, so tunnelled traffic is indistinguishable from local traffic by address alone.

BehaviourDetail
Port--share-port, else NEUROLINK_PROXY_SHARE_PORT, else main port + 1
LifecycleComes up on the first active grant, closes when the last is revoked — no restart on either edge
Poll interval15s against the grant file
Bind failureLogged once and retried; never fatal. Set --share-port to move it
Rolling replacementThe incoming worker loses the bind until the outgoing one drains, then takes it on the next poll
DisableNEUROLINK_PROXY_SHARE_LISTENER=0

neurolink proxy expose picks this port automatically. Expose it, not the main one.

Show the current proxy status.

FlagAliasTypeDefaultDescription
--formatstringtextOutput format. Choices: text, json.
--quiet-qbooleanfalseSuppress non-essential output.

Examples:

# Human-readable status
neurolink proxy status

# Machine-readable JSON (for scripts)
neurolink proxy status --format json

JSON output shape (when --format json):

{
"running": true,
"pid": 12345,
"port": 55669,
"host": "127.0.0.1",
"strategy": "fill-first",
"startTime": "2025-03-22T10:00:00.000Z",
"uptime": 3600000,
"url": "http://127.0.0.1:55669",
"autoUpdateEnabled": true,
"updaterPid": 12346,
"updaterRunning": true,
"latestVersion": "9.88.9",
"pendingRestartVersion": null,
"lastUpdateFailure": null,
"fallbackChain": [{ "provider": "google-ai", "model": "gemini-2.5-pro" }],
"stats": {
"totalAttempts": 42,
"totalAttemptErrors": 5,
"totalRequests": 31,
"totalSuccess": 29,
"totalErrors": 2,
"totalRateLimits": 3,
"totalTransientRateLimits": 2,
"totalQuotaRateLimits": 1
}
}

totalRequests, totalSuccess, and totalErrors are final request outcomes. totalAttempts, totalAttemptErrors, and the rate-limit counters describe upstream attempts, including retries that later recovered. Per-account requests, success, and errors use the same final-outcome semantics; attemptErrors and the rate-limit fields remain attempt-level diagnostics.

Manage the repo-owned local OpenObserve stack and the maintained proxy dashboard.

ActionDescription
setupStart OpenObserve + OTEL collector and import the maintained dashboard
startStart the local telemetry stack without re-importing the dashboard
stopStop the local telemetry stack
statusShow local stack health and endpoint info
logsFollow OpenObserve and collector logs
import-dashboardRe-import the dashboard and dedupe older dashboards with the same title
FlagAliasTypeDefaultDescription
--quiet-qbooleanfalseSuppress the local CLI spinner before delegating.

Examples:

neurolink proxy telemetry setup
neurolink proxy telemetry status
neurolink proxy telemetry logs

One-command setup: login + install proxy service + configure Claude Code.

FlagAliasTypeDefaultDescription
--port-pnumber55669Proxy port.
--methodstringoauthAuthentication method. Choices: oauth, api-key.
--no-servicebooleanfalseSkip launchd install, just start foreground.
--env-filestringPath to a proxy provider env file to persist.

Examples:

# Full setup with defaults (OAuth login, port 55669, launchd service)
neurolink proxy setup

# Setup on a custom port
neurolink proxy setup -p 9000

# Login + start foreground (no auto-restart service)
neurolink proxy setup --no-service

What proxy setup does:

  1. Checks for existing authenticated accounts in the TokenStore.
  2. Falls back to the legacy ~/.neurolink/anthropic-credentials.json file.
  3. If no valid accounts are found, runs the OAuth login flow.
  4. Installs as macOS launchd service (auto-restart on crash/reboot) and configures Claude Code. Use --no-service for foreground start.

Internal fail-open guard process. Spawned only by a foreground proxy start; launchd-managed proxies leave restart ownership entirely to launchd. The guard reverts stale Claude Code settings only after its parent is confirmed dead and never restarts or signals a live proxy. A launchd installation uses a separate updater-only worker which never changes client settings.

FlagTypeDefaultDescription
--hoststring127.0.0.1Proxy host to monitor.
--portnumber55669Proxy port to monitor.
--parent-pidnumber(required)PID of the parent proxy process.
--max-wait-msnumber0Maximum monitoring duration (0 = indefinite).
--failure-thresholdnumber5Consecutive health check failures before triggering cleanup.
--poll-interval-msnumber1000Interval between health checks in milliseconds.
--quietbooleantrueSuppress output (guards are silent by default).

You should never need to run this command manually.

Install the proxy as a persistent macOS launchd service. The service auto-starts on login and auto-restarts on crash (5-second throttle). Currently macOS-only.

FlagAliasTypeDefaultDescription
--port-pnumber55669Proxy port.
--hoststring127.0.0.1Proxy host/IP to bind to.
--env-filestringPath to provider env file to persist for the service.
--configstringPath to proxy routing config file to persist for the service.

Examples:

# Install with defaults (port 55669)
neurolink proxy install

# Install on custom port
neurolink proxy install -p 9000

What it does:

  1. Writes a launchd plist to ~/Library/LaunchAgents/com.neurolink.proxy.plist.
  2. Loads the service via launchctl load.
  3. The service runs neurolink proxy start --port <port> --host <host> --quiet and persists any --env-file / --config values into the managed service definition.
  4. Logs go to ~/.neurolink/logs/proxy-launchd-stdout.log and proxy-launchd-stderr.log.

Management:

# Start/stop manually
launchctl start com.neurolink.proxy
launchctl stop com.neurolink.proxy

# Remove entirely
neurolink proxy uninstall

Remove the proxy launchd background service. Unloads the service and deletes the plist file. Currently macOS-only.

No flags.

Examples:

neurolink proxy uninstall

Remove expired and disabled accounts from the token store.

FlagTypeDefaultDescription
--forcebooleanfalseSkip confirmation when removing disabled accounts.

Examples:

# Interactive cleanup (prompts before removing disabled accounts)
neurolink auth cleanup

# Force cleanup without confirmation
neurolink auth cleanup --force

What it does:

  1. Prunes expired entries that have no refresh token.
  2. Finds permanently disabled entries (e.g., accounts that failed refresh).
  3. Prompts for confirmation before removing disabled accounts (unless --force).

Re-enable a previously disabled account so it can be used by the proxy pool again.

ArgumentTypeRequiredDescription
<account>stringYesAccount key to re-enable (e.g., anthropic:1-VjRIq).

Examples:

# Re-enable a disabled account
neurolink auth enable anthropic:1-VjRIq

Run neurolink auth list to see all accounts and their current status.

Designate the proxy's primary (home) Anthropic account by email/label. Writes routing.primary-account to the proxy config YAML; a running proxy watching that exact file applies it automatically. With quota routing enabled (the default), primary is the ranking's final tiebreaker, not tried first, unless routing.prefer-primary is also set; it is tried first unconditionally only when quota routing is disabled, and under round-robin it is used as the home reference. Does not touch the encrypted token store and does not require re-OAuthing any account.

ArgumentTypeRequiredDescription
<email>stringYesEmail/label of the Anthropic account to make primary.
--configstringNoPath to the proxy config file. Default: ~/.neurolink/proxy-config.yaml.

If the email is not currently authenticated in the token store, the command still writes the field and prints a warning — the setting activates automatically once the account is added via auth login --add. For a running proxy, the command reports whether it watches the edited path, watches a different path, or predates hot-reload support.

Examples:

# Make [email protected] primary in the default config
neurolink auth set-primary [email protected]

# Use a non-default config path
neurolink auth set-primary [email protected] --config ./proxy.yaml

Note: writing YAML uses js-yaml.dump, which does not preserve comments. The command prints a warning before writing if the existing file contains comments. JSON config paths preserve everything except whitespace.

Show the proxy's currently configured primary account (and whether it is authenticated).

ArgumentTypeRequiredDescription
--configstringNoPath to the proxy config file. Default: ~/.neurolink/proxy-config.yaml.

Examples:

neurolink auth get-primary

Output (when configured and authenticated):

Configured primary: [email protected]
Status: authenticated (anthropic:[email protected] present in token store)
Source: /Users/.../.neurolink/proxy-config.yaml

Remove routing.primary-account (and routing.primaryAccount) from the proxy config. A running proxy watching that file reverts to insertion-order fallback on its next valid configuration generation.

ArgumentTypeRequiredDescription
--configstringNoPath to the proxy config file. Default: ~/.neurolink/proxy-config.yaml.

Examples:

neurolink auth clear-primary

Idempotent — clearing when no primary is configured prints No primary account was configured. and exits 0.

Lender-side controls for peer sharing. Conceptual documentation lives in Proxy peer sharing; this is the flag reference. Actions: create, provision, url, list, status, pause, resume, revoke, topup, set, link, rotate, level, note, notes, receipts, delete.

ArgumentTypeDescription
--peerstringBorrower label, or a grant id. Required by everything but list, status, url, note and notes — the coin-note actions are issued against the node, not a peer.
--presetstringspare (default), spillover, metered, open. Fills the gate set; every field stays overridable by an explicit flag.
--levelstringlive (default) or complete.
--ledgerstringcoins or unlimited. Implied coins when --coins is given.
--coinsnumberStarting balance for a metered grant. With topup, the amount to add; with set, the absolute balance.
--refillstringStanding allowance, e.g. 100/week or 50/session. Applied at the first borrowed request after the period elapses, not on a timer.
--max-slicestringCeiling as a percent of the pool: 20, or 5h=20,7d=15. Consumption is summed across the grant's reachable accounts and divided by their count.
--max-slice-per-accountstringThe same ceiling applied to each account independently. Opt-in; the pre-pool behaviour.
--reservestringHeadroom floor the borrower may never eat into: 30, or 5h=30,7d=20. Per-account by design.
--spilloverstringUse-it-or-lose-it window: 12h\<60 or 12h\<60@25 (hours before reset, utilization below, optional slice cap).
--modelsarrayTier allowlist, matched as case-insensitive substrings — sonnet,haiku covers every dated id in those tiers.
--accountsarrayWhich of your accounts this grant may draw on, by full key or bare label. Also the denominator of --max-slice.
--ratestringRequests per minute: 20/min or 20.
--concurrencynumberSimultaneous in-flight borrowed requests.
--schedulestringHour-of-day window in local time: 21-9 wraps midnight.
--expiresstringGrant expiry: 7d, 48h, 90m. A bare number means days.
--from-accountstringWhich account a complete share is minted from. Without it the drift audit has no baseline to reconcile against.
--codestringAuthorization code from your browser, to finish share provision.
--offline-gracestringHow long a complete-share borrower may run unheard-from. Default 24h.
--heartbeatstringComplete-share check-in interval. Default 15m.
--lease-ttlstringLease lifetime. Default 7d. A lease can never outlive the grant's own --expires.
--public-urlstringAddress share links are minted against. Recorded once by share url.
--clearbooleanWith share url: forget the recorded address.
--tostringWith share level: live or complete.
--notestringFree-text note kept with the grant.
--ttlstringWith share note: how long a coin note stays redeemable. Default 30d.
--memostringWith share note: text carried on the coin note itself.
--jsonbooleanEmit JSON. share create --json is the only way to capture a token programmatically — it is never stored.
--devbooleanUse the isolated dev-mode state directory.

share url also takes a positional: share url <address> records one, share url get prints the bare value (exit 1 when unset), share url clear forgets it, and a bare share url reports the current value.

Borrower-side controls. Actions: add, request, sync, receipts, net, redeem, list, status, test, remove, pause, resume, set.

ArgumentTypeDescription
--namestringLocal name for the lender. Required by everything but list and sync.
--linkstringShare link from the lender: neurolink://share/<host>#<token>. The token rides in the fragment, which is never transmitted to the host.
--urlstringLender's proxy address, when adding by hand instead of by link.
--tokenstringShare token, when adding by hand.
--prioritynumberLower is tried first. Default 100.
--claimbooleanWith peer request: collect a code the lender has authorized, and exchange it locally.
--labelstringLocal account label for a provisioned credential. Default <peer>-shared.
--notestringFree-text note kept with the peer.
--receipt-secretstringSecret this lender signs receipts with, when adding a peer by hand instead of from a link.
--reciprocalstringWith peer net: label of the grant you issued to the same person. Defaults to the peer's own name.
--coin-notestringWith peer redeem: the coin note to present.
--checkbooleanWith peer redeem: ask the issuer about the note without spending it.
--jsonbooleanEmit JSON instead of formatted text.
--devbooleanUse the isolated dev-mode state directory.

Publish this proxy over a Cloudflare tunnel, refusing to do so while the gate is off.

ArgumentTypeDescription
--portnumberLocal port to expose. No default — omit it and the gate-only share listener is picked, falling back to the running proxy's port.
--hoststringLocal host to expose. Default 127.0.0.1.
--namedstringNamed tunnel to run instead of a quick tunnel. Quick-tunnel URLs change on restart and rot every peer entry.
--forcebooleanPublish anyway when the gate probe says the proxy answers untokened requests. Publishes your subscription to anyone who finds the URL.

2. Config File (~/.neurolink/proxy-config.yaml)​

The proxy loads its configuration from a YAML (or JSON) file. The default location is ~/.neurolink/proxy-config.yaml. Override it with --config.

Runtime Reload Semantics​

The running proxy watches the resolved config path and proxy env path. Changes are debounced, parsed, validated, and converted into a complete immutable routing snapshot before one pointer swap publishes the next generation. Each request captures one generation, so a reload never changes model routing, fallback order, or account eligibility midway through that request.

The following values reload without restarting:

  • routing.strategy, unless --strategy supplied a fixed CLI override
  • model mappings, fallback chain, auto fallback, per-account admission, and passthrough models
  • primary account and account allowlist
  • quota routing, session soft limit, and reset tolerance
  • env interpolation used by those routing fields
  • NEUROLINK_PROXY_QUOTA_ROUTING, NEUROLINK_PROXY_SESSION_SOFT_LIMIT, and NEUROLINK_PROXY_SESSION_RESET_TOLERANCE_MS from the proxy env file

Malformed, invalid, or deleted previously observed files do not partially apply. The last-known-good generation remains active, and /status exposes config.generation, config.lastReloadError, and config.consecutiveFailures. neurolink proxy status presents the same information. SIGHUP requests an immediate reload; normal edits need no signal.

Listener address, port, passthrough mode, keep-alive dispatcher settings, telemetry initialization, and executable code remain startup concerns. Those require process replacement; editing their env values is intentionally not presented as a successful hot reload.

YAML parsing uses js-yaml when available; otherwise falls back to JSON.parse.

Environment Variable Interpolation​

All string values support ${VAR_NAME} and ${VAR_NAME:-default} syntax for environment variable resolution:

accounts:
anthropic:
- name: production
apiKey: "${ANTHROPIC_API_KEY}" # resolved from env
- name: backup
apiKey: "${BACKUP_KEY:-sk-fallback-123}" # with default value

Resolution order:

  1. Look up VAR_NAME in process.env.
  2. If not found, use the :-default value when present.
  3. If no default, the literal ${VAR_NAME} token is preserved (validation will catch missing keys).

Full Schema​

# ---------------------------------------------------------------------------
# Top-level fields
# ---------------------------------------------------------------------------

# Schema version (optional, default: 1)
version: 1

# Default provider applied when not specified per-account (optional)
defaultProvider: "anthropic"

# Default base URL applied to accounts that omit baseUrl (optional)
defaultBaseUrl: "https://api.anthropic.com"

# ---------------------------------------------------------------------------
# accounts (REQUIRED)
# ---------------------------------------------------------------------------
# Map of provider names to arrays of account configurations.
# At least one provider with at least one account is required.
accounts:
anthropic:
- name: "personal-pro" # Human-readable label (default: "unnamed")
apiKey: "${ANTHROPIC_KEY_1}" # API key or OAuth token (REQUIRED, non-empty)
baseUrl: "https://api.anthropic.com" # Base URL override (optional)
orgId: "org-abc123" # Organization ID (optional)
weight: 2 # Weight for weighted round-robin (default: 1)
enabled: true # Whether this account is active (default: true)
rateLimit: 60 # Max requests per minute (optional)
metadata: # Arbitrary metadata (optional)
tier: "pro"
notes: "Main account"

- name: "team-max"
apiKey: "${ANTHROPIC_KEY_2}"
weight: 3
enabled: true

# ---------------------------------------------------------------------------
# routing (optional)
# ---------------------------------------------------------------------------
# Controls model mapping, fallback chains, and routing strategy.
# Accepts both camelCase and kebab-case keys for YAML-friendliness.
routing:
# Account selection strategy: "round-robin" | "fill-first"
strategy: "fill-first"

# Quota-aware ordering controls for fill-first. Accounts with session
# headroom are ordered by soonest weekly expiry; a session at the soft limit
# is temporarily demoted until its 5h window resets. Environment variables
# with matching names take precedence when present in the proxy env file.
quota-routing: true
session-soft-limit: 0.97
session-reset-tolerance-ms: 900000

# Optional bound for concurrent upstream requests per OAuth account. Omit
# this key for unlimited admission. When set, requests try another eligible
# account first, then queue only if all are full. Valid range is 1 through 20.
# A value outside that range — a non-integer, 0, or 21+ — is rejected with a
# warning and leaves admission unlimited; it is NOT clamped to the nearest
# bound, so a typo here silently removes the cap rather than tightening it.
# max-inflight-per-account: 2

# Primary (home) account: under fill-first without quota routing this account
# is tried first. With quota routing enabled it is only the final tie-break
# in the ranking order below, unless prefer-primary (below) is set, in which
# case it is tried first whenever usable and not session-saturated. Under
# round-robin it sets the starting offset when account membership changes.
# Resolved per-request to a stable token-store key (anthropic:<email>); a
# numeric index is never persisted, so reordering accounts in the token
# store is irrelevant. When omitted the proxy falls back to insertion-order
# index 0. Accepts: primary-account (kebab) or primaryAccount (camel).
# Manage via:
# neurolink auth set-primary <email>
# neurolink auth get-primary
# neurolink auth clear-primary
primary-account: "[email protected]"

# Account routing policies (all optional; every default reproduces today's
# exact behavior). See "Account Routing Policies" below for precedence,
# interaction notes, and hot-reload semantics. Ignored under round-robin.
# Accepts kebab-case or camelCase for every key here.
#
# Ordering of usable accounts under quota routing: "expiry-first" (default,
# unchanged) or "headroom-first" (ranks by highest min(1 - sessionUsed,
# 1 - weeklyUsed) ahead of weekly reset).
account-ranking: "expiry-first"
# If primary-account is usable and not session-saturated, try it before the
# ranking order (default: false).
prefer-primary: false
# Keep a Claude Code session on the Anthropic account that served it, while
# that account stays usable and not session-saturated (default: false).
# Turning this off or on at runtime clears all existing bindings.
session-affinity: false
# How long an idle session binding survives, in ms (default: 3600000;
# valid range 60000-86400000). Only meaningful when session-affinity is on.
session-affinity-idle-ttl-ms: 3600000
# For an unbound request: if the account that would otherwise be tried
# first already has this many requests in flight, try the next account
# below that count that is usable and not session-saturated instead
# (default: 0, meaning off; valid range 0-100). Only has an effect below
# max-inflight-per-account's cap, or when no cap is set.
spill-inflight: 0

# Optional hard boundary for Anthropic credential discovery. Entries may be
# labels/emails or full anthropic:<label> keys. An empty list denies all.
# Hidden legacy/env credentials require explicit legacy-default/env entries.
# Accepts: account-allowlist (kebab) or accountAllowlist (camel).
account-allowlist:
- "[email protected]"

# Model mappings: remap incoming model names to different provider/model pairs
# Accepts: model-mappings (kebab) or modelMappings (camel)
model-mappings:
- from: "claude-sonnet-4-20250514" # Model name sent by Claude Code
to: "gemini-2.5-pro" # Target model name
provider: "google-ai" # Target provider (default: "anthropic")

- from: "claude-3-haiku-20240307"
to: "gpt-4o-mini"
provider: "openai"

# Fallback chain: when all Claude accounts are exhausted, try these in order
# Accepts: fallback-chain (kebab) or fallbackChain (camel)
fallback-chain:
- provider: "google-ai"
model: "gemini-2.5-pro"
- provider: "openai"
model: "gpt-4o"

# Disabled by default. Set true only when it is acceptable for a request to
# use an unspecified provider selected by the translation layer after the
# configured fallback chain is exhausted. It is skipped when any account
# returned a rate limit; that case still returns 429.
auto-fallback: false

# Passthrough models: model IDs that skip routing and go directly to Anthropic
# Accepts: passthrough-models (kebab) or passthroughModels (camel)
passthrough-models:
- "claude-sonnet-4-20250514"
- "claude-3-5-sonnet-20241022"
- "claude-3-haiku-20240307"

# ---------------------------------------------------------------------------
# cloaking (optional)
# ---------------------------------------------------------------------------
# Cloaking pipeline for making proxy requests indistinguishable from
# genuine Claude Code sessions.
cloaking:
# Mode: "auto" | "always" | "never"
# auto - apply cloaking only to OAuth accounts (default behavior)
# always - apply to all accounts (OAuth and API key)
# never - disable all cloaking plugins
mode: "auto"

plugins:
# Strip proxy-revealing headers (x-forwarded-for, via, etc.)
headerScrubber: true

# Generate consistent session identities per account (1-hour TTL)
sessionIdentity: true

# Inject Claude Code session context into system prompt (OAuth only)
systemPromptInjector: true

# Zero-width character insertion into sensitive words
wordObfuscator:
enabled: true
words: # Custom words to obfuscate
- "proxy"
- "neurolink"
- "load balancer"
- "round-robin"
- "failover"
- "multi-account"

# TLS fingerprint mimicry (stub/placeholder -- not yet implemented)
tlsFingerprint:
enabled: false

Field Reference Table​

Top-Level Fields​

FieldTypeDefaultRequiredDescription
versionnumber1NoConfig schema version.
defaultProviderstring(none)NoDefault provider name applied to accounts that omit it.
defaultBaseUrlstring(none)NoDefault base URL applied to accounts that omit baseUrl.
accountsRecord<string, Account[]>(none)YesMap of provider names to account arrays.
routingRoutingConfig(none)NoRouting strategy, model mappings, and fallback chain.
cloakingCloakingConfig(none)NoCloaking pipeline configuration.

Account Fields​

FieldTypeDefaultRequiredDescription
namestring"unnamed"NoHuman-readable account label.
apiKeystring(none)YesAPI key or OAuth token. Supports ${ENV_VAR} interpolation.
baseUrlstring(none)NoOverride the provider's API base URL.
orgIdstring(none)NoOrganization ID (e.g., OpenAI organizations).
weightnumber1NoWeight for weighted round-robin selection. Higher weight = more traffic.
enabledbooleantrueNoWhether this account is active. Disabled accounts are skipped.
rateLimitnumber(none)NoMaximum requests per minute for this account.
metadataRecord<string, unknown>(none)NoArbitrary metadata (tier info, notes, tags).

Routing Fields​

FieldTypeDefaultRequiredDescription
strategy"round-robin" | "fill-first"(none)NoAccount selection strategy. round-robin rotates across accounts. fill-first uses one account until exhausted.
primary-account / primaryAccountstring(none)NoEmail/label of the Anthropic account to treat as primary (home). With quota routing enabled, primary is only the final tie-break in the ranking order, unless prefer-primary is set, in which case it is tried first whenever it's usable and not session-saturated. It is tried first unconditionally only when quota routing is disabled. Resolved per-request to anthropic:<email>; falls back to insertion-order index 0 when absent or when the configured account isn't currently authenticated. Manage via neurolink auth set-primary <email>.
account-ranking / accountRanking"expiry-first" | "headroom-first"expiry-firstNoOrdering of usable accounts under quota routing. expiry-first is the existing comparator (soonest weekly reset, then other tie-breakers), unchanged. headroom-first ranks by highest headroom — min(1 - sessionUsed, 1 - weeklyUsed) — ahead of weekly reset. Ignored under strategy: round-robin.
prefer-primary / preferPrimarybooleanfalseNoIf the configured primary-account is usable and not session-saturated, it is tried before the ranking order (after any active session-affinity binding). Ignored under strategy: round-robin.
session-affinity / sessionAffinitybooleanfalseNoSticky sessions: a Claude Code session is bound to the Anthropic account that serves it and stays there while that account is usable and not session-saturated; a request that never tried the bound account (e.g. past max-inflight-per-account) keeps the binding. Otherwise, even if the bound account failed it, the session re-binds to the account that serves it; Codex/Vertex fallback or a peer borrow never re-binds. Ignored under strategy: round-robin and with one enabled account. Turning it off or on at runtime clears all bindings.
session-affinity-idle-ttl-ms / sessionAffinityIdleTtlMsinteger, 60000–864000003600000NoHow long a session binding survives without a served request before it's dropped (the prompt cache is assumed cold by then). Only meaningful when session-affinity is on.
spill-inflight / spillInflightinteger, 0–1000 (off)NoFor a request with no active session-affinity binding: if the account that would otherwise be tried first already has this many requests in flight, try the next account below that count that is usable and not session-saturated instead. Only affects unbound requests; never splits a bound session. Effective only below max-inflight-per-account's cap (or when no cap is set) — that cap applies to every account regardless of spill-inflight. Ignored under strategy: round-robin.
account-allowlist / accountAllowliststring[](none)NoAllowed Anthropic account labels/keys. When present, unlisted TokenStore, legacy, and environment credentials are excluded before loading or refresh. Empty denies all; absent is unrestricted. Special fallback labels are legacy-default and env.
model-mappings / modelMappingsModelMapping[][]NoArray of model-to-model remapping rules.
fallback-chain / fallbackChainFallbackEntry[][]NoOrdered list of alternative providers to try when primary accounts are exhausted.
auto-fallback / autoFallbackbooleanfalseNoAllows a translation-layer-selected fallback provider after configured fallbacks fail. Keep disabled to restrict requests to explicit accounts and fallback entries.
max-inflight-per-account / maxInflightPerAccountinteger(unlimited)NoOptional maximum concurrent upstream requests per OAuth account, from 1 through 20. When omitted, every request is admitted immediately. When set, requests prefer an available ordered account before queueing. Any value outside the range — non-integer, 0, or 21+ — is discarded with a warning and admission stays unlimited; values are never clamped to the nearest bound.
passthrough-models / passthroughModelsstring[][]NoModel IDs that bypass routing and go directly to Anthropic.
quota-routing / quotaRoutingbooleantrueNoEnables weekly-expiry-first quota ordering for fill-first with multiple accounts. Accounts at the session soft limit are temporarily demoted until their 5h window resets. The environment override takes precedence.
session-soft-limit / sessionSoftLimitnumber0.97NoSession utilization in (0, 1] at which quota routing proactively demotes an account.
session-reset-tolerance-ms / sessionResetToleranceMsinteger900000NoPositive reset-time bucket width used when session reset time breaks a weekly-expiry tie and when saturated accounts are ordered by recovery time.
use-overage / useOverage"auto" | "always" | "never"autoNoWhether an account may keep serving on paid extra usage once its subscription window is spent. auto follows what Anthropic reports per account; never parks the account at the subscription limit so the pool never spends credits; always keeps serving whenever the provider permits it. Only never overrides the provider — nothing here enables extra usage Anthropic has disabled. Manage via neurolink auth overage.

For fallback-chain/fallbackChain, account-allowlist/accountAllowlist, quota-routing/quotaRouting, use-overage/useOverage, auto-fallback/autoFallback, max-inflight-per-account/maxInflightPerAccount, session-soft-limit/sessionSoftLimit, and session-reset-tolerance-ms/sessionResetToleranceMs: an explicit null under either spelling means "unset" — the field's default applies as if the key were omitted, but a warning is logged. An empty YAML value (use-overage:) is null. A non-null value under one spelling is still used even when the other spelling is null. The five newer routing policy keys (account-ranking through spill-inflight) differ on purpose: an explicit null there is rejected — see Validation Rules.

ModelMapping Fields​

FieldTypeDefaultRequiredDescription
fromstring""YesIncoming model name (what Claude Code requests).
tostring""YesTarget model name at the destination provider.
providerstring"anthropic"NoTarget provider to route to.

FallbackEntry Fields​

FieldTypeDefaultRequiredDescription
providerstring""YesProvider name (e.g., google-ai, openai).
modelstring""YesModel to use at that provider.

Cloaking Fields​

FieldTypeDefaultDescription
mode"auto" | "always" | "never""auto"auto applies cloaking only to OAuth accounts. always applies to all. never disables all plugins.
plugins.headerScrubberbooleanfalseStrip proxy-revealing headers (x-forwarded-for, via, sec-ch-*, etc.).
plugins.sessionIdentitybooleanfalseGenerate consistent user_id/session_id per account with 1-hour TTL.
plugins.systemPromptInjectorbooleanfalseInject Claude Code session context (IDE metadata, timestamps) into system prompt. OAuth accounts only.
plugins.wordObfuscator.enabledbooleanfalseInsert zero-width characters into sensitive words to defeat string matching.
plugins.wordObfuscator.wordsstring[]["proxy", "neurolink", ...]Words to obfuscate. Defaults include: proxy, neurolink, load balancer, round-robin, failover, multi-account.
plugins.tlsFingerprint.enabledbooleanfalseTLS fingerprint mimicry. Currently a stub/placeholder (no-op).

Validation Rules​

The config loader validates the following:

  • accounts must be present and be a non-array object.
  • Each provider key in accounts must map to an array.
  • Each account must have a non-empty string apiKey.
  • If version is present, it must be a number.
  • routing.account-allowlist must be an array of non-empty strings when present.
  • routing.quota-routing must be a boolean when present.
  • routing.auto-fallback must be a boolean when present.
  • routing.max-inflight-per-account must be an integer from 1 through 20 when present.
  • routing.session-soft-limit must be a number in (0, 1] when present.
  • routing.session-reset-tolerance-ms must be a positive integer when present.
  • routing.use-overage must be auto, always, or never (case-insensitive) when present.
  • routing.account-ranking must be expiry-first or headroom-first when present.
  • routing.prefer-primary must be a boolean when present.
  • routing.session-affinity must be a boolean when present.
  • routing.session-affinity-idle-ttl-ms must be an integer from 60000 through 86400000 when present.
  • routing.spill-inflight must be an integer from 0 through 100 when present.
  • For those five routing policy keys, an explicit null counts as present and is rejected, whichever spelling (kebab-case or camelCase) carries it.
  • Plaintext API keys (not using ${ENV_VAR} references) trigger a warning.

An absent default config is optional. An existing config that cannot be read or validated fails proxy startup; it is never ignored in favor of unrestricted routing. A value that fails any check above rejects the whole config, never just that key: at startup the proxy does not start, and on a hot reload the last-known-good generation stays active.

Account Routing Policies​

The five keys above (account-ranking, prefer-primary, session-affinity, session-affinity-idle-ttl-ms, spill-inflight) all default to today's exact behavior — expiry-first ranking, no affinity, no primary preference, no spill — and are validated and hot-reloaded through the same runtime config snapshot as every other routing key.

strategy: round-robin ignores all five keys. They apply only under fill-first with more than one enabled account.

NEUROLINK_PROXY_QUOTA_ROUTING=off disables quota-based ranking, but session-affinity, prefer-primary and spill-inflight still apply. Turning off quota routing removes the expiry/headroom ordering step, so account-ranking has no effect; it does not disable sticky sessions, the primary preference or spill, which then work on the configured account order.

Precedence. After unusable accounts are removed, the order a request tries is:

  1. the session's bound account, when session-affinity is on and that account is usable and not session-saturated;
  2. the configured primary-account, when prefer-primary is on and it is usable and not session-saturated;
  3. the remaining accounts in account-ranking order.

Spill only changes which account is tried first, and only for requests without an active session-affinity binding — it never re-orders or moves a request that is already bound to an account. Its target must be usable and not session-saturated, like every step above; when no account under the threshold qualifies, nothing spills.

Binding. A served request binds its session to the Anthropic account that served it. The existing binding is kept only when that request never tried the bound account and the bound account is still usable and not session-saturated — a request that overflowed the bound account's max-inflight-per-account cap onto another account keeps the session on its warm prompt cache. A bound account that was tried and failed in the same request loses the session to the account that served it, even if its cooldown has already lapsed by the time the response starts. A Codex or Vertex fallback, a peer borrow, or an error response never binds.

Failure mode. If ranking, affinity or spill throws, the request uses the pre-policy order (plain expiry-first under quota routing) and the routing decision records routing_policy_error. If the binding step after a served response throws, the session binds to the serving account and the response is unaffected. Each such error is logged once per call site and distinct error (its name and message; the 32 most recent are remembered), not once per worker, so a later, unrelated failure still reaches the log.

spill-inflight and max-inflight-per-account. max-inflight-per-account is a single cap applied to every account; spill-inflight reroutes an unbound request away from an account that has reached its own threshold. As a result, spill-inflight only has an observable effect when it is set to a value below the max-inflight-per-account cap, or when no cap is configured at all — a spill-inflight at or above the cap never triggers, because max-inflight-per-account already stops admission first.

Paid overage. An account that may spend paid extra usage (see use-overage) is never marked session-saturated, so session-affinity and prefer-primary keep a session on it while it spends overage, even when other accounts still have free subscription headroom. If that matters, set use-overage: never alongside session-affinity: true.

Hot reload. Disabling session-affinity clears all existing session bindings, and re-enabling it clears them again, so it always starts from an empty binding store — including a binding recorded by a request that was already in flight when affinity was turned off. Every other policy key takes effect on the next request under the new config generation.

Observability. GET /status reports the active policy (policy: the current account-ranking, prefer-primary, session-affinity, session-affinity-idle-ttl-ms, and spill-inflight values) and the number of currently bound sessions (boundSessions); bindings idle for longer than session-affinity-idle-ttl-ms are not counted.


3. Environment Variables​

VariablePurposeUsed By
ANTHROPIC_API_KEYAnthropic API key. Used as a fallback credential when no OAuth accounts are found.Proxy routes, Anthropic provider
ANTHROPIC_OAUTH_TOKENOAuth access token for Anthropic (alternative to stored tokens).Anthropic provider, providerConfig
CLAUDE_OAUTH_TOKENAlias for ANTHROPIC_OAUTH_TOKEN. Checked as a fallback.Anthropic provider, providerConfig
NEUROLINK_SKIP_MCPSet to "true" to skip MCP server initialization. Automatically set by proxy start (tools come from Claude Code, not local MCP servers).NeuroLink constructor
NEUROLINK_LOG_LEVELLog level for the NeuroLink logger. Values: error, warn, info, debug.Logger utility
OTEL_EXPORTER_OTLP_ENDPOINTOTLP HTTP endpoint for proxy telemetry export. Written automatically to ~/.neurolink/.env by neurolink proxy telemetry setup. Example: http://localhost:14318.Proxy OTEL init (initializeProxyOpenTelemetry)
NEUROLINK_ENV_FILEPath to a .env file the proxy should load at startup. Overrides the default ~/.neurolink/.env auto-load.proxyEnv.ts (resolveProxyEnvFile)
NEUROLINK_PROXY_AUTO_UPDATEAutomatic package updates for launchd installations. Enabled by default; set to 0, off, or false to disable. Updates use the package manager owning the running install and restart only after all requests and streams are idle.Dedicated launchd updater worker
NEUROLINK_PROXY_REQUIRE_GRANTGate the main port too, refusing any request without a valid share token — including the operator's own client. Rarely needed now that the share listener exists; keep it for binding 0.0.0.0 with nothing in front. Also redacts account identity on /status. Values: 1, true, on, yes. Startup-only.Peer-sharing gate (shareGate.ts)
NEUROLINK_PROXY_SHARE_PORTPort for the gate-only share listener. Default: main port + 1. --share-port wins over this.Share listener (shareListener.ts)
NEUROLINK_PROXY_SHARE_LISTENERSet to 0, off, false or no to suppress the share listener entirely, whatever grants exist.Share listener (shareListener.ts)
NEUROLINK_PROXY_QUOTA_ROUTINGQuota-aware fill-first ordering. Enabled by default; set to 0, off, or false to disable. Reloads from the proxy env file at runtime.Runtime routing configuration
NEUROLINK_PROXY_SESSION_SOFT_LIMITSession utilization threshold in (0, 1]; defaults to 0.97. Reloads from the proxy env file at runtime.Runtime routing configuration
NEUROLINK_PROXY_SESSION_RESET_TOLERANCE_MSPositive reset-time bucket width in milliseconds; defaults to 900000. Reloads from the proxy env file at runtime.Runtime routing configuration
NEUROLINK_PROXY_SESSION_SECRETOptional secret used to produce stable, non-reversible lifecycle session hashes. When unset, the proxy creates a random mode-0600 installation key in the log directory and reuses it across restarts. Changing this value intentionally starts a new correlation domain.Lifecycle metadata logger
NEUROLINK_PACKAGE_MANAGER_PATHOptional absolute path to the npm or pnpm executable used by the updater. The candidate is still rejected unless its writable global root owns the running NeuroLink installation.Dedicated launchd updater worker
NEUROLINK_PACKAGE_MANAGEROptional npm or pnpm type for NEUROLINK_PACKAGE_MANAGER_PATH. When omitted, the updater infers the type from the executable name.Dedicated launchd updater worker
NEUROLINK_PNPM_PATHLegacy pnpm-specific updater override. Prefer NEUROLINK_PACKAGE_MANAGER_PATH for new installations.Dedicated launchd updater worker

Proxy Env File Resolution Order​

When the proxy starts, it loads env vars from a .env file using this priority:

  1. --env-file <path> CLI flag — explicit path, required to exist.
  2. NEUROLINK_ENV_FILE=<path> environment variable — explicit path, required to exist.
  3. ~/.neurolink/.env — loaded automatically if the file exists (created by neurolink proxy telemetry setup).
  4. Nothing — proxy starts without extra env vars; telemetry remains disabled unless env vars are already set in the shell, and the proxy emits a startup log explaining how to enable it unless output is suppressed.

The --env-file flag is baked into the launchd plist by proxy install, so the service always loads from the same file across reboots. The three runtime routing variables above and routing interpolation are reread transactionally; other env settings remain startup-only.

Priority for Anthropic credentials (checked in order by the proxy routes):

  1. TokenStore compound keys -- anthropic:<label> entries in ~/.neurolink/tokens.json.
  2. Legacy credentials file -- ~/.neurolink/anthropic-credentials.json (only if no compound keys exist).
  3. ANTHROPIC_API_KEY env var -- Only if no Anthropic TokenStore entries or legacy credential are present.

routing.account-allowlist filters these sources before loading or refresh. Legacy and environment fallbacks are never activated merely because existing TokenStore accounts are disabled, cooling, or unavailable.


4. Claude Code Settings​

When the proxy starts, it automatically writes to ~/.claude/settings.json:

{
"env": {
"ANTHROPIC_BASE_URL": "http://127.0.0.1:55669",
"ENABLE_TOOL_SEARCH": "true"
}
}
KeyValueDescription
ANTHROPIC_BASE_URLhttp://<host>:<port>Tells Claude Code to route all Anthropic API requests through the proxy.
ENABLE_TOOL_SEARCH"true"Enables tool search in Claude Code (required for full proxy compatibility).

Lifecycle:

  • On proxy start -- Both keys are written (or merged into existing settings).
  • On proxy stop (Ctrl+C / SIGTERM) -- Both keys are removed. Other env keys in the settings file are preserved.
  • Fail-open guard -- A foreground proxy's detached guard removes stale settings only after confirming its parent died and no replacement is healthy. launchd-managed proxies do not spawn this cleanup guard; they use a separate updater-only worker.
  • Safety -- If the ANTHROPIC_BASE_URL has been changed to a different value (e.g., another proxy), the cleanup will not overwrite it.

After starting the proxy, restart Claude Code for the new settings to take effect.


5. File Locations​

All NeuroLink proxy files are stored under ~/.neurolink/ (with 0o700 directory permissions).

FilePermissionsDescription
~/.neurolink/tokens.json0o600TokenStore -- Multi-provider OAuth token storage. Stores tokens keyed by provider:label (e.g., anthropic:personal). XOR-obfuscated by default (not plaintext).
~/.neurolink/anthropic-credentials.json0o600Legacy credentials -- Single-account OAuth tokens. Used as a fallback when no compound keys exist in tokens.json. Updated on token refresh (pre-request or on-401).
~/.neurolink/proxy-config.yamluser defaultProxy config -- YAML/JSON configuration file. Loaded and watched by proxy start (default path, overridable with --config). Valid routing changes publish a new runtime generation.
~/.neurolink/.env0o600Proxy env file — Auto-loaded and watched by the proxy. Runtime routing variables and routing interpolation reload transactionally; startup-only variables do not. Created by neurolink proxy telemetry setup. Override with --env-file or NEUROLINK_ENV_FILE.
~/.neurolink/proxy-state.json0o600Proxy state -- Runtime state persisted by the running proxy (PID, listener, strategy, fallback chain, account allowlist, config path/generation/reload error, and supervisor PIDs). Used by proxy status, auth config commands, and the fail-open guard.
~/.neurolink/logs/proxy-YYYY-MM-DD.jsonl0o600Request summary logs -- One JSONL entry per completed proxied request. Includes requestId, method, path, model, status, account label, response time, token usage, and trace correlation fields.
~/.neurolink/logs/proxy-attempts-YYYY-MM-DD.jsonl0o600Attempt logs -- One JSONL entry per upstream attempt. Useful for retry, failover, and cooldown debugging without inflating request totals.
~/.neurolink/logs/proxy-debug-YYYY-MM-DD.jsonl0o600Debug index logs -- Redacted body-capture index rows with phase, headers, status, duration, and the stored body artifact path.
~/.neurolink/logs/bodies/YYYY-MM-DD/<request-id>/*.json.gz0o600Body artifacts -- Compressed redacted request and response bodies captured for debugging.
~/.neurolink/account-quotas.json0o600Account quotas -- Cached quota/utilization data from Anthropic's unified-5h and unified-7d rate-limit headers. Flushed to disk every 5 seconds.
~/.neurolink/account-cooldowns.json0o600Account cooldowns -- Extend-only rate-limit and transient-auth recovery timestamps. Persisted atomically so a proxy restart cannot immediately retry a known-exhausted account.
~/.claude/settings.jsonuser defaultClaude Code settings -- Auto-configured with ANTHROPIC_BASE_URL and ENABLE_TOOL_SEARCH when the proxy starts. Cleaned up on shutdown.

Peer-sharing state​

Written only once this node lends or borrows capacity — see Proxy peer sharing. All follow the same 0o600 + atomic-rename discipline as the files above. In --dev mode they resolve under <cwd>/.neurolink-dev/ instead.

FileOwnerDescription
~/.neurolink/proxy-grants.jsonlenderShare grants -- One record per borrower: hashed token (sha256(salt + secret); the token itself is never stored), level, state, entitlement and the full gate set. Also holds this node's recorded publicUrl. Re-read when its mtime moves, so share pause lands without a restart.
~/.neurolink/proxy-share-ledger.jsonlenderCoin ledger and window buckets -- Settled coin spend and request counts per grantId|accountKey, plus how much of each 5h/7d window a grant has taken, keyed by that window's reset timestamp so a rollover starts fresh. In-flight holds are memory-only.
~/.neurolink/proxy-share-audit.jsonlenderDrift audit -- Last utilization observation per complete-mode grant, the consecutive-drift streak and the auto-pause marker. Cleared by share resume; deleted with the grant.
~/.neurolink/proxy-share-provisioning.jsonlenderSplit-PKCE requests -- A borrower's outstanding code challenge, its 15-minute expiry, and (between authorization and the single claim that consumes it) the authorization code. Never a verifier, never a token.
~/.neurolink/proxy-peers.jsonborrowerPeers -- Lender name, URL, share token, priority, per-reason cooldown, last observed grant status, and any outstanding provisioning verifier.
~/.neurolink/proxy-resident-grants.jsonborrowerResident grants -- The signed lease governing each credential a lender provisioned here, its lease secret, last heartbeat and unreported spend.
~/.neurolink/proxy-share-receipts.jsonlenderReceipts -- The last 500 signed settlement statements per grant, each carrying the usage its charge was computed from, plus the cumulative coins forgiven by reciprocal netting.
~/.neurolink/proxy-share-notes.jsonissuerCoin notes -- Every transferable note this node minted, and which have been redeemed, by which grant and when. The spent-set is the replay protection.

TokenStore Details​

The tokens.json file uses this internal structure (after deobfuscation):

{
"version": "2.0",
"lastModified": 1711100000000,
"providers": {
"anthropic:personal": {
"tokens": {
"accessToken": "...",
"refreshToken": "...",
"expiresAt": 1711103600000,
"tokenType": "Bearer",
"scope": "..."
},
"createdAt": 1711100000000,
"lastAccessed": 1711100000000
},
"anthropic:team": {
"tokens": { "...": "..." },
"createdAt": 1711100000000,
"lastAccessed": 1711100000000
}
}
}

The TokenStore class options:

  • encryptionEnabled (default: true) -- XOR obfuscation with a machine-derived key.
  • customStoragePath -- Override the default ~/.neurolink/tokens.json path.

Tokens are automatically refreshed 1 hour before expiration when a TokenRefresher function is registered.


6. Model Mapping Examples​

Model mappings let you reroute specific model requests to different providers. The proxy's ModelRouter checks mappings in this order:

  1. Explicit mapping -- If the requested model has a from match in model-mappings, use the corresponding to/provider.
  2. Gemini prefix -- If the requested model starts with gemini-, route to Vertex by default.
  3. Passthrough list -- If the model is in passthrough-models, route to Anthropic.
  4. Claude prefix -- Any model starting with claude- is routed to Anthropic.
  5. Unknown model -- Returns provider: null (the proxy will reject non-Claude models unless routing is configured).

Example: Route Haiku to a Cheaper Provider​

routing:
model-mappings:
- from: "claude-3-haiku-20240307"
to: "gpt-4o-mini"
provider: "openai"

Claude Code requests claude-3-haiku-20240307 but the proxy sends the request to OpenAI's gpt-4o-mini instead, translating the request format via neurolink.generate().

Example: Use Gemini for All Sonnet Requests​

routing:
model-mappings:
- from: "claude-sonnet-4-20250514"
to: "gemini-2.5-pro"
provider: "google-ai"
- from: "claude-3-5-sonnet-20241022"
to: "gemini-2.5-flash"
provider: "google-ai"

Example: Passthrough Specific Models​

routing:
passthrough-models:
- "claude-sonnet-4-20250514"
- "claude-3-opus-20240229"
model-mappings:
- from: "claude-3-haiku-20240307"
to: "gemini-2.5-flash"
provider: "google-ai"

Here, Sonnet 4 and Opus requests go directly to Anthropic (passthrough), while Haiku requests are redirected to Gemini.

Example: No Routing (Pure Multi-Account Pool)​

Omit the routing section entirely. All requests pass through to Anthropic using the configured accounts with the proxy's default fill-first strategy:

accounts:
anthropic:
- name: "account-1"
apiKey: "${ANTHROPIC_KEY_1}"
- name: "account-2"
apiKey: "${ANTHROPIC_KEY_2}"
- name: "account-3"
apiKey: "${ANTHROPIC_KEY_3}"

7. Fallback Chain Examples​

The fallback chain is tried in order when all primary Claude accounts are exhausted (rate-limited, errored, or cooling down). Each entry specifies a provider and model. The proxy translates the Claude-format request into the target provider's format. Codex entries use the native pooled Codex Responses transport; other providers use neurolink.generate() or neurolink.stream().

Example: Codex with Extra High reasoning​

routing:
fallback-chain:
- provider: codex
model: gpt-6-astra
reasoning-effort: xhigh

Authenticate a Codex account with neurolink auth login codex before using this fallback. reasoning-effort (or reasoningEffort) is optional and supported only on codex fallback entries. It is sent as reasoning.effort for both streaming and non-streaming Claude requests. Omit it to use the upstream model's default.

Accepted values are none, minimal, low, medium, high, xhigh (Extra High), and max; availability depends on the selected Codex model. Invalid values or use on another provider fail configuration validation. Changes reload with the routing config, and an invalid reload keeps the previous working configuration.

Example: Gemini then OpenAI​

routing:
fallback-chain:
- provider: "google-ai"
model: "gemini-2.5-pro"
- provider: "openai"
model: "gpt-4o"

Request flow:

  1. Try Claude accounts with the configured strategy (fill-first by default) plus retry/failover.
  2. If all exhausted, try Google AI Studio with gemini-2.5-pro.
  3. If that also fails, try OpenAI with gpt-4o.

Example: Multiple Gemini Tiers​

routing:
fallback-chain:
- provider: "google-ai"
model: "gemini-2.5-pro"
- provider: "google-ai"
model: "gemini-2.5-flash"
- provider: "openai"
model: "gpt-4o-mini"

Falls back through progressively cheaper models.

Example: Vertex AI as Primary Fallback (Enterprise)​

routing:
fallback-chain:
- provider: "google-vertex"
model: "gemini-2.5-pro"
- provider: "amazon-bedrock"
model: "anthropic.claude-3-5-sonnet-20241022-v2:0"

Uses enterprise-grade providers (Vertex AI, Bedrock) as fallbacks. Requires the corresponding provider credentials to be configured in environment variables.

Example: Full Multi-Tier Setup​

version: 1

accounts:
anthropic:
- name: "pro-personal"
apiKey: "${CLAUDE_PRO_KEY}"
weight: 1
- name: "max-team"
apiKey: "${CLAUDE_MAX_KEY}"
weight: 3

routing:
strategy: "fill-first"

passthrough-models:
- "claude-sonnet-4-20250514"

model-mappings:
- from: "claude-3-haiku-20240307"
to: "gemini-2.5-flash"
provider: "google-ai"

fallback-chain:
- provider: "google-ai"
model: "gemini-2.5-pro"
- provider: "openai"
model: "gpt-4o"

cloaking:
mode: "auto"
plugins:
headerScrubber: true
sessionIdentity: true
systemPromptInjector: true
wordObfuscator:
enabled: true
words:
- "proxy"
- "neurolink"

This configuration:

  • Pools two Claude accounts with 1:3 weighting (Max gets 3x traffic).
  • Passes Sonnet 4 requests directly to Anthropic.
  • Redirects Haiku requests to Gemini Flash.
  • Falls back to Gemini Pro, then GPT-4o when Claude accounts are exhausted.
  • Applies cloaking to OAuth accounts (header scrubbing, session identity, system prompt injection, word obfuscation).

Proxy Endpoints​

For reference, the running proxy exposes these HTTP endpoints:

MethodPathDescription
POST/v1/messagesAnthropic-compatible chat completions (main endpoint).
GET/v1/modelsList available models.
POST/v1/messages/count_tokensToken counting endpoint.
GET/healthHealth check. Returns { status, strategy, uptime }.
GET/statusDetailed status with per-account stats, total attempts, completed requests, and error rates, plus the active routing policy (policy) and bound-session count (boundSessions; see Account Routing Policies). On a gated proxy, account identity is released only to a caller holding the update-control token.
GET/limitsFresh per-account limits from Anthropic's usage API. ?account=<label> for one, ?snapshot=true for stored state. Operator-only: refused for borrowed traffic.
GET/peer/handshakePeer protocol version, capabilities and grant state. Authenticated by share token; touches no account.
GET/peer/limitsWhat the calling grant may still do: remaining coins, slice left, whether anything can serve it. Carries no account identity.
POST/peer/provisionLodge a split-PKCE code challenge (complete shares).
GET/peer/provisionCollect the authorization code the lender produced. Single-use.
POST/peer/heartbeatComplete-share check-in: report spend, receive a refreshed lease or a stop. Authenticated by the grant's lease secret.
GET/peer/receiptsSigned statements of what this grant was charged. ?since=<sequence> for the ones not yet collected.
POST/peer/netSettle one round of reciprocal netting. The claim is signed with the grant's receipt secret.
POST/peer/noteCheck a transferable coin note, or redeem it into this grant's balance. A check needs only the note; redeeming needs a grant to credit.

The /peer/* routes sit outside the request gate: they consume no capacity, and running them through it would spend the grant's rate allowance on calls that exist to ask whether spending is possible. Each authenticates itself.


Log Rotation​

Log files (proxy-*.jsonl, proxy-attempts-*.jsonl, proxy-debug-*.jsonl) and old body-capture directories are automatically cleaned up to prevent unbounded growth.

ParameterValueDescription
Max age7 daysFiles older than 7 days are deleted
Max total size500 MBIf remaining files exceed 500 MB, oldest are deleted first
Cleanup triggersStartup + hourlyRuns once at proxy start, then every 60 minutes

The cleanupLogs() function performs two passes:

  1. Age pass -- delete all files with mtime older than the cutoff.
  2. Size pass -- if remaining files exceed the size limit, delete oldest first until under the cap.

Log rotation is non-fatal. If cleanup fails, the proxy continues operating normally.


Rate Limit Headers from Anthropic​

The proxy captures and uses Anthropic's quota headers for per-account utilization tracking:

HeaderFormatDescription
anthropic-ratelimit-unified-5h-utilizationfloat (0.0-1.0)5-hour rolling session utilization
anthropic-ratelimit-unified-5h-statusstringSession status (e.g., ok, warning)
anthropic-ratelimit-unified-5h-resetinteger (epoch)When the 5-hour window resets
anthropic-ratelimit-unified-7d-utilizationfloat (0.0-1.0)7-day rolling weekly utilization
anthropic-ratelimit-unified-7d-statusstringWeekly status
anthropic-ratelimit-unified-7d-resetinteger (epoch)When the 7-day window resets
anthropic-ratelimit-fallback-percentagefloatFallback percentage threshold
anthropic-ratelimit-overage-statusstringOverage status

These headers are parsed by parseQuotaHeaders() in accountQuota.ts and cached in memory with debounced persistence to ~/.neurolink/account-quotas.json. The neurolink auth list command displays per-account 5h and 7d utilization when available.


Token Refresh​

The proxy coordinates background, pre-request, and on-401 refresh paths:

  1. Background check — One non-overlapping cycle runs every 30 seconds and considers only allowed, enabled accounts within 5 minutes of expiry.
  2. Pre-request check — Before each request, if expiresAt <= now + 5 minutes, refresh inline via POST https://api.anthropic.com/v1/oauth/token (fallback: https://console.anthropic.com/v1/oauth/token).
  3. On-401 retry — If Anthropic returns a 401, refresh and retry within the bounded account retry budget before rotating.

Concurrent callers sharing a rotating refresh token reuse one in-flight result. 400, 401, 403, and 404 refresh responses are credential rejections and disable the account until explicit login. Network failures, refresh 429s, and 5xx responses are transient and receive a bounded auth cooldown. Automatic token persistence preserves manual disable metadata.