Documentation

Switchyard docs

Switchyard is one API key in front of every model you use, and it speaks both the OpenAI and Anthropic formats. Change a base URL, rank your tracks, and each request lands on the model that fits it.

Quickstart

Four steps. There is no SDK to install. You keep the client you already use and point it somewhere new.

  1. Create a workspace and a key

    Sign up, create a workspace, and generate a key. It starts with sy_live_ and is shown once, so put it in your secret manager straight away.

  2. Add provider keys (optional)

    Paste the Anthropic, OpenAI, Google, or OpenRouter keys you already pay for. Skip this step and every request runs on the Yard engine, which needs no provider key.

  3. Rank your tracks

    Pick a model for each track, strongest first: Track 1 Express, Track 2 Mainline, Track 3 Local. Track 4 is always the Yard engine.

  4. Send a request

    Point any OpenAI- or Anthropic-compatible client at Switchyard and send switchyard as the model name. The response headers tell you who answered.

Python · OpenAI SDK
from openai import OpenAIclient = OpenAI(    base_url="https://api.switchyard.space/v1",    api_key="sy_live_...",)resp = client.chat.completions.with_raw_response.create(    model="switchyard",    messages=[{"role": "user", "content": "Explain this stack trace."}],)print(resp.headers["x-switchyard-model"])  # who answeredprint(resp.parse().choices[0].message.content)
curl
curl -i https://api.switchyard.space/v1/chat/completions \  -H "Authorization: Bearer $SWITCHYARD_KEY" \  -H "Content-Type: application/json" \  -d '{    "model": "switchyard",    "messages": [{"role": "user", "content": "Write a commit message for this diff."}]  }'

The -i flag prints response headers, so x-switchyard-model and x-switchyard-track show up above the body.

Authentication

Every request carries your Switchyard key. Send it in either header below. Both work on every endpoint.

HeaderValueUsed by
AuthorizationBearer sy_live_…OpenAI SDK, curl, most tools
x-api-keysy_live_…Anthropic SDK

Keys are stored as hashes. We show a key once, when you create it, and can't show it again. If you lose one, create a new one.

To rotate, create a second key, move your clients over, then revoke the first. A revoked key gets 401 on its next request.

Connect your tools

Switchyard speaks the OpenAI and Anthropic formats. Connecting a tool means changing two settings: the base URL and the key.

Claude Code

Set two environment variables before you start a session. Claude Code sends its requests to Switchyard, and your tracks decide who answers.

shell
export ANTHROPIC_BASE_URL=https://api.switchyard.spaceexport ANTHROPIC_AUTH_TOKEN=sy_live_...claude

Put both exports in your shell profile to keep them across sessions.

Everything else

ToolWhere to set itBase URL
CursorModels settings, OpenAI base URL overridehttps://api.switchyard.space/v1
ContinueapiBase on an openai providerhttps://api.switchyard.space/v1
Aider--openai-api-base or OPENAI_API_BASEhttps://api.switchyard.space/v1
LangChainbase_url on ChatOpenAIhttps://api.switchyard.space/v1
OpenAI SDKbase_url (Python), baseURL (TS)https://api.switchyard.space/v1
Anthropic SDKbase_url (Python), baseURL (TS)https://api.switchyard.space

API reference

Base URL is https://api.switchyard.space. Every endpoint takes your key in Authorization or x-api-key and returns JSON.

EndpointWhat it does
POST /v1/chat/completionsOpenAI Chat Completions format. Supports streaming and tool calls.
POST /v1/messagesAnthropic Messages format. Supports streaming and tool use.
POST /v1/messages/count_tokensCounts the input tokens of a Messages request without running it.
GET /v1/modelsLists the models on your tracks, plus switchyard.
GET /v1/usageToday's spend, per track.

Response headers

Every response carries these, streamed or not. Log them and you always know where a request went.

HeaderValue
x-switchyard-modelThe model that answered.
x-switchyard-track1 to 4. The track that answered.
x-switchyard-score0 to 100. The difficulty score the request was given.
x-switchyard-budget-remainingDollars left in today's budget for the track that answered.
Response headers
HTTP/2 200content-type: application/jsonx-switchyard-model: claude-sonnet-5.5x-switchyard-track: 2x-switchyard-score: 38x-switchyard-budget-remaining: 6.42

Routing

Tracks are your ordered list of models. Track 1 is the strongest model you're willing to pay for. Track 4 is the Yard engine, and it is always there.

TrackNameTakesExample models
1ExpressArchitecture, hard bugs, long refactors.Claude Fable 5.1, Claude Opus 5.5, GPT-6 Astra
2MainlineEveryday coding and tool calls once Track 1 has spent half its daily budget.Claude Sonnet 5.5, Gemini 3.1 Pro, GPT-6.1 Sol
3LocalRenames, summaries, lookups, commit messages.Claude Haiku 4.5, Gemini 3.8 Flash, DeepSeek V4.1 Flash
4Yard engineEverything, when tracks 1–3 are out. No provider key needed.gpt-oss-120b

How scoring works

Before a request leaves, a small classifier gives it a difficulty score from 0 to 100. It reads the prompt, the attached tools, the context size, and the size of any diff. Scoring adds about 40 ms.

The request then goes to the highest track that still has budget for that score. The score comes back in x-switchyard-score and shows on the departure board.

Budgets move the threshold

Each track has a daily dollar budget. Early in the day, most requests stay on Track 1. As its budget is spent, the score needed to stay there rises. Light work moves down to Mainline and Local, and the hard problems keep the strongest model.

Conversation pinning

A conversation stays on one model. Changing models halfway through a session costs consistency, so Switchyard only switches when the pinned model is out.

What counts as out

  • Budget. The track's daily dollar budget is spent.
  • Rate limit. The provider returned 429. Switchyard reads the reset headers and uses the track again once the window resets.
  • Context. The conversation no longer fits in the model's context window.

When a model is out, the request is switched to the next track down. The departure board logs every request, switched or not.

Yard engine

The Yard engine is gpt-oss-120b running on Switchyard's own GPUs. It is Track 4 on every account and works with zero provider keys, so a new workspace can send requests on day one.

It takes whatever tracks 1–3 can't. When every provider budget and rate limit is spent, your tools keep getting answers.

SettingStandardWith upgrades
Context32K tokens128K tokens (Long context)
Cold startPossible after idle time.None (Warm engine).

There is no subscription. Yard engine usage is pay-as-you-go from your credit balance: $0.40 per million input tokens and $1.60 per million output tokens. Credit comes from burning $YARD on Solana from a linked wallet, and the same credit buys the Warm engine, Long context and Wide batches upgrades.

Batch runs

Hand the Yard engine a list of prompts and collect the results later. Start, check, and cancel batches with the MCP tools batch_start, batch_status, and batch_cancel.

LimitStandardWide batches
Prompts per batch100100
Prompts in parallel416
Max output tokens per prompt8,1928,192
Results kept7 days7 days

MCP server

Switchyard runs an MCP server at https://api.switchyard.space/mcp. Connect it and your agent can route a prompt, run a batch, or check its credit without leaving the session.

Claude Code
claude mcp add --transport http switchyard https://api.switchyard.space/mcp \  --header "Authorization: Bearer sy_live_..."

Other MCP clients with HTTP transport connect the same way: that URL, plus your key as a Bearer token.

ToolWhat it does
routeSends one prompt through your tracks. Returns the answer, the model, and the track.
batch_startQueues a list of prompts on the Yard engine. Returns a batch ID.
batch_statusReports progress on a batch and returns the results that are done.
batch_cancelStops a batch. Prompts that already finished keep their results.
tracks_listLists your tracks, the model on each, and today's budget left.
credits_getReturns your Yard engine credit balance.

Limits & errors

Errors come back in the format of the endpoint you called, OpenAI or Anthropic, with a message that says what went wrong.

StatusMeaningWhat to do
401The key is missing, malformed, or revoked.Check the header and the sy_live_ prefix. Create a new key if this one was revoked.
402No credit left for the Yard engine.Add credit, or add a provider key so tracks 1–3 can take the work.
409Batch limit reached.Wait for a running batch to finish, or stop one with batch_cancel.
413The request body is over 25 MB.Trim the context or split the work across requests.
429Every track is out, the Yard engine included. Rare.Back off and retry. If it repeats, raise a track budget or add a provider.
503Your provider and the Yard engine are both down.Retry with exponential backoff.