Quickstart
Four steps. There is no SDK to install. You keep the client you already use and point it somewhere new.
Create a workspace and a key
Sign up, create a workspace, and generate a key. It starts with
sy_live_and is shown once, so put it in your secret manager straight away.Add provider keys (optional)
Paste the Anthropic, OpenAI, Google, or OpenRouter keys you already pay for. Skip this step and every request runs on the Yard engine, which needs no provider key.
Rank your tracks
Pick a model for each track, strongest first: Track 1 Express, Track 2 Mainline, Track 3 Local. Track 4 is always the Yard engine.
Send a request
Point any OpenAI- or Anthropic-compatible client at Switchyard and send
switchyardas the model name. The response headers tell you who answered.
from openai import OpenAIclient = OpenAI( base_url="https://api.switchyard.space/v1", api_key="sy_live_...",)resp = client.chat.completions.with_raw_response.create( model="switchyard", messages=[{"role": "user", "content": "Explain this stack trace."}],)print(resp.headers["x-switchyard-model"]) # who answeredprint(resp.parse().choices[0].message.content)curl -i https://api.switchyard.space/v1/chat/completions \ -H "Authorization: Bearer $SWITCHYARD_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "switchyard", "messages": [{"role": "user", "content": "Write a commit message for this diff."}] }'The -i flag prints response headers, so x-switchyard-model and x-switchyard-track show up above the body.
Authentication
Every request carries your Switchyard key. Send it in either header below. Both work on every endpoint.
| Header | Value | Used by |
|---|---|---|
| Authorization | Bearer sy_live_… | OpenAI SDK, curl, most tools |
| x-api-key | sy_live_… | Anthropic SDK |
Keys are stored as hashes. We show a key once, when you create it, and can't show it again. If you lose one, create a new one.
To rotate, create a second key, move your clients over, then revoke the first. A revoked key gets 401 on its next request.
Connect your tools
Switchyard speaks the OpenAI and Anthropic formats. Connecting a tool means changing two settings: the base URL and the key.
Claude Code
Set two environment variables before you start a session. Claude Code sends its requests to Switchyard, and your tracks decide who answers.
export ANTHROPIC_BASE_URL=https://api.switchyard.spaceexport ANTHROPIC_AUTH_TOKEN=sy_live_...claudePut both exports in your shell profile to keep them across sessions.
Everything else
| Tool | Where to set it | Base URL |
|---|---|---|
| Cursor | Models settings, OpenAI base URL override | https://api.switchyard.space/v1 |
| Continue | apiBase on an openai provider | https://api.switchyard.space/v1 |
| Aider | --openai-api-base or OPENAI_API_BASE | https://api.switchyard.space/v1 |
| LangChain | base_url on ChatOpenAI | https://api.switchyard.space/v1 |
| OpenAI SDK | base_url (Python), baseURL (TS) | https://api.switchyard.space/v1 |
| Anthropic SDK | base_url (Python), baseURL (TS) | https://api.switchyard.space |
API reference
Base URL is https://api.switchyard.space. Every endpoint takes your key in Authorization or x-api-key and returns JSON.
| Endpoint | What it does |
|---|---|
| POST /v1/chat/completions | OpenAI Chat Completions format. Supports streaming and tool calls. |
| POST /v1/messages | Anthropic Messages format. Supports streaming and tool use. |
| POST /v1/messages/count_tokens | Counts the input tokens of a Messages request without running it. |
| GET /v1/models | Lists the models on your tracks, plus switchyard. |
| GET /v1/usage | Today's spend, per track. |
Response headers
Every response carries these, streamed or not. Log them and you always know where a request went.
| Header | Value |
|---|---|
| x-switchyard-model | The model that answered. |
| x-switchyard-track | 1 to 4. The track that answered. |
| x-switchyard-score | 0 to 100. The difficulty score the request was given. |
| x-switchyard-budget-remaining | Dollars left in today's budget for the track that answered. |
HTTP/2 200content-type: application/jsonx-switchyard-model: claude-sonnet-5.5x-switchyard-track: 2x-switchyard-score: 38x-switchyard-budget-remaining: 6.42Routing
Tracks are your ordered list of models. Track 1 is the strongest model you're willing to pay for. Track 4 is the Yard engine, and it is always there.
| Track | Name | Takes | Example models |
|---|---|---|---|
| 1 | Express | Architecture, hard bugs, long refactors. | Claude Fable 5.1, Claude Opus 5.5, GPT-6 Astra |
| 2 | Mainline | Everyday coding and tool calls once Track 1 has spent half its daily budget. | Claude Sonnet 5.5, Gemini 3.1 Pro, GPT-6.1 Sol |
| 3 | Local | Renames, summaries, lookups, commit messages. | Claude Haiku 4.5, Gemini 3.8 Flash, DeepSeek V4.1 Flash |
| 4 | Yard engine | Everything, when tracks 1–3 are out. No provider key needed. | gpt-oss-120b |
How scoring works
Before a request leaves, a small classifier gives it a difficulty score from 0 to 100. It reads the prompt, the attached tools, the context size, and the size of any diff. Scoring adds about 40 ms.
The request then goes to the highest track that still has budget for that score. The score comes back in x-switchyard-score and shows on the departure board.
Budgets move the threshold
Each track has a daily dollar budget. Early in the day, most requests stay on Track 1. As its budget is spent, the score needed to stay there rises. Light work moves down to Mainline and Local, and the hard problems keep the strongest model.
Conversation pinning
A conversation stays on one model. Changing models halfway through a session costs consistency, so Switchyard only switches when the pinned model is out.
What counts as out
- Budget. The track's daily dollar budget is spent.
- Rate limit. The provider returned
429. Switchyard reads the reset headers and uses the track again once the window resets. - Context. The conversation no longer fits in the model's context window.
When a model is out, the request is switched to the next track down. The departure board logs every request, switched or not.
Yard engine
The Yard engine is gpt-oss-120b running on Switchyard's own GPUs. It is Track 4 on every account and works with zero provider keys, so a new workspace can send requests on day one.
It takes whatever tracks 1–3 can't. When every provider budget and rate limit is spent, your tools keep getting answers.
| Setting | Standard | With upgrades |
|---|---|---|
| Context | 32K tokens | 128K tokens (Long context) |
| Cold start | Possible after idle time. | None (Warm engine). |
There is no subscription. Yard engine usage is pay-as-you-go from your credit balance: $0.40 per million input tokens and $1.60 per million output tokens. Credit comes from burning $YARD on Solana from a linked wallet, and the same credit buys the Warm engine, Long context and Wide batches upgrades.
Batch runs
Hand the Yard engine a list of prompts and collect the results later. Start, check, and cancel batches with the MCP tools batch_start, batch_status, and batch_cancel.
| Limit | Standard | Wide batches |
|---|---|---|
| Prompts per batch | 100 | 100 |
| Prompts in parallel | 4 | 16 |
| Max output tokens per prompt | 8,192 | 8,192 |
| Results kept | 7 days | 7 days |
MCP server
Switchyard runs an MCP server at https://api.switchyard.space/mcp. Connect it and your agent can route a prompt, run a batch, or check its credit without leaving the session.
claude mcp add --transport http switchyard https://api.switchyard.space/mcp \ --header "Authorization: Bearer sy_live_..."Other MCP clients with HTTP transport connect the same way: that URL, plus your key as a Bearer token.
| Tool | What it does |
|---|---|
| route | Sends one prompt through your tracks. Returns the answer, the model, and the track. |
| batch_start | Queues a list of prompts on the Yard engine. Returns a batch ID. |
| batch_status | Reports progress on a batch and returns the results that are done. |
| batch_cancel | Stops a batch. Prompts that already finished keep their results. |
| tracks_list | Lists your tracks, the model on each, and today's budget left. |
| credits_get | Returns your Yard engine credit balance. |
Limits & errors
Errors come back in the format of the endpoint you called, OpenAI or Anthropic, with a message that says what went wrong.
| Status | Meaning | What to do |
|---|---|---|
| 401 | The key is missing, malformed, or revoked. | Check the header and the sy_live_ prefix. Create a new key if this one was revoked. |
| 402 | No credit left for the Yard engine. | Add credit, or add a provider key so tracks 1–3 can take the work. |
| 409 | Batch limit reached. | Wait for a running batch to finish, or stop one with batch_cancel. |
| 413 | The request body is over 25 MB. | Trim the context or split the work across requests. |
| 429 | Every track is out, the Yard engine included. Rare. | Back off and retry. If it repeats, raise a track budget or add a provider. |
| 503 | Your provider and the Yard engine are both down. | Retry with exponential backoff. |