Claude Code with Free AI Models and Automatic Failover

Project on: https://github.com/brsnet/claude-code-switcher

Claude Code is a powerful development assistant, but using premium models continuously can become expensive.

Claude Code Switcher was created with a free-first approach: use free AI providers and local models whenever possible, while keeping multiple alternatives available when a model is overloaded or reaches its rate limit.

One local endpoint, multiple free providers

Claude Code connects to a single local address:

$env:ANTHROPIC_BASE_URL = "http://127.0.0.1:8083"
$env:ANTHROPIC_AUTH_TOKEN = "local-switcher"
claude

Behind this endpoint, the switcher can route requests across services such as:

  • OpenRouter free models
  • Groq
  • Google Gemini
  • Cerebras
  • Cloudflare Workers AI
  • NVIDIA NIM
  • Ollama and other local runtimes

Free plans have quotas and availability limits, so relying on only one provider is risky. The switcher lets you create an ordered fallback route:

ROUTER_SONNET=groq,gemini,open_router,cloudflare,cerebras,ollama

If Groq is unavailable, the request can move to Gemini, OpenRouter, Cloudflare, Cerebras, or finally a local Ollama model.

Protection against accidental OpenRouter charges

Claude Code Switcher accepts only the OpenRouter free router or models explicitly marked as free:

OPENROUTER_MODEL=openrouter/free

Models without :free are rejected before the external request is made.

This safeguard reduces the risk of accidentally selecting a paid OpenRouter model.

Free services are not unlimited, however. Their quotas, models, and policies can change, which is why the switcher supports multiple providers instead of depending on only one.

Automatic failover without duplicate work

The switcher can retry temporary errors and change providers when it encounters:

  • Rate limits
  • Provider overload
  • Connection timeouts
  • Temporary server failures
  • Insufficient balance
  • Unavailable models

Failover is allowed only before the model produces meaningful output.

After text or a tool call begins, the switcher will not silently move the request to another model. This prevents duplicated text, commands, and file modifications.

Use cloud and local models together

Local models can be part of the same route as free cloud APIs:

ROUTER_HAIKU=groq,open_router,ollama

This provides a useful balance:

  • Free cloud models for speed and tool support
  • Local models for privacy and independence
  • Multiple fallbacks for better availability
  • One consistent Claude Code experience

Fewer tools, faster requests

Sending dozens of tool definitions increases token usage and can confuse smaller models.

Claude Code Switcher uses a focused default set:

TOOL_ALLOWLIST=Read,Edit,Write,Bash,Glob,Grep

This keeps requests smaller and helps free or local models focus on the tools commonly needed for coding.

Clear provider and model logs

The terminal shows what happens with every request:

ROUTE_CHAIN     candidates=['groq/...', 'gemini/...', 'open_router/openrouter/free']
ROUTE           provider=groq model='openai/gpt-oss-120b'
PROVIDER_STREAM provider=groq first_event_ms=...
REQUEST_DONE    result=success total_ms=...

When a provider fails, the logs show the error, retry, and next candidate.

This makes it easy to understand:

  • Which free provider was selected
  • Which model handled the request
  • Whether failover occurred
  • How long the response took
  • Why a provider failed

Secrets, prompts, and private file contents are not included in these operational logs.

Why use Claude Code Switcher?

The main benefits are:

  • Free-first routing: prioritize free providers and local models.
  • Cost protection: reject paid OpenRouter models.
  • Automatic failover: continue when a provider is unavailable.
  • Local model support: use Ollama, LM Studio, or llama.cpp.
  • Lower token overhead: send only essential tools.
  • Transparent operation: see the selected provider, model, and latency.
  • Provider independence: change models without changing how you use Claude Code.

Getting started

Configure your provider keys in .env, select only the providers you want, and start the gateway:

uv sync
uv run uvicorn server:app --host 0.0.0.0 --port 8083

Then point Claude Code to the local endpoint and start working.

Claude Code Switcher provides one stable interface for free cloud models, local inference, and automatic fallback.

The objective is simple: use Claude Code with more flexibility, greater resilience, and the lowest practical cost.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top