tashfeenahmed/freellmapi
7.4 billion tokens per month. 34 free LLM providers. 635 free model endpoints. All behind one /v1 endpoint, plus any custom OpenAI-compatible endpoint. Smart routing, automatic failover, encrypted keys. Personal experimentation only.
README
FreeLLMAPI
7.4 billion tokens per month. 34 free LLM providers. 635 free model endpoints. One OpenAI-compatible endpoint.
Aggregate free tiers from dozens of providers, plus custom OpenAI-compatible chat, embedding, image, and audio endpoints, behind a single /v1 API. Keys are stored encrypted. A router picks the best available model for each request, falls over to the next provider when one is rate-limited, and tracks per-key usage so you stay under every free-tier cap.
freellmapi.co · browse the full catalog: 474 model families, 635 free endpoints
English · 简体中文
Your router updates its own model catalog from a signed feed: new free models, quota changes, and compatibility fixes land without a git pull. Free installs get the monthly snapshot, so a model reaches them 30 days after it joins the live feed; premium routers get it the same day.
Go live at freellmapi.co ($19/yr, cancel anytime).
---
Contents
- Why this exists
- Supported providers
- Compatible CLIs & coding agents
- How it compares
- Features
- Quick start
- Desktop app
- Works with OpenAI-compatible clients
- Languages
- Premium (live catalog)
- Using the API
- Screenshots
- How it works
- FAQ
- Limitations
- Contributing
- Disclaimer
Why this exists
Every serious AI lab now offers a free tier, a few million tokens a month, a few thousand requests a day. On its own each tier is a toy. Stacked together, they add up to roughly 7.4 billion tokens per month of working inference capacity, across 474 model families / 635 provider endpoints from small-and-fast to reasonably capable.
The problem is that stacking them by hand is painful: thirty-four different SDKs, thirty-four different rate limits, thirty-four places a request can fail. FreeLLMAPI collapses that into one OpenAI-compatible endpoint. Point any OpenAI client library at your local server, and it routes transparently across whichever providers you've added keys for.
And the free-tier landscape shifts weekly: providers launch models, retire them, and change quotas without notice. FreeLLMAPI tracks all of that for you. The router pulls a signed model catalog from freellmapi.co on its own, so your install keeps up without a git pull. See Premium (live catalog) for how fast it keeps up.
Supported providers
![]() |
![]() Groq |
![]() Cerebras |
![]() OpenCode Zen |
![]() Mistral |
![]() OpenRouter |
![]() Cloudflare |
![]() Cohere |
![]() Z.ai (Zhipu) |
![]() NVIDIA |
![]() HuggingFace |
|
| ModelScope Qwen3 · DeepSeek V4 · GLM-5 (needs Aliyun cn binding) |
… and 22 more free providers
Plus a custom provider — point chat, embedding, image, or audio models at any OpenAI-compatible endpoint (llama.cpp, LM Studio, vLLM, a local Ollama, or a remote gateway) from the Keys page.
The full, always-current list lives at freellmapi.co/models with per-model rate limits, context windows, and free-token budgets.
Compatible CLIs & coding agents
![]() Claude Code |
![]() Codex CLI |
![]() Gemini CLI |
![]() Aider |
![]() Cline |
![]() Roo Code |
![]() Continue |
![]() OpenCode |
![]() Goose |
![]() Qwen Code |
![]() Kilo Code |
![]() Crush |
![]() Cursor |
![]() Zed |
![]() JetBrains AI |
![]() DeepSeek Harness |
![]() AtomCode |
… plus any OpenAI-compatible client, Anthropic SDK, Gemini SDK, or Ollama-capable app
Most of these configure themselves with one command — npx freellmapi setup-claude, setup-codex, setup-aider, setup-dsh (DeepSeek Harness), and eleven more generators that fetch your live catalog, back up existing config, and never clobber what's already there. Claude Code and Codex also get zero-persistence launchers (freellmapi launch, freellmapi launch-codex) that inject credentials into the child process only. Zed and JetBrains AI connect through the opt-in Ollama emulation; Gemini CLI speaks its native wire on /v1beta.
Per-tool recipes, the setup CLI reference, revocable URL tokens for headerless clients, and the MCP server all live in Clients & coding agents →
How it compares
Based on public documentation, July 2026 — corrections welcome.
Features
- Every OpenAI-style surface —
/v1/chat/completions,/v1/responses(what Codex CLI needs),/v1/completions(editor ghost-text autocomplete),/v1/images/generations,/v1/videos/generations,/v1/audio/speech,/v1/audio/transcriptions,/v1/embeddings, and/v1/models— streaming and non-streaming, from the official SDKs or any OpenAI-compatible client. API reference → - Anthropic Messages API —
/v1/messagesspeaks Anthropic's wire format over the same router, so Claude Code and the official Anthropic SDKs run against your free pool. Details → - Native Gemini + Ollama surfaces — Gemini CLI can use
/v1beta(generateContent, streaming, token counting, models), while opt-in Ollama emulation serves NDJSON chat/generate, tags, metadata, and embeddings for Zed, JetBrains, and other local-model clients. - Fusion (multi-model synthesis) — request the virtual
fusionmodel and the router fans your prompt out to a panel of diverse free models in parallel, then a judge model synthesizes one answer from the drafts. Details → - Image, video & speech generation —
/v1/images/generations,/v1/videos/generations, and/v1/audio/speechroute across the providers that serve media models; images and speech also accept custom OpenAI-compatible media endpoints. Video jobs are normalized across synchronous and queued providers and return a completed MP4. - Tool calling & structured outputs — OpenAI-style
toolsround-trip across providers (plain-text tool calls are rescued into realtool_calls), plusresponse_format,seed,logprobs, penalties, and the rest of the sampling params passed through per provider. - Smart routing, six strategies — live per-model speed/capability/reliability scores rank your chain; automatic fallover retries the next model on 429/5xx with cooldowns and key rotation. Routing in detail →
- Unified models & profiles — the same model on several providers collapses into one entry with strict in-group failover; named fallback-chain profiles (a coding chain, a vision chain) switch from the dashboard or per request via
auto:. - Per-key rate tracking — RPM/RPD/TPM/TPD counters per
(platform, model, key)that learn providers' reported ceilings, so routing always stays under every cap. - Self-updating model catalog — the router syncs a signed catalog from freellmapi.co twice a day: new models, quota changes, and provider quirk fixes land automatically. Free installs track the monthly snapshot, which each model joins 30 days after it lands in the live feed; premium routers get it same-day. Premium →
- Sticky sessions & context handoff — conversations stay on one model for 30 minutes; an optional compact handoff note keeps the thread coherent when a mid-chat switch does happen. Details →
- Prompt compression (opt-in) — a shared, fail-open request pipeline can deduplicate prompts, filter tool output, compact repeated JSON, and trim stale context before cache lookup and routing. Details →
- Encrypted keys, one token out — provider keys are AES-256-GCM encrypted in SQLite and decrypted in-memory per request; your apps only ever see a single unified
freellmapi-…bearer token. - Admin dashboard & analytics — React UI to manage keys, reorder the chain, run a playground, and read p50/p95/TTFT analytics over 24h–90d windows; login-gated, dark/light themes, 60 languages.
- MCP server & interactive docs — agents can introspect usable models, provider health, and routing strategy over
/mcp; a dependency-free OpenAPI viewer lives at/v1/docs. Coding agents → - Ops niceties — opt-in response cache, encrypted DB backups, periodic key health checks, bulk key import/export, declarative startup config. Install & deploy →
- Runs anywhere Node 20+ runs — Windows, macOS, Linux servers, or a small ARM SBC (Raspberry Pi included). ~40 MB RSS at idle behind PM2 / systemd / whatever supervisor you prefer.
Quick start
One-liner (Docker required — sets up ~/freellmapi, generates an encryption key, pulls the image, and starts the container):
curl -fsSL https://freellmapi.co/install.sh | bash
Prefer to read before you pipe to bash? The script is here. Re-running it is safe: your .env (and encryption key) is preserved and the container updates to :latest.
Open http://localhost:3001, add your provider keys on the Keys page, reorder the Fallback Chain to taste, and grab your unified API key from the Keys page header. That unified key is what you point your OpenAI SDK at.
On Windows, the easiest path is the desktop .exe installer from Releases (below). On Android, see the experimental Termux guide.
Everything else — Docker Compose, local development, declarative startup config, production builds, LAN access, and backups — is in docs/en/install/01-install.md.
Desktop app
A native menu-bar app lives in desktop/: the entire router + dashboard running locally from your tray, with a glass popover showing live request stats.
Download from Releases — the macOS .dmg and the Windows .exe installer are attached to every release. No account or password to set up: the only credential you need is the unified API key from the tray popover. Build-from-source steps and where your data lives: docs/en/install/01-install.md#desktop-app.
Works with OpenAI-compatible clients
Anything that can target an OpenAI-compatible base URL works: set it to http://localhost:3001/v1 with the unified key from the dashboard. Claude Code, Codex CLI, Cline / Roo Code, Continue (including inline autocomplete), Aider, opencode, and Cursor each have a short recipe in docs/en/clients/01-agent-clients.md — and the router doubles as an MCP server your agents can introspect mid-session.
The fastest setup is generated from the models available on your live server:
npx freellmapi setup-claude --url http://localhost:3001 --api-key
Every generator supports --dry-run, creates a timestamped backup before changing an existing file, and merges into the user's configuration. Launchers keep credentials out of config files entirely: npx freellmapi launch for Claude Code and npx freellmapi launch-codex for Codex.
| Agent | Automated setup | Base URL |
| --- | --- | --- |
| Claude Code | setup-claude | root |
| Codex CLI | setup-codex | /v1 |
| Cline | setup-cline | /v1 |
| Continue | setup-continue | /v1 |
| Aider | setup-aider | /v1 |
| OpenCode | setup-opencode | /v1 |
| Goose | setup-goose | /v1 |
| Qwen Code | setup-qwen | /v1 (or native /v1beta) |
| Roo / Kilo / Crush | setup-roo / setup-kilo / setup-crush | /v1 |
| DeepSeek Harness | setup-dsh | /v1 |
| MiMo Code | setup-mimo | /v1 |
| AtomCode | setup-atomcode | /v1 |
| Cursor | setup-cursor guide | public /v1 URL |
FreeLLMAPI is local-first and single-user by design. Your provider keys stay in your SQLite database, encrypted at rest, and requests go from your machine to the upstream providers you enabled.
Languages
The dashboard ships in 60 languages (the desktop tray menu in 6). The UI auto-detects your browser/system language on first load and you can switch any time from ⋯ → Settings; the choice is remembered. Right-to-left languages (العربية, עברית, فارسی, اردو) flip the whole layout automatically, and only the active language's dictionary is loaded — the rest never touch your bandwidth.




























