Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Reference guide for permanent free-tier LLM APIs with rate limits, model lists, and OpenAI-compatible integration patterns.
.claude/skills/leoyeai-awesome-free-llm-apis/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 126% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 117% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 167% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 258% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 472% | 0% |
> Skill by ara.so — Daily 2026 Skills collection.
A curated list of LLM providers offering permanent free tiers for text inference — no trial credits, no expiry. All endpoints listed are OpenAI SDK-compatible unless noted.
| Provider | Notable Models | Rate Limits | Region | |---|---|---|---| | Cohere | Command A, Command R+, Aya Expanse 32B | 20 RPM, 1K req/mo | 🇺🇸 | | Google Gemini | Gemini 2.5 Pro, Flash, Flash-Lite | 5–15 RPM, 100–1K RPD | 🇺🇸 (not EU/UK/CH) | | Mistral AI | Mistral Large 3, Small 3.1, Ministral 8B | 1 req/s, 1B tok/mo | 🇪🇺 | | Zhipu AI | GLM-4.7-Flash, GLM-4.5-Flash, GLM-4.6V-Flash | Undocumented | 🇨🇳 |
| Provider | Notable Models | Rate Limits | Region | |---|---|---|---| | Cerebras | Llama 3.3 70B, Qwen3 235B, GPT-OSS-120B | 30 RPM, 14,400 RPD | 🇺🇸 | | Cloudflare Workers AI | Llama 3.3 70B, Qwen QwQ 32B | 10K neurons/day | 🇺🇸 | | GitHub Models | GPT-4o, Llama 3.3 70B, DeepSeek-R1 | 10–15 RPM, 50–150 RPD | 🇺🇸 | | Groq | Llama 3.3 70B, Llama 4 Scout, Kimi K2 | 30 RPM, 1K RPD | 🇺🇸 | | Hugging Face | Llama 3.3 70B, Qwen2.5 72B, Mistral 7B | $0.10/mo free credits | 🇺🇸 | | Kluster AI | DeepSeek-R1, Llama 4 Maverick, Qwen3-235B | Undocumented | 🇺🇸 | | LLM7.io | DeepSeek R1, Flash-Lite, Qwen2.5 Coder | 30 RPM (120 with token) | 🇬🇧 | | NVIDIA NIM | Llama 3.3 70B, Mistral Large, Qwen3 235B | 40 RPM | 🇺🇸 | | Ollama Cloud | DeepSeek-V3.2, Qwen3.5, Kimi-K2.5 | 1 concurrent, light usage | 🇺🇸 | | OpenRouter | DeepSeek R1, Llama 3.3 70B, GPT-OSS-120B | 20 RPM, 50 RPD (1K with $10+) | 🇺🇸 |
Each provider has its own key management page:
bash# Store keys as environment variables — never hardcode them export GROQ_API_KEY="your_groq_key" export GEMINI_API_KEY="your_gemini_key" export OPENROUTER_API_KEY="your_openrouter_key" export MISTRAL_API_KEY="your_mistral_key" export COHERE_API_KEY="your_cohere_key" export CEREBRAS_API_KEY="your_cerebras_key" export GITHUB_TOKEN="your_github_pat" export HF_TOKEN="your_huggingface_token" export NVIDIA_API_KEY="your_nvidia_key" export CLOUDFLARE_API_TOKEN="your_cf_token" export CLOUDFLARE_ACCOUNT_ID="your_cf_account_id"
All providers (except Ollama Cloud) are OpenAI SDK-compatible — just swap the base_url and api_key.
pythonfrom openai import OpenAI # ── Groq ────────────────────────────────────────────────────────────────────── client = OpenAI( base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"], ) response = client.chat.completions.create( model="llama-3.3-70b-versatile", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) # ── Google Gemini ───────────────────────────────────────────────────────────── client = OpenAI( base_url="https://generativelanguage.googleapis.com/v1beta/openai/", api_key=os.environ["GEMINI_API_KEY"], ) response = client.chat.completions.create( model="gemini-2.0-flash", messages=[{"role": "user", "content": "Explain quantum entanglement."}], ) # ── Mistral AI ──────────────────────────────────────────────────────────────── client = OpenAI( base_url="https://api.mistral.ai/v1", api_key=os.environ["MISTRAL_API_KEY"], ) response = client.chat.completions.create( model="mistral-small-latest", messages=[{"role": "user", "content": "Write a haiku about code."}], ) # ── OpenRouter ──────────────────────────────────────────────────────────────── client = OpenAI( base_url="https://openrouter.ai/api/v1", api_key=os.environ["OPENROUTER_API_KEY"], ) response = client.chat.completions.create( model="deepseek/deepseek-r1", # free model on OpenRouter messages=[{"role": "user", "content": "What is 2+2?"}], extra_headers={ "HTTP-Referer": "https://yourapp.com", # optional but recommended "X-Title": "My App", }, ) # ── Cerebras ────────────────────────────────────────────────────────────────── client = OpenAI( base_url="https://api.cerebras.ai/v1", api_key=os.environ["CEREBRAS_API_KEY"], ) response = client.chat.completions.create( model="llama-3.3-70b", messages=[{"role": "user", "content": "Tell me a joke."}], ) # ── NVIDIA NIM ──────────────────────────────────────────────────────────────── client = OpenAI( base_url="https://integrate.api.nvidia.com/v1", api_key=os.environ["NVIDIA_API_KEY"], ) response = client.chat.completions.create( model="meta/llama-3.3-70b-instruct", messages=[{"role": "user", "content": "Summarize this text."}], ) # ── GitHub Models ───────────────────────────────────────────────────────────── client = OpenAI( base_url="https://models.inference.ai.azure.com", api_key=os.environ["GITHUB_TOKEN"], ) response = client.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": "Draft an email."}], ) # ── Cohere (OpenAI-compatible endpoint) ─────────────────────────────────────── client = OpenAI( base_url="https://api.cohere.com/compatibility/v1", api_key=os.environ["COHERE_API_KEY"], ) response = client.chat.completions.create( model="command-a-03-2025", messages=[{"role": "user", "content": "Translate to French: Hello world"}], )
typescriptimport OpenAI from "openai"; // ── Groq ────────────────────────────────────────────────────────────────────── const groq = new OpenAI({ baseURL: "https://api.groq.com/openai/v1", apiKey: process.env.GROQ_API_KEY, }); const completion = await groq.chat.completions.create({ model: "llama-3.3-70b-versatile", messages: [{ role: "user", content: "Hello!" }], }); console.log(completion.choices[0].message.content); // ── OpenRouter with free model router ──────────────────────────────────────── const openrouter = new OpenAI({ baseURL: "https://openrouter.ai/api/v1", apiKey: process.env.OPENROUTER_API_KEY, defaultHeaders: { "HTTP-Referer": "https://yourapp.com", "X-Title": "My App", }, }); // Use the free models router — automatically picks an available free model const freeCompletion = await openrouter.chat.completions.create({ model: "openrouter/free", messages: [{ role: "user", content: "What is the capital of France?" }], }); // ── Mistral ─────────────────────────────────────────────────────────────────── const mistral = new OpenAI({ baseURL: "https://api.mistral.ai/v1", apiKey: process.env.MISTRAL_API_KEY, }); const mistralCompletion = await mistral.chat.completions.create({ model: "mistral-small-latest", messages: [{ role: "user", content: "Explain async/await in JavaScript." }], });
Cloudflare uses a slightly different auth pattern:
pythonimport requests, os ACCOUNT_ID = os.environ["CLOUDFLARE_ACCOUNT_ID"] API_TOKEN = os.environ["CLOUDFLARE_API_TOKEN"] response = requests.post( f"https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai/run/" "@cf/meta/llama-3.3-70b-instruct-fp8-fast", headers={"Authorization": f"Bearer {API_TOKEN}"}, json={"messages": [{"role": "user", "content": "What is Cloudflare Workers?"}]}, ) result = response.json() print(result["result"]["response"])
typescript// Cloudflare Workers runtime (inside a Worker) export default { async fetch(request: Request, env: Env): Promise<Response> { const ai = new Ai(env.AI); const response = await ai.run("@cf/meta/llama-3.3-70b-instruct-fp8-fast", { messages: [{ role: "user", content: "Hello from Workers AI!" }], }); return Response.json(response); }, };
Ollama Cloud uses the Ollama API format, not the OpenAI format:
pythonimport requests, os response = requests.post( "https://ollama.com/api/chat", headers={"Authorization": f"Bearer {os.environ['OLLAMA_API_KEY']}"}, json={ "model": "deepseek-v3.2", "messages": [{"role": "user", "content": "What is 2 + 2?"}], "stream": False, }, ) print(response.json()["message"]["content"])
python# Using the ollama Python client import ollama, os client = ollama.Client( host="https://ollama.com", headers={"Authorization": f"Bearer {os.environ['OLLAMA_API_KEY']}"}, ) response = client.chat( model="qwen3.5", messages=[{"role": "user", "content": "Write a poem about the sea."}], ) print(response["message"]["content"])
pythonfrom openai import OpenAI import os client = OpenAI( base_url="https://router.huggingface.co/novita/v3/openai", api_key=os.environ["HF_TOKEN"], ) response = client.chat.completions.create( model="meta-llama/llama-3.3-70b-instruct", messages=[{"role": "user", "content": "Summarize the theory of relativity."}], max_tokens=512, ) print(response.choices[0].message.content)
pythonfrom openai import OpenAI import os client = OpenAI( base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"], ) with client.chat.completions.stream( model="llama-3.3-70b-versatile", messages=[{"role": "user", "content": "Write a short story about a robot."}], ) as stream: for text in stream.text_stream: print(text, end="", flush=True)
typescriptconst stream = await groq.chat.completions.create({ model: "llama-3.3-70b-versatile", messages: [{ role: "user", content: "Write a haiku." }], stream: true, }); for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); }
Cycle through providers when rate limits are hit:
pythonfrom openai import OpenAI, RateLimitError import os PROVIDERS = [ { "name": "Groq", "base_url": "https://api.groq.com/openai/v1", "api_key": os.environ.get("GROQ_API_KEY"), "model": "llama-3.3-70b-versatile", }, { "name": "Cerebras", "base_url": "https://api.cerebras.ai/v1", "api_key": os.environ.get("CEREBRAS_API_KEY"), "model": "llama-3.3-70b", }, { "name": "Mistral", "base_url": "https://api.mistral.ai/v1", "api_key": os.environ.get("MISTRAL_API_KEY"), "model": "mistral-small-latest", }, { "name": "OpenRouter", "base_url": "https://openrouter.ai/api/v1", "api_key": os.environ.get("OPENROUTER_API_KEY"), "model": "openrouter/free", }, ] def chat_with_fallback(messages: list[dict], **kwargs) -> str: for provider in PROVIDERS: if not provider["api_key"]: continue try: client = OpenAI( base_url=provider["base_url"], api_key=provider["api_key"], ) response = client.chat.completions.create( model=provider["model"], messages=messages, **kwargs, ) return response.choices[0].message.content except RateLimitError: print(f"Rate limited on {provider['name']}, trying next...") continue except Exception as e: print(f"Error on {provider['name']}: {e}, trying next...") continue raise RuntimeError("All providers exhausted.") # Usage answer = chat_with_fallback( messages=[{"role": "user", "content": "What is the speed of light?"}] ) print(answer)
OpenRouter provides a special router that automatically selects available free models:
pythonfrom openai import OpenAI import os client = OpenAI( base_url="https://openrouter.ai/api/v1", api_key=os.environ["OPENROUTER_API_KEY"], ) # Use the free router — picks from 29+ free models automatically response = client.chat.completions.create( model="openrouter/free", messages=[{"role": "user", "content": "Explain recursion."}], ) # Or use model fallbacks for priority ordering response = client.chat.completions.create( model="deepseek/deepseek-r1", messages=[{"role": "user", "content": "Explain recursion."}], extra_body={ "route": "fallback", "models": [ "deepseek/deepseek-r1", "meta-llama/llama-3.3-70b-instruct:free", "openrouter/free", ], }, )
pythonfrom langchain_openai import ChatOpenAI from langchain_core.messages import HumanMessage import os # Works with any OpenAI-compatible provider llm = ChatOpenAI( model="llama-3.3-70b-versatile", openai_api_base="https://api.groq.com/openai/v1", openai_api_key=os.environ["GROQ_API_KEY"], temperature=0.7, ) response = llm.invoke([HumanMessage(content="What are the SOLID principles?")]) print(response.content) # Gemini via LangChain gemini = ChatOpenAI( model="gemini-2.0-flash", openai_api_base="https://generativelanguage.googleapis.com/v1beta/openai/", openai_api_key=os.environ["GEMINI_API_KEY"], )
| Provider | RPM | RPD | Notes | |---|---|---|---| | Groq | 30 | 1,000 | 14,400 RPD for Llama 3.1 8B only | | Cerebras | 30 | 14,400 | — | | Gemini Flash | 15 | 1,500 | Not in EU/UK/CH | | Gemini 2.5 Pro | 5 | 25 | Not in EU/UK/CH | | GitHub Models | 10–15 | 50–150 | Varies by model tier | | OpenRouter (free) | 20 | 50 | 1K RPD after $10+ purchase | | Mistral | 1 req/s | — | 1B tokens/month cap | | NVIDIA NIM | 40 | — | — | | Cloudflare Workers AI | — | — | 10K neurons/day | | Cohere | 20 | — | 1K requests/month |
AuthenticationError
echo $GROQ_API_KEYRateLimitError
llama-3.1-8b-instant for the 14,400 RPD limitModel not found
:free suffix: meta-llama/llama-3.3-70b-instruct:free@cf/ prefix: @cf/meta/llama-3.3-70b-instruct-fp8-fastGemini free tier unavailable
Ollama Cloud not working with OpenAI SDK
ollama Python package or raw HTTPOpenRouter 50 RPD limit
openrouter/free router to distribute across all free modelsNeed highest RPD? → Cerebras (14,400 RPD)
Need smartest free model? → Gemini 2.5 Pro (if not in EU/UK/CH)
Need EU-hosted? → Mistral AI (France)
Need most model variety? → OpenRouter (29+ free models) or Cloudflare (48+ models)
Need fastest inference? → Groq (purpose-built inference chips)
Need reasoning model? → DeepSeek-R1 on Groq/OpenRouter/Kluster AI
Need vision? → Gemini Flash, Llama 4 Scout (Groq), GLM-4.6V-Flash (Zhipu)
No rate limit concern? → Cloudflare (10K neurons/day, compute-based)| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,260 | 17,920 | +17% | 1 | 1 | 0% | 2,984 | 8,983 | +201% | 0 | 0 | — |
case-02 | fail→pass | 18,710 | 14,813 | -21% | 1 | 1 | 0% | 3,554 | 8,015 | +126% | 0 | 0 | — |
case-03 | pass→pass | 5,183 | 4,769 | -8% | 1 | 1 | 0% | 884 | 6,201 | +601% | 0 | 0 | — |
case-04 | pass→pass | 13,196 | 8,192 | -38% | 1 | 1 | 0% | 2,352 | 6,921 | +194% | 0 | 0 | — |
case-05 | pass→pass | 4,101 | 9,322 | +127% | 1 | 1 | 0% | 954 | 7,123 | +647% | 0 | 0 | — |
case-06 | pass→pass | 4,646 | 4,299 | -7% | 1 | 1 | 0% | 907 | 6,182 | +582% | 0 | 0 | — |
case-07 | pass→pass | 6,177 | 4,145 | -33% | 1 | 1 | 0% | 1,233 | 6,066 | +392% | 0 | 0 | — |
case-08 | pass→pass | 4,445 | 3,279 | -26% | 1 | 1 | 0% | 836 | 5,822 | +596% | 0 | 0 | — |
case-09 | fail→pass | 15,688 | 5,803 | -63% | 1 | 1 | 0% | 2,955 | 6,402 | +117% | 0 | 0 | — |
case-10 | fail→pass | 14,733 | 5,482 | -63% | 1 | 1 | 0% | 2,364 | 6,320 | +167% | 0 | 0 | — |
case-11 | fail→pass | 11,955 | 9,241 | -23% | 1 | 1 | 0% | 1,871 | 6,698 | +258% | 0 | 0 | — |
case-12 | fail→pass | 5,693 | 2,474 | -57% | 1 | 1 | 0% | 995 | 5,688 | +472% | 0 | 0 | — |
case-13 | pass→pass | 4,222 | 3,592 | -15% | 1 | 1 | 0% | 720 | 5,852 | +713% | 0 | 0 | — |
case-14 | pass→pass | 8,693 | 2,543 | -71% | 1 | 1 | 0% | 1,412 | 5,689 | +303% | 0 | 0 | — |
case-15 | pass→pass | 7,958 | 4,259 | -46% | 1 | 1 | 0% | 1,511 | 6,099 | +304% | 0 | 0 | — |
case-16 | fail→pass | 11,201 | 7,225 | -35% | 1 | 1 | 0% | 1,951 | 6,639 | +240% | 0 | 0 | — |
case-17 | fail→fail | 7,955 | 6,808 | -14% | 1 | 1 | 0% | 1,557 | 6,673 | +329% | 0 | 0 | — |
case-18 | fail→pass | 7,936 | 4,892 | -38% | 1 | 1 | 0% | 1,547 | 6,262 | +305% | 0 | 0 | — |
case-19 | pass→pass | 14,656 | 5,767 | -61% | 1 | 1 | 0% | 2,390 | 6,177 | +158% | 0 | 0 | — |
case-20 | fail→pass | 8,428 | 6,859 | -19% | 1 | 1 | 0% | 1,476 | 6,665 | +352% | 0 | 0 | — |
case-21 | pass→pass | 5,920 | 1,864 | -69% | 1 | 1 | 0% | 1,101 | 5,610 | +410% | 0 | 0 | — |
case-22 | pass→pass | 16,479 | 12,026 | -27% | 1 | 1 | 0% | 2,869 | 7,364 | +157% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +36 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.