What can you run free right now?
Every row is a provider program a developer can use at no charge, with the numbers the provider publishes. Unknowns say so. Nothing here is a signup credit or a paid prerequisite.
- Verified programs
- 42
- OpenRouter free models
- 15
- Overdue for recheck
- 0
A program can appear under more than one filter. If the evidence does not confirm a requirement, that program stays out of that filter. Browse all programs to see everything.
Current program ledger
| Models | Tightest limit | To start | Your data | Type | Key link and details | |||
|---|---|---|---|---|---|---|---|---|
| Apodex: Apodex 1.1 Mini, Cohere: North Mini Code, Dots Studio: Dots3-Note Preview15 models | 50 requests / dayA one-time $10 purchase raises this to 1,000 requests / day | 20 requests / min | Signup, no payment | Varies by route | Standing | verified | Get a key ↗ | |
Program detailsFree accessExact :free model variants list $0 input and output pricing. Accounts with under $10 in lifetime credit purchases receive 20 requests/minute and 50 requests/day; accounts with at least $10 purchased receive 1,000 requests/day.
Child-model disclosureEvery model OpenRouter currently lists as a :free variant. The children below are read from OpenRouter's public model listing each morning, so the roster changes as models enter and leave the free catalog.
Signup, payment, data
What the provider's page says
| ||||||||
| coding-glm-5.3-free, gpt-5.5-free, gemini-3.8-flash-free +1 | 1M tokens / dayshared pool | 10 requests / min | One-time $1 deposit | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessEvery new account receives 10 trial calls, shared across all free models, with no credit card and no expiry. A one-time top-up of at least $1 permanently switches the account to daily quotas on the free catalog: 100 requests per day, 10 per minute, and one shared pool of 1M tokens per day, resetting daily. Free models bill $0 within these quotas. Some models carry a higher request weight and have lower caps, ranging down to 10 requests per day and 1 per minute.
Models disclosed60 models from 16 model authors listed at $0 input and $0 output on the provider's free catalog page on 2026-10-09. Nearly all are text models, plus a few image models (gpt-image-2-free, gemini-3.1-flash-image-preview-free). The roster changes over time. Signup, payment, data
What the provider's page says | ||||||||
| gpt-5.6-sol, gpt-5.5-2026-04-23, gpt-5.4-2026-03-05 +4 | 1M tokens / dayshared pool | Not stated | One-time $5 deposit | Trains on it | Limited-time | verified | Get a key ↗ | |
Program detailsFree accessExplicitly limited free promotion for eligible API organizations that opt in to share API inputs and outputs. Eligible Launch and Grow organizations receive up to 1,000,000 tokens/day in the large-model group and 10,000,000 tokens/day in the small-model group; eligible Build organizations receive 250,000 and 2,500,000 tokens/day respectively. Eligibility is shown on the data sharing settings page, and organizations that do not see the offer there are not eligible. Quotas are shared within each group and reset at 00:00 UTC. A request that would cross the daily quota is billed in full, and overage is billed normally. Only traffic from sharing-enabled projects qualifies. A positive account balance is required. Fine-tuned models, fine-tuning training, evals, and tool use are excluded.
Models disclosedOn 2026-10-09, the 1M/250K group listed: gpt-6.1-sol, gpt-6-astra, gpt-6-sol, gpt-6-luna, gpt-5.6-sol, gpt-5.5-2026-04-23, gpt-5.4-2026-03-05, gpt-5.2-2025-12-11, gpt-5.1-2025-11-13, gpt-5.1-codex, gpt-5-codex, gpt-5-2025-08-07, gpt-5-chat-latest, gpt-4.1-2025-04-14, gpt-4o-2024-05-13, gpt-4o-2024-08-06, gpt-4o-2024-11-20, o3-2025-04-16, o1-preview-2024-09-12, and o1-2024-12-17. On 2026-10-09, the 10M/2.5M group listed: gpt-5.6-terra, gpt-5.6-luna, gpt-5.4-mini-2026-03-17, gpt-5.4-nano-2026-03-17, gpt-5.1-codex-mini, gpt-5-mini-2025-08-07, gpt-5-nano-2025-08-07, gpt-4.1-mini-2025-04-14, gpt-4.1-nano-2025-04-14, gpt-4o-mini-2024-07-18, o4-mini-2025-04-16, o1-mini-2024-09-12, and codex-mini-latest. The page also keeps gpt-4.5-preview in its list but marks it shut down, so it is not current access. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | 10 requests / min | Signup, no payment | Does not train | Limited-time | verified | Get a key ↗ | |
Program detailsFree accessFree hosted inference while the Inference API is in experimental status; Hetzner says it will notify users in advance by email if that status changes. Rate limits are applied per API key and the API returns HTTP 429 when any of them is exceeded, but as of 2026-09-23 the Inference API docs page no longer publishes the numeric request and token figures. Performance and availability are not guaranteed, requests can queue at peak, and the platform is not for production use.
Models disclosedOn the Inference API docs page (2026-09-17): Qwen/Qwen3.6-35B-A3B-FP8 and Qwen3.8-27B, both 262,144-token context with text and image input. The models endpoint is the definitive list and the selection changes during the experiment; GLM-5.2, DeepSeek-V4-Flash-0731 and Kimi-K2.7-Code were served at launch and dropped when Hetzner paused large models in August 2026. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | Not stated | Signup, payment unknown | Not stated | Limited-time | verified | Get a key ↗ | |
Program detailsFree accessPoolside offers provider-hosted, OpenAI-compatible chat inference for its Laguna models. The models page says Laguna is free to use for a limited time and shows Laguna S 2.1 and Laguna XS 2.1 under that offer. The Laguna XS 2.1 launch post says free and paid endpoints are both available, with paid pricing of $0.10, $0.20 and $0.05 per 1M input, output and cache-read tokens. No free-endpoint quota or end date is disclosed.
Models disclosedOn 2026-10-09 the models page showed Laguna S 2.1 (poolside/laguna-s-2.1) and Laguna XS 2.1 under the limited-time free offer for the hosted API. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | Not stated | Signup, payment unknown | Not stated | Limited, ends 2027-03-31 | verified | Get a key ↗ | |
Program detailsFree accessUpstage's AI support program gives eligible K-12 schools, universities, university hospitals, nonprofits, and NGOs free access to the provider-hosted Solar Pro LLM and Document Parse for up to one year. Access is application-based and approved per organization; it is free access, not credits or a paid subscription.
Models disclosedSolar Pro 4 (LLM) and Document Parse, as named on the program page on 2026-10-09. Solar Pro 2 and Solar Pro 3 are no longer named in the program, and Upstage lists both for end of service on 2026-10-30 (KST). Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | Not stated | Card on file | Varies by route | Limited-time | verified | Get a key ↗ | |
Program detailsFree accessOpenCode Zen currently prices thirteen models at Free on its pricing table. Twelve are Free for input, output, and cached-read tokens, and Jev 1.13 Free is Free for input and output. The provider describes every free model as available for a limited time and publishes no request quota or end date. Zen's setup flow says to sign in, add billing details, and copy an API key.
Models disclosedFree rows on the Zen pricing table on 2026-10-09: Big Pickle, Space Bunny Free, LongCat 2.5 Preview Free, Step 5 Preview Free, Exo Free, MiMo-V2.6-Flash Free, MiMo-V2.5 Free, Ling 3.1 Flash Free, Ling 3.0 Flash Fin Free, Nemotron 3 Ultra Free, Nemotron 3.5 Lightning Free, Muse Spark 1.3 Contributor Free, and Jev 1.13 Free. Jev 1.13 Free is a TypeSafe AI System One model that answers typed yes/no, multiple-choice, and rubric questions instead of generating text. Union Alpha Free was no longer listed on 2026-10-09. Muse Spark 1.2 Contributor Free left the free table by 2026-09-17, and Ox Alpha Free and Hy3 Free were removed before 2026-09-01. The Nemotron endpoints are NVIDIA trial endpoints whose use is logged. Other Zen models are pay-as-you-go. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | 50 requests / dayper model | 10 requests / min | Signup, no payment | Varies by route | Standing | verified | Get a key ↗ | |
Program detailsFree accessAccounts without credit get 10 requests per minute and 60 weighted units per day on the free models; accounts holding credit get 20 requests per minute and 120 weighted units per day. The daily counter resets at 00:00 UTC, and longer prompts use more units. The free page says these limits are shared across free models, while the docs page says each free model is counted separately. An additional IP limit, a site-wide free cap of 15 requests per minute, and a limit of 2 concurrent free requests per account also apply. Past the quota, requests continue at the paid rate only when the account holds credit and paid fallback is on (a header can turn that fallback off); otherwise the request gets a 429 until the quota resets. The docs' example rate-limit header (200 requests per minute) is the paid-tier account limit, not the free quota.
Models disclosedTwo current free-quota models on 2026-10-09 (Qwen: Qwen3.7 Flash and DeepSeek: DeepSeek V4 Flash 0731free) plus auto:free routing; the catalog changes over time. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | 50 requests / dayA one-time IDR 10,000 purchase raises this to 1,000 requests / day | 5 requests / min | Signup, no payment | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessKenari offers free account registration and a Solo tier requiring no card, with 50 requests per day across the current free-model roster (14 models on 2026-10-09) at up to 5 requests per minute. A cumulative top-up of Rp 10,000 raises that to 1,000 requests per day. Its quickstart says users may skip top-up, create a Kenari API key, and call a :free text model through the OpenAI-compatible /v1/chat/completions endpoint without reducing their balance. The free lane is best-effort with no speed or availability guarantee; the daily quota returns HTTP 429 with a Retry-After header when exhausted.
Models disclosedFourteen provider-hosted :free text models per Kenari's homepage on 2026-10-09 (it read seventeen on 2026-09-17, eleven on 2026-09-04 and twelve on 2026-08-27); Kenari says the free roster changes over time with the provider catalog. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | 200 requests / hour | No signup | Varies by route | Standing | verified | Get a key ↗ | |
Program detailsFree accessKilo offers the dynamic kilo-auto/free route with no credits required and several individually identified free models. Free models are available to authenticated and anonymous users. Anonymous use is limited to 200 requests per hour per IP address; upstream providers may impose additional rate limits. Users must configure agentic interactions, autocomplete, and background tasks separately if they want every Kilo inference path to remain free.
Models disclosedKilo's gateway documentation does not publish a fixed list of free model IDs and directs users to the live model catalog for current free options. Free models are identified by a :free suffix in the model ID, and on 2026-09-22 the authentication documentation gives minimax/minimax-m2.1:free and z-ai/glm-5:free as its examples. kilo-auto/free selects a model dynamically per session from a curated set of available free models, and that mapping updates server-side as free model availability shifts. For an account with no balance, the kilo-auto/small background route resolves to google/gemma-4-26b-a4b-it:free. Signup, payment, data
What the provider's page says
| ||||||||
| MythoLite | Unlimitedper model | Not stated | Signup, no payment | Trains, opt-out | Standing | verified | Get a key ↗ | |
Program detailsFree accessMancer's live model table shows MythoLite with an input token cost and an output token cost of FREE. Mancer's prices page says that some models are free while most have a per-token credit cost, and it tells users to test on the free models before making a purchase. Free models stay usable even when an account's credit total is negative, so free usage does not depend on buying or holding credits.
Models disclosedMythoLite is the only model marked FREE for both input and output tokens on Mancer's model table as read on 2026-10-08. It is listed as a LLaMA 2 architecture model at 13B parameters, tagged Creative and Roleplay, and shown as online. Mancer does not state MythoLite's context window or maximum output length. Paid models on the same table include MythoMax, listed at 0.14 credits per input token and 0.24 credits per output token. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | 50 requests / dayper model | 60 requests / min | Signup, no payment | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessMegaNova hosts mistralai/Mistral-Small-3.2-24B-Instruct-2506 as a recurring free-quota text model, listed as free on 2026-09-27. Free Tier 1 registration requires no credit card and includes 50 requests per day for this model, within Tier 1 rate limits of 60 requests per minute and 200,000 tokens per minute. Free quota is granted per account and resets daily at 00:00 UTC. MegaNova documents that Tier 2 and above keep free quota access with higher limits. Once the daily free quota runs out, requests normally fail with a quota error unless the account has turned on the optional paid fallback, so this is genuine standing free hosted inference but not unlimited zero-cost usage.
Models disclosedmistralai/Mistral-Small-3.2-24B-Instruct-2506 is the documented recurring free-quota text model, labeled Free on its model page with a context length of 8,192 tokens on 2026-09-27. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | Not stated | Signup, no payment | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessMerge's provider-owned catalog lists nvidia/nemotron-3.5-lightning-30b-a3b as available hosted text inference with one NVIDIA host and Free / Free input and output pricing, as recorded on 2026-09-30. Merge's Free plan (pricing page, 2026-09-04) includes access to all free LLMs, automatic routing with built-in fallback, and starts without a credit card; the earlier $10 first-month credit was removed from the Free tier, while Pro adds $10 in expiring monthly credits and ordinarily charges model cost plus 5%. Merge's Gateway get started page adds that the Free plan covers its list of free models up to a daily request cap, and that adding a card moves the account to Pro and the full catalog. Merge does not publish the size of that daily cap on any of these pages. Merge also warns that its response usage.cost is provider cost rather than necessarily the final invoice, but the exact Nemotron route is explicitly cataloged as free rather than merely lacking pricing metadata.
Models disclosednvidia/nemotron-3.5-lightning-30b-a3b through one NVIDIA host at Free / Free input-output pricing. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | 50 requests / min | Signup, payment unknown | Trains, opt-out | Standing | verified | Get a key ↗ | |
Program detailsFree accessThe $0 Free plan provides access only to models currently marked free, with standard rate limits and no monthly credit balance.
Models disclosedTen catalog entries listed at $0.00/1M on 2026-10-09: inclusionAI Ling 3.0 Flash Fin; inclusionAI Ling 3.0 Flash Sante; inclusionAI Ling 3.1 Flash; Meituan LongCat 2.0; Meituan LongCat 2.5 Preview; Poolside Laguna S 2.1; Poolside Laguna XS 2.1; StepFun Step 3.7 Flash; StepFun Step 5 Preview; Upstage Solar Mini 4. On 2026-10-09 Upstage Solar Pro 4 was no longer listed at $0.00 (now $0.09 in / $0.36 out per 1M), and Tencent Hy3 remained paid at $0.13 in / $0.53 out per 1M. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | Not stated | Signup, payment unknown | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessOrcaRouter serves selected catalog models at $0 under model ids ending in -free, and the built-in orcarouter/free router covers that free tier behind a single id. Calls to these ids never touch the account wallet, and the allowance is a request rate rather than a balance, so a workspace can keep calling the free tier indefinitely while it stays inside the limits. Free traffic is capped by requests per minute and requests per day. Those ceilings are public, returned for every tier by the GET /api/free-package/public endpoint, and the provider tunes them live. They are tiered by lifetime spend: a workspace that has never topped up gets a deliberately small daily allowance and also a per-request prompt-size cap, which is the one limit the provider does not publish as a number, and crossing the lifetime-spend threshold raises both. Paying is optional and only buys higher limits. Free access additionally requires the workspace owner to have linked a GitHub account that has been registered for some time, unless the workspace has paid enough to be exempt. Model ids without the -free suffix bill at the upstream per-token price, and a free id never falls back to a paid model.
Models disclosedThe free lineup is the catalog ids ending in -free plus the orcarouter/free router, which scores each request by difficulty and picks among the workspace's free models. The provider states that the lineup rotates as models join and leave with capacity, so its documentation deliberately lists no ids: the current ids, and the paid model each one shadows, come from the public free-package endpoint. On 2026-10-01 the orcarouter/free model page listed one model in the router, deepseek/deepseek-v4-flash-free, on the light tier. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | Not stated | Signup, payment unknown | Trains on it | Limited-time | verified | Get a key ↗ | |
Program detailsFree accessLLMTR exposes poolside/laguna-xs-2.1 for hosted text inference at $0 input and $0 output. Its documentation says Poolside serves the model free, LLMTR does not deduct from the user's credit balance, and there is no separate token allowance, only LLMTR's standard request-rate limit. The offer may change if Poolside ends free access. Prompts and completions on the free route may be logged and used for product improvement.
Models disclosedpoolside/laguna-xs-2.1 hosted through LLMTR at $0 input/output. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | Not stated | One-time $5 deposit | Does not train | Standing | verified | Get a key ↗ | |
Program detailsFree accessFree model after a one-time minimum $15 first credit purchase (later top-ups have a $5 minimum), as stated by Together on 2026-10-09; a positive balance is required. Ternary Bonsai 27B lists $0 input and $0 output token prices, so its tokens do not consume the purchased balance under published pricing. Other paid usage can drain the balance, and API access is suspended at zero until credits are added. Rate limits are dynamic and best-effort, not a guaranteed unlimited allowance.
Models disclosedPrism-ML/Ternary-Bonsai-27B is listed with zero input and output token charges. Other Together models are not covered by this free listing. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | Not stated | Not stated | Signup, payment unknown | Not stated | Limited, ends 2026-11-01 | verified | Get a key ↗ | |
Program detailsFree accessRegolo's pricing page lists brick-v1-beta at €0.00 input and €0.00 output per 1M tokens under pay-as-you-go, and the models library shows the same model's input and output as Free, so a pay-as-you-go account can call it at a zero token price. Pay-as-you-go carries a €0 monthly fee and bills only for the tokens consumed, which makes this route distinct from Regolo's separate 30-day free trial. Every other model in the library carries a paid per-token rate or is billed per request, per second of audio, per pixel, or per query, so brick-v1-beta is the only zero-cost inference route here. On 2026-10-07 the pricing page also carried a notice that updated rates go live on November 1 and that current pricing rates and plan rules remain active until that date, and the reviewed pages do not state what brick-v1-beta will cost after that change.
Models disclosedOn 2026-10-07 brick-v1-beta was the only model listed at a zero token price. apertus-70b, brick-complexity-pro, gemma4-31b, glm5.2, gpt-oss-120b, gpt-oss-20b, mistral-small-4-119b, qwen3.5-122b, qwen3.5-9b, and qwen3.8-27b all carried paid per-token rates. deepseek-ocr-2, faster-whisper-large-v3, Qwen-Image, Qwen3-Embedding-8B, and Qwen3-Reranker-4B are billed per request, per second of audio, per pixel, or per query rather than per token, so none of them is a free inference route. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | 200 requests / day | 20 requests / min | Signup, no payment | Varies by route | Standing | verified | Get a key ↗ | |
Program detailsFree accessRequesty's pricing and documentation confirm hosted models with no token charges. Enrollment requires a Requesty organization or account, and the Free plan requires no credit card. New organizations receive 200 free-model requests per day and 20 requests per minute across all free models combined; paying organizations receive 1,000 per day and 60 per minute. The daily allowance resets each day, and exceeding a limit returns a rate-limit error without affecting paid models. Requesty says the models are free for now, may change pricing later, and will announce changes in its changelog.
Models disclosedRequesty's free-models page listed 13 free models on 2026-10-09: NVIDIA Nemotron 3 Ultra 550B, Nemotron 3 Super 120B, Nemotron 3 Nano 30B, Nemotron 3 Nano Omni 30B Reasoning, Nemotron 3.5 Content Safety, and Nemotron 3.5 Lightning 30B; Poolside Laguna M.1 and Laguna XS.2; Google Gemma 4 31B; Mistral Leanstral 1.5; InclusionAI Ling 3.0 Tiny and Ling 3.1 Flash; and Meta Muse Glimmer 30B. The docs page carries a shorter nine-row table. Requesty says the set is free for now and announces changes in its changelog. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | Not stated | Not stated | Signup, no card | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessStanding free-model roster on the CN pricing page, which presents itself as syncing prices in real time and showing only available models, usable at no cost within platform rate limits. Rate limits for the free models are fixed values rather than tiered by spending, and free model calls show as a cost of 0 in the account bill. Specific per-model rate limits are published on the model pages. The roster is dynamically maintained and models can be added or removed; the provider states it retains interpretation rights over which models are free.
Models disclosedFree ('免费') chat models on the CN pricing page on 2026-10-07, several of them behind the page's expand buttons: THUDM/GLM-Z1-9B-0414, THUDM/GLM-4-9B-0414, deepseek-ai/DeepSeek-R1-0528-Qwen3-8B, Qwen/Qwen3.5-4B, Qwen/Qwen3-8B, Qwen/Qwen2.5-7B-Instruct, and XingChenAGI/Xing4.0-29B from ChinaTelecom, which the model centre lists with a release date of 2026-09-17. Also free: tencent/Hunyuan-MT-7B (translation), PaddlePaddle/PaddleOCR-VL-1.5 and deepseek-ai/DeepSeek-OCR (document OCR), BAAI bge embeddings and rerankers, and several ASR and image models. No Baidu ERNIE chat model is free. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | 1 requests / min | Signup, payment unknown | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessUnoRouter marks part of its catalog free: a model carrying the :free suffix routes only to free providers and never uses an account balance, while the same model name without the suffix is the paid version billed per token. Access requires a UnoRouter account and an API key. UnoRouter's own cap is 1 request per minute, per free model, per user, layered over the upstream provider's own cap, daily token budgets on some free pools that reset at midnight UTC, tokens-per-minute caps triggered by very large prompts, and a per-user concurrency limit on parallel requests. A rate-limited route returns 429 with a Retry-After header, and a model whose free providers are all busy returns 503, which usually clears in minutes. UnoRouter's models, pricing and rate limit pages do not state whether a credit card is required. There is no announced expiry, and availability is best effort: a model leaves the catalog only when every channel is out, and it returns within minutes once a channel recovers.
Models disclosedFree routes carry a :free suffix in UnoRouter's own catalog, and the set changes as free pools drain or upstream providers change. On 2026-09-28 the newest catalog listings included space-bunny-alpha:free, mimo-v2.6-flash:free, mimo-v2.6-pro:free, axon-1.8-pro:free, axon-1.8-flash:free, axon-1.8-lightning:free, atria-dawn-preview:free, ling-3.0-flash-vl:free, nex-n2.5-mini:free and k2-horizon:free. Not all free models support tool calls or vision, so check the capability badges on the model page. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | Not stated | Not stated | Signup, payment unknown | Varies by route | Standing | verified | Get a key ↗ | |
Program detailsFree accessStanding model-specific zero price: input, cached input, cached-input storage, and output are all listed as Free for GLM-4.7-Flash and GLM-4.6V-Flash (and the legacy GLM-4.5-Flash id, which BigModel says now routes to GLM-4.7-Flash). Numeric request/token rate limits are not public on the pricing page and must be checked on the authenticated account rate-limits surface.
Models disclosedGLM-4.7-Flash is the current free text model and GLM-4.6V-Flash the current free vision-language model on the international pricing page. That page still prints GLM-4.5-Flash at Free on every price dimension, but the China-region BigModel page for GLM-4.5-Flash states the model was retired on 2026-01-30 with requests auto-routed to GLM-4.7-Flash, so treat GLM-4.5-Flash as a legacy id rather than a second free model. Other models' 'Limited-time Free' cached-input storage is not free inference and is not included in this offer. China-region BigModel additionally lists GLM-4-Flash-250414, GLM-4.1V-Thinking-Flash, GLM-4V-Flash, CogView-3-Flash, and CogVideoX-Flash alongside overlapping GLM Flash models. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Unlimitedper model | Not stated | Signup, payment unknown | Varies by route | Standing | verified | Get a key ↗ | |
Program detailsFree accessZenMux lists z-ai/glm-4.6v-flash-free and z-ai/glm-4.7-flash-free as active routes with a $0 effective price on both input and output, each marked Free and Rate Limit on 2026-09-28. Access requires a ZenMux account and an API key, created under either the Pay As You Go plan or the paid Builder Plan subscription. The public pages do not state a card requirement or a numeric free-route quota. ZenMux documents that an account must keep a non-negative balance even for free tiers, while some models can require a positive balance. Prior free routes have sunset without becoming permanent, so the lineup is dynamic.
Models disclosedOn 2026-09-23 the free lineup is two routes, both showing Active on their ZenMux model pages: z-ai/glm-4.6v-flash-free, described as the free version of GLM-4.6V, and z-ai/glm-4.7-flash-free, a 30B class model tuned for agentic coding. Each route is served by the Z.ai and BigModel providers at an effective $0 per million tokens for both input and output. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | Not stated | Signup, payment unknown | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessFree NVIDIA-hosted NIM API endpoints for prototyping, research, development, testing, learning, and evaluation through free NVIDIA Developer Program membership. Production use is excluded. NVIDIA's NIM FAQ says the API catalog is a trial experience, that the rate of requests varies per model and with concurrent users, and that the exact limits are shown in the authenticated account; no public page states a number (the build.nvidia.com FAQ that once said 40 requests per minute now returns 404). Developer Program membership also licenses self-hosted NIM microservices for non-production use on up to 16 GPUs.
Models disclosedOffer-wide hosted API coverage: 160+ leading AI models according to NVIDIA's current NIM FAQ. Self-hosted coverage is broader and is a separate benefit; NVIDIA's current NIM product page describes thousands of models and customizations supported through NIM runtimes. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | 20 requests / min | Signup, payment unknown | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessAgnes AI's pricing page lists agnes-2.5-flash and agnes-3.0-flash at $0 per 1M input, output and cached-input tokens. A new user can create an account and API key without subscribing, then call the hosted OpenAI-compatible API. Free/default text access is rate-limited to 10 requests per minute, after the provider cut free and enterprise text rate limits by 50 percent on 2026-09-23. The pricing page says offers can vary by account eligibility or promotion stage and that the account bill is authoritative.
Models disclosedOn 2026-10-04 the text model pricing table lists agnes-2.5-flash and agnes-3.0-flash at $0 input, $0 output and $0 cached input, and the page states that cached input, input tokens and output tokens are currently free for both models. agnes-2.5-pro is billed at its published Pro-tier prices, and agnes-3.0-pro is marked coming soon at those same Pro-tier prices. agnes-2.0-flash is not listed on the pricing page. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | 2 requests / min | No signup | Does not train | Standing | verified | Get a key ↗ | |
Program detailsFree accessAnonymous AI Endpoints access is available without an API key at 2 requests per minute per IP per model. Authenticated use is a separate billed path with a 400-requests-per-minute per-project per-model limit and requires an OVHcloud Public Cloud project with a payment method. Playground LLM responses are limited to 1,024 output tokens for testing.
Models disclosedThe anonymous limit is documented offer-wide, per model, across the current AI Endpoints catalog. On 2026-10-09 the public catalog listed 20 models, including Llama, Qwen, GPT-OSS, Qwen3Guard moderation, embedding, image-generation, speech-to-text and text-to-speech models. Models announced for retirement remain reachable through the API for a three-month grace period but are hidden from the catalog. On 2026-10-09 OVHcloud listed Mistral-7B-Instruct-v0.3, Mistral-Nemo-Instruct-2407 and Mistral-Small-3.2-24B-Instruct-2506 as being decommissioned. The authenticated catalog also prices some entries at Free outright: on 2026-10-09 these were the beta Qwen3Guard-Gen 0.6B and 8B moderation models, stable-diffusion-xl-base-v10, and four NVIDIA Riva text-to-speech voices. Qwen3.8-27B is paid, at 0.4 EUR per 1M input tokens and 2.7 EUR per 1M output tokens. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | Unlimitedshared pool | Not stated | No signup | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessAI Horde is a standing free, community-hosted text and image inference service backed by volunteer compute. It permits anonymous API use with the public key 0000000000 and requires no registration, credit card, or payment. Anonymous jobs receive the lowest queue priority, and the provider states that priority can only be earned by contributing compute and can never be bought or sold. Capacity and model availability depend on the volunteer workers connected at the time. An OpenAI-compatible endpoint at oai.aihorde.net also accepts the anonymous key at the lowest priority, and the provider describes that endpoint as a pilot that might be restricted or expanded in the future.
Models disclosedDynamic text-model catalog supplied by volunteer workers; availability depends on currently connected workers. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | 20K tokens / day | 15 requests / min | Signup, no payment | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree access15 RPM, 20K TPM, and 20K tokens/day on the default Free tier.
Models disclosedAion Labs hosted text models listed on its model catalog. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | 1,000 requests / month | 20 requests / min | Signup, payment unknown | Trains, opt-out | Standing | verified | Get a key ↗ | |
Program detailsFree accessStanding evaluation/trial API keys are free but limited. Trial keys are capped at 1,000 API calls per month. Chat models are limited to 20 requests per minute each. Other endpoint limits are Audio Transcriptions 5 RPM, Embed 2,000 inputs per minute, Embed Images 5 inputs per minute, EmbedJob 5 RPM, Rerank 10 RPM, Parse 500 RPM, Tokenize 100 RPM, and default endpoints 500 RPM. The documentation states a recurring monthly limit and gives no trial expiration date or one-time credit balance.
Models disclosedOffer-wide FAQ says trial keys can test all Cohere models and APIs. The current model-specific Chat table lists Command A+, Command A Reasoning, Command A Translate, Command A Vision, Command A, Command R+, Command R, Command R7B, and North Mini Code at 20 RPM. Other APIs include Audio Transcriptions, Embed, Embed Images, EmbedJob, Rerank, Tokenize, and endpoints covered by the default limit. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | Not stated | Not stated | Signup, no payment | Trains on it | Not stated | verified | Get a key ↗ | |
Program detailsFree accessThe Nova Terms of Use (last updated December 2, 2025) contain no billing or fee language, and the Amazon Nova Act user guide states that API key authentication through nova.amazon.com/act is free of charge with daily limits, while the separate AWS IAM route bills usage to an AWS account. API keys are issued from the developer surface (Terms 1.5), usage may be subject to rate limits or other restrictions at Amazon's discretion (Terms 1.3), and the Build with Nova schedule states the offering is intended for experimentation and does not support deploying production scale developer applications (S-1.1). No numeric daily limits are published publicly.
Models disclosedAmazon Nova family on the nova.amazon.com developer surface; page metadata names Amazon Nova 2 Lite, Nova Act, and Nova 2 Sonic. On 2026-10-09 the Amazon Nova Act user guide confirmed that Nova Act can be used with an API key from nova.amazon.com/act. The full per-model roster is behind Amazon sign-in and is not publicly confirmable. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | 1 concurrent requests12K-token context cap on the free route | Signup, payment unknown | Does not train | Standing | verified | Get a key ↗ | |
Program detailsFree accessArli AI lists a $0 Free plan with delayed responses that get slower with more requests, a maximum 12K-token context, one request at a time, no image-generation tools, and the exact quota wording "5 times / 2 days trial of all models." The wording does not clearly state whether this means five uses every two days or a one-time two-day trial.
Models disclosedAll text-generation models in Arli AI's current catalog are covered by the free trial allowance. Image-generation models and tools are excluded from the Free plan. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | Not stated | Not stated | Signup, no payment | Trains on it | Standing | verified | Get a key ↗ | |
Program detailsFree accessDatabricks Free Edition is a no-cost, serverless-only, quota-limited account for personal, academic, learning, experimentation, and other non-commercial or not-for-profit use. It has daily fair-use limits; exceeding quota can shut down compute for the rest of the day and, in extreme cases, the rest of the month. Model serving has a limited number of active endpoints, no GPU serving endpoints, no provisioned throughput, no custom models on GPU or batch inference, and some models are unavailable.
Models disclosedThe Free Edition supports experimenting with foundation models and deploying AI systems, but Databricks does not publish a Free-Edition-specific list of included foundation models. The limitation page explicitly says certain models are unavailable. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | Not stated | Signup, no payment | Trains on it | Standing | verified | Get a key ↗ | |
Program detailsFree accessStanding Gemini Developer API Free tier: standard input and output are marked free of charge for eligible models, subject to project-level RPM, TPM, RPD, and sometimes model-specific IPM or TPD limits. RPD resets at midnight Pacific time. Google no longer publishes one guaranteed numeric quota table: active limits vary by model, project tier, and account status and must be viewed in AI Studio; specified limits are not guaranteed. Google AI Studio usage is free of charge in available regions.
Models disclosedOn 2026-10-07 the Gemini Developer API pricing page marks standard Free tier input and output as free of charge for Gemini 3.8 Flash, Gemini 3.7 Flash, Gemini 3.6 Flash, Gemini 3.5 Flash, Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking, Gemini 3.1 Flash Live Preview, Gemini 3.5 Live Translate, Gemini 3.5 Transcribe Live, Gemini 3.5 Transcribe, Gemini 3.5 Flash-Lite, Gemini 3.1 Flash-Lite, Gemini 3.8 Flash TTS, Gemini 3.8 Flash-Lite TTS, Gemini 3.1 Flash TTS Preview, Gemini 3 Flash Preview, Gemini 2.5 Pro, Gemini 2.5 Flash, Gemini 2.5 Flash-Lite, Gemini 2.5 Flash Native Audio, Gemini 2.5 Flash Preview TTS, Gemini Embedding 2, Gemini Robotics ER 2 Preview, Gemini Robotics ER 2 Streaming Preview, and Gemma 4. The pricing page does not list Gemini 2.0 Flash, Gemini 2.0 Flash-Lite, Gemini Embedding, Gemini Robotics ER 1.6 Preview, or Gemini 2.5 Flash-Lite Preview; those names appear only in the Batch enqueued tokens table on the rate limits page, which covers paid tiers. Free access is specific to both the model and the serving mode: Gemini 3.1 Pro Preview, Gemini Omni Flash, the Nano Banana image models, Veo 3.1, and the Lyria music models show Free tier 'Not available,' and the batch, flex, and priority tables are free of charge for some models and not available for others. Signup, payment, data
What the provider's page says
| ||||||||
| openai/gpt-oss-120b, openai/gpt-oss-20b, openai/gpt-oss-safeguard-20b +1 | 200K tokens / dayper model | 30 requests / min | Signup, no payment | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessStanding organization-level Free Plan with no usage charges and model-specific recurring rate limits. Free Plan table on 2026-10-06: openai/gpt-oss-120b, openai/gpt-oss-20b, and qwen/qwen3.8-27b 30 RPM, 1K RPD, 8K TPM, 200K TPD; openai/gpt-oss-safeguard-20b 3 RPM, 1K RPD, 2K TPM, 200K TPD; meta-llama/llama-prompt-guard-2-22m and meta-llama/llama-prompt-guard-2-86m 30 RPM, 14.4K RPD, 15K TPM, 500K TPD; canopylabs/orpheus-arabic-saudi and canopylabs/orpheus-v1-english 10 RPM, 100 RPD, 1.2K TPM, 3.6K TPD; whisper-large-v3 and whisper-large-v3-turbo 20 RPM, 2K RPD, 7.2K audio seconds per hour, 28.8K audio seconds per day. Rate limits apply at the organization level rather than to individual users, and cached tokens do not count toward them. The page says the exact current limits are on the authenticated account Limits page; a valid payment method is needed only to upgrade to the Developer plan.
Models disclosedFree Plan rate-limit table as of 2026-10-06: openai/gpt-oss-120b, openai/gpt-oss-20b, openai/gpt-oss-safeguard-20b, qwen/qwen3.8-27b, meta-llama/llama-prompt-guard-2-22m, meta-llama/llama-prompt-guard-2-86m, canopylabs/orpheus-arabic-saudi, canopylabs/orpheus-v1-english, whisper-large-v3, and whisper-large-v3-turbo. Removed from the table since 2026-09-01: groq/compound, groq/compound-mini, and qwen/qwen3.6-27b. Removed from the table earlier, since 2026-08-25: llama-3.1-8b-instant, llama-3.3-70b-versatile, llama-4-scout, and llama-guard-4-12b. Signup, payment, data
What the provider's page says
| ||||||||
| Meta-Llama-3.1-8B-Instruct, Meta-Llama-3-8B-Instruct, Awanllm-Llama-3-8B-Dolfin +1 | Unlimited (rate-limited) | 20 requests / min | Signup, payment unknown | Does not train | Standing | verified | Get a key ↗ | |
Program detailsFree accessAwan LLM's Lite plan is labeled Free forever and includes unlimited tokens subject to model context limits, 20 requests per minute, 200 requests per day for small models, 10 requests per day for medium models, and 10 requests per day for large models.
Models disclosedCurrent small-model catalog: Meta-Llama-3.1-8B-Instruct, Meta-Llama-3-8B-Instruct, Awanllm-Llama-3-8B-Dolfin, and Awanllm-Llama-3-8B-Cumulus. Current large-model catalog: Meta-Llama-3.1-70B-Instruct and Meta-Llama-3-70B-Instruct. The pricing page also assigns 10 requests per day to a medium-model class, but the current models page does not display a medium-model entry. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | 1M tokens / day | 40 requests / min | Signup, no card | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessFree token access allows 1 request per second, 60 requests per minute, and 250 requests per hour on text generation, with 100,000 tokens per 24 hours counted as input tokens plus output tokens. The provider states that free-token quotas are provided at no charge and may be reduced without notice.
Models disclosedCurrent turbo-tier models exposed by the live Models API. Signup, payment, data
What the provider's page says | ||||||||
| not itemized, see details | 200 requests / month | Not stated | Signup, payment unknown | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessMorph documents a recurring $0 free tier of 200 API requests per month. On 2026-10-02 the pricing page states that each API call to Fast Apply, WarpGrep, Compact, Reflex, Embeddings, or Rerank counts as one request against that 200 per month cap, and that once the cap is exhausted, per-token or plan-credit pricing applies. The same page lists a Free subscription plan at $0 per month carrying 250,000 credits with low rate limits, where one credit is about $0.00001 and credits apply to Fast Apply, WarpGrep, Compact, Reflex, Embeddings, and Rerank. Credits do not roll over: they reset monthly on the subscription date. Morph's hosted chat models are listed with per-token prices rather than inside the free request allowance.
Models disclosedOn 2026-10-02 the free tier covers 200 API requests per month across Fast Apply (morph-v3-fast and morph-v3-large), WarpGrep, Compact, Reflex, Embeddings (morph-embedding-v4), and Rerank (morph-rerank-v4), plus 250,000 monthly credits on the $0 Free plan that apply to those same products. Morph's hosted chat models, including Kimi K3, GLM-5.3, GLM-5.3-Flash, DeepSeek V4.1 Flash, DeepSeek V4 Flash 0731, and Qwen 3.8, are priced per token. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | 15 requests / min | Signup, payment unknown | Not stated | Not stated | verified | Get a key ↗ | |
Program detailsFree accessRelace offers a hosted API free tier for testing its code-focused models. Signup is required. The API reference states that rate limits are a per-key request budget over a rolling 60 second window, that free accounts start at 15 requests per minute, and that the budget scales with account tier. No free token or request quota, renewal cadence, or expiration date is published.
Models disclosedOn 2026-10-01 the pricing page lists four Relace models with per-token prices: relace-search, relace-compact, relace-apply-3, and relace-rank. The API reference also documents Open Models endpoints for chat completions, messages, and listing models. Provider material does not state which of these models the free tier covers. Signup, payment, data
What the provider's page says
| ||||||||
| deepseek-v4-flash, deepseek-v4-flash-0731, glm-4.7-flash +5 | 50K tokens / month | 60 requests / min | Signup, no payment | Not stated | Standing | due soon | Get a key ↗ | |
Program detailsFree accessTerms of Service section 5 and the pricing page both state eligible accounts receive 50,000 free tokens per month, refreshed monthly, intended for developer evaluation and small-scale testing. Free-tier rate limit is 60 requests/min (paid tier 180). Strictly one free-tier account per human or organization; disposable-email domains prohibited. Reverification due soon; last checked 2026-10-04.
Models disclosedOn 2026-10-04 the pricing table marks eight models Eligible for the 50,000 monthly free tokens: deepseek-v4-flash, deepseek-v4-flash-0731, glm-4.7-flash, gpt-oss-120b, gpt-oss-20b, laguna-s-2.1, laguna-xs-2.1, and qwen3.5-9b. The other named models are marked Standard and are paid only: claude-opus-5, claude-sonnet-5, deepseek-v4-pro, glm-5.2, gpt-5.6-terra, gpt-5.6-sol, grok-4.5, kimi-k2.7-code, kimi-k3, and minimax-m2.7. The auto smart router is listed separately with dynamic pricing and a free tier quota that varies by the target model it selects, so an auto request routed to an eligible model can draw on the free allowance. The pricing page labels the catalog dynamic, so the roster can change. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | Not stated | 10 requests / min | Signup, payment unknown | Not stated | Standing | verified | Get a key ↗ | |
Program detailsFree accessSEA-LION advertises a free API for prototyping and proof-of-concept development, rate-limited to 10 calls per minute per user. Users create one Trial API Key through the Playground. The provider says the free APIs are not designed for production or high-throughput workloads.
Models disclosedThe API is offer-wide across the SEA-LION models returned for the user's key. Current documentation explicitly demonstrates SEA-LION v4.5 for instruct and function-calling use, SEA-LION v3.5 reasoning models, aisingapore/SEA-Guard, and SEA-LION-ModernBERT-Embedding-600M. The exact authenticated deployment list must be queried with the generated key and is not enumerated on the public offer page. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | 300K tokens / month | 2 requests / sec | Card on file | Does not train | Standing | verified | Get a key ↗ | |
Program detailsFree accessStanding watsonx.ai Runtime Lite plan with 300,000 foundation-model inference tokens per month, 20 CUH per month for supported machine-learning workloads, 100 pages per month of document text classification and extraction, and a plan-level limit of 2 inference requests per second. Custom foundation models are not available on Lite.
Models disclosedReady-to-use IBM and third-party foundation models available to the Lite instance, subject to regional and current catalog availability. IBM's supported-model surface describes IBM and third-party models; custom foundation models are not available on Lite. An exact immutable Lite model roster is not stated. Signup, payment, data
What the provider's page says
| ||||||||
| not itemized, see details | 10K neurons / day | Not stated | Signup, payment unknown | Does not train | Standing | verified | Get a key ↗ | |
Program detailsFree accessWorkers Free includes 10,000 Neurons per account per day at no charge. The allocation resets daily at 00:00 UTC. Further operations fail after the limit on Workers Free; using more requires Workers Paid, where usage above the same free allocation is billed at $0.011 per 1,000 Neurons.
Models disclosedThe free Neuron allocation is offer-wide across the Workers AI catalog except models that explicitly require paid billing. The current exclusion list is @cf/moonshotai/kimi-k2.6, @cf/moonshotai/kimi-k2.7-code, @cf/zai-org/glm-5.2, @cf/zai-org/glm-5.3, @cf/zai-org/glm-5.3-flash, @cf/deepseek-ai/deepseek-v4-flash-0731, and @cf/deepseek-ai/deepseek-v4-pro-0813 (as of 2026-09-04); those models can be reached with Workers Paid or prepaid AI Gateway credits. Model-specific Neuron consumption varies by model and input/output type. Signup, payment, data
What the provider's page says | ||||||||
| Not free, so you can stop looking. Well-known providers whose "free" is a signup credit or needs a paid balance. Shown for the record; never counted. | ||||||||
Cerebrasgpt-oss-120b, gemma-4-31bNot counted |
Credit, not free access New accounts receive a one-time $5 credit grant only after adding a verified payment method. | Not free | Evidence ↗ | |||||
FireworksAny eligible Fireworks modelNot counted |
Credit, not free access Excluded credit product: Fireworks serverless inference is pay per token/postpaid billing. | Not free | Evidence ↗ | |||||
AnthropicHosted API accessNot counted |
Credit, not free access Anthropic limits subscription OAuth to ordinary use of Claude Code and native Anthropic applications. | Not free | Evidence ↗ | |||||
DeepSeekHosted API accessNot counted |
Credit, not free access DeepSeek's current official API pricing is pay-per-token, and no standing free model, recurring inference quota, or zero-cost API tier is documented. | Not free | Evidence ↗ | |||||
xAIHosted API accessNot counted |
Paid prerequisite Hermes documents a no-key OAuth route that spends SuperGrok or X Premium+ entitlement at api.x.ai. | Not free | Evidence ↗ | |||||
No verified program matches those filters today. Free access changes weekly; follow @LLMPriceIndex to hear when that changes.
Want to know when one of these changes? A weekly digest of ended programs, new payment requirements and major new programs is coming; signup opens when delivery is proven.
Compare selected
Up to three programs. Unknown terms remain visible.
Select programs with the square control in the program list.
How a row earns its cells
Each program shows its evidence and verification date. Every number in the table is backed by a quote from the provider's own page, re-read nightly; a cell with no quote says "Not stated" rather than guessing. A program enters a filter only when the provider's published terms confirm it qualifies. Credits, paid prerequisites, and unresolved offers stay outside the count. Every evidence page is re-read nightly; a program that goes more than a week without renewed evidence leaves the filters, and more than two weeks leaves the headline count, while staying listed with its warning.
Use the data. free-access.json is the current inventory (every program, its structured access facts, terms, evidence URLs, verification date, and the OpenRouter child models). free-access-events.json is the confirmed change feed, and free-access-events.xml is the same feed as Atom for a reader or a Slack RSS app. All are rebuilt daily, are free to use with attribution, and carry the same fields this page shows.