Providers Tracked
24
Cognaptus DataHub Monitor
A daily dashboard for monitoring free AI inference providers, with curated vendor boards and a machine-refreshable OpenRouter free-model roster.
Providers Tracked
24
OpenRouter Free Models
22
Provider Families
10
Multimodal Free Models
13
Coding-Friendly Free Models
1
Capability Signals
Text 22 / Image 13 / Audio 5 / Video 4
Free inference capacity is fragmented. Stable vendor free tiers, rotating cloud quotas, and zero-cost OpenRouter routes need different monitoring logic.
This page is now structured as a monitor rather than a one-off article.
The daily job should separate three layers:
inferencer.The goal is not just to list providers. It is to help you answer a daily operating question: which free routes are realistically usable right now for text, coding, multimodal, image, speech, and video workloads?
Companies that operate their own API and offer some free or free-trial inference access.
Google AI Studio
Most broadly usable direct free tier.
GroqCloud
Fastest developer-facing free path for open models.
Cerebras Cloud
Useful as a high-throughput backup route.
Cohere
Useful for business NLP tooling.
Mistral
Strong open-weight ecosystem.
DeepSeek
Monitor for changes in API trial policy.
MiniMax
Relevant for agent and media workflows.
Moonshot / Kimi
Worth tracking for long-context offerings.
Multi-model platforms where free capacity changes often and is worth checking daily.
OpenRouter
Primary daily monitor source via inferencer.
Hugging Face Inference
Largest long-tail open-model surface.
Together AI
Good backup for open-weight text and diffusion.
Cloudflare Workers AI
Edge inference is strategically distinct.
Fireworks AI
Useful for comparing open-model economics.
Baseten
More deployment-oriented than general free inference.
Replicate
Best tracked for specialized modalities.
Fal.ai
Media generation prices move quickly.
Modal
More infra than gateway, but still relevant.
Providers best monitored by capability rather than by general LLM coverage.
ElevenLabs
Important benchmark for TTS.
Deepgram
Useful ASR baseline.
AssemblyAI
Speech-first provider.
Stability AI
Core diffusion benchmark.
Black Forest Labs
Track FLUX availability changes.
Runway
Consumer-friendly video benchmark.
Luma AI
Worth tracking for motion quality.
This block is designed for daily refresh from `inferencer::list_openrouter_models()`.
cohere/north-mini-code:free
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
dots-studio/dots-3-note-preview:free
Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is...
google/gemma-4-26b-a4b-it:free
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference - delivering near-31B quality at...
google/gemma-4-31b-it:free
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
google/lyria-3-clip-preview
30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...
google/lyria-3-pro-preview
Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...
inclusionai/ling-3.0-flash-fin:free
Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...
inclusionai/ling-3.0-flash-sante:free
Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for...
inclusionai/ling-3.0-flash-vl:free
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
liquid/lfm-2.5-2.6b:free
LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises against using it for agentic coding or...
nex-agi/nex-n2.5-mini:free
Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...
nex-agi/nex-n2.5-pro:free
Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free
NVIDIA NemotronTM 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
nvidia/nemotron-3-super-120b-a12b:free
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
nvidia/nemotron-3-ultra-550b-a55b:free
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
nvidia/nemotron-3.5-content-safety:free
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...
nvidia/nemotron-3.5-lightning:free
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
openrouter/free
The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that...
poolside/laguna-s-2.1:free
Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...
poolside/laguna-xs-2.1:free
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...
thinkingmachines/inkling-small:free
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...
thinkingmachines/inkling:free
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
This page should be refreshed daily because free-model availability changes quickly.
Category-aware free-model counts are derived from inferencer's OpenRouter category wrappers, while the provider boards remain manually curated.
The OpenRouter section is machine-generated; the provider boards are intentionally curated.
Use the dashboard as a monitor, not as a compliance promise. Vendor free-tier terms can change without warning.