Cognaptus DataHub Monitor

Free AI Inference Providers

A daily dashboard for monitoring free AI inference providers, with curated vendor boards and a machine-refreshable OpenRouter free-model roster.

Updated 2026-09-11 14:43:03 +0800 Source: Openrouter

Providers Tracked

24

OpenRouter Free Models

22

Provider Families

10

Multimodal Free Models

13

Coding-Friendly Free Models

1

Capability Signals

Text 22 / Image 13 / Audio 5 / Video 4

Why This Page Exists

Free inference capacity is fragmented. Stable vendor free tiers, rotating cloud quotas, and zero-cost OpenRouter routes need different monitoring logic.

This page is now structured as a monitor rather than a one-off article.

The daily job should separate three layers:

  1. Direct model vendors with relatively stable free quotas.
  2. Inference clouds where free capacity rotates more often.
  3. OpenRouter zero-cost models, which should be refreshed automatically from inferencer.

The goal is not just to list providers. It is to help you answer a daily operating question: which free routes are realistically usable right now for text, coding, multimodal, image, speech, and video workloads?

Direct foundation-model vendors

Companies that operate their own API and offer some free or free-trial inference access.

Google AI Studio

Most broadly usable direct free tier.

ACTIVE Free daily quota Text, image, audio, video

GroqCloud

Fastest developer-facing free path for open models.

ACTIVE Free inference tier Text, vision, speech

Cerebras Cloud

Useful as a high-throughput backup route.

ACTIVE Free quota Text

Cohere

Useful for business NLP tooling.

WATCH Developer quota Text, embeddings, rerank

Mistral

Strong open-weight ecosystem.

WATCH Trial credits / limited free access Text, embeddings

DeepSeek

Monitor for changes in API trial policy.

WATCH Free web usage and low-cost API Text, reasoning

MiniMax

Relevant for agent and media workflows.

WATCH Selective free access Text, speech, video

Moonshot / Kimi

Worth tracking for long-context offerings.

WATCH Selective free access Text

Inference clouds and gateways

Multi-model platforms where free capacity changes often and is worth checking daily.

OpenRouter

Primary daily monitor source via inferencer.

ACTIVE Zero-cost models rotate daily Text, vision, audio, video

Hugging Face Inference

Largest long-tail open-model surface.

ACTIVE Free usage quota Text, image, speech, embeddings

Together AI

Good backup for open-weight text and diffusion.

WATCH Free quota / credits Text, image

Cloudflare Workers AI

Edge inference is strategically distinct.

WATCH Bundled free allocation Text, image, speech

Fireworks AI

Useful for comparing open-model economics.

WATCH Credits / selected free access Text, image

Baseten

More deployment-oriented than general free inference.

WATCH Credits Model hosting

Replicate

Best tracked for specialized modalities.

WATCH Credits Image, video, audio

Fal.ai

Media generation prices move quickly.

WATCH Credits / trials Image, video

Modal

More infra than gateway, but still relevant.

WATCH Credits Model hosting

Specialized modality APIs

Providers best monitored by capability rather than by general LLM coverage.

ElevenLabs

Important benchmark for TTS.

WATCH Starter quota Speech

Deepgram

Useful ASR baseline.

WATCH Trial / starter credits Speech

AssemblyAI

Speech-first provider.

WATCH Credits Speech

Stability AI

Core diffusion benchmark.

WATCH Credits Image

Black Forest Labs

Track FLUX availability changes.

WATCH Testing access Image

Runway

Consumer-friendly video benchmark.

WATCH Credits Video

Luma AI

Worth tracking for motion quality.

WATCH Credits Video

OpenRouter Zero-Cost Roster

This block is designed for daily refresh from `inferencer::list_openrouter_models()`.

cohere 1 free models

cohere/north-mini-code:free

North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...

Context 256000 texttext->text
dots-studio 1 free models

dots-studio/dots-3-note-preview:free

Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the Dots 3 family and is...

Context 512000 textimagetext+image->text
google 4 free models

google/gemma-4-26b-a4b-it:free

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference - delivering near-31B quality at...

Context 262144 imagetextvideotext+image+video->text

google/gemma-4-31b-it:free

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Context 262144 imagetextvideotext+image+video->text

google/lyria-3-clip-preview

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...

Context 1.048576e+06 textimageaudiotext+image->text+audio

google/lyria-3-pro-preview

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...

Context 1.048576e+06 textimageaudiotext+image->text+audio
inclusionai 3 free models

inclusionai/ling-3.0-flash-fin:free

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

Context 262144 texttext->text

inclusionai/ling-3.0-flash-sante:free

Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for...

Context 262144 texttext->text

inclusionai/ling-3.0-flash-vl:free

Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...

Context 262144 textimagevideotext+image+video->text
liquid 1 free models

liquid/lfm-2.5-2.6b:free

LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises against using it for agentic coding or...

Context 65536 texttext->text
nex-agi 2 free models

nex-agi/nex-n2.5-mini:free

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

Context 262144 textimagetext+image->text

nex-agi/nex-n2.5-pro:free

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...

Context 262144 textimagetext+image->text
nvidia 5 free models

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free

NVIDIA NemotronTM 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

Context 256000 textaudioimagevideotext+image+audio+video->text

nvidia/nemotron-3-super-120b-a12b:free

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

Context 262144 texttext->text

nvidia/nemotron-3-ultra-550b-a55b:free

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Context 1e+06 texttext->text

nvidia/nemotron-3.5-content-safety:free

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

Context 128000 textimagetext+image->text

nvidia/nemotron-3.5-lightning:free

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

Context 1e+06 texttext->text
openrouter 1 free models

openrouter/free

The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that...

Context 200000 textimagetext+image->text
poolside 2 free models

poolside/laguna-s-2.1:free

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

Context 262144 texttext->text

poolside/laguna-xs-2.1:free

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

Context 262144 texttext->text
thinkingmachines 2 free models

thinkingmachines/inkling-small:free

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

Context 1.048576e+06 textimageaudiotext+image+audio->text

thinkingmachines/inkling:free

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Context 1.048576e+06 textimageaudiotext+image+audio->text

Operator Notes

This page should be refreshed daily because free-model availability changes quickly.

Category-aware free-model counts are derived from inferencer's OpenRouter category wrappers, while the provider boards remain manually curated.

The OpenRouter section is machine-generated; the provider boards are intentionally curated.

Use the dashboard as a monitor, not as a compliance promise. Vendor free-tier terms can change without warning.