Cognaptus DataHub Monitor

Free AI Inference Providers

A daily dashboard for monitoring free AI inference providers, with curated vendor boards and a machine-refreshable OpenRouter free-model roster.

Updated 2026-07-29 09:18:44 +0800 Source: Openrouter

Providers Tracked

24

OpenRouter Free Models

17

Provider Families

7

Multimodal Free Models

8

Coding-Friendly Free Models

2

Capability Signals

Text 17 / Image 8 / Audio 3 / Video 4

Why This Page Exists

Free inference capacity is fragmented. Stable vendor free tiers, rotating cloud quotas, and zero-cost OpenRouter routes need different monitoring logic.

This page is now structured as a monitor rather than a one-off article.

The daily job should separate three layers:

  1. Direct model vendors with relatively stable free quotas.
  2. Inference clouds where free capacity rotates more often.
  3. OpenRouter zero-cost models, which should be refreshed automatically from inferencer.

The goal is not just to list providers. It is to help you answer a daily operating question: which free routes are realistically usable right now for text, coding, multimodal, image, speech, and video workloads?

Direct foundation-model vendors

Companies that operate their own API and offer some free or free-trial inference access.

Google AI Studio

Most broadly usable direct free tier.

ACTIVE Free daily quota Text, image, audio, video

GroqCloud

Fastest developer-facing free path for open models.

ACTIVE Free inference tier Text, vision, speech

Cerebras Cloud

Useful as a high-throughput backup route.

ACTIVE Free quota Text

Cohere

Useful for business NLP tooling.

WATCH Developer quota Text, embeddings, rerank

Mistral

Strong open-weight ecosystem.

WATCH Trial credits / limited free access Text, embeddings

DeepSeek

Monitor for changes in API trial policy.

WATCH Free web usage and low-cost API Text, reasoning

MiniMax

Relevant for agent and media workflows.

WATCH Selective free access Text, speech, video

Moonshot / Kimi

Worth tracking for long-context offerings.

WATCH Selective free access Text

Inference clouds and gateways

Multi-model platforms where free capacity changes often and is worth checking daily.

OpenRouter

Primary daily monitor source via inferencer.

ACTIVE Zero-cost models rotate daily Text, vision, audio, video

Hugging Face Inference

Largest long-tail open-model surface.

ACTIVE Free usage quota Text, image, speech, embeddings

Together AI

Good backup for open-weight text and diffusion.

WATCH Free quota / credits Text, image

Cloudflare Workers AI

Edge inference is strategically distinct.

WATCH Bundled free allocation Text, image, speech

Fireworks AI

Useful for comparing open-model economics.

WATCH Credits / selected free access Text, image

Baseten

More deployment-oriented than general free inference.

WATCH Credits Model hosting

Replicate

Best tracked for specialized modalities.

WATCH Credits Image, video, audio

Fal.ai

Media generation prices move quickly.

WATCH Credits / trials Image, video

Modal

More infra than gateway, but still relevant.

WATCH Credits Model hosting

Specialized modality APIs

Providers best monitored by capability rather than by general LLM coverage.

ElevenLabs

Important benchmark for TTS.

WATCH Starter quota Speech

Deepgram

Useful ASR baseline.

WATCH Trial / starter credits Speech

AssemblyAI

Speech-first provider.

WATCH Credits Speech

Stability AI

Core diffusion benchmark.

WATCH Credits Image

Black Forest Labs

Track FLUX availability changes.

WATCH Testing access Image

Runway

Consumer-friendly video benchmark.

WATCH Credits Video

Luma AI

Worth tracking for motion quality.

WATCH Credits Video

OpenRouter Zero-Cost Roster

This block is designed for daily refresh from `inferencer::list_openrouter_models()`.

cohere 1 free models

cohere/north-mini-code:free

North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...

Context 256000 texttext->text
google 4 free models

google/gemma-4-26b-a4b-it:free

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference - delivering near-31B quality at...

Context 262144 imagetextvideotext+image+video->text

google/gemma-4-31b-it:free

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Context 262144 imagetextvideotext+image+video->text

google/lyria-3-clip-preview

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...

Context 1.048576e+06 textimageaudiotext+image->text+audio

google/lyria-3-pro-preview

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...

Context 1.048576e+06 textimageaudiotext+image->text+audio
inclusionai 1 free models

inclusionai/ling-3.0-flash:free

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

Context 262144 texttext->text
nvidia 7 free models

nvidia/nemotron-3-nano-30b-a3b:free

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

Context 256000 texttext->text

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free

NVIDIA NemotronTM 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

Context 256000 textaudioimagevideotext+image+audio+video->text

nvidia/nemotron-3-super-120b-a12b:free

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

Context 262144 texttext->text

nvidia/nemotron-3-ultra-550b-a55b:free

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Context 1e+06 texttext->text

nvidia/nemotron-3.5-content-safety:free

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

Context 128000 textimagetext+image->text

nvidia/nemotron-nano-12b-v2-vl:free

NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba's...

Context 128000 imagetextvideotext+image+video->text

nvidia/nemotron-nano-9b-v2:free

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and...

Context 128000 texttext->text
openai 1 free models

openai/gpt-oss-20b:free

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

Context 131072 texttext->text
openrouter 1 free models

openrouter/free

The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that...

Context 200000 textimagetext+image->text
poolside 2 free models

poolside/laguna-s-2.1:free

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

Context 262144 texttext->text

poolside/laguna-xs-2.1:free

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

Context 262144 texttext->text

Operator Notes

This page should be refreshed daily because free-model availability changes quickly.

Category-aware free-model counts are derived from inferencer's OpenRouter category wrappers, while the provider boards remain manually curated.

The OpenRouter section is machine-generated; the provider boards are intentionally curated.

Use the dashboard as a monitor, not as a compliance promise. Vendor free-tier terms can change without warning.