F24 SALES

Free LLM models
via API.

Free model directory

24 modelsLast scan

Z.ai

GLM-5.3-Flash

  1. ๐Ÿ† NVIDIA
CONTEXT1.05M
MAX OUTPUT131.07K
TextImage1 gateway

A multimodal Z.ai model for coding, visual understanding and tool-using agents. Its sparse architecture targets efficient inference with a long context window.

Details & access

Limits above: NVIDIA. Gateway differences are listed below.

Thinking levels
low ยท high ยท max
Default effort
max
NVIDIA

z-ai/glm-5.3-flash

API trial terms, quotas and model licenses apply.

  • basenvidia/z-ai/glm-5.3-flash

Sources

Try this model
NVIDIA ยท Not yet tested

Moonshot AI

Kimi K3

  1. ๐Ÿ† NVIDIA
CONTEXT1.05M
MAX OUTPUT131.07K
TextImageVideo1 gateway

Moonshot AI's multimodal reasoning model for long-running agent tasks. It accepts text and images and offers adjustable thinking effort.

Details & access

Limits above: NVIDIA. Gateway differences are listed below.

Thinking levels
low ยท high ยท max
Thinking controls
Can be toggled
NVIDIA

moonshotai/kimi-k3

API trial terms, quotas and model licenses apply.

  • basenvidia/moonshotai/kimi-k3

Sources

Try this model
NVIDIA ยท Not yet tested

Meituan

LongCat 2.5 Preview

  1. ๐Ÿ† OpenCode Zen
  1. ๐Ÿฅˆ Nous Portal
CONTEXT1.05M
MAX OUTPUT131.07K
Reasoning TextImage2 gateways

Meituan's multimodal coding and agent model, with image understanding and long-context reasoning. Its thinking mode can be switched on or off; named effort levels are not documented.

Details & access

Limits above: Nous Portal. Gateway differences are listed below.

Thinking controls
Can be toggled

Reviewed as the same advertised Meituan LongCat 2.5 Preview model across Nous Portal and OpenCode Zen, not LongCat 2.0. Gateway limits and access conditions remain separate.

OpenCode Zen
Sign up
API

โš ๏ธ Free access requires the OpenCode app or CLI.

longcat-2.5-preview-free

No passing direct API route in the latest scan.

  • baseopencode/longcat-2.5-preview-free
Nous Portal

meituan/longcat-2.5-preview:free

Check the portal for current free routes and limits.

  • basenous/meituan/longcat-2.5-preview:free

Sources

Try this model
Nous Portal ยท Not yet tested

Anonymous ยท stealth preview

Space Bunny Alpha

  1. ๐Ÿ† Kilo
  2. ๐Ÿ† OpenCode Zen
  3. ๐Ÿ† OpenRouter
CONTEXT1M
MAX OUTPUT524.29K
Reasoning TextImageVideo3 gateways

An anonymous preview model advertised for coding, reasoning and multimodal input. Its developer and underlying model identity have not been disclosed.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Curated Space Bunny alias family. The creator is undisclosed; matching aliases do not verify identical weights across gateways.

Kilo

stealth/space-bunny-alpha

Free availability and upstream data policies vary.

  • basekilo/stealth/space-bunny-alpha
OpenCode Zen
Sign up
API

โš ๏ธ Free access requires the OpenCode app or CLI.

space-bunny-free

Context here
1,048,576

Thinking levels here: low ยท medium ยท high ยท xhigh ยท max

  • baseopencode/space-bunny-free
OpenRouter

stealth/space-bunny-alpha

Thinking levels here: max ยท xhigh ยท high ยท medium ยท low

Free-model quotas and model-specific terms apply.

  • baseopenrouter/stealth/space-bunny-alpha

Sources

Try this model
Kilo ยท Last tested

SDAIA

ALLaM-2-7b

CONTEXT4.1K
MAX OUTPUT4.1K
Text1 gateway

SDAIA's 7-billion-parameter instruction-tuned model for Arabic and English. Trained from scratch with staged English and Arabic-English pretraining, it supports bilingual conversations, text generation and summarization.

Details & access

Limits above: Groq. Gateway differences are listed below.

Groq

allam-2-7b

A free developer tier with per-model limits.

  • basegroq/allam-2-7b

Sources

Try this model
Groq ยท Not yet tested

Kilo

Auto Free

CONTEXT256K
MAX OUTPUT32.77K
Reasoning Text1 gateway

Kilo's Auto Free router distributes requests across available free models on OpenRouter, with the pool updated as availability changes. It is not a fixed model; upstream providers may log prompts and outputs, so do not submit confidential data.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

kilo-auto/free

Free availability and upstream data policies vary.

  • basekilo/kilo-auto/free

Sources

Try this model
Kilo ยท Not yet tested

Cohere

North Mini Code

CONTEXT256K
MAX OUTPUT64K
Reasoning Text2 gateways

Cohere's compact mixture-of-experts model focused on agentic coding. It targets code changes and tool-driven software development.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

cohere/north-mini-code:free

Free availability and upstream data policies vary.

  • basekilo/cohere/north-mini-code:free
OpenRouter

cohere/north-mini-code:free

Free-model quotas and model-specific terms apply.

  • fastopenrouter/cohere/north-mini-code:free-fast
  • thinkopenrouter/cohere/north-mini-code:free-think

Sources

Try this model
Kilo ยท Not yet tested

Dots Studio

Dots3-Note Preview

CONTEXT512K
MAX OUTPUT460.8K
Reasoning TextImage2 gateways

A preview of Dots Studio's lighter Dots 3 mixture-of-experts model. It supports long-context work and configurable reasoning.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

dots-studio/dots-3-note-preview:free

Free availability and upstream data policies vary.

  • basekilo/dots-studio/dots-3-note-preview:free
OpenRouter

dots-studio/dots-3-note-preview:free

Free-model quotas and model-specific terms apply.

  • fastopenrouter/dots-studio/dots-3-note-preview:free-fast
  • thinkopenrouter/dots-studio/dots-3-note-preview:free-think

Sources

Try this model
Kilo ยท Not yet tested

Google DeepMind

Lyria 3 Pro Preview

CONTEXT1.05M
MAX OUTPUT65.54K
TextImage2 gateways

Google's music-generation model for full-length songs, not a conventional chat LLM. Zero token prices in the catalog do not mean song generation is free.

Details & access

Music generation is billed per song/clip despite zero token prices in the catalog. This is not an unconditionally free chat model.

Limits above: OpenRouter. Gateway differences are listed below.

Output
text, audio
Kilo

google/lyria-3-pro-preview

Free availability and upstream data policies vary.

No passing direct API route in the latest scan.

  • basekilo/google/lyria-3-pro-preview
OpenRouter

google/lyria-3-pro-preview

Free-model quotas and model-specific terms apply.

  • baseopenrouter/google/lyria-3-pro-preview

Sources

Try this model
OpenRouter ยท Not yet tested

OpenAI

GPT OSS 120B

CONTEXT131.07K
MAX OUTPUT65.54K
Text1 gateway

OpenAI's larger open-weight reasoning model, served here by Groq. It offers adjustable reasoning effort for text and tool-based tasks.

Details & access

Limits above: Groq. Gateway differences are listed below.

Thinking levels
low ยท medium ยท high
Groq

openai/gpt-oss-120b

A free developer tier with per-model limits.

  • basegroq/openai/gpt-oss-120b

Sources

Try this model
Groq ยท Not yet tested

OpenAI

GPT OSS 20B

CONTEXT131.07K
MAX OUTPUT65.54K
Text1 gateway

The smaller open-weight GPT-OSS reasoning model, served here by Groq. It supports low, medium and high reasoning effort.

Details & access

Limits above: Groq. Gateway differences are listed below.

Thinking levels
low ยท medium ยท high
Groq

openai/gpt-oss-20b

A free developer tier with per-model limits.

  • basegroq/openai/gpt-oss-20b

Sources

Try this model
Groq ยท Last tested

InclusionAI

Ling 3.0 Flash Sante

CONTEXT262.14K
MAX OUTPUT32.77K
Reasoning Text3 gateways

A health- and medicine-focused Ling 3.0 Flash model from InclusionAI. Domain specialization does not make its responses medical advice.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

inclusionai/ling-3.0-flash-sante:free

Free availability and upstream data policies vary.

  • basekilo/inclusionai/ling-3.0-flash-sante:free
Nous Portal

inclusionai/ling-3.0-flash-sante:free

Check the portal for current free routes and limits.

  • basenous/inclusionai/ling-3.0-flash-sante:free
OpenRouter

inclusionai/ling-3.0-flash-sante:free

Free-model quotas and model-specific terms apply.

  • fastopenrouter/inclusionai/ling-3.0-flash-sante:free-fast
  • thinkopenrouter/inclusionai/ling-3.0-flash-sante:free-think

Sources

Try this model
Kilo ยท Not yet tested

Liquid AI

LFM2.5-2.6B

CONTEXT65.54K
MAX OUTPUT8.19K
Reasoning Text2 gateways

Liquid AI's compact reasoning model for extraction, retrieval and agent workflows. Its model guidance does not position it as an agentic coding specialist.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

liquid/lfm-2.5-2.6b:free

Free availability and upstream data policies vary.

  • basekilo/liquid/lfm-2.5-2.6b:free
OpenRouter

liquid/lfm-2.5-2.6b:free

Free-model quotas and model-specific terms apply.

  • baseopenrouter/liquid/lfm-2.5-2.6b:free

Sources

Try this model
Kilo ยท Not yet tested

NVIDIA

Nemotron 3 Nano Omni

CONTEXT256K
MAX OUTPUT65.54K
Reasoning TextAudioImageVideo2 gateways

A multimodal NVIDIA model that can work with text, images, audio and video. Designed for perception and context processing in agent workflows.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free

Free availability and upstream data policies vary.

  • basekilo/nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free
OpenRouter

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free

Free-model quotas and model-specific terms apply.

  • fastopenrouter/nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free-fast
  • thinkopenrouter/nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free-think

Sources

Try this model
Kilo ยท Not yet tested

NVIDIA

Nemotron 3 Super

CONTEXT262.14K
MAX OUTPUT235.93K
Reasoning Text2 gateways

NVIDIA's hybrid mixture-of-experts reasoning model for multi-agent workflows. Only a subset of its total parameters is active per token.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

nvidia/nemotron-3-super-120b-a12b:free

Free availability and upstream data policies vary.

  • basekilo/nvidia/nemotron-3-super-120b-a12b:free
OpenRouter

nvidia/nemotron-3-super-120b-a12b:free

Thinking levels here: medium ยท low

Free-model quotas and model-specific terms apply.

  • fastopenrouter/nvidia/nemotron-3-super-120b-a12b:free-fast
  • thinkopenrouter/nvidia/nemotron-3-super-120b-a12b:free-think

Sources

Try this model
Kilo ยท Not yet tested

NVIDIA

Nemotron 3 Ultra

CONTEXT1M
MAX OUTPUT65.54K
Reasoning Text3 gateways

NVIDIA's larger Nemotron 3 model for reasoning and agent orchestration. It combines a long context window with a sparse hybrid architecture.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

nvidia/nemotron-3-ultra-550b-a55b:free

Free availability and upstream data policies vary.

  • basekilo/nvidia/nemotron-3-ultra-550b-a55b:free
OpenRouter

nvidia/nemotron-3-ultra-550b-a55b:free

Thinking levels here: high ยท medium

Free-model quotas and model-specific terms apply.

  • fastopenrouter/nvidia/nemotron-3-ultra-550b-a55b:free-fast
  • thinkopenrouter/nvidia/nemotron-3-ultra-550b-a55b:free-think
OpenCode Zen
Sign up
API

โš ๏ธ Free access requires the OpenCode app or CLI.

nemotron-3-ultra-free

Max output here
128,000

No passing direct API route in the latest scan.

  • baseopencode/nemotron-3-ultra-free

Sources

Try this model
Kilo ยท Not yet tested

NVIDIA

Nemotron 3.5 Content Safety

CONTEXT128K
MAX OUTPUT8.19K
Reasoning TextImage2 gateways

A compact multimodal safety model for evaluating prompts and model responses. This is a guardrail specialist, not a general-purpose chat assistant.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

nvidia/nemotron-3.5-content-safety:free

Free availability and upstream data policies vary.

  • basekilo/nvidia/nemotron-3.5-content-safety:free
OpenRouter

nvidia/nemotron-3.5-content-safety:free

Free-model quotas and model-specific terms apply.

  • fastopenrouter/nvidia/nemotron-3.5-content-safety:free-fast
  • thinkopenrouter/nvidia/nemotron-3.5-content-safety:free-think

Sources

Try this model
Kilo ยท Not yet tested

NVIDIA

Nemotron 3.5 Lightning

CONTEXT1M
MAX OUTPUT65.54K
Reasoning Text3 gateways

NVIDIA's lightweight mixture-of-experts model for responsive agent workflows. It balances a small active parameter count with long-context processing.

Details & access

Limits above: OpenRouter. Gateway differences are listed below.

Thinking controls
Can be toggled
Kilo

nvidia/nemotron-3.5-lightning:free

Free availability and upstream data policies vary.

No passing direct API route in the latest scan.

  • basekilo/nvidia/nemotron-3.5-lightning:free
OpenRouter

nvidia/nemotron-3.5-lightning:free

Free-model quotas and model-specific terms apply.

  • fastopenrouter/nvidia/nemotron-3.5-lightning:free-fast
  • thinkopenrouter/nvidia/nemotron-3.5-lightning:free-think
OpenCode Zen
Sign up
API

โš ๏ธ Free access requires the OpenCode app or CLI.

nemotron-3.5-lightning-free

Context here
262,144
Max output here
262,144

No passing direct API route in the latest scan.

  • baseopencode/nemotron-3.5-lightning-free

Sources

Try this model
OpenRouter ยท Not yet tested

OpenRouter

Free Models Router

CONTEXT200K
MAX OUTPUTโ€”
Reasoning TextImage2 gateways

A router that selects from OpenRouter's available free models. The actual model and provider can change between requests.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

openrouter/free

Free availability and upstream data policies vary.

  • basekilo/openrouter/free
OpenRouter

openrouter/free

Free-model quotas and model-specific terms apply.

  • fastopenrouter/openrouter/free-fast
  • thinkopenrouter/openrouter/free-think

Sources

Try this model
Kilo ยท Not yet tested

Poolside

Laguna S 2.1

CONTEXT262.14K
MAX OUTPUT32.77K
Reasoning Text3 gateways

Poolside's coding-agent model for tool-driven software engineering. Gateway-specific context and output limits are listed separately.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

poolside/laguna-s-2.1:free

Free availability and upstream data policies vary.

  • basekilo/poolside/laguna-s-2.1:free
Nous Portal

poolside/laguna-s-2.1:free

Max output here
131,072

Check the portal for current free routes and limits.

  • basenous/poolside/laguna-s-2.1:free
OpenRouter

poolside/laguna-s-2.1:free

Free-model quotas and model-specific terms apply.

  • fastopenrouter/poolside/laguna-s-2.1:free-fast
  • thinkopenrouter/poolside/laguna-s-2.1:free-think

Sources

Try this model
Kilo ยท Not yet tested

Poolside

Laguna XS 2.1

CONTEXT262.14K
MAX OUTPUT32.77K
Reasoning Text3 gateways

A smaller coding-agent model in Poolside's Laguna family. It is designed for software development workflows with a relatively small active parameter count.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

poolside/laguna-xs-2.1:free

Free availability and upstream data policies vary.

  • basekilo/poolside/laguna-xs-2.1:free
Nous Portal

poolside/laguna-xs-2.1:free

Check the portal for current free routes and limits.

  • basenous/poolside/laguna-xs-2.1:free
OpenRouter

poolside/laguna-xs-2.1:free

Free-model quotas and model-specific terms apply.

  • fastopenrouter/poolside/laguna-xs-2.1:free-fast
  • thinkopenrouter/poolside/laguna-xs-2.1:free-think

Sources

Try this model
Kilo ยท Not yet tested

Alibaba

Qwen3.8 27B

CONTEXT131.07K
MAX OUTPUT16.38K
TextImage3 gateways

A dense Qwen vision-language model for coding, visual analysis and agent tasks. Thinking controls and context limits vary by gateway.

Details & access

Limits above: Groq. Gateway differences are listed below.

Thinking levels
none ยท default ยท low ยท medium ยท high
Thinking controls
Can be toggled
Groq

qwen/qwen3.8-27b

A free developer tier with per-model limits.

  • basegroq/qwen/qwen3.8-27b
Kilo

qwen/qwen3.8-27b:free

Context here
262,144
Max output here
235,929
Reasoning here

Thinking levels here: Not documented

Free availability and upstream data policies vary.

No passing direct API route in the latest scan.

  • basekilo/qwen/qwen3.8-27b:free
OpenRouter

qwen/qwen3.8-27b:free

Context here
262,144
Max output here
235,929
Reasoning here

Thinking levels here: xhigh ยท medium ยท low

Free-model quotas and model-specific terms apply.

No passing direct API route in the latest scan.

  • fastopenrouter/qwen/qwen3.8-27b:free-fast
  • thinkopenrouter/qwen/qwen3.8-27b:free-think

Sources

Try this model
Groq ยท Not yet tested

StepFun

Step 3.7 Flash

CONTEXT262.14K
MAX OUTPUT262.14K
Reasoning TextImage2 gateways

StepFun's multimodal mixture-of-experts model with image and video understanding. It combines reasoning and tool use in a relatively efficient architecture.

Details & access

Limits above: Kilo. Gateway differences are listed below.

Kilo

stepfun/step-3.7-flash:free

Free availability and upstream data policies vary.

  • basekilo/stepfun/step-3.7-flash:free
Nous Portal

stepfun/step-3.7-flash:free

Max output here
32,768

Thinking levels here: high ยท medium ยท low

Check the portal for current free routes and limits.

No passing direct API route in the latest scan.

  • basenous/stepfun/step-3.7-flash:free

Sources

Try this model
Kilo ยท Not yet tested

Upstage

Solar Pro 4

CONTEXT524.29K
MAX OUTPUT131.07K
Reasoning Text1 gateway

Upstage's long-context model for document-heavy work and agent workflows. Its Nous listing exposes several named reasoning-effort levels.

Details & access

Limits above: Nous Portal. Gateway differences are listed below.

Thinking levels
max ยท xhigh ยท high ยท medium ยท low ยท minimal ยท none
Thinking controls
Can be toggled
Default effort
medium
Nous Portal

upstage/solar-pro4:free

Check the portal for current free routes and limits.

  • basenous/upstage/solar-pro4:free

Sources

Try this model
Nous Portal ยท Not yet tested