Model catalog

Open-source and free-tier models across eight modalities — all behind one API.

Text Generation

Chat and completion with open-source LLMs. OpenAI-compatible.

Featured

Gemini 2.5 Flash

Google's fast multimodal model with a large context window.

Text GenerationGoogle Gemini
1,000K context2 cr per 1K tokens
Featured

Llama 3.1 8B Instant

Fast, capable general-purpose LLM. Great default for most tasks.

Text GenerationGroq
131.072K context1 cr per 1K tokens
Featured

Llama 3.3 70B Versatile

High-quality reasoning and generation for complex tasks.

Text GenerationGroq
131.072K context4 cr per 1K tokens
Featured

Meta: Llama 3.3 70B Instruct (free)

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

Text GenerationOpenRouter
 1 cr per 1K tokens
Featured

OpenAI: gpt-oss-120b (free)

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

Text GenerationOpenRouter
 1 cr per 1K tokens
Featured

OpenAI: gpt-oss-20b (free)

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

Text GenerationOpenRouter
 1 cr per 1K tokens
Featured

Qwen: Qwen3 Coder 480B A35B (free)

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

Text GenerationOpenRouter
 1 cr per 1K tokens

Cohere: North Mini Code (free)

North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...

Text GenerationOpenRouter
 1 cr per 1K tokens

GPT-OSS 20B

Open-weight model with strong instruction following.

Text GenerationGroq
131.072K context2 cr per 1K tokens

Google: Gemma 4 26B A4B (free)

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Text GenerationOpenRouter
 1 cr per 1K tokens

Google: Gemma 4 31B (free)

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Text GenerationOpenRouter
 1 cr per 1K tokens

LiquidAI: LFM2.5-1.2B-Instruct (free)

LFM2.5-1.2B-Instruct is a compact, high-performance instruction-tuned model built for fast on-device AI. It delivers strong chat quality in a 1.2B parameter footprint, with efficient edge inference and broad runtime support.

Text GenerationOpenRouter
 1 cr per 1K tokens

LiquidAI: LFM2.5-1.2B-Thinking (free)

LFM2.5-1.2B-Thinking is a lightweight reasoning-focused model optimized for agentic tasks, data extraction, and RAG—while still running comfortably on edge devices. It supports long context (up to 32K tokens) and is...

Text GenerationOpenRouter
 1 cr per 1K tokens

Meta: Llama 3.2 3B Instruct (free)

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...

Text GenerationOpenRouter
 1 cr per 1K tokens

NVIDIA: Nemotron 3 Nano 30B A3B (free)

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

Text GenerationOpenRouter
 1 cr per 1K tokens

NVIDIA: Nemotron 3 Nano Omni (free)

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

Text GenerationOpenRouter
 1 cr per 1K tokens

NVIDIA: Nemotron 3 Super (free)

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

Text GenerationOpenRouter
 1 cr per 1K tokens

NVIDIA: Nemotron 3 Ultra (free)

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Text GenerationOpenRouter
 1 cr per 1K tokens

NVIDIA: Nemotron 3.5 Content Safety (free)

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

Text GenerationOpenRouter
 1 cr per 1K tokens

NVIDIA: Nemotron Nano 12B 2 VL (free)

NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba’s...

Text GenerationOpenRouter
 1 cr per 1K tokens

NVIDIA: Nemotron Nano 9B V2 (free)

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and...

Text GenerationOpenRouter
 1 cr per 1K tokens

Nex AGI: Nex-N2-Pro (free)

Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...

Text GenerationOpenRouter
 1 cr per 1K tokens

Nous: Hermes 3 405B Instruct (free)

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...

Text GenerationOpenRouter
 1 cr per 1K tokens

Poolside: Laguna M.1 (free)

Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. Designed for agentic coding workflows, it supports tool calling and reasoning, with a 256K...

Text GenerationOpenRouter
 1 cr per 1K tokens

Poolside: Laguna XS.2 (free)

Laguna XS.2 is the second-generation model in the XS size class from [Poolside](https://poolside.ai/), their efficient coding agent series. It combines tool calling and reasoning capabilities with a compact footprint, offering...

Text GenerationOpenRouter
 1 cr per 1K tokens

Qwen3 32B

Multilingual model with solid coding and math abilities.

Text GenerationGroq
131.072K context3 cr per 1K tokens

Qwen: Qwen3 Next 80B A3B Instruct (free)

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...

Text GenerationOpenRouter
 1 cr per 1K tokens

Venice: Uncensored (free)

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...

Text GenerationOpenRouter
 1 cr per 1K tokens

Image to Text

Vision understanding, captioning, and OCR from images.

Featured

Gemini 2.5 Flash (Vision)

Image understanding, captioning, and OCR.

Image to TextGoogle Gemini
 2 cr per 1K tokens

Llama 4 Scout (Vision)

Open multimodal model for visual question answering.

Image to TextGroq
 3 cr per 1K tokens

Text to Image

Generate images from prompts with FLUX and SDXL.

Featured

FLUX.1 [schnell]

Ultra-fast, high-quality text-to-image generation.

Text to ImageHugging Face
 20 cr per image

FLUX.1 [dev]

Highest-detail FLUX model for photorealistic images.

Text to ImageHugging Face
 40 cr per image

Stable Diffusion XL

Versatile open image model with broad style support.

Text to ImageHugging Face
 25 cr per image

Image to Image

Edit and transform images with a guiding prompt.

Featured

FLUX.1 Kontext [dev]

Prompt-guided image editing and transformation.

Image to ImageHugging Face
 40 cr per image

Text to Video

Create short video clips from text prompts.

Featured

LTX Video

Generate short video clips from a text prompt.

Text to VideoHugging Face
 200 cr per request

Text to Speech

Natural-sounding speech synthesis from text.

Featured

PlayAI TTS

Natural English speech synthesis.

Text to SpeechGroq
 5 cr per 1K characters

Kokoro 82M

Lightweight open-source TTS via Hugging Face.

Text to SpeechHugging Face
 4 cr per 1K characters

Speech to Text

Fast, multilingual transcription with Whisper.

Featured

Whisper Large v3 Turbo

Fast multilingual transcription (216x real-time).

Speech to TextGroq
 1 cr per second

Whisper Large v3

State-of-the-art accuracy for transcription & translation.

Speech to TextGroq
 2 cr per second

Embeddings

Vector embeddings for search and RAG (entity-to-entity).

Featured

BGE Base EN v1.5

Compact, high-quality English text embeddings.

EmbeddingsHugging Face
 1 cr per 1K tokens

Multilingual E5 Large

Multilingual embeddings for cross-language search.

EmbeddingsHugging Face
 1 cr per 1K tokens