Download a model to run on your own hardware, or run it on Ollama’s cloud.
20+ of 245 models
DeepSeek-V4.1-Flash is an advanced tool designed to enhance search capabilities, providing users with faster and more accurate results.
Z.ai's flagship model and the most capable open-weights model for coding, with major gains on long-horizon agentic tasks.
Z.ai's first natively multimodal model, approaching Claude Opus 4.8 on coding and agentic benchmarks with just 18B active parameters.
Clef-Flash is a 9B multimodal model that turns a state and a schema of typed questions into decisions.
Clef is a 27B multimodal model created by Cloudflare that turns a state and a schema of typed questions into decisions.
EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture.
Mistral Large 4 is Mistral's open-weight, general-purpose multimodal model. Its granular Mixture-of-Experts design activates 52B of 1.05T total parameters, with a 1.6B vision encoder handling images.
Laya is a 421M decision model from Convai Innovations, fine-tuned from ModernBERT-large.
A 4B decision model from Together AI for fast classification.
A 9B decision model from Bespoke Labs for fast, typed classification.
Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.
Qwen3.6 delivers substantial upgrades in agentic coding and thinking preservation than previous Qwen models.
Chirp Chirp! 🐦 We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement.
A self-improving family of open-source models for agentic coding
This experimental preview of the architecture that will underpin Qwen4.
Meta's latest open model built for always-on local agents. 30B parameters, licensed under Apache 2.0 and runs on a single GPU — tuned for tool use, long tasks, and failure recovery.
NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with 3B active parameters built for always-on agents.
MiniMax M3: Coding & Agentic Frontier. 1M context window. Native Multimodality.
NVIDIA Nemotron 3 Nano Omni is a multimodal large language model that unifies video, audio, image, and text understanding to support enterprise-grade Q&A, summarization, transcription, and document intelligence workflows.
GLM-5.2 is Z.ai’s flagship model for the era of long-horizon tasks.