Gemini 3.7 Flash is Google's fast, low-cost Gemini tier released in August 2026, aimed at coding and agentic workloads at $0.75 per million input tokens — arriving three weeks after the previous Flash release.
Latest AI Models 2026 — 93 Released
The latest AI models of 2026, in one place — a complete, continuously updated catalog of every major AI model released and available in 2026, from 33 brands across LLMs, image, video, multimodal, audio, and scientific models. Whether you're tracking new large language model releases, the major models of the year, or which reasoning / thinking-mode models just shipped, this is the full list. Click any model for full coverage; click any brand for its release timeline.
August 202615 models
DeepSeek-V4-Pro — the flagship tier of DeepSeek's trillion-scale V4 family, a MoE model aimed at frontier reasoning, coding and agentic workloads above the cheaper V4 Flash tier.
Meta's Muse Glimmer — a 30B open-weights, multimodal agentic model built for always-on local agent workflows, small enough to run on a single consumer GPU.
Meta's Muse Spark 1.2 — the August 2026 update to Meta's paid Muse Spark foundation model, powering the Muse Code terminal agent with higher coding and Terminal-Bench scores at a low per-task price.
NVIDIA's Nemotron 4 — the next major generation of the Nemotron open model family, following the Nemotron 3 and Nemotron 3.5 releases, aimed at reasoning-heavy and agentic enterprise AI workloads.
NVIDIA's Nemotron 3.5 Lightning — the speed-optimized tier of the Nemotron 3.5 open model family, built for fast, accurate specialized task execution in long-running agentic workflows.
WeatherNext is Google DeepMind's family of machine-learning weather forecasting models, which predict global atmospheric conditions and cyclone tracks up to 15 days ahead. WeatherNext 2 generates each forecast scenario in under a minute on a single TPU — far faster than traditional physics-based numerical weather prediction — and powers weather features across Google Search, Gemini and Pixel.
Alibaba Cloud's latest Qwen 3-series flagship LLM, the successor to Qwen 3.7 with improved reasoning, coding and agentic capabilities.
NVIDIA's Nemotron 3 Nano — the compact, efficiency-focused tier of the Nemotron 3 open model family, including the Nano Omni multimodal variant for long-context document, audio and video agents.
GPT-5.6 Sol — the flagship tier of OpenAI's GPT-5.6 family, positioned as its highest-capability reasoning model and benchmarked against Claude Opus 5 and Fable 5.
GPT-5.6 Luna — the low-cost, high-throughput tier of OpenAI's GPT-5.6 family, whose price cuts drove the 2026 inference price war.
GPT-5.6 Terra — the mid tier of OpenAI's GPT-5.6 family, sitting between Luna and Sol on price and capability, generally available via the OpenAI API and Amazon Bedrock.
DeepSeek-V4-Flash — the fast, low-cost tier of DeepSeek's V4 family, released at $0.28 per million tokens and upgraded (0731) with stronger agentic and coding performance.
July 202615 models
Granite is IBM's family of open, Apache 2.0 licensed language, vision and embedding models aimed at enterprise workloads.
Claude Opus 5 — Anthropic's most capable model in the Claude 5 family, the frontier successor to Opus 4.8 for the hardest reasoning, coding and agentic work.
DeepSeek V4 — DeepSeek's frontier model with a million-token context window built for long-running agentic workloads.
GLM-5.2 — Zhipu AI's flagship GLM model tuned for long-horizon, multi-step agentic tasks.
Step 3.7 — StepFun's enterprise-ready multimodal model (including the Step 3.7 Flash tier), optimised for NVIDIA GPU inference.
Inkling — Thinking Machines Lab's first model trained from scratch: a 975B-parameter open-weights multimodal Mixture-of-Experts with 41B active parameters and controllable thinking effort, released under Apache 2.0 in July 2026.
Kimi K3 — Moonshot AI's next-generation Kimi model, reported to close the gap with Anthropic's Opus 4.8 and shipped as one of the largest open-source frontier models.
Amazon Nova — Amazon's family of foundation models on AWS Bedrock, spanning text, vision and the Nova Act agentic browser model.
MiniMax M3 — MiniMax's long-context reasoning and agentic model, positioned for open-weight deployment on accelerated infrastructure.
Gemini 3.1 Pro — Google's frontier Gemini model for complex reasoning tasks, benchmarked against GPT-5.4, Claude Opus 4.6 and Grok.
Gemini Omni Flash — the fast, low-cost tier of Google's Gemini Omni video model, aimed at conversational enterprise video generation via the API.
Nano Banana 2 Pro — the flagship tier of Google's Nano Banana 2 image family, powered by Gemini 3.1 Flash Image.
Watermelon — the codename for Meta's upcoming frontier AI model, reported to match OpenAI's GPT-5.5 on key benchmarks. Meta's bid to reclaim ground in the frontier LLM race under Alexandr Wang's superintelligence group.
Nano Banana 2 Lite — Google's fastest, cheapest tier of its Nano Banana 2 (Gemini image) family, built for high-volume, low-latency image generation.
June 202613 models
Claude Sonnet 5 — Anthropic's balanced mid-tier model in the Claude 5 family, pairing strong reasoning and coding with fast, cost-efficient responses.
PixVerse V6 — PixVerse's 2026 flagship video generation model, advancing toward real-time, cinematic control across creative and agentic workflows.
PixVerse R1 — PixVerse's real-time video 'world model' with live input, subject priority, shared worlds and personalized avatars.
Meituan's 1.6-trillion-parameter open-source agentic coding model with a 1M-token context window (June 2026). The first trillion-parameter model claimed to be fully pre-trained AND served on domestic Chinese AI chips.
Mistral AI's document-intelligence (OCR) model (June 2026). Structure-aware extraction with bounding boxes, block classification and confidence scores across 170 languages; self-hostable in a single container. Tops OlmOCRBench; feeds RAG, agentic and enterprise-search pipelines.
Recraft's most advanced text-to-image model (V4.1) — design-grade image generation with strong visual taste, brand styles, and precise control.
Sakana AI model that coordinates and orchestrates multiple models, matching frontier models on some benchmarks.
The GLM (General Language Model) family from Zhipu AI / Z.ai — open-weight Chinese LLMs (GLM-4.5/4.6/5/5.1/5.2 plus GLM-OCR, Air, Flash and Turbo variants) known for strong coding and long-context performance under permissive (MIT) licenses.
NVIDIA's Nemotron 3 Ultra — the largest, highest-capability model in the Nemotron 3 open model family, tuned for advanced reasoning, agentic workflows and enterprise AI.
Anthropic's Mythos-class multimodal Claude model (text, vision, code) made safe for general use — strong at software engineering, knowledge work, long-context reasoning, scientific research and protein design, with safeguards that fall back to Claude Opus 4.8 for sensitive domains. Available via the Claude API and claude.ai. Priced at $10 / $50 per million input / output tokens.
The same underlying multimodal Claude model as Fable 5, but with safeguards lifted in some domains (e.g. cybersecurity, biology), offered to vetted users through Anthropic's Project Glasswing trusted-access program. Text, vision and code. Priced at $10 / $50 per million input / output tokens.
Google's 12-billion-parameter open model in the Gemma 4 family — a compact, efficient multimodal LLM designed to run on a single GPU.
May 202614 models
Anthropic's Claude Sonnet 4.8, launched alongside Opus 4.8 — brings Opus-class quality into the mid tier with the new Dynamic Workflows tool and improved vision workflows.
Anthropic's flagship Claude Opus 4.8 — the successor to Opus 4.7 with further improvements in advanced reasoning, coding and agentic capabilities.
Alibaba Cloud's latest Qwen 3-series flagship LLM, the successor to Qwen 3.6 with improved reasoning, coding and agentic capabilities.
Google DeepMind's mathematical-reasoning model that formally proves theorems; the AlphaProof Nexus version tackles Erdős problems.
Google DeepMind's embodied-reasoning Gemini model for real-world robotics tasks.
Physical Intelligence's Vision-Language-Action (VLA) models for general robot control (π0, π0-FAST, π0.6).
Google's fast, cost-efficient Gemini model tier, announced at Google I/O 2026.
OpenAI's image generation model, successor to DALL·E, integrated into ChatGPT and the API.
ByteDance's unified model for image and video understanding, generation and editing.
Google's multimodal model family for video generation and editing, announced at Google I/O 2026. Gemini Omni Flash creates and edits high-quality video from text, image, audio and video inputs with physics-aware generation, conversational editing, digital avatars and SynthID watermarking.
April 20265 models
Moonshot AI's flagship 1T-parameter open-weight LLM featuring 262K context window, long-horizon coding with up to 300 sub-agent swarms and 4,000 coordinated steps. Outperforms GPT-5.4 and Claude Opus 4.6 on SWE-Bench Pro (58.6). Supports multimodal input including vision.
Opus 4.7 is a notable improvement on Opus 4.6 in advanced software engineering, with particular gains on the most difficult tasks.
Meta open-source AI video generation model for creating short video clips from text and image prompts.
Most cost-effective video generation model is now available to developers in the Gemini API.
March 202630 models
Lyria 3 helps you express, explore, and experiment with high-fidelity music, using prompts to create tracks with natural flow from note to note.
Hybrid reasoning model with superior intelligence for agents, featuring a 1M context window
Claude Opus 4.6 is state-of-the-art across a wide range of coding and agentic capabilities.
Claude Haiku 4.5 is our fastest, most cost-efficient model, matching Sonnet 4’s performance on coding, computer use, and agent tasks.
January 20261 model
Seedance 2.0 adopts a unified multimodal audio-video joint generation architecture that supports text, image, audio, and video inputs, leading to the most comprehensive multimodal content reference and editing capabilities in the industry.
Frequently Asked Questions
How many AI models were released in 2026?
We are currently tracking 93 major AI models added in 2026, from 33 different brands. This catalog updates continuously as new models launch and we add them to our tracker.
Which company released the most AI models in 2026?
Google leads with 12 models. The top contributors by model count are: Google (12), OpenAI (11), Anthropic (11), NVIDIA (7), Google DeepMind (7).
What are the current AI models available in 2026?
The current, available AI models of 2026 span 57 LLM, 12 Multimodal, 10 Video, 9 Image, 4 Audio, 1 Science — 93 models in total. This page lists every one, newest first, with the company behind it.
Which 2026 AI models have a thinking or reasoning mode?
Most of the new 2026 large language models — including the latest flagship releases from the leading labs — ship a dedicated thinking / reasoning mode for harder problems. Use the LLM filter above to see the reasoning-capable models released this year.
Are upcoming or newly announced 2026 AI models included?
Yes — we add each model as soon as it is announced or launched and appears in news coverage, so just-announced and upcoming 2026 models show up here quickly. Dates reflect when the model entered our tracker, which closely matches public release timing.
How is this catalog maintained?
We add each major AI model as it appears in news coverage from our 30+ sources (research labs, tech publications, and AI communities). Dates reflect when the model entered our tracker, which closely corresponds to public launch dates for most entries. Click any model to see full news coverage and related entities.


