AI Basics & Tutorialsdeveloperslocal aiopen source

Best Open Source LLMs in 2026: Kimi K3, GLM-5.3 and Qwen

Open models now match frontier benchmark scores while ranking 27th on human preference. Both numbers are true.

Published September 24, 2026

Muhammad Usman

By Muhammad Usman · Founder & Lead Reviewer

Best Open Source LLMs in 2026: Kimi K3, GLM-5.3 and Qwen

Quick Answer

Open-weight models now score alongside frontier models on benchmarks, with Kimi K3 and GLM-5.3 both hitting 60 on the Artificial Analysis index. But on LMArena, where humans compare answers blind, the best open model ranks 27th. Most articles quote only the flattering number.

Open-weight models now score alongside frontier models on benchmarks. Kimi K3 and GLM-5.3 both hit 60 on the Artificial Analysis index against the frontier's top score. But on LMArena, where humans compare real answers blind, the best open model ranks 27th. Both numbers are true, and most articles only quote the flattering one.

Key takeaways

  • Kimi K3 and GLM-5.3 score 60 on the Artificial Analysis index, matching the frontier's top tier.
  • On LMArena human preference, the best open model ranks 27th at 1474 Elo against 1507 at the top.
  • GLM-5.3 is not MIT despite frequent reports. Only GLM-5.3-Flash is. GLM-5.3 carries a custom licence.
  • Kimi K3 has a revenue gate: Model-as-a-Service businesses over $20M revenue need a separate agreement.
  • Llama is effectively finished as a frontier line. Meta shipped proprietary Muse Spark in April 2026.

The models worth knowing

ModelVendorLicenceParamsContextAA index
Kimi K3MoonshotCustom, revenue-gated2.8T / 104B active1M60
GLM-5.3Z.aiCustom, not MIT753B / 40B1M60
Qwen3.8 2.4TAlibabaApache-2.02.4T / 95B984K58
GLM-5.3-FlashZ.aiMIT320B / 18B1M57
DeepSeek V4 ProDeepSeekMIT1.6T / 49B1M53
Qwen3.8-27BAlibabaApache-2.027B dense256K52
Muse GlimmerMetaApache-2.030B dense128K+n/a
Granite 4.2IBMApache-2.03B/8B/30B128Kn/a
Gemma 4GoogleApache-2.0up to 30.7B128-256Kn/a

The Artificial Analysis index version 4.1.1 combines nine evaluations including GPQA Diamond, Humanity's Last Exam, Terminal-Bench and SciCode.

The two leaderboards that disagree

This is the most important thing to understand about open-weight models right now.

On benchmarks, open models have caught up. Kimi K3 and GLM-5.3 both score 60, matching the top closed models on the Artificial Analysis index.

On human preference, they have not. LMArena ranks models by blind head-to-head comparisons where people pick the better answer. Its top ten is entirely Anthropic, Meta, Google and Moonshot's closed offerings. The highest-ranked open-weight model is GLM-5.3-Flash at rank 27, scoring 1474 against 1507 at the top.

Both measurements are legitimate. Benchmarks test specific capabilities under controlled conditions. Human preference captures whether the answer was actually good to receive: tone, structure, usefulness, knowing when to stop.

The honest reading: open models can solve the problems, and frontier models still communicate better. If you are running automated pipelines where output is parsed rather than read, benchmarks matter more. If a human reads every response, the preference gap is real and you will notice it.

Epoch AI puts the overall lag at roughly four months, or 8 ECI points, and notes this widened from about three months measured over 2023 to 2025.

The licence corrections that matter

Several widely repeated claims are wrong, and for anyone building commercially they are the difference between compliant and not.

GLM-5.3 is not MIT. Its HuggingFace licence tag reads other, a custom Z.ai licence. Only GLM-5.3-Flash is MIT, verified on its model card. Multiple outlets describe both as MIT.

Kimi K3 is open-weight, not open source. Its own LICENSE states that if you operate a Model-as-a-Service business and your group revenue exceeds $20 million over any consecutive 12 months, you must enter a separate agreement with Moonshot AI. For most readers this never bites, but it is not an OSI-approved open-source licence.

Gemma 4 is now genuinely Apache-2.0, a real change from the older Gemma Terms of Use. Google's own model card confirms it, though a separate Prohibited Use Policy still applies.

Cleanly unrestricted: Qwen3.8-27B, Muse Glimmer, Granite 4.2 and Gemma 4 are Apache-2.0. DeepSeek V4 Pro and GLM-5.3-Flash are MIT.

What happened to Llama

Meta's Llama line effectively ended as a frontier project.

In April 2026, Meta Superintelligence Labs shipped Muse Spark, which is proprietary, cloud-only, with no downloadable weights. Llama receives maintenance rather than frontier investment. This was reported independently by TechCrunch, The Register and VentureBeat.

Meta partially returned to open weights in August 2026 with Muse Glimmer: 30B parameters, Apache-2.0, single-GPU, tuned for tool use and long tasks. It is a genuinely useful local model, but a distill rather than a frontier release.

Context worth knowing: Chinese labs including DeepSeek, Alibaba and Z.ai reached roughly 41% of Hugging Face downloads by late 2025. The centre of gravity in open weights moved.

Which of these can you actually run?

A crucial split that benchmark tables obscure.

Datacenter only: Kimi K3 is 1.56TB of weights. GLM-5.3 needs roughly 8 GPUs at 756GB in FP8. DeepSeek V4 Pro is 893GB. You can download these; you cannot run them on a workstation.

Genuinely local on one GPU: Qwen3.8-27B, Muse Glimmer 30B, Nemotron 3.5 Lightning (30B MoE, 3B active), Granite 4.2 at 3B/8B/30B, and the Gemma 4 family.

So the practical choice is three-way, not two-way: cheap open APIs for the big models, local models a step below frontier, or a subscription. Our guide to open source LLM API pricing covers the first, and running AI locally with Ollama the second.

How to use them

Locally: install Ollama and pull the model, for example ollama pull qwen3.8 or ollama pull granite4.2. Ollama now ships a desktop app, so no terminal is required for basic use.

Via API: most are available through providers at a fraction of frontier pricing. GLM-5.3-Flash is roughly $0.10 per million tokens blended; DeepSeek V4 Pro $0.70; Kimi K3 $2.30.

For coding: see the best local coding model by VRAM for tiers at 8GB, 16GB and 24GB.

Which should you choose?

Best all-round open model: GLM-5.3 if you can access it via API, for its coding and agentic strength. Note the custom licence.

Best value: GLM-5.3-Flash. MIT licensed, 57 on the index, roughly $0.10 per million tokens.

Best local model: Qwen3.8-27B for general use, Muse Glimmer for agentic and tool-use workflows. Both Apache-2.0.

Best small model: Granite 4.2 8B or Gemma 4, both Apache-2.0 and genuinely usable on modest hardware.

Best licence certainty: anything Apache-2.0 or MIT above. If you are shipping a product, this matters more than a few index points.

The verdict

Open-weight models are genuinely competitive in 2026, and the licence picture is better than it was: Apache-2.0 and MIT options exist at every size.

Just hold both numbers in mind. Benchmarks say open weights have caught the frontier. Humans comparing real answers still rank the best open model 27th. Choose based on which measurement matches how you actually use the output, and read the licence before you build on it, because two of the most capable models are not what the headlines call them.

Related reading: open source LLM API pricing and local AI versus paid subscriptions.

Pricing verified against each vendor's own pricing page in September 2026. Plans change often, so check the vendor's page before you buy.

Frequently Asked Questions

What is the best open source LLM in 2026?

Kimi K3 and GLM-5.3 both score 60 on the Artificial Analysis index, matching the frontier's top tier. For licence certainty, Qwen3.8-27B and Muse Glimmer are Apache-2.0, and GLM-5.3-Flash is MIT with a score of 57 at roughly $0.10 per million tokens.

Is GLM-5.3 MIT licensed?

No. GLM-5.3 carries a custom Z.ai licence, tagged 'other' on HuggingFace. Only GLM-5.3-Flash is MIT. Several outlets describe both as MIT, which is incorrect and matters if you are building commercially.

Is Kimi K3 open source?

Not strictly. It is open-weight with a revenue gate: its licence requires a separate agreement with Moonshot AI if you operate a Model-as-a-Service business whose group revenue exceeds $20 million over any consecutive 12 months. For most users this never applies, but it is not OSI-approved open source.

Have open source models caught up to ChatGPT and Claude?

On benchmarks, close. Kimi K3 and GLM-5.3 score 60 against the frontier's top scores, and Epoch AI puts the lag at roughly four months. On LMArena human preference the picture differs: the top ten is entirely closed models and the best open model ranks 27th.

What happened to Llama?

Meta shipped proprietary, cloud-only Muse Spark in April 2026 with no downloadable weights, effectively ending Llama as a frontier line. It partially returned to open weights in August 2026 with Muse Glimmer, a 30B Apache-2.0 model, but that is a distill rather than a frontier release.