Meta released Muse Glimmer on August 10 under the Apache 2.0 license — a 30-billion-parameter open-weight model designed to run autonomous AI agents directly on consumer hardware. It's Meta's first fully open release since the company moved beyond its open-weight Llama family with the proprietary Muse Spark earlier this year.

Muse Glimmer is a dense causal transformer with approximately 29.6 billion total parameters across 52 layers, including a dedicated ~1.8B-parameter vision encoder. It accepts interleaved text and images, supports more than 100 languages, and has a context length of 131,072 tokens. The model was trained around the sequence of operations an autonomous agent performs: plan, call tools, interpret results, continue working, and recover from errors.

At full precision, the model requires more than 55GB of memory, beyond any single consumer GPU. Meta developed 4-bit quantized variants that shrink the language-model weights to under 20GB, fitting within a 24GB or 32GB envelope alongside the KV cache, perception encoder, and speculative-decoding drafter. The K-Quant-17GB configuration fits on Nvidia's RTX 3090 or RTX 4090 (24GB VRAM), while the K-Quant-Dynamic targets the RTX 5090's 32GB. On Apple Silicon, MacBooks with 32GB or more unified memory can run the full stack.

Meta also introduced DFlash speculative decoding, which raises generation speed on an RTX 5090 from 74.9 to 233.4 tokens per second (3.1x), and on an M5 Max from 26.6 to 50.2 tokens per second (1.8x). Average accuracy degradation is just 0.2% for the 32GB variant and 1% for the 24GB variant across 15 benchmarks, according to Meta's own measurements.

In a demo video, Glimmer autonomously discovered a Home Assistant instance on the network, queried device APIs, wrote a responsive HTML/CSS/JavaScript dashboard, and deployed a local server to verify its own work.

The release arrives amid intensifying competition in the local-model space. Google's Gemma 4 31B and Alibaba's Qwen3.6-27B are strong competitors. Glimmer leads on several agentic benchmarks including MCP Atlas (75.5), DeepSearch QA (74.6), and SWE-Bench Pro (51.2), though Qwen leads on OSWorld-Verified and TerminalBench 2.1.

Most notably, Zuckerberg promised to open-source Muse Spark 1.2 — Meta's frontier proprietary model — in the near future. If delivered, it would put a U.S. flagship frontier model into open circulation for the first time.

By May 2026, Chinese open-weight models accounted for roughly 61% of all tokens consumed on OpenRouter. Glimmer represents Meta's bid to reclaim ground in the open-source AI ecosystem.