
Meta has returned to releasing AI models that anyone can download and run. On 10 August 2026, Meta Superintelligence Labs released Muse Glimmer, a 30-billion-parameter model whose weights are free to download under the Apache 2.0 licence. It is designed to run AI agents, meaning software that carries out multi-step tasks using tools, on a Mac or PC with a single consumer graphics card. For a small or medium business, that makes a capable assistant that never sends data to an outside service a realistic option, with real limits attached.
What happened
Meta published the release in a post introducing Muse Glimmer, describing it as "optimized for always-on local agent workflows". The model's weights, the trained numbers that make up the model, are on Hugging Face, a public library for AI models, along with developer documentation.
Three points stand out in the announcement.
- It is open. The weights are released under Apache 2.0, which Meta calls a permissive licence. The Apache 2.0 licence text grants a perpetual, royalty-free right to use, modify and distribute the work, with conditions such as keeping licence notices, and it provides the work without warranties.
- It is small enough for local hardware. Meta says the full model would need more than 55 GB of memory, so it supplies compressed versions (a technique called quantisation) that bring the language model under 20 GB and let the whole system fit in 24 GB or 32 GB of graphics memory.
- It is built for agents. Meta lists local agents, function calling (the model asking other software to do something), local coding and using the model to grade other AI outputs. It accepts text and images, and Meta says it is trained to diagnose a failed tool call and retry instead of stopping.
Muse Glimmer was trained by distillation from Muse Spark, Meta's larger model: the small model learns to imitate the large one's outputs. The Register's report on the launch describes it as Meta's first open-weights model in more than a year and notes that Muse Spark itself remains proprietary for now.
Key details
| Detail | What the release says |
|---|---|
| Release date | 10 August 2026 |
| Licence | Apache 2.0, for commercial and research use |
| Size | About 30 billion parameters, including the image encoder |
| Input and output | Text and images in, text out. Audio is not supported. |
| Context window | 128K tokens by default, with longer contexts supported. This is how much text it can consider at once. |
| Hardware targets | Full precision: 64 GB of graphics memory. Compressed versions: 32 GB or 24 GB. |
| Speed on an Nvidia RTX 5090 | 74.9 tokens per second, or 233.4 with the bundled speed-up technique |
| Speed on an Apple M5 Max | 26.6 tokens per second, or 50.2 with the speed-up technique |
| Languages | Trained on data from more than 100 languages |
| Knowledge cutoff | 4 January 2026 |
The hardware, speed and cutoff figures come from the Muse Glimmer model card on Hugging Face, and the default context window from Meta's developer documentation. Meta says the smaller compressed version loses about 1 per cent accuracy on average across 15 benchmarks. Support in popular tools for running models locally, including Ollama, LM Studio, llama.cpp and MLX, is due "in the coming days".
Why it matters
A large United States lab is publishing open models again. Open-weights models give businesses and software developers something hosted services cannot: a copy they control. Nobody can change its price, withdraw it or alter its behaviour once it is downloaded. Artificial Analysis, an independent benchmarking firm, notes in its assessment of Muse Glimmer that every earlier open release from Meta shipped under the company's own Llama licence, and that this is the first under Apache 2.0, which places almost no restrictions on commercial use.
The performance claims are mixed, and the independent numbers are sobering. Meta compares Muse Glimmer with two models of similar size, Google's Gemma 4 and Alibaba's Qwen 3.6. On Meta's own table it is ahead of both on a tool-use benchmark called MCP Atlas (75.5 against 54.2 and 62.5) but behind Qwen on SWE-Bench Verified, a coding test (76.0 against 77.2), and on OSWorld-Verified, a computer-use test (65.9 against 75.6). Artificial Analysis found it strong for its size overall but weak on realistic knowledge work, and measured a hallucination rate of 82 per cent on its own knowledge test, against 49 per cent for the Qwen model. A hallucination is an answer that is stated confidently and is wrong. A model this size is a capable assistant, not a replacement for the largest hosted models.
Local means private, not safe by default. When a model runs on your own machine, your prompts and files do not travel to a provider. That is a genuine privacy advantage. But an agent that can read files and run tools on a work computer can also make mistakes there. Meta's model card recommends deploying the model as part of a larger system with guardrails, and specifically suggests a human confirmation step before any action that cannot be undone.
What this means for businesses
For most small and medium businesses, Muse Glimmer is not something to install this week. Running it well needs a computer with a high-end graphics card or a recent high-memory Mac, someone to set it up and keep it patched, and software around it that decides what the model is allowed to touch. The licence is free. The hardware, the setup and the upkeep are not.
Where it becomes interesting is in work that should not leave the building. An accounting practice sorting client documents, a clinic's front desk drafting letters from internal templates, or a builder searching years of quotes and plans could all use a local model for first drafts and lookups without sending that material to an outside service. The same applies to software built for you: a developer can now include a capable model inside a system, with no per-use fee to a model provider and no dependency on that provider's terms.
The cautions are practical ones. Check anything factual, given the error rates measured independently. Remember the knowledge cutoff: the model knows nothing after early January 2026 unless it is given the information. And treat a local agent like a new staff member with system access: limited permissions first, wider ones once it has earned them.
A short checklist:
- List the tasks where privacy is the reason you have not used AI so far. Those are the candidates for a local model.
- Check your hardware. A standard office laptop will not run this model at a usable speed.
- Decide who would own it: installation, updates, access rules and backups.
- Keep a person in the loop for anything irreversible, such as sending, deleting or paying.
- Compare the full cost of hardware and upkeep with a business plan from a hosted provider before deciding.
Comingwave is a technology company that provides IT consulting, custom software and cyber security services to small and medium businesses, including professional services firms that handle confidential client records. If you are weighing up where sensitive work should run, request a free first consultation. We reply within one business day.
Key takeaways
- Meta released Muse Glimmer on 10 August 2026 as an open-weights model under the Apache 2.0 licence.
- Compressed versions are designed to fit in 24 GB or 32 GB of graphics memory, so it can run on a single high-end PC or Mac.
- It is built for agent tasks such as tool use and coding, and accepts text and images.
- Independent testing found it strong for its size but prone to confident errors, so outputs need checking.
- Running a model locally keeps data in-house but adds hardware, setup and security work.
Frequently asked questions
What is Muse Glimmer?
Muse Glimmer is an AI model from Meta Superintelligence Labs with about 30 billion parameters. It is designed to run agent tasks on local hardware and was released with open weights on 10 August 2026.
What does open weights mean?
It means the trained model itself can be downloaded and run on your own equipment, instead of being reached only through the maker's online service. Muse Glimmer's weights are published under the Apache 2.0 licence, which permits commercial use.
Can Muse Glimmer run on a normal office computer?
Not a typical one. Meta's compressed versions target 24 GB or 32 GB of graphics memory, which means a high-end graphics card or a recent high-memory Mac. On lesser hardware it will be slow or will not fit.
Is a local AI model more private than a cloud service?
Prompts and files stay on your own machine, which removes one privacy risk. It does not remove the need for access controls, backups and checks on what an agent is allowed to do, and Meta itself recommends guardrails and human confirmation for irreversible actions.
Is Muse Glimmer free for commercial use?
The licence allows commercial use without a fee, subject to its conditions. The costs are in the hardware, the setup and the ongoing maintenance, and the licence provides the model without warranties.