Inside Muse: A Product and Technical Teardown
Muse is Meta's bet that people will hand an AI agent real authority over their accounts, money and time. This teardown covers the product decisions that make delegation workable, and the architecture underneath: a VM per user, a single permission authority called Sentinel, credentials the model never sees, and kernel-level taint tracking.
Meta launched Muse on September 8, 2026. It is a personal agent, not a chatbot. You give it a task, it works on a cloud machine of its own, and it comes back when the task is done or when it needs your sign-off. Ten days after launch it was the No. 1 app on the US App Store.
This teardown looks at Muse in two layers. The product layer answers one question: what does it take for people to delegate real authority to software? The technical layer answers the other side of it: how do you make that delegation safe when the agent itself can be manipulated? Muse's answers to the two fit together closely, and they meet at one narrow decision: when to ask you for permission.
Key takeaways
- The unit of work is a delegated task, not a reply. One persistent thread, work that continues after you close the app, and notifications when something finishes or needs you.
- Approval is the main interface. Five grant scopes, an activity log, a live view of the agent's browser, and approval dialogs that bypass the conversation entirely.
- The business model is the transaction. A free tier and two subscriptions today. A small cut of the purchases Muse completes is the stated long-term plan.
- The architecture assumes the agent will be attacked. Each user's VM is split into an agent cell and a privileged host domain. One component, Sentinel, controls every outbound request and connector action. The model never sees a real credential.
- Taint tracking ties it together. The kernel records whether a process has read your data. Requests that cannot carry your data run without interrupting you; requests that can must be approved. That is how Muse manages to be both quiet and safe.
- The open problem is the second party. Muse's security model has one principal. Purchases and meetings have at least two.
Muse at a glance
| Dimension | Detail |
|---|---|
| Launch | September 8, 2026, in the US. Built by Meta Superintelligence Labs |
| Surfaces | iOS, Android, muse.ai and WhatsApp. Announced: AI glasses, a Mac app that can operate other apps, and Muse Charm, a pocket-sized voice device |
| Model | Muse Spark. Meta's API lists muse-spark-1.3 with a 1,048,576-token context window and parallel tool calls |
| Compute | A dedicated Linux VM per user, with its own browser and file system |
| Pricing | Free tier of about 100 million tokens a week; Power at $20 a month; Maximum at $100 a month |
| Revenue plan | Subscriptions now; a small cut of completed transactions later |
| Traction | No. 1 on the US App Store (September 18) and Google Play (September 19); 2.3 to 4.3 million downloads by September 25, depending on the tracker |
Part 1: The product
A task, not a reply
Most AI assistants are built around an exchange: you ask, it answers, the session ends. Muse is built around delegation. Each user has one persistent conversation, plus side chats for larger projects. Work continues after you close the app. Results, reminders and approval requests arrive in the thread when they are ready, and Muse decides whether a result deserves a notification at all.
That choice shapes the rest of the product. If a task outlives the session, the system needs durable state, background execution, a record of what happened while you were away, and a way for you to see the work in progress. Muse provides each of these, including an activity log and a live view of the browser it is driving.
Approval is the interface
When an agent works unattended, most of your interaction with it is about authority: what it may do, and what it has done. Muse treats this as the core of the product, not as a settings page.
- Approvals are scoped. A grant can be one-time, for the session, for the task, time-bounded, or permanent. Sentinel, described below, decides which options to offer and checks that later calls stay inside the granted scope.
- Approvals are out of band. Sentinel sends the dialog straight to the app, outside your conversation with Muse, and your answer goes straight back to Sentinel. The model cannot word the request, forge it, or talk you into it.
- Money always asks. When the browser detects a checkout page on a site where your payment details are stored, Muse asks you to approve the exact purchase, every time.
- Connector access is granular. You choose which apps Muse connects to and how much access each one gets. Meta's example is removing the access to Gmail settings that normally comes bundled with Gmail.
Meta's stated goal is to put friction where consent matters and let routine work flow, and it says it expects to tune that balance as it learns from real users. That tuning problem is the central design problem of the product, and it returns in Part 2.
Proactivity builds the habit
Muse can start a conversation without being asked. The app has a feed of suggestions, an Ideas tab with example tasks, and a Goals tab that tracks longer plans in areas such as health, finance and career. Its memory files are visible and editable.
An agent that waits for instructions gets used occasionally. An agent that brings you things becomes part of the day, and standing goals give it something to work on between your requests.
The business model is the transaction
The free tier is generous. A Gizmodo writer spent a week having Muse build games, music and animations and used about 11% of the weekly allowance. Paid plans add headroom: Power at $20 a month and Maximum at $100.
The long-term plan is different. Zuckerberg has said Muse should pay for itself by making and saving people money, with Meta taking a small cut of the transactions it completes, possibly paid by the business on the other side rather than by the user. The commercial pieces are already in place: a wallet, Stripe Link at launch with Shop Pay to follow, and partners announced at Connect including Walmart, Best Buy, Sephora, Ulta, Expedia and Instacart.
Meta's advertising business earns money from attention between people and businesses. Muse aims to earn it from completed intent: the same marketplace, entered from a different side. Meta also says Muse conversations and VM data are not shared with its ad systems, and that users can opt out of having their interactions used for training.
Distribution Meta already owns
Muse runs inside WhatsApp, and Meta has announced it for its AI glasses, where you will be able to say the agent's name and have it act on what you are looking at. Muse Charm puts the agent in a pocket-sized device with a real-time voice model, and a Mac app will let it operate other applications with your permission. Meta started running house ads for Muse the day after launch. Within ten days Muse had most of Meta's in-house ad placements, and by September 22 it was among the top ten brands by ad spend.
Part 2: The architecture
Meta's security post for Muse starts from a blunt assumption. Any agent like this will make mistakes, and it will sometimes be attacked through the data it reads. So the system is designed to limit the damage rather than to trust the model. The mental model, in Meta's words, is "two isolated security domains on one box, not an LLM powered agent with root."
The VM: a cell for the agent, a host for authority
Each user gets a Linux VM. Inside it, the agent harness (internally called Hatch), your workspace and files, and every binary Muse runs live in a runtime cell, a systemd-nspawn container. Root inside the cell maps to an unprivileged user on the host, so taking over the cell does not mean taking over the machine. The cell has its own filesystem and network interface, filtered system calls (no io_uring), and restricted kernel capabilities (no CAP_SYS_PTRACE, no CAP_NET_ADMIN).
Everything that holds authority runs outside the cell, in the host domain, as separate services: Sentinel, the credential service authd, sandboxed connector workers, the Postgres database that holds durable state, and a safety service that runs its own models and classifiers over traffic to and from the core model.
The split exists because the cell is where untrusted data gets processed. An attacker who takes control of the agent ends up as root inside a box that holds no credentials and has no route out except through Sentinel.
Sentinel: one gate for everything that leaves
Sentinel is the sole authority over connector actions and over all network egress. Traffic from the cell reaches it through a forward proxy enforced with Linux networking controls (user namespaces, virtual Ethernet boundaries and eBPF filtering), so the agent cannot route around it.
For each request, Sentinel can see the hostname, the resolved and final IP addresses, the port, protocol, HTTP method and path, and the decoded request itself. It also checks that a public-looking hostname does not resolve to private infrastructure. Its answer is one of three: allow, deny, or ask the user.
Credentials the model never sees
OAuth tokens for the services you connect are stored in your VM, not in central Meta infrastructure, and they are held by authd. Code in the cell only ever receives a surrogate token, a placeholder minted by authd. Once Sentinel has authorized a request, it replaces the surrogate with the real credential at the network boundary.
This is least privilege taken literally. The model does not need to see an API key, so it never does, and it cannot leak a key it never had.
Connectors split in two
A common shortcut is to run connector tools in the same environment as the agent. The risk is that a prompt injection persuades the agent to modify the tool's code. Muse splits each built-in connector in half. A small command-line client in the cell only parses arguments and passes typed values and file descriptors over a Unix socket. The business logic runs in a sandboxed worker outside the cell, identified by its cgroup and limited to an explicit list of credentials. A calendar worker cannot obtain an email credential by changing a request parameter.
Authority is divided three ways. Privilege separation decides where credential-capable code runs. Authd decides which credentials a caller may receive. Sentinel decides whether the action may happen at all.
Tainted egress: deciding when to ask
This is the mechanism that connects the architecture to the product.
Every tool process starts clean. Once it reads user data, it is marked as tainted. Clean requests that fit a narrow auto-allow policy, and pass checks on the outbound URL, go through without interrupting you. Tainted or unverifiable processes lose auto-allow and go through the normal approval flow. The tracking is done with eBPF, small verified programs that run inside the Linux kernel, attached to cgroups for network interception and to Linux Security Module hooks.
Meta says the purpose is to balance the signal-to-noise ratio of approvals.
The common design gates on the type of action: reads go through, writes ask. It fails in both directions. It asks too often, because every write prompts you, and people who are asked about everything learn to approve without reading. It also asks too rarely, because a read can carry your data out. Fetching a URL with your account details in the query string is a read, and it leaks.
Gating on data flow asks a better question. Not "is this a write?" but "could this request carry your data?" Muse can then stay quiet for work that cannot leak anything and insist on approval for work that can. The price is control of the operating system the agent runs on, which is one reason every user gets a dedicated VM.
Defense in depth around the model
The rest of the design assumes that any single layer can fail.
- Prompt injection. The model is trained to resist injection, and Meta tracks this with dedicated evals; it describes Muse Spark 1.3 as close to state of the art on this measure. The harness labels all external content as untrusted. An ensemble of injection classifiers, trained on real-world attacks and on Meta's own agentic red-teaming, screens all external data entering the model's context. Deterministic boundaries sit underneath all of it.
- Email. Access to your inbox is access to password resets for most of your other accounts. The email connector filters out one-time codes, password-reset links and login magic links, using deterministic rules plus a classifier.
- Trust in the operator. Muse Confidential VM, planned for later this year, is meant to prevent Meta itself from reading your VM, cryptographically and verifiably. Meta says it is already running with a small group of trusted testers, that it has begun sharing the design and source code with external auditors, and that the launched system will be continuously auditable by anyone.
The model
Muse Spark is also available outside Muse, through Meta's Model API. The documentation lists a 1,048,576-token context window, parallel tool calls with streamed arguments, and reasoning that carries across turns, all of which suit long-running tasks. The architecture around it is built so that none of the safety properties above depend on the model behaving well.
What is hard to copy
The visible surface of Muse is the easiest part to reproduce: a chat thread, a set of connectors, reminders, approval cards. What is harder to copy sits underneath and around it.
- Calibrated consent at scale. Taint tracking requires control of the operating system, so Muse runs a VM per user, across millions of installs.
- Distribution. WhatsApp, glasses, a voice device, and Meta's own ad inventory.
- Commerce. A payments stack, merchant partners, and a business model built on the transactions themselves.
- Injection resistance in the model itself, trained, measured with evals, and backed by an independent classifier ensemble.
Open questions
How much friction is right? Meta says it will tune the balance between asking and auto-allowing with real usage. Too many prompts and people approve without reading. Too few and the safety case weakens.
Can users trust the operator? The Confidential VM is Meta's answer to whether Meta can see inside your VM. Until it ships, that answer rests on policy rather than cryptography.
What happens when there are two principals? Muse's security model, as published, is built around one person: one VM, and one Sentinel deciding on that person's behalf. The business model points somewhere else. A transaction has a merchant on the other side, and a meeting involves other people's calendars. Once those parties send agents of their own, no single Sentinel sees the whole exchange. Whose policy governs a conversation between two agents? Does a restriction travel with data that one person's agent hands to another's? Who audits the exchange?
These are the questions we work on at Systemind. We set out the research agenda in Open Challenges in Networked Agent Collaboration, and our direction in Organizational Intelligence.
Sources
- Meta, Introducing Muse, September 8, 2026.
- Meta Superintelligence Labs, Security and Safety for AI Agents: Our Approach with Muse, September 8, 2026.
- Meta, The Biggest News From Connect 2026.
- Meta, Model API documentation.
- TechCrunch, Meta debuts its Muse AI agent. Will consumers trust it?, September 8, 2026.
- TechCrunch, Meta is putting its muscle behind Muse as the AI app takes off, September 25, 2026.
- Sources (newsletter), Mark Zuckerberg on Muse, Meta's biggest AI bet yet.
- Runtime Wire, Meta says Muse will make money by taking a cut of transactions.
- Gizmodo, Meta's Muse let me waste a mind-boggling amount of free compute on nothing in particular.
- TestingCatalog, Meta introduces Muse as a proactive personal agent.