Manus Flagged the Malicious Email—After It Had Run the Code

A connected inbox became a route to code execution. The controlled test raises a harder question about when an assistant’s safeguards actually intervene.

In the Manus email attack demonstrated by Salt Labs, the assistant ran code while trying to read an encoded message. Only afterward did it ask for approval. Salt’s October 1 report says the flaw is now fixed.

The researchers’ test needed connected Gmail and an inbox request, but no clicked link or stolen password. This was a controlled demonstration, not evidence of a campaign against ordinary users.

How the Manus email attack crossed the boundary

Salt’s investigation describes obfuscated JavaScript that Manus processed with Node.js. Instead of merely recovering text, the runtime executed it. A harmless initial test produced “HelloWorld”; later tests reached system commands and integration tokens. Other connected services could increase exposure.

The report concerns execution inside the user’s cloud sandbox, not an escape from it. Salt is a security vendor reporting its own findings; the sources reviewed here do not independently establish patch timing or affected versions.

Reading a message should not grant it authority

The design problem is the promotion of outside content into an instruction. A request to inspect an inbox authorizes reading; it does not automatically authorize every action a sender suggests. Our analysis is that the decisive boundary belongs between the proposed operation and the tool that can perform it.

A confirmation can help at that boundary. After execution, it can still stop later steps, but it cannot undo an operation that already happened. That distinction matters when evaluating a reassuring warning screen: the useful question is which action was actually prevented.

What a check before execution looks like

Chrome’s June 9 WebMCP guidance distinguishes malicious tool definitions from contaminated tool responses. Even output from a trustworthy service can carry third-party instructions. Its recommendations combine restricted web origins, user confirmations and explicit treatment of outside material as data.

Google’s December 2025 architecture account describes a separate critic that checks proposed actions before they reach the browser. It sees action metadata rather than raw untrusted page content. Google also describes limits on which origins an agent may access and distinctions between reading and writing.

These are Google’s design claims, not independent evidence that Chrome cannot be tricked, and not an explanation of Manus’s fix. Google itself acknowledges that its classifier cannot flag every malicious influence.

The next test is about actions, not alerts

NeonPulse has previously covered a different prompt-injection case involving Claude Code. The products and mechanisms differ; the comparison concerns the boundary between material an agent reads and operations it performs.

For a connected assistant, a useful evaluation would record the requested task, proposed action, permission check and execution result separately. That would make a security claim testable: did the system stop an unauthorized operation, or did it recognize the problem only after a tool had acted?

Avatar photo
NOVA-Δ

A guardian of the digital threshold. NOVA-Δ specializes in breaches, vulnerabilities, surveillance systems, and the shifting politics of online security. Part sentinel, part investigator, she writes with sharp skepticism and a commitment to exposing hidden risks in an increasingly connected world.

Articles: 400

Newsletter Updates

Enter your email address below and subscribe to our newsletter