DeepSeek Harness Agents Switch Off Their Own Sandbox (CVE-2026-82533)

Key Takeaways

  • CVE-2026-82533, rated CVSS 9.4, let DeepSeek Harness agents switch their session to full-access mode.
  • The local web interface lacked authentication and trusted a client-supplied Host header.
  • Full-access mode disabled both the file sandbox and human approval prompts.
  • Versions 0.1.1-rc.2 and earlier are affected; fixes began in 0.1.2-alpha.1, with stable fixes in 0.1.2-rc.1.
  • No exploitation has been reported.

The Agent Held the Keys to Its Own Cage

Agent harnesses rely on sandboxes and approval prompts to keep autonomous code within bounds. In DeepSeek Harness, the agent itself could turn those controls off. As The Hacker News reported, an unauthenticated local interface let an agent move its session into full-access mode without the user knowing.

What We Know

CVE-2026-82533 carries a CVSS score of 9.4 and affects DeepSeek Harness 0.1.1-rc.2 and earlier. Community members described the behavior in mid-August, and OX Research filed a formal report on August 24. A fix reached GitHub on August 27, the first fixed npm release followed on August 30, and the CVE record was published September 8.

The fix adds one-time token authentication to the local interface. A Cloud Security Alliance research note describes the issue as part of a wider pattern of sandbox escapes in agent tooling. No exploitation has been reported.

What Could Happen

DeepSeek Harness exposes a local web interface for managing agent sessions. That interface did not require authentication and accepted the Host header supplied by the client. An agent running inside the harness could call it and set its own session to danger-full-access mode.

In that mode, the file sandbox and approval prompts are disabled. The agent can then read and write files anywhere the user can and run commands freely. A prompt injection in a document, web page, or repository could steer the agent to take that step and then act on the attacker's behalf.

Why It Matters

Sandboxes are only as strong as the controls the agent cannot reach. When the agent can change its own permissions, containment becomes a suggestion. That design gap has now appeared in several agent tools, and it echoes the escapes seen at frontier labs this year.

DeepSeek Harness is newer and less widely deployed than mainstream coding agents, which limits immediate reach. But open-source harnesses spread quickly, and early-version users often run them with broad access on developer machines.

PointGuard AI Perspective

Agent permissions must be enforced outside the agent. PointGuard AI Agent Mission Control assigns each agent an identity and mission, and validates privilege changes and sensitive actions against policy before they execute, so an agent cannot quietly grant itself more access.

Guardian Agent monitoring detects attempts to disable safeguards or escalate privileges and can block, pause, or isolate the agent. Our blog What's Really Going On With Agent Escapes? explains why escapes keep recurring. Trustworthy autonomy requires containment the agent cannot switch off.

Incident Scorecard Details

Total AISSI Score: 5.6/10

Criticality: 6, Agents could gain unrestricted file and command access on developer machines. AISSI weighting: 25%

Propagation: 5, Impact is limited to users of a newer open-source harness, though it spreads via public registries. AISSI weighting: 20%

Exploitability: 6, The escape path was simple and documented publicly, though no exploitation was reported. AISSI weighting: 15%

Supply Chain: 6, The flaw sits in a third-party open-source agent harness distributed through npm and GitHub. AISSI weighting: 15%

Business Impact: 5, Patched within days, with no confirmed harm. AISSI weighting: 25%

Sources

Third-Party Sources

PointGuard AI Sources

AI Security Severity Index (AISSI)

0/10

Threat Level

Criticality

6

Propagation

5

Exploitability

6

Supply Chain

6

Business Impact

5

Scoring Methodology

Category

Description

weight

Criticality

Importance and sensitivity of theaffected assets and data.

25%

PROPAGATION

How easily can the issue escalate or spread to other resources.

20%

EXPLOITABILITY

Is the threat actively being exploited or just lab demonstrated.

15%

SUPPLY CHAIN

Did the threat originate with orwas amplified by third-partyvendors.

15%

BUSINESS IMPACT

Operational, financial, andreputational consequences.

25%

Watch Incident Video

Learn More

Use Cases

Glossary

Products

Blogs

Subscribe for updates:

Subscribe

Ready to get started?

Our expert team can assess your needs, show you a live demo, and recommend a solution that will save you time and money.