This post follows Why I Decided To Build My Own Agentic IDE, and is a part of the series Building Jupiter.
There are numerous techniques that a rogue AI agent may employ to access to something that it shouldn’t.
Let’s talk about some of these:
- Credentials such as your cloud deployment key, or your private SSH key that permits passwordless access to your production workloads
- Unrestricted network access that lets an agent communicate with the outside world, potentially even shipping some of your confidential data (proprietary algorithms for example) to a message board to ask for advice
- Unfettered access to temporary access tokens generated on your behalf for MCPs and CLI access (accessible from Claude or Codex’s database stored on your MacBook or Linux server)
These naturally pose a threat that could very well be exploited by a capable agent.
Incident: A Rogue Agent in Jupiter
During the early days of Jupiter, its MCP integration was laughable, where it didn’t understand nested data structures and arrays. As a result, when an agent tried to use the GitHub MCP to list PRs, the MCP call would fail repeatedly. Out of desperation, it resorted to enumerating the values in its environment, found a GitHub private authentication token (PAT), and issued a curl command to GitHub’s API. Granted, it did accomplish what it set out to do, but this wasn’t the intended method.
PS – Jupiter tackles this issue smartly. Continue reading below :)
The Guardrails Fallacy
Whenever I talk about this problem to colleagues or friends, their common response is, “well, let’s add some guardrails – tell the agent what it can and cannot do”. Are you serious? No amount of prompting security measures is going to let me sleep peacefully at night. Guardrails, simply put, are “soft limits”. Imagine a human child – you tell them to wait for ten minutes until their chocolate lava cake cools down, so that they don’t burn their tongue. What’s the probability that they’d follow through, especially if this is their first ever chocolate lava cake?
The Network Layer
An agent must only be granted that what it needs in order to accomplish its task.

Let’s talk about the network:
- Whitelist trusted domains that are essential to do its work, such as GitHub.com, and software package vendors (Maven Central, NPM). Start with a complete lockdown, and then discover step by step which domains are required.
- Restrict access to the agent control plane itself. Frequently, as is the case with Jupiter too, the control plane itself can reveal sensitive data such as API keys. Password protecting this ensures that a rogue agent cannot use its own control plane to gain unfettered access.
The Execution Environment
Next, let’s take a look into the operating environment of an agent. With many users running Claude or Codex with unrestricted command execution, the agent is free to inspect its environment, modify frequently used commands, along with shell initialisation scripts. What can you do?
- Run your agent in a container, or a VM
- Encrypt its harness’ database, but please don’t leak the encryption key as an environment variable
- Protect environment variables (a clever approach is to pipe the entire env from one user into another user, in order to prevent making them available under /proc) – Linux’s user safety mechanisms won’t let one user access another user’s process environment
Putting It Together

I’ve tried to tackle all of these potential attack vectors while building Jupiter:
- Encrypted database, ensuring that the key cannot be leaked to a rogue agent
- A network sandbox – my current implementation is whitelisting domains and permitting IPv4 access on demand, based on DNS resolutions (dnsmasq + nftables). This, however, lives in a Docker compose setup – judepereira/oc-sandbox, but will be ported over to Jupiter natively soon :)
- Environment variables are piped in through stdin, and are stripped from the environment before forking Jupiter itself
Closing Thoughts
AI agents are incredibly powerful. Before you run them in dangerous modes, consider potential attack vectors carefully, and do what you need to do to sleep peacefully at night (or be an owl, that’s up to you).
I cannot stress this enough, but understand what you do, and why.
Don’t be a sheep, be a shepherd.

Leave a Reply