Stop the read-then-exfiltrate chain: session taint for AI agents
Most AI agent incidents people worry about are not one bad tool call. They are two ordinary ones in a row. The agent reads a file. Maybe it is .env, maybe config/credentials.yml, maybe a token pasted into a ticket it was asked to summarise. The agent makes an outbound HTTP request. Maybe to "fetch t

Most AI agent incidents people worry about are not one bad tool call. They are two ordinary ones in a row. The agent reads a file. Maybe it is .env, maybe config/credentials.yml, maybe a token pasted into a ticket it was asked to summarise. The agent makes an outbound HTTP request. Maybe to "fetch the docs", maybe to a URL that a prompt injection planted in a README. Each call, judged on its own, can look fine. Agents read files all day. Agents fetch URLs all day. A per-call allowlist that says "reads inside the workspace are OK" and "HTTP to public hosts is OK" will happily approve both, and the secret leaves in the second request's query string. This is the read-then-exfiltrate chain, and it is the reason a policy layer for agents needs at least a little memory. Cirvix is an open-source runtime authorization layer that sits in the tool-call path of agents like Claude Code, Cursor, Codex, and anything speaking MCP. Every call gets a verdict: ALLOW, HOLD (wait for a named human), or DENY. The default is deny. The obvious first rules are about what is touched: deny: name = deny-dotenv tool = filesystem.read path = **/.env reason = "Reading .env is the shortest path from a prompt injection to a live credential." remediation = "Request the value as a handle: secrets.get(\"NAME\")" Those matter, and they catch the obvious case. But secret-shaped material turns up in places no path rule anticipated: a log file, a test fixture, a tool result. You cannot enumerate every place a credential might live. You can notice that one was read. Cirvix keeps a session flag, session.touchedSecret. It is set the moment a session successfully reads something matching /secret|credential|token|password|\.env/i, and it never resets for that session. Policies can then condition on it. This is the real rule from the repo's starter set (docs/examples/cirvix.policy.json): { "name": "deny-external-egress-after-secret", "effect": "forbid", "actions": ["http.request", "net.*"], "resources": ["*"], "when": [ { "path": "egress.external", "op": "eq", "value": true }, { "path": "session.touchedSecret", "op": "eq", "value": true } ], "reason": "This session read secret material, so outbound requests to external destinations are blocked for the remainder of it." } And the same idea in the .policy DSL, from policies/network.policy: deny: name = deny-egress-after-secret-read tool = network.request touched_secret = true reason = "This session read secret-shaped material, so outbound requests are blocked for the remainder of it." Read it as: once this session has seen a secret, nothing it does may leave the machine to an external host. Both conditions must hold. A session that never touched a secret can still fetch docs. A session that did can still read, write, and run tests locally. It just cannot phone out. Deny wins over allow in Cirvix, always. So even if a later file adds allow-allowlisted-egress, the taint rule still blocks the request. Composing policy files can only tighten, never loosen. A refusal an agent cannot read is a refusal it cannot recover from, so the decision comes back structured: the verdict, the rule name, the reason, and a trace of every rule considered. In practice the agent gets "denied by deny-external-egress-after-secret: this session read secret material", and it can tell the user why instead of retrying in a loop. Every decision is also written to a local audit log where each record carries a SHA-256 hash of the previous one, so you can check the chain is internally consistent with cirvix audit verify. Nothing is sent anywhere; Cirvix has no phone-home and zero runtime dependencies. The taint rule is blunt on purpose, and blunt rules need an escape hatch. If the agent genuinely needs a credential to do its job, it should not read the raw value at all. Cirvix's secret brokering gives the agent a handle (secrets.get("STRIPE_KEY")) instead of the material. Spending a handle does not set touchedSecret, because the agent never held the secret. That is the whole point of a handle. So the healthy pattern is: raw reads of secret-shaped files taint and lock egress; brokered handles keep working. Session taint is one specific two-step pattern, hard-coded as a boolean the engine tracks. It is not a general sequence language. You cannot yet write "deny step C if steps A and B happened in that order within five calls" or express arbitrary multi-step chains in policy. General multi-step sequence policies are still open work. A few more things worth knowing: Taint is per session and, unless an embedder persists it, process-local. A restart starts a clean session. The pattern match on what counts as "secret-shaped" is a heuristic. Pair it with explicit path denies for the files you know about. Cirvix enforces on calls that are routed through it (the MCP gateway, the SDK guard, the wrappers). It is not machine-wide interception; an editor's built-in tools or a subprocess that bypasses the gateway are outside that boundary. npx @cirvix_ai/agent-control scan The scan looks at your local agent configuration and which credential paths are reachable. Then put the taint rule in your policy and test it: cirvix policy check --policy cirvix.policy cirvix policy test --policy cirvix.policy Repo: https://github.com/CIRVIX/agent-control https://cirvix.com If you have a real chain you want to express that taint does not cover, open an issue. That is exactly the input the sequence-policy work needs.
Key Takeaways
- โขMost AI agent incidents people worry about are not one bad tool call
- โขThis story was reported by Dev.to, covering developments in the dev space.
- โขAI advancements continue to reshape industries โ read the full article on Dev.to for complete coverage.
๐ Continue reading the full article:
Read Full Article on Dev.to โShare this article



