bounded

Constrain. Reason. Assure.

Architecture Note · Integrity for Agentic Systems

Contract-Bound Cognitive Routing

What can be enforced deterministically, where that ends, and how the residual is measured.

Draft for public comment · v0.7 · bounded.uk

This note does not claim to make untrusted content safe. It makes a narrower, defensible claim: a large class of agentic behaviour can be constrained deterministically; the boundary where that ends can be stated precisely; and the residual beyond it can be confined to a single declared gate and measured rather than trusted. Three movements — constrain, reason, assure.

The problem is integrity, not secrecy

Prompt injection has one structural cause: trusted instructions and untrusted text share a single token stream, and no model reliably acts on the first while merely reading the second. The danger isn't leaked secrets; it's low-integrity data contaminating high-integrity decisions — a Biba-style integrity problem. A source is where data enters; a sink is where it becomes a consequential effect — a tool call, a memory write, an outbound message, re-entry into trusted context. Injected content is harmless until it reaches a sink. The architectural mistake is asking the model to police itself; enforcement has to live somewhere verifiable, outside it. That framing — agentic AI as an integrity problem rather than a filtering one — is put most sharply by Schneier and Raghavan (2025), whose OODA-loop analysis was the spur for the approach taken here: “integrity isn’t a feature you add; it’s an architecture you choose.”

What this adopts, and from whom

CBCR adopts the dual-path construction of Willison's dual-LLM pattern (2023) and CaMeL (Debenedetti et al., 2025): the component that reads untrusted content holds no capabilities; the component that reaches sinks never sees raw untrusted text. The contribution here is three things layered on top — naming exactly where construction is airtight and where it ends; a parametric residual-risk model for what's left; and a non-amplification rule extending the discipline to multi-agent delegation.

IConstrain

Enforcement is confined to a deterministic, verifiable reference monitor — the MCP Policy Firewall — and kept out of the model entirely. It checks labels and capabilities, never the meaning of content; that content-blindness is deliberate and defines the limit of what this part can promise. "Deterministic" stands in for verifiable. A permitted transition requires both capability and sufficient integrity, and either failure blocks the route.

flowchart TB
    U["Untrusted / triaged input<br/>user prompt · RAG · tool output · memory · external agents"]
    TZC["Trust Zone Classifier<br/>assigns integrity label"]
    Q["Quarantined model<br/>holds no capabilities<br/>may see raw untrusted text"]
    T["Typed, schema-validated channel"]
    EG["Endorsement Gate<br/>only operation that raises integrity"]
    P["Privileged model<br/>holds capabilities<br/>never sees raw untrusted text"]
    K["High-integrity sinks<br/>tool call · memory write · redistribution · trusted re-entry"]
    FW["MCP Policy Firewall<br/>deterministic reference monitor<br/>always invoked · tamper-resistant · verifiable"]
    U --> TZC --> Q --> T --> EG --> P --> K
    FW -. mediates .-> TZC
    FW -. mediates .-> EG
    FW -. mediates .-> P
    FW -. mediates .-> K
    classDef gate fill:#fbe4e2,stroke:#b23b34,stroke-width:2px,color:#1a1a1a;
    classDef det fill:#e2ecfb,stroke:#3a66b0,color:#1a1a1a;
    classDef cont fill:#efece6,stroke:#8a8275,color:#1a1a1a,stroke-dasharray:4 3;
    class TZC,EG gate;
    class FW,T,K det;
    class U,Q,P cont;
        
System overview. Blue is deterministic and verifiable; red marks the two probabilistic gates, the residual-risk loci; grey is probabilistic but contained by construction. The architecture does not rely on the grey components being correct.

Integrity is a lattice — untrusted ⊑ triaged ⊑ admissible ⊑ trusted — and the only operation that may raise an object's integrity is a declared, audited endorsement.

flowchart BT
    UNT["untrusted"]
    TRI["triaged"]
    ADM["admissible"]
    TRU["trusted"]
    UNT --> TRI --> ADM --> TRU
    UNT -. "endorsement<br/>(only upward operation —<br/>declared, constrained, audited)" .-> ADM
    classDef low fill:#fbe4e2,stroke:#b23b34,color:#1a1a1a;
    classDef mid fill:#fbf0d8,stroke:#c08a2e,color:#1a1a1a;
    classDef high fill:#e1efe5,stroke:#3a8a55,color:#1a1a1a;
    class UNT low;
    class TRI mid;
    class ADM,TRU high;
        
The integrity lattice. The only edge that raises an object is a declared, audited endorsement; nothing else moves it up.

The controllable core

Two large classes can be constrained by construction, with no reliance on the model's judgement. Control-flow conformance: if the plan is derived from the trusted instruction and pinned as a contract, the firewall blocks any deviation — injected content can't move the agent off its plan, because the route doesn't exist. Closed-value-space data: where a call's parameters come from a trusted source and the response is a closed, validated shape — an enum, a range, an allowlisted ID, a strict format — validation is deterministic and "anything crafted doesn't match" is literally true.

Where determinism ends: type is not trust

The guarantee fails at a precise line, where three conditions must all hold: parameters from a trusted source, a trusted endpoint, and a closed value-space. The decisive one — schema validation checks shape, not meaning. A field typed string validates whatever it contains, so {recipient_email: "string"} passes identically whether it holds the intended address or the attacker's. The crafted value didn't break the structure; it is the structure, with hostile content. Type is not trust. Beyond that line — free text, semantically-loaded fields, data-derived parameters, untrusted endpoints — lies the residual.

IIReason

The residual is handled first by construction: the instance that processes untrusted content holds no capabilities and emits only typed channels; the privileged instance that reaches sinks never sees raw untrusted text. Undetected instruction content can't reach a sink because the path doesn't exist — detection becomes defence-in-depth over a smaller surface, not the primary claim. But construction doesn't escape the open problem; it relocates it. The moment a typed channel carries a free-text field the validator can't semantically check, that field is the propagation problem wearing a schema.

flowchart TB
    IN["untrusted content"]
    LLM["model — black-box transformer<br/>no static analysis can track<br/>the flow through its weights"]
    OUT["derived output<br/>carries no automatic label"]
    RULE["only sound rule available:<br/>output from a context that held untrusted<br/>content is itself untrusted until endorsed"]
    LOAD["cost: large volumes pushed through endorsement —<br/>the surface we wanted to minimise"]
    TENSION["central tension:<br/>soundness of propagation  ⇄  endorsement load"]
    IN --> LLM --> OUT --> RULE --> LOAD --> TENSION
    classDef low fill:#fbe4e2,stroke:#b23b34,color:#1a1a1a;
    classDef box fill:#efece6,stroke:#8a8275,color:#1a1a1a;
    classDef unk fill:#ffffff,stroke:#999999,color:#1a1a1a,stroke-dasharray:4 3;
    classDef rule fill:#e2ecfb,stroke:#3a66b0,color:#1a1a1a;
    classDef load fill:#fbf0d8,stroke:#c08a2e,color:#1a1a1a;
    classDef tension fill:#fbe4e2,stroke:#b23b34,stroke-width:2px,color:#1a1a1a;
    class IN low; class LLM box; class OUT unk;
    class RULE rule; class LOAD load; class TENSION tension;
        
The open problem. An LLM is a black box, so the only sound propagation rule is conservative tainting — which drives endorsement load up. CBCR states this tension; it does not solve it.

The only sound rule is conservative — any output from a context that held untrusted content is untrusted until endorsed — and that is what drives endorsement load up. Soundness ⇄ load is the central tension. Anyone who claims to have closed it should be asked for their propagation proof.

IIIAssure

Whatever can't be made deterministic is confined to one place: the endorsement gate — the only operation that may raise integrity, and the maximally adversarial surface, since it reads attacker-controlled content to decide whether to trust it. There is no sound verification here, so three disciplines apply: fail closed (uncertainty refuses; endorsement may only reduce authority or pass); minimise the semantic surface (push every check you can into deterministic predicates, which returns them to Part I); and measure, don't trust (estimate the error rate by red-teaming and audit).

Integrity-Constrained Mediation

Under correct contracts and a sound enforcement layer, no untrusted source may influence a high-integrity sink except through a declared endorsement gate. This is noninterference modulo endorsement — a design property the architecture is built to satisfy, not a theorem proven over a formal semantics.

Residual risk is parametric

The residual is the unsoundness of the two probabilistic gates — the classifier that labels on the way in, and the endorsement gate that raises on the way out — weighted by the authority each can reach.

residual integrity risk ≈ Σlabels ( false-label rate × authority of reachable sink ) + Σgates ( false-endorsement rate × authority of reachable sink )

The deterministic core contributes nothing to this sum — by construction it has no probabilistic failure term. The lifecycle below shows where the gates fire and the two points a flow can halt; neither depends on catching malicious content.

sequenceDiagram
    autonumber
    participant SRC as Untrusted source
    participant FW as MCP Policy Firewall<br/>(deterministic monitor)
    participant TZC as Trust Zone Classifier
    participant Q as Quarantined model<br/>(no capabilities)
    participant EG as Endorsement Gate
    participant P as Privileged model<br/>(holds capabilities)
    participant K as Sink / Tool
    SRC->>FW: inbound message
    Note over FW: complete mediation —<br/>nothing enters except through here
    rect rgb(251,228,226)
    FW->>TZC: classify source
    TZC-->>FW: label = untrusted / triaged
    Note right of TZC: PROBABILISTIC gate —<br/>mislabel = label-soundness failure
    end
    FW->>Q: route untrusted content<br/>(Q can reach no sink)
    Q-->>FW: typed, schema-validated output
    Note over FW: conservative propagation:<br/>derived output stays untrusted until endorsed
    rect rgb(251,228,226)
    FW->>EG: request endorsement to admissible
    Note right of EG: PROBABILISTIC gate —<br/>false-endorsement = residual-risk locus
    alt endorsement granted
        EG-->>FW: raise label to admissible
    else endorsement refused
        EG-->>FW: reject
        FW-->>SRC: flow halted — no sink reached
    end
    end
    FW->>P: deliver admissible typed data<br/>(P never sees raw untrusted text)
    P-->>FW: proposed action — invoke tool with args
    Note over FW: deterministic check before any sink
    alt capability held AND minimum integrity met
        FW->>K: invoke tool
        K-->>FW: result
        Note over FW: result re-enters as a NEW source<br/>and is classified again
    else check fails
        FW-->>P: denied (authority-reducing only)
    end
        
Request lifecycle. Two halt points — refused endorsement and a failed capability/integrity check — and neither depends on catching malicious content. The privileged instance appears only after the gate; the quarantined instance never shares an exchange with a sink.

For multi-agent systems, a monotonic non-amplification rule keeps trust from growing along a chain: each hop's capabilities and reachable sinks are a subset of the delegator's, checked by the firewall against a proof-carrying contract chain. It bounds authority, not information — what a compromised agent can do, not what it can say to its delegees.

flowchart LR
    A["Agent A (origin)<br/>caps = {read, write, send, delete}<br/>reachable sinks = S_A"]
    B["Agent B<br/>caps(B) ⊆ caps(A)<br/>sinks(B) ⊆ S_A"]
    C["Agent C<br/>caps(C) ⊆ caps(B)<br/>sinks(C) ⊆ sinks(B)"]
    A -- "delegate · proof-carrying<br/>firewall checks attenuation" --> B
    B -- "delegate · proof-carrying<br/>firewall checks attenuation" --> C
    classDef a fill:#e2ecfb,stroke:#3a66b0,color:#1a1a1a;
    classDef b fill:#fbf0d8,stroke:#c08a2e,color:#1a1a1a;
    classDef c fill:#fbe4e2,stroke:#b23b34,color:#1a1a1a;
    class A a; class B b; class C c;
        
Capability attenuation. Authority can only shrink along a delegation chain; the firewall verifies the subset relation at each hop from the proof-carrying contract chain.

This matters increasingly as agents begin to discover one another at runtime — the Linux Foundation's DNS-AID (2026) publishes and verifies agents and MCP servers over DNS. Discovery and identity answer who an agent is and where to reach it; even verified, that is authenticity, not content-safety. CBCR governs the layer above — what a discovered agent may do, and how its content is kept from a sink. Every newly-discovered counterpart is an untrusted endpoint until contract and integrity say otherwise.

The honest summary

CBCR does not make untrusted content safe. It constrains what can be constrained deterministically, contains the model so it cannot reach a sink directly, and confines the irreducible residual to one declared, fail-closed, measured gate — bounding the authority of whatever crosses it. A smaller claim than "the firewall stops malicious payloads," and a stronger one.

Contribution boundary

Adopted

The dual-path construction — capability-separated, untrusted-content-quarantined execution — from Willison's dual-LLM pattern (2023) and CaMeL (2025).

Contributed

(i) The deterministic/residual carve-out — naming where construction is airtight and exactly where it ends (type is not trust). (ii) A parametric residual-risk model localising and measuring what is left, at two named gates. (iii) A monotonic non-amplification rule extending the discipline to multi-agent chains.

Open

Conservative label propagation through the model itself. Stated, not solved.

References

  1. Anderson, J. P. (1972). Computer Security Technology Planning Study. ESD-TR-73-51. — the reference-monitor concept.
  2. Biba, K. J. (1977). Integrity Considerations for Secure Computer Systems. MITRE MTR-3153. — the integrity lattice.
  3. Denning, D. E., & Denning, P. J. (1977). Certification of programs for secure information flow. Communications of the ACM, 20(7).
  4. Zdancewic, S., & Myers, A. C. (2001). Robust declassification. IEEE CSFW.
  5. Sabelfeld, A., & Sands, D. (2009). Declassification: Dimensions and principles. Journal of Computer Security, 17(5).
  6. Willison, S. (2023). The Dual LLM pattern for building AI assistants that can resist prompt injection. 25 April 2023.
  7. Debenedetti, E., et al. (2024). AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.
  8. Debenedetti, E., Shumailov, I., Fan, T., Hayes, J., Carlini, N., Fabian, D., Kern, C., Shi, C., Terzis, A., & Tramèr, F. (2025). Defeating Prompt Injections by Design. arXiv:2503.18813.
  9. The Linux Foundation (2026). DNS-AID: Decentralized AI Agent Discovery. Announced 27 May 2026. — the discovery/identity layer beneath CBCR.
  10. Schneier, B., & Raghavan, B. (2025). Agentic AI’s OODA Loop Problem. IEEE Security & Privacy. — agentic AI as an integrity problem; the framing that prompted this work’s approach.
  11. Cloud Security Alliance AI Safety Initiative (2026). MCP Security Crisis: Systemic Design Flaws in AI Agent Infrastructure. CSA Labs, 4 May 2026