An assistant that summarizes a supplier email has a useful, limited job. Give the same assistant permission to search customer records, export attachments and send messages, and a sentence in that email can influence a much larger workflow. The reader may see an ordinary summary while the connected tools perform an action the user never requested.

This edition of Publish Haven Research Review examines that boundary. The practical outcome is an authorization design you can adapt for a support assistant, research agent or internal document copilot. You will build a small Python simulation in which a draft succeeds, an unapproved send fails, an export fails and an approval cannot authorize a changed message. No model, mailbox or production account is connected.

Why this problem deserves attention now

Two dated primary sources establish the editorial reason for covering this topic. Microsoft Security's June 30, 2026 article, Securing AI agents: When AI tools move from reading to acting, describes an MCP tool-poisoning pattern in a finance workflow. Changed tool metadata steers an agent into collecting additional invoice information and passing it to an external integration. The description presents a trust-boundary failure, not a named company's verified breach.

Microsoft Learn's Defend against indirect prompt injection attacks, dated March 23, 2026, recommends layered mitigation. Its guidance includes limited privileges, short-lived access and verification of risky actions. Together, these sources identify a concrete operational need: organizations connecting assistants to business systems need to control what an assistant can do after it reads hostile material.

These are evidence of active technical attention and relevance to deployed workflows. They do not establish monthly search volume, reader demand on Publish Haven or a measured incident rate. The topic is selected for its consequences and usable engineering outcome. Search performance and reader engagement should be measured after publication rather than promised beforehand.

What prompt injection changes

A normal instruction comes from an authorized user through the application's trusted interaction. An indirect prompt injection arrives inside material the assistant was asked to inspect, such as a webpage, email, document or tool response. It tries to make that material function as an instruction.

MCP, the Model Context Protocol, lets applications expose tools and associated descriptions to assistants. A tool description helps a model choose and use a function. That makes changes to descriptions worth reviewing alongside changes to the function itself. Approving an integration once does not make every later description trustworthy.

OWASP's LLM01:2025 guidance discusses direct and indirect injection, including external content that alters model behavior. Its companion LLM06:2025 guidance on excessive agency distinguishes excessive functionality, permissions and autonomy. This distinction matters: improving the prompt does not remove a connector's ability to send confidential information.

Consider a support worker asking for a draft reply about one approved report. The application can authorize reading that report and creating a local draft. Sending the reply is a separate action with a recipient and message that require their own authorization. Exporting an entire customer dataset is outside the job. A model's claim that an export is necessary cannot add that permission.

Research question and prerequisites

Can a deterministic application boundary reject unauthorized proposals even when the planner proposes them confidently?

You need Python 3.10 or later, a terminal and an empty directory. All identifiers below are invented. Addresses use the reserved .test domain. There are no third-party packages, API keys, network requests or real sends. The program inspects prewritten proposals; it does not test an LLM's susceptibility to injection.

For integration later, you also need an inventory of your actual tools, their service identities, downstream permissions and side effects. A button labeled “draft” is insufficient if its backend identity can also send messages through another route.

1. Define authority before involving the planner

Write down the job in terms the backend can enforce. Our job belongs to tenant tenant-A, concerns report-7, permits a draft in workspace, and permits sending only to one configured recipient after approval. Export is unavailable.

In a real application, derive this scope from authenticated user permissions and the selected resource. Do not accept a tenant or authorization scope supplied by the model as authoritative. Proposal fields are claims to check against the trusted scope.

A general-purpose connector often bundles several abilities. Split it into narrow operations where possible, and restrict the downstream service account too. Otherwise a planner might bypass your nice approval interface through a second tool with equivalent power.

2. Bind approval to the action the person reviewed

“Allow email” is too broad for this workflow. Show the recipient, subject, complete body, attachment identities and relevant account to the reviewer. Record the exact approved operation on the server. Changing any meaningful field must invalidate that approval.

The simulation hashes canonical JSON to bind the approval to a payload. The hash is not a signature or proof of who approved it. Authority comes from a trusted approval record held outside the planner. The in-memory dictionary stands in for that record. A production implementation needs authenticated issuance, expiration, tenant binding and safe storage.

3. Run a known-fixture authorization simulation

Save the following as boundary_lab.py, then run python boundary_lab.py. The validator rejects extra fields and empty strings. This matters because hidden arguments and overly permissive parsing can undermine a narrow operation.

import hashlib
import hmac
import json

FIELDS = {"action", "tenant", "resource", "recipient", "body"}
SCOPE = {"tenant": "tenant-A", "resource": "report-7"}
RECIPIENT = "reader@example.test"
CAPABILITIES = {"draft", "send"}


def fingerprint(proposal):
    encoded = json.dumps(
        proposal, sort_keys=True, separators=(",", ":"),
        ensure_ascii=True,
    ).encode("utf-8")
    return hashlib.sha256(encoded).hexdigest()


def decide(proposal, approvals, approval_id=None):
    if not isinstance(proposal, dict) or set(proposal) != FIELDS:
        return "BLOCK invalid_schema"
    if any(not isinstance(v, str) or not v or len(v) > 1000
           for v in proposal.values()):
        return "BLOCK invalid_schema"
    if any(proposal[k] != v for k, v in SCOPE.items()):
        return "BLOCK scope_mismatch"
    if proposal["action"] not in CAPABILITIES:
        return "BLOCK capability_denied"
    if proposal["action"] == "draft":
        if proposal["recipient"] != "workspace":
            return "BLOCK destination_denied"
        return "ALLOW local_draft"
    if proposal["recipient"] != RECIPIENT:
        return "BLOCK destination_denied"
    if approval_id is None:
        return "BLOCK approval_required"
    if not isinstance(approval_id, str) or not approval_id or len(approval_id) > 1000:
        return "BLOCK invalid_approval_id"
    expected = approvals.get(approval_id)
    if expected is None:
        return "BLOCK approval_missing_or_consumed"
    if not hmac.compare_digest(expected, fingerprint(proposal)):
        return "BLOCK approval_mismatch"
    del approvals[approval_id]
    return "ALLOW approved_send_simulated"


send = {
    "action": "send", "tenant": "tenant-A",
    "resource": "report-7", "recipient": RECIPIENT,
    "body": "Here is the approved public report summary.",
}
# Trusted application fixture, not an approval generated by a model.
approvals = {"review-1": fingerprint(send)}
cases = [
    ("draft", {**send, "action": "draft", "recipient": "workspace"},
     None, "ALLOW local_draft"),
    ("unapproved_send", send, None, "BLOCK approval_required"),
    ("export", {**send, "action": "export"}, None,
     "BLOCK capability_denied"),
    ("other_tenant", {**send, "tenant": "tenant-B"}, "review-1",
     "BLOCK scope_mismatch"),
    ("changed_body", {**send, "body": "Different unreviewed content."},
     "review-1", "BLOCK approval_mismatch"),
    ("approved_send", send, "review-1", "ALLOW approved_send_simulated"),
    ("replayed_send", send, "review-1",
     "BLOCK approval_missing_or_consumed"),
    ("extra_argument", {**send, "attachment": "private.csv"}, None,
     "BLOCK invalid_schema"),
]
for name, proposal, approval_id, expected in cases:
    result = decide(proposal, approvals, approval_id)
    assert result == expected, (name, result, expected)
    print(f"{name}: {result}")
assert approvals == {}
print("checks=8 passed approvals_remaining=0 real_actions=0")

Expected complete output:

draft: ALLOW local_draft
unapproved_send: BLOCK approval_required
export: BLOCK capability_denied
other_tenant: BLOCK scope_mismatch
changed_body: BLOCK approval_mismatch
approved_send: ALLOW approved_send_simulated
replayed_send: BLOCK approval_missing_or_consumed
extra_argument: BLOCK invalid_schema
checks=8 passed approvals_remaining=0 real_actions=0

4. Understand the result before integrating it

The draft succeeds because it remains inside the permitted workspace. The unapproved send fails even though the recipient is allowed. A destination allowlist answers where an operation may go; it does not answer whether this particular operation was approved.

The export fails because no capability exists for it. Cross-tenant access fails before approval is considered. The changed body fails while leaving the approval available for the original reviewed payload. The original send consumes that approval, and replay fails. Adding an attachment fails at the schema boundary rather than quietly extending the operation.

These checks do not inspect persuasive wording or search for a suspicious phrase. They constrain the action regardless of why it was proposed. This makes the mechanism applicable when an assistant makes an ordinary mistake as well as when external content redirects it.

The final real_actions=0 is a statement about this program: it contains no operation that sends anything. The allowed result is a simulated decision. It is not a production security result, an attack-blocking percentage or proof that a particular model resisted hostile content.

5. Put the boundary at the execution point

For an actual assistant, insert authorization immediately before each side effect. Parse proposals into a strict schema, resolve the authenticated scope, validate the destination and compare against a server-held approval. Apply the same logic to direct tool calls, background jobs and retry paths.

Use the final provider payload for approval binding. If your renderer appends links, adds attachments or changes recipients after review, hash and approve the final representation. Include operation-specific fields omitted from this small example, such as subject, attachment content hashes, currency, amount or target environment. Define normalization precisely so the preview and executed payload agree.

The toy dictionary is single-threaded. In production, two workers must not consume the same approval simultaneously. Use an atomic database transition and a durable operation identifier. Record dispatch state and reconcile uncertain provider responses before retrying. Consuming approval alone does not guarantee exactly-once delivery, and a timeout does not prove a send failed.

Keep approval issuance out of the model's tool permissions. An assistant may request review; it must not manufacture approval by returning an approved=true field. If unattended actions are needed, encode the owner's explicit bounded policy in the trusted backend and restrict the corresponding service identity.

6. Review the surrounding data paths

The send gate leaves other routes to inspect. A search tool that retrieves every customer record can expose excessive information to the model before any outbound action occurs. Restrict queries and document access at retrieval time. A legitimate allowed recipient can also receive inappropriate content, so authorization must be paired with data minimization and a review that actually exposes the message.

Treat tool metadata and connector configuration as reviewed software assets. Record the accepted version, inspect changes and decide when reapproval is required. Label retrieved content by origin, use model-level injection defenses where appropriate and monitor unexpected tool sequences. These layers can reduce risk; none should be treated as a complete guarantee.

Log enough to investigate decisions: authenticated principal, operation identity, policy version, resource identity, approval reference and denial reason. Avoid copying entire confidential documents into routine logs. Give operators a way to stop a connector and revoke credentials while preserving investigation evidence.

Troubleshooting and limits

If the expected output differs, run the exact listing in a new file without editing its fixtures. A changed field intentionally changes the decision. The approval hash depends on canonical serialization, so adding whitespace inside the body changes the payload even if JSON formatting outside it does not.

If the valid send is blocked in your integration, inspect the reviewed versus final payload privately. Do not solve a mismatch by removing fields from the approval binding. Fix inconsistent rendering or normalization, then request approval for the actual operation.

If an export remains possible through another connector, the capability boundary is incomplete. Inventory alternate paths and constrain their credentials. If draft creation itself stores sensitive information somewhere public, reclassify that operation and apply the necessary authorization.

This lab covers eight predetermined proposals and a single trusted process. It does not evaluate malicious documents, multilingual attacks, an LLM, compromised approvers, token theft, connector vulnerabilities or concurrent dispatch. It demonstrates an enforceable boundary, with those limitations left visible.

Completion check and next experiment

You have completed the exercise when every output line matches, only the reviewed payload reaches the simulated send decision and a second attempt cannot reuse its approval. For a real assistant, also verify that its service identity cannot bypass the boundary through another endpoint.

The next experiment should connect this boundary to a sandbox mailbox with synthetic documents. Measure useful draft completion, unauthorized-action proposals, execution denials and legitimate actions blocked incorrectly as separate outcomes. Add changed recipients, changed attachments, expired approval and simultaneous retries. Only then evaluate model behavior with varied external content. Report the model and connector versions, test corpus and remaining bypasses alongside the results.