← Applied AI

Capability Is Not Authority

Can, may, and may accept are different questions. A policy check can refuse a new action before its adapter runs, but delegation, replay, caller identity, and containment determine how far that claim reaches.

Part 4 — Make It Safe and Verifiable

The process can now act on the world. That is exactly when the safety questions arrive, and they are not questions a prompt can answer.

Three of them run through this part. Being able to perform an action is not being permitted to perform it, so permission has to be checked in code, outside the text a model generates. A check that passes is not proof, so verification has to be independent of the generator, adequate to the property that matters, and bound to the exact state being accepted. And trying again is not free, because a retry can cause a second real effect where a replay would have returned a record.

By the end of this part the system stops reporting “success” and starts reporting which kind of success it can establish — and where it refuses to continue.


Can, may, and may accept

Give an AI worker a tool that can write files and a prompt that says “inspect only”, and you have asked it not to write. You have not stopped it. An instruction shapes what a model generates. Whether the write actually happens is decided by the program around the model. Prompt injection makes the gap concrete: text hidden in retrieved data can redirect generation, so permission has to be checked somewhere generation cannot talk its way past.

can  ≠  may  ≠  may accept

Having a tool is ability. Permission is a grant checked before the tool runs. Accepting the result as finished work is a different authority again.

By the end of this chapter you will be able to put authorization at a named boundary in an AI application: grants as explicit sets, a check at the last branch before the effect, delegation that can only narrow what a parent held, authority to produce kept apart from authority to accept, and an honest list of the trust assumptions that remain.

A sensible correction, refused

Picture a review. A reviewer finds an incorrect cache setting and proposes the right replacement. Its worker can edit files. The person who opened the review asked for inspection only.

The correction’s quality is beside the point. An accurate proposal grants no permission to apply it. The presence of an editing tool grants none either.

Chapter 19 separated the actor’s report from evidence about the effect. Now move back to the moment before that effect: the process is about to call a worker that can change something. Which part of the system decides whether it may?

Who is allowed to authorize an available operation?

The answer must be more specific than “the human.” It needs an operation, a grant, an entry point that checks the grant, and a stated trust boundary around whoever supplies it.

Can, may, may accept

Each answers a different question and leaves another open:

QuestionWhat would answer it?What it leaves open
Can this worker edit the file?An available implementation with access to the filePermission for this request
May this request cause a write?A grant checked before the execution boundaryWhether the write is correct
May this actor accept the result?A separate acceptance grant and completion contractWhether the actor’s identity is authenticated

An instruction such as “do not write” can guide a model. Refusing to invoke the worker is a different mechanism. The first depends on how the instruction affects generation; the second is a branch in the surrounding program.

Greshake and colleagues demonstrated indirect prompt injection through data retrieved by LLM-integrated applications. Their attacks exploit the blurred boundary between data and instructions and can redirect application behavior. That motivates checking permission outside generated text. Showing that CodeAI’s policy layer defeats those attacks remains the book’s design argument, not the paper’s result (Greshake et al., 2023).

Saltzer and Schroeder’s principles sharpen the question. Least privilege limits grants to what the job needs. Complete mediation asks that access be checked throughout the system, including recovery paths; remembered authorization decisions need scrutiny when authority changes. These are design principles, not a certification of this implementation. Their application to CodeAI’s execution and reuse paths is the book’s mapping (Saltzer and Schroeder, 1975).

There is also a vocabulary trap. In their technical terminology, a capability is an unforgeable authorization ticket for an object. This chapter uses “can” for technical ability. An ordinary enum value must never be mistaken for such a ticket (Saltzer and Schroeder, 1975).

The grant is a set

CodeAI’s Capability enum names READ, WRITE, EXECUTE, VERSION_CONTROL, REMOTE, DESTRUCTIVE, and ACCEPT. NETWORK and VCS are absent from the members. Authority holds a set of these values; allows checks membership. WRITE stays separate from EXECUTE, and neither carries ACCEPT with it. 1

Membership, shown with the real types. These are illustrative assertions, not stage evidence: 2

from codeai.domain import Authority, Capability

review = Authority(frozenset({Capability.READ}))
edit = Authority(frozenset({Capability.WRITE}))

assert review.allows(Capability.READ)
assert not review.allows(Capability.WRITE)
assert not edit.allows(Capability.EXECUTE)
assert not edit.allows(Capability.ACCEPT)

This is deliberately smaller than a permissions system for a whole operating environment. The set says which category of operation is granted. File path, repository identity, network destination, expiry, and authenticated principal all live outside it. 3

A sensible application might give a reviewer READ, an editor WRITE, and a separate reviewer ACCEPT. Those are constructed role choices. An actor’s name, a tool’s installation, or the persuasive quality of its proposed change produces none of them automatically.

PolicyEngine.require is the small check that turns membership into refusal. If the requested capability is absent, it raises AuthorityDenied with a message naming the capability. No model is consulted. 4

Budget is another dimension. check_budget compares usage with declared limits and raises BudgetExceeded when a limit is exceeded. Permission to write and permission to spend remain distinct questions; the budget helper exists whether or not every action path invokes it. 5

Follow a new request to the worker

Current execute_action begins by appending action.requested, then records the operation’s governance source and resolves authority before either disclosing a prior result or invoking the adapter. If the request names a directive, authorize_action reconstructs the effective grant from that directive’s recorded chain; if it names none, the runtime falls back to the caller-supplied Authority and records that source honestly as caller_supplied. A denial appends action.authorization_refused and a DENIED action.completed; a grant appends action.authorized with the basis used. Later gates for decision evidence, replay identity and crash recovery sit after this authority decision and are developed in later chapters. 6

A counting worker makes invocation visible without a model or subprocess. The two requests use different keys, so both reach the new-action policy check. This is teaching code, not a pinned run: 6

from dataclasses import replace
from codeai.adapters import ActionRequest, ActionResult, ActionStatus
from codeai.domain import Authority, Capability
from codeai.ledger import SQLiteLedger
from codeai.runtime import Runtime

class CountingWorker:
    def __init__(self):
        self.calls = 0

    def execute(self, request):
        self.calls += 1
        return ActionResult(request.action_id, ActionStatus.SUCCEEDED)

runtime = Runtime(SQLiteLedger())
worker = CountingWorker()
request = ActionRequest(
    action_id="review-edit", task_id="review", capability="write",
    instruction="Apply the proposed correction.", precondition_hash=None,
    idempotency_key="review-edit", requested_by="reviewer",
    actor_id="worker", adapter="counting-worker",
)
denied = runtime.execute_action(
    request, authority=Authority(frozenset({Capability.READ})), adapter=worker,
)
assert denied.status == ActionStatus.DENIED
assert worker.calls == 0

allowed = runtime.execute_action(
    replace(request, action_id="approved-edit", idempotency_key="approved-edit"),
    authority=Authority(frozenset({Capability.WRITE})), adapter=worker,
)
assert allowed.status == ActionStatus.SUCCEEDED
assert worker.calls == 1

The counter establishes only invocation within this constructed program. The worker edits no file. Chapter 19’s independent reading is still needed to establish a resulting file state.

Notice what crosses the API in this example: the request names no directive_id, so current CodeAI deliberately uses the caller-supplied authority as a fallback. If a request does name a directive, the caller’s authority argument does not widen that directive’s grant; the runtime resolves the named recorded chain instead. But the directive ID itself still comes from the action request: current authorize_action does not derive it from task.created or prove that it is the task’s directive. Neither path authenticates the requester. action.requested contains the request, while the resolved authorization basis is recorded separately in action.authorized or action.authorization_refused. 6

The precise current claim is therefore narrower than either extreme: an action can be authorized from a durable directive chain, but the caller still chooses which directive to name, and an action that names none falls back to a caller-supplied grant. It is not yet true that every action is authorized by the task’s recorded authority. 6

The two gates side by side — permission at the moment of action, narrowing at the moment of delegation:

Delegation must not enlarge the grant

A review may be split into smaller jobs. The child needs some subset of the parent’s permissions. Passing a task to another worker is no occasion to manufacture a permission that the parent lacked.

Authority.narrows(parent) is a subset test. Equality counts as narrowing here: it means “does not expand,” not “must remove at least one capability.” Directive.validate_child checks both authority and budget narrowing. A finite parent budget cannot become an unlimited child budget under Budget.narrows. 7

Here is the validation helper, called directly, with fixture limits chosen for illustration. Registration in the ledger comes next. 8

from codeai.domain import Authority, Budget, Capability, Directive

parent = Directive(
    directive_id="review-parent", objective="Review the change",
    success_criteria=(), budget=Budget(max_tokens=1000),
    authority=Authority(frozenset({Capability.READ, Capability.WRITE})),
)
child = Directive(
    directive_id="review-child", parent_directive_id="review-parent",
    objective="Inspect only", success_criteria=(),
    budget=Budget(max_tokens=500),
    authority=Authority(frozenset({Capability.READ})),
)
parent.validate_child(child)

But a helper that exists is not a rule that runs. In the version of CodeAI the earlier demos ran on, Runtime.open_directive appended a directive without calling validate_child, even when it declared a parent. The runtime registration path left the narrowing rule unenforced. 9

The gap: child validation existed, but not on the path that registers the child.

Connect the check to registration

CodeAI’s registration path now performs that check. For directives that declare parent_directive_id, open_directive reads the recorded parent, reconstructs its budget and authority, and calls parent.validate_child before appending the child. The child event’s causation_id names the parent event used for validation. 9

Grant provenance comes from the ledger. The caller supplies the child and its parent ID, not an alternative parent object for validation. The runtime reconstructs the parent from its recorded event. 9

Current registration keeps the attempted registration as evidence too. open_directive first appends directive.registration_requested. For a declared child it then resolves the recorded parent and validates narrowing. Missing or ambiguous parents, self-parenting, duplicate child IDs and widening all append directive.registration_refused before raising; directive.opened is not appended for the refused child. 9

That is a later repair than the pinned seven-case run below. In that run, a refused child existed only in control flow and no durable refusal event was written. The current source closes that observability gap, but the historical bundle must not be read as if it measured the repair. 9

Continuing the example above, registration now performs the check itself: 9

from codeai.ledger import SQLiteLedger
from codeai.runtime import Runtime

runtime = Runtime(SQLiteLedger())
parent_event = runtime.open_directive(parent)
child_event = runtime.open_directive(child)
assert child_event.causation_id == parent_event.event_id

This is a small connection, not a complete delegation system. A caller can still submit a root directive without a parent, and root registration authenticates no grantor. Current action authorization can reconstruct a recorded directive chain, but authorize_action takes the directive ID from ActionRequest; it does not derive that ID from the task’s task.created record or verify that the named directive belongs to the task. If the action names no directive, it falls back to caller-supplied authority. Direct ledger writes remain outside these entry-point checks. 10

The registration change therefore answers one specific question: does a declared child opened through this runtime narrow its recorded parent? It does not, by itself, bind every later action to the task’s own directive. The pinned run in the next section tests registration across seven cases with a separate ledger-based verifier.

Seven attempts to widen a grant

The pinned run for this chapter used a protocol frozen before execution. Two durable parents were recorded (READ+WRITE with a 1000-token budget; READ-only with 500), and seven cases ran against them on one ledger, with a second Runtime reopening the same file for the reconstruction case. A stdlib-only verifier recomputes from the ledger and the external invocation receipts.

CaseOutcome
Valid narrowing (READ child under READ+WRITE parent)registered, with causation to the parent event
Capability widening (EXECUTE under READ+WRITE)refused at registration; no event appended
Budget expansion (2000 tokens under 1000)refused at registration; no event appended
Missing parentrefused at registration; no event appended
Caller-forged broader parent (WRITE child under the recorded READ-only parent)refused: the recorded parent governs, not the caller’s broader story
Unauthorized action (WRITE under READ authority)DENIED as a durable action.completed, 0 adapter invocations
Reopen reconstruction (grandchild under the recorded child)registered with causation intact; a denial after reopen stays denied with 0 invocations

The forged-parent case is the load-bearing one for this chapter’s question: a WRITE child that would be valid under a broader imagined parent is refused because the recorded parent holds READ. Registration reads the parent from the recorded directive.opened payload; the caller supplies only the child and a parent ID.

One limitation stays with that pinned result: at the version exercised there, a refused widening raised before appending a refusal event, so the refusal itself lived only in control flow, while action denials were durable completions. Current open_directive now records the registration request and refusal before raising; that later behavior is source-inspected here, not established by the seven-case bundle.

Permission to produce is not permission to accept

Acceptance is a different boundary. In the Stage 14 negative cases, an acceptor holding every capability except ACCEPT was refused for unauthorized, and the producing actor’s self-acceptance was refused for self_acceptance. Those are results from Chapter 14’s pinned run. 11

accept_task checks ACCEPT before its prior-acceptance lookup. Its source validation compares the acceptor label with the producing actor label. That comparison separates named roles, not authenticated people. 12

For a long time it checked that ACCEPT in the wrong place. The grant came from an Authority object the caller passed in, while the action path resolved its grant from the recorded directive chain — two halves of one question, answered by two different standards:

recorded directive  ->  action authority       resolved from the record
recorded directive  ->  acceptance authority   supplied by the caller

A later audit put it hostilely: a directive granting only WRITE, and a caller passing Authority({ACCEPT}). The task completed. Worse for reconstruction, task.accepted recorded no directive at all, so the acceptance could not be re-resolved against any chain afterwards. 13

Acceptance now resolves ACCEPT from the task’s own recorded directive — read from task.created, never from the request, because a caller who could name the directive could name a permissive one. What the caller passes is still accepted by the API and recorded as caller_claimed_capabilities: visible, attributable, and inert. 14

That is the chapter’s rule applied to itself. Authority comes from the durable process, not from the caller’s description of its own authority. The caller’s label remains what it always was:

actor label  ≠  authority  ≠  authentication

An edit can therefore be authorized without being accepted as finished work. Whether the result meets an adequate check is Chapter 21’s question; this chapter establishes why permission to produce carries no silent permission to accept.

Changing authority is not suspending it

Closing that gap raised a sharper question. A task whose directive never granted ACCEPT reaches a state where the process cannot accept its own work; the scheduler says so, and asks for a person (Chapter 28). Then the person’s acceptance is refused too — correctly, because the record grants no one that authority. The durable process has reached a state from which completion is unauthorized. That is the right answer, and for a while it was the only answer, because nothing could lawfully move it.

The wrong repair would be an accept-anyway call for humans. It would hand back exactly the hole just closed, with a politer name. The right one changes the authority itself:

WRITE-only grant -> scheduler asks for a person -> a human acceptance is refused
      -> a durable authority transition records a superseding directive
      -> re-project
      -> the ordinary acceptance rule runs, and now permits it

Human intervention may change authority. It does not bypass authority. The acceptance code is untouched; the only difference is that the chain now resolves to a different directive. 15

That makes two relationships between directives, and they must not be confused:

RelationshipDirectionWhat it may do
Delegationparent → childOnly narrow. A widening child is refused, always
Authority transitiondirective → superseding directiveAdd, remove or alter, because it records a new external decision

A successor therefore declares its own authority and takes no parent; directives are immutable, so WRITE at T1 stays readable after WRITE + ACCEPT at T2 and READ at T3, and an acceptance already granted keeps the basis it was granted under. What the record establishes is that an explicit authority change, attributed to this actor, entered the process at this point — never that the actor was entitled to make it. The trust boundary has not moved. 15

The path that returns old results

In the version the earlier demos ran on, a reused action result was returned before PolicyEngine.require ran. The adapter was not invoked again, but the current request’s authority went unchecked either, so “every request reaches the policy check” was false for that path. The authority demo’s own method note records the ordering; its fresh keys avoided the path rather than testing it. 16

Current CodeAI mediates the request before replay lookup. After action.requested and governance bookkeeping, it resolves the current grant; an unauthorized request gets action.authorization_refused and a DENIED completion without receiving the prior result. Only an authorized request reaches the same-key lookup, where Chapter 22’s operation-identity check decides whether reuse is legitimate. Complete mediation is the design standard being applied here; the book does not claim that every possible entry point outside these runtime paths is mediated. 17

The same path now has a separate evidentiary-basis gate. ActionRequest.decision_id, when present, names a Chapter 18 decision. Before replay or new execution, execute_action projects that decision’s current evidence; an unknown, defeated or otherwise inadmissible basis is recorded as action.decision_basis_refused and the adapter does not run. If no evidentiary decision is named, this particular gate is not invoked. That mechanism was deliberately absent from this chapter’s pinned run and was added later, so it is current source behavior rather than a Stage 20 measured result. 6

Keep that apart from scheduler governance. Chapter 28’s scheduler decisions select process operations such as CALL or CHECK and are bound to the process state they were derived from. ACTION is deliberately not scheduler-selectable, so a scheduler decision cannot license a write. A Chapter 18 decision can supply an evidentiary justification for an action; it still does not supply authority. 18

What this is not

  • Containment is missing. The action entry point runs the adapter with no operating-system sandbox around it. 6
  • Authentication is missing. Requester and actor fields arrive as caller-supplied strings. 19
  • A tamper-evident ledger is missing. The event definition and append path sign no records. 20
  • Prompt-injection protection is unclaimed. The cited attacks motivate the boundary; no attack-resistance experiment is claimed here, and correctness verdicts belong to the next chapter, since permission plus an observed effect still leaves verification open.

Where it is still weak

  1. Action authority is not yet task-bound. A request that names a directive resolves that recorded chain, but the directive ID comes from ActionRequest, not from the task’s task.created record; the runtime does not check that the two match. If no directive is named, authorization falls back to a caller-supplied Authority, recorded as caller_supplied. 6 10
  2. Nothing is authenticated. Root registration authenticates no grantor, requester and actor remain caller-supplied strings, and an authority transition records an attributed change without establishing that the actor was entitled to make it. 19 15
  3. The grant is coarse. The capability set is not scoped to a file path, repository object, network destination or expiry. Containment remains outside it. 3
  4. Supersession has an ambiguity boundary. Grants are changed by immutable successor directives rather than narrowed in place; two recorded successors for one authority epoch make resolution unresolvable rather than choosing one. 15

Several earlier gaps are now closed in current source without being retroactively added to the pinned run: acceptance authority resolves from the task’s recorded directive, declared-child registration refusals are durable, replay rechecks current authority, and an action that cites an evidentiary decision can be refused when that basis is no longer admissible. The caller-chosen action directive, caller-supplied fallback, unauthenticated labels and lack of containment remain real boundaries. 14 6

Do this now

Thirty minutes. Find the last branch before one effect.

  1. Choose a disposable file action. Write down the capability it requests and where its authority comes from. Do not infer either from the instruction text.
  2. Supply a grant that lacks the capability. Count adapter invocations and read the file independently. Preserve the request, supplied grant, status, counter, and bytes.
  3. Delegate to a child with fewer permissions, then attempt a wider child. Trace the actual registration path; do not stop at finding a validation function.
  4. Separate authority to edit from authority to accept. List the remaining trust assumptions: caller, actor identity, replay, file scope, and ledger access.

If you are building with an assistant:

Make authorization visible at a named application boundary.
Use explicit grants, and check new actions before invoking their adapters.
Keep technical ability, authority to act, and authority to accept separate.
Connect declared-child registration to validation against the recorded parent.
Preserve which parent was used; refuse ambiguous parentage.
Test with counters and independently read bytes, using fresh keys for new
execution and a separate replay control. Preserve failures and missing records.
Do not claim authentication, sandboxing, or complete mediation from a set
membership check. Report every path that still trusts a caller-supplied grant.

Failure modes

  • Letting a good proposal authorize itself. Evidence for a correction is not permission to apply it.
  • Treating an enum as a protected ticket. A name for an operation is not an authenticated grant.
  • Finding a validator and stopping the audit. The path that records the child must actually call it.
  • Counting cumulative calls as new calls. A denied request can follow an earlier permitted effect.
  • Giving the editor acceptance by default. Producing and accepting are separate responsibilities.
  • Auditing only the new-action path. A check that guards creation but not replay is not complete mediation.

What this chapter established

  • Can ≠ may ≠ may accept. A tool gives ability. A grant checked outside generation gives permission. Accepting the result is a separate authority. A prompt instruction is guidance, not an execution gate.
  • Mediate before the effect, including reuse paths. Refuse before the worker is invoked, preserve the authorization basis, and do not let replay bypass the current grant. That is the complete-mediation principle applied to these runtime entry points, not a claim that every possible path in the system is mediated.
  • Delegation narrows; it never widens. A child’s grant and budget must fit inside its recorded parent’s. Current registration records both the attempt and a refusal when that rule fails; the older pinned run established the narrowing behavior before refusal events were durable.
  • Producing is not accepting. Action authority and acceptance authority are resolved separately. Acceptance now derives ACCEPT from the task’s recorded directive rather than trusting the caller.
  • Keep authority distinct from justification and identity. A Chapter 18 decision can supply an evidentiary basis, a scheduler decision can govern scheduler-selectable operations, and a directive can supply authority. None authenticates the actor, and none is a sandbox.

What CodeAI showed. In the pinned seven-case run, the stage then under test refused capability widening, budget expansion, a missing parent and a caller-forged broader parent; it recorded an unauthorized action as a durable denial with zero invocations and preserved causation across reopen. A separate verifier rejected four seeded corruptions. 9 Stage 29B’s bounded control paired READ with no write and WRITE with one observed write. 21

Current source goes further than that pinned stage: an action that names a directive resolves authority from its recorded chain; acceptance resolves ACCEPT from the task’s own recorded directive; registration refusals are durable; replay is behind current authority and identity checks; and a cited evidentiary decision can be refused when its basis is defeated or unknown. The action path still trusts the caller to choose the directive ID, falls back to caller-supplied authority when none is named, uses unauthenticated labels, and provides no containment. 6 14

Evidence notes

The pinned run. The grant-provenance bundle is experiments/applied-ai/evidence/grant-provenance/2026-09-14-a1b562a/. Seeded corruptions mark where enforcement ends as well as where it holds: a deleted durable denial, a forged registration for the widened child, a rewritten grandchild causation link, and an inflated denial receipt are each rejected with the failure named.

Stage 29B authority control. Stage 29B, reported in Chapter 28, includes a pinned authority control. Under READ the resolved answer’s WRITE effect was denied with 0 apply-adapter calls, no file written, and a person-ask reason of authority_missing:write; under WRITE it ran once and wrote a configuration containing ttl_seconds = 300. The control files retain authority, invocation counts and resulting target state. It is a bounded control from another experiment: it does not test registration, which the pinned run in this chapter does, and neither establishes comprehensive authority enforcement. 21

The earlier authority demo. It calls parent.validate_child(child) explicitly before using the child’s grant, exercising the helper and a subsequent action denial, and leaving registration-performs-validation undemonstrated. It used a counting adapter appending markers to a temporary file, fresh idempotency keys, no model calls and an in-memory ledger, and its preserved output is a summary in results.json rather than a durable exported action history with the target bytes attached. 16

CaseRecorded resultScope of observation
WRITE under WRITEsucceeded, 1 adapter call, file changedFirst allowed invocation
WRITE under READdenied, 0 new invocations, file unchangedCumulative adapter count remains 1
Narrowed childHelper accepted; WRITE denied; 0 new invocationsParent READ/WRITE, child READ
Widened childHelper raised authority-narrowing errorExplicit helper call, not registration
EXECUTE under WRITEdenied, 0 new invocationsDifferent capability is not implicitly granted
Fresh permitted actionsucceeded, cumulative adapter count 2A fresh key reaches the execution path
ACCEPT separationWRITE cannot require ACCEPT; ACCEPT canPolicy-level check, not full task acceptance

Those are the preserved summary’s own values. The denied case’s adapter_calls: 1 is cumulative, while its new_invocations: 0 is the denial measurement; mixing the columns would turn a successful refusal into an apparent violation. The exact refusal messages are capability denied: write and, for the widened child, child directive authority must narrow the parent authority. Neither depends on whether the proposed edit was useful. 16

The demo’s verifier imports no CodeAI code for its predicates: it reads the summary fields, checks denied-path invocation counts, the denied-file unchanged flag, narrowing flags and ACCEPT separation, and rejects a seeded corruption that sets the denied case’s new invocation count to one. Its input is still the producer’s summary, though. An independently preserved target file goes unexamined, permission is never re-derived from exported request and grant records, and the allowed case’s file_changed flag and the fresh-action row go unchecked — narrower coverage than its own introduction suggests. 16

Next

The split matters just as much when the thing being changed is your own tooling. An assistant can propose how your software should adapt; applying that change is an authorization, not a capability, and Chapter 30 keeps it in your hands.

The process may now have permission to invoke the worker. The worker may have changed the file. Neither answers whether the resulting state satisfies the property the task actually cares about.

Continue with The Agent Cannot Grade Its Own Homework.

References

  • Jerome H. Saltzer and Michael D. Schroeder. The Protection of Information in Computer Systems. Proceedings of the IEEE 63(9):1278–1308, 1975. Paper and glossary, basic principles.
  • Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. Not what you’ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. arXiv:2302.12173, 2023. Paper.

Implementation and evidence sources: historical claims above stay attached to the stage that produced them; current-source claims were checked against the later CodeAI successor. src/codeai/domain.py: Capability, Authority, Budget, Directive.validate_child; src/codeai/authority.py: recorded-directive resolution, caller-supplied fallback, supersession and authority transitions; src/codeai/policy.py: PolicyEngine.require, check_budget, AuthorityDenied; src/codeai/runtime.py: open_directive, create_task, authorize_action, execute_action, transition_authority; src/codeai/acceptance.py: record-derived ACCEPT resolution and acceptance validation; src/codeai/adapters.py: ActionRequest; src/codeai/evidence.py: current decision-evidence projection; src/codeai/governance.py: later scheduler-decision binding; src/codeai/ledger.py: Event, SQLiteLedger.append. Relevant tests include tests/test_directive_registration.py, tests/test_authority_resolution.py, tests/test_acceptance_authority.py, tests/test_policy.py, tests/test_runtime.py, tests/test_task_acceptance.py, and Chapter 19’s tests/test_action_observation.py. Book-repository evidence includes experiments/applied-ai/evidence/authority/, experiments/applied-ai/evidence/grant-provenance/2026-09-14-a1b562a/, experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b/authority/, experiments/applied-ai/evidence/task-completion/, and the later authority/decision-binding audit artifacts named in the footnotes. Historical bundles do not establish later repairs merely because current source contains them.


  1. Read in src/codeai/domain.py — Capability, Authority. ↩︎

  2. Read in src/codeai/domain.py — Authority.allows. ↩︎

  3. src/codeai/domain.py defines Authority as a plain set wrapper. ↩︎ ↩︎

  4. The gate lives in src/codeai/policy.py — PolicyEngine.require. ↩︎

  5. Budget logic in src/codeai/policy.py — PolicyEngine.check_budget. ↩︎

  6. Traced in src/codeai/runtime.py — Runtime.execute_action. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  7. Narrowing helpers in src/codeai/domain.py — Authority.narrows, Directive.validate_child, Budget.narrows. ↩︎

  8. Directive.validate_child in src/codeai/domain.py. ↩︎

  9. Registration path in src/codeai/runtime.py — Runtime.open_directive. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  10. Entry points in src/codeai/runtime.py — Runtime.open_directive, Runtime.create_task, Runtime.execute_action. ↩︎ ↩︎

  11. Measured run: experiments/applied-ai/evidence/task-completion. ↩︎

  12. Acceptance check in src/codeai/acceptance.py — accept_task, _validate_source. ↩︎

  13. Measured run: CodeAI’s frozen Wave 1 composition audit, probe B (experiments/W1-composition-results.md, baseline runtime f7d2910). Chapter 29 reports the audit in full. ↩︎

  14. Measured run: experiments/W1-R2-authority-symmetry.md and its recheck; src/codeai/acceptance.py (resolve_acceptance_authority), matrix in tests/test_acceptance_authority.py. ↩︎ ↩︎ ↩︎

  15. Measured run: experiments/W1-E1-authority-transition.md and W1-E1-composition-check.json; src/codeai/authority.py (resolve_supersession), Runtime.transition_authority, limits in docs/seams/authority-transition.md. ↩︎ ↩︎ ↩︎ ↩︎

  16. Unpinned demonstration: experiments/applied-ai/evidence/authority. ↩︎ ↩︎ ↩︎ ↩︎

  17. Replay logic in src/codeai/runtime.py — Runtime._replay_action_result. ↩︎

  18. Measured run: experiments/W1-R4-decision-execution-binding.md; src/codeai/governance.py. ↩︎

  19. Request shape in src/codeai/adapters.py — ActionRequest. ↩︎ ↩︎

  20. Append path in src/codeai/ledger.py — Event, SQLiteLedger.append. ↩︎

  21. Measured run: experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b. ↩︎ ↩︎