Consequences Nobody Wrote Down
The hardest unfinished work was never recorded as work at all.
Chapter 9 handles work the history states: promises, assignments, follow-ups with an established expectation behind them. This chapter pushes Question 5 past conventional task tracking, to consequences nobody wrote down. The migration broke assumptions encoded in fixtures, documentation, and configuration that no session mentions. A memory that tracks only stated intentions will report the open list as empty while the project quietly rots. Whether anything can be done about that — reliably, without inventing obligations — is treated here as a difficult hypothesis, not a capability.
Explicit versus derived
The distinction organises the whole chapter.
Explicit open loop
The history directly says it:
Update the migration docs.
Commitment evidence exists, and the transition logic from Chapter 9 decides its standing. Hard problems remain — silent completion, supersession — but the obligation itself is given.
Derived open loop
The system concludes it:
The database changed, but the backup configuration still references the old database.
No utterance promised this. The conclusion combines three things the system already maintains:
current architecture (PostgreSQL is the store — Ch 8 belief)
+
historical decision (the July migration — Ch 4/5 events)
+
current artifact state (backup config still names SQLite — present fact)
That combination is not retrieval alone. It compares remembered expectations with present state and derives a candidate consequence from the difference. The experiment requires each candidate loop to be traceable to exactly that triple:
current state
+
expected state
+
difference
The triple is also the honesty criterion. An inferred obligation without an expected-state source is a guess wearing a uniform. An inferred obligation whose current-state evidence is stale — a configuration snapshot from before the fix landed — is a resolved loop reported as open, the derived-world twin of Chapter 9’s stale tasks.
The corpus_import case
A second running example carries the cases the migration cannot. The team removes a legacy domain:
Decision (adr-013, 13 January 2025):
Remove the legacy corpus_import domain.
Consequences, only two stated:
- migrate CLI callers (issue-088, opened 14 January)
- migrate web callers (issue-089, opened 14 January)
Unstated:
- update docs (docs still describe corpus_import flags)
- delete compatibility facade (facade kept "temporarily")
- remove old tests (old tests still run, still pass)
The base case presents five candidate consequences, two with explicit commitment evidence. The fixture then adds traps that distinguish genuine derived debt from plausible-looking but invalid obligations.
The facade was kept deliberately. A partner note records that one external partner still imports through it for a fixed period. So delete the facade is a false apparent consequence — acting on it breaks a commitment in the other direction.
The old tests are intentionally preserved, as regression coverage for the migration itself. Removing them destroys evidence the team chose to keep.
A deferred-cleanup note parks the docs update until after release. That makes it an explicit intention with a scope condition, not a derived loop at all.
A partial migration — four of the nine CLI call sites still import corpus_import, and the web migration has not started at all — leaves some of the underlying work discharged and some still outstanding. The gate carries no separate fractional status: in the full-scope task below, the remaining CLI and web work are scored as genuine consequences and admitted like any other.
A system that turns all three unstated candidates into obligations scores well on recall and fails the project: it deletes a facade under contract, removes tests being preserved deliberately, and duplicates a scoped intention. A system that lists none misses real rot. The operating point between them is the chapter’s entire subject, and it is why the metric must punish plausible-but-invented obligations.
Why precision dominates listing — with a safety-recall boundary
For actionable derived obligations, high recall bought with speculative todos can be worse than a narrower list.
The asymmetry is practical. Each invented obligation costs attention, and attention spent verifying phantom debts is attention taken from real ones.
Worse, a fluent justification can make an invented obligation look supported. Chapter 7’s distinction matters here: traceability cannot turn an unsupported inference into evidential support.
So precision leads for listing derived obligations, with recall reported alongside it rather than aggregated away. The harness also scores a third property the earlier questions never needed: abstention quality.
One scope note. Where missing a safety-relevant obligation is itself costly, recall may dominate instead. The precision-first rule here is scoped to acting on or inventing obligations. It is not a universal retrieval law.
On fixtures where the ledger records no derivable loop — the facade under contract, the intentionally preserved tests — the correct output is silence. And silence must score above confident invention.
Stages and evidence carry this. A derived loop arrives with the triple that produced it and a reason-coded verdict the scorer can audit. Admission requires traceable current and expected state, a stated source for the expectation rather than a guessed one, current-state evidence fresh as of the standpoint, and no valid cancelling evidence. The representation stays minimal, shown here only in outline:
def infer_open_loops(state, now):
# Illustrative: compare remembered expectations against
# present artifact state; every candidate carries its
# (current, expected, difference) triple and a gate decision.
candidates = propose_differences(state, now)
return [c for c in candidates if gate_admits(c, now)]
The implemented gate uses no scalar confidence threshold. A staged, reason-coded pipeline attributes failures where a weighted score cannot, and this chapter’s errors must be debuggable before they are trusted. A scalar would have to demonstrate value beyond decisions the stages already make, under the same replay and regression gates used below.
The experiment
Fixtures contain explicit tasks, hidden consequences, false apparent consequences (the contracted facade), intentionally preserved stale references (the regression tests), deferred cleanups with scope conditions, and partial migrations. Queries ask what remains unfinished in a scope; the ledger records which loops are derivable, which are explicitly stated, and which apparent loops are traps. Measurement extends the Question 5 family with inferred-consequence precision and recall scored separately from explicit-task metrics, plus abstention scoring on trap-only fixtures and evidence correctness on the reported triple.
Conditions hold corpus, standpoint, and candidate proposal fixed and change only the derivation mechanism: unconstrained obligation listing, the staged S0–S4 triple gate, three ablations removing one leg each, the ledger oracle ceiling, and a trivial always-abstain control. On zero-positive trap tasks, undefined precision or recall terms remain null rather than being converted to zero; macro averages use only defined values, with micro totals and the separate abstention metric alongside.
Mean over the eight tasks (macro over defined values; micro in parentheses):
| Condition | Inferred precision | Inferred recall | Harmful tasks | Abstention |
|---|---|---|---|---|
| unconstrained | 0.458 (0.476) | 1.000 (1.000) | 4/8 | 0.000 |
| staged | 1.000 (1.000) | 1.000 (1.000) | 0/8 | 1.000 |
| no-current | 0.929 (0.909) | 1.000 (1.000) | 0/8 | 1.000 |
| no-expected | 0.857 (0.833) | 1.000 (1.000) | 0/8 | 1.000 |
| no-cancel | 0.646 (0.556) | 1.000 (1.000) | 4/8 | 0.000 |
| oracle | 1.000 (1.000) | 1.000 (1.000) | 0/8 | 1.000 |
| abstain | null (null) | 0.000 (0.000) | 0/8 | 1.000 |
Frozen run: ch11-20260920T171039Z-derived-loops (staged-triple-gate-v1, standpoint 2025-02-07, eight scope tasks, zero model calls).
Recall is 1.0 for every listing condition, including the oracle. No genuine positive is missed anywhere, so the contest is fought entirely in precision and harm.
The unconstrained baseline lists everything proposed, including the facade under live contract, and pays for it: precision 0.458, with harmful listings on four of eight tasks.
The staged gate matches the oracle exactly — macro and micro 1.0/1.0 — with zero harmful listings.
On these fixtures, each leg earns its place in the ablations:
- Remove the freshness leg and the pre-fix backup snapshot is admitted. Macro precision drops to 0.929.
- Remove the stated-expectation leg and the inferred guess is admitted. Precision drops to 0.857.
- Remove the cancelling-evidence leg and the facade, the preserved tests, and the scoped deferral all come back. Precision returns to 0.646, with harmful listings on four tasks.
The full-scope task shows the mechanism working.
Six candidates are proposed. Three are genuine — docs flags, half-migrated CLI callers, unmigrated web callers. Three are traps.
The staged gate admits exactly the three genuine loops, and rejects each trap with its stage recorded: the facade on its contract, the preserved tests on their intentional preservation, the scoped deferral on its scope condition.
Two controls confirm the other legs. The stale-leg control rejects the pre-fix backup snapshot. The guess control rejects the inferred expectation.
Every rejection carries the stage and the cited legs. No candidate is explained after the fact.
The always-abstain control shows why abstention cannot be read alone: it reaches perfect abstention with zero recall and null precision because it never admits anything. The trap-only task (facade plus preserved tests, silence correct) is also abstained on correctly by the staged gate, which retains recall on the positive cases; every condition that drops the cancelling leg fails the trap-only case.
Policy revision follows the Chapter 10 discipline. The tempting loosening — drop the cancelling-evidence leg to recover the deferred-docs candidate — is proposed as an immutable gate version, replayed over all eight tasks, and rejected: primary gain −0.354 with breaches on harmful-task rate (0.0 to 0.5) and abstention rate (1.0 to 0.0). A locally correct repair that is globally harmful stays out, for the same reason Chapter 10 kept its own out.
Book result. The triple structure buys precision on traps and stale-reference fixtures while recall holds on hidden-consequence cases (staged macro/micro 1.0/1.0 against unconstrained 0.458/1.0 macro, 0.476/1.0 micro; harmful tasks 0 vs 4; oracle matched). On these fixtures, the result supports guarded derived-loop inference with an explicit gate version. Demotion clauses apply: crisp fixtures with ledger adjudication, no reader, synthetic corpus, deterministic links. The staged gate equalling the oracle here flatters crisp traps; ambiguous rot with graded evidence remains untested and is recorded as the next failure class.
What this chapter earns, and the problem it exposes
The run demonstrates a derived-state operation beyond retrieval: remembered expectations are compared with present state, and the mismatch can produce a candidate obligation. The mechanism composes with the earlier layers rather than replacing them — expectations keep their provenance, temporal distinctions remain intact, and superseded evidence stays out of the expected-state leg.
But success creates the next failure, and it arrives immediately.
Suppose the system now knows every unresolved item — the backup migration, the docs update, the fixture revision, the CLI and web caller migrations, the deferred cleanup with its scope condition.
It cannot put all of them into every context.
A migration task needs the fixture warning and the prior failure. A documentation task needs the scope condition. A release task needs the backup debt.
The open list is too large to be the answer. So the question changes, from what is unfinished? to which of it matters for what I am doing now?
What remains unsolved. Even complete knowledge of every unresolved consequence creates a selection problem. Remembering everything unfinished is not the same as bringing the right unfinished thing to the present task. Chapter 10 has already built the machinery for that selection and shown it working on stated obligations — under its Type C verdict, where selection improved and answer correctness did not follow. It can select and rank a derived consequence after this chapter has produced it; it cannot originate a consequence absent from the memory representation. The two capabilities are therefore complementary: one produces guarded candidates from maintained state, while the other decides which candidates matter to the work at hand.
Research foundations
Work on agents learning from trajectories shows why unresolved consequences may be distributed across actions, observations, and later outcomes, rather than written down as explicit tasks. ExpeL extracts reusable knowledge from experience. Reflexion carries verbal feedback between attempts. Voyager accumulates skills through interaction. SWE-bench grounds software-agent evaluation in real issues and repositories.
Together they provide precedents for deriving reusable state from relations among observations, actions, and outcomes rather than copying it from an explicit task sentence.
They do not establish that a derived consequence is correct merely because it is plausible. This chapter therefore keeps its own requirement separate: inferred loops must outperform correct abstention on trap cases. An inferred loop is a hypothesis, not an extracted fact.
The closest classical precedent is older than any of them.
Doyle’s truth maintenance system keeps each belief together with the reasons that support it, and revises the belief set by dependency-directed backtracking when an assumption changes.
That is the conceptual engine this chapter needs, in miniature. A derived loop is a maintained expectation checked against present state. When a leg changes — the snapshot refreshes, the contract lapses, the belief revises — the loop must be reconsidered or retracted, never served stale.
The system does not implement a general truth maintenance system. It borrows one discipline: recorded justifications, with retraction as a primitive.
The (current, expected, difference) triple therefore needs provenance and falsification for every leg. Current-state evidence establishes what is observed; a decision, contract, or other stated expectation licenses what should hold; the mismatch proposes the candidate loop. Trap fixtures—contracted facades, intentional stale references, scoped deferrals—supply evidence that cancels apparent mismatches. The decisive metric is inferred-consequence precision at matched recall with abstention rewarded. Unrestricted production of plausible debts is not memory.
References
- Jon Doyle, A Truth Maintenance System, Artificial Intelligence 12(3), 1979.
- ExpeL: LLM Agents Are Experiential Learners (2023).
- Reflexion: Language Agents with Verbal Reinforcement Learning (2023).
- Voyager: An Open-Ended Embodied Agent with Large Language Models (2023).
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (2023).