Skip to the finding
After the PromptJames DimachkieSay hello
Back to selected research

Receiver Reliance

Before the next agent acts.

One agent’s work becomes another agent’s input. What does the evidence actually support at that handoff?

A decision at the receiving boundaryAn exact work product, its evidence, and the intended use feed a structural check. The host releases work after a pass; otherwise it may repair or hold. Answer correctness requires a separate check.Exact workproductEvidenceIntendeduseReceiving decisionRepair or holdRelease on pass
The check makes a decision. The host supplies the facts and enforces it.

The checks changed what agents received.

In the completed 1,805-episode study, Receiver Reliance reduced the specified improper handoffs. Persistent defects also stopped useful work. Structurally acceptable but wrong answers still passed through.

That is the contribution of this preliminary study: an observed structural-release effect, with its costs and boundaries measured alongside it. It is a completed stage of an ongoing research question.

Read the one-page abstract

A record is only part of the decision.

An output can be authentic, current, and authorized for a particular use while still being wrong. RR asks a narrower, operational question: does the available evidence support this receiver’s use of this exact work product for this exact operation?

Consider a subtotal passed between agents.

The declaration might describe an older revision. It might cover a different use. A structural check can detect a specified mismatch when the host supplies the necessary facts.

If the declaration, revision, and use all match but the subtotal is wrong, the relationship can pass. Checking the arithmetic needs something else.

The full operational package checks admissible input, identities, engine predicates, lineage, and acceptance. The host gathers observations and enforces the decision. A hash binds an identity; it does not establish the truth of the outside world.

Four policies. One shared host.

The study compared the complete declared RR procedure with transparent release and two fixed static checking policies. Each arm ran its own model trajectory under the same host rules.

Tasks
Sum groups of integers; select item identifiers by grade and count.
Workflows
One receiver, a three-receiver chain, and a four-receiver fork/join.
Conditions
625 episodes with constructed administrative defects; 1,156 clean episodes; 24 separate semantic controls.
Runtime
A pinned Qwen3 14B model through Ollama on a qualified Colab A100 GPU, with an honest controller.
Comparators
Transparent release (U), the original static policy (S4), and the same static policy at a lower threshold (S1).

The static policies did not express the complete dynamic RR relation. This comparison cannot establish an advantage over an equally capable conventional implementation.

Sample design, measurement, and internal review

The initial-stage sample was amended after collection began and before endpoint analysis: the first 1,800 scheduled entries plus five remaining controls. All 24 controls were retained; 619 original assignments were deferred. This was not a publicly preregistered full-sample study.

The analysis equally weighted predefined condition, task, and workflow cells. All 1,805 episode measurements were valid under the declared checks. Task success was assessed separately.

Nineteen checkpoint archives were verified locally. An internal reviewer independently recomputed all six estimates and bounds and examined twelve predetermined traces. This was neither a second full evidence census nor independent human replication.

Read the methods and assumptions

A result with a visible tradeoff.

Four constructed conditions tested missing or mismatched declarations, an expired local grant, and delivery at a different revision. The structural endpoint counted receiver presentations with a currently evidenced defect, per episode.

Compare the measured effects

Fewer improper handoffs in the seeded conditions.

RR minus each comparator, in improper presentations per episode. Negative values mean less structural exposure under RR.

Transparent release−1.000[−1.599, −0.401]
Static policy · S4−1.000[−1.596, −0.404]
Static policy · S1−0.750[−1.334, −0.166]
Points are equally weighted estimates; lines show conditional simultaneous intervals. The six contrasts have at least 95.3125% joint coverage under the declared sampling and runtime assumptions. These are synthetic study results, not field incident rates.
All six estimates and intervals
RR minus comparator. Structural exposure is a count per episode; clean completion is in percentage points.
OutcomeComparatorEstimateConditional interval
Structural exposureU−1.000[−1.599, −0.401]
Structural exposureS4−1.000[−1.596, −0.404]
Structural exposureS1−0.750[−1.334, −0.166]
Clean completionU−0.375[−14.973, +14.224]
Clean completionS4+0.888[−13.831, +15.607]
Clean completionS1−0.429[−15.362, +14.505]

Withholding work has a cost.

RR had no observed improper presentations in the seeded population. It also correctly completed only 80 of 162 seeded tasks, compared with 150/151, 154/155, and 139/157 under the alternatives.

Correctly completed tasks
PolicyWith seeded defectsClean conditions
Transparent release150 / 151298 / 299
Static policy · S4154 / 155286 / 291
Static policy · S1139 / 157273 / 274
Receiver Reliance80 / 162290 / 292
Observed completion counts over the realized draws. These descriptive fractions differ from the equally weighted estimates above; bar lengths show the fraction completed, on a common 0–100% scale.

A first rejection allowed one administrative refetch without changing the answer. A persistent defect then held the receiver and prevented dependent work. RR recorded 162 such retries and 81 held receivers in the seeded population.

Clean-task intervals include both benefit and harm. They establish neither a utility gain nor equivalence or noninferiority. Fewer model calls on seeded tasks partly reflect stopped work; they do not establish equal-quality efficiency.

The wrong answer can still get through.

Every arm completed its six semantic-control trajectories. Every arm produced zero correct final answers.

Those controls deliberately changed an answer while keeping the structural relationships coherent. Receivers lacked the original source data needed to recompute it. The result shows propagation under that task design, not a general limit on a model’s ability to recover.

RR supplied no semantic repair. A valid relationship between an object, its evidence, and its use is not a guarantee of a correct answer.

The findings that stayed in the record.

The latest result sits within a longer investigation. Earlier experiments tested different interventions and cannot be pooled with this study.

A narrower operational relation
A conventional implementation matched RR in the earlier experiments. The new full-package result does not erase that parity.
A learned policy
In a separate ten-seed experiment, every final learned policy underperformed a simple heuristic.
A prompt-based review instruction
An eighteen-family pilot found no demonstrated benefit, with a wide interval and acknowledged grading uncertainty.

My role in the work

I directed the research, with substantial AI assistance in specification, implementation, mathematical analysis, experiment execution, review, and writing. The paper preserves the corrections and negative results alongside the positive finding.

The reviewers remained within one operator’s project. Internal checking made the work more inspectable; it is not external peer review or independent human replication.

The next question is practical value.

When is a prevented violation worth the legitimate work withheld? The answer depends on the workflow, the consequence of violating its requirements, and the cost of holding or repairing a handoff.

An informative next comparison would measure those costs in an authentic workflow against an equally expressive conventional policy. That comparison has not been performed. The historical public/adversarial program also remains unexecuted.

This study establishes a bounded effect of the tested package. It does not establish field efficacy, general AI safety, semantic correctness, component causation, a novel security primitive, or production readiness.

Read at your own depth.

Three editions of the same completed preliminary study. Start with the finding; follow the methods and earlier evidence as far as you need.

These are working-paper editions of the current local study. The earlier public software release, v1.3.0, is a separate artifact. The study is complete at this preliminary stage; the wider research remains open.

The PDFs retain source references into the local research archive, which is not included on this website. Use the links on this page to move between reading editions.

A question worth
following carefully.

Talk with me about the research