KnowledgeGuard / EGB
A controlled study asking whether a RAG system can tell how its retrieved evidence is deficient — missing, insufficient, conflicting, outdated, or absent from the corpus — and whether that diagnosis carries actionable information for selecting a repair action.
The problem
Every published method for handling deficient retrieval evidence is developed and evaluated on a corpus containing exactly one deficiency mode by construction. But the correct repairs diverge, and some are opposites: escalating retrieval helps when evidence is missing and actively harms when it is contradictory. No detector has ever been required to tell those cases apart, because no benchmark presented them together with labels — so neither their detectability nor their usefulness had been measured.
Context and constraints
The project did not start here. The original proposal was a RAG system that checks evidence sufficiency, detects gaps, re-retrieves and abstains. A literature search found that already published, clause by clause — sufficiency analysis and gap-driven re-retrieval as S2G-RAG at ACL 2026, sufficiency-guided abstention at ICLR 2025, retrieval evaluation with corrective action as CRAG, and diagnosis-conditioned repair as Doctor-RAG and D2R-RAG. Building it would have been re-implementation presented as research, so the direction changed to the question the literature had left open.
- The claim under test needs ground-truth deficiency labels, so it cannot be computed on any existing benchmark — the benchmark had to be built first.
- OUTDATED cannot be synthesised honestly, which forced the choice of the one corpus carrying real superseded values.
- No API budget: the reader is a 250M-parameter local model on CPU, so absolute scores are not comparable with published systems and every comparison had to be within-instance.
- Mixing corpora across rows would confound deficiency type with source corpus, so the second corpus is held as a separate replication and never pooled.
- The null hypothesis had to be pre-registered as a real possibility: if type-agnostic repair matched oracle routing, that is the finding.
What I built
EGB, a benchmark where evidence-deficiency type is a manipulated, labelled variable with co-occurrence cells, and a fully within-record 5 x 6 factorial over it. Four of six construction operators are purely subtractive — evidence is withheld or removed rather than fabricated. Every source record is instantiated under every deficiency type and run under every repair action, including the cells no router would ever pick, which is what makes it a factorial rather than a system comparison and makes every contrast paired.
Architecture
The system exists as the apparatus for the measurement, not as a product. A record is instantiated into a typed deficiency, retrieved against, repaired under one action, generated from, and scored — with admission gates and label verification standing between construction and the factorial.
- HoH corpus
- 18,807 indexed passages
- Deficiency operators
- Four are subtractive
- Admission gates
- Contamination, leakage
- Label verification
- Independent NLI
- Retrieval
- BM25 index
- Repair action
- Six, three families
- Local reader
- 250M parameters, CPU
- Scoring
- Four pre-registered DVs
My contribution
Sole researcher. Benchmark design and construction, experimental design and pre-registration, analysis, and the forensic re-analysis that produced the published correction.
The 5 x 6 factorial
5 × 6, fully crossed. Every cell was run. Values are token F1 against the gold answer. Select a row, column or cell to see what it represents.
| Evidence-deficiency type ↓ / Repair action → | ||||||
|---|---|---|---|---|---|---|
// select a row, column or cell
Run 2026-09-15. 47 records x 5 types x 6 actions = 1,410 cells, every cell n = 47. 2,946 generator calls, no API spend. Floor check passed: SUFFICIENT x NONE = 0.831, so the reader can use good evidence and the cells are interpretable.
Measured values, released under the project's own Tier P policy, which clears per-cell scores with passage text and prompts removed. Absolute numbers reflect a 250M-parameter local reader and are not comparable with published RAG systems; the factorial is a within-instance contrast.
What the results support
Each number is paired with what it does and does not license.
Type and action interact
partial eta-squared 0.32, permutation p = 1e-4The action profile genuinely differs by deficiency type, corroborated by a mixed-model likelihood-ratio test. This says the cells differ; it does not by itself say that knowing the type is worth anything.
Oracle typing beats the best single action
+6.6 F1 points, 95% CI [2.6, 10.5]The pre-registered null is rejected, but by a margin whose lower bound sits exactly at the frozen practical threshold of three points rather than comfortably above it. The honest statement is that typing buys roughly six points and the data are consistent with as little as three.
The number that must not be quoted
+41.5 F1 points against fixed escalationAgainst a fixed-escalation policy typing looks enormous, but almost all of that is the action main effect: escalation is simply a poor universal policy for a small reader because it dilutes the context. Reporting this as the routing benefit would be the single easiest way to overstate the result.
With a real detector the benefit reverses
predicted routing 0.064 F1 below type-agnostic [-0.120, -0.012]This is the result that matters for anyone wanting to build on it. The headroom is real and, on this evidence, unreachable: routing on a diagnosed type is worse than just picking one good action and applying it everywhere.
Constructed conflict is far easier to detect than natural conflict
recall 0.957 against 0.574A pre-registered threat to validity, now measured rather than feared. Nearly a third of natural conflicts are called sufficient — the dangerous error, because the system then answers from evidence it has not noticed contradicts itself.
Knowing when to decline is the largest single effect
+0.96 selective utility for ABSENT to ABSTAINInvisible under answer correctness, where an abstention and a confident fabrication both score zero. It appears only because a second, abstention-sensitive dependent variable was pre-registered before the run.
Correction
The study's only cell surviving multiple-comparison correction was read as arbitration resolving conflict by corroboration. A forensic re-analysis of the frozen artifacts — no model loaded, nothing modified — found otherwise. ARBITRATE scores identically to four decimal places in the conflicting, outdated and sufficient cells, and produces the same answer string as the sufficient cell on 45 of 47 records. The injected counter-passages are the only passages absent from the index that arbitration corroborates against: 47 of 47, against 0 of 1,315 gold and distractor passages. The counter-passage also quotes the whole question, giving it a query-token overlap of 1.000 against gold's 0.676, making it the strict maximum-overlap passage in 47 of 47 delivered sets. Two trivial rules — pick the maximum-overlap passage, or the one absent from the index — identify the counter-side in 47 of 47 instances with no generation at all.
The number stands; the causal reading does not. The correction establishes that the artifact exists. Whether it explains the effect is what the pre-registered replication measures, and that replication has not been run.
Technical decisions
What was chosen, why, and what it cost.
Abandon the original proposal after the literature search
- Decision
- The proposed system was found to be published work, clause by clause, and the project was redirected to the question that remained open: whether deficiency type carries actionable information when the repairs genuinely differ.
- Why
- Building it anyway would have been re-implementation presented as research. The gap that was actually open is that no benchmark presents the deficiency types together with labels, so nothing had ever been required to discriminate them.
- Trade-off
- The new question needs ground-truth labels that no existing benchmark carries, so the benchmark had to be built before the experiment could run at all.
Make the type factor oracle rather than predicted
- Decision
- The factorial uses ground-truth deficiency types, so it measures whether type carries information independently of whether any detector can recover it. Detection is a separate experiment.
- Why
- Confounding the two would make a null result uninterpretable: a failure could mean type is useless, or merely that the detector is poor.
- Trade-off
- The factorial alone cannot say anything about a deployable system. That required the third arm, which is where the benefit turned out to reverse.
Freeze the analysis before the first result existed
- Decision
- Methods were written and the analysis script — including its out-of-sample action-selection rule — was committed before the first result row was produced.
- Why
- In-sample selection of the best action guarantees a positive headroom even under a true null. Fixing the rule in advance is the only way the number means anything.
- Trade-off
- A pre-registered analysis cannot be improved after seeing the data, so a better test that becomes obvious later has to be reported as exploratory.
Publish the correction against the original rather than editing it
- Decision
- The forensic finding was added as a new section with the original record left untouched, and the superseded claims were marked in place.
- Why
- The value of a frozen record is that it cannot be quietly revised. Editing the earlier sections would have destroyed the thing that makes pre-registration meaningful.
- Trade-off
- The document now contains a claim and its refutation, which is harder to read than a corrected version would be.
Verification
How the implementation was checked, and how much of that can be shown publicly.
Pre-registered and frozen before the runevidenced
Methods frozen and written before execution; the analysis script with its out-of-sample selection rule committed before the first result row existed.
Admission and label-verification gatesevidenced
Contamination probing, independent NLI label verification with rejected records dropped entirely to keep the factorial balanced, a retrieval-ceiling admission criterion, and a leakage audit finding no ground-truth field or label string in any rendered prompt.
A gate that fails, reported rather than repairedevidenced
The surface-leakage gate does not pass for every deficiency type. The failure is reported in the results and bounds the detection claims rather than being quietly fixed.
Forensic re-analysis of the frozen artifactsevidenced
A reproduction script re-derives every number in the correction from the frozen artifacts at a named commit, with no model loaded and no artifact modified.
Results
- The interaction between deficiency type and repair action is real and large, and the pre-registered null is rejected on the primary measure.
- The benefit is concentrated rather than general: for MISSING and SUFFICIENT the best action is also the globally best action, so knowing the type bought nothing.
- Routing on a predicted type performs worse than applying a single good action everywhere, so the measured headroom is not currently reachable.
- The headline cell was subsequently found to be measuring a construction artifact, and the replication that would settle it has not been run.
Limitations and disclosure
What this project does not do, and what cannot be shown publicly.
The repository is private. Published here are measured scores, counts and statistics only — the categories the project's own release policy clears for public release with passage text and rendered prompts removed. No benchmark passage, rendered prompt, question or gold answer is reproduced. The HotpotQA replication factorial is not complete and no numbers are shown for it.