← Matrix

Improvement Lab

Research and extraction audit

First 20 complete answers in public feed order at capture; 10 pages from one formula-rich scanned PDF selected for risk. Not a random or representative sample.

20 answers checked · 9 contain challenge-page citation records · 10 PDF pages spot-checked

All 20 answers already had quality warnings. All ten selected pages show at least one transcription defect. This is not a measured overall error rate.

Inspect the 20 research records
  • Q3847 — 8 citation records; 1 challenge-page matches; 2 provider-attributed unchecked records.

    Internal-context mismatch: the question supplies a catalog version definition, but the answer treats absent public documentation as the obstacle. USGS STAC describes geospatial metadata, not verification of Network0 catalog membership.

  • Q3846 — 9 citation records; 4 challenge-page matches; 0 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3845 — 9 citation records; 0 challenge-page matches; 0 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3844 — 12 citation records; 1 challenge-page matches; 3 provider-attributed unchecked records.

    AWS documentation supports ground-station contact scheduling in general. It does not validate this project’s visibility windows, latency or compatibility.

  • Q3843 — 12 citation records; 0 challenge-page matches; 3 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3842 — 10 citation records; 0 challenge-page matches; 1 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3841 — 9 citation records; 4 challenge-page matches; 2 provider-attributed unchecked records.

    Source-selection mismatch: NIST Core Network Technologies describes its own networking research. It does not establish Network0 implementation readiness.

  • Q3840 — 10 citation records; 1 challenge-page matches; 1 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3839 — 9 citation records; 2 challenge-page matches; 2 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3838 — 12 citation records; 0 challenge-page matches; 1 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3837 — 7 citation records; 4 challenge-page matches; 3 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3836 — 10 citation records; 6 challenge-page matches; 1 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3835 — 12 citation records; 0 challenge-page matches; 2 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3834 — 9 citation records; 0 challenge-page matches; 3 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3833 — 10 citation records; 0 challenge-page matches; 3 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3832 — 9 citation records; 0 challenge-page matches; 4 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3831 — 7 citation records; 0 challenge-page matches; 0 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3830 — 6 citation records; 0 challenge-page matches; 2 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3829 — 10 citation records; 3 challenge-page matches; 0 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

  • Q3828 — 12 citation records; 0 challenge-page matches; 3 provider-attributed unchecked records.

    preliminary structural audit; full claim-support review pending

Inspect ten page-level OCR findings

BIG IDEA - engineering formulas rewrite - corrected.pdf

  1. Speed unit m/s became rn/s; squared-area notation became mA2.
  2. Absorbed-power expression lost the wavelength-bin factor and closing bracket.
  3. Heat-transfer units use mA2 and KA4 instead of powers.
  4. Friis equation lost the denominator ending and square; transfer-efficiency equals sign missing.
  5. Thrust equation lost nozzle-area factor A_e.
  6. Carnot expression lost hot-reservoir denominator T_h; inductor-energy left side missing.
  7. Powers became A2; vibration equation loses its right side across the page break.
  8. Transmissibility powers became A2; qubit notation needs token-level checking.
  9. Transpose notation became AT; Euler buckling powers became A2.
  10. Design weight w1 became WI and reciprocal numerator 1 became I.

Repair and review queue

  1. Exclude challenge-page content from evidence retrieval.
  2. Resolve internal catalog definitions from project records before web search.
  3. Independently review claim-to-passage support; do not use URL matches as accuracy.
  4. Correct equation OCR with versioned page-level annotations before calculations.
  5. Use fresh uninspected cases for future holdout; this sample is development only.

Queue only: these findings do not rewrite historical answers or automatically correct OCR.

Download audit and provenance · Download independent-review worksheet

Codex AI-assisted audit, not independent human validation. Structural flags do not prove an answer false. Existing warnings are historical system output, not new accuracy measurements. No AI accuracy, extraction error rate or before/after improvement has been established.

Measurement coverage and sources

No overall accuracy percentage is established. A source catalog is not an imported dataset or a verified claim.

Executed numerical baseline

NIST PiDigits: 3 / 3 checks passed. This tests the statistics utility on one constructed dataset, not AI answers or physical reality.

Research citation support: unmeasured · Extraction fidelity: unmeasured · Physical simulation validity: unmeasured · Before/after improvement: not established.

Baseline provenance, tolerances and results
{
  "schema": "network0-numerical-baseline-v1",
  "executed_at": "2026-09-13T15:24:49.811Z",
  "dataset": "NIST PiDigits",
  "source_url": "https://www.itl.nist.gov/div898/strd/univ/data/PiDigits.dat",
  "raw_sha256": "e9421a2f4c0a31c505099332b4d22e507e808953c3c57bddae9226f25587d7d2",
  "implementation_sha256": "09822dd1cf86dd3420c254f1c2c8e4e7fc4cfd7a0dbb607d016e8041ae529165",
  "implementation": "Welford one-pass sample statistics",
  "runtime": "v24.19.0",
  "classification": "executed numerical software reference check; constructed data, not experimental observations",
  "ai_accuracy": null,
  "before_after_improvement": null,
  "results": [
    {
      "metric": "count",
      "expected": 5000,
      "absolute_tolerance": 0,
      "actual": 5000,
      "absolute_error": 0,
      "pass": true
    },
    {
      "metric": "mean",
      "expected": 4.5348,
      "absolute_tolerance": 1e-12,
      "actual": 4.53480000000001,
      "absolute_error": 1.0658141036401503e-14,
      "pass": true
    },
    {
      "metric": "sample_standard_deviation",
      "expected": 2.86733906028871,
      "absolute_tolerance": 1e-12,
      "actual": 2.8673390602887068,
      "absolute_error": 3.1086244689504383e-15,
      "pass": true
    }
  ],
  "limits": "One public dataset, two statistics plus count. Tolerances are numerical acceptance bounds, not measurement uncertainty. No generalization or AI-quality claim. Raw file retained locally only."
}

Download sources and baseline · Download measurement-plan template

Repeat the numerical check locally

Download the NIST PiDigits file, then select it below. Processing stays in this browser; the file is not uploaded. This is a published reference check, not a blind holdout.

No local file checked.

Next: independently annotate representative citations and extraction examples; freeze separate development and holdout cases; run both versions with the same settings; publish failures, missing outputs and uncertainty alongside the score. Set a cost limit before any AI evaluation. Missing evidence stays unresolved.

Compare one candidate with the current version on the same reviewed cases. Keep failures and missing information visible.

Calculation pilot: 50 synthetic examples in five families. References need independent review. This evaluates numeric answers and units; broad factual accuracy and citation support require separate assessment.

1. Prepare the benchmark

Development: force, torque and battery calculations. Holdout: efficiency and photon energy. Do not tune against holdout results. Because answers are visible, this pilot cannot certify a blind evaluation.

Generate baseline and candidate outputs separately using the downloaded questions. Record the model, prompt and code version in each run name. Supply actual cost and latency where available. A URL match is not proof that a citation supports an answer.

2. Compare versions

Detailed results and failures
No comparison yet.

3. Decide and preserve

Inspect every regression and missing answer. Review source support separately. Consider accuracy, coverage, costs and timing together; no single score triggers deployment. Export the comparison as the evidence for a proposed change. Apply approved changes through the existing Admin review process; retain the previous version for rollback.

Local workspace: these controls do not upload your inputs or call AI. Download your workspace before leaving; unsaved edits disappear on reload. Imported reviews are user assertions, not authenticated approvals. This lab does not train a model or deploy changes.

Methodological references: Data leakage · NIST evaluation guidance