# Batch 01 Results Summary

Status: `in_progress`

This file is reserved for aggregate findings from `2026-05-17-batch-01`.

Raw output collection, local unblinded review, and the first independent frontier review have started. Fifty-two of 90 predeclared run slots have completed and have local reviews across fifteen local `agentic_coding` slices, three `frontier_reasoning` strategic repeat-1 slots, three regulated wastewater `frontier_reasoning` repeat-1 slots, and one formal proof-search `frontier_reasoning` repeat-1 slot. One additional formal proof-search `frontier_reasoning` light-structured report is visible in ChatGPT but remains blocked outside completion, review, and aggregate eligibility until a native Markdown export is harvested and locally reviewed:

- `case-strategic-gated-diligence-001` / `direct_answer` / `agentic_coding` / repeat 1
- `case-strategic-gated-diligence-001` / `light_structured` / `agentic_coding` / repeat 1
- `case-strategic-gated-diligence-001` / `full_ofone` / `agentic_coding` / repeat 1
- `case-scientific-mechanism-check-001` / `direct_answer` / `agentic_coding` / repeat 1
- `case-scientific-mechanism-check-001` / `light_structured` / `agentic_coding` / repeat 1
- `case-scientific-mechanism-check-001` / `full_ofone` / `agentic_coding` / repeat 1
- `case-regulated-wastewater-market-entry-001` / `direct_answer` / `agentic_coding` / repeat 1
- `case-regulated-wastewater-market-entry-001` / `light_structured` / `agentic_coding` / repeat 1
- `case-regulated-wastewater-market-entry-001` / `full_ofone` / `agentic_coding` / repeat 1
- `case-formal-proof-search-001` / `direct_answer` / `agentic_coding` / repeat 1
- `case-formal-proof-search-001` / `light_structured` / `agentic_coding` / repeat 1
- `case-formal-proof-search-001` / `full_ofone` / `agentic_coding` / repeat 1
- `case-public-sector-ai-policy-audit-001` / `direct_answer` / `agentic_coding` / repeat 1
- `case-public-sector-ai-policy-audit-001` / `light_structured` / `agentic_coding` / repeat 1
- `case-public-sector-ai-policy-audit-001` / `full_ofone` / `agentic_coding` / repeat 1
- `case-strategic-gated-diligence-001` / `direct_answer` / `agentic_coding` / repeat 2
- `case-strategic-gated-diligence-001` / `light_structured` / `agentic_coding` / repeat 2
- `case-strategic-gated-diligence-001` / `full_ofone` / `agentic_coding` / repeat 2
- `case-scientific-mechanism-check-001` / `direct_answer` / `agentic_coding` / repeat 2
- `case-scientific-mechanism-check-001` / `light_structured` / `agentic_coding` / repeat 2
- `case-scientific-mechanism-check-001` / `full_ofone` / `agentic_coding` / repeat 2
- `case-regulated-wastewater-market-entry-001` / `direct_answer` / `agentic_coding` / repeat 2
- `case-regulated-wastewater-market-entry-001` / `light_structured` / `agentic_coding` / repeat 2
- `case-regulated-wastewater-market-entry-001` / `full_ofone` / `agentic_coding` / repeat 2
- `case-formal-proof-search-001` / `direct_answer` / `agentic_coding` / repeat 2
- `case-formal-proof-search-001` / `light_structured` / `agentic_coding` / repeat 2
- `case-formal-proof-search-001` / `full_ofone` / `agentic_coding` / repeat 2
- `case-public-sector-ai-policy-audit-001` / `direct_answer` / `agentic_coding` / repeat 2
- `case-public-sector-ai-policy-audit-001` / `light_structured` / `agentic_coding` / repeat 2
- `case-public-sector-ai-policy-audit-001` / `full_ofone` / `agentic_coding` / repeat 2
- `case-strategic-gated-diligence-001` / `direct_answer` / `agentic_coding` / repeat 3
- `case-strategic-gated-diligence-001` / `light_structured` / `agentic_coding` / repeat 3
- `case-strategic-gated-diligence-001` / `full_ofone` / `agentic_coding` / repeat 3
- `case-scientific-mechanism-check-001` / `direct_answer` / `agentic_coding` / repeat 3
- `case-scientific-mechanism-check-001` / `light_structured` / `agentic_coding` / repeat 3
- `case-scientific-mechanism-check-001` / `full_ofone` / `agentic_coding` / repeat 3
- `case-regulated-wastewater-market-entry-001` / `direct_answer` / `agentic_coding` / repeat 3
- `case-regulated-wastewater-market-entry-001` / `light_structured` / `agentic_coding` / repeat 3
- `case-regulated-wastewater-market-entry-001` / `full_ofone` / `agentic_coding` / repeat 3
- `case-formal-proof-search-001` / `direct_answer` / `agentic_coding` / repeat 3
- `case-formal-proof-search-001` / `light_structured` / `agentic_coding` / repeat 3
- `case-formal-proof-search-001` / `full_ofone` / `agentic_coding` / repeat 3
- `case-public-sector-ai-policy-audit-001` / `direct_answer` / `agentic_coding` / repeat 3
- `case-public-sector-ai-policy-audit-001` / `light_structured` / `agentic_coding` / repeat 3
- `case-public-sector-ai-policy-audit-001` / `full_ofone` / `agentic_coding` / repeat 3
- `case-strategic-gated-diligence-001` / `direct_answer` / `frontier_reasoning` / repeat 1
- `case-strategic-gated-diligence-001` / `light_structured` / `frontier_reasoning` / repeat 1
- `case-strategic-gated-diligence-001` / `full_ofone` / `frontier_reasoning` / repeat 1
- `case-regulated-wastewater-market-entry-001` / `direct_answer` / `frontier_reasoning` / repeat 1
- `case-regulated-wastewater-market-entry-001` / `light_structured` / `frontier_reasoning` / repeat 1
- `case-regulated-wastewater-market-entry-001` / `full_ofone` / `frontier_reasoning` / repeat 1
- `case-formal-proof-search-001` / `direct_answer` / `frontier_reasoning` / repeat 1

Run 06 independently adjudicated the first slice. It accepted the direct-answer and light-structured slots for later aggregate scoring, but rejected the full-OfOne slot because the artifact identity is copied from `case-strategy-micro-001` rather than bound to `case-strategic-gated-diligence-001`.

Run 07 hardened the benchmark workflow, then the first full-OfOne slot was rerun as a remedial record:

- `case-strategic-gated-diligence-001` / `full_ofone` / `agentic_coding` / repeat 1 / remedial rerun 1

The original excluded full-OfOne run remains immutable evidence. The remedial rerun is tracked outside the original 90-slot count and can replace the excluded original only for future aggregate scoring after review.

The scientific mechanism slice completed after the remedial rerun. All three scientific `agentic_coding` repeat-1 arms passed local pre-score compliance. The full-OfOne scientific artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, rendering, and patch artifacts.

The regulated wastewater slice completed after the scientific slice. All three regulated wastewater `agentic_coding` repeat-1 arms passed local pre-score compliance. The full-OfOne regulated wastewater artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, rendering, and patch artifacts.

The formal proof-search slice completed after the regulated wastewater slice. All three formal `agentic_coding` repeat-1 arms passed local pre-score compliance. The full-OfOne formal artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, rendering, and patch artifacts.

The public-sector AI policy audit slice completed after the formal proof-search slice. All three policy-audit `agentic_coding` repeat-1 arms passed local pre-score compliance. The full-OfOne policy artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, Audit rendering, patch, review-log, and local review artifacts. The slice is published and Pages-confirmed.

The strategic gated diligence repeat-2 slice completed after the policy-audit slice. All three strategic repeat-2 `agentic_coding` arms passed local pre-score compliance. The full-OfOne strategic repeat-2 artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, Map rendering, patch, and local review artifacts. The slice is published and Pages-confirmed after public commit `732c6c8`.

The scientific mechanism repeat-2 slice completed after the strategic repeat-2 slice. All three scientific repeat-2 `agentic_coding` arms passed local pre-score compliance. The full-OfOne scientific repeat-2 artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, Map rendering, patch, and local review artifacts. The slice is published and Pages-confirmed after public commit `7694e7f`.

The regulated wastewater repeat-2 slice completed after the scientific repeat-2 slice. All three regulated wastewater repeat-2 `agentic_coding` arms passed local pre-score compliance. The full-OfOne regulated wastewater repeat-2 artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, Map rendering, patch, and local review artifacts. The slice is published and Pages-confirmed after public commit `f62f7c9`.

The formal proof-search repeat-2 slice completed after the regulated wastewater repeat-2 slice. All three formal repeat-2 `agentic_coding` arms passed local pre-score compliance. The full-OfOne formal repeat-2 artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, Map rendering, patch, and local review artifacts. The slice is published and Pages-confirmed after public commit `2287da0`.

The public-sector AI policy audit repeat-2 slice completed after the formal proof-search repeat-2 slice. All three policy-audit repeat-2 `agentic_coding` arms passed local pre-score compliance. The full-OfOne policy repeat-2 artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, Audit rendering, patch, review-log objects, and local review artifacts. The slice is published and Pages-confirmed after public commit `4499601`.

The strategic gated diligence repeat-3 slice completed after the public-sector AI policy audit repeat-2 slice. All three strategic repeat-3 `agentic_coding` arms passed local pre-score compliance. The full-OfOne strategic repeat-3 artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, Map rendering, patch, and local review artifacts. The slice is published and Pages-confirmed after public commit `83a68e9`.

The scientific mechanism repeat-3 slice completed after the strategic gated diligence repeat-3 slice. All three scientific repeat-3 `agentic_coding` arms passed local pre-score compliance. The full-OfOne scientific repeat-3 artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, Map rendering, patch, and local review artifacts. The slice is published and Pages-confirmed after public commit `b770b96`.

The regulated wastewater repeat-3 slice completed after the scientific mechanism repeat-3 slice. All three regulated wastewater repeat-3 `agentic_coding` arms passed local pre-score compliance. The full-OfOne regulated wastewater repeat-3 artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, Map rendering, patch, and local review artifacts. The slice is published and Pages-confirmed after public commit `046282e`.

The formal proof-search repeat-3 slice completed after the regulated wastewater repeat-3 slice. All three formal repeat-3 `agentic_coding` arms passed local pre-score compliance. The full-OfOne formal repeat-3 artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, Map rendering, patch, and local review artifacts. The slice is published and Pages-confirmed after public commit `35fdd24`.

The public-sector AI policy audit repeat-3 slice completed after the formal proof-search repeat-3 slice. All three policy-audit repeat-3 `agentic_coding` arms passed local pre-score compliance. The full-OfOne policy repeat-3 artifact is case-native, schema-valid, benchmark-trace-bound, and includes validator, Audit rendering, patch, review-log objects, and local review artifacts. The slice is published and Pages-confirmed after public commit `bb8b474`.

The next predeclared model family is `frontier_reasoning`. A strategic gated diligence repeat-1 packet exists at `benchmarks/runs/2026-05-17-batch-01/frontier-run-packets/2026-05-18-strategic-gated-diligence-frontier-r1.md` for separate direct-answer, light-structured, and full-OfOne frontier runs. The direct-answer arm completed in ChatGPT Deep Research at https://chatgpt.com/c/6a0e3e09-fd6c-83e8-a914-36445d70d090 with visible report metadata `Research completed in 17m`, `6 citations`, and `81 searches`. Its raw Markdown was harvested to `benchmarks/runs/2026-05-17-batch-01/outputs/2026-05-17-batch-01__case-strategic-gated-diligence-001__direct_answer__frontier_reasoning__r1.md`, locally reviewed at `benchmarks/reviews/2026-05-17-batch-01/2026-05-17-batch-01__case-strategic-gated-diligence-001__direct_answer__frontier_reasoning__r1.md`, and accepted as aggregate-eligible. The light-structured arm completed in the separate ChatGPT Deep Research conversation at https://chatgpt.com/c/6a0e7bcd-43b0-83e8-9a92-5195521c42fe with visible completed report title `Benchmark Raw Output` and run metadata `Status: completed`; its raw Markdown was harvested to `benchmarks/runs/2026-05-17-batch-01/outputs/2026-05-17-batch-01__case-strategic-gated-diligence-001__light_structured__frontier_reasoning__r1.md`, locally reviewed at `benchmarks/reviews/2026-05-17-batch-01/2026-05-17-batch-01__case-strategic-gated-diligence-001__light_structured__frontier_reasoning__r1.md`, and accepted as aggregate-eligible. The full-OfOne frontier arm completed in its separate ChatGPT Deep Research conversation at https://chatgpt.com/c/6a0e8476-9f6c-83e8-b201-ff3f97fae18b with visible metadata `Research completed in 18m`, `5 citations`, `9 searches`, title `Benchmark Raw Output`, and run metadata `Status: completed`. Its raw Markdown was harvested to `benchmarks/runs/2026-05-17-batch-01/outputs/2026-05-17-batch-01__case-strategic-gated-diligence-001__full_ofone__frontier_reasoning__r1.md`; artifact JSON, computed validator JSON, rendering, patch report, and local review were added. The full-OfOne frontier slot is completed/reviewed but excluded before aggregate scoring because computed semantic validation failed. Remedial frontier full-OfOne rerun 1 completed at https://chatgpt.com/c/6a0e8efd-2234-83e8-af43-a7e25266034d and was harvested to `benchmarks/runs/2026-05-17-batch-01/outputs/2026-05-17-batch-01__case-strategic-gated-diligence-001__full_ofone__frontier_reasoning__r1__rerun1.md`, but it was rejected before matrix insertion because it returned an advisory research report rather than the required benchmark raw output package. Remedial frontier full-OfOne rerun 2 completed at https://chatgpt.com/c/6a0ea350-3584-83e8-9d3e-ab7759c489f6 with visible metadata `Research completed in 44m`, `1 citation`, and `20 searches`; its raw Markdown, artifact JSON, computed validator JSON, rendering, patch report, and local review are preserved, but computed local validation failed missing evidence `movement_jobs` fields and a tradeoff reversal-condition defect. Rerun 2 is rejected before aggregate scoring and is not inserted into `remedial_runs`. Remedial frontier full-OfOne rerun 3 completed at https://chatgpt.com/c/6a0eb57b-6b08-83e8-a3e2-16e26adc497f with visible metadata `Research completed in 19m`, `5 citations`, `23 searches`, `21 May`, `5 sources`, and title `Benchmark Raw Output`; its raw Markdown, artifact JSON, computed validator JSON, rendering, patch report, and local review are preserved. Rerun 3 is rejected before aggregate scoring because the raw export omitted exact top-level run metadata and computed local validation failed current-schema `benchmark_trace` requirements plus relation legality for edges `X2`, `X3`, and `X4`. It remains outside `remedial_runs` and is not aggregate-eligible. No frontier aggregate comparison is supported.

Remedial frontier full-OfOne rerun 4 completed at https://chatgpt.com/c/6a0ec0a2-3814-83e8-8f86-23b625eace67 as a meta/advisory report titled `Running an Unspecified OfOne Benchmark Packet Exactly`; it is preserved as failed evidence outside artifact extraction, `remedial_runs`, aggregate eligibility, aggregate comparison, and superiority claims. Same-shape attachment-led Deep Research reruns are now barred for this slot. The Mode A controlled execution contract at `benchmarks/runs/2026-05-17-batch-01/frontier-run-packets/2026-05-21-strategic-gated-diligence-frontier-full-r1-mode-a-contract.md` produced `2026-05-17-batch-01__case-strategic-gated-diligence-001__full_ofone__frontier_reasoning__r1__rerun5` with raw output, artifact JSON, computed validator JSON, rendering, patch report, and local review. Rerun 5 is recorded in `remedial_runs` as `replace_for_aggregate_only` and aggregate-eligible replacement evidence for the excluded frontier full-OfOne repeat-1 slot only. The original excluded frontier run remains immutable, failed reruns 1-4 remain outside aggregate scoring, and no aggregate comparison or superiority claim is supported.

The regulated wastewater `frontier_reasoning` direct-answer repeat-1 arm completed in ChatGPT Deep Research at https://chatgpt.com/c/6a0ed3db-cccc-83e8-b84c-b3b1cb7b0bfa with visible report metadata `Research completed in 12m`, `16 citations`, `319 searches`, `21 May`, `16 sources`, report title `Benchmark Raw Output`, and run metadata `Status: completed`. Its raw Markdown was harvested to `benchmarks/runs/2026-05-17-batch-01/outputs/2026-05-17-batch-01__case-regulated-wastewater-market-entry-001__direct_answer__frontier_reasoning__r1.md`, locally reviewed at `benchmarks/reviews/2026-05-17-batch-01/2026-05-17-batch-01__case-regulated-wastewater-market-entry-001__direct_answer__frontier_reasoning__r1.md`, and accepted as aggregate-eligible. The regulated wastewater `frontier_reasoning` light-structured repeat-1 arm completed in the separate ChatGPT Deep Research conversation at https://chatgpt.com/c/6a0eda60-dc18-83e8-a888-e9e8ac1ab1fe with visible report metadata `Research completed in 22m`, `15 citations`, `236 searches`, `21 May`, `15 sources`, report title `Benchmark Raw Output`, and run metadata `Status: completed`. Its raw Markdown was harvested to `benchmarks/runs/2026-05-17-batch-01/outputs/2026-05-17-batch-01__case-regulated-wastewater-market-entry-001__light_structured__frontier_reasoning__r1.md`, locally reviewed at `benchmarks/reviews/2026-05-17-batch-01/2026-05-17-batch-01__case-regulated-wastewater-market-entry-001__light_structured__frontier_reasoning__r1.md`, and accepted as aggregate-eligible. The regulated wastewater `frontier_reasoning` full-OfOne repeat-1 arm completed in a third clean ChatGPT Deep Research conversation at https://chatgpt.com/c/6a0ee4ad-8854-83e8-866e-f671c12880da with visible metadata `Research completed in 45m`, `5 citations`, `22 searches`, `21 May`, `5 sources`, report title `Benchmark Raw Output`, and run metadata `Status: completed`. Its raw Markdown, artifact JSON, computed validator JSON, rendering, patch report, and local review are preserved, but computed local validation failed relation legality and relation-family checks. The full-OfOne slot is completed/reviewed but excluded before aggregate scoring. No regulated wastewater frontier aggregate comparison is supported.

The regulated wastewater frontier full-OfOne replacement path reused the governed Mode A repair protocol instead of launching another same-shape Deep Research rerun. Controlled rerun `2026-05-17-batch-01__case-regulated-wastewater-market-entry-001__full_ofone__frontier_reasoning__r1__rerun1` is case-native, benchmark-trace-bound, computed-validator-valid, locally reviewed, and recorded in `remedial_runs` as `replace_for_aggregate_only`. It is replacement evidence only; the original excluded full-OfOne frontier run remains immutable, and no aggregate comparison or superiority claim is supported.

The formal proof-search `frontier_reasoning` direct-answer repeat-1 arm completed in ChatGPT Deep Research at https://chatgpt.com/c/6a0f0a85-c75c-83e8-b0d0-4c15a041cb7b with visible report metadata `Research completed in 10m`, `8 citations`, `101 searches`, `21 May`, report title `Benchmark Raw Output`, and run metadata `Status: completed`. Its raw Markdown was harvested through the Chrome extension export flow to `benchmarks/runs/2026-05-17-batch-01/outputs/2026-05-17-batch-01__case-formal-proof-search-001__direct_answer__frontier_reasoning__r1.md`, locally reviewed at `benchmarks/reviews/2026-05-17-batch-01/2026-05-17-batch-01__case-formal-proof-search-001__direct_answer__frontier_reasoning__r1.md`, and accepted as aggregate-eligible. No formal proof-search frontier aggregate comparison is supported.

The formal proof-search `frontier_reasoning` light-structured repeat-1 arm is recorded in `execution-matrix.json` under `blocked_runs` with status `completed_report_visible`, not `completed`. The Chrome extension report at `research/deep-research-extension-report.json` records the completed ChatGPT conversation at https://chatgpt.com/c/6a0f1fe5-3494-83e8-9f92-1a2b732c4958 with visible metadata `Research completed in 9m`, `6 citations`, `120 searches`, report title `Benchmark Raw Output`, run metadata `Status: completed`, and the correct run ID visible. The raw Markdown is still unavailable because the report body and download control are inside ChatGPT's cross-origin Deep Research sandbox iframe. The manual recovery ledger at `research/deep-research-manual-recovery.json` remains the promotion gate. This slot must not be harvested, reviewed, marked complete, or counted as aggregate-eligible until a native Markdown export validates against the required run markers.

Current aggregate eligibility among reviewed local slots:

| Run slot | Eligibility | Reason |
| --- | --- | --- |
| strategic / `direct_answer` / `agentic_coding` / repeat 1 | eligible | Passes pre-score compliance gate. |
| strategic / `light_structured` / `agentic_coding` / repeat 1 | eligible | Passes pre-score compliance gate. |
| strategic / original `full_ofone` / `agentic_coding` / repeat 1 | excluded | Wrong-case copied artifact; schema-valid is not benchmark-valid. |
| strategic / `full_ofone` / `agentic_coding` / repeat 1 / remedial rerun 1 | eligible for future aggregate scoring as replacement | Case-native artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| strategic / `direct_answer` / `frontier_reasoning` / repeat 1 | eligible | Completed in ChatGPT Deep Research; harvested raw Markdown; passed pre-score compliance and local review. |
| strategic / `light_structured` / `frontier_reasoning` / repeat 1 | eligible | Completed in ChatGPT Deep Research; harvested raw Markdown; passed pre-score compliance and local review. |
| strategic / `full_ofone` / `frontier_reasoning` / repeat 1 | excluded | Case-bound artifact was harvested with validator/rendering/patch files, but computed semantic validation failed relation legality and option expected-effect checks. |
| strategic / `full_ofone` / `frontier_reasoning` / repeat 1 / remedial rerun 2 | not eligible | Package shape was harvested, but computed local validation failed evidence and tradeoff checks; the run is preserved outside aggregate scoring. |
| strategic / `full_ofone` / `frontier_reasoning` / repeat 1 / remedial rerun 3 | not eligible | Package sections were harvested, but exact top-level run metadata was missing and computed local validation failed benchmark-trace and relation-legality checks; the run is preserved outside aggregate scoring. |
| strategic / `full_ofone` / `frontier_reasoning` / repeat 1 / controlled rerun 5 | eligible for future aggregate scoring as replacement | Controlled Mode A package is case-native, validator-valid, benchmark-trace-bound, locally reviewed, and recorded as `replace_for_aggregate_only`. |
| scientific / `direct_answer` / `agentic_coding` / repeat 1 | eligible | Passes pre-score compliance gate. |
| scientific / `light_structured` / `agentic_coding` / repeat 1 | eligible | Passes pre-score compliance gate. |
| scientific / `full_ofone` / `agentic_coding` / repeat 1 | eligible | Case-native Map artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| regulated wastewater / `direct_answer` / `agentic_coding` / repeat 1 | eligible | Passes pre-score compliance gate. |
| regulated wastewater / `light_structured` / `agentic_coding` / repeat 1 | eligible | Passes pre-score compliance gate. |
| regulated wastewater / `full_ofone` / `agentic_coding` / repeat 1 | eligible | Case-native Map artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| regulated wastewater / `direct_answer` / `frontier_reasoning` / repeat 1 | eligible | Completed in ChatGPT Deep Research; harvested raw Markdown; passed pre-score compliance and local review. |
| regulated wastewater / `light_structured` / `frontier_reasoning` / repeat 1 | eligible | Completed in ChatGPT Deep Research; harvested raw Markdown; passed pre-score compliance and local review. |
| regulated wastewater / `full_ofone` / `frontier_reasoning` / repeat 1 | excluded | Completed in ChatGPT Deep Research; harvested raw/artifact/validator/rendering/patch/review, but computed semantic validation failed relation legality and relation-family checks. |
| regulated wastewater / `full_ofone` / `frontier_reasoning` / repeat 1 / controlled rerun 1 | eligible for future aggregate scoring as replacement | Controlled Mode A package is case-native, validator-valid, benchmark-trace-bound, locally reviewed, and recorded as `replace_for_aggregate_only`. |
| formal proof-search / `direct_answer` / `frontier_reasoning` / repeat 1 | eligible | Completed in ChatGPT Deep Research; harvested raw Markdown through Chrome extension export; passed pre-score compliance and local review. |
| formal proof-search / `direct_answer` / `agentic_coding` / repeat 1 | eligible | Passes pre-score compliance gate. |
| formal proof-search / `light_structured` / `agentic_coding` / repeat 1 | eligible | Passes pre-score compliance gate. |
| formal proof-search / `full_ofone` / `agentic_coding` / repeat 1 | eligible | Case-native Map artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| public-sector AI policy audit / `direct_answer` / `agentic_coding` / repeat 1 | eligible | Passes pre-score compliance gate. |
| public-sector AI policy audit / `light_structured` / `agentic_coding` / repeat 1 | eligible | Passes pre-score compliance gate. |
| public-sector AI policy audit / `full_ofone` / `agentic_coding` / repeat 1 | eligible | Case-native Audit artifact with benchmark trace binding, validator output, rendering, patch report, review-log objects, and local review. |
| strategic / `direct_answer` / `agentic_coding` / repeat 2 | eligible | Passes pre-score compliance gate. |
| strategic / `light_structured` / `agentic_coding` / repeat 2 | eligible | Passes pre-score compliance gate. |
| strategic / `full_ofone` / `agentic_coding` / repeat 2 | eligible | Case-native Map artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| strategic / `direct_answer` / `agentic_coding` / repeat 3 | eligible | Passes pre-score compliance gate. |
| strategic / `light_structured` / `agentic_coding` / repeat 3 | eligible | Passes pre-score compliance gate. |
| strategic / `full_ofone` / `agentic_coding` / repeat 3 | eligible | Case-native Map artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| scientific / `direct_answer` / `agentic_coding` / repeat 3 | eligible | Passes pre-score compliance gate. |
| scientific / `light_structured` / `agentic_coding` / repeat 3 | eligible | Passes pre-score compliance gate. |
| scientific / `full_ofone` / `agentic_coding` / repeat 3 | eligible | Case-native Map artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| regulated wastewater / `direct_answer` / `agentic_coding` / repeat 3 | eligible | Passes pre-score compliance gate. |
| regulated wastewater / `light_structured` / `agentic_coding` / repeat 3 | eligible | Passes pre-score compliance gate. |
| regulated wastewater / `full_ofone` / `agentic_coding` / repeat 3 | eligible | Case-native Map artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| scientific / `direct_answer` / `agentic_coding` / repeat 2 | eligible | Passes pre-score compliance gate. |
| scientific / `light_structured` / `agentic_coding` / repeat 2 | eligible | Passes pre-score compliance gate. |
| scientific / `full_ofone` / `agentic_coding` / repeat 2 | eligible | Case-native Map artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| regulated wastewater / `direct_answer` / `agentic_coding` / repeat 2 | eligible | Passes pre-score compliance gate. |
| regulated wastewater / `light_structured` / `agentic_coding` / repeat 2 | eligible | Passes pre-score compliance gate. |
| regulated wastewater / `full_ofone` / `agentic_coding` / repeat 2 | eligible | Case-native Map artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| formal proof-search / `direct_answer` / `agentic_coding` / repeat 2 | eligible | Passes pre-score compliance gate. |
| formal proof-search / `light_structured` / `agentic_coding` / repeat 2 | eligible | Passes pre-score compliance gate. |
| formal proof-search / `full_ofone` / `agentic_coding` / repeat 2 | eligible | Case-native Map artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| formal proof-search / `direct_answer` / `agentic_coding` / repeat 3 | eligible | Passes pre-score compliance gate. |
| formal proof-search / `light_structured` / `agentic_coding` / repeat 3 | eligible | Passes pre-score compliance gate. |
| formal proof-search / `full_ofone` / `agentic_coding` / repeat 3 | eligible | Case-native Map artifact with benchmark trace binding, validator output, rendering, patch report, and local review. |
| public-sector AI policy audit / `direct_answer` / `agentic_coding` / repeat 2 | eligible | Passes pre-score compliance gate. |
| public-sector AI policy audit / `light_structured` / `agentic_coding` / repeat 2 | eligible | Passes pre-score compliance gate. |
| public-sector AI policy audit / `full_ofone` / `agentic_coding` / repeat 2 | eligible | Case-native Audit artifact with benchmark trace binding, validator output, rendering, patch report, review-log objects, and local review. |
| public-sector AI policy audit / `direct_answer` / `agentic_coding` / repeat 3 | eligible | Passes pre-score compliance gate. |
| public-sector AI policy audit / `light_structured` / `agentic_coding` / repeat 3 | eligible | Passes pre-score compliance gate. |
| public-sector AI policy audit / `full_ofone` / `agentic_coding` / repeat 3 | eligible | Case-native Audit artifact with benchmark trace binding, validator output, rendering, patch report, review-log objects, and local review. |

No aggregate scoring or performance comparison has been completed. No performance or superiority claim is supported by this in-progress state.

Required before this file can move beyond `in_progress`:

- raw outputs for every case, arm, model family, and repeated run
- validation reports and patch artifacts for full-OfOne artifacts
- review notes using `benchmarks/reviews/2026-05-17-batch-01-review-template.md`
- aggregate score table
- excluded-run log with reasons
- limitation notes for blinding, source access, model variance, and reviewer expertise
