Chapter 6 experiment coverage ledger¶
This ledger separates runnable code, pinned external sources, and direct acceptance evidence. A repository checkout, smoke test, or mechanism demo is never counted as completion of a broader manuscript experiment.
| Experiment | Manuscript acceptance scope | Current evidence | Audit status |
|---|---|---|---|
| 6-1 | Run τ²-bench, inspect multi-turn failures, and compare the dual-control telecom design with historical τ-bench | exp6-1-openrouter-gpt41mini-telecom-20260802-v1 retains the raw five-task telecom trajectory from pinned upstream 8d005b0…, 4/5 Pass@1, exact costs, and a wrong-line failure analysis. Upstream format and trial-count checks pass; full-task coverage fails as expected because the manuscript command deliberately samples five tasks rather than a complete leaderboard submission. Historical τ-bench remains pinned at 59a200c… for the design comparison. |
Complete saved bounded campaign |
| 6-2 | Personally complete simple/medium/hard tasks from GAIA, AndroidWorld, SWE-bench Verified, τ²-bench, Terminal-Bench, and OSWorld-Verified, retaining trajectories and official verification | experiment-6-2-human-benchmark/results.json binds the preregistered 18/18 Codex-as-human cases to their per-benchmark trajectories and first official results: 13 passed, 5 failed, 0 unscored. The report explains every task, operator trajectory, score, failure, and AndroidWorld/τ² compatibility boundary. |
Complete saved bounded campaign |
| 6-3 | Four-grade precision/recall/reasoning/proactivity rubric, examples/boundaries, hallucination veto | user-memory-system-evaluation/results/full_6_3_structured_rubric_evidence.json: 60 cases, 180/180 structured judgments, full scope, complete. |
Complete |
| 6-4 | Run Advanced JSON Cards, RAG, and hybrid over the same 60 cases; compare quality, steps, tools, latency, cost, and failure boundaries | user-memory-system-evaluation/results/full_6_4_60_cases_costed.json: 180/180 real trajectories, zero errors, complete pricing coverage and failure analysis. |
Complete |
| 6-5 | Multiple TTS providers/configurations × diverse corpus; direct-audio judge scores accuracy, naturalness, emotion, and voice consistency against reference audio | mistral_multimodal_20260730 retains 8/8 content-hashed OpenAI/Fish MP3 cells over four challenge categories, the fixed reference hash, exact four-dimension Voxtral judgments, and a recomputed complete gate. Earlier Google/OpenRouter/account failures remain as historical negative evidence. |
Complete saved campaign |
| 6-6 | Online Elo from real Arena votes, Bradley-Terry comparison, win matrix, official-ranking comparison, and historical animation | exp6-6-arena-20260731-v1 binds the 2.0 GB public snapshot by SHA-256 and processes all 1,799,991 source rows (1,670,250 accepted blind votes, 129 models). Chronological K=4 Elo and deterministic-bootstrap Bradley-Terry rankings have Spearman 0.787 / Kendall 0.606 agreement and 12/20 top-model overlap; empirical/predicted matrices, 17 monthly snapshots, three plots, and the D3 animation are content-hashed and independently revalidated. |
Complete saved campaign |
| 6-7 | Hold a neutral coding harness fixed while swapping GPT/Claude models; repeat localized, cross-cutting, and contract-sensitive tasks; measure pre-edit exploration, first-patch acceptance, rework, final tests, latency, files, and tokens | model-action-threshold/results/exp6-7-action-threshold-20260731-v1/manifest.json binds 18/18 real OpenRouter cells (2 models × 3 tasks × 3 trials), zero API errors, full trajectories, and independently passing artifact hashes. Both models pass all final tests; GPT-5.6-sol averages 6.89 pre-edit tool calls / 4.67 files versus Claude Sonnet 5 at 4.56 / 3.56. |
Complete saved campaign |
| 6-8 | End-to-end multi-turn cost decomposition and measured KV-cache/context-compression A/B | agent-cost-analysis/sample_trace.json retains the real four-arm, eight-turn token/cache/latency observations; README reports per-step, percentile, component, and 2×2 results. |
Complete saved campaign |
| 6-9 | Multi-provider 8K/32K/128K × 512/2048 workload at N≥100, p50/p95/p99, thinking, pricing, same-model provider pair, rate ramp, Agent trace, and 168-hour hourly availability | Campaign and strict analyzer implemented. model-benchmark/results/manifest.json reports only 29 smoke/readiness observations, no standard N=100 cells, no rate ramp/Agent-cost phase, and no 168-hour campaign. |
Incomplete—long-running/costly campaign |
| 6-10 | Full 4 embeddings × 3 rerankers × 2 main models × 60 cases with retrieval/task metrics and interaction analysis | user-memory-system-evaluation/results/full_6_9_60_case_matrix.json: 60 cases × 24 cells = 1,440/1,440 real trajectories, zero error records, zero unpriced usage, complete retrieval/task metrics and factorial interaction analysis, top-level and completion status complete; independently rechecked by user-memory-system-evaluation/validation/verify_full_matrix_20260731.py (ALL CHECKS PASSED). Executed under documented backend substitutions (BGE-M3 and OpenAI embeddings via OpenRouter as identical models, Qwen3-Embedding-8B substituting the endpoint-gated Doubao embedding, Doubao-LLM reranker substituting the unreachable BGE cross-encoder); see user-memory-system-evaluation/results/full_matrix_backend_readiness_20260731.json and candidate_backend_probes_20260731.json. |
Complete saved campaign |
| 6-11 | Diagnose AndroidWorld, test layered hypotheses, make cost-benefit decision, rerun full suite, and iterate | Real H1/H5/H5C paired runs exist. android-world/validation/manifest.json limits evidence to four Wi-Fi tasks on Pixel 9/API-35; the 116×5 candidate rerun and reference app image are absent. |
Incomplete—emulator/app provisioning |
| 6-12 | Real OpenVLA + RoboTwin2 move_can_pot evaluation with three RGB views, 14-D proprio/action, IID/OOD seeds, timing, failures, and action-chunk ablation |
Strict preflight/launch/analyze protocol implemented against pinned SimpleVLA-RL. openvla-robotwin2-eval/results/preflight-20260729.json records missing checkpoint/RoboTwin2/8-GPU CUDA stack. |
Incomplete—hardware and external artifacts |
External source identities and commands are maintained in README.md.