# Claim ledger

Every publication claim should resolve to one of these records or a newly documented analysis. Intervals are exploratory 95% paired bootstrap intervals, not simultaneous intervals.

## C01 · Auditable experiment

17 configurations × 100 questions yield 1,693 valid answers and 3,385 full-pass grades.

Evidence: `data/experiment.json`, `data/audit.json`.

Limit: Human gold review is pending; hashes verify content identity, not human correctness.

Analysis: descriptive.

## C02 · Source composition

751 selected rows from 29 publishers; 688 Nepali and 63 English.

Evidence: `data/dataset.json`, `data/publishers.json`.

Limit: A selected benchmark corpus, not a random population sample or 29 independent accounts.

Analysis: descriptive.

## C03 · Fable leads the observed composite

Fable 92.651; Sol 91.605. The paired gap is +1.046 [−0.066, 2.240].

Evidence: `data/leaderboard.json`, `data/paired-comparisons.json`.

Limit: Uncertainty in the closest pair does not establish equivalence.

Analysis: descriptive.

## C04 · Do not group all top 15 as equivalent

Fable exceeds Gemma by +4.713 [2.717, 7.317].

Evidence: `data/paired-comparisons.json`.

Limit: Pairwise intervals are exploratory and unadjusted. Adjacent comparisons do not define equivalence classes.

Analysis: descriptive.

## C05 · Prose leader differs from composite leader

Gemini 3.1 Pro has the highest observed Nepali prose mean, 9.025/10.

Evidence: `data/leaderboard.json`.

Limit: This is prose quality, not standalone Nepali fact accuracy; fine rank differences need paired evidence.

Analysis: descriptive.

## C06 · DeepSeek Flash improves the baseline

DeepSeek − Gemini 3 Flash = +2.570 [1.149, 3.996]; 69 wins, 1 tie, 30 losses.

Evidence: `data/production-comparison.json`.

Limit: Configured baseline, not verified live deployment; all-question effective score.

Analysis: paired exploratory.

## C07 · Lower recorded cost without a median speed gain

DeepSeek costs 68.1% less than Gemini; median 3.342 versus 3.2995 s.

Evidence: `data/leaderboard.json`.

Limit: Historical retained-call cost; retries and end-to-end production latency not fully captured.

Analysis: descriptive.

## C08 · Nepali-only input gain

Across 77 Nepali-only source questions, DeepSeek gains +3.125 composite and +0.260 Nepali prose.

Evidence: `data/production-comparison.json`.

Limit: Coverage remains bilingual. This is not a randomized input-language experiment.

Analysis: paired exploratory.

## C09 · Gains depend on story slice

Lead stories show −0.040 and mixed inputs −0.051 composite change for DeepSeek versus Gemini.

Evidence: `data/production-comparison.json`.

Limit: Broad intervals cross zero. These point estimates do not prove a loss or justify a routing policy.

Analysis: descriptive.

## C10 · Numeric regressions can be severe

DeepSeek changes 2.2 billion yuan to 220 million in Q72; its composite falls 22.20 relative to Gemini.

Evidence: `data/production-comparison.json`, `editorial/CASES.md`.

Limit: Purposively inspected source-grounding example, not a human-validated prevalence estimate.

Analysis: descriptive.

## C11 · Flash outperforms Pro on this task

DeepSeek Flash − Pro = +1.923 [0.419, 3.406], with approximately 71.1% lower recorded cost.

Evidence: `data/paired-comparisons.json`, `data/leaderboard.json`.

Limit: Only these endpoints and this workload; does not prove cheap/small models generally dominate.

Analysis: paired exploratory.

## C12 · Observed cost-quality frontier

Gemma, DeepSeek Flash, Sol and Fable form the observed cost-effective-score frontier.

Evidence: `data/leaderboard.json`.

Limit: Point-estimate frontier; uncertainty, pricing and alternative utility weights can change membership.

Analysis: descriptive.

## C13 · Length association reverses

Between model means r=0.966; within each model across stories r ranges from −0.517 to −0.095.

Evidence: `data/length-analysis.json`, `data/sensitivity.json`.

Limit: Confounding by story difficulty, gold density and model behavior; no causal claim about making answers longer.

Analysis: post-hoc exploratory.

## C14 · Judges differ in thresholds

Opus flags 1.474 unsupported claims per valid answer; Gemini 0.413, about 3.57× fewer.

Evidence: `data/judge-profile.json`.

Limit: Different segmentation and support thresholds; neither judge is an independent human truth label.

Analysis: descriptive.

## C15 · Repeatable judging is not repeated generation

Calibration exact coverage agreement is 0.9636 and 0.9561; kappa 0.9394 and 0.9250.

Evidence: `data/calibration.json`.

Limit: Same saved answers regraded; label dependence and human validity remain unresolved.

Analysis: descriptive.

## C16 · Settings differ materially

Six runs use mandatory reasoning with 4,096-token caps; four use JSON-object mode.

Evidence: `data/models.json`.

Limit: Cannot infer a causal reasoning or schema effect across different models.

Analysis: descriptive.

## C17 · Reliability differs by serving event

Seven failed generation slots are preserved and scored zero; Q4 includes different client/server failure classes.

Evidence: `data/failure-events.json`.

Limit: No nationality-wide refusal claim and no service-level reliability guarantee from n=100.

Analysis: descriptive.

## C18 · Weight choice changes interpretation

Fable stays first in several weight scenarios; Sol leads the language-only scenario.

Evidence: `data/sensitivity.json`.

Limit: Post-hoc diagnostic. Original primary rubric remains unchanged.

Analysis: post-hoc exploratory.

## C19 · Conditional rank stability

Fable ranks first in 94.55% of 2,000 question resamples.

Evidence: `data/sensitivity.json`.

Limit: Not probability of universal superiority; conditional on fixed outputs/judges.

Analysis: post-hoc exploratory.

## C20 · Metadata and grade audit

All 100 user-prompt hashes verify; scores reproduce within 3e−14; 20 grade records have fact-ID anomalies.

Evidence: `data/audit.json`.

Limit: Sanitization is a sensitivity, not a rewrite of raw grades. Endpoint snapshots may refresh on resume.

Analysis: descriptive.

## C21 · Evaluation costs exceed answer generation

Retained candidate calls cost $23.36; gold, calibration and full grading cost $354.65, about 15.2 times as much.

Evidence: `data/costs.json`.

Limit: This is retained-record accounting, not the complete account bill; excluded and overwritten attempts are not recovered.

Analysis: descriptive.
