Outback Yak / Research paper 02

NepNewsBench v2.

What models retain,
invent and mistranslate.

Seventeen model configurations. One hundred Nepali news stories. A closer look at faithful coverage, bilingual prose and the cost of getting it wrong.

Outback Yak ResearchSeptember 2026Experiment 2.0

Prepared for editorial review · not yet a final release. Candidate calls: 15 September 2026 UTC. News-cluster window: 20 April–10 September 2026, Asia/Kathmandu. This evaluates consolidation of existing story clusters; clustering accuracy itself is not measured.

17API configurations
55.2Mrecorded tokens · US$378.01
751source rows / 29 publishers
3,385completed judge grades
Start with the evidence

A leaderboard with context.

Quality, price and latency answer different questions. Start with eight reference configurations, then explore the full lineup and source-language slices.

AnthropicOpenAIGoogleDeepSeekAlibaba / QwenMoonshot AIZ.AIxAI
Reference models span composite and prose leaders, price and latency.

8 configurations · 100 question slots per model · Effective score /100.

Recorded primary rubric · Both judges averaged · All 100 stories
Shown rankConfigurationEffective score /100Valid / attemptsNE /10USD /100 attemptsMedian / P95 s
01Claude Fable 5.1Anthropic92.6595% interval 91.76–93.52100 / 1000 generation failures8.995$8.89600 missing cost records19.44 / 28.18
02GPT-5.6 SolOpenAI91.6195% interval 90.33–92.77100 / 1000 generation failures8.950$0.83150 missing cost records15.67 / 20.17
03DeepSeek V4.1 FlashDeepSeek91.0395% interval 89.74–92.23100 / 1000 generation failures8.835$0.06990 missing cost records3.34 / 4.43
04Grok 4.6xAI90.9395% interval 89.62–92.17100 / 1000 generation failures8.755$0.92360 missing cost records14.30 / 21.92
05GLM 5.3Z.AI90.8195% interval 89.49–92.00100 / 1000 generation failures8.625$0.73900 missing cost records15.83 / 23.33
06Gemini 3.1 ProGoogle AI Studio90.0395% interval 88.44–91.52100 / 1000 generation failures9.025$1.12740 missing cost records6.36 / 13.26
07Gemini 3.8 FlashGoogle AI Studio89.5095% interval 88.10–90.86100 / 1000 generation failures8.815$0.30070 missing cost records3.04 / 4.90
08Gemma 4 31BVenice87.9495% interval 85.36–89.9899 / 1001 generation failures8.672$0.03820 missing cost records9.78 / 18.29

Effective = 40% core recall + 30% precision proxy + 15% Nepali prose + 15% English prose (prose scaled to /100). Generation failures score zero. Other quality measures use valid answers; prose is /10. Costs are retained September 2026 USD, normalized to 100 selected attempts; latency includes retained calls. Intervals: 5,000 question-bootstrap draws, fixed outputs and judges; unadjusted. Shown rank is within the selected configurations.

Cost, speed & quality

The Pareto frontier.

Recorded cost of 100 attempts, on a logarithmic scale, against effective score. The dashed step marks the frontier: no displayed alternative is cheaper with an equal or higher score, or scores higher at the same cost.
On the frontierOther configurationsScroll horizontally to explore →
Effective score /1008082848688909294$0.05$0.1$0.25$0.5$1$2.5$5$10Claude Fable 5.1 · 92.65/100 · $8.8960 /100 attempts · on the displayed frontierGPT-5.6 Sol · 91.61/100 · $0.8315 /100 attempts · on the displayed frontierDeepSeek V4.1 Flash · 91.03/100 · $0.0699 /100 attempts · on the displayed frontierGrok 4.6 · 90.93/100 · $0.9236 /100 attempts · comparison modelGLM 5.3 · 90.81/100 · $0.7390 /100 attempts · comparison modelGemini 3.1 Pro · 90.03/100 · $1.1274 /100 attempts · comparison modelGemini 3.8 Flash · 89.50/100 · $0.3007 /100 attempts · comparison modelGemma 4 31B · 87.94/100 · $0.0382 /100 attempts · on the displayed frontiercost of 100 attempts (USD, log scale)

Hover, tap or keyboard-focus a point for its model and exact values.

Frontier: Gemma 4 31B → DeepSeek V4.1 Flash → GPT-5.6 Sol → Fable 5.1. 100 slots per model; generation failures = zero. The axes expand to include all plotted scores. The chart updates with the selected configurations, stories and scoring view, and always plots effective score. Point estimates only; frontier membership does not imply statistically significant differences.

Judge views and exploratory scoring weights

These controls reset to all stories. Weight scenarios were chosen after seeing the results. They do not replace the primary rubric; their individual-model intervals are not computed. Single-judge prose uses available valid grades; the Gemini view has one missing Qwen 27B grade.

All 100 slots; shared-success quality uses 96 questions. Two thousand conditional rank resamples.
ModelOriginalSanitizedConsensus coreShared successConditional rank 95% range
Claude Fable 5.192.6592.65093.3292.741.0–2.0
Claude Haiku 4.584.8384.83485.0884.8515.0–16.0
Claude Opus 591.2691.26091.8592.132.0–10.0
Claude Sonnet 588.4188.41389.0289.209.0–15.0
DeepSeek V4.1 Flash91.0391.02291.5691.212.0–8.0
DeepSeek V4 Pro89.1089.10389.6789.248.0–14.0
Gemini 3.1 Pro90.0390.02891.0990.174.0–11.0
Gemini 3.8 Flash89.5089.50090.4289.496.0–13.0
Gemini 3 Flash88.4688.45489.2288.519.0–15.0
Gemma 4 31B87.9487.93888.8889.119.0–15.0
GLM 5.390.8190.80591.3390.923.0–8.0
GPT-5.6 Luna89.8089.80291.0989.885.0–13.0
GPT-5.6 Sol91.6191.60592.6291.841.0–6.0
Grok 4.690.9390.93491.6991.162.0–8.0
Kimi K389.4389.42890.0191.283.0–15.0
Qwen 3.8 27B81.5781.56382.6582.5217.0–17.0
Qwen 3.8 Max88.5888.58389.2689.468.0–15.0

Sanitization fixes malformed fact IDs; consensus core restricts core recall to three-extractor agreement. Neither establishes human validity. Rank ranges are conditional on frozen outputs and judges.

Compare two configurations

A paired interval measures the difference on the same stories. Overlapping individual-model intervals cannot answer that question.

Claude Fable 5.1versusGPT-5.6 Sol

Loading paired comparison…

Abstract

We evaluate 17 model–provider configurations on bilingual news consolidation using 100 selected story questions from K cha khabar. The questions contain 751 headline/excerpt rows from 29 publishers, including 688 Nepali and 63 English rows. Each candidate produces English and Nepali headlines and summaries, a slug, and typed entities. The candidate runs use OpenRouter with a pinned serving provider and saved endpoint metadata. A common production-derived prompt is verified by hashes; provider constraints create explicitly recorded differences in reasoning, temperature support and response format.

Three model passes construct source-linked fact sheets and a fourth pass adjudicates them. Two full-pass judges score 1,693 valid candidate outputs against 1,647 adjudicated facts, producing 3,385 completed grades. The failure-inclusive composite weights core-fact recall, a coverage-based precision proxy and two language-quality scores. We reconstruct all reported scores from raw records and report paired question-bootstrap intervals, judge-specific results, source-language slices and concrete errors.

Claude Fable 5.1 has the highest observed composite, 92.65/100. Gemini 3.1 Pro has the highest observed Nepali prose mean, 9.025/10. The observed cost–quality frontier spans Gemma, DeepSeek V4.1 Flash, GPT-5.6 Sol and Claude Fable 5.1, exposing different trade-offs between coverage, prose and recorded cost. Numeric, calendar and attribution failures remain even in high-scoring configurations. Across models, average output length strongly correlates with supporting-fact recall; within every model, the same association across stories is negative. These results describe workload-specific trade-offs and motivate further controlled evaluation. They do not establish human-level quality, identical decoding conditions or independently verified real-world truth.

Why this task needs its own evaluation

A news aggregator needs a short, faithful account of several reports that may differ in names, dates, emphasis and detail. In Nepal, it also needs parallel Nepali and English copy. A model can write natural Nepali while selecting the wrong facts. It can preserve an amount in Nepali and mistranslate it into English. It can turn an analyst's interpretation into a politician's declared intention. Those outcomes are consequential even when the paragraph sounds polished.

NepNewsBench evaluates this concrete consolidation task. It asks whether a model can use the supplied rows, retain the story's central facts, avoid additions the rows do not support, and produce usable copy in both languages. The task does not assess all Nepali capabilities. It does not measure open-domain knowledge, full-article reading, search quality, classification accuracy or the truth of the underlying news reports.

The important unit is the full model–provider configuration. A model ID alone does not describe the decoding policy, quantization, structured-output support or routing used to obtain an answer. Readers should be able to inspect those settings before interpreting a leaderboard.

From source evidence to a scored bilingual brief
  1. 01 / SelectOne story, 3–10 publisher rows; headline, timestamp and optional excerpt.
  2. 02 / GenerateEN + NE headline and summary, slug and typed entity list.
  3. 03 / Build goldThree fact extractions → one adjudicated, source-linked sheet.
  4. 04 / EvaluateTwo anonymous pointwise grades → components, failures and paired uncertainty.

The “clustering” in the earlier name describes the input unit. Neither release measures how accurately the upstream system assigns articles to clusters.

What changed from NepNewsCluster v0.3

The previous online paper (v0.3, May 2026) described 107 questions, 1,310 snippets and 15 models, with a three-axis rubric applied primarily by one judge. V2 changes the question construction, generation provenance and evaluation design. It uses 100 selected stories, a source-linked fact-sheet representation, pointwise grading of one anonymous candidate at a time, two full judging passes and repeated judge calibration. It records prompt hashes, model and provider identities, response formats, reasoning policy and per-call cost.

These changes make the experiment easier to inspect and its scoring easier to reproduce. They do not make the v2 score a direct continuation of the v0.3 score. An 88 in one version cannot be read as a gain over an 81 in the other. Models, dates, selected stories and the mathematical meaning of the score differ.

The experiment files describe an internally frozen protocol and document amendments. There is no independent public registration timestamp in this package. We therefore describe the protocol as internally frozen, rather than publicly preregistered.

DimensionNepNewsCluster v0.3NepNewsBench v2What improves / what remains
Question selection107 clusters stratified primarily by publisher coverage, from about 9,000 active clusters.100 stories stratified by salience, breadth and hard input shapes, from a stated eligibility pool of about 17,600.Targets specific production risks; remains purposive, with some related stories.
Evidence per questionUp to 15 recent snippets, excerpts limited to 280 characters; 1,310 rows, 23 publishers.3–10 rows, one per publisher, peak-48h or initial-72h windows, untruncated saved excerpts; 751 rows, 29 publishers.Improves temporal and publisher balance; fewer rows is a design change, not a larger benchmark.
Reference standardNo explicit gold fact sheets.Three independent extractions plus adjudication; 1,647 source-linked facts.Makes credited and omitted facts inspectable. Gold is still machine-generated.
Grading unitOne primary judge compared anonymized candidate options; three holistic axes.One anonymous candidate per call, against sources and gold; two full-pass judges.Reduces dependence on the candidate lineup and exposes judge disagreement. Bias is not eliminated.
Score and failuresMean of NE prose, EN prose and topic/entity quality; primarily successful-output quality.40% core recall, 30% precision proxy, 15% each prose language; failed generations count zero.Separates factual coverage from prose and includes serving failures. The /100 scales are incompatible.
Calibration and repetitionInformal 20-question cross-judge check; earlier lineup and a second DeepSeek pass supplied limited repeat evidence.Twenty-question calibration; two repeats each for Opus and Gemini, one Astra pass; explicit agreement gates.Stronger measured judge repeatability. Candidate-generation variance remains unmeasured in v2.
Run provenanceMostly OpenRouter, some native APIs; nominal common temperature; rate-table cost estimates.All 17 via OpenRouter, pinned hosts, returned IDs checked, prompt hashes and actual settings recorded, retained usage.cost.More auditable serving conditions; endpoint snapshots can refresh on resume and decoding still differs.
Protocol and auditDescriptive rubric refined after an earlier run.Internally frozen plans, dated amendments, numerical reconstruction, explicit deviations.Better traceability; no independent public preregistration or completed human gold review.

Archive note: the May online paper covers 15 models and identifies itself as v0.3. The earlier blog describes 13 models and v0.2. The legacy PDF has v0.2 in its filename but v0.3 and 15 models inside. These artifacts remain available as historical records.

Dataset and source construction

The selected cluster window is 20 April–10 September 2026, expressed in Asia/Kathmandu time. The selection account describes a pool of approximately 17,600 eligible clusters with at least three articles and three publishers, excluding recurring beats. The preserved inputs support rebuilding the selected questions; the entire original production pool is not included, so the full selection from that population is not independently reproducible from the public numerical tables alone.

The 100 questions contain 16 lead stories, 32 salient stories, 29 breadth stories and 23 hard cases. The sampling procedure uses historical daily rankings, section quotas, difficult input shapes, deduplication and caps on related saga facets. This is a purposive, stratified test collection, not a random sample of all Nepal news. Some related events still recur.

For each selected story, the builder takes a 48-hour window ending on the peak ranking day, or an initial 72-hour window when no peak is available; it can fall back to the full cluster when publisher coverage is insufficient. It selects one row per publisher, preferring excerpts, classified rows and recency. When more than ten publishers remain, it samples evenly across time and preserves minority-language coverage where available. The final rows are ordered newest first. Exact logic and selected inputs are archived.

There are 751 selected rows: 688 Nepali (91.6%) and 63 English (8.4%). The questions divide into 77 Nepali-only, 17 mixed-language and six English-only inputs. All candidates still produce both output languages. The median Nepali excerpt length is 248 characters when empty excerpts are included, and 262 among nonempty excerpts. There are 67 headline-only rows, all Nepali: 8.9% of all selected rows and 9.7% of Nepali rows. These measured rates replace the rough 20% figure retained in earlier protocol prose.

The selected rows come from 29 publishers, listed with row and question counts in the appendix. “29 publishers” does not establish 29 editorially independent accounts: syndication and shared reporting can occur. Politics appears among the tags of 61 questions, but the theme tags overlap rather than partition the corpus. The corpus contains 943 gold entities and 1,647 gold facts: 426 core, 871 supporting and 350 detail facts.

The export preserves article IDs, publisher identity, source timestamps and the rows the models saw. It does not preserve canonical article URLs. Source links must be recovered from an authorized original export if they are to appear in a public appendix; publisher homepages are not substitutes for article-level citations.

Filter the numerical source inventory by publisher, row language, month, story track and excerpt availability. Filtering rows does not redefine the leaderboard.

All 29 publishers and their representation

Counts describe selected rows, not independent editorial accounts. One publisher contributes at most one row per selected story.

PublisherRowsQuestionsNE rowsEN rowsHeadline only
Ajako Artha18181800
Artha Sarokar99900
Baahrakhari191919017
BBC News Nepali22200
Bizmandu19191900
Global Aawaj26262600
Gorkhapatra Online52525200
The Himalayan Times11010
Himal Press50505000
Khabarhub58585800
Khabarhub (English)12120120
myRepublica33030
Nagarik News66666600
Nepal Lead99908
Nepal News (English)15150150
Nepal Press53535300
Nepal Samaya99909
Nepal Views111111011
News of Nepal32323200
OnlineKhabar62626200
OnlineKhabar (English)77070
Pradesh Khabar34343400
Rajdhani National Daily33333300
Ratopati73737301
Ratopati (English)15150150
The Rising Nepal10100100
Setopati34343406
Thaha Khabar181818015
Ujyaalo Online11100

Candidate generation and provenance

Seventeen completed API configurations define the leaderboard. They span Anthropic, OpenAI, Google, xAI, DeepSeek, Alibaba/Qwen, Moonshot AI and Z.AI model families. Gemma is served by Venice. All included candidate API runs use OpenRouter. Earlier agent-transport runs and excluded candidates are not silently pooled into the leaderboard.

The runner requests a single pinned provider with fallbacks disabled and checks the returned model and provider names. Endpoint snapshots record supported parameters, prices, context length and reported quantization. Unknown quantization is left unknown. Kimi records mxfp4, GLM fp8 and Gemma bf16. A provider pin narrows the serving configuration; it does not freeze remote model weights or prove that a preview alias cannot change.

Every candidate receives the same system prompt and the same question-specific user prompt. The package verifies the system hash and all 100 exported user-prompt hashes. The intended sampling temperature is 0.2, but some endpoints do not advertise support and the run metadata records it as not honoured. This flag describes the recorded capability check, not an independent measurement of a provider's internal sampling.

Eleven configurations disable reasoning. Six use the recorded minimum mandatory reasoning policy: Fable, Gemini 3.8 Flash, Gemini 3.1 Pro, Grok and GLM use low; Qwen Max uses minimal. Reasoning-enabled candidates receive 4,096 total output tokens; the others receive 1,536. Four configurations use JSON-object mode because their pinned endpoints lack structured-output support: both DeepSeek runs, GLM and Qwen 27B. The others request strict JSON schema. Consequently, v2 compares practical supported configurations, not identical decoding conditions. It cannot isolate the causal effect of reasoning.

There is one retained generation per model-question slot. A documented exception affects Kimi K3: after an erroneous resume policy, Q74 and Q87 were each resampled once following a schema violation. Q74 failed again; Q87 succeeded. These retained outcomes remain in the reported leaderboard. This prevents describing every cell as an untouched first sample. Transient transport failures may be retried. Candidate truncations, empty outputs and parser failures count as failed answers. The parser follows production leniency: it requires three core fields and normalizes optional fields; successful parsing does not prove that all six intended fields arrived with perfect strict-schema compliance.

Saved per-call cost and latency describe retained calls. Earlier overwritten retries and waiting time are not reconstructed into an account-wide bill or end-to-end service latency. Prices are historical experiment records, not a live price comparison.

Evaluation costs much more than obtaining the candidate answers. Retained candidate generation costs total $23.36; gold construction, calibration and full grading total $354.65, about 15.2 times as much. That expense buys inspectable fact sheets, repeated judging and a second full judge, rather than additional candidate samples. The total across these retained records is $378.01; it excludes overwritten attempts and other earlier work.

Inspect the exact configuration

17 of 17 candidate configurations. All use OpenRouter, with provider fallbacks disabled. Historical settings from saved calls; no current catalogue claims.

Claude Fable 5.1Anthropic · reasoning low · 4,096 tokens
Exact requested ID
anthropic/claude-fable-5.1
Pinned provider slug
anthropic
Format / requested temperature
json_schema / 0.2 · honoured flag: no
Quantization / snapshot UTC
unknown / 2026-09-15T14:50:52.701Z
Observed reasoning tokens
100 reporting calls; mean 2.89; max 289
Evaluation roles
Candidate · gold extractor
Returned providers
{"Anthropic":100}
Returned model IDs
{"anthropic/claude-fable-5.1":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "claude-fable-5-1",
  "label": "Claude Fable 5.1",
  "transport": "openrouter",
  "modelSlug": "anthropic/claude-fable-5.1",
  "reasoning": {
    "enabled": true,
    "effort": "low"
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 10,
    "outputPricePerM": 50
  },
  "providerPolicy": {
    "order": [
      "anthropic"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T14:50:52.701Z",
    "providerSlug": "anthropic",
    "providerName": "Anthropic",
    "tag": "anthropic",
    "quantization": "unknown",
    "contextLength": 1000000,
    "pricing": {
      "prompt": 10,
      "completion": 50
    },
    "supportedParameters": [
      "max_tokens",
      "stop",
      "reasoning",
      "include_reasoning",
      "tools",
      "structured_outputs",
      "response_format",
      "verbosity",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": true,
      "supported_efforts": [
        "max",
        "xhigh",
        "high",
        "medium",
        "low"
      ],
      "default_effort": "high"
    }
  },
  "temperatureHonoured": false,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 4096,
  "requestParams": {
    "model": "anthropic/claude-fable-5.1",
    "max_tokens": 4096,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "effort": "low"
    },
    "provider": {
      "order": [
        "anthropic"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T14:50:52.703Z",
  "finishedAt": "2026-09-15T15:02:04.223Z",
  "goldContributor": true,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Anthropic": 100
  },
  "servedModelIds": {
    "anthropic/claude-fable-5.1": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 2.89,
    "maxWhenReported": 289
  }
}
Claude Haiku 4.5Anthropic · reasoning off · 1,536 tokens
Exact requested ID
anthropic/claude-haiku-4.5
Pinned provider slug
anthropic
Format / requested temperature
json_schema / 0.2 · honoured flag: yes
Quantization / snapshot UTC
unknown / 2026-09-15T13:41:28.086Z
Observed reasoning tokens
100 reporting calls; mean 0; max 0
Evaluation roles
Candidate
Returned providers
{"Anthropic":100}
Returned model IDs
{"anthropic/claude-haiku-4.5":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "claude-haiku-4-5",
  "label": "Claude Haiku 4.5",
  "transport": "openrouter",
  "modelSlug": "anthropic/claude-haiku-4.5",
  "reasoning": {
    "enabled": false
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 1,
    "outputPricePerM": 5
  },
  "providerPolicy": {
    "order": [
      "anthropic"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T13:41:28.086Z",
    "providerSlug": "anthropic",
    "providerName": "Anthropic",
    "tag": "anthropic",
    "quantization": "unknown",
    "contextLength": 200000,
    "pricing": {
      "prompt": 1,
      "completion": 5
    },
    "supportedParameters": [
      "max_tokens",
      "top_p",
      "temperature",
      "stop",
      "reasoning",
      "include_reasoning",
      "tools",
      "tool_choice",
      "top_k",
      "structured_outputs",
      "response_format"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": false
    }
  },
  "temperatureHonoured": true,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 1536,
  "requestParams": {
    "model": "anthropic/claude-haiku-4.5",
    "max_tokens": 1536,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "enabled": false
    },
    "provider": {
      "order": [
        "anthropic"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T13:41:28.087Z",
  "finishedAt": "2026-09-15T13:47:42.248Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Anthropic": 100
  },
  "servedModelIds": {
    "anthropic/claude-haiku-4.5": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
Claude Opus 5Anthropic · reasoning off · 1,536 tokens
Exact requested ID
anthropic/claude-opus-5
Pinned provider slug
anthropic
Format / requested temperature
json_schema / 0.2 · honoured flag: no
Quantization / snapshot UTC
unknown / 2026-09-15T14:42:43.433Z
Observed reasoning tokens
100 reporting calls; mean 0; max 0
Evaluation roles
Candidate · full-pass judge
Returned providers
{"Anthropic":100}
Returned model IDs
{"anthropic/claude-opus-5":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "claude-opus-5",
  "label": "Claude Opus 5",
  "transport": "openrouter",
  "modelSlug": "anthropic/claude-opus-5",
  "reasoning": {
    "enabled": false
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 5,
    "outputPricePerM": 25
  },
  "providerPolicy": {
    "order": [
      "anthropic"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T14:42:43.433Z",
    "providerSlug": "anthropic",
    "providerName": "Anthropic",
    "tag": "anthropic",
    "quantization": "unknown",
    "contextLength": 1000000,
    "pricing": {
      "prompt": 5,
      "completion": 25
    },
    "supportedParameters": [
      "max_tokens",
      "stop",
      "reasoning",
      "include_reasoning",
      "tool_choice",
      "tools",
      "structured_outputs",
      "response_format",
      "verbosity",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": false,
      "default_enabled": true,
      "supported_efforts": [
        "max",
        "xhigh",
        "high",
        "medium",
        "low"
      ],
      "default_effort": "high"
    }
  },
  "temperatureHonoured": false,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 1536,
  "requestParams": {
    "model": "anthropic/claude-opus-5",
    "max_tokens": 1536,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "enabled": false
    },
    "provider": {
      "order": [
        "anthropic"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T14:42:43.434Z",
  "finishedAt": "2026-09-15T14:50:51.399Z",
  "goldContributor": false,
  "fullPassJudge": true,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Anthropic": 100
  },
  "servedModelIds": {
    "anthropic/claude-opus-5": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
Claude Sonnet 5Anthropic · reasoning off · 1,536 tokens
Exact requested ID
anthropic/claude-sonnet-5
Pinned provider slug
anthropic
Format / requested temperature
json_schema / 0.2 · honoured flag: no
Quantization / snapshot UTC
unknown / 2026-09-15T14:35:32.117Z
Observed reasoning tokens
100 reporting calls; mean 0; max 0
Evaluation roles
Candidate
Returned providers
{"Anthropic":100}
Returned model IDs
{"anthropic/claude-sonnet-5":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "claude-sonnet-5",
  "label": "Claude Sonnet 5",
  "transport": "openrouter",
  "modelSlug": "anthropic/claude-sonnet-5",
  "reasoning": {
    "enabled": false
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 2,
    "outputPricePerM": 10
  },
  "providerPolicy": {
    "order": [
      "anthropic"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T14:35:32.117Z",
    "providerSlug": "anthropic",
    "providerName": "Anthropic",
    "tag": "anthropic",
    "quantization": "unknown",
    "contextLength": 1000000,
    "pricing": {
      "prompt": 2,
      "completion": 10
    },
    "supportedParameters": [
      "max_tokens",
      "stop",
      "reasoning",
      "include_reasoning",
      "tools",
      "tool_choice",
      "structured_outputs",
      "response_format",
      "verbosity",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": false,
      "default_enabled": true,
      "supported_efforts": [
        "max",
        "xhigh",
        "high",
        "medium",
        "low"
      ],
      "default_effort": "high"
    }
  },
  "temperatureHonoured": false,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 1536,
  "requestParams": {
    "model": "anthropic/claude-sonnet-5",
    "max_tokens": 1536,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "enabled": false
    },
    "provider": {
      "order": [
        "anthropic"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T14:35:32.118Z",
  "finishedAt": "2026-09-15T14:42:03.437Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Anthropic": 100
  },
  "servedModelIds": {
    "anthropic/claude-sonnet-5": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
DeepSeek V4.1 FlashDeepSeek · reasoning off · 1,536 tokens
Exact requested ID
deepseek/deepseek-v4.1-flash
Pinned provider slug
deepseek
Format / requested temperature
json_object / 0.2 · honoured flag: yes
Quantization / snapshot UTC
unknown / 2026-09-15T22:24:22.295Z
Observed reasoning tokens
100 reporting calls; mean 0; max 0
Evaluation roles
Candidate
Returned providers
{"DeepSeek":100}
Returned model IDs
{"deepseek/deepseek-v4.1-flash":100}

The journal notes possible upstream aliasing between the DeepSeek endpoints; different returned IDs do not prove distinct underlying weights.

Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "deepseek-v4-1-flash",
  "label": "DeepSeek V4.1 Flash",
  "transport": "openrouter",
  "modelSlug": "deepseek/deepseek-v4.1-flash",
  "reasoning": {
    "enabled": false
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 0.15,
    "outputPricePerM": 0.6
  },
  "providerPolicy": {
    "order": [
      "deepseek"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T22:24:22.295Z",
    "providerSlug": "deepseek",
    "providerName": "DeepSeek",
    "tag": "deepseek",
    "quantization": "unknown",
    "contextLength": 1048576,
    "pricing": {
      "prompt": 0.15,
      "completion": 0.6
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "max_tokens",
      "temperature",
      "top_p",
      "stop",
      "frequency_penalty",
      "presence_penalty",
      "logprobs",
      "top_logprobs",
      "tools",
      "tool_choice",
      "response_format",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": false,
      "default_enabled": true,
      "supported_efforts": [
        "max",
        "high",
        "low"
      ],
      "default_effort": "high"
    }
  },
  "temperatureHonoured": true,
  "responseFormatMode": "json_object",
  "maxTokensSent": 1536,
  "requestParams": {
    "model": "deepseek/deepseek-v4.1-flash",
    "max_tokens": 1536,
    "temperature": 0.2,
    "response_format": "{ type: \"json_object\" } — endpoint lacks structured_outputs",
    "reasoning": {
      "enabled": false
    },
    "provider": {
      "order": [
        "deepseek"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T22:24:22.296Z",
  "finishedAt": "2026-09-15T22:26:16.593Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "DeepSeek": 100
  },
  "servedModelIds": {
    "deepseek/deepseek-v4.1-flash": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
DeepSeek V4 ProDeepSeek · reasoning off · 1,536 tokens
Exact requested ID
deepseek/deepseek-v4-pro-0813
Pinned provider slug
deepseek
Format / requested temperature
json_object / 0.2 · honoured flag: yes
Quantization / snapshot UTC
unknown / 2026-09-15T22:21:26.039Z
Observed reasoning tokens
100 reporting calls; mean 0; max 0
Evaluation roles
Candidate
Returned providers
{"DeepSeek":100}
Returned model IDs
{"deepseek/deepseek-v4-pro-0813":100}

The journal notes possible upstream aliasing between the DeepSeek endpoints; different returned IDs do not prove distinct underlying weights.

Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "deepseek-v4-pro-0813",
  "label": "DeepSeek V4 Pro (0813 GA)",
  "transport": "openrouter",
  "modelSlug": "deepseek/deepseek-v4-pro-0813",
  "reasoning": {
    "enabled": false
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 0.66,
    "outputPricePerM": 1.98
  },
  "providerPolicy": {
    "order": [
      "deepseek"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T22:21:26.039Z",
    "providerSlug": "deepseek",
    "providerName": "DeepSeek",
    "tag": "deepseek",
    "quantization": "unknown",
    "contextLength": 1048576,
    "pricing": {
      "prompt": 0.66,
      "completion": 1.9800000000000002
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "max_tokens",
      "temperature",
      "top_p",
      "stop",
      "frequency_penalty",
      "presence_penalty",
      "logprobs",
      "top_logprobs",
      "tools",
      "tool_choice",
      "response_format",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": false,
      "supported_efforts": [
        "max",
        "high",
        "low"
      ],
      "default_effort": "high"
    }
  },
  "temperatureHonoured": true,
  "responseFormatMode": "json_object",
  "maxTokensSent": 1536,
  "requestParams": {
    "model": "deepseek/deepseek-v4-pro-0813",
    "max_tokens": 1536,
    "temperature": 0.2,
    "response_format": "{ type: \"json_object\" } — endpoint lacks structured_outputs",
    "reasoning": {
      "enabled": false
    },
    "provider": {
      "order": [
        "deepseek"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T22:21:26.039Z",
  "finishedAt": "2026-09-15T22:24:20.845Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "DeepSeek": 100
  },
  "servedModelIds": {
    "deepseek/deepseek-v4-pro-0813": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
Gemini 3.1 ProGoogle AI Studio · reasoning low · 4,096 tokens
Exact requested ID
google/gemini-3.1-pro-preview
Pinned provider slug
google-ai-studio
Format / requested temperature
json_schema / 0.2 · honoured flag: yes
Quantization / snapshot UTC
unknown / 2026-09-15T14:14:23.854Z
Observed reasoning tokens
100 reporting calls; mean 203.49; max 1225
Evaluation roles
Candidate
Returned providers
{"Google AI Studio":100}
Returned model IDs
{"google/gemini-3.1-pro-preview":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "gemini-3-1-pro",
  "label": "Gemini 3.1 Pro (preview)",
  "transport": "openrouter",
  "modelSlug": "google/gemini-3.1-pro-preview",
  "reasoning": {
    "enabled": true,
    "effort": "low"
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 2,
    "outputPricePerM": 12
  },
  "providerPolicy": {
    "order": [
      "google-ai-studio"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T14:14:23.854Z",
    "providerSlug": "google-ai-studio",
    "providerName": "Google AI Studio",
    "tag": "google-ai-studio",
    "quantization": "unknown",
    "contextLength": 1048576,
    "pricing": {
      "prompt": 2,
      "completion": 12
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "max_tokens",
      "temperature",
      "top_p",
      "seed",
      "response_format",
      "tools",
      "tool_choice",
      "structured_outputs",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": true,
      "supported_efforts": [
        "high",
        "medium",
        "low"
      ],
      "default_effort": "medium"
    }
  },
  "temperatureHonoured": true,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 4096,
  "requestParams": {
    "model": "google/gemini-3.1-pro-preview",
    "max_tokens": 4096,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "effort": "low"
    },
    "provider": {
      "order": [
        "google-ai-studio"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T14:14:23.854Z",
  "finishedAt": "2026-09-15T14:18:43.970Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Google AI Studio": 100
  },
  "servedModelIds": {
    "google/gemini-3.1-pro-preview": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 203.49,
    "maxWhenReported": 1225
  }
}
Gemini 3.8 FlashGoogle AI Studio · reasoning low · 4,096 tokens
Exact requested ID
google/gemini-3.8-flash
Pinned provider slug
google-ai-studio
Format / requested temperature
json_schema / 0.2 · honoured flag: yes
Quantization / snapshot UTC
unknown / 2026-09-15T13:18:52.521Z
Observed reasoning tokens
100 reporting calls; mean 0; max 0
Evaluation roles
Candidate · gold extractor · full-pass judge
Returned providers
{"Google AI Studio":100}
Returned model IDs
{"google/gemini-3.8-flash":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "gemini-3-8-flash",
  "label": "Gemini 3.8 Flash",
  "transport": "openrouter",
  "modelSlug": "google/gemini-3.8-flash",
  "reasoning": {
    "enabled": true,
    "effort": "low"
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 0.75,
    "outputPricePerM": 3.75
  },
  "providerPolicy": {
    "order": [
      "google-ai-studio"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T13:18:52.521Z",
    "providerSlug": "google-ai-studio",
    "providerName": "Google AI Studio",
    "tag": "google-ai-studio",
    "quantization": "unknown",
    "contextLength": 1048576,
    "pricing": {
      "prompt": 0.75,
      "completion": 3.75
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "max_tokens",
      "temperature",
      "top_p",
      "seed",
      "response_format",
      "structured_outputs",
      "tool_choice",
      "tools",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": true,
      "default_enabled": true,
      "supported_efforts": [
        "high",
        "medium",
        "low"
      ],
      "default_effort": "medium"
    }
  },
  "temperatureHonoured": true,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 4096,
  "requestParams": {
    "model": "google/gemini-3.8-flash",
    "max_tokens": 4096,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "effort": "low"
    },
    "provider": {
      "order": [
        "google-ai-studio"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T13:18:52.521Z",
  "finishedAt": "2026-09-15T13:20:44.643Z",
  "goldContributor": true,
  "fullPassJudge": true,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Google AI Studio": 100
  },
  "servedModelIds": {
    "google/gemini-3.8-flash": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
Gemini 3 FlashGoogle AI Studio · reasoning off · 1,536 tokens
Exact requested ID
google/gemini-3-flash-preview
Pinned provider slug
google-ai-studio
Format / requested temperature
json_schema / 0.2 · honoured flag: yes
Quantization / snapshot UTC
unknown / 2026-09-15T13:16:55.716Z
Observed reasoning tokens
100 reporting calls; mean 0; max 0
Evaluation roles
Candidate
Returned providers
{"Google AI Studio":100}
Returned model IDs
{"google/gemini-3-flash-preview":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "gemini-3-flash",
  "label": "Gemini 3 Flash (preview) — prod primary",
  "transport": "openrouter",
  "modelSlug": "google/gemini-3-flash-preview",
  "reasoning": {
    "enabled": false
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 0.5,
    "outputPricePerM": 3
  },
  "providerPolicy": {
    "order": [
      "google-ai-studio"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T13:16:55.716Z",
    "providerSlug": "google-ai-studio",
    "providerName": "Google AI Studio",
    "tag": "google-ai-studio",
    "quantization": "unknown",
    "contextLength": 1048576,
    "pricing": {
      "prompt": 0.5,
      "completion": 3
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "max_tokens",
      "temperature",
      "top_p",
      "seed",
      "response_format",
      "tools",
      "tool_choice",
      "structured_outputs",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": false,
      "supported_efforts": [
        "high",
        "medium",
        "low",
        "minimal"
      ],
      "default_effort": "medium"
    }
  },
  "temperatureHonoured": true,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 1536,
  "requestParams": {
    "model": "google/gemini-3-flash-preview",
    "max_tokens": 1536,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "enabled": false
    },
    "provider": {
      "order": [
        "google-ai-studio"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T13:16:55.717Z",
  "finishedAt": "2026-09-15T13:18:50.989Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Google AI Studio": 100
  },
  "servedModelIds": {
    "google/gemini-3-flash-preview": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
Gemma 4 31BVenice · reasoning off · 1,536 tokens
Exact requested ID
google/gemma-4-31b-it
Pinned provider slug
venice
Format / requested temperature
json_schema / 0.2 · honoured flag: yes
Quantization / snapshot UTC
bf16 / 2026-09-15T13:20:46.958Z
Observed reasoning tokens
100 reporting calls; mean 0; max 0
Evaluation roles
Candidate
Returned providers
{"Venice":100}
Returned model IDs
{"google/gemma-4-31b-it":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "gemma-4-31b",
  "label": "Gemma 4 31B (it)",
  "transport": "openrouter",
  "modelSlug": "google/gemma-4-31b-it",
  "reasoning": {
    "enabled": false
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 0.12,
    "outputPricePerM": 0.36
  },
  "providerPolicy": {
    "order": [
      "venice"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T13:20:46.958Z",
    "providerSlug": "venice",
    "providerName": "Venice",
    "tag": "venice/bf16",
    "quantization": "bf16",
    "contextLength": 256000,
    "pricing": {
      "prompt": 0.12,
      "completion": 0.36
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "max_tokens",
      "temperature",
      "top_p",
      "stop",
      "frequency_penalty",
      "presence_penalty",
      "top_k",
      "response_format",
      "structured_outputs",
      "tools",
      "tool_choice",
      "logprobs",
      "top_logprobs"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": false,
      "default_enabled": false
    }
  },
  "temperatureHonoured": true,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 1536,
  "requestParams": {
    "model": "google/gemma-4-31b-it",
    "max_tokens": 1536,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "enabled": false
    },
    "provider": {
      "order": [
        "venice"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T13:20:46.958Z",
  "finishedAt": "2026-09-15T13:27:00.336Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Venice": 100
  },
  "servedModelIds": {
    "google/gemma-4-31b-it": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
GLM 5.3Z.AI · reasoning low · 4,096 tokens
Exact requested ID
z-ai/glm-5.3
Pinned provider slug
z-ai
Format / requested temperature
json_object / 0.2 · honoured flag: yes
Quantization / snapshot UTC
fp8 / 2026-09-15T14:05:13.751Z
Observed reasoning tokens
100 reporting calls; mean 4.8; max 108
Evaluation roles
Candidate
Returned providers
{"Z.AI":100}
Returned model IDs
{"z-ai/glm-5.3":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "glm-5-3",
  "label": "GLM 5.3",
  "transport": "openrouter",
  "modelSlug": "z-ai/glm-5.3",
  "reasoning": {
    "enabled": true,
    "effort": "low"
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 1.4,
    "outputPricePerM": 4.4
  },
  "providerPolicy": {
    "order": [
      "z-ai"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T14:05:13.751Z",
    "providerSlug": "z-ai",
    "providerName": "Z.AI",
    "tag": "z-ai/fp8",
    "quantization": "fp8",
    "contextLength": 1048576,
    "pricing": {
      "prompt": 1.4,
      "completion": 4.4
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "max_tokens",
      "temperature",
      "top_p",
      "tools",
      "tool_choice",
      "top_k",
      "response_format",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": true,
      "default_enabled": true,
      "supported_efforts": [
        "max",
        "high",
        "low"
      ],
      "default_effort": "max"
    }
  },
  "temperatureHonoured": true,
  "responseFormatMode": "json_object",
  "maxTokensSent": 4096,
  "requestParams": {
    "model": "z-ai/glm-5.3",
    "max_tokens": 4096,
    "temperature": 0.2,
    "response_format": "{ type: \"json_object\" } — endpoint lacks structured_outputs",
    "reasoning": {
      "effort": "low"
    },
    "provider": {
      "order": [
        "z-ai"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T14:05:13.752Z",
  "finishedAt": "2026-09-15T14:14:21.914Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Z.AI": 100
  },
  "servedModelIds": {
    "z-ai/glm-5.3": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 4.8,
    "maxWhenReported": 108
  }
}
GPT-5.6 LunaOpenAI · reasoning off · 1,536 tokens
Exact requested ID
openai/gpt-5.6-luna
Pinned provider slug
openai
Format / requested temperature
json_schema / 0.2 · honoured flag: no
Quantization / snapshot UTC
unknown / 2026-09-15T13:27:02.062Z
Observed reasoning tokens
100 reporting calls; mean 0; max 0
Evaluation roles
Candidate
Returned providers
{"OpenAI":100}
Returned model IDs
{"openai/gpt-5.6-luna":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "gpt-5-6-luna",
  "label": "GPT-5.6 Luna",
  "transport": "openrouter",
  "modelSlug": "openai/gpt-5.6-luna",
  "reasoning": {
    "enabled": false
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 0.2,
    "outputPricePerM": 1.2
  },
  "providerPolicy": {
    "order": [
      "openai"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T13:27:02.062Z",
    "providerSlug": "openai",
    "providerName": "OpenAI",
    "tag": "openai",
    "quantization": "unknown",
    "contextLength": 1050000,
    "pricing": {
      "prompt": 0.19999999999999998,
      "completion": 1.2
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "seed",
      "max_tokens",
      "response_format",
      "structured_outputs",
      "tools",
      "tool_choice",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": false,
      "default_enabled": true,
      "supported_efforts": [
        "max",
        "xhigh",
        "high",
        "medium",
        "low",
        "none"
      ],
      "default_effort": "medium"
    }
  },
  "temperatureHonoured": false,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 1536,
  "requestParams": {
    "model": "openai/gpt-5.6-luna",
    "max_tokens": 1536,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "enabled": false
    },
    "provider": {
      "order": [
        "openai"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T13:27:02.062Z",
  "finishedAt": "2026-09-15T13:29:46.919Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "OpenAI": 100
  },
  "servedModelIds": {
    "openai/gpt-5.6-luna": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
GPT-5.6 SolOpenAI · reasoning off · 1,536 tokens
Exact requested ID
openai/gpt-5.6-sol
Pinned provider slug
openai
Format / requested temperature
json_schema / 0.2 · honoured flag: no
Quantization / snapshot UTC
unknown / 2026-09-15T13:47:43.924Z
Observed reasoning tokens
100 reporting calls; mean 0; max 0
Evaluation roles
Candidate
Returned providers
{"OpenAI":100}
Returned model IDs
{"openai/gpt-5.6-sol":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "gpt-5-6-sol",
  "label": "GPT-5.6 Sol",
  "transport": "openrouter",
  "modelSlug": "openai/gpt-5.6-sol",
  "reasoning": {
    "enabled": false
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 2,
    "outputPricePerM": 10
  },
  "providerPolicy": {
    "order": [
      "openai"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T13:47:43.924Z",
    "providerSlug": "openai",
    "providerName": "OpenAI",
    "tag": "openai",
    "quantization": "unknown",
    "contextLength": 1050000,
    "pricing": {
      "prompt": 2,
      "completion": 10
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "seed",
      "max_tokens",
      "response_format",
      "structured_outputs",
      "tools",
      "tool_choice",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": false,
      "default_enabled": true,
      "supported_efforts": [
        "max",
        "xhigh",
        "high",
        "medium",
        "low",
        "none"
      ],
      "default_effort": "medium"
    }
  },
  "temperatureHonoured": false,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 1536,
  "requestParams": {
    "model": "openai/gpt-5.6-sol",
    "max_tokens": 1536,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "enabled": false
    },
    "provider": {
      "order": [
        "openai"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T13:47:43.925Z",
  "finishedAt": "2026-09-15T13:56:40.246Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "OpenAI": 100
  },
  "servedModelIds": {
    "openai/gpt-5.6-sol": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
Grok 4.6xAI · reasoning low · 4,096 tokens
Exact requested ID
x-ai/grok-4.6
Pinned provider slug
xai
Format / requested temperature
json_schema / 0.2 · honoured flag: yes
Quantization / snapshot UTC
unknown / 2026-09-15T13:56:41.552Z
Observed reasoning tokens
100 reporting calls; mean 353.26; max 959
Evaluation roles
Candidate
Returned providers
{"xAI":100}
Returned model IDs
{"x-ai/grok-4.6":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "grok-4-6",
  "label": "Grok 4.6",
  "transport": "openrouter",
  "modelSlug": "x-ai/grok-4.6",
  "reasoning": {
    "enabled": true,
    "effort": "low"
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 2,
    "outputPricePerM": 6
  },
  "providerPolicy": {
    "order": [
      "xai"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T13:56:41.552Z",
    "providerSlug": "xai",
    "providerName": "xAI",
    "tag": "xai",
    "quantization": "unknown",
    "contextLength": 500000,
    "pricing": {
      "prompt": 2,
      "completion": 6
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "structured_outputs",
      "response_format",
      "max_tokens",
      "temperature",
      "top_p",
      "seed",
      "logprobs",
      "top_logprobs",
      "tools",
      "tool_choice",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": true,
      "default_enabled": true,
      "supported_efforts": [
        "xhigh",
        "high",
        "medium",
        "low"
      ],
      "default_effort": "high"
    }
  },
  "temperatureHonoured": true,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 4096,
  "requestParams": {
    "model": "x-ai/grok-4.6",
    "max_tokens": 4096,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "effort": "low"
    },
    "provider": {
      "order": [
        "xai"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T13:56:41.552Z",
  "finishedAt": "2026-09-15T14:05:11.948Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "xAI": 100
  },
  "servedModelIds": {
    "x-ai/grok-4.6": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 353.26,
    "maxWhenReported": 959
  }
}
Kimi K3Moonshot AI · reasoning off · 1,536 tokens
Exact requested ID
moonshotai/kimi-k3
Pinned provider slug
moonshotai
Format / requested temperature
json_schema / 0.2 · honoured flag: no
Quantization / snapshot UTC
mxfp4 / 2026-09-15T15:15:26.266Z
Observed reasoning tokens
99 reporting calls; mean 0; max 0
Evaluation roles
Candidate
Returned providers
{"Moonshot AI":99,"unrecorded":1}
Returned model IDs
{"moonshotai/kimi-k3":99,"unrecorded":1}

Protocol exception: Q74 and Q87 were resampled once after schema failures. Q87 succeeded; Q74 remained a failure. Recorded quantization is mxfp4.

Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "kimi-k3",
  "label": "Kimi K3",
  "transport": "openrouter",
  "modelSlug": "moonshotai/kimi-k3",
  "reasoning": {
    "enabled": false
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 3,
    "outputPricePerM": 15
  },
  "providerPolicy": {
    "order": [
      "moonshotai"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T15:15:26.266Z",
    "providerSlug": "moonshotai",
    "providerName": "Moonshot AI",
    "tag": "moonshotai/mxfp4",
    "quantization": "mxfp4",
    "contextLength": 1048576,
    "pricing": {
      "prompt": 3,
      "completion": 15
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "max_tokens",
      "stop",
      "frequency_penalty",
      "presence_penalty",
      "structured_outputs",
      "response_format",
      "tool_choice",
      "tools",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": false,
      "default_enabled": true,
      "supported_efforts": [
        "max",
        "high",
        "low"
      ],
      "default_effort": "max"
    }
  },
  "temperatureHonoured": false,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 1536,
  "requestParams": {
    "model": "moonshotai/kimi-k3",
    "max_tokens": 1536,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "enabled": false
    },
    "provider": {
      "order": [
        "moonshotai"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T14:42:04.847Z",
  "finishedAt": "2026-09-15T15:16:18.451Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Moonshot AI": 99,
    "unrecorded": 1
  },
  "servedModelIds": {
    "moonshotai/kimi-k3": 99,
    "unrecorded": 1
  },
  "observedReasoningTokens": {
    "reportedCalls": 99,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
Qwen 3.8 27BAlibaba · reasoning off · 1,536 tokens
Exact requested ID
qwen/qwen3.8-27b
Pinned provider slug
alibaba
Format / requested temperature
json_object / 0.2 · honoured flag: yes
Quantization / snapshot UTC
unknown / 2026-09-15T15:12:52.601Z
Observed reasoning tokens
100 reporting calls; mean 0; max 0
Evaluation roles
Candidate
Returned providers
{"Alibaba":100}
Returned model IDs
{"qwen/qwen3.8-27b":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "qwen3-8-27b",
  "label": "Qwen 3.8 27B",
  "transport": "openrouter",
  "modelSlug": "qwen/qwen3.8-27b",
  "reasoning": {
    "enabled": false
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 0.42,
    "outputPricePerM": 2.55
  },
  "providerPolicy": {
    "order": [
      "alibaba"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T15:12:52.601Z",
    "providerSlug": "alibaba",
    "providerName": "Alibaba",
    "tag": "alibaba",
    "quantization": "unknown",
    "contextLength": 1000000,
    "pricing": {
      "prompt": 0.425,
      "completion": 2.5500000000000003
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "max_tokens",
      "temperature",
      "top_p",
      "seed",
      "presence_penalty",
      "response_format",
      "top_k",
      "frequency_penalty",
      "stop",
      "tools",
      "tool_choice",
      "logprobs",
      "top_logprobs",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": false,
      "default_enabled": true,
      "supported_efforts": [
        "xhigh",
        "medium",
        "low"
      ],
      "default_effort": "xhigh"
    }
  },
  "temperatureHonoured": true,
  "responseFormatMode": "json_object",
  "maxTokensSent": 1536,
  "requestParams": {
    "model": "qwen/qwen3.8-27b",
    "max_tokens": 1536,
    "temperature": 0.2,
    "response_format": "{ type: \"json_object\" } — endpoint lacks structured_outputs",
    "reasoning": {
      "enabled": false
    },
    "provider": {
      "order": [
        "alibaba"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T13:33:25.867Z",
  "finishedAt": "2026-09-15T15:13:40.946Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Alibaba": 100
  },
  "servedModelIds": {
    "qwen/qwen3.8-27b": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 0,
    "maxWhenReported": 0
  }
}
Qwen 3.8 MaxAlibaba · reasoning minimal · 4,096 tokens
Exact requested ID
qwen/qwen3.8-max-0902
Pinned provider slug
alibaba
Format / requested temperature
json_schema / 0.2 · honoured flag: yes
Quantization / snapshot UTC
unknown / 2026-09-15T15:13:41.876Z
Observed reasoning tokens
100 reporting calls; mean 872.09; max 1474
Evaluation roles
Candidate
Returned providers
{"Alibaba":100}
Returned model IDs
{"qwen/qwen3.8-max-0902":100}
Link to this configuration

Full request, endpoint snapshot and parameter list

Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.

{
  "modelKey": "qwen3-8-max",
  "label": "Qwen 3.8 Max (0902)",
  "transport": "openrouter",
  "modelSlug": "qwen/qwen3.8-max-0902",
  "reasoning": {
    "enabled": true,
    "effort": "minimal"
  },
  "promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
  "invariants": {
    "temperature": 0.2,
    "maxTokens": 1536,
    "reasoning": "off",
    "responseFormat": "json_schema:strict",
    "articlesOrder": "most-recent-first"
  },
  "pricing": {
    "inputPricePerM": 2,
    "outputPricePerM": 6
  },
  "providerPolicy": {
    "order": [
      "alibaba"
    ],
    "allow_fallbacks": false
  },
  "endpoint": {
    "fetchedAt": "2026-09-15T15:13:41.876Z",
    "providerSlug": "alibaba",
    "providerName": "Alibaba",
    "tag": "alibaba",
    "quantization": "unknown",
    "contextLength": 1000000,
    "pricing": {
      "prompt": 2,
      "completion": 6
    },
    "supportedParameters": [
      "reasoning",
      "include_reasoning",
      "max_tokens",
      "temperature",
      "top_p",
      "seed",
      "presence_penalty",
      "response_format",
      "tools",
      "tool_choice",
      "structured_outputs",
      "logprobs",
      "top_logprobs",
      "top_k",
      "frequency_penalty",
      "stop",
      "reasoning_effort"
    ],
    "status": 0,
    "modelReasoning": {
      "mandatory": true,
      "default_enabled": true,
      "supported_efforts": [
        "xhigh",
        "high",
        "medium",
        "low",
        "minimal"
      ],
      "default_effort": "xhigh"
    }
  },
  "temperatureHonoured": true,
  "responseFormatMode": "json_schema",
  "maxTokensSent": 4096,
  "requestParams": {
    "model": "qwen/qwen3.8-max-0902",
    "max_tokens": 4096,
    "temperature": 0.2,
    "response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
    "reasoning": {
      "effort": "minimal"
    },
    "provider": {
      "order": [
        "alibaba"
      ],
      "allow_fallbacks": false
    }
  },
  "startedAt": "2026-09-15T14:18:45.741Z",
  "finishedAt": "2026-09-15T15:15:25.406Z",
  "goldContributor": false,
  "fullPassJudge": false,
  "quantizationEvidence": "endpoint snapshot; unknown is not full precision",
  "servedProviders": {
    "Alibaba": 100
  },
  "servedModelIds": {
    "qwen/qwen3.8-max-0902": 100
  },
  "observedReasoningTokens": {
    "reportedCalls": 100,
    "meanWhenReported": 872.09,
    "maxWhenReported": 1474
  }
}

Gold construction and judging use different settings

Candidate Opus has reasoning off; judge Opus uses medium effort. Gold extraction uses high effort, 16,000 tokens; adjudication uses high effort, 20,000 tokens. Full judges and calibration use medium effort, 12,000 tokens.

StageExact model IDReasoningToken capRecorded provider
extractiongold/extract/claude-fable-5-1.jsonanthropic/claude-fable-5.1high16000Anthropic
extractiongold/extract/gemini-3-8-flash.jsongoogle/gemini-3.8-flashhigh16000Google AI Studio
extractiongold/extract/gpt-6-astra.jsonopenai/gpt-6-astrahigh16000OpenAI
adjudicationgold/gold.v2.jsonopenai/gpt-6-astrahigh20,000 (archived code)OpenAI
full-judgegrades/claude-opus-5.jsonanthropic/claude-opus-5medium12000Anthropic
full-judgegrades/gemini-3-8-flash.jsongoogle/gemini-3.8-flashmedium12000Google AI Studio
calibrationgrades/calibration/claude-opus-5-r1.jsonanthropic/claude-opus-5medium12000Anthropic
calibrationgrades/calibration/claude-opus-5-r2.jsonanthropic/claude-opus-5medium12000Anthropic
calibrationgrades/calibration/gemini-3-8-flash-r1.jsongoogle/gemini-3.8-flashmedium12000Google AI Studio
calibrationgrades/calibration/gemini-3-8-flash-r2.jsongoogle/gemini-3.8-flashmedium12000Google AI Studio
calibrationgrades/calibration/gpt-6-astra-r1.jsonopenai/gpt-6-astramedium12000OpenAI
Full stage metadata
[
  {
    "stageType": "extraction",
    "artifact": "gold/extract/claude-fable-5-1.json",
    "benchmark": "NepNewsBench",
    "version": "2.0-draft",
    "stage": "gold-extract",
    "extractorKey": "claude-fable-5-1",
    "label": "Claude Fable 5.1",
    "contestant": true,
    "modelSlug": "anthropic/claude-fable-5.1",
    "endpoint": {
      "fetchedAt": "2026-09-15T22:38:37.037Z",
      "providerSlug": "anthropic",
      "providerName": "Anthropic",
      "tag": "anthropic",
      "quantization": "unknown",
      "contextLength": 1000000,
      "pricing": {
        "prompt": 10,
        "completion": 50
      },
      "supportedParameters": [
        "max_tokens",
        "stop",
        "reasoning",
        "include_reasoning",
        "tools",
        "structured_outputs",
        "response_format",
        "verbosity",
        "reasoning_effort"
      ],
      "status": 0,
      "modelReasoning": {
        "mandatory": true,
        "supported_efforts": [
          "max",
          "xhigh",
          "high",
          "medium",
          "low"
        ],
        "default_effort": "high"
      }
    },
    "reasoning": {
      "enabled": true,
      "effort": "high"
    },
    "maxTokens": 16000,
    "promptSha256": {
      "extractSystem": "5766dcc879f29bd88462078d431cde293ef3331f36146beb114e2cab54c68167",
      "extractSchema": "6b28cc3ec5a9990d43eedf8be692ed8b48130ec2da402bd9daf3613d5790fac8",
      "adjudicateSystem": "52c64292fff0a47c39f999dfd49a4a3b864f36268a46b31f86e68ca9281a6292",
      "adjudicateSchema": "8c4043c596c4a3c541f51da60ecff85b3c59b3d2f92bcef91b60633a38ef491b"
    },
    "startedAt": "2026-09-15T22:38:37.038Z",
    "finishedAt": "2026-09-16T02:13:01.244Z"
  },
  {
    "stageType": "extraction",
    "artifact": "gold/extract/gemini-3-8-flash.json",
    "benchmark": "NepNewsBench",
    "version": "2.0-draft",
    "stage": "gold-extract",
    "extractorKey": "gemini-3-8-flash",
    "label": "Gemini 3.8 Flash",
    "contestant": true,
    "modelSlug": "google/gemini-3.8-flash",
    "endpoint": {
      "fetchedAt": "2026-09-15T22:54:24.720Z",
      "providerSlug": "google-ai-studio",
      "providerName": "Google AI Studio",
      "tag": "google-ai-studio",
      "quantization": "unknown",
      "contextLength": 1048576,
      "pricing": {
        "prompt": 0.75,
        "completion": 3.75
      },
      "supportedParameters": [
        "reasoning",
        "include_reasoning",
        "max_tokens",
        "temperature",
        "top_p",
        "seed",
        "response_format",
        "structured_outputs",
        "tool_choice",
        "tools",
        "reasoning_effort"
      ],
      "status": 0,
      "modelReasoning": {
        "mandatory": true,
        "default_enabled": true,
        "supported_efforts": [
          "high",
          "medium",
          "low"
        ],
        "default_effort": "medium"
      }
    },
    "reasoning": {
      "enabled": true,
      "effort": "high"
    },
    "maxTokens": 16000,
    "promptSha256": {
      "extractSystem": "5766dcc879f29bd88462078d431cde293ef3331f36146beb114e2cab54c68167",
      "extractSchema": "6b28cc3ec5a9990d43eedf8be692ed8b48130ec2da402bd9daf3613d5790fac8",
      "adjudicateSystem": "52c64292fff0a47c39f999dfd49a4a3b864f36268a46b31f86e68ca9281a6292",
      "adjudicateSchema": "8c4043c596c4a3c541f51da60ecff85b3c59b3d2f92bcef91b60633a38ef491b"
    },
    "startedAt": "2026-09-15T22:54:24.721Z",
    "finishedAt": "2026-09-15T23:16:17.333Z"
  },
  {
    "stageType": "extraction",
    "artifact": "gold/extract/gpt-6-astra.json",
    "benchmark": "NepNewsBench",
    "version": "2.0-draft",
    "stage": "gold-extract",
    "extractorKey": "gpt-6-astra",
    "label": "GPT-6 Astra",
    "contestant": false,
    "modelSlug": "openai/gpt-6-astra",
    "endpoint": {
      "fetchedAt": "2026-09-15T22:40:05.050Z",
      "providerSlug": "openai",
      "providerName": "OpenAI",
      "tag": "openai",
      "quantization": "unknown",
      "contextLength": 1050000,
      "pricing": {
        "prompt": 10,
        "completion": 50
      },
      "supportedParameters": [
        "reasoning",
        "include_reasoning",
        "seed",
        "max_tokens",
        "response_format",
        "structured_outputs",
        "tools",
        "tool_choice",
        "reasoning_effort"
      ],
      "status": 0,
      "modelReasoning": {
        "mandatory": true,
        "default_enabled": true,
        "supported_efforts": [
          "max",
          "xhigh",
          "high",
          "medium",
          "low"
        ],
        "default_effort": "medium"
      }
    },
    "reasoning": {
      "enabled": true,
      "effort": "high"
    },
    "maxTokens": 16000,
    "promptSha256": {
      "extractSystem": "5766dcc879f29bd88462078d431cde293ef3331f36146beb114e2cab54c68167",
      "extractSchema": "6b28cc3ec5a9990d43eedf8be692ed8b48130ec2da402bd9daf3613d5790fac8",
      "adjudicateSystem": "52c64292fff0a47c39f999dfd49a4a3b864f36268a46b31f86e68ca9281a6292",
      "adjudicateSchema": "8c4043c596c4a3c541f51da60ecff85b3c59b3d2f92bcef91b60633a38ef491b"
    },
    "startedAt": "2026-09-15T22:40:05.052Z",
    "finishedAt": "2026-09-16T02:38:03.490Z"
  },
  {
    "stageType": "adjudication",
    "artifact": "gold/gold.v2.json",
    "benchmark": "NepNewsBench",
    "version": "2.0-draft",
    "stage": "gold",
    "adjudicatorKey": "gpt-6-astra",
    "modelSlug": "openai/gpt-6-astra",
    "endpoint": {
      "fetchedAt": "2026-09-15T22:55:53.317Z",
      "providerSlug": "openai",
      "providerName": "OpenAI",
      "tag": "openai",
      "quantization": "unknown",
      "contextLength": 1050000,
      "pricing": {
        "prompt": 10,
        "completion": 50
      },
      "supportedParameters": [
        "reasoning",
        "include_reasoning",
        "seed",
        "max_tokens",
        "response_format",
        "structured_outputs",
        "tools",
        "tool_choice",
        "reasoning_effort"
      ],
      "status": 0,
      "modelReasoning": {
        "mandatory": true,
        "default_enabled": true,
        "supported_efforts": [
          "max",
          "xhigh",
          "high",
          "medium",
          "low"
        ],
        "default_effort": "medium"
      }
    },
    "reasoning": {
      "enabled": true,
      "effort": "high"
    },
    "promptSha256": {
      "extractSystem": "5766dcc879f29bd88462078d431cde293ef3331f36146beb114e2cab54c68167",
      "extractSchema": "6b28cc3ec5a9990d43eedf8be692ed8b48130ec2da402bd9daf3613d5790fac8",
      "adjudicateSystem": "52c64292fff0a47c39f999dfd49a4a3b864f36268a46b31f86e68ca9281a6292",
      "adjudicateSchema": "8c4043c596c4a3c541f51da60ecff85b3c59b3d2f92bcef91b60633a38ef491b"
    },
    "extractors": [
      "claude-fable-5-1",
      "gpt-6-astra",
      "gemini-3-8-flash"
    ],
    "startedAt": "2026-09-15T22:55:53.318Z",
    "finishedAt": "2026-09-16T04:29:41.267Z"
  },
  {
    "stageType": "full-judge",
    "artifact": "grades/claude-opus-5.json",
    "benchmark": "NepNewsBench",
    "version": "2.0-draft",
    "stage": "grades",
    "judgeKey": "claude-opus-5",
    "label": "Claude Opus 5",
    "contestant": true,
    "modelSlug": "anthropic/claude-opus-5",
    "endpoint": {
      "fetchedAt": "2026-09-16T13:07:14.825Z",
      "providerSlug": "anthropic",
      "providerName": "Anthropic",
      "tag": "anthropic",
      "quantization": "unknown",
      "contextLength": 1000000,
      "pricing": {
        "prompt": 5,
        "completion": 25
      },
      "supportedParameters": [
        "max_tokens",
        "stop",
        "reasoning",
        "include_reasoning",
        "tool_choice",
        "tools",
        "structured_outputs",
        "response_format",
        "verbosity",
        "reasoning_effort"
      ],
      "status": 0,
      "modelReasoning": {
        "mandatory": false,
        "default_enabled": true,
        "supported_efforts": [
          "max",
          "xhigh",
          "high",
          "medium",
          "low"
        ],
        "default_effort": "high"
      }
    },
    "reasoning": {
      "enabled": true,
      "effort": "medium"
    },
    "maxTokens": 12000,
    "promptSha256": {
      "system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
      "schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
    },
    "goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
    "mode": "full",
    "startedAt": "2026-09-16T13:07:14.834Z",
    "finishedAt": "2026-09-16T14:16:43.782Z"
  },
  {
    "stageType": "full-judge",
    "artifact": "grades/gemini-3-8-flash.json",
    "benchmark": "NepNewsBench",
    "version": "2.0-draft",
    "stage": "grades",
    "judgeKey": "gemini-3-8-flash",
    "label": "Gemini 3.8 Flash",
    "contestant": true,
    "modelSlug": "google/gemini-3.8-flash",
    "endpoint": {
      "fetchedAt": "2026-09-16T13:07:17.245Z",
      "providerSlug": "google-ai-studio",
      "providerName": "Google AI Studio",
      "tag": "google-ai-studio",
      "quantization": "unknown",
      "contextLength": 1048576,
      "pricing": {
        "prompt": 0.75,
        "completion": 3.75
      },
      "supportedParameters": [
        "reasoning",
        "include_reasoning",
        "max_tokens",
        "temperature",
        "top_p",
        "seed",
        "response_format",
        "structured_outputs",
        "tool_choice",
        "tools",
        "reasoning_effort"
      ],
      "status": 0,
      "modelReasoning": {
        "mandatory": true,
        "default_enabled": true,
        "supported_efforts": [
          "high",
          "medium",
          "low"
        ],
        "default_effort": "medium"
      }
    },
    "reasoning": {
      "enabled": true,
      "effort": "medium"
    },
    "maxTokens": 12000,
    "promptSha256": {
      "system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
      "schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
    },
    "goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
    "mode": "full",
    "startedAt": "2026-09-16T13:07:17.246Z",
    "finishedAt": "2026-09-16T13:47:41.140Z"
  },
  {
    "stageType": "calibration",
    "artifact": "grades/calibration/claude-opus-5-r1.json",
    "benchmark": "NepNewsBench",
    "version": "2.0-draft",
    "stage": "grades",
    "judgeKey": "claude-opus-5",
    "label": "Claude Opus 5",
    "contestant": true,
    "modelSlug": "anthropic/claude-opus-5",
    "endpoint": {
      "fetchedAt": "2026-09-16T04:30:11.903Z",
      "providerSlug": "anthropic",
      "providerName": "Anthropic",
      "tag": "anthropic",
      "quantization": "unknown",
      "contextLength": 1000000,
      "pricing": {
        "prompt": 5,
        "completion": 25
      },
      "supportedParameters": [
        "max_tokens",
        "stop",
        "reasoning",
        "include_reasoning",
        "tool_choice",
        "tools",
        "structured_outputs",
        "response_format",
        "verbosity",
        "reasoning_effort"
      ],
      "status": 0,
      "modelReasoning": {
        "mandatory": false,
        "default_enabled": true,
        "supported_efforts": [
          "max",
          "xhigh",
          "high",
          "medium",
          "low"
        ],
        "default_effort": "high"
      }
    },
    "reasoning": {
      "enabled": true,
      "effort": "medium"
    },
    "maxTokens": 12000,
    "promptSha256": {
      "system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
      "schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
    },
    "goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
    "mode": "calibration",
    "sample": {
      "n": 20,
      "seed": 42,
      "questionIds": [
        1,
        13,
        20,
        21,
        29,
        36,
        37,
        39,
        41,
        44,
        45,
        53,
        54,
        64,
        66,
        79,
        84,
        86,
        88,
        94
      ],
      "run": 1
    },
    "startedAt": "2026-09-16T04:30:11.904Z",
    "finishedAt": "2026-09-16T04:50:54.612Z"
  },
  {
    "stageType": "calibration",
    "artifact": "grades/calibration/claude-opus-5-r2.json",
    "benchmark": "NepNewsBench",
    "version": "2.0-draft",
    "stage": "grades",
    "judgeKey": "claude-opus-5",
    "label": "Claude Opus 5",
    "contestant": true,
    "modelSlug": "anthropic/claude-opus-5",
    "endpoint": {
      "fetchedAt": "2026-09-16T04:50:55.801Z",
      "providerSlug": "anthropic",
      "providerName": "Anthropic",
      "tag": "anthropic",
      "quantization": "unknown",
      "contextLength": 1000000,
      "pricing": {
        "prompt": 5,
        "completion": 25
      },
      "supportedParameters": [
        "max_tokens",
        "stop",
        "reasoning",
        "include_reasoning",
        "tool_choice",
        "tools",
        "structured_outputs",
        "response_format",
        "verbosity",
        "reasoning_effort"
      ],
      "status": 0,
      "modelReasoning": {
        "mandatory": false,
        "default_enabled": true,
        "supported_efforts": [
          "max",
          "xhigh",
          "high",
          "medium",
          "low"
        ],
        "default_effort": "high"
      }
    },
    "reasoning": {
      "enabled": true,
      "effort": "medium"
    },
    "maxTokens": 12000,
    "promptSha256": {
      "system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
      "schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
    },
    "goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
    "mode": "calibration",
    "sample": {
      "n": 20,
      "seed": 42,
      "questionIds": [
        1,
        13,
        20,
        21,
        29,
        36,
        37,
        39,
        41,
        44,
        45,
        53,
        54,
        64,
        66,
        79,
        84,
        86,
        88,
        94
      ],
      "run": 2
    },
    "startedAt": "2026-09-16T04:50:55.801Z",
    "finishedAt": "2026-09-16T13:10:51.532Z"
  },
  {
    "stageType": "calibration",
    "artifact": "grades/calibration/gemini-3-8-flash-r1.json",
    "benchmark": "NepNewsBench",
    "version": "2.0-draft",
    "stage": "grades",
    "judgeKey": "gemini-3-8-flash",
    "label": "Gemini 3.8 Flash",
    "contestant": true,
    "modelSlug": "google/gemini-3.8-flash",
    "endpoint": {
      "fetchedAt": "2026-09-16T04:30:08.060Z",
      "providerSlug": "google-ai-studio",
      "providerName": "Google AI Studio",
      "tag": "google-ai-studio",
      "quantization": "unknown",
      "contextLength": 1048576,
      "pricing": {
        "prompt": 0.75,
        "completion": 3.75
      },
      "supportedParameters": [
        "reasoning",
        "include_reasoning",
        "max_tokens",
        "temperature",
        "top_p",
        "seed",
        "response_format",
        "structured_outputs",
        "tool_choice",
        "tools",
        "reasoning_effort"
      ],
      "status": 0,
      "modelReasoning": {
        "mandatory": true,
        "default_enabled": true,
        "supported_efforts": [
          "high",
          "medium",
          "low"
        ],
        "default_effort": "medium"
      }
    },
    "reasoning": {
      "enabled": true,
      "effort": "medium"
    },
    "maxTokens": 12000,
    "promptSha256": {
      "system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
      "schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
    },
    "goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
    "mode": "calibration",
    "sample": {
      "n": 20,
      "seed": 42,
      "questionIds": [
        1,
        13,
        20,
        21,
        29,
        36,
        37,
        39,
        41,
        44,
        45,
        53,
        54,
        64,
        66,
        79,
        84,
        86,
        88,
        94
      ],
      "run": 1
    },
    "startedAt": "2026-09-16T04:30:08.061Z",
    "finishedAt": "2026-09-16T04:40:14.994Z"
  },
  {
    "stageType": "calibration",
    "artifact": "grades/calibration/gemini-3-8-flash-r2.json",
    "benchmark": "NepNewsBench",
    "version": "2.0-draft",
    "stage": "grades",
    "judgeKey": "gemini-3-8-flash",
    "label": "Gemini 3.8 Flash",
    "contestant": true,
    "modelSlug": "google/gemini-3.8-flash",
    "endpoint": {
      "fetchedAt": "2026-09-16T04:40:15.680Z",
      "providerSlug": "google-ai-studio",
      "providerName": "Google AI Studio",
      "tag": "google-ai-studio",
      "quantization": "unknown",
      "contextLength": 1048576,
      "pricing": {
        "prompt": 0.75,
        "completion": 3.75
      },
      "supportedParameters": [
        "reasoning",
        "include_reasoning",
        "max_tokens",
        "temperature",
        "top_p",
        "seed",
        "response_format",
        "structured_outputs",
        "tool_choice",
        "tools",
        "reasoning_effort"
      ],
      "status": 0,
      "modelReasoning": {
        "mandatory": true,
        "default_enabled": true,
        "supported_efforts": [
          "high",
          "medium",
          "low"
        ],
        "default_effort": "medium"
      }
    },
    "reasoning": {
      "enabled": true,
      "effort": "medium"
    },
    "maxTokens": 12000,
    "promptSha256": {
      "system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
      "schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
    },
    "goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
    "mode": "calibration",
    "sample": {
      "n": 20,
      "seed": 42,
      "questionIds": [
        1,
        13,
        20,
        21,
        29,
        36,
        37,
        39,
        41,
        44,
        45,
        53,
        54,
        64,
        66,
        79,
        84,
        86,
        88,
        94
      ],
      "run": 2
    },
    "startedAt": "2026-09-16T04:40:15.681Z",
    "finishedAt": "2026-09-16T04:50:27.431Z"
  },
  {
    "stageType": "calibration",
    "artifact": "grades/calibration/gpt-6-astra-r1.json",
    "benchmark": "NepNewsBench",
    "version": "2.0-draft",
    "stage": "grades",
    "judgeKey": "gpt-6-astra",
    "label": "GPT-6 Astra",
    "contestant": false,
    "modelSlug": "openai/gpt-6-astra",
    "endpoint": {
      "fetchedAt": "2026-09-16T04:30:14.051Z",
      "providerSlug": "openai",
      "providerName": "OpenAI",
      "tag": "openai",
      "quantization": "unknown",
      "contextLength": 1050000,
      "pricing": {
        "prompt": 10,
        "completion": 50
      },
      "supportedParameters": [
        "reasoning",
        "include_reasoning",
        "seed",
        "max_tokens",
        "response_format",
        "structured_outputs",
        "tools",
        "tool_choice",
        "reasoning_effort"
      ],
      "status": 0,
      "modelReasoning": {
        "mandatory": true,
        "default_enabled": true,
        "supported_efforts": [
          "max",
          "xhigh",
          "high",
          "medium",
          "low"
        ],
        "default_effort": "medium"
      }
    },
    "reasoning": {
      "enabled": true,
      "effort": "medium"
    },
    "maxTokens": 12000,
    "promptSha256": {
      "system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
      "schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
    },
    "goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
    "mode": "calibration",
    "sample": {
      "n": 20,
      "seed": 42,
      "questionIds": [
        1,
        13,
        20,
        21,
        29,
        36,
        37,
        39,
        41,
        44,
        45,
        53,
        54,
        64,
        66,
        79,
        84,
        86,
        88,
        94
      ],
      "run": 1
    },
    "startedAt": "2026-09-16T04:30:14.052Z",
    "finishedAt": "2026-09-16T04:52:13.159Z"
  }
]

Gold, judges and score construction

Gold is a fact sheet, not an ideal reference paragraph. Fable 5.1, GPT-6 Astra and Gemini 3.8 Flash independently extract source-linked facts at high reasoning effort. GPT-6 Astra adjudicates shuffled, anonymously labelled sheets at high effort. The recorded token caps are 16,000 for extraction and 20,000 for adjudication. The adjudicator is also one of the extractors. Two extractors are contestants. That overlap is disclosed and cannot be removed merely by anonymizing labels.

The fact sheets include salience, kind, support-row indices, extractor agreement, entities and known traps. There are 1,169 facts labelled as present in all three extraction sheets, 308 with two-way agreement and 170 singleton facts. The agreement labels are adjudicator-produced provenance, not a substitute for human verification.

Opus 5 and Gemini 3.8 Flash grade each valid candidate at medium effort with a 12,000-token cap. They see the source rows, gold and one candidate with its identity hidden. Both judges are themselves contestants. GPT-6 Astra supplies an additional judge view on a 20-question calibration subset. Repeated calibration grades reassess the same saved answers; they do not repeat candidate generation.

Coverage grades each gold fact as stated (1), partial (0.5), absent (0) or contradicted (0), considering the English and Nepali summaries together. Core recall averages credits over core facts. The “precision” component is credited gold coverage across all salience levels divided by credited coverage plus the count of unsupported-claim annotations. It is a coverage-based proxy, not conventional precision over an independently enumerated set of generated claims. Supporting recall is separately reported. Nepali and English prose are scored from 1 to 10; headline acceptance and approximate entity matching are additional diagnostics.

The composite is 40 × core recall + 30 × precision proxy + 1.5 × Nepali prose + 1.5 × English prose. Judges are averaged per answer. An invalid generation contributes zero to the effective score; valid-answer quality excludes it. One Qwen 27B answer lacks the Gemini grade, so its available Opus grade determines that cell. The denominator is visible rather than silently filled.

The publication analysis uses 5,000 paired question-bootstrap resamples for primary comparisons. New rank-frequency analysis uses 2,000 resamples. These are exploratory, unadjusted intervals conditional on the saved runs and judges. They describe variation across sampled questions; they do not isolate repeated-generation variance or account fully for related stories. The analysis has no established equivalence margin, so an interval crossing zero is not evidence that two systems are equivalent.

Score = 40Rcore + 30Pproxy + 1.5QNE + 1.5QEN
Terminology and metric definitions
Story question / cluster
A saved group of publisher rows about one underlying story. Related questions can still share events.
Gold fact sheet
A machine-generated set of atomic facts with salience, kind and support-row indices. It is a reference for source fidelity, not verified real-world truth.1
Core / supporting / detail
Facts essential to a faithful summary / usually useful / optional detail, as assigned by the gold process.
Coverage credit
Stated = 1; partial = 0.5; absent or contradicted = 0. A contradiction also remains visible as its own diagnostic.
Core recall
Average coverage credit on core facts; ranges from 0 to 1. Coverage is assessed jointly across the two summaries.
Precision proxy
Credited gold coverage across all salience levels ÷ (credited coverage + unsupported-claim annotation count). This is not the fraction of generated statements that are true.
Effective / valid-answer quality
Effective averages all 100 slots, assigning generation failures zero. Quality averages only valid answers using available judge grades.
Nepali prose
Judge score, 1–10, covering the Nepali headline and summary, naturalness, register, grammar and consistency with English. It is not Nepali-only fact recall.
Unsupported annotation
A judge’s flagged addition not supported by a supplied source row. Its segmentation and support threshold depend on the judge.
Entity F1
Per-answer harmonic mean of surface/alias- and type-matched entity precision and recall, then averaged. Gold entities may exceed those needed for a concise English summary.
Paired bootstrap
Resample the same question IDs for both models and recompute their mean difference. Intervals describe question variation conditional on saved answers and judges.
P50 / P95 latency
The median / 95th percentile of stored call durations. These do not reconstruct complete retry, queue or user-perceived latency.
Pareto frontier
Configurations for which no observed alternative is both cheaper and at least as good, or better and no more expensive, on the chosen point estimates.
Language anchors, headline checks and deterministic heuristics

Prose anchors: 9–10 clean wire copy; 7–8 clean with minor awkwardness; 5–6 readable but with clear faults or one cross-language inconsistency; 3–4 hard to read or several faults; 1–2 unusable. Judges are instructed not to reward length.

The actual headline rubric checks accuracy, the main event and clickbait; it does not enforce the planned 12-word cap. Correct Gregorian conversions may receive coverage credit; wrong conversions can be contradicted. If source rows disagree with gold, judges are instructed to prefer the rows and explain the conflict.

Script ratios, digit matching, normalized entity surfaces and cross-language number comparisons are deterministic diagnostics. They do not perform semantic calendar conversion or prove cross-language factual equivalence. The parser requires headline_en, summary_en and summary_ne; it normalizes the other requested fields. A valid answer therefore does not prove perfect six-field schema compliance.

Different criteria select different leaders

Fable has the highest observed failure-inclusive composite at 92.65. GPT-5.6 Sol follows at 91.61, Opus at 91.26 and DeepSeek Flash at 91.03. Fable's gain over Sol is 1.05 points, with a paired interval of −0.07 to 2.24. This comparison remains uncertain. By contrast, Fable's 4.71-point lead over Gemma has an interval of 2.72–7.32.

This is why “the top 15 are all the same” is an unsuitable interpretation. A chain of adjacent uncertain comparisons is not an equivalence test across a group. The full pairwise matrix supports more specific comparisons. In the new conditional question resampling, Fable occupies first place in 94.55% of 2,000 resamples; this is a sampling frequency under the frozen dataset, not a probability of universal superiority or a contradiction of the FableSol interval.

Gemini 3.1 Pro has the highest observed Nepali prose mean, 9.025/10, while Fable leads the composite through its broader coverage profile. Fable's core recall is 95.55%, supporting recall 79.79% and entity recall 74.81%. Its observed cost is $8.896 per 100 retained attempts. Gemma records 93.14% precision proxy but only 47.60% supporting recall, demonstrating how a concise answer can avoid additions while omitting substantial source material.

Post-hoc weighting scenarios retain Fable first under equal components, a grounding-heavy score and a Nepali-emphasis score. A language-only score puts Sol first. These are illustrations of preference sensitivity, not alternative benchmark winners chosen after viewing results. The original score remains the primary outcome.

The smaller DeepSeek configuration wins this workload comparison

DeepSeek V4.1 Flash scores 91.03 versus 89.10 for V4 Pro 0813. The paired gain is 1.92 points, with an exploratory interval of 0.42–3.41. Its recorded generation cost is $0.0699 versus $0.2419 per 100 attempts, approximately 71.1% lower. Both runs use first-party DeepSeek through OpenRouter with reasoning disabled and JSON-object mode.

The experiment journal also records an unresolved possibility of upstream aliasing between the two DeepSeek endpoints. Their returned model IDs, outputs, prices and latency differ, but those records cannot establish that the underlying weights are different. This is a direct practical result for these two endpoints on this task. It does not identify model parameter count, establish that inexpensive models are generally better or predict performance on unrelated tasks. The observed cost-quality frontier contains Gemma, DeepSeek Flash, Sol and Fable. Membership is based on point estimates; nearby alternatives can move under new questions, different weights, provider pricing or repeated runs.

Nepali competence has several distinct meanings here

The Nepali prose score combines headline and summary language quality, including translation consistency. It is not a standalone Nepali factual-accuracy score. Coverage is graded jointly over both summaries. A fact appearing only in English can therefore receive coverage credit even if Nepali readers do not receive it. A future Nepali-only evaluation must grade fact coverage separately for each language.

The 77 Nepali-source questions assess consolidation from Nepali evidence with bilingual output. The six English-only questions add English-to-Nepali translation pressure. They are different stories, not matched translations of the same inputs, so their difference does not identify a causal language effect.

Tail rates are informative. Fable and Gemini 3.1 Pro have no valid outputs below 7/10 for averaged Nepali quality. DeepSeek Flash has two, Gemini 3 Flash one, Gemma four, Haiku 20 and Qwen 27B 28. These are observed counts, not precision estimates for all future news. Opus also has no low-scoring valid Nepali outputs but one generation failure, which must remain visible next to that statement.

Nepali-only unsupported flags and flags shared with English are reported separately. They are judge annotations. Different judges may segment or classify the same flawed sentence differently. A count of annotations is not a count of unique real-world falsehoods, and an answer flagged by both judges need not contain the same flagged claim under each judge.

All source languages; averaged judges; valid-answer prose only. Failures shown separately.
ModelNE mean /10Below 7 / validAt least 9 / validGeneration failures /100NE-only flags, mean
Gemini 3.1 Pro9.0250 / 10066 / 10000.035
Claude Fable 5.18.9950 / 10071 / 10000.055
GPT-5.6 Sol8.9504 / 10072 / 10000.070
Claude Opus 58.8790 / 9962 / 9910.066
Kimi K38.8622 / 9864 / 9820.071
DeepSeek V4.1 Flash8.8352 / 10064 / 10000.095
DeepSeek V4 Pro8.8251 / 10056 / 10000.035
Gemini 3.8 Flash8.8151 / 10053 / 10000.165
Grok 4.68.7553 / 10059 / 10000.045
Qwen 3.8 Max8.7223 / 9951 / 9910.076
Gemini 3 Flash8.6801 / 10040 / 10000.210
Gemma 4 31B8.6724 / 9948 / 9910.030
GPT-5.6 Luna8.6352 / 10050 / 10000.080
GLM 5.38.6254 / 10052 / 10000.090
Claude Sonnet 58.5563 / 9942 / 9910.136
Claude Haiku 4.57.57020 / 1007 / 10000.150
Qwen 3.8 27B7.52028 / 9916 / 9910.293

The mistakes behind the averages

These purposefully selected examples illustrate numeric, calendar, attribution and serving failures. They are source-grounding inspections and machine annotations, not a random error sample or a human validation panel.

Q72

A tenfold magnitude error

The source amount is 2.2 billion yuan. DeepSeek Flash renders it as 220 million in English and 22 crore in Nepali.

Q6

Nepali figures survive; English units do not

DeepSeek preserves the printed Nepali budget amounts, then translates kharba incorrectly into trillions in English. A bilingual average can conceal which readers receive the error.

Q29

A month changes between languages

DeepSeek’s Nepali month conflicts with the supplied source and its English answer. This is a language-specific error despite fluent wording.

Q7

Who owes whom an apology?

DeepSeek reverses the direction of an apology demand in Nepali and changes Gen Z into Janajati in English. Natural prose does not establish faithful attribution.

Q85

Preserving a date is better than guessing its conversion

Gemini 3 Flash adds an unsupported Gregorian date and misassigns the presenter credit. The source provides a Bikram Sambat date; calendar conversion introduces an avoidable failure mode.

Q39

Sparse evidence rewards restraint

Three headline-only rows say 197 savers from ten cooperatives received refunds. Gemini 3 Flash adds context about troubled cooperatives and government action that these rows do not supply.

Q69

Keep interpretation attributed

Gemini 3 Flash recasts analysts’ interpretation of Sheikh Hasina’s remarks as her own intention or appeal. Interpretation needs to remain attributed to its source.

Q21

Dense stories can exhaust the output budget

This question has 34 gold facts. Opus 5 and Sonnet 5 truncate during generation and score zero. Rich entity lists compete with bilingual prose for a finite token budget.

Q16

Calendar conversion needs separate language review

The saved Haiku answer is a calendar-error inspection case. Separately, Qwen 27B lacks its Gemini judge grade here; a missing grade is not a failed candidate answer.

Q51

Local units and an evolving loan ceiling

Luna’s headline says 50 lakh, its English summary says 500,000 rupees, and its Nepali summary says 50 thousand. The supplied rows include changing reported limits. This is a priority for native-Nepali review, not a human-confirmed error label.

Q88

English evidence, Nepali output

An English-only landslide story exposes translation quality. It belongs to a six-question slice, too small to support a general language-effect claim.

Q4

Different serving failures on one question

Kimi records a client/content-filter error; the two Qwen configurations record server errors. DeepSeek and GLM answer. These events do not support a nationality-wide refusal claim.

The public numerical release includes per-question scores and annotation counts. Full source excerpts, model outputs and judge comments remain in the private audit archive while distribution terms are reviewed.

Length and coverage: the direction changes with the comparison

Across the 17 model means, English summary length and supporting-fact recall correlate at r=0.966. Models that tend to write longer summaries capture a larger share of supporting facts. This is a useful description of the model configurations.

Within each of the 17 models, however, its longer answers across different questions have lower supporting-fact recall: correlations range from −0.517 to −0.095. On the 96 questions all models answer successfully, the within-model centered correlation is −0.273. Comparing models within the same question gives a positive centered association, r=0.690.

These patterns can coexist because stories differ in complexity and the number of facts available to recall. A dense story can elicit a longer summary while still leaving a larger fraction of its facts uncovered. The data do not support calling supporting recall “just a length proxy,” nor do they support prescribing longer output as a causal solution. A matched length-budget experiment would test that mechanism directly.

Across model means · n=17r = +0.966

Longer-writing configurations retain more supporting facts.

Within each model · 98–100 valid storiesr = −0.517 to −0.095

Longer answers to different stories retain a smaller fraction.

Inspect all within-model correlations
ModelValid nEN words vs supporting recall, Pearson rEN words vs core recall, Pearson r
Claude Fable 5.1100-0.470-0.462
Claude Haiku 4.5100-0.216-0.236
Claude Opus 599-0.346-0.330
Claude Sonnet 599-0.476-0.385
DeepSeek V4.1 Flash100-0.517-0.256
DeepSeek V4 Pro100-0.303-0.315
Gemini 3.1 Pro100-0.300-0.358
Gemini 3.8 Flash100-0.417-0.302
Gemini 3 Flash100-0.246-0.162
Gemma 4 31B99-0.180-0.245
GLM 5.3100-0.325-0.393
GPT-5.6 Luna100-0.253-0.410
GPT-5.6 Sol100-0.295-0.356
Grok 4.6100-0.156-0.175
Kimi K398-0.424-0.331
Qwen 3.8 27B99-0.095-0.262
Qwen 3.8 Max99-0.292-0.397

Judges agree about many facts and disagree about acceptable additions

Repeated coverage-label agreement on the calibration subset is 96.36% for Opus and 95.61% for Gemini. Newly computed unweighted Cohen's kappa is 0.939 and 0.925 respectively. Full-pass exact coverage agreement is 91.56%, with kappa 0.856. Labels are clustered within answers, and high consistency does not establish human correctness.

The judges' effective model rankings correlate at 0.946, but Nepali-quality rankings correlate less strongly at 0.807. Their score levels also differ. Opus records approximately 1.474 unsupported annotations per valid answer, versus Gemini's 0.413: a 3.57-fold difference. Gemini's average composite grade is about 5.48 points higher. The judges' thresholds for supported paraphrase and additions therefore materially affect score levels even when many fact labels agree.

Both judge views are available in the explorer. Averaging is a useful summary, not a way to erase disagreement. The calibration sample includes a non-contestant judge but is too small to establish the absence of family bias.

Opus 5 · 1,693 valid grades1.474

Unsupported-claim annotations per answer.

Gemini 3.8 Flash · 1,692 valid grades0.413

Different support thresholds and claim segmentation.

Calibration on 20 questions; repeated grading of the same saved answers.
Judge comparisonMatched answersLabel comparisonsExact agreementKappaRank Spearman ρ
claude-opus-5-r1vs claude-opus-5-r2338571496.4%0.9390.966
gemini-3-8-flash-r1vs gemini-3-8-flash-r2338571295.6%0.9250.958
claude-opus-5-r1vs gemini-3-8-flash-r1338571292.2%0.8690.941
claude-opus-5-r1vs gpt-6-astra-r1337569589.7%0.8350.860
gemini-3-8-flash-r1vs gpt-6-astra-r1337569589.4%0.8280.833

Calibration gates required ≥0.80 intra-judge label agreement and ≥0.70 inter-judge model-rank correlation; the completed records meet both. Astra has no second repeat and one missing calibration grade. Full-pass agreement uses 1,692 matched answers and 27,822 label comparisons; its valid-grade rank correlation (0.949) differs from the failure-inclusive rank correlation (0.946).

Reliability and serving failures

Seven of 1,700 generation slots fail: Opus and Sonnet truncate on Q21, Gemma truncates on Q30, Kimi has a client error on Q4 and a schema violation on Q74, and both Qwen configurations have server errors on Q4. The Q4 grouping is not evidence of a nationality-wide refusal policy. The recorded Kimi client error and Alibaba server errors are different failure classes; DeepSeek and Z.AI answer the same question.

One hundred attempts per model is too small to certify a service-level failure rate, especially for rare outages. A 100/100 observed success count is not proof of universal reliability. The explorer reports both the failure-inclusive score and valid-answer quality.

ConfigurationQuestionRecorded classRetained attempts field
Claude Opus 5Q21truncated1
Claude Sonnet 5Q21truncated1
Gemma 4 31BQ30truncated1
Kimi K3Q4client_error1
Kimi K3Q74schema_violation1
Qwen 3.8 27BQ4server_error3
Qwen 3.8 MaxQ4server_error3

Audit, uncertainty and remaining work

Independent reconstruction matches saved scores to a maximum numerical difference below 3×10⁻¹⁴. Prompt hashes match, and the full-pass judge files reference the same gold hash. Twenty Opus grade records have extra, duplicate or missing fact IDs. Sanitizing those lists changes model means by less than 0.007 points and does not alter the main conclusions. The original and sanitized analyses are both retained; the original raw files are not edited.

The protocol and plans contain stale descriptions, including a 16-model count followed by 17 names, generic strict-schema settings, and intended human review preceding the full pass. The status of the planned human review is recorded in the gold-review footnote. The package's methods ledger resolves the differences using saved run headers, implementation and completed records.

The main unresolved limitations are machine-generated gold, judge/contestant overlap, one candidate generation per slot, imperfect independence of stories, a selected rather than random corpus, unequal decoding constraints and incomplete accounting for overwritten attempts. Pairwise intervals and slice analyses are exploratory and unadjusted for multiple comparisons. The old benchmark's run-to-run noise figure should not be imported as a v2 significance threshold.

Reproduction has three meanings here: recomputing scores from frozen records is supported offline; rerunning the models requires paid services that may change; rebuilding the full production selection requires the original database context. The resources section distinguishes what the numerical download supports from what requires the private audit snapshot.

Website preparation independently ran the supplied offline reconstruction and verifier: 258 checks passed. The maximum score difference was 2.842 × 10−14. The paired interval tables are imported only after their input hashes verify; the separate archived bootstrap programs can recompute them.

Protocol amendments and implementation deviations
  • Provider pinning, returned model/provider checks and recorded usage cost replaced looser early transport handling. Earlier agent runs are excluded.
  • JSON-object mode replaces strict schema on four endpoints. Six mandatory-reasoning runs receive 4,096 total output tokens, versus 1,536 elsewhere.
  • Transport errors can be retried; provider finish_reason=error was added to that class. Kimi Q74/Q87 were accidentally resampled after schema failures before the resume policy was corrected.
  • Gemini 3.8 Flash replaced Gemini 3.1 Pro as the third gold extractor. Gemini 3 Flash was added to the candidate lineup. Final count: 17 configurations.
  • The planned human review is described in the gold-review footnote.
  • Judge truncation retries were added without changing the rubric. One Gemini full-pass grade remains missing.
  • The original scorer used 1,000 bootstrap draws; publication pairwise and focused comparisons use 5,000; post-hoc rank frequencies use 2,000.
  • Twenty Opus grades have extra, duplicate or missing fact IDs. The sanitation sensitivity enumerates gold IDs, ignores extras, uses the lowest duplicate credit and assigns missing IDs zero.
  • Gold-consensus, weighting, rank-frequency, kappa and within-model length analyses are post-hoc. Their original weights do not change.

The methods ledger supplies the complete planned-versus-completed comparison. Original protocol files use “pre-registered” internally; no public registration is established. The journal’s early “top 15 are one block” and causal reasoning/length interpretations are superseded by the paired and within-model analyses reported here.

What to do next

Complete the planned blinded gold review and add native-Nepali review of numeric, date and attribution examples. Correct gold with a new version and regrade affected questions under documented hashes. Repeat matched candidate generations to estimate generation variance. Grade fact coverage separately for Nepali and English. Expand matched evaluation to longer clusters and more mixed-language stories, keeping schema and retry policies explicit. Run matched length and reasoning ablations before making causal claims about those settings.

These follow-ups address uncertainties exposed by the experiment. They need not obscure the present findings: coverage, prose quality, cost and severe mistakes vary in identifiable, inspectable ways across the tested configurations.

Full results and experiment accounting

All 17 configurations, both full-pass judges. Effective score uses all 100 slots; other quality metrics use valid answers. Supporting recall and entity F1 are diagnostics outside the composite.

ConfigurationEffective /100Valid quality /100Core %Precision proxy %Supporting %NE /10EN /10Entity F1 %
Claude Fable 5.192.65192.65195.6%91.6%79.8%8.9958.97079.7%
GPT-5.6 Sol91.60591.60590.2%94.8%56.9%8.9509.12067.7%
Claude Opus 591.26092.18295.5%91.1%83.1%8.8798.90981.3%
DeepSeek V4.1 Flash91.02691.02692.3%92.8%69.8%8.8358.67575.5%
Grok 4.690.93490.93491.0%93.8%60.6%8.7558.83573.6%
GLM 5.390.80590.80593.0%91.9%68.2%8.6258.74072.8%
Gemini 3.1 Pro90.02890.02887.4%93.8%52.4%9.0258.93567.6%
GPT-5.6 Luna89.80289.80289.7%93.1%60.8%8.6358.70568.9%
Gemini 3.8 Flash89.50089.50089.9%90.5%58.7%8.8158.78067.7%
Kimi K389.42891.25393.1%91.5%74.0%8.8628.85275.8%
DeepSeek V4 Pro89.10389.10388.0%92.2%55.7%8.8258.67069.0%
Qwen 3.8 Max88.58389.47790.6%90.8%63.8%8.7228.60668.9%
Gemini 3 Flash88.45588.45590.0%88.5%61.1%8.6808.57069.3%
Claude Sonnet 588.41389.30792.3%89.6%74.2%8.5568.46574.5%
Gemma 4 31B87.93888.82686.8%93.1%47.6%8.6728.76364.3%
Claude Haiku 4.584.83484.83488.0%88.2%59.9%7.5707.88553.5%
Qwen 3.8 27B81.57082.39483.8%85.9%53.8%7.5207.88459.1%

Fact kinds and annotation types

Inspect what each model covers and where judges flag additions. These diagnostics sit outside the primary composite and do not measure error prevalence in future traffic.

Token usage and evaluation cost

Retained records report 55,183,774 tokens: 44,274,103 input and 10,909,671 output tokens across 7,175 calls with usage, out of 7,176 saved calls. Reasoning is a subset of reported completion tokens and is not added again. Provider tokenizers differ. Overwritten retries, smoke tests and superseded runs are excluded; missing usage is not imputed. Download the token accounting by artifact.

Candidate generation: $23.36. Gold, calibration and full grading: $354.65. Combined retained-record total: $378.01. This is not the complete account bill.

Stage / artifactSaved callsOKRecorded USDMissing cost
extractiongold/extract/claude-fable-5-1.json100100$25.51030
extractiongold/extract/gemini-3-8-flash.json100100$2.04000
extractiongold/extract/gpt-6-astra.json100100$28.14850
adjudicationgold/gold.v2.json100100$54.08620
full-judgegrades/claude-opus-5.json16931693$131.06740
full-judgegrades/gemini-3-8-flash.json16931692$19.03620
calibrationgrades/calibration/claude-opus-5-r1.json338338$26.53200
calibrationgrades/calibration/claude-opus-5-r2.json338338$26.47170
calibrationgrades/calibration/gemini-3-8-flash-r1.json338338$3.80270
calibrationgrades/calibration/gemini-3-8-flash-r2.json338338$3.84330
calibrationgrades/calibration/gpt-6-astra-r1.json338337$34.11391

Data, definitions and reproducibility

Candidate public numerical release

Explore the numbers behind the paper.

Version: publication-handoff-1 · experiment 2.0-draft. Forty-one files: numerical tables, model metadata, schemas, metric definitions and editable figure specifications. Excludes raw API bodies and source excerpts.

Download numerical data · 370 KiB

378,929 bytes · SHA-256

0f36167f7531f05e1a7f41976ca4afb71acdc55870aecf382da739c2ddf4847e

The separate audit archive contains the frozen source rows, raw outputs, gold, grades, prompts and offline reproducer. It remains private pending distribution terms. The numerical ZIP alone supports numerical inspection, not independent reconstruction from raw grades. No new release license or DOI has been assigned. Public source article URLs cannot be reconstructed from the saved article IDs alone.

Three meanings of reproduction

  1. Recompute frozen scores: supported offline by the private audit snapshot and standard-library Python reproducer; no paid calls.
  2. Regenerate model answers: a separate paid experiment, with changing endpoints and nondeterministic remote sampling.
  3. Rebuild the original selection: requires the original eligibility database and historical ranking context. The selected-row construction is preserved, but the full pool is not.

Draft citation and version history

Working attribution: Outback Yak Research; individual author credit and final release date remain to be confirmed. The prepared route is /whitepapers/nepnewsbench-v2/; this preview is not evidence of publication.

Outback Yak Research. NepNewsBench v2: what models retain, invent and mistranslate. Research draft, experiment 2.0-draft, September 2026.

Read the previous NepNewsCluster v0.3 paper →

17 September 2026: website review draft; frozen handoff results retained; Kimi resampling and the unresolved DeepSeek aliasing caveat added from the original journal. Legacy PDF version labels checked against its contents. No benchmark generations or paid grades were run for this page.

Lab logos: Lobe Icons · asset sources and color key · MIT license. Lab colors identify model developers; serving providers are listed separately.

Questions or collaboration: research@kchakhabar.com. Dataset reporting originates from the publishers listed above; the benchmark evaluates fidelity to supplied excerpts, not independent verification of their reporting.