Abstract
We evaluate 17 model–provider configurations on bilingual news consolidation using 100 selected story questions from K cha khabar. The questions contain 751 headline/excerpt rows from 29 publishers, including 688 Nepali and 63 English rows. Each candidate produces English and Nepali headlines and summaries, a slug, and typed entities. The candidate runs use OpenRouter with a pinned serving provider and saved endpoint metadata. A common production-derived prompt is verified by hashes; provider constraints create explicitly recorded differences in reasoning, temperature support and response format.
Three model passes construct source-linked fact sheets and a fourth pass adjudicates them. Two full-pass judges score 1,693 valid candidate outputs against 1,647 adjudicated facts, producing 3,385 completed grades. The failure-inclusive composite weights core-fact recall, a coverage-based precision proxy and two language-quality scores. We reconstruct all reported scores from raw records and report paired question-bootstrap intervals, judge-specific results, source-language slices and concrete errors.
Claude Fable 5.1 has the highest observed composite, 92.65/100. Gemini 3.1 Pro has the highest observed Nepali prose mean, 9.025/10. The observed cost–quality frontier spans Gemma, DeepSeek V4.1 Flash, GPT-5.6 Sol and Claude Fable 5.1, exposing different trade-offs between coverage, prose and recorded cost. Numeric, calendar and attribution failures remain even in high-scoring configurations. Across models, average output length strongly correlates with supporting-fact recall; within every model, the same association across stories is negative. These results describe workload-specific trade-offs and motivate further controlled evaluation. They do not establish human-level quality, identical decoding conditions or independently verified real-world truth.
Why this task needs its own evaluation
A news aggregator needs a short, faithful account of several reports that may differ in names, dates, emphasis and detail. In Nepal, it also needs parallel Nepali and English copy. A model can write natural Nepali while selecting the wrong facts. It can preserve an amount in Nepali and mistranslate it into English. It can turn an analyst's interpretation into a politician's declared intention. Those outcomes are consequential even when the paragraph sounds polished.
NepNewsBench evaluates this concrete consolidation task. It asks whether a model can use the supplied rows, retain the story's central facts, avoid additions the rows do not support, and produce usable copy in both languages. The task does not assess all Nepali capabilities. It does not measure open-domain knowledge, full-article reading, search quality, classification accuracy or the truth of the underlying news reports.
The important unit is the full model–provider configuration. A model ID alone does not describe the decoding policy, quantization, structured-output support or routing used to obtain an answer. Readers should be able to inspect those settings before interpreting a leaderboard.
- 01 / SelectOne story, 3–10 publisher rows; headline, timestamp and optional excerpt.
- 02 / GenerateEN + NE headline and summary, slug and typed entity list.
- 03 / Build goldThree fact extractions → one adjudicated, source-linked sheet.
- 04 / EvaluateTwo anonymous pointwise grades → components, failures and paired uncertainty.
The “clustering” in the earlier name describes the input unit. Neither release measures how accurately the upstream system assigns articles to clusters.
What changed from NepNewsCluster v0.3
The previous online paper (v0.3, May 2026) described 107 questions, 1,310 snippets and 15 models, with a three-axis rubric applied primarily by one judge. V2 changes the question construction, generation provenance and evaluation design. It uses 100 selected stories, a source-linked fact-sheet representation, pointwise grading of one anonymous candidate at a time, two full judging passes and repeated judge calibration. It records prompt hashes, model and provider identities, response formats, reasoning policy and per-call cost.
These changes make the experiment easier to inspect and its scoring easier to reproduce. They do not make the v2 score a direct continuation of the v0.3 score. An 88 in one version cannot be read as a gain over an 81 in the other. Models, dates, selected stories and the mathematical meaning of the score differ.
The experiment files describe an internally frozen protocol and document amendments. There is no independent public registration timestamp in this package. We therefore describe the protocol as internally frozen, rather than publicly preregistered.
| Dimension | NepNewsCluster v0.3 | NepNewsBench v2 | What improves / what remains |
|---|---|---|---|
| Question selection | 107 clusters stratified primarily by publisher coverage, from about 9,000 active clusters. | 100 stories stratified by salience, breadth and hard input shapes, from a stated eligibility pool of about 17,600. | Targets specific production risks; remains purposive, with some related stories. |
| Evidence per question | Up to 15 recent snippets, excerpts limited to 280 characters; 1,310 rows, 23 publishers. | 3–10 rows, one per publisher, peak-48h or initial-72h windows, untruncated saved excerpts; 751 rows, 29 publishers. | Improves temporal and publisher balance; fewer rows is a design change, not a larger benchmark. |
| Reference standard | No explicit gold fact sheets. | Three independent extractions plus adjudication; 1,647 source-linked facts. | Makes credited and omitted facts inspectable. Gold is still machine-generated. |
| Grading unit | One primary judge compared anonymized candidate options; three holistic axes. | One anonymous candidate per call, against sources and gold; two full-pass judges. | Reduces dependence on the candidate lineup and exposes judge disagreement. Bias is not eliminated. |
| Score and failures | Mean of NE prose, EN prose and topic/entity quality; primarily successful-output quality. | 40% core recall, 30% precision proxy, 15% each prose language; failed generations count zero. | Separates factual coverage from prose and includes serving failures. The /100 scales are incompatible. |
| Calibration and repetition | Informal 20-question cross-judge check; earlier lineup and a second DeepSeek pass supplied limited repeat evidence. | Twenty-question calibration; two repeats each for Opus and Gemini, one Astra pass; explicit agreement gates. | Stronger measured judge repeatability. Candidate-generation variance remains unmeasured in v2. |
| Run provenance | Mostly OpenRouter, some native APIs; nominal common temperature; rate-table cost estimates. | All 17 via OpenRouter, pinned hosts, returned IDs checked, prompt hashes and actual settings recorded, retained usage.cost. | More auditable serving conditions; endpoint snapshots can refresh on resume and decoding still differs. |
| Protocol and audit | Descriptive rubric refined after an earlier run. | Internally frozen plans, dated amendments, numerical reconstruction, explicit deviations. | Better traceability; no independent public preregistration or completed human gold review. |
Archive note: the May online paper covers 15 models and identifies itself as v0.3. The earlier blog describes 13 models and v0.2. The legacy PDF has v0.2 in its filename but v0.3 and 15 models inside. These artifacts remain available as historical records.
Dataset and source construction
The selected cluster window is 20 April–10 September 2026, expressed in Asia/Kathmandu time. The selection account describes a pool of approximately 17,600 eligible clusters with at least three articles and three publishers, excluding recurring beats. The preserved inputs support rebuilding the selected questions; the entire original production pool is not included, so the full selection from that population is not independently reproducible from the public numerical tables alone.
The 100 questions contain 16 lead stories, 32 salient stories, 29 breadth stories and 23 hard cases. The sampling procedure uses historical daily rankings, section quotas, difficult input shapes, deduplication and caps on related saga facets. This is a purposive, stratified test collection, not a random sample of all Nepal news. Some related events still recur.
For each selected story, the builder takes a 48-hour window ending on the peak ranking day, or an initial 72-hour window when no peak is available; it can fall back to the full cluster when publisher coverage is insufficient. It selects one row per publisher, preferring excerpts, classified rows and recency. When more than ten publishers remain, it samples evenly across time and preserves minority-language coverage where available. The final rows are ordered newest first. Exact logic and selected inputs are archived.
There are 751 selected rows: 688 Nepali (91.6%) and 63 English (8.4%). The questions divide into 77 Nepali-only, 17 mixed-language and six English-only inputs. All candidates still produce both output languages. The median Nepali excerpt length is 248 characters when empty excerpts are included, and 262 among nonempty excerpts. There are 67 headline-only rows, all Nepali: 8.9% of all selected rows and 9.7% of Nepali rows. These measured rates replace the rough 20% figure retained in earlier protocol prose.
The selected rows come from 29 publishers, listed with row and question counts in the appendix. “29 publishers” does not establish 29 editorially independent accounts: syndication and shared reporting can occur. Politics appears among the tags of 61 questions, but the theme tags overlap rather than partition the corpus. The corpus contains 943 gold entities and 1,647 gold facts: 426 core, 871 supporting and 350 detail facts.
The export preserves article IDs, publisher identity, source timestamps and the rows the models saw. It does not preserve canonical article URLs. Source links must be recovered from an authorized original export if they are to appear in a public appendix; publisher homepages are not substitutes for article-level citations.
Filter the numerical source inventory by publisher, row language, month, story track and excerpt availability. Filtering rows does not redefine the leaderboard.
All 29 publishers and their representation
Counts describe selected rows, not independent editorial accounts. One publisher contributes at most one row per selected story.
| Publisher | Rows | Questions | NE rows | EN rows | Headline only |
|---|---|---|---|---|---|
| Ajako Artha | 18 | 18 | 18 | 0 | 0 |
| Artha Sarokar | 9 | 9 | 9 | 0 | 0 |
| Baahrakhari | 19 | 19 | 19 | 0 | 17 |
| BBC News Nepali | 2 | 2 | 2 | 0 | 0 |
| Bizmandu | 19 | 19 | 19 | 0 | 0 |
| Global Aawaj | 26 | 26 | 26 | 0 | 0 |
| Gorkhapatra Online | 52 | 52 | 52 | 0 | 0 |
| The Himalayan Times | 1 | 1 | 0 | 1 | 0 |
| Himal Press | 50 | 50 | 50 | 0 | 0 |
| Khabarhub | 58 | 58 | 58 | 0 | 0 |
| Khabarhub (English) | 12 | 12 | 0 | 12 | 0 |
| myRepublica | 3 | 3 | 0 | 3 | 0 |
| Nagarik News | 66 | 66 | 66 | 0 | 0 |
| Nepal Lead | 9 | 9 | 9 | 0 | 8 |
| Nepal News (English) | 15 | 15 | 0 | 15 | 0 |
| Nepal Press | 53 | 53 | 53 | 0 | 0 |
| Nepal Samaya | 9 | 9 | 9 | 0 | 9 |
| Nepal Views | 11 | 11 | 11 | 0 | 11 |
| News of Nepal | 32 | 32 | 32 | 0 | 0 |
| OnlineKhabar | 62 | 62 | 62 | 0 | 0 |
| OnlineKhabar (English) | 7 | 7 | 0 | 7 | 0 |
| Pradesh Khabar | 34 | 34 | 34 | 0 | 0 |
| Rajdhani National Daily | 33 | 33 | 33 | 0 | 0 |
| Ratopati | 73 | 73 | 73 | 0 | 1 |
| Ratopati (English) | 15 | 15 | 0 | 15 | 0 |
| The Rising Nepal | 10 | 10 | 0 | 10 | 0 |
| Setopati | 34 | 34 | 34 | 0 | 6 |
| Thaha Khabar | 18 | 18 | 18 | 0 | 15 |
| Ujyaalo Online | 1 | 1 | 1 | 0 | 0 |
Candidate generation and provenance
Seventeen completed API configurations define the leaderboard. They span Anthropic, OpenAI, Google, xAI, DeepSeek, Alibaba/Qwen, Moonshot AI and Z.AI model families. Gemma is served by Venice. All included candidate API runs use OpenRouter. Earlier agent-transport runs and excluded candidates are not silently pooled into the leaderboard.
The runner requests a single pinned provider with fallbacks disabled and checks the returned model and provider names. Endpoint snapshots record supported parameters, prices, context length and reported quantization. Unknown quantization is left unknown. Kimi records mxfp4, GLM fp8 and Gemma bf16. A provider pin narrows the serving configuration; it does not freeze remote model weights or prove that a preview alias cannot change.
Every candidate receives the same system prompt and the same question-specific user prompt. The package verifies the system hash and all 100 exported user-prompt hashes. The intended sampling temperature is 0.2, but some endpoints do not advertise support and the run metadata records it as not honoured. This flag describes the recorded capability check, not an independent measurement of a provider's internal sampling.
Eleven configurations disable reasoning. Six use the recorded minimum mandatory reasoning policy: Fable, Gemini 3.8 Flash, Gemini 3.1 Pro, Grok and GLM use low; Qwen Max uses minimal. Reasoning-enabled candidates receive 4,096 total output tokens; the others receive 1,536. Four configurations use JSON-object mode because their pinned endpoints lack structured-output support: both DeepSeek runs, GLM and Qwen 27B. The others request strict JSON schema. Consequently, v2 compares practical supported configurations, not identical decoding conditions. It cannot isolate the causal effect of reasoning.
There is one retained generation per model-question slot. A documented exception affects Kimi K3: after an erroneous resume policy, Q74 and Q87 were each resampled once following a schema violation. Q74 failed again; Q87 succeeded. These retained outcomes remain in the reported leaderboard. This prevents describing every cell as an untouched first sample. Transient transport failures may be retried. Candidate truncations, empty outputs and parser failures count as failed answers. The parser follows production leniency: it requires three core fields and normalizes optional fields; successful parsing does not prove that all six intended fields arrived with perfect strict-schema compliance.
Saved per-call cost and latency describe retained calls. Earlier overwritten retries and waiting time are not reconstructed into an account-wide bill or end-to-end service latency. Prices are historical experiment records, not a live price comparison.
Evaluation costs much more than obtaining the candidate answers. Retained candidate generation costs total $23.36; gold construction, calibration and full grading total $354.65, about 15.2 times as much. That expense buys inspectable fact sheets, repeated judging and a second full judge, rather than additional candidate samples. The total across these retained records is $378.01; it excludes overwritten attempts and other earlier work.
Inspect the exact configuration
17 of 17 candidate configurations. All use OpenRouter, with provider fallbacks disabled. Historical settings from saved calls; no current catalogue claims.
Claude Fable 5.1Anthropic · reasoning low · 4,096 tokens
- Exact requested ID
anthropic/claude-fable-5.1- Pinned provider slug
- anthropic
- Format / requested temperature
- json_schema / 0.2 · honoured flag: no
- Quantization / snapshot UTC
- unknown / 2026-09-15T14:50:52.701Z
- Observed reasoning tokens
- 100 reporting calls; mean 2.89; max 289
- Evaluation roles
- Candidate · gold extractor
- Returned providers
- {"Anthropic":100}
- Returned model IDs
- {"anthropic/claude-fable-5.1":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "claude-fable-5-1",
"label": "Claude Fable 5.1",
"transport": "openrouter",
"modelSlug": "anthropic/claude-fable-5.1",
"reasoning": {
"enabled": true,
"effort": "low"
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 10,
"outputPricePerM": 50
},
"providerPolicy": {
"order": [
"anthropic"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T14:50:52.701Z",
"providerSlug": "anthropic",
"providerName": "Anthropic",
"tag": "anthropic",
"quantization": "unknown",
"contextLength": 1000000,
"pricing": {
"prompt": 10,
"completion": 50
},
"supportedParameters": [
"max_tokens",
"stop",
"reasoning",
"include_reasoning",
"tools",
"structured_outputs",
"response_format",
"verbosity",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low"
],
"default_effort": "high"
}
},
"temperatureHonoured": false,
"responseFormatMode": "json_schema",
"maxTokensSent": 4096,
"requestParams": {
"model": "anthropic/claude-fable-5.1",
"max_tokens": 4096,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"effort": "low"
},
"provider": {
"order": [
"anthropic"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T14:50:52.703Z",
"finishedAt": "2026-09-15T15:02:04.223Z",
"goldContributor": true,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Anthropic": 100
},
"servedModelIds": {
"anthropic/claude-fable-5.1": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 2.89,
"maxWhenReported": 289
}
}Claude Haiku 4.5Anthropic · reasoning off · 1,536 tokens
- Exact requested ID
anthropic/claude-haiku-4.5- Pinned provider slug
- anthropic
- Format / requested temperature
- json_schema / 0.2 · honoured flag: yes
- Quantization / snapshot UTC
- unknown / 2026-09-15T13:41:28.086Z
- Observed reasoning tokens
- 100 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate
- Returned providers
- {"Anthropic":100}
- Returned model IDs
- {"anthropic/claude-haiku-4.5":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "claude-haiku-4-5",
"label": "Claude Haiku 4.5",
"transport": "openrouter",
"modelSlug": "anthropic/claude-haiku-4.5",
"reasoning": {
"enabled": false
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 1,
"outputPricePerM": 5
},
"providerPolicy": {
"order": [
"anthropic"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T13:41:28.086Z",
"providerSlug": "anthropic",
"providerName": "Anthropic",
"tag": "anthropic",
"quantization": "unknown",
"contextLength": 200000,
"pricing": {
"prompt": 1,
"completion": 5
},
"supportedParameters": [
"max_tokens",
"top_p",
"temperature",
"stop",
"reasoning",
"include_reasoning",
"tools",
"tool_choice",
"top_k",
"structured_outputs",
"response_format"
],
"status": 0,
"modelReasoning": {
"mandatory": false
}
},
"temperatureHonoured": true,
"responseFormatMode": "json_schema",
"maxTokensSent": 1536,
"requestParams": {
"model": "anthropic/claude-haiku-4.5",
"max_tokens": 1536,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"enabled": false
},
"provider": {
"order": [
"anthropic"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T13:41:28.087Z",
"finishedAt": "2026-09-15T13:47:42.248Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Anthropic": 100
},
"servedModelIds": {
"anthropic/claude-haiku-4.5": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}Claude Opus 5Anthropic · reasoning off · 1,536 tokens
- Exact requested ID
anthropic/claude-opus-5- Pinned provider slug
- anthropic
- Format / requested temperature
- json_schema / 0.2 · honoured flag: no
- Quantization / snapshot UTC
- unknown / 2026-09-15T14:42:43.433Z
- Observed reasoning tokens
- 100 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate · full-pass judge
- Returned providers
- {"Anthropic":100}
- Returned model IDs
- {"anthropic/claude-opus-5":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "claude-opus-5",
"label": "Claude Opus 5",
"transport": "openrouter",
"modelSlug": "anthropic/claude-opus-5",
"reasoning": {
"enabled": false
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 5,
"outputPricePerM": 25
},
"providerPolicy": {
"order": [
"anthropic"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T14:42:43.433Z",
"providerSlug": "anthropic",
"providerName": "Anthropic",
"tag": "anthropic",
"quantization": "unknown",
"contextLength": 1000000,
"pricing": {
"prompt": 5,
"completion": 25
},
"supportedParameters": [
"max_tokens",
"stop",
"reasoning",
"include_reasoning",
"tool_choice",
"tools",
"structured_outputs",
"response_format",
"verbosity",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"default_enabled": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low"
],
"default_effort": "high"
}
},
"temperatureHonoured": false,
"responseFormatMode": "json_schema",
"maxTokensSent": 1536,
"requestParams": {
"model": "anthropic/claude-opus-5",
"max_tokens": 1536,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"enabled": false
},
"provider": {
"order": [
"anthropic"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T14:42:43.434Z",
"finishedAt": "2026-09-15T14:50:51.399Z",
"goldContributor": false,
"fullPassJudge": true,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Anthropic": 100
},
"servedModelIds": {
"anthropic/claude-opus-5": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}Claude Sonnet 5Anthropic · reasoning off · 1,536 tokens
- Exact requested ID
anthropic/claude-sonnet-5- Pinned provider slug
- anthropic
- Format / requested temperature
- json_schema / 0.2 · honoured flag: no
- Quantization / snapshot UTC
- unknown / 2026-09-15T14:35:32.117Z
- Observed reasoning tokens
- 100 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate
- Returned providers
- {"Anthropic":100}
- Returned model IDs
- {"anthropic/claude-sonnet-5":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "claude-sonnet-5",
"label": "Claude Sonnet 5",
"transport": "openrouter",
"modelSlug": "anthropic/claude-sonnet-5",
"reasoning": {
"enabled": false
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 2,
"outputPricePerM": 10
},
"providerPolicy": {
"order": [
"anthropic"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T14:35:32.117Z",
"providerSlug": "anthropic",
"providerName": "Anthropic",
"tag": "anthropic",
"quantization": "unknown",
"contextLength": 1000000,
"pricing": {
"prompt": 2,
"completion": 10
},
"supportedParameters": [
"max_tokens",
"stop",
"reasoning",
"include_reasoning",
"tools",
"tool_choice",
"structured_outputs",
"response_format",
"verbosity",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"default_enabled": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low"
],
"default_effort": "high"
}
},
"temperatureHonoured": false,
"responseFormatMode": "json_schema",
"maxTokensSent": 1536,
"requestParams": {
"model": "anthropic/claude-sonnet-5",
"max_tokens": 1536,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"enabled": false
},
"provider": {
"order": [
"anthropic"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T14:35:32.118Z",
"finishedAt": "2026-09-15T14:42:03.437Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Anthropic": 100
},
"servedModelIds": {
"anthropic/claude-sonnet-5": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}DeepSeek V4.1 FlashDeepSeek · reasoning off · 1,536 tokens
- Exact requested ID
deepseek/deepseek-v4.1-flash- Pinned provider slug
- deepseek
- Format / requested temperature
- json_object / 0.2 · honoured flag: yes
- Quantization / snapshot UTC
- unknown / 2026-09-15T22:24:22.295Z
- Observed reasoning tokens
- 100 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate
- Returned providers
- {"DeepSeek":100}
- Returned model IDs
- {"deepseek/deepseek-v4.1-flash":100}
The journal notes possible upstream aliasing between the DeepSeek endpoints; different returned IDs do not prove distinct underlying weights.
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "deepseek-v4-1-flash",
"label": "DeepSeek V4.1 Flash",
"transport": "openrouter",
"modelSlug": "deepseek/deepseek-v4.1-flash",
"reasoning": {
"enabled": false
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 0.15,
"outputPricePerM": 0.6
},
"providerPolicy": {
"order": [
"deepseek"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T22:24:22.295Z",
"providerSlug": "deepseek",
"providerName": "DeepSeek",
"tag": "deepseek",
"quantization": "unknown",
"contextLength": 1048576,
"pricing": {
"prompt": 0.15,
"completion": 0.6
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"stop",
"frequency_penalty",
"presence_penalty",
"logprobs",
"top_logprobs",
"tools",
"tool_choice",
"response_format",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"default_enabled": true,
"supported_efforts": [
"max",
"high",
"low"
],
"default_effort": "high"
}
},
"temperatureHonoured": true,
"responseFormatMode": "json_object",
"maxTokensSent": 1536,
"requestParams": {
"model": "deepseek/deepseek-v4.1-flash",
"max_tokens": 1536,
"temperature": 0.2,
"response_format": "{ type: \"json_object\" } — endpoint lacks structured_outputs",
"reasoning": {
"enabled": false
},
"provider": {
"order": [
"deepseek"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T22:24:22.296Z",
"finishedAt": "2026-09-15T22:26:16.593Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"DeepSeek": 100
},
"servedModelIds": {
"deepseek/deepseek-v4.1-flash": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}DeepSeek V4 ProDeepSeek · reasoning off · 1,536 tokens
- Exact requested ID
deepseek/deepseek-v4-pro-0813- Pinned provider slug
- deepseek
- Format / requested temperature
- json_object / 0.2 · honoured flag: yes
- Quantization / snapshot UTC
- unknown / 2026-09-15T22:21:26.039Z
- Observed reasoning tokens
- 100 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate
- Returned providers
- {"DeepSeek":100}
- Returned model IDs
- {"deepseek/deepseek-v4-pro-0813":100}
The journal notes possible upstream aliasing between the DeepSeek endpoints; different returned IDs do not prove distinct underlying weights.
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "deepseek-v4-pro-0813",
"label": "DeepSeek V4 Pro (0813 GA)",
"transport": "openrouter",
"modelSlug": "deepseek/deepseek-v4-pro-0813",
"reasoning": {
"enabled": false
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 0.66,
"outputPricePerM": 1.98
},
"providerPolicy": {
"order": [
"deepseek"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T22:21:26.039Z",
"providerSlug": "deepseek",
"providerName": "DeepSeek",
"tag": "deepseek",
"quantization": "unknown",
"contextLength": 1048576,
"pricing": {
"prompt": 0.66,
"completion": 1.9800000000000002
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"stop",
"frequency_penalty",
"presence_penalty",
"logprobs",
"top_logprobs",
"tools",
"tool_choice",
"response_format",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"supported_efforts": [
"max",
"high",
"low"
],
"default_effort": "high"
}
},
"temperatureHonoured": true,
"responseFormatMode": "json_object",
"maxTokensSent": 1536,
"requestParams": {
"model": "deepseek/deepseek-v4-pro-0813",
"max_tokens": 1536,
"temperature": 0.2,
"response_format": "{ type: \"json_object\" } — endpoint lacks structured_outputs",
"reasoning": {
"enabled": false
},
"provider": {
"order": [
"deepseek"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T22:21:26.039Z",
"finishedAt": "2026-09-15T22:24:20.845Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"DeepSeek": 100
},
"servedModelIds": {
"deepseek/deepseek-v4-pro-0813": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}Gemini 3.1 ProGoogle AI Studio · reasoning low · 4,096 tokens
- Exact requested ID
google/gemini-3.1-pro-preview- Pinned provider slug
- google-ai-studio
- Format / requested temperature
- json_schema / 0.2 · honoured flag: yes
- Quantization / snapshot UTC
- unknown / 2026-09-15T14:14:23.854Z
- Observed reasoning tokens
- 100 reporting calls; mean 203.49; max 1225
- Evaluation roles
- Candidate
- Returned providers
- {"Google AI Studio":100}
- Returned model IDs
- {"google/gemini-3.1-pro-preview":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "gemini-3-1-pro",
"label": "Gemini 3.1 Pro (preview)",
"transport": "openrouter",
"modelSlug": "google/gemini-3.1-pro-preview",
"reasoning": {
"enabled": true,
"effort": "low"
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 2,
"outputPricePerM": 12
},
"providerPolicy": {
"order": [
"google-ai-studio"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T14:14:23.854Z",
"providerSlug": "google-ai-studio",
"providerName": "Google AI Studio",
"tag": "google-ai-studio",
"quantization": "unknown",
"contextLength": 1048576,
"pricing": {
"prompt": 2,
"completion": 12
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"seed",
"response_format",
"tools",
"tool_choice",
"structured_outputs",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"supported_efforts": [
"high",
"medium",
"low"
],
"default_effort": "medium"
}
},
"temperatureHonoured": true,
"responseFormatMode": "json_schema",
"maxTokensSent": 4096,
"requestParams": {
"model": "google/gemini-3.1-pro-preview",
"max_tokens": 4096,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"effort": "low"
},
"provider": {
"order": [
"google-ai-studio"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T14:14:23.854Z",
"finishedAt": "2026-09-15T14:18:43.970Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Google AI Studio": 100
},
"servedModelIds": {
"google/gemini-3.1-pro-preview": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 203.49,
"maxWhenReported": 1225
}
}Gemini 3.8 FlashGoogle AI Studio · reasoning low · 4,096 tokens
- Exact requested ID
google/gemini-3.8-flash- Pinned provider slug
- google-ai-studio
- Format / requested temperature
- json_schema / 0.2 · honoured flag: yes
- Quantization / snapshot UTC
- unknown / 2026-09-15T13:18:52.521Z
- Observed reasoning tokens
- 100 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate · gold extractor · full-pass judge
- Returned providers
- {"Google AI Studio":100}
- Returned model IDs
- {"google/gemini-3.8-flash":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "gemini-3-8-flash",
"label": "Gemini 3.8 Flash",
"transport": "openrouter",
"modelSlug": "google/gemini-3.8-flash",
"reasoning": {
"enabled": true,
"effort": "low"
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 0.75,
"outputPricePerM": 3.75
},
"providerPolicy": {
"order": [
"google-ai-studio"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T13:18:52.521Z",
"providerSlug": "google-ai-studio",
"providerName": "Google AI Studio",
"tag": "google-ai-studio",
"quantization": "unknown",
"contextLength": 1048576,
"pricing": {
"prompt": 0.75,
"completion": 3.75
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"seed",
"response_format",
"structured_outputs",
"tool_choice",
"tools",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"default_enabled": true,
"supported_efforts": [
"high",
"medium",
"low"
],
"default_effort": "medium"
}
},
"temperatureHonoured": true,
"responseFormatMode": "json_schema",
"maxTokensSent": 4096,
"requestParams": {
"model": "google/gemini-3.8-flash",
"max_tokens": 4096,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"effort": "low"
},
"provider": {
"order": [
"google-ai-studio"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T13:18:52.521Z",
"finishedAt": "2026-09-15T13:20:44.643Z",
"goldContributor": true,
"fullPassJudge": true,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Google AI Studio": 100
},
"servedModelIds": {
"google/gemini-3.8-flash": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}Gemini 3 FlashGoogle AI Studio · reasoning off · 1,536 tokens
- Exact requested ID
google/gemini-3-flash-preview- Pinned provider slug
- google-ai-studio
- Format / requested temperature
- json_schema / 0.2 · honoured flag: yes
- Quantization / snapshot UTC
- unknown / 2026-09-15T13:16:55.716Z
- Observed reasoning tokens
- 100 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate
- Returned providers
- {"Google AI Studio":100}
- Returned model IDs
- {"google/gemini-3-flash-preview":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "gemini-3-flash",
"label": "Gemini 3 Flash (preview) — prod primary",
"transport": "openrouter",
"modelSlug": "google/gemini-3-flash-preview",
"reasoning": {
"enabled": false
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 0.5,
"outputPricePerM": 3
},
"providerPolicy": {
"order": [
"google-ai-studio"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T13:16:55.716Z",
"providerSlug": "google-ai-studio",
"providerName": "Google AI Studio",
"tag": "google-ai-studio",
"quantization": "unknown",
"contextLength": 1048576,
"pricing": {
"prompt": 0.5,
"completion": 3
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"seed",
"response_format",
"tools",
"tool_choice",
"structured_outputs",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"supported_efforts": [
"high",
"medium",
"low",
"minimal"
],
"default_effort": "medium"
}
},
"temperatureHonoured": true,
"responseFormatMode": "json_schema",
"maxTokensSent": 1536,
"requestParams": {
"model": "google/gemini-3-flash-preview",
"max_tokens": 1536,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"enabled": false
},
"provider": {
"order": [
"google-ai-studio"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T13:16:55.717Z",
"finishedAt": "2026-09-15T13:18:50.989Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Google AI Studio": 100
},
"servedModelIds": {
"google/gemini-3-flash-preview": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}Gemma 4 31BVenice · reasoning off · 1,536 tokens
- Exact requested ID
google/gemma-4-31b-it- Pinned provider slug
- venice
- Format / requested temperature
- json_schema / 0.2 · honoured flag: yes
- Quantization / snapshot UTC
- bf16 / 2026-09-15T13:20:46.958Z
- Observed reasoning tokens
- 100 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate
- Returned providers
- {"Venice":100}
- Returned model IDs
- {"google/gemma-4-31b-it":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "gemma-4-31b",
"label": "Gemma 4 31B (it)",
"transport": "openrouter",
"modelSlug": "google/gemma-4-31b-it",
"reasoning": {
"enabled": false
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 0.12,
"outputPricePerM": 0.36
},
"providerPolicy": {
"order": [
"venice"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T13:20:46.958Z",
"providerSlug": "venice",
"providerName": "Venice",
"tag": "venice/bf16",
"quantization": "bf16",
"contextLength": 256000,
"pricing": {
"prompt": 0.12,
"completion": 0.36
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"stop",
"frequency_penalty",
"presence_penalty",
"top_k",
"response_format",
"structured_outputs",
"tools",
"tool_choice",
"logprobs",
"top_logprobs"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"default_enabled": false
}
},
"temperatureHonoured": true,
"responseFormatMode": "json_schema",
"maxTokensSent": 1536,
"requestParams": {
"model": "google/gemma-4-31b-it",
"max_tokens": 1536,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"enabled": false
},
"provider": {
"order": [
"venice"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T13:20:46.958Z",
"finishedAt": "2026-09-15T13:27:00.336Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Venice": 100
},
"servedModelIds": {
"google/gemma-4-31b-it": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}GLM 5.3Z.AI · reasoning low · 4,096 tokens
- Exact requested ID
z-ai/glm-5.3- Pinned provider slug
- z-ai
- Format / requested temperature
- json_object / 0.2 · honoured flag: yes
- Quantization / snapshot UTC
- fp8 / 2026-09-15T14:05:13.751Z
- Observed reasoning tokens
- 100 reporting calls; mean 4.8; max 108
- Evaluation roles
- Candidate
- Returned providers
- {"Z.AI":100}
- Returned model IDs
- {"z-ai/glm-5.3":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "glm-5-3",
"label": "GLM 5.3",
"transport": "openrouter",
"modelSlug": "z-ai/glm-5.3",
"reasoning": {
"enabled": true,
"effort": "low"
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 1.4,
"outputPricePerM": 4.4
},
"providerPolicy": {
"order": [
"z-ai"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T14:05:13.751Z",
"providerSlug": "z-ai",
"providerName": "Z.AI",
"tag": "z-ai/fp8",
"quantization": "fp8",
"contextLength": 1048576,
"pricing": {
"prompt": 1.4,
"completion": 4.4
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"tools",
"tool_choice",
"top_k",
"response_format",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"default_enabled": true,
"supported_efforts": [
"max",
"high",
"low"
],
"default_effort": "max"
}
},
"temperatureHonoured": true,
"responseFormatMode": "json_object",
"maxTokensSent": 4096,
"requestParams": {
"model": "z-ai/glm-5.3",
"max_tokens": 4096,
"temperature": 0.2,
"response_format": "{ type: \"json_object\" } — endpoint lacks structured_outputs",
"reasoning": {
"effort": "low"
},
"provider": {
"order": [
"z-ai"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T14:05:13.752Z",
"finishedAt": "2026-09-15T14:14:21.914Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Z.AI": 100
},
"servedModelIds": {
"z-ai/glm-5.3": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 4.8,
"maxWhenReported": 108
}
}GPT-5.6 LunaOpenAI · reasoning off · 1,536 tokens
- Exact requested ID
openai/gpt-5.6-luna- Pinned provider slug
- openai
- Format / requested temperature
- json_schema / 0.2 · honoured flag: no
- Quantization / snapshot UTC
- unknown / 2026-09-15T13:27:02.062Z
- Observed reasoning tokens
- 100 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate
- Returned providers
- {"OpenAI":100}
- Returned model IDs
- {"openai/gpt-5.6-luna":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "gpt-5-6-luna",
"label": "GPT-5.6 Luna",
"transport": "openrouter",
"modelSlug": "openai/gpt-5.6-luna",
"reasoning": {
"enabled": false
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 0.2,
"outputPricePerM": 1.2
},
"providerPolicy": {
"order": [
"openai"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T13:27:02.062Z",
"providerSlug": "openai",
"providerName": "OpenAI",
"tag": "openai",
"quantization": "unknown",
"contextLength": 1050000,
"pricing": {
"prompt": 0.19999999999999998,
"completion": 1.2
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"seed",
"max_tokens",
"response_format",
"structured_outputs",
"tools",
"tool_choice",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"default_enabled": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low",
"none"
],
"default_effort": "medium"
}
},
"temperatureHonoured": false,
"responseFormatMode": "json_schema",
"maxTokensSent": 1536,
"requestParams": {
"model": "openai/gpt-5.6-luna",
"max_tokens": 1536,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"enabled": false
},
"provider": {
"order": [
"openai"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T13:27:02.062Z",
"finishedAt": "2026-09-15T13:29:46.919Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"OpenAI": 100
},
"servedModelIds": {
"openai/gpt-5.6-luna": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}GPT-5.6 SolOpenAI · reasoning off · 1,536 tokens
- Exact requested ID
openai/gpt-5.6-sol- Pinned provider slug
- openai
- Format / requested temperature
- json_schema / 0.2 · honoured flag: no
- Quantization / snapshot UTC
- unknown / 2026-09-15T13:47:43.924Z
- Observed reasoning tokens
- 100 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate
- Returned providers
- {"OpenAI":100}
- Returned model IDs
- {"openai/gpt-5.6-sol":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "gpt-5-6-sol",
"label": "GPT-5.6 Sol",
"transport": "openrouter",
"modelSlug": "openai/gpt-5.6-sol",
"reasoning": {
"enabled": false
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 2,
"outputPricePerM": 10
},
"providerPolicy": {
"order": [
"openai"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T13:47:43.924Z",
"providerSlug": "openai",
"providerName": "OpenAI",
"tag": "openai",
"quantization": "unknown",
"contextLength": 1050000,
"pricing": {
"prompt": 2,
"completion": 10
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"seed",
"max_tokens",
"response_format",
"structured_outputs",
"tools",
"tool_choice",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"default_enabled": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low",
"none"
],
"default_effort": "medium"
}
},
"temperatureHonoured": false,
"responseFormatMode": "json_schema",
"maxTokensSent": 1536,
"requestParams": {
"model": "openai/gpt-5.6-sol",
"max_tokens": 1536,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"enabled": false
},
"provider": {
"order": [
"openai"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T13:47:43.925Z",
"finishedAt": "2026-09-15T13:56:40.246Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"OpenAI": 100
},
"servedModelIds": {
"openai/gpt-5.6-sol": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}Grok 4.6xAI · reasoning low · 4,096 tokens
- Exact requested ID
x-ai/grok-4.6- Pinned provider slug
- xai
- Format / requested temperature
- json_schema / 0.2 · honoured flag: yes
- Quantization / snapshot UTC
- unknown / 2026-09-15T13:56:41.552Z
- Observed reasoning tokens
- 100 reporting calls; mean 353.26; max 959
- Evaluation roles
- Candidate
- Returned providers
- {"xAI":100}
- Returned model IDs
- {"x-ai/grok-4.6":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "grok-4-6",
"label": "Grok 4.6",
"transport": "openrouter",
"modelSlug": "x-ai/grok-4.6",
"reasoning": {
"enabled": true,
"effort": "low"
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 2,
"outputPricePerM": 6
},
"providerPolicy": {
"order": [
"xai"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T13:56:41.552Z",
"providerSlug": "xai",
"providerName": "xAI",
"tag": "xai",
"quantization": "unknown",
"contextLength": 500000,
"pricing": {
"prompt": 2,
"completion": 6
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"structured_outputs",
"response_format",
"max_tokens",
"temperature",
"top_p",
"seed",
"logprobs",
"top_logprobs",
"tools",
"tool_choice",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"default_enabled": true,
"supported_efforts": [
"xhigh",
"high",
"medium",
"low"
],
"default_effort": "high"
}
},
"temperatureHonoured": true,
"responseFormatMode": "json_schema",
"maxTokensSent": 4096,
"requestParams": {
"model": "x-ai/grok-4.6",
"max_tokens": 4096,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"effort": "low"
},
"provider": {
"order": [
"xai"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T13:56:41.552Z",
"finishedAt": "2026-09-15T14:05:11.948Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"xAI": 100
},
"servedModelIds": {
"x-ai/grok-4.6": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 353.26,
"maxWhenReported": 959
}
}Kimi K3Moonshot AI · reasoning off · 1,536 tokens
- Exact requested ID
moonshotai/kimi-k3- Pinned provider slug
- moonshotai
- Format / requested temperature
- json_schema / 0.2 · honoured flag: no
- Quantization / snapshot UTC
- mxfp4 / 2026-09-15T15:15:26.266Z
- Observed reasoning tokens
- 99 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate
- Returned providers
- {"Moonshot AI":99,"unrecorded":1}
- Returned model IDs
- {"moonshotai/kimi-k3":99,"unrecorded":1}
Protocol exception: Q74 and Q87 were resampled once after schema failures. Q87 succeeded; Q74 remained a failure. Recorded quantization is mxfp4.
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "kimi-k3",
"label": "Kimi K3",
"transport": "openrouter",
"modelSlug": "moonshotai/kimi-k3",
"reasoning": {
"enabled": false
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 3,
"outputPricePerM": 15
},
"providerPolicy": {
"order": [
"moonshotai"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T15:15:26.266Z",
"providerSlug": "moonshotai",
"providerName": "Moonshot AI",
"tag": "moonshotai/mxfp4",
"quantization": "mxfp4",
"contextLength": 1048576,
"pricing": {
"prompt": 3,
"completion": 15
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"stop",
"frequency_penalty",
"presence_penalty",
"structured_outputs",
"response_format",
"tool_choice",
"tools",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"default_enabled": true,
"supported_efforts": [
"max",
"high",
"low"
],
"default_effort": "max"
}
},
"temperatureHonoured": false,
"responseFormatMode": "json_schema",
"maxTokensSent": 1536,
"requestParams": {
"model": "moonshotai/kimi-k3",
"max_tokens": 1536,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"enabled": false
},
"provider": {
"order": [
"moonshotai"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T14:42:04.847Z",
"finishedAt": "2026-09-15T15:16:18.451Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Moonshot AI": 99,
"unrecorded": 1
},
"servedModelIds": {
"moonshotai/kimi-k3": 99,
"unrecorded": 1
},
"observedReasoningTokens": {
"reportedCalls": 99,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}Qwen 3.8 27BAlibaba · reasoning off · 1,536 tokens
- Exact requested ID
qwen/qwen3.8-27b- Pinned provider slug
- alibaba
- Format / requested temperature
- json_object / 0.2 · honoured flag: yes
- Quantization / snapshot UTC
- unknown / 2026-09-15T15:12:52.601Z
- Observed reasoning tokens
- 100 reporting calls; mean 0; max 0
- Evaluation roles
- Candidate
- Returned providers
- {"Alibaba":100}
- Returned model IDs
- {"qwen/qwen3.8-27b":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "qwen3-8-27b",
"label": "Qwen 3.8 27B",
"transport": "openrouter",
"modelSlug": "qwen/qwen3.8-27b",
"reasoning": {
"enabled": false
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 0.42,
"outputPricePerM": 2.55
},
"providerPolicy": {
"order": [
"alibaba"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T15:12:52.601Z",
"providerSlug": "alibaba",
"providerName": "Alibaba",
"tag": "alibaba",
"quantization": "unknown",
"contextLength": 1000000,
"pricing": {
"prompt": 0.425,
"completion": 2.5500000000000003
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"seed",
"presence_penalty",
"response_format",
"top_k",
"frequency_penalty",
"stop",
"tools",
"tool_choice",
"logprobs",
"top_logprobs",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"default_enabled": true,
"supported_efforts": [
"xhigh",
"medium",
"low"
],
"default_effort": "xhigh"
}
},
"temperatureHonoured": true,
"responseFormatMode": "json_object",
"maxTokensSent": 1536,
"requestParams": {
"model": "qwen/qwen3.8-27b",
"max_tokens": 1536,
"temperature": 0.2,
"response_format": "{ type: \"json_object\" } — endpoint lacks structured_outputs",
"reasoning": {
"enabled": false
},
"provider": {
"order": [
"alibaba"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T13:33:25.867Z",
"finishedAt": "2026-09-15T15:13:40.946Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Alibaba": 100
},
"servedModelIds": {
"qwen/qwen3.8-27b": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 0,
"maxWhenReported": 0
}
}Qwen 3.8 MaxAlibaba · reasoning minimal · 4,096 tokens
- Exact requested ID
qwen/qwen3.8-max-0902- Pinned provider slug
- alibaba
- Format / requested temperature
- json_schema / 0.2 · honoured flag: yes
- Quantization / snapshot UTC
- unknown / 2026-09-15T15:13:41.876Z
- Observed reasoning tokens
- 100 reporting calls; mean 872.09; max 1474
- Evaluation roles
- Candidate
- Returned providers
- {"Alibaba":100}
- Returned model IDs
- {"qwen/qwen3.8-max-0902":100}
Full request, endpoint snapshot and parameter list
Per-run fields take precedence over the historical “invariants” object. The response_format string is a description of the archived request, not executable API JSON.
{
"modelKey": "qwen3-8-max",
"label": "Qwen 3.8 Max (0902)",
"transport": "openrouter",
"modelSlug": "qwen/qwen3.8-max-0902",
"reasoning": {
"enabled": true,
"effort": "minimal"
},
"promptSha256": "4344af0ac1ba92d71558cc4ee08e396be0f199f23884c01a706b12d3792888b0",
"invariants": {
"temperature": 0.2,
"maxTokens": 1536,
"reasoning": "off",
"responseFormat": "json_schema:strict",
"articlesOrder": "most-recent-first"
},
"pricing": {
"inputPricePerM": 2,
"outputPricePerM": 6
},
"providerPolicy": {
"order": [
"alibaba"
],
"allow_fallbacks": false
},
"endpoint": {
"fetchedAt": "2026-09-15T15:13:41.876Z",
"providerSlug": "alibaba",
"providerName": "Alibaba",
"tag": "alibaba",
"quantization": "unknown",
"contextLength": 1000000,
"pricing": {
"prompt": 2,
"completion": 6
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"seed",
"presence_penalty",
"response_format",
"tools",
"tool_choice",
"structured_outputs",
"logprobs",
"top_logprobs",
"top_k",
"frequency_penalty",
"stop",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"default_enabled": true,
"supported_efforts": [
"xhigh",
"high",
"medium",
"low",
"minimal"
],
"default_effort": "xhigh"
}
},
"temperatureHonoured": true,
"responseFormatMode": "json_schema",
"maxTokensSent": 4096,
"requestParams": {
"model": "qwen/qwen3.8-max-0902",
"max_tokens": 4096,
"temperature": 0.2,
"response_format": "prod SUMMARY_RESPONSE_FORMAT (json_schema, strict)",
"reasoning": {
"effort": "minimal"
},
"provider": {
"order": [
"alibaba"
],
"allow_fallbacks": false
}
},
"startedAt": "2026-09-15T14:18:45.741Z",
"finishedAt": "2026-09-15T15:15:25.406Z",
"goldContributor": false,
"fullPassJudge": false,
"quantizationEvidence": "endpoint snapshot; unknown is not full precision",
"servedProviders": {
"Alibaba": 100
},
"servedModelIds": {
"qwen/qwen3.8-max-0902": 100
},
"observedReasoningTokens": {
"reportedCalls": 100,
"meanWhenReported": 872.09,
"maxWhenReported": 1474
}
}Gold construction and judging use different settings
Candidate Opus has reasoning off; judge Opus uses medium effort. Gold extraction uses high effort, 16,000 tokens; adjudication uses high effort, 20,000 tokens. Full judges and calibration use medium effort, 12,000 tokens.
| Stage | Exact model ID | Reasoning | Token cap | Recorded provider |
|---|---|---|---|---|
| extractiongold/extract/claude-fable-5-1.json | anthropic/claude-fable-5.1 | high | 16000 | Anthropic |
| extractiongold/extract/gemini-3-8-flash.json | google/gemini-3.8-flash | high | 16000 | Google AI Studio |
| extractiongold/extract/gpt-6-astra.json | openai/gpt-6-astra | high | 16000 | OpenAI |
| adjudicationgold/gold.v2.json | openai/gpt-6-astra | high | 20,000 (archived code) | OpenAI |
| full-judgegrades/claude-opus-5.json | anthropic/claude-opus-5 | medium | 12000 | Anthropic |
| full-judgegrades/gemini-3-8-flash.json | google/gemini-3.8-flash | medium | 12000 | Google AI Studio |
| calibrationgrades/calibration/claude-opus-5-r1.json | anthropic/claude-opus-5 | medium | 12000 | Anthropic |
| calibrationgrades/calibration/claude-opus-5-r2.json | anthropic/claude-opus-5 | medium | 12000 | Anthropic |
| calibrationgrades/calibration/gemini-3-8-flash-r1.json | google/gemini-3.8-flash | medium | 12000 | Google AI Studio |
| calibrationgrades/calibration/gemini-3-8-flash-r2.json | google/gemini-3.8-flash | medium | 12000 | Google AI Studio |
| calibrationgrades/calibration/gpt-6-astra-r1.json | openai/gpt-6-astra | medium | 12000 | OpenAI |
Full stage metadata
[
{
"stageType": "extraction",
"artifact": "gold/extract/claude-fable-5-1.json",
"benchmark": "NepNewsBench",
"version": "2.0-draft",
"stage": "gold-extract",
"extractorKey": "claude-fable-5-1",
"label": "Claude Fable 5.1",
"contestant": true,
"modelSlug": "anthropic/claude-fable-5.1",
"endpoint": {
"fetchedAt": "2026-09-15T22:38:37.037Z",
"providerSlug": "anthropic",
"providerName": "Anthropic",
"tag": "anthropic",
"quantization": "unknown",
"contextLength": 1000000,
"pricing": {
"prompt": 10,
"completion": 50
},
"supportedParameters": [
"max_tokens",
"stop",
"reasoning",
"include_reasoning",
"tools",
"structured_outputs",
"response_format",
"verbosity",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low"
],
"default_effort": "high"
}
},
"reasoning": {
"enabled": true,
"effort": "high"
},
"maxTokens": 16000,
"promptSha256": {
"extractSystem": "5766dcc879f29bd88462078d431cde293ef3331f36146beb114e2cab54c68167",
"extractSchema": "6b28cc3ec5a9990d43eedf8be692ed8b48130ec2da402bd9daf3613d5790fac8",
"adjudicateSystem": "52c64292fff0a47c39f999dfd49a4a3b864f36268a46b31f86e68ca9281a6292",
"adjudicateSchema": "8c4043c596c4a3c541f51da60ecff85b3c59b3d2f92bcef91b60633a38ef491b"
},
"startedAt": "2026-09-15T22:38:37.038Z",
"finishedAt": "2026-09-16T02:13:01.244Z"
},
{
"stageType": "extraction",
"artifact": "gold/extract/gemini-3-8-flash.json",
"benchmark": "NepNewsBench",
"version": "2.0-draft",
"stage": "gold-extract",
"extractorKey": "gemini-3-8-flash",
"label": "Gemini 3.8 Flash",
"contestant": true,
"modelSlug": "google/gemini-3.8-flash",
"endpoint": {
"fetchedAt": "2026-09-15T22:54:24.720Z",
"providerSlug": "google-ai-studio",
"providerName": "Google AI Studio",
"tag": "google-ai-studio",
"quantization": "unknown",
"contextLength": 1048576,
"pricing": {
"prompt": 0.75,
"completion": 3.75
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"seed",
"response_format",
"structured_outputs",
"tool_choice",
"tools",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"default_enabled": true,
"supported_efforts": [
"high",
"medium",
"low"
],
"default_effort": "medium"
}
},
"reasoning": {
"enabled": true,
"effort": "high"
},
"maxTokens": 16000,
"promptSha256": {
"extractSystem": "5766dcc879f29bd88462078d431cde293ef3331f36146beb114e2cab54c68167",
"extractSchema": "6b28cc3ec5a9990d43eedf8be692ed8b48130ec2da402bd9daf3613d5790fac8",
"adjudicateSystem": "52c64292fff0a47c39f999dfd49a4a3b864f36268a46b31f86e68ca9281a6292",
"adjudicateSchema": "8c4043c596c4a3c541f51da60ecff85b3c59b3d2f92bcef91b60633a38ef491b"
},
"startedAt": "2026-09-15T22:54:24.721Z",
"finishedAt": "2026-09-15T23:16:17.333Z"
},
{
"stageType": "extraction",
"artifact": "gold/extract/gpt-6-astra.json",
"benchmark": "NepNewsBench",
"version": "2.0-draft",
"stage": "gold-extract",
"extractorKey": "gpt-6-astra",
"label": "GPT-6 Astra",
"contestant": false,
"modelSlug": "openai/gpt-6-astra",
"endpoint": {
"fetchedAt": "2026-09-15T22:40:05.050Z",
"providerSlug": "openai",
"providerName": "OpenAI",
"tag": "openai",
"quantization": "unknown",
"contextLength": 1050000,
"pricing": {
"prompt": 10,
"completion": 50
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"seed",
"max_tokens",
"response_format",
"structured_outputs",
"tools",
"tool_choice",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"default_enabled": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low"
],
"default_effort": "medium"
}
},
"reasoning": {
"enabled": true,
"effort": "high"
},
"maxTokens": 16000,
"promptSha256": {
"extractSystem": "5766dcc879f29bd88462078d431cde293ef3331f36146beb114e2cab54c68167",
"extractSchema": "6b28cc3ec5a9990d43eedf8be692ed8b48130ec2da402bd9daf3613d5790fac8",
"adjudicateSystem": "52c64292fff0a47c39f999dfd49a4a3b864f36268a46b31f86e68ca9281a6292",
"adjudicateSchema": "8c4043c596c4a3c541f51da60ecff85b3c59b3d2f92bcef91b60633a38ef491b"
},
"startedAt": "2026-09-15T22:40:05.052Z",
"finishedAt": "2026-09-16T02:38:03.490Z"
},
{
"stageType": "adjudication",
"artifact": "gold/gold.v2.json",
"benchmark": "NepNewsBench",
"version": "2.0-draft",
"stage": "gold",
"adjudicatorKey": "gpt-6-astra",
"modelSlug": "openai/gpt-6-astra",
"endpoint": {
"fetchedAt": "2026-09-15T22:55:53.317Z",
"providerSlug": "openai",
"providerName": "OpenAI",
"tag": "openai",
"quantization": "unknown",
"contextLength": 1050000,
"pricing": {
"prompt": 10,
"completion": 50
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"seed",
"max_tokens",
"response_format",
"structured_outputs",
"tools",
"tool_choice",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"default_enabled": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low"
],
"default_effort": "medium"
}
},
"reasoning": {
"enabled": true,
"effort": "high"
},
"promptSha256": {
"extractSystem": "5766dcc879f29bd88462078d431cde293ef3331f36146beb114e2cab54c68167",
"extractSchema": "6b28cc3ec5a9990d43eedf8be692ed8b48130ec2da402bd9daf3613d5790fac8",
"adjudicateSystem": "52c64292fff0a47c39f999dfd49a4a3b864f36268a46b31f86e68ca9281a6292",
"adjudicateSchema": "8c4043c596c4a3c541f51da60ecff85b3c59b3d2f92bcef91b60633a38ef491b"
},
"extractors": [
"claude-fable-5-1",
"gpt-6-astra",
"gemini-3-8-flash"
],
"startedAt": "2026-09-15T22:55:53.318Z",
"finishedAt": "2026-09-16T04:29:41.267Z"
},
{
"stageType": "full-judge",
"artifact": "grades/claude-opus-5.json",
"benchmark": "NepNewsBench",
"version": "2.0-draft",
"stage": "grades",
"judgeKey": "claude-opus-5",
"label": "Claude Opus 5",
"contestant": true,
"modelSlug": "anthropic/claude-opus-5",
"endpoint": {
"fetchedAt": "2026-09-16T13:07:14.825Z",
"providerSlug": "anthropic",
"providerName": "Anthropic",
"tag": "anthropic",
"quantization": "unknown",
"contextLength": 1000000,
"pricing": {
"prompt": 5,
"completion": 25
},
"supportedParameters": [
"max_tokens",
"stop",
"reasoning",
"include_reasoning",
"tool_choice",
"tools",
"structured_outputs",
"response_format",
"verbosity",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"default_enabled": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low"
],
"default_effort": "high"
}
},
"reasoning": {
"enabled": true,
"effort": "medium"
},
"maxTokens": 12000,
"promptSha256": {
"system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
"schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
},
"goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
"mode": "full",
"startedAt": "2026-09-16T13:07:14.834Z",
"finishedAt": "2026-09-16T14:16:43.782Z"
},
{
"stageType": "full-judge",
"artifact": "grades/gemini-3-8-flash.json",
"benchmark": "NepNewsBench",
"version": "2.0-draft",
"stage": "grades",
"judgeKey": "gemini-3-8-flash",
"label": "Gemini 3.8 Flash",
"contestant": true,
"modelSlug": "google/gemini-3.8-flash",
"endpoint": {
"fetchedAt": "2026-09-16T13:07:17.245Z",
"providerSlug": "google-ai-studio",
"providerName": "Google AI Studio",
"tag": "google-ai-studio",
"quantization": "unknown",
"contextLength": 1048576,
"pricing": {
"prompt": 0.75,
"completion": 3.75
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"seed",
"response_format",
"structured_outputs",
"tool_choice",
"tools",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"default_enabled": true,
"supported_efforts": [
"high",
"medium",
"low"
],
"default_effort": "medium"
}
},
"reasoning": {
"enabled": true,
"effort": "medium"
},
"maxTokens": 12000,
"promptSha256": {
"system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
"schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
},
"goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
"mode": "full",
"startedAt": "2026-09-16T13:07:17.246Z",
"finishedAt": "2026-09-16T13:47:41.140Z"
},
{
"stageType": "calibration",
"artifact": "grades/calibration/claude-opus-5-r1.json",
"benchmark": "NepNewsBench",
"version": "2.0-draft",
"stage": "grades",
"judgeKey": "claude-opus-5",
"label": "Claude Opus 5",
"contestant": true,
"modelSlug": "anthropic/claude-opus-5",
"endpoint": {
"fetchedAt": "2026-09-16T04:30:11.903Z",
"providerSlug": "anthropic",
"providerName": "Anthropic",
"tag": "anthropic",
"quantization": "unknown",
"contextLength": 1000000,
"pricing": {
"prompt": 5,
"completion": 25
},
"supportedParameters": [
"max_tokens",
"stop",
"reasoning",
"include_reasoning",
"tool_choice",
"tools",
"structured_outputs",
"response_format",
"verbosity",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"default_enabled": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low"
],
"default_effort": "high"
}
},
"reasoning": {
"enabled": true,
"effort": "medium"
},
"maxTokens": 12000,
"promptSha256": {
"system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
"schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
},
"goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
"mode": "calibration",
"sample": {
"n": 20,
"seed": 42,
"questionIds": [
1,
13,
20,
21,
29,
36,
37,
39,
41,
44,
45,
53,
54,
64,
66,
79,
84,
86,
88,
94
],
"run": 1
},
"startedAt": "2026-09-16T04:30:11.904Z",
"finishedAt": "2026-09-16T04:50:54.612Z"
},
{
"stageType": "calibration",
"artifact": "grades/calibration/claude-opus-5-r2.json",
"benchmark": "NepNewsBench",
"version": "2.0-draft",
"stage": "grades",
"judgeKey": "claude-opus-5",
"label": "Claude Opus 5",
"contestant": true,
"modelSlug": "anthropic/claude-opus-5",
"endpoint": {
"fetchedAt": "2026-09-16T04:50:55.801Z",
"providerSlug": "anthropic",
"providerName": "Anthropic",
"tag": "anthropic",
"quantization": "unknown",
"contextLength": 1000000,
"pricing": {
"prompt": 5,
"completion": 25
},
"supportedParameters": [
"max_tokens",
"stop",
"reasoning",
"include_reasoning",
"tool_choice",
"tools",
"structured_outputs",
"response_format",
"verbosity",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": false,
"default_enabled": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low"
],
"default_effort": "high"
}
},
"reasoning": {
"enabled": true,
"effort": "medium"
},
"maxTokens": 12000,
"promptSha256": {
"system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
"schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
},
"goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
"mode": "calibration",
"sample": {
"n": 20,
"seed": 42,
"questionIds": [
1,
13,
20,
21,
29,
36,
37,
39,
41,
44,
45,
53,
54,
64,
66,
79,
84,
86,
88,
94
],
"run": 2
},
"startedAt": "2026-09-16T04:50:55.801Z",
"finishedAt": "2026-09-16T13:10:51.532Z"
},
{
"stageType": "calibration",
"artifact": "grades/calibration/gemini-3-8-flash-r1.json",
"benchmark": "NepNewsBench",
"version": "2.0-draft",
"stage": "grades",
"judgeKey": "gemini-3-8-flash",
"label": "Gemini 3.8 Flash",
"contestant": true,
"modelSlug": "google/gemini-3.8-flash",
"endpoint": {
"fetchedAt": "2026-09-16T04:30:08.060Z",
"providerSlug": "google-ai-studio",
"providerName": "Google AI Studio",
"tag": "google-ai-studio",
"quantization": "unknown",
"contextLength": 1048576,
"pricing": {
"prompt": 0.75,
"completion": 3.75
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"seed",
"response_format",
"structured_outputs",
"tool_choice",
"tools",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"default_enabled": true,
"supported_efforts": [
"high",
"medium",
"low"
],
"default_effort": "medium"
}
},
"reasoning": {
"enabled": true,
"effort": "medium"
},
"maxTokens": 12000,
"promptSha256": {
"system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
"schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
},
"goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
"mode": "calibration",
"sample": {
"n": 20,
"seed": 42,
"questionIds": [
1,
13,
20,
21,
29,
36,
37,
39,
41,
44,
45,
53,
54,
64,
66,
79,
84,
86,
88,
94
],
"run": 1
},
"startedAt": "2026-09-16T04:30:08.061Z",
"finishedAt": "2026-09-16T04:40:14.994Z"
},
{
"stageType": "calibration",
"artifact": "grades/calibration/gemini-3-8-flash-r2.json",
"benchmark": "NepNewsBench",
"version": "2.0-draft",
"stage": "grades",
"judgeKey": "gemini-3-8-flash",
"label": "Gemini 3.8 Flash",
"contestant": true,
"modelSlug": "google/gemini-3.8-flash",
"endpoint": {
"fetchedAt": "2026-09-16T04:40:15.680Z",
"providerSlug": "google-ai-studio",
"providerName": "Google AI Studio",
"tag": "google-ai-studio",
"quantization": "unknown",
"contextLength": 1048576,
"pricing": {
"prompt": 0.75,
"completion": 3.75
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"max_tokens",
"temperature",
"top_p",
"seed",
"response_format",
"structured_outputs",
"tool_choice",
"tools",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"default_enabled": true,
"supported_efforts": [
"high",
"medium",
"low"
],
"default_effort": "medium"
}
},
"reasoning": {
"enabled": true,
"effort": "medium"
},
"maxTokens": 12000,
"promptSha256": {
"system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
"schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
},
"goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
"mode": "calibration",
"sample": {
"n": 20,
"seed": 42,
"questionIds": [
1,
13,
20,
21,
29,
36,
37,
39,
41,
44,
45,
53,
54,
64,
66,
79,
84,
86,
88,
94
],
"run": 2
},
"startedAt": "2026-09-16T04:40:15.681Z",
"finishedAt": "2026-09-16T04:50:27.431Z"
},
{
"stageType": "calibration",
"artifact": "grades/calibration/gpt-6-astra-r1.json",
"benchmark": "NepNewsBench",
"version": "2.0-draft",
"stage": "grades",
"judgeKey": "gpt-6-astra",
"label": "GPT-6 Astra",
"contestant": false,
"modelSlug": "openai/gpt-6-astra",
"endpoint": {
"fetchedAt": "2026-09-16T04:30:14.051Z",
"providerSlug": "openai",
"providerName": "OpenAI",
"tag": "openai",
"quantization": "unknown",
"contextLength": 1050000,
"pricing": {
"prompt": 10,
"completion": 50
},
"supportedParameters": [
"reasoning",
"include_reasoning",
"seed",
"max_tokens",
"response_format",
"structured_outputs",
"tools",
"tool_choice",
"reasoning_effort"
],
"status": 0,
"modelReasoning": {
"mandatory": true,
"default_enabled": true,
"supported_efforts": [
"max",
"xhigh",
"high",
"medium",
"low"
],
"default_effort": "medium"
}
},
"reasoning": {
"enabled": true,
"effort": "medium"
},
"maxTokens": 12000,
"promptSha256": {
"system": "c06b6e107145044b393b6572258b42db175091e8f8bd8ecfec7070614a63ad27",
"schema": "9b22ce7692aa05c6d3e0f4a736e109e279c920244070884b46fe480184234d36"
},
"goldSha256": "0a7c39df1d485158ed48019533199d4473711892b327fc9e3c4a93929ab590ca",
"mode": "calibration",
"sample": {
"n": 20,
"seed": 42,
"questionIds": [
1,
13,
20,
21,
29,
36,
37,
39,
41,
44,
45,
53,
54,
64,
66,
79,
84,
86,
88,
94
],
"run": 1
},
"startedAt": "2026-09-16T04:30:14.052Z",
"finishedAt": "2026-09-16T04:52:13.159Z"
}
]Gold, judges and score construction
Gold is a fact sheet, not an ideal reference paragraph. Fable 5.1, GPT-6 Astra and Gemini 3.8 Flash independently extract source-linked facts at high reasoning effort. GPT-6 Astra adjudicates shuffled, anonymously labelled sheets at high effort. The recorded token caps are 16,000 for extraction and 20,000 for adjudication. The adjudicator is also one of the extractors. Two extractors are contestants. That overlap is disclosed and cannot be removed merely by anonymizing labels.
The fact sheets include salience, kind, support-row indices, extractor agreement, entities and known traps. There are 1,169 facts labelled as present in all three extraction sheets, 308 with two-way agreement and 170 singleton facts. The agreement labels are adjudicator-produced provenance, not a substitute for human verification.
Opus 5 and Gemini 3.8 Flash grade each valid candidate at medium effort with a 12,000-token cap. They see the source rows, gold and one candidate with its identity hidden. Both judges are themselves contestants. GPT-6 Astra supplies an additional judge view on a 20-question calibration subset. Repeated calibration grades reassess the same saved answers; they do not repeat candidate generation.
Coverage grades each gold fact as stated (1), partial (0.5), absent (0) or contradicted (0), considering the English and Nepali summaries together. Core recall averages credits over core facts. The “precision” component is credited gold coverage across all salience levels divided by credited coverage plus the count of unsupported-claim annotations. It is a coverage-based proxy, not conventional precision over an independently enumerated set of generated claims. Supporting recall is separately reported. Nepali and English prose are scored from 1 to 10; headline acceptance and approximate entity matching are additional diagnostics.
The composite is 40 × core recall + 30 × precision proxy + 1.5 × Nepali prose + 1.5 × English prose. Judges are averaged per answer. An invalid generation contributes zero to the effective score; valid-answer quality excludes it. One Qwen 27B answer lacks the Gemini grade, so its available Opus grade determines that cell. The denominator is visible rather than silently filled.
The publication analysis uses 5,000 paired question-bootstrap resamples for primary comparisons. New rank-frequency analysis uses 2,000 resamples. These are exploratory, unadjusted intervals conditional on the saved runs and judges. They describe variation across sampled questions; they do not isolate repeated-generation variance or account fully for related stories. The analysis has no established equivalence margin, so an interval crossing zero is not evidence that two systems are equivalent.
Terminology and metric definitions
- Story question / cluster
- A saved group of publisher rows about one underlying story. Related questions can still share events.
- Gold fact sheet
- A machine-generated set of atomic facts with salience, kind and support-row indices. It is a reference for source fidelity, not verified real-world truth.1
- Core / supporting / detail
- Facts essential to a faithful summary / usually useful / optional detail, as assigned by the gold process.
- Coverage credit
- Stated = 1; partial = 0.5; absent or contradicted = 0. A contradiction also remains visible as its own diagnostic.
- Core recall
- Average coverage credit on core facts; ranges from 0 to 1. Coverage is assessed jointly across the two summaries.
- Precision proxy
- Credited gold coverage across all salience levels ÷ (credited coverage + unsupported-claim annotation count). This is not the fraction of generated statements that are true.
- Effective / valid-answer quality
- Effective averages all 100 slots, assigning generation failures zero. Quality averages only valid answers using available judge grades.
- Nepali prose
- Judge score, 1–10, covering the Nepali headline and summary, naturalness, register, grammar and consistency with English. It is not Nepali-only fact recall.
- Unsupported annotation
- A judge’s flagged addition not supported by a supplied source row. Its segmentation and support threshold depend on the judge.
- Entity F1
- Per-answer harmonic mean of surface/alias- and type-matched entity precision and recall, then averaged. Gold entities may exceed those needed for a concise English summary.
- Paired bootstrap
- Resample the same question IDs for both models and recompute their mean difference. Intervals describe question variation conditional on saved answers and judges.
- P50 / P95 latency
- The median / 95th percentile of stored call durations. These do not reconstruct complete retry, queue or user-perceived latency.
- Pareto frontier
- Configurations for which no observed alternative is both cheaper and at least as good, or better and no more expensive, on the chosen point estimates.
Language anchors, headline checks and deterministic heuristics
Prose anchors: 9–10 clean wire copy; 7–8 clean with minor awkwardness; 5–6 readable but with clear faults or one cross-language inconsistency; 3–4 hard to read or several faults; 1–2 unusable. Judges are instructed not to reward length.
The actual headline rubric checks accuracy, the main event and clickbait; it does not enforce the planned 12-word cap. Correct Gregorian conversions may receive coverage credit; wrong conversions can be contradicted. If source rows disagree with gold, judges are instructed to prefer the rows and explain the conflict.
Script ratios, digit matching, normalized entity surfaces and cross-language number comparisons are deterministic diagnostics. They do not perform semantic calendar conversion or prove cross-language factual equivalence. The parser requires headline_en, summary_en and summary_ne; it normalizes the other requested fields. A valid answer therefore does not prove perfect six-field schema compliance.
Different criteria select different leaders
Fable has the highest observed failure-inclusive composite at 92.65. GPT-5.6 Sol follows at 91.61, Opus at 91.26 and DeepSeek Flash at 91.03. Fable's gain over Sol is 1.05 points, with a paired interval of −0.07 to 2.24. This comparison remains uncertain. By contrast, Fable's 4.71-point lead over Gemma has an interval of 2.72–7.32.
This is why “the top 15 are all the same” is an unsuitable interpretation. A chain of adjacent uncertain comparisons is not an equivalence test across a group. The full pairwise matrix supports more specific comparisons. In the new conditional question resampling, Fable occupies first place in 94.55% of 2,000 resamples; this is a sampling frequency under the frozen dataset, not a probability of universal superiority or a contradiction of the Fable–Sol interval.
Gemini 3.1 Pro has the highest observed Nepali prose mean, 9.025/10, while Fable leads the composite through its broader coverage profile. Fable's core recall is 95.55%, supporting recall 79.79% and entity recall 74.81%. Its observed cost is $8.896 per 100 retained attempts. Gemma records 93.14% precision proxy but only 47.60% supporting recall, demonstrating how a concise answer can avoid additions while omitting substantial source material.
Post-hoc weighting scenarios retain Fable first under equal components, a grounding-heavy score and a Nepali-emphasis score. A language-only score puts Sol first. These are illustrations of preference sensitivity, not alternative benchmark winners chosen after viewing results. The original score remains the primary outcome.
The smaller DeepSeek configuration wins this workload comparison
DeepSeek V4.1 Flash scores 91.03 versus 89.10 for V4 Pro 0813. The paired gain is 1.92 points, with an exploratory interval of 0.42–3.41. Its recorded generation cost is $0.0699 versus $0.2419 per 100 attempts, approximately 71.1% lower. Both runs use first-party DeepSeek through OpenRouter with reasoning disabled and JSON-object mode.
The experiment journal also records an unresolved possibility of upstream aliasing between the two DeepSeek endpoints. Their returned model IDs, outputs, prices and latency differ, but those records cannot establish that the underlying weights are different. This is a direct practical result for these two endpoints on this task. It does not identify model parameter count, establish that inexpensive models are generally better or predict performance on unrelated tasks. The observed cost-quality frontier contains Gemma, DeepSeek Flash, Sol and Fable. Membership is based on point estimates; nearby alternatives can move under new questions, different weights, provider pricing or repeated runs.
Nepali competence has several distinct meanings here
The Nepali prose score combines headline and summary language quality, including translation consistency. It is not a standalone Nepali factual-accuracy score. Coverage is graded jointly over both summaries. A fact appearing only in English can therefore receive coverage credit even if Nepali readers do not receive it. A future Nepali-only evaluation must grade fact coverage separately for each language.
The 77 Nepali-source questions assess consolidation from Nepali evidence with bilingual output. The six English-only questions add English-to-Nepali translation pressure. They are different stories, not matched translations of the same inputs, so their difference does not identify a causal language effect.
Tail rates are informative. Fable and Gemini 3.1 Pro have no valid outputs below 7/10 for averaged Nepali quality. DeepSeek Flash has two, Gemini 3 Flash one, Gemma four, Haiku 20 and Qwen 27B 28. These are observed counts, not precision estimates for all future news. Opus also has no low-scoring valid Nepali outputs but one generation failure, which must remain visible next to that statement.
Nepali-only unsupported flags and flags shared with English are reported separately. They are judge annotations. Different judges may segment or classify the same flawed sentence differently. A count of annotations is not a count of unique real-world falsehoods, and an answer flagged by both judges need not contain the same flagged claim under each judge.
| Model | NE mean /10 | Below 7 / valid | At least 9 / valid | Generation failures /100 | NE-only flags, mean |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | 9.025 | 0 / 100 | 66 / 100 | 0 | 0.035 |
| Claude Fable 5.1 | 8.995 | 0 / 100 | 71 / 100 | 0 | 0.055 |
| GPT-5.6 Sol | 8.950 | 4 / 100 | 72 / 100 | 0 | 0.070 |
| Claude Opus 5 | 8.879 | 0 / 99 | 62 / 99 | 1 | 0.066 |
| Kimi K3 | 8.862 | 2 / 98 | 64 / 98 | 2 | 0.071 |
| DeepSeek V4.1 Flash | 8.835 | 2 / 100 | 64 / 100 | 0 | 0.095 |
| DeepSeek V4 Pro | 8.825 | 1 / 100 | 56 / 100 | 0 | 0.035 |
| Gemini 3.8 Flash | 8.815 | 1 / 100 | 53 / 100 | 0 | 0.165 |
| Grok 4.6 | 8.755 | 3 / 100 | 59 / 100 | 0 | 0.045 |
| Qwen 3.8 Max | 8.722 | 3 / 99 | 51 / 99 | 1 | 0.076 |
| Gemini 3 Flash | 8.680 | 1 / 100 | 40 / 100 | 0 | 0.210 |
| Gemma 4 31B | 8.672 | 4 / 99 | 48 / 99 | 1 | 0.030 |
| GPT-5.6 Luna | 8.635 | 2 / 100 | 50 / 100 | 0 | 0.080 |
| GLM 5.3 | 8.625 | 4 / 100 | 52 / 100 | 0 | 0.090 |
| Claude Sonnet 5 | 8.556 | 3 / 99 | 42 / 99 | 1 | 0.136 |
| Claude Haiku 4.5 | 7.570 | 20 / 100 | 7 / 100 | 0 | 0.150 |
| Qwen 3.8 27B | 7.520 | 28 / 99 | 16 / 99 | 1 | 0.293 |
The mistakes behind the averages
These purposefully selected examples illustrate numeric, calendar, attribution and serving failures. They are source-grounding inspections and machine annotations, not a random error sample or a human validation panel.
A tenfold magnitude error
The source amount is 2.2 billion yuan. DeepSeek Flash renders it as 220 million in English and 22 crore in Nepali.
Nepali figures survive; English units do not
DeepSeek preserves the printed Nepali budget amounts, then translates kharba incorrectly into trillions in English. A bilingual average can conceal which readers receive the error.
A month changes between languages
DeepSeek’s Nepali month conflicts with the supplied source and its English answer. This is a language-specific error despite fluent wording.
Who owes whom an apology?
DeepSeek reverses the direction of an apology demand in Nepali and changes Gen Z into Janajati in English. Natural prose does not establish faithful attribution.
Preserving a date is better than guessing its conversion
Gemini 3 Flash adds an unsupported Gregorian date and misassigns the presenter credit. The source provides a Bikram Sambat date; calendar conversion introduces an avoidable failure mode.
Sparse evidence rewards restraint
Three headline-only rows say 197 savers from ten cooperatives received refunds. Gemini 3 Flash adds context about troubled cooperatives and government action that these rows do not supply.
Keep interpretation attributed
Gemini 3 Flash recasts analysts’ interpretation of Sheikh Hasina’s remarks as her own intention or appeal. Interpretation needs to remain attributed to its source.
Dense stories can exhaust the output budget
This question has 34 gold facts. Opus 5 and Sonnet 5 truncate during generation and score zero. Rich entity lists compete with bilingual prose for a finite token budget.
Calendar conversion needs separate language review
The saved Haiku answer is a calendar-error inspection case. Separately, Qwen 27B lacks its Gemini judge grade here; a missing grade is not a failed candidate answer.
Local units and an evolving loan ceiling
Luna’s headline says 50 lakh, its English summary says 500,000 rupees, and its Nepali summary says 50 thousand. The supplied rows include changing reported limits. This is a priority for native-Nepali review, not a human-confirmed error label.
English evidence, Nepali output
An English-only landslide story exposes translation quality. It belongs to a six-question slice, too small to support a general language-effect claim.
Different serving failures on one question
Kimi records a client/content-filter error; the two Qwen configurations record server errors. DeepSeek and GLM answer. These events do not support a nationality-wide refusal claim.
The public numerical release includes per-question scores and annotation counts. Full source excerpts, model outputs and judge comments remain in the private audit archive while distribution terms are reviewed.
Length and coverage: the direction changes with the comparison
Across the 17 model means, English summary length and supporting-fact recall correlate at r=0.966. Models that tend to write longer summaries capture a larger share of supporting facts. This is a useful description of the model configurations.
Within each of the 17 models, however, its longer answers across different questions have lower supporting-fact recall: correlations range from −0.517 to −0.095. On the 96 questions all models answer successfully, the within-model centered correlation is −0.273. Comparing models within the same question gives a positive centered association, r=0.690.
These patterns can coexist because stories differ in complexity and the number of facts available to recall. A dense story can elicit a longer summary while still leaving a larger fraction of its facts uncovered. The data do not support calling supporting recall “just a length proxy,” nor do they support prescribing longer output as a causal solution. A matched length-budget experiment would test that mechanism directly.
Longer-writing configurations retain more supporting facts.
Longer answers to different stories retain a smaller fraction.
Inspect all within-model correlations
| Model | Valid n | EN words vs supporting recall, Pearson r | EN words vs core recall, Pearson r |
|---|---|---|---|
| Claude Fable 5.1 | 100 | -0.470 | -0.462 |
| Claude Haiku 4.5 | 100 | -0.216 | -0.236 |
| Claude Opus 5 | 99 | -0.346 | -0.330 |
| Claude Sonnet 5 | 99 | -0.476 | -0.385 |
| DeepSeek V4.1 Flash | 100 | -0.517 | -0.256 |
| DeepSeek V4 Pro | 100 | -0.303 | -0.315 |
| Gemini 3.1 Pro | 100 | -0.300 | -0.358 |
| Gemini 3.8 Flash | 100 | -0.417 | -0.302 |
| Gemini 3 Flash | 100 | -0.246 | -0.162 |
| Gemma 4 31B | 99 | -0.180 | -0.245 |
| GLM 5.3 | 100 | -0.325 | -0.393 |
| GPT-5.6 Luna | 100 | -0.253 | -0.410 |
| GPT-5.6 Sol | 100 | -0.295 | -0.356 |
| Grok 4.6 | 100 | -0.156 | -0.175 |
| Kimi K3 | 98 | -0.424 | -0.331 |
| Qwen 3.8 27B | 99 | -0.095 | -0.262 |
| Qwen 3.8 Max | 99 | -0.292 | -0.397 |
Judges agree about many facts and disagree about acceptable additions
Repeated coverage-label agreement on the calibration subset is 96.36% for Opus and 95.61% for Gemini. Newly computed unweighted Cohen's kappa is 0.939 and 0.925 respectively. Full-pass exact coverage agreement is 91.56%, with kappa 0.856. Labels are clustered within answers, and high consistency does not establish human correctness.
The judges' effective model rankings correlate at 0.946, but Nepali-quality rankings correlate less strongly at 0.807. Their score levels also differ. Opus records approximately 1.474 unsupported annotations per valid answer, versus Gemini's 0.413: a 3.57-fold difference. Gemini's average composite grade is about 5.48 points higher. The judges' thresholds for supported paraphrase and additions therefore materially affect score levels even when many fact labels agree.
Both judge views are available in the explorer. Averaging is a useful summary, not a way to erase disagreement. The calibration sample includes a non-contestant judge but is too small to establish the absence of family bias.
Unsupported-claim annotations per answer.
Different support thresholds and claim segmentation.
| Judge comparison | Matched answers | Label comparisons | Exact agreement | Kappa | Rank Spearman ρ |
|---|---|---|---|---|---|
| claude-opus-5-r1vs claude-opus-5-r2 | 338 | 5714 | 96.4% | 0.939 | 0.966 |
| gemini-3-8-flash-r1vs gemini-3-8-flash-r2 | 338 | 5712 | 95.6% | 0.925 | 0.958 |
| claude-opus-5-r1vs gemini-3-8-flash-r1 | 338 | 5712 | 92.2% | 0.869 | 0.941 |
| claude-opus-5-r1vs gpt-6-astra-r1 | 337 | 5695 | 89.7% | 0.835 | 0.860 |
| gemini-3-8-flash-r1vs gpt-6-astra-r1 | 337 | 5695 | 89.4% | 0.828 | 0.833 |
Calibration gates required ≥0.80 intra-judge label agreement and ≥0.70 inter-judge model-rank correlation; the completed records meet both. Astra has no second repeat and one missing calibration grade. Full-pass agreement uses 1,692 matched answers and 27,822 label comparisons; its valid-grade rank correlation (0.949) differs from the failure-inclusive rank correlation (0.946).
Reliability and serving failures
Seven of 1,700 generation slots fail: Opus and Sonnet truncate on Q21, Gemma truncates on Q30, Kimi has a client error on Q4 and a schema violation on Q74, and both Qwen configurations have server errors on Q4. The Q4 grouping is not evidence of a nationality-wide refusal policy. The recorded Kimi client error and Alibaba server errors are different failure classes; DeepSeek and Z.AI answer the same question.
One hundred attempts per model is too small to certify a service-level failure rate, especially for rare outages. A 100/100 observed success count is not proof of universal reliability. The explorer reports both the failure-inclusive score and valid-answer quality.
| Configuration | Question | Recorded class | Retained attempts field |
|---|---|---|---|
| Claude Opus 5 | Q21 | truncated | 1 |
| Claude Sonnet 5 | Q21 | truncated | 1 |
| Gemma 4 31B | Q30 | truncated | 1 |
| Kimi K3 | Q4 | client_error | 1 |
| Kimi K3 | Q74 | schema_violation | 1 |
| Qwen 3.8 27B | Q4 | server_error | 3 |
| Qwen 3.8 Max | Q4 | server_error | 3 |
Audit, uncertainty and remaining work
Independent reconstruction matches saved scores to a maximum numerical difference below 3×10⁻¹⁴. Prompt hashes match, and the full-pass judge files reference the same gold hash. Twenty Opus grade records have extra, duplicate or missing fact IDs. Sanitizing those lists changes model means by less than 0.007 points and does not alter the main conclusions. The original and sanitized analyses are both retained; the original raw files are not edited.
The protocol and plans contain stale descriptions, including a 16-model count followed by 17 names, generic strict-schema settings, and intended human review preceding the full pass. The status of the planned human review is recorded in the gold-review footnote. The package's methods ledger resolves the differences using saved run headers, implementation and completed records.
The main unresolved limitations are machine-generated gold, judge/contestant overlap, one candidate generation per slot, imperfect independence of stories, a selected rather than random corpus, unequal decoding constraints and incomplete accounting for overwritten attempts. Pairwise intervals and slice analyses are exploratory and unadjusted for multiple comparisons. The old benchmark's run-to-run noise figure should not be imported as a v2 significance threshold.
Reproduction has three meanings here: recomputing scores from frozen records is supported offline; rerunning the models requires paid services that may change; rebuilding the full production selection requires the original database context. The resources section distinguishes what the numerical download supports from what requires the private audit snapshot.
Website preparation independently ran the supplied offline reconstruction and verifier: 258 checks passed. The maximum score difference was 2.842 × 10−14. The paired interval tables are imported only after their input hashes verify; the separate archived bootstrap programs can recompute them.
Protocol amendments and implementation deviations
- Provider pinning, returned model/provider checks and recorded usage cost replaced looser early transport handling. Earlier agent runs are excluded.
- JSON-object mode replaces strict schema on four endpoints. Six mandatory-reasoning runs receive 4,096 total output tokens, versus 1,536 elsewhere.
- Transport errors can be retried; provider finish_reason=error was added to that class. Kimi Q74/Q87 were accidentally resampled after schema failures before the resume policy was corrected.
- Gemini 3.8 Flash replaced Gemini 3.1 Pro as the third gold extractor. Gemini 3 Flash was added to the candidate lineup. Final count: 17 configurations.
- The planned human review is described in the gold-review footnote.
- Judge truncation retries were added without changing the rubric. One Gemini full-pass grade remains missing.
- The original scorer used 1,000 bootstrap draws; publication pairwise and focused comparisons use 5,000; post-hoc rank frequencies use 2,000.
- Twenty Opus grades have extra, duplicate or missing fact IDs. The sanitation sensitivity enumerates gold IDs, ignores extras, uses the lowest duplicate credit and assigns missing IDs zero.
- Gold-consensus, weighting, rank-frequency, kappa and within-model length analyses are post-hoc. Their original weights do not change.
The methods ledger supplies the complete planned-versus-completed comparison. Original protocol files use “pre-registered” internally; no public registration is established. The journal’s early “top 15 are one block” and causal reasoning/length interpretations are superseded by the paired and within-model analyses reported here.
What to do next
Complete the planned blinded gold review and add native-Nepali review of numeric, date and attribution examples. Correct gold with a new version and regrade affected questions under documented hashes. Repeat matched candidate generations to estimate generation variance. Grade fact coverage separately for Nepali and English. Expand matched evaluation to longer clusters and more mixed-language stories, keeping schema and retry policies explicit. Run matched length and reasoning ablations before making causal claims about those settings.
These follow-ups address uncertainties exposed by the experiment. They need not obscure the present findings: coverage, prose quality, cost and severe mistakes vary in identifiable, inspectable ways across the tested configurations.
Full results and experiment accounting
All 17 configurations, both full-pass judges. Effective score uses all 100 slots; other quality metrics use valid answers. Supporting recall and entity F1 are diagnostics outside the composite.
| Configuration | Effective /100 | Valid quality /100 | Core % | Precision proxy % | Supporting % | NE /10 | EN /10 | Entity F1 % |
|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 92.651 | 92.651 | 95.6% | 91.6% | 79.8% | 8.995 | 8.970 | 79.7% |
| GPT-5.6 Sol | 91.605 | 91.605 | 90.2% | 94.8% | 56.9% | 8.950 | 9.120 | 67.7% |
| Claude Opus 5 | 91.260 | 92.182 | 95.5% | 91.1% | 83.1% | 8.879 | 8.909 | 81.3% |
| DeepSeek V4.1 Flash | 91.026 | 91.026 | 92.3% | 92.8% | 69.8% | 8.835 | 8.675 | 75.5% |
| Grok 4.6 | 90.934 | 90.934 | 91.0% | 93.8% | 60.6% | 8.755 | 8.835 | 73.6% |
| GLM 5.3 | 90.805 | 90.805 | 93.0% | 91.9% | 68.2% | 8.625 | 8.740 | 72.8% |
| Gemini 3.1 Pro | 90.028 | 90.028 | 87.4% | 93.8% | 52.4% | 9.025 | 8.935 | 67.6% |
| GPT-5.6 Luna | 89.802 | 89.802 | 89.7% | 93.1% | 60.8% | 8.635 | 8.705 | 68.9% |
| Gemini 3.8 Flash | 89.500 | 89.500 | 89.9% | 90.5% | 58.7% | 8.815 | 8.780 | 67.7% |
| Kimi K3 | 89.428 | 91.253 | 93.1% | 91.5% | 74.0% | 8.862 | 8.852 | 75.8% |
| DeepSeek V4 Pro | 89.103 | 89.103 | 88.0% | 92.2% | 55.7% | 8.825 | 8.670 | 69.0% |
| Qwen 3.8 Max | 88.583 | 89.477 | 90.6% | 90.8% | 63.8% | 8.722 | 8.606 | 68.9% |
| Gemini 3 Flash | 88.455 | 88.455 | 90.0% | 88.5% | 61.1% | 8.680 | 8.570 | 69.3% |
| Claude Sonnet 5 | 88.413 | 89.307 | 92.3% | 89.6% | 74.2% | 8.556 | 8.465 | 74.5% |
| Gemma 4 31B | 87.938 | 88.826 | 86.8% | 93.1% | 47.6% | 8.672 | 8.763 | 64.3% |
| Claude Haiku 4.5 | 84.834 | 84.834 | 88.0% | 88.2% | 59.9% | 7.570 | 7.885 | 53.5% |
| Qwen 3.8 27B | 81.570 | 82.394 | 83.8% | 85.9% | 53.8% | 7.520 | 7.884 | 59.1% |
Fact kinds and annotation types
Inspect what each model covers and where judges flag additions. These diagnostics sit outside the primary composite and do not measure error prevalence in future traffic.
Token usage and evaluation cost
Retained records report 55,183,774 tokens: 44,274,103 input and 10,909,671 output tokens across 7,175 calls with usage, out of 7,176 saved calls. Reasoning is a subset of reported completion tokens and is not added again. Provider tokenizers differ. Overwritten retries, smoke tests and superseded runs are excluded; missing usage is not imputed. Download the token accounting by artifact.
Candidate generation: $23.36. Gold, calibration and full grading: $354.65. Combined retained-record total: $378.01. This is not the complete account bill.
| Stage / artifact | Saved calls | OK | Recorded USD | Missing cost |
|---|---|---|---|---|
| extractiongold/extract/claude-fable-5-1.json | 100 | 100 | $25.5103 | 0 |
| extractiongold/extract/gemini-3-8-flash.json | 100 | 100 | $2.0400 | 0 |
| extractiongold/extract/gpt-6-astra.json | 100 | 100 | $28.1485 | 0 |
| adjudicationgold/gold.v2.json | 100 | 100 | $54.0862 | 0 |
| full-judgegrades/claude-opus-5.json | 1693 | 1693 | $131.0674 | 0 |
| full-judgegrades/gemini-3-8-flash.json | 1693 | 1692 | $19.0362 | 0 |
| calibrationgrades/calibration/claude-opus-5-r1.json | 338 | 338 | $26.5320 | 0 |
| calibrationgrades/calibration/claude-opus-5-r2.json | 338 | 338 | $26.4717 | 0 |
| calibrationgrades/calibration/gemini-3-8-flash-r1.json | 338 | 338 | $3.8027 | 0 |
| calibrationgrades/calibration/gemini-3-8-flash-r2.json | 338 | 338 | $3.8433 | 0 |
| calibrationgrades/calibration/gpt-6-astra-r1.json | 338 | 337 | $34.1139 | 1 |
Data, definitions and reproducibility
Explore the numbers behind the paper.
Version: publication-handoff-1 · experiment 2.0-draft. Forty-one files: numerical tables, model metadata, schemas, metric definitions and editable figure specifications. Excludes raw API bodies and source excerpts.
Download numerical data · 370 KiB378,929 bytes · SHA-256
0f36167f7531f05e1a7f41976ca4afb71acdc55870aecf382da739c2ddf4847e- Data dictionary and integration contract — units, denominators, joins and failure policies.
- Methods and deviations ledger — intended settings versus completed runs.
- Claim-to-evidence ledger — numerical claims, source tables and interpretation limits.
- Handoff verification record — checks performed on the frozen inputs.
- Per-question numerical scores, paired comparisons, fact-kind recall, annotation counts and grade audit.
The separate audit archive contains the frozen source rows, raw outputs, gold, grades, prompts and offline reproducer. It remains private pending distribution terms. The numerical ZIP alone supports numerical inspection, not independent reconstruction from raw grades. No new release license or DOI has been assigned. Public source article URLs cannot be reconstructed from the saved article IDs alone.
Three meanings of reproduction
- Recompute frozen scores: supported offline by the private audit snapshot and standard-library Python reproducer; no paid calls.
- Regenerate model answers: a separate paid experiment, with changing endpoints and nondeterministic remote sampling.
- Rebuild the original selection: requires the original eligibility database and historical ranking context. The selected-row construction is preserved, but the full pool is not.
Draft citation and version history
Working attribution: Outback Yak Research; individual author credit and final release date remain to be confirmed. The prepared route is /whitepapers/nepnewsbench-v2/; this preview is not evidence of publication.
Outback Yak Research. NepNewsBench v2: what models retain, invent and mistranslate. Research draft, experiment 2.0-draft, September 2026.
Read the previous NepNewsCluster v0.3 paper →
17 September 2026: website review draft; frozen handoff results retained; Kimi resampling and the unresolved DeepSeek aliasing caveat added from the original journal. Legacy PDF version labels checked against its contents. No benchmark generations or paid grades were run for this page.
Lab logos: Lobe Icons · asset sources and color key · MIT license. Lab colors identify model developers; serving providers are listed separately.
Questions or collaboration: research@kchakhabar.com. Dataset reporting originates from the publishers listed above; the benchmark evaluates fidelity to supplied excerpts, not independent verification of their reporting.