Warwick Denver Local-Flair Experiment: Models, Failures, and Scores
Read the published autonomous Warwick Denver candidate.
Outcome
The experiment produced a 1,138-word, fully local Warwick Denver article. It passed the final factual audit but failed the predeclared style gate. Michael explicitly authorized publishing it as an intentionally imperfect capability-test artifact because OddByte is an experimental website, not a production travel-advice publication.
- Final factual audit:
{"findings":[],"pass":true} - Final style scores:
{"voice":6,"usefulness":9,"structure":8,"non_promotional":9,"rhythm":5,"traveler_fit":9,"factual_restraint":9,"entity_separation":10} - Style gate: failed (voice and rhythm below threshold)
- Cloud prose calls: 0
Why Warwick
Warwick was selected instead of reusing Bar Harbor to test generalization on a new property, source set, comparison set, map context, and image set. This is a harder test than an apples-to-apples Bar Harbor rewrite.
Evidence and images
- Frozen ledger: 182 literal-evidence claims; 153 approved for drafting.
- Historical independent operational claims were excluded from the 2026 drafting contract.
- Deterministic pedestrian routes were retrieved through Valhalla and labeled as dated estimates.
- Four Wikimedia Commons images were downloaded, hashed, visually reviewed, and documented with creator, license, dimensions, truthful alt text, and source page.
Writers’ room
- Qwen3.8 distilled the house voice and extracted the evidence ledger.
- Qwen3.8, Mistral Small 24B, and Gemma 4 26B generated independent candidates.
- Qwen3.6 blindly selected Mistral’s first candidate. Initial Mistral scores were voice 8, usefulness 9, structure 9, non-promotional tone 9, rhythm 8, traveler fit 9.
- Mistral completed two further revision passes. Qwen3.8 supplied a bounded line-edit plan between them.
- Qwen3.6 performed the final factual and style gates.
The stricter final judge scored voice 6 and rhythm 5. Its diagnosis was consistent with direct inspection: the article remains useful, accurate, and decisive about fit, but still relies too heavily on specification lists and repetitive sentence shapes.
Transparent local usage
- Calls: 23 (3 failed/malformed attempts, all recovered or superseded)
- Input tokens: 88,766
- Output tokens: 48,235
- Total tokens: 137,001
- Aggregate model-call wall time: 5681.394 seconds (94.69 minutes)
By model
| Model | Calls | Input | Output | Total | Wall seconds | Outcomes |
|---|---|---|---|---|---|---|
| qwen3.8:27b-mlx | 15 | 40,147 | 39,204 | 79,351 | 3721.133 | {“accepted”:12,”malformed”:3} |
| mistral-small:24b | 3 | 17,868 | 5,196 | 23,064 | 1062.310 | {“accepted”:3} |
| gemma4:26b-mlx | 1 | 3,991 | 2,206 | 6,197 | 91.132 | {“accepted”:1} |
| qwen3.6:27b | 4 | 26,760 | 1,629 | 28,389 | 806.818 | {“accepted”:4} |
Failure and retry accounting
Three Qwen3.8 stages were recorded as malformed:
- one oversized itinerary extraction hit its output cap and was recovered by splitting the packet;
- one oversized competitor extraction hit its output cap and was recovered by splitting by hotel;
- the first factual-draft call was mistakenly invoked in JSON mode for Markdown output and was rerun in text mode.
No failure was silently converted into a success, and the failed calls remain in the metrics totals.
Decision
The zero-cloud local phase reached the agreed three-cycle limit. The artifact did not satisfy the publication style threshold, but it is now public as a clearly labeled experimental result alongside this report. No cloud prose pass was performed before publication. A future compact cloud comparison, if authorized, must remain a separate, validated artifact.
