Warwick Denver Local-Flair Experiment: Models, Failures, and Scores

Experiment report: This page documents a deliberately non-production local-model writing trial, including expected failures. It is not travel advice.

Read the published autonomous Warwick Denver candidate.

Outcome

The experiment produced a 1,138-word, fully local Warwick Denver article. It passed the final factual audit but failed the predeclared style gate. Michael explicitly authorized publishing it as an intentionally imperfect capability-test artifact because OddByte is an experimental website, not a production travel-advice publication.

  • Final factual audit: {"findings":[],"pass":true}
  • Final style scores: {"voice":6,"usefulness":9,"structure":8,"non_promotional":9,"rhythm":5,"traveler_fit":9,"factual_restraint":9,"entity_separation":10}
  • Style gate: failed (voice and rhythm below threshold)
  • Cloud prose calls: 0

Why Warwick

Warwick was selected instead of reusing Bar Harbor to test generalization on a new property, source set, comparison set, map context, and image set. This is a harder test than an apples-to-apples Bar Harbor rewrite.

Evidence and images

  • Frozen ledger: 182 literal-evidence claims; 153 approved for drafting.
  • Historical independent operational claims were excluded from the 2026 drafting contract.
  • Deterministic pedestrian routes were retrieved through Valhalla and labeled as dated estimates.
  • Four Wikimedia Commons images were downloaded, hashed, visually reviewed, and documented with creator, license, dimensions, truthful alt text, and source page.

Writers’ room

  1. Qwen3.8 distilled the house voice and extracted the evidence ledger.
  2. Qwen3.8, Mistral Small 24B, and Gemma 4 26B generated independent candidates.
  3. Qwen3.6 blindly selected Mistral’s first candidate. Initial Mistral scores were voice 8, usefulness 9, structure 9, non-promotional tone 9, rhythm 8, traveler fit 9.
  4. Mistral completed two further revision passes. Qwen3.8 supplied a bounded line-edit plan between them.
  5. Qwen3.6 performed the final factual and style gates.

The stricter final judge scored voice 6 and rhythm 5. Its diagnosis was consistent with direct inspection: the article remains useful, accurate, and decisive about fit, but still relies too heavily on specification lists and repetitive sentence shapes.

Transparent local usage

  • Calls: 23 (3 failed/malformed attempts, all recovered or superseded)
  • Input tokens: 88,766
  • Output tokens: 48,235
  • Total tokens: 137,001
  • Aggregate model-call wall time: 5681.394 seconds (94.69 minutes)

By model

Model Calls Input Output Total Wall seconds Outcomes
qwen3.8:27b-mlx 15 40,147 39,204 79,351 3721.133 {“accepted”:12,”malformed”:3}
mistral-small:24b 3 17,868 5,196 23,064 1062.310 {“accepted”:3}
gemma4:26b-mlx 1 3,991 2,206 6,197 91.132 {“accepted”:1}
qwen3.6:27b 4 26,760 1,629 28,389 806.818 {“accepted”:4}

Failure and retry accounting

Three Qwen3.8 stages were recorded as malformed:

  • one oversized itinerary extraction hit its output cap and was recovered by splitting the packet;
  • one oversized competitor extraction hit its output cap and was recovered by splitting by hotel;
  • the first factual-draft call was mistakenly invoked in JSON mode for Markdown output and was rerun in text mode.

No failure was silently converted into a success, and the failed calls remain in the metrics totals.

Decision

The zero-cloud local phase reached the agreed three-cycle limit. The artifact did not satisfy the publication style threshold, but it is now public as a clearly labeled experimental result alongside this report. No cloud prose pass was performed before publication. A future compact cloud comparison, if authorized, must remain a separate, validated artifact.

Similar Posts