Can Local Models Build Reliable Hotel Source Packages?

Cross-hotel source-acquisition qualification benchmark

Artifact links: full per-run ledger and retained-source report.

Result

Remote Gemma 4 31B produced 9 extraction-ready packages from 12 hotels; remote Qwen 3.5 9B produced 4. Gemma won 11 hotel-level scores, Qwen won 0, and 1 tied.

The bounded model stages generalized: both models returned valid query plans and strict source selections for all 12 hotels. Local Playwright was still blocked by many protected official and review sites, but a bounded Firecrawl fallback recovered all 19 selected Tripadvisor pages and 6 of 8 failed Hilton pages. That raised extraction-ready packages from 4 to 9 for Gemma and from 0 to 4 for Qwen. Blocked runs still emitted validated fail-closed packages rather than weak evidence.

Aggregate grid

Metric Remote Gemma 4 31B Remote Qwen 3.5 9B
Valid query plans 12 12
Valid selections 12 12
Correct canonical official selections 12 9
Usable canonical official pages 11 8
Extraction-ready packages 9 4
Mean score 86.33 63.08
Median score 90.0 68.0
Median valid guest families 4.0 3.0
Median usable guest families 3.0 2.5
Median query-plan seconds 35.06 12.2
Median selection seconds 47.53 10.91
Total model seconds 990.44 275.73

Hotel-level grid

Hotel Category Gemma Qwen Better score
Omni Mount Washington Resort destination_resort 98.0 (ready) 88.0 (ready) Gemma
Disney’s Grand Californian Hotel & Spa destination_resort 76.0 (blocked) 70.0 (blocked) Gemma
Hotel del Coronado destination_resort 98.0 (ready) 66.0 (blocked) Gemma
Grand Wailea, A Waldorf Astoria Resort destination_resort 76.0 (blocked) 76.0 (blocked) Tie
The Chanler at Cliff Walk notable_independent_boutique 88.0 (ready) 87.0 (ready) Gemma
The Peabody Memphis notable_independent_boutique 98.0 (ready) 22.0 (blocked) Gemma
The Hermitage Hotel notable_independent_boutique 92.0 (ready) 38.0 (blocked) Gemma
The Mission Inn Hotel & Spa notable_independent_boutique 87.0 (ready) 28.0 (blocked) Gemma
Hampton Inn & Suites Madison-West ordinary_chain 62.0 (blocked) 51.0 (blocked) Gemma
Home2 Suites by Hilton Erie ordinary_chain 92.0 (ready) 88.0 (ready) Gemma
La Quinta Inn & Suites by Wyndham Flagstaff ordinary_chain 77.0 (ready) 55.0 (blocked) Gemma
Courtyard by Marriott Albuquerque Airport ordinary_chain 92.0 (ready) 88.0 (ready) Gemma

What the artifacts prove

Every one of the 24 runs ends in a versioned frozen package with local paths, hashes, model provenance, exact queries, selected IDs, rejected-source disclosures, and a controller decision. Ready packages can be consumed without web access. Blocked packages are also reliable artifacts: they state why extraction must not start.

This benchmark does not validate quotation-backed claim extraction. It validates the Step 1 handoff contract and demonstrates that retrieval—not model JSON compliance—is now the principal bottleneck.

Scale

  • 12 hotels selected from a frozen 24-hotel candidate pool with seed oddbyte-hotel-source-generalization-2026-08-26-v1
  • 24 model runs
  • 164 unique controller-executed searches; 0 search errors
  • 136 unique selected hotel/URL pairs attempted
  • 119 usable extracted pages after local rendering and fallback recovery
  • 25 of 27 bounded Firecrawl fallback requests succeeded
  • One calibration hotel per category, followed by nine unchanged qualification hotels

Decision

Gemma 4 31B is the stronger source-planning and selection model for the current controller. With local Playwright first and Firecrawl only for failed protected pages, Gemma produced extraction-ready packages for 9 of 12 hotels. The remaining three require named gap-fill or official-page recovery, so arbitrary hotels should still fail closed rather than run unattended. Qwen 3.5 9B remains much faster but produced only 4 ready packages. The next engineering task is a bounded gap-fill pass for the three blocked Gemma hotels, followed by local quotation-backed extraction on the nine ready packages—not a broader model sweep.

Similar Posts

One Comment

Comments are closed.