Distributed Hotel and Restaurant Workflow Qualification: Final Report

Sanitized final qualification report: this is an engineering report, not a hotel or restaurant review and not a claim of firsthand experience.

Executive outcome

The three-hour distributed hotel/restaurant qualification ended safely and fail-closed. The hotel path validated Mac source acquisition and Alibaba evidence extraction, then reached a bounded LAN scoring block after its single permitted sealed-source recovery. The restaurant path validated source acquisition, quotation-backed extraction, and scoring readiness, then stopped after two changed-model composition attempts failed the same immutable 600-word minimum. No hotel review and no restaurant review was published.

Hotel workflow: Disney’s Coronado Springs Resort

  • Step 1 passed: Mac Qwen 3.5 9B produced a hash-validated frozen source package and handoff in 50.30 seconds of model time.
  • Step 2 passed: Alibaba Qwen 3.8 Flash produced 54 accepted literal-quotation claims across all six required evidence types. It used 11 cloud calls and 130.71 seconds of model time.
  • Step 3 stopped correctly: LAN Qwen 3.8 used zero cloud calls. Initial scoring used 229.39 seconds; the one permitted sealed-source recovery used 181.08 seconds and added 18 accepted claims; the terminal scoring pass used 259.78 seconds.
  • Terminal state: not_scored, editorial readiness blocked. Blocking criteria were room comfort/cleanliness, service/hospitality, and amenities/facilities. The pipeline did not convert missing support into favorable assumptions.

The earlier Disney Coronado milestone remains available at Distributed Hotel Workflow: Disney Coronado Canary Result.

Restaurant workflow: 1902

  • Evidence stages passed: three frozen source families yielded 40 accepted quotation-backed claims.
  • Extraction timing: five local/LAN model calls, two bounded retries, 339.49 seconds wall time, 47,941 input tokens and 8,172 output tokens.
  • Scoring timing: one LAN call, no retry, 145.74 seconds wall time; the scorecard reached ready_for_composition.
  • Composition attempt 1: Mistral Small 24B drafted and revised locally; measured draft, critic, and revision calls took 342.12, 62.93, and 352.10 seconds. The final body was 525 words and failed the frozen 600–1,000-word gate.
  • Composition attempt 2: Qwen 3.5 9B produced 543 words and failed the same immutable minimum.
  • Terminal state: blocked after two attempts with the same root cause. The bounded retry ladder stopped; no critic or publisher was allowed to waive the length contract.

Failed and superseded lineages

  • The earlier hotel packet lineage that encountered serialized Vaultwarden/empty-password control-path failures was superseded by a fresh workspace-bound run using a parent-scoped credential session. The failed artifacts were preserved rather than rewritten.
  • The hotel Step 3 initial attempt was followed by exactly one changed-condition sealed-source recovery. Its replacement still lacked sufficient criterion-specific evidence, so that root lineage terminated blocked rather than entering another local or cloud retry.
  • The restaurant Mistral writer lineage failed the minimum-length gate. A changed-model Qwen writer lineage reproduced the same root failure; the two-same-root-cause rule then terminated composition.

Code, checks, and accounting

  • Commits: 82e2ea2 (reuse cloud credential across Step 2 packets) and ce8625f (serialize Vaultwarden credential sessions).
  • Verification gate: 297 tests passed, with Ruff and diff/whitespace checks also passing.
  • Final mission Alibaba modeled spend: $0.02242878.
  • Local and LAN Ollama work had zero modeled cloud-model spend.
  • Firecrawl was called twice for one recovered Tripadvisor source after a local harness-directory mistake. The exact Firecrawl dollar cost is unavailable; it is not estimated here.

Production-promotion decision

Do not promote the full hotel-to-publication or restaurant-to-publication workflow to unattended production yet. Narrow stages are qualified: hotel distributed routing through bounded scoring, credential serialization, restaurant source acquisition, evidence extraction, and scoring readiness. Hotel editorial readiness and restaurant composition remain unqualified. Publication controls passed by refusing to publish either review.

Exact next steps

  1. Add a bounded, pre-scoring evidence-acquisition branch for the three hotel blockers: room condition/cleanliness, service interactions, and facility condition/seasonality. Keep property identity, quote, freshness, and hash gates unchanged.
  2. Re-run the hotel from a fresh sealed handoff only after those new evidence artifacts validate; allow one scoring attempt and the existing single changed-condition recovery, with no gate weakening.
  3. Correct the restaurant composition contract or prompt so a first draft deterministically targets 600–1,000 body words before critique. Add a preflight token/section budget and preserve the hard minimum.
  4. Repeat restaurant composition on the same frozen evidence canary, then run the same frozen contract on a second unrelated venue to establish portability.
  5. Fix the Firecrawl harness working-directory preflight and add an assertion that prevents duplicate fallback calls for one source; retain explicit provider-cost telemetry.
  6. Promote end-to-end publication only after a hotel and restaurant each pass final factual revalidation, package hashing, restricted publication, and independent public read-back.

Final disposition: bounded qualification complete; useful subpaths validated; both review publications correctly withheld.

Similar Posts