Skip to main content
At a glanceEval Labs is Lucia’s quality-control system. It exists to test whether Lucia is useful, truthful, calm, semantically aware, and operationally correct before behavior becomes trusted.

Why Evaluation Exists

Lucia is not evaluated only on whether an answer is “correct.” Lucia must be evaluated on whether the answer:

Evaluation Is Product Infrastructure

Eval Labs is not a side tool. Eval Labs is part of Lucia’s intelligence stack.
Lucia behaviorEval Labsreviewrefinementsafer behavior

The gates on 2026-09-10

The standing gates for a change to Lucia’s behaviour, as they stood on 2026-09-10. Each is dated by the record that built it; none is a claim that the platform is launch-ready. Battery discipline across the fleet: green means exit code 0, never a count of failure lines; integration is by literal 40-character SHA; the Engine runs its battery on Node 26 and Node 22 (ALGO-40; company gate law as recorded on HLLC-41). Live testing of the served surfaces is recorded on a dated founder-walkthrough ledger that is internal and not published; this Library names it, not its contents.

Evaluation posture — 2026-09-10, with the 2026-08-12 diagnostic as history

The Engine evaluation bank contains 217 cases. On 2026-08-12 an exact-SHA, non-strict diagnostic against Engine commit 9a15b1da… returned 76/217; on 2026-09-03 two runs on Development read 174/217 and 192/217 (ALGO-36). None of these is a pass, a release gate, a quality baseline, or a production-readiness claim. The bank’s standing failures are open work, and the September gate is the replay harness plus identical failing sets, as in the table above. Eval Labs Production evaluates operator-facing Lucia against Engine Development at:
The canonical Engine route remains:
Eval Labs itself is a production web app at https://evaluationlabs.ai. Its release identity verified on 2026-08-12 is commit 6cbfd15f75330c414bfe79e0d2fab51ef8115102, version 0.1.0, with same-deploy runtime identity evidence. Guest verification separately targets Guest Agent Production at https://guest.hellolucia.ai. Fieldwork evaluation relays remain Development-only.
Engine and Guest Agent own their own model calls; Eval Labs owns orchestration, evidence persistence, human review, and controlled evaluation contracts. A configured or requested model name alone does not prove execution. Since 2026-09-04 the Engine’s model layer is two providers behind one gateway: the provider follows the model id, Focus composition runs on Anthropic’s Fable on Development, and the Fieldwork conversation stays on GPT (Fable is Lucia’s voice, ALGO-39; Focus answers from one brain, ALGO-40). The Guest Agent’s own path is unchanged since 2026-07-18. Eval Labs itself has had no release since 2026-07-29. The Guest Agent claim above is limited to its accepted 2026-07-18 same-deploy/provider receipt, which recorded configured and provider-resolved gpt-5.6-sol for one production request. It does not establish later Guest Agent execution.
Neither the 2026-08-12 refresh nor this 2026-09-10 pass made a fresh Engine, Guest Agent, or Fieldwork model call, inspected authenticated production counts, ran the bank, or executed a Fieldwork cohort. Bank figures above are quoted from their dated records.

Historical checkpoints

  • The 2026-05-29 178/178 checkpoint belongs to an older 178-case bank and revision, not the current 217-case baseline.
  • The May 2026 Eval Labs gate recorded 60 runs and 3,000 prompts, items, responses, and reviews. It is historical platform-lifecycle evidence, not a current count or human approval.
  • The 2026-07-10 owner verification observed gpt-5.6-sol configured and provider-resolved for one Engine Development request. It does not prove unrelated runs.
  • The accepted 2026-07-18 Guest Agent same-deploy/provider receipt proved configured and provider-resolved gpt-5.6-sol for source fbf77b662511cf004ee1f46787e05773a8a78d71 and one production request. No fresh August Guest Agent model call was made.
Eval Labs has an implemented Guest Verification Check for deterministic booked-guest scenarios. The broader first-class guest-facing battery remains requested/future, and a recorded check result still does not establish launch readiness without the applicable gate. Payment truth now requires dedicated financial-attention coverage before policy-aware payment judgment is treated as stable. Since 2026-09-03 the corpus the replay harness and the Focus bank read is the living development world: a synthetic Villa Valentin played on the real clock. See lucia-world.

Guest-facing evaluation status

Implemented: Guest Verification Check

The Guest Facing Agent Verification Check is an implemented Eval Labs surface separate from operator-facing Lucia evals. It targets Guest Agent Production at https://guest.hellolucia.ai, runs the deterministic booked-guest verification scenario pack, and exposes check results. It does not establish that the broader battery below exists or passes.

Requested/future: broader Guest-facing battery

The broader first-class Guest-Facing Lucia Eval Track remains Requested with capability status future. Purpose:
Required validation coverage includes:

Primary Evaluation Targets

1. Intent Accuracy

Did Lucia understand what the operator was asking? Examples:

2. Operational Usefulness

Did Lucia identify a useful next move? A technically accurate answer can still fail if it leaves the operator with too much work. Current booking-spine usefulness must be tested against arrivals, departures, stay windows, and the difference between Full Booking Page review and Dynamic Action Workspace completion. Current Workspace OS usefulness must also be tested against the difference between Lucia Workspace + DAW as the cockpit and Full Booking Page as the record/review surface.

Guest Identity and Linkage

Did Lucia preserve the difference between a guest claim, a candidate booking, an operator-linked booking, and a verified booking? Required guest-facing scenarios:
Expected behavior is warm, useful, and bounded.

3. Emotional Containment

Did Lucia reduce overwhelm? Good containment:
Bad containment:

4. Truth-State Discipline

Did Lucia avoid overclaiming? Lucia must not imply:
unless verified.

5. Semantic Conversational Intent

Did Lucia understand short, social, lightweight utility, and scoped context prompts by meaning rather than exact phrase? Protected families include:
The expected behavior is bounded usefulness, not open-domain chat.

6. Required Signal → Action → Save → Reminder validation coverage

The Development/live-dev validation contract requires coverage of:
Product rule to validate:
Required Guest-facing validation coverage also includes:
This checklist is required validation coverage, not blanket verified runtime truth. Each passing claim still needs dated, scoped evidence from the owning runtime or evaluation record, and no item here establishes production readiness by itself.

Payment Truth Eval Requirements

Historical Development evidence — 2026-06-19:
This is synthetic historical evidence from 2026-06-19, not current proof. Present behavior must be re-established from the owning runtime and dated evaluation evidence. Required future coverage:
This coverage should test the architecture recorded in Lucia Payment Truth Foundation.

Eval Labs Role

EvaluationLabs.ai is Lucia’s proprietary evaluation platform for shaping her human intent layer, emotional awareness, psychological understanding, natural language interpretation, warmth, empathy, judgment, and operational intelligence. Eval Labs captures:
It creates a repeatable workflow for improving Lucia’s intelligence and tone. Current surfaces include custom and generated prompt tests, controlled batch runs, Guest Agent verification, a 72-case Fieldwork maintenance baseline, controlled Fieldwork conversation v1/v2 contracts, Review Queue, Run History, Team Review, Analysis, Human Eval Research, Registry Diagnostics, Behavioral Observatory, and the owner-only Platform Runtime Registry.

Quality Bar

A passing Lucia response should be:

See Also


Upstream / Downstream

Upstream

This layer is fed by:

Downstream

This layer affects: