At a glanceEval Labs is Lucia’s quality-control system. It exists to test whether Lucia is useful, truthful, calm, semantically aware, and operationally correct before behavior becomes trusted.
Why Evaluation Exists
Lucia is not evaluated only on whether an answer is “correct.” Lucia must be evaluated on whether the answer:Evaluation Is Product Infrastructure
Eval Labs is not a side tool. Eval Labs is part of Lucia’s intelligence stack.Lucia behaviorEval Labsreviewrefinementsafer behavior
The gates on 2026-09-10
The standing gates for a change to Lucia’s behaviour, as they stood on 2026-09-10. Each is dated by the record that built it; none is a claim that the platform is launch-ready.
Battery discipline across the fleet: green means exit code 0, never a count of failure lines; integration is by literal 40-character SHA; the Engine runs its battery on Node 26 and Node 22 (ALGO-40; company gate law as recorded on HLLC-41).
Live testing of the served surfaces is recorded on a dated founder-walkthrough ledger that is internal and not published; this Library names it, not its contents.
Evaluation posture — 2026-09-10, with the 2026-08-12 diagnostic as history
The Engine evaluation bank contains 217 cases. On 2026-08-12 an exact-SHA, non-strict diagnostic against Engine commit9a15b1da… returned 76/217; on 2026-09-03 two runs on Development read 174/217 and 192/217 (ALGO-36). None of these is a pass, a release gate, a quality baseline, or a production-readiness claim. The bank’s standing failures are open work, and the September gate is the replay harness plus identical failing sets, as in the table above.
Eval Labs Production evaluates operator-facing Lucia against Engine Development at:
https://evaluationlabs.ai. Its release identity verified on 2026-08-12 is commit 6cbfd15f75330c414bfe79e0d2fab51ef8115102, version 0.1.0, with same-deploy runtime identity evidence.
Guest verification separately targets Guest Agent Production at https://guest.hellolucia.ai. Fieldwork evaluation relays remain Development-only.
gpt-5.6-sol for one production request. It does not establish later Guest Agent execution.
Historical checkpoints
- The 2026-05-29
178/178checkpoint belongs to an older 178-case bank and revision, not the current 217-case baseline. - The May 2026 Eval Labs gate recorded 60 runs and 3,000 prompts, items, responses, and reviews. It is historical platform-lifecycle evidence, not a current count or human approval.
- The 2026-07-10 owner verification observed
gpt-5.6-solconfigured and provider-resolved for one Engine Development request. It does not prove unrelated runs. - The accepted 2026-07-18 Guest Agent same-deploy/provider receipt proved configured and provider-resolved
gpt-5.6-solfor sourcefbf77b662511cf004ee1f46787e05773a8a78d71and one production request. No fresh August Guest Agent model call was made.
Guest-facing evaluation status
Implemented: Guest Verification Check
The Guest Facing Agent Verification Check is an implemented Eval Labs surface separate from operator-facing Lucia evals. It targets Guest Agent Production athttps://guest.hellolucia.ai, runs the deterministic booked-guest verification scenario pack, and exposes check results. It does not establish that the broader battery below exists or passes.
Requested/future: broader Guest-facing battery
The broader first-class Guest-Facing Lucia Eval Track remainsRequested with capability status future.
Purpose:
Primary Evaluation Targets
1. Intent Accuracy
Did Lucia understand what the operator was asking? Examples:2. Operational Usefulness
Did Lucia identify a useful next move? A technically accurate answer can still fail if it leaves the operator with too much work. Current booking-spine usefulness must be tested against arrivals, departures, stay windows, and the difference between Full Booking Page review and Dynamic Action Workspace completion. Current Workspace OS usefulness must also be tested against the difference between Lucia Workspace + DAW as the cockpit and Full Booking Page as the record/review surface.Guest Identity and Linkage
Did Lucia preserve the difference between a guest claim, a candidate booking, an operator-linked booking, and a verified booking? Required guest-facing scenarios:3. Emotional Containment
Did Lucia reduce overwhelm? Good containment:4. Truth-State Discipline
Did Lucia avoid overclaiming? Lucia must not imply:5. Semantic Conversational Intent
Did Lucia understand short, social, lightweight utility, and scoped context prompts by meaning rather than exact phrase? Protected families include:6. Required Signal → Action → Save → Reminder validation coverage
The Development/live-dev validation contract requires coverage of:Payment Truth Eval Requirements
Historical Development evidence — 2026-06-19:Eval Labs Role
EvaluationLabs.ai is Lucia’s proprietary evaluation platform for shaping her human intent layer, emotional awareness, psychological understanding, natural language interpretation, warmth, empathy, judgment, and operational intelligence. Eval Labs captures:Quality Bar
A passing Lucia response should be:See Also
- 02 - Validation Battery
- 03 - Quality Bar
- 06 - Semantic Conversational Intent Assist
- 04 - Seed Data and Test Worlds
- lucia-world
- 02 - Focus Ops Intelligence
- 04 - Lucia Workspace OS Milestone
- 07 - Guest-Facing Eval Requirements

