Skip to main content
At a glanceEval Labs is Lucia’s quality-control system. It exists to test whether Lucia is useful, truthful, calm, semantically aware, and operationally correct before behavior becomes trusted.

Why Evaluation Exists

Lucia is not evaluated only on whether an answer is “correct.” Lucia must be evaluated on whether the answer:

Evaluation Is Product Infrastructure

Eval Labs is not a side tool. Eval Labs is part of Lucia’s intelligence stack.
Lucia behaviorEval Labsreviewrefinementsafer behavior

Current v0.1.3.6 Evaluation Posture

Strict brain quality eval reached 178/178 after workspace-context awareness. This is current live-dev evidence, not a permanent guarantee. Eval Labs should validate v0.1.3.6 against:
The canonical Focus Ops route is:
v0.1.3.6 is not promoted to staging yet. Staging promotion waits until the Eval Labs dev baseline is captured and reviewed. Guest-facing Lucia now requires its own first-class Eval Labs track before guest-facing behavior is treated as launch-ready. Payment truth now requires dedicated financial-attention coverage before policy-aware payment judgment is treated as stable.

Guest-Facing Lucia Eval Track

The Guest-Facing Lucia Eval Track is separate from operator-facing Lucia evals. Purpose:
This track must cover:

Primary Evaluation Targets

1. Intent Accuracy

Did Lucia understand what the operator was asking? Examples:

2. Operational Usefulness

Did Lucia identify a useful next move? A technically accurate answer can still fail if it leaves the operator with too much work. Current booking-spine usefulness must be tested against arrivals, departures, stay windows, and the difference between Full Booking Page review and Dynamic Action Workspace completion. Current Workspace OS usefulness must also be tested against the difference between Lucia Workspace + DAW as the cockpit and Full Booking Page as the record/review surface.

Guest Identity and Linkage

Did Lucia preserve the difference between a guest claim, a candidate booking, an operator-linked booking, and a verified booking? Required guest-facing scenarios:
Expected behavior is warm, useful, and bounded.

3. Emotional Containment

Did Lucia reduce overwhelm? Good containment:
Bad containment:

4. Truth-State Discipline

Did Lucia avoid overclaiming? Lucia must not imply:
unless verified.

5. Semantic Conversational Intent

Did Lucia understand short, social, lightweight utility, and scoped context prompts by meaning rather than exact phrase? Protected families include:
The expected behavior is bounded usefulness, not open-domain chat.

6. Signal → Action → Save → Reminder Loop

Current live-dev validation must cover:
The product rule under test:
Guest-facing validation must also cover:
This is Development/live-dev runtime truth, not a production-readiness claim.

Payment Truth Eval Requirements

Current proof status:
Required future coverage:
This coverage should test the architecture recorded in Lucia Payment Truth Foundation.

Eval Labs Role

EvaluationLabs.ai is Lucia’s proprietary evaluation platform for shaping her human intent layer, emotional awareness, psychological understanding, natural language interpretation, warmth, empathy, judgment, and operational intelligence. Eval Labs captures:
It creates a repeatable workflow for improving Lucia’s intelligence and tone.

Quality Bar

A passing Lucia response should be:

See Also


Upstream / Downstream

Upstream

This layer is fed by:

Downstream

This layer affects: