Eval Labs is the dedicated platform for testing Lucia’s responses, capturing review judgments, and hardening behavior over time.
Product Role
Eval Labs is a first-class product surface.
EvaluationLabs.ai is Lucia’s proprietary evaluation platform for shaping her human intent layer, emotional awareness, psychological understanding, natural language interpretation, warmth, empathy, judgment, and operational intelligence.
It supports:
Why It Matters
Lucia’s intelligence cannot be trusted by vibes.
It needs a system that can answer:
Review Philosophy
Eval Labs should not reward polished nonsense.
It should reward:
Current Relationship to Lucia
Eval Labs is currently used to test Lucia’s operator-facing behavior, especially Focus Ops.
Over time it should support:
Current validation focus includes v0.1.3.6 Focus Ops behavior, semantic conversational intent assist, scoped entity routing, Calendar/booking-spine awareness, Workspace OS context awareness, Signal Stream active_context, active_context.workspace surface awareness, prior recommendation memory, prior offer context, saved DAW workflow truth-state, the Signal → Action → Save → Reminder loop, Resolver Matrix route correctness, Dynamic Action Workspace render correctness, Full Booking Page route correctness, GPT-5.6 Sol default-model behavior, and Lucia JSON gateway behavior through the OpenAI Responses API.
Runtime Verification and Run Provenance
Eval Labs production targets Lucia Engine Development at:
The owner-only Lucia Runtime Verification panel shows direct evidence from the same Engine response that produced Lucia’s answer:
New Eval runs persist runtimeProvenance on the case that produced the response and mirror the latest evidence in session metadata. The record includes the target endpoint/environment, Engine identity, request ID, configured/resolved models, execution-path booleans, task evidence, and runtime_verified_at.
Historical runs are not backfilled. When a run predates capture or otherwise has no record, the owner view says: Runtime provenance was not captured for this run. Missing evidence is not converted into an inferred endpoint or model.
Live owner verification on 2026-07-10 proved:
The uninterrupted Staging-target configuration began at 2026-05-04T16:21:35Z; routing was corrected and live-verified against Development on 2026-07-10. This proves the routing drift and correction, not that every historical run in the interval individually used Staging. Existing runs remain unchanged unless they already carry their own captured provenance.
Guest-facing Lucia adds a second required validation surface:
This surface should be formalized as a first-class Guest-Facing Lucia Eval Track, separate from operator-facing Lucia evals.
Purpose:
The track must cover identity orientation, orientation paths for already booked / joining someone / planning / exploring, booked-guest claim parsing, claim fragments across turns, weak vs strong claims, ambiguous/no-match cases, magic-link eligibility, email sent only to the booking email on file, no private data leakage, token consume/replay/expiry, verified session state, guest-to-operator signal creation, Admin Signal Stream visibility, Focus Ops no Luca/Nora drift, and warm hospitality tone.
Payment truth adds a separate required financial-attention validation lane:
This is future Eval Labs coverage. Current Harper Quinn proof is runtime Development proof, not broad Eval Labs certification.
v0.1.3.6 Dev Baseline Workflow
Eval Labs should validate v0.1.3.6 against:
The canonical Focus Ops route is:
Strict brain quality eval reached 178/178 after workspace-context awareness. This is evidence to capture and review, not a permanent guarantee.
v0.1.3.6 is not promoted to staging yet. Staging promotion waits until the Eval Labs dev baseline is captured and reviewed.
Current live-dev build identity under review:
Critical Review Question
The central question is:
If not, the response fails even if technically correct.
For guest-facing Lucia, the parallel question is:
If not, the response fails even if it sounds friendly.
See Also