The validation battery protects Lucia from regressions in intent, truth, tone, and operational usefulness.
Purpose
The validation battery exists to test Lucia against known scenarios before behavior is considered stable. It protects against:The gates on 2026-09-10
The standing gates for a change to Lucia’s behaviour, as they stood on 2026-09-10. Each is dated by the record that built it; none is a claim that the platform is launch-ready.
Battery discipline across the fleet: green means exit code 0, never a count of failure lines; integration is by literal 40-character SHA; the Engine runs its battery on Node 26 and Node 22 (ALGO-40; company gate law as recorded on HLLC-41).
Live testing of the served surfaces is recorded on a dated founder-walkthrough ledger that is internal and not published; this Library names it, not its contents.
Target and diagnostic posture — 2026-09-10, with the 2026-08-12 diagnostic as history
The Engine validation bank contains 217 cases. Eval Labs Production evaluates operator-facing Lucia through Engine Development at:9a15b1da… returned 76/217. On 2026-09-03 two runs on Development read 174/217 and 192/217 (ALGO-36).
Eval Labs Production currently serves commit 6cbfd15f75330c414bfe79e0d2fab51ef8115102, version 0.1.0, at https://evaluationlabs.ai, verified through same-deploy runtime identity evidence.
fbf77b662511cf004ee1f46787e05773a8a78d71 proved configured and provider-resolved gpt-5.6-sol for that request only. This August documentation review made no fresh Guest Agent model call. Eval Labs has delegated model ownership and verifies returned same-response or signed evidence.
Eval Labs has an implemented Guest Verification Check for deterministic booked-guest scenarios. The broader guest identity and verification battery remains requested/future; a recorded check result still does not establish launch readiness without the applicable gate.
Payment truth now also requires a dedicated financial-attention battery before policy-aware payment judgment is claimed.
Guest-facing evaluation status
The implemented Guest Facing Agent Verification Check is separate from operator-facing Focus Ops evals and covers a deterministic booked-guest scenario pack. The broader first-class Guest-Facing Lucia Eval Track below remains requested/future. Purpose:Fieldwork evaluation tracks
Fieldwork validation is separate from the ordinary prompt-test session graph.Communication baseline
Conversation behavior v1
Conversation execution v2
Core Test Categories
Payment Truth / Financial Attention
Historical Development evidence — 2026-06-19:Property clock
Scenarios involving every served day word and clock word (Property time is the only clock, ALGO-34, 2026-09-03):Owner language
Every served owner-facing field (title, subject, summary, reason, next action, label, why-now, guest-message draft) is scanned for engine words with no exemptions (ALGO-29, 2026-09-03). Expected: the rules R1 to R10 on Voice and Tone; failing behaviour is any engine word, any acute in a name, or any decimal duration reaching the owner.Calendar / Booking Spine
Scenarios involving:Priority Triage
Prompts like:Deferral
Prompts like:Human Utility
Prompts like:Recommended Action Destination Contract
- CTA path and destination path must match.
- Workflow-specific CTAs resolve through structured action metadata into the Dynamic Action Workspace or another safe focused workspace.
- Calendar booking clicks route to the Full Booking Page for record/review, not to Dynamic Action Workspace.
- Generic booking overview is used only as fallback.
- Admin can route directly to the focused workflow UI.
/ops/actionsis the default universal Dynamic Action Workspace for structured Focus Ops actions.- Specialized pages may remain available as supporting or future-specialized surfaces, but they are not the default Focus Ops CTA destination.
- CTA label is copy; structured action intent and metadata are routing truth.
- Failing behavior: CTA label names a specific action, but destination points to generic booking overview, causing the operator to hunt or land mid-page.
Named Guest / Service Scoping
Prompts like:Signal Stream Active Context
Prompts started from Signal Stream Chat should carry active context into Focus Ops. Expected:Prior Recommendation Memory
Short follow-ups:Workspace Context Awareness
Prompts from the Workspace Sidebar should be evaluated against the current surface. Supported surfaces include:Reminder Create / Resurface / Got It
Prompts and actions:Signal Stream Ordering / Dismiss / Move-To-Top
Expected:Resolver Matrix Route Correctness
Expected:Dynamic Action Workspace Render Correctness
Expected:DAW Save Truth-State
Expected:Semantic Conversational Intent Assist
Prompts like:Guest Identity Orientation
Scenarios:Guest Claim Strength
Scenarios:Magic Link Booking Verification
Scenarios:Guest Signal Routing
Scenarios:Maintenance Focus
Prompts or signals involving:Pass Criteria
A response passes if it is:Fail Criteria
A response fails if it:Dated evidence history
- The 2026-05-29 strict brain-quality checkpoint recorded
178/178against an older 178-case bank and revision. It is not the current 217-case baseline. - The May 2026 Eval Labs platform-readiness gate recorded 60 runs and 3,000 prompts, items, responses, and reviews. It is not a current database count or human quality approval.
- The 2026-07-10 owner verification observed
gpt-5.6-solconfigured and provider-resolved for one Engine Development request. No new model-bearing verification was run for this update. - The accepted 2026-07-18 Guest Agent same-deploy/provider receipt proved configured and provider-resolved
gpt-5.6-solfor sourcefbf77b662511cf004ee1f46787e05773a8a78d71and one production request. No fresh August Guest Agent model call was made.

