Skip to main content
The validation battery protects Lucia from regressions in intent, truth, tone, and operational usefulness.

Purpose

The validation battery exists to test Lucia against known scenarios before behavior is considered stable. It protects against:

The gates on 2026-09-10

The standing gates for a change to Lucia’s behaviour, as they stood on 2026-09-10. Each is dated by the record that built it; none is a claim that the platform is launch-ready. Battery discipline across the fleet: green means exit code 0, never a count of failure lines; integration is by literal 40-character SHA; the Engine runs its battery on Node 26 and Node 22 (ALGO-40; company gate law as recorded on HLLC-41). Live testing of the served surfaces is recorded on a dated founder-walkthrough ledger that is internal and not published; this Library names it, not its contents.

Target and diagnostic posture — 2026-09-10, with the 2026-08-12 diagnostic as history

The Engine validation bank contains 217 cases. Eval Labs Production evaluates operator-facing Lucia through Engine Development at:
The canonical Focus Ops route is:
On 2026-08-12 an exact-SHA non-strict diagnostic against Engine commit 9a15b1da… returned 76/217. On 2026-09-03 two runs on Development read 174/217 and 192/217 (ALGO-36).
None of those figures is a pass, strict gate, baseline, or release claim; the bank has standing failures that are open work. Record strictness, suite identity, case count, target, and exact SHA with every result. The September behaviour gate is the replay harness plus identical failing sets before and after (table above).
Eval Labs Production currently serves commit 6cbfd15f75330c414bfe79e0d2fab51ef8115102, version 0.1.0, at https://evaluationlabs.ai, verified through same-deploy runtime identity evidence.
Engine Development and Guest Agent Production independently own their direct model calls; since 2026-09-04 the Engine resolves the provider from the model id and composes Focus on Anthropic’s Fable on Development, while the Guest Agent’s OpenAI path is unchanged since 2026-07-18 (ALGO-39). Engine execution claims require same-response provenance. For Guest Agent Production, the accepted 2026-07-18 same-deploy/provider receipt for source fbf77b662511cf004ee1f46787e05773a8a78d71 proved configured and provider-resolved gpt-5.6-sol for that request only. This August documentation review made no fresh Guest Agent model call. Eval Labs has delegated model ownership and verifies returned same-response or signed evidence. Eval Labs has an implemented Guest Verification Check for deterministic booked-guest scenarios. The broader guest identity and verification battery remains requested/future; a recorded check result still does not establish launch readiness without the applicable gate. Payment truth now also requires a dedicated financial-attention battery before policy-aware payment judgment is claimed.

Guest-facing evaluation status

The implemented Guest Facing Agent Verification Check is separate from operator-facing Focus Ops evals and covers a deterministic booked-guest scenario pack. The broader first-class Guest-Facing Lucia Eval Track below remains requested/future. Purpose:
Required coverage:

Fieldwork evaluation tracks

Fieldwork validation is separate from the ordinary prompt-test session graph.

Communication baseline

Conversation behavior v1

Conversation execution v2

Claims, seals, receipts, review attestations, and recovery lineage are evidence. They do not independently grant release authority. No Fieldwork cohort was executed during the 2026-08-12 documentation verification.

Core Test Categories

Payment Truth / Financial Attention

Historical Development evidence — 2026-06-19:
This is an explicitly synthetic historical fixture from 2026-06-19, not current proof. Current payment behavior requires fresh evidence from the owning runtime and a dated evaluation record. Required Eval Labs / regression scenarios:
Failing behavior:
Boundary recorded at the 2026-06-19 checkpoint:

Property clock

Scenarios involving every served day word and clock word (Property time is the only clock, ALGO-34, 2026-09-03):
Expected: plain local time only, and no served string that would let an owner tell two clocks exist.

Owner language

Every served owner-facing field (title, subject, summary, reason, next action, label, why-now, guest-message draft) is scanned for engine words with no exemptions (ALGO-29, 2026-09-03). Expected: the rules R1 to R10 on Voice and Tone; failing behaviour is any engine word, any acute in a name, or any decimal duration reaching the owner.

Calendar / Booking Spine

Scenarios involving:
Expected:
Failing behavior:

Priority Triage

Prompts like:
Expected:

Deferral

Prompts like:
Expected:

Human Utility

Prompts like:
Expected:

  • CTA path and destination path must match.
  • Workflow-specific CTAs resolve through structured action metadata into the Dynamic Action Workspace or another safe focused workspace.
  • Calendar booking clicks route to the Full Booking Page for record/review, not to Dynamic Action Workspace.
  • Generic booking overview is used only as fallback.
  • Admin can route directly to the focused workflow UI.
  • /ops/actions is the default universal Dynamic Action Workspace for structured Focus Ops actions.
  • Specialized pages may remain available as supporting or future-specialized surfaces, but they are not the default Focus Ops CTA destination.
  • CTA label is copy; structured action intent and metadata are routing truth.
  • Failing behavior: CTA label names a specific action, but destination points to generic booking overview, causing the operator to hunt or land mid-page.

Named Guest / Service Scoping

Prompts like:
Expected:
Failing behavior:

Signal Stream Active Context

Prompts started from Signal Stream Chat should carry active context into Focus Ops. Expected:

Prior Recommendation Memory

Short follow-ups:
Expected:
Failing behavior:

Workspace Context Awareness

Prompts from the Workspace Sidebar should be evaluated against the current surface. Supported surfaces include:
Prompts:
Expected:
Failing behavior:

Reminder Create / Resurface / Got It

Prompts and actions:
Expected:

Signal Stream Ordering / Dismiss / Move-To-Top

Expected:

Resolver Matrix Route Correctness

Expected:
Failing behavior:

Dynamic Action Workspace Render Correctness

Expected:

DAW Save Truth-State

Expected:

Semantic Conversational Intent Assist

Prompts like:
Expected:

Guest Identity Orientation

Scenarios:
Expected:

Guest Claim Strength

Scenarios:
Expected:
Failing behavior:

Scenarios:
Expected:

Guest Signal Routing

Scenarios:
Expected:

Maintenance Focus

Prompts or signals involving:
Expected:

Pass Criteria

A response passes if it is:
A battery passes only under its declared strict gate. A partial, non-strict, source-only, or historical result remains diagnostic even when individual cases pass.

Fail Criteria

A response fails if it:

Dated evidence history

  • The 2026-05-29 strict brain-quality checkpoint recorded 178/178 against an older 178-case bank and revision. It is not the current 217-case baseline.
  • The May 2026 Eval Labs platform-readiness gate recorded 60 runs and 3,000 prompts, items, responses, and reviews. It is not a current database count or human quality approval.
  • The 2026-07-10 owner verification observed gpt-5.6-sol configured and provider-resolved for one Engine Development request. No new model-bearing verification was run for this update.
  • The accepted 2026-07-18 Guest Agent same-deploy/provider receipt proved configured and provider-resolved gpt-5.6-sol for source fbf77b662511cf004ee1f46787e05773a8a78d71 and one production request. No fresh August Guest Agent model call was made.

See Also