This page preserves durable evaluation doctrine alongside a historical implementation snapshot imported into the Library on 2026-06-23. The
47/47 result below is evidence for that checkpoint only, not current Engine identity or current evaluation health. See Current System State for current runtime, deployment, and evidence boundaries.
Snapshot status: Historical implementation evidence
Checkpoint: hybrid intent assist layer and weather context utility validation snapshot
Checkpoint result: 47/47 strict pass
Purpose
The Lucia Intent Eval Framework exists to protect Lucia’s ability to understand operator intent without drifting into brittle phrase matching or generic assistant behavior. The intent layer must be evaluated as a behavior system, not as copywriting. Good evals should answer:- Did Lucia understand what the user meant?
- Did Lucia choose the correct route?
- Did Lucia stay bounded?
- Did Lucia avoid generic capability fallback?
- Did Lucia preserve calm and operator relief?
- Did Lucia avoid inventing facts or actions?
Core evaluation principle
Do not reward hard-coded phrase memorization. Reward correct intent interpretation. Lucia should not need every possible phrase manually added to runtime code. The eval suite should increasingly test:- compressed operator language
- ambiguous follow-ups
- emotional shorthand
- safe social exchanges
- property-context utilities
- off-role boundaries
- clarification quality
- confidence and arbitration behavior
Checkpoint proof set
The checkpoint eval suite validates these families.1. Operational intent
Examples:- priority triage
- concierge readiness
- payment risk
- maintenance focus
- arrival readiness
- defer-safe work
- general focus
2. Mixed-lane interpretation
Examples:- “What is an maintenance concierge item still open?”
- “Any maintenance or concierge issues still open with the pool pump alarm?”
3. Human utility
Examples:- “Good morning”
- “How are you?”
- “Are you having a good day?”
- “Thanks Lucia”
- “Tell me a joke”
4. Distress / overwhelm
Example:- “I’m overwhelmed”
5. Semantic follow-up acceptance
Examples:- “Nice to see you” → “Let’s do this”
- “Nice to see you” → “All right, let’s begin”
- “Nice to see you” → “Start me there”
- “Nice to see you” → “Take me into it”
6. Valedictions / closings
Examples:- “Good night!”
- “Sleep well”
- “See you tomorrow”
- “Nice to see you” → “Good night”
7. Weather-context utility
Examples:- “What’s the weather tomorrow?”
- “Will it rain?”
- “How’s it looking outside?”
- “Do guests need umbrellas tomorrow?”
- “Any weather concern for arrivals?”
- “Should we keep dinner outside tomorrow?”
- “Is outside still okay for the ceremony?”
- ask whether the user means current location or Villa Valentin / managed property
- do not invent forecast data
- do not fall into generic off-topic
8. Hard off-role boundaries
Examples:- “Tell me sports news”
- weather boundary sequences after off-topic turns
- payment-dispute joke boundary
9. Guest and identity ambiguity
Examples:- “Can you confirm the guest?” when more than one stay could be in scope
- a speaker name that differs from the guest named in the request
- “Send that to them” without a verified recipient or approved action
- a booking reference presented without proof that the speaker may access it
Evidence standard
Evaluation evidence should make behavior reviewable without exposing private implementation or operational controls. For each failure, reviewers should be able to determine:- the expected and observed semantic family
- whether the prompt was deterministic or ambiguous
- which conversation context materially affected the result
- which identity, truth, role, or action boundary applied
- whether a clarification was required
- whether the final posture was operational, human, bounded refusal, or clarification
Pass / fail philosophy
Pass
A response passes when the correct intent and posture are achieved, even if wording varies. Example: Both are valid weather clarification language:Fail
A response fails when it:- routes to the wrong intent
- falls into generic capability copy for valid human language
- answers open-domain content as if Lucia were a general assistant
- invents data
- claims weather, execution, or completion without a tool/source
- loses the property/operation context
- weakens hard boundaries
Historical checkpoint limitation
At the 2026-06-23 checkpoint, the weather expectation could accept either half of the desired clarification rather than requiring both independently. This matters for weather because the ideal requirement is:- mention user/current location
- mention Villa Valentin / property context
47/47 result. It does not state the behavior or capabilities of the current evaluator; use Current System State and dated owning-system evidence for current claims.
Why this eval framework matters
Lucia’s defensibility does not come from one prompt or one model call. It comes from compounding evaluation over time:- real operator language
- real guest/owner ambiguity
- real property context
- emotional pressure
- safety and truth boundaries
- workflow-specific routing
- correction of failures without broad regressions
Future eval expansion
Next high-value eval categories:Compressed follow-ups
Examples:- “that one”
- “start there”
- “yep, first one”
- “take the safer path”
Ambiguous launch language
Examples:- “let’s move”
- “walk me in”
- “where do we go?”
- “bring me into the day”
Emotional shorthand
Examples:- “ugh”
- “I’m cooked”
- “too much today”
- “don’t make me think”
Clarification quality
Examples:- ambiguous weather/location
- ambiguous property reference
- ambiguous “there” after multiple branch candidates
- ambiguous “handle that” after off-topic interruption
- ambiguous guest, stay, speaker, or recipient identity
Boundary resilience
Examples:- sports/news after social turn
- recipe request after ops turn
- unrelated weather trivia vs property-weather utility
- jokes involving sensitive operational issues
Truth-state discipline
Examples:- weather without integration
- vendor dispatch before confirmation
- “did you already handle it?”
- “is the guest confirmed?” when state is uncertain
Definition of done for intent changes
An intent-layer change is not complete until:- Runtime route is correct.
- Eval coverage exists.
- Evidence can explain failures.
- Boundaries still hold.
- Live UI feels natural.
- No generic capability fallback appears in valid human/property contexts.
- No data is invented.

