Skip to main content
Eval Labs is Lucia’s role-based human evaluation and platform-evidence infrastructure. It lets authorized people test, score, annotate, analyze, compare, oversee, and improve Lucia’s behavior over time.

Definition

Eval Labs is the internal product and workflow used to evaluate Lucia’s responses with role-based human judgment. It is not simply a prompt runner. It is first-class Lucia intelligence infrastructure: the place where behavior is tested, inspected, reviewed, analyzed, and improved. Human judgment remains the quality authority; a successful platform run is evidence that a workflow operated, not proof that Lucia is approved. Eval Labs captures prompts, responses, sources, suites, runtime provenance when returned, human ratings, notes, review state, finalization, adjudication metadata, role and ownership scope, and durable run, item, and review records after Supabase persistence succeeds. Browser-local suite, resume, and Guest verification state is convenience state, not durable evidence.

Why Eval Labs exists

Generic AI benchmarks are not enough for Lucia. Lucia is built to help hospitality operators stay oriented, make good decisions, and trust the system under real pressure. That requires repeated, inspectable human evaluation of calm, warmth, trust, intent accuracy, operational usefulness, containment, truth-state discipline, and operator cognitive load. The system helps reviewers ask whether Lucia understood the user, chose the right mode, created calm rather than noise, preserved truth and trust, and gave the right next move.

Current product surfaces

As of 2026-08-12, Eval Labs supports:
  • Custom Prompt Test, Auto-generated Prompt Test, and saved custom suites
  • saved-suite deep links for evaluator and tester roles as well as privileged roles
  • Guest Facing Agent Verification Check and results
  • Controlled Batch Runner and specialized Fieldwork evaluation contracts
  • Run History, Running, Review Queue, and lifecycle finalization
  • keyboard review for every recognized role inside a run that role may already review
  • Team Review for owner/admin oversight
  • Analysis and single-run diagnostic surfaces for owner/admin
  • Registry Diagnostics, Behavioral Observatory, and owner-only Platform Runtime Registry
  • role-aware exports
The current analytics label is Analysis, not Global Analysis. Read Eval Labs Current System State and the Eval Labs Roles and Access Matrix.

Evidence domains are distinct

Eval Labs production delegates ordinary Lucia responses to Engine Development. Guest verification targets Guest Agent Production. Eval Labs itself does not own a configured or resolved OpenAI model claim.

Human judgment and suggestions

Eval Labs separates employee reaction from canonical meaning:
Current suggested selections and Registry Diagnostics mappings are heuristic aids. They are not a persisted AI judge and do not become canonical labels by appearing in the interface. Human reviewers intentionally save review evidence; senior review and adjudication assign canonical meaning where required. Behavioral Observatory differs from a transient suggestion: it preserves intentional reviewer labels when persistence succeeds.

Current access posture

  • owner has every declared surface, including Platform Runtime Registry.
  • admin has broad privileged testing, oversight, shared evidence, and Analysis access, but not Platform Runtime Registry.
  • evaluator has the evaluator workbench, evaluator-safe launchers, saved-suite deep links, and own scoped evidence.
  • tester has the narrower prompt-testing lane, saved-suite deep links, and own allowed evidence.
  • missing or unrecognized roles fail closed for protected routes.
Keyboard shortcuts never widen route, run, ownership, or RLS scope.

What Eval Labs is not

Eval Labs is not a generic chat app, one-off prompt playground, rubber-stamp review form, replacement for product judgment or Lucia doctrine, claim that Lucia is human-approved, or claim that visible UI proves durable authorization. It is Lucia-native behavioral judgment infrastructure with explicit evidence boundaries.

Dated readiness evidence

In May 2026, an AI-reviewed platform-readiness gate recorded 60 completed runs and 3,000 prompts, run items, Lucia responses, and reviews. Those figures are historical gate evidence. They are not current database counts, a fresh runtime observation, or human approval of Lucia’s quality.
The platform is implemented for controlled role-based human onboarding, with evaluator workspace polish, role-specific verification, and cohort guidance in active hardening.