Skip to main content
Eval Labs is the internal evaluation system used to test and improve Lucia’s behavior. Evaluators help decide whether Lucia is useful for humans, not whether the platform merely ran.

The plain version

Eval Labs helps the team test Lucia against real behavioral expectations. It is first-class human evaluation infrastructure, not a score-collection exercise. It captures prompts, Lucia responses, human scores and notes, final review state, role and ownership scope, runtime provenance when the response source provides it, and Supabase-backed evidence when persistence succeeds. The goal is reliable evidence about whether Lucia is improving. A visible response or local browser state is not durable evidence until the relevant save succeeds.

What evaluators are judging

You are judging whether Lucia worked for the human situation in front of her. Ask:
  • Did Lucia understand the prompt?
  • Was the response truthful, useful, and clear?
  • Was the tone right for the moment?
  • Did it reduce confusion or add to it?
  • Would a real operator trust Lucia more after reading it?

Your current workspace

As of 2026-08-12, evaluator access includes the evaluator workbench, evaluator-safe test types, saved custom suites and their deep links, and your own scoped run, review, and history routes. The Keyboard Review Mode reference is available to every recognized role. Its shortcuts apply only inside a Review Queue you are already authorized to use; keyboard review never widens run ownership, route access, or RLS scope. Evaluators do not receive Team Review, Analysis, Registry Diagnostics, Behavioral Observatory, Platform Runtime Registry, Fieldwork Communication Baseline, cross-user evidence, or owner/admin tools. The current product label is Analysis, not Global Analysis. See Your Role and Access for evaluator routes, the Eval Labs Roles and Access Matrix for canonical grants, and Eval Labs Current System State for dated platform truth.

Evidence boundaries

Ordinary run evidence is separate from the owner-only Platform Runtime Registry and from specialized Fieldwork and Guest target protocols. A Guest verification result is evidence about the Guest Agent target. A Fieldwork cohort or receipt belongs to its declared Fieldwork contract. Neither silently becomes ordinary run evidence or proof about every Lucia surface. Suggested selections are heuristic review aids, not a persisted AI judge. Treat a suggestion as guidance only. Your deliberately saved human review is the judgment record when persistence succeeds.

What evaluators are not judging

You are not being asked to approve the whole product, debug infrastructure, decide product strategy, infer quality from platform counts, or work around access boundaries. You are reviewing Lucia responses inside your assigned Eval Labs workflow.

Dated readiness evidence

The May 2026 AI-reviewed platform-readiness gate recorded 60 completed runs and 3,000 prompts, run items, Lucia responses, and reviews. Those figures are historical gate evidence, not current database counts and not human Lucia-quality approval. Evaluator workbench access is implemented. Onboarding and workspace polish remain in active hardening. That distinction matters every time you review.