Skip to main content
Eval Labs is a deployed, role-based human evaluation platform for Lucia. Its source and configuration define controlled human onboarding, durable-evidence storage, owner/admin oversight, and evaluator-safe workflows, while authenticated access and persistence remain rollout gates rather than claims from this documentation check.
Freshness note — 2026-09-10. No Eval Labs release has shipped since 2026-07-29. On 2026-09-10 the documentation audit read lucia-ai-eval-labs main and the production deploy at evaluationlabs.ai, both still at commit 6cbfd15f75330c414bfe79e0d2fab51ef8115102, the same identity verified on 2026-08-12. Claims on this page about the systems Eval Labs targets (the Engine Development deployment, its model configuration, the Guest Agent receipt) were not re-verified on 2026-09-10 and stay dated to the evidence they cite.

Current platform truth

Status: deployed source/config contract implemented; authenticated behavior not re-proven in this documentation check.
Verified 2026-08-12: https://evaluationlabs.ai serves Eval Labs Production commit 6cbfd15f75330c414bfe79e0d2fab51ef8115102, version 0.1.0. The runtime identity reports production and same-deploy build evidence.
Eval Labs is no longer only a founder or AI-agent testing tool. Its deployed source/config contract defines Lucia’s role-based human evaluation platform:
  • Clerk authentication integration and protected routes.
  • Role-aware client behavior for the owner, admin, evaluator, and tester roles.
  • Durable server and data authorization intended to enforce role and ownership boundaries independently of the interface.
  • Durable storage for ordinary runs and human reviews; separate Runtime Registry and Fieldwork protocols retain their own claims, seals, receipts, reviews, and lineage.
  • Browser-local custom suites, review-resume state, and Guest verification history are convenience state and do not replace durable evidence.
  • Production configuration targeting Lucia Engine Development at https://api-dev.hellolucia.ai/admin/operator-focus.
  • Guest verification configuration targeting Guest Agent Production at https://guest.hellolucia.ai, with Fieldwork relay configuration targeting Engine Development.
  • An owner-only Lucia Runtime Verification panel designed to read runtime/model provenance from the same Engine response that produced Lucia’s answer.
  • A durable-run write contract that retains captured runtime provenance for new runs and reports uncaptured provenance explicitly for historical runs.
  • An owner/admin access contract for shared durable Eval Labs evidence.
  • The access contract scopes evaluator and tester data to their own work except where owner/admin oversight applies.
  • Team Review is defined as the owner/admin oversight surface.
  • Platform Runtime Registry is defined as an owner-only deploy/model-evidence surface.
  • Human Eval Research is designed to render sanitized, reviewable snapshots and does not automatically change Lucia behavior.
  • Staged-hydration source loads run summaries first, then recent and deeper evidence, so dashboards can render without synthetic metrics.

Runtime and model ownership

Engine and Guest Agent own their OpenAI Responses API calls. Eval Labs owns orchestration, provenance assessment, evidence persistence, human review, and release-gated evaluation contracts. A requested or configured model name alone is not execution proof.

Current roles

Status: implemented. Current roles:
  • owner
  • admin
  • evaluator
  • tester
  • unassigned or missing role
Read the canonical matrix: Eval Labs Roles and Access Matrix.

Current test surfaces

Status: implemented. Current test surfaces:
  1. Custom Prompt Test
  2. Auto-generated Prompt Test
  3. Guest Facing Agent Verification Check
  4. Controlled Batch Runner
  5. Fieldwork Communication Baseline — 72 fixed cases, owner/admin only
  6. Fieldwork Conversation Behavior v1 — 40 cases across three runs
  7. Fieldwork Conversation Execution v2 — explicit 120-entry cohort and receipt contract
Tester access is intentionally narrower than evaluator access. Tester is for clean prompt-testing onboarding cohorts. Evaluator is for the full evaluator workbench and evaluator-safe test types. The Fieldwork v1/v2 controls are specialized interfaces, not ordinary Launcher sessions, and nothing runs automatically on page load.

Oversight and analysis

Status: implemented. The deployed role contract gives owner and admin broad privileged access to shared durable evidence, Team Review, Analysis, and all Launcher test surfaces. They are not identical: Platform Runtime Registry is owner-only. This was not re-proven through authenticated production interaction on 2026-08-12. Team Review exists for owner/admin oversight of human evaluation work: evidence quality, reviewer activity, review gaps, and escalation readiness. Analysis is owner/admin-only platform-wide evidence inspection. It is not a tester or evaluator onboarding surface. The owner-only runtime verification panel shows the resolved endpoint, classified environment, Engine commit SHA, configured model, provider-resolved model, model-invoked state, deterministic-path state, fallback state, and verification timestamp.

Dated history, not current proof

  • The May 2026 AI-reviewed platform-readiness gate recorded 60 completed runs and 3,000 prompts, run items, Lucia responses, and reviews. Those counts describe that gate, not the current database.
  • A Verified Engine Development observation on 2026-07-10 showed gpt-5.6-sol configured and resolved, model invocation true, deterministic path false, and fallback false. It proves that request only.
  • The latest accepted Guest Agent same-deploy/provider receipt is dated 2026-07-18 for source fbf77b662511cf004ee1f46787e05773a8a78d71. It proves that Guest deployment and request only; no fresh August Guest model call was made for this documentation check.
  • The continuous Staging-target configuration began at 2026-05-04T16:21:35Z and was corrected and live-verified against Development on 2026-07-10. This is routing history, not proof that every run used Staging.
Historical records are not rewritten or backfilled with invented provenance.
The 2026-08-12 documentation verification did not authenticate, exercise durable role/ownership authorization, verify a durable write/read, make a model-bearing request, inspect production counts, or execute a Fieldwork cohort. Do not present deployed source/config policy as live authentication or persistence proof, a new provider observation, or a completed specialized cohort.

Human onboarding posture

Status: active hardening. The deployed source/config contract supports preparation for controlled human onboarding by role and assignment. The cohort remains gated on the authenticated checks below. Do not describe the platform as broadly production-mature or open-access. Do not describe Lucia as human-approved because the May 2026 AI-reviewed platform-readiness gate passed. Use:
Avoid softer labels that imply more maturity than the source state proves.

Active hardening

These areas are implemented but still being tightened, polished, or verified for rollout:
  • evaluator onboarding/workspace polish
  • first human cohort instructions
  • role-specific route verification
  • end-to-end Clerk authentication and durable authorization verification after access-control changes
  • staged hydration behavior across large evidence sets
  • clear tester-vs-evaluator assignment guidance
  • fresh authenticated route/access proof after material role changes
  • fresh runtime/model observations when a release requires them

Deferred

Deferred means intentionally outside the current access model:
  • tester access to Verification Check
  • tester access to Verification Results
  • tester access to Controlled Batch Runner
  • tester access to Team Review
  • tester access to Analysis
  • tester access to Registry Diagnostics
  • tester access to Behavioral Observatory
  • evaluator access to Team Review
  • evaluator access to Analysis
  • admin access to Platform Runtime Registry

Future

Future means possible later, not current behavior:
  • broader public or external evaluator rollout
  • expanded assignment management
  • additional owner/admin management tooling
  • deeper cohort analytics beyond current oversight surfaces
  • more final evaluator UX polish

First human onboarding readiness criteria

Before a first human onboarding cohort starts:
  1. Confirm every participant can authenticate through Clerk.
  2. Confirm each participant has the intended Eval Labs application role.
  3. Confirm durable server and data authorization enforce that role and the expected ownership scope.
  4. Confirm visible routes match the access matrix.
  5. Run a real prompt test and verify that it saves and reloads within the expected durable-data scope.
  6. Confirm owner/admin can see shared persisted evidence where oversight applies.
  7. Keep testers limited to their supported Custom Prompt Test, Auto-generated Prompt Test, history, keyboard-review, and eligible owned-run routes.
  8. Give evaluators only assignments that match evaluator-safe surfaces.
  9. Name any active-hardening caveats before the work begins.
  10. Repeat that AI-reviewed platform readiness is not human Lucia-quality approval.