Skip to main content
This page records the May 2026 review-layer evolution: shared run launchers, employee review, suggested selections, Human Guidance Evaluation, adjudication-ready schema, exports, queue filters, lifecycle finalization, and the later platform-readiness split between normal testing and controlled batch gates. For current platform truth, see Current System State.

May 2026 review-layer milestone

Eval Labs evolved from a prompt runner into a layered review product. The key change:
custom or automated runshared Review Queuesuggested selections plus human reviewlifecycle finalization and export

Changes shipped in May 2026

Adjudication-ready schema

The review contract separated draft and saved judgment, app suggestions, human labels, employee signal, Human Guidance Evaluation, senior adjudication, canon-candidate routing, and run lifecycle state. At the May 2026 checkpoint, those meanings were preserved through browser-local draft continuity, durable payload persistence, dirty-state detection, and exports. The release outcome matters here; internal storage keys and composition remain implementation detail.

Employee Review layer

Added guided employee-review fields:
At that checkpoint, these fields replaced freeform taxonomy collection for non-expert reviewers. The app could also suggest Employee Review answers from prompt/response heuristics. The suggestion was visible as suggested signal; the reviewer still saved the human review.

Suggested review layer

Added app-suggested review values for:
At that checkpoint, these suggestions came from prompt text, Lucia response text, run status, run errors, and simple response-quality heuristics such as clear next move, calming language, list-heavy output, robotic language, fake empathy, and overclaiming. They are not canonical truth.

Review Queue UX

At the May 2026 checkpoint, the Review Queue favored guided employee judgment:
  • single-column Quick Review flow
  • numbered question cards
  • separate selection boxes
  • suggested selections
  • reduced freeform text burden
  • senior-review routing
  • canon-candidate routing
  • Human Guidance Evaluation
  • Save / Save & Next / Save & next flagged flows
  • search and workflow filters
  • JSON, CSV, and Markdown export controls
  • finalization after all prompts are reviewed

Semantic confidence sliders

The “How did Lucia do?” scoring section moved from 1–10 button rows to stepped semantic confidence sliders. The final design direction:
The sliders should feel like native OS controls: calm, premium, tactile, and low-friction.

Adjudication queue filters

Added workflow queue filters for:
These filters let senior review focus on the cases that mattered most. The May 2026 release supported adjudication routing, metadata, and exports. It did not depend on a separate senior-adjudication editing screen.

Exports

At the May 2026 checkpoint, JSON, CSV, and Markdown exports preserved structured review, suggested review, Employee Review, Human Guidance Evaluation, adjudication metadata, lifecycle state, tester identity, and prompt dirty/completion state.

Supabase persistence

At that checkpoint, Supabase persistence kept run lifecycle, case context, and prompt-level review evidence together closely enough for a saved review to hydrate without relying on browser-local state. Separate review records remained available for review history and oversight. This historical outcome does not publish the current database layout and does not prove that a live write is authorized today.

Doctrine established in May 2026

This release established a new Eval Labs principle:
This should be protected in future product work.

May 2026 product-surface refinement

After the May 2026 review-layer release, Eval Labs was refined into a clearer product surface:
  • top app shell owned page identity
  • in-page blog-style mastheads were removed from the app
  • Custom Prompt Test, Auto-generated Prompt Test, and Controlled Batch Runner were separate surfaces
  • Controlled Batch Runner was controlled readiness tooling; access at this checkpoint was owner/admin/evaluator, not tester
  • Auto-generated Prompt Test was the normal 50-prompt generated tester
  • Run History rows used a standardized two-zone layout
  • copy controls used Copy Session ID / Copy Deep Link patterns across key surfaces
  • Single Run Analysis gave read-only run-level evidence outside the Review Queue

May 2026 readiness doctrine

The May 2026 AI-reviewed platform readiness gate passed after 60 completed runs and 3,000 reviewed prompts. This established the following review-layer doctrine:
Protect this distinction in future release notes and onboarding language.