Skip to main content
Global Analysis, Single Run Analysis, Run History, and Team Review are evidence surfaces for completed Eval Labs work. They support platform and behavior inspection, but they do not replace human Lucia-quality review.

Global Analysis

Canonical route:
Legacy alias:
Global Analysis is owner/admin-only in the current access model. It is read-only behavioral/analytics evidence. It is AI-analyzed platform evidence, not human quality approval. Owner/admin should see shared persisted Eval Labs evidence here when Supabase hydration and RLS scope allow it.

Single Run Analysis

Canonical route:
Single Run Analysis is read-only analysis of one completed run/session. It can include:
  • run metadata
  • behavioral summaries
  • item rows
  • item-level review links
  • Copy Session ID
  • Copy Deep Link
When only compact local state is available, summary counts can render before full item-level cloud hydration.

Run History

Canonical route:
Run History is the scoped run ledger. It records completed/finalized run truth and may show scoped operational state. Run History truth means the UI agrees with the persisted run lifecycle and scoped account context. Owner/admin can inspect shared/global persisted run evidence. Evaluator and tester users should only see runs scoped to their own allowed work.

Team Review

Canonical route:
Team Review is the owner/admin oversight surface. It exists to inspect evaluator activity, review quality, missing checks, flags, recent work, and evidence that needs owner/admin attention. Team Review is not available to evaluator, tester, or unassigned users.

Registry Diagnostics

Canonical route:
Registry Diagnostics is read-only diagnostic evidence. It derives Dataset Registry membership suggestions and Human Review Queue 2.0 lane suggestions from existing Eval Labs data. It does not save labels or prove human approval.

Behavioral Observatory

Canonical route:
Behavioral Observatory is the saved behavioral label surface. It may start from derived run/review context, but its durable truth comes only from saved labels that reload from Supabase.

Truth distinction

Use this distinction everywhere:
The 60-run AI-reviewed gate proved Run History and Global Analysis were aligned with Supabase and local compact state for that readiness scope. It did not prove human Lucia-quality approval.

Staged hydration

Current dashboards should hydrate in stages:
This is a performance and truth-state requirement. Fast dashboards must still be source-backed. If deep evidence has not loaded yet, the UI or documentation should name that limitation instead of filling the gap with fake metrics. The same rule applies to Team Review and Global Analysis.