Skip to main content
This map reflects the deployed Eval Labs product at evaluationlabs.ai. The source and deployed identity verified for this reset is 6cbfd15f75330c414bfe79e0d2fab51ef8115102.

Declared route map

/sign-in and /sign-up are authentication portal routes. They are not evaluation workspaces.
Unknown paths fall back to the home route in the current client. Do not treat an undeclared path as a new product surface.

Current launcher set

The Launcher shows only cards allowed for the signed-in role: Custom-suite save, load, delete, and /lucia/custom/suites/:suiteId deep-link behavior is declared for evaluator and tester roles as well as owner/admin. The suite does not widen access to another user’s run evidence.

Surface definitions

Keyboard Review Mode

The shortcut reference at /keyboard-review-mode is available to every recognized role. Keyboard controls operate inside the Review Queue only when that role can access the referenced run. The shortcut surface does not bypass ownership or RLS boundaries.

Eval History and Review Queue

Eval History is the scoped run ledger. Review Queue is the prompt/item scoring workflow. Owner/admin may inspect shared evidence where privileged policy allows it; evaluator and tester routes remain limited to their own allowed work.

Analysis

The current product label is Analysis. /analysis is canonical and /experiments is a legacy inbound alias. Analysis is read-only platform evidence for owner/admin; it is not human quality approval. Single-run pages render Single Eval Analysis when complete analysis evidence is present and Eval Diagnostics when only inspection evidence is available.

Registry Diagnostics

Registry Diagnostics derives dataset membership and review-lane suggestions from existing eval evidence. Suggestions are not saved labels and do not prove final dataset membership. /registry-diagnostics is canonical. /dataset-diagnostics is a legacy alias.

Platform Runtime Registry

Platform Runtime Registry is a separate owner-only surface for immutable deploy and model-verification observations. Admin access to other privileged surfaces does not grant access here.

Behavioral Observatory

Behavioral Observatory persists structured human labels only after the Supabase save succeeds. Derived suggestions are starting context, not reviewer judgment.

Guest verification

Guest verification runs the booked-guest verification scenario pack and exposes its results. It is a distinct launcher and evidence surface; it is not the generic Auto-generated Prompt Test.

Fieldwork evaluation

The Launcher also declares an owner/admin Fieldwork Communication Baseline. Fieldwork evaluation uses dedicated durable contracts in addition to the generic run/item/review topology. Do not describe it as a normal custom or generated prompt suite.

Evidence addressability

Copy Session ID, Copy Eval ID, and Copy Deep Link controls make runs and review items addressable. Addressability supports debugging and handoff; it does not by itself prove that evidence is durable or authorized.

Surface distinction rule

Do not collapse these surfaces into one evidence claim.