Eval Labs production is
evaluationlabs.ai, deployed from commit 6cbfd15f75330c414bfe79e0d2fab51ef8115102 as of 2026-08-12. Its Engine Development and Guest production targets are separate systems and must be verified independently.Freshness note — 2026-09-10. No Eval Labs release has shipped since 2026-07-29. On 2026-09-10 the documentation audit read
lucia-ai-eval-labs main and the production deploy at evaluationlabs.ai, both still at commit 6cbfd15f75330c414bfe79e0d2fab51ef8115102, the same identity verified on 2026-08-12. Claims on this page about the systems Eval Labs targets (the Engine Development deployment, its model configuration, the Guest Agent receipt) were not re-verified on 2026-09-10 and stay dated to the evidence they cite.Current target map
Eval Labs uses the Engine Development endpoint for Lucia and Fieldwork evaluation work:Active dev endpoint
Engine Development environment variable
Eval Labs uses a Vite environment variable:Guest production boundary
The deployed Guest verification runner targets the canonical Guest Agent production service athttps://guest.hellolucia.ai/.
Its live-run permissions and outbound-effect safeguards are private server-side controls. Before any authorized live verification, confirm their current state through the approved operator evidence boundary. Do not infer them from an Eval Labs commit, a client bundle, or the public Guest hostname.
The Guest Agent service and the Engine Development endpoint remain separate targets and require separate provenance.
Cross-origin boundary
The Development Engine must accept requests from the canonical Eval Labs production origin. The full cross-origin policy is an operator-controlled configuration and is not published here. Verify this boundary with a captured preflight observation. A successful preflight proves only that the observed origin, method, and request shape were accepted at that time.Live-path proof doctrine
A current live-path observation must bind the resolved target and environment, requesting origin, deployment identity, preflight result, application result, and timestamp. Gather that evidence through the authorized runtime-verification boundary so the observation remains attributable without publishing an operator runbook. If the production app points to Staging outside an explicitly authorized promoted-validation window, treat it as a target mismatch and do not trust the run as Development evidence. If a cross-origin failure is observed, retain the resolved endpoint, requesting origin, preflight status, deployment identity, and timestamp in the protected evidence record. Treat the run as unverified until an authorized operator reconciles the policy and a fresh observation succeeds. AnOPTIONS 204 observation proves only that the captured CORS preflight was accepted. A separate POST 200 can prove application-request routing and response delivery, but neither status alone proves that a model was invoked or identifies the provider-resolved model.
For current platform runtime/model proof, use the authorized runtime-verification boundary. The Engine Development and Guest Agent observations are independent; keep their provenance separate.
An Engine observation can show:
Verified, with configured and resolved model gpt-5.6-sol, model invoked true, deterministic path false, and fallback false.
That is a dated observation, not current model proof. No model-bearing verification or authenticated database-count query was performed on 2026-08-12. The current Eval Labs deployment SHA proves product identity only.
Historical CORS preflight observation — 2026-04-29
An archived DevTools capture recorded anOPTIONS request to the Engine Development operator-focus route returning 204 No Content.
That dated capture proves CORS preflight acceptance only. It contains no captured POST response, Lucia output, persistence result, or model provenance, and it is not proof of current configuration. The archived images are excluded from publication because they expose unnecessary browser and infrastructure detail.
