Find
Facts are reachable, current, consistent and available in an agent-usable form.
Versioned scenarios, deterministic assertions and reviewable evidence keep each result accountable.
A readiness signal combines discoverability, successful task completion, repeatability and efficiency.
Facts are reachable, current, consistent and available in an agent-usable form.
Controls, policies, prices and states have unambiguous meaning.
The allowed workflow reaches an independently verified final state.
Validation, auth, session and inventory failures expose a workable next action.
A missing agent-native surface cannot erase a successful browser result, and the reverse is also true.
Crawl access, citation surfaces, content consistency and machine-readable context.
Rendered interaction, semantic controls, keyboard path, sessions and responsive behavior.
Schemas, authorization, errors, idempotency, permissions and confirmation boundaries.
Unauthenticated surfaces with irreversible actions excluded.
Approved test accounts, fixtures and redacted evidence.
Complete journeys using controlled data and reversible operations.
Prompt wording alone never defines success. The contract states the goal, permissions, stop gate and assertions.
Session evidence explains whether a nominal success was fast, stable, repeatable and autonomous.
Page-level Cumulative Layout Shift is useful, but target stability is also inspected around each planned click or input.
Standard tasks run three times. Critical tasks run five times. Agent, model, tools, viewport, locale, auth state and policy restrictions remain attached to every run.
Results distinguish website, agent-specific, policy and infrastructure failure before scoring.
Readiness reflects observed capability. Confidence reflects coverage, repetitions, evidence quality and unresolved uncertainty.
P0 blocks a critical task or creates material safety risk.
P1 creates major repeatable friction or unreliable completion.
Artifacts are stored with stable IDs, timestamps, environment and retention rules.
WebMCP, MCP, OpenAPI, llms.txt, llms-full.txt and other agent surfaces are detected and lifecycle-labeled. Their absence is not automatically a failure.
Agent products, policies and models change. Results are versioned observations from a defined environment, not permanent compatibility guarantees.
Start with a scoped audit and keep every conclusion connected to evidence.