Матеріали

Agency audit delivery

Проводьте доказові client audits повністю.

Перетворіть один client journey на versioned baseline, reviewable evidence, owner-ready backlog і comparable retest.

Aiscovery Research16 хв читанняОстання перевірка: 26 серпня 2026 р.
Один authorized client website journey проходить baseline evidence, cause classification, engineering handoff і verified retest

Що це допоможе оцінити

  • Як перетворити client authorization на bounded technical audit contract
  • Яке evidence робить findings перевірюваними для account та engineering teams
  • Як порівнювати baseline і retest без приховування model або environment changes

Визначте authorized client engagement

Почніть із named client owner, owned domains, approved environments, business-critical journey, target users і рішення, яке має підтримати audit. До active run зафіксуйте included та excluded routes, public, authenticated або sandbox mode, test dates, credentials owner, data classification, production stop gate і responsible disclosure path.

Перетворіть engagement на один immutable manifest. Manifest має визначати persona, prompt, fixtures, locale, viewport, browser, agent profiles, allowed actions, expected terminal state і deterministic assertions. Portfolio-level scan ніколи не розширює authorization за межі domains та actions, явно погоджених для цього клієнта.

  • OWNERSHIP: визначте, хто може authorize domain, accounts і test data.
  • BOUNDARY: розділіть passive discovery, browser interaction і state-changing sandbox work.
  • STOP: назвіть exact action, яку production не повинен submit.
  • VERIFY: визначте state, що незалежно доводить task completion або safe termination.

Створіть versioned scenario pack

Повторно використовуйте contract shape, а не припущення про client stack. Кожен scenario потребує stable ID, goal, prerequisites, test data, allowed actions, stop gate, expected result, criticality та evidence requirements. Зберігайте domain-specific fixtures і assertions поруч зі scenario, а не ховайте їх у executor prompt.

Версіонуйте website release або deployment, environment, source snapshots, agent product і model, tool access, policy restrictions, viewport, locale, consent state та authentication fixture. Плануйте repeated runs до execution, щоб retry не перетворився непомітно на звичайний first-pass success.

Запустіть evidence-safe baseline

Почніть із deterministic browser baseline, потім запускайте named agent profiles для того самого task contract. Ізолюйте cookies, local storage, credentials і mutable fixtures між runs. Зберігайте successful і failed attempts, щоб клієнт відрізняв repeatability від одного favorable outcome.

Фіксуйте session recording, ordered actions, screenshots, accessibility tree, DOM snapshots, console і network signals, source citations, final-state assertion, completion time, waits, retries, interventions, cost, Cumulative Layout Shift і target movement. Приховуйте secrets та client data до потрапляння artifact у shareable report, а потім застосовуйте погоджені access і retention policy.

  • Використовуйте synthetic identities та reversible fixtures для authenticated або sandbox runs.
  • Фіксуйте observation і action timing навколо unstable controls.
  • Доведіть відсутність unintended payment, message, booking, deletion або external notification.

Розділяйте cause і confidence

Класифікуйте primary cause як website, agent-specific, agent policy, infrastructure або inconclusive. Deterministic baseline, який проходить при failure одного named agent, відрізняється від rendered state, що блокує кожного executor. Зберігайте exact step, decision, response або policy boundary, що підтримує classification.

Звітуйте readiness і confidence окремо. Repetitions, fixture quality, evidence completeness та environment stability підвищують confidence, а retries і human interventions залишаються видимим friction. Flaky pass не дорівнює clean first-pass result, а untested area не є evidence успіху.

Створіть owner-ready backlog

Пов'яжіть кожен priority finding з affected customer task, observed result, expected result, reproduction steps, evidence artifact, primary cause і acceptance criteria. Додайте impact, effort, likely owner і dependencies, щоб account teams пояснили business consequence, а implementation teams перевірили fix.

Відокремлюйте website remediation від agent compatibility notes, policy limitations та infrastructure incidents. Описуйте optional technologies на кшталт llms.txt, OpenAPI, MCP або WebMCP через observed state і task impact. Не перетворюйте їхню наявність на maturity claim, а відсутність на automatic failure.

  • Пріоритезуйте unsafe side effects, blocked critical journeys і data exposure до minor friction.
  • Залишайте confidence, limitations і not-run areas видимими в executive summary.
  • Уникайте guaranteed visibility, universal agent support та unsupported automation claims.

Проведіть comparable retest

Пов'яжіть кожен retest із baseline manifest та approved findings. Повторно використовуйте той самий scenario ID, fixtures, stop gate і assertions, де це можливо. Фіксуйте кожну intentional change у website, environment, browser, agent version, policy або test data, щоб comparison не приписав moving executor до client fix.

Порівнюйте verified outcome, repetitions, completion time, waits, retries, interventions, CLS, target movement і side effects. Finding є resolved лише після проходження acceptance criteria з reviewable evidence. New regression стає окремим finding, а не ховається всередині improved aggregate score.

Керуйте portfolio evidence і claims

Ізолюйте client workspaces, credentials, recordings та exports. Визначте, хто може переглядати raw traces, отримувати redacted reports, approve excerpts і request deletion. Portfolio summaries мають агрегувати compatible statuses без розкриття client URLs, prompts, screenshots або business data за межами agreed audience.

Показуйте лише capabilities, які service справді підтримує. White-label reports, automated exports, scheduled monitoring і cross-client benchmarks залишаються explicit product states, а не implied promises. Кожен client-facing claim має вести до dated scope, tested surface, executor profile і confidence level.

  • Використовуйте standard report structure, зберігаючи client-specific limitations.
  • Не порівнюйте scores із incompatible scopes без clear qualification.
  • Зберігайте evidence протягом agreed retest window, а потім виконуйте deletion rules.

Agency audit delivery checklist

  • Зафіксуйте client owner, authorized domains, environments, routes та excluded surfaces.
  • Оголосіть public, authenticated або sandbox mode, allowed actions і production stop gate.
  • Визначте stable task ID з persona, fixtures, expected state та assertions.
  • Версіонуйте website release, sources, browser, agent, tools, policy, viewport і locale.
  • Запустіть isolated deterministic baseline до named agent profiles.
  • Зберігайте кожні repetition, retry та intervention без переписування first-pass result.
  • Фіксуйте session, accessibility, DOM, console, network, timing, CLS і final-state evidence.
  • Знеособлюйте secrets та client data, потім застосовуйте evidence access і retention rules.
  • Окремо класифікуйте website, agent-specific, policy, infrastructure та inconclusive causes.
  • Звітуйте readiness, confidence, limitations і not-run areas як distinct fields.
  • Дайте кожному finding reproduction steps, task impact, owner і acceptance criteria.
  • Фіксуйте emerging technologies як observed states, а не automatic scoring requirements.
  • Пов'яжіть retest з baseline manifest і розкрийте кожну changed variable.
  • Доведіть resolved findings через matching assertions і зберігайте new regressions окремо.
  • Публікуйте лише supported agency capabilities та authorized portfolio evidence.

Первинні delivery і testing sources

Перевірте сигнал у реальному journey.

Наявність технології важлива лише тоді, коли покращує verified outcome. Пов'яжіть сигнал із task, evidence і final-state assertion.