Authority · evidence · effect · recovery

An agent should earn authority with evidence.

BitEvo audits the action layer of AI systems: what a workflow is allowed to change, what evidence must exist before it acts, how the external result is confirmed, and what the system does when certainty disappears.

Browser-local mapping first · staging/test by default · written Rules of Engagement before testing
DECISION TRACE / SYNTHETIC LAB-001 SYNTHETIC WORKED CASE
01
INTENTexternal change proposed
captured
02
AUTHORITYscope + approval matched
bounded
03
EVIDENCEsource freshness stopped advancing
insufficient
04
DECISIONdo not expand authority
constrain
WORKED FINDING CLASSCapability exceeded the evidence available to justify the action.

Synthetic worked scenario only. No customer environment, production execution or observed private infrastructure is represented.

Start with the smallest step

Map. Challenge. Then scope.

You do not need to buy an audit to understand the model. Draft one action chain locally, inspect worked evidence, and only turn the unresolved gates into scope when the decision is worth testing.

OBJECTOne action-capable workflow
QUESTIONWhat may it do — and why?
TEST10–20 agreed failure scenarios
DECISIONExpand · constrain · repair · retest
BitEvo vocabulary

Three concepts define the operating philosophy.

The terminology is intentionally narrow. It describes recurring control problems in action-capable workflows without pretending every problem is a vulnerability or every passing test is a safety guarantee.

Read the full doctrine →
01

Authority Budget

Every write permission, integration, autonomous step and retry path expands the authority surface. Capability does not automatically justify permission.

02

Evidence Before Effect

Critical effects should depend on evidence available at decision time. Logs after the fact can reconstruct an event; they cannot retroactively justify it.

03

False Green

A workflow can look healthy or successful while required evidence is stale, missing, attached to the wrong object or never confirmed externally.

The operating problem

Capability grows faster than proof.

Every new tool, write permission and autonomous step increases the authority surface. If the evidence chain does not grow with it, the team accumulates authority it cannot reliably justify, reconstruct or contain. BitEvo treats that gap as an engineering problem.

01

Before authority expands

A workflow is about to gain a new write-capable tool, integration, approval path or external side effect.

02

When evidence lags capability

The agent can do more than the team can reliably reconstruct, confirm or explain after the fact.

03

Before a pilot becomes infrastructure

A useful workflow is becoming operationally important, but duplicate execution, stale state or ambiguous approval would still be hard to contain.

The control loop

Permission is only the first gate.

A reliable workflow needs four linked answers: was the action authorized, was the evidence sufficient, did the intended external effect occur, and can the system recover when any of those answers becomes uncertain?

01

Authority

Define what the workflow is permitted to change, which object that permission applies to, and where human authorization is required.

02

Evidence

Identify the minimum evidence required before an action is allowed: inputs, approvals, source freshness, tool state and version context.

03

Effect

Verify the external result instead of treating an internal run status or tool response as proof that the intended change actually happened.

04

Recovery

Test what happens when certainty disappears: retry, interruption, partial execution, stale state, duplicate actions and rollback.

Proof discipline

A finding must change a decision.

BitEvo does not sell a generic “AI safety” score. A useful finding connects an observed failure to the authority boundary, the evidence available at decision time, the external effect and the owner action that follows.

Decision-artifact spine

A1–A4 structure the owner decision.

Authority Ledger, Evidence Contract, Finding Record and Decision Memo are the canonical decision-artifact spine. They organize the method; they are not the complete commercial delivery package.

01

Authority Ledger

A bounded map of critical actions, tools, objects, approvals and explicitly prohibited transitions.

02

Evidence Contract

The evidence that must exist before, during and after each critical action for the result to be trusted.

03

Finding Record

Reproducible findings tied to observed consequence, evidence quality and recovery behavior — with hypotheses kept separate.

04

Decision Memo

A bounded owner decision with supporting evidence, residual uncertainty, required control change and retest criterion.

Complete Primary Audit delivery package

Eight deliverables. One bounded audit.

The Primary Audit commercial package is exhaustive at eight deliverable classes; the A1–A4 artifact spine remains the internal decision structure used across them.

01

Executive report

02

Authority / effect map

03

Test inventory and scenario results

04

Reproducible evidence pack

05

Finding cards with impact, evidence, limitations, and confidence

06

Prioritized repair backlog

07

One retest

08

Evidence manifest / hashes where applicable

5 working days after complete evidence/access + written scope

Move from “it works” to “we know why it may act.”

No testing begins from a public form alone. Scope, access approval, allowed tests, prohibited actions, data handling and replay conditions are agreed first. The phases below describe the audit method, not a guaranteed day-by-day allocation.

See the full method →
  1. Phase A
    Map authority

    Identify actor and identity, tools and credentials, read/write surfaces, approval gates, target binding, scope/expiry and recovery path.

  2. Phase B
    Agree failure scenarios

    Select the bounded safe scenarios that challenge approval, object binding, scope, replay, external effects and recovery without turning the engagement into an unbounded pentest.

  3. Phase C
    Collect evidence

    Preserve the minimum necessary scope, configuration, identity, execution, state, approval, recovery, finding and retest receipts for material scenarios.

  4. Phase D
    Make the owner decision

    Translate accepted findings into observable behavior, impact, evidence class, limitations, repair recommendations, retest criteria and residual risk.

  5. After repair
    Retest the claim

    Re-run the agreed finding criteria after relevant remediation is available. One retest is included in the Primary Audit.

Primary engagement

Agent Authority & Evidence Audit

For teams that need to decide whether one write-capable workflow is ready for more authority — or needs to be constrained first.

$4,900fixed · 5 working days after complete evidence/access + written scope
Object1 staging/test workflow
Authority BudgetUp to 3 tools / APIs / MCP servers
Failure plan10–20 agreed scenarios
EvidenceAuthority Ledger + Evidence Contract + reproducible Finding Records
DecisionExpand / constrain / repair backlog
VerificationOne retest
FAQ

The boundary is part of the product.

What exactly is being audited?

A concrete agent workflow with an external effect. The object of the audit is the action chain — not an abstract model score.

What is an Authority Budget?

A BitEvo concept: consequential capability is treated as a finite permission surface that should expand only when evidence, confirmation and recovery controls justify the expansion.

What is a False Green?

A BitEvo term for an operational evidence mismatch: the workflow appears healthy or successful while a required source, confirmation or state condition is no longer trustworthy.

Is this a penetration test?

No. The standard engagement is a bounded engineering review in an agreed staging/test workflow. Production penetration testing is not included by default.

Do you certify an agent as safe?

No. The result is scoped engineering evidence and owner decisions — not certification, a universal trust score, or a guarantee of security or defect absence.

Do you need production secrets?

No. Public intake excludes credentials and customer secrets. Access, data boundaries and replay conditions are agreed separately in writing.

PRIMARY AUDIT$4,900 · 5 working days after complete evidence/access + written scopePrepare scope ↗