Authority Budget
Every write permission, integration, autonomous step and retry path expands the authority surface. Capability does not automatically justify permission.
BitEvo audits the action layer of AI systems: what a workflow is allowed to change, what evidence must exist before it acts, how the external result is confirmed, and what the system does when certainty disappears.
Synthetic worked scenario only. No customer environment, production execution or observed private infrastructure is represented.
You do not need to buy an audit to understand the model. Draft one action chain locally, inspect worked evidence, and only turn the unresolved gates into scope when the decision is worth testing.
Draft the Authority Ledger and Evidence Contract in the browser. No testing authorization is created.
02 · PROOFInspect three synthetic worked action classes and see how evidence changes an owner decision.
03 · SCOPEWhen the unresolved gates matter, prepare a bounded scope brief. The public intake still does not authorize testing.
The terminology is intentionally narrow. It describes recurring control problems in action-capable workflows without pretending every problem is a vulnerability or every passing test is a safety guarantee.
Read the full doctrine →Every write permission, integration, autonomous step and retry path expands the authority surface. Capability does not automatically justify permission.
Critical effects should depend on evidence available at decision time. Logs after the fact can reconstruct an event; they cannot retroactively justify it.
A workflow can look healthy or successful while required evidence is stale, missing, attached to the wrong object or never confirmed externally.
Every new tool, write permission and autonomous step increases the authority surface. If the evidence chain does not grow with it, the team accumulates authority it cannot reliably justify, reconstruct or contain. BitEvo treats that gap as an engineering problem.
A workflow is about to gain a new write-capable tool, integration, approval path or external side effect.
The agent can do more than the team can reliably reconstruct, confirm or explain after the fact.
A useful workflow is becoming operationally important, but duplicate execution, stale state or ambiguous approval would still be hard to contain.
A reliable workflow needs four linked answers: was the action authorized, was the evidence sufficient, did the intended external effect occur, and can the system recover when any of those answers becomes uncertain?
Define what the workflow is permitted to change, which object that permission applies to, and where human authorization is required.
Identify the minimum evidence required before an action is allowed: inputs, approvals, source freshness, tool state and version context.
Verify the external result instead of treating an internal run status or tool response as proof that the intended change actually happened.
Test what happens when certainty disappears: retry, interruption, partial execution, stale state, duplicate actions and rollback.
BitEvo does not sell a generic “AI safety” score. A useful finding connects an observed failure to the authority boundary, the evidence available at decision time, the external effect and the owner action that follows.
In this worked scenario, a service surface is treated as apparently healthy while source freshness stops advancing. The modeled finding is not “the agent is unsafe.” It is that the workflow could appear operable after the evidence required to justify action had become stale. The modeled owner decision is therefore to constrain the Authority Budget until freshness and confirmation are made explicit.
Inspect evidence strength →Synthetic worked packs plus an explicit evidence-strength ladder showing what method, build and observed evidence can — and cannot — establish.
↗Internal dogfoodBounded internal self-audit evidence from BitEvo’s own control stack — not customer proof, independent certification, or production-wide security evidence.
↗Build + 10/10 StandardThe public engineering index and reusable quality contract, with machine gates kept separate from human visual review.
↗BitEvo DoctrineAuthority Budget, Evidence Before Effect and False Green — the vocabulary behind the audit model.
↗Audit methodologyThe four-plane model, finding anatomy, failure plan and Rules of Engagement.
↗Reviewed researchA curated public set of engineering notes on agent reliability, isolation, drift and tool boundaries.
↗BitEvo UniverseA static map of current public projects and proof surfaces — not a live operations dashboard.
↗Authority Ledger, Evidence Contract, Finding Record and Decision Memo are the canonical decision-artifact spine. They organize the method; they are not the complete commercial delivery package.
A bounded map of critical actions, tools, objects, approvals and explicitly prohibited transitions.
The evidence that must exist before, during and after each critical action for the result to be trusted.
Reproducible findings tied to observed consequence, evidence quality and recovery behavior — with hypotheses kept separate.
A bounded owner decision with supporting evidence, residual uncertainty, required control change and retest criterion.
The Primary Audit commercial package is exhaustive at eight deliverable classes; the A1–A4 artifact spine remains the internal decision structure used across them.
No testing begins from a public form alone. Scope, access approval, allowed tests, prohibited actions, data handling and replay conditions are agreed first. The phases below describe the audit method, not a guaranteed day-by-day allocation.
See the full method →Identify actor and identity, tools and credentials, read/write surfaces, approval gates, target binding, scope/expiry and recovery path.
Select the bounded safe scenarios that challenge approval, object binding, scope, replay, external effects and recovery without turning the engagement into an unbounded pentest.
Preserve the minimum necessary scope, configuration, identity, execution, state, approval, recovery, finding and retest receipts for material scenarios.
Translate accepted findings into observable behavior, impact, evidence class, limitations, repair recommendations, retest criteria and residual risk.
Re-run the agreed finding criteria after relevant remediation is available. One retest is included in the Primary Audit.
For teams that need to decide whether one write-capable workflow is ready for more authority — or needs to be constrained first.
A concrete agent workflow with an external effect. The object of the audit is the action chain — not an abstract model score.
A BitEvo concept: consequential capability is treated as a finite permission surface that should expand only when evidence, confirmation and recovery controls justify the expansion.
A BitEvo term for an operational evidence mismatch: the workflow appears healthy or successful while a required source, confirmation or state condition is no longer trustworthy.
No. The standard engagement is a bounded engineering review in an agreed staging/test workflow. Production penetration testing is not included by default.
No. The result is scoped engineering evidence and owner decisions — not certification, a universal trust score, or a guarantee of security or defect absence.
No. Public intake excludes credentials and customer secrets. Access, data boundaries and replay conditions are agreed separately in writing.