The gate blocks an AI-written change when required continuous integration evidence is missing, incomplete or unverifiable. Uncertainty does not count as approval.
An AI-written change can appear complete before it is safe to advance. The implementation may be coherent and review comments resolved. These signals do not establish how the candidate behaves in the delivery pipeline.
Our gate therefore fails closed. A positive result must be independently observed.
Not verified means not ready.
Evidence must sit outside authorship
The authoring process asks whether the requested change has been understood and implemented plausibly. Continuous integration (CI) asks whether the candidate satisfies the repository’s executable checks in a controlled environment.
An authoring agent can inspect its reasoning, edits and local tool output. It is not independent of the artefact under assessment. Creation and authorisation must not occupy the same trust boundary.
For CI evidence to support advancement, it must:
- Be attributable to the candidate under consideration.
- Cover the required checks.
- Be observed independently of the author’s claim.
A successful local command, an expectation that tests will pass or a partial pipeline result cannot fill a gap in this evidence.
The source data reports that 93.2% of evidence records were verified and that all verification was independent. It does not define an evidence record, the verification criteria or the system boundaries that enforce independence. The figures should therefore be read as reported rates rather than independently reproducible measurements.
We use claude-opus-5 to verify work at any risk. Model judgement cannot turn missing pipeline evidence into a pass.
Failing closed preserves the actual state
A fail-open gate converts uncertainty into approval. If evidence is unavailable, incomplete, no longer attributable to the candidate or cannot be independently verified, the change continues as though the required checks had succeeded.
A fail-closed gate blocks advancement. This can pause flow, but it prevents infrastructure faults and missing signals from being reported as engineering success.
It also lets leaders distinguish changes that passed the configured path from those waiting for evidence. Assurance does not vary silently between cases.
The source data reports uneven coverage of repository readiness controls:
| Readiness control | Reported rate |
|---|---|
| Presence of continuous integration | 85.7% |
| Presence of an agent trigger | 85.7% |
| Committed-secret checks | 28.6% |
| Default-branch protection | 14.3% |
The source also states that observation of CI checks passed throughout the assessment. It does not explain how that claim relates to the reported 85.7% rate for CI presence. No conclusion should be drawn from the relationship between those two measures without that definition.
These results do not justify weakening the gate. They show why the evidence boundary must remain explicit.
Repair has a defined limit
Most changes required no repair cycle. The source data reports that:
- 3.8% required a further pass.
- 3.8% exhausted the configured repair budget.
The source does not state the repair-budget size or the terminal state after exhaustion.
A repair limit is a safety boundary. Once acceptable evidence cannot be produced within that boundary, repeated agent activity does not count as increased confidence.
Review and planning cover different risks
Review, planning and CI address separate failure modes. Review can identify questionable logic, unsafe assumptions and mismatches with requested behaviour. Planning can improve the clarity of the intended implementation. Neither demonstrates that the resulting candidate builds, tests and integrates successfully.
The source data reports that automated reviewers raised all recorded findings. Of those findings:
- 57.3% were blocking.
- 29.8% were major.
- 12.9% were in categories not specified by the source.
Resolving a finding requires fresh CI evidence for the resulting candidate.
An approved plan covered 48.1% of started changes, with no reported revocations. First-pass acceptance was 77.8% where an approved plan existed and 72.7% without one. This is an association, not proof that planning caused the difference.
The source data also reports that verified evidence increased from 90.5% in the early build phase to 93.7% in the current phase. Over the same period, the share requiring a further pass fell from 5.1% to 2.5%. These outcomes do not establish causation.
Reporting limits
The source does not provide:
- The assessment period.
- Sample sizes or denominators.
- The assessed repository population.
- The definition of an evidence record.
- The criteria for marking evidence as verified.
- The candidate identifier or mechanism binding a candidate to a CI run.
- The rounding method.
The reported percentages indicate observed rates within the assessment. They are not sufficient for independent reproduction.
What we are changing next
We are concentrating readiness work on the controls with the weakest reported coverage:
- Default-branch protection.
- Committed-secret checks.
- Agent tooling.
We are also tightening the classification of evidence gaps and review findings. The aim is to make unknown states easier to diagnose without weakening the gate.
What this means
An AI-written change advances only when the required CI result is present, attributable and independently verified. If that evidence is missing or unverifiable, the operational state is blocked, not passed.
Talk to us
Building something adjacent to this?
Lender, broker, dealer group or technology partner — if this is the kind of problem you’re working on, we should compare notes.
support@autofintech.co.uk