The agent saying “done” isn’t evidence that the work is done.
Engineering With AI separates planning, Build, standards, tests, independent review, Delivery, Manual QA and learning. Each stage has a purpose, an evidence contract and a visible human boundary.
Fourteen governed stages
Conditional independent-validation stages stay explicit when a provider is unavailable. Missing review isn’t silently relabelled as passed.
Fast code is useful. Unreviewable code is a debt with good timing.
A vibe-coded feature can look complete while permissions, failure handling, performance, accessibility, operations and the original acceptance criteria remain untouched.
The harness gives each concern a place in the route. It can reduce ceremony for low-risk work and deepen it for sensitive changes, but it doesn’t let enthusiasm collapse every decision into one green check.
Build begins only after an exact scope is approved. Manual QA and release remain human decisions after the code and automated evidence exist.
The things that stop a successful demo becoming a fragile product.
Guarded fourteen-stage delivery
Intent, reconciliation, plan, pattern validation, test design, independent review, Build, standards sweep, test execution, code validation, Delivery and Retro each retain their own evidence and transition rules.
Exact Build approval
A named person approves a specific plan and test boundary. The agent can’t infer permission from “carry on”, a passing plan check or an earlier conversation.
Dependency-aware execution leases
Parallel workers claim bounded tasks against current state, with conflicts, expiry and recovery handled by the delivery contract rather than optimism.
Deterministic standards evidence
Project rules are routed into the plan and checked again across the complete change. An attractive implementation can’t quietly sidestep an inconvenient standard.
Persona-driven test scenarios
Relevant roles challenge journeys, failure paths and acceptance. Only human-reviewed scenarios become test oracles; a persona doesn’t approve its own expectations.
Solution Readiness Review
The harness assembles what is present and missing across intent, impact, standards, tests, Manual QA, security, technology, operations and guidance. It prepares a named human decision; it doesn’t certify or release the system.
Illustrative release review
The tests passed. Would you release it?
What the evidence supports
- Implementation: the changed code matches the approved scope.
- Automated tests: the agreed checks passed against this revision.
- Code review: the reviewer’s findings and responses are recorded.
What is still missing
- Security: the required independent review didn’t run.
- Manual QA: nobody has yet completed the real user journey.
How security findings join the review
A required profile identifies the evidence, freshness and reviewer. Configured providers such as Agentic Security, DeepSec or a VVAH hand-off feed a common evidence record. Findings then need an explicit decision: remediate, accept the risk, reject a false positive or escalate. Stale evidence and unresolved blockers remain visible.
What you can ask when the code looks finished.
Delivery state binds the current plan, tests, standards and approval fingerprints rather than trusting a summary written after the fact.
Configured external review records availability, cycles and output. Unsupported or unavailable review stays visible.
Manual QA, specialist judgement, representative-user evidence, risk acceptance, business acceptance and release authority remain separate named routes.
A valid security finding can become its own governed remediation intent. Accepted risk and false-positive decisions remain accountable and append-only.
Make “done” something the team can explain.
Keep the speed. Add the evidence, independent challenge and human decisions that let the result survive beyond the demo.