Conversational CodingA Backstory project Install the harness →
Menu

The agent saying “done” isn’t evidence that the work is done.

Engineering With AI separates planning, Build, standards, tests, independent review, Delivery, Manual QA and learning. Each stage has a purpose, an evidence contract and a visible human boundary.

Fourteen governed stages

IntentReconcilePlanPattern validationTest planBuildStandardsTest executeCode reviewDeliveryRetro

Conditional independent-validation stages stay explicit when a provider is unavailable. Missing review isn’t silently relabelled as passed.

Fast code is useful. Unreviewable code is a debt with good timing.

A vibe-coded feature can look complete while permissions, failure handling, performance, accessibility, operations and the original acceptance criteria remain untouched.

The harness gives each concern a place in the route. It can reduce ceremony for low-risk work and deepen it for sensitive changes, but it doesn’t let enthusiasm collapse every decision into one green check.

Build begins only after an exact scope is approved. Manual QA and release remain human decisions after the code and automated evidence exist.

The things that stop a successful demo becoming a fragile product.

01

Guarded fourteen-stage delivery

Intent, reconciliation, plan, pattern validation, test design, independent review, Build, standards sweep, test execution, code validation, Delivery and Retro each retain their own evidence and transition rules.

02

Exact Build approval

A named person approves a specific plan and test boundary. The agent can’t infer permission from “carry on”, a passing plan check or an earlier conversation.

03

Dependency-aware execution leases

Parallel workers claim bounded tasks against current state, with conflicts, expiry and recovery handled by the delivery contract rather than optimism.

04

Deterministic standards evidence

Project rules are routed into the plan and checked again across the complete change. An attractive implementation can’t quietly sidestep an inconvenient standard.

05

Persona-driven test scenarios

Relevant roles challenge journeys, failure paths and acceptance. Only human-reviewed scenarios become test oracles; a persona doesn’t approve its own expectations.

06

Solution Readiness Review

The harness assembles what is present and missing across intent, impact, standards, tests, Manual QA, security, technology, operations and guidance. It prepares a named human decision; it doesn’t certify or release the system.

Illustrative release review

The tests passed. Would you release it?

What the evidence supports

  • Implementation: the changed code matches the approved scope.
  • Automated tests: the agreed checks passed against this revision.
  • Code review: the reviewer’s findings and responses are recorded.

What is still missing

  • Security: the required independent review didn’t run.
  • Manual QA: nobody has yet completed the real user journey.
How security findings join the review

A required profile identifies the evidence, freshness and reviewer. Configured providers such as Agentic Security, DeepSec or a VVAH hand-off feed a common evidence record. Findings then need an explicit decision: remediate, accept the risk, reject a false positive or escalate. Stale evidence and unresolved blockers remain visible.

What you can ask when the code looks finished.

Did we build the approved thing?

Delivery state binds the current plan, tests, standards and approval fingerprints rather than trusting a summary written after the fact.

Did an independent reviewer actually run?

Configured external review records availability, cycles and output. Unsupported or unavailable review stays visible.

What still needs a person?

Manual QA, specialist judgement, representative-user evidence, risk acceptance, business acceptance and release authority remain separate named routes.

What happens after a finding?

A valid security finding can become its own governed remediation intent. Accepted risk and false-positive decisions remain accountable and append-only.

Make “done” something the team can explain.

Keep the speed. Add the evidence, independent challenge and human decisions that let the result survive beyond the demo.