# Implementation result: durable pure human wait (Engine)

Scope: `vl-workflow-engine` 4.21.8, isolated library copy. Engine-only implementation
of the pure-wait part of `HOST-COLD-RESUME-PROPOSAL.en.md`: a custom handler that
only waits for a human can survive a process death and be re-entered on a NEW Engine
with the SAME run, the SAME request token, and the answer delivered once. No
bound-effect kind, no token replacement, no quota receipt machinery.

This is not a release claim and not Matrix L05 acceptance. The tests are engine
fixtures driving `Engine` directly with an in-memory host store; Flow's durable
approval integration is the lead's separate work.

## Changed files

| File | Change |
| --- | --- |
| `lib/human-wait.js` | New, dependency-free helper: record shape, structural re-entry checks, taint rule, validator descriptor and verdict check, canonical JSON + pure-JS SHA-256 digest. No Node `crypto` import (browser build unaffected). |
| `lib/engine.js` | `options.humanWaitValidator`; `bindHumanWait`, `resolveHumanWait`, `digestHumanWaitRequest` (static + instance); `_releaseHumanWait` around the custom handler call; planner hook `_planHumanWait` in the `started` branch of `_planRecovery`; `_planLegacyHumanWaits` for the no-frame resume path; `_verifyHumanWaits` awaited in `executeFrom` before `workflow_start`; `humanWaits` restored from the checkpoint. Module also exports `digestHumanWaitRequest`. |
| `lib/types.js` | `ExecutionContext`: `_humanWaits` (restored from `humanWaits`), transient `_activeHumanWait`; taint observation in both `emitEvent` bodies; additive `humanWaits` field in `checkpoint()`; branch-local `events` counter on handler chains in `observeRecoveryLifecycle`; `ChildExecutionContext` delegates `_humanWaits` and owns `_activeHumanWait`. |
| `test/host-human-wait.js` | Focused suite, 16 cases (12 original + 4 regression cases for review findings 1-4, see `HUMAN-WAIT-FIXES.en.md`). |
| `test/host-human-wait-final.js` | Final-review regression suite, 13 cases (F1-F5, see `../../ENGINE-FINAL-FIXES.en.md`). |

Review corrections (2026-09-13, `HUMAN-WAIT-REVIEW.en.md` findings 1-4) are
described in `HUMAN-WAIT-FIXES.en.md`: monotonic `status` with the host delivering
the actual answer, host-confirmed `payloadDigest` for a resolved record, taint over
the whole handler span, and the owner identity of a direct branch nested inside a
parallel branch. Finding 5 (schema gate on the legacy path) is not implemented here;
the root added strict recovery schema/shape validation before any Engine state
mutation separately.

Final-review corrections (2026-09-13, `FINAL-REVIEW.en.md` findings F1-F5) are
described in `../../ENGINE-FINAL-FIXES.en.md` and are folded into the sections
below: the answer content is required on both resolved paths (F1), an interrupted
wait survives a cooperative handler RETURN (F2), a validated re-entry is not
re-gated by `if:` or by the result cache and a dropped record drops its verdict
(F3), purity is an explicit neutral classification of engine event types (F4), and
the record carries its binding `runId` with the branch identity re-asserted at the
hand-over (F5).

Not touched: cancellation logic, `lib/runtime-snapshot.js`, malformed-checkpoint
validation, `completedBranches` scoping, versions, `scripts/`, existing tests,
`index.js`, `package.json`, dependencies. Checkpoint `_version` stays 4; `humanWaits`
and `events` (on handler chains) are additive fields.

## Test results

```
node test/host-human-wait.js         16 passed, 0 failed
node test/host-human-wait-final.js   13 passed, 0 failed   (2 passed, 11 failed before the F1-F5 fixes)
node scripts/run-all-tests.mjs       21 suites, 0 failed
```

The two `lib/runtime-snapshot.js` cache failures recorded in earlier revisions of
this document are no longer present in this baseline: `test/retry-cache.js` and
`test/cache-resume.js` pass. All 21 suites are local builtins and need no network,
credentials or provider calls.

## Host API (exact usage)

Handler side. The binding must be the handler's first action on this branch: no
event emitted and no engine sub-step run on the handler chain before it.

```js
const engine = new Engine(workflow, {
  customHandlers, onEvent, onCheckpoint,
  humanWaitValidator: (descriptor) => host.validateHumanWait(descriptor)   // code, registered at construction
});

async function humanWaitHandler(engine, ctx, step) {
  const request = { tool: toolName, stepId: step.id, contract };   // what the human is asked; digested, not stored
  const wait = engine.bindHumanWait(ctx, step, {
    token: `hin_${step.id}_${Date.now()}`,        // used only on first entry; ignored on re-entry
    request,
    allowedHostEvents: ['tool_start', 'tool_done'] // host event types that may be emitted while pending
  });
  // wait is one of:
  //   { created: true,   token }                        first entry; record checkpointed before this returns
  //   { reattached: true, token, verdict }              cold re-entry, host still pending: SAME token; do not mint a new request
  //   { resolved: true,  token, resolution: { requestId, payload } }   the durable answer: map it, do not ask.
  //       Either resolveHumanWait persisted it before the crash, or the host validator delivered it
  //       (answer persisted by the host, process died before resolveHumanWait); in both cases the
  //       record is `resolved` and checkpointed before this returns.
  if (wait.resolved) return applyAnswer(ctx, step, wait.resolution.payload);
  if (wait.created) await host.requests.upsert({ runId: ctx.workflowID, stepId: step.id, waitToken: wait.token, request });
  ctx.emitEvent({ type: 'tool_start', stepID: step.id, payload: { waitToken: wait.token, resumed: Boolean(wait.reattached) } });
  const answer = await host.waitForAnswer(wait.token);          // host delivers Base's durable answer { requestId, ... }
  engine.resolveHumanWait(ctx, step, { requestId: answer.requestId, payload: answer });   // durable BEFORE outputs
  ctx.emitEvent({ type: 'tool_done', stepID: step.id, payload: {} });   // declared in allowedHostEvents: the handler is still in flight
  return applyAnswer(ctx, step, answer);
}
```

A handler may also leave cooperatively on shutdown — `if (!answer) return;` when
`ctx.signal` fires or the run is paused — instead of throwing. That is treated
identically: the record survives, the step is not closed, and the same request is
re-attached on the next resume (see **Record lifecycle** below).

Validator side. The descriptor is frozen and carries no payload. The validator
reports what the host ACTUALLY persisted for that token (it must not echo
`d.status` or `d.resolution` back — and for a resolved record it now cannot: the
required `payload` is not in the descriptor):

```js
// descriptor: { runId, stepID, kind: 'pure-wait', token, status: 'pending'|'resolved', requestDigest,
//               owner, allowedHostEvents, boundAt, reattachCount,
//               resolution: { requestId, payloadDigest, resolvedAt } | null }
async function validateHumanWait(d) {
  const req = await base.humanRequests.find({ runId: d.runId, stepId: d.stepID, waitToken: d.token });
  if (!req) return { verdict: 'refuse', reason: 'no_such_request' };
  if (!PURE_WAIT_TOOLS.has(req.tool)) return { verdict: 'refuse', reason: 'not_a_pure_wait_tool' };
  if (Engine.digestHumanWaitRequest({ tool: req.tool, stepId: d.stepID, contract: req.contract }) !== d.requestDigest) {
    return { verdict: 'refuse', reason: 'request_mismatch' };
  }
  const verdict = { verdict: 'reattach', token: d.token, requestDigest: d.requestDigest, status: req.status };
  if (req.status === 'resolved') {
    verdict.requestId = req.answerId;                                   // the durable answer id (Base)
    verdict.payload = req.answer;                                       // the persisted answer content
    verdict.payloadDigest = Engine.digestHumanWaitRequest(req.answer);  // digest of what the host stores
    // verdict.resolvedAt = req.answeredAt;                             // optional ISO string
  }
  return verdict;
}
```

Verdict rules (`humanWaitVerdictRefusal`; the first failing rule is the `detail`):

| Checkpoint record | Host verdict | Result |
| --- | --- | --- |
| any | not an object / `verdict !== 'reattach'` | `verdict_not_object` / the host's `reason` |
| any | other `token` / `requestDigest` | `token_mismatch` / `request_digest_mismatch` |
| `pending` | `status: 'pending'` | accepted: `bindHumanWait` returns `reattached` |
| `pending` | `status: 'resolved'` + `requestId` + `payload` + matching `payloadDigest` | accepted: the host answer becomes the record's resolution at `bindHumanWait` (checkpointed), which returns `resolved` |
| `pending` | `status: 'resolved'` without `requestId` / `payloadDigest` / `payload`, or `digest(payload) !== payloadDigest` | `request_id_missing` / `payload_digest_missing` / `payload_missing` / `payload_mismatch` |
| `resolved` | `status: 'pending'` (or absent request) | `status_mismatch`: the checkpoint claims an answer the host does not hold |
| `resolved` | `status: 'resolved'` + same `requestId` + same `payloadDigest` + `payload` digesting to it | accepted: `bindHumanWait` returns `resolved` with the recorded payload (content proven by the host) |
| `resolved` | other `requestId` / other `payloadDigest` / `payload` not digesting to `payloadDigest` / no `payload` | `request_id_mismatch` / `payload_digest_mismatch` / `payload_mismatch` / `payload_missing` |
| any | any other `status` | `status_mismatch` |

`status` is monotonic: pending → resolved is accepted, resolved → pending never.
**A resolved verdict must carry `payload` on BOTH paths** (F1). The descriptor
publishes a resolved record's `requestId` and `payloadDigest` but never its content,
so identity alone can be produced by a validator that answers from the descriptor
instead of from its own store; requiring the content makes the resolved path
host-authoritative in the same way the pending path already was. Errors thrown by
the validator are refusals with the message as `detail`. A refused verdict never
mutates the record.

**Residual trust (not removable by omitting descriptor fields).** `humanWaitValidator`
is CODE the host registers at construction. A hostile or broken validator can
fabricate any verdict, including a payload it never held; no descriptor shape
prevents that. What the engine guarantees is narrower and unchanged: checkpoint data
alone never authorizes re-entry or the answer content, a truthful validator must
produce evidence it can only have from its own store, and the first failing rule
refuses without mutating anything.

Error codes thrown by the API: `HUMAN_WAIT_SPEC` (bad token/request/allowlist,
missing `requestId`, file targets on the step, request differs on re-entry,
re-entered on a different branch than the bound one),
`HUMAN_WAIT_UNTRACKED` (not a tracked custom handler step, or called outside the
in-flight handler), `HUMAN_WAIT_NOT_FIRST_ACTION`, `HUMAN_WAIT_LIVE` (record exists
and no accepted verdict in this `executeFrom`: second attach or second executor),
`HUMAN_WAIT_UNBOUND`, `HUMAN_WAIT_RESOLVE_CONFLICT` (a resolved record receives a
different answer id, or the same answer id with different content; only the
identical id and content is an idempotent no-op).

## Resume contract (first failure wins)

1. Existing planner rules unchanged (schema, malformed frames, missing step, orphan
   frames, `in_flight_error_handler`, Loop re-entry via `replaySafe`, which still
   returns false for any custom step: a wait inside an interrupted iteration is
   `in_flight_iteration`).
2. A `started` custom step needs a record in `humanWaits` for its id; otherwise
   `in_flight_custom_handler`, exactly as before.
3. `human_wait_invalid` with `detail`: record shape (`kind`, `token`, 64-hex
   `requestDigest`, `status`, `run_id_shape`, resolution shape and payload digest),
   `not_custom`, `step`, `run` (a record that names a `runId` is re-entered only by
   that run; records written before the field keep today's behaviour),
   `owner` (record owner must equal what the binding context saw: the branch
   record for a branch on its own child context; for a direct single-child branch the
   owner of its fan-out frame, i.e. `null` at the root or the enclosing branch record
   when the fan-out itself sits inside a parallel branch), `nested_steps` (handler
   chain must be empty), `tainted:<event>` (any effect while the handler was in
   flight, before or after the answer was persisted), `file_targets`.
4. Records the planner did not reach (`human_wait_orphan`) fail the whole resume.
5. No validator registered: `human_wait_unverified`.
6. Validator did not satisfy the verdict rules above: `human_wait_refused` with
   `verdict` and `detail`.
7. Only then does execution start. `bindHumanWait` consumes the one-shot verdict on
   re-entry, re-checks the request digest and re-asserts that the context taking the
   wait over is the branch the planner validated (`HUMAN_WAIT_SPEC` otherwise);
   `HUMAN_WAIT_LIVE` when there is no verdict. A pending record with a resolved
   verdict takes the host's answer as its resolution at that point (checkpointed
   before the handler receives it).
8. A step whose wait was accepted is re-entered as an IN-FLIGHT step, not as a fresh
   entry: `step.if` is not re-evaluated and a result-cache entry does not stand in
   for the step. Either gate would end the step without reaching the handler, which
   would leave the record unreleased (every later resume of that run then fails
   closed with `human_wait_orphan`) and would answer — or silently abandon — a human
   request that is still open at the host. Same reasoning as a Branch continuation,
   which records `selectedID` rather than re-evaluating its condition on resume. The
   breakpoint gate is not part of this rule: it is evaluated before it and cannot
   fire at a wait re-entry anyway (a parallel branch context exposes no `breakpoints`
   set, and on the legacy path the wait step is `currentStepID`, which `executeFrom`
   one-shot-releases).

Legacy resume (no open fan-out frame, `currentStepID` re-executed): a record is
admitted only for `currentStepID` and only if the persisted root cursor shows that
step `started` with its handler chain; a wait under a Loop, Branch or onError
continuation is `human_wait_orphan`. The same validator step applies.

Unresolved resumes pause with `cold_resume_unresolved` and emit a checkpoint that
still carries the record, so the host can register the validator and retry.

## Evidence the engine records (branch-local)

- **First action.** `observeRecoveryLifecycle` counts events emitted on the branch's
  own handler chain (`nested.events`), excluding `var_changed`. Siblings emit on
  their own chains and never move it; no global `eventSeq` comparison is used.
  A handler may `await` before binding as long as its own branch did nothing.
- **Taint while the handler is in flight.** Classification fails closed on both arms,
  from the binding until the handler leaves, including after `resolveHumanWait`:
  - an **engine** event type is clean only if it is in `HUMAN_WAIT_CLEAN_ENGINE_EVENTS`
    (`workflow_start|done|failed|cancelled|paused`, `breakpoint_hit`, `step_skipped`,
    `step_print`, the three `workflow_preflight_*` and the two `step_guard_*` types).
    Every other engine type sets `tainted` + `taintEvent` — including the ones an
    older denylist exempted by omission (`channel_recv`, which destructively consumes
    a checkpointed message; `actor_done` / `actor_failed` / `actor_message`;
    `child_run_done` / `child_run_failed`; `interact_resolved`; `human_gate_decided`;
    `review_timeout`; `llm_error`; `llm_thinking`; `segment_verification_receipt`;
    `pause_resumed`; `pause_timeout`; `pause_rejected`; `review_rejected`; every
    `swarm_*` and `loop_branch_*` type) and every engine event type added in the
    future;
  - a **host** event type is clean only if the binding declared it in
    `allowedHostEvents`. That declaration governs host types only: naming an engine
    type in it does not buy a replay permission for an engine effect.

  `var_changed` is neutral on both arms. The step's own `step_done` is emitted after
  the release and never taints. A tainted wait still completes in-process; it is
  refused on cold resume.
- **Nested engine steps** land on the handler chain (`nested.steps`) and are refused.
- **Record lifecycle.** While the run is still Running the step is finished either
  way: a normal return deletes the record (the next Loop iteration binds afresh) and
  so does a throw (a genuine failure; a retry binds afresh). Deleting the record also
  drops its one-shot verdict, so a spent permission can never authorize a NEW request
  at the same step id. If the run was interrupted (aborted, paused, stopped) the
  record is KEPT for a cold resume **however the handler left** — a cooperative
  handler that RETURNS on `ctx.signal` has an open human request exactly like one
  that throws, so the engine's durability must not depend on the throw. In that case
  the step is also not closed: no `step_done`, no `completedSteps` entry, no `next`
  and no enclosing join; the frame stays `started` with its handler nesting, which is
  the state `_planHumanWait` re-enters. Custom handlers that hold no wait record are
  unaffected and close their step exactly as before.

## Unresolved edges

1. **Purity is asserted by host code.** A native action that emits no event is
   invisible; the validator's tool allowlist is the second check. Unchanged from the proposal.
2. **`var_changed` is neutral.** A handler may set variables before binding or while
   pending without tainting; those writes are repeated on re-entry. Pure, but a host
   must not treat a variable write as a side effect.
3. **Untracked contexts cannot bind.** Parallel-Loop branches (`_loopBranchState`) and
   strict-segment candidate contexts are not recovery-tracked; `bindHumanWait` throws
   `HUMAN_WAIT_UNTRACKED` there. A host handler that may run in those contexts must
   catch the code and fall back to an unbound wait (which fails closed on resume, as today).
4. **Re-entry mints and discards a token.** The handler's fresh `token` argument is
   ignored on re-entry; the request must digest to the bound one, so a workflow edit
   that changes the contract refuses re-entry (`HUMAN_WAIT_SPEC`) rather than re-asking.
5. **Resolved record before `step_done`** re-runs the handler's own output mapping on
   re-entry; pure only because file targets are refused at bind and at plan time and
   because any event-emitting effect after the resolve taints the record.
6. **Legacy path and `human_wait_orphan` for a resolved record** at a step the run has
   moved past cannot occur through the API (normal return deletes the record) but a
   hand-edited checkpoint produces a fail-closed pause, not a repair.
7. **Two executors.** The second `executeFrom` on the same checkpoint gets its own
   verdict from the host; the engine cannot see the other process. Refusing the second
   attach is the host store's job (the proposal's `HANDLER_BINDING_LIVE` at Authority);
   the engine refuses only a second attach within one `executeFrom`.
8. **Exports.** `index.js` and `package.json` `exports` were not edited, so hosts reach
   the digest via `Engine.digestHumanWaitRequest` or `require('vl-workflow-engine/engine')`.
   Both suites are registered in `scripts/run-all-tests.mjs`.
9. **Run identity is defence in depth, not the authority.** The record's `runId` lets
   the engine refuse a record that travelled into a forked or renamed run before the
   validator is asked, but it is written from `ctx.workflowID` — a host that forks a
   run while KEEPING the same `workflowID` is indistinguishable to the engine. The
   host store remains the authority (Flow's validator compares
   `instance.state.workflow.runId` with `descriptor.runId` and requires exactly one
   matching approval task). Records written before the field carry no `runId` and are
   compared as before, so no checkpoint migration is needed.
10. **Branch identity at the hand-over** is asserted in `bindHumanWait`, but no
   natural path is known that reaches it on a branch the planner did not validate;
   the guard exists so that plan-time validation and the actual hand-over cannot
   diverge silently. It is exercised by fault injection, not by a reproduced defect.
11. **The breakpoint entry gate** cannot fire at a wait re-entry today, for two
   independent structural reasons (no `breakpoints` set on a parallel branch context;
   `currentStepID` — the only step the legacy path admits a record for — is
   one-shot-released by `executeFrom`). It is therefore left exactly as it is. If a
   future change gives branch contexts breakpoints, a breakpoint pause at an
   `entered` frame would strand the record the same way the `if:` skip did, because
   `_planHumanWait` is only reachable from a `started` frame with a handler chain.
