Current local evidence
The implementation record dated October 8, 2026 reports debug and release builds of all three arm64 executables with Swift 6 and SDK 27, plus 43 passing Swift Testing cases (23 retained tests and 20 new security contract tests). Native rendering checks cover 20 light/dark cases, including demo/real separation, pending/expired review, and denied notifications.
Scene launch and relocated bundle checks validate a visible dashboard and packaged resources. Full Xcode archive validation, manual VoiceOver coverage, and actual notification delivery acceptance remain pending. These are development results, not efficacy claims.
Run the relevant checks
./scripts/test.sh
./scripts/smoke-test.sh
./scripts/bundle-smoke-test.sh
./scripts/ui-smoke-test.shScripts build and exercise harmless fixtures. Native render artifacts are under build/ui-smoke/. The UI renderer checks appearance and state behavior, not pixel equality with an approved design.
The checked-in implementation status explains the Command Line Tools adjustments and remaining release gates.
Required live-phase tests
- Malformed, oversized, null, invalid UTF-8, deeply nested, and concurrent host input.
- Host output encoding and preservation of native permissions.
- Path normalization, shell ambiguity, quoting, symlinks, and benign cleanup near misses.
- Secrets in strings, nested objects, fragmented tokens, Unicode, logs, and errors.
- Model errors, deadlines, budgets, hallucination, and evidence precedence.
- Exact approval scope, expiry boundaries, replay, double-click, disconnect, restart, and simultaneous reviews.
- SQLite migrations, foreign keys, permissions, retention, and transactional races.
- Installer ownership, idempotency, preservation, concurrent edits, backups, and interrupted writes.
Prove the body never ran
Both Claude Code and Codex must exercise a real local hook on exact verified releases. A harmless denied operation in a disposable directory must leave no execution marker. Codex tests require user-reviewed trust before the callback is counted.
Review tests must cover Block, Allow once with host permissions intact, expiry denial, and an old notification. Failure tests kill the UI and service, remove the helper, disable hooks, force timeouts, and confirm honest degraded status. Privacy tests revoke keys, simulate network errors, and prove unsafe payloads are never sent.
Never use real secrets, destructive targets, or actual exfiltration as test material.
The release evaluation corpus
MVP1 requires at least 40 sanitized scenarios: 10 catastrophic, 10 high-risk review, 10 benign near-misses, and 10 prompt-injection or drift cases. Each defines the gold outcome, evidence, task anchor source, and false-positive tolerance.
Version prompts and schemas, and rerun the corpus after changes. Track measured precision, missed high-risk cases, latency, calls per 100 actions, and token use. Do not publish a security efficacy percentage without measured evidence.
Performance targets, not promises
| Measure | Engineering target |
|---|---|
| Ordinary local no-override, added time | MVP2 warmed-path target: p95 under 150 ms |
| Local rule decision | p95 ≤ 25 ms |
| Contextual review | 8 s internal model deadline |
| High-risk approval | 45 s default, within verified host timeout |
| Concurrent model reviews | Default maximum 3 |
These targets need on-device verification and are not observed production measurements or network SLAs.