Policy & protection MVP1 / v1.0

Local policy & rule families

Understand deterministic evidence, severity routing, and the role of contextual analysis.

Development previewLive hooks, enforcement, and Anthropic BYOK are pending. This guide distinguishes current behavior from the MVP1 design.

The MVP1 rule taxonomy

Rule IDEvidence or purposeDefault response
TR-CRED-EXFILSensitive credential path + outbound transfer in the same effective commandCritical denial
TR-SENSITIVE-READUnusual read of SSH, cloud, Keychain, or token locationsHigh review if unrelated
TR-DESTRUCT-OUTSIDERecursive destruction targeting home, system, or outside the repositoryCritical denial
TR-SECURITY-CONFIGSecurity controls, authentication, or persistence changesHigh / critical
TR-REMOTE-EXECFetch-and-execute from an external destinationHigh review
TR-PRIV-ESCPrivilege or permission expansion; privileged mutationHigh review
TR-UNKNOWN-EGRESSNew outbound destination with sensitive inputsHigh review
TR-PUBLISH-DEPLOYPublish, force-push, release, deploy, or production mutationHigh when task mismatches
TR-DOTENV-ACCESSSensitive environment/token access outside the taskMedium / high
TR-POLICY-TAMPERVerified disabling or alteration of app-owned safeguardsCritical when verified
TR-ENCODED-COMMANDObfuscation combined with a high-impact operationHigh review
TR-REPEAT-DENIEDA repeated previously denied actionHigh + drift review
TR-REPO-INSTRUCTIONUntrusted instructions attempting to redirect intentAsynchronous medium / high
TR-TASK-DRIFTAggregate behavior inconsistent with the task anchorAsynchronous medium / high

These are implementation requirements for Phase 3 and Phase 4, not a claim that the current demo has a live detector.

Scores explain routing

RangeSeverityPolicy
0–39LowAllow and minimally record
40–69MediumAllow and notify
70–89HighHuman review
90–100CriticalRequires concrete catastrophic deterministic evidence

Review and warning thresholds default to 70 and 40. Automatic critical blocking depends on an individually enabled catastrophic rule, not a configurable numeric threshold alone. A risk score is neither a probability of attack nor a measured security efficacy percentage.

Evidence has precedence

  1. Protection disabled means no new TraceRook policy decisions and an explicit off status.
  2. A trusted catastrophic deterministic match denies immediately.
  3. An exact, scoped, unexpired exception can release a non-catastrophic action.
  4. Deterministic high-risk evidence requests review.
  5. A contextual model recommendation with evidence can request review or warn.
  6. Medium risk warns while allowing; low risk records with no override.

The model cannot lower a deterministic critical denial. A quick Allow once control must not create a catastrophic exception. Such an exception would require a separate deliberate advanced-settings flow.

Conservative command interpretation

Rules must consider paths, quoting, redirection, pipelines, embedded interpreter flags, normalized relative components, and known symlinks when safe. A regex is not a shell sandbox. Ambiguous high-impact commands should move to review rather than receive a silent safe classification.

Generic build cleanup, disposable workspace operations, and legitimate dependency work need benign near-miss tests. Correlated features and task scope matter: a sensitive path alone is insufficient evidence for the same denial as a credential upload.

Unsafe actions and agent behavior

Unsafe-action findings focus on the proposed operation. Agent-behavior findings focus on task drift, prompt injection, repeated overrides, or suspicious delegation. A finding can belong to both categories.

Asynchronous review can influence future risk posture and warn about drift, but cannot undo an earlier action. Task unknown should reduce drift confidence. All findings should retain evidence and explain limitations.

Based on the MVP1 architecture specification, version 1.0 · October 8, 2026.