The MVP1 rule taxonomy
| Rule ID | Evidence or purpose | Default response |
|---|---|---|
| TR-CRED-EXFIL | Sensitive credential path + outbound transfer in the same effective command | Critical denial |
| TR-SENSITIVE-READ | Unusual read of SSH, cloud, Keychain, or token locations | High review if unrelated |
| TR-DESTRUCT-OUTSIDE | Recursive destruction targeting home, system, or outside the repository | Critical denial |
| TR-SECURITY-CONFIG | Security controls, authentication, or persistence changes | High / critical |
| TR-REMOTE-EXEC | Fetch-and-execute from an external destination | High review |
| TR-PRIV-ESC | Privilege or permission expansion; privileged mutation | High review |
| TR-UNKNOWN-EGRESS | New outbound destination with sensitive inputs | High review |
| TR-PUBLISH-DEPLOY | Publish, force-push, release, deploy, or production mutation | High when task mismatches |
| TR-DOTENV-ACCESS | Sensitive environment/token access outside the task | Medium / high |
| TR-POLICY-TAMPER | Verified disabling or alteration of app-owned safeguards | Critical when verified |
| TR-ENCODED-COMMAND | Obfuscation combined with a high-impact operation | High review |
| TR-REPEAT-DENIED | A repeated previously denied action | High + drift review |
| TR-REPO-INSTRUCTION | Untrusted instructions attempting to redirect intent | Asynchronous medium / high |
| TR-TASK-DRIFT | Aggregate behavior inconsistent with the task anchor | Asynchronous medium / high |
These are implementation requirements for Phase 3 and Phase 4, not a claim that the current demo has a live detector.
Scores explain routing
| Range | Severity | Policy |
|---|---|---|
| 0–39 | Low | Allow and minimally record |
| 40–69 | Medium | Allow and notify |
| 70–89 | High | Human review |
| 90–100 | Critical | Requires concrete catastrophic deterministic evidence |
Review and warning thresholds default to 70 and 40. Automatic critical blocking depends on an individually enabled catastrophic rule, not a configurable numeric threshold alone. A risk score is neither a probability of attack nor a measured security efficacy percentage.
Evidence has precedence
- Protection disabled means no new TraceRook policy decisions and an explicit off status.
- A trusted catastrophic deterministic match denies immediately.
- An exact, scoped, unexpired exception can release a non-catastrophic action.
- Deterministic high-risk evidence requests review.
- A contextual model recommendation with evidence can request review or warn.
- Medium risk warns while allowing; low risk records with no override.
The model cannot lower a deterministic critical denial. A quick Allow once control must not create a catastrophic exception. Such an exception would require a separate deliberate advanced-settings flow.
Conservative command interpretation
Rules must consider paths, quoting, redirection, pipelines, embedded interpreter flags, normalized relative components, and known symlinks when safe. A regex is not a shell sandbox. Ambiguous high-impact commands should move to review rather than receive a silent safe classification.
Generic build cleanup, disposable workspace operations, and legitimate dependency work need benign near-miss tests. Correlated features and task scope matter: a sensitive path alone is insufficient evidence for the same denial as a credential upload.
Unsafe actions and agent behavior
Unsafe-action findings focus on the proposed operation. Agent-behavior findings focus on task drift, prompt injection, repeated overrides, or suspicious delegation. A finding can belong to both categories.
Asynchronous review can influence future risk posture and warn about drift, but cannot undo an earlier action. Task unknown should reduce drift confidence. All findings should retain evidence and explain limitations.