Agent-security traces and evaluation
The repository publishes sanitized traces from controlled autonomous-agent experiments, together with a study, an animated timeline, and offline validation and scoring tools. They help examine how authority crosses service boundaries and how evidence accumulates during long-running work.
- Study and experiment overview
- Study PDF
- Animated event timeline
- Cross-boundary capability-escape dataset
- Package-service boundary-escape dataset
Gensee and Tclone ran in observe-only mode in these trials. The data supports behavioral analysis, observability, correlation, and replay; it does not establish active-defense effectiveness. Dataset methodology, sanitization, checksums, schemas, and licensing are included with each release.
For product tooling that reconstructs an integrity-checked timeline from source evidence without re-executing agent actions, see authenticated replay. For an explicit enforcing runtime demonstration, use the boundary conformance proof and end-to-end operation demo.