Skip to content

Agent-security traces and evaluation

The repository publishes sanitized traces from controlled autonomous-agent experiments, together with a study, an animated timeline, and offline validation and scoring tools. They help examine how authority crosses service boundaries and how evidence accumulates during long-running work.

Gensee and Tclone ran in observe-only mode in these trials. The data supports behavioral analysis, observability, correlation, and replay; it does not establish active-defense effectiveness. Dataset methodology, sanitization, checksums, schemas, and licensing are included with each release.

For product tooling that reconstructs an integrity-checked timeline from source evidence without re-executing agent actions, see authenticated replay. For an explicit enforcing runtime demonstration, use the boundary conformance proof and end-to-end operation demo.

Released under the Apache 2.0 License.