Choose a bounded question
Define the decision before collection: performance, generated-code review, compatibility, or extension audit.
Evidence methodology
AgentSight is not a promise that every low-level event explains itself. The method is to scope one question, collect the relevant system boundary, retain provenance, and make uncertainty visible.
Define the decision before collection: performance, generated-code review, compatibility, or extension audit.
Observe the selected process family and retain enough context to attribute model and system activity.
Connect turns, calls, commands, paths, destinations, resource phases, and exit status without erasing provenance.
Aggregate the trace into a causal profile or review artifact that answers the original question.
Separate observed effects, supported inferences, unavailable evidence, and questions that need another experiment.
Evidence levels
A useful report tells the reviewer which statements came from a recorded event, which combine multiple events, and which remain hypotheses.
A command, process, file operation, network destination, model call, or resource phase exists in the selected record.
Multiple events are connected to a run, process family, turn, or bounded task through recorded identifiers and timing.
A likely explanation is supported by evidence but still needs a follow-up run, source inspection, or controlled comparison.
The collector, runtime packaging, encryption boundary, permissions, or retained data does not support a conclusion.
Privacy boundary
Agent sessions can contain prompts, responses, repository paths, headers, commands, and network targets. Keep raw databases and exports local, redact artifacts before sharing, and publish only the minimum evidence required for the decision.
$ sudo agentsight record -- claude
$ agentsight report audit --json
$ agentsight report export -o review.jsonReproducibility
Browse the run library to see how the methodology becomes a public-safe engineering artifact.