“The trace never left my laptop” sounds reassuring. It says nothing about whether the trace is safe to paste into an issue, attach to a pull request, or send to a colleague.
Agent traces can contain prompts, model outputs, tool arguments, retrieved documents, URLs, identifiers, credentials, or personal data. A local-first capture architecture reduces automatic transmission; it does not remove the responsibility to inspect what was captured.
This is a general logging problem as much as an AI-specific one. The OWASP Logging Cheat Sheet recommends excluding, masking, sanitizing, hashing, or encrypting sensitive fields such as access tokens, passwords, connection strings, keys, and protected personal data. A recent HackerNoon article on Spring AI agent observability explores a different architecture—masking telemetry at export. The workflow here applies the same underlying caution to local evidence artifacts.
I maintain AgentInspect, an open-source TypeScript toolkit for inspecting agent executions locally. This article explains the safety model I designed around evidence sharing: assess the source, redact into a separate artifact, assess the artifact, package it, and verify its integrity. The examples were checked against [email protected].
The central boundary is worth stating early:
AgentInspect provides best-effort local safety checks. It does not certify privacy, security, or regulatory compliance, and it cannot guarantee that every sensitive value has been detected.
A system can be local-first and still create risky artifacts.
Suppose a trace event contains:
{
"runId": "demo-pii",
"kind": "TOOL",
"name": "send-confirmation",
"metadata": {
"email": "[email protected]",
"apiKey": "sk-synthetic-not-a-real-secret"
}
}
This is synthetic data. Even so, it demonstrates two distinct decisions:
Those decisions may have different answers. A developer may be authorized to inspect a local test trace but not to upload the same fields to a public issue.
That is why I avoided treating “export” as a simple format conversion.
The workflow is easier to reason about as stages:
raw local trace
|
v
source assessment
|
v
redaction to a separate artifact
|
v
post-redaction artifact assessment
|
v
evidence bundle + manifest + hashes
|
v
integrity verification by recipient
Each stage answers a different question.
No single green status answers all five.
The CLI exposes a local safety check:
npx agent-inspect verify-safe demo-pii \
--dir .agent-inspect
Against the synthetic fixture, the result is:
Safety status: SAFE WITH WARNINGS
Format: agent-inspect-v1.0-jsonl
Run: demo-pii
Source assessment: SAFE WITH WARNINGS
Artifact assessment: SAFE WITH WARNINGS
Redaction: profile=share, findings=2
Summary: 2 finding(s), 2 warning(s), 0 error(s)
The two findings identify the synthetic metadata.apiKey and metadata.email fields. The full output also includes the explicit reminder that this is best-effort verification, not a compliance or regulatory certification.
Why call the status SAFE WITH WARNINGS instead of simply PASS? Because evidence handling should preserve ambiguity. A warning is neither an automatic release nor an automatic block. It requires a policy decision and often a human review.
The available high-level outcomes are:
SAFE
SAFE WITH WARNINGS
UNSAFE
UNKNOWN
UNKNOWN is important. A tool that cannot confidently assess an artifact should say so rather than converting uncertainty into safety.
Redaction should not silently mutate the debugging source. The original trace may be needed for an authorized investigation, and in-place editing would destroy evidence about what was actually captured.
Instead, create a separate artifact:
npx agent-inspect redact demo-pii \
--dir .agent-inspect \
--profile share \
--out ./artifacts/demo-pii.share.jsonl
This separation creates a useful audit boundary:
.agent-inspect/demo-pii.jsonl # source; restricted
artifacts/demo-pii.share.jsonl # derived candidate for sharing
AgentInspect includes redaction profiles such as local, share, and strict. These names describe redaction behavior. They are not the same as safety verification policies, which also use development/share/strict-style contexts. Keeping the two concepts separate prevents a common mistake: assuming that selecting a profile automatically proves the resulting artifact meets an organization’s policy.
The artifact gate must evaluate the post-redaction output. If a source contains a secret and the redactor correctly removes it, the source can remain risky while the derived artifact may be acceptable. If redaction misses a custom identifier, the artifact must not inherit a green status merely because a transformation ran.
This is the model:
source risk != artifact risk
redaction executed != artifact is safe
It also explains why custom metadata deserves special attention. Generic detectors cannot understand every company-specific account ID, internal URL, customer code, or regulated data category. Applications should minimize what they record, apply organization-specific redaction where required, and keep a human reviewer in the sharing loop.
Once the candidate artifact passes the chosen review process, it can be packaged:
npx agent-inspect bundle demo-pii \
--dir .agent-inspect \
--profile share \
--out ./evidence
For the synthetic fixture, the CLI reports:
Bundle written to ./evidence
Format: directory
Safe status: SAFE_WITH_WARNINGS
Runs: demo-pii
Files: 12
The directory contains the redacted trace and derived evidence such as reports and check results. Its evidence.json manifest records the evidence format, creation metadata, policy and assessment information, the source run, file roles, and SHA-256 hashes.
A simplified excerpt looks like this:
{
"files": [
{
"path": "assets/runs/demo-pii.jsonl",
"role": "redacted-trace",
"sha256": "1d8bcfd1..."
},
{
"path": "check-results.json",
"role": "checks",
"sha256": "cb7f2939..."
}
]
}
The shortened hashes are for readability in this article. The bundle records complete values.
A recipient can verify the bundle:
npx agent-inspect bundle verify ./evidence
The synthetic bundle produces:
Evidence verify: pass (11 file(s) checked)
Root: ./evidence
Assessment: SAFE WITH WARNINGS
Why does bundle creation report 12 files while verification reports 11 files checked? The manifest is part of the directory, but it records hashes for the other bundle files; it cannot meaningfully include a stable hash of itself while containing that hash.
A passing verification means those recorded files still match the manifest. It does not prove:
The release does not add digital signatures or a key-based chain of custody. If authorship and non-repudiation are required, use an external signing and artifact-governance system.
Redaction is a useful control, but the safest sensitive value is the one never captured.
For agent instrumentation, I recommend deciding field by field:
| Data | Debugging value | Default capture approach |
|---|---|---|
| operation name and status | high | capture |
| parent/child identifiers | high | capture |
| duration and token counts | high | capture |
| full prompts and responses | context-dependent | opt in deliberately |
| tool arguments and results | context-dependent | minimize or summarize |
| credentials and auth headers | none | never capture |
| personal or customer identifiers | usually low | omit, tokenize, or redact at source |
This table is guidance, not a universal policy. The correct decision depends on the system and applicable obligations. The architectural lesson is broader: treat metadata as data. Sensitive values often arrive through labels and convenience fields rather than through the obvious prompt body.
For organization-wide planning, NIST's log-management project treats collection, access, retention, and improvement as a lifecycle rather than a single tooling choice. A local evidence bundle should fit into that larger governance process, not bypass it.
Before attaching agent evidence to an external issue or sending it outside the team, ask:
There is no single command that turns arbitrary agent telemetry into universally safe evidence. Safety emerges from capture choices, redaction rules, artifact inspection, access controls, reviewer judgment, and the context in which the material will be used.
AgentInspect’s role is to make those stages explicit and produce evidence that can be inspected and integrity-checked. Its best-effort detectors help find common risks. Its status model preserves warnings and uncertainty. Its bundle makes the shared artifact reviewable. None of those features removes human accountability.
That is the design principle behind share-checked evidence: keep raw traces local by default, treat sharing as a separate security decision, and never let a successful hash check masquerade as proof that the underlying data is safe or true.
The tagged release and synthetic evidence fixture used here are available on GitHub. Reports of missed detector cases, confusing status language, or unsafe defaults are valuable contributions to this part of the project.