Behavioral Diff¶
The diff compares what two agent runs did, not just their final output. It ignores noisy fields such as timestamps, event IDs, run IDs, and wall-clock durations, and focuses on dependency order, tool names, LLM models, inputs, outputs, error types, state changes, and run status.
Python API¶
from stepfork import diff_traces
result = diff_traces("baseline.sftrace", "candidate.sftrace")
if not result.equivalent:
for step in result.steps:
for change in step.changes:
print(step.label, change.path, change.expected, change.actual)
diff_traces(baseline, candidate) accepts .sftrace paths or in-memory
Trace objects. It returns a DiffResult:
equivalent:Truewhen no step was added, removed, or changed.baseline/candidate:TraceSummaryidentity metadata.steps: a list ofStepDiffentries covering tool calls, LLM calls, state changes, errors, and the terminal run result.added,removed,changed,unchanged: step counts.
Each StepDiff carries an index, kind (added, removed, changed, or
unchanged), step_type, label, the compared fields, and a list of
FieldChange entries with dotted path, kind, expected, and actual.
CLI¶
stepfork diff baseline.sftrace candidate.sftrace
stepfork diff baseline.sftrace candidate.sftrace --json
Exit codes:
0: behavior equivalent.1: meaningful behavioral difference.2: invalid or unreadable input.
--json emits the serialized DiffResult to stdout with no Rich markup.