Skip to content

Stepfork

Your agent failed. Make the failure a test.

Stepfork is a local-first, open source tool for behavioral regression testing of AI agents. It records what an agent actually did during a run (instrumented tool calls, instrumented LLM requests, the final output), then turns the run into a pytest regression test that encodes the corrected behavior.

The problem

AI agents fail in ways unit tests miss. A flight-search agent books the wrong flight. A refund agent denies an eligible claim. When you diagnose such a failure you usually hand-write a fix and hope the regression is covered.

The Stepfork answer

Stepfork workflow: instrument, record, inspect, validate, frozen replay, compare, export a pytest test, and verify the fix

failed agent run
      ↓
record        capture tools, LLM calls, output, and failure
      ↓
inspect       review events, integrity, and error details
      ↓
frozen replay rerun the entrypoint without touching dependencies
      ↓
behavioral diff
      ↓
pytest regression test
  • Record a buggy run into a portable .sftrace bundle, including runs that raise.
  • Replay it with dependencies frozen: instrumented tool and LLM calls are answered from the recording, so their real bodies do not run again.
  • Diff the buggy behavior against a corrected run to see exactly what changed.
  • Export a pytest regression test that fails on the buggy code and passes on the fix.

Nothing is sent to a remote service. Traces, diffing, replay, and test export all run locally.

Key features

  • Typed .sftrace v0.1 trace format with canonical JSON persistence
  • Best-effort secret redaction
  • SHA-256 integrity verification
  • Frozen replay that never executes instrumented dependency bodies
  • Behavioral diffing across equivalent dependency calls
  • Executable pytest export with corrected expectations
  • Python 3.11, 3.12, and 3.13 support

Get started

Stepfork is on PyPI: pip install --pre stepfork. Then run a 60-second example and learn the core workflow in the Getting started guide. The examples page walks through the five runnable demo agents in this repository.

Explore

Stepfork is experimental. See the changelog and the v0.1.0a6 release notes for what shipped.