Getting Started¶
This guide installs Stepfork and walks through the full workflow with a
60-second example. A complete, runnable version of the example ships in
examples/quickstart.
Install¶
Install the current alpha from PyPI:
pip install --pre stepfork
The --pre flag is required while Stepfork is a pre-release. To install from
the tagged GitHub repository instead:
pip install "git+https://github.com/utsab345/stepfork.git@v0.1.0a6"
Requires Python 3.11, 3.12, or 3.13. This installs the stepfork CLI and the
Python package.
Prefer uv? The same flags and URLs work with uv pip install. For
contributors working from a checkout, use uv sync in the repository root
instead.
60-second example¶
The agent below looks up an order status with an instrumented tool and decides
whether to notify the customer. The buggy version only notifies when an order
is delivered, so a shipped order is skipped. Save this as quickstart.py:
from stepfork import record, trace_tool
@trace_tool(name="order_status")
def order_status(order_id: str) -> dict:
# Stand-in for your real dependency. Offline and deterministic.
return {"order_id": order_id, "status": "shipped", "carrier": "FedEx"}
def run_agent(order_id: str = "ORD-1001") -> dict:
status = order_status(order_id)
should_notify = status["status"] == "delivered" # bug: shipped orders skipped
return {
"order_id": order_id,
"should_notify": should_notify,
"carrier": status["carrier"],
}
def run_agent_fixed(order_id: str = "ORD-1001") -> dict:
status = order_status(order_id)
should_notify = status["status"] == "shipped" # fix: notify on shipment
return {
"order_id": order_id,
"should_notify": should_notify,
"carrier": status["carrier"],
}
if __name__ == "__main__":
with record("notify-agent", output="failure.sftrace") as session:
result = run_agent()
session.set_output(result)
print(result)
Record the failure:
python quickstart.py
{'order_id': 'ORD-1001', 'should_notify': False, 'carrier': 'FedEx'}
Define the corrected behavior and export a regression test for both entrypoints:
printf '{"order_id": "ORD-1001", "should_notify": true, "carrier": "FedEx"}' > expected.json
stepfork export failure.sftrace --pytest \
--entrypoint quickstart:run_agent \
--expect-output expected.json \
--output test_notify_buggy.py --overwrite
stepfork export failure.sftrace --pytest \
--entrypoint quickstart:run_agent_fixed \
--expect-output expected.json \
--output test_notify_fixed.py --overwrite
expected.json holds the behavior you want after the fix. The generated test
fails on the buggy entrypoint and passes on the fixed one. pytest is not a
runtime dependency, so install it first if needed:
pip install pytest # only needed to run the exported regression tests
python -m pytest -q test_notify_buggy.py # 1 failed (behavior mismatch)
python -m pytest -q test_notify_fixed.py # 1 passed
The core workflow¶
failed agent run
↓
record capture tools, LLM calls, output, and failure
↓
inspect review events, integrity, and error details
↓
frozen replay rerun the entrypoint without touching dependencies
↓
behavioral diff
↓
pytest regression test
- Record the failing run. Wrap the agent call in
stepfork.recordand mark each external dependency with@trace_tool(and, for LLM calls,stepfork.llm_request). The run is saved as a.sftracebundle, including when the block raises. - Replay the failure with frozen dependencies. Stepfork returns the recorded responses for instrumented tool and LLM calls instead of calling those dependencies again, so a replayed run does not re-execute those real services.
- Fix the agent, then diff the recorded failure against a corrected run to see exactly which behaviors changed.
- Export a pytest regression test from the recorded
failure. Supply the corrected output with
--expect-output; the generated test fails on the buggy code and passes on the fix.
Next steps¶
- Learn the vocabulary in Concepts.
- Run the four example agents in Examples.
- Read Troubleshooting for common errors.