Examples¶
Six runnable, fully offline example agents ship in this repository. No API keys, no network access: each uses a fake provider that stands in for a real model or service, so the demos are deterministic and repeatable.
Clone the repository and sync dependencies once:
git clone https://github.com/utsab345/stepfork.git
cd stepfork
uv sync
Quickstart agent¶
examples/quickstart is the minimal entry point. A single tool
(order_status) and a single decision (should we notify the customer?). The
buggy implementation only notifies when an order is delivered; the fix
notifies when it is shipped.
uv run python examples/quickstart/demo.py
DEMO COMPLETE: failure -> frozen trace -> regression test
buggy agent: FAIL (rc=1), fixed agent: PASS
See the quickstart README for the files and the workflow it demonstrates.
Booking agent¶
examples/booking_agent pairs a fake LLM provider with a flight-search tool.
The agent plans a trip and books the wrong flight. The demo records the buggy
run, freeze-replays it without touching the search tool, diffs it against the
corrected run, and exports a regression test.
uv run python examples/booking_agent/demo.py
Refund agent¶
examples/refund_agent pairs a fake LLM provider with a refund-policy-check
tool. The agent misreads the policy and denies an eligible refund. The demo
runs the full pipeline: record both sides, validate with integrity, replay
frozen, diff, and export a regression test that denies the buggy run and
approves the fixed one.
uv run python examples/refund_agent/demo.py
OpenAI SDK-style adapter¶
examples/openai_chat exercises the optional OpenAI Python SDK adapter through
a fake OpenAI-shaped synchronous client. It records a chat-completions response,
replays it without executing the SDK call, and exports a regression test for a
buggy sentiment decision.
uv run python examples/openai_chat/demo.py
LangGraph agent¶
examples/langgraph_agent builds a real LangGraph StateGraph around a refund
triage decision and wires it to Stepfork through the optional
stepfork.integrations.langgraph adapter. The buggy node inverts the refund
eligibility comparison. The demo records the buggy run, freezes and replays it
with zero model and tool executions, diffs it against the corrected run, and
proves the exported regression test fails for the buggy agent and passes for
the fixed one.
uv run python examples/langgraph_agent/demo.py
See the LangGraph guide and the example README for details.
What every demo proves¶
The original five demos end with the same assertions:
- the buggy implementation FAILS the generated regression test,
- the corrected implementation PASSES the same test,
- frozen replay never executed any external dependency call.
Model-driven prompt decision¶
examples/model_decision uses a LangGraph workflow and a deterministic fake
chat provider that makes the final refund decision from its prompt and tool
results. It records a denial, shows that frozen replay repeats the denial and
rejects a changed prompt, then runs selected fake-model calls in hybrid replay
while the tools stay frozen. It exports a failing test for the buggy prompt and
a passing test for the corrected prompt.
uv run python examples/model_decision/demo.py
This tests the replay and assertion mechanism. The fake provider is not evidence that a real model would improve with the same prompt change. See the example README.
The demos print real results from real subprocesses; they never fabricate output.