Skip to content

Stepfork v0.1.0a6 - Behavioral Replay

This alpha adds trajectory assertions, async instrumented boundaries, and an opt-in hybrid mode for selected live LLM calls. Existing .sftrace bundles remain compatible; frozen replay remains the default.

Changes

  • Assert tool calls, absence, counts, JSON arguments, order, sequences, and a maximum dependency-step count against the execution that actually replayed. Generated pytest tests can include a reviewed JSON trajectory expectation.
  • Record and frozen-replay async tools, LLM requests, and agent entrypoints. Supported OpenAI and LangGraph adapters now include async paths.
  • Re-execute selected LLMs in hybrid replay while recorded tools stay frozen. The offline model-decision example uses a deterministic fake provider and shows a failing regression before a prompt fix and a passing one afterward. It does not establish real-model quality improvement.
  • Live tool execution now requires allow_live_tools={"name"} or --allow-live-tool name. Previously, mode="live" ran every matching tool. See migration guidance.

Hybrid replay is not deterministic. Selected live models can incur network calls, token costs, and latency. The tool allowlist controls instrumented tool bodies by name; it does not sandbox arbitrary code or uninstrumented side effects. Concurrent async calls still match one strict invocation sequence, while task scheduling remains outside Stepfork's control.

Installation

pip install --pre --upgrade stepfork

Requires Python 3.11, 3.12, or 3.13.

Verification

  • 444 tests passed on Python 3.11, 3.12, and 3.13. Python 3.12 combined statement/branch coverage was 92%.
  • Ruff, MyPy, strict documentation builds, six offline demos, source and wheel builds, and an installed-wheel smoke of the new APIs passed locally.

Feedback

https://github.com/utsab345/stepfork/issues