Skip to content

Hybrid replay foundation

Hybrid replay is an opt-in Python API for evaluating a changed LLM request against reviewed output and trajectory expectations. It is not deterministic: the selected model runs again and can produce a different response. It may incur network traffic, token cost, and model latency.

from stepfork import ReplaySession
from stepfork.assertions import assert_tool_not_called

with ReplaySession.from_trace(
    "failure.sftrace", mode="hybrid", live_llms={"openai/gpt-example"}
) as replay:
    result = run_agent_with_changed_prompt()
    replay.verify_complete()

assert result == {"approved": True}  # reviewed expectation
assert_tool_not_called(replay, "charge_card")
assert [call.action for call in replay.executed_calls] == ["executed", "substituted"]

The example labels and expected output are illustrative. The live LLM label must exist in the trace and is the recorded provider/model pair (or the model name when no provider was recorded). The caller supplies the actual model client. Stepfork does not create a client or perform a network request on its own. Frozen replay remains the default.

Hybrid replay checks the selected LLM's provider and model but permits its request input to change. It checks that the stored original input fingerprint is valid, then runs the selected LLM body. All tool calls remain frozen and must match their recorded names, arguments, and order. A new or reordered tool call raises before the instrumented tool body executes. Uninstrumented calls and side effects remain outside Stepfork's control.

replay.executed_calls and replay.matched mark each response as executed or substituted. For live calls, ExecutedCall.status is None because the recorded status does not describe the new response. Catching a replay mismatch does not make trajectory assertions valid.

For an executable reviewed test, use run_regression_case(..., mode="hybrid", live_llms={...}, expectation=..., trajectory_expectation=...) or pass the same arguments to generate_pytest_source, which is importable from the public stepfork.export package:

from stepfork.export import generate_pytest_source

Keep live tests outside ordinary CI unless their provider is a deterministic fake; real-model evaluations have variable cost, latency, and behavior.

This foundation does not allow live tools in hybrid mode. It does not fork at a trace step, choose individual repeated LLM occurrences, or realign a changed tool path to a different recorded response. Those features need a separate reviewed design. The older mode="live" API re-executes matching LLM calls. Matching tools require an explicit allow_live_tools name allowlist; see live-tool authorization. Authorized tools can cause real side effects. Neither live nor hybrid replay is deterministic.