- Evaluate agent work: compare models by recorded decisions, code retention, and automated checks.
- Reuse context: give agents selected messages, tool calls, and results from a previous Session.
- Build training datasets: export examples for fine-tuning, preference training, and reinforcement learning.
Quickstart
Capture your first agent session on one machine in about five minutes.
How it works
One evidence store for evaluation, agent context, and training datasets.
Evaluate agent work
Compare model outcomes, code retention, and rework.
Reuse context
Select captured evidence for an agent continuing a task.
Deploy
Gateway capture, forge webhooks, per-machine setup, and verification.
Build training datasets
DPO, SFT, diff-SFT, Recovery, and RLVR rows.