sediment derive to build a reviewed, immutable bundle of Attributed
completions and Rollouts, and then export training rows from that bundle.
How Derivation works explains the
snapshot, policy, and bundle contracts.
Run the commands from the operator shell.
Set up an operator shell shows it
for EC2, and the same section for your
own host. sediment derive reads PostgreSQL and the Git mirrors, and derives
the organization in SEDIMENT_ORG_ID.
Build a bundle
Choose a destination that doesn’t exist, and run:skipped, fragmented, and excluded counts, the policy
digest, and the bundle’s as_of time.
The bundle holds manifest.json and six JSONL files. The directory has mode
0700, and each file has mode 0600. Sediment never overwrites a bundle, so
each run needs an unused destination.
Missing mirrors don’t stop the run. Sediment counts them under skipped, and
the bundle can hold fewer Attributions.
Select a cohort
To limit the bundle to some users and a time window, run:--since is inclusive and --until is exclusive. Both need a timezone, and
both select by Inference call observation time. Sediment keeps a whole Rollout
when any of its Inference calls falls in the window. It excludes and counts a
Rollout that mixes allowed and other users.
Set Derivation policy
To change a default, write a TOML file with only the fields that you want to override. These values are the defaults:sediment derive:
eval_fraction assigns about that share of Sessions to evaluation, by a hash
of the Session identifier. Set it to 0.0 to write no split. The split keeps
each Session on one side, but the same repository, prompt, or task can appear
on both sides. If your evaluation needs unseen repositories or tasks, assign
them in your own versioned benchmark manifest before training.
Inspect the bundle
To print the firstN Attributed completions and Rollouts after the write,
add --sample N:
skipped: each closed reason says why an input didn’t qualify. For example,unmatched_decision_call_idmeans that a Developer decision names no captured Inference call. A narrow cohort appears underexcluded, notskipped.fragmented: each Turn stays in a Segment. Readprior_output_absent,input_history_changed, andprior_output_not_replayedbefore you treat separate Segments as one trajectory.conflicting_run_identityandambiguous_workflow_verdicts: CI outcomes that exports can’t resolve to one verdict.
record_json string. Use the
bundle reader, which validates the whole bundle, rather than parsing lines
yourself. Validation proves the bundle’s internal relationships. It can’t prove
that an external producer supplied a complete or truthful Fact population.
Export from the bundle
Project the same bundle into each objective that you need:Compare two policies
Build each policy into its own destination:policy_digest, scope, as_of, and the skipped, fragmented, and
excluded counts before you compare rows.
Troubleshoot
A bundle carries the organization’s complete call and repository identity
populations, each capped at 50,000 rows, even for a small cohort. A narrower
cohort doesn’t help. Don’t trim identity files, quarantine valid Facts, or
split the organization to get under the cap.
Open an issue with the
failed command and its
as_of.
If sediment derive stops before it reports success, check the destination.
If the destination exists, it holds a complete bundle. If it doesn’t, rerun
into an unused path. A process that ends abruptly can leave a private staging
directory named .<destination>.<suffix> beside the destination. Remove it
after the process has ended.
In a container where /tmp is memory-backed, set TMPDIR to a private
disk-backed directory before a large run. The Compose operator service
already does.
Use the bundle from Python
build_derived_bundle_context derives a file-backed bundle, and
open_derived_bundle validates and opens one from disk. Use both as context
managers, and keep the context open until you finish reading. For a small
bundle, build_derived_bundle and read_derived_bundle load it into memory,
up to 64 MiB of encoded payload.