sft.jsonl and diff_sft.jsonl, or split train and eval
files. Both select sft_curated version 1 by default. For shared prerequisites,
reviewed-bundle steps, output inspection, and holdout guidance, see Choose a
training export.
Choose an Evidence recipe
SFT and diff-SFT use the same selected Evidence recipe andSFTPolicy.
Diff-SFT doesn’t redefine eligibility.
sft_curated version 1 is the default. It admits a human-explicit accept or an
accepted edit with edit_retention_score >= 0.8. A human-explicit reject,
abandonment, or resolved workflow failure vetoes eligibility. The veto includes
an ambiguous aggregate with one failing workflow. A resolved CI pass can affect
Confidence and CI reliability, but it can’t create curated eligibility.
To admit only completions with a clean resolved CI pass, select the verified
recipe:
sft_verified version 1 requires explicit selection. It admits only a clean
resolved CI pass. Ambiguous verdicts, non-verdict-only evidence, and suspected
flakes don’t create eligibility.
Both recipes require a resolved Confidence of at least
SFTPolicy.min_confidence, which defaults to 0.6.
Export SFT rows
The SFT export writes at most one row per inference call that the selected recipe says is safe to imitate. An inference call can have one eligible Attributed completion per attributed file. The projection keeps the one with the highest Confidence, then uses the repository-qualified evidence identity to break ties. It counts identical extra copies asduplicate_completion. If copies of one complete evidence identity
carry conflicting payloads, the projection declines the entire Inference call
once under conflicting_evidence. A higher-ranked sibling cannot hide the
conflict. This includes conflicts in split, Provenance, and source Fact content.
Read an SFT row
Example SFT row:
content, readable reasoning in thinking, and tool
calls in tool_calls. Arguments remain objects. Tool results retain their order;
structured results become deterministic JSON strings. Earlier assistant messages
stay in prompt, outside the completion loss boundary.
Interpret skipped SFT inputs
Inspectskipped before training. The SFT export reports every skipped input
under this closed vocabulary: conflicting_run_identity,
ambiguous_workflow_verdicts, repository_identity_absent,
repository_identity_conflict, repository_identity_unresolved,
repository_mirror_identity_unresolved, repository_source_absent,
non_finite_number, unrepresentable_unicode, completionless,
duplicate_tool_call_id, empty_message, non_string_tool_result,
unrepresentable_part_order, unresolved_tool_call,
unsupported_completion_role, unsupported_message_role,
unsupported_role_part, abandoned, explicit_reject, resolved_ci_failure,
no_eligibility_source, unreliable_ci_resolution, no_reward_signal,
below_confidence_floor, inference_call_not_found, model_absent,
duplicate_completion, and conflicting_evidence.
Export diff-SFT rows
The diff-SFT projection writes one row per(inference call, commit) group and
reads the concrete per-file edits from the mirror. It never synthesizes a diff.
The projection skips abandonment rows before grouping or mirror access. If any
remaining member is ineligible or explicitly rejected, the projection skips
the whole group. Group Confidence is the min() across members.
Read a diff-SFT row
The
completion value uses this shape, with the mirror text preserved exactly:
files[].added_lines. The shared
Git parser decodes quoted filenames while preserving the original patch text,
including rename, deletion, empty-file, and no-newline-marker sections. Invalid
sections never contribute partial patches. A declined section counts once per
commit even when several groups use that commit; valid selected sections can
still produce rows.
Interpret skipped diff-SFT inputs
The diff-SFT export reports every skipped input under this closed vocabulary:conflicting_run_identity, ambiguous_workflow_verdicts,
repository_identity_absent, repository_identity_conflict,
repository_identity_unresolved, repository_mirror_identity_unresolved,
repository_source_absent, non_finite_number, unrepresentable_unicode,
completionless, duplicate_tool_call_id, empty_message,
non_string_tool_result, unrepresentable_part_order, unresolved_tool_call,
unsupported_completion_role, unsupported_message_role,
unsupported_role_part, abandoned, explicit_reject, resolved_ci_failure,
no_eligibility_source, unreliable_ci_resolution, no_reward_signal,
below_confidence_floor, inference_call_not_found, model_absent,
duplicate_completion, conflicting_evidence, unsupported_diff_section,
malformed_diff_section, repo_mismatch, split_mismatch,
eligibility_source_mismatch, mirror_absent, commit_diff_unavailable,
empty_patch, and file_diff_unavailable. unsupported_diff_section and
malformed_diff_section identify unusable Git evidence.
Keep Attribution source metadata and observation IDs with each row. Version 1
recipe eligibility permits inferred Attribution; an empty observation list means
that the row carries no observed Session-to-commit evidence. See
the source contract.