Skip to main content

training-data-pipeline

Build training datasets for LLM specialization from production data, frontier model distillation, and synthetic bootstrapping. Use when formatting production logs into SFT data, distilling from frontier APIs, or preparing data for fine-tuning. Covers JSONL formatting, data quality validation, deduplication, and train/eval splitting.

Jump to install

Source facts

Repository
synthetic-sciences/openscience
Last source activity
July 4, 2026 at 06:25
Detected SKILL.md language
English
Stars
3,362
Forks
454

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.