| name | naive_reverse_grouper |
| description | 将批处理的样本拆分为独立样本。当用户提到样本拆分、拆分批次、批处理展开、反向分组等需求时使用此skill。
即使用户没有明确说出"拆分",只要任务涉及将批处理数据(字段值为数组)展开为独立样本,就应该使用此skill。
|
| name_zh | 将批处理的样本拆分为独立样本算子 |
| input_params | [{"name":"input_path","type":"string","required":true,"description":"输入JSON文件路径"},{"name":"output_path","type":"string","required":true,"description":"输出JSON文件路径"},{"name":"export_path","type":"string","required":false,"description":"可选,导出批次元数据到JSONL文件"}] |
| output_params | [{"name":"output_path","type":"json_file","description":"拆分后的JSON文件,每个数组元素展开为独立样本"}] |
| tag | 数据衍生 |
| publisher | COMMUNITY |
Naive Reverse Grouper
将批处理的样本拆分为独立样本。
本SKILL使用依赖data_juicer,请在调用前安装好python环境并安装data_juicer,你可用同以下指令进行安装:
pip install py-data-juicer
核心参数
| 参数 | 类型 | 必填 | 默认值 | 说明 |
|---|
| input_path | string | 是 | - | 输入JSON文件路径 |
| output_path | string | 是 | - | 输出JSON文件路径 |
| export_path | string | 否 | - | 可选,导出批次元数据到JSONL文件 |
使用方法
python scripts/run_naive_reverse_grouper.py --input_path <input_path> --output_path <output_path> [--export_path <path>]
实现原理
参照测试代码 test_naive_reverse_grouper.py 中的 _run_helper 函数:
dataset = Dataset.from_list(samples)
op = NaiveReverseGrouper()
new_dataset = op.run(dataset)
输入输出格式
输入格式 (JSON数组,字段值为数组)
[
{"text": ["Sample 1", "Sample 2", "Sample 3"]}
]
输出格式 (JSON数组,拆分为独立样本)
[
{"text": "Sample 1"},
{"text": "Sample 2"},
{"text": "Sample 3"}
]
示例
示例1:基本拆分
python scripts/run_naive_reverse_grouper.py --input_path example_input.json --output_path output.json
示例2:同时导出批次元数据
python scripts/run_naive_reverse_grouper.py --input_path example_input.json --output_path output.json --export_path batch_meta.jsonl
注意事项
- 此算子无需额外参数(export_path可选)
- 输入样本的字段值为数组形式
- 输出会将数组展开为多个独立样本
- 如果传入 export_path,会将批次元数据导出为 JSONL 格式