Skip to main content

OpenDCAI/DataFlow-Skills

SkillsMP 已收集 OpenDCAI/DataFlow-Skills 中的 20 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
20
GitHub 星标
33
GitHub Forks
8

这个仓库中的 skills

2 个职业分类 · 已分类 100%

已展示 20 / 20 个已收集 Skill。

职业分类
软件开发工程师
描述

DataFlow 开发专家上下文加载器。当用户在 DataFlow 仓库中进行开发时触发, 涵盖:新建算子/Pipeline/Prompt、诊断报错、规范审查、 以及感知仓库变更并建议更新知识库。 Trigger: user is developing in DataFlow repo, asks to create operator/pipeline/prompt, encounters errors, wants code review, or asks about operators.

更新
职业分类
软件开发工程师
描述

Reasoning-guided pipeline planner that generates standard DataFlow pipeline code

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the BenchDatasetEvaluatorQuestion operator. Extended version of BenchDatasetEvaluator with question and subquestion support. Use when: evaluating answers with question context or multiple subquestions.

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the BenchDatasetEvaluator operator. Covers the constructor, two comparison modes (match/semantic), and pipeline usage. Use when: comparing predicted answers against ground truth answers in benchmark evaluation.

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the PromptedEvaluator operator. Use when: scoring text quality with LLM without filtering rows.

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the Text2QASampleEvaluator operator. Use when: evaluating QA pair quality across multiple dimensions.

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the UnifiedBenchDatasetEvaluator operator. Use when: evaluating model answers on benchmark datasets.

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the GeneralFilter operator. Covers the constructor, rule-based filtering logic, and pipeline usage notes. Use when: filtering rows based on column value conditions that can be expressed as lambda functions without LLM calls.

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the KCenterGreedyFilter operator. Covers the constructor, K-Center Greedy algorithm behavior, embedding serving requirements, and pipeline usage notes. Use when: downsampling a large dataset by semantic diversity using embedding…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Reference documentation for the PromptedFilter operator. Covers the constructor, actual scoring and filtering behavior, and pipeline usage notes. Use when: filtering rows based on LLM semantic quality judgment rather than simple rule-based conditions.

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the PandasOperator operator. Use when: applying custom DataFrame transformations without LLM.

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the PromptedRefiner operator. Use when: refining text with LLM, overwriting original column.

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the BenchAnswerGenerator operator. Covers the constructor, full run() signature, actual generation behavior, and integration notes for unified bench evaluation pipelines. Use when: generating model answers from benchmark question…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Reference documentation for the ChunkedPromptedGenerator operator. Covers the constructor, file-path based chunking flow, actual prompt construction, and output file writing behavior. Use when: the dataframe stores file paths, the file content may exceed a…

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the EmbeddingGenerator operator. Covers the constructor, embedding serving requirements, actual dataframe flow, and runnable pipeline usage. Use when: converting one text column in a dataframe into embedding vectors for retrieval,…

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the FormatStrPromptedGenerator operator. Covers the constructor, prompt template restrictions, placeholder-to-column mapping, actual prompt-building logic, and runnable example usage. Use when: one generation task needs multiple…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Reference documentation for the PromptedGenerator operator. Covers constructor parameters, run() signature, actual row-processing behavior, and pipeline usage notes. Use when: integrating PromptedGenerator into a DataFlow pipeline for single-field LLM…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Reference documentation for the RandomDomainKnowledgeRowGenerator operator. [Purpose] Calls an LLM repeatedly with the same domain-generation prompt and writes the generated results into one column of an existing DataFrame. [When to use] Use it when you…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Reference documentation for the RetrievalGenerator operator. [Purpose] Reads one text column from storage, forwards every non-empty row to `llm_serving.generate_from_input(...)`, and writes the returned list into a new output column. [Default backend] Use…

原文语言:英语

更新
职业分类
数据科学家
描述

Reference documentation for the Text2MultiHopQAGenerator operator. [Purpose] Generates multi-hop QA pairs from one text column and writes two output columns: one for `qa_pairs` and one for metadata. [When to use] Use it when you want reasoning-style QA pairs…

原文语言:英语

更新
已展示 20 / 20 个已收集 Skill。