Skip to main content

relation-extraction-pipeline

Prepare relation extraction datasets and inspect the TensorFlow PCNN training workflow for Agriculture_KnowledgeGraph.

설치로 이동

소스 정보

저장소
VectorSpaceLab/AREX-Skill
최근 소스 활동
2026년 8월 26일 16:31
감지된 SKILL.md 언어
영어
스타
12
포크
2

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
7 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
relation-extraction-pipeline
description
Prepare relation extraction datasets and inspect the TensorFlow PCNN training workflow for Agriculture_KnowledgeGraph.
disable-model-invocation
true
metadata
{"disco-role":"operating"}
license
GPL 3.0
# Relation Extraction Pipeline Use this sub-skill when the task is to create, validate, split, or troubleshoot the remote-supervised relation-extraction dataset, or to understand the TensorFlow PCNN training stack in this agriculture knowledge-graph repository. ## Route here for - Turning aligned Wikidata/Wikipedia sentence rows into `rel2id.json`, `entity2id.json`, `dataset.json`, `word2vec.json`, `train_dataset.json`, and `test_dataset.json`. - Validating six-column training TSV rows and the JSON schemas expected by the PCNN data loader. - Diagnosing Fire command names, working-directory-sensitive paths, stale `_processed_data`, bad entity positions, relation-label filters, or NA-sample split failures. - Reviewing the PCNN model configuration, TensorFlow 1.x assumptions, GPU settings, and large word-vector limits before any expensive training run. ## Route elsewhere - General Wikidata crawling, Wikidata JSON-to-CSV conversion, Neo4j import CSV creation, Hudong/weather crawlers, and network collection pipelines belong to the crawler/Wikidata workflow, not this sub-skill. - Django relation-label annotation UI, Mongo-backed tagging pages, and web-service startup belong to the web-app workflow, not this sub-skill. - Neo4j graph querying or Cypher import plans belong to graph query/data management. ## Operating map 1. For the end-to-end data build sequence, required source artifacts, Fire commands, and safe replacement commands, read [dataset preparation](references/dataset-preparation.md). 2. For exact row and JSON schemas, path-sensitive generated files, and validator usage, read [data formats](references/data-formats.md). 3. Before training or editing `config.py`/`train.py`, read [PCNN training](references/pcnn-training.md). 4. For known failure modes and recovery actions, read [troubleshooting](references/troubleshooting.md). ## Bundled scripts - [scripts/relation_dataset_schema_check.py](scripts/relation_dataset_schema_check.py) validates tiny training TSV, `rel2id`, `entity2id`, dataset JSON, and optional word-vector JSON files without importing TensorFlow. - [scripts/deduplicate_training_rows.sh](scripts/deduplicate_training_rows.sh) deduplicates training rows with explicit input and output paths and a safe `--help` path. ## Safe default Do not launch PCNN training, live Neo4j alignment, Mongo annotation, network crawling, or large word-vector conversion as a first check. First run the schema checker or a tiny deduplication fixture, then escalate only after data files, working directory, and TensorFlow 1.x/GPU constraints are explicit.
GitHub에서 보기