Skip to main content

graphify

Turn any input (code, docs, papers, images, videos) into a navigable knowledge graph, then answer questions against it. Use when the user wants to build/map/index a codebase or corpus, OR asks any natural-language question about a codebase, documents, or project content ("how does X work", "what calls Y", "trace the data flow", "map this repo", "comment fonctionne X", "construis un graphe", "indexe ce projet", "qu'est-ce qui appelle Y") - especially if graphify-out/ exists, treat the question as a /graphify query rather than reading files by hand. Triggers on /graphify and on codebase/corpus questions when a graph is present or wanted.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
agentik-os/claude-code-skills
آخر نشاط في المصدر
١٧ سبتمبر ٢٠٢٦ في ٢١:٤٢
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٠
التفرعات
٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.

مستكشف الملفات
2 ملفات

عرض SKILL.md

SKILL.md
تعليمات المصدر · معاينة للقراءة فقط
name
graphify
description
Turn any input (code, docs, papers, images, videos) into a navigable knowledge graph, then answer questions against it. Use when the user wants to build/map/index a codebase or corpus, OR asks any natural-language question about a codebase, documents, or project content ("how does X work", "what calls Y", "trace the data flow", "map this repo", "comment fonctionne X", "construis un graphe", "indexe ce projet", "qu'est-ce qui appelle Y") - especially if graphify-out/ exists, treat the question as a /graphify query rather than reading files by hand. Triggers on /graphify and on codebase/corpus questions when a graph is present or wanted.
trigger
/graphify
# /graphify Turn any folder of files into a navigable knowledge graph with community detection, an honest audit trail, and three outputs: interactive HTML, GraphRAG-ready JSON, and a plain-language GRAPH_REPORT.md. ## Usage ``` /graphify # full pipeline on current directory → Obsidian vault /graphify <path> # full pipeline on specific path /graphify https://github.com/<owner>/<repo> # clone repo then run full pipeline on it /graphify https://github.com/<owner>/<repo> --branch <branch> # clone a specific branch /graphify <url1> <url2> ... # clone multiple repos, build each, merge into one cross-repo graph /graphify <path> --mode deep # thorough extraction, richer INFERRED edges /graphify <path> --update # incremental - re-extract only new/changed files /graphify <path> --directed # build directed graph (preserves edge direction: source→target) /graphify <path> --whisper-model medium # use a larger Whisper model for better transcription accuracy /graphify <path> --cluster-only # rerun clustering on existing graph /graphify <path> --no-viz # skip visualization, just report + JSON /graphify <path> --html # (HTML is generated by default - this flag is a no-op) /graphify <path> --svg # also export graph.svg (embeds in Notion, GitHub) /graphify <path> --graphml # export graph.graphml (Gephi, yEd) /graphify <path> --neo4j # generate graphify-out/cypher.txt for Neo4j /graphify <path> --neo4j-push bolt://localhost:7687 # push directly to Neo4j /graphify <path> --mcp # start MCP stdio server for agent access /graphify <path> --watch # watch folder, auto-rebuild on code changes (no LLM needed) /graphify <path> --wiki # build agent-crawlable wiki (index.md + one article per community) /graphify <path> --obsidian --obsidian-dir ~/vaults/my-project # write vault to custom path (e.g. existing vault) /graphify add <url> # fetch URL, save to ./raw, update graph /graphify add <url> --author "Name" # tag who wrote it /graphify add <url> --contributor "Name" # tag who added it to the corpus /graphify query "<question>" # BFS traversal - broad context /graphify query "<question>" --dfs # DFS - trace a specific path /graphify query "<question>" --budget 1500 # cap answer at N tokens /graphify path "AuthModule" "Database" # shortest path between two concepts /graphify explain "SwinTransformer" # plain-language explanation of a node ``` ## What graphify is for Drop any folder of code, docs, papers, images, or video into graphify and get a queryable knowledge graph. Persistent across sessions, honest audit trail (EXTRACTED/INFERRED/AMBIGUOUS), community detection surfaces cross-document connections you wouldn't think to ask about. ## Dynamic Workflow orchestration graphify IS a fan-out workflow. The natural unit of work is a **file chunk** (20-25 files; one image = one chunk). Do not read files yourself one-by-one — that is 5-10x slower and forbidden in Step 3B. Orchestrate: 1. **Plan from the runtime, not a guess.** `detect` returns the real corpus (file counts, words, types). Size the fan-out from it: `agents = ceil(uncached_non_code_files / 22)`. AST (deterministic, free) and semantic (LLM) operate on disjoint file types — dispatch both in the SAME message so they run in parallel (R-SCOPE: disjoint scopes, safe to parallelize). 2. **Fan out parallel extractors.** Dispatch ALL semantic subagents in ONE response (`subagent_type="general-purpose"`, never `Explore` — it can't write). Group same-directory files into one chunk so cross-file edges survive. Each writes its fragment to an absolute `CHUNK_PATH`. 3. **Adversarially verify, don't trust the "done".** A subagent's return text is an input, never the verdict (R-VERIFY). The success signal is the chunk file existing on disk with valid `nodes`+`edges`. Apply ≥3 skeptic checks before accepting the merge: (a) **disk-presence** — every dispatched chunk produced a file; missing ⇒ likely read-only agent, warn + re-run; (b) **schema validity** — `file_type` ∈ the six legal values, `confidence_score` present on every edge, no `0.5` defaults, IDs match the `{parent_dir}_{stem}_{entity}` rule (mismatched IDs create ghost-duplicate orphans); (c) **edge honesty** — direction of `calls` is caller→callee, INFERRED edges carry a rubric value, uncertain ⇒ AMBIGUOUS not omitted. **2-of-3 consensus gate:** if more than half the chunks fail any check, STOP and tell the user to re-run with `general-purpose` — do not ship a half-extracted graph. 4. **Loop-until-dry for unknown-size discovery.** The cache (Step B0) makes this incremental: each `--update` / `--watch` / commit-hook pass re-extracts only changed files and merges. Keep looping until `detect_incremental` reports zero new/changed files — that is the "dry" terminal state, not a fixed pass count. 5. **Synthesize yourself.** The merged graph + community labels + god nodes are raw material; YOU write the community names, the report narrative, and the query answers. Never paste a subagent's fragment as the verdict — compose the cross-community story from the assembled graph. **Single-voice exception:** the post-pipeline `query` / `path` / `explain` answers are NOT fan-out — they are one coherent explanation grounded in the graph. Answer in your own voice using only graph evidence (see Honesty Rules), one writer, no sub-dispatch. ## What You Must Do When Invoked If the user invoked `/graphify --help` or `/graphify -h` (with no other arguments), print the contents of the `## Usage` section above verbatim and stop. Do not run any commands, do not detect files, do not default the path to `.`. Just print the Usage block and return. **Fast path — existing graph:** Before doing anything else, check whether `graphify-out/graph.json` exists. The expected location is `graphify-out/graph.json` relative to the **current working directory** (i.e. the project root where you are running commands). If it exists AND the user's request is a natural-language question about the codebase (e.g. "How does X work?", "What calls Y?", "Trace the data flow through Z") and NOT an explicit rebuild command (`--update`, `--cluster-only`, or a bare path/URL that implies fresh extraction): **skip Steps 1–5 entirely and jump straight to `## For /graphify query`.** Run `graphify query "<question>"` immediately. Do not run detect. Do not check corpus size. Do not ask the user to narrow. The graph is already built — use it. If no path was given, use `.` (current directory). Do not ask the user for a path. If the path argument starts with `https://github.com/` or `http://github.com/`, treat it as a GitHub URL - run Step 0 before anything else, then continue with the resolved local path. Follow these steps in order. Do not skip steps. ### Step 0 - Clone GitHub repo(s) (only if a GitHub URL was given) **Single repo:** ```bash LOCAL_PATH=$(graphify clone <github-url> [--branch <branch>]) # Use LOCAL_PATH as the target for all subsequent steps ``` **Multiple repos (cross-repo graph):** ```bash # Clone each repo, run the full pipeline on each, then merge graphify clone <url1> # → ~/.graphify/repos/<owner1>/<repo1> graphify clone <url2> # → ~/.graphify/repos/<owner2>/<repo2> # Run /graphify on each local path to produce their graph.json files # Then merge: graphify merge-graphs \ ~/.graphify/repos/<owner1>/<repo1>/graphify-out/graph.json \ ~/.graphify/repos/<owner2>/<repo2>/graphify-out/graph.json \ --out graphify-out/cross-repo-graph.json ``` Graphify clones into `~/.graphify/repos/<owner>/<repo>` and reuses existing clones on repeat runs. Each node in the merged graph carries a `repo` attribute so you can filter by origin. **Multiple local subfolders (monorepo or multi-service layout):** The skill pipeline writes all intermediate and final outputs to `graphify-out/` in the current working directory. Running the skill on each subfolder separately will clobber the same output dir. Instead, use the CLI directly for each subfolder — it places `graphify-out/` *inside* the scanned path: ```bash graphify extract ./core/ # → ./core/graphify-out/graph.json graphify extract ./service/ # → ./service/graphify-out/graph.json graphify extract ./platform/ # → ./platform/graphify-out/graph.json # Add --backend gemini|kimi|openai|deepseek|claude-cli depending on which API key you have set # Then merge at the project root: graphify merge-graphs \ ./core/graphify-out/graph.json \ ./service/graphify-out/graph.json \ ./platform/graphify-out/graph.json \ --out graphify-out/graph.json ``` Once `graphify-out/graph.json` exists, the fast path above takes over: any codebase question runs `graphify query` directly on the merged graph — no re-extraction, no size gate. ### Step 1 - Ensure graphify is installed ```bash # Detect the correct Python interpreter (handles uv tool, pipx, venv, system installs) PYTHON="" GRAPHIFY_BIN=$(which graphify 2>/dev/null) # 1. uv tool installs — most reliable on modern Mac/Linux if [ -z "$PYTHON" ] && command -v uv >/dev/null 2>&1; then _UV_PY=$(uv tool run graphifyy python -c "import sys; print(sys.executable)" 2>/dev/null) if [ -n "$_UV_PY" ]; then PYTHON="$_UV_PY"; fi fi # 2. Read shebang from graphify binary (pipx and direct pip installs) if [ -z "$PYTHON" ] && [ -n "$GRAPHIFY_BIN" ]; then _SHEBANG=$(head -1 "$GRAPHIFY_BIN" | tr -d '#!') case "$_SHEBANG" in *[!a-zA-Z0-9/_.-]*) ;; *) "$_SHEBANG" -c "import graphify" 2>/dev/null && PYTHON="$_SHEBANG" ;; esac fi # 3. Fall back to python3 if [ -z "$PYTHON" ]; then PYTHON="python3"; fi if ! "$PYTHON" -c "import graphify" 2>/dev/null; then if command -v uv >/dev/null 2>&1; then uv tool install --upgrade graphifyy -q 2>&1 | tail -3 _UV_PY=$(uv tool run graphifyy python -c "import sys; print(sys.executable)" 2>/dev/null) if [ -n "$_UV_PY" ]; then PYTHON="$_UV_PY"; fi else "$PYTHON" -m pip install graphifyy -q 2>/dev/null \ || "$PYTHON" -m pip install graphifyy -q --break-system-packages 2>&1 | tail -3 fi fi # Write interpreter path for all subsequent steps (persists across invocations) mkdir -p graphify-out "$PYTHON" -c "import sys; open('graphify-out/.graphify_python', 'w', encoding='utf-8').write(sys.executable)" # Save scan root so `graphify update` (no args) knows where to look next time echo "$(cd INPUT_PATH && pwd)" > graphify-out/.graphify_root ``` If the import succeeds, print nothing and move straight to Step 2. **In every subsequent bash block, replace `python3` with `$(cat graphify-out/.graphify_python)` to use the correct interpreter.** ### Step 2 - Detect files ```bash $(cat graphify-out/.graphify_python) -c " import json from graphify.detect import detect from pathlib import Path result = detect(Path('INPUT_PATH')) print(json.dumps(result, ensure_ascii=False)) " > graphify-out/.graphify_detect.json ``` Replace INPUT_PATH with the actual path the user provided. Do NOT cat or print the JSON - read it silently and present a clean summary instead: ``` Corpus: X files · ~Y words code: N files (.py .ts .go ...) docs: N files (.md .txt ...) papers: N files (.pdf ...) images: N files video: N files (.mp4 .mp3 ...) ``` Omit any category with 0 files from the summary. Then act on it: - If `total_files` is 0: stop with "No supported files found in [path]." - If `skipped_sensitive` is non-empty: mention file count skipped, not the file names. - If `total_words` > 2,000,000 OR `total_files` > 500: show the warning. Then compute the top 5 first-level subdirectories by file count: - Read `scan_root` from the detect JSON (always an absolute path to the resolved INPUT_PATH). - Concatenate all file lists across all types (`code`, `document`, `paper`, `image`, `video`). - Filter out any path that starts with `scan_root + "/graphify-out/"` to exclude converted sidecars. - For each file, strip the `scan_root` prefix and take the first path component. Files directly in `scan_root` with no subdirectory count as `(root)`. - If all files are in `(root)` with no subdirectories, do not ask to narrow — no subfolders exist. Instead suggest `--no-cluster` to skip the expensive clustering step and proceed. - Otherwise rank by count, show the top 5 with file counts, then ask which subfolder to run on. Wait for the user's answer before proceeding. - Otherwise: proceed directly to Step 2.5 if video files were detected, or Step 3 if not. ### Step 2.5 - Transcribe video / audio files (only if video files detected) Skip this step entirely if `detect` returned zero `video` files. Video and audio files cannot be read directly. Transcribe them to text first, then treat the transcripts as doc files in Step 3. **Strategy:** Read the god nodes from `graphify-out/.graphify_detect.json` (or the analysis file if it exists from a previous run). You are already a language model — write a one-sentence domain hint yourself from those labels. Then pass it to Whisper as the initial prompt. No separate API call needed. **However**, if the corpus has *only* video files and no other docs/code, use the generic fallback prompt: `"Use proper punctuation and paragraph breaks."` **Step 1 - Write the Whisper prompt yourself.** Read the top god node labels from detect output or analysis, then compose a short domain hint sentence, for example: - Labels: `transformer, attention, encoder, decoder` → `"Machine learning research on transformer architectures and attention mechanisms. Use proper punctuation and paragraph breaks."` - Labels: `kubernetes, deployment, pod, helm` → `"DevOps discussion about Kubernetes deployments and Helm charts. Use proper punctuation and paragraph breaks."` Set it as `WHISPER_PROMPT` to use in the next command. **Step 2 - Transcribe:** ```bash GRAPHIFY_WHISPER_MODEL=base # or whatever --whisper-model the user passed $(cat graphify-out/.graphify_python) -c " import json, os from pathlib import Path from graphify.transcribe import transcribe_all detect = json.loads(Path('graphify-out/.graphify_detect.json').read_text(encoding=\"utf-8\")) video_files = detect.get('files', {}).get('video', []) prompt = os.environ.get('GRAPHIFY_WHISPER_PROMPT', 'Use proper punctuation and paragraph breaks.') transcript_paths = transcribe_all(video_files, initial_prompt=prompt) print(json.dumps(transcript_paths, ensure_ascii=False)) " > graphify-out/.graphify_transcripts.json ``` After transcription: - Read the transcript paths from `graphify-out/.graphify_transcripts.json` - Add them to the docs list before dispatching semantic subagents in Step 3B - Print how many transcripts were created: `Transcribed N video file(s) -> treating as docs` - If transcription fails for a file, print a warning and continue with the rest **Whisper model:** Default is `base`. If the user passed `--whisper-model <name>`, set `GRAPHIFY_WHISPER_MODEL=<name>` in the environment before running the command above. ### Step 3 - Extract entities and relationships **Before starting:** note whether `--mode deep` was given. You must pass `DEEP_MODE=true` to every subagent in Step B2 if it was. Track this from the original invocation - do not lose it. This step has two parts: **structural extraction** (deterministic, free) and **semantic extraction** (LLM, costs tokens). **Before dispatching subagents:** check whether `GEMINI_API_KEY` or `GOOGLE_API_KEY` is set. If neither is set, print this one-liner to the user: > Tip: set `GEMINI_API_KEY` or `GOOGLE_API_KEY` to use Gemini for semantic extraction (`pip install 'graphifyy[gemini]'`).
عرض على GitHub
ملف SKILL.md هذا كبير جدا، لذلك يعرض SkillsMP القسم الاول فقط هنا. عرض على GitHub