Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

unstructured-eda

النجوم٢
التفرعات٠
آخر تحديث٢٥ يونيو ٢٠٢٦ في ١١:٥٨

Profiles a raw TEXT or IMAGE corpus before modeling — document/token length distributions, language + encoding mix (mojibake/mixed-script), exact + near-duplicate rate, label coverage + imbalance, vocabulary/OOV, PII flag (text); resolution/aspect/channel/bit-depth distribution, corrupt-file audit, class balance, cross-split duplicate leakage, EXIF/source-batch effects, brightness/blur outliers (image). Emits a corpus-EDA report with per-dimension GO / FIX-FIRST verdicts. Use before NLP or CV modeling, when asked to "profile this text corpus", "audit this image dataset", "EDA on documents/images", or before /nlp-pipeline or /computer-vision. Owns pre-modeling unstructured-corpus profiling; defers tabular columns to /eda, temporal structure to /time-series-eda, dedup mechanics to /dedup (reports rate only), and PII detail to /pii-scan (flags presence only).

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

SKILL.md
readonly