Skip to main content

media-ocr-ai

Modern AI OCR with open-source + commercial-safe models: PaddleOCR (Apache 2.0, Baidu, 80+ languages, layout analysis, tables), EasyOCR (Apache 2.0, JaidedAI, 80+ languages, easiest install), Tesseract 5 (Apache 2.0, mature LSTM backend, 100+ languages), TrOCR (MIT, Microsoft transformer, the only one that really handles cursive handwriting). Extract text from images and PDFs, structured layout (headers/paragraphs/tables), multilingual documents, handwriting, receipts, invoices, screenshots, scanned forms, signage, whiteboards. Use when the user asks to OCR an image, read text from a picture, extract text from a scanned PDF, parse a receipt or invoice, detect table structure, transcribe handwriting, process a multilingual document (English/Japanese/Chinese/Arabic/etc.), handle CJK or RTL scripts, or pick between PaddleOCR vs EasyOCR vs Tesseract vs TrOCR.

インストールへ移動

ソース情報

リポジトリ
damionrashford/media-os
ソースの最終更新活動
2026年4月18日 01:58
検出された SKILL.md の言語
英語
スター
17
フォーク
4

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。