con un clic
vision-ocr
Extract text from images using OCR and vision AI
Instalar con Codex o Claude Copia este prompt, pégalo en Codex, Claude u otro asistente, y deja que revise la página de la skill y la instale por ti.
Menú
Extract text from images using OCR and vision AI
Instalar con Codex o Claude Copia este prompt, pégalo en Codex, Claude u otro asistente, y deja que revise la página de la skill y la instale por ti.
Basado en la clasificación ocupacional SOC
Edit, transform, and enhance images using AI models
Neural text-to-speech via Amazon Polly
Speaker diarization — identifies and tracks who is speaking at each moment in an audio stream
Semantic endpoint detection — uses an LLM to classify whether the user's utterance is a complete thought, reducing false turn boundaries on mid-sentence pauses
Batch speech-to-text via Google Cloud Speech-to-Text API
Text-to-speech synthesis via Google Cloud Text-to-Speech API
| name | vision-ocr |
| description | Extract text from images using OCR and vision AI |
| version | 1.0.0 |
| tags | ["vision","ocr","text-extraction","document","handwriting"] |
| tools_required | ["vision-pipeline"] |
Extract text from images, documents, and handwritten notes using a progressive 3-tier pipeline: local OCR (PaddleOCR) -> local vision models (TrOCR, Florence-2) -> cloud vision (GPT-4o, Claude).
"Read the text from this receipt" "What does this handwritten note say?" "Extract the table data from this PDF page"