원클릭으로
image-to-text
Parse an image from a file path to obtain a description, enabling non-multimodal LLMs to have vision capabilities.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Parse an image from a file path to obtain a description, enabling non-multimodal LLMs to have vision capabilities.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
Schedule reminders and recurring tasks. Actions: add, list, remove or set_context.
Periodic wake-up service that checks HEARTBEAT.md for pending tasks and executes them automatically. Use this skill to write or update HEARTBEAT.md.
Private knowledge base for indexing multimodal files or folders into a knowledge graph, supporting multi-hop graph retrieval
Generate wiki docs + Mermaid diagrams for any codebase. Use when the user asks to document a codebase, generate architecture diagrams, or create a structured wiki for a repo.
When the user needs to transcribe video (such as .mp4, .mkv, .avi) into text, use the python_repl tool to generate text.
Karpathy's LLM Wiki: build/query interlinked markdown KB.
| name | image_to_text |
| description | Parse an image from a file path to obtain a description, enabling non-multimodal LLMs to have vision capabilities. |
Parse a regular image:
from skills.builtin.core.image_to_text.scripts import itt
if __name__ == '__main__':
user_text: str = "{placeholder}" # <- replace with the absolute path of the input image
image_path: str = "{placeholder}" # <- input the user's question about the image
res = itt(image_path=image_path, user_text=user_text)
print(res)