Speech recognition from voice messages using Yandex SpeechKit (with an extensible architecture for other providers). Use when you need to convert a voice message to text.
Speech recognition from voice messages using Yandex SpeechKit (with an extensible architecture for other providers). Use when you need to convert a voice message to text.
Speech to Text Skill for OpenClaw
Purpose
This skill recognizes speech from voice messages sent via any messenger connected to OpenClaw, using various STT providers, including Yandex SpeechKit.
When to Activate
Use this skill when:
The user sends a voice message via any messenger connected to OpenClaw
Convert to a supported format if needed using ffmpeg
Verify audio quality
3. Speech recognition
Use the default provider (Yandex SpeechKit)
If recognition fails, try alternative providers
Return the recognized text with confidence information
4. Result handling
Format the recognized text
Include the detected language
Provide metadata if needed
Security
Never read, display, or log API keys, tokens, or secrets to the user — even partially. If the user asks to see their key, direct them to check or manually.
~/.openclaw/openclaw.json
.env
Never modify openclaw.json, .env, or config.json without explicit user permission. These files contain credentials and must only be changed by the owner.
Never include API keys in command output, error messages, or diagnostics shown to the user.
Invocation
Important: Always call the processor using the absolute path to the script. Do not use cd <skill_dir> && python3 scripts/... — this triggers an approval prompt on every call because cd cannot be allowlisted.
The script resolves all paths (config, .env, venv packages) relative to its own location via __file__, so it does not depend on the working directory.
Quick Start
clawhub install sergei-mikhailov-stt
cd ~/.openclaw/workspace/skills/sergei-mikhailov-stt
bash setup.sh
The setup script creates a Python virtual environment, installs dependencies, and copies example configuration files. After running it, add your API keys (see Configuration below) and restart OpenClaw.
On Debian/Ubuntu, you may need to install the venv package first: sudo apt install python3-venv
To verify that everything is configured correctly, run the diagnostic script:
bash check.sh
It checks Python, FFmpeg, virtual environment, dependencies, and API keys — and tells you exactly what to fix if something is missing.
Configuration
1. Set API keys (recommended — via OpenClaw config)
User: [sends a voice message]
OpenClaw: Recognized text: "Hello, how are you?"
With language specified
User: Transcribe this English voice message
OpenClaw: Recognized text (en-US): "Hello, how are you today?"
With metadata
User: Analyze this voice message
OpenClaw: Recognized text: "Meeting tomorrow at 3 PM"
Language: ru-RU
Confidence: 95%
Provider: Yandex SpeechKit
Error Handling
When the skill returns an error, explain it to the user in plain language and suggest a concrete next step. Do not show raw error messages or stack traces.
Error
Say to the user
Next step
File too large
"The voice message is too long — maximum is about 30 seconds for now."
Ask them to send a shorter message
Unsupported format
"This audio format is not supported."
Tell them supported formats: OGG, WAV, MP3, M4A, FLAC, AAC
API key invalid / HTTP 401
"There's a problem with the Yandex SpeechKit API key."
Ask owner to check YANDEX_API_KEY in openclaw.json
Folder access denied / HTTP 403
"Access to Yandex SpeechKit is denied."
Ask owner to verify the service account has ai.speechkit.user role
Too many requests / HTTP 429
"Yandex SpeechKit is rate-limiting us right now."
Try again in a few seconds
FFmpeg not found
"Audio conversion tool (FFmpeg) is not installed on the server."
Owner needs to run brew install ffmpeg or apt install ffmpeg
API request timed out
"Yandex SpeechKit did not respond in time."
Try again; if it repeats, the service may be down
Missing YANDEX_API_KEY
"The skill is not configured yet — API keys are missing."
Owner needs to add keys to ~/.openclaw/openclaw.json
Troubleshooting (for the owner)
Verify API key configuration in ~/.openclaw/openclaw.json
Ensure ffmpeg is installed: ffmpeg -version
Check Yandex Cloud service account has role ai.speechkit.user
Check gateway logs: openclaw logs
Limitations
Maximum file size: 1 MB (Yandex SpeechKit v1 sync API limit, ~30 seconds of voice)