AssemblyAI Core Workflow A — Async Transcription
Overview
Primary money-path workflow: submit audio for async transcription with audio intelligence features. The SDK handles file upload (for local files), queues the transcription job, and polls until completion.
Prerequisites
assemblyai package installed
- API key configured in
ASSEMBLYAI_API_KEY
Instructions
Step 1: Basic Async Transcription
import { AssemblyAI } from 'assemblyai';
const client = new AssemblyAI({
apiKey: process.env.ASSEMBLYAI_API_KEY!,
});
const transcript = await client.transcripts.transcribe({
audio: 'https://example.com/meeting-recording.mp3',
});
console.log(transcript.text);
console.log(`Duration: ${transcript.audio_duration}s`);
console.log(`Words: ${transcript.words?.length}`);
Step 2: Local File Upload
const transcript = await client.transcripts.transcribe({
audio: './recordings/interview.wav',
});
import fs from 'fs';
const buffer = fs.readFileSync('./recordings/interview.wav');
const transcript2 = await client.transcripts.transcribe({
audio: buffer,
});
Step 3: Speaker Diarization
const transcript = await client.transcripts.transcribe({
audio: audioUrl,
speaker_labels: true,
speakers_expected: 3,
});
for (const utterance of transcript.utterances ?? []) {
console.log(`Speaker ${utterance.speaker}: ${utterance.text}`);
}
Step 4: Full Audio Intelligence Stack
const transcript = await client.transcripts.transcribe({
audio: audioUrl,
speaker_labels: true,
sentiment_analysis: true,
entity_detection: true,
auto_highlights: true,
iab_categories: true,
content_safety: true,
summarization: true,
summary_model: 'informative',
summary_type: 'bullets',
punctuate: true,
format_text: true,
language_code: 'en',
word_boost: ['AssemblyAI', 'LeMUR', 'transcription'],
boost_param: 'high',
});
for (const s of transcript.sentiment_analysis_results ?? []) {
console.log(`[${s.sentiment}] `);
}
( e transcript. ?? []) {
.();
}
( h transcript.?. ?? []) {
.();
}
categories = transcript.?. ?? {};
( [category, relevance] .(categories)) {
((relevance ) > ) {
.();
}
}
( result transcript.?. ?? []) {
( label result.) {
.();
}
}
.(, transcript.);
Step 5: PII Redaction
const transcript = await client.transcripts.transcribe({
audio: audioUrl,
redact_pii: true,
redact_pii_policies: [
'email_address',
'phone_number',
'person_name',
'credit_card_number',
'social_security_number',
'date_of_birth',
],
redact_pii_sub: 'hash',
redact_pii_audio: true,
});
console.log(transcript.text);
if (transcript.redact_pii_audio_quality) {
const redactedAudio = await client.transcripts.redactedAudio(transcript.id);
console.log('Redacted audio URL:', redactedAudio.redacted_audio_url);
}
Step 6: Manage Transcripts
const page = await client.transcripts.list({ limit: 20 });
for (const t of page.transcripts) {
console.log(`${t.id} | ${t.status} | ${t.audio_duration}s`);
}
const existing = await client.transcripts.get('transcript-id');
await client.transcripts.delete('transcript-id');
Supported Audio Formats
MP3, WAV, FLAC, M4A, OGG, WebM, MP4, AAC. Max file size: 5 GB. Max duration: 10 hours (async). The SDK auto-detects format.
Output
- Complete transcript with word-level timestamps and confidence scores
- Speaker-labeled utterances (with
speaker_labels: true)
- Sentiment analysis, entity detection, key phrases, topic categories
- PII-redacted text and audio
- Content safety labels for moderation
Error Handling
| Error | Cause | Solution |
|---|
transcript.status === 'error' | Corrupted audio or unsupported format | Verify audio file plays locally |
download_url must be accessible | Private/expired URL | Use a publicly accessible URL or upload locally |
Could not process audio | File too short (<200ms) or silent | Ensure audio has speech content |
word_boost has no effect | Misspelled terms or wrong model | Check spelling; word boost works with Best model tier |
Resources
Next Steps
For real-time streaming transcription, see assemblyai-core-workflow-b.
For LLM-powered analysis of transcripts, see assemblyai-sdk-patterns (LeMUR examples).