Skip to main content

e2

Agent E2 - Qualitative Coding Specialist - Systematic coding and theme development. Covers codebook development, coding strategies, saturation assessment, and CAQDAS guidance.

Jump to install

Source facts

Repository
brycewang-stanford/Auto-Empirical-Research-Skills
Last source activity
April 3, 2026 at 02:07
Detected SKILL.md language
English
Stars
4,389
Forks
532

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
e2
description
Agent E2 - Qualitative Coding Specialist - Systematic coding and theme development. Covers codebook development, coding strategies, saturation assessment, and CAQDAS guidance.
version
12.0.1
## ⛔ Prerequisites (v8.2 — MCP Enforcement) `diverga_check_prerequisites("e2")` → must return `approved: true` If not approved → AskUserQuestion for each missing checkpoint (see `.claude/references/checkpoint-templates.md`) ### Checkpoints During Execution - 🟠 CP_CODING_APPROACH → `diverga_mark_checkpoint("CP_CODING_APPROACH", decision, rationale)` - 🟠 CP_THEME_VALIDATION → `diverga_mark_checkpoint("CP_THEME_VALIDATION", decision, rationale)` ### Fallback (MCP unavailable) Read `.research/decision-log.yaml` directly to verify prerequisites. Conversation history is last resort. --- # E2: Qualitative Coding Specialist ## Role Expert in systematic qualitative data coding, codebook development, theme identification, and saturation assessment. Guides researchers through rigorous coding processes for thematic analysis, grounded theory, and content analysis. ## Core Capabilities ### 1. Codebook Development Approaches #### Deductive (A Priori) Coding **When to Use:** - Literature-driven research - Theory-testing studies - Structured content analysis - Pre-defined frameworks (e.g., SDT, TPB) **Process:** ```yaml deductive_coding: step_1_literature_review: action: "Extract key constructs from theoretical framework" output: "Initial code list with definitions" step_2_operationalization: action: "Define codes with inclusion/exclusion criteria" output: "Structured codebook" step_3_pilot_coding: action: "Test codebook on 10-20% of data" output: "Refined codebook" step_4_reliability_check: action: "Calculate inter-rater reliability (Kappa)" output: "Reliability metrics, codebook adjustments" ``` **Example Deductive Codebook (Self-Determination Theory):** ```yaml code: "autonomy_support" definition: "Teacher actions that support student self-direction and choice" when_to_use: - "Teacher offers choices" - "Teacher solicits student input" - "Teacher acknowledges feelings" when_not_to_use: - "Teacher gives commands without rationale" - "Generic praise without choice element" example_quotes: - "The teacher said 'you can choose to work alone or in pairs'" - "She asked us what topics we wanted to explore" related_codes: ["autonomy_thwarting", "intrinsic_motivation"] parent_theme: "motivational_climate" ``` #### Inductive (Emergent) Coding **When to Use:** - Exploratory research - Phenomenological studies - Grounded theory - Under-researched phenomena **Process:** ```yaml inductive_coding: phase_1_open_coding: approach: "Line-by-line, no preconceptions" output: "100-200 initial codes" phase_2_axial_coding: approach: "Group codes by similarity, identify patterns" output: "30-50 focused codes" phase_3_selective_coding: approach: "Identify core categories and relationships" output: "8-15 themes with subthemes" ``` **Example Inductive Code Evolution:** ```yaml evolution: open_codes: - "student_mentions_chatbot_patience" - "student_appreciates_no_judgment" - "student_feels_safe_making_errors" focused_code: "psychological_safety" theme: "non-judgmental_learning_environment" definition: "Learners perceive AI chatbot as safe space for practice without fear of negative evaluation" ``` #### Hybrid (Deductive + Inductive) **Best Practice for Social Science:** ```yaml hybrid_approach: step_1: "Start with literature-derived codes (deductive)" step_2: "Remain open to emergent codes (inductive)" step_3: "Track code sources (deductive vs. emergent)" step_4: "Report both a priori and emergent themes" example: deductive_codes: ["engagement", "motivation", "self-efficacy"] emergent_codes: ["technical_frustration", "privacy_concern", "gamification_appeal"] ``` ### 2. Coding Strategies by Methodology #### Thematic Analysis (Braun & Clarke, 2006) **Six-Phase Process:** ```yaml phase_1_familiarization: activities: - "Read and re-read entire dataset" - "Note initial ideas and patterns" - "Highlight interesting passages" tools: ["Annotation software", "Memo writing"] output: "Annotated transcripts, research journal notes" time_estimate: "20-30% of total coding time" phase_2_initial_coding: activities: - "Systematic line-by-line coding" - "Create code labels" - "Organize data extracts by code" tools: ["CAQDAS", "Excel", "Index cards"] output: "Initial codebook (50-150 codes typical)" quality_check: - "Every data item coded" - "Data extracts retain context" - "Codes are specific enough to be meaningful" phase_3_theme_searching: activities: - "Collate codes into potential themes" - "Mind mapping of relationships" - "Create theme tables" techniques: - "Post-it note sorting" - "Mind maps" - "Thematic tables" output: "Candidate themes (8-15 typical)" phase_4_theme_review: level_1_review: action: "Check themes against coded data extracts" criteria: "Internal homogeneity (coherence within theme)" level_2_review: action: "Check themes against entire dataset" criteria: "External heterogeneity (distinction between themes)" output: "Refined themes, thematic map" phase_5_theme_defining: activities: - "Name each theme" - "Write theme descriptions (2-3 paragraphs)" - "Identify essence of each theme" - "Define subthemes if needed" output: - "Final theme definitions" - "Thematic structure" quality_criteria: - "Theme names are concise and informative" - "Definitions capture unique contribution" - "No significant overlap between themes" phase_6_report_writing: activities: - "Select vivid quotations" - "Write analytic narrative" - "Link themes to research question" - "Situate findings in literature" output: "Findings section with theme-based structure" ``` **Thematic Analysis Quality Checklist:** ```yaml quality_criteria: data_engagement: - "[ ] Transcripts read multiple times" - "[ ] Coding checked against transcripts" - "[ ] Themes grounded in data extracts" coding_rigor: - "[ ] Each data item coded" - "[ ] Coding systematic and thorough" - "[ ] Similar codes collated" theme_coherence: - "[ ] Themes internally consistent" - "[ ] Themes distinct from each other" - "[ ] Thematic structure logical" transparency: - "[ ] Coding process described" - "[ ] Code-to-theme process explained" - "[ ] Sufficient quotations provided" ``` #### Grounded Theory (Charmaz, 2006) ```yaml grounded_theory_coding: phase_1_initial_coding: approach: "Open coding - line-by-line analysis" coding_style: - "Use gerunds (verbs ending in -ing)" - "Stay close to data" - "Avoid premature interpretation" example: data: "I felt nervous talking to real people, but the chatbot didn't judge me" codes: - "feeling_nervous_with_humans" - "perceiving_chatbot_as_non_judgmental" - "comparing_human_vs_AI_interaction" phase_2_focused_coding: approach: "Select most frequent/significant codes" activities: - "Synthesize initial codes" - "Test codes against data" - "Develop categories" example: initial_codes: ["feeling_anxious", "fearing_judgment", "avoiding_speaking"] focused_code: "social_anxiety_in_language_learning" phase_3_axial_coding: approach: "Identify relationships between categories" framework: conditions: "When/why category occurs" actions_interactions: "How people respond" consequences: "What happens as result" example: category: "chatbot_psychological_safety" conditions: "High speaking anxiety + fear of peer judgment" actions: "Increased practice with AI, risk-taking in language use" consequences: "Gradual confidence building" phase_4_theoretical_coding: approach: "Integrate categories into theory" output: "Core category + theoretical model" example: core_category: "scaffolded_confidence_development" theoretical_model: "AI → Safe practice → Risk-taking → Competence → Human interaction" ``` **Grounded Theory Memos:** ```yaml memo_types: code_memo: purpose: "Define and elaborate codes" example: | Memo: "Perceiving chatbot as non-judgmental" Date: 2024-10-15 This code captures participants' descriptions of chatbots as lacking evaluative judgment. Unlike human interlocutors, AI doesn't show disappointment, frustration, or impatience. This perception creates psychological safety. Properties: - Non-verbal judgment absent (no eye-rolling, sighs) - Consistent tone regardless of errors - No social comparison with peers Related codes: "social_anxiety", "fear_of_negative_evaluation" theoretical_memo: purpose: "Develop conceptual relationships" example: | Theoretical Memo: Anxiety-Safety-Practice Loop Emerging pattern: High speaking anxiety → Preference for AI practice → Increased practice volume → Gradual confidence → Willingness to speak with humans. This suggests AI serves as transitional object/space for anxious learners. Not replacement for human interaction but scaffold toward it. operational_memo: purpose: "Track methodological decisions" example: | Operational Memo: Saturation assessment After 18 interviews, no new codes emerging for "chatbot affordances" category. Last 3 interviews yielded only variations on existing codes. Consider saturation reached for this category. ``` #### Content Analysis (Descriptive + Interpretive) ```yaml content_analysis_coding: manifest_content: definition: "Surface-level, visible content" approach: "Objective, countable" examples: - "Frequency of 'chatbot' mentions" - "Number of positive vs. negative adjectives" - "Presence/absence of specific themes" reliability: "High inter-rater reliability possible (Kappa > 0.80)" latent_content: definition: "Underlying meaning, interpretive" approach: "Subjective, inferential" examples: - "Implicit attitudes toward AI" - "Underlying emotional tone" - "Power dynamics in human-AI interaction" reliability: "More challenging (Kappa 0.60-0.80 acceptable)" coding_units: word_level: "Individual words (e.g., AI, anxiety, practice)" phrase_level: "Meaningful phrases (e.g., 'felt less judged')" sentence_level: "Complete thoughts" paragraph_level: "Thematic segments" document_level: "Whole interview" ``` ### 3. Code Quality Criteria **High-Quality Code Entry Template:** ```yaml code_template: code_name: "clear_descriptive_name" definition: conceptual: "Abstract definition of construct" operational: "How it manifests in data" when_to_use: inclusion_criteria: - "Criterion 1" - "Criterion 2" boundary_conditions: "Where code applies" when_not_to_use: exclusion_criteria: - "What this code is NOT" - "Common misapplications" edge_cases: "Ambiguous situations" example_quotes: typical_examples: - "Quote 1 [Participant 3, Line 45]" - "Quote 2 [Participant 7, Line 112]" boundary_examples: - "Borderline case [P5, L89] - coded because..." counter_examples: - "Quote that seems similar but isn't [P2, L34] - not coded because..." related_codes: parent_code: "Higher-level category" sibling_codes: ["Related codes at same level"] child_codes: ["More specific sub-codes"] code_metadata: date_created: "2024-10-15" created_by: "Researcher initials" source: "deductive/inductive" frequency: "Number of times applied" ``` **Example High-Quality Codebook Entry:** ```yaml code_name: "perceived_judgment_anxiety" definition: conceptual: "Psychological discomfort arising from anticipation of negative evaluation by others during language production" operational: "Participant explicitly mentions fear, worry, nervousness, or discomfort about being judged, evaluated, or criticized by others when speaking" when_to_use: inclusion_criteria: - "Explicit mention of judgment, evaluation, or criticism from others" - "Affective states (fear, anxiety, nervousness) linked to social evaluation" - "Comparisons between human vs. AI interaction where judgment is factor" boundary_conditions: "Must be specific to language speaking context, not general social anxiety" when_not_to_use: exclusion_criteria: - "Generic nervousness without reference to being judged" - "Task difficulty anxiety (not social evaluation)" - "Performance anxiety about grades (use 'grade_anxiety' code)" edge_cases: "Self-judgment (internal criticism) → use 'self_critical_perfectionism' instead" example_quotes: typical_examples: - "I was scared my classmates would laugh at my pronunciation" [P3, L45] - "The chatbot doesn't judge me, but people do" [P7, L112] - "I felt nervous because the teacher would notice my mistakes" [P11, L201] boundary_examples: - "I was worried about saying the wrong thing" [P5, L89] - coded because implies judgment from listener counter_examples: - "I was nervous because the vocabulary was difficult" [P2, L34] - NOT coded (task difficulty, not judgment) - "I felt anxious before the test" [P9, L156] - NOT coded (test anxiety, use 'evaluation_anxiety') related_codes: parent_code: "affective_barriers" sibling_codes: ["speaking_anxiety_general", "fear_of_mistakes"] child_codes: ["peer_judgment_anxiety", "teacher_judgment_anxiety"] code_metadata: date_created: "2024-10-15" created_by: "HY" source: "deductive (Foreign Language Anxiety Scale)" frequency: 27 prevalence: "18/24 participants (75%)" ``` ### 4. Inter-Rater Reliability **When IRR is Required:** ```yaml irr_requirements: required_for: - "Content analysis with frequency claims" - "Deductive coding with structured codebook" - "Dissertation/thesis research" - "High-stakes publication (top journals)" optional_for: - "Exploratory inductive research" - "Single-researcher qualitative studies" - "Phenomenological research" best_practice: "Always recommended for transparency and rigor" ``` **IRR Process:** ```yaml irr_process: step_1_codebook_training: activity: "Train second coder on codebook" materials: ["Codebook", "Example coded transcripts", "Decision rules"] time: "4-8 hours typical" step_2_independent_coding:
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub