Skip to main content 홈 크리에이터 freedomintelligence openclaw-medical-skills bio-consensus-sequences
bio-consensus-sequences Generate consensus FASTA sequences by applying VCF variants to a reference using bcftools consensus. Use when creating sample-specific reference sequences or reconstructing haplotypes.
설치로 이동 Skills Marketplace 커뮤니티가 만든 AI 스킬을 발견하고 탐색하세요.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
직접 명령은 검토 Prompt를 거치지 않습니다. 실행하기 전에 소스를 확인하세요.
npx skills add https://github.com/FreedomIntelligence/OpenClaw-Medical-Skills --skill bio-consensus-sequences명령은 한 줄로 유지됩니다. 복사하기 전에 가로로 스크롤해 전체 내용을 확인하세요.
로컬 사본을 원하시나요? SkillsMP에서 현재 제공할 수 있는 파일을 다운로드하세요.
Zip 다운로드 다운로드 중... 이 저장소의 다른 Skills Convert raw Nanopore signal data (FAST5/POD5) to nucleotide sequences using Dorado basecaller. Covers model selection, GPU acceleration, modified base detection, and quality filtering. Use when processing raw Nanopore data before alignment. Guppy is deprecated; use Dorado for all new analyses.
Production-ready PDF processing with forms, tables, OCR, validation, and batch operations. Use when working with complex PDF workflows in production environments, processing large volumes of PDFs, or requiring robust error handling and validation.
Extract text and tables from PDF files, fill forms, merge documents. Use when working with PDF files or when the user mentions PDFs, forms, or document extraction.
FreedomIntelligence
FreedomIntelligence/OpenClaw-Medical-Skills
GitHub 저장소 열기 name bio-consensus-sequences description Generate consensus FASTA sequences by applying VCF variants to a reference using bcftools consensus. Use when creating sample-specific reference sequences or reconstructing haplotypes. tool_type cli primary_tool bcftools
Version Compatibility
Reference examples tested with: BioPython 1.83+, bcftools 1.19+, bedtools 2.31+, minimap2 2.26+, samtools 1.19+
Before using code patterns, verify installed versions match. If versions differ:
Python: pip show <package> then help(module.function) to check signatures
CLI: <tool> --version then <tool> --help to confirm flags
If code throws ImportError, AttributeError, or TypeError, introspect the installed
package and adapt the example to match the actual API rather than retrying.
Consensus Sequences
"Generate a consensus sequence from my VCF" → Apply called variants to a reference FASTA, producing a sample-specific genome with optional haplotype selection and low-coverage masking.
CLI: bcftools consensus -f reference.fa input.vcf.gz
Python: cyvcf2 + Bio.SeqIO for simple SNP-only cases
Basic Usage
Generate Consensus
bcftools consensus -f reference.fa input.vcf.gz > consensus.fa
Specify Sample bcftools consensus -f reference.fa -s sample1 input.vcf.gz > sample1.fa
Output to File bcftools consensus -f reference.fa -o consensus.fa input.vcf.gz
Haplotype Selection
First Haplotype Only bcftools consensus -f reference.fa -H 1 input.vcf.gz > haplotype1.fa
Second Haplotype Only bcftools consensus -f reference.fa -H 2 input.vcf.gz > haplotype2.fa
Haplotype Options Option Description -H 1First haplotype -H 2Second haplotype -H AApply all ALT alleles -H RApply REF alleles where heterozygous -IApply IUPAC ambiguity codes (separate flag)
IUPAC Codes for Heterozygous Sites bcftools consensus -f reference.fa -I input.vcf.gz > consensus_iupac.fa
Heterozygous sites encoded with IUPAC ambiguity codes:
A/G → R
C/T → Y
A/C → M
G/T → K
A/T → W
C/G → S
Missing Data Handling
Mark Missing as N bcftools consensus -f reference.fa -M N input.vcf.gz > consensus.fa
Mark Low Coverage as N
samtools depth input.bam | awk '$3<10 {print $1"\t"$2-1"\t"$2}' > low_coverage.bed
bcftools consensus -f reference.fa -m low_coverage.bed input.vcf.gz > consensus.fa
Mask Options Option Description -m FILEMask regions in BED file with N -M CHARCharacter for masked regions (default N)
Region Selection
Specific Region bcftools consensus -f reference.fa -r chr1:1000-2000 input.vcf.gz > region.fa
Multiple Regions Use with BED file to extract multiple regions.
Chain Files
Generate Chain File bcftools consensus -f reference.fa -c chain.txt input.vcf.gz > consensus.fa
Chain files map coordinates between reference and consensus:
Useful for liftover of annotations
Required when indels change sequence length
Chain File Format chain score ref_name ref_size ref_strand ref_start ref_end query_name query_size query_strand query_start query_end id
Sample-Specific Consensus
For Each Sample for sample in $(bcftools query -l input.vcf.gz); do
bcftools consensus -f reference.fa -s "$sample " input.vcf.gz > "${sample} .fa"
done
Both Haplotypes sample="sample1"
bcftools consensus -f reference.fa -s "$sample " -H 1 input.vcf.gz > "${sample} _hap1.fa"
bcftools consensus -f reference.fa -s "$sample " -H 2 input.vcf.gz > "${sample} _hap2.fa"
Filtering Before Consensus
PASS Variants Only bcftools view -f PASS input.vcf.gz | \
bcftools consensus -f reference.fa > consensus.fa
High-Quality Variants Only bcftools filter -i 'QUAL>=30 && INFO/DP>=10' input.vcf.gz | \
bcftools consensus -f reference.fa > consensus.fa
SNPs Only bcftools view -v snps input.vcf.gz | \
bcftools consensus -f reference.fa > consensus_snps.fa
Sequence Naming
Default Naming Output uses reference sequence names.
Custom Prefix bcftools consensus -f reference.fa -p "sample1_" input.vcf.gz > consensus.fa
Sequences named: sample1_chr1, sample1_chr2, etc.
Common Workflows Goal: Generate consensus sequences for downstream analyses like phylogenetics, viral surveillance, or gene-level comparison.
Approach: Filter variants to high-quality calls, apply per-sample consensus generation, mask low-coverage regions with N, then combine for multi-sample workflows.
Phylogenetic Analysis Preparation
mkdir -p consensus
for sample in $(bcftools query -l cohort.vcf.gz); do
bcftools view -s "$sample " cohort.vcf.gz | \
bcftools view -c 1 | \
bcftools consensus -f reference.fa > "consensus/${sample} .fa"
done
cat consensus/*.fa > all_samples.fa
Viral Genome Assembly
bcftools filter -i 'QUAL>=30 && INFO/DP>=20' variants.vcf.gz | \
bcftools view -f PASS | \
bcftools consensus -f reference.fa -M N > consensus.fa
Gene-Specific Consensus
bcftools consensus -f reference.fa -r chr1:1000000-1010000 \
-s sample1 variants.vcf.gz > gene.fa
Masked Low-Coverage Regions
samtools depth -a input.bam | \
awk '$3<5 {print $1"\t"$2-1"\t"$2}' | \
bedtools merge > low_coverage.bed
bcftools consensus -f reference.fa -m low_coverage.bed \
variants.vcf.gz > consensus.fa
Verify Consensus
Check Differences
minimap2 -a reference.fa consensus.fa | samtools view -bS > alignment.bam
diff <(grep -v "^>" reference.fa) <(grep -v "^>" consensus.fa) | head
Count Changes
bcftools view -H input.vcf.gz | wc -l
Handling Overlapping Variants bcftools consensus handles overlapping variants automatically:
Applies variants in order
Warns about conflicts
bcftools consensus -f reference.fa input.vcf.gz 2>&1 | grep -i warn
cyvcf2 Consensus (Simple Cases)
Manual Consensus Generation from cyvcf2 import VCF
from Bio import SeqIO
ref_dict = {rec.id : str (rec.seq) for rec in SeqIO.parse('reference.fa' , 'fasta' )}
vcf = VCF('input.vcf.gz' )
changes = {}
for variant in vcf:
if variant.is_snp and len (variant.ALT) == 1 :
chrom = variant.CHROM
pos = variant.POS - 1
if chrom not in changes:
changes[chrom] = {}
changes[chrom][pos] = variant.ALT[0 ]
for chrom, positions in changes.items():
seq = list (ref_dict[chrom])
for pos, alt in positions.items():
seq[pos] = alt
ref_dict[chrom] = '' .join(seq)
with open ('consensus.fa' , 'w' ) as f:
for chrom, seq in ref_dict.items():
f.write(f'>{chrom} \n{seq} \n' )
Note: Use bcftools consensus for production - handles indels and edge cases properly.
Quick Reference Task Command Basic consensus bcftools consensus -f ref.fa in.vcf.gzSpecific sample bcftools consensus -f ref.fa -s sample in.vcf.gzHaplotype 1 bcftools consensus -f ref.fa -H 1 in.vcf.gzIUPAC codes bcftools consensus -f ref.fa -I in.vcf.gzWith mask bcftools consensus -f ref.fa -m mask.bed in.vcf.gzGenerate chain bcftools consensus -f ref.fa -c chain.txt in.vcf.gzSpecific region bcftools consensus -f ref.fa -r chr1:1-1000 in.vcf.gz
Common Errors Error Cause Solution not indexedVCF not indexed Run bcftools index sequence not foundChromosome mismatch Check chromosome names overlapping recordsVariants overlap Usually OK, check warnings REF does not matchWrong reference Use same reference as caller
Related Skills
variant-calling - Generate VCF for consensus
filtering-best-practices - Filter variants before consensus
variant-normalization - Normalize indels first
alignment-files/reference-operations - Reference manipulation