원클릭으로
alignment-files
alignment_files skill
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
alignment_files skill
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
| name | alignment_files |
| description | alignment_files skill |
| metadata | {"short-description":"alignment_files skill","category":"utilities","source":"claude-code-templates"} |
Pysam provides the AlignmentFile class for reading and writing SAM/BAM/CRAM formatted files containing aligned sequence data. BAM/CRAM files support compression and random access through indexing.
Specify format via mode qualifier:
"rb" - Read BAM (binary)"r" - Read SAM (text)"rc" - Read CRAM (compressed)"wb" - Write BAM"w" - Write SAM"wc" - Write CRAMimport pysam
# Reading
samfile = pysam.AlignmentFile("example.bam", "rb")
# Writing (requires template or header)
outfile = pysam.AlignmentFile("output.bam", "wb", template=samfile)
Use "-" as filename for stdin/stdout operations:
# Read from stdin
infile = pysam.AlignmentFile('-', 'rb')
# Write to stdout
outfile = pysam.AlignmentFile('-', 'w', template=infile)
Important: Pysam does not support reading/writing from true Python file objects—only stdin/stdout streams are supported.
Header Information:
references - List of chromosome/contig nameslengths - Corresponding lengths for each referenceheader - Complete header as dictionarysamfile = pysam.AlignmentFile("example.bam", "rb")
print(f"References: {samfile.references}")
print(f"Lengths: {samfile.lengths}")
Retrieves reads overlapping specified genomic regions using 0-based coordinates.
# Fetch specific region
for read in samfile.fetch("chr1", 1000, 2000):
print(read.query_name, read.reference_start)
# Fetch entire contig
for read in samfile.fetch("chr1"):
print(read.query_name)
# Fetch without index (sequential read)
for read in samfile.fetch(until_eof=True):
print(read.query_name)
Important Notes:
until_eof=True for non-indexed files or sequential readingfetch("*") or until_eof=TrueWhen using multiple iterators on the same file:
samfile = pysam.AlignmentFile("example.bam", "rb", multiple_iterators=True)
iter1 = samfile.fetch("chr1", 1000, 2000)
iter2 = samfile.fetch("chr2", 5000, 6000)
Without multiple_iterators=True, a new fetch() call repositions the file pointer and breaks existing iterators.
# Count all reads
num_reads = samfile.count("chr1", 1000, 2000)
# Count with quality filter
num_quality_reads = samfile.count("chr1", 1000, 2000, quality=20)
Returns four arrays (A, C, G, T) with per-base coverage:
coverage = samfile.count_coverage("chr1", 1000, 2000)
a_counts, c_counts, g_counts, t_counts = coverage
Each read is represented as an AlignedSegment object with these key attributes:
query_name - Read name/IDquery_sequence - Read sequence (bases)query_qualities - Base quality scores (ASCII-encoded)query_length - Length of the readreference_name - Chromosome/contig namereference_start - Start position (0-based, inclusive)reference_end - End position (0-based, exclusive)mapping_quality - MAPQ scorecigarstring - CIGAR string (e.g., "100M")cigartuples - CIGAR as list of (operation, length) tuplesImportant: cigartuples format differs from SAM specification. Operations are integers:
flag - SAM flag as integeris_paired - Is read paired?is_proper_pair - Is read in a proper pair?is_unmapped - Is read unmapped?mate_is_unmapped - Is mate unmapped?is_reverse - Is read on reverse strand?mate_is_reverse - Is mate on reverse strand?is_read1 - Is this read1?is_read2 - Is this read2?is_secondary - Is secondary alignment?is_qcfail - Did read fail QC?is_duplicate - Is read a duplicate?is_supplementary - Is supplementary alignment?get_tag(tag) - Get value of optional fieldset_tag(tag, value) - Set optional fieldhas_tag(tag) - Check if tag existsget_tags() - Get all tags as list of tuplesfor read in samfile.fetch("chr1", 1000, 2000):
if read.has_tag("NM"):
edit_distance = read.get_tag("NM")
print(f"{read.query_name}: NM={edit_distance}")
header = {
'HD': {'VN': '1.0'},
'SQ': [
{'LN': 1575, 'SN': 'chr1'},
{'LN': 1584, 'SN': 'chr2'}
]
}
outfile = pysam.AlignmentFile("output.bam", "wb", header=header)
# Create new read
a = pysam.AlignedSegment()
a.query_name = "read001"
a.query_sequence = "AGCTTAGCTAGCTACCTATATCTTGGTCTTGGCCG"
a.flag = 0
a.reference_id = 0 # Index into header['SQ']
a.reference_start = 100
a.mapping_quality = 20
a.cigar = [(0, 35)] # 35M
a.query_qualities = pysam.qualitystring_to_array("IIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIII")
# Write to file
outfile.write(a)
# BAM to SAM
infile = pysam.AlignmentFile("input.bam", "rb")
outfile = pysam.AlignmentFile("output.sam", "w", template=infile)
for read in infile:
outfile.write(read)
infile.close()
outfile.close()
The pileup() method provides column-wise (position-by-position) analysis across a region:
for pileupcolumn in samfile.pileup("chr1", 1000, 2000):
print(f"Position {pileupcolumn.pos}: coverage = {pileupcolumn.nsegments}")
for pileupread in pileupcolumn.pileups:
if not pileupread.is_del and not pileupread.is_refskip:
# Query position is the position in the read
base = pileupread.alignment.query_sequence[pileupread.query_position]
print(f" {pileupread.alignment.query_name}: {base}")
Key attributes:
pileupcolumn.pos - 0-based reference positionpileupcolumn.nsegments - Number of reads covering positionpileupread.alignment - The AlignedSegment objectpileupread.query_position - Position in the read (None for deletions)pileupread.is_del - Is this a deletion?pileupread.is_refskip - Is this a reference skip (N in CIGAR)?Important: Keep iterator references alive. The error "PileupProxy accessed after iterator finished" occurs when iterators go out of scope prematurely.
Critical: Pysam uses 0-based, half-open coordinates (Python convention):
reference_start is 0-based (first base is 0)reference_end is exclusive (not included in range)Exception: Region strings in fetch() and pileup() follow samtools conventions (1-based):
# These are equivalent:
samfile.fetch("chr1", 999, 2000) # Python style: 0-based
samfile.fetch("chr1:1000-2000") # samtools style: 1-based
Create BAM index:
pysam.index("example.bam")
Or use command-line interface:
pysam.samtools.index("example.bam")
pileup() for column-wise analysis instead of repeated fetch operationsfetch(until_eof=True) for sequential reading of non-indexed filescount() for simple counting instead of iterating and counting manuallyfetch() returns reads that overlap region boundaries—implement explicit filtering if exact boundaries are neededquery_qualities in place after modifying query_sequence. Create a copy first: quals = read.query_qualitiesfetch() without until_eof=True requires an index fileInvoke this skill with:
$alignment_files [arguments]
Or let Codex auto-select based on your prompt.
Long autonomous task execution with iteration control. Use for multi-hour refactors, TDD workflows, batch operations, or any task requiring sustained autonomous work.
Persistent memory system across Codex sessions. Use when you need to remember facts, recall information, or maintain context between sessions.
Create, manage, and merge git worktrees for parallel development. Use when starting parallel features, running multiple Codex instances, or for isolated development.
3p-updates skill
Web accessibility specialist for WCAG compliance, ARIA implementation, and inclusive design. Use when auditing websites for accessibility issues, implementing WCAG 2.1 AA/AAA standards, testing with screen readers, or ensuring ADA compliance. Expert in semantic HTML, keyboard navigation, and assistive technology compatibility.
Cloud laboratory platform for automated protein testing and validation. Use when designing proteins and needing experimental validation including binding assays, expression testing, thermostability measurements, enzyme activity assays, or protein sequence optimization. Also use for submitting experiments via API, tracking experiment status, downloading results, optimizing protein sequences for better expression using computational tools (NetSolP, SoluProt, SolubleMPNN, ESM), or managing protein design workflows with wet-lab validation.