Skip to main content

bio-clinical-biostatistics-cdisc-data

Reads, validates, and prepares CDISC SDTM and ADaM clinical trial data for analysis. Covers SDTM domain joins (DM, AE, EX, VS, LB, DS), ADaM architecture (ADSL, BDS, OCCDS, ADTTE) with traceability, treatment-emergent AE conventions, baseline derivation, SUPPQUAL/NSV handling, Define-XML 2.1, and Pinnacle 21 / CORE validation. Use when working with clinical trial datasets in CDISC SDTM/ADaM format, preparing analysis-ready data, or validating for regulatory submission.

ソース情報

リポジトリ
GPTomics/bioSkills
ソースの最終更新活動
2026年7月18日 11:42
検出された SKILL.md の言語
英語
スター
1,209
フォーク
251

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。

ファイルエクスプローラー
3 ファイル

SKILL.md を表示中

SKILL.md
ソースの指示 · 読み取り専用プレビュー
name
bio-clinical-biostatistics-cdisc-data
description
Reads, validates, and prepares CDISC SDTM and ADaM clinical trial data for analysis. Covers SDTM domain joins (DM, AE, EX, VS, LB, DS), ADaM architecture (ADSL, BDS, OCCDS, ADTTE) with traceability, treatment-emergent AE conventions, baseline derivation, SUPPQUAL/NSV handling, Define-XML 2.1, and Pinnacle 21 / CORE validation. Use when working with clinical trial datasets in CDISC SDTM/ADaM format, preparing analysis-ready data, or validating for regulatory submission.
tool_type
python
primary_tool
pyreadstat
goal_approach_exempt
true
## Version Compatibility Reference examples tested with: pyreadstat 1.2+, pandas 2.1+, numpy 1.26+. CDISC standards referenced: SDTM 2.0 / SDTMIG 3.4 (SDTM 3.0 / SDTMIG 4.0 in public review through April 2026); ADaMIG v1.3 (2021); OCCDS v1.1 (Nov 2021); BDS-for-TTE v1.0; Define-XML 2.1 (FDA-recommended for studies starting on/after March 15, 2023); Dataset-JSON v1.1 (Dec 2024; FDA Federal Register notice April 2025); Pinnacle 21 Community 4.0+; CORE (CDISC Open Rules Engine, 2021). Define-XML 2.1 FDA support began March 15, 2021 and is required for studies starting on/after March 15, 2023. Before using code patterns, verify installed versions match. If versions differ: - Python: `pip show <package>` then `help(module.function)` to check signatures - R packages cited (essential for ADaM derivation): admiral (Roche/openpharma), metacore, metatools, xportr If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying. # CDISC SDTM and ADaM Data Handling **"Load clinical trial data"** -> Parse CDISC SDTM domain files; build or consume ADaM analysis-ready datasets; preserve subject-level and event-level structure; respect traceability and validation expectations for regulatory submission. - Python: `pyreadstat.read_xport()`, `pd.read_sas()`, `pd.merge()` - R: `haven::read_xpt()`, `admiral` for ADaM derivation, `Pinnacle21` or `CORE` for validation ## Aggregation Strategy Taxonomy -- Choose the Right Question | Strategy | Scientific question answered | Example endpoint | Fails when | |----------|------------------------------|------------------|------------| | Any event (binary) | Does treatment change probability of experiencing the event at all? | Had any serious AE: Yes/No | Treatment changes event burden but not anyone-event probability | | Event count | Does treatment change burden of events per patient? | Total AE count per subject | Subjects with 1 vs 10 events treated equivalently | | Maximum severity | Does treatment shift patients toward more severe manifestations? | Worst AESEV per subject | Confounded with event count (more events -> higher chance of severe) | | First event + time | Does treatment delay onset of the event? | Time to first serious AE (TTE) | Multiple events per subject ignored | | Rate (events per person-time) | What is the per-time-unit rate? | AEs per subject-year | Requires exposure-time tracking; differential dropout biases rates | | Composite (per ICH E9 R1) | Event becomes part of endpoint definition | Death = treatment failure | Direction of components conflict; needs hierarchy | These are NOT interchangeable. A drug might not change the proportion with AEs (binary: no effect) but increase events per patient (count: harmful). **The choice must be pre-specified in the SAP based on the scientific question, not analytic convenience.** ## Decision Tree by Scenario | Scenario | Recommended aggregation | Why | |----------|------------------------|-----| | Primary safety endpoint, single SAE event | Any event (binary); analyse with logistic | Standard regulatory; cite FDA Safety Reporting Guidance | | Total adverse-event burden across study | Event count per subject; analyse with Poisson or negative binomial | Captures all events; sandwich SE recommended | | Toxicity grade comparison across arms | Max severity per subject; ordinal logistic with PO check | Preserves grade ordering; cite Brant test for PO | | Time-to-first AE (Kaplan-Meier visualisation) | First event + time; censor non-events | See clinical-biostatistics/survival-analysis | | Rate of exacerbations per patient-year | Rate via negative binomial with offset for exposure time | Standard in COPD/asthma trials | | Composite endpoint (e.g., MACE) | Component-level definition with hierarchy | Pre-specify per ICH E9(R1) composite strategy | | Stratification factor extraction | Use STRATA1, STRATA2 from RANDB or DM SUPP | Must appear in analysis (Kahan-Morris 2012) | | Baseline value derivation | VSBLFL='Y' / LBBLFL='Y'; derive latest pre-dose only if flag missing | Trust SDTM flag when present | ## SDTM vs ADaM -- The Regulatory Layer Cake | Layer | Standard | Purpose | Granularity | Examples | |-------|----------|---------|-------------|----------| | Source CRF | EDC system | Raw data capture | Form/page | Rave, Medidata, Veeva | | SDTM | SDTM 2.0 / SDTMIG 3.4 | Tabulation; "what happened" | One row per observation | DM, AE, EX, VS, LB, DS | | ADaM | ADaMIG v1.3 (2021); v3.0 in development | Analysis-ready; "one-PROC-away from CSR table" | Subject (ADSL), parameter-timepoint (BDS), occurrence (OCCDS) | ADSL, ADAE, ADLB, ADTTE, ADRS | | TLF | Sponsor SAS / R / Python | Tables, listings, figures for CSR | Output | Statistical methods section, demographic table, primary efficacy | **The ADaM Fundamental Principles (the "ROT" document):** 1. Analysis-ready (one procedure call -> the analysis result) 2. Traceability (every value links back to SDTM via metadata) 3. Clear/unambiguous communication via Define-XML 4. Naming conventions (PARAM, PARAMCD, AVAL, AVALC, BASE, CHG, PCHG, ABLFL, ANL01FL, ...) 5. Structural rules (ADSL one row per subject; BDS one row per subject/parameter/timepoint/analysis flag) **Postdoc reading:** the ADaM IG v1.3 PDF (cdisc.org), ADaM ROT, FDA Study Data Technical Conformance Guide (current 2024 version), Pinnacle 21 validation rule catalog, PHUSE Connect 2023-2025 conference proceedings. ## SDTM Domain Overview | Domain | Level | Description | Key Variables | |--------|-------|-------------|---------------| | DM | Subject | Demographics (one row per subject) | USUBJID, ARM, ARMCD, ACTARM, ACTARMCD, AGE, SEX, RACE, RFSTDTC, RFXSTDTC, RFENDTC | | AE | Event | Adverse events (multiple per subject) | USUBJID, AETERM, AEDECOD, AEBODSYS, AESEV, AESER, AESTDTC, AEENDTC | | EX | Event | Drug exposure/dosing | USUBJID, EXTRT, EXDOSE, EXSTDTC, EXENDTC | | VS | Event | Vital signs | USUBJID, VSTESTCD, VSSTRESN, VSBLFL, VISIT | | LB | Event | Lab results | USUBJID, LBTESTCD, LBSTRESN, LBSTRESC, LBORRES, LBBLFL, LBSPEC | | DS | Event | Disposition | USUBJID, DSDECOD, DSSTDTC | | SE | Event | Subject elements (treatment epochs) | USUBJID, ETCD, SESTDTC, SEENDTC | | MH | Event | Medical history | USUBJID, MHDECOD, MHCAT | | CM | Event | Concomitant medications | USUBJID, CMDECOD, CMSTDTC, CMENDTC | USUBJID = STUDYID-SITEID-SUBJID is the universal merge key. Subject-level domains (DM) have one row per USUBJID; event-level domains have multiple. **ARM vs ACTARM:** ARM is planned treatment from randomisation; ACTARM is actual treatment received. **In crossover designs, ARM differs from ACTARM by definition;** in parallel-arm trials, they diverge when subjects are randomised to one arm but receive another (per-protocol violations). Primary analyses use ARM (ITT); safety uses ACTARM. **RFSTDTC vs RFXSTDTC:** RFSTDTC is "first study activity date" (typically screening start); RFXSTDTC is "first treatment date." For treatment-emergent adverse event (TEAE) calculations, ALWAYS use RFXSTDTC (per ICH E2A) — RFSTDTC includes screening AEs which are not treatment-emergent. ## Reading .xpt Files ```python import pyreadstat import pandas as pd # pyreadstat (recommended -- handles SAS metadata) dm, meta = pyreadstat.read_xport('dm.xpt') # meta.column_names, meta.column_labels, meta.variable_value_labels # pandas built-in (SAS XPORT v5) dm = pd.read_sas('dm.xpt', format='xport', encoding='utf-8') # CSV fallback (common in academic datasets) dm = pd.read_csv('DM.csv') ``` When pyreadstat is available, the metadata object provides column labels, value labels, and format information lost with other readers. **Critical for analysis-dataset derivation:** the metadata carries the controlled-terminology codelist, essential for handling values like AESEV ('MILD'/'MODERATE'/'SEVERE') with semantic ordering. ## SAS XPT v5 vs Dataset-JSON -- The 2025-2026 Transition **SAS XPT v5** is the current FDA submission format but dates to 1995, with constraints: - 8-character variable names (so `LBSTRESN` is a max-length name) - 200-character text values - No UTF-8 (ASCII only) -> problematic for multilingual trials - Single dataset per file **Dataset-JSON v1.1 (CDISC, December 2024; FDA Federal Register notice April 2025)** is the modern replacement. PHUSE-CDISC-FDA pilot has demonstrated drop-in feasibility. FDA adoption timeline pending as of mid-2026; EMA and PMDA exploring in parallel. **Pragmatic position:** for the next ~2 years, SAS XPT v5 will remain the de facto submission format; sponsors should architect for Dataset-JSON migration but maintain XPT compliance. ## Joining Domains -- The Right Way ```python import pandas as pd dm = pd.read_csv('DM.csv') ae = pd.read_csv('AE.csv') # WRONG: merging event-level directly onto subject-level inflates rows # RIGHT: aggregate first, then merge any_serious = ae.groupby('USUBJID')['AESER'].apply(lambda x: (x == 'Y').any()).reset_index() any_serious.columns = ['USUBJID', 'HAD_SERIOUS_AE'] analysis = dm.merge(any_serious, on='USUBJID', how='left') analysis['HAD_SERIOUS_AE'] = analysis['HAD_SERIOUS_AE'].fillna(False) ``` **Always use `how='left'` when merging onto DM** to preserve all randomised subjects, even those with no events. Fill missing event indicators with 0 or False. ### Aggregation strategy must follow the scientific question | Strategy | Scientific question | Example | |----------|---------------------|---------| | Any event (binary) | Does treatment increase probability of experiencing the event at all? | Had any serious AE: yes/no | | Event count | Does treatment increase event burden per patient? | Total AE count per subject | | Maximum severity | Does treatment shift toward more severe manifestations? | Worst AESEV per subject | | First event + time | Does treatment delay onset? | Time to first serious AE | | Rate (events per person-time) | What is the per-time-unit rate? | AEs per subject-year | These are NOT interchangeable. A drug might not change the proportion with AEs (binary: no effect) but increase events per patient (count: harmful). The choice must follow the SAP, not analytic convenience. ```python # Count events per subject ae_counts = ae.groupby('USUBJID').size().reset_index(name='AE_COUNT') # Maximum severity per subject (map to numeric first -- string max is unreliable) severity_map = {'MILD': 1, 'MODERATE': 2, 'SEVERE': 3} ae['AESEV_NUM'] = ae['AESEV'].map(severity_map) max_severity = ae.groupby('USUBJID')['AESEV_NUM'].max().reset_index() # Specific event: COVID-19 adverse event covid_ae = ae[ae['AEDECOD'] == 'COVID-19'] covid_ae['AESEV_NUM'] = covid_ae['AESEV'].map(severity_map) had_covid = covid_ae.groupby('USUBJID')['AESEV_NUM'].max().reset_index() had_covid.columns = ['USUBJID', 'COVID_SEVERITY'] analysis = dm.merge(had_covid, on='USUBJID', how='left') analysis['HAD_COVID'] = analysis['COVID_SEVERITY'].notna().astype(int) ``` ## ADaM Architecture -- The Postdoc Deep Dive ### ADSL (Subject-Level) -- The Spine **Exactly one row per subject.** Every other ADaM dataset must merge to ADSL on USUBJID. Standard variables: - **USUBJID** -- universal subject ID - **TRT01A / TRT01P / ACTARMCD / ARMCD** -- planned and actual treatment, period 1 - **TRTSDT / TRTEDT** -- treatment start/end dates (derived from EX, not SDTM) - **AGE, SEX, RACE, ETHNIC** -- demographics from DM - **RANDDT** -- randomisation date - **DCSREAS / DCSREASP / DCSREASCD** -- discontinuation reason (coded + verbatim) - **Population flags:** ITTFL, FASFL, SAFFL, PPROTFL, EFFFL (Y/N flags for analysis populations) - **Stratification factors:** STRATA1, STRATA2 (from randomisation) - **Baseline covariates** that will be used as model covariates downstream ### BDS (Basic Data Structure) -- Long Format Analysis Data **One row per subject per parameter per analysis timepoint per analysis flag.** Used for ADVS, ADLB, ADEFF, ADQS, ADTTE. Required variables: - **USUBJID, STUDYID** -- merge keys - **PARAM, PARAMCD, PARAMN** -- parameter name, code, number - **AVISIT, AVISITN** -- analysis visit name, number - **ADT, ADY** -- analysis date, analysis day (relative to TRTSDT) - **AVAL, AVALC** -- analysis value (numeric, character) - **BASE** -- baseline value (replicated per subject/parameter) - **CHG, PCHG** -- change from baseline, percent change - **ABLFL** -- 'Y' for the baseline record - **ANL01FL, ANL02FL** -- analysis flags for primary/secondary analyses - **DTYPE** -- derivation type ('LOCF', 'WOCF', 'AVERAGE', 'BOCF', or null for original) - **BASETYPE** -- when multiple baselines per subject/parameter (crossover) - **EPOCH** -- study period (SCREENING, TREATMENT, FOLLOW-UP) ```python # Example: derive ADLB BDS structure from LB SDTM import pandas as pd lb = pd.read_csv('LB.csv') adsl = pd.read_csv('ADSL.csv') # Filter to active tests adlb = lb[lb['LBTESTCD'].isin(['ALT', 'AST', 'CREAT', 'HGB'])].copy() adlb['AVAL'] = adlb['LBSTRESN'] adlb['PARAM'] = adlb['LBTEST'] adlb['PARAMCD'] = adlb['LBTESTCD'] # Merge subject-level treatment from ADSL adlb = adlb.merge(adsl[['USUBJID', 'TRT01A', 'TRTSDT']], on='USUBJID') # Compute analysis day adlb['ADT'] = pd.to_datetime(adlb['LBDTC'], errors='coerce') adlb['TRTSDT'] = pd.to_datetime(adlb['TRTSDT'], errors='coerce') adlb['ADY'] = (adlb['ADT'] - adlb['TRTSDT']).dt.days + 1 # Day 1 = first dose # Set ABLFL from LBBLFL
GitHubで見る
この SKILL.md は非常に大きいため、SkillsMP では最初のセクションだけを表示しています。 GitHubで見る