- name
- bio-clinical-biostatistics-cdisc-data
- description
- Reads, validates, and prepares CDISC SDTM and ADaM clinical trial data for analysis. Covers SDTM domain joins (DM, AE, EX, VS, LB, DS), ADaM architecture (ADSL, BDS, OCCDS, ADTTE) with traceability, treatment-emergent AE conventions, baseline derivation, SUPPQUAL/NSV handling, Define-XML 2.1, and Pinnacle 21 / CORE validation. Use when working with clinical trial datasets in CDISC SDTM/ADaM format, preparing analysis-ready data, or validating for regulatory submission.
- tool_type
- python
- primary_tool
- pyreadstat
- goal_approach_exempt
- true
## Version Compatibility
Reference examples tested with: pyreadstat 1.2+, pandas 2.1+, numpy 1.26+. CDISC standards referenced: SDTM 2.0 / SDTMIG 3.4 (SDTM 3.0 / SDTMIG 4.0 in public review through April 2026); ADaMIG v1.3 (2021); OCCDS v1.1 (Nov 2021); BDS-for-TTE v1.0; Define-XML 2.1 (FDA-recommended for studies starting on/after March 15, 2023); Dataset-JSON v1.1 (Dec 2024; FDA Federal Register notice April 2025); Pinnacle 21 Community 4.0+; CORE (CDISC Open Rules Engine, 2021). Define-XML 2.1 FDA support began March 15, 2021 and is required for studies starting on/after March 15, 2023.
Before using code patterns, verify installed versions match. If versions differ:
- Python: `pip show <package>` then `help(module.function)` to check signatures
- R packages cited (essential for ADaM derivation): admiral (Roche/openpharma), metacore, metatools, xportr
If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
# CDISC SDTM and ADaM Data Handling
**"Load clinical trial data"** -> Parse CDISC SDTM domain files; build or consume ADaM analysis-ready datasets; preserve subject-level and event-level structure; respect traceability and validation expectations for regulatory submission.
- Python: `pyreadstat.read_xport()`, `pd.read_sas()`, `pd.merge()`
- R: `haven::read_xpt()`, `admiral` for ADaM derivation, `Pinnacle21` or `CORE` for validation
## Aggregation Strategy Taxonomy -- Choose the Right Question
| Strategy | Scientific question answered | Example endpoint | Fails when |
|----------|------------------------------|------------------|------------|
| Any event (binary) | Does treatment change probability of experiencing the event at all? | Had any serious AE: Yes/No | Treatment changes event burden but not anyone-event probability |
| Event count | Does treatment change burden of events per patient? | Total AE count per subject | Subjects with 1 vs 10 events treated equivalently |
| Maximum severity | Does treatment shift patients toward more severe manifestations? | Worst AESEV per subject | Confounded with event count (more events -> higher chance of severe) |
| First event + time | Does treatment delay onset of the event? | Time to first serious AE (TTE) | Multiple events per subject ignored |
| Rate (events per person-time) | What is the per-time-unit rate? | AEs per subject-year | Requires exposure-time tracking; differential dropout biases rates |
| Composite (per ICH E9 R1) | Event becomes part of endpoint definition | Death = treatment failure | Direction of components conflict; needs hierarchy |
These are NOT interchangeable. A drug might not change the proportion with AEs (binary: no effect) but increase events per patient (count: harmful). **The choice must be pre-specified in the SAP based on the scientific question, not analytic convenience.**
## Decision Tree by Scenario
| Scenario | Recommended aggregation | Why |
|----------|------------------------|-----|
| Primary safety endpoint, single SAE event | Any event (binary); analyse with logistic | Standard regulatory; cite FDA Safety Reporting Guidance |
| Total adverse-event burden across study | Event count per subject; analyse with Poisson or negative binomial | Captures all events; sandwich SE recommended |
| Toxicity grade comparison across arms | Max severity per subject; ordinal logistic with PO check | Preserves grade ordering; cite Brant test for PO |
| Time-to-first AE (Kaplan-Meier visualisation) | First event + time; censor non-events | See clinical-biostatistics/survival-analysis |
| Rate of exacerbations per patient-year | Rate via negative binomial with offset for exposure time | Standard in COPD/asthma trials |
| Composite endpoint (e.g., MACE) | Component-level definition with hierarchy | Pre-specify per ICH E9(R1) composite strategy |
| Stratification factor extraction | Use STRATA1, STRATA2 from RANDB or DM SUPP | Must appear in analysis (Kahan-Morris 2012) |
| Baseline value derivation | VSBLFL='Y' / LBBLFL='Y'; derive latest pre-dose only if flag missing | Trust SDTM flag when present |
## SDTM vs ADaM -- The Regulatory Layer Cake
| Layer | Standard | Purpose | Granularity | Examples |
|-------|----------|---------|-------------|----------|
| Source CRF | EDC system | Raw data capture | Form/page | Rave, Medidata, Veeva |
| SDTM | SDTM 2.0 / SDTMIG 3.4 | Tabulation; "what happened" | One row per observation | DM, AE, EX, VS, LB, DS |
| ADaM | ADaMIG v1.3 (2021); v3.0 in development | Analysis-ready; "one-PROC-away from CSR table" | Subject (ADSL), parameter-timepoint (BDS), occurrence (OCCDS) | ADSL, ADAE, ADLB, ADTTE, ADRS |
| TLF | Sponsor SAS / R / Python | Tables, listings, figures for CSR | Output | Statistical methods section, demographic table, primary efficacy |
**The ADaM Fundamental Principles (the "ROT" document):**
1. Analysis-ready (one procedure call -> the analysis result)
2. Traceability (every value links back to SDTM via metadata)
3. Clear/unambiguous communication via Define-XML
4. Naming conventions (PARAM, PARAMCD, AVAL, AVALC, BASE, CHG, PCHG, ABLFL, ANL01FL, ...)
5. Structural rules (ADSL one row per subject; BDS one row per subject/parameter/timepoint/analysis flag)
**Postdoc reading:** the ADaM IG v1.3 PDF (cdisc.org), ADaM ROT, FDA Study Data Technical Conformance Guide (current 2024 version), Pinnacle 21 validation rule catalog, PHUSE Connect 2023-2025 conference proceedings.
## SDTM Domain Overview
| Domain | Level | Description | Key Variables |
|--------|-------|-------------|---------------|
| DM | Subject | Demographics (one row per subject) | USUBJID, ARM, ARMCD, ACTARM, ACTARMCD, AGE, SEX, RACE, RFSTDTC, RFXSTDTC, RFENDTC |
| AE | Event | Adverse events (multiple per subject) | USUBJID, AETERM, AEDECOD, AEBODSYS, AESEV, AESER, AESTDTC, AEENDTC |
| EX | Event | Drug exposure/dosing | USUBJID, EXTRT, EXDOSE, EXSTDTC, EXENDTC |
| VS | Event | Vital signs | USUBJID, VSTESTCD, VSSTRESN, VSBLFL, VISIT |
| LB | Event | Lab results | USUBJID, LBTESTCD, LBSTRESN, LBSTRESC, LBORRES, LBBLFL, LBSPEC |
| DS | Event | Disposition | USUBJID, DSDECOD, DSSTDTC |
| SE | Event | Subject elements (treatment epochs) | USUBJID, ETCD, SESTDTC, SEENDTC |
| MH | Event | Medical history | USUBJID, MHDECOD, MHCAT |
| CM | Event | Concomitant medications | USUBJID, CMDECOD, CMSTDTC, CMENDTC |
USUBJID = STUDYID-SITEID-SUBJID is the universal merge key. Subject-level domains (DM) have one row per USUBJID; event-level domains have multiple.
**ARM vs ACTARM:** ARM is planned treatment from randomisation; ACTARM is actual treatment received. **In crossover designs, ARM differs from ACTARM by definition;** in parallel-arm trials, they diverge when subjects are randomised to one arm but receive another (per-protocol violations). Primary analyses use ARM (ITT); safety uses ACTARM.
**RFSTDTC vs RFXSTDTC:** RFSTDTC is "first study activity date" (typically screening start); RFXSTDTC is "first treatment date." For treatment-emergent adverse event (TEAE) calculations, ALWAYS use RFXSTDTC (per ICH E2A) — RFSTDTC includes screening AEs which are not treatment-emergent.
## Reading .xpt Files
```python
import pyreadstat
import pandas as pd
# pyreadstat (recommended -- handles SAS metadata)
dm, meta = pyreadstat.read_xport('dm.xpt')
# meta.column_names, meta.column_labels, meta.variable_value_labels
# pandas built-in (SAS XPORT v5)
dm = pd.read_sas('dm.xpt', format='xport', encoding='utf-8')
# CSV fallback (common in academic datasets)
dm = pd.read_csv('DM.csv')
```
When pyreadstat is available, the metadata object provides column labels, value labels, and format information lost with other readers. **Critical for analysis-dataset derivation:** the metadata carries the controlled-terminology codelist, essential for handling values like AESEV ('MILD'/'MODERATE'/'SEVERE') with semantic ordering.
## SAS XPT v5 vs Dataset-JSON -- The 2025-2026 Transition
**SAS XPT v5** is the current FDA submission format but dates to 1995, with constraints:
- 8-character variable names (so `LBSTRESN` is a max-length name)
- 200-character text values
- No UTF-8 (ASCII only) -> problematic for multilingual trials
- Single dataset per file
**Dataset-JSON v1.1 (CDISC, December 2024; FDA Federal Register notice April 2025)** is the modern replacement. PHUSE-CDISC-FDA pilot has demonstrated drop-in feasibility. FDA adoption timeline pending as of mid-2026; EMA and PMDA exploring in parallel.
**Pragmatic position:** for the next ~2 years, SAS XPT v5 will remain the de facto submission format; sponsors should architect for Dataset-JSON migration but maintain XPT compliance.
## Joining Domains -- The Right Way
```python
import pandas as pd
dm = pd.read_csv('DM.csv')
ae = pd.read_csv('AE.csv')
# WRONG: merging event-level directly onto subject-level inflates rows
# RIGHT: aggregate first, then merge
any_serious = ae.groupby('USUBJID')['AESER'].apply(lambda x: (x == 'Y').any()).reset_index()
any_serious.columns = ['USUBJID', 'HAD_SERIOUS_AE']
analysis = dm.merge(any_serious, on='USUBJID', how='left')
analysis['HAD_SERIOUS_AE'] = analysis['HAD_SERIOUS_AE'].fillna(False)
```
**Always use `how='left'` when merging onto DM** to preserve all randomised subjects, even those with no events. Fill missing event indicators with 0 or False.
### Aggregation strategy must follow the scientific question
| Strategy | Scientific question | Example |
|----------|---------------------|---------|
| Any event (binary) | Does treatment increase probability of experiencing the event at all? | Had any serious AE: yes/no |
| Event count | Does treatment increase event burden per patient? | Total AE count per subject |
| Maximum severity | Does treatment shift toward more severe manifestations? | Worst AESEV per subject |
| First event + time | Does treatment delay onset? | Time to first serious AE |
| Rate (events per person-time) | What is the per-time-unit rate? | AEs per subject-year |
These are NOT interchangeable. A drug might not change the proportion with AEs (binary: no effect) but increase events per patient (count: harmful). The choice must follow the SAP, not analytic convenience.
```python
# Count events per subject
ae_counts = ae.groupby('USUBJID').size().reset_index(name='AE_COUNT')
# Maximum severity per subject (map to numeric first -- string max is unreliable)
severity_map = {'MILD': 1, 'MODERATE': 2, 'SEVERE': 3}
ae['AESEV_NUM'] = ae['AESEV'].map(severity_map)
max_severity = ae.groupby('USUBJID')['AESEV_NUM'].max().reset_index()
# Specific event: COVID-19 adverse event
covid_ae = ae[ae['AEDECOD'] == 'COVID-19']
covid_ae['AESEV_NUM'] = covid_ae['AESEV'].map(severity_map)
had_covid = covid_ae.groupby('USUBJID')['AESEV_NUM'].max().reset_index()
had_covid.columns = ['USUBJID', 'COVID_SEVERITY']
analysis = dm.merge(had_covid, on='USUBJID', how='left')
analysis['HAD_COVID'] = analysis['COVID_SEVERITY'].notna().astype(int)
```
## ADaM Architecture -- The Postdoc Deep Dive
### ADSL (Subject-Level) -- The Spine
**Exactly one row per subject.** Every other ADaM dataset must merge to ADSL on USUBJID. Standard variables:
- **USUBJID** -- universal subject ID
- **TRT01A / TRT01P / ACTARMCD / ARMCD** -- planned and actual treatment, period 1
- **TRTSDT / TRTEDT** -- treatment start/end dates (derived from EX, not SDTM)
- **AGE, SEX, RACE, ETHNIC** -- demographics from DM
- **RANDDT** -- randomisation date
- **DCSREAS / DCSREASP / DCSREASCD** -- discontinuation reason (coded + verbatim)
- **Population flags:** ITTFL, FASFL, SAFFL, PPROTFL, EFFFL (Y/N flags for analysis populations)
- **Stratification factors:** STRATA1, STRATA2 (from randomisation)
- **Baseline covariates** that will be used as model covariates downstream
### BDS (Basic Data Structure) -- Long Format Analysis Data
**One row per subject per parameter per analysis timepoint per analysis flag.** Used for ADVS, ADLB, ADEFF, ADQS, ADTTE.
Required variables:
- **USUBJID, STUDYID** -- merge keys
- **PARAM, PARAMCD, PARAMN** -- parameter name, code, number
- **AVISIT, AVISITN** -- analysis visit name, number
- **ADT, ADY** -- analysis date, analysis day (relative to TRTSDT)
- **AVAL, AVALC** -- analysis value (numeric, character)
- **BASE** -- baseline value (replicated per subject/parameter)
- **CHG, PCHG** -- change from baseline, percent change
- **ABLFL** -- 'Y' for the baseline record
- **ANL01FL, ANL02FL** -- analysis flags for primary/secondary analyses
- **DTYPE** -- derivation type ('LOCF', 'WOCF', 'AVERAGE', 'BOCF', or null for original)
- **BASETYPE** -- when multiple baselines per subject/parameter (crossover)
- **EPOCH** -- study period (SCREENING, TREATMENT, FOLLOW-UP)
```python
# Example: derive ADLB BDS structure from LB SDTM
import pandas as pd
lb = pd.read_csv('LB.csv')
adsl = pd.read_csv('ADSL.csv')
# Filter to active tests
adlb = lb[lb['LBTESTCD'].isin(['ALT', 'AST', 'CREAT', 'HGB'])].copy()
adlb['AVAL'] = adlb['LBSTRESN']
adlb['PARAM'] = adlb['LBTEST']
adlb['PARAMCD'] = adlb['LBTESTCD']
# Merge subject-level treatment from ADSL
adlb = adlb.merge(adsl[['USUBJID', 'TRT01A', 'TRTSDT']], on='USUBJID')
# Compute analysis day
adlb['ADT'] = pd.to_datetime(adlb['LBDTC'], errors='coerce')
adlb['TRTSDT'] = pd.to_datetime(adlb['TRTSDT'], errors='coerce')
adlb['ADY'] = (adlb['ADT'] - adlb['TRTSDT']).dt.days + 1 # Day 1 = first dose
# Set ABLFL from LBBLFL
GitHubで見る