Skip to main content

sequence-validity

Validate temporal ordering of date variables in a QUIQ-format table. Uses Claude CLI to automatically identify start/end date variable pairs, then checks whether start_date <= end_date for each matched record. No API key required. Use for LYDUS data quality assessment of chronological consistency.

Jump to install

Source facts

Repository
28sungmin/m4-add-skills
Last source activity
June 19, 2026 at 00:33
Detected SKILL.md language
Mixed languages
Stars
0
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

File Explorer
3 files

Showing SKILL.md

SKILL.md
Source instructions ยท Read-only preview
name
sequence-validity
description
Validate temporal ordering of date variables in a QUIQ-format table. Uses Claude CLI to automatically identify start/end date variable pairs, then checks whether start_date <= end_date for each matched record. No API key required. Use for LYDUS data quality assessment of chronological consistency.
tier
community
category
lydus
parameters
{"quiq_path":{"description":"Path to QUIQ-format CSV file. Must contain rows where Mapping_info_1 contains 'date'.","type":"string"},"save_path":{"description":"Directory path to save output files.","type":"string"}}
# Sequence Validity Validates the **chronological ordering** of date variables in a QUIQ-format table. Claude CLI automatically identifies meaningful (start, end) date pairs (e.g., admission โ†’ discharge), then checks whether `start_date โ‰ค end_date` for each matched patient record. ## When to Use This Skill - After QUIQ conversion, to detect records where events occur in impossible or illogical order (e.g., discharge before admission) - To assess chronological consistency of temporal data - As part of LYDUS quality management assessment ## SQL Support **Not applicable.** Claude CLI is required to identify date-variable pairs. ## Filtering Logic | Condition | Value | |-----------|-------| | `Mapping_info_1` | contains `date` (case-insensitive) | | `Value` | parsed as datetime | ## Pipeline 1. **Extract date rows** โ€” filter `Mapping_info_1` contains `date`, parse `Value` as datetime 2. **Collect unique identifiers** โ€” `Original_table_name - Variable_name` for all date variables 3. **LLM pair identification** โ€” send identifier list to Claude CLI; receive `timepoint_pairs` list 4. **LLM exclusion rules**: - Non-time variables excluded - Unpaired time variables excluded - Sensitive/complex variables excluded (death_time, year_of_birth, diagnosis_date, etc.) - Additional-context-required pairs excluded 5. **Validation** โ€” for each pair: merge on `(Patient_id, Original_table_name, Primary_key)`, check `Start_date โ‰ค End_date` 6. **Summary** โ€” per-(table, start_var, end_var): Total_num, Invalid_num, Sequence_Validity (%) ## Output | File | Description | |------|-------------| | `sequence_validity_total.txt` | Overall Sequence Validity (%), Total Num, Invalid Num | | `sequence_validity_summary.csv` | Per-(table, start_var, end_var): counts and Sequence_Validity (%) | | `sequence_validity_detail.csv` | Per-record: Start_date, End_date, Is_valid | ## How to Run ```python import pandas as pd from scripts.sequence_validity import get_sequence_validity quiq = pd.read_csv("/path/to/quiq.csv") df_total, df_summary = get_sequence_validity( quiq=quiq ) total_num = df_summary['Total_num'].sum() invalid_num = df_summary['Invalid_num'].sum() seq_validity = round((total_num - invalid_num) / total_num * 100, 2) print(f"Sequence Validity (%) = {seq_validity}") print(df_summary) ``` ### As a script with config ```yaml # config.yaml quiq_path: /path/to/quiq.csv save_path: /path/to/output ``` ```bash python scripts/sequence_validity.py --config config.yaml ``` ## Critical Notes 1. **Same-table constraint** โ€” pairs spanning different tables are skipped. Start and end variables must be from the same `Original_table_name`. 2. **LLM ์‘๋‹ต ํŒŒ์‹ฑ** โ€” Claude CLI๋Š” `timepoint_pairs = [(...), ...]` ํ˜•์‹์œผ๋กœ ์‘๋‹ตํ•ด์•ผ ํ•จ. `=` ๊ธฐ์ค€์œผ๋กœ ๋ถ„๋ฆฌ ํ›„ `ast.literal_eval` ํŒŒ์‹ฑ. ํ˜•์‹ ๋ถˆ์ผ์น˜ ์‹œ `ValueError` ๋ฐœ์ƒ. ์‘๋‹ต์ด ๋งˆํฌ๋‹ค์šด ์ฝ”๋“œ๋ธ”๋ก์„ ํฌํ•จํ•˜๋ฉด ํŒŒ์‹ฑ ์‹คํŒจํ•  ์ˆ˜ ์žˆ์œผ๋‹ˆ system prompt์˜ output format ์˜ˆ์‹œ๋ฅผ ๊ทธ๋Œ€๋กœ ๋”ฐ๋ฅด๋„๋ก ์„ค๊ณ„๋จ. 3. **๋‚ ์งœ ๋ณ€ํ™˜ ์‹คํŒจ** โ€” `pd.to_datetime(..., errors='coerce')`๋กœ ๋ณ€ํ™˜ ๋ถˆ๊ฐ€ํ•œ ๊ฐ’์€ NaT โ†’ `dropna()` ๋กœ ์ œ์™ธ๋จ. 4. **์›๋ณธ ์ฝ”๋“œ ๊ฐœ์„  ์‚ฌํ•ญ**: - `combined_time_df['Value'] = ...` SettingWithCopyWarning โ†’ `.copy()` ํ›„ ํ• ๋‹น - LLM ์‘๋‹ต ํŒŒ์‹ฑ `try/except` ์ถ”๊ฐ€ (์›๋ณธ์€ ํŒŒ์‹ฑ ์‹คํŒจ ์‹œ unhandled exception) - `_validate_sequence` ๋‚ด `pd.concat` loop โ†’ list ์ˆ˜์ง‘ ํ›„ ํ•œ ๋ฒˆ์— concat - `os.path.join` ์‚ฌ์šฉ (๋ฌธ์ž์—ด ์—ฐ๊ฒฐ ๋Œ€์‹ ) - `required=True` for `--config` 5. **Dependencies** โ€” `pandas`, `numpy` (LLM: Claude CLI via subprocess, timeout 180s) ## References - LYDUS ํ’ˆ์งˆ๊ด€๋ฆฌ ํ”„๋กœ๊ทธ๋žจ ํ™œ์šฉ ๊ฐ€์ด๋“œ๋ผ์ธ (๋น„๊ณต๊ฐœ ๋‚ด๋ถ€ ๋ฌธ์„œ) - Original Python implementation: LYDUS_Sequence_Validity.py (์ด์„ฑ๋ฏผ ์ž‘์„ฑ)
View on GitHub