| name | medcat |
| description | Medical Concept Annotation Toolkit. Trainable NLP for extracting clinical concepts from unstructured text. Supports ICD-10, SNOMED CT, RxNorm, UMLS. Active learning for custom medical ontologies. |
| tags | ["clinical-nlp","medical-entity-extraction","icd10","snomed","umls","healthcare","zorai"] |
Overview
MedCAT trains NLP models for extracting clinical concepts from unstructured text. Supports ICD-10, SNOMED CT, RxNorm, UMLS, and custom ontologies with active learning.
Installation
uv pip install medcat
Pre-trained Model
from medcat.cat import CAT
cat = CAT.load_model_pack("medcat_model_pack.dat")
text = "Patient with type 2 diabetes and hypertension, prescribed metformin 500mg BID."
doc = cat(text)
for entity in doc.entities:
print(f"{entity.name:<25} {entity.cui:<10} confidence={entity.confidence:.2f}")
Active Learning
cat.add_cui_to_category("D003920", "Diabetes Mellitus")
cat.train(text="Patient has diabetes", cui="D003920", value="Diabetes Mellitus")
unmatched = cat.get_unmatched_concepts()
Workflow
- Load a pre-trained model pack
- Annotate clinical text -> extract UMLS CUIs
- Map concepts to ICD-10/SNOMED/RxNorm
- Train with active learning: correct errors, add concepts
- Export and deploy trained model