As input, this model expects a sequence of coded medical events that have been mapped to Standard Concepts within the
OMOP-CDM vocabulary
. The model generates representations of patients which can then be used for downstream prediction tasks.
Input patients should be provided in the
MEDS
schema.
This model is intended to generate representations for patients based on the structured data within their electronic health record.
These representations can then be used for downstream tasks such as predicting diagnoses, detecting anomalies, or doing propensity score matching for causal inference.
Direct Use
You will likely want to tune the model for your downstream use case.
Out-of-Scope Use
This model is for research purposes only. It is not for use in any real-world decision making that impacts patients, providers, or hospital operations.
Bias, Risks, and Limitations
This model was trained on a corpus of 2.57 million patients from Stanford Medicine.
The model will thus reflect the patterns of how care is delivered at Stanford Medicine, in addition to the racial and socioeconomic makeup of Stanford Medicine's patient base.
This model may not generalize well to other hospitals and demographic mixes.
While this is technically a generative model, we have not tested its generative abilities and thus do not anticipate it being used to generate synthetic EHR records.
We aim to explore its generative abilities in future work.
The model is trained on 2.57 million patients from the
Stanford Medicine Research Data Repository (STARR)
, which contains EHR data from both Stanford Health Care (primarily adult care)
and Lucile Packard Children’s Hospital (primarily pediatric care).
The dataset contains only structured data (i.e. no clinical text or images) and covers demographics (e.g. age, sex, race), diagnoses, procedures, laboratory results, medication prescriptions, and other coded clinical observations.
The data is formatted according to the
Observational Medical Outcomes Partnership Common Data Model (OMOP-CDM)
.
All data that we work with is deidentified.
Training Procedure
We train our model using an autoregressive next code prediction objective, i.e. predict the next code in a patient's timeline given their previous codes.
Preprocessing
We use the
FEMR
Python library for data preprocessing.
Information on this benchmark, tasks, and results are detailed in
Wornow et al. 2023
Technical Specifications
This model uses the CLMBR architecture from
(Steinberg et al. 2021)
.
The objective is an autoregressive next token prediction task.
Please see
Wornow et al. 2023
for more details on the specific model architecture.
Vocabulary
CLMBR is a language model and requires defining a token vocabulary
V
. However, unlike natural languages, the vocabulary of a structured EHR language model is defined by
medical codes
. Here tokens map to standardized concepts in medical ontologies. Since the union of all tokens from all ontologies,
V_all
, results in a prohibitively large vocabuary, we derive
~V
by filtering to the top
k
most frequent codes as follows:
Knowledge Graphs (G):
A set of
n
medical ontologies (knowledge graphs),
G = ({G_1, G_2, ..., G_n})
, defined by
Athena's OMOP Vocabulary List
.
Medical Codes as Tokens:
Each knowledge graph
G_i
has a set of unique medical codes
M_i
. The union of all these codes serve as the tokens in our complete vocabulary
V_all = M_1 ∪ M_2 ∪ ... ∪ M_n
. Our final, filtered vocabulary is then
~V = sort_freq(V_all)[1:k]
where frequency is calculated over our
STARR EHR OMOP
dataset.
CLMBR Vocabulary Summary
21 Source Ontologies/Knowledge Graphs
65,536 tokens (the max value of
uint16_t
)
PREFIX
SOURCE
SIZE
EXAMPLE TOKENS
LOINC
Logical Observation Identifiers Names and Codes (Regenstrief Institute)
37,590
31790-9, 20449-5
SNOMED
Systematic Nomenclature of Medicine - Clinical Terms (IHTSDO)
18,174
105013009, 200755008
RxNorm
RxNorm (NLM)
4,678
2375327, 372375
CPT4
Current Procedural Terminology version 4 (AMA)
3,730
00790, 36818
RxNorm Extension
OMOP RxNorm Extension
255
OMOP358911, OMOP2153393
ICD10PCS
ICD-10 Procedure Coding System (CMS)
233
10907ZC, 4A0234Z
ICD9Proc
International Classification of Diseases, Ninth Revision, Clinical Modification, Volume 3 (NCHS)
International Classification of Diseases for Oncology, Third Edition (WHO)
52
NULL-C34.8, C56.9
CVX
CDC Vaccine Administered CVX (NCIRD)
41
151, 158
Domain
OMOP
27
OMOP generated
Race
Race and Ethnicity Code Set (USBC)
5
5, 4
OMOP Extension
OMOP Extension (OHDSI)
3
OMOP5160861, OMOP4912978
Gender
OMOP Gender
2
F, M
Ethnicity
OMOP Ethnicity
2
Not Hispanic, Hispanic
CMS Place of Service
Place of Service Codes for Professional Claims (CMS)
2
OMOP4822036, 02
Medicare Specialty
Medicare provider/supplier specialty codes (CMS)
1
A0
Condition Type
OMOP
1
OMOP4822053
CARE_SITE
STANFORD_CUSTOM
396
7930934, 7929373
Visit
STANFORD_CUSTOM
6
ERIP, ER
Citation
BibTeX:
@article{wornow2023ehrshot,
title={EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models},
author={Michael Wornow and Rahul Thapa and Ethan Steinberg and Jason Fries and Nigam Shah},
booktitle={Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track},
year={2023}
}
Model Card Authors
Michael Wornow, Ethan Steinberg, Rahul Thapa, Jason Fries, Nigam H. Shah
clmbr-t-base huggingface.co is an AI model on huggingface.co that provides clmbr-t-base's model effect (), which can be used instantly with this StanfordShahLab clmbr-t-base model. huggingface.co supports a free trial of the clmbr-t-base model, and also provides paid use of the clmbr-t-base. Support call clmbr-t-base model through api, including Node.js, Python, http.
clmbr-t-base huggingface.co is an online trial and call api platform, which integrates clmbr-t-base's modeling effects, including api services, and provides a free online trial of clmbr-t-base, you can try clmbr-t-base online for free by clicking the link below.
StanfordShahLab clmbr-t-base online free url in huggingface.co:
clmbr-t-base is an open source model from GitHub that offers a free installation service, and any user can find clmbr-t-base on GitHub to install. At the same time, huggingface.co provides the effect of clmbr-t-base install, users can directly use clmbr-t-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.