Publication Link

Why This Matters Link to heading

Predictive models built on intensive-care unit (ICU) clinical notes are increasingly used to forecast patient mortality. But the pre-processing of those free-text notes (cleaning, tokenizing, stemming, TF-IDF weighting, n-gram creation, etc.) is usually treated as a routine step. This study asks a simple question: does the choice of note-preparation strategy actually matter for model performance?

Core Objectives Link to heading

  1. Evaluate several common text-pre-processing pipelines (raw, cleaned, stemmed, TF-IDF, n-grams).
  2. Quantify their impact on the discriminative ability (AUROC) of mortality-prediction models.
  3. Test robustness across three algorithm families: penalized logistic regression, feed-forward neural networks, and random-forest classifiers.

Data & Experimental Design Link to heading

  • Cohort: Adult ICU admissions from the University of California, San Francisco (UCSF), and externally validated on Beth Israel Deaconess Medical Center (BIDMC).
  • Outcome: In-hospital mortality.
  • Features: Unstructured clinical note text (e.g., progress notes, discharge summaries).
  • Pre-processing variants:
    • Raw text (no processing)
    • Cleaned text (removing PHI, punctuation, stop-words)
    • Stemming (Porter stemmer)
    • TF-IDF vectorization
    • N-gram generation (bi-/tri-grams).
  • Model training: 10-fold cross-validation on the UCSF dataset.
  • Evaluation metric: Area Under the Receiver Operating Characteristic curve (AUROC).

Key Findings Link to heading

AUROC of models trained on UCSF and validated on BIDMC.

Pre-processing (Stacked)Logistic-Regression AUROCNeural-Network AUROCRandom-Forest AUROC
Raw text0.720.760.67
Cleaned text0.750.780.68
Stemming0.770.800.71
TF-IDF0.830.810.79
N-grams0.800.780.77

Takeaway: TF-IDF vectorization consistently yielded the highest AUROC improvement across all three model families.

Interpretation Link to heading

TF-IDF works here because weighted term frequencies capture both term importance and sparsity, which helps linear models and tree-based ensembles alike. The gain held across penalized logistic regression, deep neural nets, and random forests, so the benefit comes from the input representation, not the architecture. In practice: if your pipeline consumes clinical notes, run at least a cleaning step followed by TF-IDF (optionally with bi-/tri-grams) before the data hits the model.

Limitations & Future Directions Link to heading

  • Trained on single-center data: Results are derived from UCSF ICU records; external validation on other hospitals is needed to confirm generalizability.
  • Scope of outcomes: The study focused solely on mortality; other clinically relevant predictions (e.g., length of stay, sepsis onset) may respond differently to preprocessing choices.

Bottom Line for Practitioners Link to heading

If you’re building a predictive model that ingests free-text ICU notes, don’t settle for “just clean the text.” Use TF-IDF (with optional n-grams) as your baseline preprocessing pipeline. It gives a measurable AUROC lift across all three algorithm families and costs almost nothing compared to deep contextual embeddings.