GastroAGI Logo
OverviewBlogsAbout
Trending TopicsDaily BriefConference
Topics/Artificial Intelligence /Evaluation of NLP Algorithms for Identifying IBD from Clinical Records
43

Evaluation of NLP Algorithms for Identifying IBD from Clinical Records

Clinical knowledge base written and curated by GastroAGI Team from primary medical literatureLast updated October 1, 2025

The study conducted a comprehensive evaluation of 15 natural language processing (NLP) algorithms to identify patients with inflammatory bowel disease (IBD) from free-text secondary care clinical records. Here are the detailed findings related to the evaluation:

1. Study Objectives

The primary goal was to compare the performance of various NLP algorithms spanning 50 years of evolution, ranging from traditional rule-based models to advanced transformer-based and large language models (LLMs). The evaluation focused on multiple aspects, including accuracy, fairness, cost-efficiency, explainability, and environmental impact.


2. Dataset

The study utilized a robust dataset of 9311 labeled clinical documents from 1612 patients. This dataset ensured high diversity and statistical power, allowing for reliable model comparisons.


3. Algorithm Scope

The algorithms evaluated included:

  • Traditional NLP models: Rule-based regex, TF-IDF, and Bag-of-Words (BoW).
  • Transformer-based models: BERT, DistilBERT, sBERT.
  • Large language models (LLMs): DeepSeek, Mistral, Qwen.

4. Performance Results

Top Performer

  • DistilBERT_IBD: Achieved the highest accuracy with a micro F1 score of 93.54%, closely followed by sBERT with a score of 93.05%. However, both models exhibited moderate specificity limitations.

LLM Performance

  • LLMs like Mistral, Qwen, and DeepSeek performed well without exposure to training data, achieving micro F1 scores between 86.47% and 92.20%. However, they required longer runtimes and higher computational resources.

Traditional Models

  • Regex and spaCy negation models were fast and cost-effective but had poor specificity, making them unsuitable for precise patient identification.
  • TF-IDF and BoW demonstrated solid baseline performance (~92% F1) but lacked contextual understanding for complex medical text.

Document vs Patient-Level Accuracy

  • Document-level accuracy was higher compared to patient-level accuracy due to misalignment between document predictions and overall patient diagnoses, which reduced recall and calibration.

5. Fairness Findings

  • All models exhibited biases, particularly against women, wealthier individuals, and patients of African ethnicity.
  • LLMs were more balanced in terms of gender bias but displayed mild age-related bias, overpredicting IBD in older age groups.

6. Economic and Environmental Analysis

Cost Efficiency

  • Simple models like regex and TF-IDF were the fastest and cheapest, with processing times under 2 minutes and carbon emissions below 2g CO₂.
  • In contrast, the 70B parameter LLM consumed over 33,000g CO₂, highlighting sustainability concerns for large-scale deployment.

Energy Consumption

  • Transformer-based and LLM models required significantly more energy (up to 163 kWh) compared to traditional machine learning methods.

7. Explainability

  • Explainability tools such as SHAP and LIME were used to analyze model decision-making patterns and detect overfitting.
  • LLMs posed unique challenges for interpretability, requiring the development of new frameworks beyond traditional feature attribution methods.

8. Calibration

  • sBERT-Base demonstrated the best calibration with the lowest Brier scores, indicating reliable probability predictions compared to other models.

9. Document Contribution

  • Clinic letters were found to be the most predictive of IBD, with higher odds ratios (22.69–23.38) compared to endoscopy or pathology reports, reflecting their richer contextual information.

10. Clinical Implications

The study’s findings have significant implications for healthcare:

  • Transformer-based models like BERT and DistilBERT offer the best trade-off between accuracy, efficiency, and interpretability, making them ideal for clinical NLP deployment.
  • Open-source availability on platforms like Hugging Face and GitHub allows healthcare institutions to replicate or fine-tune models for their local electronic health record (EHR) systems.

11. Future Outlook

The study predicts that once issues related to cost, performance, and bias are addressed, LLMs will become the standard for clinical data retrieval and patient identification within the next decade.


12. Study Strengths

The study is notable for its:

  • Transparency and open-source code.
  • Multi-metric evaluation encompassing accuracy, fairness, cost, energy consumption, and explainability.
  • Robust dataset and diverse algorithm comparison, making it a benchmark for clinical NLP validation.

In summary, the evaluation highlighted the strengths and limitations of different NLP algorithms in identifying IBD from clinical records. While transformer-based models like DistilBERT and sBERT emerged as top performers, challenges such as fairness, environmental impact, and interpretability remain critical areas for improvement.

Related Q&A

44

Large language models (LLMs) with script concordance testing (SCT)

The benchmarking study that evaluated the clinical reasoning performance of large language models (LLMs) using Script Concordance Testing (SCT) provided several insights into their capabilities and limitations in...

45

TRIALSCOPE

TRIALSCOPE is a cutting-edge framework designed for clinical trial simulation using real-world data (RWD). It leverages advanced artificial intelligence (AI) and causal inference techniques to extract, clean, and...

46

CADe and CRC

Computer-Aided Detection (CADe) and Colorectal Cancer (CRC) are connected through the use of artificial intelligence (AI) technologies to improve the detection of polyps during colonoscopy procedures, which is...

47

AI-Based Computer-Aided Detection in Colonoscopy

AI-based computer-aided detection (CADe) in colonoscopy refers to the use of artificial intelligence (AI) systems to assist healthcare professionals in identifying abnormalities, such as polyps or adenomas, in...

01

AI in Medical Publishing: Assistance Is Acceptable, Authorship Is Human: BMJ Future Health | August 2026

Introduction: Generative AI is rapidly entering medical writing, research, clinical documentation, and evidence synthesis. The central question is no longer whether researchers will use AI, but where its...

02

Translating Artificial Intelligence in Pathology into Real Clinical Practice: NEJM AI | August 2026

Introduction: Artificial intelligence (AI) has demonstrated remarkable ability to extract diagnostic, prognostic, and molecular information from routine hematoxylin and eosin (H&E)-stained pathology slides. Despite this rapid scientific progress,...

GastroAGI Logo

We are pioneers in clinical intelligence, dedicated to helping gastroenterologists harness the power of artificial intelligence to drive precision, efficiency, and patient growth.

For You

For StudentsFor CliniciansFor ResearchersFor Patients

Core Tools

MELD-Na ScoreChild-PughFIB-4 IndexGlasgow-BlatchfordBISAP Score

Explore

OverviewAboutCalculators
Trending Topics
Conference Briefings
Blog Insights
©GastroAGI 2026
Privacy PolicyTerms of UseMedical Disclaimer