Javascript must be enabled to continue!
DeepSeek for Pathology Report Understanding: A Benchmark Study of Cancer Type Extraction, AJCC Staging, and Prognosis Prediction
View through CrossRef
Abstract
Background:
Pathology reports contain clinically important information for cancer diagnosis, staging, and outcome research, but their unstructured format limits scalable reuse. We evaluated whether DeepSeek can support pathology report understanding across cancer type extraction, AJCC stage prediction, and prognosis classification.
Methods:
We conducted an independent benchmark on PathRep-Bench style pathology report tasks derived from The Cancer Genome Atlas. The dataset contained 9,523 reports split into 7,618 training, 953 validation, and 952 test reports. DeepSeek V4 Flash and DeepSeek V4 Pro were evaluated using JSON-constrained prompting. Cancer type identification used non-reasoning prompting, whereas AJCC staging and prognosis prompting used reasoning-mode prompting. We also trained local TF-IDF logistic regression prognosis models and tested whether DeepSeek-extracted structured variables improved supervised prognosis modeling.
Results:
DeepSeek V4 Flash achieved 0.9800 accuracy (95% CI, 0.9695--0.9884) and 0.9778 macro F1 (95% CI, 0.9660--0.9867) for cancer type identification. For AJCC stage prediction, DeepSeek V4 Flash achieved 0.8165 accuracy (95% CI, 0.7845--0.8468) and 0.7841 macro F1 (95% CI, 0.7421--0.8201) on 594 labeled test reports. DeepSeek V4 Pro did not significantly improve paired accuracy over Flash for either task and produced 14 empty final outputs during AJCC staging. DeepSeek prognosis prompting was weak on a 100-report subset, improving from 0.5083 to 0.5647 macro F1 with eight examples. A supervised TF-IDF text-only logistic regression classifier achieved the highest prognosis point estimate on the full test set, with 0.8571 accuracy and 0.8543 macro F1. Adding DeepSeek-extracted structured variables did not produce a statistically significant improvement.
Conclusions:
DeepSeek is effective for pathology information extraction and provides strong AJCC staging support, but prognosis prediction is better framed as supervised outcome modeling under the evaluated framework. These findings support a hybrid architecture in which DeepSeek performs extraction and staging support, while supervised text-based models handle prognosis prediction.
Title: DeepSeek for Pathology Report Understanding: A Benchmark Study of Cancer Type Extraction, AJCC Staging, and Prognosis Prediction
Description:
Abstract
Background:
Pathology reports contain clinically important information for cancer diagnosis, staging, and outcome research, but their unstructured format limits scalable reuse.
We evaluated whether DeepSeek can support pathology report understanding across cancer type extraction, AJCC stage prediction, and prognosis classification.
Methods:
We conducted an independent benchmark on PathRep-Bench style pathology report tasks derived from The Cancer Genome Atlas.
The dataset contained 9,523 reports split into 7,618 training, 953 validation, and 952 test reports.
DeepSeek V4 Flash and DeepSeek V4 Pro were evaluated using JSON-constrained prompting.
Cancer type identification used non-reasoning prompting, whereas AJCC staging and prognosis prompting used reasoning-mode prompting.
We also trained local TF-IDF logistic regression prognosis models and tested whether DeepSeek-extracted structured variables improved supervised prognosis modeling.
Results:
DeepSeek V4 Flash achieved 0.
9800 accuracy (95% CI, 0.
9695--0.
9884) and 0.
9778 macro F1 (95% CI, 0.
9660--0.
9867) for cancer type identification.
For AJCC stage prediction, DeepSeek V4 Flash achieved 0.
8165 accuracy (95% CI, 0.
7845--0.
8468) and 0.
7841 macro F1 (95% CI, 0.
7421--0.
8201) on 594 labeled test reports.
DeepSeek V4 Pro did not significantly improve paired accuracy over Flash for either task and produced 14 empty final outputs during AJCC staging.
DeepSeek prognosis prompting was weak on a 100-report subset, improving from 0.
5083 to 0.
5647 macro F1 with eight examples.
A supervised TF-IDF text-only logistic regression classifier achieved the highest prognosis point estimate on the full test set, with 0.
8571 accuracy and 0.
8543 macro F1.
Adding DeepSeek-extracted structured variables did not produce a statistically significant improvement.
Conclusions:
DeepSeek is effective for pathology information extraction and provides strong AJCC staging support, but prognosis prediction is better framed as supervised outcome modeling under the evaluated framework.
These findings support a hybrid architecture in which DeepSeek performs extraction and staging support, while supervised text-based models handle prognosis prediction.
Related Results
Educators’ Perspectives on DeepSeek in ELT: A Qualitative Case Study of Pedagogical Potentials and Pitfalls in Chinese Higher Education
Educators’ Perspectives on DeepSeek in ELT: A Qualitative Case Study of Pedagogical Potentials and Pitfalls in Chinese Higher Education
Aim/Purpose: This study aimed to investigate the perspectives of English Language Teaching (ELT) educators on DeepSeek, emphasizing its pedagogical value, practical challenges, and...
Hydatid Disease of The Brain Parenchyma: A Systematic Review
Hydatid Disease of The Brain Parenchyma: A Systematic Review
Abstarct
Introduction
Isolated brain hydatid disease (BHD) is an extremely rare form of echinococcosis. A prompt and timely diagnosis is a crucial step in disease management. This ...
Breast Carcinoma within Fibroadenoma: A Systematic Review
Breast Carcinoma within Fibroadenoma: A Systematic Review
Abstract
Introduction
Fibroadenoma is the most common benign breast lesion; however, it carries a potential risk of malignant transformation. This systematic review provides an ove...
Abstract 3463: Beyond staging systems in cutaneous squamous cell carcinoma: Validation of two thresholds for a model-estimated metastatic risk
Abstract 3463: Beyond staging systems in cutaneous squamous cell carcinoma: Validation of two thresholds for a model-estimated metastatic risk
Abstract
Background: Cutaneous squamous cell carcinoma (cSCC) is the second most common skin cancer, therefore causing a death toll comparable to that of melanoma, d...
A Survey of DeepSeek Models
A Survey of DeepSeek Models
Advances in artificial intelligence (AI) rely on systems capable of human-like reasoning, a limitation for conventional Large Language Models (LLMs), which struggle with multi-step...
Algorithmic implementation of pancreatic cancer staging guidelines: comparison with a retrieval-augmented large language model
Algorithmic implementation of pancreatic cancer staging guidelines: comparison with a retrieval-augmented large language model
Abstract
Purpose
To implement a comprehensive knowledge-based algorithm (KBA) for pancreatic cancer staging based on the curren...
Abstract P4-02-11: Can preoperative axillary staging replace sentinel node biopsy? Comparison of preoperative axillary and final histologic nodal findings in 2108 patients with primary breast cancer
Abstract P4-02-11: Can preoperative axillary staging replace sentinel node biopsy? Comparison of preoperative axillary and final histologic nodal findings in 2108 patients with primary breast cancer
Abstract
Background Axillary staging is an integral part of the preoperative work-up in patients with breast cancer. Given the increasing number of patients treated ...
Are Cervical Ribs Indicators of Childhood Cancer? A Narrative Review
Are Cervical Ribs Indicators of Childhood Cancer? A Narrative Review
Abstract
A cervical rib (CR), also known as a supernumerary or extra rib, is an additional rib that forms above the first rib, resulting from the overgrowth of the transverse proce...

