2026
-
Sparse Concept Attribution for Histomorphological Hypothesis Generation from Whole-Slide Classifiers
2026
Abstract
Histology images contain rich morphological information and can provide insights into pathological processes. However, deriving hypotheses relating morphological phenotypes to clinical attributes is bottlenecked by a manual image interpretation step. Here, we demonstrate that this process can be automated through interpretable deep learning. We present SCOPE, a method to interpret slide-level classifiers by combining pathology-specific vision-language models with sparse concept attribution onto a generalist histomorphological concept bank. To measure whether such explanations recover known morphology, we introduce MorphoRecoveryBench, a benchmark of seven tasks with pathologist-curated reference descriptions. On this benchmark, dense concept attribution is indistinguishable from a random baseline, whereas sparse attribution recovers substantial known morphology; decomposing the pooled slide embedding reaches similar explanation correctness at a fraction of the computational cost. Post-hoc interpretation of whole-slide classifiers can thus generate morphological hypotheses at scale, for expert validation.
BibTeX
@misc{lazardSparseConceptAttribution2026, title = {Sparse Concept Attribution for Histomorphological Hypothesis Generation from Whole-Slide Classifiers}, author = {Lazard, Tristan and Bouzid, Kenza and Hense, Julius and Bannur, Shruthi and {Coelho de Castro}, Daniel and Shao, Daniel and Jena, Rajesh and Williamson, Drew and Hyland, Stephanie}, year = {2026}, month = sep, eprint = {2609.02985}, archiveprefix = {arXiv}, primaryclass = {q-bio.QM}, doi = {10.48550/arXiv.2609.02985}, }
2025
-
Self-Supervised Learning to Predict Intrahepatic Cholangiocarcinoma Transcriptomic Classes on Routine Histology
JHEP Reports · 2025
Abstract
This study applies self-supervised learning to predict transcriptomic classes of intrahepatic cholangiocarcinoma from routine whole-slide histological images, demonstrating that SSL representations can effectively capture molecular subtypes from morphology alone.
BibTeX
@article{beaufrereSSLCholangiocarcinoma2024, title = {Self-Supervised Learning to Predict Intrahepatic Cholangiocarcinoma Transcriptomic Classes on Routine Histology}, author = {Lazard*, Tristan and Beaufr{\`e}re*, Aur{\'e}lie and Gardrat, Sophie and Couchy, Giuliana and Zucman-Rossi, Jessica and Decenci{\`e}re, Etienne and Walter, Thomas and Paradis, Val{\'e}rie}, year = {2025}, journal = {JHEP Reports}, pages = {101675}, doi = {10.1016/j.jhepr.2025.101675}, issn = {2589-5559}, } -
NOVA: An Agentic Framework for Automated Histopathology Analysis and Discovery
2025
Abstract
We present NOVA, an agentic framework that automates complex histopathology analyses by translating scientific queries into executable pipelines. NOVA introduces the SlideQuest benchmark for multi-step computational pathology evaluation and demonstrates strong performance for scalable discovery in digital pathology.
BibTeX
@misc{vaidyaNOVAAgenticFramework2025, title = {NOVA: An Agentic Framework for Automated Histopathology Analysis and Discovery}, author = {Vaidya, Anurag J. and Meissen, Felix and Castro, Daniel C. and Bannur, Shruthi and Lazard, Tristan and Williamson, Drew F. K. and Mahmood, Faisal and Alvarez-Valle, Javier and Hyland, Stephanie L. and Bouzid, Kenza}, year = {2025}, eprint = {2511.11324}, archiveprefix = {arXiv}, primaryclass = {cs.CV}, doi = {10.48550/arXiv.2511.11324}, }
2023
-
Democratizing Computational Pathology: Optimized Whole Slide Image Representations for The Cancer Genome Atlas
bioRxiv · 2023
Abstract
Automatic analysis of hematoxylin and eosin (H&E) stained Whole Slide Images (WSI) bears great promise for computer assisted diagnosis and biomarker discovery. However, scarcity of annotated datasets leads to underperforming models. Furthermore, the size and complexity of the image data limit their integration into bioinformatic workflows and thus their adoption by the bioinformatics community. Here, we present Giga-SSL, a self-supervised method for learning WSI representations without any annotation. We show that applying a simple linear classifier on the Giga-SSL representations improves classification performance over the fully supervised alternative on five benchmarked tasks and across different datasets. Moreover, we observe a substantial performance increase for small datasets (average gain of 7 AUC point) and a doubling of the number of mutations predictable from WSIs in a pan-cancer setting (from 45 to 93). We make the WSI representations available, compressing the TCGA-FFPE images from 12TB to 23MB and enabling fast analysis on a laptop CPU. We hope this resource will facilitate multimodal data integration in order to analyze WSI in their genomic and transcriptomic context.
BibTeX
@misc{lazardDemocratizingComputationalPathology2023, title = {Democratizing Computational Pathology: Optimized {{Whole Slide Image}} Representations for {{The Cancer Genome Atlas}}}, shorttitle = {Democratizing Computational Pathology}, author = {Lazard, Tristan and Lerousseau, Marvin and Gardrat, Sophie and {Vincent-Salomon}, Anne and Stern, Marc-Henri and Rodrigues, Manuel and Decenci{\`e}re, Etienne and Walter, Thomas}, year = {2023}, month = dec, primaryclass = {Confirmatory Results}, pages = {2023.12.04.569894}, publisher = {{bioRxiv}}, doi = {10.1101/2023.12.04.569894}, urldate = {2024-01-19}, archiveprefix = {bioRxiv}, chapter = {Confirmatory Results}, copyright = {{\textcopyright} 2023, Posted by Cold Spring Harbor Laboratory. This pre-print is available under a Creative Commons License (Attribution 4.0 International), CC BY 4.0, as described at http://creativecommons.org/licenses/by/4.0/}, langid = {english}, file = {/Users/trislaz/Documents/syncthings/bibliographie/cbio/my_papers/Lazard et al_2023_Democratizing computational pathology.pdf} } -
Giga-SSL: Self-Supervised Learning for Gigapixel Images
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) · 2023
BibTeX
@inproceedings{lazardGigaSSLSelfSupervisedLearning2023a, title = {Giga-{{SSL}}: {{Self-Supervised Learning}} for {{Gigapixel Images}}}, shorttitle = {Giga-{{SSL}}}, booktitle = {2023 {{IEEE}}/{{CVF Conference}} on {{Computer Vision}} and {{Pattern Recognition Workshops}} ({{CVPRW}})}, author = {Lazard, Tristan and Lerousseau, Marvin and Decenci{\`e}re, Etienne and Walter, Thomas}, year = {2023}, month = jun, pages = {4305--4314}, publisher = {{IEEE}}, address = {{Vancouver, BC, Canada}}, doi = {10.1109/CVPRW59228.2023.00453}, urldate = {2023-11-15}, isbn = {9798350302493}, langid = {english}, file = {/Users/trislaz/Zotero/storage/S3N2YNNI/Lazard et al. - 2023 - Giga-SSL Self-Supervised Learning for Gigapixel I.pdf} } -
Automatic Grading of Cervical Biopsies by Combining Full and Self-supervision
Computer Vision – ECCV 2022 Workshops · 2023
Abstract
In computational pathology, predictive models from Whole Slide Images (WSI) mostly rely on Multiple Instance Learning (MIL), where the WSI are represented as a bag of tiles, each of which is encoded by a Neural Network (NN). Slide-level predictions are then achieved by building models on the agglomeration of these tile encodings. The tile encoding strategy thus plays a key role for such models. Current approaches include the use of encodings trained on unrelated data sources, full supervision or self-supervision. While self-supervised learning (SSL) exploits unlabeled data, it often requires large computational resources to train. On the other end of the spectrum, fully-supervised methods make use of valuable prior knowledge about the data but involve a costly amount of expert time. This paper proposes a framework to reconcile SSL and full supervision, showing that a combination of both provides efficient encodings, both in terms of performance and in terms of biological interpretability. On a recently organized challenge on grading Cervical Biopsies, we show that our mixed supervision scheme reaches high performance (weighted accuracy (WA): 0.945), outperforming both SSL (WA: 0.927) and transfer learning from ImageNet (WA: 0.877). We further shed light upon the internal representations that trigger classification results, providing a method to reveal relevant phenotypic patterns for grading cervical biopsies. We expect that the combination of full and self-supervision is an interesting strategy for many tasks in computational pathology and will be widely adopted by the field.
BibTeX
@inproceedings{lubranoAutomaticGradingCervical2023, title = {Automatic {{Grading}} of~{{Cervical Biopsies}} by~{{Combining Full}} and~{{Self-supervision}}}, booktitle = {Computer {{Vision}} {\textendash} {{ECCV}} 2022 {{Workshops}}}, author = {Lubrano, M{\'e}lanie and Lazard, Tristan and Balezo, Guillaume and {Bellahsen-Harrar}, Ya{\"e}lle and Badoual, C{\'e}cile and Berlemont, Sylvain and Walter, Thomas}, editor = {Karlinsky, Leonid and Michaeli, Tomer and Nishino, Ko}, year = {2023}, series = {Lecture {{Notes}} in {{Computer Science}}}, pages = {408--423}, publisher = {{Springer Nature Switzerland}}, address = {{Cham}}, doi = {10.1007/978-3-031-25082-8_27}, isbn = {978-3-031-25082-8}, langid = {english}, keywords = {Histopathology,Mixed supervision,Self-supervised learning,Whole-slide classification}, file = {/Users/trislaz/Documents/syncthings/bibliographie/cbio/Democratizing/Lubrano et al_2023_Automatic Grading of Cervical Biopsies by Combining Full and Self-supervision.pdf} }
2022
-
Deep Learning Identifies Morphological Patterns of Homologous Recombination Deficiency in Luminal Breast Cancers from Whole Slide Images
Cell Reports Medicine · 2022
Abstract
Homologous recombination DNA-repair deficiency (HRD) is becoming a well-recognized marker of platinum salt and polyADP-ribose polymerase inhibitor chemotherapies in ovarian and breast cancers. While large-scale screening for HRD using genomic markers is logistically and economically challenging, stained tissue slides are routinely acquired in clinical practice. With the objectives of providing a robust deep-learning method for HRD prediction from tissue slides and identifying related morphological phenotypes, we first show that digital pathology workflows are sensitive to potential biases in the training set, then we propose a method to overcome the influence of these biases, and we develop an interpretation method capable of identifying complex phenotypes. Application to our carefully curated in-house dataset allows us to predict HRD with high accuracy (area under the receiver-operator characteristics curve 0.86) and to identify morphological phenotypes related to HRD. In particular, the presence of laminated fibrosis and clear tumor cells associated with HRD open new hypotheses regarding its phenotypic impact.
BibTeX
@article{lazardDeepLearningIdentifies2022a, title = {Deep Learning Identifies Morphological Patterns of Homologous Recombination Deficiency in Luminal Breast Cancers from Whole Slide Images}, author = {Lazard, Tristan and Bataillon, Guillaume and Naylor, Peter and Popova, Tatiana and Bidard, Fran{\c c}ois-Cl{\'e}ment and {Stoppa-Lyonnet}, Dominique and Stern, Marc-Henri and Decenci{\`e}re, Etienne and Walter, Thomas and {Vincent-Salomon}, Anne}, year = {2022}, month = dec, journal = {Cell Reports Medicine}, volume = {3}, number = {12}, pages = {100872}, issn = {2666-3791}, doi = {10.1016/j.xcrm.2022.100872}, urldate = {2023-11-15}, keywords = {bias,breast cancer,computational pathology,deep learning,homologous recombination deficiency,interpretability,molecular subtype,prediction,self-supervised learning,whole slide images}, file = {/Users/trislaz/Documents/syncthings/bibliographie/cbio/Democratizing/Lazard et al_2022_Deep learning identifies morphological patterns of homologous recombination.pdf;/Users/trislaz/Zotero/storage/FJ6UTQJ5/S2666379122004360.html} } -
Prediction of Treatment Response in Triple Negative Breast Cancer From Whole Slide Images
Frontiers in Signal Processing · 2022
Abstract
The automatic analysis of stained histological sections is becoming increasingly popular. Deep Learning is today the method of choice for the computational analysis of such data, and has shown spectacular results for large datasets for a large variety of cancer types and prediction tasks. On the other hand, many scientific questions relate to small, highly specific cohorts. Such cohorts pose serious challenges for Deep Learning, typically trained on large datasets. In this article, we propose a modification of the standard nested cross-validation procedure for hyperparameter tuning and model selection, dedicated to the analysis of small cohorts. We also propose a new architecture for the particularly challenging question of treatment prediction, and apply this workflow to the prediction of response to neoadjuvant chemotherapy for Triple Negative Breast Cancer.
BibTeX
@article{naylorPredictionTreatmentResponse2022, title = {Prediction of {{Treatment Response}} in {{Triple Negative Breast Cancer From Whole Slide Images}}}, author = {Naylor, Peter and Lazard, Tristan and Bataillon, Guillaume and La{\'e}, Marick and {Vincent-Salomon}, Anne and Hamy, Anne-Sophie and Reyal, Fabien and Walter, Thomas}, year = {2022}, journal = {Frontiers in Signal Processing}, volume = {2}, issn = {2673-8198}, doi = {10.3389/frsip.2022.851809}, file = {/Users/trislaz/Documents/syncthings/bibliographie/cbio/Democratizing/Naylor et al_2022_Prediction of Treatment Response in Triple Negative Breast Cancer From Whole.pdf} }