Researchers from Paige and Microsoft Research, including Eugene Vorontsov, Thomas J. Fuchs and Nicolo Fusi, described PRISM2, a multimodal slide-level foundation model for computational pathology, in an arXiv preprint first submitted 16 June 2025 and revised 31 October 2025. The model was trained on 700,000 diagnostic specimen-report pairs covering 2.3 million whole slide images and 14 million question-answer pairs, which the authors call the largest vision-and-language histopathology dataset to date. Supervision came from clinical dialogue, aligning histomorphologic features with diagnostic reasoning language so the model supports both direct diagnostic question-answering and transferable slide-level embeddings. The authors report that without additional training PRISM2 matches or exceeds the cancer-detection performance of clinical-grade products, and that task-specific finetuning on a large dataset beats dedicated survival-prediction models.
- arxiv.org2026-08-04