Improving perturbation-based explanations by understanding the role of uncertainty calibration
Published in: 39th Conference on Neural Information Processing Systems (NeurIPS 2025)
Thomas Decker¹⸍²⸍³, Volker Tresp²⸍³, Florian Büttner⁴⸍⁵⸍⁶
(¹ Siemens, ² LMU Munich, ³ Munich Center for Machine Learning (MCML), ⁴ Goethe University Frankfurt, ⁵ German Cancer Research Center, ⁶ German Cancer Consortium)
Perturbation-based explanations enhance ML transparency, but unknown model behavior under perturbations compromises their reliability. We prove that models produce unreliable probability estimates under these perturbations, directly degrading local and global explanation quality. We propose ReCalX to recalibrate models for better explanations without altering original predictions, improving robustness.
View at publisher's page
Incremental uncertainty-aware performance monitoring with active labeling intervention
Published in: Proceedings of The 28th International Conference on Artificial Intelligence and Statistics, PMLR 258:2188‑2196, 2025
Alexander Koebler, Thomas Decker, Ingo Thon, Volker Tresp, Florian Buettner
We study monitoring ML models under gradual distribution shifts, which slowly degrade accuracy undetected. We propose incremental uncertainty-aware performance monitoring (IUPM), a label-free method that estimates performance changes using optimal transport. IUPM quantifies prediction uncertainty and uses active labeling under a limited budget to restore estimates, outperforming baselines in shift scenarios.
View at publisher's page
Wiki-TabNER: Integrating named entity recognition into Wikipedia tables
Published in: 2025 International ACM SIGIR Conference on Research and Development in Information Retrieval
Aneta Koleva, Martin Ringsquandl, Ahmed Hatem, Thomas Runkler, Volker Tresp
Existing table interpretation benchmarks often oversimplify real-world data. To address this, we introduce Wiki-TabNER, a challenging dataset featuring complex Wikipedia tables with multiple entities per cell, annotated using DBpedia classes. Designed for in-table named entity recognition (NER) and entity linking, we detail its features, labeling and an LLM evaluation framework, followed by qualitative analysis of model performance and dataset limitations.
View at publisher's page
Lightning UQ Box: Uncertainty quantification for neural networks
Published in: Journal of Machine Learning Research 26 (2025)
Nils Lehmann, Nina Maria Gottschling, Jakob Gawlikowski, Adam J. Stewart, Stefan Depeweg, Eric Nalisnick
Despite deep learning's success, its black-box nature and missing confidence estimates cause skepticism in critical fields like medicine. To make uncertainty quantification (UQ) more accessible and reusable, we introduce Lightning UQ Box, a PyTorch-based library powered by PyTorch Lightning. It supports classification, regression and segmentation across diverse UQ methods, giving practitioners scalable, ready-to-use tools.
View at publisher's page
Fine-grained uncertainty decomposition in large language models: A spectral approach
Published in: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI26)
Nassim Walha¹⸍²⸍⁴, Sebastian G. Gruber⁵, Thomas Decker⁶⸍⁷⸍⁸, Yinchong Yang⁶, Alireza Javanmardi⁷⸍⁸, Eyke Hüllermeier⁷⸍⁸⸍⁹, Florian Buettner¹⸍²⸍³⸍⁴
(¹ German Cancer Research Center (DKFZ), ² German Cancer Consortium (DKTK), ³ Frankfurt Cancer Institute, Germany, ⁴ Goethe University Frankfurt, Germany, ⁵ ESAT-PSI, KU Leuven, Belgium, ⁶ Siemens AG, ⁷ LMU Munich, ⁸ Munich Center for Machine Learning (MCML), ⁹ German Center for Artificial Intelligence (DFKI), Kaiserslautern, Germany)
As LLMs integrate into diverse applications, obtaining reliable predictive uncertainty measures is critical. Distinguishing between aleatoric data ambiguity and epistemic model limitations is essential to address each source. We introduce spectral uncertainty, an approach using quantum information theory’s Von Neumann entropy to rigorously separate total uncertainty into aleatoric and epistemic components. By leveraging semantic similarity, it outperforms state-of-the-art methods across diverse models and benchmarks.
View at publisher's page