EDBT 2026 Demo / reviewers in the wild / expert
Zhen Lin 0001
dblp:26/346-1
· DBLP profile ↗
11ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0001-8673-6868ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Trustworthy machine learning · 78% Deep learning architectures and training · 9% Language models and text generation · 6% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Medical and health informatics · 50% Bioinformatics and computational biology · 50% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 67% Privacy and data protection · 33% |
Topics — the 23 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
uncertainty estimation |
4.2 | 7 | 2025 | Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey · KDD (2) 2025 Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation · EMNLP 2024 Fast Online Value-Maximizing Prediction Sets with Conformal Cost Control · ICML 2023 |
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction |
2.2 | 5 | 2024 | Fast Online Value-Maximizing Prediction Sets with Conformal Cost Control · ICML 2023 Conformal Prediction with Temporal Quantile Adjustments · NeurIPS 2022 Locally Valid and Discriminative Prediction Intervals for Deep Learning Models · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › calibration
confidence calibration |
0.9 | 1 | 2025 | Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey · KDD (2) 2025 |
Natural language and speech › Language models and text generation › trustworthy language model
large language model reliability |
0.9 | 1 | 2025 | Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey · KDD (2) 2025 |
Machine learning › Trustworthy machine learning › uncertainty estimation
confidence estimation |
0.8 | 1 | 2024 | Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation · EMNLP 2024 |
Security and privacy of machine learning
adversarial robustness |
0.8 | 1 | 2024 | Certifiably Byzantine-Robust Federated Conformal Prediction · ICML 2024 |
Security and privacy of machine learning › federated learning defense
byzantine-robust federated learning |
0.8 | 1 | 2024 | Certifiably Byzantine-Robust Federated Conformal Prediction · ICML 2024 |
Privacy and data protection
privacy-preserving machine learning |
0.8 | 1 | 2024 | Certifiably Byzantine-Robust Federated Conformal Prediction · ICML 2024 |
Machine learning › Trustworthy machine learning
calibration |
0.7 | 1 | 2023 | Taking a Step Back with KCal: Multi-Class Kernel-Based Calibration for Deep Neural Networks · ICLR 2023 |
Machine learning › Learning paradigms
multi-label classification |
0.7 | 1 | 2023 | Fast Online Value-Maximizing Prediction Sets with Conformal Cost Control · ICML 2023 |
Bioinformatics and computational biology
drug discovery |
0.7 | 1 | 2023 | CoDrug: Conformal Drug Property Prediction with Density Estimation under Covariate Shift · NeurIPS 2023 |
Medical and health informatics › electronic health records
electronic health record analysis |
0.7 | 1 | 2023 | PyHealth: A Deep Learning Toolkit for Healthcare Applications · KDD 2023 |
Bioinformatics and computational biology
molecular property prediction |
0.7 | 1 | 2023 | CoDrug: Conformal Drug Property Prediction with Density Estimation under Covariate Shift · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification |
0.6 | 1 | 2022 | SCRIB: Set-Classifier with Class-Specific Risk Bounds for Blackbox Models · AAAI 2022 |
Machine learning › Trustworthy machine learning › uncertainty estimation
prediction intervals |
0.5 | 1 | 2021 | Locally Valid and Discriminative Prediction Intervals for Deep Learning Models · NeurIPS 2021 |
Machine learning › Deep learning architectures and training
equivariant neural network |
0.3 | 1 | 2018 | Clebsch-Gordan Nets: a Fully Fourier Space Spherical Convolutional Neural Network · NeurIPS 2018 |
Machine learning › Deep learning architectures and training › equivariant neural network
group equivariant neural network |
0.3 | 1 | 2018 | Clebsch-Gordan Nets: a Fully Fourier Space Spherical Convolutional Neural Network · NeurIPS 2018 |
Machine learning › Deep learning architectures and training › equivariant neural network
spherical CNN |
0.3 | 1 | 2018 | Clebsch-Gordan Nets: a Fully Fourier Space Spherical Convolutional Neural Network · NeurIPS 2018 |
Machine learning › Trustworthy machine learning
interpretability |
0.3 | 1 | 2025 | Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey · KDD (2) 2025 |
Machine learning › Trustworthy machine learning › calibration
neural network calibration |
0.2 | 1 | 2023 | Taking a Step Back with KCal: Multi-Class Kernel-Based Calibration for Deep Neural Networks · ICLR 2023 |
Machine learning › Trustworthy machine learning › fairness
fairness and robustness |
0.2 | 1 | 2022 | SCRIB: Set-Classifier with Class-Specific Risk Bounds for Blackbox Models · AAAI 2022 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
regression |
0.2 | 1 | 2022 | Conformal Prediction with Temporal Quantile Adjustments · NeurIPS 2022 |
Computer vision › 3D vision › invariant feature extraction
rotation invariance |
0.1 | 1 | 2018 | Clebsch-Gordan Nets: a Fully Fourier Space Spherical Convolutional Neural Network · NeurIPS 2018 |
Methods — techniques the papers use, named apart from their topics
conformal prediction · 5.1malicious client estimation · 1.5byzantine failure modeling · 1.5deep learning · 1.3taxonomy · 0.9survey · 0.9sequence likelihood · 0.8attention-based weighting · 0.8online update mechanism · 0.7kernel-based calibration · 0.7kernel density estimation · 0.7energy-based model · 0.7covariate shift correction · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Uncertainty Quantification and Confidence Calibration in Large Language Models: A SurveyabstractUncertainty quantification (UQ) enhances the reliability of Large Language Models (LLMs) by estimating confidence in outputs, enabling risk mitigation and selective prediction. However, traditional UQ methods struggle with LLMs due to computational constraints and decoding inconsistencies. Moreover, LLMs introduce unique uncertainty sources, such as input ambiguity, reasoning path divergence, and decoding stochasticity, that extend beyond classical aleatoric and epistemic uncertainty. To address this, we introduce a new taxonomy that categorizes UQ methods based on computational efficiency and uncertainty dimensions, including input, reasoning, parameter, and prediction uncertainty. We evaluate existing techniques, summarize existing benchmarks and metrics for UQ, assess their real-world applicability, and identify open challenges, emphasizing the need for scalable, interpretable, and robust UQ approaches to enhance LLM reliability. Xiaoou Liu, Tiejin Chen, Longchao Da, Chacha Chen, Zhen Lin 0001, Hua Wei 0001 |
KDD (2) | 5 |
| 2024 | Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language GenerationabstractThe advent of large language models (LLMs) has dramatically advanced the state-of-the-art in numerous natural language generation tasks.For LLMs to be applied reliably, it is essential to have an accurate measure of their confidence.Currently, the most commonly used confidence score function is the likelihood of the generated sequence, which, however, conflates semantic and syntactic components.For instance, in question-answering (QA) tasks, an awkward phrasing of the correct answer might result in a lower probability prediction.Additionally, different tokens should be weighted differently depending on the context.In this work, we propose enhancing the predicted sequence probability by assigning different weights to various tokens using attention values elicited from the base LLM.By employing a validation set, we can identify the relevant attention heads, thereby significantly improving the reliability of the vanilla sequence probability confidence measure.We refer to this new score as the Contextualized Sequence Likelihood (CSL).CSL is easy to implement, fast to compute, and offers considerable potential for further improvement with task-specific prompts.Across several QA datasets and a diverse array of LLMs, CSL has demonstrated significantly higher reliability than state-of-the-art baselines in predicting generation quality, as measured by the AUROC or AUARC. Zhen Lin 0001, Shubhendu Trivedi, Jimeng Sun 0001 |
EMNLP | 1 |
| 2024 | Certifiably Byzantine-Robust Federated Conformal PredictionabstractConformal prediction has shown impressive capacity in constructing statistically rigorous prediction sets for machine learning models with exchangeable data samples. The siloed datasets, coupled with the escalating privacy concerns related to local data sharing, have inspired recent innovations extending conformal prediction into federated environments with distributed data samples. However, this framework for distributed uncertainty quantification is susceptible to Byzantine failures. A minor subset of malicious clients can significantly compromise the practicality of coverage guarantees. To address this vulnerability, we introduce a novel framework Rob-FCP, which executes robust federated conformal prediction, effectively countering malicious clients capable of reporting arbitrary statistics with the conformal calibration process. We theoretically provide the conformal coverage bound of Rob-FCP in the Byzantine setting and show that the coverage of Rob-FCP is asymptotically close to the desired coverage level. We also propose a malicious client number estimator to tackle a more challenging setting where the number of malicious clients is unknown to the defender and theoretically shows its effectiveness. We empirically demonstrate the robustness of Rob-FCP against diverse proportions of malicious clients under a variety of Byzantine attacks on five standard benchmark and real-world healthcare datasets. Mintong Kang, Zhen Lin 0001, Jimeng Sun 0001, Cao Xiao, Bo Li 0026 |
ICML | 2 |
| 2023 | Taking a Step Back with KCal: Multi-Class Kernel-Based Calibration for Deep Neural Networks
Zhen Lin 0001, Shubhendu Trivedi, Jimeng Sun 0001 |
ICLR | 1 |
| 2023 | Fast Online Value-Maximizing Prediction Sets with Conformal Cost ControlabstractMany real-world multi-label prediction problems involve set-valued predictions that must satisfy specific requirements dictated by downstream usage. We focus on a typical scenario where such requirements, separately encoding *value* and *cost*, compete with each other. For instance, a hospital might expect a smart diagnosis system to capture as many severe, often co-morbid, diseases as possible (the value), while maintaining strict control over incorrect predictions (the cost). We present a general pipeline, dubbed as FavMac, to maximize the value while controlling the cost in such scenarios. FavMac can be combined with almost any multi-label classifier, affording distribution-free theoretical guarantees on cost control. Moreover, unlike prior works, FavMac can handle real-world large-scale applications via a carefully designed online update mechanism, which is of independent interest. Our methodological and theoretical contributions are supported by experiments on several healthcare tasks and synthetic datasets - FavMac furnishes higher value compared with several variants and baselines while maintaining strict cost control. Zhen Lin 0001, Shubhendu Trivedi, Cao Xiao, Jimeng Sun 0001 |
ICML | 1 |
| 2023 | PyHealth: A Deep Learning Toolkit for Healthcare ApplicationsabstractDeep learning (DL) has emerged as a promising tool in healthcare applications. However, the reproducibility of many studies in this field is limited by the lack of accessible code implementations and standard benchmarks. To address the issue, we create PyHealth, a comprehensive library to build, deploy, and validate DL pipelines for healthcare applications. PyHealth supports various data modalities, including electronic health records (EHRs), physiological signals, medical images, and clinical text. It offers various advanced DL models and maintains comprehensive medical knowledge systems. The library is designed to support both DL researchers and clinical data scientists. Upon the time of writing, PyHealth has received 633 stars, 130 forks, and 15k+ downloads in total on GitHub. Chaoqi Yang, Zhenbang Wu, Patrick Jiang, Zhen Lin 0001, Benjamin P. Danek, Jimeng Sun 0001 |
KDD | 4 |
| 2023 | CoDrug: Conformal Drug Property Prediction with Density Estimation under Covariate ShiftabstractIn drug discovery, it is vital to confirm the predictions of pharmaceutical properties from computational models using costly wet-lab experiments. Hence, obtaining reliable uncertainty estimates is crucial for prioritizing drug molecules for subsequent experimental validation. Conformal Prediction (CP) is a promising tool for creating such prediction sets for molecular properties with a coverage guarantee. However, the exchangeability assumption of CP is often challenged with covariate shift in drug discovery tasks: Most datasets contain limited labeled data, which may not be representative of the vast chemical space from which molecules are drawn. To address this limitation, we propose a method called CoDrug that employs an energy-based model leveraging both training data and unlabelled data, and Kernel Density Estimation (KDE) to assess the densities of a molecule set. The estimated densities are then used to weigh the molecule samples while building prediction sets and rectifying for distribution shift. In extensive experiments involving realistic distribution drifts in various small-molecule drug discovery tasks, we demonstrate the ability of CoDrug to provide valid prediction sets and its utility in addressing the distribution shift arising from de novo drug design models. On average, using CoDrug can reduce the coverage gap by over 35% when compared to conformal prediction sets not adjusted for covariate shift. Siddhartha Laghuvarapu, Zhen Lin 0001, Jimeng Sun 0001 |
NeurIPS | 2 |
| 2022 | SCRIB: Set-Classifier with Class-Specific Risk Bounds for Blackbox ModelsabstractDespite deep learning (DL) success in classification problems, DL classifiers do not provide a sound mechanism to decide when to refrain from predicting. Recent works tried to control the overall prediction risk with classification with rejection options. However, existing works overlook the different significance of different classes. We introduce Set-classifier with class-specific RIsk Bounds (SCRIB) to tackle this problem, assigning multiple labels to each example. Given the output of a black-box model on the validation set, SCRIB constructs a set-classifier that controls the class-specific prediction risks. The key idea is to reject when the set classifier returns more than one label. We validated SCRIB on several medical applications, including sleep staging on electroencephalogram(EEG) data, X-ray COVID image classification, and atrial fibrillation detection based on electrocardiogram (ECG) data.SCRIB obtained desirable class-specific risks, which are 35%-88% closer to the target risks than baseline methods. Zhen Lin 0001, Lucas Glass, M. Brandon Westover, Cao Xiao, Jimeng Sun 0001 |
AAAI | 1 |
| 2022 | Conformal Prediction with Temporal Quantile AdjustmentsabstractWe develop Temporal Quantile Adjustment (TQA), a general method to construct efficient and valid prediction intervals (PIs) for regression on cross-sectional time series data. Such data is common in many domains, including econometrics and healthcare. A canonical example in healthcare is predicting patient outcomes using physiological time-series data, where a population of patients composes a cross-section. Reliable PI estimators in this setting must address two distinct notions of coverage: cross-sectional coverage across a cross-sectional slice, and longitudinal coverage along the temporal dimension for each time series. Recent works have explored adapting Conformal Prediction (CP) to obtain PIs in the time series context. However, none handles both notions of coverage simultaneously. CP methods typically query a pre-specified quantile from the distribution of nonconformity scores on a calibration set. TQA adjusts the quantile to query in CP at each time $t$, accounting for both cross-sectional and longitudinal coverage in a theoretically-grounded manner. The post-hoc nature of TQA facilitates its use as a general wrapper around any time series regression model. We validate TQA's performance through extensive experimentation: TQA generally obtains efficient PIs and improves longitudinal coverage while preserving cross-sectional coverage. Zhen Lin 0001, Shubhendu Trivedi, Jimeng Sun 0001 |
NeurIPS | 1 |
| 2021 | Locally Valid and Discriminative Prediction Intervals for Deep Learning ModelsabstractCrucial for building trust in deep learning models for critical real-world applications is efficient and theoretically sound uncertainty quantification, a task that continues to be challenging. Useful uncertainty information is expected to have two key properties: It should be valid (guaranteeing coverage) and discriminative (more uncertain when the expected risk is high). Moreover, when combined with deep learning (DL) methods, it should be scalable and affect the DL model performance minimally. Most existing Bayesian methods lack frequentist coverage guarantees and usually affect model performance. The few available frequentist methods are rarely discriminative and/or violate coverage guarantees due to unrealistic assumptions. Moreover, many methods are expensive or require substantial modifications to the base neural network. Building upon recent advances in conformal prediction [13, 33] and leveraging the classical idea of kernel regression, we propose Locally Valid and Discriminative prediction intervals (LVD), a simple, efficient, and lightweight method to construct discriminative prediction intervals (PIs) for almost any DL model. With no assumptions on the data distribution, such PIs also offer finite-sample local coverage guarantees (contrasted to the simpler marginal coverage). We empirically verify, using diverse datasets, that besides being the only locally valid method for DL, LVD also exceeds or matches the performance (including coverage rate and prediction accuracy) of existing uncertainty quantification methods, while offering additional benefits in scalability and flexibility. Zhen Lin 0001, Shubhendu Trivedi, Jimeng Sun 0001 |
NeurIPS | 1 |
| 2018 | Clebsch-Gordan Nets: a Fully Fourier Space Spherical Convolutional Neural NetworkabstractRecent work by Cohen et al. has achieved state-of-the-art results for learning spherical images in a rotation invariant way by using ideas from group representation theory and noncommutative harmonic analysis. In this paper we propose a generalization of this work that generally exhibits improved performace, but from an implementation point of view is actually simpler. An unusual feature of the proposed architecture is that it uses the Clebsch--Gordan transform as its only source of nonlinearity, thus avoiding repeated forward and backward Fourier transforms. The underlying ideas of the paper generalize to constructing neural networks that are invariant to the action of other compact groups. Risi Kondor, Zhen Lin 0001, Shubhendu Trivedi |
NeurIPS | 2 |