EDBT 2026 Demo / reviewers in the wild / expert
Yangyi Li
dblp:326/2983
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language ModelsabstractThe rapid advancements in artificial intelligence (AI) have primarily focused on the process of learning from data to acquire knowledgeable learning systems. As these systems are increasingly deployed in critical areas, ensuring their privacy and alignment with human values is paramount. Recently, selective forgetting (also known as machine unlearning) has shown promise for privacy and data removal tasks, and has emerged as a transformative paradigm shift in the field of AI. It refers to the ability of a model to selectively erase the influence of previously seen data, which is especially important for compliance with modern data protection regulations and for aligning models with human values. Despite its promise, selective forgetting raises significant privacy concerns, especially when the data involved come from sensitive domains. While new unlearning-induced privacy attacks are continuously proposed, each is shown to outperform its predecessors using different experimental settings, which can lead to overly optimistic and potentially unfair assessments that may disproportionately favor one particular attack over the others. In this work, we present the first comprehensive benchmark for evaluating privacy vulnerabilities in selective forgetting. We extensively investigate privacy vulnerabilities of machine unlearning techniques and benchmark privacy leakage across a wide range of victim data, state-of-the-art unlearning privacy attacks, unlearning methods, and model architectures. We systematically evaluate and identify critical factors related to unlearning-induced privacy leakage. With our novel insights, we aim to provide a standardized tool for practitioners seeking to deploy customized unlearning applications with faithful privacy assessments. Yangyi Li, Mengdi Huai |
AAAI | 3 |
| 2026 | Quantifying and Understanding Uncertainty in Large Reasoning ModelsabstractLarge Reasoning Models (LRMs) have recently demonstrated significant improvements in complex reasoning.While quantifying generation uncertainty in LRMs is crucial, traditional methods are often insufficient because they do not provide finite-sample guarantees for reasoning-answer generation.Conformal prediction (CP) stands out as a distributionfree and model-agnostic methodology that constructs statistically rigorous uncertainty sets.However, existing CP methods ignore the logical connection between the reasoning trace and the final answer.Additionally, prior studies fail to interpret the origins of uncertainty coverage for LRMs as they typically overlook the specific training factors driving valid reasoning.Notably, it is challenging to disentangle reasoning quality from answer correctness when quantifying uncertainty, while simultaneously establishing theoretical guarantees for computationally efficient explanation methods.To address these challenges, we first propose a novel methodology that quantifies uncertainty in the reasoning-answer structure with statistical guarantees.Subsequently, we develop a unified example-to-step explanation framework using Shapley values that identifies a provably sufficient subset of training examples and their key reasoning steps to preserve the guarantees.We also provide theoretical analyses of our proposed methods.Extensive experiments on challenging reasoning datasets verify the effectiveness of the proposed methods. Yangyi Li, Mengdi Huai |
ACL (1) | 1 |
| 2026 | Exploiting Variable-Dimensional LDPC Coding to Improve NAND Flash Memory System PerformanceabstractSolid state drives (SSDs) based on NAND flash technology are steadily gaining popularity and mass market adoption due to their increased storage capacity and density. However, because of the more bits in each cell and the reduced cell spacing, they are experiencing a decline in reliability. The most efficient way to ensure reliability of data is to use low-density parity-check (LDPC) codes. Nevertheless, using a hybrid decoding technique for LDPC codes results in a significant decoding latency, which exacerbates performance issues. In this paper, we propose a variable-dimensional LDPC coding scheme, called VDLDPC, to reduce the high decoding latency and thus improve read performance of NAND flash memory on hot read data. One of the crucial designs in the VDLDPC scheme is the two-dimensional LDPC (TD-LDPC) algorithm. TD-LDPC implements row and column encoding separately when writing data to the flash memory by using sub-LDPC codes. Errors in the data arise after a period of retention. When the data is read out, TD-LDPC performs row and column decoding using sub-LDPC codes, and the column decoding result can be re-decoded as a new round of row decoding input. Simulation results show that the proposed VDLDPC scheme has the advantage in decoding latency and reduces the flash memory read response time by up to 12.0% (5.8% on average across all workloads) compared to the current LDPC code scheme. The proposed VDLDPC scheme ensures reliability while improving NAND flash system read performance on hot read data. Meng Zhang 0014, Wei Li 0312, Yangyi Li, Tianwei Gui, Changsheng Xie 0001, Fei Wu 0005 |
DATE | 3 |
| 2026 | Uncertainty-Aware Language Guidance for Concept Bottleneck Models
Yangyi Li, Mengdi Huai |
PAKDD (4) | 1 |
| 2026 | SiDTBF: Merging Soft Information With Dynamic Threshold Bit Flipping LDPC Decoding for 3-D NAND flash memoryabstractThrough stacking and multi-bit technology, three-dimensional (3D) flash memory enhances storage capacity and density; nevertheless, the reduction in noise margin results in an increase in raw bit error rate (RBER) and a decrease in data reliability. Low-density parity-check (LDPC) codes are widely used in flash memory for improving data reliability because of its strong error correction capability. In the early stages of 3D flash memory use, the RBER is low, and hard decision decoding (e.g., bit flipping decoding) is generally invoked for error correction. Existing LDPC codes with dynamic threshold bit flipping (DTBF) decoding algorithms cannot correct bit errors when the gradually increasing RBER exceeds its error correction threshold, resulting in an increase in decoding latency. To enhance error correction capability and reduce decoding latency, this paper proposes SiDTBF: merging soft information with DTBF LDPC decoding for 3D NAND flash memory. First, the read reference voltage of various interval lengths is applied in accordance with the threshold voltage distribution drift characteristics of the 3D flash memory cell to get the decoding soft information of each bit. Second, the strong and weak bits are distinguished using the soft information. In contrast to weak bits, which are more likely to be erroneous, strong bits are more likely to be correct. Finally, all the strong and weak bits are input as initial values for bit-flip iterative decoding. Using the column weight of the parity-check matrix, the threshold for the number of flipped weak bits is determined in the first decoding iteration process. The portion of the weak bits that exceeds the threshold is flipped. In the ensuing iteration phase, the DTBF decoding algorithm is used. SiDTBF improves decoding error correction performance by fusing each bit’s soft information with the DTBF algorithm during the decoding phase. Simulation results show that compared with current DTBF, SiDTBF significantly improves bit flipping decoding error correction capability and reduces decoding latency. Yangyi Li, Meng Zhang 0014, Wei Li 0312, Tianwei Gui, Changsheng Xie 0001, Fei Wu 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2026 | DyLDPC: A Dynamic LDPC Code with Variable Correction Capability to Improve Decoding Performance for 3D NAND Flash Memoryabstract3D NAND flash memory is currently the mainstream storage medium due to high density and large capacity. However, the high raw bit error rate (RBER) poses challenges to data reliability. Low-density parity-check (LDPC) codes, known for strong error correction capabilities, are widely used to ensure data reliability. Traditional error correction schemes employ a single, fixed LDPC code, resulting in suboptimal performance–excess decoding overhead at low RBER and insufficient correction capability at high RBER. To address these limitations, we propose DyLDPC: a dynamic LDPC code with variable correction capability to improve decoding performance for 3D NAND flash memory. DyLDPC dynamically adjusts the error correction capability of LDPC codes based on the temporal and spatial variations of RBER in 3D NAND flash memory. Temporally, RBER increases with retention time and program/erase (P/E) cycles. Spatially, RBER varies across layers and pages. DyLDPC predicts RBER under varying conditions and allocates appropriate LDPC codes accordingly, effectively reducing ECC storage overhead, extending flash memory lifespan, and improving decoding efficiency. While ensuring data reliability, it optimizes error correction performance. Evaluations indicate that DyLDPC reduces decoding iterations by 8.1% and latency by 74% on average compared to static schemes. Additionally, using differentiated LDPC codes for most significant bit (MSB) and least significant bit (LSB) pages in multi-level cell (MLC) NAND further reduces LSB decoding latency by 27.8%. Geyang Ren, Meng Zhang 0014, Yangyi Li, Ruifeng Tu, Shaoqi Gao, Lingyan Fan, Changsheng Xie 0001, Fei Wu 0005 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | Towards Unveiling Predictive Uncertainty Vulnerabilities in the Context of the Right to Be ForgottenabstractCurrently, various uncertainty quantification methods have been proposed to provide certainty and probability estimates for deep learning models' label predictions. Meanwhile, with the growing demand for the right to be forgotten, machine unlearning has been extensively studied as a means to remove the impact of requested sensitive data from a pre-trained model without retraining the model from scratch. However, the vulnerabilities of such generated predictive uncertainties with regard to dedicated malicious unlearning attacks remain unexplored. To bridge this gap, for the first time, we propose a new class of malicious unlearning attacks against predictive uncertainties, where the adversary aims to cause the desired manipulations of specific predictive uncertainty results. We also design novel optimization frameworks for our attacks and conduct extensive experiments, including black-box scenarios. Notably, our extensive experiments show that our attacks are more effective in manipulating predictive uncertainties than traditional attacks that focus on label misclassifications, and existing defenses against conventional attacks are ineffective against our attacks. Yangyi Li, Wenqian Ye, Mengdi Huai |
CIKM | 3 |
| 2024 | Towards Modeling Uncertainties of Self-Explaining Neural Networks via Conformal PredictionabstractDespite the recent progress in deep neural networks (DNNs), it remains challenging to explain the predictions made by DNNs. Existing explanation methods for DNNs mainly focus on post-hoc explanations where another explanatory model is employed to provide explanations. The fact that post-hoc methods can fail to reveal the actual original reasoning process of DNNs raises the need to build DNNs with built-in interpretability. Motivated by this, many self-explaining neural networks have been proposed to generate not only accurate predictions but also clear and intuitive insights into why a particular decision was made. However, existing self-explaining networks are limited in providing distribution-free uncertainty quantification for the two simultaneously generated prediction outcomes (i.e., a sample's final prediction and its corresponding explanations for interpreting that prediction). Importantly, they also fail to establish a connection between the confidence values assigned to the generated explanations in the interpretation layer and those allocated to the final predictions in the ultimate prediction layer. To tackle the aforementioned challenges, in this paper, we design a novel uncertainty modeling framework for self-explaining networks, which not only demonstrates strong distribution-free uncertainty modeling performance for the generated explanations in the interpretation layer but also excels in producing efficient and effective prediction sets for the final predictions based on the informative high-level basis explanations. We perform the theoretical analysis for the proposed framework. Extensive experimental evaluation demonstrates the effectiveness of the proposed uncertainty framework. Yangyi Li, Fenglong Ma, Chao Zhang 0014, Mengdi Huai |
AAAI | 3 |
| 2024 | Data Poisoning Attacks against Conformal PredictionabstractThe efficient and theoretically sound uncertainty quantification is crucial for building trust in deep learning models. This has spurred a growing interest in conformal prediction (CP), a powerful technique that provides a model-agnostic and distribution-free method for obtaining conformal prediction sets with theoretical guarantees. However, the vulnerabilities of such CP methods with regard to dedicated data poisoning attacks have not been studied previously. To bridge this gap, for the first time, we in this paper propose a new class of black-box data poisoning attacks against CP, where the adversary aims to cause the desired manipulations of some specific examples’ prediction uncertainty results (instead of misclassifications). Additionally, we design novel optimization frameworks for our proposed attacks. Further, we conduct extensive experiments to validate the effectiveness of our attacks on various settings (e.g., the full and split CP settings). Notably, our extensive experiments show that our attacks are more effective in manipulating uncertainty results than traditional poisoning attacks that aim at inducing misclassifications, and existing defenses against conventional attacks are ineffective against our proposed attacks. Yangyi Li, Aobo Chen 0002, Divya Lidder, Mengdi Huai |
ICML | 1 |
| 2024 | Rethinking Adversarial Robustness in the Context of the Right to be ForgottenabstractThe past few years have seen an intense research interest in the practical needs of the "right to be forgotten", which has motivated researchers to develop machine unlearning methods to unlearn a fraction of training data and its lineage. While existing machine unlearning methods prioritize the protection of individuals’ private data, they overlook investigating the unlearned models’ susceptibility to adversarial attacks and security breaches. In this work, we uncover a novel security vulnerability of machine unlearning based on the insight that adversarial vulnerabilities can be bolstered, especially for adversarially robust models. To exploit this observed vulnerability, we propose a novel attack called Adversarial Unlearning Attack (AdvUA), which aims to generate a small fraction of malicious unlearning requests during the unlearning process. AdvUA causes a significant reduction of adversarial robustness in the unlearned model compared to the original model, providing an entirely new capability for adversaries that is infeasible in conventional machine learning pipelines. Notably, we also show that AdvUA can effectively enhance model stealing attacks by extracting additional decision boundary information, further emphasizing the breadth and significance of our research. We also conduct both theoretical analysis and computational complexity of AdvUA. Extensive numerical studies are performed to demonstrate the effectiveness and efficiency of the proposed attack. Yangyi Li, Aobo Chen 0002, Mengdi Huai |
ICML | 3 |
| 2024 | Modeling and Understanding Uncertainty in Medical Image Classification
Aobo Chen 0002, Yangyi Li, Kathy Morse, Chenglin Miao, Mengdi Huai |
MICCAI (10) | 2 |
| 2024 | A regionally coordinated allocation strategy for medical resources based on multidimensional uncertain information
Xinxin Wang 0001, Yangyi Li, Zeshui Xu |
Inf. Sci. | 2 |
| 2022 | Nested information representation of multi-dimensional decision: An improved PROMETHEE method based on NPLTSs
Xinxin Wang 0001, Yangyi Li, Zeshui Xu, Yuyan Luo |
Inf. Sci. | 2 |