VLDB 2026 Research / reviewers in the wild / expert
Yiwei Fu
dblp:176/8274
· DBLP profile ↗
18ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Edge Self-Adversarial Augmentation Enhances Graph Contrastive Learning Against Neighborhood InconsistencyabstractRecent studies have shown that unsupervised graph contrastive learning (GCL) is vulnerable to adversarial attacks. Automatic adversarial augmentation techniques are proposed to improve both the effectiveness and robustness of GCL. Existing methods typically regard unsupervised contrastive loss as the adversarial goal, essentially aiming to maximize inter-view instance-wise discrepancies between adversarial and original views. However, such attacks overlook intra-view neighborhood inconsistency, which hinders the robustness of GCL models against local neighborhood noises, resulting in performance degradation on low-homophily graphs. To tackle this issue, we propose a novel adversarial contrastive paradigm, named Edge self-aDversarial Augmentation for Graph Contrastive Learning (EDA-GCL). We theoretically establish that the adversarial objective of the intra-view neighborhood is equivalent to maximizing the discrepancy between bidirectional edge features. Hence, we build our adversarial framework based on edge self-adversarial learning. It generates pairwise adversarial augmentations from the original view by learning distinct neighborhood connectivity structures. The learned pairwise adversarial views are utilized for GCL model training in the minimization stage. Notably, this edge-level adversarial approach reduces the computational complexity to the level of the edge number. Experiments on various graph tasks and complex noise scenarios demonstrate the superiority and robustness of our EDA-GCL. Chunchun Chen, Chenrun Wang, Yiwei Fu, Xin Sun 0003, Rui Fan 0001, Wei Ye 0001 |
AAAI | 5 |
| 2026 | Calibrating Inference Time Alignment with Sequence-level Risk AccumulationabstractShanwen Tan, Ziyang Dong, Wei Ju, Yiwei Fu, Hao Wu, Kun Wang, Yifan Wang, Ziyue Qiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shanwen Tan, Ziyang Dong, Wei Ju 0001, Yiwei Fu, Kun Wang 0056, Yifan Wang 0014, Ziyue Qiao |
ACL (1) | 4 |
| 2026 | When Context Bites: Detecting RAG Poisoning via Document-Level Attention CollapseabstractRetrieval-augmented generation (RAG) is indispensable for enhancing large language models. However, RAGs are increasingly susceptible to poisoning attacks, in which adversarial documents are injected to manipulate generator outputs. Previous methods rely on output-side signals such as perplexity and consistency checks to detect such attacks. Nevertheless, our analysis reveals that deliberate attacks often induce false confidence, where poisoned outputs exhibit even lower perplexity than benign ones, rendering uncertainty-based detection ineffective. To address this challenge, we explore the internal dynamics of the generator and identify a distinctive signature termed Attention Collapse. Unlike the dispersed attention in benign generations, attacked generations exhibit a decrease in entropy as attention concentrates on poisoned documents. Building on these findings, we propose D-SCAN (Document-level Signal Collapse Analysis), a lightweight detection framework that monitors attention dynamics to identify attacked generations. Extensive experiments on multiple attack benchmarks demonstrate the effectiveness of our method. Moreover, D-SCAN can detect attacks even when they fail to alter the final answer. Code is available at https://github.com/yingtaoren/D-Scan.git. Yingtao Ren, Yiwei Fu, Xiao Luo 0001, Chin-Teng Lin |
SIGIR | 3 |
| 2026 | Long-Tailed Recognition of Evidential Experts for Graph-level ClassificationabstractGraph-level classification involves analyzing the property of the whole graph, which is typically solved by using graph neural networks (GNNs). Existing efforts generally assume a balanced class distribution. However, real-world data often exhibit long-tailed distributions, i.e., tail classes have significantly fewer samples than head classes, and thus directly applying GNNs is eventually biased toward the head classes, resulting in limited generalization over the tail classes. Moreover, the predictions of existing algorithms are usually not trustworthy, and the trained classifiers remain ignorant to their predictive confidence. Towards this end, in this paper we develop a principled framework called GraphEVER for long-tailed graph-level classification. Technically, GraphEVER incorporates the beliefs of multiple experts and leverages the idea of subjective logic within the Dempster-Shafer Evidence Theory (DST). It can provide the evidence and uncertainty estimation for each expert, where the evidence is parameterized by a Dirichlet distribution to model class probability distribution, and the uncertainty is quantified via a well-defined theoretical framework. In this way, diverse experts can be integrated under DST to endow the classifier with both reliability and robustness. Moreover, we propose an evidence-based routing mechanism to dynamically assign experts, such that the tail classes can receive more attention, while the head classes can reduce redundant engaged experts, further cutting down the computational cost and improving the efficiency. Extensive experiments on seven datasets verify the superiority of our proposed framework. Wei Ju 0001, Siyu Yi, Zhengyang Mao, Yifang Qin, Yifan Wang 0014, Zhiping Xiao 0001, Yiwei Fu, Ziyue Qiao, Ming Zhang 0004 |
WWW | 7 |
| 2026 | Out-of-distribution generalization enhances protein function annotation for low-homology sequencesabstractUnderstanding protein functions in biological processes is pivotal for disease elucidation and drug discovery. Despite notable progress, existing approaches primarily focus on function transfer under in-distribution (ID) settings, where training and test proteins exhibit high sequence similarity. As a result, their performance often degrades when applied to novel, diverse, and low-homology protein sequences, posing a major challenge for out-of-distribution (OOD) generalization encountered in practice. Towards this end, we develop ProteinScore, a graph transformer approach tailored to improve protein function prediction in OOD settings. ProteinScore integrates a label-invariant variational subgraph generator with self-supervised contrastive learning, thereby identifying meaning substructures within proteins. By highlighting informative features while filtering out redundant ones, ProteinScore improves generalization to diverse and low-homology sequences. Experiments on datasets with both experimentally resolved and AlphaFold2-predicted structures demonstrate that ProteinScore consistently outperforms strong baselines and provides biologically meaningful interpretability through accurately identifying binding sites. In addition, ProteinScore generalizes effectively to two additional downstream tasks, drug-target interaction classification and subcellular localization prediction, achieving superior predictive performance and reliable interpretability. Yiwei Fu, Jiaxiao Chen, Haoyu Lin, Zhonghui Gu, Qingqing Long, Xiao Luo 0001, Minghua Deng |
Briefings Bioinform. | 1 |
| 2026 | HGOOD-D: Hyperbolic Hierarchical Exploration for Graph Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection has garnered increasing concern for identifying test samples that exhibit a distributional shift from the training dataset in practical deep learning applications. With the significant advancements in graph deep learning for graph representation, graph OOD detection has emerged as a research problem. Graph contrastive learning (GCL) is applied to graph OOD detection due to its capacity for learning discriminative representations in a self-supervised manner, thereby eliminating the need for time-consuming and labor-intensive label information. However, existing methods often neglect the explicit consideration of underlying semantics behind graph data distribution for OOD detection. We argue that simple data augmentations for GCL may risk disrupting the intrinsic graph structure while retaining redundant structural information, which hinders semantic discrimination between graphs. Additionally, Euclidean space embedding struggles to maintain hierarchical structural consistency, making it challenging to meaningfully capture the hierarchical semantic distribution of graph data. In response to these issues, we propose a novel framework termed HGOOD-D, which aims to explore latent semantic hierarchies in hyperbolic space for graph OOD detection. Specifically, we design a bottleneck graph extractor grounded in the information bottleneck (IB) principle, which captures the minimal sufficient information to distinguish graph patterns. Based on this, we introduce hierarchical contrastive learning to capture the hierarchical semantics within graph data distribution. These methods are based on hyperbolic space embedding that can preserve complex inter-relationships in graph hierarchies, thereby mitigating data distortion. Comprehensive evaluations on ten widely used benchmark datasets show that HGOOD-D consistently surpasses current state-of-the-art approaches in graph OOD detection. Yuntai Ding, Tao Ren 0002, Yiwei Fu, Yifan Wang 0014, Chong Chen 0002, Wei Ju 0001, Xiao Luo 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign RelearningabstractMachine unlearning is a promising approach to mitigate undesirable memorization of training data in ML models. However, in this work we show that existing approaches for unlearning in LLMs are surprisingly susceptible to a simple set of benign relearning attacks. With access to only a small and potentially loosely related set of data, we find that we can “jog” the memory of unlearned models to reverse the effects of unlearning. For example, we show that relearning on public medical articles can lead an unlearned LLM to output harmful knowledge about bioweapons, and relearning general wiki information about the book series Harry Potter can force the model to output verbatim memorized text. We formalize this unlearning-relearning pipeline, explore the attack across three popular unlearning benchmarks, and discuss future directions and guidelines that result from our study. Our work indicates that current approximate unlearning methods simply suppress the model outputs and fail to robustly forget target knowledge in the LLMs. Shengyuan Hu 0001, Yiwei Fu, Steven Z. Wu, Virginia Smith |
ICLR | 2 |
| 2025 | Sparse Causal Discovery with Generative Intervention for Unsupervised Graph Domain AdaptationabstractUnsupervised Graph Domain Adaptation (UGDA) leverages labeled source domain graphs to achieve effective performance in unlabeled target domains despite distribution shifts. However, existing methods often yield suboptimal results due to the entanglement of causal-spurious features and the failure of global alignment strategies. We propose SLOGAN (Sparse Causal Discovery with Generative Intervention), a novel approach that achieves stable graph representation transfer through sparse causal modeling and dynamic intervention mechanisms. Specifically, SLOGAN first constructs a sparse causal graph structure, leveraging mutual information bottleneck constraints to disentangle sparse, stable causal features while compressing domain-dependent spurious correlations through variational inference. To address residual spurious correlations, we innovatively design a generative intervention mechanism that breaks local spurious couplings through cross-domain feature recombination while maintaining causal feature semantic consistency via covariance constraints. Furthermore, to mitigate error accumulation in target domain pseudo-labels, we introduce a category-adaptive dynamic calibration strategy, ensuring stable discriminative learning. Extensive experiments on multiple real-world datasets demonstrate that SLOGAN significantly outperforms existing baselines. Junyu Luo 0002, Yuhao Tang, Yiwei Fu, Xiao Luo 0001, Zhizhuo Kou, Zhiping Xiao 0001, Wei Ju 0001, Wentao Zhang 0001, Ming Zhang 0004 |
ICML | 3 |
| 2025 | A Physical Attack for Segmentation-Based Visual Foundation ModelsabstractWe formalize a threat model for general patch-based attacks on visual foundation models (VFMs), highlighting the strong assumptions made by previous approaches regarding the proximity of adversarial patches to objects of interest. We demonstrate that these methods often fail to generalize to physical domains. To address these limitations, we propose a patch-based physical attack method that targets VFMs without relying on these assumptions. By operating in the feature space of VFMs and excluding the adversarial patch in the optimization objective, we ensure a more robust attack. Our method achieves the first successful physical adversarial attacks on segmentation-based VFMs, with empirical results showing that transformer-based segmentation VFMs, including SAM, SEEM, and CLIPSeg, are vulnerable to our attack in real-world scenarios. Mengqi He, Jinhong Ni, Zhaoyuan Yang, Yiwei Fu, John N. Karigiannis |
IJCNN | 4 |
| 2025 | Audio-Visual Deepfake Detection via Multi-Teacher Distillation of Content and Semantic ConsistencyabstractThe advent of deepfake technology enables the synthesis of highly realistic audio-visual content, posing severe challenges such as identity impersonation and public opinion manipulation. Existing detection methods primarily focus on unimodal analysis or shallow consistency modeling, making it difficult to effectively address cross-modal forgeries and well-synchronized fake samples. These limitations result in poor generalization and high misclassification rates. To tackle these issues, this paper proposes an innovative multi-teacher knowledge distillation detection framework based on content consistency and semantic consistency. By leveraging high-level supervision from speech recognition and semantic comprehension, the framework guides the detection model to simultaneously learn cross-modal alignment in terms of both content and semantics. Additionally, a Transformer-based architecture is integrated to enhance the audio temporal modeling capability of the detection model, improving its ability to capture long-term temporal dependencies in deepfake speech. Comprehensive evaluations conducted on the FakeAVCeleb dataset demonstrate that the proposed method outperforms existing methods in terms of accuracy and AUC, while also exhibiting fewer parameters with higher computational efficiency, highlighting strong practical applicability. Beijia Sun, Haiqing Du, Yiwei Fu, Tingya Dong, Wenzhe Lu |
VCIP | 3 |
| 2024 | One Masked Model is All You Need for Sensor Fault Detection, Isolation and AccommodationabstractAccurate and reliable sensor measurements are critical for ensuring the safety and longevity of complex engineering systems such as wind turbines. In this paper, we propose a novel framework for sensor fault detection, isolation, and accommodation (FDIA) using masked models and self-supervised learning. Our proposed approach is a general time series modeling approach that can be applied to any neural network (NN) model capable of sequence modeling, and captures the complex spatio-temporal relationships among different sensors. During training, the proposed masked approach creates a random mask, which acts like a fault, for one or more sensors, making the training and inference task unified: finding the faulty sensors and correcting them. We validate our proposed technique on both a public dataset and a real-world dataset from GE offshore wind turbines, and demonstrate its effectiveness in detecting, diagnosing and correcting sensor faults. The masked model not only simplifies the overall FDIA pipeline, but also outperforms existing approaches. Our proposed technique has the potential to significantly improve the accuracy and reliability of sensor measurements in complex engineering systems in real-time, and could be applied to other types of sensors and engineering systems in the future. We believe that our proposed framework can contribute to the development of more efficient and effective FDIA techniques for a wide range of applications. Yiwei Fu, Weizhong Yan |
IJCNN | 1 |
| 2024 | Continually adapting pre-trained language model to universal annotation of single-cell RNA-seq dataabstractMOTIVATION: Cell-type annotation of single-cell RNA-sequencing (scRNA-seq) data is a hallmark of biomedical research and clinical application. Current annotation tools usually assume the simultaneous acquisition of well-annotated data, but without the ability to expand knowledge from new data. Yet, such tools are inconsistent with the continuous emergence of scRNA-seq data, calling for a continuous cell-type annotation model. In addition, by their powerful ability of information integration and model interpretability, transformer-based pre-trained language models have led to breakthroughs in single-cell biology research. Therefore, the systematic combining of continual learning and pre-trained language models for cell-type annotation tasks is inevitable. RESULTS: We herein propose a universal cell-type annotation tool, called CANAL, that continuously fine-tunes a pre-trained language model trained on a large amount of unlabeled scRNA-seq data, as new well-labeled data emerges. CANAL essentially alleviates the dilemma of catastrophic forgetting, both in terms of model inputs and outputs. For model inputs, we introduce an experience replay schema that repeatedly reviews previous vital examples in current training stages. This is achieved through a dynamic example bank with a fixed buffer size. The example bank is class-balanced and proficient in retaining cell-type-specific information, particularly facilitating the consolidation of patterns associated with rare cell types. For model outputs, we utilize representation knowledge distillation to regularize the divergence between previous and current models, resulting in the preservation of knowledge learned from past training stages. Moreover, our universal annotation framework considers the inclusion of new cell types throughout the fine-tuning and testing stages. We can continuously expand the cell-type annotation library by absorbing new cell types from newly arrived, well-annotated training datasets, as well as automatically identify novel cells in unlabeled datasets. Comprehensive experiments with data streams under various biological scenarios demonstrate the versatility and high model interpretability of CANAL. AVAILABILITY: An implementation of CANAL is available from https://github.com/aster-ww/CANAL-torch. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Journal Name online. Musu Yuan, Yiwei Fu, Minghua Deng |
Briefings Bioinform. | 3 |
| 2022 | MAD: Self-Supervised Masked Anomaly Detection Task for Multivariate Time SeriesabstractIn this paper, we introduce Masked Anomaly Detection (MAD), a general self-supervised learning task for multivariate time series anomaly detection. With the increasing availability of sensor data from industrial systems, being able to detecting anomalies from streams of multivariate time series data is of significant importance. Given the scarcity of anomalies in real-world applications, the majority of literature has been focusing on modeling normality. The learned normal representations can empower anomaly detection as the model has learned to capture certain key underlying data regularities. A typical formulation is to learn a predictive model, i.e., use a window of time series data to predict future data values. In this paper, we propose an alternative self-supervised learning task. By randomly masking a portion of the inputs and training a model to estimate them using the remaining ones, MAD is an improvement over the traditional left-to-right next step prediction (NSP) task. Our experimental results demonstrate that MAD can achieve better anomaly detection rates over traditional NSP approaches when using exactly the same neural network (NN) base models, and can be modified to run as fast as NSP models during test time on the same hardware, thus making it an ideal upgrade for many existing NSP-based NN anomaly detection models. Yiwei Fu |
IJCNN | 1 |
| 2022 | A Dynamically Stabilized Recurrent Neural Network
Samer Saab 0002, Yiwei Fu, Asok Ray, Michael Hauser |
Neural Process. Lett. | 2 |
| 2020 | Spatiotemporal Representation Learning with GAN Trained LSTM-LSTM NetworksabstractLearning robot behaviors in unstructured environments often requires handcrafting the features for a given task. In this paper, we present and evaluate an unsupervised representation learning architecture, Layered Spatiotemporal Memory Long Short-Term Memory (LSTM-LSTM), that learns the underlying representation without knowledge of the task. The goal of this architecture is to learn the dynamics of the environment from high-dimensional raw video inputs. Using a Generative Adversarial Network (GAN) framework with the proposed network, this architecture is able to learn a spatiotemporal representation in its lower-dimensional latent space directly from raw input sequences. We show that our approach learns the spatial and temporal information simultaneously as opposed to a two-stage learning approach of alternating between training a Convolutional Neural Network (ConvNet) and a Long Short-Term Network (LSTM). Furthermore, by using LSTM-LSTM cells that shrink in size with the increase in the number of layers, the network learns a hierarchical representation with a low-dimensional representation at the top layer. We show that this architecture achieves state-of-the-art results with a substantially lower-dimensional representation than existing methods. We evaluate our approach on a video prediction task with standard benchmark datasets like Moving MNIST and KTH Action, as well as a simulated robot dataset. Yiwei Fu, Shiraj Sen, Johan Reimann, Charles Theurer |
ICRA | 1 |
| 2020 | DeepReturn: A deep neural network can learn how to detect previously-unseen ROP payloads without using any heuristicsabstractReturn-oriented programming (ROP) is a code reuse attack that chains short snippets of existing code to perform arbitrary operations on target machines. Existing detection methods against ROP exhibit unsatisfactory detection accuracy and/or have high runtime overhead. In this paper, we present DeepReturn, which innovatively combines address space layout guided disassembly and deep neural networks to detect ROP payloads. The disassembler treats application input data as code pointers and aims to find any potential gadget chains, which are then classified by a deep neural network as benign or malicious. Our experiments show that DeepReturn has high detection rate (99.3%) and a very low false positive rate (0.01%). DeepReturn successfully detects all of the 100 real-world ROP exploits that are collected in-the-wild, created manually or created by ROP exploit generation tools. DeepReturn is non-intrusive and does not incur any runtime overhead to the protected program. Zhisheng Hu, Yiwei Fu, Ping Chen 0003, Peng Liu 0005 |
J. Comput. Secur. | 4 |
| 2018 | Bayesian Nonparametric Regression Modeling of Panel Data for Sequential ClassificationabstractThis paper proposes a Bayesian nonparametric regression model of panel data for sequential pattern classification. The proposed method provides a flexible and parsimonious model that allows both time-independent spatial variables and time-dependent exogenous variables to be predictors. Not only this method improves the accuracy of parameter estimation for limited data, but also it facilitates model interpretation by identifying statistically significant predictors with hypothesis testing. Moreover, as the data length approaches infinity, posterior consistency of the model is guaranteed for general data-generating processes under regular conditions. The resulting model of panel data can also be used for sequential classification. The proposed method has been tested by numerical simulation, then validated on an econometric public data set, and subsequently validated for detection of combustion instabilities with experimental data that have been generated in a laboratory environment. Sihan Xiong, Yiwei Fu, Asok Ray |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Placement mitigation techniques for power grid electromigrationabstractIn advanced technology nodes, power grid metal wires are prone to electromigration (EM) failures due to small wire sizes and high unidirectional current densities. Power grid EM failures usually happen around weak power grid connections delivering current to high power-consuming regions. Previously, power grid EM was mostly addressed at the post-routing stage, which may be too late for a large number of EM violations in modern designs. In this paper, we propose a new set of incremental placement techniques to mitigate power grid EM, including cell move, single row placement, and single tile placement. Experimental results demonstrate the proposed placement techniques can effectively reduce EM violations with negligible wirelength and placement density impacts. Wei Ye 0008, Yibo Lin, Wuxi Li, Yiwei Fu, Yongsheng Sun, Canhui Zhan, David Z. Pan |
ISLPED | 5 |