VLDB 2026 Research / reviewers in the wild / expert
Yifan Fu
dblp:60/9922
· DBLP profile ↗
17ranked-venue papers
12as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
3 papers |
Mathematical optimization · 98% Graph algorithms and graph theory · 2% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Machine learning and data management · 67% Data mining · 33% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computational science and engineering › scientific machine learning
neural operator |
0.9 | 1 | 2025 | HEAP: Hyper Extended A-PDHG Operator for Constrained High-dim PDEs · ICML 2025 |
Computational science and engineering
scientific machine learning |
0.9 | 1 | 2025 | HEAP: Hyper Extended A-PDHG Operator for Constrained High-dim PDEs · ICML 2025 |
Mathematical optimization
continuous optimization |
0.9 | 1 | 2025 | HEAP: Hyper Extended A-PDHG Operator for Constrained High-dim PDEs · ICML 2025 |
Mathematical optimization › primal-dual method
primal-dual hybrid gradient |
0.9 | 1 | 2025 | HEAP: Hyper Extended A-PDHG Operator for Constrained High-dim PDEs · ICML 2025 |
Mathematical optimization › continuous optimization › nonlinear optimization
quadratic programming |
0.9 | 1 | 2025 | HEAP: Hyper Extended A-PDHG Operator for Constrained High-dim PDEs · ICML 2025 |
Machine learning and data management
active learning |
0.2 | 1 | 2014 | Active Learning without Knowing Individual Instance Labels: A Pairwise Label Homogeneity Query Approach · IEEE Trans. Knowl. Data Eng. 2014 |
Data mining › predictive modeling
regression |
0.2 | 1 | 2014 | Tensor Regression Based on Linked Multiway Parameter Analysis · ICDM 2014 |
Machine learning and data management › tensor learning
tensor regression |
0.2 | 1 | 2014 | Tensor Regression Based on Linked Multiway Parameter Analysis · ICDM 2014 |
Machine learning › Efficient and distributed learning
active learning |
0.1 | 1 | 2011 | Optimal Subset Selection for Active Learning · AAAI 2011 |
Machine learning › Efficient and distributed learning
subset selection |
0.1 | 1 | 2011 | Optimal Subset Selection for Active Learning · AAAI 2011 |
Graph algorithms and graph theory › minimum cut
max-flow min-cut |
0.1 | 1 | 2014 | Active Learning without Knowing Individual Instance Labels: A Pairwise Label Homogeneity Query Approach · IEEE Trans. Knowl. Data Eng. 2014 |
Mathematical optimization
semidefinite programming |
0.0 | 1 | 2011 | Optimal Subset Selection for Active Learning · AAAI 2011 |
Methods — techniques the papers use, named apart from their topics
quadratic programming · 1.7neural operator · 1.7adaptive primal-dual hybrid gradient · 1.7min-cut · 0.4max-flow · 0.4confidence-based data selection · 0.4uncertainty estimation · 0.2semidefinite programming · 0.2correlation matrix · 0.2tucker decomposition · 0.2sparsity regularization · 0.2coordinate descent · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multibranch Dendritic Computing for Federated Graph Neural Networks: Maintaining Interpretability Under Privacy ConstraintsabstractGraph Neural Networks (GNNs) have demonstrated remarkable performance in modeling complex relational data. However, their deployment in privacy-sensitive domains remains challenging due to the inherent tension between privacy protection and model interpretability. This paper introduces a novel Multi-Branch Dendritic Computing framework for Federated Graph Neural Networks (MBD-FGNN) that addresses this fundamental challenge. Our approach integrates biologically-inspired dendritic computing with federated learning paradigms, enabling distributed training across decentralized graph data while preserving both privacy and interpretability. The multi-branch architecture mimics dendritic computation in biological neurons,where each branch processes patterns at different receptive field scales (1-hop, 2-hop, and 4-hop neighborhoods), providing natural model explanations through branch contribution analysis.To ensure rigorous privacy protection, we develop a branch-adaptive differential privacy mechanism with gradient clipping and calibrated Gaussian noise injection, satisfying (ϵ, δ)-differential privacy. We theoretically analyze the framework’s ability to maintain interpretation stability under differential privacy constraints and derive formal guarantees on the privacy-interpretability trade-off. Extensive experiments on benchmark graph datasets demonstrate that MBD-FGNN outperforms state-of-the-art federated GNN approaches in terms of accuracy, privacy preservation, and explanation quality. Notably, our method achieves robust explanation stability even under privacy constraints, maintaining higher explanation consistency compared to conventional approaches where explanations significantly degrade as privacy protection increases. The proposed framework opens new avenues for deploying interpretable graph learning systems in privacy-critical applications such as healthcare networks, financial transaction graphs, and social network analysis. The code and dataset are publicly available at https://github.com/ZhaoGuan-nuist/MBD-FGNN. Yifan Fu, Min Xia 0002, Liguo Weng, Haifeng Lin |
IEEE Internet Things J. | 1 |
| 2025 | HEAP: Hyper Extended A-PDHG Operator for Constrained High-dim PDEsabstractNeural operators have emerged as a promising approach for solving high-dimensional partial differential equations (PDEs). However, existing neural operators often have difficulty in dealing with constrained PDEs, where the solution must satisfy additional equality or inequality constraints beyond the governing equations. To close this gap, we propose a novel neural operator, Hyper Extended Adaptive PDHG (HEAP) for constrained high-dim PDEs, where the learned operator evolves in the parameter space of PDEs. We first show that the evolution operator learning can be formulated as a quadratic programming (QP) problem, then unroll the adaptive primal-dual hybrid gradient (APDHG) algorithm as the QP-solver into the neural operator architecture. This allows us to improve efficiency while retaining theoretical guarantees of the constrained optimization. Empirical results on a variety of high-dim PDEs show that HEAP outperforms the state-of-the-art neural operator model. Mingquan Feng, Weixin Liao, Yifan Fu, Qifu Zheng, Junchi Yan |
ICML | 4 |
| 2025 | Feature extraction and fusion algorithm for infrared visible light images based on residual and generative adversarial network
Naigong Yu, Yifan Fu, Qiusheng Xie, Qiming Cheng, Mohammad Mehedi Hasan |
Image Vis. Comput. | 2 |
| 2025 | FSMT: Few-shot object detection via Multi-Task DecoupledabstractWith the advancement of object detection technology, few-shot object detection (FSOD) has become a research hotspot. Existing methods face two major challenges: base models have limited generalization to unseen categories, especially with limited few-shot data, where the shared feature representation fails to meet the distinct needs of classification and regression tasks; FSOD is susceptible to overfitting during training. To address these issues, this paper proposes a Multi-Task Decoupled Method (MTDM), which enhances the model’s generalization to new categories by separating the feature extraction processes for different tasks. Additionally, a dynamic adjustment strategy is adopted, which adaptively modifies the IOU threshold and loss function parameters based on variations in the training data, reducing the risk of overfitting and maximizing the utilization of limited data resources. Experimental results show that the proposed hybrid model performs well on multiple few-shot datasets, effectively overcoming the challenges posed by limited annotated data. Jiahui Qin, Yang Xu 0006, Yifan Fu, Zebin Wu 0001, Zhihui Wei |
Pattern Recognit. Lett. | 3 |
| 2025 | Dual Prediction-Guided Distillation for Object Detection in Remote Sensing ImagesabstractKnowledge distillation, which intends to transfer the expertise from a complex teacher model to a concise student model, has achieved impressive success in object detection. However, many existing distillation methods are designed for object detection tasks on natural images and perform poorly on more challenging remote sensing images. In this study, we first attribute the deficiencies of knowledge distillation in remote sensing object detection to two core reasons: 1) lack of relation distillation among different instances and pixels and 2) significant differences in feature magnitude between the teacher-student pair. Then, we propose a dual prediction-guided knowledge distillation framework, which includes relation distillation and output distillation to address two issues, respectively. Prediction-guided relation distillation (PGRD) is proposed to capture the knowledge of global relation at both instance-wise and pixel-wise, then allowing students to better understand and depict the features distribution across different categories. Prediction-guided output distillation (PGOD) is proposed to mitigate the impact of feature magnitude inconsistencies on distillation with classification and location knowledge, then allowing students to directly capture task-relevant information. Finally, experimental results have indicated the consistent effectiveness of our method across anchor-based two-stage, one-stage, and anchor-free detectors with 11 comparison knowledge distillation methods on three remote sensing detection datasets. Source codes are available athttps://github.com/RQ-W/DPGD.git. Ruiqing Wang, Yifan Fu, Yang Xu 0006, Zebin Wu 0001, Zhihui Wei |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | A haptic feedback glove for virtual piano interactionabstractHaptic feedback plays a crucial role in virtual reality (VR) interaction, helping to improve the precision of user operation and enhancing the immersion of the user experience. Instrumental haptic feedback in virtual environments is primarily realized using grounded force or vibration feedback devices. However, improvements are required in terms of the active space and feedback realism. We propose a lightweight and flexible haptic feedback glove that can haptically render objects in VR environments via kinesthetic and vibration feedback, thereby enabling users to enjoy a rich virtual piano-playing experience. The kinesthetic feedback of the glove relies on a cable-pulling mechanism that rotates the mechanism and pulls the two cables connected to it, thereby changing the amount of force generated to simulate the hardness or softness of the object. Vibration feedback is provided by small vibration motors embedded in the bottom of the fingertips of the glove. We designed a piano-playing scenario in the virtual environment and conducted user tests. The evaluation metrics were clarity, realism, enjoyment, and satisfaction. A total of 14 subjects participated in the test, and the results showed that our proposed glove scored significantly higher on the four evaluation metrics than the no-feedback and vibration feedback methods. Our proposed glove significantly enhances the user experience when interacting with virtual objects. Yifan Fu, Xiaoying Sun |
Virtual Real. Intell. Hardw. | 1 |
| 2016 | Prognosis of chip-loss failure in high-power IGBT module by self-testingabstractReliability is of great importance for power converter systems, such as metro, high-speed train and offshore wind turbine. Bonding wire lift-off is a common failure in IGBT module due to the thermomechanical fatigue under cyclical temperature and power swings. Consequently, this will lead to the chip-loss failure gradually in a high power multi-chip IGBT module during long-term operation. To prevent a catastrophic failure from happening, a self-testing technique for IGBT module is proposed in this paper to detect the chip-loss failure at initial stage. During the test, a shoot-through current is imposed on the device under test (DUT) by gate control and its peak value is selected as the failure criterion for chip-loss failure prognosis. Since the IGBT switching behavior changes after failure, the shoot-through current will deviate from its initial value, which can be detected for failure prognosis. For safety, the self-testing is conducted intermittently on-site during system down time and the shoot-through current is limited less than the IGBT rated current at lowered DC voltage. To clarify the method, its test scheme, working principle and implementation are firstly discussed in the paper. Then, experimental work was carried out on a test rig with 3.3kV/800A IGBT modules (emulating a metro traction drive system) to verify the method. Preliminary test results demonstrate that the proposed self-testing method is a simple and effective technique for detecting the gradually developed IGBT chip-loss failure. Yeke Liu, Dawei Xiang, Yifan Fu |
IECON | 3 |
| 2016 | Tensor LRR and Sparse Coding-Based Subspace ClusteringabstractSubspace clustering groups a set of samples from a union of several linear subspaces into clusters, so that the samples in the same cluster are drawn from the same linear subspace. In the majority of the existing work on subspace clustering, clusters are built based on feature information, while sample correlations in their original spatial structure are simply ignored. Besides, original high-dimensional feature vector contains noisy/redundant information, and the time complexity grows exponentially with the number of dimensions. To address these issues, we propose a tensor low-rank representation (TLRR) and sparse coding-based (TLRRSC) subspace clustering method by simultaneously considering feature information and spatial structures. TLRR seeks the lowest rank representation over original spatial structures along all spatial directions. Sparse coding learns a dictionary along feature spaces, so that each sample can be represented by a few atoms of the learned dictionary. The affinity matrix used for spectral clustering is built from the joint similarities in both spatial and feature spaces. TLRRSC can well capture the global structure and inherent feature information of data and provide a robust subspace segmentation from corrupted data. Experimental results on both synthetic and real-world data sets show that TLRRSC outperforms several established stateof- the-art methods. Yifan Fu, Junbin Gao, David Tien, Zhouchen Lin, Xia Hong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Low Rank Representation on Riemannian Manifold of Symmetric Positive Definite MatricesabstractSparse coding aims to find a more compact representation based on a set of dictionary atoms. A well-known technique looking at 2D sparsity is the low rank representation (LRR). However, in many computer vision applications, data often originate from a manifold, which is equipped with some Riemannian geometry. In this case, the existing LRR becomes inappropriate for modeling and incorporating the intrinsic geometry of the manifold that is potentially important and critical to applications. In this paper, we generalize the LRR over the Euclidean space to the LRR model over a specific Rimannian manifold—the manifold of symmetric positive matrices (SPD). Experiments on several computer vision datasets showcase its noise robustness and superior performance on classification and segmentation compared with state-of-the-art approaches. Yifan Fu, Junbin Gao, Xia Hong 0001, David Tien |
SDM | 1 |
| 2014 | Tensor Regression Based on Linked Multiway Parameter AnalysisabstractClassical regression methods take vectors as covariates and estimate the corresponding vectors of regression parameters. When addressing regression problems on covariates of more complex form such as multi-dimensional arrays (i.e. Tensors), traditional computational models can be severely compromised by ultrahigh dimensionality as well as complex structure. By exploiting the special structure of tensor covariates, the tensor regression model provides a promising solution to reduce the model's dimensionality to a manageable level, thus leading to efficient estimation. Most of the existing tensor-based methods independently estimate each individual regression problem based on tensor decomposition which allows the simultaneous projections of an input tensor to more than one direction along each mode. As a matter of fact, multi-dimensional data are collected under the same or very similar conditions, so that data share some common latent components but can also have their own independent parameters for each regression task. Therefore, it is beneficial to analyse regression parameters among all the regressions in a linked way. In this paper, we propose a tensor regression model based on Tucker Decomposition, which identifies not only the common components of parameters across all the regression tasks, but also independent factors contributing to each particular regression task simultaneously. Under this paradigm, the number of independent parameters along each mode is constrained by a sparsity-preserving regulariser. Linked multiway parameter analysis and sparsity modeling further reduce the total number of parameters, with lower memory cost than their tensor-based counterparts. The effectiveness of the new method is demonstrated on real data sets. Yifan Fu, Junbin Gao, Xia Hong 0001, David Tien |
ICDM | 1 |
| 2014 | Joint multiple dictionary learning for Tensor sparse codingabstractTraditional dictionary learning algorithms are used for finding a sparse representation on high dimensional data by transforming samples into a one-dimensional (ID) vector. This ID model loses the inherent spatial structure property of data. An alternative solution is to employ Tensor Decomposition for dictionary learning on their original structural form - a tensor - by learning multiple dictionaries along each mode and the corresponding sparse representation in respect to the Kronecker product of these dictionaries. To learn tensor dictionaries along each mode, all the existing methods update each dictionary iteratively in an alternating manner. Because atoms from each mode dictionary jointly make contributions to the spar sity of tensor, existing works ignore atoms correlations between different mode dictionaries by treating each mode dictionary independently. In this paper, we propose a joint multiple dictionary learning method for tensor sparse coding, which explores atom correlations for sparse representation and updates multiple atoms from each mode dictionary simultaneously. In this algorithm, the Frequent-Pattern Tree (FP-tree) mining algorithm is employed to exploit frequent atom patterns in the sparse representation. Inspired by the idea of K-SVD, we develop a new dictionary update method that jointly updates elements in each pattern. Experimental results demonstrate our method outperforms other tensor based dictionary learning algorithms. Yifan Fu, Junbin Gao, Xia Hong 0001 |
IJCNN | 1 |
| 2014 | Tensor LRR based subspace clusteringabstractSubspace clustering groups a set of samples (vectors) into clusters by approximating this set with a mixture of several linear subspaces, so that the samples in the same cluster are drawn from the same linear subspace. In majority of existing works on subspace clustering, samples are simply regarded as being independent and identically distributed, that is, arbitrarily ordering samples when necessary. However, this setting ignores sample correlations in their original spatial structure. To address this issue, we propose a tensor low-rank representation (TLRR) for subspace clustering by keeping available spatial information of data. TLRR seeks a lowest-rank representation over all the candidates while maintaining the inherent spatial structures among samples, and the affinity matrix used for spectral clustering is built from the combination of similarities along all data spatial directions. TLRR better captures the global structures of data and provides a robust subspace segmentation from corrupted data. Experimental results on both synthetic and real-world datasets show that TLRR outperforms several established state-of-the-art methods. Yifan Fu, Junbin Gao, David Tien, Zhouchen Lin |
IJCNN | 1 |
| 2014 | Active Learning without Knowing Individual Instance Labels: A Pairwise Label Homogeneity Query ApproachabstractTraditional active learning methods require the labeler to provide a class label for each queried instance. The labelers are normally highly skilled domain experts to ensure the correctness of the provided labels, which in turn results in expensive labeling cost. To reduce labeling cost, an alternative solution is to allow nonexpert labelers to carry out the labeling task without explicitly telling the class label of each queried instance. In this paper, we propose a new active learning paradigm, in which a nonexpert labeler is only asked “whether a pair of instances belong to the same class”, namely, a pairwise label homogeneity. Under such circumstances, our active learning goal is twofold: (1) decide which pair of instances should be selected for query, and (2) how to make use of the pairwise homogeneity information to improve the active learner. To achieve the goal, we propose a “Pairwise Query on Max-flow Paths” strategy to query pairwise label homogeneity from a nonexpert labeler, whose query results are further used to dynamically update a Min-cut model (to differentiate instances in different classes). In addition, a “Confidence-based Data Selection” measure is used to evaluate data utility based on the Min-cut model’s prediction results. The selected instances, with inferred class labels, are included into the labeled set to form a closed-loop active learning process. Experimental results and comparisons with state-of-the-art methods demonstrate that our new active learning paradigm can result in good performance with nonexpert labelers. Yifan Fu, Bin Li 0015, Xingquan Zhu 0001, Chengqi Zhang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2013 | A survey on instance selection for active learning
Yifan Fu, Xingquan Zhu 0001, Bin Li 0015 |
Knowl. Inf. Syst. | 1 |
| 2013 | Active Learning With Optimal Instance Subset SelectionabstractActive learning (AL) traditionally relies on some instance-based utility measures (such as uncertainty) to assess individual instances and label the ones with the maximum values for training. In this paper, we argue that such approaches cannot produce good labeling subsets mainly because instances are evaluated independently without considering their interactions, and individuals with maximal ability do not necessarily form an optimal instance subset for learning. Alternatively, we propose to achieve AL with optimal subset selection (ALOSS), where the key is to find an instance subset with a maximum utility value. To achieve the goal, ALOSS simultaneously considers the following: 1) the importance of individual instances and 2) the disparity between instances, to build an instance-correlation matrix. As a result, AL is transformed to a semidefinite programming problem to select a k-instance subset with a maximum utility value. Experimental results demonstrate that ALOSS outperforms state-of-the-art approaches for AL. Yifan Fu, Xingquan Zhu 0001, Ahmed K. Elmagarmid |
IEEE Trans. Cybern. | 1 |
| 2011 | Optimal Subset Selection for Active LearningabstractActive learning traditionally relies on instance based utility measures to rank and select instances for labeling, which may result in labeling redundancy. To address this issue, we explore instance utility from two dimensions: individual uncertainty and instance disparity, using a correlation matrix. The active learning is transformed to a semi-definite programming problem to select an optimal subset with maximum utility value. Experiments demonstrate the algorithm performance in comparison with baseline approaches. Yifan Fu, Xingquan Zhu 0001 |
AAAI | 1 |
| 2011 | Do they belong to the same class: active learning by querying pairwise label homogeneityabstractTraditional active learning methods request experts to provide ground truths to the queried instances, which can be expensive in practice. An alternative solution is to ask nonexpert labelers to do such labeling work, which can not tell the definite class labels. In this paper, we propose a new active learning paradigm, in which a nonexpert labeler is only asked "whether a pair of instances belong to the same class". To instantiate the proposed paradigm, we adopt the MinCut algorithm as the base classifier. We first construct a graph based on the pairwise distance of all the labeled and unlabeled instances and then repeatedly update the unlabeled edge weights on the max-flow paths in the graph. Finally, we select an unlabeled subset of nodes with the highest prediction confidence as the labeled data, which are included into the labeled data set to learn a new classifier for the next round of active learning. The experimental results and comparisons, with state-of-the-art methods, demonstrate that our active learning paradigm can result in good performance with nonexpert labelers. Yifan Fu, Bin Li 0015, Xingquan Zhu 0001, Chengqi Zhang |
CIKM | 1 |