VLDB 2026 Research / reviewers in the wild / expert
Guihong Wan
dblp:245/3380
· DBLP profile ↗
20ranked-venue papers
13as first author
16since 2021 · last 2026
0000-0003-1100-4018ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 9 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A survey on computational pathology foundation models: datasets, adaptation strategies, and evaluation tasksabstractAbstract Computational pathology foundation models (CPathFMs) have emerged as a powerful approach for analyzing histopathological data, leveraging self-supervised learning to extract robust feature representations from unlabeled whole-slide images. These models, categorized into uni-modal and multi-modal frameworks, have demonstrated promise in automating complex pathology tasks such as segmentation, classification, and biomarker discovery. However, the development of CPathFMs presents significant challenges, such as limited data accessibility, high variability across datasets, the necessity for domain-specific adaptation, and the lack of standardized evaluation benchmarks. This survey provides a comprehensive review of CPathFMs in computational pathology, focusing on pre-training datasets, adaptation strategies, and evaluation tasks. We analyze key techniques, such as contrastive learning, masked image modeling and multi-modal integration, and highlight existing gaps in current research. Finally, we explore future directions from four perspectives for advancing CPathFMs. This survey serves as a valuable resource for researchers, clinicians, and AI practitioners, guiding the advancement of CPathFMs toward robust and clinically applicable AI-driven pathology solutions. Dong Li 0034, Guihong Wan, Xintao Wu, Yi He 0007, Zhong Chen 0003, Ajit Johnson Nirmal, Christine G. Lian, Peter K. Sorger, Yevgeniy R. Semenov, Chen Zhao 0010 |
Knowl. Inf. Syst. | 2 |
| 2025 | Multi-View Unsupervised Column Subset Selection via Combinatorial Search (Student Abstract)abstractGiven a data matrix, unsupervised column subset selection refers to the problem of identifying a subset of columns that can be used to linearly approximate the original data matrix. This problem has many applications, such as feature selection and representative selection, but solving it optimally is known to be NP-hard. We consider multi-view unsupervised column subset selection, which extends the concept of (single-view) column subset selection to data represented in multiple views or modalities. We introduce a combinatorial search algorithm for this generalized problem. One variant of the algorithm is guaranteed to compute an optimal solution in a setting similar to the classical A* algorithm. Other suboptimal variants, in a setting similar to the weighted A* algorithm, are much faster and provide a solution along with a bound on its quality. Guihong Wan, Ninghui Hao, Crystal Maung, Haim Schweitzer, Chen Zhao 0010, Kun-Hsing Yu, Yevgeniy R. Semenov |
AAAI | 1 |
| 2025 | FADE: Towards Fairness-aware Data Generation for Domain Generalization via Classifier-Guided Score-based Diffusion ModelsabstractFairness-aware domain generalization (FairDG) has emerged as a critical challenge for deploying trustworthy AI systems, particularly in scenarios involving distribution shifts. Traditional methods for addressing fairness have failed in domain generalization due to their lack of consideration for distribution shifts. Although disentanglement has been used to tackle FairDG, it is limited by its strong assumptions. To overcome these limitations, we propose Fairness-aware Classifier-Guided Score-based Diffusion Models (FADE) as a novel approach to effectively address the FairDG issue. Specifically, we first pre-train a score-based diffusion model (SDM) and two classifiers to equip the model with strong generalization capabilities across different domains. Then, we guide the SDM using these pre-trained classifiers to effectively eliminate sensitive information from the generated data. Finally, the generated fair data is used to train downstream classifiers, ensuring robust performance under new data distributions. Extensive experiments on three real-world datasets demonstrate that FADE not only enhances fairness but also improves accuracy in the presence of distribution shifts. Additionally, FADE outperforms existing methods in achieving the best accuracy-fairness trade-offs. Dong Li 0034, Minglai Shao 0001, Guihong Wan, Chen Zhao 0010 |
IJCAI | 4 |
| 2024 | Graph Clustering Methods Derived from Column Subset Selection (Student Abstract)abstractSpectral clustering is a powerful clustering technique. It leverages the spectral properties of graphs to partition data points into meaningful clusters. The most common criterion for evaluating multi-way spectral clustering is NCut. Column Subset Selection is an important optimization technique in the domain of feature selection and dimension reduction which aims to identify a subset of columns of a given data matrix that can be used to approximate the entire matrix. We show that column subset selection can be used to compute spectral clustering and use this to obtain new graph clustering algorithms. Guihong Wan, Haim Schweitzer |
AAAI | 2 |
| 2024 | Equivalence between Graph Spectral Clustering and Column Subset Selection (Student Abstract)abstractThe common criteria for evaluating spectral clustering are NCut and RatioCut. The seemingly unrelated column subset selection (CSS) problem aims to compute a column subset that linearly approximates the entire matrix. A common criterion is the approximation error in the Frobenius norm (ApproxErr). We show that any algorithm for CSS can be viewed as a clustering algorithm that minimizes NCut by applying it to a matrix formed from graph edges. Conversely, any clustering algorithm can be seen as identifying a column subset from that matrix. In both cases, ApproxErr and NCut have the same value. Analogous results hold for RatioCut with a slightly different matrix. Therefore, established results for CSS can be mapped to spectral clustering. We use this to obtain new clustering algorithms, including an optimal one that is similar to A*. This is the first nontrivial clustering algorithm with such an optimality guarantee. A variant of the weighted A* runs much faster and provides bounds on the accuracy. Finally, we use the results from spectral clustering to prove the NP-hardness of CSS from sparse matrices. Guihong Wan, Yevgeniy R. Semenov, Haim Schweitzer |
AAAI | 1 |
| 2024 | Pass-Efficient Algorithms for Graph Spectral Clustering (Student Abstract)abstractGraph spectral clustering is a fundamental technique in data analysis, which utilizes eigenpairs of the Laplacian matrix to partition graph vertices into clusters. However, classical spectral clustering algorithms require eigendecomposition of the Laplacian matrix, which has cubic time complexity. In this work, we describe pass-efficient spectral clustering algorithms that leverage recent advances in randomized eigendecomposition and the structure of the graph vertex-edge matrix. Furthermore, we derive formulas for their efficient implementation. The resulting algorithms have a linear time complexity with respect to the number of vertices and edges and pass over the graph constant times, making them suitable for processing large graphs stored on slow memory. Experiments validate the accuracy and efficiency of the algorithms. Boshen Yan, Guihong Wan, Haim Schweitzer, Zoltan Maliga, Sara Khattab, Kun-Hsing Yu, Peter K. Sorger, Yevgeniy R. Semenov |
AAAI | 2 |
| 2024 | SpatialCells: automated profiling of tumor microenvironments with spatially resolved multiplexed single-cell dataabstractCancer is a complex cellular ecosystem where malignant cells coexist and interact with immune, stromal and other cells within the tumor microenvironment (TME). Recent technological advancements in spatially resolved multiplexed imaging at single-cell resolution have led to the generation of large-scale and high-dimensional datasets from biological specimens. This underscores the necessity for automated methodologies that can effectively characterize molecular, cellular and spatial properties of TMEs for various malignancies. This study introduces SpatialCells, an open-source software package designed for region-based exploratory analysis and comprehensive characterization of TMEs using multiplexed single-cell data. The source code and tutorials are available at https://semenovlab.github.io/SpatialCells. SpatialCells efficiently streamlines the automated extraction of features from multiplexed single-cell data and can process samples containing millions of cells. Thus, SpatialCells facilitates subsequent association analyses and machine learning predictions, making it an essential tool in advancing our understanding of tumor growth, invasion and metastasis. Guihong Wan, Zoltan Maliga, Boshen Yan, Tuulia Vallius, Yingxiao Shi, Sara Khattab, Crystal T. Chang, Ajit Johnson Nirmal, Kun-Hsing Yu, David S. L. Wei, Christine G. Lian, Mia S. Desimone, Peter K. Sorger, Yevgeniy R. Semenov |
Briefings Bioinform. | 1 |
| 2024 | The art of centering without centering for robust principal component analysis
Guihong Wan, Baokun He, Haim Schweitzer |
Data Min. Knowl. Discov. | 1 |
| 2023 | Electrophysiological Brain Source Imaging via Combinatorial Search with Provable OptimalityabstractElectrophysiological Source Imaging (ESI) refers to reconstructing the underlying brain source activation from non-invasive Electroencephalography (EEG) and Magnetoencephalography (MEG) measurements on the scalp. Estimating the source locations and their extents is a fundamental tool in clinical and neuroscience applications. However, the estimation is challenging because of the ill-posedness and high coherence in the leadfield matrix as well as the noise in the EEG/MEG data. In this work, we proposed a combinatorial search framework to address the ESI problem with a provable optimality guarantee. Specifically, by exploiting the graph neighborhood information in the brain source space, we converted the ESI problem into a graph search problem and designed a combinatorial search algorithm under the framework of A* to solve it. The proposed algorithm is guaranteed to give an optimal solution to the ESI problem. Experimental results on both synthetic data and real epilepsy EEG data demonstrated that the proposed algorithm could faithfully reconstruct the source activation in the brain. Guihong Wan, Meng Jiao, Xinglong Ju, Haim Schweitzer |
AAAI | 1 |
| 2022 | Extended Electrophysiological Source Imaging with Spatial Graph Filters
Feng Liu 0011, Guihong Wan, Yevgeniy R. Semenov, Patrick L. Purdon |
MICCAI (1) | 2 |
| 2021 | Accelerated Combinatorial Search for Outlier Detection with Provable Bound on Sub-OptimalityabstractOutliers negatively affect the accuracy of data analysis. In this paper we are concerned with their influence on the accuracy of Principal Component Analysis (PCA). Algorithms that attempt to detect outliers and remove them from the data prior to applying PCA are sometimes called Robust PCA, or Robust Subspace Recovery algorithms. We propose a new algorithm for outlier detection that combines two ideas. The first is "chunk recursive elimination" that was used effectively to accelerate feature selection, and the second is combinatorial search, in a setting similar to A*. Our main result is showing how to combine these two ideas. One variant of our algorithm is guaranteed to compute an optimal solution according to some natural criteria, but its running time makes it impractical for large datasets. Other variants are much faster and come with provable bounds on sub-optimality. Experimental results show the effectiveness of the proposed approach. Guihong Wan, Haim Schweitzer |
AAAI | 1 |
| 2021 | A New Robust Subspace Recovery Algorithm (Student Abstract)abstractA common task in data analysis is to compute an approximate embedding of the data in a low dimensional subspace. This is used, for example, for dimensionality reduction. Robust Subspace Recovery computes the embedding by ignoring a fraction of the data considered as outliers. Its performance can be evaluated by how accurate the inliers (non-outliers) are represented. We propose a new algorithm that outperforms the current state of the art when the data is dominated by outliers. The main idea is to rank each point by evaluating the change in the global PCA error when that point is considered as an outlier. We show that this lookahead procedure can be implemented efficiently by centered rank-one modifications. Guihong Wan, Haim Schweitzer |
AAAI | 1 |
| 2021 | Edge Sparsification for Graphs via Meta-LearningabstractWe present a novel edge sparsification approach for semi-supervised learning on undirected and attributed graphs. The main challenge is to retain few edges while minimizing the loss of node classification accuracy. The task can be mathematically formulated as a bi-level optimization problem. We propose to use meta-gradients, which have traditionally been used in meta-learning, to solve the optimization problem, specifically, treating the graph adjacency matrix as hyperparameters to optimize. Experimental results show the effectiveness of the proposed approach. Remarkably, with the resulting sparse and light graph, in many cases the classification accuracy is significantly improved. Guihong Wan, Haim Schweitzer |
ICDE | 1 |
| 2021 | A Lookahead Algorithm for Robust Subspace RecoveryabstractA common task in the analysis of data is to compute an approximate embedding of the data in a low-dimensional subspace. The standard algorithm for computing this subspace is the well-known Principal Component Analysis (PCA). PCA can be extended to the case where some data points are viewed as “outliers” that can be ignored, allowing the remaining data points (inliers”) to be more tightly embedded. We develop a new algorithm that detects outliers so that they can be removed prior to applying PCA. The main idea is to rank each point by looking ahead and evaluating the change in the global PCA error if that point is considered as an outlier. Our technical contribution is showing that this lookahead procedure can be implemented efficiently, producing an accurate algorithm with running time not much above the running time of standard PCA algorithms. Guihong Wan, Haim Schweitzer |
ICDM | 1 |
| 2021 | Heuristic Search for Approximating One Matrix in Terms of Another MatrixabstractWe study the approximation of a target matrix in terms of several selected columns of another matrix, sometimes called "a dictionary". This approximation problem arises in various domains, such as signal processing, computer vision, and machine learning. An optimal column selection algorithm for the special case where the target matrix has only one column is known since the 1970's, but most previously proposed column selection algorithms for the general case are greedy. We propose the first nontrivial optimal algorithm for the general case, using a heuristic search setting similar to the classical A* algorithm. We also propose practical sub-optimal algorithms in a setting similar to the classical Weighted A* algorithm. Experimental results show that our sub-optimal algorithms compare favorably with the current state-of-the-art greedy algorithms. They also provide bounds on how close their solutions are to the optimal solution. Guihong Wan, Haim Schweitzer |
IJCAI | 1 |
| 2021 | A Fast Algorithm for Simultaneous Sparse Approximation
Guihong Wan, Haim Schweitzer |
PAKDD (3) | 1 |
| 2020 | A Bias Trick for Centered Robust Principal Component Analysis (Student Abstract)abstractOutlier based Robust Principal Component Analysis (RPCA) requires centering of the non-outliers. We show a “bias trick” that automatically centers these non-outliers. Using this bias trick we obtain the first RPCA algorithm that is optimal with respect to centering. Baokun He, Guihong Wan, Haim Schweitzer |
AAAI | 2 |
| 2020 | Fast Distance Metrics in Low-dimensional Space for Neighbor Search ProblemsabstractWe consider popular dimension reduction techniques that project data on a low dimensional subspace. They include Principal Component Analysis, Column Subset Selection, and Johnson-Lindenstrauss projections. These techniques have been classically used to efficiently compute various approximations. We propose the following three-step procedure for enhancing the accuracy of such approximations: 1. Unknown quantities in the approximation are replaced with random variables. 2. The Maximum Entropy method is applied to infer the most likely probability distribution. 3. Expected values of the random variables are used to compute the enhanced estimates. Our use of the Maximum Entropy method requires knowledge of vector norms that can be easily computed during the dimension reduction. We demonstrate significant enhancements in average accuracy for Euclidean distance and Mahalanobis distance, and improvements in evaluating k-nearest neighbors and k-furthest neighbors by using the enhanced Euclidean distance formula. Guihong Wan, Crystal Maung, Haim Schweitzer |
ICDM | 1 |
| 2019 | Heuristic Search Algorithm for Dimensionality Reduction Optimally Combining Feature Selection and Feature ExtractionabstractThe following are two classical approaches to dimensionality reduction: 1. Approximating the data with a small number of features that exist in the data (feature selection). 2. Approximating the data with a small number of arbitrary features (feature extraction). We study a generalization that approximates the data with both selected and extracted features. We show that an optimal solution to this hybrid problem involves a combinatorial search, and cannot be trivially obtained even if one can solve optimally the separate problems of selection and extraction. Our approach that gives optimal and approximate solutions uses a “best first” heuristic search. The algorithm comes with both an a priori and an a posteriori optimality guarantee similar to those that can be obtained for the classical weighted A* algorithm. Experimental results show the effectiveness of the proposed approach. Baokun He, Swair Shah, Crystal Maung, Gordon Arnold, Guihong Wan, Haim Schweitzer |
AAAI | 5 |
| 2019 | Improving the Accuracy of Principal Component Analysis by the Maximum Entropy MethodabstractClassical Principal Component Analysis (PCA) approximates data in terms of projections on a small number of orthogonal vectors. There are simple procedures to efficiently compute various functions of the data from the PCA approximation. The most important function is arguably the Euclidean distance between data items. This can be used, for example, to solve the approximate nearest neighbor problem. We use random variables to model the inherent uncertainty in such approximations, and apply the Maximum Entropy Method to infer the underlying probability distribution. We propose using the expected values of distances between these random variables as improved estimates of the distance. We show experimentally that in most cases results obtained by our method are more accurate than what is obtained by the classical approach. This improves the accuracy of a classical technique that have been used with little change for over 100 years. Guihong Wan, Crystal Maung, Haim Schweitzer |
ICTAI | 1 |