VLDB 2026 Research / reviewers in the wild / expert
Pengfei Gu
dblp:215/4376
· DBLP profile ↗
23ranked-venue papers
8as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Theory of computation · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Integrating Multi-scale and Multi-filtration Topological Features for Medical Image Classification*abstractModern deep neural networks have shown remarkable performance in medical image classification. However, such networks either emphasize pixel-intensity features instead of fundamental anatomical structures (e.g., those encoded by topological invariants), or they capture only simple topological features via single-parameter persistence. In this paper, we propose a new topology-guided classification framework that extracts multi-scale and multi-filtration persistent topological features and integrates them into vision classification backbones. For an input image, we first compute cubical persistence diagrams (PDs) across multiple image resolutions/scales. We then develop a "vineyard" algorithm that consolidates these PDs into a single, stable diagram capturing signatures at varying granularities, from global anatomy to subtle local irregularities that may indicate early-stage disease. To further exploit richer topological representations produced by multiple filtrations, we design a cross-attention-based neural network that directly processes the consolidated final PDs. The resulting topological embeddings are fused with feature maps from CNNs or Transformers. By integrating multi-scale and multi-filtration topologies into an end-to-end architecture, our approach enhances the model’s capacity to recognize complex anatomical structures. Evaluations on three public datasets show consistent, considerable improvements over strong baselines and state-of-the-art methods, demonstrating the value of our comprehensive topological perspective for robust and interpretable medical image classification. Pengfei Gu, Haoteng Tang, Dongkuan Xu, Erik Enriquez, DongChul Kim, Danny Ziyi Chen |
WACV | 1 |
| 2025 | Self Pre-Training with Topology- and Spatiality-Aware Masked Autoencoders for 3D Medical Image Segmentation
Pengfei Gu, Yejia Zhang, Chaoli Wang 0001, Danny Ziyi Chen |
BIBM | 1 |
| 2025 | Topo-VM-UNetV2: Encoding Topology Into Vision Mamba UNet for Polyp SegmentationabstractConvolutional neural network (CNN) and Transformer-based architectures are two dominant deep learning models for polyp segmentation. However, CNNs have limited capability for modeling long-range dependencies, while Transformers incur quadratic computational complexity. Recently, State Space Models such as Mamba have been recognized as a promising approach for polyp segmentation because they not only model long-range interactions effectively but also maintain linear computational complexity. However, Mamba-based architectures still struggle to capture topological features (e.g., connected components, loops, voids), leading to inaccurate boundary delineation and polyp segmentation. To address these limitations, we propose a new approach called Topo-VM-UNetV2, which encodes topological features into the Mamba-based state-of-the-art polyp segmentation model, VM-UNetV2. Our method consists of two stages: Stage 1: VM-UNetV2 is used to generate probability maps (PMs) for the training and test images, which are then used to compute topology attention maps. Specifically, we first compute persistence diagrams of the PMs, then we generate persistence score maps by assigning persistence values (i.e., the difference between death and birth times) of each topological feature to its birth location, finally we transform persistence scores into attention weights using the sigmoid function. Stage 2: These topology attention maps are integrated into the semantics and detail infusion (SDI) module of VM-UNetV2 to form a topologyguided semantics and detail infusion (Topo-SDI) module for enhancing the segmentation results. Extensive experiments on five public polyp segmentation datasets demonstrate the effectiveness of our proposed method. The code will be made publicly available. Diego Adame, Jose Angel Nuñez, Fabian Vazquez, Nayeli Gurrola, Haoteng Tang, Pengfei Gu |
CBMS | 8 |
| 2025 | Adapting a Segmentation Foundation Model for Medical Image ClassificationabstractRecent advancements in foundation models, such as the Segment Anything Model (SAM), have shown strong performance in various vision tasks, particularly image segmentation, due to their impressive zero-shot segmentation capabilities. However, effectively adapting such models for medical image classification is still a less explored topic. In this paper, we introduce a new framework to adapt SAM for medical image classification. First, we utilize the SAM image encoder as a feature extractor to capture segmentation-based features that convey important spatial and contextual details of the image, while freezing its weights to avoid unnecessary overhead during training. Next, we propose a novel Spatially Localized Channel Attention (SLCA) mechanism to compute spatially localized attention weights for the feature maps. The features extracted from SAM's image encoder are processed through SLCA to compute attention weights, which are then integrated into deep learning classification models to enhance their focus on spatially relevant or meaningful regions of the image, thus improving classification performance. Experimental results on three public medical image classification datasets demonstrate the effectiveness and dataefficiency of our approach. Pengfei Gu, Haoteng Tang, Islam Akef Ebeid, Jose Angel Nuñez, Fabian Vazquez, Diego Adame, Marcus Zhan, Danny Ziyi Chen |
CBMS | 1 |
| 2025 | TopoImages: Incorporating Local Topology Encoding into Deep Learning Models for Medical Image ClassificationabstractTopological structures in image data, such as connected components and loops, play a crucial role in understanding image content (e.g., biomedical objects). Despite remarkable successes of numerous image processing methods that rely on appearance information, these methods often lack sensitivity to topological structures when used in general deep learning (DL) frameworks. In this paper, we introduce a new general approach, called TopoImages (for Topology Images), which computes a new representation of input images by encoding local topology of patches. In TopoImages, we leverage persistent homology (PH) to encode geometric and topological features inherent in image patches. Our main objective is to capture topological information in local patches of an input image into a vectorized form. Specifically, we first compute persistence diagrams (PDs) of the patches, and then vectorize and arrange these PDs into long vectors for pixels of the patches. The resulting multi-channel image-form representation is called a TopoImage. TopoImages offers a new perspective for data analysis. To garner diverse and significant topological features in image data and ensure a more comprehensive and enriched representation, we further generate multiple TopoImages of the input image using various filtration functions, which we call multi-view TopoImages. The multi-view TopoImages are fused with the input image for DL-based classification, with considerable improvement. Our TopoImages approach is highly versatile and can be seamlessly integrated into common DL frameworks. Experiments on three public medical image classification datasets demonstrate noticeably improved accuracy over state-of-the-art methods. Pengfei Gu, Yejia Zhang, Chaoli Wang 0001, Danny Ziyi Chen |
ACM Multimedia | 1 |
| 2025 | Sli2Vol+: Segmenting 3D Medical Images Based on an Object Estimation Guided Correspondence Flow NetworkabstractDeep learning (DL) methods have shown remarkable successes in medical image segmentation, often using large amounts of annotated data for model training. However, acquiring a large number of diverse labeled 3D medical image datasets is highly difficult and expensive. Recently, mask propagation DL methods were developed to reduce the annotation burden on 3D medical images. For example, Sli2Vol [59] proposed a self-supervised framework (SSF) to learn correspondences by matching neighboring slices via slice reconstruction in the training stage; the learned correspondences were then used to propagate a labeled slice to other slices in the test stage. But, these methods are still prone to error accumulation due to the inter-slice propagation of reconstruction errors. Also, they do not handle discontinuities well, which can occur between consecutive slices in 3D images, as they emphasize exploiting object continuity. To address these challenges, in this work, we propose a new SSF, called Sli2Vol+, for segmenting any anatomical structures in 3D medical images using only a single annotated slice per training and testing volume. Specifically, in the training stage, we first propagate an annotated 2D slice of a training volume to the other slices, generating pseudo-labels (PLs). Then, we develop a novel Object Estimation Guided Correspondence Flow Network to learn reliable correspondences between consecutive slices and corresponding PLs in a self-supervised manner. In the test stage, such correspondences are utilized to propagate a single annotated slice to the other slices of a test volume. We demonstrate the effectiveness of our method on various medical image segmentation tasks with different datasets, showing better generalizability across different organs, modalities, and modals. Code is available at https://github.com/adlsn/Sli2Volplus Delin An, Pengfei Gu, Milan Sonka, Chaoli Wang 0001, Danny Ziyi Chen |
WACV | 2 |
| 2024 | Gradient Reactivation Enhanced Causal Attention for Out-Of-Distribution Generalizable Graph ClassificationabstractSeeking for generalizable graph representations becomes hot spot in the area of graph learning. Recently, causality theory has been applied for extracting the causal relations between graph data and labels, which are generalizable under distribution shift and result in better OOD generalization. In this paper, for more accurately capturing causal representation of graph data, we propose a gradient reactivation enhanced causal subgraph extraction method. The proposed model utilizes attention mechanism to extract the causal features and attenuates the confounding effect of shortcut features. For ensuring stability of extracted causal features, we propose a novel gradient reactivation method to filter features with greater effect on making prediction. Extensively experimental result proves the effectiveness of the proposed model. Xu Wang 0029, Pengfei Gu, Yudong Zhang 0005, Binwu Wang, Pengkun Wang 0001, Yang Wang 0015 |
ICASSP | 2 |
| 2024 | IHCSurv: Effective Immunohistochemistry Priors for Cancer Survival Analysis in Gigapixel Multi-stain Whole Slide Images
Yejia Zhang, Hanqing Chao, Zhongwei Qiu, Nishchal Sapkota, Pengfei Gu, Danny Ziyi Chen, Le Lu 0001, Ke Yan 0006, Dakai Jin, Yun Bian |
MICCAI (4) | 7 |
| 2024 | Combining Segment Anything Model with Domain-Specific Knowledge for Semi-Supervised Learning in Medical Image Segmentation
Yizhe Zhang 0001, Tao Zhou 0002, Ye Wu 0001, Pengfei Gu, Shuo Wang 0011 |
PRCV (14) | 4 |
| 2024 | FCNR: Fast Compressive Neural Representation of Visualization ImagesabstractWe present FCNR, a fast compressive neural representation for tens of thousands of visualization images under varying viewpoints and timesteps. The existing NeRVI solution, albeit enjoying a high compression ratio, incurs slow speeds in encoding and decoding. Built on the recent advances in stereo image compression, FCNR assimilates stereo context modules and joint context transfer modules to compress image pairs. Our solution significantly improves encoding and decoding speed while maintaining high reconstruction quality and satisfying compression ratio. To demonstrate its effectiveness, we compare FCNR with state-of-the-art neural compression methods, including E-NeRV, HNeRV, NeRVI, and ECSIC. The source code can be found at https://github.com/YunfeiLu0112/FCNR. Yunfei Lu, Pengfei Gu, Chaoli Wang 0001 |
IEEE VIS | 2 |
| 2024 | TestFit: A plug-and-play one-pass test time method for medical image segmentation
Yizhe Zhang 0001, Tao Zhou 0002, Yuhui Tao, Shuo Wang 0011, Ye Wu 0001, Benyuan Liu, Pengfei Gu, Qiang Chen 0004, Danny Ziyi Chen |
Medical Image Anal. | 7 |
| 2024 | RVIO: An Effective Localization Algorithm for Range-Aided Visual-Inertial Odometry SystemabstractThis paper presents an efficient and accurate range-aided visual-inertial odometry (RVIO) system for the global positioning system denied environment. In particular, the ultra-wideband (UWB) measurements are integrated to reduce the long-term drift of the visual-inertial odometry (VIO) system. Our approach starts with a filter-based scheme to localize the unknown UWB anchor in the local world frame. In particular, a novel surface-based particle filter is proposed to localize the UWB anchors efficiently. When the initialization is complete, the UWB location information is utilized to support the subsequent long-term robot positioning. An observability-constrained optimization approach is developed to combine the visual, inertial, and UWB range measurements. Such a framework takes advantage of both VIO and UWB measurements and is feasible even when the number of observed UWB anchors is below four. Experiments on both simulated and real-world scenes demonstrate the validity and superiority of the proposed system. Jun Wang 0067, Pengfei Gu, Lei Wang 0162, Ziyang Meng 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Searching Lottery Tickets in Graph Neural Networks: A Dual Perspective
Kun Wang 0056, Yuxuan Liang 0002, Pengkun Wang 0001, Xu Wang 0029, Pengfei Gu, Junfeng Fang, Yang Wang 0015 |
ICLR | 5 |
| 2023 | SwIPE: Efficient and Robust Medical Image Segmentation with Implicit Patch Embeddings
Yejia Zhang, Pengfei Gu, Nishchal Sapkota, Danny Ziyi Chen |
MICCAI (5) | 2 |
| 2023 | NeRVI: Compressive neural representation of visualization images for communicating volume visualization results
Pengfei Gu, Danny Ziyi Chen, Chaoli Wang 0001 |
Comput. Graph. | 1 |
| 2022 | Scalar2Vec: Translating Scalar Fields to Vector Fields via Deep LearningabstractWe introduce Scalar2Vec, a new deep learning solution that translates scalar fields to velocity vector fields for scientific visualization. Given multivariate or ensemble scalar field volumes and their velocity vector field counterparts, Scalar2Vec first identifies suitable variables for scalar-to-vector translation. It then leverages a k-complete bipartite translation network (kCBT-Net) to complete the translation task. kCBT-Net takes a set of sampled scalar volumes of the same variable as input, extracts their multi -scale information, and learns to synthesize the corresponding vector volumes. Ground-truth vector fields and their derived quantities are utilized for loss computation and network training. After training, Scalar2Vec can infer unseen velocity vector fields of the same data set directly from their scalar field counterparts. We demonstrate the effectiveness of Scalar2Vec with quantitative and qualitative results on multiple data sets and compare it with three other state-of-the-art deep learning methods. Pengfei Gu, Jun Han 0010, Danny Ziyi Chen, Chaoli Wang 0001 |
PacificVis | 1 |
| 2022 | Keep Your Friends Close & Enemies Farther: Debiasing Contrastive Learning with Spatial Priors in 3D Radiology ImagesabstractUnderstanding of spatial attributes is central to effective 3D radiology image analysis where crop-based learning is the de facto standard. Given an image patch, its core spatial properties (e.g., position & orientation) provide helpful priors on expected object sizes, appearances, and structures through inherent anatomical consistencies. Spatial correspondences, in particular, can effectively gauge semantic similarities between inter-image regions, while their approximate extraction requires no annotations or overbearing computational costs. However, recent 3D contrastive learning approaches either neglect correspondences or fail to maximally capitalize on them. To this end, we propose an extensible 3D contrastive framework (Spade, for Spa tial De biasing) that leverages extracted correspondences to select more effective positive & negative samples for representation learning. Our method learns both globally invariant and locally equivariant representations with downstream segmentation in mind. We also propose separate selection strategies for global & local scopes that tailor to their respective representational requirements. Compared to recent state-of-the-art approaches, Spade shows notable improvements on three downstream segmentation tasks (CT Abdominal Organ, CT Heart, MR Heart). Yejia Zhang, Nishchal Sapkota, Pengfei Gu, Yaopeng Peng, Hao Zheng 0006, Danny Ziyi Chen |
BIBM | 3 |
| 2022 | Real-Time Visual Inertial Odometry with a Resource-Efficient Harris Corner Detection Accelerator on FPGA PlatformabstractVisual Inertial Odometry (VIO) is a widely studied localization technique in robotics. State-of-the-art VIO algorithms are composed of two parts: a frontend which performs visual perception and inertial measurement pre-processing, and a backend which fuses vision and inertial measurements to estimate the robot's pose. Both image processing in the frontend and sensor fusion in the backend are computationally expensive, making it very challenging to run the VIO algorithm, especially the optimization-based VIO algorithm in real time on embedded platforms with limited power budget. In this paper, a real-time optimization-based monocular VIO algorithm is proposed based on algorithm-and-hardware co-design and successfully implemented on an embedded platform with only 2.6W processor power consumption. In particular, the time-consuming Harris corner detection (HCD) is accelerated on Field Programmable Gate Array (FPGA), achieving an average 16 × processing time reduction compared with the ARM implementation. Compared with the state-of-the-art HCD accelerator provided by Xilinx, the hardware resource required of our accelerator is largely reduced without any compromise in speed, thanks to the proposed dedicated pruning and paral-lelization techniques. Finally, experiment on the public dataset demonstrates that the proposed real-time VIO algorithm on the FPGA-based platform has comparable accuracy with respect to the existing state-of-the-art VIO algorithm on the desktop, and 3 × faster frontend processing speed over the ARM-based implementation. Pengfei Gu, Ziyang Meng 0001, Pengkun Zhou |
IROS | 1 |
| 2021 | kCBAC-Net: Deeply Supervised Complete Bipartite Networks with Asymmetric Convolutions for Medical Image Segmentation
Pengfei Gu, Hao Zheng 0006, Yizhe Zhang 0001, Chaoli Wang 0001, Danny Ziyi Chen |
MICCAI (1) | 1 |
| 2021 | Approximate set union via approximate randomization
Pengfei Gu |
Theor. Comput. Sci. | 2 |
| 2020 | Polyhedral Circuits and Their Applications
Pengfei Gu |
AAIM | 2 |
| 2020 | InTracker: An Integrated Detector-Tracker Framework for Cell Detection and TrackingabstractAutomatic tracking of moving cells in time-lapse image sequences plays an important role in studying many biological processes in development and diseases. Large variations in cell appearances, limited image resolution, and various cell behaviors (e.g., division, apoptosis, deformation, clustering, and migration in or out of the imaging window) make cell tracking a challenging task. However, known cell tracking methods were designed for and tailored to specific cell image sequences and behaviors, thus having limited applicability to various cell image sequences. Aiming toward more robust cell tracking, we propose a new detector-tracker approach for detection and association based cell tracking. First, we propose a new deep learning based detector to detect cells in each image frame and assign division/non-division labels to them. Second, we carefully design an Earth Mover's Distance (EMD) based hierarchical tracker to associate detected cells through the image sequence and form moving cell trajectories. The tracker is able to correct possible detection errors made by the detector. Evaluated on several open challenge datasets, our approach outperforms state-of-the-art cell tracking methods for determining cell trajectories. Peixian Liang, Jianxu Chen 0001, Yizhe Zhang 0001, Hao Zheng 0006, Pengfei Gu, Danny Ziyi Chen |
CBMS | 6 |
| 2020 | Approximate Set Union via Approximate Randomization
Pengfei Gu |
COCOON | 2 |