Haoteng Tang

dblp:244/8016 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0003-0323-1755ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Integrating Multi-scale and Multi-filtration Topological Features for Medical Image Classification*
abstract
Modern deep neural networks have shown remarkable performance in medical image classification. However, such networks either emphasize pixel-intensity features instead of fundamental anatomical structures (e.g., those encoded by topological invariants), or they capture only simple topological features via single-parameter persistence. In this paper, we propose a new topology-guided classification framework that extracts multi-scale and multi-filtration persistent topological features and integrates them into vision classification backbones. For an input image, we first compute cubical persistence diagrams (PDs) across multiple image resolutions/scales. We then develop a "vineyard" algorithm that consolidates these PDs into a single, stable diagram capturing signatures at varying granularities, from global anatomy to subtle local irregularities that may indicate early-stage disease. To further exploit richer topological representations produced by multiple filtrations, we design a cross-attention-based neural network that directly processes the consolidated final PDs. The resulting topological embeddings are fused with feature maps from CNNs or Transformers. By integrating multi-scale and multi-filtration topologies into an end-to-end architecture, our approach enhances the model’s capacity to recognize complex anatomical structures. Evaluations on three public datasets show consistent, considerable improvements over strong baselines and state-of-the-art methods, demonstrating the value of our comprehensive topological perspective for robust and interpretable medical image classification.
Pengfei Gu, Haoteng Tang, Dongkuan Xu, Erik Enriquez, DongChul Kim, Danny Ziyi Chen
WACV3
2025 Topo-VM-UNetV2: Encoding Topology Into Vision Mamba UNet for Polyp Segmentation
abstract
Convolutional neural network (CNN) and Transformer-based architectures are two dominant deep learning models for polyp segmentation. However, CNNs have limited capability for modeling long-range dependencies, while Transformers incur quadratic computational complexity. Recently, State Space Models such as Mamba have been recognized as a promising approach for polyp segmentation because they not only model long-range interactions effectively but also maintain linear computational complexity. However, Mamba-based architectures still struggle to capture topological features (e.g., connected components, loops, voids), leading to inaccurate boundary delineation and polyp segmentation. To address these limitations, we propose a new approach called Topo-VM-UNetV2, which encodes topological features into the Mamba-based state-of-the-art polyp segmentation model, VM-UNetV2. Our method consists of two stages: Stage 1: VM-UNetV2 is used to generate probability maps (PMs) for the training and test images, which are then used to compute topology attention maps. Specifically, we first compute persistence diagrams of the PMs, then we generate persistence score maps by assigning persistence values (i.e., the difference between death and birth times) of each topological feature to its birth location, finally we transform persistence scores into attention weights using the sigmoid function. Stage 2: These topology attention maps are integrated into the semantics and detail infusion (SDI) module of VM-UNetV2 to form a topologyguided semantics and detail infusion (Topo-SDI) module for enhancing the segmentation results. Extensive experiments on five public polyp segmentation datasets demonstrate the effectiveness of our proposed method. The code will be made publicly available.
Diego Adame, Jose Angel Nuñez, Fabian Vazquez, Nayeli Gurrola, Haoteng Tang, Pengfei Gu
CBMS6
2025 Adapting a Segmentation Foundation Model for Medical Image Classification
abstract
Recent advancements in foundation models, such as the Segment Anything Model (SAM), have shown strong performance in various vision tasks, particularly image segmentation, due to their impressive zero-shot segmentation capabilities. However, effectively adapting such models for medical image classification is still a less explored topic. In this paper, we introduce a new framework to adapt SAM for medical image classification. First, we utilize the SAM image encoder as a feature extractor to capture segmentation-based features that convey important spatial and contextual details of the image, while freezing its weights to avoid unnecessary overhead during training. Next, we propose a novel Spatially Localized Channel Attention (SLCA) mechanism to compute spatially localized attention weights for the feature maps. The features extracted from SAM's image encoder are processed through SLCA to compute attention weights, which are then integrated into deep learning classification models to enhance their focus on spatially relevant or meaningful regions of the image, thus improving classification performance. Experimental results on three public medical image classification datasets demonstrate the effectiveness and dataefficiency of our approach.
Pengfei Gu, Haoteng Tang, Islam Akef Ebeid, Jose Angel Nuñez, Fabian Vazquez, Diego Adame, Marcus Zhan, Danny Ziyi Chen
CBMS2
2025 BPEN: Brain Posterior Evidential Network for trustworthy brain imaging analysis
Kai Ye 0002, Haoteng Tang, Siyuan Dai, Igor Fortel, Paul M. Thompson, Scott Mackin, Alex D. Leow, Heng Huang 0001, Liang Zhan
Neural Networks2
2024 Interpretable Spatio-Temporal Embedding for Brain Structural-Effective Network with Ordinary Differential Equation
Haoteng Tang, Siyuan Dai, Kai Ye 0002, Kun Zhao 0007, Wenlu Wang, Carl Yang 0001, Lifang He 0001, Alex D. Leow, Paul M. Thompson, Heng Huang 0001, Liang Zhan
MICCAI (2)1
2024 Contrastive Brain Network Learning via Hierarchical Signed Graph Pooling Model
abstract
Recently, brain networks have been widely adopted to study brain dynamics, brain development, and brain diseases. Graph representation learning techniques on brain functional networks can facilitate the discovery of novel biomarkers for clinical phenotypes and neurodegenerative diseases. However, current graph learning techniques have several issues on brain network mining. First, most current graph learning models are designed for unsigned graph, which hinders the analysis of many signed network data (e.g., brain functional networks). Meanwhile, the insufficiency of brain network data limits the model performance on clinical phenotypes' predictions. Moreover, few of the current graph learning models are interpretable, which may not be capable of providing biological insights for model outcomes. Here, we propose an interpretable hierarchical signed graph representation learning (HSGPL) model to extract graph-level representations from brain functional networks, which can be used for different prediction tasks. To further improve the model performance, we also propose a new strategy to augment functional brain network data for contrastive learning. We evaluate this framework on different classification and regression tasks using data from human connectome project (HCP) and open access series of imaging studies (OASIS). Our results from extensive experiments demonstrate the superiority of the proposed model compared with several state-of-the-art techniques. In addition, we use graph saliency maps, derived from these prediction tasks, to demonstrate detection and interpretation of phenotypic biomarkers.
Haoteng Tang, Guixiang Ma, Lei Guo 0028, Xiyao Fu, Heng Huang 0001, Liang Zhan
IEEE Trans. Neural Networks Learn. Syst.1
2023 Bidirectional Mapping with Contrastive Learning on Multimodal Neuroimaging Data
Kai Ye 0002, Haoteng Tang, Siyuan Dai, Lei Guo 0028, Johnny Yuehan Liu, Yalin Wang 0001, Alex D. Leow, Paul M. Thompson, Heng Huang 0001, Liang Zhan
MICCAI (3)2
2023 Signed graph representation learning for functional-to-structural brain network mapping
Haoteng Tang, Lei Guo 0028, Xiyao Fu, Yalin Wang 0001, Scott Mackin, Olusola Ajilore, Alex D. Leow, Paul M. Thompson, Heng Huang 0001, Liang Zhan
Medical Image Anal.1
2021 CommPOOL: An interpretable graph pooling framework for hierarchical graph representation learning
Haoteng Tang, Guixiang Ma, Lifang He 0001, Heng Huang 0001, Liang Zhan
Neural Networks1
2020 Vulnerability vs. Reliability: Disentangled Adversarial Examples for Cross-Modal Learning
abstract
The vulnerability of deep neural networks has gained a great upsurge of research attention, which engages well-designed examples through adding little perturbations to fool a well-performed network. Meanwhile, a progress has been made in leveraging adversarial examples to boost the robustness of deep cross-modal networks. However, for cross-modal learning, both the causes of adversarial examples and their latent advantages in learning cross-modal correlations are under-explored. In this paper, we propose novel Disentangled Adversarial examples for Cross-Modal learning, dubbed DACM. Specifically, we first divide cross-modal data into two aspects, namely modality-related component and modality-unrelated counterpart, and then learn to improve the reliability of network using the modality-related component. To achieve this goal, we apply the generation of adversarial perturbations to strengthen cross-modal correlations, wherein the modality-related component is acquired through gradually detaching the modality-unrelated component. Finally, the proposed DACM is employed to create modality-related examples towards the application of cross-modal hashing retrieval. Extensive experiments carried out on two cross-modal benchmarks show that the adversarial examples learned by DACM are efficient at fooling a target deep cross-modal hashing network. On the other hand, training this target model by merely leveraging our created modality-related examples in turn significantly promotes the robustness of this model itself.
Chao Li 0033, Haoteng Tang, Cheng Deng 0002, Liang Zhan, Wei Liu 0005
KDD2