Baoming Zhang

dblp:79/7838 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AdaInertia: Dynamic Tool Usage Inertia Control via Deep Q-Learning for Efficient LLM Agent Reasoning
Yanxi Hu, Guangqi Dong, Baoming Zhang, Mingyu Di, Chong-Jun Wang
ICIC (2)5
2026 HALF: A homophily-aware loss fusion for robust learning under label noise in heterophilic graphs
Shuangjie Li, Baoming Zhang, Meng Cao 0004, Jianqing Song, Chong-Jun Wang
Inf. Sci.2
2026 TGSL: Trade-off graph structure learning via multifaceted graph information bottleneck
abstract
Graph neural networks (GNNs) are prominent for their effectiveness in processing graph-structured data for semi-supervised node classification tasks. Most existing GNNs perform message passing directly based on the observed graph structure. However, in real-world scenarios, the observed structure is often suboptimal due to multiple factors, significantly degrading the performance of GNNs. To address this challenge, we first conduct an empirical analysis showing that different graph structures significantly impact empirical risk and classification performance. Motivated by our observations, we propose a novel method named Trade-off Graph Structure Learning (TGSL), guided by the multifaceted Graph Information Bottleneck (GIB) principle based on Mutual Information (MI). The key idea behind TGSL is to learn a minimal sufficient graph structure that minimizes empirical risk while maintaining performance. Specifically, we introduce global feature augmentation to capture the structural roles of nodes, and global structure augmentation to uncover global relationships between nodes. The augmented graphs are then processed by structure estimators with different parameters for refinement and redefinition, respectively. Additionally, we innovatively leverage multifaceted GIB as the optimization objective by maximizing the MI between the labels and the representation derived from the final structure, while constraining the MI between this representation and that based on the redefined structures. This trade-off helps avoid capturing irrelevant information from the redefined structures and enhances the final representation for node classification. We conduct extensive experiments across a range of datasets under clean and attacked conditions. The results demonstrate the outstanding performance and robustness of TGSL over state-of-the-art baselines.
Shuangjie Li, Baoming Zhang, Jianqing Song, Gaoli Ruan, Chong-Jun Wang, Junyuan Xie
Neural Networks2
2025 Robust Logit Adjustment for Learning with Long-Tailed Noisy Data
abstract
Learning with noisy labels (LNL) methods have enabled the deployment of machine learning systems with imperfectly labeled data. However, these methods often struggle to identify noise in the presence of long-tailed (LT) class distributions, where the memorization effect becomes class-dependent. Conversely, LT methods are suboptimal under label noise, as it hinders access to accurate label frequency statistics. This study aims to address the long-tailed noisy data by bridging the methodological gap between LNL and LT approaches. We propose a direct solution, termed Robust Logit Adjustment, which estimates ground-truth labels through label refurbishment, thereby mitigating the impact of label noise. Simultaneously, our method incorporates the distribution of training-time corrected target labels into the LT method logit adjustment, providing class-rebalanced supervision. Extensive experiments on both synthetic and real-world long-tailed noisy datasets demonstrate the superior performance of our method.
Mingcai Chen, Yuntao Du 0001, Baoming Zhang, Yi Xin 0003, Chong-Jun Wang
AAAI4
2025 PALM: Pushing Adaptive Learning Rate Mechanisms for Continual Test-Time Adaptation
abstract
Real-world vision models in dynamic environments face rapid shifts in domain distributions, leading to decreased recognition performance. Using unlabeled test data, continuous test-time adaptation (CTTA) directly adjusts a pre-trained source discriminative model to these changing domains. A highly effective CTTA method involves applying layer-wise adaptive learning rates for selectively adapting pre-trained layers. However, it suffers from the poor estimation of domain shift and the inaccuracies arising from the pseudo-labels. This work aims to overcome these limitations by identifying layers for adaptation via quantifying model prediction uncertainty without relying on pseudo-labels. We utilize the magnitude of gradients as a metric, calculated by backpropagating the KL divergence between the softmax output and a uniform distribution, to select layers for further adaptation. Subsequently, for the parameters exclusively belonging to these selected layers, with the remaining ones frozen, we evaluate their sensitivity to approximate the domain shift and adjust their learning rates accordingly. We conduct extensive image classification experiments on CIFAR-10C, CIFAR-100C, and ImageNet-C, demonstrating the superior efficacy of our method compared to prior approaches.
Sarthak Kumar Maharana, Baoming Zhang, Yunhui Guo
AAAI2
2025 Normalize Then Propagate: Efficient Homophilous Regularization for Few-Shot Semi-Supervised Node Classification
abstract
Graph Neural Networks (GNNs) have demonstrated remarkable ability in semi-supervised node classification. However, most existing GNNs rely heavily on a large amount of labeled data for training, which is labor-intensive and requires extensive domain knowledge. In this paper, we first analyze the restrictions of GNNs generalization from the perspective of supervision signals in the context of few-shot semi-supervised node classification. To address these challenges, we propose a novel algorithm named NormProp, which utilizes the homophily assumption of unlabeled nodes to generate additional supervision signals, thereby enhancing the generalization against label scarcity. The key idea is to efficiently capture both the class information and the consistency of aggregation during message passing, via decoupling the direction and Euclidean norm of node representations. Moreover, we conduct a theoretical analysis to determine the upper bound of Euclidean norm, and then propose homophilous regularization to constraint the consistency of unlabeled nodes. Extensive experiments demonstrate that NormProp achieve state-of-the-art performance under low-label rate scenarios with low computational complexity.
Baoming Zhang, Mingcai Chen, Jianqing Song, Shuangjie Li, Jie Zhang 0152, Chong-Jun Wang
AAAI1
2025 BATCLIP: Bimodal Online Test-Time Adaptation for CLIP
Sarthak Kumar Maharana, Baoming Zhang, Leonid Karlinsky, Rogério Feris, Yunhui Guo
ICCV2
2025 Test-Time Selective Adaptation for Uni-Modal Distribution Shift in Multi-Modal Data
abstract
Modern machine learning applications are characterized by the increasing size of deep models and the growing diversity of data modalities. This trend underscores the importance of efficiently adapting pre-trained multi-modal models to the test distribution in real time, i.e., multi-modal test-time adaptation. In practice, the magnitudes of multi-modal shifts vary because multiple data sources interact with the impact factor in diverse manners. In this research, we investigate the the under-explored practical scenario uni-modal distribution shift, where the distribution shift influences only one modality, leaving the others unchanged. Through theoretical and empirical analyses, we demonstrate that the presence of such shift impedes multi-modal fusion and leads to the negative transfer phenomenon in existing test-time adaptation techniques. To flexibly combat this unique shift, we propose a selective adaptation schema that incorporates multiple modality-specific adapters to accommodate potential shifts and a “router” module that determines which modality requires adaptation. Finally, we validate the effectiveness of our proposed method through extensive experimental evaluations. Code available at https://github.com/chenmc1996/Uni-Modal-Distribution-Shift.
Mingcai Chen, Baoming Zhang, Zongbo Han, Yanmeng Wang, Yuntao Du 0001, Bing-Kun Bao
ICML2
2025 <tt>AVROBUSTBENCH</tt>: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time
Sarthak Kumar Maharana, Saksham Singh Kushwaha, Baoming Zhang, Adrian Rodriguez, Songtao Wei, Yapeng Tian, Yunhui Guo
NeurIPS3
2025 LaplaceConfidence: A graph-based approach for learning with noisy labels
abstract
Real-world machine learning applications seldom provide perfect labeled data, posing a challenge in developing models robust to noisy labels. Recent methods prioritize noise filtering based on the discrepancies between model predictions and the provided noisy labels, assuming samples with minimal classification losses to be clean. In this work, we capitalize on the consistency between the learned model and the complete noisy dataset, employing the data’s rich representational and topological information. We introduce LaplaceConfidence, a method that to obtain label confidence (i.e., clean probabilities) utilizing the Laplacian energy. Specifically, it first constructs graphs based on the feature representations of all noisy samples and minimizes the Laplacian energy to produce a low-energy graph. Clean labels should fit well into the low-energy graph while noisy ones should not, allowing our method to determine data’s clean probabilities. Furthermore, LaplaceConfidence is embedded into a holistic method for robust training, where co-training technique generates unbiased label confidence and label refurbishment technique better utilizes it. We also explore the dimensionality reduction technique to accommodate our method on large-scale noisy datasets. Our experiments demonstrate that LaplaceConfidence outperforms state-of-the-art methods on benchmark datasets under both synthetic and real-world noise. Code available at https://github.com/chenmc1996/LaplaceConfidence .
Mingcai Chen, Yuntao Du 0001, Baoming Zhang, Chong-Jun Wang
Intell. Data Anal.4
2025 Graph Neural Networks with Coarse- and Fine-Grained Division for mitigating label noise and sparsity
Shuangjie Li, Baoming Zhang, Jianqing Song, Gaoli Ruan, Chong-Jun Wang, Junyuan Xie
Neural Networks2
2024 Seeking Similarities While Removing Differences: Graph Neural Networks Based on Node Correlation
abstract
Graph neural networks (GNNs) have proven highly effective in handling graph-structured data. However, most existing GNNs rely on the homophily assumption, hindering their performance on heterophilic graphs. This limitation is partially due to aggregation containing irrelevant nodes. In this work, we propose a novel GNN model based on node correlation called NoC-GNN, to address the deficiencies of existing techniques. NoC-GNN retains relevant nodes while removing irrelevant nodes, enhancing the effectiveness of nodes in the mixed state. NoC-GNN first constructs a new graph structure based on the k-nearest neighbor (kNN) graph to aggregate relevant nodes, and then constructs a matrix based on the new graph structure to remove possibly irrelevant nodes. Finally, the attention mechanism is used to adaptively integrate relevant, irrelevant, and self-information to model both homophilic and heterophilic graphs. Experimental results demonstrate that NoC-GNN achieves superior performance across a wide range of semi-supervised node classification tasks.
Shuangjie Li, Baoming Zhang, Jianqing Song, Junyuan Xie, Chong-Jun Wang
ICASSP2
2024 Similarity-Navigated Conformal Prediction for Graph Neural Networks
abstract
Graph Neural Networks have achieved remarkable accuracy in semi-supervised node classification tasks. However, these results lack reliable uncertainty estimates. Conformal prediction methods provide a theoretical guarantee for node classification tasks, ensuring that the conformal prediction set contains the ground-truth label with a desired probability (e.g., 95\%). In this paper, we empirically show that for each node, aggregating the non-conformity scores of nodes with the same label can improve the efficiency of conformal prediction sets while maintaining valid marginal coverage. This observation motivates us to propose a novel algorithm named $\textit{Similarity-Navigated Adaptive Prediction Sets}$ (SNAPS), which aggregates the non-conformity scores based on feature similarity and structural neighborhood. The key idea behind SNAPS is that nodes with high feature similarity or direct connections tend to have the same label. By incorporating adaptive similar nodes information, SNAPS can generate compact prediction sets and increase the singleton hit ratio (correct prediction sets of size one). Moreover, we theoretically provide a finite-sample coverage guarantee of SNAPS. Extensive experiments demonstrate the superiority of SNAPS, improving the efficiency of prediction sets and singleton hit ratio while maintaining valid coverage.
Jianqing Song, Jianguo Huang, Baoming Zhang, Shuangjie Li, Chong-Jun Wang
NeurIPS4
2023 FairHELP: Fairness-Aware Heterogeneous Information Network Embedding for Link Prediction
Meng Cao 0004, Jianqing Song, Jinliang Yuan, Baoming Zhang, Chong-Jun Wang
DASFAA (3)4
2023 Self-supervised short text classification with heterogeneous graph neural networks
abstract
Abstract Short text classification has been a fundamental task in natural language processing, which benefits various applications, such as sentiment analysis, news tagging, and intent recommendation. However, classifying short texts is challenging due to the information sparsity in the text corpus. Besides, the performance of existing machine learning classification models largely relies on sufficient training data, yet labels can be scarce and expensive to obtain in real‐world text classification scenarios. In this article, we propose a novel self‐supervised short text classification method. Specifically, we first model the short text corpus as a heterogeneous graph to address the information sparsity problem. Then, we introduce a self‐attention‐based heterogeneous graph neural network model to learn short text embeddings. In addition, we adopt a self‐supervised learning framework to exploit internal and external similarities among short texts. Experiments on five real‐world short text benchmarks validate the effectiveness of our proposed method compared with the state‐of‐the‐art methods.
Meng Cao 0004, Jinliang Yuan, Hualei Yu, Baoming Zhang, Chong-Jun Wang
Expert Syst. J. Knowl. Eng.4
2022 LSEGNN: Encode Local Topology Structure in Graph Neural Networks
abstract
Learning robust representations for nodes in graphs is crucial for graph learning tasks. Graph Neural Networks(GNNs) attract much attention recently as the frameworks achieve great success in node representation learning. Existing state-of-the-art GNN methods (like GCN) aggregate messages from neighbor nodes through message passing neural network to update representations for nodes. However, the message passing strategy fails to capture the structural similarity between nodes. Besides, it assumes that neighbor nodes are independent and ignores abundant local neighbor structures around nodes in real networks. This weakness may hurt the performance of GNNs in some classification tasks. To capture the overlooked information, in the experimental investigation, we found that co-occurrence probabilities based on random walks can preserve local neighbor structures among nodes well. Furthermore, we propose a novel but effective method to encode local structure information into node features by co-occurrence probabilities. We call this method Local Structure Enhanced Graph Neural Network, short as LSEGNN. Extensive experiments are conducted in benchmark datasets and the results show the effectiveness of our method.
Ming Xu 0014, Baoming Zhang, Meng Cao 0004, Hualei Yu, Chong-Jun Wang
IPCCC2
2017 SRTM DEM-Aided Mapping Satellite-1 Image Geopositioning Without Ground Control Points
abstract
A Shuttle Radar Topography Mission (SRTM) digital elevation model (DEM)-aided geopositioning method is proposed to solve the problem of geopositioning without ground control points for Mapping Satellite-1 imagery. The method comprises coarse and accurate correction stages, and it compensates errors gradually. DEM extraction and DEM matching are important steps in both the stages, the objectives of which are to compensate the relative and absolute errors in an image, respectively. The SRTM DEM is integrated into all the processes to take full advantage of its consistent and high accuracy. Experimental results showed that this method could greatly improve geometry accuracy and obtain stable and highly accurate geopositioning for Mapping Satellite-1 images, regardless of the land area proportion (LAP) or the production mode. The planimetric and vertical accuracies were better than 8.1 and 5.2 m, respectively, which could satisfy the accuracy requirements of mapping at 1:50 000 scale. The computational efficiency depends on the LAP and target DEM resolution.
Xiaowei Chen 0015, Baoming Zhang, Minyi Cen, Haitao Guo, Tonggang Zhang
IEEE Geosci. Remote. Sens. Lett.2