VLDB 2026 Research / reviewers in the wild / expert
Hao Cheng 0016
dblp:09/5158-16
· DBLP profile ↗
12ranked-venue papers
6as first author
11since 2021 · last 2025
0000-0003-4823-0908ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive Decision Boundary for Few-Shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) aims to continuously learn new classes from a limited set of training samples without forgetting knowledge of previously learned classes. Conventional FSCIL methods typically build a robust feature extractor during the base training session with abundant training samples and subsequently freeze this extractor, only fine-tuning the classifier in subsequent incremental phases. However, current strategies primarily focus on preventing catastrophic forgetting, considering only the relationship between novel and base classes, without paying attention to the specific decision spaces of each class. To address this challenge, we propose a plug-and-play Adaptive Decision Boundary Strategy (ADBS), which is compatible with most FSCIL methods. Specifically, we assign a specific decision boundary to each class and adaptively adjust these boundaries during training to optimally refine the decision spaces for the classes in each session. Furthermore, to amplify the distinctiveness between classes, we employ a novel inter-class constraint loss that optimizes the decision boundaries and prototypes for each class. Extensive experiments on three benchmarks, namely CIFAR100, miniImageNet, and CUB200, demonstrate that incorporating our ADBS method with existing FSCIL techniques significantly improves performance, achieving overall state-of-the-art results. Linhao Li, Yongzhang Tan, Siyuan Yang 0001, Hao Cheng 0016, Yongfeng Dong |
AAAI | 4 |
| 2025 | Reconciling Stochastic and Deterministic Strategies for Zero-shot Image Restoration using Diffusion Model in DualabstractPlug-and-play (PnP) methods offer an iterative strategy for solving image restoration (IR) problems in a zero-shot manner, using a learned discriminative denoiser as the implicit prior. More recently, a sampling-based variant of this approach, which utilizes a pre-trained generative diffusion model, has gained great popularity for solving IR problems through stochastic sampling. The IR results using PnP with a pre-trained diffusion model demonstrate distinct advantages compared to those using discriminative denoisers, i.e.,improved perceptual quality while sacrificing the data fidelity. The unsatisfactory results are due to the lack of integration of these strategies in the IR tasks. In this work, we propose a novel zero-shot IR scheme, dubbed Reconciling Diffusion Model in Dual (RDMD), which leverages only a single pre-trained diffusion model to construct two complementary regularizers. Specifically, the diffusion model in RDMD will iteratively perform deterministic denoising and stochastic sampling, aiming to achieve highfidelity image restoration with appealing perceptual quality. RDMD also allows users to customize the distortion-perception tradeoff with a single hyperparameter, enhancing the adaptability of the restoration process in different practical scenarios. Extensive experiments on several IR tasks demonstrate that our proposed method could achieve superior results compared to existing approaches on both the FFHQ and ImageNet datasets. Code is available at https://github.com/chongwang1024/rdmd. Chong Wang 0011, Lanqing Guo, Zixuan Fu, Siyuan Yang 0001, Hao Cheng 0016, Alex Chichung Kot, Bihan Wen |
CVPR | 5 |
| 2025 | Weakly Supervised Bilinear Convolutional Neural Network for Fine-Grained Vehicle ClassificationabstractFine-grained vehicle classification, which is a key technology within intelligent transportation systems, has been gaining increasing importance with the burgeoning growing number of vehicles. Previous studies have predominantly focused on intricate and distinctive local features. However, in various tasks, it has been proven that global features are of significant importance when they can be effectively integrated with local features in a harmonious manner. So, we consider that a comprehensive consideration of both local and global features is crucial for enhancing classification decisions. Consequently, the paper designs a novel architecture for the task, which combines global and local features to improve classification performance. The architecture consists of two components: the local-feature net and the global-feature net. Specially, for the local feature, we propose an Essential Part Locator module that uses global feature-weighted attention masks to obtain local features, and a Cross-Part Feature Transformer that boosts interactions between local features. Meanwhile, our architecture processes the entire image through an encoder to capture global features and then integrates both global and local features. Experimental results on the Stanford Cars, CompCars, and BoxCars116K datasets demonstrate that the proposed approach surpasses state-of-the-art methods, achieving accuracies of 97.5%, 96.4%, and 92.1%, respectively. Linhao Li, Han Zang, Xiaojuan Fan, Hao Cheng 0016, Yongfeng Dong |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Progressive Divide-and-Conquer via Subsampling Decomposition for Accelerated MRIabstractDeep unfolding networks (DUN) have emerged as a pop-ular iterative framework for accelerated magnetic reso-nance imaging (MRI) reconstruction. However, conventional DUN aims to reconstruct all the missing information within the entire null space in each iteration. Thus it could be challenging when dealing with highly ill-posed degradation, often resulting in subpar reconstruction. In this work, we propose a Progressive Divide-And-Conquer (PDAC) strategy, aiming to break down the subsampling process in the actual severe degradation and thus per-form reconstruction sequentially. Starting from decomposing the original maximum-a-posteriori problem of accel-erated MRI, we present a rigorous derivation of the pro-posed PDAC framework, which could be further unfolded into an end-to-end trainable network. Each PDAC iter-ation specifically targets a distinct segment of moderate degradation, based on the decomposition. Furthermore, as part of the PDAC iteration, such decomposition is adaptively learned as an auxiliary task through a degradation predictor which provides an estimation of the decomposed sampling mask. Following this prediction, the sampling mask is further integrated via a severity conditioning mod-ule to ensure awareness of the degradation severity at each stage. Extensive experiments demonstrate that our pro-posed method achieves superior performance on the pub-licly available fastMRI and Stanford2D FSE datasets in both multi-coil and single-coil settings. Code is available at https://github.com/ChongWang1024/PDAC. Chong Wang 0011, Lanqing Guo, Yufei Wang 0006, Hao Cheng 0016, Yi Yu 0011, Bihan Wen |
CVPR | 4 |
| 2024 | STSP: Spatial-Temporal Subspace Projection for Video Class-Incremental Learning
Hao Cheng 0016, Siyuan Yang 0001, Chong Wang 0011, Joey Tianyi Zhou, Alex Chichung Kot, Bihan Wen |
ECCV (28) | 1 |
| 2024 | Learning-Based Human Detection via Radar for Dynamic and Cluttered Indoor EnvironmentsabstractRadar-based human detection draws significant attention in response to growing safety concerns driven by advances in factory automation and smart home technologies. However, much of this research typically operates in controlled environments characterized by minimal clutter and noise, which limits their effectiveness in real-world scenarios such as urban areas, factories, and smart homes. In this study, we address this limitation by collecting a real-world radar dataset in dynamic and cluttered indoor environments, spanning five distinctive environments. To simulate non-human targets, we introduce a moving trolley. Subsequently, we propose a system including simple Radar Signal Processing steps and a learning-based model using unsupervised domain adaptation to enhance its adaptability to unseen environments. Through a series of comprehensive experiments employing popular learning-based methods on our dataset, we demonstrate the model’s efficacy in mitigating environmental interference and successfully adapting to previously unseen environments. Songnan Lin, Hao Cheng 0016, Weixian Liu, Bihan Wen |
ISCAS | 3 |
| 2024 | Disentangled Feature Representation for Few-Shot Image ClassificationabstractLearning the generalizable feature representation is critical to few-shot image classification. While recent works exploited task-specific feature embedding using meta-tasks for few-shot learning, they are limited in many challenging tasks as being distracted by the excursive features such as the background, domain, and style of the image samples. In this work, we propose a novel disentangled feature representation (DFR) framework, dubbed DFR, for few-shot learning applications. DFR can adaptively decouple the discriminative features that are modeled by the classification branch, from the class-irrelevant component of the variation branch. In general, most of the popular deep few-shot learning methods can be plugged in as the classification branch, thus DFR can boost their performance on various few-shot tasks. Furthermore, we propose a novel FS-DomainNet dataset based on DomainNet, for benchmarking the few-shot domain generalization (DG) tasks. We conducted extensive experiments to evaluate the proposed DFR on general, fine-grained, and cross-domain few-shot classification, as well as few-shot DG, using the corresponding four benchmarks, i.e., mini-ImageNet, tiered-ImageNet, Caltech-UCSD Birds 200-2011 (CUB), and the proposed FS-DomainNet. Thanks to the effective feature disentangling, the DFR-based few-shot classifiers achieved state-of-the-art results on all datasets. Hao Cheng 0016, Yufei Wang 0006, Haoliang Li, Alex Chichung Kot, Bihan Wen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | ShadowFormer: Global Context Helps Shadow RemovalabstractRecent deep learning methods have achieved promising results in image shadow removal. However, most of the existing approaches focus on working locally within shadow and non-shadow regions, resulting in severe artifacts around the shadow boundaries as well as inconsistent illumination between shadow and non-shadow regions. It is still challenging for the deep shadow removal model to exploit the global contextual correlation between shadow and non-shadow regions. In this work, we first propose a Retinex-based shadow model, from which we derive a novel transformer-based network, dubbed ShandowFormer, to exploit non-shadow regions to help shadow region restoration. A multi-scale channel attention framework is employed to hierarchically capture the global information. Based on that, we propose a Shadow-Interaction Module (SIM) with Shadow-Interaction Attention (SIA) in the bottleneck stage to effectively model the context correlation between shadow and non-shadow regions. We conduct extensive experiments on three popular public datasets, including ISTD, ISTD+, and SRD, to evaluate the proposed method. Our method achieves state-of-the-art performance by using up to 150X fewer model parameters. Lanqing Guo, Siyu Huang, Ding Liu 0001, Hao Cheng 0016, Bihan Wen |
AAAI | 4 |
| 2023 | Frequency Guidance Matters in Few-Shot LearningabstractFew-shot classification aims to learn a discriminative feature representation to recognize unseen classes with few labeled support samples. While most few-shot learning methods focus on exploiting the spatial information of image samples, frequency representation has also been proven essential in classification tasks. In this paper, we investigate the effect of different frequency components on the few-shot learning tasks. To enhance the performance and generalizability of few-shot methods, we propose a novel Frequency-Guided Few-shot Learning framework (dubbed FGFL), which leverages the task-specific frequency components to adaptively mask the corresponding image information, with a novel multi-level metric learning strategy including a triplet loss among original, masked and unmasked image as well as a contrastive loss between masked and original support and query sets to exploit more discriminative information. Extensive experiments on four benchmarks under several few-shot scenarios, i.e., standard, cross-dataset, cross-domain, and coarse-to-fine annotated classification, are conducted. Both qualitative and quantitative results show that our proposed FGFL scheme can attend to the class-discriminative frequency components, thus integrating those information towards more effective and generalizable few-shot learning. Hao Cheng 0016, Siyuan Yang 0001, Joey Tianyi Zhou, Lanqing Guo, Bihan Wen |
ICCV | 1 |
| 2023 | Reconciliation of statistical and spatial sparsity for robust visual classification
Hao Cheng 0016, Kim-Hui Yap, Bihan Wen |
Neurocomputing | 1 |
| 2023 | Graph Neural Networks With Triple Attention for Few-Shot LearningabstractRecent advances in Graph Neural Networks (GNNs) have achieved superior results in many challenging tasks, such as few-shot learning. Despite its capacity to learn and generalize a model from only a few annotated samples, GNN is limited in scalability, as deep GNN models usually suffer from severe over-fitting and over-smoothing. In this work, we propose a novel GNN framework with atriple-attention mechanism,i.e.node self-attention, neighbor attention, and layer memory attention, to tackle these challenges. We provide both theoretical analysis and illustrations to explain why the proposed attentive modules can improve GNN scalability for few-shot learning tasks. Our experiments show that the proposed Attentive GNN model outperforms the state-of-the-art few-shot learning methods using both GNN and non-GNN approaches. The improvement is consistent over the mini-ImageNet, tiered-ImageNet, CUB-200-2011, and Flowers-102 benchmarks, using both ConvNet-4 and ResNet-12 backbones, and under both the inductive and transductive settings. Furthermore, we demonstrate the superiority of our method for few-shot fine-grained and semi-supervised classification tasks with extensive experiments. The code for this work is publicly available athttps://github.com/chenghao-ch94/AGNN. Hao Cheng 0016, Joey Tianyi Zhou, Wee-Peng Tay, Bihan Wen |
IEEE Trans. Multim. | 1 |
| 2020 | Joint Statistical and Spatial Sparse Representation for Robust Image and Image-Set ClassificationabstractRecent image classification schemes, by learning deep features from large-scale dataset, have achieved the significantly better results comparing to classic feature-based approaches. However, there are still challenges in practice, such as classifying noisy image-set queries and training over limited-scale dataset. Instead of applying generic deep features, the model-based approaches can be more effective for robust image and image-set classification tasks, as we need various image priors to exploit the inter- and intra-set data variations while prevent over-fitting. In this work, we propose a novel joint statistical and spatial sparse representation, dubbed J3S, to model the image or image-set data, by exploiting both their local patch structures and global Gaussian distribution into Riemannian manifold. To the best of our knowledge, no work to date utilized both global statistics and local patch structures jointly via sparse representation. We propose to solve a co-regularized sparse coding problem based on the J3S model, by coupling the local and global representations using joint sparsity. The learned J3S models are used for robust image and image-set classification. Experiments show that the proposed J3S-based image classification scheme outperforms the popular or state-of-the-art competing methods. Hao Cheng 0016, Bihan Wen |
ICIP | 1 |