VLDB 2026 Research / reviewers in the wild / expert
Huixia Ben
dblp:310/1536
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0001-7946-8199ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Controllable Generation via Hybrid-grained CacheabstractControllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation computational requirements, resulting in generally low generation efficiency. To address this issue, we propose a Hybrid-Grained Cache (HGC) approach that reduces computational overhead by adopting cache strategies with different granularities at different computational stages. Specifically, (1) we use a coarse-grained cache (block-level) based on feature reuse to dynamically bypass redundant computations in encoder-decoder blocks between each step of model reasoning. (2) We design a fine-grained cache (prompt-level) that acts within a module, where the fine-grained cache reuses cross-attention maps within consecutive reasoning steps and extends them to the corresponding module computations of adjacent steps. These caches of different granularities can be seamlessly integrated into each computational link of the controllable generation process. We verify the effectiveness of HGC on four benchmark datasets, especially its advantages in balancing generation efficiency and visual quality. For example, on the COCO-Stuff segmentation benchmark, our HGC significantly reduces the computational cost (MACs) by 63% (from 18.22T → 6.70T↓), while keeping the loss of semantic fidelity (quantized performance degradation) within 1.5%. Huixia Ben, Shuo Wang 0008, Jinda Lu, Junxiang Qiu, Shengeng Tang, Yanbin Hao |
AAAI | 2 |
| 2025 | Mixture of Multimodal Adapters for Sentiment AnalysisabstractKezhou Chen, Shuo Wang, Huixia Ben, Shengeng Tang, Yanbin Hao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kezhou Chen, Shuo Wang 0008, Huixia Ben, Shengeng Tang, Yanbin Hao |
NAACL (Long Papers) | 3 |
| 2025 | Cross-modal Feature Enhancement and Contrastive Alignment for Micro-gesture Recognition
Tuyun Shang, Yanbin Hao, Ming Pei, Huixia Ben |
PRCV (7) | 5 |
| 2025 | Interventional Feature Generation for Few-shot LearningabstractFew-shot learning (FSL) aims to classify a novel object into a specific category under limited training samples. This is a challenging task since (1) the features expressed by pre-trained knowledge introduce perceived bias and then constrain the classification space, and (2) the use of general hallucination techniques based on global features fails to escape the limited classification space, resulting in sub-optimal improvements. To solve these issues, this article proposes an interventional feature generation (IFG) method. Specifically, we first use the relations of the categories or instances as interventional operations to implicitly constrain the feature representations (pre-trained knowledge) into different classification subsets. Then, we employ a parameter-free feature generation strategy to enrich each subset’s training samples of the support category. In other words, IFG provides a multi-subsets learning strategy to reduce the influence of perceived bias, enrich the diversity of generated features, and improve the robustness of the few-shot classifier. We apply our method to four benchmark datasets and observe state-of-the-art performance across all experiments. Specifically, compared to the baseline on the Mini-ImageNet dataset, our approach yields accuracy improvements of 6.03% and 3.46% for 1 and 5 support training samples, respectively. Furthermore, the proposed interventional feature generation technique can improve classifier performance in other FSL methods, demonstrating its versatility and potential for broader applications. The code is available at https://github.com/ShuoWangCS/IFG-FSL/ . Shuo Wang 0008, Jinda Lu, Huixia Ben, Yanbin Hao, Xingyu Gao 0001, Meng Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Pseudo Content Hallucination for Unpaired Image CaptioningabstractUnpaired Image Captioning (UIC) is designed to describe an image without relying on matched vision-language training data. It is a challenging task since (1) the implicit and unpaired vision-language data nature of the training task limits the captioning model's ability to represent diverse scene representations, and (2) it is difficult for the captioning model to discern the intrinsic relationships among objects, potentially leading to misinterpretation of the image con- tent. To solve these issues, we propose pseudo content hallucination (PCH) to help the captioning model enlarge the perception of the ob- jects and capture the relations between the objects. Specifically, we select similar objects from different images as pseudo content and then hallucinate new visual content for training. This hallucinated content contains a similar scene but with a different representation, thus enriching the diversity of the training samples. Meanwhile, we utilize the relationships among these objects to improve the generated captions as a textual content hallucination and construct pseudo image-sentence pairs to refine the captioning model. These hallucinated sentences are beneficial for the captioning model as they enable the capture of additional semantics from the image, ultimately enhancing the sentence generation ability. Extensive experiments on the two benchmarks, i.e., MSCOCO, and Flickr30k, show the effectiveness of our method. The results show a significant improvement compared to the baseline in the MSCOCO dataset, with 1.5 increase in the CIDEr score. Huixia Ben, Shuo Wang 0008, Meng Wang 0001, Richang Hong |
ICMR | 1 |
| 2023 | Boosting Hyperspectral Image Classification with Dual Hierarchical LearningabstractHyperspectral image (HSI) classification aims at predicting the pixel-wise labels in an image, where there are only a few labeled pixel samples (hard labels) for training. It is a challenging task since the classification process is susceptible to over-fitting under training with limited samples. To relieve this problem, we propose a method based on dual hierarchical learning. First, we employ a connectionist hyperspectral convolution (HC) network to capture the representations of the pixels from different receptive fields. Specifically, an HC is designed to learn the correlation among adjacent pixels and is further extended to a connectionist hierarchical structure. These operations use the correlation to enhance one-pixel learning from multiple receptive fields. Second, we analyze the properties in the hyperspectral image and introduce a hierarchical pseudo label generation algorithm to enrich the supervision of the label information. Finally, we design a dual hierarchical learning strategy to help all HC layers learn from both the hard labels and the hierarchical pseudo labels. In other words, it addresses the HSI classification problem from different views. For inference, we employ two fusion strategies to find a better prediction. The experimental results on four popular HSI benchmarks, i.e., Salinas-A, IndianPines, PaviaU, and PaviaC, demonstrate the effectiveness of the proposed method. Our code is publicly available on GitHub: https://github.com/ShuoWangCS/HSI-DHL. Shuo Wang 0008, Huixia Ben, Yanbin Hao, Xiangnan He 0001, Meng Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Unpaired Image Captioning With semantic-Constrained Self-LearningabstractImage captioning has been an emerging and fast-developing research topic. Nevertheless, most existing works heavily rely on large amounts of image-sentence pairs and therefore hinder the practical applications of captioning in the wild. In this paper, we present a novel Semantic-Constrained Self-learning (SCS) framework that explores an iterative self-learning strategy to learn an image captioner with only unpaired image and text data. Technically, SCS consists of two stages, i.e., pseudo pair generation and captioner re-training, iteratively producing "pseudo" image-sentence pairs via a pre-trained captioner and re-training the captioner with the pseudo pairs, respectively. Particularly, both stages are guided by the recognized objects in the image, that act as semantic constraint to strengthen the semantic alignment between the input image and the output sentence. We leverage a semantic-constrained beam search for pseudo pair generation to regularize the decoding process with the recognized objects via forcing the inclusion/exclusion of the recognized/irrelevant objects in output sentence. For captioner re-training, a self-supervised triplet loss is utilized to preserve the relative semantic similarity ordering among generated sentences with regard to the input image triplets. Moreover, an object inclusion reward and an adversarial reward are adopted to encourage the inclusion of the predicted objects in the output sentence and pursue the generation of more realistic sentences during self-critical training, respectively. Experiments conducted on both dependent and independent unpaired data validate the superiority of SCS. More remarkably, we obtain the best published CIDEr score to-date of 74.7\% on COCO Karpathy test split for unpaired image captioning. Huixia Ben, Yingwei Pan, Yehao Li, Ting Yao 0003, Richang Hong, Meng Wang 0001, Tao Mei 0001 |
IEEE Trans. Multim. | 1 |