Xiong Gao

dblp:36/10140 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 APENet: Task-aware adaptation prototype evolution network for few-shot semantic segmentation
Zhaobin Chang, Xiong Gao, Dongliang Chang, Yande Li, Yonggang Lu
Expert Syst. Appl.2
2025 MIP-CLIP: Multimodal Independent Prompt CLIP for Action Recognition
abstract
Recently, the Contrastive Language Image Pre-training (CLIP) model has shown significant generalizability by optimizing the distance between visual and text features. The mainstream CLIP-based action recognition methods mitigate the low “zero-shot” generalization of the 1-of-N paradigm but also lead to a significant degradation in supervised performance. Therefore, powerful supervision and competitive “zero-shot” need to be effectively traded off. In this work, a Multimodal Independent Prompt CLIP (MIP-CLIP) model is proposed to address this challenge. On the visual side, we propose novel Video Motion Prompt (VMP) to empower the visual encoder with motion perception, which performs short- and long-term motion modelling via temporal difference operation. Next, the visual classification branch is introduced to improve the discrimination of visual features. Specifically, the temporal difference and visual classification operations of the 1-of-N paradigm are extended to CLIP to satisfy the need for strong supervised performance. On the text side, we design Class-Agnostic text prompt Template (CAT) under the constraint of Semantic Alignment (SA) module to solve the label semantic dependency problem. Finally, a Dual-branch Feature Reconstruction (DFR) module is proposed to complete cross-modal interactions for better feature matching, which uses the class confidence of the visual classification branch as input. The experiments are conducted on four widely used benchmarks (HMDB-51, UCF-101, Jester, and Kinetics-400). The results demonstrate that our method achieves excellent supervised performance while preserving competitive generalizability.
Xiong Gao, Zhaobin Chang, Dongyi Kong, Huiyu Zhou 0001, Yonggang Lu
IEEE Trans. Multim.1
2025 Multi-prototype collaborative perception enhancement network for few-shot semantic segmentation
Zhaobin Chang, Xiong Gao, Dongyi Kong, Yonggang Lu
Vis. Comput.2
2024 CANet: Comprehensive Attention Network for video-based action recognition
Xiong Gao, Zhaobin Chang, Xingcheng Ran, Yonggang Lu
Knowl. Based Syst.1
2024 DRNet: Disentanglement and Recombination Network for Few-Shot Semantic Segmentation
abstract
Few-shot semantic segmentation (FSS) aims to segment novel classes with only a few annotated samples. Existing methods to FSS generally combine the annotated mask and the corresponding support image to generate the class-specific representation, and perform the segmentation for the query image by matching the features of the query image to these representations. However, the segmentation performance could be fragile for the lack of an effective method to handle the inappropriate use of query features and the neglection of correlation between features in support and query images. In this work, we propose a novel Disentanglement and Recombination Network (DRNet) to alleviate this problem. Concretely, we first apply the self-attention on both support foreground features and query foreground features. Then, the foreground features of the support and query branches are recombined using the cross-attention after self-attention computation, which can encourage the foreground feature alignment between branches. Finally, the prototypes are generated from the recombined foreground features and support background features, and are utilized to guide the segmentation for given images. Considering the sensitivity of prototypes related to the subtle differences among objects from different classes and the same class, we further introduce a joint learning strategy to derive accurate segmentation of both seen and unseen objects in the support image and the query image respectively. Extensive experiments on the PASCAL-5iand COCO-20idatasets demonstrate the superiority of our DRNet comparing with the recent popular methods. The code is released on https://github.com/GS-Chang-Hn/DRNet-fss.
Zhaobin Chang, Xiong Gao, Huiyu Zhou 0001, Yonggang Lu
IEEE Trans. Circuits Syst. Video Technol.2
2023 Simple yet effective joint guidance learning for few-shot semantic segmentation
Zhaobin Chang, Yonggang Lu, Xingcheng Ran, Xiong Gao
Appl. Intell.4
2023 Few-shot semantic segmentation: a review on recent approaches
Zhaobin Chang, Yonggang Lu, Xingcheng Ran, Xiong Gao
Neural Comput. Appl.4
2022 Accelerating spiking neural networks using quantum algorithm with high success probability and high calculation accuracy
Cen Wang, Hongxiang Guo, Xiong Gao, Jian Wu 0010
Neurocomputing4
2022 Fully automatic image segmentation based on FCN and graph cuts
Zhaobin Wang, Xiong Gao, Runliang Wu, Jianfang Kang, Yaonan Zhang
Multim. Syst.2
2022 Local feature fusion and SRC-based decision fusion for ear recognition
Zhaobin Wang, Xiong Gao, Qizhen Yan, Yaonan Zhang
Multim. Syst.2
2022 An Edge Storage Acceleration Service for Collaborative Mobile Devices
abstract
Fueled by the advances in the Internet of Things, and the growing capacity of smart mobile devices at the edge of the Internet, we have witnessed a growing trend in research and development for edge computing and edge storage, which extends the abilities of single mobile device on the edge through on-demand collaboration among multiple geographically distributed mobile devices. In this article, we address several technical challenges that are unique for collaborative storage at the edge due to the unique characteristics of mobile devices. First, we formalize the collaborative storage problem as an optimization problem. Second, we design an Acceleration Algorithm for Collaborative Storage, called A2CS, based on the architecture of Alternating Direction Method of Multipliers (ADMM). Specifically, we use the Nesterov’s Acceleration strategy and the step size rules in the process of updating variables and determining the optimal speed of convergence. We develop a novel collaborative storage policy in order to guide the whole lifecycle of collaborative storage. Finally, we conduct a series of experiments for acceleration performance analysis and validation. We show that A2CS delivers a better convergence performance with different step size rules, compared with two existing approaches: the ADMM baseline and the ADMM-OR (ADMM with Over-Relaxation), achieving the acceleration percentage by at least 25.33 percent and at most 64.01 percent. In addition, by conducting the utility performance comparison analysis with the existing Average Distribution Strategy (ADS) and the existing Distance Preferred Distribution Strategy (DPDS), we show the advantage of A2CS over both ADS and DPDS with respect to the total utility and energy consumption.
Xiong Gao, Weidong Bao 0001, Xiaomin Zhu 0001, Guanlin Wu, Ling Liu 0001
IEEE Trans. Serv. Comput.1
2021 AKG: automatic kernel generation for neural processing units using polyhedral transformations
abstract
Existing tensor compilers have proven their effectiveness in deploying deep neural networks on general-purpose hardware like CPU and GPU, but optimizing for neural processing units (NPUs) is still challenging due to the heterogeneous compute units and complicated memory hierarchy.
Jie Zhao 0002, Bojie Li, Wang Nie, Zhen Geng, Renwei Zhang, Xiong Gao, Zheng Li 0035, Peng Di, Xuefeng Jin 0004
PLDI6
2016 CSE: Conceptual Sentence Embeddings based on Attention Model
abstract
Most sentence embedding models typically represent each sentence only using word surface, which makes these models indiscriminative for ubiquitous homonymy and polysemy.In order to enhance representation capability of sentence, we employ conceptualization model to assign associated concepts for each sentence in the text corpus, and then learn conceptual sentence embedding (CSE).Hence, this semantic representation is more expressive than some widely-used text representation models such as latent topic model, especially for short-text.Moreover, we further extend CSE models by utilizing a local attention-based model that select relevant words within the context to make more efficient prediction.In the experiments, we evaluate the CSE models on two tasks, text classification and information retrieval.The experimental results show that the proposed models outperform typical sentence embed-ding models.
Yashen Wang, Heyan Huang, Chong Feng 0001, Jiahui Gu, Xiong Gao
ACL (1)6