Shuai Gong

dblp:88/7525 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 AVM: Towards Structure-Preserving Neural Response Modeling in the Visual Cortex Across Stimuli and Individuals
abstract
While deep learning models have shown strong performance in simulating neural responses, they often fail to clearly separate stable visual encoding from condition-specific adaptation, which limits their ability to generalize across stimuli and individuals. We introduce the Adaptive Visual Model (AVM), a structure-preserving framework that enables condition-aware adaptation through modular subnetworks, without modifying the core representation. AVM keeps a Vision Transformer-based encoder frozen to capture consistent visual features, while independently trained modulation paths account for neural response variations driven by stimulus content and subject identity. We evaluate AVM in three experimental settings, including stimulus-level variation, cross-subject generalization, and cross-dataset adaptation, all of which involve structured changes in inputs and individuals. Across two large-scale mouse V1 datasets, AVM outperforms the state-of-the-art V1T model by approximately 2% in predictive correlation, demonstrating robust generalization, interpretable condition-wise modulation, and high architectural efficiency. Specifically, AVM achieves a 9.1% improvement in explained variance (FEVE) under the cross-dataset adaptation setting. These results suggest that AVM provides a unified framework for adaptive neural modeling across biological and experimental conditions, offering a scalable solution under structural constraints. Its design may inform future approaches to cortical modeling in both neuroscience and biologically inspired AI systems.
Qi Xu 0008, Shuai Gong, Xuming Ran, Haihua Luo, Yangfan Hu
AAAI2
2026 Semantic structure fusion graph for abstractive dialogue summarization
Furui Wang, Zhenfang Zhu, Qiang Lu 0006, Shuai Gong, Hongli Pei, Zhenrui Fu, Dawei Zhao 0001
Neurocomputing4
2026 Collaborative Model and Data Adaptation at Test Time
Chunyun Zhang, Fujun Yang, Chaoran Cui, Shuai Gong, Wenna Wang, Xue Lin 0003, Yonggang Qi, Lei Zhu 0002
IEEE Trans. Circuits Syst. Video Technol.4
2026 Federated Domain Generalization via Prompt Learning and Aggregation
abstract
Federated domain generalization (FedDG) aims to improve the global model’s generalization ability in unseen domains by addressing data heterogeneity under privacy-preserving constraints. A common strategy in existing FedDG studies involves sharing domain-specific knowledge among clients, such as spectrum information, class prototypes, and data styles. However, this knowledge is extracted directly from local client samples, and sharing such sensitive information poses a potential risk of data leakage, which might not fully meet the FedDG requirements. In this paper, we introduce prompt learning to adapt pretrained vision-language models (VLMs) in the FedDG scenario, and leverage locally learned prompts as a more secure bridge to facilitate knowledge transfer among clients. Specifically, we propose a novel FedDG framework through Prompt Learning and AggregatioN (PLAN), which comprises two training stages to collaboratively generate local prompts and global prompts at each federated round. First, each client performs both text and visual prompt learning using their own data, with local prompts indirectly synchronized by regarding the global prompts as a common reference. Second, all domain-specific local prompts are exchanged among clients and selectively aggregated into global prompts using lightweight attention-based aggregators. The global prompts are finally applied to adapt the VLMs to unseen target domains. As our PLAN framework requires training only a limited number of prompts and lightweight aggregators, it offers notable advantages in terms of computational and communication efficiency for FedDG. Extensive experiments demonstrate the superior generalization ability of PLAN across four benchmark datasets. We have released our code at https://github.com/GongShuai8210/PLAN.
Shuai Gong, Chaoran Cui, Chunyun Zhang, Wenna Wang, Xiushan Nie, Lei Zhu 0002
IEEE Trans. Inf. Forensics Secur.1
2026 Token-Level Prompt Mixture With Parameter-Free Routing for Federated Domain Generalization
abstract
Federated Domain Generalization (FedDG) aims to train a globally generalizable model on data from decentralized, heterogeneous clients. While recent work has adapted vision-language models for FedDG using prompt learning, the prevailing "one-prompt-fits-all" paradigm struggles with sample diversity, causing a marked performance decline on personalized samples. The Mixture of Experts (MoE) architecture offers a promising solution for specialization. However, existing MoE-based prompt learning methods suffer from two key limitations: coarse image-level expert assignment and high communication costs from parameterized routers. To address these limitations, we propose TRIP, a Token-level pRompt mIxture with Parameter-free routing framework for FedDG. TRIP treats prompts as multiple experts, and assigns individual tokens within an image to distinct experts, facilitating the capture of fine-grained visual patterns. To ensure communication efficiency, TRIP introduces a parameter-free routing mechanism based on capacity-aware clustering and Optimal Transport (OT). First, tokens are grouped into capacity-aware clusters to ensure balanced workloads. These clusters are then assigned to experts via OT, stabilized by mapping cluster centroids to static, non-learnable keys. The final instance-specific prompt is synthesized by aggregating experts, weighted by the number of tokens assigned to each. Extensive experiments across four benchmarks demonstrate that TRIP achieves optimal generalization results, with communicating as few as 1K parameters. Our code is available at https://github.com/GongShuai8210/TRIP.
Shuai Gong, Chaoran Cui, Xiaolin Dong, Xiushan Nie, Lei Zhu 0002, Xiaojun Chang
IEEE Trans. Image Process.1
2025 Black-Box Test-Time Prompt Tuning for Vision-Language Models
abstract
Test-time prompt tuning (TPT) aims to adjust the vision-language models (e.g., CLIP) with learnable prompts during the inference phase. However, previous works overlooked that pre-trained models as a service (MaaS) have become a noticeable trend due to their commercial usage and potential risk of misuse. In the context of MaaS, users can only design prompts in inputs and query the black-box vision-language models through inference APIs, rendering the previous paradigm of utilizing gradient for prompt tuning is infeasible. In this paper, we propose black-box test-time prompt tuning (B²TPT), a novel framework that addresses the challenge of optimizing prompts without gradients in an unsupervised manner. Specifically, B²TPT designs a consistent or confident (CoC) pseudo-labeling strategy to generate high-quality pseudo-labels from the outputs. Subsequently, we propose to optimize low-dimensional intrinsic prompts using a derivative-free evolution algorithm and to project them onto the original text and vision prompts. This strategy addresses the gradient-free challenge while reducing complexity. Extensive experiments across 15 datasets demonstrate the superiority of B²TPT. The results show that B²TPT not only outperforms CLIP's zero-shot inference at test time, but also surpasses other gradient-based TPT methods.
Fan'an Meng, Chaoran Cui, Hongjun Dai, Shuai Gong
AAAI4
2025 Adversarial Topic-Aware Prompt-Tuning for Cross-Topic Automated Essay Scoring
abstract
Cross-topic automated essay scoring (AES) aims to develop a transferable model capable of effectively evaluating essays on a target topic. A significant challenge in this domain arises from the inherent discrepancies between topics. While existing methods predominantly focus on extracting topic-shared features through distribution alignment of source and target topics, they often neglect topic-specific features, limiting their ability to assess critical traits such as topic adherence. To address this limitation, we propose an Adversarial TOpic-aware Prompt-tuning (ATOP), a novel method that jointly learns topic-shared and topic-specific features to improve cross-topic AES. ATOP achieves this by optimizing a learnable topic-aware prompt—comprising both shared and specific components—to elicit relevant knowledge from pre-trained language models (PLMs). To enhance the robustness of topic-shared prompt learning and mitigate feature scale sensitivity introduced by topic alignment, we incorporate adversarial training within a unified regression and classification framework. In addition, we employ a neighbor-based classifier to model the local structure of essay representations and generate pseudo-labels for target-topic essays. These pseudo-labels are then used to guide the supervised learning of topic-specific prompts tailored to the target topic. Extensive experiments on the publicly available ASAP++ dataset demonstrate that ATOP significantly outperforms existing state-of-the-art methods in both holistic and multi-trait essay scoring. The implementation of our method is publicly available at: https://github.com/zhaohy777/ATOP.
Chunyun Zhang, Chaoran Cui, Qilong Song, Zhiqing Lu, Shuai Gong, Kailin Liu
ECAI6
2025 Dynamic prompt allocation and tuning for continual test-time adaptation
Chaoran Cui, Yongrui Zhen, Shuai Gong, Chunyun Zhang, Hui Liu 0016, Yilong Yin
Sci. China Inf. Sci.3
2025 Consistency-guided Multi-Source-Free Domain Adaptation
Chaoran Cui, Chunyun Zhang, Fan'an Meng, Shuai Gong, Muzhi Xi, Lei Li 0008
Eng. Appl. Artif. Intell.5
2025 When Adversarial Training Meets Prompt Tuning: Adversarial Dual Prompt Tuning for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) aims to adapt models learned from a well-annotated source domain to a target domain, where only unlabeled samples are available. To this end, adversarial training is widely used in conventional UDA methods to reduce the discrepancy between source and target domains. Recently, prompt tuning has emerged as an efficient way to adapt large pre-trained vision-language models like CLIP to a variety of downstream tasks. In this paper, we present a novel method named Adversarial DuAl Prompt Tuning (ADAPT) for UDA, which employs text prompts and visual prompts to guide CLIP simultaneously. Rather than simply performing a joint optimization of text prompts and visual prompts, we integrate text prompt tuning and visual prompt tuning into a collaborative framework where they engage in an adversarial game: text prompt tuning focuses on distinguishing between source and target images, whereas visual prompt tuning seeks to align source and target domains. Unlike most existing adversarial training-based UDA approaches, ADAPT does not require explicit domain discriminators for domain alignment. Instead, the objective is effectively achieved at both global and category levels through modeling the joint probability distribution of images on domains and categories. Extensive experiments on four benchmark datasets demonstrate the effectiveness of our ADAPT method for UDA. We have released our code at https://github.com/Liuziyi1999/ADAPT.
Chaoran Cui, Shuai Gong, Lei Zhu 0002, Chunyun Zhang, Hui Liu 0016
IEEE Trans. Image Process.3
2024 Geometric Feature Extraction of Ship Target Based on Multi-Resolution Sar Images
abstract
As one of the key performance indicators of SAR sensors, resolution, which is determined by the SAR signal bandwidth and plays a very important role in the interpretation of SAR images. The transformed images of high-resolution SAR images at different resolutions contain different layers of information, which can provide complementary resolvability for ship target extraction. In this paper, experiments and analyses are conducted on the effects of multi-resolution SAR images on the extraction accuracy of four parameters: length, width, principal axis angle and geometric center of ship targets. It is concluded that there is an influence on the geometric feature parameters of the ship targets of different resolution SAR images, and for a certain type of typical targets, there exists a critical resolution, which makes the extraction of the relevant parameters with the highest accuracy and the most stable performance.
Shuai Gong, Bing Sun 0002, Jingwen Li 0003, Yunfei Xi
IGARSS1
2024 Clutter Space-Time Distribution and Suppression of GEO-SBR
abstract
With the advantages of short revisit time and wide coverage, geosynchronous orbit spaceborne radar (GEO-SBR) is expected to be utilized for early warning tasks. However, GEO-SBR also encounters challenges, such as the "stop-and-go" assumption not being applicable and the significant impact of Earth’s rotation. This paper introduces the geometry under the "non-stop-and-go" assumption, analyzes the clutter space-time distribution considering the Earth’s rotation effects, and proposes a clutter suppression method based on attitude steering. Simulation results represents the clutter space-time distribution, and validates the effectiveness of the clutter suppression method.
Yunfei Xi, Bing Sun 0002, Shuai Gong
IGARSS4
2024 Accelerating Domain Adaptation with Cascaded Adaptive Vision Transformer
Qilin Jiang, Chaoran Cui, Chunyun Zhang, Yongrui Zhen, Shuai Gong, Fan'an Meng
PRCV (1)5
2024 DSAMR: Dual-Stream Attention Multi-hop Reasoning for knowledge-based visual question answering
Yanhan Sun, Zhenfang Zhu, Zicheng Zuo, Kefeng Li 0003, Shuai Gong, Jiangtao Qi
Expert Syst. Appl.5
2024 HTPosum:Heterogeneous Tree Structure augmented with Triplet Positions for extractive Summarization of scientific papers
Zhenfang Zhu, Shuai Gong, Jiangtao Qi, Chunling Tong
Expert Syst. Appl.2
2024 Adversarial Source Generation for Source-Free Domain Adaptation
abstract
Unsupervised domain adaptation aims to transfer the knowledge learned from a labeled source domain to an unlabeled target domain with different data distributions. However, in practice, source samples are not always available due to privacy protection and storage resource limitations. To address this concern, Source-Free Domain Adaptation (SFDA) has recently attracted growing research attention, as it only needs a pre-trained source model without direct access to source data. In this paper, we propose a novel Adversarial SOurce GEneration (ASOGE) method for SFDA, which introduces an additional generative module to produce synthetic labeled source samples and uses them to facilitate cross-domain adaptation. Unlike early studies that train the generator independently and perform the adaptation only after the generator is finished, ASOGE integrates the generation and adaptation stages within a collaborative framework by making them play an adversarial game. In the generation stage, the labeled source samples are not produced blindly; instead, they are hard-to-align samples that provide knowledge more worth learning for the adaptation stage. To achieve a fine-grained domain alignment, a class-aware discrepancy between source and target domains is measured via contrastive learning. Extensive experiments on benchmark datasets demonstrate the effectiveness of ASOGE compared to the state-of-the-art methods.
Chaoran Cui, Fan'an Meng, Chunyun Zhang, Lei Zhu 0002, Shuai Gong, Xue Lin 0003
IEEE Trans. Circuits Syst. Video Technol.6
2024 Gtpsum: guided tensor product framework for abstractive summarization
Jingan Lu, Zhenfang Zhu, Kefeng Li 0003, Shuai Gong, Hongli Pei, Wenling Wang
J. Supercomput.4
2023 Improving End-to-End Modeling For Mandarin-English Code-Switching Using Lightweight Switch-Routing Mixture-of-Experts
Fengyun Tan, Chaofeng Feng, Tao Wei 0003, Shuai Gong, Jinqiang Leng, Jun Ma 0018, Jing Xiao 0006
INTERSPEECH4
2023 SeburSum: a novel set-based summary ranking strategy for summary-level extractive summarization
Shuai Gong, Zhenfang Zhu, Jiangtao Qi, Wenqing Wu 0002, Chunling Tong
J. Supercomput.1
2022 SatSOT: A Benchmark Dataset for Satellite Video Single Object Tracking
abstract
By imaging a specific area continuously, satellite video shows excellent capability in various applications such as surveillance and traffic management. Although object tracking has made significant progress in recent years, development in satellite object tracking is limited by the lack of open-source satellite datasets. It is thus essential to establish a satellite video object-tracking benchmark to fill the gap and advance the research. In this work, we present SatSOT, the first densely annotated satellite video single object-tracking benchmark dataset. SatSOT consists of 105 sequences with 27664 frames, 11 attributes, and four categories of typical moving targets in satellite videos: car, plane, ship, and train. Based on the proposed dataset and the significant challenges in satellite video object tracking, such as small targets, background interference, and severe occlusion, detailed evaluation and analysis are performed on 15 among the best and most representative tracking algorithms, which provides a basis for further research on satellite video object tracking.
Manqi Zhao, Shengyang Li, Shiyu Xuan, Longxuan Kou, Shuai Gong
IEEE Trans. Geosci. Remote. Sens.5
2019 Image quality guided biology application for genetic analysis
Shuai Gong, Mingjiu Luo
J. Vis. Commun. Image Represent.2
2014 Discovering Diversity Corrections for Incompatible Web Services
abstract
The increasing amount of web services over the Internet enable users composing them to satisfy the users'needs efficiently. Such service composing is prone to errors. Automatically detecting incompatible web services interaction and correcting them will largely improve users' experience on service composing. When correcting the errors, two major issues need to be addressed: First, how to satisfy diverse correction requirements of different users, Second, how to find the corrections efficiently. This paper proposes an approach to discovering maximum diversity corrections to reduce the risk of unsatisfying different end users' needs when presenting correction plans to them. To solve the problem efficiently, this paper proposes an approximate algorithm to find diverse correction plans. Furthermore, two pruning strategies are adopted to reduce the runtime of the algorithm. Experiments illustrate that our approach outperforms the baseline on the diversity of correction plans, and the two pruning strategies reduce the runtime significantly.
Shuai Gong, Jinhua Xiong, Zhiyong Liu 0002, Manfred Wojciechowski
ICWS1
2013 Identifying Semantic-Related Search Tasks in Query Log
Shuai Gong, Jinhua Xiong, Zhiyong Liu 0002
APWeb1
2012 Continuous Query for QoS-Aware Automatic Service Composition
abstract
Current QoS-aware automatic service composition queries over a network of Web services are often one-time innature. After a network of Web services is built, such queries are issued once, and answers are found from the scratch. The underlying assumption is that the participating Web services are rather static so that their functional and non-functional parameters seldom change. However, such an assumption is often baseless. New services come and go, service APIs change gradually, and QoS values fluctuate. Therefore, a support for efficiently handling "continuous" service composition queries is desired. In this paper, we propose an event driven continuous query algorithm for QoS-aware automatic service composition problem to cope with different types of dynamic services. Moreover, we integrated this algorithm in our service composition system, QSynth. Finally, we evaluate our proposal using both real QoS data and synthetic Web service data and show the superior performance of ours, compared to the state-of-the art solution which won the performance championship of Web Service Challenge in 2009 and 2010.
Wei Jiang 0028, Songlin Hu 0001, Dongwon Lee 0001, Shuai Gong, Zhiyong Liu 0002
ICWS4