Jingsheng Gao

dblp:275/7262 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0001-6271-0903ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DiverSeed: Integrating Active Learning for Target Domain Data Generation in Instruction Tuning
Jingsheng Gao, Mengnan Qi, Suncheng Xiang, Ke Ji, Jiacheng Ruan, Ting Liu 0016, Yuzhuo Fu
Mach. Learn.1
2025 TTE: Two Tokens Are Enough to Improve Parameter-Efficient Tuning
abstract
Existing fine-tuning paradigms are predominantly characterized by Full Parameter Tuning (FPT) and Parameter-Efficient Tuning (PET). FPT fine-tunes all parameters of a pre-trained model on downstream tasks, whereas PET freezes the pre-trained model and employs only a minimal number of learnable parameters for fine-tuning. However, both approaches face issues of overfitting, especially in scenarios where downstream samples are limited. This issue has been thoroughly explored in FPT, but less so in PET. To this end, this paper investigates overfitting in PET, representing a pioneering study in the field. Specifically, across 19 image classification datasets, we employ three classic PET methods (e.g., VPT, Adapter/Adaptformer, and LoRA) and explore various regularization techniques to mitigate overfitting. Regrettably, the results suggest that existing regularization techniques are incompatible with the PET process and may even lead to performance degradation. Consequently, we introduce a new framework named TTE (Two Tokens are Enough), which effectively alleviates overfitting in PET through a novel constraint function based on the learnable tokens. Experiments conducted on 24 datasets across image and few-shot classification tasks demonstrate that our fine-tuning framework not only mitigates overfitting but also significantly enhances PET's performance. Notably, our TTE framework surpasses the highest-performing FPT framework (DR-Tune), utilizing significantly fewer parameters (0.15M vs. 85.84M) and achieving an improvement of 1%.
Jiacheng Ruan, Mingye Xie, Jingsheng Gao, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu
AAAI3
2025 SmartRAG: Jointly Learn RAG-Related Tasks From the Environment Feedback
abstract
RAG systems consist of multiple modules to work together. However, these modules are usually separately trained. We argue that a system like RAG that incorporates multiple modules should be jointly optimized to achieve optimal performance. To demonstrate this, we design a specific pipeline called SmartRAG that includes a policy network and a retriever. The policy network can serve as 1) a decision maker that decides when to retrieve, 2) a query rewriter to generate a query most suited to the retriever and 3) an answer generator that produces the final response with/without the observations. We then propose to jointly optimize the whole system using a reinforcement learning algorithm, with the reward designed to encourage the system to achieve the highest performance with minimal retrieval cost. When jointly optimized, each module can be aware of how other modules are working and thus find the best way to work together as a complete system. Empirical results demonstrate that the jointly optimized system can achieve better performance than separately optimized counterparts.
Jingsheng Gao, Linxu Li, Ke Ji, Weiyuan Li, Yixin Lian, Yuzhuo Fu
ICLR1
2025 MPI-CD: Multi-Path Information Contrastive Decoding for Mitigating Hallucinations in Large Vision-Language Models
abstract
In recent years, despite substantial advancements in large vision-language models (LVLMs), they still encounter the issue of ''hallucinations''-where generated results appears reasonable but often deviates from the visual input or actual facts. In contrast, the human cognitive system, when processing visual input, initially relies on visual perception to distinguish between the salient region and non-salient region, integrating relevant information. Subsequently, it recalls pertinent memory details, ultimately generating a comprehensive cognitive outcome. Inspired by this process, we propose a novel, training-free decoding approach, dubbed as Multi-Path Information Contrastive Decoding (MPI-CD). Specifically, to simulate the human information integration process, we design a three-branch structure called the Tri-Branch Integrator (TBI), which contrasts the original, salient region, and non-salient region images to effectively improve the reliability of the LVLMs' output. Furthermore, to mimic the human memory recall mechanism, we further investigate the importance of hidden layer features and propose the Memory Recall Module (MRM). This module adaptively extracts meaningful memory information from the hidden layers and incorporates it into the decoding process, thereby effectively alleviating the hallucination issue. We conduct extensive experiments on three widely used benchmarks (e.g. POPE, AMBER, and MME) using two classic LVLMs. The experimental results demonstrate that our MPI-CD significantly mitigates hallucinations in LVLMs without requiring additional training.
Jiacheng Ruan, Zongyun Zhang, Jingsheng Gao, Wenzhen Yuan 0002, Ting Liu 0016, Yuzhuo Fu
ACM Multimedia3
2025 Learning multi-axis representation in frequency domain for medical image segmentation
Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Suncheng Xiang
Mach. Learn.2
2025 Learning Visual-Semantic Embedding for Generalizable Person Re-Identification: A Unified Perspective
abstract
Generalizable person Re-Identification (Re-ID) is a very hot research topic in machine learning and computer vision, which plays a significant role in realistic scenarios due to its various applications in public security and video surveillance. However, previous methods mainly focus on the visual representation learning, while neglect to explore the potential of semantic features during training, which easily leads to poor generalization capability when adapted to the new domain. In this article, we present a unified perspective called MMET for more robust visual-semantic embedding learning on generalizable Re-ID. To further enhance the robust feature learning in the context of transformer, a dynamic masking mechanism called Masked Multimodal Modeling (MMM) strategy is introduced to mask both the image patches and the text tokens, which can jointly work on multimodal or unimodal data and significantly boost the performance of generalizable person Re-ID. Extensive experiments on benchmark datasets demonstrate the competitive performance of our method over previous approaches. We hope this method could advance the research towards visual-semantic representation learning. Our source code is also publicly available at https://github.com/JeremyXSC/MMET .
Suncheng Xiang, Jingsheng Gao, Mingye Xie, Mengyuan Guan, Jiacheng Ruan, Yuzhuo Fu
ACM Trans. Multim. Comput. Commun. Appl.2
2024 LAMM: Label Alignment for Multi-Modal Prompt Learning
abstract
With the success of pre-trained visual-language (VL) models such as CLIP in visual representation tasks, transferring pre-trained models to downstream tasks has become a crucial paradigm. Recently, the prompt tuning paradigm, which draws inspiration from natural language processing (NLP), has made significant progress in VL field. However, preceding methods mainly focus on constructing prompt templates for text and visual inputs, neglecting the gap in class label representations between the VL models and downstream tasks. To address this challenge, we introduce an innovative label alignment method named \textbf{LAMM}, which can dynamically adjust the category embeddings of downstream datasets through end-to-end training. Moreover, to achieve a more appropriate label distribution, we propose a hierarchical loss, encompassing the alignment of the parameter space, feature space, and logits space. We conduct experiments on 11 downstream vision datasets and demonstrate that our method significantly improves the performance of existing multi-modal prompt learning models in few-shot scenarios, exhibiting an average accuracy improvement of 2.31(\%) compared to the state-of-the-art methods on 16 shots. Moreover, our methodology exhibits the preeminence in continual learning compared to other prompt tuning methods. Importantly, our method is synergistic with existing prompt tuning methods and can boost the performance on top of them. Our code and dataset will be publicly available at https://github.com/gaojingsheng/LAMM.
Jingsheng Gao, Jiacheng Ruan, Suncheng Xiang, Zefang Yu, Ke Ji, Mingye Xie, Ting Liu 0016, Yuzhuo Fu
AAAI1
2024 From Raw Video to Pedagogical Insights: A Unified Framework for Student Behavior Analysis
abstract
Understanding student behavior in educational settings is critical in improving both the quality of pedagogy and the level of student engagement. While various AI-based models exist for classroom analysis, they tend to specialize in limited tasks and lack generalizability across diverse educational environments. Additionally, these models often fall short in ensuring student privacy and in providing actionable insights accessible to educators. To bridge this gap, we introduce a unified, end-to-end framework by leveraging temporal action detection techniques and advanced large language models for a more nuanced student behavior analysis. Our proposed framework provides an end-to-end pipeline that starts with raw classroom video footage and culminates in the autonomous generation of pedagogical reports. It offers a comprehensive and scalable solution for student behavior analysis. Experimental validation confirms the capability of our framework to accurately identify student behaviors and to produce pedagogically meaningful insights, thereby setting the stage for future AI-assisted educational assessments.
Zefang Yu, Mingye Xie, Jingsheng Gao, Ting Liu 0016, Yuzhuo Fu
AAAI3
2024 Learning to Floorplan like Human Experts via Reinforcement Learning
abstract
Deep reinforcement learning (RL) has gained popularity for automatically generating placements in modern chip design. However, the visual style of the fioorplans generated by these RL models is significantly different from the manual layouts' style, for RL placers usually only adopt metrics like wirelength and routing congestion as the reward in reinforcement learning, ignoring the complex and fine-grained layout experience of human experts. In this paper, we propose a placement scorer to rate the quality of layouts and apply abnormal detection to the fioorplanning task. In addition, we add the output of this scorer as a part of the reward for reinforcement learning of the placement process. Experimental results on ISPD 2005 benchmark show that our proposed placement quality scorer can evaluate the layouts according to human craft style efficiently, and that adding this scorer into reinforcement learning reward helps generating placements with shorter wirelength than previous methods for some circuit designs.
Binjie Yan, Zefang Yu, Mingye Xie, Wei Ran, Jingsheng Gao, Yuzhuo Fu, Ting Liu 0016
DATE6
2024 Oceanship: A Large-Scale Dataset for Underwater Audio Target Recognition
Suncheng Xiang, Jingsheng Gao, Jiacheng Ruan, Yanping Hu, Ting Liu 0016, Yuzhuo Fu
ICIC (4)4
2024 iDAT: inverse Distillation Adapter-Tuning
abstract
Adapter-Tuning (AT) method involves freezing a pre-trained model and introducing trainable adapter modules to acquire downstream knowledge, thereby calibrating the model for better adaptation to downstream tasks. This paper proposes a distillation framework for the AT method instead of crafting a carefully designed adapter module, which aims to improve fine-tuning performance. For the first time, we explore the possibility of combining the AT method with knowledge distillation. Via statistical analysis, we observe significant differences in the knowledge acquisition between adapter modules of different models. Leveraging these differences, we propose a simple yet effective framework called inverse Distillation Adapter-Tuning (iDAT). Specifically, we designate the smaller model as the teacher and the larger model as the student. The two are jointly trained, and online knowledge distillation is applied to inject knowledge of different perspective to student model, and significantly enhance the fine-tuning performance on downstream tasks. Extensive experiments on the VTAB-1K benchmark with 19 image classification tasks demonstrate the effectiveness of iDAT. The results show that using existing AT method within our iDAT framework can further yield a 2.66% performance gain, with only an additional 0.07M trainable parameters. Our approach compares favorably with state-of-the-arts without bells and whistles. Our code is available at https://github.com/JCruan519/iDAT.
Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Daize Dong, Suncheng Xiang, Ting Liu 0016, Yuzhuo Fu
ICME2
2024 Domain-Hierarchy Adaptation via Chain of Iterative Reasoning for Few-shot Hierarchical Text Classification
Ke Ji, Peng Wang 0004, Wenjun Ke 0002, Jiajun Liu 0005, Jingsheng Gao, Ziyu Shang
IJCAI6
2024 GIST: Improving Parameter Efficient Fine-Tuning via Knowledge Interaction
abstract
Recently, the Parameter Efficient Fine-Tuning (PEFT) method, which adjusts or introduces fewer trainable parameters to calibrate pre-trained models on downstream tasks, has been a hot research topic. However, existing PEFT methods within the traditional fine-tuning framework have two main shortcomings: 1) They overlook the explicit association between trainable parameters and downstream knowledge. 2) They neglect the interaction between the intrinsic task-agnostic knowledge of pre-trained models and the task-specific knowledge of downstream tasks. These oversights lead to insufficient utilization of knowledge and suboptimal performance. To address these issues, we propose a novel fine-tuning framework, named GIST, that can be seamlessly integrated into the current PEFT methods in a plug-and-play manner. Specifically, our framework first introduces a trainable token, called the Gist token, when applying PEFT methods on downstream tasks. This token serves as an aggregator of the task-specific knowledge learned by the PEFT methods and builds an explicit association with downstream tasks. Furthermore, to facilitate explicit interaction between task-agnostic and task-specific knowledge, we introduce the concept of knowledge interaction via a Bidirectional Kullback-Leibler Divergence objective. As a result, PEFT methods within our framework can enable the pre-trained model to understand downstream tasks more comprehensively by fully leveraging both types of knowledge. Extensive experiments on the 35 datasets demonstrate the universality and scalability of our framework. Notably, the PEFT method within our GIST framework achieves up to a 2.25% increase on the VTAB-1K benchmark with an addition of just 0.8K parameters (0.009‰ of ViT-B/16). The code is available at https://github.com/JCruan519/GIST.
Jiacheng Ruan, Jingsheng Gao, Mingye Xie, Suncheng Xiang, Zefang Yu, Ting Liu 0016, Yuzhuo Fu, Xiaoye Qu
ACM Multimedia2
2024 Toward an end-to-end implicit addressee modeling for dialogue disentanglement
Jingsheng Gao, Suncheng Xiang, Zhuowei Wang 0003, Ting Liu 0016, Yuzhuo Fu
Multim. Tools Appl.1
2024 Rethinking Person Re-Identification via Semantic-based Pretraining
abstract
Pretraining is a dominant paradigm in computer vision. Generally, supervised ImageNet pretraining is commonly used to initialize the backbones of person re-identification (Re-ID) models. However, recent works show a surprising result that CNN-based pretraining on ImageNet has limited impacts on Re-ID system due to the large domain gap between ImageNet and person Re-ID data. To seek an alternative to traditional pretraining, here we investigate semantic-based pretraining as another method to utilize additional textual data against ImageNet pretraining. Specifically, we manually construct a diversified FineGPR-C caption dataset for the first time on person Re-ID events. Based on it, a pure semantic-based pretraining approach named VTBR is proposed to adopt dense captions to learn visual representations with fewer images. We train convolutional neural networks from scratch on the captions of FineGPR-C dataset, and then transfer them to downstream Re-ID tasks. Comprehensive experiments conducted on benchmark datasets show that our VTBR can achieve competitive performance compared with ImageNet pretraining—despite using up to 1.4× fewer images, revealing its potential in Re-ID pretraining. Our source code is also publicly available at https://github.com/JeremyXSC/VTBR .
Suncheng Xiang, Dahong Qian, Jingsheng Gao, Ting Liu 0016, Yuzhuo Fu
ACM Trans. Multim. Comput. Commun. Appl.3
2023 LiveChat: A Large-Scale Personalized Dialogue Dataset Automatically Constructed from Live Streaming
abstract
Open-domain dialogue systems have made promising progress in recent years.While the state-of-the-art dialogue agents are built upon large-scale text-based social media data and large pre-trained models, there is no guarantee these agents could also perform well in fast-growing scenarios, such as live streaming, due to the bounded transferability of pretrained models and biased distributions of public datasets from Reddit and Weibo, etc.To improve the essential capability of responding and establish a benchmark in the live opendomain scenario, we introduce the LiveChat dataset, composed of 1.33 million real-life Chinese dialogues with almost 3800 average sessions across 351 personas and fine-grained profiles for each persona.LiveChat is automatically constructed by processing numerous live videos on the Internet and naturally falls within the scope of multi-party conversations, where the issues of Who says What to Whom should be considered.Therefore, we target two critical tasks of response modeling and addressee recognition and propose retrieval-based baselines grounded on advanced techniques.Experimental results have validated the positive effects of leveraging persona profiles and larger average sessions per persona.In addition, we also benchmark the transferability of advanced generation-based models on LiveChat and pose some future directions for current challenges. 1
Jingsheng Gao, Yixin Lian, Yuzhuo Fu, Baoyuan Wang
ACL (1)1
2023 Hierarchical Verbalizer for Few-Shot Hierarchical Text Classification
abstract
Due to the complex label hierarchy and intensive labeling cost in practice, the hierarchical text classification (HTC) suffers a poor performance especially when low-resource or fewshot settings are considered.Recently, there is a growing trend of applying prompts on pretrained language models (PLMs), which has exhibited effectiveness in the few-shot flat text classification tasks.However, limited work has studied the paradigm of prompt-based learning in the HTC problem when the training data is extremely scarce.In this work, we define a path-based few-shot setting and establish a strict path-based evaluation metric to further explore few-shot HTC tasks.To address the issue, we propose the hierarchical verbalizer ("Hi-erVerb"), a multi-verbalizer framework treating HTC as a single-or multi-label classification problem at multiple layers and learning vectors as verbalizers constrained by hierarchical structure and hierarchical contrastive learning.In this manner, HierVerb fuses label hierarchy knowledge into verbalizers and remarkably outperforms those who inject hierarchy through graph encoders, maximizing the benefits of PLMs.Extensive experiments on three popular HTC datasets under the few-shot settings demonstrate that prompt with HierVerb significantly boosts the HTC performance, meanwhile indicating an elegant way to bridge the gap between the large pre-trained model and downstream hierarchical classification tasks. 1
Ke Ji, Yixin Lian, Jingsheng Gao, Baoyuan Wang
ACL (1)3
2023 DynaSlim: Dynamic Slimming for Vision Transformers
abstract
Vision transformers (ViTs) have achieved significant performance on various vision tasks. However, high computational and memory costs hinder their edge deployment. Existing compression methods employ static constraints between accuracy and efficiency during sparsification. The static constraints restrict the sparsification efficiency and their initialization relies heavily on human expertise. We propose a dynamic slimming strategy for ViT, DynaSlim, to achieve an adaptive accuracy-efficiency constraint during sparsification. We first equip fine-grained, adjustable sparsity weights, the scaling factor between accuracy and efficiency, for multiple dimensions, including input tokens, Multihead Self-Attention (MSA) and Multilayer Perceptron (MLP). We then employ the heuristic search for these non-differentiable factors and combine the search with regularization-based sparsification to obtain the optimal sparsed model. Finally, we compress and retrain the sparsed model under various budgets to get our resulting submodels. Experiments show that our DynaSlim outperforms previous state-of-the-art methods under different budgets. For example, we reduce both parameters and FLOPs of DeiT-B by 39% while increasing its accuracy by 1.9% on ImageNet-1K. Moreover, we demonstrate the transferability of our compressed models on several downstream datasets.
Da Shi, Jingsheng Gao, Ting Liu 0016, Yuzhuo Fu
ICME2
2023 EGE-UNet: An Efficient Group Enhanced UNet for Skin Lesion Segmentation
Jiacheng Ruan, Mingye Xie, Jingsheng Gao, Ting Liu 0016, Yuzhuo Fu
MICCAI (4)3
2021 Focus on Inherent Attributes for Temporal Knowledge Graph Completion
abstract
In the last few years, the availability of temporal knowledge graphs has stimulated extensive research in temporal knowledge graph completion (TKGC) and temporal knowledge graph embedding (TKGE), where temporal information is added to static knowledge graphs that have been widely applied previously. However, most existing methods, such as current state-of-the-art DE-SimplE and TeRo, learn embeddings of temporal-evolving attributes, overlooking the inherent attributes inside entities, where some essential and inherent features are included. In this paper, we introduce a novel method utilizing Inherent Attributes with a Graph Attention network (IAGAT) for TKGC. Our IAGAT extracts inherent attributes from sufficient features corresponding to various facts at different time stamps, to obtain the inherent embeddings. And we take advantages of previous rotation based methods to obtain the temporal-evolving embed-dings. Through extensive experiments and sufficient comparisons, we demonstrate our model outperforms the current state-of-the-art models on link prediction task. Furthermore, we evaluate and prove the necessity of the inherent attributes in performance improvement, and study how our model functions in extracting inherent features.
Kai Chen 0020, Aiping Li, Jingsheng Gao, Sixia Ma
IJCNN4
2020 Location Privacy-Preserving Truth Discovery in Mobile Crowd Sensing
abstract
Truth discovery techniques are commonly used in mobile crowd sensing (MCS) applications to infer accurate aggregated results based on quality-aware data aggregation. However, the location information of participants may be exposed when they upload their sensitive geo-tagged sensory data to relative platforms. While there are considerable existing privacy preserving truth discovery schemes for MCS, they mainly focus on protecting the privacy of sensory data, neglecting the tagged location information which is of equal if not higher importance for the privacy of participants. In this paper, we propose a novel and efficient location privacy preserving truth discovery (LoPPTD) mechanism, which can achieve data aggregation with high accuracy, while protecting both location privacy and data privacy of users. By structuring multi-dimensional sensory data obtained at different locations and exploiting homomorphic Paillier encryption, our approach can prevent leakage of both sensory data and tagged locations effectively. Also, super-increasing sequence techniques are employed in Lo-PPTD to ensure efficiency and feasibility. Theoretical analysis and thorough experiments performed on real-world datasets demonstrate that the proposed scheme can achieve high aggregation accuracy while providing complete privacy protection for users.
Jingsheng Gao, Shaojing Fu, Yuchuan Luo, Tao Xie 0012
ICCCN1