Chongyang Gao

dblp:259/8515 · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0002-2358-4710ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 2 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploiting Label-Independent Regularization from Spatial Patterns for Whole Slide Image Analysis
abstract
Whole slide images, with their gigapixel-scale panoramas of tissue samples, are pivotal for precise disease diagnosis. However, their analysis is hindered by immense data size and scarce annotations. Existing MIL methods face challenges due to the fundamental imbalance where a single bag-level label must guide the learning of numerous patch-level features. This sparse supervision makes it difficult to reliably identify discriminative patches during training, leading to unstable optimization and suboptimal solutions. We propose a spatially regularized MIL framework that leverages inherent spatial relationships among patch features as label-independent regularization signals. Our approach learns a shared representation space by jointly optimizing feature-induced spatial reconstruction and label-guided classification objectives, enforcing consistency between intrinsic structural patterns and supervisory signals. Experimental results on multiple public datasets demonstrate significant improvements over state-of-the-art methods, offering a promising direction. The code is available at https://github.com/wwyi1828/SRMIL.
Weiyi Wu, Xinwen Xu, Chongyang Gao, Xingjian Diao, Siting Li, Jiang Gui
WACV3
2025 GODDS: The Global Online Deepfake Detection System
abstract
Fake audios, videos, and images are now proliferating widely. We developed GODDS, the Global Online Deepfake Detection system, for a specific user community, namely journalists. GODDS leverages an ensemble of deepfake detectors, along with a human in the loop, to provide a deepfake report on each submitted video/image/audio or VIA artifact submitted to the system. To date, VIA artifacts submitted by over 50 journalists from outlets such as the New York Times, Wall Street Journal, CNN, Agence France Press, and others have been run through GODDS. Unlike other deepfake detection systems, GODDS doesn't just focus on the submitted artifact but automatically derives context about the subject of the VIA artifact. Because context is not always available on all subjects, GODDS focuses on alleged deepfakes of high profile individuals, organizations, and events, where there is likely to be considerable contextual information.
Marco Postiglione, Julian Baldwin, Natalia Denisenko, Luke Fosdick, Chongyang Gao, Isabel Gortner, Chiara Pulice, Sarit Kraus, V. S. Subrahmanian
AAAI5
2025 On Large Language Model Continual Unlearning
abstract
While large language models have demonstrated impressive performance across various domains and tasks, their security issues have become increasingly severe. Machine unlearning has emerged as a representative approach for model safety and security by removing the influence of undesired data on the target model. However, these methods do not sufficiently consider that unlearning requests in real-world scenarios are continuously emerging, especially in the context of LLMs, which may lead to accumulated model utility loss that eventually becomes unacceptable. Moreover, existing LLM unlearning methods often ignore previous data access limitations due to privacy concerns and copyright protection. Without previous data, the utility preservation during unlearning is much harder. To overcome these challenges, we propose the \OOO{} framework that includes an \underline{\textit{O}}rthogonal low-rank adapter (LoRA) for continually unlearning requested data and an \underline{\textit{O}}ut-\underline{\textit{O}}f-Distribution (OOD) detector to measure the similarity between input and unlearning data. The orthogonal LoRA achieves parameter disentanglement among continual unlearning requests. The OOD detector is trained with a novel contrastive entropy loss and utilizes a glocal-aware scoring mechanism. During inference, our \OOO{} framework can decide whether and to what extent to load the unlearning LoRA based on the OOD detector's predicted similarity between the input and the unlearned knowledge. Notably, \OOO{}'s effectiveness does not rely on any retained data. We conducted extensive experiments on \OOO{} and state-of-the-art LLM unlearning methods across three tasks and seven datasets. The results indicate that \OOO{} consistently achieves the best unlearning effectiveness and utility preservation, especially when facing continuous unlearning requests. The source codes can be found at \url{https://github.com/GCYZSL/O3-LLM-UNLEARNING}.
Chongyang Gao, Lixu Wang, Kaize Ding, Chenkai Weng, Xiao Wang 0012, Qi Zhu 0002
ICLR1
2024 How to Configure Good In-Context Sequence for Visual Question Answering
abstract
Inspired by the success of Large Language Models in dealing with new tasks via In-Context Learning (ICL) in NLP, researchers have also developed Large Vision-Language Models (LVLMs) with ICL capabilities. However, when implementing ICL using these LVLMs, researchers usually resort to the simplest way like random sampling to configure the in-context sequence, thus leading to sub-optimal results. To enhance the ICL performance, in this study, we use Visual Question Answering (VQA) as case study to explore diverse in-context configurations to find the powerful ones. Additionally, through observing the changes of the LVLM outputs by altering the in-context sequence, we gain insights into the inner properties of LVLMs, improving our understanding of them. Specifically, to explore in-context configurations, we design diverse retrieval methods and employ different strategies to manipulate the retrieved demonstrations. Through exhaustive experiments on three VQA datasets: VQAv2, VizWiz, and OK-VQA, we uncover three important inner properties of the applied LVLM and demonstrate which strategies can consistently improve the ICL VQA performance. Our code is provided in: https://github.com/GaryJiajia/OFv2_ICL_VQA.
Jiawei Peng 0001, Huiyi Chen, Chongyang Gao, Xu Yang 0021
CVPR4
2024 AlphaLoRA: Assigning LoRA Experts Based on Layer Training Quality
abstract
Parameter-efficient fine-tuning methods, such as Low-Rank Adaptation (LoRA), are known to enhance training efficiency in Large Language Models (LLMs).Due to the limited parameters of LoRA, recent studies seek to combine LoRA with Mixture-of-Experts (MoE) to boost performance across various tasks.However, inspired by the observed redundancy in traditional MoE structures, prior studies find that LoRA experts within the MoE architecture also exhibit redundancy, suggesting a need to vary the allocation of LoRA experts across different layers.In this paper, we leverage Heavy-Tailed Self-Regularization (HT-SR) Theory to design a fine-grained allocation strategy.Our analysis reveals that the number of experts per layer correlates with layer training quality, which exhibits significant variability across layers.Based on this, we introduce AlphaLoRA, a theoretically principled and training-free method for allocating LoRA experts to reduce redundancy further.Experiments on three models across ten language processing and reasoning benchmarks demonstrate that AlphaLoRA achieves comparable or superior performance over all baselines.Our code is available at https://github.com/morelife2017/alphalora.
Peijun Qing, Chongyang Gao, Yefan Zhou, Xingjian Diao, Yaoqing Yang 0002, Soroush Vosoughi
EMNLP2
2024 Knowledge from Large-Scale Protein Contact Prediction Models Can Be Transferred to the Data-Scarce RNA Contact Prediction Task
Yiren Jian, Chongyang Gao, Yunjie Zhao, Soroush Vosoughi
ICPR (10)2
2024 GEM: Generating Engaging Multimodal Content
Chongyang Gao, Yiren Jian, Natalia Denisenko, Soroush Vosoughi, V. S. Subrahmanian
IJCAI1
2024 Enhancing Network Role Modeling: Introducing Attributed Multiplex Structural Role Embedding for Complex Networks
Chenghan Huang, Ruiye Yao, Chongyang Gao, Soroush Vosoughi
PAKDD (2)4
2024 ${\sf FakeDB}$FakeDB: Generating Fake Synthetic Databases
abstract
Health care providers may wish to share limited information with researchers. Manufacturing companies may want to share some but not all data with regulators or partners. Since the emergence of generative adversarial networks (GANs), efforts have been made to generate synthetic data that preserves semantic properties on the one hand and distributions on the other hand. However, all past efforts focus on a single table at a time. We propose${\sf FakeDB}$, a general framework to generate synthetic data that preserves a a wide variety of semantic integrity constraints as well as a broad set of statistical properties, across an entire relational database. We compare${\sf FakeDB}$with natural extensions of prior work on 8 well known relational databases as well as on a synthetically generated dataset, and show that${\sf FakeDB}$outperforms them. We also show that${\sf FakeDB}$runs in reasonable amounts of time, making it a practical solution to the problem of generating synthetic data.
Chongyang Gao, Sushil Jajodia, Andrea Pugliese 0001, V. S. Subrahmanian
IEEE Trans. Dependable Secur. Comput.1
2023 Improving Representation Learning for Histopathologic Images with Cluster Constraints
abstract
Recent advances in whole-slide image (WSI) scanners and computational capabilities have significantly propelled the application of artificial intelligence in histopathology slide analysis. While these strides are promising, current supervised learning approaches for WSI analysis come with the challenge of exhaustively labeling high-resolution slides-a process that is both labor-intensive and timeconsuming. In contrast, self-supervised learning (SSL) pretraining strategies are emerging as a viable alternative, given that they don't rely on explicit data annotations. These SSL strategies are quickly bridging the performance disparity with their supervised counterparts. In this context, we introduce an SSL framework. This framework aims for transferable representation learning and semantically meaningful clustering by synergizing invariance loss and clustering loss in WSI analysis. Notably, our approach outperforms common SSL methods in downstream classification and clustering tasks, as evidenced by tests on the Camelyon16 and a pancreatic cancer dataset. The code and additional details are accessible at https://github.com/wwyi1828/CluSiam.
Weiyi Wu, Chongyang Gao, Joseph DiPalma, Soroush Vosoughi, Saeed Hassanpour
ICCV2
2023 Bootstrapping Vision-Language Learning with Decoupled Language Pre-training
abstract
We present a novel methodology aimed at optimizing the application of frozen large language models (LLMs) for resource-intensive vision-language (VL) pre-training. The current paradigm uses visual features as prompts to guide language models, with a focus on determining the most relevant visual features for corresponding text. Our approach diverges by concentrating on the language component, specifically identifying the optimal prompts to align with visual features. We introduce the Prompt-Transformer (P-Former), a model that predicts these ideal prompts, which is trained exclusively on linguistic data, bypassing the need for image-text pairings. This strategy subtly bifurcates the end-to-end VL training process into an additional, separate stage. Our experiments reveal that our framework significantly enhances the performance of a robust image-to-text baseline (BLIP-2), and effectively narrows the performance gap between models trained with either 4M or 129M image-text pairs. Importantly, our framework is modality-agnostic and flexible in terms of architectural design, as validated by its successful application in a video learning task using varied base modules. The code will be made available at https://github.com/yiren-jian/BLIText.
Yiren Jian, Chongyang Gao, Soroush Vosoughi
NeurIPS2
2023 Joint Latent Topic Discovery and Expectation Modeling for Financial Markets
Chenghan Huang, Chongyang Gao, Soroush Vosoughi
PAKDD (3)3
2023 Learning to Collocate Visual-Linguistic Neural Modules for Image Captioning
Xu Yang 0021, Hanwang Zhang, Chongyang Gao, Jianfei Cai 0001
Int. J. Comput. Vis.3
2023 Linking Terrorist Network Structure to Lethality: Algorithms and Analysis of Al Qaeda and ISIS
abstract
Without measures of the lethality of terrorist networks, it is very difficult to assess if capturing or killing a terrorist is effective. We present the predictive lethality analysis of terrorist organization () algorithm, which merges machine learning with techniques from graph theory and social network analysis to predict the number of attacks that a terrorist network will carry out based on a network structure alone. We show that is highly accurate on two novel datasets, which cover Al Qaeda (AQ) and the Islamic State (ISIS). Using both machine learning and statistical methods, we show that the most significant macrofeatures for predicting AQ’s lethality are related to their public communications (PCs) and logistical subnetworks, while the leadership and operational subnetworks are most impactful for predicting ISISs lethality. Across both groups, the average degree and the diameters of the strongly connected components (SCCs) within these networks are strongly linked with lethality.
Youdinghuan Chen, Chongyang Gao, Daveed Gartenstein-Ross, Kevin T. Greene, Karin Kalif, Sarit Kraus, Francesco Parisi, Chiara Pulice, Anja Subasic, V. S. Subrahmanian
IEEE Trans. Comput. Soc. Syst.2
2022 Non-Parallel Text Style Transfer with Self-Parallel Supervision
Ruibo Liu, Chongyang Gao, Chenyan Jia, Guangxuan Xu, Soroush Vosoughi
ICLR2
2022 Knowledge Infused Decoding
Ruibo Liu, Guoqing Zheng, Radhika Gaonkar, Chongyang Gao, Soroush Vosoughi, Milad Shokouhi, Ahmed Awadallah 0001
ICLR5
2022 Embedding Hallucination for Few-shot Language Fine-tuning
abstract
Few-shot language learners adapt knowledge from a pre-trained model to recognize novel classes from a few-labeled sentences.In such settings, fine-tuning a pre-trained language model can cause severe over-fitting.In this paper, we propose an Embedding Hallucination (EmbedHalluc) method, which generates auxiliary embedding-label pairs to expand the finetuning dataset.The hallucinator is trained by playing an adversarial game with the discriminator, such that the hallucinated embedding is indiscriminative to the real ones in the finetuning dataset.By training with the extended dataset, the language learner effectively learns from the diverse hallucinated embeddings to overcome the over-fitting issue.Experiments demonstrate that our proposed method is effective in a wide range of language tasks, outperforming current fine-tuning methods.Further, we show that EmbedHalluc outperforms other methods that address this over-fitting problem, such as common data augmentation, semi-supervised pseudo-labeling, and regularization.
Yiren Jian, Chongyang Gao, Soroush Vosoughi
NAACL-HLT2
2022 Contrastive Learning for Prompt-based Few-shot Language Learners
abstract
The impressive performance of GPT-3 using natural language prompts and in-context learning has inspired work on better fine-tuning of moderately-sized models under this paradigm.Following this line of work, we present a contrastive learning framework that clusters inputs from the same class for better generality of models trained with only limited examples.Specifically, we propose a supervised contrastive framework that clusters inputs from the same class under different augmented "views" and repel the ones from different classes.We create different "views" of an example by appending it with different language prompts and contextual demonstrations.Combining a contrastive loss with the standard masked language modeling (MLM) loss in prompt-based few-shot learners, the experimental results show that our method can improve over the state-of-the-art methods in a diverse set of 15 language tasks.Our framework makes minimal assumptions on the task or the base model, and can be applied to many recent methods with little modification.
Yiren Jian, Chongyang Gao, Soroush Vosoughi
NAACL-HLT2
2022 Non-Linguistic Supervision for Contrastive Learning of Sentence Embeddings
abstract
Semantic representation learning for sentences is an important and well-studied problem in NLP. The current trend for this task involves training a Transformer-based sentence encoder through a contrastive objective with text, i.e., clustering sentences with semantically similar meanings and scattering others. In this work, we find the performance of Transformer models as sentence encoders can be improved by training with multi-modal multi-task losses, using unpaired examples from another modality (e.g., sentences and unrelated image/audio data). In particular, besides learning by the contrastive loss on text, our model clusters examples from a non-linguistic domain (e.g., visual/audio) with a similar contrastive loss at the same time. The reliance of our framework on unpaired non-linguistic data makes it language-agnostic, enabling it to be widely applicable beyond English NLP. Experiments on 7 semantic textual similarity benchmarks reveal that models trained with the additional non-linguistic (images/audio) contrastive objective lead to higher quality sentence embeddings. This indicates that Transformer models are able to generalize better by doing a similar task (i.e., clustering) with \textit{unpaired} examples from different modalities in a multi-task fashion. The code is available at https://github.com/yiren-jian/NonLing-CSE.
Yiren Jian, Chongyang Gao, Soroush Vosoughi
NeurIPS2
2021 Embedding Heterogeneous Networks into Hyperbolic Space Without Meta-path
abstract
Networks found in the real-world are numerous and varied. A common type of network is the heterogeneous network, where the nodes (and edges) can be of different types. Accordingly, there have been efforts at learning representations of these heterogeneous networks in low-dimensional space. However, most of the existing heterogeneous network embedding suffers from the following two drawbacks: (1) The target space is usually Euclidean. Conversely, many recent works have shown that complex networks may have hyperbolic latent anatomy, which is non-Euclidean. (2) These methods usually rely on meta-paths, which requires domain-specific prior knowledge for meta-path selection. Additionally, different down-streaming tasks on the same network might require different meta-paths in order to generate task-specific embeddings. In this paper, we propose a novel self-guided random walk method that does not require meta-path for embedding heterogeneous networks into hyperbolic space. We conduct thorough experiments for the tasks of network reconstruction and link prediction on two public datasets, showing that our model outperforms a variety of well-known baselines across all tasks.
Chongyang Gao, Chenghan Huang, Ruibo Liu, Soroush Vosoughi
AAAI2
2021 Auto-Parsing Network for Image Captioning and Visual Question Answering
abstract
We propose an Auto-Parsing Network (APN) to discover and exploit the input data’s hidden tree structures for improving the effectiveness of the Transformer-based vision-language systems. Specifically, we impose a Probabilistic Graphical Model (PGM) parameterized by the attention operations on each self-attention layer to incorporate sparse assumption. We use this PGM to softly segment an input sequence into a few clusters where each cluster can be treated as the parent of the inside entities. By stacking these PGM constrained self-attention layers, the clusters in a lower layer compose into a new sequence, and the PGM in a higher layer will further segment this sequence. Iteratively, a sparse tree can be implicitly parsed, and this tree’s hierarchical knowledge is incorporated into the transformed embeddings, which can be used for solving the target vision-language tasks. Specifically, we showcase that our APN can strengthen Transformer based networks in two major vision-language tasks: Captioning and Visual Question Answering. Also, a PGM probability-based parsing algorithm is developed by which we can discover what the hidden structure of input is during the inference.
Xu Yang 0021, Chongyang Gao, Hanwang Zhang, Jianfei Cai 0001
ICCV2
2021 MetaPix: Domain transfer for semantic segmentation by meta pixel weighting
Yiren Jian, Chongyang Gao
Image Vis. Comput.2
2020 Hierarchical Scene Graph Encoder-Decoder for Image Paragraph Captioning
abstract
When we humans tell a long paragraph about an image, we usually first implicitly compose a mental "script'' and then comply with it to generate the paragraph. Inspired by this, we render the modern encoder-decoder based image paragraph captioning model such ability by proposing Hierarchical Scene Graph Encoder-Decoder (HSGED) for generating coherent and distinctive paragraphs. In particular, we use the image scene graph as the "script" to incorporate rich semantic knowledge and, more importantly, the hierarchical constraints into the model. Specifically, we design a sentence scene graph RNN (SSG-RNN) to generate sub-graph level topics, which constrain the word scene graph RNN (WSG-RNN) to generate the corresponding sentences. We propose irredundant attention in SSG-RNN to improve the possibility of abstracting topics from rarely described sub-graphs and inheriting attention in WSG-RNN to generate more grounded sentences with the abstracted topics, both of which give rise to more distinctive paragraphs. An efficient sentence-level loss is also proposed for encouraging the sequence of generated sentences to be similar to that of the ground-truth paragraphs. We validate HSGED on Stanford image paragraph dataset and show that it not only achieves a new state-of-the-art 36.02 CIDEr-D, but also generates more coherent and distinctive paragraphs under various metrics.
Xu Yang 0021, Chongyang Gao, Hanwang Zhang, Jianfei Cai 0001
ACM Multimedia2
2020 Sequence in sequence for video captioning
Huiyun Wang, Chongyang Gao, Yahong Han
Pattern Recognit. Lett.2