Kaixiong Gong

dblp:289/0124 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0002-7510-9356ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Probing Audio-Visual Reasoning in Multimodal Language Models through the Lens of Audio
abstract
Kaixiong Gong, Kaituo Feng, Bohao Li, Yibing Wang, Mofan Cheng, Shijia Yang, Jiaming Han, Benyou Wang, Yutong Bai, Zhuoran Yang, Xiangyu Yue. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Kaixiong Gong, Kaituo Feng, Mofan Cheng, Shijia Yang, Jiaming Han, Benyou Wang, Yutong Bai, Zhuoran Yang, Xiangyu Yue 0001
ACL (1)1
2026 Learning While Staying Curious: Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models
abstract
Hao Wang, Hao Gu, Hongming Piao, Kaixiong Gong, Yuxiao Ye, Xiangyu Yue, Sirui Han, Yike Guo, Dapeng Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hao Wang 0193, Hao Gu 0001, Hongming Piao, Kaixiong Gong, Yuxiao Ye, Xiangyu Yue 0001, Sirui Han, Yike Guo, Dapeng Oliver Wu
ACL (1)4
2025 Video-R1: Reinforcing Video Reasoning in MLLMs
abstract
Inspired by DeepSeek-R1's success in eliciting reasoning abilities through rule-based reinforcement learning (RL), we introduce Video-R1 as the first attempt to systematically explore the R1 paradigm for incentivizing video reasoning within multimodal large language models (MLLMs). However, directly applying RL training with the GRPO algorithm to video reasoning presents two primary challenges: (i) a lack of temporal modeling for video reasoning, and (ii) the scarcity of high-quality video-reasoning data. To address these issues, we first propose the T-GRPO algorithm, which encourages models to utilize temporal information in videos for reasoning. Additionally, instead of relying solely on video data, we incorporate high-quality image-reasoning data into the training process. We have constructed two datasets: Video-R1-CoT-165k for SFT cold start and Video-R1-260k for RL training, both comprising image and video data. Experimental results demonstrate that Video-R1 achieves significant improvements on video reasoning benchmarks such as VideoMMMU and VSI-Bench, as well as on general video benchmarks including MVBench and TempCompass, etc. Notably, Video-R1-7B attains a 37.1\% accuracy on video spatial reasoning benchmark VSI-bench, surpassing the commercial proprietary model GPT-4o. All code, models, and data will be released.
Kaituo Feng, Kaixiong Gong, Zonghao Guo, Tianshuo Peng, Junfei Wu, Benyou Wang, Xiangyu Yue 0001
NeurIPS2
2025 Source-Free Active Domain Adaptation via Augmentation-Based Sample Query and Progressive Model Adaptation
abstract
Active domain adaptation (ADA), which enormously improves the performance of unsupervised domain adaptation (UDA) at the expense of annotating limited target data, has attracted a surge of interest. However, in real-world applications, the source data in conventional ADA are not always accessible due to data privacy and security issues. To alleviate this dilemma, we introduce a more practical and challenging setting, dubbed as source-free ADA (SFADA), where one can select a small quota of target samples for label query to assist the model learning, but labeled source data are unavailable. Therefore, how to query the most informative target samples and mitigate the domain gap without the aid of source data are two key challenges in SFADA. To address SFADA, we propose a unified method SQAdapt via augmentation-based ample uery and progressive model Adapt ation. In specific, an active selection module (ASM) is built for target label query, which exploits data augmentation to select the most informative target samples with high predictive sensitivity and uncertainty. Then, we further introduce a classifier adaptation module (CAM) to leverage both the labeled and unlabeled target data for progressively calibrating the classifier weights. Meanwhile, the source-like target samples with low selection scores are taken as source surrogates to realize the distribution alignment in the source-free scenario by the proposed distribution alignment module (DAM). Moreover, as a general active label query method, SQAdapt can be easily integrated into other source-free UDA (SFUDA) methods, and improve their performance. Comprehensive experiments on multiple benchmarks have shown that SQAdapt can achieve superior performance and even surpass most of the ADA methods.
Shuang Li 0008, Rui Zhang 0113, Kaixiong Gong, Mixue Xie, Wenxuan Ma 0001, Guangyu Gao
IEEE Trans. Neural Networks Learn. Syst.3
2024 Text-to-3D Generation with Bidirectional Diffusion Using Both 2D and 3D Priors
abstract
Most 3D generation research focuses on up-projecting 2D foundation models into the 3D space, either by minimizing 2D Score Distillation Sampling (SDS) loss or fine-tuning on multi-view datasets. Without explicit 3D priors, these methods often lead to geometric anomalies and multi-view inconsistency. Recently, researchers have attempted to improve the genuineness of 3D objects by directly training on 3D datasets, albeit at the cost of low-quality texture generation due to the limited texture diversity in 3D datasets. To harness the advantages of both approaches, we propose Bidirectional Diffusion (BiDiff), a unified framework that incorporates both a 3D and a 2D diffusion process, to preserve both 3D fidelity and 2D texture richness, respectively. Moreover, as a simple combination may yield inconsistent generation results, we further bridge them with novel bidirectional guidance. In addition, our method can be used as an initialization of optimization-based models to further improve the quality of 3D models and the efficiency of optimization, reducing the process from 3.4 hours to 20 minutes. Experimental results have shown that our model achieves high-quality, diverse, and scalable 3D generation. Project website https://bidiff.github.io/.
Lihe Ding, Shaocong Dong, Zhanpeng Huang, Zibin Wang, Kaixiong Gong, Dan Xu 0002, Tianfan Xue
CVPR6
2024 OneLLM: One Framework to Align All Modalities with Language
abstract
Multimodal large language models (MLLMs) have gained significant attention due to their strong multimodal understanding capability. However, existing works rely heavily on modality-specific encoders, which usually differ in architecture and are limited to common modalities. In this paper, we present OneLLM, an MLLM that aligns eight modalities to language using a unified framework. We achieve this through a unified multimodal encoder and a progressive multimodal alignment pipeline. In detail, we first train an image projection module to connect a vision encoder with LLM. Then, we build a universal projection module (UPM) by mixing multiple image projection modules and dynamic routing. Finally, we progressively align more modalities to LLM with the UPM. To fully leverage the potential of OneLLM in following instructions, we also curated a comprehensive multimodal instruction dataset, including 2M items from image, audio, video, point cloud, depth/normal map, IMU and fMRI brain activity. OneLLM is evaluated on 25 diverse benchmarks, encompassing tasks such as multimodal captioning, question answering and reasoning, where it delivers excellent performance. Code, data, model and online demo are available at https://github.com/csuhan/OneLLM.
Jiaming Han, Kaixiong Gong, Jiaqi Wang 0003, Kaipeng Zhang, Dahua Lin, Yu Qiao 0001, Peng Gao 0007, Xiangyu Yue 0001
CVPR2
2024 Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
abstract
We propose to improve transformers of a specific modality with irrelevant data from other modalities, e.g., improve an ImageNet model with audio or point cloud datasets. We would like to highlight that the data samples of the target modality are irrelevant to the other modalities, which distinguishes our method from other works utilizing paired (e.g., CLIP) or interleaved data of different modalities. We propose a methodology named Multimodal Pathway - given a target modality and a transformer designed for it, we use an auxiliary transformer trained with data of another modality and construct pathways to connect components of the two models so that data of the target modality can be processed by both models. In this way, we utilize the universal sequence-to-sequence modeling abilities of transformers obtained from two modalities. As a concrete implementation, we use a modality-specific tokenizer and task-specific head as usual but utilize the transformer blocks of the auxiliary model via a proposed method named Cross-Modal Re-parameterization, which exploits the auxiliary weights without any inference costs. On the image, point cloud, video, and audio recognition tasks, we observe significant and consistent performance improvements with irrelevant data from other modalities. The code and models are available at https://github.com/AILab-CVC/M2PT.
Xiaohan Ding, Kaixiong Gong, Yixiao Ge, Ying Shan, Xiangyu Yue 0001
CVPR3
2024 Reg-TTA3D: Better Regression Makes Better Test-Time Adaptive 3D Object Detection
Jiakang Yuan, Bo Zhang 0069, Kaixiong Gong, Xiangyu Yue 0001, Botian Shi, Yu Qiao 0001, Tao Chen 0003
ECCV (43)3
2024 Subjective Topic meets LLMs: Unleashing Comprehensive, Reflective and Creative Thinking through the Negation of Negation
abstract
Large language models (LLMs) exhibit powerful reasoning capacity, as evidenced by prior studies focusing on objective topics that with unique standard answer such as arithmetic and commonsense reasoning.However, the reasoning to definite answers emphasizes more on logical thinking, and falls short in effectively reflecting the comprehensive, reflective, and creative thinking that is also critical for the overall reasoning prowess of LLMs.In light of this, we build a dataset SJTP comprising diverse SubJective ToPics with free responses, as well as three evaluation indicators to fully explore LLM's reasoning ability.We observe that a sole emphasis on logical thinking falls short in effectively tackling subjective challenges.Therefore, we introduce a framework grounded in the principle of the Negation of Negation (NeoN) to unleash the potential comprehensive, reflective, and creative thinking abilities of LLMs.Comprehensive experiments on SJTP demonstrate the efficacy of NeoN, and the enhanced performance on various objective reasoning tasks unequivocally underscores the benefits of stimulating LLM's subjective thinking in augmenting overall reasoning capabilities.
Fangrui Lv, Kaixiong Gong, Jian Liang 0002, Xinyu Pang, Changshui Zhang
EMNLP2
2024 Bifröst: 3D-Aware Image Compositing with Language Instructions
Kaixiong Gong, Wei-Hong Li 0001, Xili Dai, Tao Chen 0003, Xiangyu Yue 0001
NeurIPS2
2024 Adapting Across Domains via Target-Oriented Transferable Semantic Augmentation Under Prototype Constraint
Mixue Xie, Shuang Li 0008, Kaixiong Gong, Yulin Wang 0002, Gao Huang 0001
Int. J. Comput. Vis.3
2023 Critical Classes and Samples Discovering for Partial Domain Adaptation
abstract
Partial domain adaptation (PDA) attempts to learn transferable models from a large-scale labeled source domain to a small unlabeled target domain with fewer classes, which has attracted a recent surge of interest in transfer learning. Most conventional PDA approaches endeavor to design delicate source weighting schemes by leveraging target predictions to align cross-domain distributions in the shared class space. Accordingly, two crucial issues are overlooked in these methods. First, target prediction is a double-edged sword, and inaccurate predictions will result in negative transfer inevitably. Second, not all target samples have equal transferability during the adaptation; thus, "ambiguous" target data predicted with high uncertainty should be paid more attentions. In this article, we propose a critical classes and samples discovering network (CSDN) to identify the most relevant source classes and critical target samples, such that more precise cross-domain alignment in the shared label space could be enforced by co-training two diverse classifiers. Specifically, during the training process, CSDN introduces an adaptive source class weighting scheme to select the most relevant classes dynamically. Meanwhile, based on the designed target ambiguous score, CSDN emphasizes more on ambiguous target samples with larger inconsistent predictions to enable fine-grained alignment. Taking a step further, the weighting schemes in CSDN can be easily coupled with other PDA and DA methods to further boost their performance, thereby demonstrating its flexibility. Extensive experiments verify that CSDN attains excellent results compared to state of the arts on four highly competitive benchmark datasets.
Shuang Li 0008, Kaixiong Gong, Binhui Xie, Chi Harold Liu, Weipeng Cao, Song Tian
IEEE Trans. Cybern.2
2023 End-to-End Transferable Anomaly Detection via Multi-Spectral Cross-Domain Representation Alignment
abstract
Anomaly detection (AD) aims to distinguish abnormal instances from what is defined as normal, which strongly correlates with the safe and robust applications of machine learning. A well-performed anomaly detector often relies on the training on massive labeled data, while it is of high cost to annotate data in practice. Fortunately, this dilemma can be solved by transferring the knowledge of a label-rich dataset (source domain) to assist the learning on the label-scarce dataset (target domain), which is known as domain adaptation in transfer learning. In this paper, we propose a Multi-spectral Cross-domain Representation Alignment (MsRA) method for the anomaly detection in the domain adaptation setting, where we can only access normal source data andlimitednormal target data. Specifically, MsRA first constructs multi-spectral feature representations by fusing different frequency components of the original features, which mitigates the information scarcity due to limited target training data by capturing richer input pattern information. Then we employ the adversarial training strategy to learn domain-invariant features and force the features of normal data to be more compact by the center clustering. Finally, the distance of each sample to the prototype of normal class can be used as its anomaly score, where the prototype is the center of both source and target data. In this way, we achieve anomaly detection in an end-to-end manner, without two-stage training for feature extraction and anomaly detection. Comprehensive experiments on cross-domain anomaly detection benchmarks validate the effectiveness of MsRA.
Shuang Li 0008, Shugang Li 0002, Mixue Xie, Kaixiong Gong, Jianxin Zhao 0001, Chi Harold Liu, Guoren Wang
IEEE Trans. Knowl. Data Eng.4
2022 Improving Transferability for Domain Adaptive Detection Transformers
abstract
DETR-style detectors stand out amongst in-domain scenarios, but their properties in domain shift settings are under-explored. This paper aims to build a simple but effective baseline with a DETR-style detector on domain shift settings based on two findings. For one, mitigating the domain shift on the backbone and the decoder output features excels in getting favorable results. For another, advanced domain alignment methods in both parts further enhance the performance. Thus, we propose the Object-Aware Alignment (OAA) module and the Optimal Transport based Alignment (OTA) module to achieve comprehensive domain alignment on the outputs of the backbone and the detector. The OAA module aligns the foreground regions identified by pseudo-labels in the backbone outputs, leading to domain-invariant base features. The OTA module utilizes sliced Wasserstein distance to maximize the retention of location information while minimizing the domain gap in the decoder outputs. We implement the findings and the alignment modules into our adaptation method, and it benchmarks the DETR-style detector on the domain shift settings. Experiments on various domain adaptive scenarios validate the effectiveness of our method.
Kaixiong Gong, Shuang Li 0008, Shugang Li 0002, Rui Zhang 0113, Chi Harold Liu
ACM Multimedia1
2021 MetaSAug: Meta Semantic Augmentation for Long-Tailed Visual Recognition
abstract
Real-world training data usually exhibits long-tailed distribution, where several majority classes have a significantly larger number of samples than the remaining minority classes. This imbalance degrades the performance of typical supervised learning algorithms designed for balanced training sets. In this paper, we address this issue by augmenting minority classes with a recently proposed implicit semantic data augmentation (ISDA) algorithm [37], which produces diversified augmented samples by translating deep features along many semantically meaningful directions. Importantly, given that ISDA estimates the class-conditional statistics to obtain semantic directions, we find it ineffective to do this on minority classes due to the insufficient training data. To this end, we propose a novel approach to learn transformed semantic directions with meta-learning automatically. In specific, the augmentation strategy during training is dynamically optimized, aiming to minimize the loss on a small balanced validation set, which is approximated via a meta update step. Extensive empirical results on CIFAR-LT-10/100, ImageNet-LT, and iNaturalist 2017/2018 validate the effectiveness of our method.
Shuang Li 0008, Kaixiong Gong, Chi Harold Liu, Yulin Wang 0002, Feng Qiao 0001, Xinjing Cheng
CVPR2
2021 Transferable Semantic Augmentation for Domain Adaptation
abstract
Domain adaptation has been widely explored by transferring the knowledge from a label-rich source domain to a related but unlabeled target domain. Most existing domain adaptation algorithms attend to adapting feature representations across two domains with the guidance of a shared source-supervised classifier. However, such classifier limits the generalization ability towards unlabeled target recognition. To remedy this, we propose a Transferable Semantic Augmentation (TSA) approach to enhance the classifier adaptation ability through implicitly generating source features towards target semantics. Specifically, TSA is inspired by the fact that deep feature transformation towards a certain direction can be represented as meaningful semantic altering in the original input space. Thus, source features can be augmented to effectively equip with target semantics to train a more transferable classifier. To achieve this, for each class, we first use the inter-domain feature mean difference and target intra-class feature covariance to construct a multivariate normal distribution. Then we augment source features with random directions sampled from the distribution class-wisely. Interestingly, such source augmentation is implicitly implemented through an expected transferable cross-entropy loss over the augmented source distribution, where an upper bound of the expected loss is derived and minimized, introducing negligible computational overhead. As a light-weight and general technique, TSA can be easily plugged into various domain adaptation methods, bringing remarkable improvements. Comprehensive experiments on cross-domain benchmarks validate the efficacy of TSA.
Shuang Li 0008, Mixue Xie, Kaixiong Gong, Chi Harold Liu, Yulin Wang 0002, Wei Li 0111
CVPR3
2021 Pareto Domain Adaptation
abstract
Domain adaptation (DA) attempts to transfer the knowledge from a labeled source domain to an unlabeled target domain that follows different distribution from the source. To achieve this, DA methods include a source classification objective to extract the source knowledge and a domain alignment objective to diminish the domain shift, ensuring knowledge transfer. Typically, former DA methods adopt some weight hyper-parameters to linearly combine the training objectives to form an overall objective. However, the gradient directions of these objectives may conflict with each other due to domain shift. Under such circumstances, the linear optimization scheme might decrease the overall objective value at the expense of damaging one of the training objectives, leading to restricted solutions. In this paper, we rethink the optimization scheme for DA from a gradient-based perspective. We propose a Pareto Domain Adaptation (ParetoDA) approach to control the overall optimization direction, aiming to cooperatively optimize all training objectives. Specifically, to reach a desirable solution on the target domain, we design a surrogate loss mimicking target classification. To improve target-prediction accuracy to support the mimicking, we propose a target-prediction refining mechanism which exploits domain labels via Bayes’ theorem. On the other hand, since prior knowledge of weighting schemes for objectives is often unavailable to guide optimization to approach the optimal solution on the target domain, we propose a dynamic preference mechanism to dynamically guide our cooperative optimization by the gradient of the surrogate loss on a held-out unlabeled target dataset. Our theoretical analyses show that the held-out data can guide but will not be over-fitted by the optimization. Extensive experiments on image classification and semantic segmentation benchmarks demonstrate the effectiveness of ParetoDA
Fangrui Lv, Jian Liang 0002, Kaixiong Gong, Shuang Li 0008, Chi Harold Liu, Han Li 0005, Di Liu 0029, Guoren Wang
NeurIPS3