Yuxiang Nie

dblp:247/9594 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0009-0001-4197-1079ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
abstract
Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers provide no learning signal, particularly in challenging tasks. To address this,we propose Multi-Expert Mutual Learning GRPO (MEML-GRPO), an innovative framework that utilizes diverse expert prompts as system prompts to generate a broader range of responses, substantially increasing the likelihood of identifying correct solutions. Additionally, we introduce an inter-expert mutual learning mechanism that facilitates knowledge sharing and transfer among experts, further boosting the model’s performance through RLVR. Extensive experiments across multiple reasoning benchmarks show that MEML-GRPO delivers significant improvements, achieving an average performance gain of 4.89% with Qwen and 11.33% with Llama, effectively overcoming the core limitations of traditional RLVR methods.
Weitao Jia, Jinghui Lu, Haiyang Yu 0004, Guozhi Tang, An-Lan Wang, Weijie Yin, Dingkang Yang, Yuxiang Nie, Bin Shan, Hao Feng 0009, Irene Li, Kun Yang 0010, Jingqun Tang, Teng Fu 0001, Changhong Jin, Xiaohui Lv, Can Huang 0002
AAAI9
2025 Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
abstract
The application of Large Vision-Language Models (LVLMs) for analyzing images and videos is an exciting and rapidly evolving field. In recent years, we've seen significant growth in high-quality image-text datasets for fine-tuning image understanding, but there is still a lack of comparable datasets for videos. Additionally, many VideoLLMs are extensions of single-image VLMs, which may not efficiently handle the complexities of longer videos. In this study, we introduce a large-scale synthetic dataset created from proprietary models, using carefully designed prompts to tackle a wide range of questions. We also explore a dynamic visual token compression architecture that strikes a balance between computational efficiency and performance. Our proposed \model{} achieves state-of-the-art results across various video tasks and shows impressive generalization, setting new baselines in multi-image understanding. Notably, \model{} delivers an absolute improvement of 2.7\% over LLaVA-OneVision on VideoMME and 10.7\% on MuirBench. Codes are available at https://github.com/Hon-Wong/ByteVideoLLM
Yuxiang Nie, Yongjie Ye, Haiyang Yu 0004, Jinghui Lu, Can Huang 0002
ICCV2
2024 SciMRC: Multi-perspective Scientific Machine Reading Comprehension
abstract
Scientific Machine Reading Comprehension (SMRC) aims to facilitate the understanding of scientific texts through human-machine interactions. While existing dataset has significantly contributed to this field, it predominantly focus on single-perspective question-answer pairs, thereby overlooking the inherent variation in comprehension levels among different readers. To address this limitation, we introduce a novel multi-perspective scientific machine reading comprehension dataset, SciMRC, which incorporates perspectives from beginners, students, and experts. Our dataset comprises 741 scientific papers and 6,057 question-answer pairs, with 3,306, 1,800, and 951 pairs corresponding to beginners, students, and experts respectively. Extensive experiments conducted on SciMRC using pre-trained models underscore the importance of considering diverse perspectives in SMRC and highlight the challenging nature of our scientific machine comprehension tasks.
Xiao Zhang 0036, Heqi Zheng, Yuxiang Nie, Heyan Huang, Xianling Mao
LREC/COLING3
2024 Elysium: Exploring Object-Level Perception in Videos via MLLM
Yongjie Ye, Yuxiang Nie
ECCV (22)4
2024 Mix-Initiative Response Generation with Dynamic Prefix Tuning
abstract
Yuxiang Nie, Heyan Huang, Xian-Ling Mao, Lizi Liao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yuxiang Nie, Heyan Huang, Xianling Mao, Lizi Liao
NAACL-HLT1
2023 Adapting Object Size Variance and Class Imbalance for Semi-supervised Object Detection
abstract
Semi-supervised object detection (SSOD) attracts extensive research interest due to its great significance in reducing the data annotation effort. Collecting high-quality and category-balanced pseudo labels for unlabeled images is critical to addressing the SSOD problem. However, most of the existing pseudo-labeling-based methods depend on a large and fixed threshold to select high-quality pseudo labels from the predictions of a teacher model. Considering different object classes usually have different detection difficulty levels due to scale variance and data distribution imbalance, conventional pseudo-labeling-based methods are arduous to explore the value of unlabeled data sufficiently. To address these issues, we propose an adaptive pseudo labeling strategy, which can assign thresholds to classes with respect to their “hardness”. This is beneficial for ensuring the high quality of easier classes and increasing the quantity of harder classes simultaneously. Besides, label refinement modules are set up based on box jittering for guaranteeing the localization quality of pseudo labels. To further improve the algorithm’s robustness against scale variance and make the most of pseudo labels, we devise a joint feature-level and prediction-level consistency learning pipeline for transferring the information of the teacher model to the student model. Extensive experiments on COCO and VOC datasets indicate that our method achieves state-of-the-art performance. Especially, it brings mean average precision gains of 2.08 and 1.28 on MS-COCO dataset with 5% and 10% labeled images, respectively.
Yuxiang Nie, Chaowei Fang, Lechao Cheng, Liang Lin 0004, Guanbin Li
AAAI1
2023 Reinforced Target-driven Conversational Promotion
abstract
The ability to proactively engage with users towards pitching products is highly desired for conversational assistants.However, existing conversational recommendation methods overemphasize on acquiring user preferences while ignore the strategic planning for nudging users towards accepting a designated item.Hence, these methods fail to promote specified items with engaging responses.In this work, we propose a Reinforced Target-driven Conversational Promotion (RTCP) framework for conversational promotion.Specifically, RTCP integrates short-term and long-term planning via a balanced gating mechanism.Inside which, the dialogue strategies are predicted via knowledge-integrated multi-head attention and guided via reinforcement learning rewards.RTCP then employs an action-guided prefix tuning method to generate relevant responses.Experimental results demonstrate that our model outperforms state-of-the-art models on both automatic metrics and human evaluation.Moreover, RTCP has a strong capability in quickly adapting to unseen scenarios just by updating prefix parameters without re-training the whole model.Code and data are here 1 .
Huy Dao, Lizi Liao, Dung D. Le, Yuxiang Nie
EMNLP4
2023 Overview of NLPCC Shared Task 2: Multi-perspective Scientific Machine Reading Comprehension
Xiao Zhang 0036, Heqi Zheng, Yuxiang Nie, Xianling Mao
NLPCC (3)3
2022 Unsupervised Question Answering via Answer Diversifying
abstract
Unsupervised question answering is an attractive task due to its independence on labeled data. Previous works usually make use of heuristic rules as well as pre-trained models to construct data and train QA models. However, most of these works regard named entity (NE) as the only answer type, which ignores the high diversity of answers in the real world. To tackle this problem, we propose a novel unsupervised method by diversifying answers, named DiverseQA. Specifically, the proposed method is composed of three modules: data construction, data augmentation and denoising filter. Firstly, the data construction module extends the extracted named entity into a longer sentence constituent as the new answer span to construct a QA dataset with diverse answers. Secondly, the data augmentation module adopts an answer-type dependent data augmentation process via adversarial training in the embedding level. Thirdly, the denoising filter module is designed to alleviate the noise in the constructed data. Extensive experiments show that the proposed method outperforms previous unsupervised models on five benchmark datasets, including SQuADv1.1, NewsQA, TriviaQA, BioASQ, and DuoRC. Besides, the proposed method shows strong performance in the few-shot learning setting.
Yuxiang Nie, Heyan Huang, Zewen Chi, Xianling Mao
COLING1
2022 Capturing Global Structural Information in Long Document Question Answering with Compressive Graph Selector Network
abstract
Long document question answering is a challenging task due to its demands for complex reasoning over long text.Previous works usually take long documents as non-structured flat texts or only consider the local structure in long documents.However, these methods usually ignore the global structure of the long document, which is essential for long-range understanding.To tackle this problem, we propose Compressive Graph Selector Network (CGSN) to capture the global structure in a compressive and iterative manner.The proposed model mainly focuses on the evidence selection phase of long document question answering.Specifically, it consists of three modules: local graph network, global graph network and evidence memory network.Firstly, the local graph network builds the graph structure of the chunked segment in token, sentence, paragraph and segment levels to capture the short-term dependency of the text.Secondly, the global graph network selectively receives the information of each level from the local graph, compresses them into the global graph nodes and applies graph attention to the global graph nodes to build the long-range reasoning over the entire text in an iterative way.Thirdly, the evidence memory network is designed to alleviate the redundancy problem in the evidence selection by saving the selected result in the previous steps.Extensive experiments show that the proposed model outperforms previous methods on two datasets.
Yuxiang Nie, Heyan Huang, Wei Wei 0002, Xianling Mao
EMNLP1
2022 Double-Check Soft Teacher for Semi-Supervised Object Detection
abstract
In the semi-supervised object detection task, due to the scarcity of labeled data and the diversity and complexity of objects to be detected, the quality of pseudo-labels generated by existing methods for unlabeled data is relatively low, which severely restricts the performance of semi-supervised object detection. In this paper, we revisit the pseudo-labeling based Teacher-Student mutual learning framework for semi-supervised object detection and identify that the inconsistency of the location and feature of the candidate object proposals between the Teacher and the Student branches are the fatal cause of the low quality of the pseudo labels. To address this issue, we propose a simple yet effective technique within the mainstream teacher-student framework, called Double Check Soft Teacher, to overcome the harm caused by insufficient quality of pseudo labels. Specifically, our proposed method leverages teacher model to generate pseudo labels for the student model. Especially, the candidate boxes generated by the student model based on the pseudo label will be sent to the teacher model for "double check", and then the teacher model will output probabilistic soft label with background class for those candidate boxes, which will be used to train the student model. Together with a pseudo labeling mechanism based on the sum of the TOP-K prediction score, which improves the recall rate of pseudo labels, Double Check Soft Teacher consistently surpasses state-of-the-art methods by significant margins on the MS-COCO benchmark, pushing the new state-of-the-art. Source codes are available at https://github.com/wkfdb/DCST.
Yuxiang Nie, Chaowei Fang, Chengzhi Han, Xuewen Wu, Liang Lin 0004, Guanbin Li
IJCAI2