Nuo Chen 0002

dblp:135/5622-2 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0001-6563-1215ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ExtendAttack: Attacking Servers of LRMs via Extending Reasoning
abstract
Large Reasoning Models (LRMs) have demonstrated promising performance in complex tasks. However, the resource-consuming reasoning processes may be exploited by attackers to maliciously occupy the resources of the servers, leading to a crash, like the DDoS attack in cyber. To this end, we propose a novel attack method on LRMs termed ExtendAttack to maliciously occupy the resources of servers by stealthily extending the reasoning processes of LRMs. Concretely, we systematically obfuscate characters within a benign prompt, transforming them into a complex, poly-base ASCII representation. This compels the model to perform a series of computationally intensive decoding sub-tasks that are deeply embedded within the semantic structure of the query itself. Extensive experiments demonstrate the effectiveness of our proposed ExtendAttack. Remarkably, it significantly increases response length and latency, with the former increasing by over 2.7 times for the o3 model on the HumanEval benchmark. Besides, it preserves the original meaning of the query and achieves comparable answer accuracy, showing the stealthiness.
Zhenhao Zhu, Yue Liu 0008, Yingwei Ma, Hongcheng Gao, Nuo Chen 0002, Yanpei Guo, Wenjie Qu 0001, Zifeng Kang, Xinzhong Zhu, Jiaheng Zhang
AAAI6
2026 XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration
abstract
Nuo Chen, Andre Lin HuiKai, Jiaying Wu, Junyi Hou, Zining Zhang, Qian Wang, Xidong Wang, Bingsheng He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Nuo Chen 0002, Andre Huikai Lin, Junyi Hou, Zining Zhang 0001, Qian Wang 0002, Xidong Wang, Bingsheng He
ACL (1)1
2026 MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
abstract
Xueqing Peng, Lingfei Qian, Yan Wang, Ruoyu Xiang, Yueru He, Yang Ren, Mingyang Jiang, Vincent Jim Zhang, Yuqing Guo, Jeff Zhao, Huan He, Yi Han, Yun Feng, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Xiaoyu Wang, Penglei Gao, Shengyuan Lin, Keyi Wang, Shanshan Yang, Yilun Zhao, Zhiwei Liu, Peng Lu, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen, Junichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xueqing Peng, Lingfei Qian, Yan Wang 0015, Ruoyu Xiang, Yueru He, Mingyang Jiang, Vincent Jim Zhang, Jeff Zhao, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Penglei Gao, Shengyuan Lin, Yilun Zhao 0001, Zhiwei Liu 0003, Peng Lu 0006, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen 0002, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E. Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen 0003, Jun'ichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie
ACL (1)31
2025 Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts
abstract
Adapting medical Large Language Models to local languages can reduce barriers to accessing healthcare services, but data scarcity remains a significant challenge, particularly for low-resource languages. To address this, we first construct a high-quality medical dataset and conduct analysis to ensure its quality. In order to leverage the generalization capability of multilingual LLMs to efficiently scale to more resource-constrained languages, we explore the internal information flow of LLMs from a multilingual perspective using Mixture of Experts (MoE) modularity. Technically, we propose a novel MoE routing method that employs language-specific experts and cross-lingual routing. Inspired by circuit theory, our routing analysis revealed a \textit{``Spread Out in the End``} information flow mechanism: while earlier layers concentrate cross-lingual information flow, the later layers exhibit language-specific divergence. This insight directly led to the development of the Post-MoE architecture, which applies sparse routing only in the later layers while maintaining dense others. Experimental results demonstrate that this approach enhances the generalization of multilingual models to other languages while preserving interpretability. Finally, to efficiently scale the model to 50 languages, we introduce the concept of \textit{language family} experts, drawing on linguistic priors, which enables scaling the number of languages without adding additional parameters.
Guorui Zheng, Xidong Wang, Juhao Liang, Nuo Chen 0002, Yuping Zheng, Benyou Wang
ICLR4
2025 Is Your LLM Outdated? A Deep Look at Temporal Generalization
abstract
Chenghao Zhu, Nuo Chen, Yufei Gao, Yunyi Zhang, Prayag Tiwari, Benyou Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
ChenghaoZhu ChenghaoZhu, Nuo Chen 0002, Prayag Tiwari, Benyou Wang
NAACL (Long Papers)2
2025 MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
abstract
Wentao Ge, Shunian Chen, Hardy Chen, Nuo Chen, Junying Chen, Zhihong Chen, Wenya Xie, Shuo Yan, Chenghao Zhu, Ziyue Lin, Dingjie Song, Xidong Wang, Anningzhe Gao, Zhang Zhiyi, Jianquan Li, Xiang Wan, Benyou Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Wentao Ge, Shunian Chen, Hardy Chen, Nuo Chen 0002, Wenya Xie, ChenghaoZhu ChenghaoZhu, Ziyue Lin, Dingjie Song, Xidong Wang, Anningzhe Gao, Zhiyi Zhang 0007, Benyou Wang
NAACL (Long Papers)4
2024 Make Prompt-based Black-Box Tuning Colorful: Boosting Model Generalization from Three Orthogonal Perspectives
abstract
Large language models (LLMs) have shown increasing power on various natural language processing (NLP) tasks. However, tuning these models for downstream tasks usually needs exorbitant costs or is unavailable due to commercial considerations. Recently, black-box tuning has been proposed to address this problem by optimizing task-specific prompts without accessing the gradients and hidden representations. However, most existing works have yet fully exploited the potential of gradient-free optimization under the scenario of few-shot learning. In this paper, we describe BBT-RGB, a suite of straightforward and complementary techniques for enhancing the efficiency and performance of black-box optimization. Specifically, our method includes three plug-and-play components: (1) Two-stage derivative-free optimization strategy that facilitates fast convergence and mitigates overfitting; (2) Automatic verbalizer construction with its novel usage under few-shot settings; (3) Better prompt initialization policy based on instruction search and auto-selected demonstration. Extensive experiments across various tasks on natural language understanding and inference demonstrate the effectiveness of our method. Our codes are available at https://github.com/QiushiSun/BBT-RGB.
Qiushi Sun, Chengcheng Han 0004, Nuo Chen 0002, Renyu Zhu, Jingyang Gong, Xiang Li 0067, Ming Gao 0001
LREC/COLING3
2024 TransCoder: Towards Unified Transferable Code Representation Learning Inspired by Human Skills
abstract
Code pre-trained models (CodePTMs) have recently demonstrated a solid capacity to process various code intelligence tasks, e.g., code clone detection, code translation, and code summarization. The current mainstream method that deploys these models to downstream tasks is to fine-tune them on individual tasks, which is generally costly and needs sufficient data for large models. To tackle the issue, in this paper, we present TransCoder, a unified Transferable fine-tuning strategy for Code representation learning. Inspired by human inherent skills of knowledge generalization, TransCoder drives the model to learn better code-related knowledge like human programmers. Specifically, we employ a tunable prefix encoder to first capture cross-task and cross-language transferable knowledge, subsequently applying the acquired knowledge for optimized downstream adaptation. Besides, our approach confers benefits for tasks with minor training sample sizes and languages with smaller corpora, underscoring versatility and efficacy. Extensive experiments conducted on representative datasets clearly demonstrate that our method can lead to superior performance on various code-related tasks and encourage mutual reinforcement, especially in low-resource scenarios. Our codes are available at https://github.com/QiushiSun/TransCoder.
Qiushi Sun, Nuo Chen 0002, Jianing Wang 0002, Ming Gao 0001, Xiang Li 0067
LREC/COLING2
2024 Structure-aware Fine-tuning for Code Pre-trained Models
abstract
Over the past few years, we have witnessed remarkable advancements in Code Pre-trained Models (CodePTMs). These models achieved excellent representation capabilities by designing structure-based pre-training tasks for code. However, how to enhance the absorption of structural knowledge when fine-tuning CodePTMs still remains a significant challenge. To fill this gap, in this paper, we present SAT, a novel structure-enhanced and plug-and-play fine-tuning method for CodePTMs. We first propose a structure loss to quantify the difference between the information learned by CodePTMs and the knowledge extracted from code structure. Specifically, we use the attention scores from Transformer layer as the learned information, and the shortest path length between leaves in abstract syntax trees as the structural knowledge. Subsequently, multi-task learning is introduced to improve the performance of fine-tuning. Experiments conducted on four pre-trained models and two generation tasks demonstrate the effectiveness of our proposed method as a plug-and-play solution. Furthermore, we observed that SAT can benefit CodePTMs more with limited training data.
Jiayi Wu 0001, Renyu Zhu, Nuo Chen 0002, Qiushi Sun, Xiang Li 0067, Ming Gao 0001
LREC/COLING3
2024 CryptoTrade: A Reflective LLM-based Agent to Guide Zero-shot Cryptocurrency Trading
abstract
The utilization of Large Language Models (LLMs) in financial trading has primarily been concentrated within the stock market, aiding in economic and financial decisions.Yet, the unique opportunities presented by the cryptocurrency market, noted for its on-chain data's transparency and the critical influence of offchain signals like news, remain largely untapped by LLMs.This work aims to bridge the gap by developing an LLM-based trading agent, CryptoTrade, which uniquely combines the analysis of on-chain and off-chain data.This approach leverages the transparency and immutability of on-chain data, as well as the timeliness and influence of off-chain signals, providing a comprehensive overview of the cryptocurrency market.CryptoTrade incorporates a reflective mechanism specifically engineered to refine its daily trading decisions by analyzing the outcomes of prior trading decisions.This research makes two significant contributions.Firstly, it broadens the applicability of LLMs to the domain of cryptocurrency trading.Secondly, it establishes a benchmark for cryptocurrency trading strategies.Through extensive experiments, CryptoTrade has demonstrated superior performance in maximizing returns compared to time-series baselines, but not compared to traditional trading signals, across various cryptocurrencies and market conditions.Our code and data are available at https://github. com/Xtra-Computing/CryptoTrade.CryptoTrade makes day-to-day trading decisions.
Yuan Li 0032, Bingqiao Luo, Qian Wang 0002, Nuo Chen 0002, Xu Liu 0014, Bingsheng He
EMNLP4
2024 Rethinking the Role of Structural Information: How It Enhances Code Representation Learning?
abstract
Code pre-trained models (CodePTMs) have recently exhibited remarkable accomplishments in the realm of software engineering. However, there are still limited advancements in understanding the inner mechanism of these models, as well as their sensitivity to samples of varying quality. Codes have a more rigid and structured syntax compared to natural languages; hence, leveraging and understanding structural information becomes essential for analyzing, interpreting, and utilizing CodePTMs. While previous studies have verified models’ ability to acquire knowledge from code structure through techniques such as attention analysis and probing tasks, the specific roles it plays in downstream tasks have yet to be explored. In this work, we propose a set of novel and practical methods for probing and exploiting the structural information within the code. In particular, dataflow perturbation experiments are first employed to explore the sensitivity of models with varying levels of structural information when confronted with input changes. Based on our findings, structure-aware exemplars selection strategies are proposed for both code generation and understanding, aiming to recover the model performance at minimal cost under perturbed conditions. Moreover, efficient fine-tuning can be achieved by utilizing exemplars instead of full fine-tuning.
Qiushi Sun, Nuo Chen 0002, Jianing Wang 0002, Xiaoli Li 0001
IJCNN2
2023 HugNLP: A Unified and Comprehensive Library for Natural Language Processing
abstract
In this paper, we introduce HugNLP, a unified and comprehensive library for natural language processing (NLP) with the prevalent backend of Hugging Face Transformers, which is designed for NLP researchers to easily utilize off-the-shelf algorithms and develop novel methods with user-defined models and tasks in real-world scenarios. HugNLP consists of a hierarchical structure including models, processors and applications that unifies the learning process of pre-trained language models (PLMs) on different NLP tasks. Additionally, we present some featured NLP applications to show the effectiveness of HugNLP, such as knowledge-enhanced PLMs, universal information extraction, low-resource mining, and code understanding and generation, etc. The source code will be released on GitHub (https://github.com/HugAILab/HugNLP).
Jianing Wang 0002, Nuo Chen 0002, Qiushi Sun, Wenkang Huang, Chengyu Wang 0001, Ming Gao 0001
CIKM2