Xiaoyu Shen 0001

dblp:85/7634-1 · DBLP profile ↗
← Back
52ranked-venue papers
9as first author
35since 2021 · last 2026
0000-0002-0217-2469ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 9 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Learning A Bank of Transferable Prompts for Vision-Language Models
Zhongwei Huang, Chong Wang 0001, Endai Huang, Ran Zhou 0002, Haitao Gan, Yingying Zhu 0001, Xiaoyu Shen 0001
ICMR8
2026 Understanding domain-specific attribute constraints in multimodal sentiment analysis via ensemble multimodal large language models with structured multistage prompts
Rongfei Chen, Junlong Tong, Xiaoyu Shen 0001, Wei Zhang 0185
Expert Syst. Appl.5
2025 InternLM-Law: An Open-Sourced Chinese Legal Large Language Model
abstract
We introduce InternLM-Law, a large language model (LLM) tailored for addressing diverse legal tasks related to Chinese laws. These tasks range from responding to standard legal questions (e.g., legal exercises in textbooks) to analyzing complex real-world legal situations. Our work contributes to Chinese Legal NLP research by (1) conducting one of the most extensive evaluations of state-of-the-art general-purpose and legal-specific LLMs to date that involves an automatic evaluation on the 20 legal NLP tasks in LawBench, a human evaluation on a challenging version of the Legal Consultation task, and an automatic evaluation of a model’s ability to handle very long legal texts; (2) presenting a methodology for training a Chinese legal LLM that offers superior performance to all of its counterparts in our extensive evaluation; and (3) facilitating future research in this area by making all of our code and model publicly available at https://github.com/InternLM/InternLM-Law.
Zhiwei Fei, Songyang Zhang 0001, Xiaoyu Shen 0001, Xiao Wang 0042, Jidong Ge, Vincent Ng 0001
COLING3
2025 Multimodal Language Models See Better When They Look Shallower
abstract
Multimodal large language models (MLLMs) typically extract visual features from the final layers of a pretrained Vision Transformer (ViT).This widespread deep-layer bias, however, is largely driven by empirical convention rather than principled analysis.While prior studies suggest that different ViT layers capture different types of information-shallower layers focusing on fine visual details and deeper layers aligning more closely with textual semantics, the impact of this variation on MLLM performance remains underexplored.We present the first comprehensive study of visual layer selection for MLLMs, analyzing representation similarity across ViT layers to establish shallow, middle, and deep layer groupings.Through extensive evaluation of MLLMs (1.4B-7B parameters) across 10 benchmarks encompassing 60+ tasks, we find that while deep layers excel in semantic-rich tasks like OCR, shallow and middle layers significantly outperform them on fine-grained visual tasks including counting, positioning, and object localization.Building on these insights, we propose a lightweight feature fusion method that strategically incorporates shallower layers, achieving consistent improvements over both single-layer and specialized fusion baselines.Our work offers the first principled study of visual layer selection in MLLMs, showing that MLLMs can often see better when they look shallower.
Junyan Lin, Xinghao Chen 0009, Jianfeng Dong, Xin Jin 0014, Hui Su, Jinlan Fu, Xiaoyu Shen 0001
EMNLP9
2025 VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs
abstract
Multimodal Large Language Models (MLLMs) have achieved strong performance across vision-language tasks, but suffer from significant computational overhead due to the quadratic growth of attention computations with the number of multimodal tokens.Though efforts have been made to prune tokens in MLLMs, they lack a fundamental understanding of how MLLMs process and fuse multimodal information.Through systematic analysis, we uncover a three-stage cross-modal interaction process: (1) Shallow layers recognize task intent, with visual tokens acting as passive attention sinks; (2) Cross-modal fusion occurs abruptly in middle layers, driven by a few critical visual tokens; (3) Deep layers discard vision tokens, focusing solely on linguistic refinement.Based on these findings, we propose VisiPruner, a training-free pruning framework that reduces up to 99% of visionrelated attention computations and 53.9% of FLOPs on LLaVA-v1.5 7B.It significantly outperforms existing token pruning methods and generalizes across diverse MLLMs.Beyond pruning, our insights further provide actionable guidelines for training efficient MLLMs by aligning model architecture with its intrinsic layer-wise processing dynamics.
Yingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong, Hui Su, Yijie Pan, Wei Zhang 0185, Xiaoyu Shen 0001
EMNLP8
2025 PricingLogic: Evaluating LLMs Reasoning on Complex Tourism Pricing Tasks
abstract
We present PricingLogic, the first benchmark that probes whether Large Language Models (LLMs) can reliably automate tourism-related prices when multiple, overlapping fare rules apply.Travel agencies are eager to offload this error-prone task onto AI systems; however, deploying LLMs without verified reliability could result in significant financial losses and erode customer trust.PricingLogic comprises 300 natural-language questions based on booking requests derived from 42 real-world pricing policies, spanning two levels of difficulty: (i) basic customer-type pricing and (ii) bundled-tour calculations involving interacting discounts.Evaluations of a line of LLMs reveal a steep performance drop on the harder tier, exposing systematic failures in rule interpretation and arithmetic reasoning.These results highlight that, despite their general capabilities, today's LLMs remain unreliable in revenuecritical applications without further safeguards or domain adaptation.
Yunuo Liu, Zena Al-Khalili, Dai Cheng, Yanjun Chen 0001, Dietrich Klakow, Wei Zhang 0185, Xiaoyu Shen 0001
EMNLP8
2025 Context Guided Transformer Entropy Modeling for Video Compression
abstract
Conditional entropy models effectively leverage spatio-temporal contexts to reduce video redundancy. However, incorporating temporal context often introduces additional model complexity and increases computational cost. In parallel, many existing spatial context models lack explicit modeling the ordering of spatial dependencies, which may limit the availability of relevant context during decoding. To address these issues, we propose the Context Guided Transformer (CGT) entropy model, which estimates probability mass functions of the current frame conditioned on resampled temporal context and dependency-weighted spatial context. A temporal context resampler learns predefined latent queries to extract critical temporal information using transformer encoders, reducing downstream computational overhead. Meanwhile, a teacher-student network is designed as dependency-weighted spatial context assigner to explicitly model the dependency of spatial context order. The teacher generates an attention map to represent token importance and an entropy map to reflect prediction certainty from randomly masked inputs, guiding the student to select the weighted top-k tokens with the highest spatial dependency. During inference, only the student is used to predict undecoded tokens based on high-dependency context. Experimental results demonstrate that our CGT model reduces entropy modeling time by approximately 65% and achieves an 11% BD-Rate reduction compared to the previous state-of-the-art conditional entropy model.
Junlong Tong, Wei Zhang 0185, Yaohui Jin, Xiaoyu Shen 0001
ICCV4
2025 CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
abstract
Multimodal Large Language Models (MLLMs) still struggle with hallucinations despite their impressive capabilities. Recent studies have attempted to mitigate this by applying Direct Preference Optimization (DPO) to multimodal scenarios using preference pairs from text-based responses. However, our analysis of representation distributions reveals that multimodal DPO struggles to align image and text representations and to distinguish between hallucinated and non-hallucinated descriptions. To address these challenges, In this work, we propose a Cross-modal Hierarchical Direct Preference Optimization (CHiP) to address these limitations. We introduce a visual preference optimization module within the DPO framework, enabling MLLMs to learn from both textual and visual preferences simultaneously. Furthermore, we propose a hierarchical textual preference optimization module that allows the model to capture preferences at multiple granular levels, including response, segment, and token levels. We evaluate CHiP through both quantitative and qualitative analyses, with results across multiple benchmarks demonstrating its effectiveness in reducing hallucinations. On the Object HalBench dataset, CHiP outperforms DPO in hallucination reduction, achieving improvements of 52.7% and 55.5% relative points based on the base model Muffin and LLaVA models, respectively. We make all our datasets and code publicly available.
Jinlan Fu, Shenzhen Huangfu, Hao Fei 0001, Xiaoyu Shen 0001, Bryan Hooi, Xipeng Qiu, See-Kiong Ng
ICLR4
2025 Beyond Content Relevance: Evaluating Instruction Following in Retrieval Models
abstract
Instruction-following capabilities in large language models (LLMs) have progressed significantly, enabling more complex user interactions through detailed prompts. However, retrieval systems have not matched these advances, most of them still relies on traditional lexical and semantic matching techniques that fail to fully capture user intent. Recent efforts have introduced instruction-aware retrieval models, but these primarily focus on intrinsic content relevance, which neglects the importance of customized preferences for broader document-level attributes. This study evaluates the instruction-following capabilities of various retrieval models beyond content relevance, including LLM-based dense retrieval and reranking models. We develop InfoSearch, a novel retrieval evaluation benchmark spanning six document-level attributes: Audience, Keyword, Format, Language, Length, and Source, and introduce novel metrics -- Strict Instruction Compliance Ratio (SICR) and Weighted Instruction Sensitivity Evaluation (WISE) to accurately assess the models' responsiveness to instructions. Our findings indicate that although fine-tuning models on instruction-aware retrieval datasets and increasing model size enhance performance, most models still fall short of instruction compliance. We release our dataset and code on https://github.com/EIT-NLP/InfoSearch.
Jianqun Zhou, Yuanlei Zheng, Zeyuan Shang, Wei Zhang 0185, Xiaoyu Shen 0001
ICLR8
2025 SkipGPT: Each Token is One of a Kind
abstract
Large language models (LLMs) achieve remarkable performance across tasks but incur substantial computational costs due to their deep, multi-layered architectures. Layer pruning has emerged as a strategy to alleviate these inefficiencies, but conventional static pruning methods overlook two critical dynamics inherent to LLM inference: (1) *horizontal dynamics*, where token-level heterogeneity demands context-aware pruning decisions, and (2) *vertical dynamics*, where the distinct functional roles of MLP and self-attention layers necessitate component-specific pruning policies. We introduce **SkipGPT**, a dynamic layer pruning framework designed to optimize computational resource allocation through two core innovations: (1) global token-aware routing to prioritize critical tokens and (2) decoupled pruning policies for MLP and self-attention components. To mitigate training instability, we propose a two-stage optimization paradigm: first, a disentangled training phase that learns routing strategies via soft parameterization to avoid premature pruning decisions, followed by parameter-efficient LoRA fine-tuning to restore performance impacted by layer removal. Extensive experiments demonstrate that SkipGPT reduces over 40% model parameters while matching or exceeding the performance of the original dense model across benchmarks. By harmonizing dynamic efficiency with preserved expressivity, SkipGPT advances the practical deployment of scalable, resource-aware LLMs. Our code is publicly available at: https://github.com/EIT-NLP/SkipGPT.
Anhao Zhao, Fanghua Ye 0001, Yingqi Fan, Junlong Tong, Zhiwei Fei, Hui Su, Xiaoyu Shen 0001
ICML8
2025 MAER-Nav: Bidirectional Motion Learning Through Mirror-Augmented Experience Replay for Robot Navigation
abstract
Deep Reinforcement Learning (DRL) based navigation methods have demonstrated promising results for mobile robots, but suffer from limited action flexibility in confined spaces. Conventional DRL approaches predominantly learn forward-motion policies, causing robots to become trapped in complex environments where backward maneuvers are necessary for recovery. This paper presents MAER-Nav (Mirror-Augmented Experience Replay for Robot Navigation), a novel framework that enables bidirectional motion learning without requiring explicit failure-driven hindsight experience replay or reward function modifications. Our approach integrates a mirror-augmented experience replay mechanism with curriculum learning to generate synthetic backward navigation experiences from successful trajectories. Experimental results in both simulation and real-world environments demonstrate that MAER-Nav significantly outperforms state-of-the-art methods while maintaining strong forward navigation capabilities. The framework effectively bridges the gap between the comprehensive action space utilization of traditional planning methods and the environmental adaptability of learning-based approaches, enabling robust navigation in scenarios where conventional DRL methods consistently fail.
Shanze Wang, Mingao Tan, Biao Huang 0016, Xiaoyu Shen 0001, Hailong Huang 0001, Wei Zhang 0185
IROS5
2025 Enhancing Deep Reinforcement Learning-based Robot Navigation Generalization through Scenario Augmentation
abstract
This work focuses on enhancing the generalization performance of deep reinforcement learning-based robot navigation in unseen environments. We present a novel data augmentation approach called scenario augmentation, which enables robots to navigate effectively across diverse settings without altering the training scenario. The method operates by mapping the robot’s observation into an imagined space, generating an imagined action based on this transformed observation, and then remapping this action back to the real action executed in simulation. Through scenario augmentation, we conduct extensive comparative experiments to investigate the underlying causes of suboptimal navigation behaviors in unseen environments. Our analysis indicates that limited training scenarios represent the primary factor behind these undesired behaviors. Experimental results confirm that scenario augmentation substantially enhances the generalization capabilities of deep reinforcement learning-based navigation systems. The improved navigation framework demonstrates exceptional performance by producing near-optimal trajectories with significantly reduced navigation time in real-world applications.
Shanze Wang, Mingao Tan, Xianghui Wang, Xiaoyu Shen 0001, Hailong Huang 0001, Wei Zhang 0185
IROS5
2025 MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text Generation
Jinlan Fu, Shenzhen Huangfu, Hao Fei 0001, Yichong Huang, Xiaoyu Shen 0001, Xipeng Qiu, See-Kiong Ng
ACM Multimedia5
2025 Large Language Models Empowered Personalized Web Agents
abstract
Web agents have emerged as a promising direction to automate Web task completion based on user instructions, significantly enhancing user experience. Recently, Web agents have evolved from traditional agents to Large Language Models (LLMs)-based Web agents. Despite their success, existing LLM-based Web agents overlook the importance of personalized data (e.g., user profiles and historical Web behaviors) in assisting the understanding of users' personalized instructions and executing customized actions.
Hongru Cai, Yongqi Li 0001, Wenjie Wang 0007, Fengbin Zhu, Xiaoyu Shen 0001, Wenjie Li 0002, Tat-Seng Chua
WWW5
2025 Canvas: Compositional Generation for Art Painting With Seamless Subject-Driven Infusion
Yunnan Wang, Lexiang Lv, Zequn Zhang, Xiaoyu Shen 0001, Xin Jin 0014, Wenjun Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Experimental Evaluation of Parameter-Efficient Fine-Tuning for Software Engineering Tasks
abstract
Pre-trained models (PTMs) have succeeded in various software engineering (SE) tasks following the “pre-train then fine-tune” paradigm. As fully fine-tuning all parameters of PTMs can be computationally expensive, a potential solution is parameter-efficient fine-tuning (PEFT), which freezes PTMs while introducing extra parameters. Although PEFT methods have been applied to SE tasks, researchers often focus on specific scenarios and lack a comprehensive comparison of PTMs from different aspects such as field, size, and architecture. To fill this gap, we have conducted an empirical study on six PEFT methods, eight PTMs, and four SE tasks. The experimental results reveal several noteworthy findings. For example, model architecture has little impact on PTM performance when using PEFT methods. Additionally, we provide a comprehensive discussion of PEFT methods from three perspectives. First, we analyze the effectiveness and efficiency of PEFT methods. Second, we explore the impact of the scaling factor hyperparameter. Finally, we investigate the application of PEFT methods on the latest open source large language model, Llama 3.2. These findings provide valuable insights to guide future researchers in effectively applying PEFT methods to SE tasks.
Wentao Zou, Zongwen Shen, Jidong Ge, Chuanyi Li, Xiang Chen 0005, Xiaoyu Shen 0001, LiGuo Huang, Bin Luo 0003
ACM Trans. Softw. Eng. Methodol.7
2024 SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects
abstract
David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba O. Alabi, Yanke Mao, Haonan Gao, En-Shiun Annie Lee. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen 0001, Nikita Vassilyev, Jesujoba O. Alabi, Yanke Mao, Haonan Gao, Annie En-Shiun Lee
EACL (1)3
2024 The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
abstract
Reinforcement Learning from Human Feedback significantly enhances Natural Language Processing by aligning language models with human expectations.A critical factor in this alignment is the strength of reward models used during training.This study explores whether stronger reward models invariably lead to better language models.In this paper, through experiments on relevance, factuality, and completeness tasks using the QA-FEEDBACK dataset and reward models based on Longformer, we uncover a surprising paradox: language models trained with moderately accurate reward models outperform those guided by highly accurate ones.This challenges the widely held belief that stronger reward models always lead to better language models, and opens up new avenues for future research into the key factors driving model performance and how to choose the most suitable reward models.
Yanjun Chen 0001, Yirong Sun, Xinghao Chen 0009, Wei Zhang 0185, Xiaoyu Shen 0001
EMNLP6
2024 LawBench: Benchmarking Legal Knowledge of Large Language Models
abstract
Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Alan Huang, Songyang Zhang, Kai Chen, Zhixin Yin, Zongwen Shen, Jidong Ge, Vincent Ng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Zhiwei Fei, Xiaoyu Shen 0001, Fengzhe Zhou, Zhuo Han, Alan Huang, Songyang Zhang 0001, Kai Chen 0026, Zhixin Yin, Zongwen Shen, Jidong Ge, Vincent Ng 0001
EMNLP2
2024 To Preserve or To Compress: An In-Depth Study of Connector Selection in Multimodal Large Language Models
abstract
In recent years, multimodal large language models (MLLMs) have garnered significant attention from both industry and academia.However, there is still considerable debate on constructing MLLM architectures, particularly regarding the selection of appropriate connectors for perception tasks of varying granularities.This paper systematically investigates the impact of connectors on MLLM performance.Specifically, we classify connectors into feature-preserving and featurecompressing types.Utilizing a unified classification standard, we categorize sub-tasks from three comprehensive benchmarks, MM-Bench, MME, and SEED-Bench, into three task types: coarse-grained perception, fine-grained perception, and reasoning, and evaluate the performance.Our findings reveal that featurepreserving connectors excel in fine-grained perception tasks due to their ability to retain detailed visual information.In contrast, featurecompressing connectors, while less effective in fine-grained perception tasks, offer significant speed advantages and perform comparably in coarse-grained perception and reasoning tasks.These insights are crucial for guiding MLLM architecture design and advancing the optimization of MLLM architectures.
Junyan Lin, Xiaoyu Shen 0001
EMNLP4
2024 Assessing "Implicit" Retrieval Robustness of Large Language Models
abstract
Retrieval-augmented generation has gained popularity as a framework to enhance large language models with external knowledge.However, its effectiveness hinges on the retrieval robustness of the model.If the model lacks retrieval robustness, its performance is constrained by the accuracy of the retriever, resulting in significant compromises when the retrieved context is irrelevant.In this paper, we evaluate the "implicit" retrieval robustness of various large language models, instructing them to directly output the final answer without explicitly judging the relevance of the retrieved context.Our findings reveal that fine-tuning on a mix of gold and distracting context significantly enhances the model's robustness to retrieval inaccuracies, while still maintaining its ability to extract correct answers when retrieval is accurate.This suggests that large language models can implicitly handle relevant or irrelevant retrieved context by learning solely from the supervision of the final answer in an end-toend manner.Introducing an additional process for explicit relevance judgment can be unnecessary and disrupts the end-to-end approach.1
Xiaoyu Shen 0001, Rexhina Blloshmi, Jiahuan Pei, Wei Zhang 0185
EMNLP1
2024 Unveiling In-Context Learning: A Coordinate System to Understand Its Working Mechanism
abstract
Large language models (LLMs) exhibit remarkable in-context learning (ICL) capabilities.However, the underlying working mechanism of ICL remains poorly understood.Recent research presents two conflicting views on ICL: One emphasizes the impact of similar examples in the demonstrations, stressing the need for label correctness and more shots.The other attributes it to LLMs' inherent ability of task recognition, deeming label correctness and shot numbers of demonstrations as not crucial.In this work, we provide a Two-Dimensional Coordinate System that unifies both views into a systematic framework.The framework explains the behavior of ICL through two orthogonal variables: whether similar examples are presented in the demonstrations (perception) and whether LLMs can recognize the task (cognition).We propose the peak inverse rank metric to detect the task recognition ability of LLMs and study LLMs' reactions to different definitions of similarity.Based on these, we conduct extensive experiments to elucidate how ICL functions across each quadrant on multiple representative classification tasks.Finally, we extend our analyses to generation tasks, showing that our coordinate system can also be used to interpret ICL for generation tasks effectively.
Anhao Zhao, Fanghua Ye 0001, Jinlan Fu, Xiaoyu Shen 0001
EMNLP4
2024 Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?
abstract
Traditionally, success in multilingual machine translation can be attributed to three key factors in training data: large volume, diverse translation directions, and high quality.In the current practice of fine-tuning large language models (LLMs) for translation, we revisit the importance of these factors.We find that LLMs display strong translation capability after being fine-tuned on as few as 32 parallel sentences and that fine-tuning on a single translation direction enables translation in multiple directions.However, the choice of direction is critical: fine-tuning LLMs with only English on the target side can lead to task misinterpretation, which hinders translation into non-English languages.Problems also arise when noisy synthetic data is placed on the target side, especially when the target language is wellrepresented in LLM pre-training.Yet interestingly, synthesized data in an under-represented language has a less pronounced effect.Our findings suggest that when adapting LLMs to translation, the requirement on data quantity can be eased but careful considerations are still crucial to prevent an LLM from exploiting unintended data biases.
Pinzhen Chen, Miaoran Zhang, Barry Haddow, Xiaoyu Shen 0001, Dietrich Klakow
EMNLP5
2024 A Preference-driven Paradigm for Enhanced Translation with Large Language Models
abstract
Dawei Zhu, Sony Trenous, Xiaoyu Shen, Dietrich Klakow, Bill Byrne, Eva Hasler. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Sony Trenous, Xiaoyu Shen 0001, Dietrich Klakow, William J. Byrne, Eva Hasler
NAACL-HLT3
2023 Weaker Than You Think: A Critical Look at Weakly Supervised Learning
abstract
Weakly supervised learning is a popular approach for training machine learning models in low-resource settings.Instead of requesting high-quality yet costly human annotations, it allows training models with noisy annotations obtained from various weak sources.Recently, many sophisticated approaches have been proposed for robust training under label noise, reporting impressive results.In this paper, we revisit the setup of these approaches and find that the benefits brought by these approaches are significantly overestimated.Specifically, we find that the success of existing weakly supervised learning approaches heavily relies on the availability of clean validation samples which, as we show, can be leveraged much more efficiently by simply training on them.After using these clean labels in training, the advantages of using these sophisticated approaches are mostly wiped out.This remains true even when reducing the size of the available clean data to just five samples per class, making these approaches impractical.To understand the true value of weakly supervised learning, we thoroughly analyze diverse NLP datasets and tasks to ascertain when and why weakly supervised approaches work.Based on our findings, we provide recommendations for future research.1
Xiaoyu Shen 0001, Marius Mosbach, Andreas Stephan, Dietrich Klakow
ACL (1)2
2023 Meta Self-Refinement for Robust Learning with Weak Supervision
abstract
Training deep neural networks (DNNs) under weak supervision has attracted increasing research attention as it can significantly reduce the annotation cost.However, labels from weak supervision can be noisy, and the high capacity of DNNs enables them to easily overfit the label noise, resulting in poor generalization.Recent methods leverage self-training to build noiseresistant models, in which a teacher trained under weak supervision is used to provide highly confident labels for teaching the students.Nevertheless, the teacher derived from such frameworks may have fitted a substantial amount of noise and therefore produce incorrect pseudolabels with high confidence, leading to severe error propagation.In this work, we propose Meta Self-Refinement (MSR), a noise-resistant learning framework, to effectively combat label noise from weak supervision.Instead of relying on a fixed teacher trained with noisy labels, we encourage the teacher to refine its pseudolabels.At each training step, MSR performs a meta gradient descent on the current mini-batch to maximize the student performance on a clean validation set.Extensive experimentation on eight NLP benchmarks demonstrates that MSR is robust against label noise in all settings and outperforms state-of-the-art methods by up to 11.4% in accuracy and 9.26% in F1 score.
Xiaoyu Shen 0001, Michael A. Hedderich, Dietrich Klakow
EACL2
2022 RoCBert: Robust Chinese Bert with Multimodal Contrastive Pretraining
abstract
Large-scale pretrained language models have achieved SOTA results on NLP tasks.However, they have been shown vulnerable to adversarial attacks especially for logographic languages like Chinese.In this work, we propose ROCBERT: a pretrained Chinese Bert that is robust to various forms of adversarial attacks like word perturbation, synonyms, typos, etc.It is pretrained with the contrastive learning objective which maximizes the label consistency under different synthesized adversarial examples.The model takes as input multimodal information including the semantic, phonetic and visual features.We show all these features are important to the model robustness since the attack can be performed in all the three forms.Across 5 Chinese NLU tasks, ROCBERT outperforms strong baselines under three blackbox adversarial algorithms without sacrificing the performance on clean testset.It also performs the best in the toxic content detection task under human-made attacks. * Equal contribution.
Hui Su, Xiaoyu Shen 0001, Xiao Zhou 0004, Tuo Ji, Jiarui Fang, Jie Zhou 0016
ACL (1)3
2022 AST-Trans: Code Summarization with Efficient Tree-Structured Attention
abstract
Code summarization aims to generate brief natural language descriptions for source codes. The state-of-the-art approaches follow a transformer-based encoder-decoder architecture. As the source code is highly structured and follows strict grammars, its Abstract Syntax Tree (AST) is widely used for encoding structural information. However, ASTs are much longer than the corresponding source code. Existing approaches ignore the size constraint and simply feed the whole linearized AST into the encoders. We argue that such a simple process makes it difficult to extract the truly useful dependency relations from the overlong input sequence. It also incurs significant computational overhead since each node needs to apply self-attention to all other nodes in the AST. To encode the AST more effectively and efficiently, we propose AST-Trans in this paper which exploits two types of node relationships in the AST: ancestor-descendant and sibling relationships. It applies the tree-structured attention to dynamically allocate weights for relevant nodes and exclude irrelevant nodes based on these two relationships. We further propose an efficient implementation to support fast parallel computation for tree-structure attention. On the two code summarization datasets, experimental results show that AST-Trans significantly outperforms the state-of-the-arts while being times more efficient than standard transformers1.
Ze Tang 0002, Xiaoyu Shen 0001, Chuanyi Li, Jidong Ge, LiGuo Huang, Zheling Zhu, Bin Luo 0003
ICSE2
2022 A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation
abstract
David Adelani, Jesujoba Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Emezue, Colin Leong, Michael Beukman, Shamsuddeen Muhammad, Guyo Jarso, Oreen Yousuf, Andre Niyongabo Rubungo, Gilles Hacheme, Eric Peter Wairagala, Muhammad Umair Nasir, Benjamin Ajibade, Tunde Ajayi, Yvonne Gitau, Jade Abbott, Mohamed Ahmed, Millicent Ochieng, Anuoluwapo Aremu, Perez Ogayo, Jonathan Mukiibi, Fatoumata Ouoba Kabore, Godson Kalipe, Derguene Mbaye, Allahsera Auguste Tapo, Victoire Memdjokam Koagne, Edwin Munkoh-Buabeng, Valencia Wagner, Idris Abdulmumin, Ayodele Awokoya, Happy Buzaaba, Blessing Sibanda, Andiswa Bukula, Sam Manthalu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
David Ifeoluwa Adelani, Jesujoba O. Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen 0001, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Rabiu Gwadabe, Sackey Freshia, Bonaventure F. P. Dossou, Chris C. Emezue, Colin Leong, Michael Beukman, Shamsuddeen Hassan Muhammad, Guyo Dub Jarso, Oreen Yousuf, Rubungo Andre Niyongabo, Gilles Hacheme, Eric Peter Wairagala, Muhammad Umair Nasir, Benjamin Ajibade, Tunde Ajayi, Yvonne Wambui, Jade Z. Abbott, Millicent Ochieng, Aremu Anuoluwapo, Perez Ogayo, Jonathan Mukiibi, Fatoumata Ouoba Kabore, Godson Kalipe, Derguene Mbaye, Allahsera Tapo, Victoire Memdjokam Koagne, Edwin Munkoh-Buabeng, Valencia Wagner, Idris Abdulmumin, Ayodele Awokoya, Happy Buzaaba, Blessing K. Sibanda, Andiswa Bukula, Sam Manthalu
NAACL-HLT5
2021 Question Rewriting for Open-Domain Conversational QA: Best Practices and Limitations
abstract
Open-domain conversational QA (ODCQA) calls for effective question rewriting (QR), as the questions in a conversation typically lack proper context for the QA model to interpret. In this paper, we compare two types of QR approaches, generative and expansive QR, in end-to-end ODCQA systems with recently released QReCC and OR-QuAC benchmarks. While it is common practice to apply the same QR approach for both the retriever and the reader in the QA system, our results show such strategy is generally suboptimal and suggest expansive QR is better for the sparse retriever and generative QR is better for the reader. Furthermore, while conversation history modeling with dense representations outperforms QR, we show the advantages to apply both jointly, as QR boosts the performance especially when limited history turns are considered.
Marco Del Tredici, Gianni Barlacchi, Xiaoyu Shen 0001, Weiwei Cheng, Adrià de Gispert
CIKM3
2021 Neural Data-to-Text Generation with LM-based Text Augmentation
abstract
For many new application domains for datato-text generation, the main obstacle in training neural models consists of a lack of training data.While usually large numbers of instances are available on the data side, often only very few text samples are available.To address this problem, we here propose a novel fewshot approach for this setting.Our approach automatically augments the data available for training by (i) generating new text samples based on replacing specific values by alternative ones from the same category, (ii) generating new text samples based on GPT-2, and (iii) proposing an automatic method for pairing the new text samples with data samples.As the text augmentation can introduce noise to the training data, we use cycle consistency as an objective, in order to make sure that a given data sample can be correctly reconstructed after having been formulated as text (and that text samples can be reconstructed from data).On both the E2E and WebNLG benchmarks, we show that this weakly supervised training paradigm is able to outperform fully supervised seq2seq models with less than 10% annotations.By utilizing all annotated data, our model can boost the performance of a standard seq2seq model by over 5 BLEU points, establishing a new state-of-the-art on both datasets. * Work done prior to joining Amazon.The Blue Spice is a restaurant that serves English cuisine.
Ernie Chang, Xiaoyu Shen 0001, Vera Demberg, Hui Su
EACL2
2021 Preventing Author Profiling through Zero-Shot Multilingual Back-Translation
abstract
Documents as short as a single sentence may inadvertently reveal sensitive information about their authors, including e.g.their gender or ethnicity.Style transfer is an effective way of transforming texts in order to remove any information that enables author profiling.However, for a number of current state-of-theart approaches the improved privacy is accompanied by an undesirable drop in the downstream utility of the transformed data.In this paper, we propose a simple, zero-shot way to effectively lower the risk of author profiling through multilingual back-translation using off-the-shelf translation models.We compare our models with five representative text style transfer models on three datasets across different domains.Results from both an automatic and a human evaluation show that our approach achieves the best overall performance while requiring no training data.We are able to lower the adversarial prediction of gender and race by up to 22% while retaining 95% of the original utility on downstream tasks.
David Ifeoluwa Adelani, Miaoran Zhang, Xiaoyu Shen 0001, Ali Davody, Thomas Kleinbauer, Dietrich Klakow
EMNLP (1)3
2021 The SelectGen Challenge: Finding the Best Training Samples for Few-Shot Neural Text Generation
abstract
We propose a shared task on training instance selection for few-shot neural text generation.Large-scale pretrained language models have led to dramatic improvements in few-shot text generation.Nonetheless, almost all previous work simply applies random sampling to select the few-shot training instances.Little to no attention has been paid to the selection strategies and how they would affect model performance.The study of the selection strategy can help us to ( 1) make the most use of our annotation budget in downstream tasks and (2) better benchmark few-shot text generative models.We welcome submissions that present their selection strategies and the effects on the generation quality.
Ernie Chang, Xiaoyu Shen 0001, Alex Marin, Vera Demberg
INLG2
2021 AST-Transformer: Encoding Abstract Syntax Trees Efficiently for Code Summarization
abstract
Code summarization aims to generate brief natural language descriptions for source code. As source code is highly structured and follows strict programming language grammars, its Abstract Syntax Tree (AST) is often leveraged to inform the encoder about the structural information. However, ASTs are usually much longer than the source code. Current approaches ignore the size limit and simply feed the whole linearized AST into the encoder. To address this problem, we propose AST-Transformer to efficiently encode tree-structured ASTs. Experiments show that AST-Transformer outperforms the state-of-arts by a substantial margin while being able to reduce 90 ~ 95% of the computational complexity in the encoding process.
Ze Tang 0002, Chuanyi Li, Jidong Ge, Xiaoyu Shen 0001, Zheling Zhu, Bin Luo 0003
ASE4
2021 Learning Fine-Grained Fact-Article Correspondence in Legal Cases
abstract
Automatically recommending relevant law articles to a given legal case has attracted much attention as it can greatly release human labor from searching over the large database of laws. However, current researches only support coarse-grained recommendation where all relevant articles are predicted as a whole without explaining which specific fact each article is relevant with. Since one case can be formed of many supporting facts, traversing over them to verify the correctness of recommendation results can be time-consuming. We believe that learning fine-grained correspondence between each single fact and law articles is crucial for an accurate and trustworthy AI system. With this motivation, we perform a pioneering study and create a corpus with manually annotated fact-article correspondences. We treat the learning as a text matching task and propose a multi-level matching network to address it. To help the model better digest the content of law articles, we parse articles in form of premise-conclusion pairs with random forest. Experiments show that the parsed form yielded better performance and the resulting model surpassed other popular text matching baselines. Furthermore, we compare with previous researches and find that establishing the fine-grained fact-article correspondences can improve the recommendation accuracy by a large margin. Our best system reaches an F1 score of 96.3%, making it of great potential for practical use. It can also significantly boost the downstream task of legal decision prediction, increasing the F1 score by up to 12.7%. The dataset and code will be released upon acceptance. Code and dataset are available at https://github.com/gjdnju/MLMN.
Jidong Ge, Yunyun Huang, Xiaoyu Shen 0001, Chuanyi Li, Wei Hu 0007
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Neural Data-to-Text Generation via Jointly Learning the Segmentation and Correspondence
abstract
The neural attention model has achieved great success in data-to-text generation tasks. Though usually excelling at producing fluent text, it suffers from the problem of information missing, repetition and "hallucination". Due to the black-box nature of the neural attention architecture, avoiding these problems in a systematic way is non-trivial. To address this concern, we propose to explicitly segment target text into fragment units and align them with their data correspondences. The segmentation and correspondence are jointly learned as latent variables without any human annotations. We further impose a soft statistical constraint to regularize the segmental granularity. The resulting architecture maintains the same expressive power as neural attention models, while being able to generate fully interpretable outputs with several times less computational cost. On both E2E and WebNLG benchmarks, we show the proposed model consistently outperforms its neural attention counterparts.
Xiaoyu Shen 0001, Ernie Chang, Hui Su, Cheng Niu, Dietrich Klakow
ACL1
2020 Diversifying Dialogue Generation with Non-Conversational Text
abstract
Neural network-based sequence-to-sequence (seq2seq) models strongly suffer from the lowdiversity problem when it comes to opendomain dialogue generation.As bland and generic utterances usually dominate the frequency distribution in our daily chitchat, avoiding them to generate more interesting responses requires complex data filtering, sampling techniques or modifying the training objective.In this paper, we propose a new perspective to diversify dialogue generation by leveraging non-conversational text.Compared with bilateral conversations, nonconversational text are easier to obtain, more diverse and cover a much broader range of topics.We collect a large-scale nonconversational corpus from multi sources including forum comments, idioms and book snippets.We further present a training paradigm to effectively incorporate these text via iterative back translation.The resulting model is tested on two conversational datasets and is shown to produce significantly more diverse responses without sacrificing the relevance with context.
Hui Su, Xiaoyu Shen 0001, Sanqiang Zhao, Xiao Zhou 0004, Pengwei Hu 0001, Randy Zhong, Cheng Niu, Jie Zhou 0016
ACL2
2020 MovieChats: Chat like Humans in a Closed Domain
abstract
Being able to perform in-depth chat with humans in a closed domain is a precondition before an open-domain chatbot can ever be claimed.In this work, we take a close look at the movie domain and present a large-scale high-quality corpus with fine-grained annotations in hope of pushing the limit of moviedomain chatbots.We propose a unified, readily scalable neural approach which reconciles all subtasks like intent prediction and knowledge retrieval.The model is first pretrained on the huge general-domain data, then finetuned on our corpus.We show this simple neural approach trained on high-quality data is able to outperform commercial systems replying on complex rules.On both the static and interactive tests, we find responses generated by our system exhibits remarkably good engagement and sensibleness close to human-written ones.We further analyze the limits of our work and point out potential directions for future work 1 .
Hui Su, Xiaoyu Shen 0001, Xiao Zhou 0004, Ernie Chang, Cheng Niu, Jie Zhou 0016
EMNLP (1)2
2019 Improving Multi-turn Dialogue Modelling with Utterance ReWriter
abstract
Recent research has achieved impressive results in single-turn dialogue modelling. In the multi-turn setting, however, current models are still far from satisfactory. One major challenge is the frequently occurred coreference and information omission in our daily conversation, making it hard for machines to understand the real intention. In this paper, we propose rewriting the human utterance as a pre-process to help multi-turn dialgoue modelling. Each utterance is first rewritten to recover all coreferred and omitted information. The next processing steps are then performed based on the rewritten utterance. To properly train the utterance rewriter, we collect a new dataset with human annotations and introduce a Transformer-based utterance rewriting architecture using the pointer network. We show the proposed architecture achieves remarkably good performance on the utterance rewriting task. The trained utterance rewriter can be easily integrated into online chatbots and brings general improvement over different domains.
Hui Su, Xiaoyu Shen 0001, Rongzhi Zhang, Fei Sun 0001, Pengwei Hu 0001, Cheng Niu, Jie Zhou 0016
ACL (1)2
2019 Unsupervised Rewriter for Multi-Sentence Compression
abstract
Multi-sentence compression (MSC) aims to generate a grammatical but reduced compression from multiple input sentences while retaining their key information.Previous dominating approach for MSC is the extractionbased word graph approach.A few variants further leveraged lexical substitution to yield more abstractive compression.However, two limitations exist.First, the word graph approach that simply concatenates fragments from multiple sentences may yield nonfluent or ungrammatical compression.Second, lexical substitution is often inappropriate without the consideration of context information.To tackle the above-mentioned issues, we present a neural rewriter for multisentence compression that does not need any parallel corpus.Empirical studies have shown that our approach achieves comparable results upon automatic evaluation and improves the grammaticality of compression based on human evaluation.A parallel corpus with more than 140,000 (sentence group, compression) pairs is also constructed as a by-product for future research.
Xiaoyu Shen 0001, Wei Bi, Akiko Aizawa
ACL (1)2
2019 Select and Attend: Towards Controllable Content Selection in Text Generation
abstract
Xiaoyu Shen, Jun Suzuki, Kentaro Inui, Hui Su, Dietrich Klakow, Satoshi Sekine. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xiaoyu Shen 0001, Jun Suzuki 0001, Kentaro Inui, Hui Su, Dietrich Klakow, Satoshi Sekine
EMNLP/IJCNLP (1)1
2019 Improving Latent Alignment in Text Summarization by Generalizing the Pointer Generator
abstract
Xiaoyu Shen, Yang Zhao, Hui Su, Dietrich Klakow. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xiaoyu Shen 0001, Hui Su, Dietrich Klakow
EMNLP/IJCNLP (1)1
2018 Towards Better Variational Encoder-Decoders in Seq2Seq Tasks
abstract
Variational encoder-decoders have shown promising results in seq2seq tasks. However, the training process is known difficult to be controlled because latent variables tend to be ignored while decoding. In this paper, we thoroughly analyze the reason behind this training difficulty, compare different ways of alleviating it and propose a new framework that helps significantly improve the overall performance.
Xiaoyu Shen 0001, Hui Su
AAAI1
2018 Improving Variational Encoder-Decoders in Dialogue Generation
abstract
Variational encoder-decoders (VEDs) have shown promising results in dialogue generation. However, the latent variable distributions are usually approximated by a much simpler model than the powerful RNN structure used for encoding and decoding, yielding the KL-vanishing problem and inconsistent training objective. In this paper, we separate the training step into two phases: The first phase learns to autoencode discrete texts into continuous embeddings, from which the second phase learns to generalize latent representations by reconstructing the encoded embedding. In this case, latent variables are sampled by transforming Gaussian noise through multi-layer perceptrons and are trained with a separate VED model, which has the potential of realizing a much more flexible distribution. We compare our model with current popular models and the experiment demonstrates substantial improvement in both metric-based and human evaluations.
Xiaoyu Shen 0001, Hui Su, Shuzi Niu, Vera Demberg
AAAI1
2018 Dialogue Generation With GAN
abstract
This paper presents a Generative Adversarial Network (GAN) to model multiturn dialogue generation, which trains a latent hierarchical recurrent encoder-decoder simultaneously with a discriminative classifier that make the prior approximate to the posterior. Experiments show that our model achieves better results.
Hui Su, Xiaoyu Shen 0001, Pengwei Hu 0001, Wenjie Li 0002
AAAI2
2018 Nexus Network: Connecting the Preceding and the Following in Dialogue Generation
abstract
Sequence-to-Sequence (seq2seq) models have become overwhelmingly popular in building end-to-end trainable dialogue systems.Though highly efficient in learning the backbone of human-computer communications, they suffer from the problem of strongly favoring short generic responses.In this paper, we argue that a good response should smoothly connect both the preceding dialogue history and the following conversations.We strengthen this connection through mutual information maximization.To sidestep the nondifferentiability of discrete natural language tokens, we introduce an auxiliary continuous code space and map such code space to a learnable prior distribution for generation purpose.Experiments on two dialogue datasets validate the effectiveness of our model, where the generated responses are closely related to the dialogue context and lead to more interactive conversations.* Indicates equal contribution.X. Shen focuses on algorithm and H. Su is responsible for experiments.
Xiaoyu Shen 0001, Hui Su, Wenjie Li 0002, Dietrich Klakow
EMNLP1
2018 A comprehensive study: Sentence compression with linguistic knowledge-enhanced gated neural network
Xiaoyu Shen 0001, Hajime Senuma, Akiko Aizawa
Data Knowl. Eng.2
2018 Simulating the Large-Scale Erosion of Genomic Privacy Over Time
abstract
The dramatically decreasing costs of DNA sequencing have triggered more than a million humans to have their genotypes sequenced. Moreover, these individuals increasingly make their genomic data publicly available, thereby creating privacy threats for themselves and their relatives because of their DNA similarities. More generally, an entity that gains access to a significant fraction of sequenced genotypes might be able to infer even the genomes of unsequenced individuals. In this paper, we propose a simulation-based model for quantifying the impact of continuously sequencing and publicizing personal genomic data on a population's genomic privacy. Our simulation probabilistically models data sharing and takes into account events such as migration and interracial mating. We exemplarily instantiate our simulation with a sample population of 1,000 individuals and evaluate the privacy under multiple settings over 6,000 genomic variants and a subset of phenotype-related variants. Our findings demonstrate that an increasing sharing rate in the future entails a substantial negative effect on the privacy of all older generations. Moreover, we find that mixed populations face a less severe erosion of privacy over time than more homogeneous populations. Finally, we demonstrate that genomic-data sharing can be much more detrimental for the privacy of the phenotype-related variants.
Michael Backes 0001, Pascal Berrang, Mathias Humbert, Xiaoyu Shen 0001, Verena Wolf 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2017 Wake-Sleep Variational Autoencoders for Language Modeling
Xiaoyu Shen 0001, Hui Su, Shuzi Niu, Dietrich Klakow
ICONIP (1)1
2017 DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
abstract
We develop a high-quality multi-turn dialog dataset, DailyDialog, which is intriguing in several aspects. The language is human-written and less noisy. The dialogues in the dataset reflect our daily communication way and cover various topics about our daily life. We also manually label the developed dataset with communication intention and emotion information. Then, we evaluate existing approaches on DailyDialog dataset and hope it benefit the research field of dialog systems. The dataset is available on http://yanran.li/dailydialog
Yanran Li, Hui Su, Xiaoyu Shen 0001, Wenjie Li 0002, Ziqiang Cao, Shuzi Niu
IJCNLP(1)3
2017 Estimation of Gap Between Current Language Models and Human Performance
Xiaoyu Shen 0001, Youssef Oualil, Clayton Greenberg, Mittul Singh, Dietrich Klakow
INTERSPEECH1
2017 Gated Neural Network for Sentence Compression Using Linguistic Knowledge
Hajime Senuma, Xiaoyu Shen 0001, Akiko Aizawa
NLDB3