Longyue Wang

dblp:127/3421 · DBLP profile ↗
← Back
76ranked-venue papers
10as first author
55since 2021 · last 2026
0000-0002-9062-6183ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 62 · 9 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs
abstract
Large Language Models (LLMs) frequently exhibit strong translation abilities, even without task-specific fine-tuning. However, the internal mechanisms governing this innate capability remain largely opaque. To demystify this process, we leverage Sparse Autoencoders (SAEs) and introduce a novel framework for identifying task-specific features. Our method first recalls features that are frequently co-activated on translation inputs and then filters them for functional coherence using a PCA-based consistency metric. This framework successfully isolates a small set of "translation initiation" features. Causal interventions demonstrate that amplifying these features steers the model towards correct translation, while ablating them induces hallucinations and off-task outputs, confirming they represent a core component of the model's innate translation competency. Moving from analysis to application, we leverage this mechanistic insight to propose a new data selection strategy for efficient fine-tuning. Specifically, we prioritize training on "mechanistically hard" samples—those that fail to naturally activate the translation initiation features. Experiments show this approach significantly improves data efficiency and suppresses hallucinations. Furthermore, we find these mechanisms are transferable to larger models of the same family. Our work not only decodes a core component of the translation mechanism in LLMs but also provides a blueprint for using internal model mechanism to create more robust and efficient models.
Xinwei Wu 0001, Yuqi Ren, Linlong Xu, Longyue Wang, Deyi Xiong, Weihua Luo, Kaifu Zhang
AAAI6
2026 HSCodeComp: A Realistic and Expert-level Agent Benchmark for Hierarchical Rule Application
abstract
Tian Lan, Yiqian Yang, Qianghuai Jia, Li Zhu, Hui Jiang, Hang Zhu, Weihua Luo, Longyue Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yiqian Yang, Qianghuai Jia, Weihua Luo, Longyue Wang
ACL (1)8
2026 Structured Episodic Event Memory
abstract
Current approaches to memory in Large Language Models (LLMs) predominantly rely on static Retrieval-Augmented Generation (RAG), which often results in scattered retrieval and fails to capture the structural dependencies required for complex reasoning.For autonomous agents, these passive and flat architectures lack the cognitive organization necessary to model the dynamic and associative nature of longterm interaction.To address this, we propose Structured Episodic Event Memory (SEEM), a hierarchical framework that synergizes a graph memory layer for relational facts with a dynamic episodic memory layer for narrative progression.Grounded in cognitive frame theory, SEEM transforms interaction streams into structured Episodic Event Frames (EEFs) anchored by precise provenance pointers.Furthermore, we introduce an agentic associative fusion and Reverse Provenance Expansion (RPE) mechanism to reconstruct coherent narrative contexts from fragmented evidence.Experimental results on the LoCoMo and Long-MemEval benchmarks demonstrate that SEEM significantly outperforms baselines, enabling agents to maintain superior narrative coherence and logical consistency.
Zhengxuan Lu, Dongfang Li 0002, Yukun Shi, Beilun Wang, Longyue Wang, Baotian Hu
ACL (1)5
2026 CAML: A Conflict-Aware Molecular Language Model Merging Framework for Multi-Constraint Molecular Generation
abstract
Xuanbai Ren, Luoda Tan, Pei Liu, Tengfei Ma, Xiangzheng Fu, Longyue Wang, Yiping Liu, Xiangxiang Zeng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xuanbai Ren, Luoda Tan, Pei Liu 0008, Tengfei Ma 0002, Xiangzheng Fu, Longyue Wang, Xiangxiang Zeng
ACL (1)6
2026 GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models
abstract
Zhiwen Ruan, Yichao Du, Jianjie Zheng, Longyue Wang, Yun Chen, Peng Li, Jinsong Su, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhiwen Ruan, Yichao Du, Jianjie Zheng, Longyue Wang, Yun Chen 0007, Peng Li 0030, Jinsong Su, Yang Liu 0005, Guanhua Chen 0001
ACL (1)4
2026 From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models
abstract
Ling Shi, Xinwei Wu, Xiaohu Zhao, Hao Wang, Heng Liu, Yangyang Liu, Linlong Xu, Longyue Wang, Deyi Xiong, Weihua Luo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ling Shi 0004, Xinwei Wu 0001, Linlong Xu, Longyue Wang, Deyi Xiong, Weihua Luo
ACL (1)8
2026 M²PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation
abstract
Hao Wang, Linlong Xu, Heng Liu, Yangyang Liu, Xiaohu Zhao, Bo Zeng, Liangying Shao, Yichen Dong, Xinwei Wu, Jiang Zhou, Tianyu Dong, Xiangxiang Zeng, Longyue Wang, Weihua Luo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Linlong Xu, Liangying Shao, Yichen Dong, Xinwei Wu 0001, Tianyu Dong, Xiangxiang Zeng, Longyue Wang, Weihua Luo
ACL (1)13
2026 DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection
abstract
The effective detection and governance of Large Language Model (LLM) generated content has become increasingly critical due to the growing risk of misuse. Despite the impressive performance of existing detectors, their reliability and potential in multilingual, real-world scenarios remain largely underexplored.In this study, we introduce DetectRL-X, a comprehensive multilingual benchmark designed to evaluate advanced detectors across 8 dimensions. The benchmark encompasses 8 languages commonly used in commercial contexts and collects human-written texts from 6 domains highly susceptible to LLM misuse. To better aligned with real-world applications, We create LLM-generated texts using 4 popular commercial LLMs, and include typical AI-assisted writing operations such as polishing, expanding, and condensing to capture authentic usage patterns. Furthermore, we develop a multilingual framework for paraphrasing and perturbation attacks to simulate diverse human modifications and writing noise, enabling stress testing of detectors across languages.Experimental results on DetectRL-X reveal the strengths and limitations of current state-of-the-art detectors when applied to diverse linguistic resources. We further analyze how domains, generators, attack strategies, text length, and refinement operations influence performance in different languages, underscoring DetectRL-X as an effective benchmark for strengthening multilingual and language-specific detectors.
Junchao Wu, Yefeng Liu, Chenyu Zhu, Tianqi Shi, Yichao Du, Longyue Wang, Weihua Luo, Jinsong Su, Derek F. Wong
ACL (1)8
2026 Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation
abstract
Jiang Zhou, Xiaohu Zhao, Xinwei Wu, Tianyu Dong, Hao Wang, Yangyang Liu, Heng Liu, Linlong Xu, Longyue Wang, Weihua Luo, Deyi Xiong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xinwei Wu 0001, Tianyu Dong, Linlong Xu, Longyue Wang, Weihua Luo, Deyi Xiong
ACL (1)9
2026 New Trends for Modern Machine Translation with Large Reasoning Models
abstract
Recent advances in Large Reasoning Models (LRMs), particularly those leveraging Chain-of-Thought reasoning (CoT), have opened brand new possibility for Machine Translation (MT). This position paper argues that LRMs substantially transformed traditional neural MT as well as LLMs-based MT paradigms by reframing translation as a dynamic reasoning task that requires contextual, cultural, and linguistic understanding and reasoning. We identify three foundational shifts: 1) contextual coherence, where LRMs resolve ambiguities and preserve discourse structure through explicit reasoning over cross-sentence and complex context or even lack of context; 2) cultural intentionality, enabling models to adapt outputs by inferring speaker intent, audience expectations, and socio-linguistic norms; 3) self-reflection, LRMs can perform self-reflection during the inference time to correct the potential errors in translation especially extremely noisy cases, showing better robustness compared to simply mapping X->Y translation. We explore various scenarios in translation including stylized translation, document-level translation and multimodal translation by showcasing empirical examples that demonstrate the superiority of LRMs in translation. We also identify several interesting phenomenons for LRMs for MT including auto-pivot translation as well as the critical challenges such as over-localisation in translation and inference efficiency. In conclusion, we think that LRMs redefine translation systems not merely as text converters but as multilingual cognitive agents capable of reasoning about meaning beyond the text. This paradigm shift reminds us to think of problems in translation beyond traditional translation scenarios in a much broader context with LRMs - what we can achieve on top of it.
Sinuo Liu, Chenyang Lyu, Minghao Wu, Zifu Shang, Longyue Wang, Weihua Luo, Kaifu Zhang
LREC5
2026 BloodPatrol: Revolutionizing Blood Cancer Diagnosis - Advanced Real-Time Detection Leveraging Deep Learning & Cloud Technologies
abstract
Cloud computing and Internet of Things (IoT) technologies are gradually becoming the technological changemakers in cancer diagnosis. Blood cancer is an aggressive disease affecting the blood, bone marrow, and lymphatic system, and its early detection is crucial for subsequent treatment. Flow cytometry has been widely studied as a commonly used method for detecting blood cancer. However, the high computation and resource consumption severely limit its practical application, especifically in regions with limited medical and computational resources. In this study, with the help of cloud computing and IoT technologies, we develop a novel blood cancer dynamic monitoring diagnostic model named BloodPatrol based on an intelligent feature weight fusion mechanism. The proposed model is capable of capturing the dual-view importance relationship between cell samples and features, greatly improving prediction accuracy and significantly surpassing previous models. Besides, benefiting from the powerful processing ability of cloud computing, BloodPatrol can run on a distributed network to efficiently process large-scale cell data, which provides immediate and scalable blood cancer diagnostic services.
Jinhang Wei, Longyue Wang, Zhecheng Zhou, Linlin Zhuo, Xiangxiang Zeng, Xiangzheng Fu, Quan Zou 0001, Keqin Li 0001, Zhongjun Zhou
IEEE J. Biomed. Health Informatics2
2025 A Unified Agentic Framework for Evaluating Conditional Image Generation
abstract
Conditional image generation has gained significant attention for its ability to personalize content. However, the field faces challenges in developing task-agnostic, reliable, and explainable evaluation metrics. This paper introduces CIGEval, a unified agentic framework for comprehensive evaluation of conditional image generation tasks. CIGEval utilizes large multimodal models (LMMs) as its core, integrating a multi-functional toolbox and establishing a fine-grained evaluation framework. Additionally, we synthesize evaluation trajectories for fine-tuning, empowering smaller LMMs to autonomously select appropriate tools and conduct nuanced analyses based on tool outputs. Experiments across seven prominent conditional image generation tasks demonstrate that CIGEval (GPT-4o version) achieves a high correlation of 0.4625 with human assessments, closely matching the inter-annotator correlation of 0.47. Notably, when implemented with 7B open-source LMMs using only 2.3K training trajectories, CIGEval surpasses the previous GPT-4o-based state-of-the-art method. These findings indicate that CIGEval holds great potential for automating evaluation of image generation tasks while maintaining human-level reliability.
Jifang Wang, Yangxue, Longyue Wang, Zhenran Xu, Yaowei Wang 0001, Weihua Luo, Kaifu Zhang, Baotian Hu, Min Zhang 0005
ACL (1)3
2025 Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models
abstract
Large Reasoning Models (LRMs) such as OpenAI o1 and DeepSeek-R1 have shown remarkable reasoning capabilities by scaling test-time compute and generating long Chain-of-Thought (CoT). Distillation post-training on LRMs-generated data is a straightforward yet effective method to enhance the reasoning abilities of smaller models, but faces a critical bottleneck: we found that distilled long CoT data poses learning difficulty for small models and leads to the inheritance of biases (i.e., formalistic long-time thinking) when using Supervised Fine-tuning (SFT) and Reinforcement Learning (RL) methods. To alleviate this bottleneck, we propose constructing data from scratch using Monte Carlo Tree Search (MCTS). We then exploit a set of CoT-aware approaches, including Thoughts Length Balance, Fine-grained DPO, and Joint Post-training Objective, to enhance SFT and RL on the MCTS data. We conducted evaluation on various benchmarks such as math (GSM8K, MATH, AIME). instruction-following (Multi-IF) and planning (Blocksworld), results demonstrate our CoT-aware approaches substantially improve the reasoning performance of distilled models compared to standard distilled models via reducing the hallucinations in long-time thinking.
Huifeng Yin, Minghao Wu, Xuanfan Ni, Tianqi Shi, Liangying Shao, Chenyang Lyu, Longyue Wang, Weihua Luo, Kaifu Zhang
ACL (1)10
2025 Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language
abstract
Bo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng, Minghao Wu, Xuanfan Ni, Tianqi Shi, Yu Zhao, Yefeng Liu, Chenyu Zhu, Ruizhe Li, Jiahui Geng, Qing Li, Yu Tong, Longyue Wang, Weihua Luo, Kaifu Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chenyang Lyu, Sinuo Liu, Mingyan Zeng, Minghao Wu, Xuanfan Ni, Tianqi Shi, Yefeng Liu, Chenyu Zhu, Ruizhe Li 0001, Jiahui Geng, Longyue Wang, Weihua Luo, Kaifu Zhang
ACL (1)15
2025 Large Language and Protein Assistant for Protein-Protein Interactions Prediction
abstract
Peng Zhou, Pengsen Ma, Jianmin Wang, Xibao Cai, Haitao Huang, Wei Liu, Longyue Wang, Lai Hou Tim, Xiangxiang Zeng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Peng Zhou 0011, Pengsen Ma, Jianmin Wang 0016, Xibao Cai, Wei Liu 0005, Longyue Wang, Lai Hou Tim, Xiangxiang Zeng
ACL (1)7
2025 Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites
abstract
Detoxifying offensive language while preserving the speaker's original intent is a challenging yet critical goal for improving the quality of online interactions.Although large language models (LLMs) show promise in rewriting toxic content, they often default to overly polite rewrites, distorting the emotional tone and communicative intent.This problem is especially acute in Chinese, where toxicity often arises implicitly through emojis, homophones, or discourse context.We present TOXIREWRITECN, the first Chinese detoxification dataset explicitly designed to preserve sentiment polarity.The dataset comprises 1,556 carefully annotated triplets, each containing a toxic sentence, a sentiment-aligned non-toxic rewrite, and labeled toxic spans.It covers five real-world scenarios: standard expressions, emoji-induced and homophonic toxicity, as well as single-turn and multi-turn dialogues.We evaluate 17 LLMs, including commercial and open-source models with variant architectures, across four dimensions: detoxification accuracy, fluency, content preservation, and sentiment polarity.Results show that while commercial and MoE models perform best overall, all models struggle to balance safety with emotional fidelity in more subtle or context-heavy settings such as emoji, homophone, and dialogue-based inputs.We release TOXIREWRITECN to support future research on controllable, sentiment-aware detoxification for Chinese.Caution: This paper contains examples of violent or offensive language that may be disturbing to some readers.
Xintong Wang 0001, Jingheng Pan, Liang Ding 0006, Longyue Wang, Chris Biemann
EMNLP5
2025 Enhancing Video-Text Matching via Sparse Stratified Sampling
abstract
Video-text matching is a critical task in multimedia retrieval, but traditional methods often fail to capture the diversity and depth of video content due to inefficient and inaccurate frame sampling. We propose a novel sparse stratified sampling technique that can substantially improve the video-text matching process by segmenting video content into clusters based on relevant features and selectively sampling representative frames. Our method further introduces a threshold for the feature metric used to divide clusters, eliminating video frames with low relevance. We propose two variants of our approach: an offline approach that performs sampling before training, and an online approach that dynamically conducts sampling based on the relevance between video frames and the text query during training. Extensive experiments on datasets like MSRVTT and AVSD for video retrieval and multiple-choice VideoQA datasets, including AVQA and Music-AVQA, demonstrate the superiority of our method over previous state-of-the-art approaches. Our sparse stratified sampling technique achieves improvements of over 1.2% on MSRVTT and 1.7% on AVSD for R@1 in video retrieval tasks. For multiple-choice VideoQA tasks, our approach achieves significant improvements of 1.8% accuracy on AVQA and 3.9% on Music-AVQA, strongly supporting its effectiveness in enhancing video-text matching systems.
Chenyang Lyu, Wenxi Li, Tianbo Ji, Liting Zhou, Pintu Lohar, Yi Yu 0001, Longyue Wang
ICASSP7
2025 D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models
Zhongwei Wan, Xinjian Wu, Yu Zhang 0133, Yi Xin 0003, Chaofan Tao, Zhihong Zhu 0001, Xin Wang 0120, Longyue Wang, Mi Zhang 0002
ICLR10
2025 Retrieval-Augmented Multi-Modal Chain-of-Thoughts Reasoning for Large Language Models
abstract
The advancement of Large Language Models (LLMs) has brought substantial attention to the Chain of Thought (CoT) approach, primarily due to its ability to enhance the capability of LLMs on complex reasoning tasks. Moreover, the CoT approach extends to multi-modal tasks of LLMs. However, the selection of optimal CoT demonstration examples for LLMs in multi-modal reasoning remains less explored due to the inherent complexity of multi-modal examples. In this paper, we introduce a novel approach that addresses this challenge by using retrieval mechanisms to dynamically select demonstration examples based on cross-modal and intra-modal similarities. Furthermore, we employ a Group Selection method to select examples containing rationales from different retrieval directions to promote the diversity of demonstration examples. To the best of our knowledge, we are the first to apply Retrieval-Augmented Generation (RAG) with CoT to complex multi-modal reasoning tasks. Through a series of experiments on two popular benchmarks, ScienceQA and MathVista, we demonstrate that our approach significantly improves the performance of GPT-4 by 6% on ScienceQA and 12.9% on MathVista. Additionally, it enhances the performance of GPT-4V on these two datasets by 2.7%, respectively, further advancing the capabilities of the most advanced LLMs and Large Multimodal Models (LMMs) for complex multi-modal reasoning tasks.
Bingshuai Liu, Chenyang Lyu, Zijun Min, Zhanyu Wang, Jinsong Su, Longyue Wang
IJCNN6
2025 EditEval: Towards Comprehensive and Automatic Evaluation for Text-guided Video Editing
abstract
Recently, video editing task has gained widespread attention due to its practical applications and rapid advancements. However, current automatic evaluation metrics for video editing are mostly poorly aligned with human judgments. Thus, researchers heavily rely on human evaluation, which is not only labor-intensive but also difficult to ensure consistency and objectivity. To address these issues, we propose EditEval, the largest-ever video editing benchmark to comprehensively evaluate the performance of video editing models in three aspects: Textual Faithfulness, Frame Consistency, and Video Fidelity. It includes 200 video clips and 1,010 text prompts, from which 160 instances are sampled to generate 1,280 edited videos using eight open-source video editing models, accompanied by human annotations. Furthermore, we propose EditScore, leveraging the advanced reasoning and comprehension capabilities of Multi-modal Large Language Models (MLLMs) as evaluators to assess edited videos across the aforementioned aspects. Experiments show that the best-performing video editing model only reaches an average score of 3.16 (out of a perfect 5), highlighting the challenge of EditEval. Besides, results from more than 10 MLLMs demonstrate the great potential of utilizing EditScore for automatic evaluation. Notably, for textual faithfulness, EditScore equipped with LLaVA-OneVision-7B achieves a significantly higher Pearson Correlation score compared to previous methods based on CLIP (0.50 vs 0.22). The code and dataset are available at: https://github.com/XMUDeepLIT/EditEval
Bingshuai Liu, Ante Wang, Zijun Min, Chenyang Lyu, Longyue Wang, Xu Han 0007, Peng Li 0030, Jinsong Su
ACM Multimedia5
2025 Alleviating Hallucinations in Large Language Models through Multi-Model Contrastive Decoding and Dynamic Hallucination Detection
abstract
Despite their outstanding performance in numerous applications, large language models (LLMs) remain prone to hallucinations, generating content inconsistent with their pretraining corpora. Currently, almost all contrastive decoding approaches alleviate hallucinations by introducing a model susceptible to hallucinations and appropriately widening the contrastive logits gap between hallucinatory tokens and target tokens. However, although existing contrastive decoding methods mitigate hallucinations, they lack enough confidence in the factual accuracy of the generated content. In this work, we propose Multi-Model Contrastive Decoding (MCD), which integrates a pretrained language model with an evil model and a truthful model for contrastive decoding. Intuitively, a token is assigned a high probability only when deemed potentially hallucinatory by the evil model while being considered factual by the truthful model. This decoding strategy significantly enhances the model’s confidence in its generated responses and reduces potential hallucinations. Furthermore, we introduce a dynamic hallucination detection mechanism that facilitates token-by-token identification of hallucinations during generation and a tree-based revision mechanism to diminish hallucinations further. Extensive experimental evaluations demonstrate that our MCD strategy effectively reduces hallucinations in LLMs and outperforms state-of-the-art methods across various benchmarks.
Chenyu Zhu, Yefeng Liu, Aowen Wang, Yangxue, Guanhua Chen 0001, Longyue Wang, Weihua Luo, Kaifu Zhang
NeurIPS7
2025 AniMaker: Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
abstract
Despite rapid advancements in video generation models, generating coherent, long-form storytelling videos that span multiple scenes and characters remains challenging. Current methods often rigidly convert pre-generated keyframes into fixed-length clips, resulting in disjointed narratives and pacing issues. Furthermore, the inherent instability of video generation models means that even a single low-quality clip can significantly degrade the entire output animation’s logical coherence and visual continuity. To overcome these obstacles, we introduce AniMaker, a multi-agent framework enabling efficient multi-candidate clip generation and storytelling-aware clip selection, thus creating globally consistent and story-coherent animation solely from text input. The framework is structured around specialized agents, including the Director Agent for storyboard generation, the Photography Agent for video clip generation, the Reviewer Agent for evaluation, and the Post-Production Agent for editing and voiceover, collectively realizing multi-character, multi-scene animation. Central to AniMaker’s approach are two key technical components: MCTS-Gen in Photography Agent, an efficient Monte Carlo Tree Search (MCTS)-inspired strategy that intelligently navigates the candidate space to generate high-potential clips while optimizing resource usage; and AniEval in Reviewer Agent, the first framework specifically designed for multi-shot animation evaluation, which assesses critical aspects such as story-level consistency, action completion, and animation-specific features by considering each clip in the context of its preceding and succeeding clips. Experiments demonstrate that AniMaker achieves superior quality as measured by popular metrics including VBench and our proposed AniEval framework, while significantly improving the efficiency of multi-candidate generation, pushing AI-generated storytelling animation closer to production standards. Code and data for this paper are at https://animaker-dev.github.io/
Yunxin Li, Xinyu Chen 0003, Longyue Wang, Baotian Hu, Min Zhang 0005
SIGGRAPH Asia4
2025 DrugAssist: a large language model for molecule optimization
abstract
Recently, the impressive performance of large language models (LLMs) on a wide range of tasks has attracted an increasing number of attempts to apply LLMs in drug discovery. However, molecule optimization, a critical task in the drug discovery pipeline, is currently an area that has seen little involvement from LLMs. Most of existing approaches focus solely on capturing the underlying patterns in chemical structures provided by the data, without taking advantage of expert feedback. These non-interactive approaches overlook the fact that the drug discovery process is actually one that requires the integration of expert experience and iterative refinement. To address this gap, we propose DrugAssist, an interactive molecule optimization model which performs optimization through human-machine dialogue by leveraging LLM's strong interactivity and generalizability. DrugAssist has achieved leading results in both single and multiple property optimization, simultaneously showcasing immense potential in transferability and iterative optimization. In addition, we publicly release a large instruction-based dataset called 'MolOpt-Instructions' for fine-tuning language models on molecule optimization tasks. We have made our code and data publicly available at https://github.com/blazerye/DrugAssist, which we hope to pave the way for future research in LLMs' application for drug discovery.
Geyan Ye, Xibao Cai, Houtim Lai, Xing Wang 0007, Junhong Huang, Longyue Wang, Wei Liu 0005, Xiangxiang Zeng
Briefings Bioinform.6
2025 Widening the bottleneck of lexical choice for non-autoregressive translation
Liang Ding 0006, Longyue Wang, Siyou Liu, Weihua Luo, Kaifu Zhang
Comput. Speech Lang.2
2025 Uni-MoE: Scaling Unified Multimodal LLMs With Mixture of Experts
abstract
Recent advancements in Multimodal Large Language Models (MLLMs) underscore the significance of scalable models and data to boost performance, yet this often incurs substantial computational costs. Although the Mixture of Experts (MoE) architecture has been employed to scale large language or visual-language models efficiently, these efforts typically involve fewer experts and limited modalities. To address this, our work presents the pioneering attempt to develop a unified MLLM with the MoE architecture, named Uni-MoE that can handle a wide array of modalities. Specifically, it features modality-specific encoders with connectors for a unified multimodal representation. We also implement a sparse MoE architecture within the LLMs to enable efficient training and inference through modality-level data parallelism and expert-level model parallelism. To enhance the multi-expert collaboration and generalization, we present a progressive training strategy: 1) Cross-modality alignment using various connectors with different cross-modality data, 2) Training modality-specific experts with cross-modality instruction data to activate experts' preferences, and 3) Tuning the whole Uni-MoE framework utilizing Low-Rank Adaptation (LoRA) on mixed multimodal instruction data. We evaluate the instruction-tuned Uni-MoE on a comprehensive set of multimodal datasets. The extensive experimental results demonstrate Uni-MoE's principal advantage of significantly reducing performance bias in handling mixed multimodal datasets, alongside improved multi-expert collaboration and generalization.
Yunxin Li, Shenyuan Jiang, Baotian Hu, Longyue Wang, Wanqi Zhong, Wenhan Luo, Lin Ma 0002, Min Zhang 0005
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models
abstract
Abstract The evolution of Neural Machine Translation (NMT) has been significantly influenced by six core challenges (Koehn and Knowles, 2017) that have acted as benchmarks for progress in this field. This study revisits these challenges, offering insights into their ongoing relevance in the context of advanced Large Language Models (LLMs): domain mismatch, amount of parallel data, rare word prediction, translation of long sentences, attention model as word alignment, and sub-optimal beam search. Our empirical findings show that LLMs effectively reduce reliance on parallel data for major languages during pretraining and significantly improve translation of long sentences containing approximately 80 words, even translating documents up to 512 words. Despite these improvements, challenges in domain mismatch and rare word prediction persist. While NMT-specific challenges like word alignment and beam search may not apply to LLMs, we identify three new challenges in LLM-based translation: inference efficiency, translation of low-resource languages during pretraining, and human-aligned evaluation.
Jianhui Pang, Fanghua Ye 0001, Derek F. Wong, Dian Yu 0001, Shuming Shi 0001, Zhaopeng Tu, Longyue Wang
Trans. Assoc. Comput. Linguistics7
2025 (Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts
abstract
Abstract Literary translations remains one of the most challenging frontiers in machine translation due to the complexity of capturing figurative language, cultural nuances, and unique stylistic elements. In this work, we introduce TransAgents, a novel multi-agent framework that simulates the roles and collaborative practices of a human translation company, including a CEO, Senior Editor, Junior Editor, Translator, Localization Specialist, and Proofreader. The translation process is divided into two stages: a preparation stage where the team is assembled and comprehensive translation guidelines are drafted, and an execution stage that involves sequential translation, localization, proofreading, and a final quality check. Furthermore, we propose two innovative evaluation strategies: Monolingual Human Preference (MHP), which evaluates translations based solely on target language quality and cultural appropriateness, and BLP, which leverages large language models like gpt-4 for direct text comparison. Although TransAgents achieves lower d-BLEU scores, due to the limited diversity of references, its translations are significantly better than those of other baselines and are preferred by both human evaluators and LLMs over traditional human references and gpt-4 translations. Our findings highlight the potential of multi-agent collaboration in enhancing translation quality, particularly for longer texts.1
Minghao Wu, Yulin Yuan, Gholamreza Haffari, Longyue Wang, Weihua Luo, Kaifu Zhang
Trans. Assoc. Comput. Linguistics5
2025 Molecular Dynamics-Powered Hierarchical Geometric Deep Learning Framework for Protein-Ligand Interaction
abstract
Accurate prediction of the drug binding between proteins and ligands can significantly advance the development of structure-based drug design. Recent advances have shown great potential in applying equivariant graph neural network (EGNN) -based methods to learn representations of protein-ligand (PL) complexes. However, most of them typically focus on atom-level graph representations and omit the residue-level information in PL complexes, which are considered essential for understanding the binding mechanism. In this article, we develop a SO(3)-equivariant hierarchical graph neural network (EHGNN) that effectively captures the intrinsic hierarchy of biomolecular structures to enhance the predictive performance of PL interactions. Based on the SO(3)-EHGNN, we further propose a molecular dynamics-powered and energy-guided deep learning framework, called Dynamics-PLI, to capture the spatial structures and energetic information inside molecular dynamic (MD) trajectories. Extensive experimental results show significant improvements over current state-of-the-art methods, with a decrease of 4.03% in RMSE for the binding affinity problem and an average increase of 3.95% in AUROC and AUPRC for the ligand efficacy problem, demonstrating the superiority of Dynamics-PLI for PL interaction prediction. Our findings indicate that the SO(3)-EHGNN exhibits enhanced performance without the necessity of pre-training, emphasizing the inherent analytical strength of SO(3)-EHGNN.
Mingquan Liu, Shuting Jin, Houtim Lai, Longyue Wang, Jianmin Wang 0016, Zhixiang Cheng, Xiangxiang Zeng
IEEE Trans. Comput. Biol. Bioinform.4
2024 MAGE: Machine-generated Text Detection in the Wild
abstract
Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, Yue Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi 0001, Yue Zhang 0004
ACL (1)6
2024 A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models
abstract
Machine Translation (MT) has greatly advanced over the years due to the developments in deep neural networks. However, the emergence of Large Language Models (LLMs) like GPT-4 and ChatGPT is introducing a new phase in the MT domain. In this context, we believe that the future of MT is intricately tied to the capabilities of LLMs. These models not only offer vast linguistic understandings but also bring innovative methodologies, such as prompt-based techniques, that have the potential to further elevate MT. In this paper, we provide an overview of the significant enhancements in MT that are influenced by LLMs and advocate for their pivotal role in upcoming MT research and implementations. We highlight several new MT directions, emphasizing the benefits of LLMs in scenarios such as Long-Document Translation, Stylized Translation, and Interactive Translation. Additionally, we address the important concern of privacy in LLM-driven MT and suggest essential privacy-preserving strategies. By showcasing practical instances, we aim to demonstrate the advantages that LLMs offer, particularly in tasks like translating extended documents. We conclude by emphasizing the critical role of LLMs in guiding the future evolution of MT and offer a roadmap for future exploration in the sector.
Chenyang Lyu, Zefeng Du, Jitao Xu 0003, Yitao Duan, Minghao Wu, Teresa Lynn, Alham Fikri Aji, Derek F. Wong, Longyue Wang
LREC/COLING9
2024 On the Cultural Gap in Text-to-Image Generation
abstract
One challenge in text-to-image (T2I) generation is the inadvertent reflection of culture gaps present in the training data, which signifies the disparity in generated image quality when the cultural elements of the input text are rarely collected in the training set. Although various T2I models have shown impressive but arbitrary examples, there is no benchmark to systematically evaluate a T2I model’s ability to generate cross-cultural images. To bridge the gap, we propose a Challenging Cross-Cultural (C3) benchmark with comprehensive evaluation criteria, which can assess how well-suited a model is to a target culture. By analyzing the flawed images generated by the Stable Diffusion model on the C3 benchmark, we find that the model often fails to generate certain cultural objects. Accordingly, we propose a novel multi-modal metric that considers object-text alignment to filter the fine-tuning data in the target culture, which is used to fine-tune a T2I model to improve cross-cultural generation. Experimental results show that our multi-modal metric provides stronger data selection performance on the C3 benchmark than existing metrics, in which the object-text alignment is crucial. We release the benchmark, data, code, and generated images to facilitate future research on culturally diverse T2I generation.
Bingshuai Liu, Longyue Wang, Chenyang Lyu, Yong Zhang 0034, Jinsong Su, Shuming Shi 0001, Zhaopeng Tu
ECAI2
2024 Reassessing Non-Autoregressive Neural Machine Translation with a Fine-Grained Error Taxonomy
abstract
Non-autoregressive neural machine translation (NAT) has made remarkable progress since it is proposed. The performance of NAT in terms of BLEU has approached or even matched that of autoregressive neural machine translation (AT). However, other evaluation metrics show that NAT still lags behind. Unfortunately, these metrics only provide a numerical difference, and it is unclear how the translations produced by NAT differ from those produced by AT. In addition, the multimodality problem is always a significant issue in NAT. To assess whether NAT models are fully capable of solving the multimodality problem and achieving the performance of AT, we specifically design an error taxonomy to annotate errors in translations. The taxonomy is grounded on a systematic and hierarchical error analysis. We carry out an extensive annotation with professional annotators and analyze four NAT models and two AT models. Our analysis and experiments show that (1) the number of errors in NAT translations marked by annotators is 1.54 times that of AT translations, (2) the multimodality problem of NAT affects translations from lexical to syntactic levels, and even up to discourse, and (3) the four NAT models cannot fully eradicate the multimodality problem despite mitigation efforts.
Longyue Wang, Zhaopeng Tu, Deyi Xiong
ECAI2
2024 Alternate Diverse Teaching for Semi-supervised Medical Image Segmentation
Zhen Zhao 0001, Zicheng Wang 0012, Longyue Wang, Dian Yu 0001, Yixuan Yuan, Luping Zhou
ECCV (5)3
2024 Semantic Enrichment for Video Question Answering with Gated Graph Neural Networks
abstract
Video Question Answering (VideoQA) is a complex task that requires a deep understanding of a video to accurately answer questions. Existing methods often struggle to effectively integrate the visual and language-based semantic information, subsequently leading to an incomplete understanding of video content and sub-optimal performance. To address the challenge, we introduce a novel approach in this paper to enrich the semantics of video frames, questions, and answer candidates. Specifically, we parse video frames and questions into semantic graphs - visual semantic graph and question semantic graph, which captures information about objects, their attributes, and relationships. These graphs are then encoded using a Gated Graph Neural Network (GGNN). For answer candidates, we propose to verbalize them using Large Language Models (LLMs) to further inject more semantic information from visual and acoustic aspects. We evaluate our approach on benchmark VideoQA datasets: AVQA and Music-AVQA. Experimental results show that our approach outperforms competitive baseline models, achieving state-of-the-art performance on various question types.
Chenyang Lyu, Wenxi Li, Tianbo Ji, Yi Yu 0001, Longyue Wang
ICASSP5
2024 VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
abstract
Large Multimodal Models (LMMs) have achieved impressive success in visual reasoning, particularly in visual mathematics. However, problem-solving capabilities in graph theory remain less explored for LMMs, despite being a crucial aspect of mathematical reasoning that requires an accurate understanding of graphical structures and multi-step reasoning on visual graphs. To step forward in this direction, we are the first to design a benchmark named VisionGraph, used to explore the capabilities of advanced LMMs in solving multimodal graph theory problems. It encompasses eight complex graph problem tasks, from connectivity to shortest path problems. Subsequently, we present a Description-Program-Reasoning (DPR) chain to enhance the logical accuracy of reasoning processes through graphical structure description generation and algorithm-aware multi-step reasoning. Our extensive study shows that 1) GPT-4V outperforms Gemini Pro in multi-step graph reasoning; 2) All LMMs exhibit inferior perception accuracy for graphical structures, whether in zero/few-shot settings or with supervised fine-tuning (SFT), which further affects problem-solving performance; 3) DPR significantly improves the multi-step graph reasoning capabilities of LMMs and the GPT-4V (DPR) agent achieves SOTA performance.
Yunxin Li, Baotian Hu, Wei Wang 0164, Longyue Wang, Min Zhang 0005
ICML5
2024 GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation
Zhanyu Wang, Longyue Wang, Zhen Zhao 0001, Minghao Wu, Chenyang Lyu, Deng Cai 0002, Luping Zhou, Shuming Shi 0001, Zhaopeng Tu
ACM Multimedia2
2024 Benchmarking LLMs via Uncertainty Quantification
abstract
The proliferation of open-source Large Language Models (LLMs) from various institutions has highlighted the urgent need for comprehensive evaluation methods. However, current evaluation platforms, such as the widely recognized HuggingFace open LLM leaderboard, neglect a crucial aspect -- uncertainty, which is vital for thoroughly assessing LLMs. To bridge this gap, we introduce a new benchmarking approach for LLMs that integrates uncertainty quantification. Our examination involves nine LLMs (LLM series) spanning five representative natural language processing tasks. Our findings reveal that: I) LLMs with higher accuracy may exhibit lower certainty; II) Larger-scale LLMs may display greater uncertainty compared to their smaller counterparts; and III) Instruction-finetuning tends to increase the uncertainty of LLMs. These results underscore the significance of incorporating uncertainty in the evaluation of LLMs. Our implementation is available at https://github.com/smartyfh/LLM-Uncertainty-Bench.
Fanghua Ye 0001, Jianhui Pang, Longyue Wang, Derek F. Wong, Emine Yilmaz, Shuming Shi 0001, Zhaopeng Tu
NeurIPS4
2024 Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
abstract
demands substantial human effort and incurs high training
Yunxin Li, Baotian Hu, Longyue Wang, Jiashun Zhu, Jinyi Xu, Zhen Zhao 0001, Min Zhang 0005
SIGGRAPH Asia4
2024 Learning to Denoise Biomedical Knowledge Graph for Robust Molecular Interaction Prediction
abstract
Molecular interaction prediction plays a crucial role in forecasting unknown interactions between molecules, such as drug-target interaction (DTI) and drug-drug interaction (DDI), which are essential in the field of drug discovery and therapeutics. Although previous prediction methods have yielded promising results by leveraging the rich semantics and topological structure of biomedical knowledge graphs (KGs), they have primarily focused on enhancing predictive performance without addressing the presence of inevitable noise and inconsistent semantics. This limitation has hindered the advancement of KG-based prediction methods. To address this limitation, we propose BioKDN (BiomedicalKnowledge GraphDenoisingNetwork) for robust molecular interaction prediction. BioKDN refines the reliable structure of local subgraphs by denoising noisy links in a learnable manner, providing a general module for extracting task-relevant interactions. To enhance the reliability of the refined structure, BioKDN maintains consistent and robust semantics by smoothing relations around the target interaction. By maximizing the mutual information between reliable structure and smoothed relations, BioKDN emphasizes informative semantics to enable precise predictions. Experimental results on real-world datasets show that BioKDN surpasses state-of-the-art models in DTI and DDI prediction tasks, confirming the effectiveness and robustness of BioKDN in denoising unreliable interactions within contaminated KGs.
Tengfei Ma 0002, Yujie Chen 0002, Wen Tao, Dashun Zheng, Xuan Lin, Patrick Pang 0001, Yijun Wang 0002, Longyue Wang, Bosheng Song, Xiangxiang Zeng, Philip S. Yu
IEEE Trans. Knowl. Data Eng.9
2023 A Survey on Zero Pronoun Translation
abstract
Zero pronouns (ZPs) are frequently omitted in pro-drop languages (e.g.Chinese, Hungarian, and Hindi), but should be recalled in nonpro-drop languages (e.g.English).This phenomenon has been studied extensively in machine translation (MT), as it poses a significant challenge for MT systems due to the difficulty in determining the correct antecedent for the pronoun.This survey paper highlights the major works that have been undertaken in zero pronoun translation (ZPT) after the neural revolution so that researchers can recognize the current state and future directions of this field.We provide an organization of the literature based on evolution, dataset, method, and evaluation.In addition, we compare and analyze competing models and evaluation metrics on different benchmarks.We uncover a number of insightful findings such as: 1) ZPT is in line with the development trend of large language model; 2) data limitation causes learning bias in languages and domains; 3) performance improvements are often reported on single benchmarks, but advanced methods are still far from realworld use; 4) general-purpose metrics are not reliable on nuances and complexities of ZPT, emphasizing the necessity of targeted metrics; 5) apart from commonly-cited errors, ZPs will cause risks of gender bias.
Longyue Wang, Siyou Liu, Mingzhou Xu, Linfeng Song, Shuming Shi 0001, Zhaopeng Tu
ACL (1)1
2023 Document-Level Machine Translation with Large Language Models
abstract
Large language models (LLMs) such as Chat-GPT can produce coherent, cohesive, relevant, and fluent answers for various natural language processing (NLP) tasks.Taking documentlevel machine translation (MT) as a testbed, this paper provides an in-depth evaluation of LLMs' ability on discourse modeling.The study focuses on three aspects: 1) Effects of Context-Aware Prompts, where we investigate the impact of different prompts on document-level translation quality and discourse phenomena; 2) Comparison of Translation Models, where we compare the translation performance of Chat-GPT with commercial MT systems and advanced document-level MT methods; 3) Analysis of Discourse Modelling Abilities, where we further probe discourse knowledge encoded in LLMs and shed light on impacts of training techniques on discourse modeling.By evaluating on a number of benchmarks, we surprisingly find that LLMs have demonstrated superior performance and show potential to become a new paradigm for document-level translation: 1) leveraging their powerful long-text modeling capabilities, GPT-3.5 and GPT-4 outperform commercial MT systems in terms of human evaluation; 1 2) GPT-4 demonstrates a stronger ability for probing linguistic knowledge than GPT-3.5.This work highlights the challenges and opportunities of LLMs for MT, which we hope can inspire the future design and evaluation of LLMs. 2 * Equal contribution.
Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu 0001, Shuming Shi 0001, Zhaopeng Tu
EMNLP1
2023 More Than Spoken Words: Nonverbal Message Extraction and Generation
abstract
Nonverbal messages (NM) such as speakers' facial expressions and speed of speech are essential for face-to-face communication, and they can be regarded as implicit knowledge as they are usually not included in existing dialogue understanding or generation tasks.This paper introduces the task of extracting NMs in written text and generating NMs for spoken text.Previous studies merely focus on extracting NMs from relatively small-scale well-structured corpora such as movie scripts wherein NMs are enclosed in parentheses by scriptwriters, which greatly decreases the difficulty of extraction.To enable extracting NMs from unstructured corpora, we annotate the first NM extraction dataset for Chinese based on novels and develop three baselines to extract single-span or multi-span NM of a target utterance from its surrounding context.Furthermore, we use the extractors to extract 749K (context, utterance, NM) triples from Chinese novels and investigate whether we can use them to improve NM generation via semi-supervised learning.Experimental results demonstrate that the automatically extracted triples can serve as high-quality augmentation data of clean triples extracted from scripts to generate more relevant, fluent, valid, and factually consistent 1 NMs than the purely supervised generator, and the resulting generator can in turn help Chinese dialogue understanding tasks such as dialogue machine reading comprehension and emotion classification by simply adding the predicted "unspoken" NM to each utterance or narrative in inputs.
Dian Yu 0001, Xiaoyang Wang 0001, Wanshun Chen, Longyue Wang, Haitao Mi, Dong Yu 0001
EMNLP5
2023 Towards a Unified Training for Levenshtein Transformer
abstract
Levenshtein Transformer (LevT) is a widely-used text-editing model, which generates a sequence based on editing operations (deletion and insertion) in a non-autoregressive manner. However, it is challenging to train the key refinement components of LevT due to training-inference discrepancy. By carefully designing experiments, our work reveals that the deletion module is under-trained while the insertion module is over-trained due to the imbalance training signals for the two refinement modules. Based on these observations, we further propose a dual learning approach that can remedy the imbalance training by feeding an initial input to both refinement modules, which is consistent with the process in inference. Experimental results on three representative NLP tasks demonstrate the effectiveness and universality of the proposed approach.1
Kangjie Zheng, Longyue Wang, Binqi Chen, Ming Zhang 0004, Zhaopeng Tu
ICASSP2
2023 Prompt-Learning for Cross-Lingual Relation Extraction
abstract
Relation Extraction (RE) is a crucial task in Information Extraction, which entails predicting relationships between entities within a given sentence. However, extending pre-trained RE models to other languages is challenging, particularly in real-world scenarios where Cross-Lingual Relation Extraction (XRE) is required. Despite recent advancements in Prompt-Learning, which involves transferring knowledge from Multilingual Pre-trained Language Models (PLMs) to diverse downstream tasks, there is limited research on the effective use of multilingual PLMs with prompts to improve XRE. In this paper, we present a novel XRE algorithm based on Prompt-Tuning, referred to as Prompt-Xre. To evaluate its effectiveness, we design and implement several prompt templates, including hard, soft, and hybrid prompts, and empirically test their performance on competitive multilingual PLMs, specifically mBART. Our extensive experiments, conducted on the low-resource ACE05 benchmark across multiple languages, demonstrate that our Prompt-Xre algorithm significantly outperforms both vanilla multilingual PLMs and other existing models, achieving state-of-the-art performance in XRE. To further show the generalization of our Prompt-XRE on larger data scales, we construct and release a new XRE dataset-WMTI7-EnZh XRE, containing 0.9M English-Chinese pairs extracted from WMT 2017 parallel corpus. Experiments on WMTI7-EnZh XRE also show the effectiveness of our Prompt-XRE against other competitive baselines. The code and newly constructed dataset are freely available at httus://2ithub.com/HSU-CHIA-MING/Promut-XRE.
Chiaming Hsu, Changtong Zan, Liang Ding 0006, Longyue Wang, Weifeng Liu 0001, Wenbin Hu 0001
IJCNN4
2023 How Does Pretraining Improve Discourse-Aware Translation?
Longyue Wang, Siyou Liu, Derek F. Wong
INTERSPEECH2
2023 Graph-Based Video-Language Learning with Multi-Grained Audio-Visual Alignment
abstract
Video-language learning has attracted significant attention in the fields of multimedia, computer vision and natural language processing in recent years. One of the key challenges in this area is how to effectively integrate visual and linguistic information to enable machines to understand video content and query information. In this work, we leverage graph-based representations and multi-grained audio-visual alignment to address this challenge. First, our approach starts by transforming video and query inputs into visual-scene graphs and semantic role graphs using a visual-scene parser and semantic role labeler respectively. These graphs are then encoded using graph neural networks to obtain enriched representations and combined to obtain a video-query joint representation that enhances the semantic expressivity of the inputs. Second, to achieve accurate matching of relevant parts of audio and visual features, we propose a multi-grained alignment module that aligns the audio and visual features at multiple scales. This enables us to effectively fuse the audio and visual information in a way that is consistent with the semantic-level information captured by the graph-based representations. Experiments on five representative datasets collected for Video Retrieval and Video Question Answering tasks show that our approach outperforms the literature on several metrics. Our extensive ablation studies demonstrate the effectiveness of graph-based representation and multi-grained audio-visual alignment.
Chenyang Lyu, Wenxi Li, Tianbo Ji, Longyue Wang, Liting Zhou, Cathal Gurrin, Linyi Yang, Yi Yu 0001, Yvette Graham, Jennifer Foster
ACM Multimedia4
2023 Interactive Story Visualization with Multiple Characters
abstract
Accurate Story visualization requires several necessary elements, such as identity consistency across frames, the alignment between plain text and visual content, and a reasonable layout of objects in images. Most previous works endeavor to meet these requirements by fitting a text-to-image (T2I) model on a set of videos in the same style and with the same characters, e.g., the FlintstonesSV dataset. However, the learned T2I models typically struggle to adapt to new characters, scenes, and styles, and often lack the flexibility to revise the layout of the synthesized images. This paper proposes a system for generic interactive story visualization, capable of handling multiple novel characters and supporting the editing of layout and local structure. It is developed by leveraging the prior knowledge of large language and T2I models, trained on massive corpora. The system comprises four interconnected components: story-to-prompt generation (S2P), text-to-layout generation (T2L), controllable text-to-image generation (C-T2I), and image-to-video animation (I2V). First, the S2P module converts concise story information into detailed prompts required for subsequent stages. Next, T2L generates diverse and reasonable layouts based on the prompts, offering users the ability to adjust and refine the layout to their preferences. The core component, C-T2I, enables the creation of images guided by layouts, sketches, and actor-specific identifiers to maintain consistency and detail across visualizations. Finally, I2V enriches the visualization process by animating the generated images. Extensive experiments and a user study are conducted to validate the effectiveness and flexibility of interactive editing of the proposed system.
Yuan Gong 0002, Youxin Pang, Xiaodong Cun, Menghan Xia, Yingqing He, Haoxin Chen, Longyue Wang, Yong Zhang 0034, Xintao Wang 0002, Ying Shan, Yujiu Yang 0001
SIGGRAPH Asia7
2023 Search-engine-augmented dialogue response generation with cheaply supervised query production
Ante Wang, Linfeng Song, Qi Liu 0049, Haitao Mi, Longyue Wang, Zhaopeng Tu, Jinsong Su, Dong Yu 0001
Artif. Intell.5
2022 Redistributing Low-Frequency Words: Making the Most of Monolingual Data in Non-Autoregressive Translation
abstract
Knowledge distillation (KD) is the preliminary step for training non-autoregressive translation (NAT) models, which eases the training of NAT models at the cost of losing important information for translating low-frequency words.In this work, we provide an appealing alternative for NAT -monolingual KD, which trains NAT student on external monolingual data with AT teacher trained on the original bilingual data.Monolingual KD is able to transfer both the knowledge of the original bilingual data (implicitly encoded in the trained AT teacher model) and that of the new monolingual data to the NAT student model.Extensive experiments on eight WMT benchmarks over two advanced NAT models show that monolingual KD consistently outperforms the standard KD by improving lowfrequency word translation, without introducing any computational cost.Monolingual KD enjoys desirable expandability, which can be further enhanced (when given more computational budget) by combining with the standard KD, a reverse monolingual KD, or enlarging the scale of monolingual data.Extensive analyses demonstrate that these techniques can be used together profitably to further recall the useful information lost in the standard KD.Encouragingly, combining with standard KD, our approach achieves 30.4 and 34.1 BLEU points on the WMT14 English-German and German-English datasets, respectively.Our code and trained models are freely available at https://github.com/ alphadl/RLFW-NAT.mono.
Liang Ding 0006, Longyue Wang, Shuming Shi 0001, Dacheng Tao, Zhaopeng Tu
ACL (1)2
2022 ngram-OAXE: Phrase-Based Order-Agnostic Cross Entropy for Non-Autoregressive Machine Translation
abstract
Recently, a new training oaxe loss has proven effective to ameliorate the effect of multimodality for non-autoregressive translation (NAT), which removes the penalty of word order errors in the standard cross-entropy loss. Starting from the intuition that reordering generally occurs between phrases, we extend oaxe by only allowing reordering between ngram phrases and still requiring a strict match of word order within the phrases. Extensive experiments on NAT benchmarks across language pairs and data scales demonstrate the effectiveness and universality of our approach. Further analyses show that ngram noaxe indeed improves the translation of ngram phrases, and produces more fluent translation with a better modeling of sentence structure.
Cunxiao Du, Zhaopeng Tu, Longyue Wang
COLING3
2022 GuoFeng: A Benchmark for Zero Pronoun Recovery and Translation
abstract
Mingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu, Linfeng Song, Lidia S. Chao, Shuming Shi, Zhaopeng Tu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Mingzhou Xu, Longyue Wang, Derek F. Wong, Hongye Liu, Linfeng Song, Lidia S. Chao, Shuming Shi 0001, Zhaopeng Tu
EMNLP2
2021 Rejuvenating Low-Frequency Words: Making the Most of Parallel Data in Non-Autoregressive Translation
abstract
Liang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong, Dacheng Tao, Zhaopeng Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Liang Ding 0006, Longyue Wang, Xuebo Liu 0002, Derek F. Wong, Dacheng Tao, Zhaopeng Tu
ACL/IJCNLP (1)2
2021 Understanding and Improving Encoder Layer Fusion in Sequence-to-Sequence Learning
Xuebo Liu 0002, Longyue Wang, Derek F. Wong, Liang Ding 0006, Lidia S. Chao, Zhaopeng Tu
ICLR2
2021 Understanding and Improving Lexical Choice in Non-Autoregressive Translation
Liang Ding 0006, Longyue Wang, Xuebo Liu 0002, Derek F. Wong, Dacheng Tao, Zhaopeng Tu
ICLR2
2021 Context-aware Self-Attention Networks for Natural Language Processing
Baosong Yang, Longyue Wang, Derek F. Wong, Shuming Shi 0001, Zhaopeng Tu
Neurocomputing2
2020 Go From the General to the Particular: Multi-Domain Translation with Domain Transformation Networks
abstract
The key challenge of multi-domain translation lies in simultaneously encoding both the general knowledge shared across domains and the particular knowledge distinctive to each domain in a unified model. Previous work shows that the standard neural machine translation (NMT) model, trained on mixed-domain data, generally captures the general knowledge, but misses the domain-specific knowledge. In response to this problem, we augment NMT model with additional domain transformation networks to transform the general representations to domain-specific representations, which are subsequently fed to the NMT decoder. To guarantee the knowledge transformation, we also propose two complementary supervision signals by leveraging the power of knowledge distillation and adversarial learning. Experimental results on several language pairs, covering both balanced and unbalanced multi-domain translation, demonstrate the effectiveness and universality of the proposed approach. Encouragingly, the proposed unified model achieves comparable results with the fine-tuning approach that requires multiple models to preserve the particular knowledge. Further analyses reveal that the domain transformation networks successfully capture the domain-specific knowledge as expected.1
Yong Wang 0032, Longyue Wang, Shuming Shi 0001, Victor O. K. Li, Zhaopeng Tu
AAAI2
2020 Self-Attention with Cross-Lingual Position Representation
abstract
Position encoding (PE), an essential part of self-attention networks (SANs), is used to preserve the word order information for natural language processing tasks, generating fixed position indices for input sequences.However, in cross-lingual scenarios, e.g., machine translation, the PEs of source and target sentences are modeled independently.Due to word order divergences in different languages, modeling the cross-lingual positional relationships might help SANs tackle this problem.In this paper, we augment SANs with crosslingual position representations to model the bilingually aware latent structure for the input sentence.Specifically, we utilize bracketing transduction grammar (BTG)-based reordering information to encourage SANs to learn bilingual diagonal alignments.Experimental results on WMT'14 English⇒German, WAT'17 Japanese⇒English, and WMT'17 Chinese⇔English translation tasks demonstrate that our approach significantly and consistently improves translation quality over strong baselines.Extensive analyses confirm that the performance gains come from the cross-lingual information.
Liang Ding 0006, Longyue Wang, Dacheng Tao
ACL2
2020 How Does Selective Mechanism Improve Self-Attention Networks?
abstract
Self-attention networks (SANs) with selective mechanism has produced substantial improvements in various NLP tasks by concentrating on a subset of input words.However, the underlying reasons for their strong performance have not been well explained.In this paper, we bridge the gap by assessing the strengths of selective SANs (SSANs), which are implemented with a flexible and universal Gumbel-Softmax.Experimental results on several representative NLP tasks, including natural language inference, semantic role labelling, and machine translation, show that SSANs consistently outperform the standard SANs.Through well-designed probing experiments, we empirically validate that the improvement of SSANs can be attributed in part to mitigating two commonly-cited weaknesses of SANs: word order encoding and structure modeling.Specifically, the selective mechanism improves SANs by paying more attention to content words that contribute to the meaning of the sentence.The code and data are released at https://github.com/xwgeng/SSAN.
Xinwei Geng, Longyue Wang, Xing Wang 0007, Bing Qin 0001, Ting Liu 0001, Zhaopeng Tu
ACL2
2020 Context-Aware Cross-Attention for Non-Autoregressive Translation
abstract
Non-autoregressive translation (NAT) significantly accelerates the inference process by predicting the entire target sequence.However, due to the lack of target dependency modelling in the decoder, the conditional generation process heavily depends on the cross-attention.In this paper, we reveal a localness perception problem in NAT cross-attention, for which it is difficult to adequately capture source context.To alleviate this problem, we propose to enhance signals of neighbour source tokens into conventional cross-attention.Experimental results on several representative datasets show that our approach can consistently improve translation quality over strong NAT baselines.Extensive analyses demonstrate that the enhanced cross-attention achieves better exploitation of source contexts by leveraging both local and global information.
Liang Ding 0006, Longyue Wang, Dacheng Tao, Zhaopeng Tu
COLING2
2020 On the Sparsity of Neural Machine Translation Models
abstract
Modern neural machine translation (NMT) models employ a large number of parameters, which leads to serious over-parameterization and typically causes the underutilization of computational resources.In response to this problem, we empirically investigate whether the redundant parameters can be reused to achieve better performance.Experiments and analyses are systematically conducted on different datasets and NMT architectures.We show that: 1) the pruned parameters can be rejuvenated to improve the baseline model by up to +0.8 BLEU points; 2) the rejuvenated parameters are reallocated to enhance the ability of modeling low-level lexical information.
Yong Wang 0032, Longyue Wang, Victor O. K. Li, Zhaopeng Tu
EMNLP (1)2
2019 Dynamic Layer Aggregation for Neural Machine Translation with Routing-by-Agreement
abstract
With the promising progress of deep neural networks, layer aggregation has been used to fuse information across layers in various fields, such as computer vision and machine translation. However, most of the previous methods combine layers in a static fashion in that their aggregation strategy is independent of specific hidden states. Inspired by recent progress on capsule networks, in this paper we propose to use routing-by-agreement strategies to aggregate layers dynamically. Specifically, the algorithm learns the probability of a part (individual layer representations) assigned to a whole (aggregated representations) in an iterative way and combines parts accordingly. We implement our algorithm on top of the state-of-the-art neural machine translation model TRANSFORMER and conduct experiments on the widely-used WMT14 sh⇒German and WMT17 Chinese⇒English translation datasets. Experimental results across language pairs show that the proposed approach consistently outperforms the strong baseline model and a representative static aggregation model.
Zi-Yi Dou, Zhaopeng Tu, Xing Wang 0007, Longyue Wang, Shuming Shi 0001, Tong Zhang 0001
AAAI4
2019 Exploiting Sentential Context for Neural Machine Translation
abstract
In this work, we present novel approaches to exploit sentential context for neural machine translation (NMT).Specifically, we first show that a shallow sentential context extracted from the top encoder layer only, can improve translation performance via contextualizing the encoding representations of individual words.Next, we introduce a deep sentential context, which aggregates the sentential context representations from all the internal layers of the encoder to form a more comprehensive context representation.Experimental results on the WMT14 English⇒German and English⇒French benchmarks show that our model consistently improves performance over the strong TRANSFORMER model (Vaswani et al., 2017), demonstrating the necessity and effectiveness of exploiting sentential context for NMT.
Xing Wang 0007, Zhaopeng Tu, Longyue Wang, Shuming Shi 0001
ACL (1)3
2019 Assessing the Ability of Self-Attention Networks to Learn Word Order
abstract
Self-attention networks (SAN) have attracted a lot of interests due to their high parallelization and strong performance on a variety of NLP tasks, e.g. machine translation.Due to the lack of recurrence structure such as recurrent neural networks (RNN), SAN is ascribed to be weak at learning positional information of words for sequence modeling.However, neither this speculation has been empirically confirmed, nor explanations for their strong performances on machine translation tasks when "lacking positional information" have been explored.To this end, we propose a novel word reordering detection task to quantify how well the word order information learned by SAN and RNN.Specifically, we randomly move one word to another position, and examine whether a trained model can detect both the original and inserted positions.Experimental results reveal that: 1) SAN trained on word reordering detection indeed has difficulty learning the positional information even with the position embedding; and 2) SAN trained on machine translation learns better positional information than its RNN counterpart, in which position embedding plays a critical role.Although recurrence structure make the model more universally-effective on learning word order, learning objectives matter more in the downstream tasks such as machine translation.
Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, Zhaopeng Tu
ACL (1)2
2019 Towards Understanding Neural Machine Translation with Word Importance
abstract
Shilin He, Zhaopeng Tu, Xing Wang, Longyue Wang, Michael Lyu, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shilin He, Zhaopeng Tu, Xing Wang 0007, Longyue Wang, Michael R. Lyu, Shuming Shi 0001
EMNLP/IJCNLP (1)4
2019 One Model to Learn Both: Zero Pronoun Prediction and Translation
abstract
Longyue Wang, Zhaopeng Tu, Xing Wang, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Longyue Wang, Zhaopeng Tu, Xing Wang 0007, Shuming Shi 0001
EMNLP/IJCNLP (1)1
2019 Self-Attention with Structural Position Representations
abstract
Xing Wang, Zhaopeng Tu, Longyue Wang, Shuming Shi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xing Wang 0007, Zhaopeng Tu, Longyue Wang, Shuming Shi 0001
EMNLP/IJCNLP (1)3
2018 Translating Pro-Drop Languages With Reconstruction Models
abstract
Pronouns are frequently omitted in pro-drop languages, such as Chinese, generally leading to significant challenges with respect to the production of complete translations. To date, very little attention has been paid to the dropped pronoun (DP) problem within neural machine translation (NMT). In this work, we propose a novel reconstruction-based approach to alleviating DP translation problems for NMT models. Firstly, DPs within all source sentences are automatically annotated with parallel information extracted from the bilingual training corpus. Next, the annotated source sentence is reconstructed from hidden representations in the NMT model. With auxiliary training objectives, in the terms of reconstruction scores, the parameters associated with the NMT model are guided to produce enhanced hidden representations that are encouraged as much as possible to embed annotated DP information. Experimental results on both Chinese-English and Japanese-English dialogue translation tasks show that the proposed approach significantly and consistently improves translation performance over a strong NMT baseline, which is directly built on the training data annotated with DPs.
Longyue Wang, Zhaopeng Tu, Shuming Shi 0001, Tong Zhang 0001, Yvette Graham, Qun Liu 0001
AAAI1
2018 Learning to Jointly Translate and Predict Dropped Pronouns with a Shared Reconstruction Mechanism
abstract
Pronouns are frequently omitted in pro-drop languages, such as Chinese, generally leading to significant challenges with respect to the production of complete translations.Recently, Wang et al. (2018) proposed a novel reconstruction-based approach to alleviating dropped pronoun (DP) translation problems for neural machine translation models.In this work, we improve the original model from two perspectives.First, we employ a shared reconstructor to better exploit encoder and decoder representations.Second, we jointly learn to translate and predict DPs in an end-to-end manner, to avoid the errors propagated from an external DP prediction model.Experimental results show that our approach significantly improves both translation performance and DP prediction accuracy.
Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001
EMNLP1
2018 Chinese-Portuguese Machine Translation: A Study on Building Parallel Corpora from Comparable Texts
Siyou Liu, Longyue Wang, Chao-Hong Liu
LREC2
2018 IDEA: An Interactive Dialogue Translation Demo System Using Furhat Robots
Jinhua Du, Darragh Blake, Longyue Wang, Clare Conran, Declan McKibben, Andy Way
ECML/PKDD (3)3
2017 Exploiting Cross-Sentence Context for Neural Machine Translation
abstract
In translation, considering the document as a whole can help to resolve ambiguities and inconsistencies.In this paper, we propose a cross-sentence context-aware approach and investigate the influence of historical contextual information on the performance of neural machine translation (NMT).First, this history is summarized in a hierarchical way.We then integrate the historical representation into NMT in two strategies: 1) a warm-start of encoder and decoder states, and 2) an auxiliary context source for updating decoder states.Experimental results on a large Chinese-English translation task show that our approach significantly improves upon a strong attention-based NMT system by up to +2.1 BLEU points.
Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001
EMNLP1
2017 A novel and robust approach for pro-drop language translation
abstract
A significant challenge for machine translation (MT) is the phenomena of dropped pronouns (DPs), where certain classes of pronouns are frequently dropped in the source language but should be retained in the target language. In response to this common problem, we propose a semi-supervised approach with a universal framework to recall missing pronouns in translation. Firstly, we build training data for DP generation in which the DPs are automatically labelled according to the alignment information from a parallel corpus. Secondly, we build a deep learning-based DP generator for input sentences in decoding when no corresponding references exist. More specifically, the generation has two phases: (1) DP position detection, which is modeled as a sequential labelling task with recurrent neural networks; and (2) DP prediction, which employs a multilayer perceptron with rich features. Finally, we integrate the above outputs into our statistical MT (SMT) system to recall missing pronouns by both extracting rules from the DP-labelled training data and translating the DP-generated input sentences. To validate the robustness of our approach, we investigate our approach on both Chinese–English and Japanese–English corpora extracted from movie subtitles. Compared with an SMT baseline system, experimental results show that our approach achieves a significant improvement of $$+$$ 1.58 BLEU points in translation performance with 66% F-score for DP generation accuracy for Chinese–English, and nearly $$+$$ 1 BLEU point with 58% F-score for Japanese–English. We believe that this work could help both MT researchers and industries to boost the performance of MT systems between pro-drop and non-pro-drop languages.
Longyue Wang, Zhaopeng Tu, Siyou Liu, Hang Li 0001, Andy Way, Qun Liu 0001
Mach. Transl.1
2016 Dropped pronoun generation for dialogue machine translation
abstract
Dropped pronoun (DP) is a common problem in dialogue machine translation, in which pronouns are frequently dropped in the source sentence and thus are missing in its translation. In response to this problem, we propose a novel approach to improve the translation of DPs for dialogue machine translation. Firstly, we build a training data for DP generation, in which the DPs are automatically added according to the alignment information from a parallel corpus. Then we model the DP generation problem as a sequence labelling task, and develop a generation model based on recurrent neural networks and language models. Finally, we apply the DP generator to machine translation task by completing the source sentences with the missing pronouns. Experimental results show that our approach achieves a significant improvement of 1.7 BLEU points by recalling possible DPs in the source sentences.
Longyue Wang, Zhaopeng Tu, Hang Li 0001, Qun Liu 0001
ICASSP1
2016 Automatic Construction of Discourse Corpora for Dialogue Translation
Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001
LREC1
2016 A Novel Approach to Dropped Pronoun Translation
abstract
Longyue Wang, Zhaopeng Tu, Xiaojun Zhang, Hang Li, Andy Way, Qun Liu. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Longyue Wang, Zhaopeng Tu, Hang Li 0001, Andy Way, Qun Liu 0001
HLT-NAACL1
2015 Linguistically-augmented perplexity-based data selection for language models
abstract
This paper explores the use of linguistic information for the selection of data to train language models. We depart from the state-of-the-art method in perplexity-based data selection and extend it in order to use word-level linguistic units (i.e. lemmas, named entity categories and part-of-speech tags) instead of surface forms. We then present two methods that combine the different types of linguistic knowledge as well as the surface forms (1, naïve selection of the top ranked sentences selected by each method; 2, linear interpolation of the datasets selected by the different methods). The paper presents detailed results and analysis for four languages with different levels of morphologic complexity (English, Spanish, Czech and Chinese). The interpolation-based combination outperforms the purely statistical baseline in all the scenarios, resulting in language models with lower perplexity. In relative terms the improvements are similar regardless of the language, with perplexity reductions achieved in the range 7.72–13.02%. In absolute terms the reduction is higher for languages with high type-token ratio (Chinese, 202.16) or rich morphology (Czech, 81.53) and lower for the remaining languages, Spanish (55.2) and English (34.43 on the English side of the same parallel dataset as for Czech and 61.90 on the same parallel dataset as for Spanish).
Antonio Toral, Pavel Pecina, Longyue Wang, Josef van Genabith
Comput. Speech Lang.3