Jian Luan 0001

dblp:61/3233-1 · DBLP profile ↗
← Back
59ranked-venue papers
2as first author
51since 2021 · last 2026
0000-0002-2383-226XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 2 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 2 first-author · 21 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
abstract
Sound effect editing—modifying audio by adding, removing, or replacing elements—remains constrained by existing approaches that rely solely on low-level signal processing or coarse text prompts, often resulting in limited flexibility and suboptimal audio quality. To address this, we propose AV-Edit, a generative sound effect editing framework that enables fine-grained editing of existing audio tracks in videos by jointly leveraging visual, audio, and text semantics. Specifically, the proposed method employs a specially designed contrastive audio-visual masking autoencoder (CAV-MAE-Edit) for multimodal pre-training, learning aligned cross-modal representations. These representations are then used to train an editorial Multimodal Diffusion Transformer (MM-DiT) capable of removing visually irrelevant sounds and generating missing audio elements consistent with video content through a correlation-based feature gating training strategy. Furthermore, we construct a dedicated video-based sound editing dataset as an evaluation benchmark. Experiments demonstrate that the proposed AV-Edit generates high-quality audio with precise modifications based on visual content, achieving state-of-the-art performance in the field of sound effect editing and exhibiting strong competitiveness in the domain of audio generation.
Xinyue Guo 0001, Lipan Zhang, Jianxuan Yang, Jian Luan 0001
AAAI6
2026 End-to-End Optimization of LLM-Driven Multi-Agent Search Systems via Heterogeneous-Group-Based Reinforcement Learning
abstract
Large language models (LLMs) are versatile, yet their deployment in complex real-world settings is limited by static knowledge cutoffs and the difficulty of producing controllable behavior within a single inference.Multi-agent search systems (MASS), which coordinate specialized LLM agents equipped with search tools, mitigate these issues via task decomposition and retrieval-augmented problem solving.However, optimizing LLMs for agent-specific roles remains labor-intensive with prompt engineering or supervised fine-tuning, motivating automated end-to-end training.Existing multiagent reinforcement learning (MARL) methods such as Multi-Agent Proximal Policy Optimization (MAPPO) typically depend on large critic networks to evaluate joint actions, leading to instability and high memory costs.We introduce Multi-Agent Heterogeneous Group Policy Optimization (MHGPO), which updates policies by estimating relative advantages across heterogeneous groups of multi-agent rollouts, shifting the optimization focus from local agent performance to global system success.We further study three group rollout sampling strategies to trade off sample efficiency and optimization quality.Experiments show that MHGPO captures implicit inter-agent dependencies and consistently outperforms strong baselines in both task performance and computational efficiency.
Shaoxiong Yang, Chao Li 0033, Wei Liu 0302, Jian Luan 0001, Zenglin Xu
ACL (1)5
2026 VecInfer: Efficient LLM Inference with Low-Bit KV Cache via Outlier-Suppressed Vector Quantization
abstract
Dingyu Yao, Chenxu Yang, Zhengyang Tong, Zheng Lin, Wei Liu, Jian Luan, Weiping Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Dingyu Yao, Chenxu Yang, Zhengyang Tong, Zheng Lin 0001, Wei Liu 0302, Jian Luan 0001, Weiping Wang 0005
ACL (1)6
2026 Attention Basin: Why Contextual Position Matters in Large Language Models
abstract
Zihao Yi, Zhenqing Ling, Delong Zeng, Haohao Luo, Zhe Xu, Wei Liu, Jian Luan, Wanxia Cao, Ying Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zihao Yi, Zhenqing Ling, Delong Zeng, Haohao Luo, Zhe Xu 0009, Wei Liu 0302, Jian Luan 0001, Wanxia Cao, Ying Shen 0001
ACL (1)7
2026 Doc-V^*: Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
abstract
Yuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang, Yuyi Zhang, Wenyu Ruan, Xiaojin Zhang, Zhongyu Wei, Zhenbo Luo, Jian Luan, Wei Chen, Xiang Bai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuanlei Zheng, Pei Fu, Hang Li 0001, Wenyu Ruan, Xiaojin Zhang 0002, Zhongyu Wei, Zhenbo Luo, Jian Luan 0001, Wei Chen 0088, Xiang Bai
ACL (1)10
2026 It Takes Two: Embracing Sparsity and Speculative Decoding for Efficient LLM Inference
Haolin Chu, Changyu Chen, Jian Luan 0001, Jiabin Deng, Huadong Ma, Xiaolong Zheng 0002
IWQoS4
2026 Mobile GUI Agents under Real-world Threats: Are We There Yet?
abstract
Recent years have witnessed a rapid development of mobile GUI agents powered by large language models (LLMs), which can autonomously execute diverse device-control tasks based on natural language instructions. The increasing accuracy of these agents on standard benchmarks has raised expectations for large-scale real-world deployment, and there are already several commercial agents released and used by early adopters. However, are we really ready for GUI agents integrated into our daily devices as system building blocks? We argue that an important pre-deployment validation is missing to examine whether the agents can maintain their performance under real-world threats. Specifically, unlike existing common benchmarks that are based on simple static app contents (they have to do so to ensure environment consistency between different tests), real-world apps are filled with contents from untrustworthy third parties, such as advertisement emails, user-generated posts and medias, etc. These contents may inevitably appear in the agents' observation space and influence the task execution process. Systematic investigation of this problem is challenging since the real-world app contents are significantly skewed—testing on normal real-world apps usually cannot uncover any potential risk since most app contents are benign. To this end, we introduce a scalable app content instrumentation framework to enable flexible and targeted content modifications within existing applications. Leveraging this framework, we create a test suite comprising both a dynamic task execution environment and a static dataset of challenging GUI states. The dynamic environment encompasses 122 reproducible tasks, and the static dataset consists of over 3,000 scenarios constructed from commercial apps. We perform experiments on both open-source and commercial GUI agents. Our findings reveal that all examined agents can be significantly degraded due to third-party contents, with an average misleading rate of 42.0% and 36.1% in dynamic and static environments respectively. The framework and benchmark has been released at https://agenthazard.github.io.
Guohong Liu 0002, Jialei Ye, Wei Liu 0302, Pengzhi Gao, Jian Luan 0001, Yuanchun Li 0003, Yunxin Liu 0001
MobiSys6
2026 C2-Cite: Contextual-Aware Citation Generation for Attributed Large Language Models
abstract
The attribution technique enhances the credibility of LLMs by adding citations to the generated sentences, enabling users to trace back to the original sources and verify the reliability of the output. However, existing instruction-tuned attributed LLMs often fail to properly interpret the contextual semantics of citation symbols (e.g., [i]) during text generation. This shortcoming arises from their insufficient awareness of the context information surrounding citation markers, which in turn leads to disjointed references and poor integration of retrieved knowledge into the generated content. To address this issue, we propose a novel Contextual-aware Citation generation framework (C²-Cite) that explicitly integrates the semantic relationships between citation markers and their referenced content. Specifically, a contextual citation alignment mechanism is adopted: it first encodes the retrieved document contexts into the symbol representation of citations, then aligns the marker numbers by decoding information from a citation router function. This mechanism enables the transformation of citation markers from generic placeholders into active knowledge pointers that link to the referenced source information. Experimental results on the ALCE benchmark across three datasets validate our framework C²-Cite++: it outperforms the SOTA baseline by an average of 5.8% in citation quality and 17.4% in response correctness. The implementation is publicly available at https://github.com/BAI-LAB/c2cite
Yue Yu 0007, Ting Bai 0004, Hengzhi Lan, Jie Wu 0017, Wei Liu 0302, Jian Luan 0001, Chuan Shi 0001
WSDM8
2026 ZoneSep: A Lightweight End-to-End Neural Beamformer With Post-Mask Decoder for In-Vehicle Multi-Zone Speech Separation
abstract
Recently, the task of in-vehicle multi-zone speech separation (IMSS) has attracted significant interest. However, a persistent issue in this area is the cross-zone leakage problem. To address this, we present ZoneSep, a lightweight, fully end-to-end neural beamformer equipped with a post-mask decoder to suppress leakage effectively. Additionally, we propose a novel loss function, SC-SI-SDR loss, which enables simultaneous optimization of both non-silent and silent zones. Experiments on the IMSS dataset demonstrate the effectiveness of our approaches, showing that ZoneSep significantly outperforms recent advanced methods while requiring only 0.65M parameters and 4.02 GMACs. The source code is available athttps://github.com/EeLLJ/ZoneSep/.
Longjie Luo, Junnan Wu, Lichun Fan, Zhenbo Luo, Jian Luan 0001, Qingyang Hong, Lin Li 0032
IEEE Signal Process. Lett.5
2025 Stability and Generalization of Zeroth-Order Decentralized Stochastic Gradient Descent with Changing Topology
abstract
Zeroth-order (ZO) optimization as the gradient-free method has become a powerful tool when the first-order gradient is unavailable or expensive to obtain, especially in decentralized learning scenarios where data and computational resources are distributed across multiple clients. There have been many efforts to analyze the optimization convergence rate of zeroth-order decentralized stochastic gradient descent (ZO-DSGD) algorithms. However, the generalization of these methods has not been well studied. In this paper, we provide a generalization analysis of ZO-DSGD with changing topology, where the clients run zeroth-order SGD with local data and communicate with each other according to time-varying topology. We systematically analyze the generalization error in convex, strongly convex, and non-convex cases. The obtained results in the convex and strongly convex cases with zeroth-order oracles recover the results of SGD. Moreover, the generalization bounds derived in non-convex cases align with that of DSGD. To capture the influence of communication topology on the generalization performance, we analyze local generalization bounds concerning local models held at different clients. The obtained results reflect the influence of the number of clients, local sample size, and topology on the generalization error. To the best of our knowledge, this is the first work that provides a generalization analysis of zeroth-order decentralized stochastic gradient descent methods and recovers the results of SGD.
Xiaolin Hu 0001, Zixuan Gong, Gengze Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Yong Liu 0020
AAAI5
2025 HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
abstract
Many positional encodings (PEs) are designed to exhibit long-term decay, based on an entrenched and long-standing inductive opinion: tokens farther away from the current position carry less relevant information. We argue that long-term decay is outdated in the era of LLMs, as LLMs are now applied to tasks demanding precise retrieval of in-context information from arbitrary positions. Firstly, we present empirical analyses on various PEs, demonstrating that models inherently learn attention with only a local-decay pattern while forming a U-shape pattern globally, contradicting the principle of long-term decay. Furthermore, we conduct a detailed analysis of rotary position encoding (RoPE, a prevalent relative positional encoding in LLMs), and found that the U-shape attention is caused by some learned components, which are also the key factor limiting RoPE’s expressiveness and extrapolation. Inspired by these insights, we propose High-frequency rotary Position Encoding (HoPE). HoPE replaces the specific components in RoPE with position-independent ones, retaining only high-frequency signals, which also breaks the principle of long-term decay in theory. HoPE achieves two major advantages: (1) Without constraints imposed by long-term decay, contradictory factors that limit attention optimization are removed. Thus, the model’s context awareness is enhanced. (2) HoPE exhibits greater robustness to the out-of-distribution behavior in attention patterns during extrapolation. The effectiveness of HoPE is validated through extensive experiments and with a large language model of up to 3 billion parameters.
Yuhan Chen 0001, Ang Lv, Jian Luan 0001, Bin Wang 0004, Wei Liu 0302
ACL (1)3
2025 Demystifying Small Language Models for Edge Deployment
abstract
Small language models (SLMs) have emerged as a promising solution for deploying resource-constrained devices, such as smartphones and Web of Things. This work presents the first comprehensive study of over 60 SLMs such as Microsoft Phi and Google Gemma that are publicly accessible. Our findings show that state-of-the-art SLMs outperform 7B models in general tasks, proving their practical viability. However, SLMs’ in-context learning capabilities remain limited, and their efficiency has significant optimization potential. We identify key SLM optimization opportunities, including dynamic task-specific routing, model-hardware co-design, and vocabulary/KV cache compression. Overall, we expect the work to reveal an all-sided landscape of SLMs, benefiting the research community across algorithm, model, system, and hardware levels.
Zhenyan Lu, Xiang Li 0067, Dongqi Cai 0001, Rongjie Yi, Fangming Liu, Wei Liu 0302, Jian Luan 0001, Nicholas D. Lane, Mengwei Xu 0001
ACL (1)7
2025 Global Eye: Breaking the "Fixed Thinking Pattern" during the Instruction Expansion Process
abstract
An extensive high-quality instruction dataset is crucial for the instruction tuning process of Large Language Models (LLMs).Recent instruction expansion methods have demonstrated their capability to improve the quality and quantity of existing datasets, by prompting high-performance LLM to generate multiple new instructions from the original ones.However, existing methods focus on constructing multi-perspective prompts (e.g., increasing complexity or difficulty) to expand instructions, overlooking the "Fixed Thinking Pattern" issue of LLMs.This issue arises when repeatedly using the same set of prompts, causing LLMs to rely on a limited set of certain expressions to expand all instructions, potentially compromising the diversity of the final expanded dataset.This paper theoretically analyzes the causes of the "Fixed Thinking Pattern", and corroborates this phenomenon through multi-faceted empirical research.Furthermore, we propose a novel method based on dynamic prompt updating: Global Eye.Specifically, after a fixed number of instruction expansions, we analyze the statistical characteristics of newly generated instructions and then update the prompts.Experimental results show that our method enables LLaMA3-8B and LLaMA2-13B to surpass the performance of open-source LLMs and GPT3.5 across various metrics.
Wenxuan Lu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Songhao Jiang, Tianning Zang
ACL (1)3
2025 Browsing Like Human: A Multimodal Web Agent with Experiential Fast-and-Slow Thinking
abstract
Automating web navigation which aims to build a web agent that follows user instructions to complete tasks like booking flights by interacting with websites, has received increasing attention due to its practical value.Although existing web agents are mostly equipped with visual perception, planning, and memory abilities, their reasoning process are still deviate from human cognition.In this work, we study the human thought pattern to empower agent with more human-like abilities in web navigation.To tackle this problem, we propose a novel multimodal web agent framework called WebExperT, which is designed to emulate the human planning process of "thinking fast and slow" to effectively decompose complex user instructions.Furthermore, WebExperT leverages experiential learning by reflecting from failure for continuously refining planning and decision-making outcomes.Experimental results on the MIND2WEB benchmark demonstrate the superiority of WebExperT in both supervised and unsupervised settings.
Haohao Luo, Jiayi Kuang, Wei Liu 0302, Ying Shen 0001, Jian Luan 0001, Yang Deng 0002
ACL (1)5
2025 Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
abstract
Vision-language models (VLMs) achieve remarkable success in single-image tasks.However, real-world scenarios often involve intricate multi-image inputs, leading to a notable performance decline as models struggle to disentangle critical information scattered across complex visual features.In this work, we propose Focus-Centric Visual Chain, a novel paradigm that enhances VLMs' perception, comprehension, and reasoning abilities in multi-image scenarios.To facilitate this paradigm, we propose Focus-Centric Data Synthesis, a scalable bottom-up approach for synthesizing high-quality data with elaborate reasoning paths.Through this approach, We construct VISC-150K, a large-scale dataset with reasoning data in the form of Focus-Centric Visual Chain, specifically designed for multi-image tasks.Experimental results on seven multi-image benchmarks demonstrate that our method achieves average performance gains of 3.16% and 2.24% across two distinct model architectures, without compromising the general vision-language capabilities.Our study represents a significant step toward more robust and capable vision-language systems that can handle complex visual scenarios: VISC. * Corresponding authors.Which of the following images contains the same object as the first image and shares the same attribute weight?
Juntian Zhang, Chuanqi Cheng, Yuhan Liu 0023, Wei Liu 0302, Jian Luan 0001, Rui Yan 0001
ACL (1)5
2025 More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives
abstract
Xiaoqing Zhang, Ang Lv, Yuhan Liu, Flood Sung, Wei Liu, Jian Luan, Shuo Shang, Xiuying Chen, Rui Yan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xiaoqing Zhang 0017, Ang Lv, Yuhan Liu 0023, Flood Sung, Wei Liu 0302, Jian Luan 0001, Shuo Shang, Xiuying Chen, Rui Yan 0001
ACL (1)6
2025 DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization
abstract
Songs, as a central form of musical art, exemplify the richness of human intelligence and creativity. While recent advances in generative modeling have enabled notable progress in long-form song generation, current systems for fulllength song synthesis still face major challenges, including data imbalance, insufficient controllability, and inconsistent musical quality. DiffRhythm, a pioneering diffusion-based model, advanced the field by generating full-length songs with expressive vocals and accompaniment. However, its performance was constrained by an unbalanced model training dataset and limited controllability over musical style, resulting in noticeable quality disparities and restricted creative flexibility. To address these limitations, we propose DiffRhythm+, an enhanced diffusionbased framework for controllable and flexible full-length song generation. DiffRhythm+ leverages a substantially expanded and balanced training dataset to mitigate issues such as repetition and omission of lyrics, while also fostering the emergence of richer musical skills and expressiveness. The framework introduces a multi-modal style conditioning strategy, enabling users to precisely specify musical styles through both descriptive text and reference audio, thereby significantly enhancing creative control and diversity. We further introduce direct performance optimization aligned with user preferences, guiding the model toward consistently preferred outputs across evaluation metrics. Extensive experiments demonstrate that DiffRhythm+ achieves significant improvements in naturalness, arrangement complexity, and listener satisfaction over previous systems. Audio samples are available at https://longwaytog0.github.io/DiffRhythmPlus/.
Huakang Chen, Yuepeng Jiang, Guobin Ma, Chunbo Hao, Shuai Wang 0016, Jixun Yao, Ziqian Ning, Jian Luan 0001, Lei Xie 0001
ASRU9
2025 PMSS: Pretrained Matrices Skeleton Selection for LLM Fine-tuning
abstract
Low-rank adaptation (LoRA) and its variants have recently gained much interest due to their ability to avoid excessive inference costs. However, LoRA still encounters the following challenges: (1) Limitation of low-rank assumption; and (2) Its initialization method may be suboptimal. To this end, we propose PMSS(Pre-trained Matrices Skeleton Selection), which enables high-rank updates with low costs while leveraging semantic and linguistic information inherent in pre-trained weight. It achieves this by selecting skeletons from the pre-trained weight matrix and only learning a small matrix instead. Experiments demonstrate that PMSS outperforms LoRA and other fine-tuning methods across tasks with much less trainable parameters. We demonstrate its effectiveness, especially in handling complex tasks such as DROP benchmark(+3.4%/+5.9% on LLaMA2-7B/13B) and math reasoning (+12.89%/+5.61%/+3.11% on LLaMA2-7B, Mistral-7B and Gemma-7B of GSM8K).The code and model will be released soon.
Qibin Wang, Xiaolin Hu 0001, Weikai Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004
COLING5
2025 MAKAR: a Multi-Agent framework based Knowledge-Augmented Reasoning for Grounded Multimodal Named Entity Recognition
abstract
Xinkui Lin, Yuhui Zhang, Yongxiu Xu, Kun Huang, Hongzhang Mu, Yubin Wang, Gaopeng Gou, Li Qian, Li Peng, Wei Liu, Jian Luan, Hongbo Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Xinkui Lin, Yongxiu Xu, Hongzhang Mu, Gaopeng Gou, Wei Liu 0302, Jian Luan 0001
EMNLP11
2025 BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
abstract
Graphical User Interface (GUI) agents have gained substantial attention due to their impressive capabilities to complete tasks through multiple interactions within GUI environments.However, existing agents primarily focus on enhancing the accuracy of individual actions and often lack effective mechanisms for detecting and recovering from errors.To address these shortcomings, we propose the BacktrackAgent, a robust framework that incorporates a backtracking mechanism to improve task completion efficiency.BacktrackAgent includes verifier, judger, and reflector components as modules for error detection and recovery, while also applying judgment rewards to further enhance the agent's performance.Additionally, we develop a training dataset specifically designed for the backtracking mechanism, which considers the outcome pages after action executions.Experimental results show that BacktrackAgent has achieved performance improvements in both task success rate and step accuracy on Mobile3M and Auto-UI benchmarks.Our data and code will be released upon acceptance.
Qinzhuo Wu, Pengzhi Gao, Wei Liu 0302, Jian Luan 0001
EMNLP4
2025 LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
abstract
Recent advances in large vision-language models (LVLMs) typically employ vision encoders based on the Vision Transformer (ViT) architecture. The division of the images into patches by ViT results in a fragmented perception, thereby hindering the visual understanding capabilities of LVLMs. In this paper, we propose an innovative enhancement to address this limitation by introducing a Scene Graph Expression (SGE) module in LVLMs. This module extracts and structurally expresses the complex semantic information within images, thereby improving the foundational perception and understanding abilities of LVLMs. Extensive experiments demonstrate that integrating our SGE module significantly enhances the LVLM’s performance in vision-language tasks, indicating its effectiveness in preserving intricate semantic details and facilitating better visual understanding.
Jianzhong Ju, Jian Luan 0001, Zhidong Deng
ICASSP3
2025 Q-Frame: Query-Aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
abstract
Multimodal Large Language Models (MLLMs) have demonstrated significant success in visual understanding tasks. However, challenges persist in adapting these models for video comprehension due to the large volume of data and temporal complexity. Existing Video-LLMs using uniform frame sampling often struggle to capture the query-related crucial spatiotemporal clues of videos effectively. In this paper, we introduce Q-Frame, a novel approach for adaptive frame selection and multi-resolution scaling tailored to the video's content and the specific query. Q-Frame employs a training-free, plug-and-play strategy generated by a text-image matching network like CLIP, utilizing the Gumbel-Max trick for efficient frame selection. Q-Frame allows Video-LLMs to process more frames without exceeding computational limits, thereby preserving critical temporal and spatial information. We demonstrate Q-Frame's effectiveness through extensive experiments on benchmark datasets, including MLVU, LongVideoBench, and Video-MME, illustrating its superiority over existing methods and its applicability across various video understanding tasks.
Shaojie Zhang 0004, Jianqin Yin, Zhenbo Luo, Jian Luan 0001
ICCV5
2025 Theoretical Insights into Fine-Tuning Attention Mechanism: Generalization and Optimization
abstract
Large Language Models (LLMs), built on Transformer architectures, exhibit remarkable generalization across a wide range of tasks. However, fine-tuning these models for specific tasks remains resource-intensive due to their extensive parameterization. In this paper, we explore two remarkable phenomena related to the attention mechanism during the fine-tuning of LLMs (where Wq, Wk, and Wv denote the weights of the query, key, and value layers, respectively). The first phenomenon, termed “Unequal Importance of Attention Matrices”, highlights the impact of fine-tuning different weight matrices. It shows that optimizing the Wv matrix yields significantly better performance than optimizing the Wk matrix. Fine-tuning only the Wq and Wv matrices is computationally efficient while delivering results comparable to, or even better than fine-tuning all three matrices (Wq, Wk, and Wv). The second phenomenon, “Attention Matrices with Customized Learning Rate Lead to Better Convergence”, emphasizes the importance of assigning distinct learning rates to these matrices. Specifically, a higher learning rate for the Wv matrix compared to Wq and Wk accelerates convergence and improves performance. Building on these insights, we propose a new strategy that improves fine-tuning efficiency in terms of both storage and time. Experimental results on benchmark datasets validate the effectiveness of this approach, supporting our theoretical findings. Our analysis lays the theoretical groundwork for configuring and improving algorithms in LLMs fine-tuning.
Xinhao Yao, Hongjin Qian, Xiaolin Hu 0001, Gengze Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Yong Liu 0020
IJCAI6
2025 GLCLAP: A Novel Contrastive Learning Pre-trained Model for Contextual Biasing in ASR
Yuxiang Kong, Fan Cui, Liyong Guo, Heinrich Dinkel, Lichun Fan, Jian Luan 0001
INTERSPEECH7
2025 StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
Fengjin Li, Yadong Niu, Jian Luan 0001, Zhiyong Wu 0001
INTERSPEECH6
2025 Text-Enhanced Audio Encoder for Large Language Model based Speech Recognition via Cross-Modality Pre-training with Unpaired Audio-Text Data
Yuxiang Kong, Lichun Fan, Jian Luan 0001
INTERSPEECH4
2025 Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders
Xingwei Sun, Heinrich Dinkel, Yadong Niu, Linzhang Wang, Jian Luan 0001
INTERSPEECH6
2025 X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
Heinrich Dinkel, Yadong Niu, Anbei Zhao, Jian Luan 0001
INTERSPEECH7
2025 MobileSteward: Integrating Multiple App-Oriented Agents with Self-Evolution to Automate Cross-App Instructions
abstract
Mobile phone agents can assist people in automating daily tasks on their phones, which have emerged as a pivotal research spotlight. However, existing procedure-oriented agents struggle with cross-app instructions, due to the following challenges: (1) complex task relationships, (2) diverse app environment, and (3) error propagation and information loss in multi-step execution. Drawing inspiration from object-oriented programming principles, we recognize that object-oriented solutions is more suitable for cross-app instruction. To address these challenges, we propose a self-evolving multi-agent framework named MobileSteward which integrates multiple app-oriented StaffAgents coordinated by a centralized StewardAgent. We design three specialized modules in MobileSteward: (1) Dynamic Recruitment generates a scheduling graph guided by information flow to explicitly associate tasks among apps. (2) Assigned Execution assigns the task to app-oriented StaffAgents, each equipped with app-specialized expertise to address the diversity between apps. (3) Adjusted Evaluation conducts evaluation to provide reflection tips or deliver key information, which alleviates error propagation and information loss during multi-step execution. To continuously improve the performance of MobileSteward, we develop a Memory-based Self-evolution mechanism, which summarizes the experience from successful execution, to improve the performance of MobileSteward. We establish the first English Cross-APP Benchmark (CAPBench) in the real-world environment to evaluate the agents' capabilities of solving complex cross-app instructions. Experimental results demonstrate that MobileSteward achieves the best performance compared to both single-agent and multi-agent frameworks, highlighting the superiority of MobileSteward in better handling user instructions with diverse complexity.
Yuxuan Liu 0009, Hongda Sun 0001, Wei Liu 0302, Jian Luan 0001, Bo Du 0001, Rui Yan 0001
KDD (1)4
2025 Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
abstract
Menglong Cui, Pengzhi Gao, Wei Liu, Jian Luan, Bin Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Menglong Cui, Pengzhi Gao, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004
NAACL (Long Papers)4
2025 ReachAgent: Enhancing Mobile Agent via Page Reaching and Operation
abstract
Qinzhuo Wu, Wei Liu, Jian Luan, Bin Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Qinzhuo Wu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004
NAACL (Long Papers)3
2025 Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
abstract
Large Language Models (LLMs) achieve superior performance through Chain-of-Thought (CoT) reasoning, but these token-level reasoning chains are computationally expensive and inefficient. In this paper, we introduce Compressed Latent Reasoning (CoLaR), a novel framework that dynamically compresses reasoning processes in latent space through a two-stage training approach. First, during supervised fine-tuning, CoLaR extends beyond next-token prediction by incorporating an auxiliary next compressed embedding prediction objective. This process merges embeddings of consecutive tokens using a compression factor $c$ randomly sampled from a predefined range, and trains a specialized latent head to predict distributions of subsequent compressed embeddings. Second, we enhance CoLaR through reinforcement learning (RL) that leverages the latent head's non-deterministic nature to explore diverse reasoning paths and exploit more compact ones. This approach enables CoLaR to: i) **perform reasoning at a dense latent level** (i.e., silently), substantially reducing reasoning chain length, and ii) **dynamically adjust reasoning speed** at inference time by simply prompting the desired compression factor. Extensive experiments across four mathematical reasoning datasets demonstrate that CoLaR achieves 14.1% higher accuracy than latent-based baseline methods at comparable compression ratios, and reduces reasoning chain length by 53.3% with only 4.8% performance degradation compared to explicit CoT method. Moreover, when applied to more challenging mathematical reasoning tasks, our RL-enhanced CoLaR demonstrates performance gains of up to 5.4% while dramatically reducing latent reasoning chain length by 82.8%. The code and models will be released upon acceptance.
Wenhui Tan, Jianzhong Ju, Zhenbo Luo, Ruihua Song, Jian Luan 0001
NeurIPS6
2025 Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
abstract
Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vision-Language Models (LVLMs) have shown early promise in tackling TVG through supervised fine-tuning (SFT), their ability to generalize remains limited. To address this, we propose a novel post-training framework that enhances the generalization capabilities of LVLMs via reinforcement learning (RL). Specifically, our contributions span three key directions: (1) Time-R1: we introduce a reasoning-guided post-training framework via RL with verifiable reward to enhance capabilities of LVLMs on the TVG task. (2) TimeRFT: we explore post-training strategies on our curated RL-friendly dataset, which trains the model to progressively comprehend more difficult samples, leading to better generalization. (3) TVGBench: we carefully construct a small but comprehensive and balanced benchmark suitable for LVLM evaluation, which is sourced from available public benchmarks. Extensive experiments demonstrate that Time-R1 achieves state-of-the-art performance across multiple downstream datasets using significantly less training data than prior LVLM approaches, while improving its general video understanding capabilities. Project Page: https://xuboshen.github.io/Time-R1/.
Boshen Xu, Yang Du 0011, Kejun Lin, Zihan Xiao 0001, Zihao Yue, Jianzhong Ju, Dingyi Yang, Xiangnan Fang, Zewen He, Zhenbo Luo, Wenxuan Wang 0001, Junqi Lin, Jian Luan 0001, Qin Jin
NeurIPS16
2025 BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
abstract
In the field of AI-driven human-GUI interaction automation, while rapid advances in multimodal large language models and reinforcement fine-tuning techniques have yielded remarkable progress, a fundamental challenge persists: their interaction logic significantly deviates from natural human-GUI communication patterns. To address this gap, we propose Blink–Think–Link (BTL), a brain-inspired framework for human-GUI interaction that mimics the human cognitive process between users and graphical interfaces. The system decomposes interactions into three biologically plausible phases: (1) \textbf{Blink} - rapid detection and attention to relevant screen areas, analogous to saccadic eye movements; (2) \textbf{Think} - higher-level reasoning and decision-making, mirroring cognitive planning; and (3) \textbf{Link} - generation of executable commands for precise motor control, emulating human action selection mechanisms. Additionally, we introduce two key technical innovations for BTL framework: (1) Blink Data Generation - an automated annotation pipeline specifically optimized for blink data, and (2) {BTL Reward – the first rule-based reward mechanism that enables reinforcement learning driven by both process and outcome.} Building upon this framework, we develop a GUI agent model named BTL-UI, which demonstrates competitive performance across both static GUI understanding and dynamic interaction tasks in comprehensive benchmarks. These results provide conclusive empirical validation of the framework's efficacy in developing advanced GUI agents.
Shaojie Zhang 0004, Ruoceng Zhang, Pei Fu, Shiqi Cui, Zhenbo Luo, Jian Luan 0001
NeurIPS11
2025 DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models
Xiaolin Hu 0001, Xiang Cheng 0007, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Yong Liu 0020
PAKDD (5)5
2024 Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
abstract
Shihan Deng, Weikai Xu, Hongda Sun, Wei Liu, Tao Tan, Jianfeng Liu, Ang Li, Jian Luan, Bin Wang, Rui Yan, Shuo Shang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shihan Deng, Weikai Xu, Hongda Sun 0001, Wei Liu 0302, Tao Tan 0005, Jianfeng Liu 0005, Jian Luan 0001, Bin Wang 0004, Rui Yan 0001, Shuo Shang
ACL (1)8
2024 DetermLR: Augmenting LLM-based Logical Reasoning from Indeterminacy to Determinacy
abstract
Hongda Sun, Weikai Xu, Wei Liu, Jian Luan, Bin Wang, Shuo Shang, Ji-Rong Wen, Rui Yan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Hongda Sun 0001, Weikai Xu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004, Shuo Shang, Ji-Rong Wen, Rui Yan 0001
ACL (1)4
2024 ToolRerank: Adaptive and Hierarchy-Aware Reranking for Tool Retrieval
abstract
Tool learning aims to extend the capabilities of large language models (LLMs) with external tools. A major challenge in tool learning is how to support a large number of tools, including unseen tools. To address this challenge, previous studies have proposed retrieving suitable tools for the LLM based on the user query. However, previously proposed methods do not consider the differences between seen and unseen tools, nor do they take the hierarchy of the tool library into account, which may lead to suboptimal performance for tool retrieval. Therefore, to address the aforementioned issues, we propose ToolRerank, an adaptive and hierarchy-aware reranking method for tool retrieval to further refine the retrieval results. Specifically, our proposed ToolRerank includes Adaptive Truncation, which truncates the retrieval results related to seen and unseen tools at different positions, and Hierarchy-Aware Reranking, which makes retrieval results more concentrated for single-tool queries and more diverse for multi-tool queries. Experimental results show that ToolRerank can improve the quality of the retrieval results, leading to better execution results generated by the LLM.
Yuanhang Zheng, Peng Li 0021, Wei Liu 0302, Yang Liu 0005, Jian Luan 0001, Bin Wang 0004
LREC/COLING5
2024 SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM
abstract
While Large Language Models (LLMs) have achieved remarkable success in various fields, the efficiency of training and inference remains a major challenge. To address this issue, we propose SUBLLM, short for Subsampling-Upsampling-Bypass Large Language Model, an innovative architecture that extends the core decoder-only framework by incorporating subsampling, upsampling, and bypass modules. The subsampling modules are responsible for shortening the sequence, while the upsampling modules restore the sequence length, and the bypass modules enhance convergence. In comparison to LLaMA, the proposed SUBLLM exhibits significant enhancements in both training and inference speeds as well as memory usage, while maintaining competitive few-shot performance. During training, SUBLLM increases speeds by 26% and cuts memory by 10GB per GPU. In inference, it boosts speeds by up to 37% and reduces memory by 1GB per GPU. The training and inference speeds can be enhanced by 34% and 52% respectively when the context window is expanded to 8192. Our code is available at https://github.com/XiaoMi/subllm.
Quandong Wang, Xiaoyu Yang 0005, Ruike Zhang, Wei Liu 0302, Jian Luan 0001, Daniel Povey, Bin Wang 0004
ECAI7
2024 ToolPlanner: A Tool Augmented LLM for Multi Granularity Instructions with Path Planning and Feedback
abstract
Recently, tool-augmented LLMs have gained increasing attention.Given an instruction, toolaugmented LLMs can interact with various external tools in multiple rounds and provide a final answer.However, previous LLMs were trained on overly detailed instructions, which included API names or parameters, while real users would not explicitly mention these API details.This leads to a gap between trained LLMs and real-world scenarios.In addition, most works ignore whether the interaction process follows the instruction.To address these issues, we constructed a training dataset called MGToolBench, which contains statement and category-level instructions to better reflect realworld scenarios.In addition, we propose Tool-Planner, a two-stage reinforcement learning framework that utilizes path planning and two feedback mechanisms to enhance the LLM's task completion and instruction-following capabilities.Experimental results show that Tool-Planner significantly improves the Match Rate, Pass Rate and Win Rate by 26.8%, 20.2%, and 5.6% compared to the SOTA model.Human evaluation verifies that the multi-granularity instructions can better align with users' usage habits.Our data and code are available at https://github.com/XiaoMi/toolplanner.
Qinzhuo Wu, Wei Liu 0302, Jian Luan 0001, Bin Wang 0004
EMNLP3
2023 BERT-ERC: Fine-Tuning BERT Is Enough for Emotion Recognition in Conversation
abstract
Previous works on emotion recognition in conversation (ERC) follow a two-step paradigm, which can be summarized as first producing context-independent features via fine-tuning pretrained language models (PLMs) and then analyzing contextual information and dialogue structure information among the extracted features. However, we discover that this paradigm has several limitations. Accordingly, we propose a novel paradigm, i.e., exploring contextual information and dialogue structure information in the fine-tuning step, and adapting the PLM to the ERC task in terms of input text, classification structure, and training strategy. Furthermore, we develop our model BERT-ERC according to the proposed paradigm, which improves ERC performance in three aspects, namely suggestive text, fine-grained classification module, and two-stage training. Compared to existing methods, BERT-ERC achieves substantial improvement on four datasets, indicating its effectiveness and generalization capability. Besides, we also set up the limited resources scenario and the online prediction scenario to approximate real-world scenarios. Extensive experiments demonstrate that the proposed paradigm significantly outperforms the previous one and can be adapted to various scenes.
Xiangyu Qin, Zhiyu Wu, Yanran Li, Jian Luan 0001, Bin Wang 0004, Li Wang 0114, Jinshi Cui
AAAI5
2023 Exploring Better Text Image Translation with Multimodal Codebook
abstract
Zhibin Lan, Jiawei Yu, Xiang Li, Wen Zhang, Jian Luan, Bin Wang, Degen Huang, Jinsong Su. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Zhibin Lan, Xiang Li 0104, Wen Zhang 0015, Jian Luan 0001, Bin Wang 0004, Degen Huang, Jinsong Su
ACL (1)5
2023 Exploring All-In-One Knowledge Distillation Framework for Neural Machine Translation
abstract
Conventional knowledge distillation (KD) approaches are commonly employed to compress neural machine translation (NMT) models.However, they only obtain one lightweight student each time.Consequently, we have to conduct KD multiple times when different students are required at the same time, which could be resource-intensive.Additionally, these students are individually optimized, and thus lack interactions with each other, leading to their potential not being fully exerted.In this work, we propose a novel All-In-One Knowledge Distillation (AIO-KD) framework for NMT, which generates multiple satisfactory students at once.Under AIO-KD, we first randomly extract fewer-layer subnetworks from the teacher as the sample students.Then, we jointly optimize the teacher and these students, where the students simultaneously learn the knowledge from the teacher and interact with other students via mutual learning.When utilized, we re-extract the candidate students, satisfying the specifications of various devices.Particularly, we adopt carefully-designed strategies for AIO-KD: 1) we dynamically detach gradients to prevent poorly-performed students from negatively affecting the teacher during the knowledge transfer, which could subsequently impact other students; 2) we design a twostage mutual learning strategy, which alleviates the negative impacts of poorly-performed students on the early-stage student interactions.Extensive experiments and in-depth analyses on three benchmarks demonstrate the effectiveness and eco-friendliness of AIO-KD.Our source code is available at https://github. com/DeepLearnXMU/AIO-KD.
Zhongjian Miao, Wen Zhang 0015, Jinsong Su, Xiang Li 0104, Jian Luan 0001, Yidong Chen 0001, Bin Wang 0004, Min Zhang 0005
EMNLP5
2023 Joint Training and Decoding for Multilingual End-to-End Simultaneous Speech Translation
abstract
Recent studies on end-to-end speech translation(ST) have facilitated the exploration of multilingual end-to-end ST and end-to-end simultaneous ST. In this paper, we investigate end-to-end simultaneous speech translation in a one-to-many multilingual setting which is closer to applications in real scenarios. We explore a separate decoder architecture and a unified architecture for joint synchronous training in this scenario. To further explore knowledge transfer across languages, we propose an asynchronous training strategy on the proposed unified decoder architecture. A multi-way aligned multilingual end-to-end ST dataset was curated as a benchmark testbed to evaluate our methods. Experimental results demonstrate the effectiveness of our models on the collected dataset. Our codes and data are available at: https://github.com/XiaoMi/TED-MMST.
Wuwei Huang, Renren Jin, Wen Zhang 0015, Jian Luan 0001, Bin Wang 0004, Deyi Xiong
ICASSP4
2023 Rethinking the Reasonability of the Test Set for Simultaneous Machine Translation
abstract
Simultaneous machine translation (SimulMT) models start translation before the end of the source sentence, making the translation monotonically aligned with the source sentence. However, the general full-sentence translation test set is acquired by offline translation of the entire source sentence, which is not designed for SimulMT evaluation, making us rethink whether this will underestimate the performance of SimulMT models. In this paper, we manually annotate a monotonic test set based on the MuST-C English-Chinese test set, denoted as SiMuST-C. Our human evaluation confirms the acceptability of our annotated test set. Evaluations on three different SimulMT models verify that the underestimation problem can be alleviated on our test set. Further experiments show that finetuning on an automatically extracted monotonic training set improves SimulMT models by up to 3 BLEU points.
Mengge Liu, Wen Zhang 0015, Xiang Li 0104, Jian Luan 0001, Bin Wang 0004, Yuhang Guo 0001, Shuoying Chen
ICASSP4
2023 LightClone: Speaker-guided Parallel Subnet Selection for Few-shot Voice Cloning
Jie Wu 0017, Jian Luan 0001
INTERSPEECH2
2023 Improving Bilingual TTS Using Language And Phonology Embedding With Embedding Strength Modulator
Fengyu Yang 0002, Jian Luan 0001
INTERSPEECH2
2023 Overview of the NLPCC 2023 Shared Task 9: User Feedback Prediction and Response Generation
Hanlin Teng, Hongda Sun 0001, Wei Liu 0302, Shuang Dong, Rui Yan 0001, Jian Luan 0001, Bin Wang 0004
NLPCC (3)6
2022 PAMA-TTS: Progression-Aware Monotonic Attention for Stable SEQ2SEQ TTS with Accurate Phoneme Duration Control
abstract
Sequence expansion between encoder and decoder is a critical challenge in sequence-to-sequence TTS. Attention-based methods achieve great naturalness but suffer from unstable issues like missing and repeating phonemes, not to mention accurate duration control. Duration-informed methods, on the contrary, seem to easily adjust phoneme duration but show obvious degradation in speech naturalness. This paper proposes PAMA-TTS to address the problem. It takes the advantage of both flexible attention and explicit duration models. Based on the monotonic attention mechanism, PAMA-TTS also leverages token duration and relative position of a frame, especially countdown information, i.e. in how many future frames the present phoneme will end. They help the attention to move forward along the token sequence in a soft but reliable control. Experimental results prove that PAMA-TTS achieves the highest naturalness, while has on-par or even better duration controllability than the duration-informed model.
Yunchao He, Jian Luan 0001
ICASSP2
2022 MSDTRON: A High-Capability Multi-Speaker Speech Synthesis System for Diverse Data Using Characteristic Information
abstract
In multi-speaker speech synthesis, data from a number of speakers usually tend to have great diversity due to the fact that the speakers may differ largely in ages, speaking styles, emotions, and so on. It is important but challenging to improve the modeling capabilities for multi-speaker speech synthesis. To address the issue, this paper proposes a high-capability speech synthesis system, called Msdtron, in which 1) a representation of the harmonic structure of speech, called excitation spectrogram, is designed to directly guide the learning of harmonics in mel-spectrogram. 2) conditional gated LSTM (CGLSTM) is proposed to control the flow of text content information through the network by re-weighting the gates of LSTM using speaker information. The experiments show a significant reduction in reconstruction error of mel-spectrogram in the training of the multi-speaker model, and a great improvement is observed in the subjective evaluation of speaker adapted model.
Quanbo Shen, Jian Luan 0001
ICASSP3
2022 Improving Emotional Speech Synthesis by Using SUS-Constrained VAE and Text Encoder Aggregation
abstract
Learning emotion embedding from reference audio is a straightforward approach for multi-emotion speech synthesis in encoder-decoder systems. But how to get better emotion embedding and how to inject it into TTS acoustic model more effectively are still under investigation. In this paper, we propose an innovative constraint to help VAE extract emotion embedding with better cluster cohesion. Besides, the obtained emotion embedding is used as query to aggregate latent representations of all encoder layers via attention. Moreover, the queries from encoder layers themselves are also helpful. Experiments prove the proposed methods can enhance the encoding of comprehensive syntactic and semantic information and produce more expressive emotional speech.
Fengyu Yang 0002, Jian Luan 0001
ICASSP2
2020 Transfer Learning for Improving Singing-Voice Detection in Polyphonic Instrumental Music
abstract
Detecting singing-voice in polyphonic instrumental music is critical to music information retrieval.To train a robust vocal detector, a large dataset marked with vocal or non-vocal label at frame-level is essential.However, frame-level labeling is time-consuming and labor expensive, resulting there is little well-labeled dataset available for singing-voice detection (S-VD).Hence, we propose a data augmentation method for S-VD by transfer learning.In this study, clean speech clips with voice activity endpoints and separate instrumental music clips are artificially added together to simulate polyphonic vocals to train a vocal /non-vocal detector.Due to the different articulation and phonation between speaking and singing, the vocal detector trained with the artificial dataset does not match well with the polyphonic music which is singing vocals together with the instrumental accompaniments.To reduce this mismatch, transfer learning is used to transfer the knowledge learned from the artificial speech-plus-music training set to a small but matched polyphonic dataset, i.e., singing vocals with accompaniments.By transferring the related knowledge to make up for the lack of well-labeled training data in S-VD, the proposed data augmentation method by transfer learning can improve S-VD performance with an F-score improvement from 89.5% to 93.2%.
Yuanbo Hou, Frank K. Soong, Jian Luan 0001, Shengchen Li
INTERSPEECH3
2020 XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System
abstract
This paper presents XiaoiceSing, a high-quality singing voice synthesis system which employs an integrated network for spectrum, F0 and duration modeling.We follow the main architecture of FastSpeech while proposing some singing-specific design: 1) Besides phoneme ID and position encoding, features from musical score (e.g.note pitch and length) are also added.2) To attenuate off-key issues, we add a residual connection in F0 prediction.3) In addition to the duration loss of each phoneme, the duration of all the phonemes in a musical note is accumulated to calculate the syllable duration loss for rhythm enhancement.Experiment results show that XiaoiceSing outperforms the baseline system of convolutional neural networks by 1.44 MOS on sound quality, 1.18 on pronunciation accuracy and 1.38 on naturalness respectively.In two A/B tests, the proposed F0 and duration modeling methods achieve 97.3% and 84.3% preference rate over baseline respectively, which demonstrates the overwhelming advantages of XiaoiceSing.
Peiling Lu, Jie Wu 0017, Jian Luan 0001, Xu Tan 0003
INTERSPEECH3
2020 Adversarially Trained Multi-Singer Sequence-to-Sequence Singing Synthesizer
abstract
This paper presents a high quality singing synthesizer that is able to model a voice with limited available recordings.Based on the sequence-to-sequence singing model, we design a multisinger framework to leverage all the existing singing data of different singers.To attenuate the issue of musical score unbalance among singers, we incorporate an adversarial task of singer classification to make encoder output less singer dependent.Furthermore, we apply multiple random window discriminators (MRWDs) on the generated acoustic features to make the network be a GAN.Both objective and subjective evaluations indicate that the proposed synthesizer can generate higher quality singing voice than baseline (4.12 vs 3.53 in MOS).Especially, the articulation of high-pitched vowels is significantly enhanced.
Jie Wu 0017, Jian Luan 0001
INTERSPEECH2
2020 Re-Weighted Interval Loss for Handling Data Imbalance Problem of End-to-End Keyword Spotting
Zhiyong Wu 0001, Daode Yuan, Jian Luan 0001, Jia Jia 0001, Helen M. Meng, Binheng Song
INTERSPEECH4
2020 DeepSinger: Singing Voice Synthesis with Data Mined From the Web
abstract
In this paper, we develop DeepSinger, a multi-lingual multi-singer singing voice synthesis (SVS) system, which is built from scratch using singing training data mined from music websites. The pipeline of DeepSinger consists of several steps, including data crawling, singing and accompaniment separation, lyrics-to-singing alignment, data filtration, and singing modeling. Specifically, we design a lyrics-to-singing alignment model to automatically extract the duration of each phoneme in lyrics starting from coarse-grained sentence level to fine-grained phoneme level, and further design a multi-lingual multi-singer singing model based on a feed-forward Transformer to directly generate linear-spectrograms from lyrics, and synthesize voices using Griffn-Lim. DeepSinger has several advantages over previous SVS systems: 1) to the best of our knowledge, it is the first SVS system that directly mines training data from music websites, 2) the lyrics-to-singing alignment model further avoids any human efforts for alignment labeling and greatly reduces labeling cost, 3) the singing model based on a feed-forward Transformer is simple and efficient, by removing the complicated acoustic feature modeling in parametric synthesis and leveraging a reference encoder to capture the timbre of a singer from noisy singing data, and 4) it can synthesize singing voices in multiple languages and multiple singers. We evaluate DeepSinger on our mined singing dataset that consists of about 92 hours data from 89 singers on three languages (Chinese, Cantonese and English). The results demonstrate that with the singing data purely mined from the Web, DeepSinger can synthesize high-quality singing voices in terms of both pitch accuracy and voice naturalness. Our audio samples are shown in https://speechresearch.github.io/deepsinger/.
Yi Ren 0006, Xu Tan 0003, Tao Qin 0001, Jian Luan 0001, Zhou Zhao 0001, Tie-Yan Liu
KDD4
2019 Vocal Pitch Extraction in Polyphonic Music Using Convolutional Residual Network
Mingye Dong, Jie Wu 0017, Jian Luan 0001
INTERSPEECH3
2012 Expand CRF to Model Long Distance Dependencies in Prosodic Break Prediction
Jian Luan 0001
INTERSPEECH1
2010 Improvement on plural unit selection and fusion
Jian Luan 0001
INTERSPEECH1