EDBT 2026 Demo / reviewers in the wild / expert
Ang Lv
dblp:326/5506
· DBLP profile ↗
19ranked-venue papers
6as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Union-of-Experts: Neurons in Mixture-of-Experts are Secretly RoutersabstractMixture-of-Experts (MoE) models rely on an external router to assign tokens to experts.This design inherently separates the routing decision from each expert's internal capabilities, leading to suboptimal performance.In this work, we address this limitation with Union-of-Experts (UoE), an MoE variant that performs "expertautonomous routing".The core mechanism of UoE is to pre-designate a minute fraction of neurons within each expert as routing neurons.Experts autonomously select relevant tokens by comparing the activation intensity of these neurons, aligning routing decisions with each expert's functional profile.To prevent the waste of activations from unselected experts, we aggregate all routing neuron outputs and sum them into the final layer output.This aggregation acts as a novel virtual shared expert whose parameters are distributed across the individual experts, and improves overall parameter efficiency.We pre-train UoE models with up to 3B parameters, demonstrating that they outperform traditional MoEs with matched efficiency.Furthermore, our analysis of the routing neurons provides valuable insights into expert-autonomous selection and advances the understanding of MoE routing. Songhao Wu, Ang Lv, Ruobing Xie, Samm Sun, Di Wang 0052, Rui Yan 0001, Yankai Lin 0001 |
ACL (1) | 2 |
| 2026 | Data Pollination: An Emergent Ecological Process Driving AI Population EvolutionabstractAI development is often framed as the outcome of isolated research and engineering efforts, yet evidence from deployed systems suggests that language models interact through a shared data ecosystem.While the optimization of individual models is extensively studied, the emergent properties of this interconnected population remain largely unexplored, limiting our ability to predict long-term ecosystem trajectories.We term this process data pollination, the unintentional circulation of synthetic model outputs through shared online platforms and web-scale training corpora, and formalize it as a population-based evolutionary framework to investigate stability dynamics under synthetic data training.Our theoretical analysis and controlled experiments involving 320 language models demonstrate that population dynamics can mitigate the model collapse observed in single-lineage recursive training, yielding stable or improving performance across diverse benchmarks.Crucially, we find that ecological diversity functions as a fundamental resilience mechanism that safeguards the ecosystem against collapse, highlighting the critical importance of maintaining model diversity for sustainable AI development. Shufang Xie 0003, Qizhi Pei, Ang Lv, Jingyang Hu, Lijun Wu 0003, Rui Yan 0001 |
ACL (1) | 3 |
| 2026 | StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to ReasonabstractReinforcement learning with verifiable rewards (RLVR) is a promising approach for improving the complex reasoning abilities of large language models (LLMs).However, current RLVR methods face two significant challenges: the near-miss reward problem, where a small mistake can invalidate an otherwise correct reasoning process, greatly hindering training efficiency; and exploration stagnation, where models tend to focus on solutions within their "comfort zone," lacking the motivation to explore potentially more effective alternatives.To address these challenges, we propose StepHint, a novel RLVR algorithm that utilizes multi-level stepwise hints to help models explore the solution space more effectively.StepHint partitions valid reasoning chains into reasoning steps using our proposed adaptive partitioning method.The initial few steps are used as hints, and simultaneously, multiple-level hints (each comprising a different number of steps) are provided to the model.This approach directs the model's exploration toward a promising solution subspace while preserving its flexibility for independent exploration.By providing hints, StepHint mitigates the near-miss reward problem, thereby improving training efficiency.Additionally, the external reasoning pathways help the model develop better reasoning abilities, enabling it to move beyond its "comfort zone" and mitigate exploration stagnation.StepHint outperforms competitive RLVR enhancement methods across six mathematical benchmarks and two out-of-domain benchmarks. 1 Ang Lv, Jinpeng Li 0003, Feng Wang 0023, Haoyuan Hu, Rui Yan 0001 |
ACL (1) | 2 |
| 2026 | A multi-feature fusion network for transmission line channel scene classification based on superpixel information extraction and fine-grained information selection
Tangfei Tao, Ang Lv, Maohui Tang, Lanjun Xu |
Expert Syst. Appl. | 3 |
| 2025 | HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and ExtrapolationabstractMany positional encodings (PEs) are designed to exhibit long-term decay, based on an entrenched and long-standing inductive opinion: tokens farther away from the current position carry less relevant information. We argue that long-term decay is outdated in the era of LLMs, as LLMs are now applied to tasks demanding precise retrieval of in-context information from arbitrary positions. Firstly, we present empirical analyses on various PEs, demonstrating that models inherently learn attention with only a local-decay pattern while forming a U-shape pattern globally, contradicting the principle of long-term decay. Furthermore, we conduct a detailed analysis of rotary position encoding (RoPE, a prevalent relative positional encoding in LLMs), and found that the U-shape attention is caused by some learned components, which are also the key factor limiting RoPE’s expressiveness and extrapolation. Inspired by these insights, we propose High-frequency rotary Position Encoding (HoPE). HoPE replaces the specific components in RoPE with position-independent ones, retaining only high-frequency signals, which also breaks the principle of long-term decay in theory. HoPE achieves two major advantages: (1) Without constraints imposed by long-term decay, contradictory factors that limit attention optimization are removed. Thus, the model’s context awareness is enhanced. (2) HoPE exhibits greater robustness to the out-of-distribution behavior in attention patterns during extrapolation. The effectiveness of HoPE is validated through extensive experiments and with a large language model of up to 3 billion parameters. Yuhan Chen 0001, Ang Lv, Jian Luan 0001, Bin Wang 0004, Wei Liu 0302 |
ACL (1) | 2 |
| 2025 | More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting ObjectivesabstractXiaoqing Zhang, Ang Lv, Yuhan Liu, Flood Sung, Wei Liu, Jian Luan, Shuo Shang, Xiuying Chen, Rui Yan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiaoqing Zhang 0017, Ang Lv, Yuhan Liu 0023, Flood Sung, Wei Liu 0302, Jian Luan 0001, Shuo Shang, Xiuying Chen, Rui Yan 0001 |
ACL (1) | 2 |
| 2025 | Autonomy-of-Experts ModelsabstractMixture-of-Experts (MoE) models mostly use a router to assign tokens to specific expert modules, activating only partial parameters and often outperforming dense models. We argue that the separation between the router’s decision-making and the experts’ execution is a critical yet overlooked issue, leading to suboptimal expert selection and learning. To address this, we propose Autonomy-of-Expert (AoE), a novel MoE paradigm in which experts autonomously select themselves to process inputs. AoE is based on the insight that an expert is aware of its own capacity to effectively process a token, an awareness reflected in the scale of its internal activations. In AoE, routers are removed; instead, experts pre-compute internal activations for inputs and are ranked based on their activation norms. Only the top-ranking experts proceed with the forward pass, while the others abort. The overhead of pre-computing activations is reduced through a low-rank weight factorization. This self-evaluating-then-partner-comparing approach ensures improved expert selection and effective learning. We pre-train language models having 700M up to 4B parameters, demonstrating that AoE outperforms traditional MoE models with comparable efficiency. Ang Lv, Ruobing Xie, Songhao Wu, Xingwu Sun, Zhanhui Kang, Di Wang 0052, Rui Yan 0001 |
ICML | 1 |
| 2025 | GETMusic: Generating Music Tracks with a Unified Representation and Diffusion FrameworkabstractSymbolic music generation aims to create musical notes, which can help users compose music, such as generating target instrument tracks based on provided source tracks. In practical scenarios where there’s a predefined ensemble of tracks and various composition needs, an efficient and effective generative model that can generate any target tracks based on the other tracks becomes crucial. However, previous efforts have fallen short in addressing this necessity due to limitations in their music representations and models. In this paper, we introduce a framework known as GETMusic, with ``GET'' standing for ``GEnerate music Tracks.'' This framework encompasses a novel music representation ``GETScore'' and a diffusion model ``GETDiff.'' GETScore represents musical notes as tokens and organizes tokens in a 2D structure, with tracks stacked vertically and progressing horizontally over time. At a training step, each track of a music piece is randomly selected as either the target or source. The training involves two processes: In the forward process, target tracks are corrupted by masking their tokens, while source tracks remain as the ground truth; in the denoising process, GETDiff is trained to predict the masked target tokens conditioning on the source tracks. Our proposed representation, coupled with the non-autoregressive generative model, empowers GETMusic to generate music with any arbitrary source-target track combinations.Our experiments demonstrate that the versatile GETMusic outperforms prior works proposed for certain specific composition tasks. Ang Lv, Xu Tan 0003, Peiling Lu, Wei Ye 0004, Shikun Zhang, Jiang Bian 0002, Rui Yan 0001 |
IJCAI | 1 |
| 2025 | PolarQuant: Leveraging Polar Transformation for Key Cache Quantization and Decoding AccelerationabstractThe increasing demand for long-context generation has made the KV cache in large language models a bottleneck in memory consumption. Quantizing the cache to lower bit widths is an effective way to reduce memory costs; however, previous methods struggle with key cache quantization due to outliers, resulting in suboptimal performance. We propose a novel quantization approach PolarQuant, which provides a new perspective for key cache quantization and efficiently addresses the outlier dilemma. We observe that the distribution of the key states reveals well-structured patterns under polar transformation. Outliers generally appear in only one of the two dimensions, which are rotated together by a specific angle when rotary position embeddings are applied. When represented as two-dimensional vectors, these dimensions exhibit well-organized patterns, with radii and angles smoothly distributed in polar space. This alleviates the channel-wise outliers, making them well-suited for key cache quantization. PolarQuant divides key vectors into groups of two-dimensional sub-vectors, encoding them as the quantized radius and the polar angle, rather than quantizing original key vectors directly. PolarQuant achieves the superior efficiency in KV cache quantization and accelerates the decoding process by turning the query-key inner product into a table lookup, all while maintaining the downstream performance of full-precision models.
Our code is available at https://github.com/ericshwu/PolarQuant. Songhao Wu, Ang Lv, Guojun Yin, Rui Yan 0001 |
NeurIPS | 2 |
| 2025 | PEAR: Position-Embedding-Agnostic Attention Re-weighting Enhances Retrieval-Augmented Generation with Zero Inference OverheadabstractLarge language models (LLMs) enhanced with retrieval-augmented generation (RAG) have introduced a new paradigm for web search. However, the limited context awareness of LLMs degrades their performance on RAG tasks. Existing methods to enhance context awareness are often inefficient, incurring time or memory overhead during inference, and many are tailored to specific position embeddings. In this paper, we propose Position-Embedding-Agnostic attention Re-weighting (PEAR), which enhances the context awareness of LLMs with zero inference overhead. Specifically, on a proxy task focused on context copying, we first detect heads which suppress the models' context awareness, thereby diminishing RAG performance. To weaken the impact of these heads, we re-weight their outputs with learnable coefficients. The LLM (with frozen parameters) is optimized by adjusting these coefficients to minimize loss on the proxy task. During inference, the optimized coefficients are fixed to re-weight these heads, regardless of the specific task at hand. Our proposed PEAR offers two major advantages over previous approaches: (1) It introduces zero additional inference overhead in terms of memory usage or inference time, while outperforming competitive baselines in accuracy and efficiency across various RAG tasks. (2) It is independent of position embedding algorithms, ensuring broader applicability. Our code is available at https://github.com/TTArch/PEAR-RAG. Tao Tan 0005, Ang Lv, Hongzhan Lin 0002, Songhao Wu, Feng Wang 0023, Jingtong Wu, Rui Yan 0001 |
WWW | 3 |
| 2024 | Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool UseabstractYuhan Chen, Ang Lv, Ting-En Lin, Changyu Chen, Yuchuan Wu, Fei Huang, Yongbin Li, Rui Yan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yuhan Chen 0001, Ang Lv, Ting-En Lin, Changyu Chen, Yuchuan Wu, Fei Huang 0002, Yongbin Li 0001, Rui Yan 0001 |
ACL (1) | 2 |
| 2024 | Masked Thought: Simply Masking Partial Reasoning Steps Can Improve Mathematical Reasoning Learning of Language ModelsabstractChangyu Chen, Xiting Wang, Ting-En Lin, Ang Lv, Yuchuan Wu, Xin Gao, Ji-Rong Wen, Rui Yan, Yongbin Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Changyu Chen, Xiting Wang, Ting-En Lin, Ang Lv, Yuchuan Wu, Xin Gao 0001, Ji-Rong Wen, Rui Yan 0001, Yongbin Li 0001 |
ACL (1) | 4 |
| 2024 | Mixture-of-Modules: Reinventing Transformers as Dynamic Assemblies of ModulesabstractIs it always necessary to compute tokens from shallow to deep layers in Transformers?The continued success of vanilla Transformers and their variants suggests an undoubted "yes".In this work, however, we attempt to break the depth-ordered convention by proposing a novel architecture dubbed mixture-of-modules (MoM), which is motivated by an intuition that any layer, regardless of its position, can be used to compute a token as long as it possesses the needed processing capabilities.The construction of MoM starts from a finite set of modules defined by multi-head attention and feed-forward networks, each distinguished by its unique parameterization.Two routers then iteratively select attention modules and feedforward modules from the set to process a token.The selection dynamically expands the computation graph in the forward pass of the token, culminating in an assembly of modules.We show that MoM provides not only a unified framework for Transformers and their numerous variants but also a flexible and learnable approach for reducing redundancy in Transformer parameterization.We pre-train various MoMs using OpenWebText.Empirical results demonstrate that MoMs, of different parameter counts, consistently outperform vanilla transformers on both GLUE and XSUM benchmarks.More interestingly, with a fixed parameter budget, MoM-large enables an over 38% increase in depth for computation graphs compared to GPT-2-large, resulting in absolute gains of 1.4 on GLUE and 1 on XSUM.On the other hand, MoM-large also enables an over 60% reduction in depth while involving more modules per layer, yielding a 16% reduction in TFLOPs and a 43% decrease in memory usage compared to GPT-2-large, while maintaining comparable performance.1 * Equal Contributions.† Corresponding authors. 1 Code is available at https://github.com/gzhch/Mixture-of- Modules Zhuocheng Gong, Ang Lv, Jian Guan 0002, Wei Wu 0014, Huishuai Zhang, Minlie Huang, Dongyan Zhao 0001, Rui Yan 0001 |
EMNLP | 2 |
| 2024 | An Analysis and Mitigation of the Reversal CurseabstractRecent research observed a noteworthy phenomenon in large language models (LLMs), referred to as the "reversal curse."The reversal curse is that when dealing with two entities, denoted as a and b, connected by their relation R and its inverse R -1 , LLMs excel in handling sequences in the form of "aRb," but encounter challenges when processing "bR -1 a," whether in generation or comprehension.For instance, GPT-4 can accurately respond to the query "Tom Cruise's mother is?" with "Mary Lee Pfeiffer," but it struggles to provide a satisfactory answer when asked "Mary Lee Pfeiffer's son is?"In this paper, we undertake the first-ever study of how the reversal curse happens in LLMs.Our investigations reveal that the reversal curse can stem from the specific training objectives, which become particularly evident in the widespread use of next-token prediction within most causal language models.We hope this initial investigation can draw more attention to the reversal curse, as well as other underlying limitations in current LLMs. 1 Ang Lv, Shufang Xie 0003, Quan Tu, Yuhan Chen 0001, Ji-Rong Wen, Rui Yan 0001 |
EMNLP | 1 |
| 2024 | Re-creation of Creations: A New Paradigm for Lyric-to-Melody Generation
Ang Lv, Xu Tan 0003, Tao Qin 0001, Tie-Yan Liu, Rui Yan 0001 |
IJCAI | 1 |
| 2024 | Mixture of In-Context Experts Enhance LLMs' Long Context AwarenessabstractMany studies have revealed that large language models (LLMs) exhibit uneven awareness of different contextual positions. Their limited context awareness can lead to overlooking critical information and subsequent task failures. While several approaches have been proposed to enhance LLMs' context awareness, achieving both effectiveness and efficiency remains challenging. In this paper, for LLMs utilizing RoPE as position embeddings, we introduce a novel method called "Mixture of In-Context Experts" (MoICE) to address this challenge. MoICE comprises two key components: a router integrated into each attention head within LLMs and a lightweight router-only training optimization strategy:(1) MoICE views each RoPE angle as an 'in-context' expert, demonstrated to be capable of directing the attention of a head to specific contextual positions. Consequently, each attention head flexibly processes tokens using multiple RoPE angles dynamically selected by the router to attend to the needed positions. This approach mitigates the risk of overlooking essential contextual information. (2) The router-only training strategy entails freezing LLM parameters and exclusively updating routers for only a few steps. When applied to open-source LLMs including Llama and Mistral, MoICE surpasses prior methods across multiple tasks on long context understanding and generation, all while maintaining commendable inference efficiency. Hongzhan Lin 0002, Ang Lv, Yuhan Chen 0001, Chen Zhu 0003, Yang Song 0021, Hengshu Zhu, Rui Yan 0001 |
NeurIPS | 2 |
| 2023 | Envisioning Future from the Past: Hierarchical Duality Learning for Multi-Turn Dialogue GenerationabstractIn this paper, we define a widely neglected property in dialogue text, duality, which is a hierarchical property that is reflected in human behaviours in daily conversations: Based on the logic in a conversation (or a sentence), people can infer follow-up utterances (or tokens) based on the previous text, and vice versa.We propose a hierarchical duality learning for dialogue (HDLD) to simulate this human cognitive ability, for generating high quality responses that connect both previous and follow-up dialogues.HDLD utilizes hierarchical dualities at token hierarchy and utterance hierarchy.HDLD maximizes the mutual information between past and future utterances.Thus, even if the future text is invisible during inference, HDLD is capable of estimating future information implicitly based on dialogue history and generates both coherent and informative responses.In contrast to previous approaches that solely utilize future text as auxiliary information to encode during training, HDLD leverages duality to enable interaction between dialogue history and the future.This enhances the utilization of dialogue data, leading to the improvement in both automatic and human evaluation. Ang Lv, Jinpeng Li 0003, Shufang Xie 0003, Rui Yan 0001 |
ACL (1) | 1 |
| 2023 | DialoGPS: Dialogue Path Sampling in Continuous Semantic Space for Data Augmentation in Multi-Turn ConversationsabstractIn open-domain dialogue generation tasks, contexts and responses in most datasets are oneto-one mapped, violating an important manyto-many characteristic: a context leads to various responses, and a response answers multiple contexts.Without such patterns, models poorly generalize and prefer responding safely.Many attempts have been made in either multiturn settings from a one-to-many perspective or in a many-to-many perspective but limited to single-turn settings.The major challenge to many-to-many augment multi-turn dialogues is that discretely replacing each turn with semantic similarity breaks fragile context coherence.In this paper, we propose DialoGue Path Sampling (DialoGPS) method in continuous semantic space, the first many-to-many augmentation method for multi-turn dialogues.Specifically, we map a dialogue to our extended Brownian Bridge, a special Gaussian process.We sample latent variables to form coherent dialogue paths in the continuous space.A dialogue path corresponds to a new multi-turn dialogue and is used as augmented training data.We show the effect of DialoGPS with both automatic and human evaluation. Ang Lv, Jinpeng Li 0003, Yuhan Chen 0001, Ji Zhang 0011, Rui Yan 0001 |
ACL (1) | 1 |
| 2022 | Target-Side Input Augmentation for Sequence to Sequence Generation
Shufang Xie 0003, Ang Lv, Yingce Xia, Lijun Wu 0003, Tao Qin 0001, Tie-Yan Liu, Rui Yan 0001 |
ICLR | 2 |