Yuhan Chen 0001

dblp:155/2863-1 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0009-0001-8752-9411ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 46% Language models and text generation · 29% Multi-agent systems · 14%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 72% Computational finance and economics · 28%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems
context awareness
1.622025
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation · ACL (1) 2025
Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use · ACL (1) 2024
Machine learning › Deep learning architectures and training
positional encoding
1.622025
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation · ACL (1) 2025
Mixture of In-Context Experts Enhance LLMs' Long Context Awareness · NeurIPS 2024
Machine learning › Deep learning architectures and training
attention mechanism
1.522024
Mixture of In-Context Experts Enhance LLMs' Long Context Awareness · NeurIPS 2024
Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use · ACL (1) 2024
Natural language and speech › Language models and text generation › language modeling
long-context language modeling
0.912025
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation · ACL (1) 2025
Machine learning › Deep learning architectures and training › positional encoding
rotary position embedding
0.912025
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation · ACL (1) 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context understanding
0.812024
Mixture of In-Context Experts Enhance LLMs' Long Context Awareness · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model reasoning › reasoning robustness
reversal curse
0.812024
An Analysis and Mitigation of the Reversal Curse · EMNLP 2024
Machine learning › Deep learning architectures and training
training objective
0.812024
An Analysis and Mitigation of the Reversal Curse · EMNLP 2024
Machine learning › Deep learning architectures and training
data augmentation
0.712023
DialoGPS: Dialogue Path Sampling in Continuous Semantic Space for Data Augmentation in Multi-Turn Conversations · ACL (1) 2023
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.712023
DialoGPS: Dialogue Path Sampling in Continuous Semantic Space for Data Augmentation in Multi-Turn Conversations · ACL (1) 2023
Computational finance and economics › financial market prediction › stock prediction
stock movement prediction
0.712023
PEN: Prediction-Explanation Network to Forecast Stock Price Movement with Better Explainability · AAAI 2023
Medical and health informatics
clinical decision support
0.612022
Debiased, Longitudinal and Coordinated Drug Recommendation through Multi-Visit Clinic Records · NeurIPS 2022
Medical and health informatics
electronic health records
0.612022
Debiased, Longitudinal and Coordinated Drug Recommendation through Multi-Visit Clinic Records · NeurIPS 2022
Medical and health informatics › clinical decision support
medication recommendation
0.612022
Debiased, Longitudinal and Coordinated Drug Recommendation through Multi-Visit Clinic Records · NeurIPS 2022
Natural language and speech › Language models and text generation › compositional generalization › length generalization
length extrapolation
0.312025
HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

shared representation learning · 1.3salient vector · 1.3logit lens · 0.9empirical attention analysis · 0.9router-only training · 0.8mixture of experts · 0.8attention enhancement · 0.8latent variable sampling · 0.7gaussian process · 0.7brownian bridge · 0.7satisfiability solving · 0.6front-door adjustment · 0.6causal inference · 0.6
YearPublicationVenuePosition
2025 HoPE: A Novel Positional Encoding Without Long-Term Decay for Enhanced Context Awareness and Extrapolation
abstract
Many positional encodings (PEs) are designed to exhibit long-term decay, based on an entrenched and long-standing inductive opinion: tokens farther away from the current position carry less relevant information. We argue that long-term decay is outdated in the era of LLMs, as LLMs are now applied to tasks demanding precise retrieval of in-context information from arbitrary positions. Firstly, we present empirical analyses on various PEs, demonstrating that models inherently learn attention with only a local-decay pattern while forming a U-shape pattern globally, contradicting the principle of long-term decay. Furthermore, we conduct a detailed analysis of rotary position encoding (RoPE, a prevalent relative positional encoding in LLMs), and found that the U-shape attention is caused by some learned components, which are also the key factor limiting RoPE’s expressiveness and extrapolation. Inspired by these insights, we propose High-frequency rotary Position Encoding (HoPE). HoPE replaces the specific components in RoPE with position-independent ones, retaining only high-frequency signals, which also breaks the principle of long-term decay in theory. HoPE achieves two major advantages: (1) Without constraints imposed by long-term decay, contradictory factors that limit attention optimization are removed. Thus, the model’s context awareness is enhanced. (2) HoPE exhibits greater robustness to the out-of-distribution behavior in attention patterns during extrapolation. The effectiveness of HoPE is validated through extensive experiments and with a large language model of up to 3 billion parameters.
Yuhan Chen 0001, Ang Lv, Jian Luan 0001, Bin Wang 0004, Wei Liu 0302
ACL (1)1
2025 MIN: Multi-Channel Interaction Network for Drug-Target Interaction With Protein Distillation
abstract
Traditional drug discovery processes are both time-consuming and require extensive professional expertise. With the accumulation of drug-target interaction (DTI) data from experimental studies, leveraging modern machine-learning techniques to discern patterns between drugs and target proteins has become increasingly feasible. In this paper, we introduce the Multi-channel Interaction Network (MIN), a novel framework designed to predict DTIs through two primary components: a representation learning module and a multi-channel interaction module. The representation learning module features a C-Score Predictor-assisted screening mechanism, which selects critical residues to enhance prediction accuracy and reduce noise. The multi-channel interaction module incorporates a structure-agnostic channel, a structure-aware channel, and an extended-mixture channel, facilitating the identification of interaction patterns at various levels for optimal complementarity. Additionally, contrastive learning is utilized to harmonize the representations of diverse data types. Our experimental evaluations on public datasets demonstrate that MIN surpasses other strong DTI prediction methods. Furthermore, the case study reveals a high overlap between the residues selected by the C-Score Predictor and those in actual binding pockets, underscoring MIN's explainability capability. These findings affirm that MIN is not only a potent tool for DTI prediction but also offers fresh insights into the prediction of protein binding sites.
Shuqi Li 0001, Shufang Xie 0003, Hongda Sun 0001, Yuhan Chen 0001, Tao Qin 0001, Tianjun Ke, Rui Yan 0001
IEEE Trans. Comput. Biol. Bioinform.4
2024 Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use
abstract
Yuhan Chen, Ang Lv, Ting-En Lin, Changyu Chen, Yuchuan Wu, Fei Huang, Yongbin Li, Rui Yan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yuhan Chen 0001, Ang Lv, Ting-En Lin, Changyu Chen, Yuchuan Wu, Fei Huang 0002, Yongbin Li 0001, Rui Yan 0001
ACL (1)1
2024 An Analysis and Mitigation of the Reversal Curse
abstract
Recent research observed a noteworthy phenomenon in large language models (LLMs), referred to as the "reversal curse."The reversal curse is that when dealing with two entities, denoted as a and b, connected by their relation R and its inverse R -1 , LLMs excel in handling sequences in the form of "aRb," but encounter challenges when processing "bR -1 a," whether in generation or comprehension.For instance, GPT-4 can accurately respond to the query "Tom Cruise's mother is?" with "Mary Lee Pfeiffer," but it struggles to provide a satisfactory answer when asked "Mary Lee Pfeiffer's son is?"In this paper, we undertake the first-ever study of how the reversal curse happens in LLMs.Our investigations reveal that the reversal curse can stem from the specific training objectives, which become particularly evident in the widespread use of next-token prediction within most causal language models.We hope this initial investigation can draw more attention to the reversal curse, as well as other underlying limitations in current LLMs. 1
Ang Lv, Shufang Xie 0003, Quan Tu, Yuhan Chen 0001, Ji-Rong Wen, Rui Yan 0001
EMNLP5
2024 Mixture of In-Context Experts Enhance LLMs' Long Context Awareness
abstract
Many studies have revealed that large language models (LLMs) exhibit uneven awareness of different contextual positions. Their limited context awareness can lead to overlooking critical information and subsequent task failures. While several approaches have been proposed to enhance LLMs' context awareness, achieving both effectiveness and efficiency remains challenging. In this paper, for LLMs utilizing RoPE as position embeddings, we introduce a novel method called "Mixture of In-Context Experts" (MoICE) to address this challenge. MoICE comprises two key components: a router integrated into each attention head within LLMs and a lightweight router-only training optimization strategy:(1) MoICE views each RoPE angle as an 'in-context' expert, demonstrated to be capable of directing the attention of a head to specific contextual positions. Consequently, each attention head flexibly processes tokens using multiple RoPE angles dynamically selected by the router to attend to the needed positions. This approach mitigates the risk of overlooking essential contextual information. (2) The router-only training strategy entails freezing LLM parameters and exclusively updating routers for only a few steps. When applied to open-source LLMs including Llama and Mistral, MoICE surpasses prior methods across multiple tasks on long context understanding and generation, all while maintaining commendable inference efficiency.
Hongzhan Lin 0002, Ang Lv, Yuhan Chen 0001, Chen Zhu 0003, Yang Song 0021, Hengshu Zhu, Rui Yan 0001
NeurIPS3
2024 Learning to Generate Style-Specific Adapters for Stylized Dialogue Generation
Jinpeng Li 0003, Yuhan Chen 0001, Pengfei Wu 0003, Yingce Xia, Shufang Xie 0003, Dongyan Zhao 0001, Rui Yan 0001
NLPCC (1)2
2023 PEN: Prediction-Explanation Network to Forecast Stock Price Movement with Better Explainability
abstract
Nowadays explainability in stock price movement prediction is attracting increasing attention in banks, hedge funds and asset managers, primarily due to audit or regulatory reasons. Text data such as financial news and social media posts can be part of the reasons for stock price movement. To this end, we propose a novel framework of Prediction-Explanation Network (PEN) jointly modeling text streams and price streams with alignment. The key component of the PEN model is an shared representation learning module that learns which texts are possibly associated with the stock price movement by modeling the interaction between the text data and stock price data with a salient vector characterizing their correlation. In this way, the PEN model is able to predict the stock price movement by identifying and utilizing abundant messages while on the other hand, the selected text messages also explain the stock price movement. Experiments on real-world datasets demonstrate that we are able to kill two birds with one stone: in terms of accuracy, the proposed PEN model outperforms the state-of-art baseline; on explainability, the PEN model are demonstrated to be far superior to attention mechanism, capable of picking out the crucial texts with a very high confidence.
Shuqi Li 0001, Weiheng Liao, Yuhan Chen 0001, Rui Yan 0001
AAAI3
2023 DialoGPS: Dialogue Path Sampling in Continuous Semantic Space for Data Augmentation in Multi-Turn Conversations
abstract
In open-domain dialogue generation tasks, contexts and responses in most datasets are oneto-one mapped, violating an important manyto-many characteristic: a context leads to various responses, and a response answers multiple contexts.Without such patterns, models poorly generalize and prefer responding safely.Many attempts have been made in either multiturn settings from a one-to-many perspective or in a many-to-many perspective but limited to single-turn settings.The major challenge to many-to-many augment multi-turn dialogues is that discretely replacing each turn with semantic similarity breaks fragile context coherence.In this paper, we propose DialoGue Path Sampling (DialoGPS) method in continuous semantic space, the first many-to-many augmentation method for multi-turn dialogues.Specifically, we map a dialogue to our extended Brownian Bridge, a special Gaussian process.We sample latent variables to form coherent dialogue paths in the continuous space.A dialogue path corresponds to a new multi-turn dialogue and is used as augmented training data.We show the effect of DialoGPS with both automatic and human evaluation.
Ang Lv, Jinpeng Li 0003, Yuhan Chen 0001, Ji Zhang 0011, Rui Yan 0001
ACL (1)3
2022 Debiased, Longitudinal and Coordinated Drug Recommendation through Multi-Visit Clinic Records
abstract
AI-empowered drug recommendation has become an important task in healthcare research areas, which offers an additional perspective to assist human doctors with more accurate and more efficient drug prescriptions. Generally, drug recommendation is based on patients' diagnosis results in the electronic health records. We assume that there are three key factors to be addressed in drug recommendation: 1) elimination of recommendation bias due to limitations of observable information, 2) better utilization of historical health condition and 3) coordination of multiple drugs to control safety. To this end, we propose DrugRec, a causal inference based drug recommendation model. The causal graphical model can identify and deconfound the recommendation bias with front-door adjustment. Meanwhile, we model the multi-visit in the causal graph to characterize a patient's historical health conditions. Finally, we model the drug-drug interactions (DDIs) as the propositional satisfiability (SAT) problem, and solving the SAT problem can help better coordinate the recommendation. Comprehensive experiment results show that our proposed model achieves state-of-the-art performance on the widely used datasets MIMIC-III and MIMIC-IV, demonstrating the effectiveness and safety of our method.
Hongda Sun 0001, Shufang Xie 0003, Shuqi Li 0001, Yuhan Chen 0001, Ji-Rong Wen, Rui Yan 0001
NeurIPS4