Kewei Liao

dblp:425/7958 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 43% Trustworthy machine learning · 37% Image recognition and object detection · 20%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability › representation engineering
activation editing
1.922026
Query-Routed Activation Editing with Truth-hierarchical Preference Optimization · AAAI 2026
Token-Aware Editing of Internal Activations for Large Language Model Alignment · EMNLP 2025
Natural language and speech › Language models and text generation
hallucination mitigation
1.012026
Query-Routed Activation Editing with Truth-hierarchical Preference Optimization · AAAI 2026
Computer vision › Image recognition and object detection › object detection
prohibited item detection
1.012026
Towards universal X-ray security inspection: a benchmark and stereoscopic-aware oriented prohibited item detection framework · Sci. China Inf. Sci. 2026
Natural language and speech › Language models and text generation › alignment
inference-time alignment
0.912025
Token-Aware Editing of Internal Activations for Large Language Model Alignment · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

stereoscopic-aware detection · 2.0truth-hierarchical preference optimization · 1.0query-routed activation editing · 1.0mutual information-guided graph aggregation · 0.9adaptive intervention · 0.9
YearPublicationVenuePosition
2026 Query-Routed Activation Editing with Truth-hierarchical Preference Optimization
abstract
Hallucination has emerged as a pivotal challenge of Large Language Models (LLMs) that generate plausible yet non‑factual content, significantly impeding the trustworthy AI applications in real-world scenarios like medical diagnosis and autonomous driving. Editing the internal activations of LLMs during inference has shown promising effectiveness in mitigating hallucinations with minimal cost. However, previous editing approaches neglect the query‑specific inference pathways that require tailored truthful steering vectors, resulting in suboptimal hallucination mitigation. To address these issues, we propose the Query-Routed Activation Editing (QRAE) framework, which comprises Divergence-sensitive Head Routing (DHR) and Truth-hierarchical Preference Steering (TPS), to fully leverage query-specific semantics for adaptive activation editing. Specifically, DHR is proposed to establish a query-aware head selection criterion, thereby dynamically routing to truth-critical attention heads. Subsequently, TPS introduces a query-specific steering vector calibration policy with the guidance of progressive truth-preferred optimization, enabling precise and adaptive editing for each distinct query. Extensive experiments on the widely recognized TruthfulQA benchmark demonstrate that QRAE outperforms SOTA methods by up to 13.2% in MC1. Meanwhile, QRAE demonstrates strong generalization to out-of-distribution TriviaQA and Natural Questions benchmarks.
Kewei Liao, Yuqing Ma, Zhange Zhang, Zhicheng Geng, Jiakai Wang, Xianglong Liu 0001
AAAI1
2026 Towards universal X-ray security inspection: a benchmark and stereoscopic-aware oriented prohibited item detection framework
Kewei Liao, Zhange Zhang, Yuqing Ma, Hongping Zhi, Aishan Liu, Ruihao Gong, Xianglong Liu 0001
Sci. China Inf. Sci.2
2025 Token-Aware Editing of Internal Activations for Large Language Model Alignment
abstract
Intervening the internal activations of large language models (LLMs) provides an effective inference-time alignment approach to mitigate undesirable behaviors, such as generating erroneous or harmful content, thereby ensuring safe and reliable applications of LLMs.However, previous methods neglect the misalignment discrepancy among varied tokens, resulting in deviant alignment direction and inflexible editing strength.To address these issues, we propose a token-aware editing (TAE) approach to fully utilize token-level alignment information in the activation space, therefore realizing superior post-intervention performance.Specifically, a Mutual Information-guided Graph Aggregation (MIG) module first develops an MI-guided graph to exploit the tokens' informative interaction for activation enrichment, thus improving alignment probing and facilitating intervention.Subsequently, Misalignment-aware Adaptive Intervention (MAI) comprehensively perceives the token-level misalignment degree from token representation and prediction to guide the adaptive adjustment of editing strength, thereby enhancing final alignment performance.Extensive experiments on three alignment capabilities demonstrate the efficacy of TAE, notably surpassing baseline by 25.8% on the primary metric of truthfulness with minimal cost. 1 MHSA
Yuqing Ma, Kewei Liao, Chengzhao Yang, Zhange Zhang, Jiakai Wang, Xianglong Liu 0001
EMNLP3