Yong Li 0004

dblp:l/YongLi4 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
10since 2021 · last 2026
0009-0005-1664-6425ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 47% Question answering and dialogue systems · 15% Trustworthy machine learning · 9%
Software engineering, system software, and programming languages
3 papers
Program synthesis and code generation · 31% Empirical software engineering · 18% Debugging and program repair · 16%
Human-computer interaction and pervasive computing
3 papers
Human-AI interaction · 50% Accessibility and assistive technology · 22% User interface design and tools · 22%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 20 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Human-AI interaction
GUI agent
2.022026
ProBench: Benchmarking GUI Agents with Accurate Process Information · AAAI 2026
History-Aware Reasoning for GUI Agents · AAAI 2026
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
1.012026
Code-Based English Models Reveal Surprising Performance on Chinese QA Pair Extraction Task · SIGIR 2026
Machine learning › Trustworthy machine learning › interpretability
faithful reasoning
1.012026
RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning Models · ACL (1) 2026
Natural language and speech › Language models and text generation
hallucination detection
1.012026
RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning Models · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model fine-tuning
1.012026
Code-Based English Models Reveal Surprising Performance on Chinese QA Pair Extraction Task · SIGIR 2026
Natural language and speech › Language models and text generation › large language model
large reasoning model
1.012026
RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning Models · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
question-answer extraction
1.012026
Code-Based English Models Reveal Surprising Performance on Chinese QA Pair Extraction Task · SIGIR 2026
Information retrieval › interactive information retrieval
conversational information seeking
1.012026
Efficient Memory Alignment for Long-term Conversational Information Seeking · SIGIR 2026
Computer vision › Vision and language › vision-language model › multimodal large language model
GUI agent
0.912025
PG-Agent: An Agent Powered by Page Graph · ACM Multimedia 2025
Natural language and speech › Language models and text generation › LLM agents
multimodal large language model agent
0.912025
PG-Agent: An Agent Powered by Page Graph · ACM Multimedia 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
PG-Agent: An Agent Powered by Page Graph · ACM Multimedia 2025
Compilers and program optimization
code efficiency optimization
0.912025
PEACE: Towards Efficient Project-Level Efficiency Optimization via Hybrid Code Editing · ASE 2025
Program synthesis and code generation
code generation with language models
0.912025
Issue Localization via LLM-Driven Iterative Code Graph Searching · ASE 2025
Debugging and program repair
fault localization
0.912025
Issue Localization via LLM-Driven Iterative Code Graph Searching · ASE 2025
Software maintenance and evolution › issue management
issue resolution
0.912025
Issue Localization via LLM-Driven Iterative Code Graph Searching · ASE 2025
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.312026
Efficient Memory Alignment for Long-term Conversational Information Seeking · SIGIR 2026
Natural language and speech › Language models and text generation
LLM agents
0.312026
History-Aware Reasoning for GUI Agents · AAAI 2026
Natural language and speech › Question answering and dialogue systems › personalized dialogue
persona-grounded dialogue
0.312026
Efficient Memory Alignment for Long-term Conversational Information Seeking · SIGIR 2026
Interaction techniques and input
mobile interaction
0.312026
ProBench: Benchmarking GUI Agents with Accurate Process Information · AAAI 2026
Machine learning › Graph learning
graph neural network
0.312025
Towards an Inclusive Mobile Web: A Dataset and Framework for Focusability in UI Accessibility · WWW 2025

Methods — techniques the papers use, named apart from their topics

large language model · 4.7retrieval · 2.0reflection-guided memory editing · 2.0vocabulary expansion · 1.0multimodal model · 1.0large language model fine-tuning · 1.0hidden-state cosine similarity · 1.0entropy-guided stepwise scaling · 1.0attention routing analysis · 1.0user study · 0.9task decomposition · 0.9retrieval-augmented generation · 0.9reflection mechanism · 0.9page graph · 0.9graph neural network · 0.9formative study · 0.9code graph search · 0.9code editing · 0.9
YearPublicationVenuePosition
2026 History-Aware Reasoning for GUI Agents
Leyang Yang, Xiaoxuan Tang, Sheng Zhou 0004, Dajun Chen, Wei Jiang 0041, Yong Li 0004
AAAI7
2026 ProBench: Benchmarking GUI Agents with Accurate Process Information
abstract
With the deep integration of artificial intelligence and interactive technology, Graphical User Interface (GUI) Agent, as the carrier connecting goal-oriented natural language and real-world devices, has received widespread attention from the community. Contemporary benchmarks aim to evaluate the comprehensive capabilities of GUI agents in GUI operation tasks, generally determining task completion solely by inspecting the final screen state. However, GUI operation tasks consist of multiple chained steps while not all critical information is presented in the final few pages. Although a few research has begun to incorporate intermediate steps into evaluation, accurately and automatically capturing this process information still remains an open challenge. To address this weakness, we introduce ProBench, a comprehensive mobile benchmark with over 200 challenging GUI tasks covering widely-used scenarios. Remaining the traditional State-related Task evaluation, we extend our dataset to include Process-related Task and design a specialized evaluation method. A newly introduced Process Provider automatically supplies accurate process information, enabling presice assessment of agent's performance. Our evaluation of advanced GUI agents reveals significant limitations for real-world GUI scenarios. These shortcomings are prevalent across diverse models, including both large-scale generalist models and smaller, GUI-specific models. A detailed error analysis further exposes several universal problems, outlining concrete directions for future improvements.
Leyang Yang, Xiaoxuan Tang, Sheng Zhou 0004, Dajun Chen, Wei Jiang 0041, Yong Li 0004
AAAI7
2026 RFS-Guard: Detecting Reasoning Hallucinations via Cross-Phase Routing Focus in Large Reasoning Models
abstract
Large reasoning models (LRMs) achieve strong performance on complex tasks by generating intermediate reasoning before the final answer, yet they remain prone to reasoning hallucinations such as subtle arithmetic or constraintviolation errors.Prior hallucination detectors often rely on external verification or local tokenlevel signals, which are limited for LRMs and largely overlook whether the cross-phase information flow from reasoning to answering is structurally robust.We propose Routing Focus Score (RFS), a step-level indicator that measures how strongly cross-step attention routing aligns with semantic proximity derived from hidden-state cosine similarity.We further design RFS-Guard, a lightweight hallucination detection framework based on RFS.Empirically, we observe that higher reasoning-answer RFS is consistently associated with higher hallucination risk, suggesting a routing-collapse failure mode where models might prefer selfconfirmation loops and suppress the ability to audit their own generations.Experimental results across multiple domains and models demonstrate the superiority of RFS-Guard for detecting and localizing hallucinations in LRMs without requiring external tools or repeated sampling.
Zihang Liu 0001, Zhouhua Fang, Yong Li 0004, Haishuai Wang
ACL (1)5
2026 EGSS: Entropy-guided Stepwise Scaling for Reliable Software Engineering
abstract
Chenhui Mao, Yuanting Lei, Zhixiang Wei, Ming Liang, Zhixiang Wang, Jingxuan Xu, Dajun Chen, Wei Jiang, Yong Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chenhui Mao, Yuanting Lei, Zhixiang Wei, Jingxuan Xu, Dajun Chen, Wei Jiang 0041, Yong Li 0004
ACL (1)9
2026 Efficient Memory Alignment for Long-term Conversational Information Seeking
abstract
Long-term conversational agents rely on personal memory to maintain coherence and personalization, yet practical systems must operate under context budgets and cope with evolving or contradictory user information. We frame persona memory as a retrieval problem over a growing memory store, and propose REMAP, a reflection-guided memory editing approach for online alignment of persona facts that selectively writes and revises memory entries based on the current dialogue evidence and retrieved related items. The method aims to preserve salient facts while reducing redundancy and resolving apparent conflicts, enabling more efficient context utilization over extended interaction horizons. Experiments on multi-session dialogue datasets show consistent gains in persona-consistent retrieval and response continuity over commonly used memory strategies, while achieving more selective memory updates under comparable operational overhead.
Qingyang Xu, Xiao Liu 0045, Zhouhua Fang, Yong Li 0004, Vincent Lee, Haishuai Wang
SIGIR4
2026 Code-Based English Models Reveal Surprising Performance on Chinese QA Pair Extraction Task
abstract
This paper explores advancements in automated Question-Answer (QA) extraction using large language models (LLMs), addressing challenges in transforming unstructured text into high-quality, retrievable QA pairs. Traditional approaches, whether through segmented question and answer generation or end-to-end extraction, often struggle with efficiency, dataset limitations, and performance consistency. Leveraging recent progress in LLMs, we constructed a large-scale Chinese QA extraction dataset with 143,846 documents and evaluated multiple fine-tuned models on public and private datasets. Surprisingly, code-based English LLMs outperformed Chinese-specialized models on Chinese text with a lower hallucination rate. Building upon this finding, we enhanced the best-performing code-based model with an expanded Chinese vocabulary, creating Code Llama-M, which achieved better results. Integrating Code Llama-M into our internal assistant, Luo Ying, demonstrated notable user satisfaction gains, affirming its practical impact. Key contributions include: (i) creation of a robust Chinese QA extraction instruction dataset; (ii) evidence of cross-lingual efficacy of code-based LLMs for Chinese QA tasks, further enhanced through Code Llama-M's expanded Chinese vocabulary; and (iii) successful application of the fine-tuned LLM in a live assistant system, enhancing user experience.
Jiajun Yu, Linghan Zheng, Jiayuan Dong, Yaozhen Liang, Yong Li 0004, Haishuai Wang
SIGIR8
2025 Issue Localization via LLM-Driven Iterative Code Graph Searching
abstract
Issue solving aims to generate patches to fix re-ported issues in real-world code repositories according to issue descriptions. Issue localization forms the basis for accurate issue solving. Recently, large language model (LLM) based issue localization methods have demonstrated state-of-the-art performance. However, these methods either search from files mentioned in issue descriptions or in the whole repository and struggle to balance the breadth and depth of the search space to converge on the target efficiently. Moreover, they allow LLM to explore whole repositories freely, making it challenging to control the search direction to prevent the LLM from searching for incorrect targets. Meanwhile, because LLMs may not correctly produce the required interaction formats with the environment, they suffer from search failures.This paper introduces COSIL, an LLM-driven, powerful function-level issue localization method without training or indexing. To balance search breadth and depth, COSIL employs a two-phase code graph search strategy. It first conducts broad exploration at the file level using dynamically constructed module call graphs, and then performs in-depth analysis at the function level by expanding the module call graph into a function call graph and executing iterative searches. To precisely control the search direction, COSIL designs a pruner to filter unrelated directions and irrelevant contexts. To avoid incorrect interaction formats in long contexts, COSIL introduces a reflection mechanism that uses additional independent queries in short contexts to enhance formatted abilities. Experiment results demonstrate that COSIL achieves a Top-1 localization accuracy of 43.3% and 44.6% on SWE-bench Lite and SWE-bench Verified, respectively, with Qwen2.5-Coder-32B, average outperforming the state-of-the-art methods by 96.04%. When COSIL is integrated into an issue-solving method, Agentless, the issue resolution rate improves by 2.98%–30.5%.
Zhonghao Jiang, Xiaoxue Ren, Meng Yan 0001, Wei Jiang 0041, Yong Li 0004, Zhongxin Liu 0002
ASE5
2025 PEACE: Towards Efficient Project-Level Efficiency Optimization via Hybrid Code Editing
abstract
Large Language Models (LLMs) have demonstrated significant capability in code generation, but their potential in code efficiency optimization remains underexplored. Previous LLM-based code efficiency optimization approaches exclusively focus on function-level optimization and overlook interaction between functions, failing to generalize to real-world development scenarios. Code editing techniques show great potential for conducting project-level optimization, yet they face challenges associated with invalid edits and suboptimal internal functions. To address these gaps, we propose PEACE, a novel hybrid framework for Project-Level code Efficiency optimization through Automatic Code Editing, which also ensures the overall correctness and integrity of the project. PEACE integrates three key phases: dependency-aware optimizing function sequence construction, valid associated edits identification, and efficiency optimization editing iteration. To rigorously evaluate the effectiveness of PEACE, we construct PEACEXEC, the first benchmark comprising 146 real-world optimization tasks from 47 high-impact GitHub Python projects, along with highly qualified test cases and executable environments. Extensive experiments demonstrate PEACE’s superiority over the state-of-the-art baselines, achieving a 69.2% correctness rate (pass@1), +46.9% opt rate, and 0.840 speedup in execution efficiency. Notably, our PEACE outperforms all baselines by significant margins, particularly in complex optimization tasks with multiple functions. Moreover, extensive experiments are also conducted to validate the contributions of each component in PEACE, as well as the rationale and effectiveness of our hybrid framework design.
Xiaoxue Ren, Yun Peng 0003, Zhongxin Liu 0002, Dajun Chen, Wei Jiang 0041, Yong Li 0004
ASE8
2025 PG-Agent: An Agent Powered by Page Graph
abstract
Graphical User Interface (GUI) agents possess significant commercial and social value, and GUI agents powered by advanced multimodal large language models (MLLMs) have demonstrated remarkable potential. Currently, existing GUI agents usually utilize sequential episodes of multi-step operations across pages as the prior GUI knowledge, which fails to capture the complex transition relationship between pages, making it challenging for the agents to deeply perceive the GUI environment and generalize to new scenarios. Therefore, we design an automated pipeline to transform the sequential episodes into page graphs, which explicitly model the graph structure of the pages that are naturally connected by actions. To fully utilize the page graphs, we further introduce Retrieval-Augmented Generation (RAG) technology to effectively retrieve reliable perception guidelines of GUI from them, and a tailored multi-agent framework PG-Agent with task decomposition strategy is proposed to be injected with the guidelines so that it can generalize to unseen scenarios. Extensive experiments on various benchmarks demonstrate the effectiveness of PG-Agent, even with limited episodes for page graph construction. Our codes will be publicly available at https://github.com/chenwz-123/PG-Agent.
Weizhi Chen, Leyang Yang, Sheng Zhou 0004, Xiaoxuan Tang, Jiajun Bu, Yong Li 0004, Wei Jiang 0041
ACM Multimedia7
2025 Towards an Inclusive Mobile Web: A Dataset and Framework for Focusability in UI Accessibility
abstract
The rapid growth of mobile web technologies has revolutionized how people manage daily activities, emphasizing the critical need for accessible mobile user interfaces (UIs) that accommodate users with disabilities and situational impairments. Current AI-driven UI understanding methods show promise but primarily target general UI modeling, neglecting nuanced, user-centric accessibility requirements. To bridge this gap, we first conducted a formative study with 12 visually impaired participants. Our study uncovers selective-accessible issues, a new class of accessibility challenges requiring finer granularity and selective focus on UI components, which existing methods largely overlook. Our findings also reveal that the severity of issues varies across interaction stages, with earlier stages posing a more significant impact. Building on these insights, we propose a comprehensive framework of three accessibility stages: focusability, information, and functionality (FIF), encompassing 12 sub-tasks under 3 overarching tasks. Identifying UI element focusability prediction (UFP) as a pivotal yet underexplored task within FIF, hindered by the absence of dedicated datasets, we introduce a new dataset (NOS) with 117,480 annotated components addressing accessibility issues comprehensively. To further enhance UFP, we introduce Graph-based UI Focusability Prediction (GIFT), a method leveraging graph neural networks to model UFP-targeted UI relationships. User studies validate the dataset's quality, while experiments show GIFT's effectiveness in improving UFP outcomes. Our code and datasets are publicly available to support further web inclusivity advancements at https://github.com/eaglelab-zju/NOS.
Ming Gu 0014, Sheng Zhou 0004, Ming Shen 0003, Zirui Gao, Wei Jiang 0041, Yong Li 0004, Jiajun Bu
WWW10
2003 Group undo framework and algorithms in real-time collaborative image editing systems
abstract
The ability to undo operations is an indispensable feature of single user editing systems, but supporting group undo in real-time collaborative editing systems is still a difficult problem. In this paper, we propose an undo framework and algorithms to achieve group undo in image-based collaborative graphics editing systems. The basic idea is to interpret an undo command as a concurrent inverse operation by means of image operation transformation algorithm, so that an operation is always undoable under its current context. Through exploiting relations among operations, space cost for operation preservation is greatly reduced. The global undo, local undo and selective undo mode are supported in our solution. The undo algorithms are also applicable in single-user applications. The algorithms are implemented in CoDesign-a multi-level collaborative graphics designing system, which aims at supporting both object-based and image-based collaborative pattern design.
Xianghua Xu, Jiajun Bu, Chun Chen 0001, Yong Li 0004
SMC4
2003 Consistency maintenance in real-time collaborative image editing systems
Xianghua Xu, Chun Chen 0001, Jiajun Bu, Yong Li 0004
SMC4