VLDB 2026 Research / reviewers in the wild / expert
Wei Jiang 0041
dblp:21/3839-41
· DBLP profile ↗
15ranked-venue papers
0as first author
15since 2021 · last 2026
0009-0003-6605-9793ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | History-Aware Reasoning for GUI Agents
Leyang Yang, Xiaoxuan Tang, Sheng Zhou 0004, Dajun Chen, Wei Jiang 0041, Yong Li 0004 |
AAAI | 6 |
| 2026 | ProBench: Benchmarking GUI Agents with Accurate Process InformationabstractWith the deep integration of artificial intelligence and interactive technology, Graphical User Interface (GUI) Agent, as the carrier connecting goal-oriented natural language and real-world devices, has received widespread attention from the community. Contemporary benchmarks aim to evaluate the comprehensive capabilities of GUI agents in GUI operation tasks, generally determining task completion solely by inspecting the final screen state. However, GUI operation tasks consist of multiple chained steps while not all critical information is presented in the final few pages. Although a few research has begun to incorporate intermediate steps into evaluation, accurately and automatically capturing this process information still remains an open challenge. To address this weakness, we introduce ProBench, a comprehensive mobile benchmark with over 200 challenging GUI tasks covering widely-used scenarios. Remaining the traditional State-related Task evaluation, we extend our dataset to include Process-related Task and design a specialized evaluation method. A newly introduced Process Provider automatically supplies accurate process information, enabling presice assessment of agent's performance. Our evaluation of advanced GUI agents reveals significant limitations for real-world GUI scenarios. These shortcomings are prevalent across diverse models, including both large-scale generalist models and smaller, GUI-specific models. A detailed error analysis further exposes several universal problems, outlining concrete directions for future improvements. Leyang Yang, Xiaoxuan Tang, Sheng Zhou 0004, Dajun Chen, Wei Jiang 0041, Yong Li 0004 |
AAAI | 6 |
| 2026 | EGSS: Entropy-guided Stepwise Scaling for Reliable Software EngineeringabstractChenhui Mao, Yuanting Lei, Zhixiang Wei, Ming Liang, Zhixiang Wang, Jingxuan Xu, Dajun Chen, Wei Jiang, Yong Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chenhui Mao, Yuanting Lei, Zhixiang Wei, Jingxuan Xu, Dajun Chen, Wei Jiang 0041, Yong Li 0004 |
ACL (1) | 8 |
| 2025 | An Empirical Study on Commit Message Generation Using LLMs via In-Context LearningabstractCommit messages concisely describe code changes in natural language and are important for software maintenance. Several approaches have been proposed to automatically generate commit messages, but they still suffer from critical limitations, such as time-consuming training and poor generalization ability. To tackle these limitations, we propose to borrow the weapon of large language models (LLMs) and in-context learning (ICL). Our intuition is based on the fact that the training corpora of LLMs contain extensive code changes and their pairwise commit messages, which makes LLMs capture the knowledge about commits, while ICL can exploit the knowledge hidden in the LLMs and enable them to perform downstream tasks without model tuning. However, it remains unclear how well LLMs perform on commit message generation via ICL. In this paper, we conduct an empirical study to investigate the capability of LLMs to generate commit messages via ICL. Specifically, we first explore the impact of different settings on the performance of ICL-based commit message generation. We then compare ICL-based commit message generation with state-of-the-art approaches on a popular multilingual dataset and a new dataset we created to mitigate potential data leakage. The results show that ICL-based commit message generation significantly outperforms state-of-the-art approaches on subjective evaluation and achieves better generalization ability. We further analyze the root causes for LLM's underperformance and propose several implications, which shed light on future research directions for using LLMs to generate commit messages. Yifan Wu 0002, Ying Li 0012, Siyu Yu, Wei Jiang 0041 |
ICSE | 7 |
| 2025 | Issue Localization via LLM-Driven Iterative Code Graph SearchingabstractIssue solving aims to generate patches to fix re-ported issues in real-world code repositories according to issue descriptions. Issue localization forms the basis for accurate issue solving. Recently, large language model (LLM) based issue localization methods have demonstrated state-of-the-art performance. However, these methods either search from files mentioned in issue descriptions or in the whole repository and struggle to balance the breadth and depth of the search space to converge on the target efficiently. Moreover, they allow LLM to explore whole repositories freely, making it challenging to control the search direction to prevent the LLM from searching for incorrect targets. Meanwhile, because LLMs may not correctly produce the required interaction formats with the environment, they suffer from search failures.This paper introduces COSIL, an LLM-driven, powerful function-level issue localization method without training or indexing. To balance search breadth and depth, COSIL employs a two-phase code graph search strategy. It first conducts broad exploration at the file level using dynamically constructed module call graphs, and then performs in-depth analysis at the function level by expanding the module call graph into a function call graph and executing iterative searches. To precisely control the search direction, COSIL designs a pruner to filter unrelated directions and irrelevant contexts. To avoid incorrect interaction formats in long contexts, COSIL introduces a reflection mechanism that uses additional independent queries in short contexts to enhance formatted abilities. Experiment results demonstrate that COSIL achieves a Top-1 localization accuracy of 43.3% and 44.6% on SWE-bench Lite and SWE-bench Verified, respectively, with Qwen2.5-Coder-32B, average outperforming the state-of-the-art methods by 96.04%. When COSIL is integrated into an issue-solving method, Agentless, the issue resolution rate improves by 2.98%–30.5%. Zhonghao Jiang, Xiaoxue Ren, Meng Yan 0001, Wei Jiang 0041, Yong Li 0004, Zhongxin Liu 0002 |
ASE | 4 |
| 2025 | PEACE: Towards Efficient Project-Level Efficiency Optimization via Hybrid Code EditingabstractLarge Language Models (LLMs) have demonstrated significant capability in code generation, but their potential in code efficiency optimization remains underexplored. Previous LLM-based code efficiency optimization approaches exclusively focus on function-level optimization and overlook interaction between functions, failing to generalize to real-world development scenarios. Code editing techniques show great potential for conducting project-level optimization, yet they face challenges associated with invalid edits and suboptimal internal functions. To address these gaps, we propose PEACE, a novel hybrid framework for Project-Level code Efficiency optimization through Automatic Code Editing, which also ensures the overall correctness and integrity of the project. PEACE integrates three key phases: dependency-aware optimizing function sequence construction, valid associated edits identification, and efficiency optimization editing iteration. To rigorously evaluate the effectiveness of PEACE, we construct PEACEXEC, the first benchmark comprising 146 real-world optimization tasks from 47 high-impact GitHub Python projects, along with highly qualified test cases and executable environments. Extensive experiments demonstrate PEACE’s superiority over the state-of-the-art baselines, achieving a 69.2% correctness rate (pass@1), +46.9% opt rate, and 0.840 speedup in execution efficiency. Notably, our PEACE outperforms all baselines by significant margins, particularly in complex optimization tasks with multiple functions. Moreover, extensive experiments are also conducted to validate the contributions of each component in PEACE, as well as the rationale and effectiveness of our hybrid framework design. Xiaoxue Ren, Yun Peng 0003, Zhongxin Liu 0002, Dajun Chen, Wei Jiang 0041, Yong Li 0004 |
ASE | 7 |
| 2025 | PG-Agent: An Agent Powered by Page GraphabstractGraphical User Interface (GUI) agents possess significant commercial and social value, and GUI agents powered by advanced multimodal large language models (MLLMs) have demonstrated remarkable potential. Currently, existing GUI agents usually utilize sequential episodes of multi-step operations across pages as the prior GUI knowledge, which fails to capture the complex transition relationship between pages, making it challenging for the agents to deeply perceive the GUI environment and generalize to new scenarios. Therefore, we design an automated pipeline to transform the sequential episodes into page graphs, which explicitly model the graph structure of the pages that are naturally connected by actions. To fully utilize the page graphs, we further introduce Retrieval-Augmented Generation (RAG) technology to effectively retrieve reliable perception guidelines of GUI from them, and a tailored multi-agent framework PG-Agent with task decomposition strategy is proposed to be injected with the guidelines so that it can generalize to unseen scenarios. Extensive experiments on various benchmarks demonstrate the effectiveness of PG-Agent, even with limited episodes for page graph construction. Our codes will be publicly available at https://github.com/chenwz-123/PG-Agent. Weizhi Chen, Leyang Yang, Sheng Zhou 0004, Xiaoxuan Tang, Jiajun Bu, Yong Li 0004, Wei Jiang 0041 |
ACM Multimedia | 8 |
| 2025 | Towards an Inclusive Mobile Web: A Dataset and Framework for Focusability in UI AccessibilityabstractThe rapid growth of mobile web technologies has revolutionized how people manage daily activities, emphasizing the critical need for accessible mobile user interfaces (UIs) that accommodate users with disabilities and situational impairments. Current AI-driven UI understanding methods show promise but primarily target general UI modeling, neglecting nuanced, user-centric accessibility requirements. To bridge this gap, we first conducted a formative study with 12 visually impaired participants. Our study uncovers selective-accessible issues, a new class of accessibility challenges requiring finer granularity and selective focus on UI components, which existing methods largely overlook. Our findings also reveal that the severity of issues varies across interaction stages, with earlier stages posing a more significant impact. Building on these insights, we propose a comprehensive framework of three accessibility stages: focusability, information, and functionality (FIF), encompassing 12 sub-tasks under 3 overarching tasks. Identifying UI element focusability prediction (UFP) as a pivotal yet underexplored task within FIF, hindered by the absence of dedicated datasets, we introduce a new dataset (NOS) with 117,480 annotated components addressing accessibility issues comprehensively. To further enhance UFP, we introduce Graph-based UI Focusability Prediction (GIFT), a method leveraging graph neural networks to model UFP-targeted UI relationships. User studies validate the dataset's quality, while experiments show GIFT's effectiveness in improving UFP outcomes. Our code and datasets are publicly available to support further web inclusivity advancements at https://github.com/eaglelab-zju/NOS. Ming Gu 0014, Sheng Zhou 0004, Ming Shen 0003, Zirui Gao, Wei Jiang 0041, Yong Li 0004, Jiajun Bu |
WWW | 9 |
| 2024 | CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window ExtendingabstractSelf-attention and position embedding are two crucial modules in transformer-based Large Language Models (LLMs).However, the potential relationship between them is far from well studied, especially for long context window extending.In fact, anomalous behaviors that hinder long context extrapolation exist between Rotary Position Embedding (RoPE) and vanilla self-attention.Incorrect initial angles between Q and K can cause misestimation in modeling rotary position embedding of the closest tokens.To address this issue, we propose Collinear Constrained Attention mechanism, namely CoCA.Specifically, we enforce a collinear constraint between Q and K to seamlessly integrate RoPE and self-attention.While only adding minimal computational and spatial complexity, this integration significantly enhances long context window extrapolation ability.We provide an optimized implementation, making it a drop-in replacement for any existing transformer-based models.Extensive experiments demonstrate that CoCA excels in extending context windows.A CoCAbased GPT model, trained with a context length of 512, can extend the context window up to 32K (60×) without any fine-tuning.Additionally, incorporating CoCA into LLaMA-7B achieves extrapolation up to 32K within a training length of only 2K.Our code is publicly available at: https://github.com/codefuse- ai/Collinear-Constrained-Attention Shiyi Zhu, Wei Jiang 0041, Siqiao Xue, Yifan Wu 0002 |
ACL (1) | 3 |
| 2024 | RepoGenix: Dual Context-Aided Repository-Level Code Completion with Language ModelsabstractThe success of language models in code assistance has spurred the proposal of repository-level code completion as a means to enhance prediction accuracy, utilizing the context from the entire codebase. However, this comprehensive context comes at a cost: while it enhances model performance, it also increases inference latency. This balance between improved accuracy and computational efficiency poses a significant challenge in real-world applications. We present RepoGenix, a solution that enhances repository-level code completion without increased latency. RepoGenix combines analogous context and relevant context, using Context-Aware Selection technology to efficiently compress these contexts into limited-size prompts. Our experiments on CrossCodeEval demonstrate that RepoGenix not only achieves a substantial 48.41% reduction in inference time, but also yields improvement in performance compared to baseline methods. We have successfully implemented and tested RepoGenix within AntGroup's development environments. This approach is being extended to multiple programming languages and will be open-sourced, aiming to enhance code completion efficiency for the broader developer community. Xiaoheng Xie, Gehao Zhang, Xunjin Zheng, Peng Di, Wei Jiang 0041, Chengpeng Wang 0001, Gang Fan |
ASE | 6 |
| 2024 | MFTCoder: Boosting Code LLMs with Multitask Fine-TuningabstractCode LLMs have emerged as a specialized research field, with remarkable studies dedicated to enhancing model's coding capabilities through fine-tuning on pre-trained models. Previous fine-tuning approaches were typically tailored to specific downstream tasks or scenarios, which meant separate fine-tuning for each task, requiring extensive training resources and posing challenges in terms of deployment and maintenance. Furthermore, these approaches failed to leverage the inherent interconnectedness among different code-related tasks. To overcome these limitations, we present a multi-task fine-tuning framework, MFTCoder, that enables simultaneous and parallel fine-tuning on multiple tasks. By incorporating various loss functions, we effectively address common challenges in multi-task learning, such as data imbalance, varying difficulty levels, and inconsistent convergence speeds. Extensive experiments have conclusively demonstrated that our multi-task fine-tuning approach outperforms both individual fine-tuning on single tasks and fine-tuning on a mixed ensemble of tasks. Moreover, MFTCoder offers efficient training capabilities, including efficient data tokenization modes and parameter efficient fine-tuning (PEFT) techniques, resulting in significantly improved speed compared to traditional fine-tuning methods. MFTCoder seamlessly integrates with several mainstream open-source LLMs, such as CodeLLama and Qwen. Our MFTCoder fine-tuned CodeFuse-DeepSeek-33B claimed the top spot on the Big Code Models Leaderboard ranked by WinRate as of January 30, 2024. MFTCoder is open-sourced at https://github.com/codefuse-ai/MFTCOder Bingchang Liu, Chaoyu Chen, Zi Gong, Cong Liao, Zhichao Lei, Dajun Chen, Hailian Zhou, Wei Jiang 0041, Hang Yu 0002 |
KDD | 11 |
| 2024 | MMDL-Based Data Augmentation with Domain Knowledge for Time Series Classification
Xiaosheng Li, Yifan Wu 0002, Wei Jiang 0041, Ying Li 0012 |
ECML/PKDD (3) | 3 |
| 2024 | Understanding and Improving Change Risk Detection in PracticeabstractChanges are inevitable and frequent in large-scale online service systems, which has been one of the leading causes that induce incidents. Change risk detection (CRD) aims to help engineers detect high-risk changes so that proactive actions can be taken to avoid incidents, which is vital for the availability and reliability of online service systems. Though some efforts have been dedicated to CRD, their performances are still far from satisfactory in practice. To better understand the practical challenges of CRD, we conducted the first empirical study on a large-scale online service system in Ant Group. Through this study, we identified four critical challenges, including poor interpretability, adaptation to diverse change types, indirect anomaly factors, and expected but false alarm anomalies. To address these challenges, we propose an effective and eXplainable Change Risk Detection framework named XCRD. XCRD can detect change-induced unexpected anomalies using multi-source data and provide explainable alerts for engineers to facilitate anomaly diagnosis and mitigation. We have successfully deployed XCRD in Ant Group for the past 14 months, demonstrating a significant performance improvement in CRD. We also discuss some successful cases and lessons learned during our study. To our knowledge, we are the first to deeply investigate CRD in industrial scenarios. We believe that our work can provide valuable insights for engineers and researchers to understand and improve CRD in practice. Yifan Wu 0002, Ying Li 0012, Bingxu Chai, Wei Jiang 0041 |
SANER | 6 |
| 2024 | DeepScaling: Autoscaling Microservices With Stable CPU Utilization for Large Scale Production Cloud SystemsabstractCloud service providers often provision excessive resources to meet the desired Service Level Objectives (SLOs), by setting lower CPU utilization targets. This can result in a waste of resources and a noticeable increase in power consumption in large-scale cloud deployments. To address this issue, this paper presents DeepScaling, an innovative solution for minimizing resource cost while ensuring SLO requirements are met in a dynamic, large-scale production microservice-based system. We propose DeepScaling, which introduces three innovative components to adaptively refine the target CPU utilization of servers in the data center, and we maintain it at a stable value to meet SLO constraints while using minimum amount of system resources. First, DeepScaling forecasts workloads for each service using a Spatio-temporal Graph Neural Network. Secondly, it estimates CPU utilization with a Deep Neural Network, considering factors such as periodic tasks and traffic. Finally, it uses a modified Deep Q-Network (DQN) to generate an autoscaling policy that controls service resources to maximize service stability while meeting SLOs. Evaluation of DeepScaling in Ant Group’s large-scale cloud environment shows that it outperforms state-of-the-art autoscaling approaches in terms of maintaining stable performance and resource savings. The deployment of DeepScaling in the real-world environment of 1900+ microservices saves the provisioning of over 100,000 CPU cores per day, on average. Shiyi Zhu, Wei Jiang 0041, K. K. Ramakrishnan, Meng Yan 0001, Xiaohong Zhang 0002, Alex X. Liu |
IEEE/ACM Trans. Netw. | 4 |
| 2022 | DeepScaling: microservices autoscaling for stable CPU utilization in large scale cloud systemsabstractCloud service providers conservatively provision excessive resources to ensure service level objectives (SLOs) are met. They often set lower CPU utilization targets to ensure service quality is not degraded, even when the workload varies significantly. Not only does this potentially waste resources, but it can also consume excessive power in large-scale cloud deployments. This paper aims to minimize resource costs while ensuring SLO requirements are met in a dynamically varying, large-scale production microservice environment. We propose DeepScaling, which introduces three innovative components to adaptively refine the target CPU utilization to a level that is maintained at a stable value to meet SLO constraints while using minimum resources. First, DeepScaling forecasts the workload for each service using a Spatio-temporal Graph Neural Network. Second, DeepScaling estimates the CPU utilization by mapping the workload intensity to an estimated CPU utilization with a Deep Neural Network, while taking into account multiple factors in the cloud environment (e.g., periodic tasks and traffic). Third, DeepScaling generates an autoscaling policy for each service based on an improved Deep Q Network (DQN). The adaptive autoscaling policy updates the target CPU utilization to be a maximum, stable value, while ensuring SLOs is not violated. We compare DeepScaling with state-of-the-art autoscaling approaches in the large-scale production cloud environment of the Ant Group. It shows that DeepScaling outperforms other approaches both in terms of maintaining stable service performance, and saving resources, by a significant margin. The deployment of DeepScaling in Ant Group's real production environment with 135 microservices saves the provisioning of over 30,000 CPU cores per day, on average. Shiyi Zhu, Wei Jiang 0041, K. K. Ramakrishnan, Yangfei Zheng, Meng Yan 0001, Xiaohong Zhang 0002, Alex X. Liu |
SoCC | 4 |