VLDB 2026 Research / reviewers in the wild / expert
Junxiao Han
dblp:231/8546
· DBLP profile ↗
17ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-7630-8667ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CodeGlance: Understanding Code Reasoning Challenges in LLMs through Multi-Dimensional Feature AnalysisabstractIn modern software development, developers frequently need to understand code behavior at a glance—whether reviewing pull requests, debugging issues, or navigating unfamiliar codebases. This ability to reason about dynamic program behavior is fundamental to effective software engineering and increasingly supported by Large Language Models (LLMs). However, existing studies on code reasoning focus primarily on isolated code snippets, overlooking the complexity of real-world scenarios involving external API interactions and unfamiliar functions. This gap hinders our understanding of what truly makes code reasoning challenging for LLMs across diverse programming contexts. Yunkun Wang, Xuanhe Zhang, Junxiao Han, Chen Zhi, Shuiguang Deng |
ICPC | 3 |
| 2026 | An Entropy-Based Privacy-Preserving Federated Deep Reinforcement Learning Framework for Task Offloading in Vehicular Edge Computing NetworksabstractWith the rapid evolution of 5G and the ongoing development of 6G technologies, the Internet of Vehicles (IoV) is expected to play a critical role in next-generation intelligent transportation systems. Applications such as autonomous driving, augmented reality, and smart mobility not only require ultra-low latency and high computational efficiency, but also demand enhanced trustworthiness and privacy assurance. To address these demands, Vehicular Edge Computing (VEC) has emerged as a foundational paradigm for 6G-IoT, enabling intelligent services by offloading tasks from vehicles to edge nodes. However, task offloading in IoV-VEC systems still faces critical challenges, including the need for responsible AI decision-making under dynamic network conditions and the protection of sensitive vehicular data. This paper proposes FedVTO, a privacy-preserving federated vehicle task offloading framework that integrates Federated Learning (FL) and Deep Reinforcement Learning (DRL) to optimize task offloading decisions and resource allocation strategies in VEC networks. By incorporating information entropy models and dynamically adjusting weighting parameters using an entropy-based method within a three-tier architecture (vehicles, roadside units, and cloud server), FedVTO minimizes latency, energy consumption, and privacy leakage. Experimental results show that FedVTO significantly improves task offloading efficiency and mitigates privacy risks compared to traditional methods in dynamic VEC environments. Yishan Chen 0001, Wenshuo Dai, Junxiao Han, Miaojiang Chen, Zhiquan Liu 0001, Ahmed Farouk |
IEEE Internet Things J. | 4 |
| 2025 | ProvBench: A Benchmark of Legal Provision Recommendation for Contract Auto-ReviewingabstractXiuxuan Shen, Zhongyuan Jiang, Junsan Zhang, Junxiao Han, Yao Wan, Chengjie Guo, Bingcheng Liu, Jie Wu, Renxiang Li, Philip S. Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiuxuan Shen, Zhongyuan Jiang, Junsan Zhang, Junxiao Han, Yao Wan 0001, Chengjie Guo, Bingcheng Liu, Renxiang Li, Philip S. Yu |
ACL (1) | 4 |
| 2025 | Detecting Intent Drift in Continuous Conversation via Temporal Transition AccumulationabstractAs large language models (LLMs)-driven conversational systems have advanced, users have become accustomed to engaging in standalone, continuous interactions. In such interactions, users may abruptly change their intent across turns. However, most existing intent detection methods focus on accurately extracting intents and slots from individual utterances, without considering the broader conversational dynamics. This makes them ill-equipped to handle long, evolving conversations where arbitrary intent drift can occur. To address this challenge, we define the intent drift detection task in continuous conversations. We then propose a differentiable method, termed DriftHunter, that enables neural networks to understand how user intent shifts as the conversation progresses via dynamically accumulating the temporal transition across turns. Unlike existing methods, our proposed method incrementally captures global and local transition patterns between intents and slots without relying on prior statistical results. Moreover, our model sequentially accumulates transition patterns across conversation turns. This allows it to learn temporal accumulated dynamics, enabling neural network models to better focus on the most trending user intents during continuous interaction. Experimental evaluations on real-world datasets demonstrate that the proposed method outperforms state-of-the-art baselines in both intent drift detection, intent identification, and slot-filling downstream tasks. Our case study analysis reveals that the learned temporal transition patterns explain the predicted intent drifts.11The source code and dataset are available at https://github.com/FDHTJ/DriftHunter Yue Wang 0014, Dehang Fu, Junxiao Han, Yao Wan 0001, Lixin Cui, Lu Bai 0001, Philip S. Yu |
ICDM | 4 |
| 2025 | RAG4GFM: Bridging Knowledge Gaps in Graph Foundation Models through Graph Retrieval Augmented GenerationabstractGraph Foundation Models (GFMs) have demonstrated remarkable potential across graph learning tasks but face significant challenges in knowledge updating and reasoning faithfulness. To address these issues, we introduce the Retrieval-Augmented Generation (RAG) paradigm for GFMs, which leverages graph knowledge retrieval. We propose RAG4GFM, an end-to-end framework that seamlessly integrates multi-level graph indexing, task-aware retrieval, and graph fusion enhancement.
RAG4GFM implements a hierarchical graph indexing architecture, enabling multi-granular graph indexing while achieving efficient logarithmic-time retrieval. The task-aware retriever implements adaptive retrieval strategies for node, edge, and graph-level tasks to surface structurally and semantically relevant evidence.
The graph fusion enhancement module fuses retrieved graph features with query features and augments the topology with sparse adjacency links that preserve structural and semantic proximity, yielding a fused graph for GFM inference.
Extensive experiments conducted across diverse GFM applications demonstrate that RAG4GFM significantly enhances both the efficiency of knowledge updating and reasoning faithfulness\footnote{Code: \url{https://github.com/Matrixmax/RAG4GFM}.}. Xingliang Wang, Junxiao Han, Shuiguang Deng |
NeurIPS | 3 |
| 2025 | HeSQLNet: A Heterogeneous graph neural network for SQL-to-Text generation
Junsan Zhang, Ao Lu, Junxiao Han, Yudie Yan, Juncai Guo 0003, Yao Wan 0001 |
Inf. Softw. Technol. | 3 |
| 2025 | Model-Oriented Training With Two-Stage Hierarchical Knowledge Distillation Under Non-IID Conditions in Federated Edge-Cloud CollaborationabstractWith the continuous rolling-out of wireless edge cloud networks, Federated Learning (FL) has emerged as a promising solution for decentralized model training without exposing raw data. However, conventional centralized FL faces several limitations in resource-constrained mobile environments, including limited privacy-preserving capabilities and substantial communication overhead, which can lead to privacy leakage. Moreover, in non-independent and identically distributed (Non IID) data environments, FL faces the critical challenge of “client drift”, which leads to performance degradation. To address these challenges, this paper proposes TWHFL, a two-stage hierarchical knowledge distillation framework for Non-IID federated learning, designed to enhance terminal privacy protection and improve model personalization under heterogeneous data distributions. Specifically, in the cloud-edge collaboration stage, edge servers generate pseudo “hard samples” for all sub-MEC centers by optimizing noise inputs guided by feature distribution statistics (e.g., batch normalization running means and variances). To alleviate label distribution skew, both the label proportions and the volume of pseudo data are dynamically adapted based on the real-time operational state of each sub-MEC center. In the edge-terminal collaboration stage, each sub-MEC center conducts localized training using both real and synthetic data without external communication, thereby significantly reducing the risk of privacy leakage. Furthermore, a joint optimization problem is formulated to determine optimal configurations of pruning rates, CPU frequencies, up-link power, and bandwidth allocation, while jointly considering constraints on convergence rate, energy consumption, and latency. Experimental results show that the proposed TWHFL framework can effectively balance privacy protection and model performance in Non-IID settings. Yishan Chen 0001, Wenshuo Dai, Junxiao Han, Zhen Qin 0004, Shuiguang Deng |
IEEE Trans. Cloud Comput. | 3 |
| 2024 | Exploring Parameter-Efficient Fine-Tuning of Large Language Model on Automated Program RepairabstractAutomated Program Repair (APR) aims to fix bugs by generating patches. And existing work has demonstrated that "pre-training and fine-tuning" paradigm enables Large Language Models (LLMs) improve fixing capabilities on APR. However, existing work mainly focuses on Full-Model Fine-Tuning (FMFT) for APR and limited research has been conducted on the execution-based evaluation of Parameter-Efficient Fine-Tuning (PEFT) for APR. Comparing to FMFT, PEFT can reduce computing resource consumption without compromising performance and has been widely adopted to other software engineering tasks. Guochang Li, Chen Zhi, Junxiao Han, Shuiguang Deng |
ASE | 4 |
| 2024 | Sustainability Forecasting for Deep Learning PackagesabstractDeep Learning (DL) technologies have been widely adopted to tackle various tasks. In this process, through software dependencies, a multi-layer DL supply chain (SC) is formed, with DL frameworks acting as the root, DL packages acting as the bridge nodes, and downstream DL projects acting as the periphery. However, most Open Source Software (OSS) projects may fail. Considering the crucial position of DL packages in the DL SC, to foster the sustainable development of DL SCs and DL packages, we aim to forecast the long-term sustainability of DL packages. Here, sustained activity is adopted as the main proxy of sustainability, and the sustainability status is classified as “sus-tainable” or “dormant”. Relatedly, a DL package is considered as “sustainable” if it has sustained activity in its last 12 months. Otherwise, it is deemed as “dormant”. To this end, we propose an approach that begins with obtaining longitudinal features for each DL package in each month. Then, we develop a model to forecast the sustainability of DL packages by incorporating the longitudinal features, which can aptly predict sustainability with an accuracy of up to 0.81. Subsequently, an interpretable module is developed to interpret the determinants (i.e., important features) that impact the sustainability of DL packages. Finally, we generate sustainability trajectories for each DL package to better understand the monthly changes of their sustainability status. Our findings uncover that for most DL packages, fewer but more centralized developers and a balanced collaboration are more likely to help sustain the DL packages. Furthermore, although some DL packages are sustainable, their sustainability trajectories present statistically decreasing trends over time. Based on the findings, we shed light on the dynamic sustainability of DL packages, highlight future research directions, and provide practical suggestions to DL package maintainers, developers, users, and software engineering researchers. Junxiao Han, Yunkun Wang, Zhongxin Liu 0002, Lingfeng Bao, David Lo 0001, Shuiguang Deng |
SANER | 1 |
| 2024 | A Game-Theoretic Approach-Based Task Offloading and Resource Pricing Method for Idle Vehicle Devices Assisted VECabstractVehicle Edge Computing (VEC), as an emerging computing paradigm, aims to achieve the high efficiencies and quality of service by distributing computation tasks to vehicles and cloud-edge servers. The resource pricing problem focuses on how to reasonably price the resources of VEC to encourage their allocation and utilization. However, VEC server overloading may lead to performance degradation, especially in urban congested areas. Meanwhile, idle resources near VEC roads, such as parked vehicles and RSUs, are underutilized and can provide additional computation and communication resources to the system. Inspired by this, this paper introduces a model to assist vehicle edge computing by attracting Idle Vehicles (IVs) to share resources. We use a two-stage Stackelberg game model to address the resource pricing and task offloading problem, analyzing the interaction between requesting vehicles and cloud-edge servers. Through a backward induction method, we transform the problem into a convex optimization problem and theoretically prove the existence of a unique Nash equilibrium. In the first stage, optimal offloading ratio strategy is solved using convex optimization theory. In the second stage, the original problem is decomposed into 2N sub-problems and solved using the Lagrangian dual method and Karush-Kuhn-Tucker (KKT) conditions for optimal resource pricing. Additionally, a price incentive mechanism and a task-vehicle stable matching game model are employed to recruit idle vehicles around the roads to spontaneously participate in the task offloading process. Finally, simulation results reveal our solution effectively reduces offloading costs, latency, energy use, and enhances task completion compared to others. Yishan Chen 0001, Junxiao Han, Hailiang Zhao, Shuiguang Deng |
IEEE Internet Things J. | 3 |
| 2024 | On the sustainability of deep learning projects: Maintainers' perspectiveabstractAbstract Deep learning (DL) techniques have grown in leaps and bounds in both academia and industry over the past few years. Despite the growth of DL projects, there has been little study on how DL projects evolve, whether maintainers in this domain encounter a dramatic increase in workload and whether or not existing maintainers can guarantee the sustained development of projects. To address this gap, we perform an empirical study to investigate the sustainability of DL projects, understand maintainers' workloads and workloads growth in DL projects, and compare them with traditional open‐source software (OSS) projects. In this regard, we first investigate how DL projects grow, then, understand maintainers' workload in DL projects, and explore the workload growth of maintainers as DL projects evolve. After that, we mine the relationships between maintainers' activities and the sustainability of DL projects. Eventually, we compare it with traditional OSS projects. Our study unveils that although DL projects show increasing trends in most activities, maintainers' workloads present a decreasing trend. Meanwhile, the proportion of workload maintainers conducted in DL projects is significantly lower than in traditional OSS projects. Moreover, there are positive and moderate correlations between the sustainability of DL projects and the number of maintainers' releases, pushes, and merged pull requests. Our findings shed lights that help understand maintainers' workload and growth trends in DL and traditional OSS projects and also highlight actionable directions for organizations, maintainers, and researchers. Junxiao Han, David Lo 0001, Chen Zhi, Yishan Chen 0001, Shuiguang Deng |
J. Softw. Evol. Process. | 1 |
| 2024 | Understanding Newcomers' Onboarding Process in Deep Learning ProjectsabstractAttracting and retaining newcomers are critical for the sustainable development of Open Source Software (OSS) projects. Considerable efforts have been made to help newcomers identify and overcome barriers in the onboarding process. However, fewer studies focus on newcomers’ activities before their successful onboarding. Given the rising popularity of deep learning (DL) techniques, we wonder what the onboarding process of DL newcomers is, and if there exist commonalities or differences in the onboarding process for DL and non-DL newcomers. Therefore, we reported a study to understand the growth trends of DL and non-DL newcomers, mine DL and non-DL newcomers’ activities before their successful onboarding (i.e., past activities), and explore the relationships between newcomers’ past activities and their first commit patterns and retention rates. By analyzing 20 DL projects with 9,191 contributors and 20 non-DL projects with 9,839 contributors, and conducting email surveys with contributors, we derived the following findings: 1) DL projects have attracted and retained more newcomers than non-DL projects. 2) Compared to non-DL newcomers, DL newcomers encounter more deployment, documentation, and version issues before their successful onboarding. 3) DL newcomers statistically require more time to successfully onboard compared to non-DL newcomers, and DL newcomers with more past activities (e.g., issues, issue comments, and watch) are prone to submit an intensive first commit (i.e., a commit with many source code and documentation files being modified). Based on the findings, we shed light on the onboarding process for DL and non-DL newcomers, highlight future research directions, and provide practical suggestions to newcomers, researchers, and projects. Junxiao Han, David Lo 0001, Xin Xia 0001, Shuiguang Deng, Minghui Wu 0001 |
IEEE Trans. Software Eng. | 1 |
| 2022 | Towards automatic detection and prioritization of pre-logging overhead: a case study of hadoop ecosystem
Chen Zhi, Shuiguang Deng, Junxiao Han, Jianwei Yin |
Autom. Softw. Eng. | 3 |
| 2020 | A Preliminary Study on Sensitive Information Exposure Through LoggingabstractLogging is a common practice to collect valuable runtime information about software systems. However, information written to log files can be sensitive and give valuable guidance to attackers. In fact, information exposure through logging is not uncommon. Even large-scale online services (e.g., Facebook and Twitter) have reported exposing sensitive information via log files, and hundreds of millions of users are affected. Despite the severity of such vulnerabilities, there is no existing work that studies such vulnerabilities in the real-world context, and we have little knowledge about them. To fill this gap, we conduct a preliminary study on 413 real-world vulnerabilities to investigate the exploitability and root causes of such vulnerabilities. By analyzing these vulnerabilities, we find that 1) about two-third (67.8%) vulnerabilities can be exploited via the network, and a significant amount (89.3%) of vulnerabilities can be exploited with low efforts; 2) malicious users and insiders can use about half of (46.9%) proof-of-concept exploits to launch attacks without any expertise; 3) the top three common root causes for the vulnerabilities are insecure whole-object logging (43.4%), incorrect permission assignment (17.5%), and improper implementation of sanitization (11.2%). Based on the findings, we also discuss the implications for researchers and practitioners. We believe our work can inspire further work on detecting and fixing the vulnerabilities. Chen Zhi, Jianwei Yin, Junxiao Han, Shuiguang Deng |
APSEC | 3 |
| 2020 | An Empirical Study of the Dependency Networks of Deep Learning LibrariesabstractDeep Learning techniques have been prevalent in various domains, and more and more open source projects in GitHub rely on deep learning libraries to implement their algorithms. To that end, they should always keep pace with the latest versions of deep learning libraries to make the best use of deep learning libraries. Aptly managing the versions of deep learning libraries can help projects avoid crashes or security issues caused by deep learning libraries. Unfortunately, very few studies have been done on the dependency networks of deep learning libraries. In this paper, we take the first step to perform an exploratory study on the dependency networks of deep learning libraries, namely, Tensorflow, PyTorch, and Theano. We study the project purposes, application domains, dependency degrees, update behaviors and reasons as well as version distributions of deep learning projects that depend on Tensorflow, PyTorch, and Theano. Our study unveils some commonalities in various aspects (e.g., purposes, application domains, dependency degrees) of deep learning libraries and reveals some discrepancies as for the update behaviors, update reasons, and the version distributions. Our findings highlight some directions for researchers and also provide suggestions for deep learning developers and users. Junxiao Han, Shuiguang Deng, David Lo 0001, Chen Zhi, Jianwei Yin, Xin Xia 0001 |
ICSME | 1 |
| 2020 | What do Programmers Discuss about Deep Learning Frameworks
Junxiao Han, Emad Shihab, Zhiyuan Wan, Shuiguang Deng, Xin Xia 0001 |
Empir. Softw. Eng. | 1 |
| 2019 | Characterization and Prediction of Popular Projects on GitHubabstractGitHub is a large and popular open source project platform, which hosts various open source projects. Despite the prevalence of GitHub platform, not every project has gained high popularity. Identification of popular projects on GitHub can help developers choose proper projects to follow or contribute to, as well as provide guidance in building a popular project. In this paper, we propose an approach to predict the popularity of GitHub projects. We first conducted online surveys with GitHub users to determine the threshold (the number of stars of a project) of popular and unpopular projects. Next, we extract 35 features from both GitHub and Stack Overflow, which are divided into three dimensions: project, evolutionary, and project owner. A random forest classifier is built based on these features to identify popular GitHub projects. To evaluate the performance of our approach, we collect a large-scale dataset from GitHub which contains a total of 409,784 GitHub projects and 174,784 GitHub users. Our model achieves an average AUC of 0.76, which statistically significantly improves state-of-the-art by a substantial margin. We also study which features are of the most importance in distinguishing popular projects from unpopular ones. Experimental results show that number of branches, number of open issues, and number of contributors play the most important roles in identification of popular projects, and all of them have large effect size. Junxiao Han, Shuiguang Deng, Xin Xia 0001, Dongjing Wang, Jianwei Yin |
COMPSAC (1) | 1 |