Linghao Zhang

dblp:49/10934 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 ComLQ: Benchmarking Complex Logical Queries in Information Retrieval
abstract
Information retrieval (IR) systems play a critical role in navigating information overload across various applications. Existing IR benchmarks primarily focus on simple queries that are semantically analogous to single- and multi-hop relations, overlooking complex logical queries involving first-order logic operations such as conjunction (∧), disjunction (∨), and negation (¬). Thus, these benchmarks can not be used to sufficiently evaluate the performance of IR models on complex queries in real-world scenarios. To address this problem, we propose a novel method leveraging large language models (LLMs) to construct a new IR dataset ComLQ for Complex Logical Queries, which comprises 2,909 queries and 11,251 candidate passages. A key challenge in constructing the dataset lies in capturing the underlying logical structures within unstructured text. Therefore, by designing the subgraph-guided prompt with the subgraph indicator, an LLM (such as GPT-4o) is guided to generate queries with specific logical structures based on selected passages. All query-passage pairs in ComLQ are ensured structure conformity and evidence distribution through expert annotation. To better evaluate whether retrievers can handle queries with negation, we further propose a new evaluation metric, Log-Scaled Negation Consistency (LSNC@K). As a supplement to standard relevance-based metrics (such as nDCG and mAP), LSNC@K measures whether top-K retrieved passages violate negation conditions in queries. Our experimental results under zero-shot settings demonstrate existing retrieval models' limited performance on complex logical queries, especially on queries with negation, exposing their inferior capabilities of modeling exclusion. In summary, our ComLQ offers a comprehensive and fine-grained exploration, paving the way for future research on complex logical queries in IR.
Ganlin Xu, Zhitao Yin, Linghao Zhang, Jiaqing Liang, Weijia Lu, Zhifei Yang 0005, Sihang Jiang 0001, Deqing Yang
AAAI3
2026 CoRaCMG: Contextual retrieval-augmented framework for commit message generation
Linghao Zhang, Zongen Ren, Chong Wang 0004, Peng Liang 0001
Inf. Softw. Technol.2
2026 Convergence Analysis and Resource Allocation for Hierarchical Split Federated Learning Over Space-Air-Ground Integrated Networks
abstract
Federated Learning (FL) confronts challenges such as resource constraints and unbalanced data distribution in the Space-Air-Ground Integrated Network (SAGIN). This paper proposes a Hierarchical Split Federated Learning (HSFL) framework considering satellite handoff and derives its upper bound of loss function affected by model splitting and data distribution. To minimize the weighted sum of training loss and latency, we formulate a joint optimization problem that integrates device association, model split layer selection, and resource allocation. We decompose the original problem into several subproblems, where an iterative optimization algorithm incorporating closed-form solutions and brute-force split point search is proposed. Simulation results demonstrate that the proposed algorithm can balance training efficiency and model accuracy for FL in SAGIN.
Haitao Zhao 0004, Bo Xu 0020, Jinlong Sun, Linghao Zhang
IEEE Signal Process. Lett.5
2026 Time Updatable Policy-Based Chameleon Hash for Traceable and Accountable Redactable Blockchain
abstract
Ateniese et al. (EuroS&P 2017) proposed the notion of redactable blockchains (RBs), in which a designated party uses a secret key to modify blockchain history without causing a hard fork. Nevertheless, redactions may be performed mistakenly or maliciously due to misbehavior or operational errors. From a regulatory perspective, any RB design must therefore incorporate accountability and traceability mechanisms to ensure that redactions are non-abusive and publicly verifiable. As a countermeasure, we propose the notion of time-updatable policy-based chameleon hash (TPCH). This construction addresses regulatory concerns by enabling publicly verifiable proofs of redaction and traceable user identities. Our basic building block, termed time-updatable chameleon hash (TUCH), provides redaction accountability through an intrinsic property formalized as Type-2 Trapdoor Collisions. TUCH is functionally versatile and achieves acceptably efficient performance compared to peer chameleon hash schemes. Following the heuristics of Camenisch et al. (PKC 2017) and Derler et al. (NDSS 2019), we further extend TUCH by integrating attribute-based encryption (ABE) to obtain a time-updatable, policy-based variant, namely TPCH. The resulting scheme overcomes the limitations of coarse-grained redaction and the impracticality of specifying the exact modifier in advance. Overall, TPCH provides a secure, efficient, and comprehensive solution for accountable and traceable redactable blockchains under practical regulatory requirements. Our systematic analysis further demonstrates the suitability of TPCH for small scale deployment.
Ke Huang 0002, Xiong Li 0002, Fatemeh Rezaeibagha, Linghao Zhang, Xiaosong Zhang 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Solving the Min-Max Multiple Traveling Salesmen Problem via Learning-Based Path Generation and Optimal Splitting
abstract
This study addresses the Min-Max Multiple Traveling Salesmen Problem (m3-TSP), which aims to coordinate tours for multiple salesmen such that the length of the longest tour is minimized. Due to its NP-hard nature, exact solvers become impractical under the assumption that P ≠ NP. As a result, learning-based approaches have gained traction for their ability to rapidly generate high-quality approximate solutions. Among these, two-stage methods combine learning-based components with classical solvers, simplifying the learning objective. However, this decoupling often disrupts consistent optimization, potentially degrading solution quality. To address this issue, we propose a novel two-stage framework named Generate-and-Split (GaS), which integrates reinforcement learning (RL) with an optimal splitting algorithm in a joint training process. The splitting algorithm offers near-linear scalability with respect to the number of cities and guarantees optimal splitting in Euclidean space for any given path. To facilitate the joint optimization of the RL component with the algorithm, we adopt an LSTM-enhanced model architecture to address partial observability. Extensive experiments show that the proposed GaS framework significantly outperforms existing learning-based approaches in both solution quality and transferability.
Xiangchen Wu, Liang Wang 0006, Hao Hu 0001, XianPing Tao, Linghao Zhang
ECAI6
2025 Contextual Code Retrieval for Commit Message Generation: A Preliminary Study
abstract
Background: A commit message describes the main code changes in a commit and plays a crucial role in software maintenance. Existing commit message generation (CMG) approaches typically frame it as a direct mapping which inputs a code diff and produces a brief descriptive sentence as output. Aims: Since the raw code diff lacks the context related to the code itself, we intend to supplement the relevant code as input of CMG to generate high-quality and informative commit messages. Method: We propose a contextual code retrieval-based method called C3Gen to enhance CMG by retrieving commit-relevant code snippets from the repository and incorporating them into the model input to provide richer contextual information at the repository scope. In the experiments, we evaluated the effectiveness of C3Gen across various models using four objective and three subjective metrics. Meanwhile, we design and conduct a human evaluation to investigate how C3Gen-generated commit messages are perceived by human developers. Results & Conclusions: By incorporating contextual code into the input, C3Gen enables models to leverage additional information to generate more comprehensive and informative commit messages with greater practical value in real-world development scenarios. Further analysis underscores concerns about the reliability of similarity-based metrics and provides empirical insights for CMG.
Linghao Zhang, Chong Wang 0004, Peng Liang 0001
ESEM2
2025 RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
abstract
Multimodal Large Language Models (MLLMs) increasingly excel at perception,understanding, and reasoning. However, current benchmarks inadequately evaluate their ability to perform these tasks continuously in dynamic, real-world environments. To bridge this gap, we introduce RT V-Bench, a fine-grained benchmark for MLLM real-time video analysis. RTV-Bench includes three key principles: (1) Multi-Timestamp Question Answering (MTQA), where answers evolve with scene changes; (2) Hierarchical Question Structure, combining basic and advanced queries; and (3) Multi-dimensional Evaluation, assessing the ability of continuous perception, understanding, and reasoning. RTV-Bench contains 552 diverse videos (167.2 hours) and 4,631 high-quality QA pairs. We evaluated leading MLLMs, including proprietary (GPT-4o, Gemini 2.0), open-source offline (Qwen2.5-VL, VideoLLaMA3), and open-source real-time (VITA-1.5, InternLM-XComposer2.5-OmniLive) models. Experiment results show open-source real-time models largely outperform offline ones but still trail top proprietary models. Our analysis also reveals that larger model size or higher frame sampling rates do not significantly boost RTV-Bench performance, sometimes causing slight decreases.This underscores the need for better model architectures optimized for video stream processing and long sequences to advance real-time video analysis with MLLMs.
Shuhang Xun, Sicheng Tao, Jungang Li, Yibo Shi, Zhixin Lin, Zhanhui Zhu, Hanqian Li, Linghao Zhang, Shikang Wang, Hanbo Zhang, Xuming Hu
NeurIPS9
2025 SWE-bench Goes Live!
abstract
The issue-resolving task, where a model generates patches to fix real-world bugs, has emerged as a key benchmark for evaluating the capabilities of large language models (LLMs). While SWE-bench has become the dominant benchmark in this domain, it suffers from several limitations: it has not been updated since its release, is restricted to only 12 repositories, and relies heavily on manual effort for constructing test instances and setting up executable environments, significantly limiting its scalability. We present SWE-bench-Live, a live-updatable benchmark designed to address these limitations. SWE-bench-Live currently includes 1,890 tasks derived from real GitHub issues created since 2024, spanning 223 repositories. Each task is accompanied by a dedicated Docker image to ensure reproducible execution. Additionally, we introduce an automated curation pipeline that streamlines the entire process from instance creation to environment setup, removing manual bottlenecks and enabling scalability and continuous updates. We evaluate a range of state-of-the-art models and agent frameworks on SWE-bench-Live, offering detailed empirical insights into their real-world bug-fixing capabilities. By providing a fresh, diverse, and executable benchmark grounded in live repository activity, SWE-bench-Live supports reliable, large-scale assessment of code LLMs and code agents in realistic development settings.
Linghao Zhang, Shilin He, Chaoyun Zhang, Yu Kang 0006, Bowen Li 0002, Chengxing Xie, Maoquan Wang, Yufan Huang, Shengyu Fu, Elsie Nallipogu, Qingwei Lin, Yingnong Dang, Saravan Rajmohan, Dongmei Zhang 0001
NeurIPS1
2025 Submodular Optimization Based Co-Inference in Space-Air-Ground Integrated Vehicular Networks
abstract
Space-air-ground integrated vehicular networks (SAGVN) play a crucial role in the 6G system, offering global coverage and ultra-wide-area broadband access. Meanwhile, advancements in artificial intelligence (AI) have led to a significant increase in model inference demands, which come with stringent latency requirements. In this paper, considering the intelligent services in SAGVN, we explore the co-inference problem and improve the inference efficiency by performing model splitting, vehicle association, and resource allocation. This problem is solved iteratively by decomposing it into multiple sub-problems. In particular, the model splitting problem is addressed by systematically searching for the optimal split points. Besides, the joint vehicle association and resource allocation problem are reformulated into a monotone submodular function without satellites, and then a low-complexity submodular optimization algorithm is proposed. To further utilize the computing power of the satellite, we further introduce a resource reallocation algorithm based on the existing optimization results. These two subproblems can iterate alternately until the optimization goal converges. Simulation results show that the proposed algorithm can achieve better latency performance in several inference tasks.
Suyao Huang, Bo Xu 0020, Guijin Tang, Haotong Cao, Linghao Zhang, Haitao Zhao 0004
PIMRC5
2025 Mobility-Aware Task Offloading in Industrial Fog Networks: A Submodular-Based MARL Approach
abstract
The development of Industrial Internet of Things (IIoT) applications presents a critical challenge in terms of latency limitation, particularly considering the limited availability of resources that prevent a single fog device from fully executing large-scale computing tasks. In such scenarios, enabling distributed computing across multiple fog servers or collaborating with cloud servers holds promising potential. To improve the efficiency of task offloading while accounting for the crucial role of movable fog devices (e.g., robots and unmanned cars), we formulate a joint optimization problem as a partially observable Markov decision process (POMDP), incorporating offloading decisions, computing resource allocation, and trajectory optimization under constraints related to available resources and collision avoidance. Due to the nondeterministic polynomial-time hardness (NP-hardness) in the problems of task offloading and resource allocation, we reformulate a matroid-constrained submodular maximization problem and propose an iterative low-complexity algorithm to find solutions. Subsequently, extracting better solutions from submodular optimization, we propose a multiagent reinforcement learning (MARL)-based algorithm to solve the trajectory optimization problem for the movable fog devices acting as agents, making decisions based on their local observations. Finally, simulation results have validated that the proposed scheme has a superior performance compared to the baselines.
Bo Xu 0020, Haitao Zhao 0004, Haotong Cao, Jinlong Sun, Linghao Zhang, Hongbo Zhu 0002
IEEE Internet Things J.6
2024 Using Large Language Models for Commit Message Generation: A Preliminary Study
abstract
A commit message is a textual description of the code changes in a commit, which is a key part of the Git version control system (VCS). It captures the essence of software updating. Therefore, it can help developers understand code evolution and facilitate efficient collaboration between developers. However, it is time-consuming and labor-intensive to write good and valuable commit messages. Some researchers have conducted extensive studies on the automatic generation of commit messages and proposed several methods for this purpose, such as generation-based and retrieval-based models. However, seldom studies explored whether large language models (LLMs) can be used to generate commit messages automatically and effectively. To this end, this paper designed and conducted a series of experiments to comprehensively evaluate the performance of popular open-source and closed-source LLMs, i.e., Llama 2 and ChatGPT, in commit message generation. The results indicate that considering the BLEU and Rouge-L metrics, LLMs surpass the existing methods in certain indicators but lag behind in others. After human evaluations, however, LLMs show a distinct advantage over all these existing methods. Especially, in 78 % of the 366 samples, the commit messages generated by LLMs were evaluated by humans as the best. This work not only reveals the promising potential of using LLMs to generate commit messages, but also explores the limitations of commonly used metrics in evaluating the quality of auto-generated commit messages.
Linghao Zhang, Jingshu Zhao, Chong Wang 0004, Peng Liang 0001
SANER1
2024 Fault diagnosis of power equipment based on variational autoencoder and semi-supervised learning
abstract
Summary The issue of fault diagnosis in power equipment is receiving increasing attention from scholars. Due to the important role played by bearings in power equipment, bearing faults have become the main factor causing the shutdown of wind turbines units. Therefore, this paper takes bearing equipment as an example for research. In order to solve the problem of insufficient and unbalanced fault sample data of wind turbines bearings, a fault diagnosis (FD) method based on variational autoencoder and semi‐supervised learning is proposed in this paper. Firstly, based on Label Propagation‐random forests (LP‐RFs) and a small number of labeled fault samples, a semi‐supervised learning algorithm is proposed to label the original data samples. Secondly, a small number of training samples are preprocessed by the variational autoencoder to reduce the imbalance of the fault samples. Then, the RFs‐based method is adopted to train the processed fault samples to obtain a mature FD classifier. Finally, the proposed method is applied to FD for bearings, and the results show that the proposed method can realize bearings fault diagnosis (BFD). And meanwhile, the proposed method can also be applied for fault diagnosis in power transmission and transformation systems.
Linghao Zhang, Zhengwei Chang, Sayina Bodanbai
Concurr. Comput. Pract. Exp.3
2024 FSD-CLCD: Functional semantic distillation graph learning for cross-language code clone detection
Linghao Zhang, Senlin Luo, Limin Pan, Zhouting Wu, Kun Gong
Eng. Appl. Artif. Intell.1
2024 LogETA: Time-aware cross-system log-based anomaly detection with inter-class boundary optimization
Kun Gong, Senlin Luo, Limin Pan, Linghao Zhang
Future Gener. Comput. Syst.4
2024 A Privacy-Preserving Matching Service Scheme for Power Data Trading
abstract
Currently, power data trading typically relies on Web pages as the conventional mode of mediation. Nevertheless, dishonest trading Web may secretly resell the data sets of grid companies or have no way of knowing what the buyer has done with the power data, thereby compromising the privacy of power user. This article proposes a privacy-preserving supply-demand consistency matching service scheme (PPMSE) to address the problem of whether the power data provided by the seller aligns with the requirements of the buyer in power data trading. The scheme utilizes enhanced public-key searchable encryption (PKSE) to establish a matching environment that fulfills privacy protection needs, thereby facilitating consistency matching between supply and demand, all while preserving user privacy. Then, the PPMSE ensures that matching service can only occur within the designated platform by equipping the power data trading cloud platform with public and private keys. Additionally, by applying ciphertext policy attribute-based encryption (CP-ABE) to the data processing tasks of the buyer, the scheme enables the seller to decrypt and obtain what the buyer has done with the data after successful matching and meeting specific attributes. Ultimately, a comprehensive analysis and performance evaluation are provided, validating the feasibility and superiority of the proposed scheme.
Zewei Liu 0001, Chunqiang Hu, Conghao Ruan, Linghao Zhang, Pengfei Hu 0001, Tao Xiang 0001
IEEE Internet Things J.4
2024 Grouped federated learning for time-sensitive tasks in industrial IoTs
Jiangshan Hao, Linghao Zhang, Yanchao Zhao
Peer Peer Netw. Appl.2
2024 Towards Effective Long-Term Wind Power Forecasting: A Deep Conditional Generative Spatio-Temporal Approach
abstract
Accurately forecasting long-term future wind power is critical to achieve safe power grid integration. This problem is quite challenging due to wind power's high volatility and randomness. In this paper, we propose a novel time series forecasting method, namely Deep Conditional Generative Spatio-Temporal model (DCGST), and its high accuracy is achieved by tackling two critical issues simultaneously: a proper handling of the non-stationarity of multiple wind power time series, and a fine-grained modeling of their complicated yet dynamic spatio-temporal dependencies. Specifically, we first formally define theSpatio-Temporal Concept Drift(STCD) problem of wind power, and then we propose a novel deep conditional generative model to learn probabilistic distributions of future wind power values under STCD. Three different tailored neural networks are designed for distributions parameterization, including a graph-based prior network, an attention-based recognition network, and a stochastic seq2seq-based generation network. They are able to encode the dynamic spatio-temporal dependencies of multiple wind power time series and infer one-to-many mappings for future wind power generation. Compared to existing methods, DCGST can learn better spatio-temporal representations of wind power data and learn better uncertainties of data distribution to generate future values. Comprehensive experiments on real-world datasets including the largest public turbine-level wind power dataset verify the effectiveness, efficiency, generality and scalability of our method.
Peiyu Yi, Zhifeng Bao, Feihu Huang 0002, Jince Wang, Jian Peng 0002, Linghao Zhang
IEEE Trans. Knowl. Data Eng.6
2023 EPPSQ: Achieving efficient and privacy-preserving statistics queries over encrypted data in smart grids
Beibei Li 0002, Linghao Zhang, Zhengwei Chang, Liang Zhao 0020, Arun Kumar 0006
Future Gener. Comput. Syst.3
2023 Feature extraction method of HPLC communication signal based on genetic algorithm
abstract
Abstract Communication quality is a key factor affecting the effectiveness of distribution network operation status management. Therefore, it is required that the communication performance of the distribution network management system be good, and the communication signal can directly reflect the communication quality. Therefore, a genetic algorithm (GA) based HPLC communication signal feature extraction method is proposed. The process of constructing reference samples, training recognition models, constructing noise samples, self‐coding denoising processing, and denoising sample recognition has completed the identification of power line channel transmission characteristics. The state feature quantity of the topological line is based on the positioning results of the diagnostic function, and the calculated value of the reflection coefficient calculation model at the head end of the topological line is compared with the measured value. By optimizing the feature quantities containing topological line states through GAs, abnormal feature quantities close to the true state of topological lines are obtained. Experimental analysis shows that this method can correctly identify the position of reactive power compensation in the transmission line, and can also correctly identify the extraction of topological anomaly features. Therefore, the research method is of great significance in improving the safe operation of power systems and extending the service life of topology lines.
Zhengwei Chang, Huihui Liang, Linghao Zhang
IET Commun.4
2020 A mixed attributes oriented dynamic SOM fuzzy cluster algorithm for mobile user classification
Guangxia Xu, Linghao Zhang
Inf. Sci.2
2014 Low-disruptive dynamic updating of Java applications
Tianxiao Gu, Chun Cao, Chang Xu 0001, Xiaoxing Ma, Linghao Zhang, Jian Lu 0001
Inf. Softw. Technol.5
2013 Challenges in developing software for cyber-physical systems
abstract
Cyber-physical systems are systems that integrate the digital computational world with the real physical world, often using sensors and actuators as interfaces. There exist many application domains of cyber-physical systems such as autonomous systems, process control systems, robotic systems, and context-aware systems. The physical world is a complex and continuous world that changes in real-time while the computational world is a simplified and discrete world that often stores a delayed, likely inaccurate image of the physical world using sensory data. The mismatch between these two worlds poses unique challenges of developing software for cyber-physical systems.
Linghao Zhang, Xiaoxing Ma, Chang Xu 0001, Jian Lu 0001
Internetware1
2012 Javelus: A Low Disruptive Approach to Dynamic Software Updates
abstract
Practical software systems are subject to frequent updates for fixing their bugs or addressing new requirements. Updating a software system without stopping and restarting it is desired, as this helps reduce the redeployment cost as well as achieving the high availability. Existing techniques for dynamically updating Java programs may introduce noticeable pauses during which these programs are unable to function. We in this paper present Javelus, a dynamic Java update system with greatly reduced pausing time but without sacrificing update flexibility and system efficiency. Different from previous approaches, Javelus uses a lazy update mechanism with which an object-to-update will not be updated until it is really used. We implemented Javelus on top of an industry-strength OpenJDK HotSpot VM. We evaluated Javelus with real updates to Tomcat 7 and the same micro array benchmark used in evaluating Jvolve and DCE VM. The experiments report promising results that Javelus only incurred a pausing time two orders of magnitude smaller than those of Jvolve and DCE VM.
Tianxiao Gu, Chun Cao, Chang Xu 0001, Xiaoxing Ma, Linghao Zhang, Jian Lu 0001
APSEC5
2012 Resynchronizing Model-Based Self-Adaptive Systems with Environments
abstract
Self-adaptive systems are attractive due to their ability of adapting to changeable environments automatically. However, such systems may be subject to runtime failures when all environmental dynamics cannot be adequately considered at design time. When such failures occur at runtime, a system's internal adaptation logic usually has become inconsistent with its environment, according to our observation. We call this inconsistency sync-loss error. From our project experiences, we empirically identified a strong correlation between sync-loss error and system failure. This motivated us to fix sync-loss error in order to reduce failure for self-adaptive systems. In this paper, we formulate the problem of detecting sync-loss error, and present a framework ReSync to automatically fix sync-loss errors by desynchronizing a system with its environment. We experimentally evaluated ReSync on real robot cars with 20 different system versions. The evaluation reported promising results that ReSync can automatically recover our robot car systems from sync-loss errors, and significantly reduce the failure rate from 90.9% to 11.7-28.8%.
Linghao Zhang, Chang Xu 0001, Xiaoxing Ma, Tianxiao Gu, Xuezhi Hong, Chun Cao, Jian Lu 0001
APSEC1
2012 ConsView: Towards Application-Specific Consistent Context Views
abstract
Detecting and resolving context inconsistency is critical to pervasive computing applications and infrastructures. Context inconsistency occurs when an application perceives contexts that breach predefined consistency constraints. This can drive an application to behave abnormally or even cause failure. Existing work commonly assumes the presence of a single application suffering from context inconsistency, such that specific repair actions can be taken to resolve the inconsistency for this application. However, when multiple applications run on the same infrastructure, they may impose conflicting requirements on resolving context inconsistency. In this paper, we propose a novel view-based approach ConsView to address such conflicting requirements. In ConsView, each application has a specific view to its own contexts that satisfy its own requirement on resolving context inconsistency. Such views are called consistent context views. We discuss the challenges of doing so and our ideas for addressing them. We implemented a prototype infrastructure supporting consistent context views, and evaluated it experimentally with simulated applications of real-life settings. The results confirmed the effectiveness and efficiency of our ConsView approach.
Haibin Yang, Chang Xu 0001, Xiaoxing Ma, Linghao Zhang, Chun Cao, Jian Lu 0001
COMPSAC4