EDBT 2026 Demo / reviewers in the wild / expert
Zheng Li 0026
dblp:10/1143-26
· DBLP profile ↗
19ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0001-7840-8303ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language ModelsabstractLarge Language Models (LLMs) have revolutionized intelligent interactions, enabling mobile applications such as personal assistants on edge devices for local execution. Speculative decoding (SD) has emerged as a promising paradigm to accelerate LLM inference without compromising generation quality, employing a draft-then-verify manner. However, due to the constrained computing and memory resources on edge devices, existing SD works heavily rely on an auxiliary draft model that incurs additional memory burden and hinders the adaptability, as well as static token trees that yield suboptimal inference performance. To this end, we propose DIAA, a Decoding-efficient Inference Acceleration Approach for on-device LLMs. DIAA achieves plug-and-play and model-agnostic inference speedup with memory and computation efficiency for edge devices. Specifically, a pair of lightweight look-up tables (LUTs) is constructed by Top-K token sampling to cache historical tokens and probabilities for rapid candidate drafting. DIAA integrates a dynamic token tree with prior LUTs enabling paralleled verification, updated during decoding process, to adapt the online context. A computation overlap is then employed to pipeline the update operations of token tree, LUTs, and KV cache to improve the computational efficiency. Finally, through extensive experiments implemented on edge platform NVIDIA Jetson, DIAA outperforms existing baselines in generation speed and inference wall-clock time, while incurring minimal memory overhead. Hao Tian 0012, Fuwen Tian, Guangming Cui, Zheng Li 0026, Xuyun Zhang, Quan Z. Sheng, Wan-Chun Dou |
AAAI | 5 |
| 2026 | CoST: An Efficient and Error-Bounded Compression Algorithm for Streaming Trajectories
Zihang Xu, Zheng Li 0026, Huajun He, Yu Zheng 0004 |
WoWMoM | 3 |
| 2025 | A Cost-Aware Approach for Collaborating Large Language Models and Small Language ModelsabstractThe emerging reasoning ability of large language models (LLMs) and accompanying commercial applications offer a promising path for service providers to deploy intelligent agents on their own products through API calls. However, the black-box nature of LLMs has driven providers to try prompt tuning to improve reasoning quality for competitiveness, while the generated reasoning logic results in additional service costs. Although some works have proposed collaborating LLMs and Small Language Models (SLMs) to reduce the frequency of LLM calls, most overlook the actual number of tokens interacting with the LLMs, which results in a potentially high cost still. Furthermore, directly compressing the prompt to reduce tokens often leads to a significant accuracy loss. To address the above challenges, we propose a cost-aware approach for collaborating LLMs and SLMs, named Coco. In our method, a confidence-based task assignment method is designed which leverages the result confidence of SLMs to assess task complexity and determine whether LLM involvement is necessary. For complex tasks, the SLM adapts the input by compressing unnecessary information according to confidence. Considering the potential loss of accuracy, prompt tuning-based reasoning optimization methods are introduced to guide the LLM in generating both the reasoning logic sketch and the final result. Finally, logic alignment is applied to fuse sketches from both models, ensuring the rationality of the reasoning logic. Experimental results on three open-source datasets demonstrate that our approach effectively reduces the cost of API calls to LLMs while ensuring the reasoning accuracy and the reasonableness of generated logic. Zheng Li 0026, Xuyun Zhang, Hao Tian 0012, Wan-Chun Dou |
CIKM | 1 |
| 2025 | A Dynamic Inference Method for Autoregressive Transformer ModelsabstractAutoregressive transformer models have achieved state-of-the-art performance in advanced services such as text generation and machine translation. Given the significant computational bottlenecks of model inference, layer-wise skipping has emerged as a promising method to accelerate inference by bypassing redundant layers. However, existing methods face challenges, including sub-optimal performance resulting from the premature skipping of critical layers and an unbalanced focus on either multi-head attention or feed-forward sub-blocks, ultimately leading to global performance degradation. In light of the above challenges, we propose a Dynamic Inference Method, named DIM, for autoregressive transformer models. DIM dy-namically selects sub-blocks from both multi-head attention and feed-forward networks through the importance score alignment, ensuring a balanced selection that optimizes both efficiency and model performance. To further mitigate the potential performance loss of skipped sub-blocks, a lightweight adjustment is developed to approximate the computations of skipped sub-blocks. Finally, extensive experiments using several benchmarks validate that DIM outperforms existing inference methods. Hao Tian 0012, Zheng Li 0026, Wan-Chun Dou |
ICWS | 4 |
| 2025 | Adaptive Encoding Strategies for Lossless Floating-Point CompressionabstractLossless floating-point time series compression is crucial for a wide range of critical scenarios, such as data transmission in the Internet of Things. In this paper, we propose a lossless floating-point compression method Elf. that employs a set of optimizations for the encodings of significand counts, leading zeros, trailing zeros and sharing conditions. Specifically, we first devise a Huffman-based method for the significand counts along with erasing flags. Then, we develop a dynamic programming algorithm with a set of pruning strategies to efficiently compute the adaptive approximation rules for leading zeros and trailing zeros, respectively. Next, we propose an adaptive sharing condition for the counts of leading zeros and trailing zeros. We further extend Elf. to Streaming Elf., i.e., SElf., which achieves almost the same compression ratio as Elf., while enjoying higher efficiency. We compare Elf. and SElf. with 9 competitors using 14 datasets, demonstrating the powerful performance of both Elf* and SElf*. All the source codes and datasets are publicly released. Zheng Li 0026, Xiaolong Xu 0001, Chao Chen 0004, Tong Liu 0001, Jiaxing Shang, Yu Zheng 0004 |
IEEE Internet Things J. | 1 |
| 2024 | An Inference Acceleration Approach for Boosting DNN Cold Start in Cloud-Edge Computing
Hao Tian 0012, Haolong Xiang, Tingtong Zhu, Siyuan Wu 0002, Zheng Li 0026, Mingxu Jiang, Wan-Chun Dou |
ADMA (1) | 6 |
| 2024 | A Crowdsensing Service Pricing Method in Vehicular Edge ComputingabstractThe rapid advancements in vehicular networking technology have enabled onboard users to access a variety of emerging services like precise navigation and real-time hazard avoidance. However, as the diversity and volume of required data expand and demands intensify, most of vehicular services are unable to satisfy user expectations. Although crowdsensing architectures based on edge computing have been proposed in vehicular networks, how to optimally allocate the sensing service time of roadside units and set appropriate prices for services to maximize user benefits still remain significant challenges. To solve the above issue, in this paper, a crowdsensing service pricing method is proposed in vehicular edge computing. Specifically, a vehicular networking crowdsensing framework based on the edge computing is designed. Then, the optimization problem of perception time pricing during the crowdsensing process is modeled into a Stackelberg game and further the multi-agent deep deterministic policy gradient-based algorithm is employed to obtain the optimal pricing strategy with user profits maximization. Finally, simulation results demonstrate that the proposed method significantly increases the payoff for all participants during the crowdsensing process. Zheng Li 0026, Sizhe Tang, Hao Tian 0012, Haolong Xiang, Xiaolong Xu 0001, Wan-Chun Dou |
ISPA | 1 |
| 2024 | Integrated CNN and Federated Learning for COVID-19 Detection on Chest X-Ray ImagesabstractCurrently, Coronavirus Disease 2019 (COVID-19) is still endangering world health and safety and deep learning (DL) is expected to be the most powerful method for efficient detection of COVID-19. However, patients' privacy concerns prohibit data sharing between medical institutions, leading to unexpected performance of deep neural network (DNN) models. Fortunately, federated learning (FL), as a novel paradigm, allows participating clients to collaboratively train models without exposing source data outside original location. Nevertheless, the current FL-based COVID-19 detection methods prefer optimizing secondary objectives including delay, energy consumption and privacy, while few works focus on improving the model accuracy and stability. In this paper, we propose a federated learning framework with dynamic focus for COVID-19 detection on CXR images, named FedFocus. Specifically, to improve the training efficiency and accuracy, the training loss of each model is taken as the basis for parameter aggregation weights. As training layer deepens, a constantly updated dynamic factor is designed to stabilize the aggregation process. In addition, to highly restore the real dataset, the training sets in our experiments are divided based on the population and the infection of three real cities. Extensive experiments conducted on the real-world CXR images dataset demonstrate that FedFocus outperforms the baselines in model training efficiency, accuracy and stability. Zheng Li 0026, Xiaolong Xu 0001, Xuefei Cao, Yiwen Zhang 0001, Dehua Chen, Haipeng Dai 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | A Knowledge-Driven Anomaly Detection Framework for Social Production SystemabstractIn the social production system, image data are rapidly generated from almost all fields such as factories, hospitals, and transportation, promoting higher requirements for image anomaly detection technologies, including low consumption, higher adaptability, and accuracy. However, existing anomaly detection methods are fragile to heterogeneous image data generated by complex social production systems and tend to require strong computing power and resource support. To address the above problems, a knowledge-driven anomaly detection framework is proposed, in which a local feature enhancement method is designed to strengthen the knowledge representation of the initial features extracted from images. The attention mechanism in deep learning is introduced to adjust the feature attention dynamically according to prior knowledge, which solves the problem of feature loss in the cascade training. To verify the effectiveness of the proposed framework, extensive experiments on social production datasets are conducted. The results demonstrate that our framework outperforms the selected methods on image datasets with different complexity and sample distributions. Zheng Li 0026, Xiaolong Xu 0001, Tian Hang, Haolong Xiang, Yan Cui 0007, Lianyong Qi, Xiaokang Zhou |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Energy-Efficient Task Offloading in UAV-Enabled MEC via Multi-agent Reinforcement Learning
Jiakun Gao, Jie Zhang 0053, Xiaolong Xu 0001, Lianyong Qi, Yuan Yuan 0004, Zheng Li 0026, Wan-Chun Dou |
GPC (2) | 6 |
| 2023 | OPECE: Optimal Placement of Edge Servers in Cloud Environment
Fengmei Chen, Shengjun Xue, Zheng Li 0026, Yachong Tian, Xianyi Cheng |
GPC (2) | 4 |
| 2023 | Blockchain-Assisted Server Placement With Elitist Preserved Genetic Algorithm in Edge ComputingabstractThe distribution of edge resources in the edge computing (EC) environment has an important impact on the Quality of Service (QoS) of edge services. Unreasonable server placement will inevitably lead to problems, such as server overload or underload, deteriorating workload balancing and service wait time. Therefore, the key issue to be addressed in server placement is how to enhance the QoS of edge services through efficient edge server (ES) placement strategies under multiple requirements, such as average task wait time and data privacy. EC-assisted with blockchain technology was argued to be the most potential solution. In this article, we propose a blockchain-assisted secure ES placement algorithm named ETS_GA. ETS_GA is based on the elite-preserving genetic algorithm (EGA), which is proven to converge. The premature problem of traditional genetic algorithm (GA) is effectively solved by using tabu search (TS) and niche sharing (NS). In addition, we construct an adaptive state supervising machine (ASM) to realize real-time algorithm supervision and adaptively iterate the optimization strategy. Blockchain-based privacy protection methods are also deployed in the placed servers to provide real-time privacy protection. Finally, our proposed method is experimentally compared with four baselines using the real Shanghai Telecom base station data set, whose results demonstrate the superiority of ETS_GA in terms of convergence and global search capability. Zheng Li 0026, Guosheng Li, Muhammad Bilal 0003, Dongqing Liu, Xiaolong Xu 0001 |
IEEE Internet Things J. | 1 |
| 2023 | Elf: Erasing-based Lossless Floating-Point CompressionabstractThere are a prohibitively large number of floating-point time series data generated at an unprecedentedly high rate. An efficient, compact and lossless compression for time series data is of great importance for a wide range of scenarios. Most existing lossless floating-point compression methods are based on the XOR operation, but they do not fully exploit the trailing zeros, which usually results in an unsatisfactory compression ratio. This paper proposes an Erasing-based Lossless Floating-point compression algorithm, i.e., Elf. The main idea of Elf is to erase the last few bits (i.e., set them to zero) of floating-point values, so the XORed values are supposed to contain many trailing zeros. The challenges of the erasing-based method are three-fold. First, how to quickly determine the erased bits? Second, how to losslessly recover the original data from the erased ones? Third, how to compactly encode the erased data? Through rigorous mathematical analysis, Elf can directly determine the erased bits and restore the original values without losing any precision. To further improve the compression ratio, we propose a novel encoding strategy for the XORed values with many trailing zeros. Elf works in a streaming fashion. It takes only O ( N ) (where N is the length of a time series) in time and O (1) in space, and achieves a notable compression ratio with a theoretical guarantee. Extensive experiments using 22 datasets show the powerful performance of Elf compared with 9 advanced competitors. Zheng Li 0026, Chao Chen 0004, Yu Zheng 0004 |
Proc. VLDB Endow. | 2 |
| 2023 | Network Dynamic GCN Influence Maximization Algorithm With Leader Fake Labeling MechanismabstractInfluence maximization is an important technique for its significant value on various social network applications, such as viral marketing, advertisement, and recommendation. Traditional heuristic algorithms for influence maximization suffer from different issues, such as accuracy descent and information loss during network exploration. Besides, recently proposed deep learning-based approaches cannot extract the in-depth structural information of social networks well. In light of these problems, we propose a novel heuristic method called the network dynamic GCN influence maximization algorithm based on the leader fake labeling mechanism. To exploit the in-depth network topology information for the influence maximization task, we design a network dynamic GCN that owns adaptive layer numbers in terms of different network scales to obtain node representations. Then, considering that there are few labels in the network or the labels are irrelevant to the task of influence maximization, we establish a leader fake labeling mechanism to automatically generate node labels that are helpful to seed nodes selecting for model training. Finally, a heuristic method based on the Mahalanobis distance is developed to quickly select influential seed nodes with the learned node representations. Three real-world datasets are used in our experiments, and the experimental results demonstrate that our algorithm has a better performance for seed set identification under the premise of high efficiency compared with some latest heuristic influence maximization algorithms. Weimin Li 0001, Dingmei Wei, Zheng Li 0026 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2023 | Federated Learning-Based Cross-Enterprise Recommendation With Graph Neural NetworksabstractRecommender systems are technology-driven marketing solutions for businesses that analyze user behavior data. However, collaborative data sharing between enterprises is often prohibited by privacy protection regulations, leading to insufficient data for graph neural networks (GNNs) training. Fortunately, federated learning (FL), a collaborative training framework without exposing source data, can be applied congruently. Nevertheless, most of FL-based GNN model training methods adopt federated averaging, which performs poorly on highly heterogeneous graph data. To solve this problem, a FL-based GNN Model Training framework for cross-enterprise recommendation, named FL-GMT, is proposed. Specifically, a GNN-based recommendation model is deployed as the local training model. Then, considering the performance inequity caused by uneven sample quality, a loss-based federated aggregation algorithm is designed, effectively improving the performance of disadvantaged participants. To improve the system stability at the end of the aggregation, a dynamic update method of loss attention is designed. Extensive experiments on benchmark datasets demonstrate that FL-GMT outperforms baselines in terms of system fairness, stability, and accuracy. Zheng Li 0026, Muhammad Bilal 0003, Xiaolong Xu 0001, Jielin Jiang, Yan Cui 0007 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Safe: Synergic Data Filtering for Federated Learning in Cloud-Edge ComputingabstractWith the increasing data scale in the Industrial Internet of Things, edge computing coordinated with machine learning is regarded as an effective way to raise the novel latency-sensitive services. To ensure the data privacy for frequent service access, federated learning (FL), as a privacy-preserving distributed framework, is integrated into edge computing, enabling user data invisible to the training process. However, sophisticated network attacks threaten deep learning (DL) models by data poison and malicious reasoning, making the DL-based system untrustworthy. To this end, a synergic data filtering method, named Safe, is proposed to deal with the poisoning attacks. Specifically, considering that the distributed support vector machine is at risk of being attacked due to its distribution and openness to communication, edge-cloud empowered FL framework is designed. Then, the alternating direction method of multipliers is deployed to detect attacked devices whose training processes will be interrupted. Moreover, due to the untrustworthiness of label data, the poisoned data in the attacked devices are figured out and filtered by clustering the trusted data withK-means clustering algorithm. Eventually, extensive experiment results proved that the Safe outperforms correlation methods in detection accuracy and trustworthiness. Xiaolong Xu 0001, Zheng Li 0026, Xiaokang Zhou |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Information diffusion across cyber-physical-social systems in smart city: A survey
Xiaokang Zhou, Shaohua Li 0004, Zheng Li 0026, Weimin Li 0001 |
Neurocomputing | 3 |
| 2021 | Influence maximization algorithm based on Gaussian propagation modelabstractThe influence of each entity in a network is a crucial index of the network information dissemination. Greedy influence maximization algorithms suffer from time efficiency and scalability issues. In contrast, heuristic influence maximization algorithms improve efficiency, but they cannot guarantee accurate results. Considering this, this paper proposes a Gaussian propagation model based on the social networks. Multi-dimensional space modeling is constructed by offset, motif, and degree dimensions for propagation simulation. This space’s circumstances are controlled by some influence diffusion parameters. An influence maximization algorithm is proposed under this model, and this paper uses an improved CELF algorithm to accelerate the influence maximization algorithm. Further, the paper evaluates the effectiveness of the influence maximization algorithm based on the Gaussian propagation model supported by theoretical proofs. Extensive experiments are conducted to compare the effectiveness and efficiency of a series of influence maximization algorithms. The results of the experiments demonstrate that the proposed algorithm shows significant improvement in both effectiveness and efficiency. Weimin Li 0001, Zheng Li 0026, Alex Munyole Luvembe, Chao Yang 0015 |
Inf. Sci. | 2 |
| 2020 | A Review of Techniques and Methods for IoT Applications in Collaborative Cloud-Fog EnvironmentabstractCloud computing is widely used for its powerful and accessible computing and storage capacity. However, with the development trend of Internet of Things (IoTs), the distance between cloud and terminal devices can no longer meet the new requirements of low latency and real-time interaction of IoTs. Fog has been proposed as a complement to the cloud which moves servers to the edge of the network, making it possible to process service requests of terminal devices locally. Despite the fact that fog computing solves many obstacles for the development of IoT, there are still many problems to be solved for its immature technology. In this paper, the concepts and characteristics of cloud and fog computing are introduced, followed by the comparison and collaboration between them. We summarize main challenges IoT faces in new application requirements (e.g., low latency, network bandwidth constraints, resource constraints of devices, stability of service, and security) and analyze fog-based solutions. The remaining challenges and research directions of fog after integrating into IoT system are discussed. In addition, the key role that fog computing based on 5G may play in the field of intelligent driving and tactile robots is prospected. Jielin Jiang, Zheng Li 0026, Yuan Tian 0003, Najla Al-Nabhan |
Secur. Commun. Networks | 2 |