Cheng Zhang 0019

dblp:82/6384-19 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-1744-9571ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Computer networks · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DynEformer: A Unified Framework for Robust Workload Prediction Under Dynamic Environment
abstract
Workload prediction in multi-tenant edge cloud platforms (MT-ECP) is crucial for efficient application deployment and resource provisioning. However, the heterogeneous application patterns, variable infrastructure performance, and frequent deployments in MT-ECP pose significant challenges for accurate prediction. Existing clustering-based methods often incur excessive costs due to maintaining multiple data clusters and models, while end-to-end time-series prediction methods struggle with dynamic environments. To address these challenges, we perform a comprehensive analysis on a large-scale workload dataset in real-world MT-ECP and propose DynEformer, an end to-end framework with global pooling and static context aware ness, offering a unified workload prediction scheme for dynamic MT-ECP. Meticulously designed global pooling and information merging mechanisms can effectively identify and utilize global application patterns to drive local workload predictions. The integration of static content-aware mechanisms enhances model robustness in real-world scenarios. We also extend DynEformer's capabilities to Long-term workload forecasting (LTLF) and Long-period service (LPS) tasks. Experiments on six real-world datasets demonstrate that DynEformer achieves state-of-the-art performance, with a 32% relative improvement on nine baselines and a 52% improvement in application switching and new entity scenarios. Additional experiments on long-term prediction and online learning further confirm its effectiveness for LTLF and LPS tasks.
Shaoyuan Huang, Zheng Wang 0001, Heng Zhang 0032, Xiaofei Wang 0001, Cheng Zhang 0019
IEEE Trans. Knowl. Data Eng.5
2026 Delay-Aware and Energy-Efficient Integrated Optimization System for 5G Networks
abstract
To meet the demands of high-capacity and low-delay services, Fifth Generation (5G) Base Stations (BSs) are typically deployed in ultra-dense configurations, especially in urban areas. While this densification enhances coverage and service quality, it also leads to substantially increased energy consumption. However, the dense deployment pattern makes BS workloads more responsive to the spatiotemporal variations in user behavior, offering opportunities for energy-saving strategies that dynamically adjust BS operation states. In this context, we propose a Delay-aware and Energy-efficient Integrated Optimization System (DEIS) based on Deep Reinforcement Learning (DRL), which jointly optimizes energy consumption and network delay while maintaining user satisfaction. DEIS leverages a real-world dataset collected from operational 5G BSs provided by partner network operators, containing both BS deployment data and high-volume user request logs. Extensive simulations demonstrate that DEIS can achieve a 41% reduction in energy consumption while ensuring reliable delay performance.
Jingchao Tan, Tiancheng Zhang 0009, Cheng Zhang 0019, Chenyang Wang 0001, Chao Qiu, Xiaofei Wang 0001, Mohsen Guizani
IEEE Trans. Netw. Serv. Manag.3
2025 Dynamic Interaction-Driven Intent Evolver with Semantic Probability Distributions
abstract
Accurately capturing a user's dynamic search intent based on her/his interactions with the system is crucial for improving the performance of session-based search. Existing methods often require the entire interaction sequence within a session to be recomputed continuously at each interaction step, and the token-level interactions are either captured within an overall transformer structure or simply ignored. As a consequence, the current approaches suffer from an increased computation burden and fall short of accurately capturing the dynamic evolution of user intent. In this paper, we propose a novel representation approach which treats both search intent and candidate documents as dimension-specific probability distributions of token embedding representations. Based on this representation, we propose an Dynamic Interaction-Driven intent Evolver (DIDE) for dynamically updating the user's search intent throughout a session with a lightweight similarity calculation method for document ranking. Comprehensive experimental results demonstrate that DIDE adeptly captures the dynamic nature of session-based search and significantly outperforms a range of strong baseline models across three different datasets.
Zelin Li 0001, Cheng Zhang 0019, Dawei Song 0001
WSDM2
2025 Unravelling the semantic mysteries of transformers layer by layer
abstract
Abstract Despite the significant success of transformer models and their successors in various natural language processing (NLP) applications, their internal workings are still not fully understood. Much of the current interpretability research has focused primarily on numerical components, often missing the complex semantic layers within these models. To fill this gap, this study explores the interpretability of the transformer model, a cornerstone of modern NLP, by addressing the semantic complexities of its multi-layer architecture. We identify three key questions: (i) the influence of the multi-layer structure on semantic processing, (ii) the unique contributions of each layer to model performance, and (iii) methodologies for determining optimal layer counts for the encoder and decoder. To tackle these issues, we introduce the semantic interpreter for transformer hierarchy, an innovative framework that employs convex hull metrics to visualize and assess semantic quality and quantity. Our contributions include novel methods for semantic assessment, a dual analytical framework that integrates cumulative and layer-to-layer perspectives, and insights into the dynamics of encoding and decoding. This comprehensive approach aims to enhance the understanding of Transformer models, ultimately guiding their refinement for improved interpretability and effectiveness in natural language tasks.
Cheng Zhang 0019, Jinxin Lv, Jingxu Cao, Jiachuan Sheng, Dawei Song 0001, Tiancheng Zhang 0009
Comput. J.1
2025 Cloud-edge-end integrated Artificial intelligence based on ensemble learning
Zhen Gao 0005, Daning Su, Chenyang Wang 0001, Cheng Zhang 0019, Xiaofei Wang 0001, Tarik Taleb
Comput. Commun.6
2025 Enabling Real-Time Video Detection With Adaptive and Distributed Scheduling in Mobile Edge Computing
abstract
Real-time video detection is essential for many mobile visual applications, which brings the heavy computational burden of deep neural networks. Mobile edge computing offers a promising solution by deploying computational resources near mobile devices. However, achieving efficient video detection on mobile devices requires addressing challenges such as different performance requirements, diverse computing and network conditions, and system dynamics. We propose a realtime video detection framework in mobile edge computing, where multiple video streams from mobile devices are processed while balancing key performance metrics with consideration of grouping. A joint optimization problem of task scheduling, model selection, and resource provisioning is formulated for the system, where decisions are made on two timescales. To this end, we propose a window controller to unify decision-making at the time-slot level. We design an online scheduling algorithm based on multi-agent deep reinforcement learning to enable adaptive and distributed scheduling, while a masking-enhanced attention mechanism enables efficient explicit information exchange between mobile devices. Experimental evaluations across different numbers of mobile devices demonstrate that, in terms of average reward, the proposed algorithm outperforms local processing by 14.600%, fixed offloading by 10.007%, and four learning-based scheduling baselines by an average of 2.267%.
Yilan Wang, Chao Qiu, Cheng Zhang 0019, Xiaofei Wang 0001, Mianxiong Dong
IEEE Trans. Mob. Comput.5
2025 Task Allocation With Geography-Context-Capacity Awareness in Distributed Burstable Billing Edge-Cloud Systems
abstract
The new real-time interactive services, such as virtual and augmented reality, demand significantly higher network bandwidth and quality, which the traditional centralized cloud struggles to meet. In addition, centralized optimization management becomes inefficient as the scale of the scene continues to expand. In response, edge cloud systems have emerged, but distributed geographic locations, burstable billing business models, and large numbers of servers in large-scale scenarios pose new challenges for resource management. In this article, we proposeGeoCC, a novel strategy to save bandwidth overhead in burstable billing edge cloud systems.GeoCCaddresses challenges through a dual approach. First, a geography-aware graph construction and partitioning algorithm is used to organize server resources, and a large number of servers are reasonably divided into multiple server pools for parallel processing. Second, it introduces an enhanced burstable billing optimization mechanism that considers contextual factors and adaptive bandwidth capacity. Experiments based on real data from an edge cloud operator demonstrate the effectiveness ofGeoCC. Compared with the baseline,GeoCCcan effectively reduce bandwidth peaks, decreasing bandwidth costs by an average of 28.30% and up to 81.83% at the 95th percentile billing.
Shihao Shen, Chenfei Gu, Yuanze Li, Chao Qiu, Xiaofei Wang 0001, Rui Tan 0001, Cheng Zhang 0019
IEEE Trans. Serv. Comput.7
2024 SITH: Semantic Interpreter for Transformer Hierarchy
abstract
While Transformers and their derivatives have shown strong performance in various NLP tasks, understanding their internal mechanisms remains challenging. Mainstream interpretability research often focuses solely on numerical attributes, neglecting the complex semantic structure inherent in the model. We have developed the SITH(Semantic Interpreter for Transformer Hierarchy) framework to address this issue. We focus on creating universal text representation methods and uncovering the semantic principles of the Transformer's hierarchical structure. We use the convex hull method to represent sequence semantics in an n-dimensional Semantic Euclidean space and define different evaluation indicators through convex hull to analyze semantic quality and quantity changes. Our analysis takes a dual perspective: a multi-layer cumulative perspective and an individual layer-to-layer shift perspective. When applied to machine translation, our results reveal potential semantic processes and emphasize the effectiveness of stacking and hierarchical differences. These insights are valuable for fine-tuning hyperparameters at the encoder and decoder layers.
Cheng Zhang 0019, Jinxin Lv, Jingxu Cao, Jiachuan Sheng, Dawei Song 0001, Tiancheng Zhang 0009
ICTAI1
2024 MTEE: Multiscale Temporal Entropy Evaluation Paradigm for Heterogeneous Complex Datasets
Ledong An, Chenyang Wang 0001, Shaoyuan Huang, Cheng Zhang 0019, Chao Qiu, Xiaofei Wang 0001
NPC (1)5
2024 AnaNET: Anatomical Network for Aggregated Time Series Forecasting in Multi-layered Architecture
Tiancheng Zhang 0009, Cheng Zhang 0019, Shuren Liu, Xiaofei Wang 0001, Shaoyuan Huang
NPC (1)2
2024 MoEI: Mobility-Aware Edge Inference Based on Model Partition and Service Migration
abstract
Deep neural networks are the cornerstone of many mobile intelligent systems, and their inference processes bring about computation-intensive tasks. Device-edge cooperative inference in mobile edge computing provides a fine-grained processing method to migrate the burden of inference computation. However, the geographical dispersion of resources and the mobility pattern of devices pose scheduling issues to be considered. In this paper, we propose a task scheduling framework for such device-edge systems to improve the pipeline time of model inference. First, we consider the resource provisioning strategy with a pre-fetching service migration setting in the environment of multiple mobile devices and edge nodes. Then, we leverage game theory to analyze the property of the decision-making process and propose an offline algorithm under complete information. Next, we propose an algorithm based on proximal policy optimization to enable mobile devices to make decisions in a distributed online manner. Further, we adopt a memory mechanism into the online algorithm to improve the decision-makers' understanding of the system environment. Experiments demonstrate the effectiveness of the two algorithms. The average pipeline time of the proposed online algorithm is only 61.44% of that of local processing, which is 1.196 times that of the proposed offline algorithm.
Mianxiong Dong, Xiaofei Wang 0001, Chao Qiu, Cheng Zhang 0019
IEEE Trans. Mob. Comput.6
2023 Quicklayer: A Layer-Stack-Oriented Accelerating Middleware for Fast Deployment in Edge Clouds
abstract
Containers are gaining popularity in edge computing due to their standardization and low overhead. This trend has brought new technologies such as container engines and container orchestration platforms (COPs). However, fast and effective container deployment remains a challenge, especially at the edge. Prior work, which was designed for cloud datacenters, is no longer suitable for container deployment in edge clouds due to bandwidth limitations, fluctuating network performance, resource constraints, and geo-distributed organization. These edge features make rapid deployment on the edge difficult. Additionally, integrating with COPs is crucial for successful deployment.
Yicheng Feng, Shihao Shen, Cheng Zhang 0019, Xiaofei Wang 0001
APNet3
2023 Which Words Pillar the Semantic Expression of a Sentence?
abstract
In the realm of machine learning, a profound understanding of sentence semantics holds paramount importance for various applications, notably text classification. Traditionally, this comprehension has been entrusted to deep learning models, despite their computationally intensive nature, particularly when dealing with lengthy sequences. The nuanced impact of individual words within a sentence on semantic expression necessitates a strategic removal of less pertinent words to alleviate the computational burden of the model. Presently, prevailing approaches for word removal predominantly employ methods such as truncation, stop-word elimination and attention mechanisms. Regrettably, these techniques often lack a robust theoretical foundation concerning semantics and interpretability. To bridge this conceptual gap, our study introduces the concept of ‘Semantic Pillar Words’ (SPW) within a sentence, anchored in a Semantic Euclidean space. Here, the semantics of a word are represented as a constellation of semantic points, with a text sequence encapsulating the convex hull of these semantic points of words. We propose a novel method for Semantic Pillar Word extraction, known as ‘SPW-Conv’, which dynamically and interpretably prunes text segments, striving to preserve the semantic pillars inherent in the original text. Our extensive experimentation encompasses three diverse text classification datasets, revealing that SPW-Conv outperforms existing methods. Remarkably, it becomes evident that retaining less than 80% of the words within a sentence suffices to capture its semantics adequately, all while achieving classification accuracy levels comparable to those obtained using the entire original text.
Cheng Zhang 0019, Jingxu Cao, Dongmei Yan, Dawei Song 0001, Jinxin Lv
ICTAI1
2020 Assessing the Memory Ability of Recurrent Neural Networks
abstract
It is known that Recurrent Neural Networks (RNNs) can remember, in their hidden layers, part of the semantic information expressed by a sequence (e.g., a sentence) that is being processed. Different types of recurrent units have been designed to enable RNNs to remember information over longer time spans. However, the memory abilities of different recurrent units are still theoretically and empirically unclear, thus limiting the development of more effective and explainable RNNs. To tackle the problem, in this paper, we identify and analyze the internal and external factors that affect the memory ability of RNNs, and propose a Semantic Euclidean Space to represent the semantics expressed by a sequence. Based on the Semantic Euclidean Space, a series of evaluation indicators are defined to measure the memory abilities of different recurrent units and analyze their limitations (Code is available at https://github.com/chzhang/Assessing-the-Memory-Ability-of-RNNs). These evaluation indicato...
Cheng Zhang 0019, Qiuchi Li, Lingyu Hua, Dawei Song 0001
ECAI1
2019 SCSS-LIE: A Novel Synchronous Collaborative Search System with a Live Interactive Engine
abstract
Synchronous collaborative search systems (SCSS) refer to systems which support two or more users with similar information need to search together simultaneously. Generally, SCSS provide a social engine to enable users to communicate. However, when the number of users in the social engine is insufficient to collaborate on the search task, the social engine will encounter the cold start problem and can not perform collaborative search well. In this paper, we present a novel Synchronous Collaborative Search System with a Live Interactive Engine (SCSS-LIE). SCSS-LIE proposes to apply a ring topology to add an intelligent auxiliary robot, Infobot, into the social engine to support real-time interaction between users and the search engine to address the cold start problem of the social engine. The reading comprehension model BiDAF (Bi-Directional Attention Flow) is employed in the Infobot in the process of interacting with the search engine to obtain answers to facilitate the acquisition of information. SCSS-LIE can not only allow users with similar information need to be grouped into one chat channel to communicate, but also enable them to conduct real-time interaction with the search engine to improve search efficiency.
Peng Zhang 0002, Cheng Zhang 0019, Dawei Song 0001
SIGIR3
2018 Text Classification with Enriched Word Features
Jingda Xu, Cheng Zhang 0019, Peng Zhang 0002, Dawei Song 0001
PRICAI2
2016 SECC: A Novel Search Engine Interface with Live Chat Channel
abstract
Traditional information retrieval systems rank documents according to their relevance to users' input queries. State of the art commercial search engines (SEs) train ranking models and suggest query refinements by exploiting collective intelligence implicitly using global users' query logs. However, they do not provide an explicit channel for users to communicate with each other in the search process. By asking or discussing with other users on the fly, a user could find relevant information more conveniently and gain a better search experience. In this paper, we present a demo of novel Search Engine with a live Chat Channel (SECC). SECC can group users automatically based on their input queries and allow them to communicate with each other in real time through a chat interface.
Cheng Zhang 0019, Peng Zhang 0002, Jingfei Li, Dawei Song 0001
SIGIR1