Ying Zheng 0004

dblp:71/5417-4 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0001-8823-2460ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Is the Attention Matrix Really the Key to Self-Attention in Multivariate Long-Term Time Series Forecasting?
abstract
In multivariate long-term time series forecasting, the success of self-attention is commonly attributed to the attention matrix that encodes token interactions.In this paper, we provide evidence that challenges this view.Through extensive experiments on three classic and three latest Transformer models, we find that dotproduct attention can be replaced by elementwise operations without token interaction, such as the addition and Hadamard product, while maintaining or even improving accuracy.This motivates our central hypothesis: the effectiveness of self-attention in this task arises not from the dynamic attention matrix, but from the multi-branch feature extraction enabled by the parallel Query, Key, and Value projections and their fusion.To validate this hypothesis, we construct a minimalist multi-branch MLP that isolates the 'multi-branch mapping with element-wise operation' structure from the Transformer and show that it achieves competitive performance.Our findings indicate that the source of performance in self-attention is often misinterpreted, as its actual advantage stems from the architectural principle of multi-branch mapping and fusion, rather than the attention matrix.
Xinyu Li 0014, Kexi Chen, Jiajie Shen, Ying Zheng 0004, Hong Lu 0001, Jin Zhao 0001, Xin Wang 0002
ACL (1)4
2026 Modeling Point-to-Point Dependency for High-Dimensional Long-Term Series Forecasting
Xinyu Li 0014, Kexi Chen, Ying Zheng 0004, Zhiyi Yao, Yi Xie 0003, Jihan Dai, Lei Bai 0001, Jin Zhao 0001, Jiajie Shen, Yunqi Cai, Hong Lu 0001, Xin Wang 0002
WWW3
2026 Online Request Scheduling for Quality-Aware Diffusion-Based AIGC Services
abstract
Artificial Intelligence-Generated Content (AIGC) has been gaining significant traction for automatic generation of diverse content. Due to the GPU-intensive generation process and the high costs associated with purchasing and operating GPUs, users often prefer to submit requests to a nearby edge cloud, maintained by an AIGC cloud service provider. Efficiently scheduling AIGC requests in the edge cloud faces non-trivial challenges. First, AIGC requests emphasize the quality of generated content, yet conventional scheduling algorithms often overlook this aspect. Second, when the volume of incoming requests exceeds the capacity of the cloud, the AIGC service provider needs to select appropriate requests to execute, which is further complicated by the online arrival pattern of requests and the constraints imposed by request deadlines. Third, users dynamically submit multiple requests at different times. To manage costs, each user operates within a pre-allocated budget for a given time period. For the AIGC cloud service provider, it is highly non-trivial to identify valuable requests and judiciously balance different user budgets. To tackle the above challenges, we target the online AIGC request scheduling problem with the new objective of maximizing the overall content generation quality. We first conduct real experiments to establish the quality model between inference steps and the quality of generated content. Then, based on this quality model, we formulate the problem into an integer linear program, which is proven NP-hard. Under a primal-dual framework, we carefully design the update of multiple dual variables, to flexibly control the consumption of edge resources and user budgets. We rigorously analyze the performance of the proposed algorithm and prove a theoretical performance guarantee on its competitive ratio. Extensive real-world trace-driven experiments manifest that our proposed method improves the state-of-the-art by up to 25.3% in overall content generation quality.
Ying Zheng 0004, Lei Jiao 0002, Yuedong Xu 0001, Zongpeng Li
IEEE Trans. Netw.2
2024 Online Scheduling and Pricing for Multi-LoRA Fine-Tuning Tasks
abstract
Fine-tuning pre-trained models with task-specific data can produce customized models effective for downstream tasks. However, operating large-scale such fine-tuning tasks in real time in the data center faces non-trivial challenges, including unpredictable task arrival and system environment dynamics, complex deadline-driven fine-tuning scheduling, and intertwined task pricing and cost management. In this paper, targeting the popular Low-Rank Adaptation (LoRA) fine-tuning technique, we present the design and study of a novel auction-based mechanism to jointly schedule and price LoRA tasks in an online manner. We first model the social welfare maximization problem as an integer program for the fine-tuning service provider, capturing all the aforementioned challenges. Then, to solve this NP-hard problem online, we equivalently reformulate this original problem into a schedule selection problem, where each schedule corresponds to a concrete pre-specified operation plan over time for a task. We can thus design a polynomial-time online approximation algorithm via the online primal-dual method to determine the schedule, and with the dual variables, also determine the pricing for each admitted task. We rigorously prove the competitiveness of our online approach against the offline optimum, and prove the economic properties of truthfulness and individual rationality regarding pricing. Finally, we conduct extensive experiments and have validated the substantial advantages of our approach compared to existing methods.
Ying Zheng 0004, Lei Jiao 0002, Lulu Chen, Yuedong Xu 0001, Xin Wang 0003, Zongpeng Li
ICPP1
2024 Scheduling Generative-AI Job DAGs with Model Serving in Data Centers
abstract
Scheduling generative-AI jobs in the edge computing environment faces multiple non-trivial challenges, including the Directed Acyclic Graph (DAG) dependency among tasks, the intrinsic intertwinement between task scheduling and model selection, and the dynamic unpredictable arrival of job DAGs. In this work, we capture all such challenges and formulate a non-linear integer program to optimize the long-term profit of the generative-AI service provider, i.e., service revenue of the admitted jobs minus system costs of executing the tasks contained in such job DAGs. This problem is NP-hard even in the offline setting. To solve it, we first reformulate it into an equivalent schedule selection problem using generated schedules to tackle complex constraints. Then, we design a new online scheduling method through the online primal-dual technique. Experimental results confirm that our approach can increase the total service profit by up to 41.2% compared to existing algorithms.
Ying Zheng 0004, Lei Jiao 0002, Yuedong Xu 0001, Bo An 0001, Xin Wang 0003, Zongpeng Li
IWQoS1
2022 Enabling Robust DRL-Driven Networking Systems via Teacher-Student Learning
abstract
The past few years have witnessed a surge of interest towards deep reinforcement learning (DRL) in computer networks. With extraordinary ability of feature extraction, DRL has the potential to re-engineer the fundamental resource allocation problems in networking without relying on pre-programmed models or assumptions about dynamic environments. However, such black-box systems suffer from poor robustness, showing high performance variance and poor tail performance. In this work, we propose a unified Teacher-Student learning framework that harnesses rich domain knowledge to improve robustness. The domain-specific algorithms, less performant but more trustable than DRL, play the role of teachers providing advice at critical states; the student neural network is steered to maximize the expected reward as usual and mimic the teacher’s advice meanwhile. The Teacher-Student method comprises of three modules where the confidence check module locates wrong decisions and risky decisions, the reward shaping module designs a new updating function to stimulate the learning of student network, and the prioritized experience replay module to effectively utilize the advised actions. We further implement our Teacher-Student framework in existing video streaming (Pensieve), load balancing (DeepLB), and TCP congestion control (Aurora). Experimental results manifest that the proposed approach reduces the performance standard deviation of DeepLB by 37%; it improves the 90th, 95th, and 99th tail performance of Pensieve by 7.6%, 8.8%, and 10.7% respectively; and it accelerates the growth rate of Aurora by 2x at the initial stage, and achieves a more stable performance in dynamic environments.
Ying Zheng 0004, Lixiang Lin, Qingyang Duan, Yuedong Xu 0001, Xin Wang 0002
IEEE J. Sel. Areas Commun.1
2021 Leveraging Domain Knowledge for Robust Deep Reinforcement Learning in Networking
abstract
The past few years has witnessed a surge of interest towards deep reinforcement learning (Deep RL) in computer networks. With extraordinary ability of feature extraction, Deep RL has the potential to re-engineer the fundamental resource allocation problems in networking without relying on pre-programmed models or assumptions about dynamic environments. However, such black-box systems suffer from poor robustness, showing high performance variance and poor tail performance. In this work, we propose a unified Teacher-Student learning framework that harnesses rich domain knowledge to improve robustness. The domain-specific algorithms, less performant but more trustable than Deep RL, play the role of teachers providing advice at critical states; the student neural network is steered to maximize the expected reward as usual and mimic the teacher's advice meanwhile. The Teacher-Student method comprises of three modules where the confidence check module locates wrong decisions and risky decisions, the reward shaping module designs a new updating function to incentive the learning of student network, and the prioritized experience replay module to effectively utilize the advised actions. We further implement our Teacher-Student framework in existing video streaming (Pensieve), load balancing (DeepLB) and TCP congestion control (Aurora). Experimental results manifest that the proposed approach reduces the performance standard deviation of DeepLB by 37%; it improves the 90th, 95th and 99th tail performance of Pensieve by 7.6%, 8.8%, 10.7% respectively; and it accelerates the rate of growth of Aurora by 2x at the initial stage, and achieves a more stable performance in dynamic environments.
Ying Zheng 0004, Qingyang Duan, Lixiang Lin, Yiyang Shao, Wei Wang 0334, Xin Wang 0002, Yuedong Xu 0001
INFOCOM1
2018 Demystifying Deep Learning in Networking
abstract
We are witnessing a surge of efforts in networking community to develop deep neural networks (DNNs) based approaches to networking problems. Most results so far have been remarkably promising, which is arguably surprising given how intensively these problems have been studied before. Despite these promises, there has not been much systematic work to understand the inner workings of these DNNs trained in networking settings, their generalizability in different workloads, and their potential synergy with domain-specific knowledge. The problem of model opacity would eventually impede the adoption of DNN-based solutions in practice. This position paper marks the first attempt to shed light on the interpretability of DNNs used in networking problems. Inspired by recent research in ML towards interpretable ML models, we call upon this community to similarly develop techniques and leverage domain-specific insights to demystify the DNNs trained in networking settings, and ultimately unleash the potential of DNNs in an explainable and reliable way.
Ying Zheng 0004, Xinyu You, Yuedong Xu 0001, Junchen Jiang
APNet1