Huijing Yang

dblp:120/1181 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Reinforcement Learning-Driven Adaptive Prefetch Aggressiveness Control for Enhanced Performance in Parallel System Architectures
abstract
In modern parallel system architectures, prefetchers are essential to mitigating the performance challenges posed by long memory access latencies. These architectures rely heavily on efficient memory access patterns to maximize system throughput and resource utilization. Prefetch aggressiveness is a central parameter in managing these access patterns; although increased prefetch aggressiveness can enhance performance for certain applications, it often risks causing cache pollution and bandwidth contention, leading to significant performance degradation in other workloads. While many existing prefetchers rely on static or simple built-in aggressiveness controllers, a more flexible, adaptive approach based on system-level feedback is essential to achieving optimal performance across parallel computing environments. In this paper, we introduce an Adaptive Prefetch Aggressiveness Control (APAC) framework that leverages Reinforcement Learning (RL) to dynamically manage prefetch aggressiveness in parallel system architectures. The APAC controller operates as an RL agent, which optimizes prefetch aggressiveness by dynamically responding to system feedback on prefetch accuracy, timeliness, and cache pollution. The agent receives a reward signal that reflects the impact of each adjustment on both performance and memory bandwidth, learning to adapt its control strategy based on workload characteristics. This data-driven adaptability makes APAC particularly well-suited for parallel architectures, where efficient resource management across cores is essential to scaling system performance. Our evaluation with the ChampSim simulator demonstrates that APAC effectively adapts to diverse workloads and system configurations, achieving performance gains of 6.73$\%$in multi-core systems compared to traditional Feedback Directed Prefetching (FDP). By improving memory bandwidth utilization, reducing cache pollution, and minimizing inter-core interference, APAC significantly enhances prefetching performance in multi-core processors. These results underscore APAC’s potential as a robust solution for performance optimization in parallel system architectures, where efficient resource management is paramount for scaling modern processing environments.
Huijing Yang, Juan Fang 0004, Yumin Hou, Xing Su 0001, Naixue Xiong
IEEE Trans. Parallel Distributed Syst.1
2024 Attention Mechanism-Aided Deep Reinforcement Learning for Dynamic Edge Caching
abstract
The dynamic mechanism of joint proactive caching and cache replacement, which involves placing content items close to cache-enabled edge devices ahead of time until they are requested, is a promising technique for enhancing traffic offloading and relieving heavy network loads. However, due to limited edge cache capacity and wireless transmission resources, accurately predicting users’ future requests and performing dynamic caching is crucial to effectively utilizing these limited resources. This paper investigates joint proactive caching and cache replacement strategies in a general mobile edge computing (MEC) network with multiple users under a cloud-edge-device collaboration architecture. The joint optimization problem is formulated as a markov decision process (MDP) problem with an infinite range of average network load costs, aiming to reduce network load traffic while efficiently utilizing the limited available transport resources. To address this issue, we design an Attention Weighted Deep Deterministic Policy Gradient (AWD2PG) model, which uses attention weights to allocate the number of channels from server to user, and applies deep deterministic policies on both user and server sides for Cache decision-making, so as to achieve the purpose of reducing network traffic load and improving network and cache resource utilization. We verify the convergence of the corresponding algorithms and demonstrate the effectiveness of the proposed AWD2PG strategy and benchmark in reducing network load and improving hit rate.
Ziyi Teng, Juan Fang 0004, Huijing Yang, Huijie Chen, Wei Xiang 0001
IEEE Internet Things J.3
2024 RL-CoPref: a reinforcement learning-based coordinated prefetching controller for multiple prefetchers
abstract
Abstract Modern processors employ data prefetchers to alleviate the impact of long memory access latency. However, current prefetchers are designed for specific memory access patterns, which perform poorly on mixed applications with multiple memory access patterns. To address these issues, RL-CoPref, a reinforcement learning (RL)-based coordinated prefetching controller for multiple prefetchers, is proposed in this paper. RL-CoPref takes diverse program context information as the input, learns to maximize cumulative rewards, and evaluates prefetch quality based on prefetch hits/misses and memory bandwidth utilization. It can dynamically adjust the prefetch activation and prefetch degree, enabling multiple prefetchers to complement each other on mixed applications. Our extensive evaluation, utilizing the ChampSim simulator, demonstrates that RL-CoPref can effectively adapt to various workloads and system configurations, optimizing prefetch control. On average, RL-CoPref achieves 76.15% prefetch coverage, having 35.50% IPC improvement, outperforming state-of-the-art individual prefetchers by 5.91–16.54% and outperforming SBP, a state-of-the-art (non-RL) prefetch controller, by 4.64%.
Huijing Yang, Juan Fang 0004, Xing Su 0001, Zhi Cai, Yuening Wang
J. Supercomput.1
2023 DPBC-VCP: A Network-On-Chip Prioritization Mechanism Combined with VCP for CPU-GPU Heterogeneous Systems
abstract
When executing CPU and GPU applications in CPU-GPU heterogeneous systems, a common phenomenon arises where CPU applications performance is often interfered by GPU applications. This study substantiates this observation through an analysis of resource contention and identifies the limitations of the Virtual Channel Partitioning (VCP) approach in the crossbar switch allocation stage. In response to the resource contention problem in crossbar switch allocation stage, we propose a Probability-Based CPU-first Arbitration Strategy that enhances the priority of CPU packets in contention through specific probabilities. Furthermore, we introduce a Dynamic Probability-Based CPU-first Arbitration Strategy (DPBC) that dynamically selects probability values based on application execution phases to strike a balance between optimal CPU and GPU performance. Moreover, we combine this dynamic strategy with VCP to further enhance CPU performance, propose the DPBC-VCP method. The DPBC-VCP, a combination of network partitioning and prioritization techniques, yields an average enhancement of 48% in CPU performance compared to the baseline, with only a marginal 2.45% reduction in GPU performance.
Haoyu Cheng, Zhichao Wei, Huijing Yang
ICPADS4
2023 A low offset low power CMOS dynamic comparator for analog to digital converters
Huijing Yang, Mingyuan Ren
Integr.1
2023 A Prefetch-Adaptive Intelligent Cache Replacement Policy Based on Machine Learning
Huijing Yang, Min Cai, Zhi Cai
J. Comput. Sci. Technol.1
2023 A perceptual and predictive batch-processing memory scheduling strategy for a CPU-GPU heterogeneous system
abstract
When multiple central processing unit (CPU) cores and integrated graphics processing units (GPUs) share off-chip main memory, CPU and GPU applications compete for the critical memory resource. This causes serious resource competition and has a negative impact on the overall performance of the system. We describe the competition for shared-memory resources in a CPU-GPU heterogeneous multi-core architecture, and a shared-memory request scheduling strategy based on perceptual and predictive batch-processing is proposed. By sensing the CPU and GPU memory request conditions in the request buffer, the proposed scheduling strategy estimates the GPU latency tolerance and reduces mutual interference between CPU and GPU by processing CPU or GPU memory requests in batches. According to the simulation results, the scheduling strategy improves CPU performance by 8.53% and reduces mutual interference by 10.38% with low hardware complexity.
Juan Fang 0004, Huijing Yang, Yixiang Xu, Xing Su 0001
Frontiers Inf. Technol. Electron. Eng.3
2023 WSMP: a warp scheduling strategy based on MFQ and PPF
abstract
Abstract Normally, threads in a warp do not severely interfere with each other. However, the scheduler must wait until all the threads within complete before scheduling the next warp, resulting in memory divergence. The crux of the problem is scheduling the warp in a more reasonable order. Therefore, we propose a new warp scheduling strategy called WSMP, which is based on multi-level feedback queue (MFQ) and perceptron-based prefetch filtering (PPF). All the warps are sorted beforehand according to the latency tolerance of the warps and pushed into a certain queue in MFQ. We also remold PPF to enhance the modified underlying prefetcher. We are able to strike a balance between cache hit rate and prefetch coverage then. We verify its feasibility using GPGPU-Sim, along with exclusive GPGPU workload. The results show that compared to the baseline, WSMP improves IPC by 26.45% and reduces L2 cache miss rate by 9.54% on average.
Li'ang Zhao, Min Cai, Huijing Yang
J. Supercomput.4
2021 Deep Reinforcement Learning for RAN Optimization and Control
abstract
Due to the high variability of the traffic in the radio access network (RAN), fixed network configurations are not flexible enough to achieve optimal performance. Our vendors provide several settings of the eNodeB to optimize the RAN performance, such as media access control scheduler, loading balance, etc. But the detailed mechanisms of the eNodeB configurations are usually very complicated and not disclosed, not to mention the large key performance indicators (KPIs) space needed to be considered. These make constructing a simulator, offline tuning, or rule-based solutions difficult. We aim to build an intelligent controller without strong assumption or domain knowledge about the RAN and can run 24/7 without supervision. To achieve this goal, we first build a closed-loop control testbed RAN in a lab environment with one eNodeB provided by one of the largest wireless vendors and four smartphones. Next, we build a double Q network agent trained with the live feedback of the key performance indicators from the RAN. Our work proved the effectiveness of applying deep reinforcement learning to improve network performance in a real RAN network environment.
Yu Chen 0019, Ganesh Krishnamurthi, Huijing Yang, Huahui Wang
WCNC4