EDBT 2026 Demo / reviewers in the wild / expert
Fan Jiang 0015
dblp:58/3794-15
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2024
0000-0003-2625-6111ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SCNoCs: An Adaptive Heterogeneous Multi-NoC with Selective Compression and Power GatingabstractIn-network compression has been proposed recently to support efficient communication. However, we find employing compression blindly cannot always pay off since de/compression leads to extra packet transmission delay. We thereby propose selective compression which compresses data adaptively based on network state and predicted compression ratio. Moreover, we observe that simply applying selective compression in a conventional single network is not energy efficient. Therefore, we propose SCNoCs, a heterogeneous Multi-NoC (Main-Net and HelperNet) architecture with the support of selective compression and power gating. SCNoCs can dynamically adjust the policy of selective compression and the utilization degree of the Helper-Net according to the network state at runtime. Experimental results show that our selective compression outperforms conventional compression by 1.5$ \times $. Besides, our proposed SCNoCs achieves comparable performance while reducing energy consumption by 43.4%, compared with the baseline. Fan Jiang 0015, Chengeng Li, Lin Chen 0029, Wei Zhang 0012, Jiang Xu 0001 |
ASPDAC | 1 |
| 2024 | Collaborative Coalescing of Redundant Memory Access for GPU SystemabstractGPU-based computing serves as the primary solution driving the performance of HPC systems. However, modern GPU systems encounter performance bottlenecks resulting from heavy memory access traffic and insufficient NoC bandwidth. In this work, we propose a collaborative coalescing mechanism aimed at eliminating redundant memory access and boosting GPU system performance. To achieve this, we design a coalescing unit for each memory partition, effectively merging requests from both inter-cluster and intra-cluster SMs. Additionally, we introduce a hierarchical multicast module to replicate and distribute the coalesced reply messages to multiple destination SMs. Experimental results show that our method achieves 20.6% improvement on performance and 27.1% reduction on NoC traffic over the baseline. Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
ASPDAC | 1 |
| 2024 | NEOCNN: NTT-Enabled Optical Convolution Neural Network AcceleratorabstractIn the realm of neural network computation, optical neural network accelerators (ONNs) have emerged as a promising solution, leveraging the inherent speed and parallelism of optical systems. Despite their potential, current ONN designs often fall short due to inefficient data movement and reliance on traditional electronics-based dataflows. Yinyi Liu, Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
ICS | 3 |
| 2024 | PC-oriented Prediction-based Runtime Power Management for GPGPU using Knowledge TransferabstractAs Moore's law slows down, computing systems must prioritize higher energy efficiency to sustain performance scaling. GPUs have emerged as the primary workhorses of computing resources, making the achievement of high energy efficiency in GPUs a critical concern. However, implementing runtime power management on GPUs poses significant challenges due to the high variations and complexities arising from workloads and hardware configurations, which render offline optimization and reactive-based methods less effective. In this paper, we present a program counter (PC)-oriented prediction-based power management approach for GPGPUs. Our approach leverages the benefits of prediction to address online variations while enhancing prediction capability through knowledge transfer across different levels of resources. Experiments conducted on realistic applications demonstrate that our proposed method achieves the maximum energy savings under a user-defined performance constraint compared to state-of-the-art designs. Lin Chen 0029, Xiao Li 0038, Shixi Chen, Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
SPAA | 4 |
| 2024 | Deep Reinforcement Learning-Based Power Management for Chiplet-Based Multicore SystemsabstractChiplet technology has emerged as a promising solution to address the increasing demand for high-performance computing in light of the slowdown of Moore’s law. While chiplet-based multicore systems offer higher performance through heterogeneous integration, they also pose challenges for power delivery system (PDS) design. The integration of additional vertical and inter-chiplet connections, along with higher power density, impose stringent requirements on power delivery. Moreover, PDS efficiency is affected by workload variations at runtime, necessitating the need to design and manage PDSs and processors as a whole to improve system energy efficiency while balancing performance. In this article, we propose an offline-online co-design optimization methodology that combines offline PDS design optimization with online power management. To address the power consumption and delivery mismatch, we introduce a centralized deep Q-network (DQN)-based online control scheme for power co-management in chiplet-based multicore systems. By carefully designing the state space and reward functions, our approach achieves workload-aware adaptive control to reduce the energy-delay-product (EDP) while maintaining PDS efficiency under a given performance target (PT). We conduct evaluations on realistic applications to validate the effectiveness of our approach. For 64-core systems, our method achieves an average EDP reduction of 67% while meeting a 90% PT, surpassing state-of-the-art modular Q-learning (MQL)-based and heuristic-based approaches by up to 4% and 16%, respectively. Additionally, our approach demonstrates wiser action selection policies, higher control stability, and lower implementation overhead compared to the MQL-based approach. Xiao Li 0038, Lin Chen 0029, Shixi Chen, Fan Jiang 0015, Chengeng Li, Wei Zhang 0012, Jiang Xu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | Smart Knowledge Transfer-based Runtime Power ManagementabstractAs Moore's law slows down, computing systems must pivot towards higher energy efficiency to continue scaling performance. Reinforcement learning (RL) performs more adaptively than conventional methods in runtime power management under varied hardware configurations and varying software workloads. However, prior works on either model-free or model-based RL approaches face a non-negligible challenge: relearning the policies to adapt to the new environment is unacceptably time-consuming, especially when encountering significant variances in workloads or hardware configurations. Moreover, existing research on accelerating learning has focused on the speedup while largely ignoring the efficiency degradation of the results. In this paper, we present a smart transfer-enabled Q-learning (STQL) approach to boost the learning process and guarantee the learning efficiency through a contradiction checking mechanism, which wisely evicts inappropriate transferred knowledge. Experiments on realistic applications show that the proposed method can speed up the learning process to up to 2.3x and achieve a 6.2% energy-delay product (EDP) reduction compared to the state-of-the-art design. Lin Chen 0029, Xiao Li 0038, Fan Jiang 0015, Chengeng Li, Jiang Xu 0001 |
DATE | 3 |
| 2023 | RONet: Scaling GPU System with Silicon Photonic ChipletabstractModern GPU systems integrate hundreds of SMs on a single die, and future scaling envisions even more SMs being incorporated. However, the limited number of transistors per die constrains this growth. While current chiplet technology shows promise, its performance is limited by the bandwidth and energy efficiency of existing chiplet interconnect technologies. In contrast, optical interconnects offer ultra-high bandwidth and energy efficiency, making them ideal for high-performance chiplet-based GPUs. This work proposes a novel region-based optical network, called RONet, that divides a chiplet-based GPU system with a 2D Mesh layout into multiple row and column regions, where each region is connected by a separate optical link. Additionally, RONet employs a tuning-free transmission mechanism to further enhance inter-chiplet bandwidth. Experimental results show that RONet achieves 43% improvement on performance and 25.4% reduction on system energy consumption over the baseline. Chengeng Li, Fan Jiang 0015, Shixi Chen, Yinyi Liu, Lin Chen 0029, Xiao Li 0038, Jiang Xu 0001 |
ICCAD | 2 |
| 2022 | Accelerating Cache Coherence in Manycore Processor through Silicon Photonic ChipletabstractCache coherence overhead in manycore systems is becoming prominent with the increase of system scale. However, traditional electrical networks restrict the efficiency of cache coherence transactions in the system due to the limited bandwidth and long latency. Optical network promises high bandwidth and low latency, and supports both efficient unicast and multicast transmission, which can potentially accelerate cache coherence in manycore systems. This work proposes a novel photonic cache coherence network with a physically centralized logically distributed directory called PCCN for chiplet-based manycore systems. PCCN adopts a channel sharing method with a contention solving mechanism for efficient long-distance coherence-related packet transmission. Experiment results show that compared to state-of-the-art proposals, PCCN can speed up application execution time by 1.32x, reduce memory access latency by 26%, and improve energy efficiency by 1.26x, on average, in a 128-core system. Chengeng Li, Fan Jiang 0015, Shixi Chen, Yinyi Liu, Jiang Xu 0001 |
ICCAD | 2 |
| 2022 | Fast and Accurate Statistical Simulation of Shared-Memory Applications on Multicore SystemsabstractDetailed cycle-accurate simulation of multicore systems is naturally slow. Statistical simulation is one alternative that permits trading off simulation speed for accuracy. However, there is a lack of effective memory locality models for multicore applications. Hence, existing statistical simulators neglect data-sharing between threads. Additionally, the standard method to speed up statistical simulations is to blindly reduce the trace length to be synthesized. While this gives good control over the speedup, it leaves the simulation error unbounded. In this work, we introduce a novel statistical simulation methodology for exploration of shared-memory multicore systems. It includes a newsharing-localitymodel (Shalom) that captures and reproduces data-sharing in multithread applications. Furthermore, we propose a method to bound the simulation error for a particular metric while maximizing speedup. The technique works by monitoring the convergence of the statistical synthesis. It is referred to asconvergence-deterministicsimulation (Condens). The combination ofShalomandCondensis around 130x faster than cycle-accurate simulations with reasonable accuracy loss. Our approach is also 5x faster than state-of-the-art sampling simulation under the same accuracy level. Compared to previous statistical simulators ignoring sharing, our approach is 2x more accurate for performance metrics and 8x more accurate for cache miss estimations. Fan Jiang 0015, Rafael Kioji Vivas Maeda, Jun Feng 0008, Shixi Chen, Lin Chen 0029, Xiao Li 0038, Jiang Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |