VLDB 2026 Research / reviewers in the wild / expert
Jia Zhan
dblp:130/1362
· DBLP profile ↗
13ranked-venue papers
9as first author
2since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 8 first-authorComputer networks · 3 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
3 papers |
Physical-layer communications · 78% Cellular and mobile networks · 14% Internet of things and sensor networks · 8% | |
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Memory systems · 38% Energy-efficient computing · 19% Interconnection networks and networks-on-chip · 18% | |
| Theoretical computer science
2 papers |
Coding theory · 64% Information theory · 36% |
Topics — the 30 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Physical-layer communications › modulation › coded modulation
bit-interleaved coded modulation |
1.0 | 1 | 2026 | Analysis and Design of Irregular-Mapped BICM-ID System With Non-Uniform Sources · IEEE Trans. Commun. 2026 |
Physical-layer communications
channel coding |
1.0 | 1 | 2026 | Analysis and Design of Irregular-Mapped BICM-ID System With Non-Uniform Sources · IEEE Trans. Commun. 2026 |
Physical-layer communications › channel coding › decoding algorithms › iterative decoding
density evolution |
1.0 | 1 | 2026 | Analysis and Design of Irregular-Mapped BICM-ID System With Non-Uniform Sources · IEEE Trans. Commun. 2026 |
Cellular and mobile networks
integrated sensing and communication |
1.0 | 1 | 2026 | Advanced Spatial Modulation-Aided Integrated Sensing and Backscatter Communication: System Design and Performance Analysis · IEEE Trans. Commun. 2026 |
Physical-layer communications › MIMO › spatial modulation
quadrature spatial modulation |
1.0 | 1 | 2026 | Advanced Spatial Modulation-Aided Integrated Sensing and Backscatter Communication: System Design and Performance Analysis · IEEE Trans. Commun. 2026 |
Physical-layer communications › MIMO
spatial modulation |
1.0 | 1 | 2026 | Advanced Spatial Modulation-Aided Integrated Sensing and Backscatter Communication: System Design and Performance Analysis · IEEE Trans. Commun. 2026 |
Information theory › signal processing › modulation
probabilistic shaping |
1.0 | 1 | 2026 | Analysis and Design of Irregular-Mapped BICM-ID System With Non-Uniform Sources · IEEE Trans. Commun. 2026 |
Coding theory
source coding |
1.0 | 1 | 2026 | Analysis and Design of Irregular-Mapped BICM-ID System With Non-Uniform Sources · IEEE Trans. Commun. 2026 |
Interconnection networks and networks-on-chip › network-on-chip design
energy-efficient noc |
0.6 | 3 | 2015 | DimNoC: a dim silicon approach towards power-efficient on-chip network · DAC 2015 NoC-Sprinting: Interconnect for Fine-Grained Sprinting in the Dark Silicon Era · DAC 2014 Designing energy-efficient NoC for real-time embedded systems through slack optimization · DAC 2013 |
Energy-efficient computing › energy-constrained computing
dark silicon |
0.5 | 3 | 2015 | Core vs. uncore: the heart of darkness · DAC 2015 NoC-Sprinting: Interconnect for Fine-Grained Sprinting in the Dark Silicon Era · DAC 2014 DimNoC: a dim silicon approach towards power-efficient on-chip network · DAC 2015 |
Internet of things and sensor networks
backscatter communication |
0.3 | 1 | 2026 | Advanced Spatial Modulation-Aided Integrated Sensing and Backscatter Communication: System Design and Performance Analysis · IEEE Trans. Commun. 2026 |
Internet of things and sensor networks › RFID systems
tag identification |
0.3 | 1 | 2026 | Advanced Spatial Modulation-Aided Integrated Sensing and Backscatter Communication: System Design and Performance Analysis · IEEE Trans. Commun. 2026 |
Physical-layer communications › spread spectrum
chaos-based communication |
0.3 | 1 | 2017 | A Differential Chaotic Bit-Interleaved Coded Modulation System Over Multipath Rayleigh Channels · IEEE Trans. Commun. 2017 |
Physical-layer communications › spread spectrum › chaos-based communication
differential chaos shift keying |
0.3 | 1 | 2017 | A Differential Chaotic Bit-Interleaved Coded Modulation System Over Multipath Rayleigh Channels · IEEE Trans. Commun. 2017 |
Coding theory › error-correcting codes › coded modulation
bit-interleaved coded modulation |
0.3 | 1 | 2017 | A Differential Chaotic Bit-Interleaved Coded Modulation System Over Multipath Rayleigh Channels · IEEE Trans. Commun. 2017 |
Coding theory › error-correcting codes
coded modulation |
0.3 | 1 | 2017 | A Differential Chaotic Bit-Interleaved Coded Modulation System Over Multipath Rayleigh Channels · IEEE Trans. Commun. 2017 |
Memory systems
cache |
0.2 | 1 | 2016 | OSCAR: Orchestrating STT-RAM cache traffic for heterogeneous CPU-GPU architectures · MICRO 2016 |
Memory systems
in-memory computing |
0.2 | 1 | 2016 | A unified memory network architecture for in-memory computing in commodity servers · MICRO 2016 |
Memory systems › memory disaggregation
memory network |
0.2 | 1 | 2016 | A unified memory network architecture for in-memory computing in commodity servers · MICRO 2016 |
Memory systems
non-volatile memory |
0.2 | 1 | 2015 | DimNoC: a dim silicon approach towards power-efficient on-chip network · DAC 2015 |
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM |
0.2 | 1 | 2015 | DimNoC: a dim silicon approach towards power-efficient on-chip network · DAC 2015 |
Processor architecture and microarchitecture
uncore components |
0.2 | 1 | 2015 | Core vs. uncore: the heart of darkness · DAC 2015 |
Embedded and real-time systems › real-time embedded systems
hard real-time systems |
0.2 | 1 | 2014 | Optimizing the NoC Slack Through Voltage and Frequency Scaling in Hard Real-Time Embedded Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014 |
Energy-efficient computing
voltage and frequency scaling |
0.2 | 1 | 2014 | Optimizing the NoC Slack Through Voltage and Frequency Scaling in Hard Real-Time Embedded Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2014 |
Embedded and real-time systems
real-time scheduling |
0.2 | 1 | 2013 | Designing energy-efficient NoC for real-time embedded systems through slack optimization · DAC 2013 |
Embedded and real-time systems › real-time analysis
worst-case delay analysis |
0.2 | 1 | 2013 | Designing energy-efficient NoC for real-time embedded systems through slack optimization · DAC 2013 |
Coding theory › error-correcting codes
LDPC codes |
0.1 | 1 | 2017 | A Differential Chaotic Bit-Interleaved Coded Modulation System Over Multipath Rayleigh Channels · IEEE Trans. Commun. 2017 |
Coding theory › error-correcting codes › LDPC codes
protograph LDPC codes |
0.1 | 1 | 2017 | A Differential Chaotic Bit-Interleaved Coded Modulation System Over Multipath Rayleigh Channels · IEEE Trans. Commun. 2017 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
CPU-GPU heterogeneous architecture |
0.1 | 1 | 2016 | OSCAR: Orchestrating STT-RAM cache traffic for heterogeneous CPU-GPU architectures · MICRO 2016 |
Memory systems
DRAM |
0.1 | 1 | 2016 | A unified memory network architecture for in-memory computing in commodity servers · MICRO 2016 |
Methods — techniques the papers use, named apart from their topics
gaussian mixture approximation · 2.0density evolution · 2.0achievable rate analysis · 2.0compressive sensing · 1.0bit error probability analysis · 1.0PEXIT analysis · 0.6network calculus · 0.4bit-error rate simulation · 0.3bit error rate simulation · 0.3simulation · 0.2priority-based allocation · 0.2power gating · 0.2event-driven simulation · 0.2distance-aware selective compression · 0.2asynchronous batch scheduling · 0.2non-volatile memory · 0.2drowsy SRAM · 0.23d integration · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Analysis and Design of Irregular-Mapped BICM-ID System With Non-Uniform SourcesabstractBit-interleaved coded modulation with iterative decoding (BICM-ID) is a key technology enabling communication systems to realize high-throughput and high-reliability. In contrast to previous studies focused exclusively on uniformly distributed sources, this work investigates irregular mapping (IM)-based BICM-ID systems for sources with non-uniform distributions. First, a density evolution (DE) algorithm based on Gaussian weighted mixture approximation (GWMA) is put forward to estimate the convergence performance of the IM-BICM-ID systems with non-uniform sources, where the probability density function of the log-likelihood ratio messages output from the bit-level channels in this system is neither Gaussian nor symmetric. Then, by employing the GWMA-DE algorithm and the achievable system rates analysis, a design strategy for optimal IM labelings is established, which utilizes the inherent redundancy in non-uniform sources to achieve shaping gains. Furthermore, channel codes tailored to the proposed IM labelings are designed to further improve the bit error rate performance. Both analytical results and simulations indicate that the proposed labelings and codes outperform existing counterparts in IM-BICM-ID systems with non-uniform sources. Chen Chen 0060, Jiaofan Cai, Qiwang Chen, Jia Zhan, Yi Fang 0005 |
IEEE Trans. Commun. | 4 |
| 2026 | Advanced Spatial Modulation-Aided Integrated Sensing and Backscatter Communication: System Design and Performance AnalysisabstractThis paper proposes an innovative framework that integrates electromagnetic inverse scattering with improved quadrature spatial modulation (IQSM) to simultaneously accomplish sensing, identification, and backscatter communication. Specifically tailored for energy- and spectrum-efficient wireless operations in highly cluttered environments, the proposed system employs distinct load impedance modulation at the tags to effectively separate structural and antenna mode signatures. Structural mode signals are processed using inverse scattering combined with compressive sensing techniques, enabling precise localization of both tags and surrounding clutter. Concurrently, antenna mode signals are utilized for accurate tag identification. In addition, the antenna mode enables the acquisition of the reliable channel state information (CSI), facilitating the integration of IQSM schemes into the backscatter communication module. Furthermore, we present a theoretical analysis by deriving the bit error probability (BEP) for the proposed system. The proposed system is validated through a proof-of-concept experimental setup consisting of transmit and receive arrays, each configured as a 3×3 uniform linear antenna array (ULA). Both simulation and experimental results confirm that the inverse scattering-based approach achieves high-precision sensing and accurate identification of tags individually or in combination. Additionally, the integration of IQSM significantly enhances spectral efficiency (SE) and data throughput, compared with the state-of-the-art systems that do not incorporate spatial modulation (SM) techniques. These findings highlight the effectiveness and practical viability of the proposed integrated sensing and communication system for challenging clutter-rich scenarios. Dingfei Ma, Jia Zhan, Chi Zhang 0111, Yi Fang 0005 |
IEEE Trans. Commun. | 3 |
| 2017 | A Differential Chaotic Bit-Interleaved Coded Modulation System Over Multipath Rayleigh ChannelsabstractIn this paper, a novel differential chaotic bit-interleaved coded modulation (DC-BICM) system is proposed for band-limited transmission. This system combines protograph-based low density parity check codes with constellation-based M-ary differential chaos shift keying (DCSK) modulation by one bitwise interleaving. Bit error rate simulation results show that the system has higher coding gain compared with the constellation-based M-ary DCSK modulation system with the same spectral efficiency over multipath Rayleigh fading channels. At the same time, several simulations and P-EXIT analysis are used to analyze the performance of the proposed system. It is found that there is a lot of room for optimization of the system by comparing decoding thresholds and simulation results. Moreover, the system with only partial channel state information has better performance and lower complexity compared with the traditional bit-interleaved coded modulation (BICM) directsequence-spread-spectrum system. As a result, the DC-BICM system is a good candidate for band-limited transmission. Jia Zhan, Lin Wang 0003, Marcos D. Katz, Guanrong Chen |
IEEE Trans. Commun. | 1 |
| 2016 | Scalable memory fabric for silicon interposer-based multi-core systemsabstractThree-dimensional (3D) integration is considered as a solution to overcome capacity, bandwidth, and performance limitations of memories. However, due to thermal challenges and cost issues, industry embraced 2.5D implementation for integrating die-stacked memories with large-scale designs, which is enabled by silicon interposer technology that integrates processors and multiple modules of 3D-stacked memories in the same package. Previous work has adopted Network-on-Chip (NoC) concepts for the communication fabric of 3D designs, but the design of a scalable processor-memory interconnect for 2.5D integration remains elusive. Therefore, in this work, we first explore different network topologies for integrating CPUs and memories in a silicon interposer-based multi-core system and reveal that simple point-to-point connections cannot reach the full potential of the memory performance due to bandwidth limitations, especially as more and more memory modules are needed to enable emerging applications with high memory capacity and bandwidth demand, such as in-memory computing. To overcome this scaling problem, we propose a memory network design to directly connect all the memory modules, utilizing the existing routing resource of silicon interposers in 2.5D designs. Observing the unique network traffic in our design, we present a design space exploration that evaluates network topologies and routing algorithms, taking process node and interposer technology design decisions into account. We implement an event-driven simulator to evaluate our proposed memory network in silicon interposer (MemNiSI) design with synthetic traffic as well as real in-memory computing workloads. Our experimental results show that compared to baseline designs, MemNiSI topology reduces the average packet latency by up to 15.3% and Choose Fastest Path (CFP) algorithm further reduces by up to 8.0%. Our scheme can utilize the potential of integrated stacked memory effectively while providing better scalability and infrastructure for large-scale silicon interposer-based 2.5D designs. Itir Akgun, Jia Zhan, Yuangang Wang, Yuan Xie 0001 |
ICCD | 2 |
| 2016 | A unified memory network architecture for in-memory computing in commodity serversabstractIn-memory computing is emerging as a promising paradigm in commodity servers to accelerate data-intensive processing by striving to keep the entire dataset in DRAM. To address the tremendous pressure on the main memory system, discrete memory modules can be networked together to form a memory pool, enabled by recent trends towards richer memory interfaces (e.g. Hybrid Memory Cubes, or HMCs). Such an inter-memory network provides a scalable fabric to expand memory capacity, but still suffers from long multi-hop latency, limited bandwidth, and high power consumption — problems that will continue to exacerbate as the gap between interconnect and transistor performance grows. Moreover, inside each memory module, an intra-memory network (NoC) is typically employed to connect different memory partitions. Without careful design, the back-pressure inside the memory modules can further propagate to the inter-memory network to cause a performance bottleneck. To address these problems, we propose co-optimization of intra- and inter-memory network. First, we re-organize the intra-memory network structure, and provide a smart I/O interface to reuse the intra-memory NoC as the network switches for inter-memory communication, thus forming a unified memory network. Based on this architecture, we further optimize the inter-memory network for both high performance and lower energy, including a distance-aware selective compression scheme to drastically reduce communication burden, and a light-weight power-gating algorithm to turn off under-utilized links while guaranteeing a connected graph and deadlock-free routing. We develop an event-driven simulator to model our proposed architectures. Experiment results based on both synthetic traffic and real big-data workloads show that our unified memory network architecture can achieve 75.1% average memory access latency reduction and 22.1% total memory energy saving. Jia Zhan, Itir Akgun, Jishen Zhao, Al Davis, Paolo Faraboschi, Yuangang Wang, Yuan Xie 0001 |
MICRO | 1 |
| 2016 | OSCAR: Orchestrating STT-RAM cache traffic for heterogeneous CPU-GPU architecturesabstractAs we integrate data-parallel GPUs with general-purpose CPUs on a single chip, the enormous cache traffic generated by GPUs will not only exhaust the limited cache capacity, but also severely interfere with CPU requests. Such heterogeneous multicores pose significant challenges to the design of shared last-level cache (LLC). This problem can be mitigated by replacing SRAM LLC with emerging non-volatile memories like Spin-Transfer Torque RAM (STT-RAM), which provides larger cache capacity and near-zero leakage power. However, without careful design, the slow write operations of STT-RAM may offset the capacity benefit, and the system may still suffer from contention in the shared LLC and on-chip interconnects. While there are cache optimization techniques to alleviate such problems, we reveal that the true potential of STT-RAM LLC may still be limited because now that the cache hit rate has been improved by the increased capacity, the on-chip network can become a performance bottleneck. CPU and GPU packets contend with each other for the shared network bandwidth. Moreover, the mixed-criticality read/write packets to STT-RAM add another layer of complexity to the network resource allocation. Therefore, being aware of the disparate latency tolerance of CPU/GPU applications and the asymmetric read/write latency of STT-RAM, we propose OSCAR to Orchestrate STT-RAM Caches traffic for heterogeneous ARchitectures. Specifically, an integration of asynchronous batch scheduling and priority based allocation for on-chip interconnect is proposed to maximize the potential of STT-RAM based LLC. Simulation results on a 28-GPU and 14-CPU system demonstrate an average of 17.4% performance improvement for CPUs, 10.8% performance improvement for GPUs, and 28.9% LLC energy saving compared to SRAM based LLC design. Jia Zhan, Onur Kayiran, Gabriel H. Loh, Chita R. Das, Yuan Xie 0001 |
MICRO | 1 |
| 2016 | Hybrid Drowsy SRAM and STT-RAM Buffer Designs for Dark-Silicon-Aware NoCabstractThe breakdown of Dennard scaling prevents us from powering all transistors simultaneously, leaving a large fraction of dark silicon. This crisis has led to innovative work on power-efficient core and memory architecture designs. However, the research for addressing dark silicon challenges with network-on-chip (NoC), which is a major contributor to the total chip power consumption, is largely unexplored. In this paper, we comprehensively examine the network power consumers and the drawbacks of the conventional power-gating techniques. To overcome the dark silicon issue from the NoC's perspective, we propose DimNoC, a dim silicon scheme, which leverages recent drowsy SRAM design and spin-transfer torque RAM (STT-RAM) technology to replace pure SRAM-based NoC buffers. In particular, we propose two novel hybrid buffer architectures: 1) a hierarchical buffer architecture, which divides the input buffers into a set of levels with different power states and 2) a banked buffer architecture, which organizes the drowsy SRAM and the STT-RAM in different banks, and accesses them in an interleaved fashion to hide the long write latency of STT-RAM. In addition, our hybrid buffer design enables NoC data retention mechanism by storing packets in drowsy SRAM and nonvolatile STT-RAM in a lossless manner. Combined with flow control schemes, the NoC data retention mechanism can improve network performance and power simultaneously. Our experiments over real workloads show that DimNoC can achieve 30.9% network energy saving, 20.3% energy-delay product reduction, and 7.6% router area reduction compared with pure SRAM-based NoC design. Jia Zhan, Jin Ouyang, Fen Ge, Jishen Zhao, Yuan Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | Core vs. uncore: the heart of darknessabstractEven though Moore's Law continues to provide increasing transistor counts, the rise of the utilization wall limits the number of transistors that can be powered on and results in a large region of dark silicon. Prior studies have proposed energy-efficient core designs to address the "dark silico" problem. Nevertheless, the research for addressing dark silicon challenges in uncore components, such as shared cache, on-chip interconnect, etc, that contribute significant on-chip power consumption is largely unexplored. In this paper, we first illustrate that the power consumption of uncore components cannot be ignored to meet the chip's power constraint. We then introduce techniques to design energy-efficient uncore components, including shared cache and on-chip interconnect. The design challenges and opportunities to exploit 3D techniques and non-volatile memory (NVM) in dark-silicon-aware architecture are also discussed. Hsiang-Yun Cheng, Jia Zhan, Jishen Zhao, Yuan Xie 0001, Jack Sampson, Mary Jane Irwin |
DAC | 2 |
| 2015 | DimNoC: a dim silicon approach towards power-efficient on-chip networkabstractThe diminishing momentum of Dennard scaling leads to the ever increasing power density of integrated circuits, and a decreasing portion of transistors on a chip that can be switched on simultaneously---a problem recently discovered and known as dark silicon. There has been innovative work to address the "dark silicon" problem in the fields of power-efficient core and cache system. However, dark silicon challenges with Network-on-Chip (NoC) are largely unexplored. To address this issue, we propose DimNoC, a "dim silicon" approach, which leverages drowsy SRAM and STT-RAM technologies to replace pure SRAM-based NoC buffers. Specifically, we propose two novel hybrid buffer architectures: 1) a Hierarchical Buffer (HB) architecture, which divides the input buffers into a hierarchy of levels with different memory technologies operating at various power states; 2) a Banked Buffer (BB) architecture, which organizes drowsy SRAM and STT-RAM into separate banks in order to hide the long write-latency of STT-RAM. Our experiments show that the proposed DimNoC can achieve 30.9% network energy saving, 20.3% energy-delay product (EDP) reduction, and 7.6% router area decrease compared with the baseline SRAM-based NoC design. Jia Zhan, Jin Ouyang, Fen Ge, Jishen Zhao, Yuan Xie 0001 |
DAC | 1 |
| 2014 | NoΔ: Leveraging delta compression for end-to-end memory access in NoC based multicoresabstractAs the number of on-chip processing elements increases, the interconnection backbone bears bursty traffic from memory and cache accesses. In this paper, we propose a compression technique called NoΔ, which leverages delta compression to compress network traffic. Specifically, it conducts data encoding prior to packet injection and decoding before ejection in the network interface. The key idea of NoΔ is to store a data packet in the Network-on-Chip as a common base value plus an array of relative differences (Δ). It can improve the overall network performance and achieve energy savings because of the decreased network load. Moreover, this scheme does not require modifications of the cache storage design and can be seamlessly integrated with any optimization techniques for the on-chip interconnect. Our experiments reveal that the proposed NoΔ incurs negligible hardware overhead and outperforms state-of-the-art zero-content compression and frequent-value compression. Jia Zhan, Matthew Poremba, Yuan Xie 0001 |
ASP-DAC | 1 |
| 2014 | NoC-Sprinting: Interconnect for Fine-Grained Sprinting in the Dark Silicon EraabstractThe rise of utilization wall limits the number of transistors that can be powered on in a single chip and results in a large region of dark silicon. While such phenomenon has led to disruptive innovation in computation, little work has been done for the Network-on-Chip (NoC) design. NoC not only directly influences the overall multi-core performance, but also consumes a significant portion of the total chip power. In this paper, we first reveal challenges and opportunities of designing power-efficient NoC in the dark silicon era. Then we propose NoC-Sprinting: based on the workload characteristics, it explores fine-grained sprinting that allows a chip to flexibly activate dark cores for instantaneous throughput improvement. In addition, it investigates topological/routing support and thermal-aware floorplanning for the sprinting process. Moreover, it builds an efficient network power-management scheme that can mitigate the dark silicon problems. Experiments on performance, power, and thermal analysis show that NoC-sprinting can provide tremendous speedup, increase sprinting duration, and meanwhile reduce the chip power significantly. Jia Zhan, Yuan Xie 0001, Guangyu Sun 0003 |
DAC | 1 |
| 2014 | Optimizing the NoC Slack Through Voltage and Frequency Scaling in Hard Real-Time Embedded SystemsabstractHard real-time embedded systems impose a strict latency requirement on interconnection subsystems. In the case of network-on-chip (NoC), this means each packet of a traffic stream has to be delivered within a time interval. In addition, with the increasing complexity of NoC, it consumes a significant portion of total chip power, which boosts the power footprint of such chips. In this paper, we propose a methodology to minimize the energy consumption of NoC without violating the prespecified latency deadlines of real-time applications. First, we develop a formal approach based on network calculus to obtain the worst-case delay bound of all packets, from which we derive a safe estimate of the number of cycles that a packet can be further delayed in the network without violating its deadline-the worst-case slack. With this information, we then develop an optimization algorithm that trades the slacks for lower NoC energy. Our algorithm recognizes the distribution of slacks for different traffic streams, and assigns different voltages and frequencies to different routers to achieve NoC energy-efficiency, while meeting the deadlines for all packets. Furthermore, we design a feedback-control strategy to enable dynamic frequency and voltage scaling on the network routers in conjunction with the energy optimization algorithm. It can flexibly improve the energy-efficiency of the overall network in response to sporadic traffic patterns at runtime. Jia Zhan, Nikolay Stoimenov, Jin Ouyang, Lothar Thiele, Narayanan Vijaykrishnan, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Designing energy-efficient NoC for real-time embedded systems through slack optimizationabstractHard real-time embedded systems impose a strict latency requirement on interconnection subsystems. In the case of network-on-chip (NoC), this means each packet of a traffic stream has to be delivered within a time interval. In addition, with the increasing complexity of NoC, it consumes a significant portion of total chip power, which boosts the power footprint of such chips. In this work, we propose a methodology to minimize the energy consumption of NoC without violating the pre-specified latency deadlines of real-time applications. First, we develop a formal approach based on network calculus to obtain the worst-case delay bound of all packets, from which we derive a safe estimate of the number of cycles that a packet can be further delayed in the network without violating its deadline---the worst-case slack. With this information, we then develop an optimization algorithm that trades the slacks for lower NoC energy. Our algorithm recognizes the distribution of slacks for different traffic streams, and assigns different voltages and frequencies to different routers to achieve NoC energy-efficiency, while meeting the deadlines for all packets. Jia Zhan, Nikolay Stoimenov, Jin Ouyang, Lothar Thiele, Narayanan Vijaykrishnan, Yuan Xie 0001 |
DAC | 1 |