EDBT 2026 Demo / reviewers in the wild / expert
Chenjie Yu
dblp:99/2186
· DBLP profile ↗
9ranked-venue papers
5as first author
0since 2021 · last 2012
0000-0002-6734-6247ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 68% Processor architecture and microarchitecture · 26% Hardware reliability and fault tolerance · 6% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › cache management
cache partitioning |
0.1 | 1 | 2010 | Off-chip memory bandwidth minimization through cache partitioning for multi-core platforms · DAC 2010 |
Memory systems › memory bandwidth management
memory bandwidth reduction |
0.1 | 1 | 2010 | Off-chip memory bandwidth minimization through cache partitioning for multi-core platforms · DAC 2010 |
Compilers and program optimization
register allocation |
0.1 | 1 | 2008 | Compiler-driven register re-assignment for register file power-density and temperature reduction · DAC 2008 |
Memory systems
cache coherence |
0.1 | 1 | 2008 | Latency and bandwidth efficient communication through system customization for embedded multiprocessors · DAC 2008 |
Processor architecture and microarchitecture › chip multiprocessor
inter-core communication |
0.1 | 1 | 2008 | Latency and bandwidth efficient communication through system customization for embedded multiprocessors · DAC 2008 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2010 | Off-chip memory bandwidth minimization through cache partitioning for multi-core platforms · DAC 2010 |
Compilers and program optimization
program transformation |
0.0 | 1 | 2008 | Latency and bandwidth efficient communication through system customization for embedded multiprocessors · DAC 2008 |
Hardware reliability and fault tolerance › reliability analysis
thermal reliability |
0.0 | 1 | 2008 | Compiler-driven register re-assignment for register file power-density and temperature reduction · DAC 2008 |
Methods — techniques the papers use, named apart from their topics
cross-layer customization · 0.2compiler-driven code transformation · 0.2algorithmic heuristic · 0.2NP-hardness proof · 0.2application-driven partitioning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Evaluating Power-Monitoring Capabilities on IBM Blue Gene/P and Blue Gene/QabstractPower consumption is becoming a critical factor as we continue our quest toward exascale computing. Yet, actual power utilization of a complete system is an insufficiently studied research area. Estimating the power consumption of a large scale system is a nontrivial task because a large number of components are involved and because power requirements are affected by the (unpredictable) workloads. Clearly needed is a power-monitoring infrastructure that can provide timely and accurate feedback to system developers and application writers so that they can optimize the use of this precious resource. Many existing large-scale installations do feature power-monitoring sensors, however, those are part of environmental- and health monitoring sub systems and were not designed with application level power consumption measurements in mind. In this paper, we evaluate the existing power monitoring of IBM Blue Gene systems, with the goal of understanding what capabilities are available and how they fare with respect to spatial and temporal resolution, accuracy, latency, and other characteristics. We find that with a careful choice of dedicated micro benchmarks, we can obtain meaningful power consumption data even on Blue Gene/P, where the interval between available data points is measured in minutes. We next evaluate the monitoring subsystem on Blue Gene/Q, and are able to study the power characteristics of FPU and memory subsystems of Blue Gene/Q. We find the monitoring subsystem capable of providing second-scale resolution of power data conveniently separated between node components with seven seconds latency. This represents a significant improvement in power monitoring infrastructure, and hope future systems will enable real-time power measurement in order to better understand application behavior at a finer granularity. Kazutomo Yoshii, Kamil Iskra, Rinku Gupta, Pete Beckman, Venkatram Vishwanath, Chenjie Yu, Susan Coghlan |
CLUSTER | 6 |
| 2010 | Off-chip memory bandwidth minimization through cache partitioning for multi-core platformsabstractWe present a methodology for off-chip memory bandwidth minimization through application-driven L2 cache partitioning in multi-core systems. A major challenge with multi-core system design is the widening gap between the memory demand generated by the processor cores and the limited off-chip memory bandwidth and memory service speed. This severely restricts the number of cores that can be integrated into a multi-core system and the parallelism that can be actually achieved and efficiently exploited for not only memory demanding applications, but also for workloads consisting of many tasks utilizing a large number of cores and thus exceeding the available off-chip bandwidth. Chenjie Yu, Peter Petrov |
DAC | 1 |
| 2010 | Energy- and Performance-Efficient Communication Framework for Embedded MPSoCs through Application-Driven Release ConsistencyabstractWe present a framework for performance-, bandwidth-, and energy-efficient intercore communication in embedded MultiProcessor Systems-on-a-Chip (MPSoC). The methodology seamlessly integrates compiler, operating system, and hardware support to achieve a low-cost communication between synchronized producers and consumers. The technique is especially beneficial for data-streaming applications exploiting pipeline parallelism with computational phases mapped to separate cores. Code transformations utilizing a simple ISA support ensure that producer writes are propagated to consumers with a single interconnect transaction per cache block just prior to the producer exiting its synchronization region. Furthermore, in order to completely eliminate misses to shared data caused by interference with private data and also to minimize the cache energy, we integrate to the proposed framework a cache way partitioning policy based on a simple cache configurability support, which isolates the shared buffers from other cache traffic. This mechanism results in significant power savings since only a subset of the cache ways needs to be looked up for each cache access. The end result of the proposed framework is a single communication transaction per shared cache block between a producer and a consumer with no coherence misses on the consumer caches. Our experiments demonstrate significant reductions in interconnect traffic, cache misses, and energy for a set of multiprocessor benchmarks. Chenjie Yu, Peter Petrov |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2010 | Low-Cost and Energy-Efficient Distributed Synchronization for Embedded MultiprocessorsabstractWe present a framework for a distributed and lowcost implementation of synchronization mechanisms for embedded shared-memory multiprocessors. The proposed architecture effectively implements the queued-lock semantics in a completely decentralized manner through low-cost and distributed synchronization controllers performing distributed synchronization management protocols. The proposed approach achieves three major benefits. First, it completely eliminates the overwhelming bus contention traffic when multiple cores compete for a synchronization variable. Second, it exhibits extremely low best-case latency of lock acquisition (with zero bus transactions). Third, the approach enables multiple venues for high energy efficiency as the local synchronization controllers can efficiently determine, without any bus transactions or local cache spinning, the exact timing of when a lock is made available to or a barrier enabled at the local processor. It becomes possible for the system software or the thread library to employ various low-power policies. Chenjie Yu, Peter Petrov |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | Temperature-aware register reallocation for register file power-density minimizationabstractIncreased chip temperature has been known to cause severe reliability problems and to significantly increase leakage power. The register file has been previously shown to exhibit the highest temperature compared to all other hardware components in a modern high-end embedded processor, which makes it particularly susceptible to faults and elevated leakage power. We show that this is mostly due to the highly clustered register file accesses where a set of few registers physically placed close to each other are accessed with very high frequency. We propose compile-time temperature-aware register reallocation methodologies for breaking such groups of registers and to uniformly distribute the accesses to the register file. This is achieved with no performance and no hardware overheads . We show that the underlying problem is NP-hard, and subsequently introduce and evaluate two efficient algorithmic heuristics. Our extensive experimental study demonstrates the efficiency of the proposed methodology. Xiangrong Zhou, Chenjie Yu, Peter Petrov |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2009 | Low-Power Snoop Architecture for Synchronized Producer-Consumer Embedded MultiprocessingabstractWe introduce a cross-layer customization methodology where application knowledge regarding data sharing in producer-consumer relationships is used in order to aggressively eliminate unnecessary and predictable snoop-induced cache lookups even for references to shared data, thus, achieving significant power reductions with minimal hardware cost. The technique exploits application-specific information regarding the exactproducer-consumerrelationshipsbetween tasks as well as information regardingtheprecisetimingofsynchronizedaccessesto shared memory buffers by their corresponding producers and/or consumers. Snoop-induced cache lookups for accesses to the shared data are eliminated when it is ensured that such lookups will not result in extra knowledge regarding the cache state in respect to the other caches and the memory. Our experiments show average power reductions of more than 80% compared to a general-purpose snoop protocol. Chenjie Yu, Peter Petrov |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2008 | Latency and bandwidth efficient communication through system customization for embedded multiprocessorsabstractWe present a cross-layer customization methodology for latency and bandwidth efficient inter-core communication in embedded multiprocessors. The methodology integrates compiler, operating system, and hardware support to achieve a bandwidth efficient, snoop-free, and coherence cache miss-free shared memory communication between synchronized producer and consumers cores. A compiler-driven code transformation is introduced that utilizes a simple ISA support in the form of a special write-through store instruction. It ensures that producer writes are propagated to the consumers with a single bus transaction per cache block when the producer performs the last write to that cache line before exiting its synchronization region. Information regarding the shared buffers involved in the communications is captured by the OS and provided to the cores with the purpose of filtering bus traffic and performing remote updates when necessary. The end result of the proposed methodology is a single bus transaction per shared cache block and snoop-free communication between a producer and a set of consumers with no intervening coherence misses on the consumer caches. Our experiments demonstrate the significant reductions in both bus traffic and cache misses for a set of multiprocessor benchmarks. Chenjie Yu, Peter Petrov |
DAC | 1 |
| 2008 | Compiler-driven register re-assignment for register file power-density and temperature reductionabstractTemperature hot-spots have been known to cause severe reliability problems and to significantly increase leakage power. The register file has been previously shown to exhibit the highest temperature compared to all other hardware components in a modern high-end embedded processor, which makes it particularly susceptible to faults and elevated leakage power. We show that this is mostly due to the highly clustered register file accesses where a set of few registers physically placed close to each other are accessed with very high frequency. In this paper we propose a compiler-based register reassignment methodology, which purpose is to break such groups of registers and to uniformly distribute the accesses to the register file. This is achieved with no performance and no hardware overheads. We show that the underlying problem is NP-hard, and subsequently introduce an efficient algorithmic heuristic. Xiangrong Zhou, Chenjie Yu, Peter Petrov |
DAC | 2 |
| 2008 | Application-aware snoop filtering for low-power cache coherence in embedded multiprocessorsabstractMaintaining local caches coherently in shared-memory multiprocessors results in significant power consumption. The customization methodology we propose exploits the fact that in embedded systems, important knowledge is available to the system designers regarding memory sharing between tasks. We demonstrate how the snoop-induced cache probings can be significantly reduced by identifying and exploiting in a deterministic way the shared memory regions between the processors. Snoop activity is enabled only for the accesses referring to known shared regions. The hardware support is not only cost efficient, but also software programmable, which allows for reprogrammability and customization across different tasks and applications. Xiangrong Zhou, Chenjie Yu, Alokika Dash, Peter Petrov |
ACM Trans. Design Autom. Electr. Syst. | 2 |