VLDB 2026 Research / reviewers in the wild / expert
Wei Wu 0024
dblp:95/6985-24
· DBLP profile ↗
14ranked-venue papers
4as first author
0since 2021 · last 2012
0000-0003-0401-7363ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 4 first-authorSoftware engineering, systems software and programming languages · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Hardware reliability and fault tolerance · 45% Energy-efficient computing · 30% Memory systems · 15% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware reliability and fault tolerance › error correction
error-correcting codes |
0.3 | 3 | 2011 | Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011 Reducing cache power with low-cost, multi-bit error-correcting codes · ISCA 2010 Improving cache lifetime reliability at ultra-low voltages · MICRO 2009 |
Hardware reliability and fault tolerance › memory reliability
cache reliability |
0.2 | 2 | 2011 | Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011 Improving cache lifetime reliability at ultra-low voltages · MICRO 2009 |
Energy-efficient computing
voltage scaling |
0.2 | 2 | 2011 | Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011 Adaptive Cache Design to Enable Reliable Low-Voltage Operation · IEEE Trans. Computers 2011 |
Memory systems
cache design |
0.1 | 1 | 2011 | Adaptive Cache Design to Enable Reliable Low-Voltage Operation · IEEE Trans. Computers 2011 |
Hardware reliability and fault tolerance
error correction |
0.1 | 1 | 2011 | Adaptive Cache Design to Enable Reliable Low-Voltage Operation · IEEE Trans. Computers 2011 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.1 | 1 | 2010 | Reducing cache power with low-cost, multi-bit error-correcting codes · ISCA 2010 |
Memory systems › DRAM › DRAM refresh
eDRAM refresh |
0.1 | 1 | 2010 | Reducing cache power with low-cost, multi-bit error-correcting codes · ISCA 2010 |
Hardware reliability and fault tolerance › error correction
multi-bit error correction |
0.1 | 1 | 2010 | Reducing cache power with low-cost, multi-bit error-correcting codes · ISCA 2010 |
Hardware reliability and fault tolerance
soft errors |
0.1 | 2 | 2011 | Energy-efficient cache design using variable-strength error-correcting codes · ISCA 2011 Improving cache lifetime reliability at ultra-low voltages · MICRO 2009 |
Energy-efficient computing › thermal management
dynamic thermal management |
0.1 | 1 | 2006 | Fast Thermal Simulation for Runtime Temperature Tracking and Management · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Energy-efficient computing
power modeling |
0.1 | 1 | 2006 | A systematic method for functional unit power estimation in microprocessors · DAC 2006 |
Performance modeling and evaluation
simulation |
0.1 | 1 | 2006 | Fast Thermal Simulation for Runtime Temperature Tracking and Management · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Energy-efficient computing
thermal management |
0.1 | 1 | 2006 | Fast Thermal Simulation for Runtime Temperature Tracking and Management · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Performance modeling and evaluation › simulation
thermal simulation |
0.1 | 1 | 2006 | Fast Thermal Simulation for Runtime Temperature Tracking and Management · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2006 |
Energy-efficient computing
low-voltage operation |
0.0 | 1 | 2011 | Adaptive Cache Design to Enable Reliable Low-Voltage Operation · IEEE Trans. Computers 2011 |
Energy-efficient computing › memory energy efficiency
low-power cache design |
0.0 | 1 | 2010 | Reducing cache power with low-cost, multi-bit error-correcting codes · ISCA 2010 |
Memory systems › DRAM › DRAM refresh
refresh energy reduction |
0.0 | 1 | 2010 | Reducing cache power with low-cost, multi-bit error-correcting codes · ISCA 2010 |
Methods — techniques the papers use, named apart from their topics
variable-strength ECC · 0.1hardware mechanism for OS control · 0.1error-correcting codes · 0.1multi-bit error-correcting codes · 0.1thermal moment matching · 0.1spectrum analysis · 0.1power phase analysis · 0.1pole searching · 0.1piecewise constant power input · 0.1linear system of equations · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Direct Compare of Information Coded With Error-Correcting CodesabstractThere are situations in a computing system where incoming information needs to be compared with a piece of stored data to locate the matching entry, e.g., cache tag array lookup and translation look-aside buffer matching. If the stored data is protected with error-correcting codes (ECC) for reliability reason, the previous solution is to access the stored information, decode and correct if necessary before it is used to compare with the incoming data. The decoding and correcting step increases the total access time, which is often critical. In this paper, we propose a method to improve the compare latency for information encoded with ECC. We use the cache tag array look-up as an example, and results show that 30% gate count reduction and 12% latency reduction are achieved. Wei Wu 0024, Dinesh Somasekhar, Shih-Lien Lu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | Energy-efficient cache design using variable-strength error-correcting codesabstractVoltage scaling is one of the most effective mechanisms to improve microprocessors' energy efficiency. However, processors cannot operate reliably below a minimum voltage, Vccmin, since hardware structures may fail. Cell failures in large memory arrays (e.g., caches) typically determine Vccmin for the whole processor. We observe that most cache lines exhibit zero or one failures at low voltages. However, a few lines, especially in large caches, exhibit multi-bit failures and increase Vccmin. Previous solutions either significantly reduce cache capacity to enable uniform error correction across all lines, or significantly increase latency and bandwidth overheads when amortizing the cost of error-correcting codes (ECC) over large lines. Alaa R. Alameldeen, Ilya Wagner, Zeshan Chishti, Wei Wu 0024, Chris Wilkerson, Shih-Lien Lu |
ISCA | 4 |
| 2011 | Adaptive Cache Design to Enable Reliable Low-Voltage OperationabstractThe performance/energy trade-off is widely acknowledged as a primary design consideration for modern processors. A less discussed, though equally important, trade-off is the reliability/energy trade-off. Many design features that increase reliability (e.g., redundancy, error detection, and correction) have the side effect of consuming more energy. Many energy-saving features (e.g., voltage scaling) have the side effect of making systems less reliable. In this paper, we propose an adaptive cache design that enables the operating system to optimize for performance or energy efficiency without sacrificing reliability. Our proposed mechanism enables a cache with a wide operating range, where the cache can use a variable part of its data array to store error-correcting codes. A reliable, energy-efficient cache can use up to half of its data array to store error-correcting codes so that it can reliably operate at a low voltage to reduce energy. A reliable high-performance cache uses its whole data array, but operates at a higher voltage to improve reliability while sacrificing energy. We propose a hardware mechanism that allows the operating system to choose different points within that operating range based on the desired levels of performance, energy, and reliability. Alaa R. Alameldeen, Zeshan Chishti, Chris Wilkerson, Wei Wu 0024, Shih-Lien Lu |
IEEE Trans. Computers | 4 |
| 2010 | Reducing cache power with low-cost, multi-bit error-correcting codesabstractTechnology advancements have enabled the integration of large on-die embedded DRAM (eDRAM) caches. eDRAM is significantly denser than traditional SRAMs, but must be periodically refreshed to retain data. Like SRAM, eDRAM is susceptible to device variations, which play a role in determining refresh time for eDRAM cells. Refresh power potentially represents a large fraction of overall system power, particularly during low-power states when the CPU is idle. Future designs need to reduce cache power without incurring the high cost of flushing cache data when entering low-power states. In this paper, we show the significant impact of variations on refresh time and cache power consumption for large eDRAM caches. We propose Hi-ECC, a technique that incorporates multi-bit error-correcting codes to significantly reduce refresh rate. Multi-bit error-correcting codes usually have a complex decoder design and high storage cost. Hi-ECC avoids the decoder complexity by using strong ECC codes to identify and disable sections of the cache with multi-bit failures, while providing efficient single-bit error correction for the common case. Hi-ECC includes additional optimizations that allow us to amortize the storage cost of the code over large data words, providing the benefit of multi-bit correction at same storage cost as a single-bit error-correcting (SECDED) code (2 % overhead). Our proposal achieves a 93 % reduction in refresh power vs. a baseline eDRAM cache without error correcting capability, and a 66 % reduction in refresh power vs. a system using SECDED codes. Chris Wilkerson, Alaa R. Alameldeen, Zeshan Chishti, Wei Wu 0024, Dinesh Somasekhar, Shih-Lien Lu |
ISCA | 4 |
| 2009 | Improving cache lifetime reliability at ultra-low voltagesabstractVoltage scaling is one of the most effective mechanisms to reduce microprocessor power consumption. However, the increased severity of manufacturing-induced parameter variations at lower voltages limits voltage scaling to a minimum voltage, Vccmin, below which a processor cannot operate reliably. Memory cell failures in large memory structures (e.g., caches) typically determine the Vccmin for the whole processor. Memory failures can be persistent (i.e., failures at time zero which cause yield loss) or non-persistent (e.g., soft errors or erratic bit failures). Both types of failures increase as supply voltage decreases and both need to be addressed to achieve reliable operation at low voltages. In this paper, we propose a novel adaptive technique to improve cache lifetime reliability and enable low voltage operation. This technique, multi-bit segmented ECC (MS-ECC) addresses both persistent and non-persistent failures. Like previous work on mitigating persistent failures, MS-ECC trades off cache capacity for lower voltages. However, unlike previous schemes, MS-ECC does not rely on testing to identify and isolate defective bits, and therefore enables error tolerance for nonpersistent failures like erratic bits and soft errors at low voltages. Furthermore, MS-ECC’s design can allow the operating system to adaptively change the cache size and ECC capability to adjust to system operating conditions. Compared to current designs with single-bit correction, the most aggressive implementation for MS-ECC enables a 30 % reduction in supply voltage, reducing power by 71 % and energy per instruction by 42%. Zeshan Chishti, Alaa R. Alameldeen, Chris Wilkerson, Wei Wu 0024, Shih-Lien Lu |
MICRO | 4 |
| 2008 | FEKIS: a fast architecture-level thermal analyzer for online thermal regulationabstractOwning to increasing power consumption and the corresponding heat dissipated on die, efficient on-chip temperature regulation becomes imperative for today's high performance microprocessors. Temperature tracking based on the on-chip thermal sensors is not sufficient as the temperature hot spots keep changing with the load. One way to mitigate this problem is by means of software sensors, where temperature of any location is computed based on realtime power information and calibrated with the physical sensors. In this paper, we present a very efficient numerical thermal analyzer, which is suitable for fast temperature tracking and online thermal regulation. The proposed method, called FEKIS, combines two existing numerical techniques: extended Krylov subspace reduction technique to reduce the thermal circuit complexity and large-step integration method to exploits the piecewise constant power input traces, which is typical in the power traces at the architecture level. Experimental results show that FEKIS runs 10X faster than the precise time-step integration method only and $1000X$ faster than the traditional numerical integration method with high accuracy. Pu Liu, Sheldon X.-D. Tan, Wei Wu 0024, Murli Tirumala |
ACM Great Lakes Symposium on VLSI | 3 |
| 2007 | Improving the reliability of on-chip data caches under process variationsabstractOn-chip caches take a large portion of the chip area. They are much more vulnerable to parameter variation than smaller units. As leakage current becomes a significant component of the total power consumption, the leakage current variations induced thermal and reliability problem to the on-chip caches become an important design concern. This paper studies the impact of process variations, particular the leakage variations, on the temperature and reliability of on-chip caches. Our statistical simulation shows that, under process variation, 85% of the caches see shortened lifetime, with average lifetime being 81.6% of the ideal cache. At runtime, unevenly distributed dynamic power and the corresponding thermal variation would further deteriorate the situation. To mitigate this problem, we propose a dynamic cache subarray permutation scheme that can alleviate the thermal stress on a high-leakage area to improve the reliability of the caches. Experiments on 17 Spec2k benchmarks show that our scheme can extend the cache lifetime by up to 20.3%, and reduce the peak temperature by 7 degrees on average and more on data-intensive applications. Wei Wu 0024, Sheldon X.-D. Tan, Jun Yang 0002, Shih-Lien Lu |
ICCD | 1 |
| 2007 | Efficient power modeling and software thermal sensing for runtime temperature monitoringabstractThe evolution of microprocessors has been hindered by increasing power consumption and heat dissipation on die. An excessive amount of heat creates reliability problems, reduces the lifetime of a processor, and elevates the cost of cooling and packaging considerably. It is therefore imperative to be able to monitor the temperature variations across the die in a timely and accurate manner. Most current techniques rely on on-chip thermal sensors to report the temperature of the processor. Unfortunately, significant variation in chip temperature both spatially and temporally exposes the limitation of the sensors. We present a compensating approach to tracking chip temperature through an OS resident software module that generates live power and thermal profiles of the processor. We developed such a software thermal sensor (STS) in a Linux system with a Pentium 4 Northwood core. We employed highly efficient numerical methods in our model to minimize the overhead of temperature calculation. We also developed an efficient algorithm for functional unit power modeling. Our power and thermal models are calibrated and validated against on-chip sensor readings, thermal images of the Northwood heat spreader, and the thermometer measurements on the package. The resulting STS offers detailed power and temperature breakdowns of each functional unit at runtime, enabling more efficient online power and thermal monitoring and management at a higher level, such as the operating system. Wei Wu 0024, Lingling Jin, Jun Yang 0002, Pu Liu, Sheldon X.-D. Tan |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2006 | A systematic method for functional unit power estimation in microprocessorsabstractWe present a new method for mathematically estimating the active unit power of functional units in modern microprocessors such as the Pentium 4 family. Our method leverages the phasic behavior in power consumption of programs, and captures as many power phases as possible to form a linear system of equations such that the functional unit power can be solved. Our experiment results on a real Pentium 4 processor show that power estimations attained as such agree with the measured power very well, with deviations less than 5% only. Wei Wu 0024, Lingling Jin, Jun Yang 0002, Pu Liu, Sheldon X.-D. Tan |
DAC | 1 |
| 2006 | Reduce Register Files Leakage Through Discharging CellsabstractWe propose a low-leakage register file cell design based on the observation that the physical registers in a superscalar processor have very short life cycles. When a register is dead, we discharge its cells to '0' to greatly reduce the leakage current from the read bitlines to the ground. Our design has no impact to critical register read access path. Projected to future 45 nm technology, our design yields additional 38% and 47% leakage power savings on top of the existing low-leakage cell designs for 64-bit and 32-bit datapath, respectively. Taking into the account of dynamic energy savings due to the elimination of write '0' operations, our design saves nearly 20% of total energy. Lingling Jin, Wei Wu 0024, Jun Yang 0002, Chuanjun Zhang, Youtao Zhang |
ICCD | 2 |
| 2006 | Fast Thermal Simulation for Runtime Temperature Tracking and ManagementabstractAs the power density increases exponentially, the runtime regulation of operating temperature by dynamic thermal management (DTM) becomes necessary. This paper proposes two novel approaches to the thermal analysis at the chip architecture level for efficient DTM. The first method, i.e., thermal moment matching with spectrum analysis, is based on observations that the power consumption of architecture-level modules in microprocessors running typical workloads presents a strong nature of periodicity. Such a feature can be exploited by fast spectrum analysis in the frequency domain for computing steady-state response. The second method, i.e., thermal moment matching based on piecewise constant power inputs, is based on the observation that the average power consumption of architecture-level modules in microprocessors running typical workloads determines the trend of temperature variations. As a result, using piecewise constant average power inputs can further speed up the thermal analysis. To obtain transient temperature changes due to the initial condition and constant/average power inputs, numerically stable moment matching methods with enhanced pole searching are carried out to speed up online temperature tracking with high accuracy and low overhead. The resulting thermal analysis algorithm has a linear time complexity in runtime setting when the average power inputs are applied. Experimental results show that the resulting thermal analysis algorithms lead to 10times-100times speedup over the traditional integration-based transient analysis with small accuracy loss Pu Liu, Lingling Jin, Wei Wu 0024, Sheldon X.-D. Tan, Jun Yang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2005 | Assertion-Based Design Exploration of DVS in Network Processor ArchitecturesabstractWith the scaling of technology and higher requirements on performance and functionality, power dissipation is becoming one of the major design considerations in the development of network processors. We use an assertion-based methodology for system-level power/performance analysis to study two dynamic voltage scaling (DVS) techniques, traffic-based DVS and execution-based DVS, in a network processor model. Using the automatically generated distribution analyzers, we analyze the power and performance distributions and study their trade-offs for the two DVS policies with different parameter settings, such as threshold values and window sizes. We discuss the optimal configurations of the two DVS policies under different design requirements. By a set of experiments, we show that the assertion-based trace analysis methodology is an efficient tool that can help a designer easily compare and study optimal architectural configurations in a large design space. Jia Yu 0008, Wei Wu 0024, Xi Chen 0024, Harry Hsieh, Jun Yang 0002, Felice Balarin |
DATE | 2 |
| 2005 | Fast thermal simulation for architecture level dynamic thermal managementabstractAs power density increases exponentially, runtime regulation of operating temperature by dynamic thermal managements becomes necessary. This paper proposes a novel approach to the thermal analysis at chip architecture level for efficient dynamic thermal management. Our new approach is based on the observation that the power consumption of architecture level modules in microprocessors running typical workloads presents strong nature of periodicity. Such a feature can be exploited by fast spectrum analysis in frequency domain for computing steady state response. To obtain the transient temperature changes due to initial condition and constant power inputs, numerically stable moment matching approach is carried out. The total transient responses is the addition of the two simulation results. The resulting fast thermal analysis algorithm leads to at least 10/spl times/-100/spl times/ speedup over traditional integration-based transient analysis with small accuracy loss. Pu Liu, Zhenyu Qi 0002, Lingling Jin, Wei Wu 0024, Sheldon X.-D. Tan, Jun Yang 0002 |
ICCAD | 5 |
| 2005 | Efficient Thermal Simulation for Run-Time Temperature Tracking and ManagementabstractAs power density increases exponentially, run-time regulation of operating temperature by dynamic thermal management becomes imperative. This paper proposes a novel approach to real-time thermal estimation at chip level for efficient dynamic thermal management in lieu of the thermal sensors, which are erroneous and having longer delays. Our new approach is based on the observation that the average power consumption of architecture level modules in microprocessors running typical workloads determines the trend of temperature variations. Such a feature can be exploited by applying fast moment matching technique in frequency domain. To obtain the transient temperature changes due to initial condition and constant power input pattern, numerically stable moment matching approach is carried out to speed up on-line temperature tracking with high accuracy and low overhead. The resulting fast thermal analysis algorithm has linear time complexity in run-time setting and leads to about two orders of magnitude speed-up over traditional integration-based transient analysis. The average maximum error under running typical benchmarks is only about 0.37/spl deg/C as compared to other well-accepted simulation tools. Pu Liu, Zhenyu Qi 0002, Lingling Jin, Wei Wu 0024, Sheldon X.-D. Tan, Jun Yang 0002 |
ICCD | 5 |