EDBT 2026 Demo / reviewers in the wild / expert
Shih-Chang Lai
dblp:57/304 · also Shih-Chang Kevin Lai
· DBLP profile ↗
6ranked-venue papers
2as first author
0since 2021 · last 2011
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-authorArtificial intelligence and machine learning · 1Security and privacy · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 57% Processor architecture and microarchitecture · 43% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › memory access latency
cache access latency |
0.1 | 2 | 2002 | Dynamic addressing memory arrays with physical locality · MICRO 2002 Direct load: dependence-linked dataflow resolution of load address and cache coordinate · MICRO 2001 |
Processor architecture and microarchitecture › register file
register file design |
0.0 | 1 | 2002 | Dynamic addressing memory arrays with physical locality · MICRO 2002 |
Memory systems
cache |
0.0 | 1 | 2001 | Direct load: dependence-linked dataflow resolution of load address and cache coordinate · MICRO 2001 |
Memory systems › memory access latency
load latency |
0.0 | 1 | 2001 | Direct load: dependence-linked dataflow resolution of load address and cache coordinate · MICRO 2001 |
Processor architecture and microarchitecture
out-of-order execution |
0.0 | 1 | 2001 | Direct load: dependence-linked dataflow resolution of load address and cache coordinate · MICRO 2001 |
Processor architecture and microarchitecture
instruction scheduling |
0.0 | 1 | 2002 | Dynamic addressing memory arrays with physical locality · MICRO 2002 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 1 | 2001 | Direct load: dependence-linked dataflow resolution of load address and cache coordinate · MICRO 2001 |
Processor architecture and microarchitecture › special-purpose processor › application-specific processor design
instruction path coprocessor |
0.0 | 1 | 2001 | Direct load: dependence-linked dataflow resolution of load address and cache coordinate · MICRO 2001 |
Methods — techniques the papers use, named apart from their topics
circuit simulation · 0.0address decoding reconfiguration · 0.0stride-based prediction · 0.0simplescalar · 0.0dependence-linked dataflow · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | A SOMO-based approach to the operating room scheduling problem
Mu-Chun Su, Shih-Chang Lai, Pa-Chun Wang, Yi-Zeng Hsieh, Shih-Chieh Lin |
Expert Syst. Appl. | 2 |
| 2003 | Hardware-based Pointer Data PrefetcherabstractEffective prefetching of data from a lower memory hierarchy to higher level is a helpful way to combat the increasing memory latency that impedes the performance improvement. This paper presents a hardware-based prefetching technique to alleviate load misses caused by irregular access patterns common in linked-list structures. Different from previous works we identify a pointer load using its architecture source register. A table called target register bitmap (TRB) is maintained. By looking up this table we can identify if a load is a pointer load. We remember the base addresses of consumer load operations in a cache and prefetch the data pointed by speculative virtual addresses to a prefetch buffer, which is smaller than the data cache. Whenever a load is encountered, both the prefetch buffer and the data cache are looked up. SPEC2000 and Olden benchmarks are used to evaluate this method. This technique is able to predict accurately over 80% of the pointer load address. A system using this technique having 8KB LI data cache plus a 1KB address cache and a 1KB prefetch buffer gives an average of around 6% performance improvement over a system with 16KB LI data cache. Shih-Chang Lai, Shih-Lien Lu |
ICCD | 1 |
| 2002 | Ditto ProcessorabstractConcentration of design effort for current single-chip commercial-off-the-shelf (COTS) microprocessors has been directed towards performance. Reliability has not been the primary focus. As supply voltage scales to accommodate technology scaling and to lower power consumption, transient errors are more likely to be introduced. The basic idea behind any error tolerance scheme involves some type of redundancy. Redundancy techniques can be categorized in three general categories: (1) hardware redundancy, (2) information redundancy, and (3) time redundancy. Existing time redundant techniques for improving reliability of a superscalar processor utilize the otherwise unused hardware resources as much as possible to hide the overhead of program re-execution and verification. However, our study reveals that re-executing of long latency operations contributes to performance loss. We suggest a method to handle short and long latency instructions in slightly different ways to reduce the performance degradation. Our goal is to minimize the hardware overhead and performance degradation while maximizing the fault detection coverage. Experimental studies through microarchitecture simulation are used to compare performance lost due to the proposed scheme with non-fault tolerant design and different existing time redundant fault tolerant schemes. Fourteen integer and floating-point benchmarks are simulated with 1.8/spl sim/13.3% performance loss when compared with non-fault-tolerant superscalar processor. Shih-Chang Lai, Shih-Lien Lu, Jih-Kwon Peir |
DSN | 1 |
| 2002 | Bloom filtering cache misses for accurate data speculation and prefetchingabstractA processor must know a load instruction's latency to schedule the load's dependent instructions at the correct time. Unfortunately, modern processors do not know this latency until well after the dependent instructions should have been scheduled to avoid pipeline bubbles between themselves and the load. One solution to this problem is to predict the load's latency, by predicting whether the load will hit or miss in the data cache. Existing cache hit/miss predictors, however, can only correctly predict about 50% of cache misses.This paper introduces a new hit/miss predictor that uses a Bloom Filter to identify cache misses early in the pipeline. This early identification of cache misses allows the processor to more accurately schedule instructions that are dependent on loads and to more precisely prefetch data into the cache. Simulations using a modified SimpleScalar model show that the proposed Bloom Filter is nearly perfect, with a prediction accuracy greater than 99% for the SPECint2000 benchmarks. IPC (Instructions Per Cycle) performance improved by 19% over a processor that delayed the scheduling of instructions dependent on a load until the load latency was known, and by 6% and 7% over a processor that always predicted a load would hit the cache and with a counter-based hit/miss predictor respectively. This IPC reaches 99.7% of the IPC of a processor with perfect scheduling. Jih-Kwon Peir, Shih-Chang Lai, Shih-Lien Lu, Jared Stark, Konrad Lai |
ICS | 2 |
| 2002 | Dynamic addressing memory arrays with physical localityabstractAs pipeline width and depth grow to improve performance, memory arrays in microprocessors are growing in entries and ports. Arrays will increase in physical size, which prolongs the access time due to wiring delay. In order to boost clock frequency, these memory arrays must take multiple cycles to complete an access. This delays the scheduling of dependent instructions and affects overall performance. This paper proposes a different circuit organization to enable fast and slow accesses solely dependent on physical locality. Since the access time depends on a fixed physical location, it is pre-determined to scheduling dependent instructions. Furthermore, this paper presents a mechanism to re-configure the address decoding of the physical register file to increase the occurrence of fast accesses. Detailed circuit simulation using this proposed method determines the access cycle time. Reduction in average access cycle time for the register file and the first level data cache recovers 73% of the IPC degradation. Steven Hsu, Shih-Lien Lu, Shih-Chang Lai, Ram Krishnamurthy 0001, Konrad Lai |
MICRO | 3 |
| 2001 | Direct load: dependence-linked dataflow resolution of load address and cache coordinateabstractAn increasing cache latency in future processors incurs profound performance impacts in spite of advanced out-of-order execution techniques. In this paper, we describe an early address resolution mechanism that accurately resolves both regular and irregular load addresses. The basic idea is to build dynamic dependence links from the instruction that updates the base register to the consumer load instructions. Once a new base address is available, it triggers calculations of the new load addresses for dependent loads. Furthermore, the exact cache location of the requested data is predicted based on the newly resolved load address. As a result, this direct load can access the data cache directly to achieve a zero-cycle load latency. Performance evaluation using SPEC integer programs shows that the dynamic dependence links can be established accurately. Combined with a stride-based predictor, the proposed early address resolution achieves about 97% average accuracy with less than 1% misprediction. Based on a modified SimpleScalar model, the proposed method can potentially improve the IPC by about 18%. Byung-Kwon Chung, Jinsuo Zhang, Jih-Kwon Peir, Shih-Chang Lai, Konrad Lai |
MICRO | 4 |