EDBT 2026 Demo / reviewers in the wild / expert
Jinson Koppanalil
dblp:62/3708
· DBLP profile ↗
3ranked-venue papers
2as first author
0since 2021 · last 2004
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Processor architecture and microarchitecture · 91% Memory systems · 9% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
speculative execution |
0.0 | 1 | 2004 | A Simple Mechanism for Detecting Ineffectual Instructions in Slipstream Processors · IEEE Trans. Computers 2004 |
Processor architecture and microarchitecture › out-of-order execution
instruction window |
0.0 | 1 | 2002 | A Large, Fast Instruction Window for Tolerating Cache Misses · ISCA 2002 |
Processor architecture and microarchitecture › out-of-order execution
issue queue |
0.0 | 1 | 2002 | A Large, Fast Instruction Window for Tolerating Cache Misses · ISCA 2002 |
Processor architecture and microarchitecture
out-of-order execution |
0.0 | 1 | 2002 | A Large, Fast Instruction Window for Tolerating Cache Misses · ISCA 2002 |
Processor architecture and microarchitecture
register file |
0.0 | 1 | 2004 | A Simple Mechanism for Detecting Ineffectual Instructions in Slipstream Processors · IEEE Trans. Computers 2004 |
Memory systems
cache |
0.0 | 1 | 2002 | A Large, Fast Instruction Window for Tolerating Cache Misses · ISCA 2002 |
Memory systems › cache
cache miss tolerance |
0.0 | 1 | 2002 | A Large, Fast Instruction Window for Tolerating Cache Misses · ISCA 2002 |
Methods — techniques the papers use, named apart from their topics
speculative program monitoring · 0.0backward slicing · 0.0simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2004 | A Simple Mechanism for Detecting Ineffectual Instructions in Slipstream ProcessorsabstractA slipstream processor accelerates a program by speculatively removing repeatedly ineffectual instructions. Detecting the roots of ineffectual computation: unreferenced writes, nonmodifying writes, and correctly predicted branches, is straightforward. On the other hand, detecting ineffectual instructions in the backward slices of these root instructions currently requires complex back-propagation circuitry. We observe that, by logically monitoring the speculative program (instead of the original program), back-propagation can be reduced to detecting unreferenced writes. That is, once root instructions are actually removed, instructions at the next higher level in the backward slice become newly exposed unreferenced writes in the speculative program. This new algorithm, called implicit back-propagation, eliminates complex hardware and achieves an average performance improvement of 11.8 percent, only marginally lower than the 12.3 percent improvement achieved with explicit back-propagation. We further simplify the hardware component by electing not to detect ineffectual memory writes, focusing only on ineffectual register writes. A minimal implementation consisting of only a register-indexed table (similar to an architectural register file) achieves a good balance between complexity and performance (11.2 percent average performance improvement with implicit back-propagation and without detection of ineffectual memory writes). Jinson Koppanalil, Eric Rotenberg |
IEEE Trans. Computers | 1 |
| 2002 | A case for dynamic pipeline scalingabstractEnergy consumption can be reduced by scaling down frequency when peak performance is not needed. A lower frequency permits slower circuits, and hence a lower supply voltage. Energy reduction comes from voltage reduction, a technique called Dynamic Voltage Scaling (DVS).This paper makes the case that the useful frequency range of DVS is limited because there is a lower bound on voltage. Lowering frequency permits voltage reduction until the lowest voltage is reached. Beyond that point, lowering frequency further does not save energy because voltage is constant.However, there is still opportunity for energy reduction outside the influence of DVS. If frequency is lowered enough, pairs of pipeline stages can be merged to form a shallower pipeline. The shallow pipeline has better instructions-per-cycle (IPC) than the deep pipeline. Since energy also depends on IPC, energy is reduced for a given frequency. Accordingly, we propose Dynamic Pipeline Scaling (DPS). A DPS-enabled deep pipeline can merge adjacent pairs of stages by making the intermediate latches transparent and disabling corresponding feedback paths. Thus, a DPS-enabled pipeline has a deep mode for higher frequencies within the influence of DVS, and a shallow mode for lower frequencies. Shallow mode extends the frequency range for which energy reduction is possible. For frequencies outside the influence of DVS, a DPS-enabled deep pipeline consumes from 23% to 40% less energy than a rigid deep pipeline. Jinson Koppanalil, Prakash Ramrakhyani, Sameer Desai, Anu Vaidyanathan, Eric Rotenberg |
CASES | 1 |
| 2002 | A Large, Fast Instruction Window for Tolerating Cache MissesabstractInstruction window size is an important design parameter for many modern processors. This paper presents a new instruction window design targeted at achieving the latency tolerance of large windows with the clock cycle time of small windows. The key observation is that instructions dependent on a long latency operation (e.g., cache miss) cannot execute until that source operation completes. These instructions are moved out of the conventional, small, issue queue to a much larger waiting instruction buffer (WIB). When the long latency operation completes, the instructions are reinserted into the issue queue. In this paper, we focus specifically on load cache misses and their dependent instructions. Simulations reveal that, for an 8-way processor, a 2K-entry WIB with a 32-entry issue queue can achieve speedups of 20%, 84%, and 50% over a conventional 32-entry issue queue for a subset of the SPEC CINT2000, SPEC CFP2000, and Olden benchmarks, respectively. Alvin R. Lebeck, Tong Li 0003, Eric Rotenberg, Jinson Koppanalil, Jaidev P. Patwardhan |
ISCA | 4 |