EDBT 2026 Demo / reviewers in the wild / expert
Niket Kumar Choudhary
dblp:62/7751 · also Niket K. Choudhary
· DBLP profile ↗
8ranked-venue papers
2as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Processor architecture and microarchitecture · 56% Memory systems · 20% Integrated circuit design · 18% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › multicore design › heterogeneous multicore
asymmetric multicore |
0.1 | 1 | 2012 | Efficiently exploiting memory level parallelism on asymmetric coupled cores in the dark silicon era · ACM Trans. Archit. Code Optim. 2012 |
Memory systems › memory access optimization
memory-level parallelism |
0.1 | 1 | 2012 | Efficiently exploiting memory level parallelism on asymmetric coupled cores in the dark silicon era · ACM Trans. Archit. Code Optim. 2012 |
Processor architecture and microarchitecture › multicore design
heterogeneous multicore |
0.1 | 1 | 2011 | FabScalar: composing synthesizable RTL designs of arbitrary cores within a canonical superscalar template · ISCA 2011 |
Integrated circuit design › digital system design
register-transfer level design |
0.1 | 1 | 2011 | FabScalar: composing synthesizable RTL designs of arbitrary cores within a canonical superscalar template · ISCA 2011 |
Processor architecture and microarchitecture
superscalar processor |
0.1 | 1 | 2011 | FabScalar: composing synthesizable RTL designs of arbitrary cores within a canonical superscalar template · ISCA 2011 |
Energy-efficient computing › energy-constrained computing
dark silicon |
0.0 | 1 | 2012 | Efficiently exploiting memory level parallelism on asymmetric coupled cores in the dark silicon era · ACM Trans. Archit. Code Optim. 2012 |
Methods — techniques the papers use, named apart from their topics
design space exploration · 0.1application steering · 0.1template-based design composition · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Design-effort alloy: Boosting a highly tuned primary core with untuned alternate coresabstractA commercial flagship superscalar core is a highly tuned machine. Designers spend significant effort to tune the register-transfer-level (RTL) model, circuits, and layout to optimize performance and power. Nonetheless, the one-size-fits-all microarchitecture still suffers from suboptimal performance and power on individual applications. A single-ISA heterogeneous multi-core, with its multiple diverse core designs, has potential to exploit application diversity. However, tuning multiple core types will incur insurmountable design effort. This paper proposes a new class of single-ISA heterogeneous multi-core processor, called design-effort alloy (DEA). Only one of the core types, called the high-effort core (HEC), is tuned using a high-effort design flow. Much less effort is spent on tuning other core types, called low-effort cores (LECs). We begin with synthesizable RTL designs of a palette of out-of-order superscalar core types. A LEC and HEC is designed for each core type: the LEC is based on design automation and the HEC is derived from its LEC counterpart, using frequency and energy scaling factors that account for RTL, circuit, and layout optimizations. The resulting HECs have more than a 2x frequency advantage with only a 1.3× increase in energy consumption compared to their corresponding LECs. From the palette of core types, we find the best 4-core-type DEA processor for 179 SPEC SimPoints (program phases). Our study yielded the following key results: 1) The DEA processor's HEC is the same core type in the best high-effort homogeneous multi-core, owing to most program phases demonstrating “average” instruction-level behavior and favoring this balanced core. 2) The DEA processor yields a speedup in BIPS3/W of 1%-87%, and a geometric-mean speedup of 25%, on 20 out of 179 SimPoints over the best high-effort homogeneous multi-core. Thus, untuned LECs operating at less than half the frequency of the HEC nonetheless accelerate program phases with “outlier” instruction-level behavior. Elliott Forbes, Niket Kumar Choudhary, Brandon H. Dwiel, Eric Rotenberg |
ICCD | 2 |
| 2013 | A unified view of non-monotonic core selection and application steering in heterogeneous chip multiprocessorsabstractA single-ISA heterogeneous chip multiprocessor (HCMP) is an attractive substrate to improve single-thread performance and energy efficiency in the dark silicon era. We consider HCMPs comprised of non-monotonic core types where each core type is performance-optimized to different instruction-level behavior and hence cannot be ranked - different program phases achieve their highest performance on different cores. Although non-monotonic heterogeneous designs offer higher performance potential than either monotonic heterogeneous designs or homogeneous designs, steering applications to the best-performing core is challenging due to performance ambiguity of core types. Sandeep Navada, Niket Kumar Choudhary, Salil V. Wadhavkar, Eric Rotenberg |
PACT | 2 |
| 2012 | FPGA modeling of diverse superscalar processorsabstractThere is increasing interest in using Field Programmable Gate Arrays (FPGAs) as platforms for computer architecture simulation. This paper is concerned with modeling superscalar processors with FPGAs. To be transformative, the FPGA modeling framework should meet three criteria. (1) Configurable: The framework should be able to model diverse superscalar processors, like a software model. In particular, it should be possible to vary superscalar parameters such as fetch, issue, and retire widths, depths of pipeline stages, queue sizes, etc. (2) Automatic: The framework should be able to automatically and efficiently map any one of its superscalar processor configurations to the FPGA. (3) Realistic: The framework should model a modern superscalar microarchitecture in detail, ideally with prototype quality, to enable a new era and depth of microarchitecture research. A framework that meets these three criteria will enjoy the convenience of a software model, the speed of an FPGA model, and the experience of a prototype. This paper describes FPGA-Sim, a configurable, automatically FPGA-synthesizable, and register-transfer-level (RTL) model of an out-of-order superscalar processor. FPGA-Sim enables FPGA modeling of diverse superscalar processors out-of-the-box. Moreover, its direct RTL implementation yields the fidelity of a hardware prototype. Brandon H. Dwiel, Niket Kumar Choudhary, Eric Rotenberg |
ISPASS | 2 |
| 2012 | A physical design study of fabscalar-generated superscalar coresabstractAbstract—FabScalar is a recently published tool for automat-ically generating superscalar cores, of different pipeline widths, depths and sizes. The output of FabScalar is a synthesizable register-transfer-level (RTL) description of the desired core. While this capability makes sophisticated cores more accessible to designers and researchers, meaningful applications require reducing RTL descriptions to physical designs. This paper presents the first systematic physical design study of FabScalar-generated superscalar cores. I. Niket Kumar Choudhary, Brandon H. Dwiel, Eric Rotenberg |
VLSI-SoC | 1 |
| 2012 | Efficiently exploiting memory level parallelism on asymmetric coupled cores in the dark silicon eraabstractExtracting high memory-level parallelism (MLP) is essential for speeding up single-threaded applications which are memory bound. At the same time, the projected amount of dark silicon (the fraction of the chip powered off) on a chip is growing. Hence, Asymmetric Multicore Processors (AMP) offer a unique opportunity to integrate many types of cores, each powered at different times, in order to optimize for different regions of execution. In this work, we quantify the potential for exploiting core customization to speedup programs during regions of high MLP. Based on a careful design space exploration, we discover that an AMP that includes a narrow and fast specialized core has the potential to efficiently exploit MLP. Using the results of our analysis, we design an AMP with both an MLP and ILP specialized core, and we propose a hardware-level, application steering mechanism called Symbiotic Core Execution (SCE). SCE detects MLP phases by monitoring the L2 miss rate of the application, and it uses that information to steer the application to the best core. Interestingly, we show that L2 miss rates are important for deciding when an MLP region begins and when it ends. As a program runs, its execution migrates to a core customized for MLP during regions of high MLP; when the region ends, it is re-scheduled on the core that fits the application characteristics. Compared to a monolithic core optimized for both modes of operation, our AMP design provides a harmonic mean performance improvement of 5.3% and 6.6% for SPEC2000 and SPEC2006, respectively, with a maximum speedup of 14.5%. For the same study, it achieves a 18.3% and 21.1% energy delay 2 reduction for SPEC2000 and SPEC2006, respectively. Our findings yield an important message for designing AMPs with specialized cores: core customization enables efficient exploitation of MLP, and application steering mechanisms for MLP are simple to implement and effective. George Patsilaras, Niket Kumar Choudhary, James Tuck 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2011 | FabScalar: composing synthesizable RTL designs of arbitrary cores within a canonical superscalar templateabstractA growing body of work has compiled a strong case for the single-ISA heterogeneous multi-core paradigm. A single-ISA heterogeneous multi-core provides multiple, differently-designed superscalar core types that can streamline the execution of diverse programs and program phases. No prior research has addressed the 'Achilles' heel of this paradigm: design and verification effort is multiplied by the number of different core types. Niket Kumar Choudhary, Salil V. Wadhavkar, Tanmay A. Shah, Hiran Mayukh, Jayneel Gandhi, Brandon H. Dwiel, Sandeep Navada, Hashem Hashemi Najaf-abadi, Eric Rotenberg |
ISCA | 1 |
| 2010 | Criticality-driven superscalar design space explorationabstractIt has become increasingly difficult to perform design space exploration (DSE) of computer systems with a short turnaround time because of exploding design spaces, increasing design complexity and long-running workloads. Researchers have used classical search/optimization techniques like simulated annealing, genetic algorithms, etc., to accelerate the DSE. While these techniques are better than an exhaustive search, a substantial amount of time must still be dedicated to DSE. This is a serious bottleneck in reducing research/development time. These techniques do not perform the DSE quickly enough, primarily because they do not leverage any insight as to how the different design parameters of a computer system interact to increase or degrade performance at a design point and treat the computer system as a "black-box". Sandeep Navada, Niket Kumar Choudhary, Eric Rotenberg |
PACT | 2 |
| 2009 | Core-Selectability in Chip MultiprocessorsabstractThe centralized structures necessary for the extraction of instruction-level parallelism (ILP) are consuming progressively smaller portions of the total die area of chip multiprocessors (CMP). The reason for this is that scaling these structures does not enhance general performance as much as scaling the cache and interconnect. However, the fact that these structures now consume less proportional die area opens an avenue to enhancing their performance through truly overcoming the one-size-fits-all approach to their design. This paper proposes core-selectability - incorporating differently-designed cores that can be toggled into active employment. This enables differently customized ILP-extracting structures to be at hand in the system while not dramatically adding to the interconnect complexity. The design verification effort is minimized by separating the complexity of different core designs. Moreover, contrary to alternative approaches, the performance and power efficiency of the core designs are not compromised. Evaluation results are presented that show that, even when limiting the diversity between core designs to only the sizing of microarchitectural structures, core-selectability has the potential to provide notable performance enhancement (with an average of 10%) to scalable multithreaded applications, without increased concurrency. In addition, it can provide significantly greater throughput to multiprogrammed workloads by providing the potential for the system to transform into a heterogeneous design. Hashem Hashemi Najaf-abadi, Niket Kumar Choudhary, Eric Rotenberg |
PACT | 2 |