Tejas Karkhanis

dblp:14/3171 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Processor architecture and microarchitecture · 55% Performance modeling and evaluation · 24% Hardware accelerators and domain-specific architectures · 14%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
processor performance modeling
0.232009
A mechanistic performance model for superscalar out-of-order processors · ACM Trans. Comput. Syst. 2009
A performance counter architecture for computing accurate CPI components · ASPLOS 2006
A First-Order Superscalar Processor Model · ISCA 2004
Hardware accelerators and domain-specific architectures › pattern matching accelerator
regular expression matching accelerator
0.112012
Accelerating business analytics applications · HPCA 2012
Processor architecture and microarchitecture
SIMD
0.112012
Accelerating business analytics applications · HPCA 2012
Processor architecture and microarchitecture
superscalar processor
0.122007
Automated design of application specific superscalar processors: an analytical approach · ISCA 2007
A First-Order Superscalar Processor Model · ISCA 2004
Processor architecture and microarchitecture › superscalar processor
superscalar out-of-order processor
0.122009
A mechanistic performance model for superscalar out-of-order processors · ACM Trans. Comput. Syst. 2009
A performance counter architecture for computing accurate CPI components · ASPLOS 2006
Processor architecture and microarchitecture › special-purpose processor
application-specific processor design
0.112007
Automated design of application specific superscalar processors: an analytical approach · ISCA 2007
Electronic design automation
design space exploration
0.112007
Automated design of application specific superscalar processors: an analytical approach · ISCA 2007
Processor architecture and microarchitecture › out-of-order execution
instruction window
0.012004
A First-Order Superscalar Processor Model · ISCA 2004
Information retrieval
text analysis
0.012012
Accelerating business analytics applications · HPCA 2012
Processor architecture and microarchitecture
instruction fetch
0.012003
Energy Efficient Co-Adaptive Instruction Fetch and Issue · ISCA 2003
Processor architecture and microarchitecture › out-of-order execution
issue queue
0.012003
Energy Efficient Co-Adaptive Instruction Fetch and Issue · ISCA 2003
Performance modeling and evaluation
simulation
0.012009
A mechanistic performance model for superscalar out-of-order processors · ACM Trans. Comput. Syst. 2009
Performance modeling and evaluation › performance monitoring
hardware performance counters
0.012006
A performance counter architecture for computing accurate CPI components · ASPLOS 2006
Energy-efficient computing
power management
0.012003
Energy Efficient Co-Adaptive Instruction Fetch and Issue · ISCA 2003

Methods — techniques the papers use, named apart from their topics

analytical modeling · 0.4SIMD · 0.3interval analysis · 0.2mechanistic modeling · 0.1pareto optimization · 0.1trace-driven modeling · 0.0
YearPublicationVenuePosition
2012 Accelerating business analytics applications
abstract
Business text analytics applications have seen rapid growth, driven by the mining of data for various decision making processes. Regular expression processing is an important component of these applications, consuming as much as 50% of their total execution time. While prior work on accelerating regular expression processing has focused on Network Intrusion Detection Systems, business analytics applications impose different requirements on regular expression processing efficiency. We present an analytical model of accelerators for regular expression processing, which includes memory bus-, I/O bus-, and network-attached accelerators with a focus on business analytics applications. Based on this model, we advocate the use of vector-style processing for regular expressions in business analytics applications, leveraging the SIMD hardware available in many modern processors. In addition, we show how SIMD hardware can be enhanced to improve regular expression processing even further. We demonstrate a realized speedup better than 1.8 for the entire range of data sizes of interest. In comparison, the alternative strategies deliver only marginal improvement for large data sizes, while performing worse than the SIMD solution for small data sizes.
Valentina Salapura, Tejas Karkhanis, Priya Nagpurkar, José E. Moreira
HPCA2
2009 A mechanistic performance model for superscalar out-of-order processors
abstract
A mechanistic model for out-of-order superscalar processors is developed and then applied to the study of microarchitecture resource scaling. The model divides execution time into intervals separated by disruptive miss events such as branch mispredictions and cache misses. Each type of miss event results in characterizable performance behavior for the execution time interval. By considering an interval's type and length (measured in instructions), execution time can be predicted for the interval. Overall execution time is then determined by aggregating the execution time over all intervals. The mechanistic model provides several advantages over prior modeling approaches, and, when estimating performance, it differs from detailed simulation of a 4-wide out-of-order processor by an average of 7%. The mechanistic model is applied to the general problem of resource scaling in out-of-order superscalar processors. First, we use the model to determine size relationships among microarchitecture structures in a balanced processor design. Second, we use the mechanistic model to study scaling of both pipeline depth and width in balanced processor designs. We corroborate previous results in this area and provide new results. For example, we show that at optimal design points, the pipeline depth times the square root of the processor width is nearly constant. Finally, we consider the behavior of unbalanced, overprovisioned processor designs based on insight gained from the mechanistic model. We show that in certain situations an overprovisioned processor may lead to improved overall performance. Designs where a processor's dispatch width is wider than its issue width are of particular interest.
Stijn Eyerman, Lieven Eeckhout, Tejas Karkhanis, James E. Smith 0001
ACM Trans. Comput. Syst.3
2007 Automated design of application specific superscalar processors: an analytical approach
abstract
Analytical modeling is applied to the automated design of application-specific superscalar processors. Using an analytical method bridges the gap between the size of the design space and the time required for detailed cycle-accurate simulations. The proposed design framework takes as inputs the design targets (upper bounds on execution time, area, and energy), design alternatives, and one or more application programs. The output is the set of out-of-order superscalar processors that are Pareto-optimal with respect to performance-energy-area. The core of the new design framework is made up of analytical performance and energy activity models, and an analytical model-based design optimization process.
Tejas Karkhanis, James E. Smith 0001
ISCA1
2006 A performance counter architecture for computing accurate CPI components
abstract
A common way of representing processor performance is to use Cycles per Instruction (CPI) `stacks' which break performance into a baseline CPI plus a number of individual miss event CPI components. CPI stacks can be very helpful in gaining insight into the behavior of an application on a given microprocessor; consequently, they are widely used by software application developers and computer architects. However, computing CPI stacks on superscalar out-of-order processors is challenging because of various overlaps among execution and miss events (cache misses, TLB misses, and branch mispredictions).This paper shows that meaningful and accurate CPI stacks can be computed for superscalar out-of-order processors. Using interval analysis, a novel method for analyzing out-of-order processor performance, we gain understanding into the performance impact of the various miss events. Based on this understanding, we propose a novel way of architecting hardware performance counters for building accurate CPI stacks. The additional hardware for implementing these counters is limited and comparable to existing hardware performance counter architectures while being significantly more accurate than previous approaches.
Stijn Eyerman, Lieven Eeckhout, Tejas Karkhanis, James E. Smith 0001
ASPLOS3
2004 A First-Order Superscalar Processor Model
abstract
A proposed performance model for superscalar processors consists of: 1) a component that models the relationship between instructions issued per cycle and the size of the instruction window under ideal conditions; and 2) methods for calculating transient performance penalties due to branch mispredictions, instruction cache misses, and data cache misses. Using trace-derived data dependence information, data and instruction cache miss rates, and branch miss-prediction rates as inputs, the model can arrive at performance estimates for a typical superscalar processor that are within 5.8% of detailed simulation on average and within 13% in the worst case. The model also provides insights into the workings of superscalar processors and long-term microarchitecture trends such as pipeline depths and issue widths.
Tejas Karkhanis, James E. Smith 0001
ISCA1
2003 Energy Efficient Co-Adaptive Instruction Fetch and Issue
abstract
Front-end instruction delivery accounts for a significant fraction of the energy consumed in a dynamic superscalar processor. The issue queue in these processors serves two crucial roles: it bridges the front and back ends of the processor and serves as the window of instructions for the out-of-order engine. A mismatch between the front end producer rate and back end consumer rate, and between the supplied instruction window from the front end, and the required instruction window to exploit the level of application parallelism, results in additional front-end energy, and increases the issue queue utilization. While the former increases overall processor energy consumption, the latter aggravates the issue queue hot spot problem.We propose a complementary combination of fetch gating and issue queue adaptation to address both of these issues. We introduce an issue-centric fetch gating scheme based on issue queue utilization and application parallelism characteristics. Our scheme attempts to provide an instruction window size that matches the current parallelism characteristics of the application while maintaining enough queue entries to avoid back-end starvation. Compared to a conventional fetch gating scheme based on flow-rate matching, we demonstrate 20% better overall energy-delay with a 44% additional reduction in issue queue energy. We identify Icache energy savings as the largest contributor to the overall savings and quantify the sources of savings in this structure. We then couple this issue-driven fetch gating approach with an issue queue adaptation scheme based on queue utilization. While the fetch gating scheme provides a window of issue queue instructions appropriate to the level of program parallelism, the issue queue adaptation approach shuts down the remaining underutilized issue queue entries. Used in tandem, these complementary techniques yield a 20% greater issue queue energy savings than the addition of the savings from each technique applied in isolation. The result of this combined approach is a 6% overall energy-delay savings coupled with a 54% reduction in issue queue energy.
Alper Buyuktosunoglu, Tejas Karkhanis, David H. Albonesi, Pradip Bose
ISCA2
2002 Saving energy with just in time instruction delivery
abstract
Just-In-Time instruction delivery is a general method for saving energy in a microprocessor by dynamically limiting the number of in-flight instructions. The goal is to save energy by 1) fetching valid instructions no sooner than necessary, avoiding cycles stalled in the pipeline -- especially the issue queue, and 2) reducing the number of fetches and subsequent processing of mis-speculated instructions. A simple algorithm monitors performance and adjusts the maximum number of in-flight instructions at fairly long intervals, 100K instructions in this study. The proposed JIT instruction delivery scheme provides the combined benefits of more targeted schemes proposed previously. With only a 3% performance degradation, energy savings in the fetch, decode pipe, and issue queue are 10%, 12%, and 40%, respectively.
Tejas Karkhanis, James E. Smith 0001, Pradip Bose
ISLPED1