EDBT 2026 Demo / reviewers in the wild / expert
Sanjeev Banerjia
dblp:20/1478
· DBLP profile ↗
7ranked-venue papers
2as first author
0since 2021 · last 2000
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Processor architecture and microarchitecture · 87% Memory systems · 13% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 73% Runtime systems and virtual machines · 27% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
instruction scheduling |
0.0 | 2 | 1998 | Unified Assign and Schedule: A New Approach to Scheduling for Clustered Register File Microarchitectures · MICRO 1998 Treegion Scheduling for Wide Issue Processors · HPCA 1998 |
Processor architecture and microarchitecture
instruction scheduling |
0.0 | 2 | 1998 | MPS: Miss-Path Scheduling for Multiple-Issue Processors · IEEE Trans. Computers 1998 Unified Assign and Schedule: A New Approach to Scheduling for Clustered Register File Microarchitectures · MICRO 1998 |
Processor architecture and microarchitecture › superscalar processor
wide-issue processor |
0.0 | 2 | 1998 | MPS: Miss-Path Scheduling for Multiple-Issue Processors · IEEE Trans. Computers 1998 Treegion Scheduling for Wide Issue Processors · HPCA 1998 |
Compilers and program optimization
dynamic optimization |
0.0 | 1 | 2000 | Dynamo: a transparent dynamic optimization system · PLDI 2000 |
Runtime systems and virtual machines › binary translation
dynamic translation |
0.0 | 1 | 2000 | Dynamo: a transparent dynamic optimization system · PLDI 2000 |
Processor architecture and microarchitecture › clustered architecture
clustered microarchitecture |
0.0 | 1 | 1998 | Unified Assign and Schedule: A New Approach to Scheduling for Clustered Register File Microarchitectures · MICRO 1998 |
Processor architecture and microarchitecture › microprocessor design › processor core design
in-order core |
0.0 | 1 | 1998 | MPS: Miss-Path Scheduling for Multiple-Issue Processors · IEEE Trans. Computers 1998 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.0 | 1 | 1998 | Treegion Scheduling for Wide Issue Processors · HPCA 1998 |
Processor architecture and microarchitecture
instruction fetch |
0.0 | 1 | 1996 | Instruction Fetch Mechanisms for VLIW Architectures with Compressed Encodings · MICRO 1996 |
Processor architecture and microarchitecture
instruction set architecture |
0.0 | 1 | 1996 | Instruction Fetch Mechanisms for VLIW Architectures with Compressed Encodings · MICRO 1996 |
Processor architecture and microarchitecture › instruction set architecture
object code compatibility |
0.0 | 1 | 1996 | A Persistent Rescheduled-page Cache for Low Overhead Object Code Compatibility in VLIW Architectures · MICRO 1996 |
Processor architecture and microarchitecture › instruction-level parallelism
VLIW |
0.0 | 1 | 1996 | A Persistent Rescheduled-page Cache for Low Overhead Object Code Compatibility in VLIW Architectures · MICRO 1996 |
Memory systems › cache › CPU cache
instruction cache |
0.0 | 2 | 1998 | MPS: Miss-Path Scheduling for Multiple-Issue Processors · IEEE Trans. Computers 1998 Instruction Fetch Mechanisms for VLIW Architectures with Compressed Encodings · MICRO 1996 |
Memory systems › cache management
cache miss handling |
0.0 | 1 | 1998 | MPS: Miss-Path Scheduling for Multiple-Issue Processors · IEEE Trans. Computers 1998 |
Memory systems
cache |
0.0 | 1 | 1996 | Instruction Fetch Mechanisms for VLIW Architectures with Compressed Encodings · MICRO 1996 |
Memory systems › cache management › storage caching
disk cache |
0.0 | 1 | 1996 | A Persistent Rescheduled-page Cache for Low Overhead Object Code Compatibility in VLIW Architectures · MICRO 1996 |
Memory systems › virtual memory management
page replacement algorithms |
0.0 | 1 | 1996 | A Persistent Rescheduled-page Cache for Low Overhead Object Code Compatibility in VLIW Architectures · MICRO 1996 |
Methods — techniques the papers use, named apart from their topics
trace optimization · 0.1dynamic profiling · 0.1optimal scheduling comparison · 0.0heuristic scheduling · 0.0greedy heuristic · 0.0simulation · 0.0trace-driven simulation · 0.0silo cache · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2000 | Dynamo: a transparent dynamic optimization systemabstractWe describe the design and implementation of Dynamo, a software dynamic optimization system that is capable of transparently improving the performance of a native instruction stream as it executes on the processor. The input native instruction stream to Dynamo can be dynamically generated (by a JIT for example), or it can come from the execution of a statically compiled native binary. This paper evaluates the Dynamo system in the latter, more challenging situation, in order to emphasize the limits, rather than the potential, of the system. Our experiments demonstrate that even statically optimized native binaries can be accelerated Dynamo, and often by a significant degree. For example, the average performance of -O optimized SpecInt95 benchmark binaries created by the HP product C compiler is improved to a level comparable to their -O4 optimized version running without Dynamo. Dynamo achieves this by focusing its efforts on optimization opportunities that tend to manifest only at runtime, and hence opportunities that might be difficult for a static compiler to exploit. Dynamo's operation is transparent in the sense that it does not depend on any user annotations or binary instrumentation, and does not require multiple runs, or any special compiler, operating system or hardware support. The Dynamo prototype presented here is a realistic implementation running on an HP PA-8000 workstation under the HPUX 10.20 operating system. Vasanth Bala, Evelyn Duesterwald, Sanjeev Banerjia |
PLDI | 3 |
| 1998 | Treegion Scheduling for Wide Issue ProcessorsabstractInstruction scheduling is one of the most important phases of compilation for high-performance processors. A compiler typically divides a program into multiple regions of code and then schedules each region. Many past efforts have focused on linear regions such as traces and superblocks. The linearity of these regions can limit speculation, leading to under-utilization of processor resources, especially on wide-issue machines. A type of non-linear region called a treegion is presented in this paper. The formation and scheduling of treegions takes into account multiple execution paths, and the larger scope of treegions allows more speculation, leading to higher utilization and better performance. Multiple scheduling heuristics for treegions are compared against scheduling for several types of linear regions. Empirical results illustrate that instruction scheduling using treegions treegion scheduling-holds promise. Treegion scheduling using the global weight heuristic outperforms the next highest performing region-superblocks by up to 20%. William A. Havanki, Sanjeev Banerjia, Thomas M. Conte |
HPCA | 2 |
| 1998 | Unified Assign and Schedule: A New Approach to Scheduling for Clustered Register File MicroarchitecturesabstractRecently, there has been a trend towards clustered microarchitectures to reduce the cycle time for wide issue microprocessors. In such processors, the register file and functional units are partitioned and grouped into clusters. Instruction scheduling for a clustered machine requires assignment and scheduling of operations to the clusters. In this paper, a new scheduling algorithm named unified-assign-and-schedule (UAS) is proposed for clustered, statically-scheduled architectures. UAS merges the cluster assignment and instruction scheduling phases in a natural and straightforward fashion. We compared the performance of UAS with various heuristics to the well-known Bottom-up Greedy (BUG) algorithm and to an optimal cluster scheduling algorithm, measuring the schedule lengths produced by all of the schedulers. Our results show that UAS gives better performance than the BUG algorithm and is quite close to optimal. Emre Ozer 0001, Sanjeev Banerjia, Thomas M. Conte |
MICRO | 2 |
| 1998 | MPS: Miss-Path Scheduling for Multiple-Issue ProcessorsabstractMany contemporary multiple issue processors employ out-of-order scheduling hardware in the processor pipeline. Such scheduling hardware can yield good performance without relying on compile-time scheduling. The hardware can also schedule around unexpected run-time occurrences such as cache misses. As issue widths increase, however, the complexity of such scheduling hardware increases considerably and can have an impact on the cycle time of the processor. This paper presents the design of a multiple issue processor that uses an alternative approach called miss path scheduling or MPS. Scheduling hardware is removed from the processor pipeline altogether and placed on the path between the instruction cache and the next level of memory. Scheduling is performed at cache miss time as instructions are received from memory. Scheduled blocks of instructions are issued to an aggressively clocked in-order execution core. Details of a hardware scheduler that can perform speculation are outlined and shown to be feasible. Performance results from simulations are presented that highlight the effectiveness of an MPS design. Sanjeev Banerjia, Sumedh W. Sathaye, Kishore N. Menezes, Thomas M. Conte |
IEEE Trans. Computers | 1 |
| 1997 | Treegion Scheduling for Highly Parallel Processors
Sanjeev Banerjia, William A. Havanki, Thomas M. Conte |
Euro-Par | 1 |
| 1996 | Instruction Fetch Mechanisms for VLIW Architectures with Compressed EncodingsabstractVLIW architectures use very wide instruction words in conjunction with high bandwidth to the instruction cache to achieve multiple instruction issue. This report uses the TINKER experimental testbed to examine instruction fetch and instruction cache mechanisms for VLIWs. A compressed instruction encoding for VLIWs is defined and a classification scheme for i-fetch hardware for such an encoding is introduced. Several interesting cache and i-fetch organizations are described and evaluated through trace-driven simulations. A new i-fetch mechanism using a silo cache is found to have the best performance. Thomas M. Conte, Sanjeev Banerjia, Sergei Y. Larin, Kishore N. Menezes, Sumedh W. Sathaye |
MICRO | 2 |
| 1996 | A Persistent Rescheduled-page Cache for Low Overhead Object Code Compatibility in VLIW ArchitecturesabstractObject-code compatibility between processor generations is an open issue for VLIW architectures. A potential solution is a technique termed dynamic rescheduling, which performs run-time software rescheduling at the first-time page faults. The time required for rescheduling the pages constitutes a large portion of the overhead of this method. A disk caching scheme that uses a persistent rescheduled-page cache (PRC) is presented. The scheme reduces the overhead associated with dynamic rescheduling by saving rescheduled pages on disk, across program executions. Operating system support is required for dynamic rescheduling and management of the PRC. The implementation details for the PRC are discussed. Results of simulations used to gauge the effectiveness of PRC indicate that: the PRC is effective in reducing the overhead of dynamic rescheduling; and due to different overhead requirements of programs, a split PRC organization performs better than a unified PRC. The unified PRC was studied for two different page replacement policies: LRU and overhead-based replacement. It was found that with LRU replacement, all the programs consistently perform better with increasing PRC sizes, but the high-overhead programs take a consistent performance hit compared to the low-overhead programs. With overhead-based replacement, the performance of high-overhead programs improves substantially, while the low-overhead programs perform only slightly worse than in the case of the LRU replacement. Thomas M. Conte, Sumedh W. Sathaye, Sanjeev Banerjia |
MICRO | 3 |