EDBT 2026 Demo / reviewers in the wild / expert
Mani Azimi
dblp:44/6303
· DBLP profile ↗
8ranked-venue papers
2as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6Software engineering, systems software and programming languages · 2 · 1 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Interconnection networks and networks-on-chip · 26% GPUs and heterogeneous computing · 24% Memory systems · 15% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing › GPU cache
GPU cache hierarchy |
0.2 | 1 | 2015 | GREEN Cache: Exploiting the Disciplined Memory Model of OpenCL on GPUs · IEEE Trans. Computers 2015 |
Energy-efficient computing › memory energy efficiency
low-power cache design |
0.2 | 1 | 2015 | GREEN Cache: Exploiting the Disciplined Memory Model of OpenCL on GPUs · IEEE Trans. Computers 2015 |
Memory systems › cache design
non-inclusive cache |
0.2 | 1 | 2015 | GREEN Cache: Exploiting the Disciplined Memory Model of OpenCL on GPUs · IEEE Trans. Computers 2015 |
Interconnection networks and networks-on-chip › network topology › mesh network
2d mesh |
0.2 | 1 | 2014 | MoDe-X: Microarchitecture of a Layout-Aware Modular Decoupled Crossbar for On-Chip Interconnects · IEEE Trans. Computers 2014 |
Interconnection networks and networks-on-chip › router architecture
network-on-chip router microarchitecture |
0.2 | 1 | 2014 | MoDe-X: Microarchitecture of a Layout-Aware Modular Decoupled Crossbar for On-Chip Interconnects · IEEE Trans. Computers 2014 |
Interconnection networks and networks-on-chip › network topology › network topology design
network-on-chip topology |
0.2 | 1 | 2014 | MoDe-X: Microarchitecture of a Layout-Aware Modular Decoupled Crossbar for On-Chip Interconnects · IEEE Trans. Computers 2014 |
Parallel and multicore computing › task allocation
computation-to-core mapping |
0.2 | 1 | 2013 | Application-to-core mapping policies to reduce memory system interference in multi-core systems · HPCA 2013 |
GPUs and heterogeneous computing
control flow divergence |
0.2 | 1 | 2013 | SIMD divergence optimization through intra-warp compaction · ISCA 2013 |
Memory systems
memory controller |
0.2 | 1 | 2013 | Application-to-core mapping policies to reduce memory system interference in multi-core systems · HPCA 2013 |
Processor architecture and microarchitecture
multicore design |
0.2 | 1 | 2013 | Application-to-core mapping policies to reduce memory system interference in multi-core systems · HPCA 2013 |
Processor architecture and microarchitecture
SIMD |
0.2 | 1 | 2013 | SIMD divergence optimization through intra-warp compaction · ISCA 2013 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.1 | 1 | 2015 | GREEN Cache: Exploiting the Disciplined Memory Model of OpenCL on GPUs · IEEE Trans. Computers 2015 |
Energy-efficient computing
power gating |
0.1 | 1 | 2014 | MoDe-X: Microarchitecture of a Layout-Aware Modular Decoupled Crossbar for On-Chip Interconnects · IEEE Trans. Computers 2014 |
Parallel and multicore computing › data parallelism
data-parallel applications |
0.0 | 1 | 2013 | SIMD divergence optimization through intra-warp compaction · ISCA 2013 |
Reconfigurable computing and FPGAs
FPGA prototyping |
0.0 | 1 | 2010 | FPGA-based prototyping of a 2D MESH / TORUS on-chip interconnect (abstract only) · FPGA 2010 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.4power estimation · 0.2workload characterization · 0.2synthetic traffic generation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | GREEN Cache: Exploiting the Disciplined Memory Model of OpenCL on GPUsabstractAs various graphics processing unit architectures are deployed across broad computing spectrum from a hand-held or embedded device to a high-performance computing server, OpenCL becomes the de facto standard programming environment for general-purpose computing on graphics processing units. Unlike its CPU counterpart, OpenCL has several distinct features such as its disciplined memory model, which is partially inherited from conventional 3D graphics programming models. On the other hand, due to ever increasing memory bandwidth pressure and low power requirement, the capacity of on-chip caches in GPUs keeps increasing overtime. Given such trends, we believe that we have interesting programming model/architecture co-optimization opportunities, in particular, how to energy-efficiently utilize large on-chip caches for GPUs. In this paper, as a showcase, we study the characteristics of the OpenCL memory model and propose a technique called GPU Region-aware energy-efficient non-inclusive cache hierarchy, or GREEN cache hierarchy. With the GREEN cache, our simulation results show that we can save 56 percent of dynamic energy in the L1 cache, 39 percent of dynamic energy in the L2 cache, and 50 percent of leakage energy in the L2 cache with practically no performance degradation and off-chip access increases. Jaekyu Lee, Dong Hyuk Woo, Hyesoon Kim, Mani Azimi |
IEEE Trans. Computers | 4 |
| 2014 | MoDe-X: Microarchitecture of a Layout-Aware Modular Decoupled Crossbar for On-Chip InterconnectsabstractThe number of cores in a single chip keeps increasing with process technology scaling, requiring a scalable interconnection network topology. Buffered wormhole-switched interconnect architectures are attractive for such multicore architectures. The 2D mesh on-chip interconnect provides a scalable, cost-efficient, flexible, and reliable next-generation interconnect topology in this context. In this paper, we provide a microarchitecture for a power and area efficient router for a 2D mesh interconnect. We propose an efficient crossbar implementation, called MoDe-X, that uses a reasonable power-performance tradeoff. The MoDe-X router uses a Modular-Decoupled Crossbar (MoDe-X) that incorporates dimensional decomposition and segmentation to achieve power and area savings. However, unlike most prior work that considers only logical representation of the crossbars, MoDe-X is a physically aware router accounting for the actual layout of router components to reflect practical design requirements. Our simulation results and power estimate show that the MoDe-X router architectures can reduce the overall router area by up to 40 percent and power consumption by up to 35 percent with very little performance impact that occurs only at higher loads. Further, by applying aggressive power gating techniques the net power reductions can be as much as 99 percent for some workloads with no additional performance impact. Dongkook Park, Aniruddha S. Vaidya, Mani Azimi |
IEEE Trans. Computers | 4 |
| 2013 | Application-to-core mapping policies to reduce memory system interference in multi-core systemsabstractFuture many-core processors are likely to concurrently execute a large number of diverse applications. How these applications are mapped to cores largely determines the interference between these applications in critical shared hardware resources. This paper proposes new application-to-core mapping policies to improve system performance by reducing inter-application interference in the on-chip network and memory controllers. The major new ideas of our policies are to: 1) map network-latency-sensitive applications to separate parts of the network from network-bandwidth-intensive applications such that the former can make fast progress without heavy interference from the latter, 2) map those applications that benefit more from being closer to the memory controllers close to these resources. Our evaluations show that, averaged over 128 multiprogrammed workloads of 35 different benchmarks running on a 64-core system, our final application-to-core mapping policy improves system throughput by 16.7% over a state-of-the-art baseline, while also reducing system unfairness by 22.4% and average interconnect power consumption by 52.3%. Reetuparna Das, Rachata Ausavarungnirun, Onur Mutlu, Mani Azimi |
HPCA | 5 |
| 2013 | SIMD divergence optimization through intra-warp compactionabstractSIMD execution units in GPUs are increasingly used for high performance and energy efficient acceleration of general purpose applications. However, SIMD control flow divergence effects can result in reduced execution efficiency in a class of GPGPU applications, classified as divergent applications. Improving SIMD efficiency, therefore, has the potential to bring significant performance and energy benefits to a wide range of such data parallel applications. Aniruddha S. Vaidya, Anahita Shayesteh, Dong Hyuk Woo, Roy Saharoy, Mani Azimi |
ISCA | 5 |
| 2012 | Application-to-core mapping policies to reduce memory interference in multi-core systemsabstractHow applications running on a many-core system are mapped to cores largely determines the interference between these applications in critical shared resources. This paper proposes application-to-core mapping policies to improve system performance by reducing inter-application interference in the on-chip network and memory controllers. The major new ideas of our policies are to: 1) map network-latency-sensitive applications to separate parts of the network from network-bandwidth-intensive applications such that the former can make fast progress without heavy interference from the latter, 2) map those applications that benefit more from being closer to the memory controllers close to these resources. Our evaluations show that both ideas significantly improve system throughput, fairness and interconnect power efficiency. Reetuparna Das, Rachata Ausavarungnirun, Onur Mutlu, Mani Azimi |
PACT | 5 |
| 2010 | FPGA-based prototyping of a 2D MESH / TORUS on-chip interconnect (abstract only)abstractMany-core chip multiprocessors can be expected to scale to tens of cores and beyond in the near future. Existing and emerging workloads on general-purpose many-core processors typically exhibit fast-changing, unpredictable on-chip communication traffic full of burstiness and jitters between different functional blocks. To provide high sustainable performance, scalable interconnects with a rich feature set including support for adaptive and flexible communication, performance isolation, and fault-tolerance are needed. 2D mesh and torus are attractive choices because they are physical layout friendly and scale more gracefully in network latency and bisection bandwidth than other simple interconnects such as buses or rings. However, the adoption of 2D mesh/torus in many-core processor designs is dependent on a verifiable and robust micro-architecture and a validated set of features. FPGA based systems have recently become a cost-effective, rapid prototyping vehicle for chip multiprocessor architectures. In this paper we present an FPGA based prototype of 2D on-die interconnect architecture. Our prototype is a highly configurable full-scale design that supports options selecting many different micro-architectural features and routing algorithms. The prototype incorporates a synthetic traffic generator to exercise and evaluate our design. To facilitate evaluation and characterization, a rich development environment and novel software capabilities including a very detailed performance visualization infrastructure has been developed. We demonstrate the experiment results of several configurations on a 6x6 2D network emulator setup in this paper. Donglai Dai, Aniruddha S. Vaidya, Roy Saharoy, Seungjoon Park, Dongkook Park, Hariharan L. Thantry, Ralf Plate, Elmar Maas, Mani Azimi |
FPGA | 10 |
| 2003 | Experience with Applying Formal Methods to Protocol Specification and System Architecture
Mani Azimi, Ching-Tsun Chou, Victor W. Lee, Phanindra K. Mannava, Seungjoon Park |
Formal Methods Syst. Des. | 1 |
| 1990 | A software approach to multiprocessor address trace generationabstractThe authors describe a technique for generating architecture-independent multiprocessor data address traces on a widely available RISC (reduced instruction set computer) uniprocessor for a specific class of parallel applications. Automatic modification of the application assembly language enables run-time recording of the virtual address and data for loads and stores. Barrier synchronization events are captured in the traces. The tracing technique (called the Tracer) is relatively fast, portable, and does not require access to a multiprocessor. The generality of the traces and the slow-down by a factor of 10 when generating traces compares favourably with other address tracing methods. The Tracer has proved useful in the evaluation of a hierarchical shared bus multiprocessor. The Tracer can be used to gather statistics on programs for use in stochastic models such as queuing networks. Additionally, the visualization of memory access patterns that can be made with the traces is a useful tool in studying parallel applications on shared memory multiprocessors.> Mani Azimi, Carl Erickson |
COMPSAC | 1 |