EDBT 2026 Demo / reviewers in the wild / expert
Naraig Manjikian
dblp:07/2439
· DBLP profile ↗
14ranked-venue papers
7as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 6 first-authorComputer networks · 2Software engineering, systems software and programming languages · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Parallel and multicore computing · 54% Memory systems · 20% Processor architecture and microarchitecture · 11% | |
| Software engineering, system software, and programming languages
2 papers |
Compilers and program optimization · 100% |
Topics — the 12 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization › memory optimization
data locality optimization |
0.0 | 1 | 2001 | Exploiting Wavefront Parallelism on Large-Scale Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2001 |
Compilers and program optimization › loop optimization
loop tiling |
0.0 | 1 | 2001 | Exploiting Wavefront Parallelism on Large-Scale Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2001 |
Parallel and multicore computing › parallel scheduling
loop scheduling |
0.0 | 1 | 2001 | Exploiting Wavefront Parallelism on Large-Scale Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2001 |
Parallel and multicore computing › parallel algorithms › parallel algorithm design
wavefront parallelism |
0.0 | 1 | 2001 | Exploiting Wavefront Parallelism on Large-Scale Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2001 |
Memory systems › non-uniform memory access
CC-NUMA |
0.0 | 1 | 1998 | Design and Implementation of the NUMAchine Multiprocessor · DAC 1998 |
Processor architecture and microarchitecture
multiprocessor architecture |
0.0 | 1 | 1998 | Design and Implementation of the NUMAchine Multiprocessor · DAC 1998 |
Memory systems
data locality |
0.0 | 1 | 1997 | Fusion of Loops for Parallelism and Locality · IEEE Trans. Parallel Distributed Syst. 1997 |
Parallel and multicore computing
load balancing |
0.0 | 1 | 2001 | Exploiting Wavefront Parallelism on Large-Scale Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2001 |
Performance modeling and evaluation
scheduling policy |
0.0 | 1 | 2001 | Exploiting Wavefront Parallelism on Large-Scale Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2001 |
Parallel and multicore computing › multiprocessor system
shared-memory multiprocessor |
0.0 | 1 | 2001 | Exploiting Wavefront Parallelism on Large-Scale Shared-Memory Multiprocessors · IEEE Trans. Parallel Distributed Syst. 2001 |
Parallel and multicore computing › parallel computing › parallel software engineering
parallel application development |
0.0 | 1 | 1998 | Design and Implementation of the NUMAchine Multiprocessor · DAC 1998 |
Compilers and program optimization
loop transformation |
0.0 | 1 | 1997 | Fusion of Loops for Parallelism and Locality · IEEE Trans. Parallel Distributed Syst. 1997 |
Methods — techniques the papers use, named apart from their topics
static scheduling · 0.1experimental evaluation · 0.1dynamic self-scheduling · 0.1synchronization minimization · 0.0loop nest transformation · 0.0loop fusion · 0.0CAD tools · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Enhanced Bloom filter utilisation scheme for string matching using a splitting approachabstractBloom filters (BFs) are widely utilised to speed up string matching in crucial network applications such as real‐time intrusion detection and spam filters. This study introduces a new approach to improve the efficiency of BFs for string matching functions. The approach splits each target string into two substrings and considers the second substring for programming the BF. The objective is to minimise the false positive rate by maximising the common hash signatures from the second substring. Results show that compared to the traditional means of using BFs, the proposed approach reduces the false positive rate by averages of 76 and 88% for 32 and 64 Kb BFs, respectively. Moreover, a complete string matching architecture has been developed in hardware based on the proposed approach. Results demonstrate the advantages of this new architecture compared to similar previous works. Shervin Vakili, J. M. Pierre Langlois, Yvon Savaria, Naraig Manjikian |
IET Commun. | 4 |
| 2010 | Reliability- and process variation-aware placement for FPGAsabstractNegative bias temperature instability (NBTI) significantly affects nanoscale integrated circuit performance and reliability. The degradation in threshold voltage (Vth) due to NBTI is further affected by the initial value of Vthfrom fabrication-induced process variation (PV). Addressing these challenges in embedded FPGA designs is possible, as FPGA reconfigurablility can be exploited to measure the exact timing degradation of an FPGA due to the joint effect of NBTI and PV at run time with low overhead. The gathered information can then be used to improve the run-time performance and reliability of FPGA designs without targeting the pessimistic worst case. In this paper, we present joint NBTI/PV-aware placement techniques for FPGAs, including NBTI/PV-aware timing analysis, region-based delay estimation, and a new move-acceptance procedure. To evaluate the proposed techniques, we combine PV measurements from 15 Xilinx Virtex-II Pro FPGAs with a model of NBTI. The proposed techniques reduce the effect of NBTI/PV by more than 60% for over 60% of the 15 FPGA chips used in the experiments, with a typical run-time overhead of 1.4-1.8X. The standalone move-acceptance procedure also produces good results with negligible run-time overhead, making it suitable for online FPGA compilation and optimization flows. Assem A. M. Bsoul, Naraig Manjikian |
DATE | 2 |
| 2006 | Enhanced Architectural Support for Variable-Length DecodingabstractThis paper proposes a new architecture for efficient variable-length decoding (VLD) of entropy-coded data for multimedia applications on general-purpose processors. It improves on earlier proposals for low-complexity performance-enhancing hardware structures that exploit prefix/suffix properties of variable-length codes for common multimedia formats. The enhanced architecture is compared to the previous architectures in terms of complexity and operating speed for FPGA implementation, and also in terms of area requirements, power consumption, and operating speed for a 0.18-mum ASIC fabrication process. Simulation results are reported for a pipelined processor with caches executing MPEG-4 software where VLD performance is doubled by incorporating the proposed architecture Mohanarajah Sinnathamby, Subramania Sudharsanan, Naraig Manjikian |
ICME | 3 |
| 2004 | Architecture and Implementation of Chip Multiprocessors: Custom Logic Components and Software for Rapid PrototypingabstractThis work describes components and software tools in support of rapid prototyping in programmable logic for research on chip multiprocessors. Contemporary programmable logic chips offer considerable on-chip logic and memory resources. Prototyping of systems in programmable logic chips is faster and less costly than full-custom chip design. The first contribution that is described in this paper is a collection of original research-oriented logic components that provides processor, memory, and interconnect functionality for rapid prototyping. Because these are original components, and not proprietary vendor-supplied components, they may be arbitrarily extended and modified to suit research needs. The second contribution is a set of enhanced software tools for generating executable code. The third contribution is user-configurable software for testing and evaluating prototype chip multiprocessor implementations in hardware. In addition to describing these contributions, this paper provides results from implementing and testing prototype components and complete chip multiprocessors, including simulation waveforms, logic chip resource utilization, and observations of hardware operation. Naraig Manjikian, Huang Jin, James Reed, Nathan Cordeiro |
ICPP | 1 |
| 2001 | Parallel simulation of multiprocessor execution: implementation and results for simplescalarabstractIn research that relies on simulation in order to predict and compare the performance of proposed computing architectures, multiprocessor simulations have inherent concurrency that can be exploited for parallelization in order to reduce the execution time for a simulation. This paper describes the initial experiences in first introducing multiprocessor simulation support for the detailed out-of-order target simulatorfiom the popular Simplescalar tool set, and then parallelizing the resulting simulator for execution on a multiprocessor host system. The extended simulator provides the basis for further detailed modeling of target systems with multiple out-of-order processors through parallel simulation on a multiprocessor host. For experiments conducted on a Sun Enterprise 3500 platform, the measured speedup for the initial version of the parallelized simulator reached 4.4 on 6processors for a selected application from the SPLASH-2 benchmark. Naraig Manjikian |
ISPASS | 1 |
| 2001 | Hardware/software tradeoffs for IP-over-ATM frame reassembly in an integrated architecture
Peter M. Ewert, Naraig Manjikian |
Comput. Commun. | 2 |
| 2001 | Exploiting Wavefront Parallelism on Large-Scale Shared-Memory MultiprocessorsabstractWavefront parallelism, in which parallelism is limited to hyperplanes in an iteration space, can arise when compilers apply tiling to loop nests to enhance locality. Previous approaches for scheduling wavefront parallelism focused on maximizing parallelism; balancing workloads, and reducing synchronization. In this paper, we show that on large-scale shared-memory multiprocessors, locality is a crucial factor. We make the distinction between intratile and intertile locality and show that as the number of processors grows, intertile locality becomes more important. We consider and experimentally evaluate existing strategies for scheduling wavefront parallelism. We show that dynamic self-scheduling can be efficiently used on a small number of processors, but performs poorly at large scale because it does not enhance intertile locality. By contrast, static scheduling strategies enhance intertile locality for small tiles, maintaining parallelism and resulting in better performance at large scale. Results from a Convex SPP1000 multiprocessor demonstrate the importance of taking intertile locality into account. Static scheduling outperforms dynamic self-scheduling by a factor of up to 2.3 on 30 processors. Naraig Manjikian, Tarek S. Abdelrahman |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2000 | A Vector Multiprocessor for Real-Time Multi-User Detection in Spread-Spectrum CommunicationabstractMulti-user detection in spread-spectrum communications involves repeatedly solving a system of equations to decouple multiple-access interference among users communicating with a basestation. This paper proposes a vector multiprocessor to provide the necessary computing performance. Two vector processors with parallel vector pipelines that are optimized for vector multiply-add operations achieve sustained gigaflop-level performance with a 100 MHz internal clock. Performance analysis of vector assembly-language code on eight vector units indicates that matrix-vector multiplication for a 32/spl times/32 problem can be performed in less than 2 /spl mu/s, and that matrix inversion can be performed in less than 200 /spl mu/s. Naraig Manjikian |
ASAP | 1 |
| 2000 | The NUMAchine MultiprocessorabstractSmall-scale multiprocessors are becoming increasingly economical and common, whereas larger multiprocessors continue to have higher per-node costs. The NUMAchine multiprocessor project seeks to make large-scale multiprocessors more economical while maintaining high performance by exploring architectural and hardware features for low-cost, modular multiprocessors. To demonstrate our approach, we have implemented a prototype system that is scalable to 128 processors. An efficient directory-based cache coherence protocol exploits our hierarchical ring-based interconnect and supports sequential consistency. This paper documents the design choices and the resulting performance of the system using both simulation results and measurements on the prototype hardware. R. Grindley, Tarek S. Abdelrahman, Stephen Brown 0003, S. Caranci, D. DeVries, Benjamin Gamsa, A. Grbic, M. Gusat, R. Ho, Orran Krieger, Guy Lemieux, K. Loveless, Naraig Manjikian, P. McHardy, Sinisa Srbljic, Michael Stumm, Zvonko G. Vranesic, Zeljko Zilic |
ICPP | 13 |
| 1998 | Design and Implementation of the NUMAchine MultiprocessorabstractThis paper describes the design and implementation of the NUMAchine multiprocessor. As the market for CC-NUMA multiprocessors expands, this research project provides a timely architectural design and cost-effective prototype. The key to the successful implementation of our 48-processor prototype is the use of off-the-shelf components and programmable logic devices. Since this machine will serve as a research vehicle for parallel software development, a number of hardware features to enhance experimentation have been included in the design. A. Grbic, Stephen Brown 0003, S. Caranci, R. Grindley, M. Gusat, Guy Lemieux, K. Loveless, Naraig Manjikian, Sinisa Srbljic, Michael Stumm, Zvonko G. Vranesic, Zeljko Zilic |
DAC | 8 |
| 1997 | Combining Loop Fusion with Prefetching on Shared-memory MultiprocessorsabstractThe performance of programs consisting of parallel loops on shared-memory multiprocessors is limited by long memory latencies as processor speeds increase more rapidly than memory speeds. Two complementary techniques for addressing memory latency and improving performance are: (a) cache locality enhancement for latency reduction and (b) data prefetching for latency tolerance. This paper studies the benefit of combining loop fusion for locality enhancement with prefetching. Experimental results are reported for multiprocessors with support for prefetching. For a complete application on an SGI Power Challenge R10000, combining loop fusion with prefetching improves parallel speedup by 46%. Naraig Manjikian |
ICPP | 1 |
| 1997 | Fusion of Loops for Parallelism and LocalityabstractLoop fusion improves data locality and reduces synchronization in data-parallel applications. However, loop fusion is not always legal. Even when legal, fusion may introduce loop-carried dependences which prevent parallelism. In addition, performance losses result from cache conflicts in fused loops. In this paper, we present new techniques to: (1) allow fusion of loop nests in the presence of fusion-preventing dependences, (2) maintain parallelism and allow the parallel execution of fused loops with minimal synchronization, and (3) eliminate cache conflicts in fused loops. We describe algorithms for implementing these techniques in compilers. The techniques are evaluated on a 56-processor KSR2 multiprocessor and on a 18-processor Convex SPP-1000 multiprocessor. The results demonstrate performance improvements for both kernels and complete applications. The results also indicate that careful evaluation of the profitability of fusion is necessary as more processors are used. Naraig Manjikian, Tarek S. Abdelrahman |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1996 | Experience in Designing a Large-scale Multiprocessor using Field-Programmable Devices and Advanced CAD ToolsabstractThis paper provides a case study that shows how a demanding application stresses the capabilities of today's CAD tools, especially in the integration of products from multiple vendors.We relate our experiences in the design of a large, high-speed multiprocessor computer, using state of the art CAD tools.All logic circuitry is targeted to field-programmable devices (FPDs).This choice amplifies the difficulties associated with achieving a highspeed design, and places extra requirements on the CAD tools.Two main CAD systems are discussed in the paper: Cadence Logic Workbench (LWB) is employed for board-level design, and Altera MAX+plusII is used for implementation of logic circuits in FPDs.Each of these products is of great value for our project, but the integration of the two is less than satisfactory.The paper describes a custom procedure that we developed for integrating sub-designs realized in FPDs (via MAX+plusII) into our board-level designs in LWB.We also discuss experiences with Logic Modelling Smart Models, for simulation of FPDs and other types of chips. Stephen Brown 0003, Naraig Manjikian, Zvonko G. Vranesic, S. Caranci, A. Grbic, R. Grindley, M. Gusat, K. Loveless, Zeljko Zilic, Sinisa Srbljic |
DAC | 2 |
| 1995 | Fusion of Loops for Parallelism and Locality
Naraig Manjikian, Tarek S. Abdelrahman |
ICPP (2) | 1 |