VLDB 2026 Research / reviewers in the wild / expert
Mojtaba Mehrara
dblp:21/2430
· DBLP profile ↗
9ranked-venue papers
5as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-authorSoftware engineering, systems software and programming languages · 5 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Parallel and multicore computing · 60% GPUs and heterogeneous computing · 12% Memory systems · 10% | |
| Software engineering, system software, and programming languages
5 papers |
Compilers and program optimization · 63% Concurrent programming · 33% Programming languages and type systems · 4% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel programming models
automatic parallelization |
0.3 | 3 | 2011 | Dynamic parallelization of JavaScript applications using an ultra-lightweight speculation mechanism · HPCA 2011 Parallelizing sequential applications on commodity hardware using a low-cost software transactional memory · PLDI 2009 Uncovering hidden loop level parallelism in sequential applications · HPCA 2008 |
Parallel and multicore computing › speculative parallelization
thread-level speculation |
0.2 | 2 | 2009 | Parallelizing sequential applications on commodity hardware using a low-cost software transactional memory · PLDI 2009 Uncovering hidden loop level parallelism in sequential applications · HPCA 2008 |
Compilers and program optimization › dynamic optimization
adaptive compilation |
0.1 | 1 | 2012 | Adaptive input-aware compilation for graphics engines · PLDI 2012 |
GPUs and heterogeneous computing
GPU compilation |
0.1 | 1 | 2012 | Adaptive input-aware compilation for graphics engines · PLDI 2012 |
Compilers and program optimization › parallelization
dynamic parallelization |
0.1 | 1 | 2011 | Dynamic parallelization of JavaScript applications using an ultra-lightweight speculation mechanism · HPCA 2011 |
Parallel and multicore computing
speculative parallelization |
0.1 | 1 | 2011 | Dynamic parallelization of JavaScript applications using an ultra-lightweight speculation mechanism · HPCA 2011 |
Concurrent programming › concurrency control
conflict detection |
0.1 | 1 | 2009 | Transactional memory with strong atomicity using off-the-shelf memory protection hardware · PPoPP 2009 |
Concurrent programming › atomicity
strong atomicity |
0.1 | 1 | 2009 | Transactional memory with strong atomicity using off-the-shelf memory protection hardware · PPoPP 2009 |
Concurrent programming
transactional memory |
0.1 | 1 | 2009 | Transactional memory with strong atomicity using off-the-shelf memory protection hardware · PPoPP 2009 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2009 | Parallelizing sequential applications on commodity hardware using a low-cost software transactional memory · PLDI 2009 |
Parallel and multicore computing › transactional memory
software transactional memory |
0.1 | 1 | 2009 | Parallelizing sequential applications on commodity hardware using a low-cost software transactional memory · PLDI 2009 |
Compilers and program optimization
dependence analysis |
0.1 | 1 | 2008 | Uncovering hidden loop level parallelism in sequential applications · HPCA 2008 |
Compilers and program optimization
program transformation |
0.1 | 1 | 2008 | Uncovering hidden loop level parallelism in sequential applications · HPCA 2008 |
Compilers and program optimization › parallelization
speculative parallelization |
0.1 | 1 | 2008 | Uncovering hidden loop level parallelism in sequential applications · HPCA 2008 |
Parallel and multicore computing › parallelization strategies
loop parallelism |
0.1 | 1 | 2008 | Uncovering hidden loop level parallelism in sequential applications · HPCA 2008 |
Hardware reliability and fault tolerance
memory fault tolerance |
0.1 | 1 | 2008 | Exploiting selective placement for low-cost memory protection · ACM Trans. Archit. Code Optim. 2008 |
Memory systems
memory protection |
0.1 | 1 | 2008 | Exploiting selective placement for low-cost memory protection · ACM Trans. Archit. Code Optim. 2008 |
GPUs and heterogeneous computing
GPU performance optimization |
0.0 | 1 | 2012 | Adaptive input-aware compilation for graphics engines · PLDI 2012 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.0 | 1 | 2012 | Adaptive input-aware compilation for graphics engines · PLDI 2012 |
Programming languages and type systems › dynamic languages
javascript |
0.0 | 1 | 2011 | Dynamic parallelization of JavaScript applications using an ultra-lightweight speculation mechanism · HPCA 2011 |
Compilers and program optimization › dynamic optimization
profile-guided optimization |
0.0 | 1 | 2009 | Parallelizing sequential applications on commodity hardware using a low-cost software transactional memory · PLDI 2009 |
Parallel and multicore computing › parallel architecture
multicore performance |
0.0 | 1 | 2008 | Uncovering hidden loop level parallelism in sequential applications · HPCA 2008 |
Processor architecture and microarchitecture
chip multiprocessor |
0.0 | 1 | 2007 | Architectural implications of brick and mortar silicon manufacturing · ISCA 2007 |
Methods — techniques the papers use, named apart from their topics
streaming language · 0.3runtime specialization · 0.3adaptive compilation · 0.3reference counting · 0.2conflict detection · 0.2checkpointing · 0.2software transactional memory · 0.2profile-guided loop parallelization · 0.2page-level memory protection · 0.2dynamic code update · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Adaptive input-aware compilation for graphics enginesabstractWhile graphics processing units (GPUs) provide low-cost and efficient platforms for accelerating high performance computations, the tedious process of performance tuning required to optimize applications is an obstacle to wider adoption of GPUs. In addition to the programmability challenges posed by GPU's complex memory hierarchy and parallelism model, a well-known application design problem is target portability across different GPUs. However, even for a single GPU target, changing a program's input characteristics can make an already-optimized implementation of a program perform poorly. In this work, we propose Adaptic, an adaptive input-aware compilation system to tackle this important, yet overlooked, input portability problem. Using this system, programmers develop their applications in a high-level streaming language and let Adaptic undertake the difficult task of input portable optimizations and code generation. Several input-aware optimizations are introduced to make efficient use of the memory hierarchy and customize thread composition. At runtime, a properly optimized version of the application is executed based on the actual program input. We perform a head-to-head comparison between the Adaptic generated and hand-optimized CUDA programs. The results show that Adaptic is capable of generating codes that can perform on par with their hand-optimized counterparts over certain input ranges and outperform them when the input falls out of the hand-optimized programs' "comfort zone". Furthermore, we show that input-aware results are sustainable across different GPU targets making it possible to write and optimize applications once and run them anywhere. Mehrzad Samadi, Amir Hormati, Mojtaba Mehrara, Janghaeng Lee, Scott A. Mahlke |
PLDI | 3 |
| 2011 | Dynamically accelerating client-side web applications through decoupled executionabstractThe emergence and wide adoption of Web applications have moved the client-side component, often written in JavaScript, to the forefront of computing on the Web. Web application developers try to move more computation to the client side to avoid unnecessary network traffic and make the applications more responsive. Therefore, JavaScript applications are becoming larger and more computation intensive. Trace-based just-in-time compilation have been proposed to address the performance bottleneck in these applications. In this paper, we exploit the extra processing power in multicore systems to further improve the performance of trace-based execution of JavaScript programs. In trace-based engines, a considerable portion of execution time is spent on running guards which are operations inserted in the native code to check if the properties assumed by the compiled code actually hold during execution. We introduce ParaGuard to off-load these guards to another thread, while speculatively executing the main trace. In a manner similar to what happens in current trace-based JITs, if a check fails, ParaGuard aborts the native trace execution and reverts back to interpreting the JavaScript bytecode. We also propose several optimizations including guard branch aggregation and profile-based snapshot elimination to further improve the performance of our technique. We show that ParaGuard can achieve an average of 15% performance improvement over current trace-based compilers using an extra processor on commodity multicore processors. Mojtaba Mehrara, Scott A. Mahlke |
CGO | 1 |
| 2011 | Dynamic parallelization of JavaScript applications using an ultra-lightweight speculation mechanismabstractAs the web becomes the platform of choice for execution of more complex applications, a growing portion of computation is handed off by developers to the client side to reduce network traffic and improve application responsiveness. Therefore, the client-side component, often written in JavaScript, is becoming larger and more compute-intensive, increasing the demand for high performance JavaScript execution. This has led to many recent efforts to improve the performance of JavaScript engines in the web browsers. Furthermore, considering the wide-spread deployment of multi-cores in today's computing systems, exploiting parallelism in these applications is a promising approach to meet their performance requirement. However, JavaScript has traditionally been treated as a sequential language with no support for multithreading, limiting its potential to make use of the extra computing power in multicore systems. In this work, to exploit hardware concurrency while retaining traditional sequential programming model, we develop ParaScript, an automatic runtime parallelization system for JavaScript applications on the client's browser. First, we propose an optimistic runtime scheme for identifying parallelizable regions, generating the parallel code on-the-fly, and speculatively executing it. Second, we introduce an ultra-lightweight software speculation mechanism to manage parallel execution. This speculation engine consists of a selective checkpointing scheme and a novel runtime dependence detection mechanism based on reference counting and range-based array conflict detection. Our system is able to achieve an average of 2.18× speedup over the Firefox browser using 8 threads on commodity multi-core systems, while performing all required analyses and conflict detection dynamically at runtime. Mojtaba Mehrara, Po-Chun Hsu, Mehrzad Samadi, Scott A. Mahlke |
HPCA | 1 |
| 2009 | Parallelizing sequential applications on commodity hardware using a low-cost software transactional memoryabstractMulticore designs have emerged as the mainstream design paradigm for the microprocessor industry. Unfortunately, providing multiple cores does not directly translate into performance for most applications. The industry has already fallen short of the decades-old performance trend of doubling performance every 18 months. An attractive approach for exploiting multiple cores is to rely on tools, both compilers and runtime optimizers, to automatically extract threads from sequential applications. However, despite decades of research on automatic parallelization, most techniques are only effective in the scientific and data parallel domains where array dominated codes can be precisely analyzed by the compiler. Thread-level speculation offers the opportunity to expand parallelization to general-purpose programs, but at the cost of expensive hardware support. In this paper, we focus on providing low-overhead software support for exploiting speculative parallelism. We propose STMlite, a light-weight software transactional memory model that is customized to facilitate profile-guided automatic loop parallelization. STMlite eliminates a considerable amount of checking and locking overhead in conventional software transactional memory models by decoupling the commit phase from main transaction execution. Further, strong atomicity requirements for generic transactional memories are unnecessary within a stylized automatic parallelization framework. STMlite enables sequential applications to extract meaningful performance gains on commodity multicore hardware. Mojtaba Mehrara, Jeff Hao, Po-Chun Hsu, Scott A. Mahlke |
PLDI | 1 |
| 2009 | Transactional memory with strong atomicity using off-the-shelf memory protection hardwareabstractThis paper introduces a new way to provide strong atomicity in an implementation of transactional memory. Strong atomicity lets us offer clear semantics to programs, even if they access the same locations inside and outside transactions. It also avoids differences between hardware-implemented transactions and software-implemented ones. Our approach is to use off-the-shelf page-level memory protection hardware to detect conflicts between normal memory accesses and transactional ones. This page-level mechanism ensures correctness but gives poor performance because of the costs of manipulating memory protection settings and receiving notifications of access violations. However, in practice, we show how a combination of careful object placement and dynamic code update allows us to eliminate almost all of the protection changes. Existing implementations of strong atomicity in software rely on detecting conflicts by conservatively treating some non-transactional accesses as short transactions. In contrast, our page-level mechanism lets us be less conservative about how non-transactional accesses are treated; we avoid changes to non-transactional code until a possible conflict is detected dynamically, and we can respond to phase changes where a given instruction sometimes generates conflicts and sometimes does not. We evaluate our implementation with C# versions of many of the STAMP benchmarks, and show how it performs within 25% of an implementation with weak atomicity on all the benchmarks we have studied. It avoids pathological cases in which other implementations of strong atomicity perform poorly. Martín Abadi, Tim Harris 0001, Mojtaba Mehrara |
PPoPP | 3 |
| 2008 | Uncovering hidden loop level parallelism in sequential applicationsabstractAs multicore systems become the dominant mainstream computing technology, one of the most difficult challenges the industry faces is the software. Applications with large amounts of explicit thread-level parallelism naturally scale performance with the number of cores, but single-threaded applications realize little to no gains with additional cores. One solution to this problem is automatic parallelization that frees the programmer from the difficult task of parallel programming and offers hope for handling the vast amount of legacy single-threaded software. There is a long history of automatic parallelization for scientific applications, but the techniques have generally failed in the context of general-purpose software. Thread-level speculation overcomes the problem of memory dependence analysis by speculating unlikely dependences that serialize execution. However, this approach has lead to only modest performance gains. In this paper, we take another look at exploiting loop-level parallelism in single-threaded applications. We show that substantial amounts of loop-level parallelism is available in general-purpose applications, but it lurks beneath the surface and is often obfuscated by a small number of data and control dependences. We adapt and extend several code transformations from the instruction-level and scientific parallelization communities to uncover the hidden parallelism. Our results show that 61% of the dynamic execution of studied benchmarks can be parallelized with our techniques compared to 27% using traditional thread-level speculation techniques, resulting in a speedup of 1.84 on a four core system compared to 1.41 without transformations. Hongtao Zhong, Mojtaba Mehrara, Steven A. Lieberman, Scott A. Mahlke |
HPCA | 2 |
| 2008 | Exploiting selective placement for low-cost memory protectionabstractMany embedded processing applications, such as those found in the automotive or medical field, require hardware designs that are at the same time low cost and reliable. Traditionally, reliable memory systems have been implemented using coded storage techniques, such as ECC. While these designs can effectively detect and correct memory faults such as transient errors and single-bit defects, their use bears a significant cost overhead. In this article, we propose a novel partial memory protection scheme that provides high-coverage fault protection for program code and data, but with much lower cost than traditional approaches. Our approach profiles program code and data usage to assess which program elements are most critical to maintaining program correctness. Critical code and variables are then placed into a limited protected storage resources. To ensure high coverage of program elements, our placement technique considers all program components simultaneously, including code, global variables, stack frames, and heap variables. The fault coverage of our approach is gauged using Monte Carlo fault-injection experiments, which confirm that our technique provides high levels of fault protection (99% coverage) with limited memory protection resources (36% protected area). Mojtaba Mehrara, Todd M. Austin |
ACM Trans. Archit. Code Optim. | 1 |
| 2007 | Low-cost protection for SER upsets and silicon defects
Mojtaba Mehrara, Mona Attariyan, Smitha Shyam, Kypros Constantinides, Valeria Bertacco, Todd M. Austin |
DATE | 1 |
| 2007 | Architectural implications of brick and mortar silicon manufacturingabstractWe introduce a novel chip fabrication technique called "brick and mortar", in which chips are made from small, pre-fabricated ASIC bricks and bonded in a designer-specified arrangement to an inter-brick communication backbone chip. The goal of brick and mortar assembly is to provide a low-overhead method to produce custom chips, yet with performance that tracks an ASIC more closely than an FPGA. This paper examines the architectural design choices in this chip-design system. These choices include the definition of reasonable bricks, both in functionality and size, as well as the communication interconnect that the I/O cap provides. To do this we synthesize candidate bricks, analyze their area and bandwidth demands, and present an architectural design for the inter-brick communication network. We discuss a sample chip design, a 16-way CMP, and analyze the costs and benefits of designing chips with brick and mortar. We find that this method of producing chips incurs only a small performance loss (8%) compared to a fully custom ASIC, which is significantly less than the degradation seen from other low-overhead chip options, such as FPGAs. Finally, we measure the effect that architectural design decisions have on the behavior of the proposed physical brick assembly technique, fluidic self-assembly. Martha Mercaldi Kim, Mojtaba Mehrara, Mark Oskin, Todd M. Austin |
ISCA | 2 |