EDBT 2026 Demo / reviewers in the wild / expert
Nick P. Johnson
dblp:16/9118
· DBLP profile ↗
12ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 8 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
High-performance computing · 47% Hardware accelerators and domain-specific architectures · 26% Parallel and multicore computing · 13% | |
| Software engineering, system software, and programming languages
5 papers |
Compilers and program optimization · 76% Program analysis · 16% Runtime systems and virtual machines · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 50% Bioinformatics and computational biology · 50% |
Topics — the 16 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › scientific computing accelerator
molecular dynamics accelerator |
0.5 | 1 | 2021 | Anton 3: twenty microseconds of molecular dynamics simulation before lunch · SC 2021 |
High-performance computing › scientific computing systems
molecular dynamics simulation |
0.5 | 1 | 2021 | Anton 3: twenty microseconds of molecular dynamics simulation before lunch · SC 2021 |
High-performance computing
scientific computing systems |
0.5 | 1 | 2021 | Anton 3: twenty microseconds of molecular dynamics simulation before lunch · SC 2021 |
Electronic design automation
high-level synthesis |
0.2 | 1 | 2014 | CGPA: Coarse-Grained Pipelined Accelerators · DAC 2014 |
Compilers and program optimization
dependence analysis |
0.2 | 1 | 2013 | Fast condensation of the program dependence graph · PLDI 2013 |
Program analysis › dynamic analysis
dynamic information-flow tracking |
0.2 | 1 | 2013 | Practical automatic loop specialization · ASPLOS 2013 |
Compilers and program optimization › dependence analysis
program dependence graph |
0.2 | 1 | 2013 | Fast condensation of the program dependence graph · PLDI 2013 |
Compilers and program optimization
program specialization |
0.2 | 1 | 2013 | Practical automatic loop specialization · ASPLOS 2013 |
Bioinformatics and computational biology › molecular informatics › molecular modeling
biomolecular simulation |
0.1 | 1 | 2021 | Anton 3: twenty microseconds of molecular dynamics simulation before lunch · SC 2021 |
Computational science and engineering › computational chemistry
molecular simulation |
0.1 | 1 | 2021 | Anton 3: twenty microseconds of molecular dynamics simulation before lunch · SC 2021 |
Compilers and program optimization › parallelization
automatic parallelization |
0.1 | 1 | 2012 | Speculative separation for privatization and reductions · PLDI 2012 |
Parallel and multicore computing › parallel programming models
automatic parallelization |
0.1 | 1 | 2012 | Speculative separation for privatization and reductions · PLDI 2012 |
GPUs and heterogeneous computing › GPU communication
CPU-GPU communication |
0.1 | 1 | 2011 | Automatic CPU-GPU communication management and optimization · PLDI 2011 |
Hardware accelerators and domain-specific architectures
irregular application acceleration |
0.1 | 1 | 2014 | CGPA: Coarse-Grained Pipelined Accelerators · DAC 2014 |
Runtime systems and virtual machines › interpreter
interpreter optimization |
0.0 | 1 | 2013 | Practical automatic loop specialization · ASPLOS 2013 |
Compilers and program optimization
parallelization |
0.0 | 1 | 2011 | Commutative set: a language extension for implicit parallel programming · PLDI 2011 |
Methods — techniques the papers use, named apart from their topics
custom chip design · 1.0application-specific hardware · 1.0speculative execution · 0.3runtime validation · 0.3semantic commutativity assertions · 0.2runtime management · 0.2compiler transformation · 0.2commutative set · 0.2high-level synthesis · 0.2coarse-grained pipeline parallelism · 0.2invariant-induced pattern specialization · 0.2graph algorithms · 0.2dynamic information flow tracking · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Anton 3: twenty microseconds of molecular dynamics simulation before lunchabstractAnton 3 is the newest member in a family of supercomputers specially designed for atomic-level simulation of molecules relevant to biology (e.g., DNA, proteins, and drug molecules). Anton 3 achieves order-of-magnitude improvements in time-to-solution over its predecessor, Anton 2 (the current state of the art), and is over 100-fold faster than any other currently available supercomputer, thereby enabling broad new avenues of research on critical questions in biology and drug discovery. This speedup means that a 512-node Anton 3 simulates a million atoms at over 100 microseconds per day. Furthermore, Anton 3 attains this performance while consuming an order of magnitude less energy per simulated microsecond than any other machine. Like its predecessors, Anton 3 was designed from the ground up around a new custom chip to best exploit the capabilities offered by new technologies. We present here the main architectural and algorithmic developments that were necessary to achieve such significant advances. David E. Shaw, Peter J. Adams, Asaph Azaria, Joseph A. Bank, Brannon Batson, Alistair Bell, Michael Bergdorf, Jhanvi Bhatt, J. Adam Butts, Timothy Correia, Robert M. Dirks, Ron O. Dror, Michael P. Eastwood, Bruce Edwards, Amos Even, Peter Feldmann, Michael Fenn, Christopher H. Fenton, Anthony Forte, Joseph Gagliardo, Gennette Gill, Maria Gorlatova, Brian Greskamp, J. P. Grossman, Justin Gullingsrud, Anissa Harper, William Hasenplaugh, Mark Heily, Benjamin Colin Heshmat, Jeremy Hunt, Doug Ierardi, Lev Iserovich, Bryan L. Jackson, Nick P. Johnson, Mollie M. Kirk, John L. Klepeis, Jeffrey Kuskin, Kenneth M. Mackenzie, Roy J. Mader, Richard McGowen, Adam McLaughlin, Mark A. Moraes, Mohamed H. Nasr, Lawrence J. Nociolo, Lief O'Donnell, Jon L. Peticolas, Goran Pocina, Cristian Predescu, Terry Quan, John K. Salmon, Carl Schwink, Keun Sup Shim, Naseer Siddique, Jochen Spengler, Tamas Szalay, Raymond Tabladillo, Reinhard Tartler, Andrew G. Taube, Michael Theobald, Brian Towles, William Vick, Stanley C. Wang, Michael Wazlowski, Madeleine J. Weingarten, John M. Williams, Kevin A. Yuh |
SC | 34 |
| 2017 | A Generalized Framework for Automatic Scripting Language ParallelizationabstractComputational scientists are typically not expert programmers, and thus work in easy to use dynamic languages. However, they have very high performance requirements, due to their large datasets and experimental setups. Thus, the performance required for computational science must be extracted from dynamic languages in a manner that is transparent to the programmer. Current approaches to optimize and parallelize dynamic languages, such as just-in-time compilation and highly optimized interpreters, require a huge amount of implementation effort and are typically only effective for a single language. However, scientists in different fields use different languages, depending upon their needs.This paper presents techniques to enable automatic extraction of parallelism within scripts that are universally applicable across multiple different dynamic scripting languages. The key insight is that combining a script with its interpreter, through program specialization techniques, will embed any parallelism within the script into the combined program that can then be extracted via automatic parallelization techniques. Additionally, this paper presents several enhancements to existing speculative automatic parallelization techniques to handle the dependence patterns created by the specialization process. A prototype of the proposed technique, called Partial Evaluation with Parallelization (PEP), is evaluated against two open-source script interpreters with 6 input linear algebra kernel scripts each. The resulting geomean speedup of 5.10× on a 24-core machine shows the potential of the generalized approach in automatic extraction of parallelism in dynamic scripting languages. Taewook Oh, Stephen R. Beard, Nick P. Johnson, Sergiy Popovych, David I. August |
PACT | 3 |
| 2017 | A collaborative dependence analysis framework
Nick P. Johnson, Jordan Fix, Stephen R. Beard, Taewook Oh, Thomas B. Jablin, David I. August |
CGO | 1 |
| 2014 | CGPA: Coarse-Grained Pipelined AcceleratorsabstractHigh-level synthesis (HLS) tools dramatically reduce the nonrecurring engineering cost of creating specialized hardware accelerators. Existing HLS tools are successful in synthesizing efficient accelerators for program kernels with regular memory accesses and simple control flows. For other programs, however, these tools yield poor performance because they invoke computation units for instructions sequentially, without exploiting parallelism. To address this problem, this paper proposes Coarse-Grained Pipelined Accelerators (CGPA), an HLS framework that utilizes coarsegrained pipeline parallelism techniques to synthesize efficient specialized accelerator modules from irregular C/C++ programs without requiring any annotations. Compared to the sequential method, CGPA shows speedups of 3.0x--3.8x for 5 kernels from programs in different domains. Soumyadeep Ghosh, Nick P. Johnson, David I. August |
DAC | 3 |
| 2013 | Practical automatic loop specializationabstractProgram specialization optimizes a program with respect to program invariants, including known, fixed inputs. These invariants can be used to enable optimizations that are otherwise unsound. In many applications, a program input induces predictable patterns of values across loop iterations, yet existing specializers cannot fully capitalize on this opportunity. To address this limitation, we present Invariant-induced Pattern based Loop Specialization (IPLS), the first fully-automatic specialization technique designed for everyday use on real applications. Using dynamic information-flow tracking, IPLS profiles the values of instructions that depend solely on invariants and recognizes repeating patterns across multiple iterations of hot loops. IPLS then specializes these loops, using those patterns to predict values across a large window of loop iterations. This enables aggressive optimization of the loop; conceptually, this optimization reconstructs recurring patterns induced by the input as concrete loops in the specialized binary. IPLS specializes real-world programs that prior techniques fail to specialize without requiring hints from the user. Experiments demonstrate a geomean speedup of 14.1% with a maximum speedup of 138% over the original codes when evaluated on three script interpreters and eleven scripts each. Taewook Oh, Hanjun Kim 0001, Nick P. Johnson, Jae W. Lee, David I. August |
ASPLOS | 3 |
| 2013 | Automatically exploiting cross-invocation parallelism using runtime informationabstractAutomatic parallelization is a promising approach to producing scalable multi-threaded programs for multicore architectures. Many existing automatic techniques only parallelize iterations within a loop invocation and synchronize threads at the end of each loop invocation. When parallel code contains many loop invocations, synchronization can easily become a performance bottleneck. Some automatic techniques address this problem by exploiting cross-invocation parallelism. These techniques use static analysis to partition iterations among threads to avoid crossthread dependences. However, this partitioning is not always achievable at compile-time, because program input determines dependence patterns at run-time. By contrast, this paper proposes DOMORE, the first automatic parallelization technique that uses runtime information to exploit additional cross-invocation parallelism. Instead of partitioning iterations statically, DOMORE dynamically detects crossthread dependences and synchronizes only when necessary. DOMORE consists of a compiler and a runtime library. At compile time, DOMORE automatically parallelizes loops and inserts a custom runtime engine into programs. At run-time, the engine observes dependences and synchronizes iterations only when necessary. For six programs, DOMORE achieves a geomean loop speedup of 2.1× over parallel execution without cross-invocation parallelization and of 3.2 × over sequential execution on eight cores. Jialu Huang, Thomas B. Jablin, Stephen R. Beard, Nick P. Johnson, David I. August |
CGO | 4 |
| 2013 | Fast condensation of the program dependence graphabstractAggressive compiler optimizations are formulated around the Program Dependence Graph (PDG). Many techniques, including loop fission and parallelization are concerned primarily with dependence cycles in the PDG. The Directed Acyclic Graph of Strongly Connected Components (DAGSCC) represents these cycles directly. The naive method to construct the DAGSCC first computes the full PDG. This approach limits adoption of aggressive optimizations because the number of analysis queries grows quadratically with program size, making DAGSCC construction expensive. Consequently, compilers optimize small scopes with weaker but faster analyses. Nick P. Johnson, Taewook Oh, Ayal Zaks, David I. August |
PLDI | 1 |
| 2012 | Automatic speculative DOALL for clustersabstractAutomatic parallelization for clusters is a promising alternative to time-consuming, error-prone manual parallelization. However, automatic parallelization is frequently limited by the imprecision of static analysis. Moreover, due to the inherent fragility of static analysis, small changes to the source code can significantly undermine performance. By replacing static analysis with speculation and profiling, automatic parallelization becomes more robust and applicable. A naïve automatic speculative parallelization does not scale for distributed memory clusters, due to the high bandwidth required to validate speculation. This work is the first automatic speculative DOALL (Spec-DOALL) parallelization system for clusters. We have implemented a prototype automatic parallelization system, called Cluster Spec-DOALL, which consists of a Spec-DOALL parallelizing compiler and a speculative runtime for clusters. Since the compiler optimizes communication patterns, and the runtime is optimized for the cases in which speculation succeeds, Cluster Spec-DOALL minimizes the communication and validation overheads of the speculative runtime. Across 8 benchmarks, Cluster Spec-DOALL achieves a geomean speedup of 43.8x on a 120-core cluster, whereas DOALL without speculation achieves only 4.5x speedup. This demonstrates that speculation makes scalable fully-automatic parallelization for clusters possible. Hanjun Kim 0001, Nick P. Johnson, Jae W. Lee, Scott A. Mahlke, David I. August |
CGO | 2 |
| 2012 | Speculative separation for privatization and reductionsabstractAutomatic parallelization is a promising strategy to improve application performance in the multicore era. However, common programming practices such as the reuse of data structures introduce artificial constraints that obstruct automatic parallelization. Privatization relieves these constraints by replicating data structures, thus enabling scalable parallelization. Prior privatization schemes are limited to arrays and scalar variables because they are sensitive to the layout of dynamic data structures. This work presents Privateer, the first fully automatic privatization system to handle dynamic and recursive data structures, even in languages with unrestricted pointers. To reduce sensitivity to memory layout, Privateer speculatively separates memory objects. Privateer's lightweight runtime system validates speculative separation and speculative privatization to ensure correct parallel execution. Privateer enables automatic parallelization of general-purpose C/C++ applications, yielding a geomean whole-program speedup of 11.4x over best sequential execution on 24 cores, while non-speculative parallelization yields only 0.93x. Nick P. Johnson, Hanjun Kim 0001, Prakash Prabhu, Ayal Zaks, David I. August |
PLDI | 1 |
| 2011 | Automatic CPU-GPU communication management and optimizationabstractThe performance benefits of GPU parallelism can be enormous, but unlocking this performance potential is challenging. The applicability and performance of GPU parallelizations is limited by the complexities of CPU-GPU communication. To address these communications problems, this paper presents the first fully automatic system for managing and optimizing CPU-GPU communcation. This system, called the CPU-GPU Communication Manager (CGCM), consists of a run-time library and a set of compiler transformations that work together to manage and optimize CPU-GPU communication without depending on the strength of static compile-time analyses or on programmer-supplied annotations. CGCM eases manual GPU parallelizations and improves the applicability and performance of automatic GPU parallelizations. For 24 programs, CGCM-enabled automatic GPU parallelization yields a whole program geomean speedup of 5.36x over the best sequential CPU-only execution. Thomas B. Jablin, Prakash Prabhu, James A. Jablin, Nick P. Johnson, Stephen R. Beard, David I. August |
PLDI | 4 |
| 2011 | Commutative set: a language extension for implicit parallel programmingabstractSequential programming models express a total program order, of which a partial order must be respected. This inhibits parallelizing tools from extracting scalable performance. Programmer written semantic commutativity assertions provide a natural way of relaxing this partial order, thereby exposing parallelism implicitly in a program. Existing implicit parallel programming models based on semantic commutativity either require additional programming extensions, or have limited expressiveness. This paper presents a generalized semantic commutativity based programming extension, called Commutative Set (COMMSET), and associated compiler technology that enables multiple forms of parallelism. COMMSET expressions are syntactically succinct and enable the programmer to specify commutativity relations between groups of arbitrary structured code blocks. Using only this construct, serializing constraints that inhibit parallelization can be relaxed, independent of any particular parallelization strategy or concurrency control mechanism. COMMSET enables well performing parallelizations in cases where they were inapplicable or non-performing before. By extending eight sequential programs with only 8 annotations per program on average, COMMSET and the associated compiler technology produced a geomean speedup of 5.7x on eight cores compared to 1.5x for the best non-COMMSET parallelization. Prakash Prabhu, Soumyadeep Ghosh, Yun Zhang 0005, Nick P. Johnson, David I. August |
PLDI | 4 |
| 2010 | DAFT: decoupled acyclic fault toleranceabstractHigher transistor counts, lower voltage levels, and reduced noise margin increase the susceptibility of multicore processors to transient faults. Redundant hardware modules can detect such errors, but software transient fault detection techniques are more appealing for their low cost and flexibility. Recent software proposals double register pressure or memory usage, or are too slow in the absence of hardware extensions, preventing widespread acceptance. This paper presents DAFT, a fast, safe, and memory efficient transient fault detection framework for commodity multicore systems. DAFT replicates computation across multiple cores and schedules fault detection off the critical path. Where possible, values are speculated to be correct and only communicated to the redundant thread at essential program points. DAFT is implemented in the LLVM compiler framework and evaluated using SPEC CPU2000 and SPEC CPU2006 benchmarks on a commodity multicore system. Results demonstrate DAFT's high performance and broad fault coverage. Speculation allows DAFT to reduce the perfor- mance overhead of software redundant multithreading from an average of 200% to 38% with no degradation of fault coverage. Yun Zhang 0005, Jae W. Lee, Nick P. Johnson, David I. August |
PACT | 3 |