EDBT 2026 Demo / reviewers in the wild / expert
Moriyoshi Ohara
dblp:03/2600
· DBLP profile ↗
12ranked-venue papers
2as first author
2since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 1 first-author · 1 since 2021Systems, architecture and hardware · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Hardware accelerators and domain-specific architectures · 56% Emerging computing paradigms · 26% Performance modeling and evaluation · 6% | |
| Theoretical computer science
1 paper |
Computational complexity · 100% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 2 | 2021 | RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021 Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator |
0.5 | 1 | 2021 | RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021 |
Emerging computing paradigms
approximate computing |
0.4 | 1 | 2020 | Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020 |
Hardware accelerators and domain-specific architectures
approximate computing accelerator |
0.4 | 1 | 2020 | Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020 |
Emerging computing paradigms › approximate computing
cross-layer approximate computing |
0.4 | 1 | 2020 | Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020 |
Performance modeling and evaluation
workload characterization |
0.2 | 2 | 2014 | Faster Set Intersection with SIMD instructions by Reducing Branch Mispredictions · Proc. VLDB Endow. 2014 The SPLASH-2 Programs: Characterization and Methodological Considerations · ISCA 1995 |
Processor architecture and microarchitecture › branch prediction
branch misprediction |
0.2 | 1 | 2014 | Faster Set Intersection with SIMD instructions by Reducing Branch Mispredictions · Proc. VLDB Endow. 2014 |
Computational complexity › communication complexity › two-party communication
set intersection |
0.2 | 1 | 2014 | Faster Set Intersection with SIMD instructions by Reducing Branch Mispredictions · Proc. VLDB Endow. 2014 |
Energy-efficient computing › energy-efficient architecture
energy-efficient accelerator |
0.1 | 1 | 2021 | RaPiD: AI Accelerator for Ultra-low Precision Training and Inference · ISCA 2021 |
Machine learning › Deep learning architectures and training
neural network inference |
0.1 | 1 | 2020 | Efficient AI System Design With Cross-Layer Approximate Computing · Proc. IEEE 2020 |
Performance modeling and evaluation
benchmarking |
0.0 | 1 | 1995 | The SPLASH-2 Programs: Characterization and Methodological Considerations · ISCA 1995 |
Parallel and multicore computing › parallel computing › parallel program analysis
parallel application characterization |
0.0 | 1 | 1995 | The SPLASH-2 Programs: Characterization and Methodological Considerations · ISCA 1995 |
Methods — techniques the papers use, named apart from their topics
quantization · 0.9pruning · 0.9mixed-precision arithmetic · 0.9custom number representation · 0.9performance modeling · 0.5merge-based algorithm · 0.4SIMD · 0.4simulation · 0.0measurement · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Detecting Layered Bottlenecks in MicroservicesabstractWe propose a method to detect both software and hardware bottlenecks in a web service consisting of microservices. A bottleneck is a resource that limits the maximum performance of the entire web service. Bottlenecks often include both software resources such as threads, locks, and channels, and hardware resources such as processors, memories, and disks. Bottlenecks form a layered structure since a single request can utilize multiple software resources and a hardware resource simultaneously. The microservice architecture makes the detection of layered bottlenecks challenging due to the lack of a uniform analysis perspective across languages, libraries, frameworks, and middle-ware.We detect layered bottlenecks in microservices by profiling numbers and status of working threads in each microservice and dependency among microservices via network connections. Our approach can be applied to various programming languages since it relies only on standard debugging tools. Nevertheless, our approach not only detects which microservice is a bottleneck but also enables us to understand why it becomes a bottleneck. This is enabled by a novel visualization method to show layered bottlenecks in microservices at a glance. We demonstrate that our approach successfully detects and visualizes layered bottlenecks in the state-of-the-art microservice benchmarks, DeathStarBench and Acme Air microservices. This enables us to optimize the microservices themselves to achieve a higher throughput per re-source utilization rate compared with simply scaling the number of replicas of microservices. Tatsushi Inagaki, Yohei Ueda, Moriyoshi Ohara, Sunyanan Choochotkaew, Marcelo Amaral, Scott Trent, Tatsuhiro Chiba, Qi Zhang 0009 |
CLOUD | 3 |
| 2021 | RaPiD: AI Accelerator for Ultra-low Precision Training and InferenceabstractThe growing prevalence and computational demands of Artificial Intelligence (AI) workloads has led to widespread use of hardware accelerators in their execution. Scaling the performance of AI accelerators across generations is pivotal to their success in commercial deployments. The intrinsic error-resilient nature of AI workloads present a unique opportunity for performance/energy improvement through precision scaling. Motivated by the recent algorithmic advances in precision scaling for inference and training, we designed RaPiD1, a 4-core AI accelerator chip supporting a spectrum of precisions, namely, 16 and 8-bit floating-point and 4 and 2-bit fixed-point. The 36mm2RaPiD chip fabricated in 7nm EUV technology delivers a peak 3.5 TFLOPS/W in HFP8 mode and 16.5 TOPS/W in INT4 mode at nominal voltage. Using a performance model calibrated to within 1% of the measurement results, we evaluated DNN inference using 4-bit fixed-point representation for a 4-core 1 RaPiD chip system and DNN training using 8-bit floating point representation for a 768 TFLOPs AI system comprising 4 32-core RaPiD chips. Our results show INT4 inference for batch size of 1 achieves 3 - 13.5 (average 7) TOPS/W and FP8 training for a mini-batch of 512 achieves a sustained 102 - 588 (average 203) TFLOPS across a wide range of applications. Swagath Venkataramani, Vijayalakshmi Srinivasan, Wei Wang 0333, Sanchari Sen, Ankur Agrawal, Monodeep Kar, Shubham Jain 0004, Alberto Mannari, Hoang Tran, Eri Ogawa, Kazuaki Ishizaki, Hiroshi Inoue, Marcel Schaal, Mauricio J. Serrano, Jungwook Choi, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Allison Allain, James Bonanno, Nianzheng Cao, Robert Casatuta, Matthew Cohen, Bruce M. Fleischer, Michael Guillorn, Howard Haynie, Jinwook Jung, Mingu Kang, Kyu-Hyoun Kim, Siyu Koswatta, Sae Kyu Lee, Martin Lutz, Silvia M. Müller, Jinwook Oh, Ashish Ranjan 0001, Zhibin Ren, Scot Rider, Kerstin Schelm, Michael Scheuermann, Joel Silberman, Vidhi Zalani, Xin Zhang 0025, Ching Zhou, Matthew M. Ziegler, Vinay Shah, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Leland Chang, Kailash Gopalakrishnan |
ISCA | 49 |
| 2020 | Efficient AI System Design With Cross-Layer Approximate ComputingabstractAdvances in deep neural networks (DNNs) and the availability of massive real-world data have enabled superhuman levels of accuracy on many AI tasks and ushered the explosive growth of AI workloads across the spectrum of computing devices. However, their superior accuracy comes at a high computational cost, which necessitates approaches beyond traditional computing paradigms to improve their operational efficiency. Leveraging the application-level insight of error resilience, we demonstrate how approximate computing (AxC) can significantly boost the efficiency of AI platforms and play a pivotal role in the broader adoption of AI-based applications and services. To this end, we present RaPiD, a multi-tera operations per second (TOPS) AI hardware accelerator core (fabricated at 14-nm technology) that we built from the ground-up using AxC techniques across the stack including algorithms, architecture, programmability, and hardware. We highlight the workload-guided systematic explorations of AxC techniques for AI, including custom number representations, quantization/pruning methodologies, mixed-precision architecture design, instruction sets, and compiler technologies with quality programmability, employed in the RaPiD accelerator. Swagath Venkataramani, Xiao Sun 0013, Naigang Wang, Chia-Yu Chen, Jungwook Choi, Mingu Kang, Ankur Agarwal, Jinwook Oh, Shubham Jain 0004, Tina Babinsky, Nianzheng Cao, Thomas W. Fox, Bruce M. Fleischer, George Gristede, Michael Guillorn, Howard Haynie, Hiroshi Inoue, Kazuaki Ishizaki, Michael J. Klaiber, Shih-Hsien Lo, Gary W. Maier, Silvia M. Müller, Michael Scheuermann, Eri Ogawa, Marcel Schaal, Mauricio J. Serrano, Joel Silberman, Christos Vezyrtzis, Wei Wang 0333, Fanchieh Yee, Matthew M. Ziegler, Ching Zhou, Moriyoshi Ohara, Pong-Fei Lu, Brian W. Curran, Sunil Shukla, Vijayalakshmi Srinivasan, Leland Chang, Kailash Gopalakrishnan |
Proc. IEEE | 34 |
| 2019 | Profile-based Detection of Layered BottlenecksabstractDetection of software bottlenecks which hinder utilizing hardware resources is a classic but complex problem due to the layered structures of the software bottlenecks. However, model-based approaches require a performance model given, which is impractical to maintain under today's agile development environment, and profile-based approaches do not handle the layered structures of the software bottlenecks. Tatsushi Inagaki, Yohei Ueda, Takuya Nakaike, Moriyoshi Ohara |
ICPE | 4 |
| 2017 | Accelerating Spark Datasets by Inlining DeserializationabstractApache Spark is a framework for distributed computing that supports the map-reduce programming model. The SQL module of Spark contains Datasets, i.e., distributed collections of records stored in a serialized low-level format in a manually managed chunk of memory. However, the functions users provide to the map-reduce computations expect Java objects. Datasets perform an additional deserialization step beforehand to support the user-provided function, which increases the overhead. We tackled this problem by replacing map functions with their counterparts that accepted the serialized data. This allowed us to skip the unnecessary part of deserialization and achieve faster data processing speeds. Jan Wroblewski, Kazuaki Ishizaki, Hiroshi Inoue, Moriyoshi Ohara |
IPDPS | 4 |
| 2017 | Performance competitiveness of a statically compiled language for server-side Web applicationsabstractWeb developers generally prefer coding in dynamically compiled or scripting languages, such as Java and JavaScript, because those languages allow them to deploy applications quickly to deliver new features and bug fixes in a timely manner. In this paper, we compared the Go language, a popular statically compiled language, with two dynamically compiled languages, JavaScript and Java. We evaluated the Acme Air benchmark for three implementations - one in each language. Our experimental results have shown that the Go implementation achieved a 3.8x and 2.4x higher throughput than the JavaScript and Java implementations respectively, after a simple tuning in the server configuration. Our detailed analysis indicated that this is primarily because Go suffers less from polymorphism due to static typing than JavaScript, and because the Web framework for Go causes less overhead to process RESTful Web service requests than that for Java. We argue that statically compiled languages will play more significant roles for Web applications because of their performance advantages and also because of emerging continuous integration and deployment methodologies which eliminate some of the shortcomings in statically compiled languages over dynamically compiled and scripting languages. Yohei Ueda, Moriyoshi Ohara |
ISPASS | 2 |
| 2016 | Re-constructing high-level information for language-specific binary re-optimizationabstractIn this paper, we show a binary optimizer can achieve competitive performance relative to a state-of-the-art source code compiler by re-constructing high-level information (HLI) from binaries. Recent advances in compiler technologies have resulted in a large performance gap between binaries compiled with old compilers and those compiled with latest ones. This motivated us to develop a binary optimizer for old binaries using a compiler engine for a latest source code compiler. However, a traditional approach to naively convert machine instructions into an intermediate representation (IR) of the compiler engine, does not allow us to take full advantage of optimization techniques available in the compiler. This is because the HLI, such as information about variables and their data types, is not available in such an IR. To address this issue, we have devised a technique to re-construct the HLI from binaries by using contextual information. This contextual information is a set of knowledge about specific compilation technologies, such as the conventions of data structures, the patterns of instruction sequences, and the semantics of runtime routines. With this technique, our binary optimizer has improved the performance of binaries generated from an older compiler by 40.1% on average in the CPU time for a set of benchmarks, which is close to the one due to a source-code recompilation with the same compiler engine, 55.2% on average. Toshihiko Koju, Reid Copeland, Motohiro Kawahito, Moriyoshi Ohara |
CGO | 4 |
| 2014 | Faster Set Intersection with SIMD instructions by Reducing Branch MispredictionsabstractSet intersection is one of the most important operations for many applications such as Web search engines or database management systems. This paper describes our new algorithm to efficiently find set intersections with sorted arrays on modern processors with SIMD instructions and high branch misprediction penalties. Our algorithm efficiently exploits SIMD instructions and can drastically reduce branch mispredictions. Our algorithm extends a merge-based algorithm by reading multiple elements, instead of just one element, from each of two input arrays and compares all of the pairs of elements from the two arrays to find the elements with the same values. The key insight for our improvement is that we can reduce the number of costly hard-to-predict conditional branches by advancing a pointer by more than one element at a time. Although this algorithm increases the total number of comparisons, we can execute these comparisons more efficiently using the SIMD instructions and gain the benefits of the reduced branch misprediction overhead. Our algorithm is suitable to replace existing standard library functions, such as std::set_intersection in C++, thus accelerating many applications, because the algorithm is simple and requires no preprocessing to generate additional data structures. We implemented our algorithm on Xeon and POWER7+. The experimental results show our algorithm outperforms the std::set_intersection implementation delivered with gcc by up to 5.2x using SIMD instructions and by up to 2.1x even without using SIMD instructions for 32-bit and 64-bit integer datasets. Our SIMD algorithm also outperformed an existing algorithm that can leverage SIMD instructions. Hiroshi Inoue, Moriyoshi Ohara, Kenjiro Taura |
Proc. VLDB Endow. | 2 |
| 2012 | Optimizing indirect branches in a system-level dynamic binary translatorabstractA dynamic binary translator (DBT) is a runtime system that translates binary code on the fly, for example to emulate the execution of the binary code on a processor with a different instruction set. One of the major sources of the overhead is the resolution of the branch target addresses for indirect branch instructions. Previous work has addressed this problem for a single virtual address space, but none has addressed it for multiple virtual address spaces in the context of the system-level DBT. This is challenging for compiler optimizations because the compiler cannot compute the virtual addresses of the branch targets for indirect branches at compile-time since they are affected by the runtime states of the emulated TLB. In this paper, we propose a new compiler optimization technique to address the problem for a system-level DBT. Our key idea is to use an offset from the virtual address of each page that contains a branch instruction, since this offset is not affected by the emulated TLB. We found that the compiler can often compute the offset using compile-time constants and that this approach significantly simplifies the guard code necessary for an indirect branch. We implemented this technique in a compiler of a system-level DBT for the z/Architecture. Our experimental results showed our technique can reduce the execution times of the CBW2 benchmarks, part of the standard LSPR benchmark, by up to 5.9% and 2.5% on average. Our analysis indicated that our technique was able to optimize 3.8% of the total dynamic instructions in the original binary code, while completely removing the guard code for 98.9% of these indirect branches. Toshihiko Koju, Ali Ijaz Sheikh, Moriyoshi Ohara, Toshio Nakatani |
SYSTOR | 4 |
| 2009 | The data-centricity of Web 2.0 workloads and its impact on server performanceabstractAdvances in network performance and browser technologies, coupled with the ubiquity of internet access and proliferation of users, have lead to the emergence of a new class of web applications, called Web 2.0. Web 2.0 technologies enable easy collaboration and sharing by allowing users to contribute, modify, and aggregate content using applications like Wikis, Blogs, Social Networking communities, and Mashups. Web 2.0 applications also make heavy use of Ajax, which allows asynchronous communication between client and server, to provide a richer user experience. In this paper, we analyze the effect of these new features on the infrastructure that hosts these workloads. In particular, we focus on the data-centricity, inherent in many Web 2.0 applications, and study its impact on the persistence layer in an application server context. Our experimental results reveal some important performance characteristics; we show that frequent Ajax requests, and other requests arising from the participatory nature of Web 2.0, often retrieve and update persistent data. This can lead to frequent database accesses, lock contention, and reduced performance. We also show that problems in the persistence layer, arising from the data-intensive nature of Web 2.0 applications, can lead to poor scalability that can inhibit us from exploiting current and future multicore architectures. Moriyoshi Ohara, Priya Nagpurkar, Yohei Ueda, Kazuaki Ishizaki |
ISPASS | 1 |
| 2007 | Accelerating Mutual-Information-Based Linear Registration on the Cell Broadband Engine ProcessorabstractEmerging multi-core processors are able to accelerate medical imaging applications by exploiting the parallelism available in their algorithms. We have implemented a mutual-information-based 3D linear registration algorithm on the Cell Broadband Enginetrade processor. By exploiting the highly parallel architecture and its high memory bandwidth, our implementation with two CBE processors can register a pair of 256x256x30 3D images in one second. This implementation is significantly faster than a conventional one on a traditional microprocessor or even faster than a previously reported custom-hardware implementation. In addition to parallelizing the code for multiple cores and organizing the data structure for reducing the amount of the memory traffic, it is also critical to optimize the code for the SIMD pipeline structure. We note that code optimization for the SIMD pipeline alone results in a 4.2x-8.7x acceleration for the computation of small kernels. Further, SIMD optimization alone results in a 4.5x end-end application speedup. Moriyoshi Ohara, Hangu Yeo, Frank Savino, Giridharan Iyengar, Leiguang Gong, Hiroshi Inoue, Hideaki Komatsu, Vadim Sheinin, Shahrokh Daijavad |
ICME | 1 |
| 1995 | The SPLASH-2 Programs: Characterization and Methodological ConsiderationsabstractThe SPLASH-2 suite of parallel applications has recently been released to facilitate the study of centralized and distributed shared-address-space multiprocessors. In this context, this paper has two goals. One is to quantitatively characterize the SPLASH-2 programs in terms of fundamental properties and architectural interactions that are important to understand them well. The properties we study include the computational load balance, communication to computation ratio and traffic needs, important working set sizes, and issues related to spatial locality, as well as how these properties scale with problem size and the number of processors. The other, related goal is methodological: to assist people who will use the programs in architectural evaluations to prune the space of application and machine parameters in an informed and meaningful way. For example, by characterizing the working sets of the applications, we describe which operating points in terms of cache size and problem size are representative of realistic situations, which are not, and which re redundant. Using SPLASH-2 as an example, we hope to convey the importance of understanding the interplay of problem size, number of processors, and working sets in designing experiments and interpreting their results. Steven Cameron Woo, Moriyoshi Ohara, Evan Torrie, Jaswinder Pal Singh, Anoop Gupta |
ISCA | 2 |