EDBT 2026 Demo / reviewers in the wild / expert
Hirotaka Ogawa
dblp:83/5001
· DBLP profile ↗
17ranked-venue papers
4as first author
2since 2021 · last 2021
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6Systems, architecture and hardware · 6 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 74% Parallel and multicore computing · 26% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel programming models |
0.5 | 2 | 2021 | Scalable FBP decomposition for cone-beam CT reconstruction · SC 2021 OMPI: Optimizing MPI Programs using Partial Evaluation · SC 1996 |
High-performance computing
domain decomposition |
0.5 | 1 | 2021 | Scalable FBP decomposition for cone-beam CT reconstruction · SC 2021 |
High-performance computing
scientific computing systems |
0.5 | 1 | 2021 | Scalable FBP decomposition for cone-beam CT reconstruction · SC 2021 |
High-performance computing › scientific computing systems
X-ray CT reconstruction |
0.5 | 1 | 2021 | Scalable FBP decomposition for cone-beam CT reconstruction · SC 2021 |
High-performance computing
supercomputing |
0.0 | 1 | 1997 | Multi-client LAN/WAN Performance Analysis of Ninf: a High-Performance Global Computing System · SC 1997 |
Compilers and program optimization
partial evaluation |
0.0 | 1 | 1996 | OMPI: Optimizing MPI Programs using Partial Evaluation · SC 1996 |
Parallel and multicore computing
MPI |
0.0 | 1 | 1996 | OMPI: Optimizing MPI Programs using Partial Evaluation · SC 1996 |
Methods — techniques the papers use, named apart from their topics
filtered backprojection · 0.5task-parallel execution · 0.0data-parallel libraries · 0.0benchmarking · 0.0template functions · 0.0partial evaluation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Performance portable back-projection algorithms on CPUs: agnostic data locality and vectorization optimizationsabstractComputed Tomography (CT) is a key 3D imaging technology that fundamentally relies on the compute-intense back-projection operation to generate 3D volumes. GPUs are typically used for back-projection in production CT devices. However, with the rise of power-constrained micro-CT devices, and also the emergence of CPUs comparable in performance to GPUs, back-projection for CPUs could become favorable. Unlike GPUs, extracting parallelism for back-projection algorithms on CPUs is complex given that parallelism and locality are not explicitly defined and controlled by the programmer, as is the case when using CUDA for instance. We propose a collection of novel back-projection algorithms that reduce the arithmetic computation, robustly enable vectorization, enforce a regular memory access pattern, and maximize the data locality. We also implement the novel algorithms as efficient back-projection kernels that are performance portable over a wide range of CPUs. Performance evaluation using a variety of CPUs from different vendors and generations demonstrates that our back-projection implementation achieves on average 5.2 times speedup over the multi-threaded implementation of the most widely used, and optimized, open library. With a state‐of‐the‐art CPU, we reach performance that rivals top-performing GPUs. Peng Chen 0035, Mohamed Wahib, Xiao Wang 0004, Shin'ichiro Takizawa, Takahiro Hirofuchi, Hirotaka Ogawa, Satoshi Matsuoka |
ICS | 6 |
| 2021 | Scalable FBP decomposition for cone-beam CT reconstructionabstractFiltered Back-Projection (FBP) is a fundamental compute intense algorithm used in tomographic image reconstruction. Cone-Beam Computed Tomography (CBCT) devices use a cone-shaped X-ray beam, in comparison to the parallel beam used in older CT generations. Distributed image reconstruction of cone-beam datasets typically relies on dividing batches of images into different nodes. This simple input decomposition, however, introduces limits on input/output sizes and scalability. Peng Chen 0035, Mohamed Wahib, Xiao Wang 0004, Takahiro Hirofuchi, Hirotaka Ogawa, Ander Biguri, Richard P. Boardman, Thomas Blumensath, Satoshi Matsuoka |
SC | 5 |
| 2020 | Building and Evaluation of Cloud Storage and Datasets Services on AI and HPC Converged InfrastructureabstractAI Bridging Cloud Infrastructure (ABCI) is a world-leading open AI computing infrastructure, for accelerating R&D activities of artificial intelligence. In order to share and reuse AI software assets with ease, ABCI supports container-based application deployment and fine-grained resource allocation on top of the conventional HPC architecture, and provides tens of peta-bytes of high performance storage. One of the on-going major challenges in ABCI, however, is to more efficiently and flexibly exchange and share machine learning models and data related to AI, with other services deployed outside of ABCI in the real world. Our new services called as ABCI Cloud Storage and ABCI Public Datasets are designed for tackling the challenge and taking a role of "Data Harbor" of ABCI. The services allow users to store input and output data of jobs to be run on the ABCI compute nodes, and to share them with not only ABCI users but also non-ABCI users. This paper presents our design and integration of the services to conventional HPC architecture, as a case of ABCI, and reports performance evaluation of them. Based on our attempt and experience, the paper finally summarizes discussion about future direction of the S3 based front data/storage service of the AI and HPC converged system. Yusuke Tanimura, Shin'ichiro Takizawa, Hirotaka Ogawa, Takahiro Hamanishi |
IEEE BigData | 3 |
| 2017 | Understanding and improving disk-based intermediate data caching in SparkabstractApache Spark is a parallel data processing framework that executes fast for iterative calculations and interactive processing, by caching intermediate data in memory with a lineage-based data recovery from faults. The Spark system can also manage data sets larger than memory capacity by placing some cache or all of them on disks on processing nodes. However, the disadvantage is potential performance degradation due to disk I/O and/or serialization. This study aims to clarify efficient/inefficient use of disks in intermediate data caching in Spark and also to improve the usability of disks for end users. In order to achieve the purpose, influence of disk use in data caching was firstly investigated in various aspects, such as caching options, data abstractions and storage devices. The results indicate that serialization cost is dominant rather than disk I/O in most cases. Secondly, a method of combined use of memory and disk was further evaluated under a high memory pressure. Then the method was improved to avoid an excessive re-caching problem, which achieved at most 20-30% reduction of total execution time under a high memory pressure and did not degrade the performance under a low memory pressure, in our experiment with 4 machine learning benchmarks. Finally, this paper summarizes important factors and potential improvements for efficiently using disks in data caching in Spark. Kaihui Zhang, Yusuke Tanimura, Hidemoto Nakada, Hirotaka Ogawa |
IEEE BigData | 4 |
| 2017 | Stinuum: A Holistic Visual Analysis of Moving Objects with Open Source SoftwareabstractWith the development of position tracking technologies and the increasing usage of mobile devices, the analysis of moving objects, such as pedestrians, vehicles, drones, and hurricanes has become an important topic in various applications including intelligent transportation, disaster management, and urban planning. Many of existing studies have focused on managing and analyzing only time-varying locations of point-based objects. However, real-world moving phenomena are space-time continua occupying volumes, having an area at a time; even more, they contain dynamic attributes depending on time and space, such as the velocity of vehicles or the average of wind speed of hurricanes. In this demonstration, we introduce a comprehensive data format to represent various types of temporal geometries and dynamic properties of moving objects based on OGC® Moving Features. Moreover, we present a visual extension of Cesium to visualize moving objects in a space-time cube by cooperating with a data server that manages moving objects in a Cassandra database via RESTful APIs. This demonstration presents how to analyze a correlation between typhoon trajectories and geo-tagged Twitter messages with our systems. Kyoung-Sook Kim 0001, Hyemi Jeong, Hirotaka Ogawa |
SIGSPATIAL/GIS | 4 |
| 2017 | Optimal KD-Partitioning for the Local Outlier Detection in Geo-Social Points
Teerawat Kumrai, Kyoung-Sook Kim 0001, Mianxiong Dong, Hirotaka Ogawa |
ISNN (1) | 4 |
| 2017 | Understanding human perceptual experience in unstructured data on the webabstractComputing for human experience has become more important for understanding all of aspects of any interaction of human beings in the cyber, physical, and social environments. In particular, artificial intelligent technologies based on big data enable to understand natural language, enhance day to day human experience, and make a better decision. In this paper, we propose a method to classify unstructured text data on the Web into the five types of sensation features: sight (ophthalmoception), hearing (audioception), touch (tactioception), smell (olfacception), and taste (gustaoception). Even though sensation is the first process of human experience against the environments, the study of sensation information extraction is neglected due to lack of sensory expression and knowledge comparing with the sentimental analysis or opinion mining. We first define the sensation measurement that is assigned to each feature. Then, we identify which sensation feature has a strong influence on human perceptual experience in a specific topic of corpus. Finally, we evaluate our method by comparing with several baselines in terms of the accuracy. Jun Lee 0002, Kyoung-Sook Kim 0001, Yongjin Kwon, Hirotaka Ogawa |
WI | 4 |
| 2016 | I/O chunking and latency hiding approach for out-of-core sorting acceleration using GPU and flash NVMabstractWe propose an out-of-core sorting acceleration technique, called xtr2sort, that deals with multi-level memory hierarchies of device memory (GPU), host memory (CPU), and semi-external non-volatile memory (Flash NVM) for leveraging the high computational performance and memory bandwidth of GPUs, while offloading bandwidth-oblivious operations onto semi-external memory in order to significantly increasing the memory capacity available for the sort data, well beyond the that of the GPU as well as of the CPU. xtr2sort splits the input records into several chunks to fit in GPU device memory and overlaps (1) I/O operations between semi-external and host memory, (2) data transfers between host and device memory, and (3) sorting on the GPU device in an asynchronous manner for hiding latency. Experimental results show that xtr2sort can sort records up to 256 times larger than is possible with in-core GPU sorting and 16 times larger than is possible with in-core CPU sorting. xtr2sort also achieves 4.39 times faster than out-of-core CPU sorting using 72 threads on 204.8 giga records with int32_t, even though the input records could not fit in the host memory, let alone the GPU device memory. These results indicate that I/O chunking and latency hiding/overlapping maintains sorting performance, despite slow Flash NVM performance, by utilizing GPUs along with good algorithms. Such an approach is viable for accelerating future computing systems with deep memory hierarchies. Hitoshi Sato, Ryo Mizote, Satoshi Matsuoka, Hirotaka Ogawa |
IEEE BigData | 4 |
| 2016 | Performance Prediction of Memory Access Intensive Apps with Delay Insertion: A VisionabstractPredicting performance of a given program on a given machine is highly important because the environment where the program is developed and the one where it is actually executed are often different. However, this prediction is also difficult because the performance of the same program on different machines is not the same, due to the different balances in performance of the various computer components (e.g. CPU, memory, etc.). Although many studies tackle this problem by modelling the target program and/or the target machine, model-based techniques can only provide what they model and cannot leverage existing performance analysis tools. In this paper, we tackle this problem by actually executing the target program in an emulated environment, where the performance balance of the CPU and the memory subsystem is virtually tweaked using a dynamic binary instrumentation technique. We show that this approach can emulate the total execution time of a memory-access-intensive application on different machines, and provide a vision of the future, showing how our approach can outperform existing model-based approaches. Soramichi Akiyama, Takahiro Hirofuchi, Hirotaka Ogawa |
CloudCom | 3 |
| 2016 | Discovery of local topics by using latent spatio-temporal relationships in geo-social mediaabstractSocial networks have played a crucial role as information channels for people to understanding their daily lives beyond merely being communication tools. In particular, coupling social networks with geographic location has boosted the worth of social media to not only enable comprehension of the effects of natural phenomena such as global warming and disasters, but also the social patterns of human societies. However, the high rate of social data generation and the large amounts of noisy data makes it difficult to directly apply social media to decision-making processes. This article proposes a new system of analyzing the spatio-temporal patterns of social phenomena in real time and the discovery of local topics based on their latent spatio-temporal relationships. We will first describe a model that represents the local patterns of populations of geo-tagged social media. We will then define a local topic whose keywords share a region in space and time and present a system implementation based on existing open source technologies. We evaluated the model of local topics with several ways of visualization in experiments and demonstrated a certain social pattern from a dataset of daily Twitter streams. The results obtained from experiments revealed certain keywords had a strong spatio-temporal proximity even though they did not occur in the same message. Kyoung-Sook Kim 0001, Isao Kojima, Hirotaka Ogawa |
Int. J. Geogr. Inf. Sci. | 3 |
| 2012 | Stream processing with BigData: SSS-MapReduceabstractWe propose a Map Reduce based stream processing system, called SSS, which is capable of processing stream along with large scale static data. Unlike the existing stream processing systems that can work only on the relatively small on-memory data-set, SSS can process incoming streamed data consulting the stored data. SSS processes streamed data with continuous Mappers and Reducers, which are periodically invoked by the system. It also supports merge operation on two sets of data, which enables stream data processing with large static data. This poster shows overview of SSS stream processing and preliminary evaluation results. Hidemoto Nakada, Hirotaka Ogawa, Tomohiro Kudoh |
CloudCom | 2 |
| 2010 | SSS: An Implementation of Key-Value Store Based MapReduce FrameworkabstractMapReduce has been very successful in implementing large-scale data-intensive applications. Because of its simple programming model, MapReduce has also begun being utilized as a programming tool for more general distributed and parallel applications, e.g., HPC applications. However, its applicability is limited due to relatively inefficient runtime performance and hence insufficient support for flexible workflow. In particular, the performance problem is not negligible in iterative MapReduce applications. On the other hand, today, HPC community is going to be able to utilize very fast and energy-efficient Solid State Drives (SSDs) with 10 Gbit/sec-class read/write performance. This fact leads us to the possibility to develop "High-Performance MapReduce'', so called. From this perspective, we have been developing a new MapReduce framework called "SSS'' based on distributed key-value store (KVS). In this paper, we first discuss the limitations of existing MapReduce implementations and present the design and implementation of SSS. Although our implementation of SSS is still in a prototype stage, we conduct two benchmarks for comparing the performance of SSS and Hadoop. The results indicate that SSS performs 1-10 times faster than Hadoop. Hirotaka Ogawa, Hidemoto Nakada, Ryousei Takano, Tomohiro Kudoh |
CloudCom | 1 |
| 2009 | A Live Storage Migration Mechanism over WAN for Relocatable Virtual Machine Services on CloudsabstractIaaS (Infrastructure-as-a-Service) is an emerging concept of cloud computing, which allows users to obtain hardware resources from virtualized data centers. Although many commercial IaaS clouds have recently been launched, dynamic virtual machine (VM) migration is not possible among service providers; users are locked into a particular provider, and cannot transparently relocate their VMs to another one for the best cost-effectiveness. In this paper, we propose an advanced storage access mechanism that strongly supports live VM migration over WAN. It rapidly relocates VM disks between source and destination sites with the minimum impact on I/O performance. The proposed mechanism addresses I/O consistency of virtual disks before/after migration, which is the major issue regarding wide-area live migration. The proposed mechanism works as a storage server of a block-level storage I/O protocol (e.g.,iSCSI and NBD). Two key techniques (on-demand fetching and background copying) move on-line virtual disks among remote sites, transparently and efficiently. Our prototype system works perfectly for Xen and KVM without any modification to them. Experiments showed the prototype system also worked successfully for an emulated WAN environment. Takahiro Hirofuchi, Hirotaka Ogawa, Hidemoto Nakada, Satoshi Itoh, Satoshi Sekiguchi |
CCGRID | 2 |
| 2007 | GridASP: an ASP framework for Grid utility computingabstractAbstract One of the greatest evolutions brought about by Grid technology is ‘Grid utility computing’, which utilizes various kinds of IT resources and applications across multiple organizations and enterprises, and integrates them into a comprehensive and valuable service. Since 2004, we proposed and have been developing the GridASP framework, which realizes Grid‐enabled application service providers (ASPs) in order to realize Grid utility computing. GridASP can bind application providers, resource providers, and service providers together and provide application execution services with security and anonymity to enterprise/science users. In this article, we report the conceptual idea of GridASP and the detail of the framework now being developed. Information on GridASP can be found at http://www.gridasp.org . Copyright © 2006 John Wiley & Sons, Ltd. Hirotaka Ogawa, Satoshi Itoh, Tetsuya Sonoda, Satoshi Sekiguchi |
Concurr. Comput. Pract. Exp. | 1 |
| 2000 | OpenJIT: An Open-Ended, Reflective JIT Compiler Framework for Java
Hirotaka Ogawa, Kouya Shimura, Satoshi Matsuoka, Fuyuhiko Maruyama, Yukihiko Sohda, Yasunori Kimura |
ECOOP | 1 |
| 1997 | Multi-client LAN/WAN Performance Analysis of Ninf: a High-Performance Global Computing SystemabstractRapid increase in speed and availability of network of supercomputers is making high-performance global computing possible, including our Ninf system. However, critical issues regarding system performance characteristics in global computing have been little investigated, especially under multi-client, multi-site WAN settings. In order to investigate the feasibility of Ninf and similar systems, we conducted benchmarks under various LAN and WAN environments, and observed the following results: 1) Given sufficient communication bandwidth, Ninf performance quickly overtakes client local performance, 2) current supercomputers are sufficient platforms for supporting Ninf and similar systems in terms of performance and OS fault resiliency, 3) for a vector-parallel machine (Cray J90), employing optimized data-parallel library is a better choice compared to conventional task-parallel execution employed for non-numerical data servers, 4) computationally intensive tasks such as EP can readily be supported under the current Ninf infrastructure, and 5) for communication-intensive applications such as Linpack, server CPU utilization dominates LAN performance, while communication bandwidth dominates WAN performance, and furthermore, aggregate bandwidth could be sustained for multiple clients located at different Internet sites; as a result, distribution of multiple tasks to computing servers on different networks would be essential for achieving higher client-observed performance. Our results are not necessarily restricted to the Ninf system, but rather, would be applicable to other similar global computing systems. Atsuko Takefusa, Satoshi Matsuoka, Hirotaka Ogawa, Hidemoto Nakada, Hiromitsu Takagi, Mitsuhisa Sato, Satoshi Sekiguchi, Umpei Nagashima |
SC | 3 |
| 1996 | OMPI: Optimizing MPI Programs using Partial EvaluationabstractMPI is gaining acceptance as a standard for message-passing in high-performance computing, due to its powerful and flexible support of various communication styles. However, the complexity of its API poses significant software overhead, and as a result, applicability of MPI has been restricted to rather regular, coarse-grained computations. Our OMPI (Optimizing MPI) system removes much of the excess overhead by employing partial evaluation techniques, which exploit static information of MPI calls. Because partial evaluation alone is insufficient, we also utilize template functions for further optimization. To validate the effectiveness for our OMPI system, we performed baseline as well as more extensive benchmarks on a set of application cores with different communication characteristics, on the 64-node Fujitsu AP1000 MPP. Benchmarks show that OMPI improves execution efficiency by as much as factor of two for communication-intensive application core with minimal code increase. It also performs significantly better than previous dynamic optimization technique. Hirotaka Ogawa, Satoshi Matsuoka |
SC | 1 |