EDBT 2026 Demo / reviewers in the wild / expert
Hiroaki Kobayashi
dblp:24/1746
· DBLP profile ↗
61ranked-venue papers
8as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 39 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-authorHuman-computer interaction and ubiquitous computing · 4 · 3 first-authorComputer networks · 2Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive Parallelization Based on Frame-Level and Tile-Level Parallelisms for VVC EncodingabstractABSTRACT To meet the growing demand for high‐efficiency video compression standards, Versatile Video Coding (VVC) has been developed as a successor to High Efficiency Video Coding (HEVC). VVC offers approximately a 50% reduction in bitrate compared to HEVC while maintaining comparable visual quality. However, the improved performance of VVC comes at the cost of a significantly higher computational complexity, resulting in longer encoding times. Consequently, accelerating the VVC encoding process remains a critical challenge for its practical deployment. Leveraging the evolution of multi‐core processors and encoding tools designed for parallel processing, this study introduces an adaptive parallelization method that integrates frame‐level and tile‐level parallelisms. This method dynamically selects the number of concurrently processed frames and tiles by considering reference dependencies and their effects on coding efficiency. Moreover, the proposed method employs a content‐aware, cyclic adjustment of tile configurations to further reduce the encoding time. The evaluation results demonstrate that the proposed method can achieve a 10.58× speedup over the baseline single‐threaded implementation on average, while limiting the increase in BD‐BR to 3.23% and the degradation in BD‐PSNR to only −0.058 dB. These findings confirm that the method substantially decreases the encoding time without degrading coding efficiency. Furthermore, the results also demonstrate that the scalability of the proposed method is better than that of the conventional parallel method, Wavefront parallel processing. Karin Onouchi, Masayuki Sato 0001, Hiroe Iwasaki, Kazuhiko Komatsu, Hiroaki Kobayashi |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | Performance Evaluation of Tsunami Evacuation Route Planning on Multiple Annealing MachinesabstractThis paper focuses on the performance evaluation of annealing machines based on quantum annealing and simulated annealing, i.e., D-wave advantage, Fixstars Amplify Annealing Engine, D-wave Neal, and Vector Annealing, by solving a combinatorial optimization problem, tsunami evacuation route planning. First, the problem is modeled into mathematical formulas and converted into the Quadratic Unconstrained Binary Optimization (QUBO) form as an input, and will be executed on the four annealing machines. Then, the relationship between constraint weight in the objective function and the constraint compliance conditions is also discussed. Finally, the evaluation and discussion are conducted from three aspects, the maximum number of evacuees, Time to Solution (TTS), and the solutions quality. Yihui Liu, Kazuhiko Komatsu, Masahito Kumagai, Masayuki Sato 0001, Hiroaki Kobayashi |
CF | 5 |
| 2023 | Multi-scale Loss based Electron Microscopic Image Pair Matching MethodabstractNanodiffraction Imaging (NDI), a novel imaging technique based on the scanning transmission electron mi-croscopy (STEM), helps the understanding of the relationships between micro-structure and macro-properties. However, the analysis requires image pair matching tasks through meticulous and time-consuming observation and selection by domain experts. Therefore, this paper proposes an image pair matching method for NDI images. The proposed method adopts a two-step training approach. The first step is to perform pre-training by contrastive learning on specialized NDI images. The second step is fine-tuning by a customized model on image pair matching tasks. In the second step, the training is performed with the incorporation of a special loss function, OriDist loss. This loss function is designed to focus on orientation and distribution of multi-scale features. The evaluation results demonstrate the ability of the proposed method to achieve high-accuracy NDI image pair matching, and efficiently reduce the search space of candidate matching images, resulting in a significant reduction in the human workload. Through an extensive ablation study, each component of the proposed method shows positive contributions to the overall performance. Chunting Duan, Kazuhiko Komatsu, Masayuki Sato 0001, Hiroaki Kobayashi |
ICMLA | 4 |
| 2023 | A dynamic parameter tuning method for SpMM parallel executionabstractSummary Sparse matrix‐matrix multiplication (SpMM) is a basic kernel that is used by many algorithms. Several researches focus on various optimizations for SpMM parallel execution. However, a division of a task for parallelization is not well considered yet. Generally, a matrix is equally divided into blocks for processes even though the sparsities of input matrices are different. The parameter that divides a task into multiple processes for parallelization is fixed. As a result, load imbalance among the processes occurs. To balance the loads among the processes, this article proposes a dynamic parameter tuning method by analyzing the sparsities of input matrices. The experimental results show that the proposed method improves the performance of SpMM for examined matrices by up to 39.5% on a single vector engine and 3.49 on a single CPU. Bin Qi 0004, Kazuhiko Komatsu, Masayuki Sato 0001, Hiroaki Kobayashi |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | Analysis of Precision Vectors for Ising-Based Linear Regression
Kaho Aoyama, Kazuhiko Komatsu, Masahito Kumagai, Hiroaki Kobayashi |
PDCAT | 4 |
| 2022 | A Partitioned Memory Architecture with Prefetching for Efficient Video Encoders
Masayuki Sato 0001, Yuya Omori, Ryusuke Egawa, Ken Nakamura, Hiroe Iwasaki, Kazuhiko Komatsu, Hiroaki Kobayashi |
PDCAT | 8 |
| 2021 | Register Flush-free Runahead Execution for Modern Vector ProcessorsabstractModern vector processors have been designed to achieve high sustained performance, especially in HPC applications, because of their powerful instruction set oriented to data-level parallelism. Additionally, the latest vector processor adopts the out-of-order execution of the vector instructions to exploit instruction-level parallelism due to a significant gap in latency between vector arithmetic instructions and vector load/store instructions. In spite of the effort, this gap still brings a deterioration of sustained performance of the modern vector processors. This paper proposes a runahead execution mechanism for the modern vector processors to fill the latency gap by further exploiting instruction-level parallelism. If the processor stalls due to a long latency instruction, the conventional runahead execution mechanism changes the processor state from a normal mode to a runahead mode, and the processor speculatively executes the subsequent instructions that can cause stalls and their dependencies. However, the conventional runahead execution mechanisms flush the registers' values calculated in the runahead mode after finishing this mode and cannot reuse them in the subsequent normal mode. Since the vector processors have many values even in one vector register, these flushes and re-executions waste the bandwidth between cores and caches. Thus, to solve this problem of the conventional runahead mechanism, our proposed mechanism leaves the registers containing the results in the runahead mode in order for the processor to use the registers even after returning to the normal mode. For correctly using these registers after exiting the runahead mode, the proposed mechanism newly realizes functions to inherit the commit order information and the register aliasing information of the runahead-executed instructions into the normal mode. The evaluation results show that the proposed mechanism improves the performance by up to 20% and 3% on average by the conventional mechanism. Hikaru Takayashiki, Masayuki Sato 0001, Kazuhiko Komatsu, Hiroaki Kobayashi |
SBAC-PAD | 4 |
| 2021 | VGL: a high-performance graph processing framework for the NEC SX-Aurora TSUBASA vector architecture
Ilya V. Afanasyev, Vladimir V. Voevodin, Kazuhiko Komatsu, Hiroaki Kobayashi |
J. Supercomput. | 4 |
| 2020 | A Dynamic Parameter Tuning Method for High Performance SpMM
Bin Qi 0004, Kazuhiko Komatsu, Masayuki Sato 0001, Hiroaki Kobayashi |
PDCAT | 4 |
| 2018 | Proposal of Detour Path Suppression Method in PS Reinforcement Learning and Its Application to Altruistic Multi-agent Environment
Daisuke Shiraishi, Kazuteru Miyazaki, Hiroaki Kobayashi |
PRIMA | 3 |
| 2018 | Performance evaluation of a vector supercomputer SX-aurora TSUBASA
Kazuhiko Komatsu, Shintaro Momose, Yoko Isobe, Akihiro Musa, Mitsuo Yokokawa, Toshikazu Aoyama, Masayuki Sato 0001, Hiroaki Kobayashi |
SC | 9 |
| 2018 | Real-time tsunami inundation forecast system for tsunami disaster prevention and mitigationabstractThe tsunami disasters that occurred in Indonesia, Chile, and Japan have inflicted serious casualties and damaged social infrastructures. Tsunami forecasting systems are thus urgently required worldwide. We have developed a real-time tsunami inundation forecast system that can complete a tsunami inundation and damage forecast for coastal cities at the level of 10-m grid size in less than 20 min. As the tsunami inundation and damage simulation is a vectorizable memory-intensive program, we incorporate NEC’s vector supercomputer SX-ACE. In this paper, we present an overview of our system. In addition, we describe an implementation of the program on SX-ACE and evaluate its performance of SX-ACE in comparison with the cases using an Intel Xeon-based system and the K computer. Then, we clarify that the fulfillment of a real-time tsunami inundation forecast system requires a system with high-performance cores connected to the memory subsystem at a high memory bandwidth such as SX-ACE. Akihiro Musa, Hiroshi Matsuoka, Hiroaki Hokari, Takuya Inoue, Yoichi Murashima, Yusaku Ohta, Ryota Hino, Shunichi Koshimura, Hiroaki Kobayashi |
J. Supercomput. | 10 |
| 2017 | Performance and Power Analysis of SX-ACE Using HP-X Benchmark ProgramsabstractAs the SIMD width of modern microprocessors has been widening for keeping up with the computational demand for HPC systems, recently the vector architecture comes back to spotlight. Besides, a modern vector architecture that has been keeping a large SIMD width and a high B/F ratio has survived and evolved in the HPC community. In this paper, to clarify the potential of the modern vector architecture, we present the performance and power analysis of a modern vector supercomputer SX-ACE using HP-X benchmark programs (HPL, HPCG, and HPGMG). Furthermore, the implementation and optimization of these benchmarks on SX-ACE are discussed. The evaluation results show that SX-ACE achieves the highest efficiencies in the HPGMG and HPCG ranking lists. These facts clearly indicate that the powerful vector processing mechanism with a high B/F ratio is mandatory to achieve a high sustained performance in the future HPC systems. Ryusuke Egawa, Kazuhiko Komatsu, Yoko Isobe, Toshihiro Kato, Soya Fujimoto, Hiroyuki Takizawa, Akihiro Musa, Hiroaki Kobayashi |
CLUSTER | 8 |
| 2017 | Vectorization-Aware Loop Optimization with User-Defined Code TransformationsabstractThe cost of maintaining an application code would significantly increase if the application code is branched into multiple versions, each of which is optimized for a different architecture. In this work, default and vector versions of a realworld application code are refactored to be a single version, and the differences between the versions are expressed as user-defined code transformations. As a result, application developers can maintain only the single version, and transform it to its vector version just before the compilation. Although code optimizations for a vector processor are sometimes different from those for other processors, application developers can enjoy the performance of the vector processor without increasing the code complexity. Evaluation results demonstrate that vectorization-aware loop optimization for a vector processor can be expressed as user-defined code transformation rules, and thereby significantly improve the performance of a vector processor without major code modifications. Hiroyuki Takizawa, Thorsten Reimann, Kazuhiko Komatsu, Takashi Soga, Ryusuke Egawa, Akihiro Musa, Hiroaki Kobayashi |
CLUSTER | 7 |
| 2017 | Performance Evaluation of Quantum ESPRESSO on NEC SX-ACEabstractIn recent years, a lot of computer simulation codes have been developed as open-source software. Meanwhile major processors adopt a concept of a vector processing in high performance computing. Hence, the computer simulation codes need to follow a vector processing manner to have a benefit of a computational potential of the vector processing. Our study is evaluation and analysis of performance of various simulation codes developed as open-source software on several vector architectures. In this paper, we evaluate one package of Quantum ESPRESSO as an open-source software code in materials science. Quantum ESPRESSO makes use of several numerical libraries and it is known that parallel parameters called parallelization levels affect the performance in parallel execution. We discuss adjustability of the code to a vector architecture. Moreover, we clarify that the performance of PWscf, which is one of major packages of Quantum ESPRESSO, depends on numerical libraries and parallelization levels of PWscf. For evaluation of performance, we use a vector-parallel supercomputer system named NEC SX-ACE and an Intel Xeon-based cluster system named NEC LX 406Re-2. From this evaluation, we confirm that the code is suitable for vector architectures. Additionally, we clarify the effectiveness of applying optimum numerical libraries to each architecture with appropriate parallel parameters to obtain the high performance. Akihiro Musa, Hiroaki Hokari, Shivanshu Kumar Singh, Raghunandan Mathur, Hiroaki Kobayashi |
CLUSTER | 6 |
| 2017 | Potential of a modern vector supercomputer for practical applications: performance evaluation of SX-ACEabstractAchieving a high sustained simulation performance is the most important concern in the HPC community. To this end, many kinds of HPC system architectures have been proposed, and the diversity of the HPC systems grows rapidly. Under this circumstance, a vector-parallel supercomputer SX-ACE has been designed to achieve a high sustained performance of memory-intensive applications by providing a high memory bandwidth commensurate with its high computational capability. This paper examines the potential of the modern vector-parallel supercomputer through the performance evaluation of SX-ACE using practical engineering and scientific applications. To improve the sustained simulation performances of practical applications, SX-ACE adopts an advanced memory subsystem with several new architectural features. This paper discusses how these features, such as MSHR, a large on-chip memory, and novel vector processing mechanisms, are beneficial to achieve a high sustained performance for large-scale engineering and scientific simulations. Evaluation results clearly indicate that the high sustained memory performance per core enables the modern vector supercomputer to achieve outstanding performances that are unreachable by simply increasing the number of fine-grain scalar processor cores. This paper also discusses the performance of the HPCG benchmark to evaluate the potentials of supercomputers with balanced memory and computational performance against heterogeneous and cutting-edge scalar parallel systems. Ryusuke Egawa, Kazuhiko Komatsu, Shintaro Momose, Yoko Isobe, Akihiro Musa, Hiroyuki Takizawa, Hiroaki Kobayashi |
J. Supercomput. | 7 |
| 2015 | Design of tendon-driven mechanisms for fault tolerance from tendon-breaking by using centroid vectorsabstractTendon-driven mechanisms (TDMs) are mechanisms driven by more tendons than joints. This tendon redundancy is often used for changing the hardware compliance. In contrast, in this paper, the redundancy is used for coping with unexpected tendon breaking. We analyze kinematics of a class of TDMs with asymmetrically spanned tendons and modify it to design a TDM that preserves the controllability even if at least one of the tendons is broken. We validate the proposed method to design several TDMs. Ryuta Ozawa, Hiroaki Kobayashi, Kazuhito Hyodo |
ICRA | 2 |
| 2015 | A Visualization Technique to Support Searching and Comparing Features of Multivariate DatasetsabstractIn exploratory analysis of multivariate datasets, performing an analytical task is often necessary. Such tasks may include extracting characteristic subsets and comparing them. Therefore, we support searching and comparing features of multivariate datasets. We developed Blade Graph, which is a visualization technique for comparing distributions by emphasizing coloring according to the size of the difference. In addition, we developed a visual analysis tool with representations for comparing data distributions. In a case study of our analysis tool, we analyzed collective tendencies from a social media dataset. Hiroaki Kobayashi, Hiroko Suzuki, Kazuo Misue |
IV | 1 |
| 2014 | Design and control methodology for fine grain power gating based on energy characterization and code profiling of microprocessorsabstractThis paper presents a design and control scheme of a microprocessor whose internal function units are power gated at instruction-by-instruction basis. Enabling/disabling the power gating is adaptively controlled under the support of on-chip leakage monitors and the operating system to minimize energy overhead due to sleep-in and wakeup. Measured results of the fabricated chip in the 65nm CMOS technology demonstrated that our approach reduces energy to 21-35% in the range of 25-85°C as compared to the non power-gated case. Energy dissipation was reduced by up to 15% as compared to the conventional fine-grain power gating technique in the same temperature range. Kimiyoshi Usami, Masaru Kudo, Kensaku Matsunaga, Tsubasa Kosaka, Yoshihiro Tsurui, Hideharu Amano, Hiroaki Kobayashi, Ryuichi Sakamoto, Mitaro Namiki, Masaaki Kondo, Hiroshi Nakamura |
ASP-DAC | 8 |
| 2014 | Design and evaluation of fine-grained power-gating for embedded microprocessorsabstractPower-performance efficiency is still remaining a primary concern for microprocessor designers. One of the sources of power inefficiency for recent LSI chips is increasing leakage power consumption. Power-gating is a well known technique to reduce leakage power consumption by switching off the power supply to idle logic blocks. Recently, fine-grained power-gating is emerged as a technique to minimize leakage current during the active processor cycles by switching on and off a logic blocks in much finer temporal/spatial granularity. Though fine-grained power-gating is useful, a comprehensive evaluation and analysis has not been conducted on a real LSI chips. In this paper, we evaluate fine-grained run-time power-gating for microprocessors' functional units using a real embedded microprocessor. We also introduce an architecture and compiler co-operative power-gating scheme which mitigates negative power reduction caused by the energy overhead associated with finegrained power-gating. The experimental results with a fabricated core shows that a hardware-based scheme saves power consumption of functional units by 44% and hardware compiler co-operative scheme further improves power efficiency by 5.9% when core temperature is 25 ˚C. Masaaki Kondo, Hiroaki Kobayashi, Ryuichi Sakamoto, Motoki Wada, Jun Tsukamoto, Mitaro Namiki, Hideharu Amano, Kensaku Matsunaga, Masaru Kudo, Kimiyoshi Usami, Toshiya Komoda, Hiroshi Nakamura |
DATE | 2 |
| 2014 | Xevolver: An XML-based code translation framework for supporting HPC application migrationabstractThis paper proposes an extensible programming framework to separate platform-specific optimizations from application codes. The framework allows programmers to define their own code translation rules for special demands of individual systems, compilers, libraries, and applications. Code translation rules associated with user-defined compiler directives are defined in an external file, and the application code is just annotated by the directives. For code transformations based on the rules, the framework exposes the abstract syntax tree (AST) of an application code as an XML document to expert programmers. Hence, the XML document of an AST can be transformed using any XML-based technologies. Our case studies using real applications demonstrate that the framework is effective to separate platform-specific optimizations from application codes, and to incrementally improve the performance of an existing application without messing up the code. Hiroyuki Takizawa, Shoichi Hirasawa, Yasuharu Hayashi, Ryusuke Egawa, Hiroaki Kobayashi |
HiPC | 5 |
| 2014 | Parallel Box: Visually Comparable Representation for Multivariate Data AnalysisabstractIn visual analytics, data comparison is a means of analyzing data. We developed Parallel Box to support the visual analysis of multivariate data by facilitating the flexible comparison of numerous multivariate items. To compare the data distributions of multivariate data, we combine cumulative bar charts and box plots, tools that are widely used in statistics. Using shadow expression based on the visual Gestalt principles of grouping, Parallel Box enables a direct comparison between either datasets or variables. We performed a social media analysis as a case study of Parallel Box. The results of our analysis confirm that Parallel Box is useful for visual analysis. Hiroaki Kobayashi, Tadanobu Furukawa, Kazuo Misue |
IV | 1 |
| 2014 | Analysis, Classification, and Design of Tendon-Driven MechanismsabstractThis paper analyzes tendon-driven mechanisms (TDMs) with active and passive tendons and proposes a method for designing TDMs. First, we group TDMs into six classes according to their controllability and the number of driving degrees of freedom. In this classification system, the conventional underactuated mechanisms are grouped into three classes, two of which have often previously been grouped together although they have different manipulation abilities. Next, we analyze bias forces to separate and decouple a given TDM into several smaller TDMs. Finally, we propose a design method for combining smaller TDMs into an appropriate TDM. Using this method, we can easily determine an appropriate tendon transmission that meets the requirements for the number of tendons, the hardware of the tendon routing, and arbitrary joint constraint that has useful applications in prosthetic and biomimetic hands. Numerical examples show that the proposed method guarantees a nonsingular series actuation transmission and that the designed underactuated TDMs could achieve arbitrary stiffness independent of the actuation effort compared with the conventional underactuated TDM with torsion springs. Ryuta Ozawa, Hiroaki Kobayashi, Kazunori Hashirii |
IEEE Trans. Robotics | 2 |
| 2013 | Colored Mosaic Matrix: Visualization Technique for High-Dimensional DataabstractOwing to a limited display resolution, it may be difficult to obtain an overview of high-dimensional data in the display area used for visualization. In this paper, we aimed to obtain an overview of high-dimensional data in a limited screen area. We developed Colored Mosaic Matrix as a method to obtain a data overview. Colored Mosaic Matrix is a visualization method for high-dimensional categorical data that uses a color representation of the features. By representing quantitative data in category units, the proposed method enables the visualization of data containing a large number of records. As a result of an experimental investigation of its readability, we found our method to be useful in obtaining a data overview. Hiroaki Kobayashi, Kazuo Misue, Jiro Tanaka |
IV | 1 |
| 2012 | Evaluation of the Improved Penalty Avoiding Rational Policy Making Algorithm in Real World Environment
Kazuteru Miyazaki, Masaki Itou, Hiroaki Kobayashi |
ACIIDS (1) | 3 |
| 2012 | GPU implementation of phase-based stereo correspondence and its applicationabstractThis paper proposes a Graphics Processing Unit (GPU) implementation of the stereo correspondence matching using Phase-Only Correlation (POC). The use of high-accuracy stereo correspondence matching based on POC makes it possible to measure accurate 3D shape of the object using stereo vision, while the drawback of POC-based approach is its high computational cost. Addressing this problem, we propose a GPU implementation of POC-based correspondence matching. Through a set of experiments using a variety of GPUs, we demonstrate that the proposed implementation is high-speed and high-efficiency compared with the CPU implementation. We also apply the proposed approach to a real-time 3D measurement system. Mamoru Miura, Kinya Fudano, Koichi Ito 0001, Takafumi Aoki, Hiroyuki Takizawa, Hiroaki Kobayashi |
ICIP | 6 |
| 2011 | CheCL: Transparent Checkpointing and Process Migration of OpenCL ApplicationsabstractIn this paper, we propose a new transparent checkpoint/restart (CPR) tool, named CheCL, for high-performance and dependable GPU computing. CheCL can perform CPR on an OpenCL application program without any modification and recompilation of its code. A conventional check pointing system fails to checkpoint a process if the process uses OpenCL. Therefore, in CheCL, every API call is forwarded to another process called an API proxy, and the API proxy invokes the API function, two processes, an application process and an API proxy, are launched for an OpenCL application. In this case, as the application process is not an OpenCL process but a standard process, it can be safely check pointed. While CheCL intercepts all API calls, it records the information necessary for restoring OpenCL objects. The application process does not hold any OpenCL handles, but CheCL handles to keep such information. Those handles are automatically converted to OpenCL handles and then passed to API functions. Upon restart, OpenCL objects are automatically restored based on the recorded information. This paper demonstrates the feasibility of transparent check pointing of OpenCL programs including MPI applications, and quantitatively evaluates the runtime overheads. It is also discussed that CheCL can enable process migration of OpenCL applications among distinct nodes, and among different kinds of compute devices such as a CPU and a GPU. Hiroyuki Takizawa, Kentaro Koyama, Katsuto Sato, Kazuhiko Komatsu, Hiroaki Kobayashi |
IPDPS | 5 |
| 2011 | A History-Based Performance Prediction Model with Profile Data Classification for Automatic Task Allocation in Heterogeneous Computing SystemsabstractIn this paper, we propose a runtime performance prediction model for automatic selection of accelerators to execute kernels in OpenCL. The proposed method is a history-based approach that uses profile data for performance prediction. The profile data are classified into some groups, from each of which its own performance model is derived. As the execution time of a kernel depends on some runtime parameters such as kernel arguments, the proposed method first identifies parameters affecting the execution time by calculating the correlation between each parameter and the execution time. A parameter with weak correlation is used for the classification of the profile data and the selection of the performance prediction model. A parameter with strong correlation is used for building a linear model for the prediction of the kernel execution time by using only the classified profile data. Experimental results clearly indicate that the proposed method can achieve more accurate performance prediction than conventional history-based approaches because of the profile data classification. Katsuto Sato, Kazuhiko Komatsu, Hiroyuki Takizawa, Hiroaki Kobayashi |
ISPA | 4 |
| 2010 | A Load-Forwarding Mechanism for the Vector Architecture in Multimedia ApplicationsabstractNowadays, multimedia applications (MMAs) form an important workload for general purpose processors. Although the vector architecture is considered the most potential candidate for media processing, the traditional vector architecture has inefficiencies to execute MMAs. This paper proposes a media-oriented vector architecture, which improves the traditional one with a load-forwarding mechanism. The load-forwarding mechanism overcomes the inefficiency on utilization of the memory bandwidth. As a result, the proposed architecture achieves a higher performance with lower hardware cost than the traditional one. This paper evaluates the proposed architecture with architectural design parameters and finds out the most efficient size for the vector architecture when performing MMAs. Ryusuke Egawa, Hiroyuki Takizawa, Hiroaki Kobayashi |
DSD | 4 |
| 2010 | A voting-based working set assessment scheme for dynamic cache resizing mechanismsabstractConsidering the trade-off between performance and power consumption has become significantly important in multi-core processor design. Under this situation, one promising approach is to employ a power-aware dynamic cache partitioning mechanism. This mechanism individually manages activation of each cache way, and exclusively allocates the minimum number of required ways to each thread. In the mechanism, an appropriate number of ways for a thread is decided based on locality assessment. However, sampling results of cache accesses that are used for locality assessment are disturbed by exceptional behaviors of cache accesses, which happen in a very short period. Such sampling results may change locality assessment results to ones that are not along with the overall trend in a long access-sampling period. These assessment results will excessively adapt the cache to exceptional behaviors, and deteriorate energy efficiency. To avoid such excessive adaptation by the exceptional behaviors, this paper proposes a voting-based working set assessment scheme, in which the number of activated ways is adjusted based on majority voting of locality assessment of several short sampling periods. By using the majority voting, the proposed scheme can identify the periods including exceptional behaviors, and ignore the assessment results of these periods. As a result, the proposed scheme makes the cache resizing mechanism more stable and robust. The experimental results indicate that the proposed scheme can reduce energy consumption by up to 24%, and 10% on an average without significant performance degradation in multi-thread execution on a 2-core CMP. Masayuki Sato 0001, Ryusuke Egawa, Hiroyuki Takizawa, Hiroaki Kobayashi |
ICCD | 4 |
| 2009 | Design and control of underactuated tendon-driven mechanismsabstractMany robotic hands or prosthetic hands have been developed in the last several decades, and many use tendon-driven mechanisms for their transmissions. Robotic hands are now built with underactuated mechanisms, which have fewer actuators than degrees of freedom, to reduce mechanical complexity or to realize a biomimetic motion such as flexion of an index finger. The design is heuristic and it is useful to develop design methods for the underactuated mechanisms. This paper classifies mechanisms driven by tendons into three classes, and proposes a design method for them. The two classes are related to underactuated tendon-driven mechanisms, and these have been used without distinction so far. An index finger robot, which has four active tendons and two passive tendons, is developed and controlled with the proposed method. Ryuta Ozawa, Kazunori Hashirii, Hiroaki Kobayashi |
ICRA | 3 |
| 2009 | CheCUDA: A Checkpoint/Restart Tool for CUDA ApplicationsabstractIn this paper, a tool named CheCUDA is designed to checkpoint CUDA applications that use GPUs as accelerators. As existing checkpoint/restart implementations do not support checkpointing the GPU status, CheCUDA hooks a part of basic CUDA driver API calls in order to record the status changes on the main memory. At checkpointing, CheCUDA stores the status changes in a file after copying all necessary data in the video memory to the main memory and then disabling the CUDA runtime. At restarting, CheCUDA reads the file, re-initializes the CUDA runtime, and recovers the resources on GPUs so as to restart from the stored status. This paper demonstrates that a prototype implementation of CheCUDA can correctly checkpoint and restart a CUDA application written with basic APIs. This also indicates that CheCUDA can migrate a process from one PC to another even if the process uses a GPU. Accordingly, CheCUDA is useful not only to enhance the dependability of CUDA applications but also to enable dynamic task scheduling of CUDA applications required especially on heterogeneous GPU cluster systems. This paper also shows the timing overhead for checkpointing. Hiroyuki Takizawa, Katsuto Sato, Kazuhiko Komatsu, Hiroaki Kobayashi |
PDCAT | 4 |
| 2009 | Performance evaluation of NEC SX-9 using real science and engineering applicationsabstractThis paper describes a new-generation vector parallel supercomputer, NEC SX-9 system. The SX-9 processor has an outstanding core to achieve over 100Gflop/s, and a software-controllable on-chip cache to keep the high ratio of the memory bandwidth to the floating-point operation rate. Moreover, its large SMP nodes of 16 vector processors with 1.6Tflop/s performance and 1TB memory are connected with dedicated network switches, which can achieve inter-node communication at 128GB/s per direction. The sustained performance of the SX-9 processor is evaluated using six practical applications in comparison with conventional vector processors and the latest scalar processor such as Nehalem-EP. Based on the results, this paper discusses the performance tuning strategies for new-generation vector systems. An SX-9 system of 16 nodes is also evaluated by using the HPC challenge benchmark suite and a CFD code. Those evaluation results clarify the highest sustained performance and scalability of the SX-9 system. Takashi Soga, Akihiro Musa, Yoichi Shimomura, Ryusuke Egawa, Ken'ichi Itakura, Hiroyuki Takizawa, Koki Okabe, Hiroaki Kobayashi |
SC | 8 |
| 2008 | A Performance Study of Secure Data Mining on the Cell ProcessorabstractThis paper examines the potential of the Cell processor as a platform for secure data mining on the future volunteer computing systems. Volunteer computing platforms have the potential to provide massive computing power. However, privacy and security concerns prevent using volunteer computing for data mining of sensitive data. The Cell processor comes with a hardware security feature. The secure volunteer data mining can be achieved by using this hardware security feature. In this paper, we present a general security scheme for the volunteer computing, and a secure parallelized K-Means clustering algorithm for the Cell processor. We also evaluate the performance of the algorithm on the Cell secure system simulator. Evaluation results indicate that the proposed secure data clustering outperforms a non-secure clustering algorithm on the general purpose CPU, but incurs a huge performance overhead introduced by the decryption process of the Cell security features. Hong Wang 0006, Hiroyuki Takizawa, Hiroaki Kobayashi |
CCGRID | 3 |
| 2008 | SPRAT: Runtime processor selection for energy-aware computingabstractA commodity personal computer (PC) can be seen as a hybrid computing system equipped with two different kinds of processors, i.e. CPU and a graphics processing unit (GPU). Since the superiorities of GPUs in the performance and the power efficiency strongly depend on the system configuration and the data size determined at the runtime, a programmer cannot always know which processor should be used to execute a certain kernel. Therefore, this paper presents a runtime environment that dynamically selects an appropriate processor so as to improve the energy efficiency. The evaluation results clearly indicate that the runtime processor selection at executing each kernel with given data streams is promising for energy-aware computing on a hybrid computing system. Hiroyuki Takizawa, Katsuto Sato, Hiroaki Kobayashi |
CLUSTER | 3 |
| 2008 | Implementation and evaluation of a distributed and cooperative load-balancing mechanism for dependable volunteer computingabstractThis paper proposes a P2P-based dynamic load balancing mechanism to increase the dependability of volunteer computing. The proposed mechanism is incorporated into a volunteer computing middleware, called the Berkeley Open Infrastructure for Network Computing(BOINC). The proposed mechanism provides two additional features: decentralized load balancing and proxy download. The former feature reduces the variation of the execution times for individual tasks, which are usually aggravated by dynamic and unpredictable load changes on volunteer computing resources. The latter offers another way to assign tasks to idle computing resources when the BOINC project server fails in the task assignment. Using a prototype implementation, this paper examines the effect of the proposed mechanism on the performance of a real volunteer computing system. The experimental results show that the proposed mechanism can reduce the maximum turnaround time by 42% and further improve the total throughput of the volunteer computing system by 27%. Yoshitomo Murata, Tsutomu Inaba, Hiroyuki Takizawa, Hiroaki Kobayashi |
DSN | 4 |
| 2008 | A robotic finger equipped with an optical three-axis tactile sensorabstractIn a previous paper we developed an optical three-axis tactile sensor that can acquire normal and shearing forces to be mounted on a robotic finger. Normal and shearing forces applied to the sensing element were detected separately; when we examined the repeatability of the present tactile sensor with 1,000 loading-unloading cycles, the respective error of the normal forces was 2%. In the present paper, the three-axis tactile sensor is mounted on a robotic finger of three degrees of freedom to evaluate it for dexterous hands. A series of three kinds of experiments were performed. First, the robotic hand touches and scans flat specimens to evaluate the sensing ability of the friction coefficient. Second, it detects the contour of parallelepiped and cylindrical objects. Finally, it manipulates a parallelepiped case put on a table by sliding it on the table. Since the present robotic hand was able to perform the above three tasks with appropriate precision, we expected that it would be applicable to dexterous hands in subsequent studies. Masahiro Ohka, Nobuyuki Morisawa, Hirofumi Suzuki, Jumpei Takata, Hiroaki Kobayashi, Hanafiah Yussof |
ICRA | 5 |
| 2008 | Effects of MSHR and Prefetch Mechanisms on an On-Chip Cache of the Vector ArchitectureabstractVector supercomputers have been encountering the memory wall problem and their memory bandwidth per flop/s rate has decreased. To cover the insufficient memory bandwidth per flop/s rate, an on-chip vector cache has been proposed for the vector processors. Although vector caching is effective to increase the sustained performance to a certain degree, it still needs software and hardware supporting mechanisms to extract its potential. To this end, we propose miss status handling registers (MSHR) and a prefetch mechanism. This paper evaluates the performance of the vector cache with the MSHR and the prefetch mechanism on the vector supercomputer across three leading scientific applications. The MSHR is an effective mechanism for handling subsequent vector loads of the same data, which frequently appear in different schemes. The experimental results indicate that the MSHR can improve the computational performance of scientific applications by 1.45×. Moreover, we examine the performance of the prefetch mechanism on the vector cache. The prefetch mechanism increases the computational performance by 1.6×. Accordingly, the MSHR and the prefetching mechanism are very effective optimization options for vector caching of future vector supercomputers even if the vector supercomputers cannot maintain the current memory bandwidth per flop/s rate. Akihiro Musa, Yoshiei Sato, Takashi Soga, Ryusuke Egawa, Hiroyuki Takizawa, Koki Okabe, Hiroaki Kobayashi |
ISPA | 7 |
| 2008 | A Utility-Based Double Auction Mechanism for Efficient Grid Resource AllocationabstractIn Grid Computing, harnessing the power of idle resources in a distributed environment is one of the important features. However, to fully benefit from this computing model, an appropriated resource allocation method needs to be carefully chosen and deployed. A number of studies have been done on this area and one of the promising approaches is to adopt a marketing scheme called the auction model, which has been drawing much attention during past several years. In this paper, we propose a new utility-aware resource allocation protocol to make external scheduling decision in Grid. Users and service providers specify one or more weight values, and then, an auctioneer uses these values for calculating both userspsila and service providerspsila utility values which reflect preference upon the matched members in the different group. Then, we map the scheduling problem with these utility values into the problem in a weighted bipartite graph, and propose a new matching algorithm based on the existing SMP (Stable Marriage Problem) matching algorithm. Finally, the performance of this auctionpsilas awarding technique is evaluated. Chainan Satayapiwat, Ryusuke Egawa, Hiroyuki Takizawa, Hiroaki Kobayashi |
ISPA | 4 |
| 2007 | Multi-core data streaming architecture for ray tracingabstractRay tracing is a computer graphics technique to generate photo-realistic images. All though it can generate precise and realistic images, it requires a large amount of computation. The intersection test between rays and objects is one of the dominant factors of the ray tracing speed. We propose a new parallel processing architecture, named R PL S, for accelerating the ray tracing computation. R PL S boosts the speed of the intersection test by using a new algorithm based on ray-casting through planes, data streaming architecture offers highly efficient data provision in a multi-core environment. We estimate the performance of a future SoC implementation of R PL S by software simulation, and show 600 times speedup over a conventional CPU implementation. Yoshiyuki Kaeriyama, Daichi Zaitsu, Ken-Ichi Suzuki, Hiroaki Kobayashi, Nobuyuki Ohba |
ICCD | 4 |
| 2007 | A dependable Peer-to-Peer computing platform
Hong Wang 0006, Hiroyuki Takizawa, Hiroaki Kobayashi |
Future Gener. Comput. Syst. | 3 |
| 2007 | Partial distortion entropy maximization for online data clustering
Hiroyuki Takizawa, Hiroaki Kobayashi |
Neural Networks | 2 |
| 2006 | Implications of Memory Performance for Highly Efficient Supercomputing of Scientific Applications
Akihiro Musa, Hiroyuki Takizawa, Koki Okabe, Takashi Soga, Hiroaki Kobayashi |
ISPA | 5 |
| 2006 | Sensing Precision of an Optical Three-axis Tactile Sensor for a Robotic FingerabstractWe are developing an optical three-axis tactile sensor capable of acquiring normal and shearing force, with the aim of mounting it on a robotic finger. The tactile sensor is based on the principle of an optical waveguide-type tactile sensor, which is composed of an acrylic hemispherical dome, a light source, an array of rubber sensing elements, and a CCD camera. The sensing element of silicone rubber comprises one columnar feeler and eight conical feelers. The contact areas of the conical feelers, which maintain contact with the acrylic dome, detect the three-axis force applied to the tip of the sensing element. Normal and shearing forces are then calculated from integration and centroid displacement of the gray-scale value derived from the conical feeler's contacts. To evaluate the present tactile sensor, we have conducted a series of experiments using a y-z stage, a rotational stage, and a force gauge, and have found that although the relationship between the integrated gray-scale value and normal force depends on the sensor's latitude on the hemispherical surface, it is easy to modify the sensitivity according to the latitude, and that the centroid displacement of the gray-scale value is proportional to the shearing force. When we examined repeatability of the present tactile sensor with 1,000 load-unload cycles, the respective error of the normal and shearing forces was 2 and 5% Masahiro Ohka, Hiroaki Kobayashi, Jumpei Takata, Yasunaga Mitsuya |
RO-MAN | 2 |
| 2006 | Hierarchical parallel processing of large scale data clustering on a PC cluster with GPU co-processing
Hiroyuki Takizawa, Hiroaki Kobayashi |
J. Supercomput. | 2 |
| 2005 | Text Detection in Color Scene Images based on Unsupervised Clustering of Multi-channel Wavelet FeaturesabstractTexts in natural scenes provide us with much useful information. In order to use such information automatically, it is necessary to make computers detect text regions in the images. Gllavata et. al. proposed a method based on unsupervised classification of high frequency wavelet coefficients for text detection in video frames [Gllavata et. al. (2004)]. Although the method is very accurate, it does not work so well with some color images, since it lacks the ability of discriminating color difference. This paper proposes an enhanced version of the method. We develop a new unsupervised clustering technique for the classification of multi-channel wavelet features to deal with color images. Experimental results show that the new method yields better results for color scene images. Tomoyuki Saoi, Hideaki Goto, Hiroaki Kobayashi |
ICDAR | 3 |
| 2005 | Sensing characteristics of an optical three-axis tactile sensor mounted on a multi-fingered robotic handabstractTo develop a new three-axis tactile sensor for mounting on multi-fingered robotic hands, in this work we optimize sensing elements on the basis of our previous works concerning optical three-axis tactile sensors with a flat sensing surface. The present tactile sensor is based on the principle of an optical waveguide-type tactile sensor, which is composed of an acrylic hemispherical dome, a light source, an array of rubber sensing elements, and a CCD camera. The sensing element of the present tactile sensor comprises one columnar feeler and eight conical feelers. The contact areas of the conical feelers, which maintain contact with the acrylic dome, detect the three-axis force applied to the tip of the sensing element. Normal and shearing forces are then calculated from integration and centroid displacement of the gray-scale value derived from the conical feeler's contacts. To evaluate the present tactile sensor, we have conducted a series of experiments using a y-z stage, a rotational stage and a force gauge, and have found that although the relationship between integrated gray-scale value and normal force depends on the latitude on the hemispherical surface, it is easy to modify the sensitivity according to the latitude, and that the centroid displacement of the gray-scale value is proportional to the shearing force. Finally, to verify the present tactile sensor, we performed a series of scanning tests using a robotic manipulator equipped with the present tactile sensor to have the manipulator scan surfaces of fine abrasive papers. Results show that the obtained shearing force increased with an increase in the particle diameter of aluminium dioxide contained in the abrasive paper, and decreased with an increase in the scanning velocity of the manipulator over the abrasive paper. Because these results are consistent with tribology, we conclude that the present tactile sensor has sufficient dynamic sensing capability to detect normal and shearing forces. Masahiro Ohka, Hiroaki Kobayashi, Yasunaga Mitsuya |
IROS | 2 |
| 2005 | A Workflow Management Mechanism for Peer-to-Peer Computing Platforms
Hong Wang 0006, Hiroyuki Takizawa, Hiroaki Kobayashi |
ISPA | 3 |
| 2004 | Multi-grain Parallel Processing of Data-Clustering on Programmable Graphics Hardware
Hiroyuki Takizawa, Hiroaki Kobayashi |
ISPA | 2 |
| 2004 | Efficient parallel processing of competitive learning algorithms
Kentaro Sano, Shintaro Momose, Hiroyuki Takizawa, Hiroaki Kobayashi, Tadao Nakamura |
Parallel Comput. | 4 |
| 2003 | A new impedance control concept for elastic joint robots -a case of a 1 DOF robot with programmable linear passive impedanceabstractThe purpose of this paper is to propose a new impedance control concept for elastic joint robots with programmable passive impedance devices in the transmission. The concept allows us to use the same index both for free motion and for contact task. We apply it to a one-DOF elastic joint robot and derive an adjustment law for the robot. The numerical simulations show that implementation of the concept can be realized without any dynamic models. Ryuta Ozawa, Hiroaki Kobayashi |
ICRA | 2 |
| 2001 | 3DCGiRAM: An Intelligent Memory Architecture for Photo-Realistic Image SynthesisabstractThis paper proposes an intelligent memory architecture for photo-realistic image synthesis, named 3DCGiRAM. The 3DCGiRAM has a hardware-accelerated 3D line generator which finds objects that are likely to intersect traced rays. It also has functional memory cells, each of which is composed of graphics logic and its local memory to detect intersecting objects and to calculate intensities. A distributed frame buffer is employed to alleviate the access conflicts of functional cells to the frame buffer as well as to compose globally illuminated intensities at screen pixels. As the graphics processing capability is localized to data in the 3DCGiRAM through memory-logic merged LSI technology, a scalability and modularity similar to those of conventional memory modules can be expected. The experimental results show that a single 3DCGiRAM module running at 200 MHz with a memory bandwidth of 6.4 GB will be able to synthesize a ray-traced walk-through animation at a rate of one frame per second. Hiroaki Kobayashi, Ken-Ichi Suzuki, Kentaro Sano, Yoshiyuki Kaeriyama, Yasumasa Saida, Nobuyuki Ohba, Tadao Nakamura |
ICCD | 1 |
| 2000 | Reconfigurable synchronized dataflow processorabstractNo abstract available. Hitoshi Maruyama, Hideaki Tsukioka, Nobuyoshi Shoji, Hiroaki Kobayashi, Tadao Nakamura |
ASP-DAC | 5 |
| 1999 | A self-organizing network system forming memory from nonstationary probability distributionsabstractWe propose an artificial neural system that forms memory by receiving input vectors obeying an unknown nonstationary probability density function (PDF). The system consists of a set of neural vector quantizer (NVQs), each of which can approximate nonstationary PDFs. Each NVQ exclusively learns a stationary piece of the nonstationary PDF and stores its approximated representation, where the nonstationary PDF consists of some stationary pieces. Experimental results show that the system has functions "memorization", "retention", and "recall" of information which is required in memory systems. The results also illustrate that the system receives inputs from a nonstationary PDF and stores statistical information by distributing it equally over the system. The system can also be used to model nonstationary phenomena. This ability is desirable for various applications, for example, process control, economical modeling, etc. Taira Nakajima, Hiroyuki Takizawa, Hiroaki Kobayashi, Tadao Nakamura |
IJCNN | 3 |
| 1998 | Automated Design of Wave Pipelined Multiport Register FilesabstractRecent high-performance microprocessors have two or more functional units (FUs) to exploit instruction-level parallelism. To make full use of this capability, multiport register files are generally used. However, conventional multiport register files need a considerable amount of hardware. This paper proposes a multiport register file scheme, which uses time-division multiplexing with wave pipelining in order to save the needed hardware resources. For adjusting propagation delay timings, we develop a tool which automatically inserts dummy buffers into combinatorial logic. Kouji Takano, Takehito Sasaki, Nobuyuki Ohba, Hiroaki Kobayashi, Tadao Nakamura |
ASP-DAC | 4 |
| 1997 | A Cached Frame Buffer System for Object-Space parallel Processing SystemabstractThe object space parallel processing for global illumination models is one of the most promising approaches to fast photorealistic image synthesis. However, there is a potential bottleneck between processing elements and a frame buffer in massively parallel processing systems based on the object space parallel processing, and this factor may restrict their scalable performance. To solve this problem, the paper presents a novel frame buffer system, named a cached frame buffer system. By adopting the cached frame buffer system into the object space parallel processing systems, the overhead of the frame buffer access due to conflicts and long latency can be reduced, and the potential of the object space parallel processing system with a large number of processing elements will be fully exploited. Hiroaki Kobayashi |
Computer Graphics International | 1 |
| 1993 | An Adaptive Network Routing Method by Electrical-Circuit ModelingabstractA routing control method called potential routing is proposed for packet communication in computer networks. Potential routing models a computer network as an electrical circuit, and performs packet routing according to the potential differences between adjacent nodes. The node potentials are first given by Kirchhoff's law and are then dynamically adjusted according to the traffic situation. Potential routing can be applied to arbitrary network topologies; it takes account of the global network topology in determining the route. The routing table is easily and therefore quickly computed by Kirchhoff's law, by solving simple simultaneous equations; no convergence problem arises. Moreover, potential routing does not involve the ping-pong (loop) problem. It is verified by simulation that potential routing shortens transmission delays, especially when the traffic is heavy or unbalanced.> Nobuyuki Ohba, Hiroaki Kobayashi, Tadao Nakamura |
INFOCOM | 2 |
| 1988 | Load balancing strategies for a parallel ray-tracing system based on constant subdivision
Hiroaki Kobayashi, Satoshi Nishimura, Hideyuki Kubota, Tadao Nakamura, Yoshiharu Shigei |
Vis. Comput. | 1 |
| 1987 | Parallel processing of an object space for image synthesis using ray tracing
Hiroaki Kobayashi, Tadao Nakamura, Yoshiharu Shigei |
Vis. Comput. | 1 |
| 1986 | Grasping and manipulation of objects by articulated handsabstractRelations between active joints and displacements of grasped objects and relations between joint torques/forces and external and internal forces are obtained in the consistent form for articulated hands. The reconstructability of the pose of objects grasped by hand also discussed. This is useful for the case where hands grasp objects put on a table in accurately or the case where finger tips have slipped on the surface of grasped objects. Hiroaki Kobayashi |
ICRA | 1 |
| 1984 | A Language Processor of an Intelligent Link System
Tadao Nakamura, Hiroaki Kobayashi, Jun Miyajima, Noboru Endo, Yoshiharu Shigei |
ICC (2) | 2 |