EDBT 2026 Demo / reviewers in the wild / expert
Hiroshi Yamaguchi
dblp:74/4594
· DBLP profile ↗
15ranked-venue papers
5as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 1Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Reconstructing, Understanding, and Analyzing Relief Type Cultural Heritage from a Single Old Photo
Jiao Pan, Liang Li 0002, Hiroshi Yamaguchi, Kyoko Hasegawa, Fadjar I. Thufail, Brahmantara |
ACM Multimedia | 3 |
| 2023 | Effective switchless inter-FPGA memory networks
Truong Thao Nguyen, Kien Trung Pham, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi |
J. Parallel Distributed Comput. | 3 |
| 2022 | A Scalable Distributed Radix Sorter for FPGA Clusters using High-Bandwidth Memory NetworksabstractA modern FPGA card can be equipped with high bandwidth memory, such as HBM2. Since the amount of the memory is limited on an FPGA, highly parallel data processing becomes crucial on tightly coupled FPGAs by high-density optical integration, e.g., onboard Si-Photonics transceivers. This study presents a scalable distributed radix sorter, and implements it on an eight-FPGA cluster. Each custom Stratix10 MX2100 FPGA card has 819-Gbps memory bandwidth with two HBM2 memories and 800-Gbps network bandwidth with eight custom embedded optical modules. Existing FPGA sorter typically relies on a merge sort. However, it has a severe performance bottleneck at the final stage of data merge, which cannot make the best use of the high memory-to-memory bandwidth on the FPGA cluster. Instead, we implement a radix sort for a 32-bit key range consisting of eight 4-bit counting sorts optimized to the memory-network structure. Each counting sort needs memory read/write access only once through global and local pipelines. We demonstrated a sorting throughput of 37.2 GB/s. Yutaka Urino, Takanori Shimizu, Hiroshi Yamaguchi, Kenji Mizutani, Shigeru Nakamura, Tatsuya Usuki, Michihiro Koibuchi |
FCCM | 3 |
| 2022 | Scalable Low-Latency Inter-FPGA NetworksabstractA cutting-edge FPGA card can be equipped with many high-bandwidth I/Os by means of high-density optical integration, e.g., onboard Si-photonics transceivers, to provide high network bandwidth for memory-to-memory inter-FPGA communication. This study presents its scalable switchless net-work architecture by exploiting an indirect path, consisting of two one-hop paths, for enabling a diameter-2 network topology. It then takes a Kautz network topology with a diameter of two for connecting d(d + 1) FPGAs with a degree of$d$, which is close to the theoretical upper bound. The Kautz network topologies have bi-directional links and uni-directional links which form triangles. Uni-directional links introduce difficulty in avoiding channel buffer overflow because the existing link-level flow control assumes a bi-directional link. This study presents an indirect flow control along a uni-directional triangle embedded in the Kautz network topology. It then develops a combination of unicasts that forms multi-port collective communications to mitigate the influence of the startup latency on the execution time. Since a high-degree FPGA card introduces difficulty in storing many I/O ports at the panel of a 1- U compute server, we propose using WDM (Wavelength Division Multiplexing) as an alternative and present its efficient mapping onto arrayed waveguide grating (AWG). The required number of wavelengths becomes d on d+ 1 AWG equipments. Based on our experimental results with OPTWEB of custom Stratix10 FPGA cards, SimGrid simulation results show that our collective communication is 7 × faster than that of Dragonfly with 272 FPGAs. Kien Trung Pham, Truong Thao Nguyen, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi |
IPDPS | 3 |
| 2022 | Aerial Clutter Suppression in a Wind Profiler Radar With Antenna SubarraysabstractUndesired echo from a flying object (aerial clutter) significantly contaminates the received signal of a wind profiler radar (WPR) because it has high intensity and spreads over a wide Doppler velocity range. In this study, results of aerial clutter mitigation obtained by applying adaptive clutter suppression (ACS) to a 1.3-GHz WPR are shown. The 1.3-GHz WPR used in this study has a main antenna comprising 13 antenna subarrays (MSAs). five-element Yagi-Uda antennas were also used as antenna subarrays for detecting clutters from low elevation angles (CSAs). The CSAs were used only in reception and installed so that they covered most of the horizontal directions and the horizontal and vertical polarizations. The directionally constrained minimization of power (DCMP) method was used as the adaptive signal processing to mitigate clutter. By the DCMP method, the weighted sum of the signals collected by 13 MSAs and 11 CSAs was computed so that the power of output signals was minimized under the constraint of constant gain in the antenna beam direction. Results of a case study for an aerial clutter from a low elevation angle at 17:04:37 on October 1 2020 showed that an overlap of the aerial clutter over a desired echo (i.e., clear-air echo) was solved by decreasing the aerial clutter whose peak intensity was ~24 dB greater than that of the clear-air echo. In a case study at 09:30:27 on September 18 2020, effects of the DCMP method on the processed results were discussed. Masayuki K. Yamamoto, Seiji Kawamura, Katsuyuki Imai, Hiroshi Yamaguchi, Koji Saito, Koji Nishimura |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | OPTWEB: A Lightweight Fully Connected Inter-FPGA Network for Efficient CollectivesabstractModern FPGA accelerators can be equipped with many high-bandwidth network I/Os, e.g., 64 x 50 Gbps, enabled by onboard optics or co-packaged optics. Some dozens of tightly coupled FPGA accelerators form an emerging computing platform for distributed data processing. However, a conventional indirect packet network using Ethernet's Intellectual Properties imposes an unacceptably large amount of the logic for handling such high-bandwidth interconnects on an FPGA. Besides the indirect network, another approach builds a direct packet network. Existing direct inter-FPGA networks have a low-radix network topology, e.g., 2-D torus. However, the low-radix network has the disadvantage of a large diameter and large average shortest path length that increases the latency of collectives. To mitigate both problems, we propose a lightweight, fully connected inter-FPGA network called OPTWEB for efficient collectives. Since all end-to-end separate communication paths are statically established using onboard optics, raw block data can be transferred with simple link-level synchronization. Once each source FPGA assigns a communication stream to a path by its internal switch logic between memory-mapped and stream interfaces for remote direct memory access (RDMA), a one-hop transfer is provided. Since each FPGA performs input/output of the remote memory access between all FPGAs simultaneously, multiple RDMAs efficiently form collectives. The OPTWEB network provides 0.71-μsec start-up latency of collectives among multiple Intel Stratix 10 MX FPGA cards with onboard optics. The OPTWEB network consumes 31.4 and 57.7 percent of adaptive logic modules for aggregate 400-Gbps and 800-Gbps interconnects on a custom Stratix 10 MX 2100 FPGA, respectively. The OPTWEB network reduces by 40 percent the cost compared to a conventional packet network. Kenji Mizutani, Hiroshi Yamaguchi, Yutaka Urino, Michihiro Koibuchi |
IEEE Trans. Computers | 2 |
| 2015 | Privacy Preserving Data ProcessingabstractA data processing functions are expected as a key-issue of knowledge-intensive service functions in the Cloud computing environment. Cloud computing is a technology that evolved from technologies of the field of virtual machine and distributed computing. However, these unique technologies brings unique privacy and security problems concerns for customers and service providers due to involvement of expertise (such as knowledge, experience, idea, etc.) in data to be processed. We propose the cryptographic protocols preserving the privacy of users and confidentiality of the problem solving servers. Hiroshi Yamaguchi, Masahito Gotaishi, Phillip C.-Y. Sheu, Shigeo Tsujii |
AINA | 1 |
| 2015 | Inverse macro in ScalaabstractWe propose a new variant of typed syntactic macro systems named inverse macro, which improves the expressiveness of macro systems. The inverse macro system enables to implement operators with complex side-effects, such as lazy operators and delimited continuation operators, which are beyond the power of existing macro systems. We have implemented the inverse macro system as an extension to Scala 2.11. We also show the expressiveness of the inverse macro system by comparing two versions of shift/reset, bundled in Scala 2.11 and implemented with the inverse macro system. Hiroshi Yamaguchi, Shigeru Chiba |
GPCE | 1 |
| 2012 | Text Segmentation by Language Using Minimum Description Length
Hiroshi Yamaguchi, Kumiko Tanaka-Ishii |
ACL (1) | 1 |
| 2012 | Scheme overcoming incompatibility of privacy and utilization of personal data
Shigeo Tsujii, Kohtaro Tadaki, Ryou Fujita, Hiroshi Yamaguchi, Masahito Gotaishi, Yukiyasu Tsunoo, Takahiko Syouji, Norihisa Doi |
ISITA | 4 |
| 2011 | Parallel Association Rule Mining for Medical ApplicationsabstractFor real-time applications that consist of massive number of rules, partitioning of the rules to support parallel processing is important. This paper proposes a suite of algorithms called GAPCM for parallel processing of massive number of rules. By considering even distribution, minimal waiting time and minimal inter-processor communication, we propose three algorithms for subnet allocation, and apply these algorithms to association rule mining. G. G. Zhang, C. Z. Xu, Phillip C.-Y. Sheu, Hiroshi Yamaguchi |
BIBE | 4 |
| 2007 | Evaluating HDR rendering algorithmsabstractA series of three experiments has been performed to test both the preference and accuracy of high dynamic-range (HDR) rendering algorithms in digital photography application. The goal was to develop a methodology for testing a wide variety of previously published tone-mapping algorithms for overall preference and rendering accuracy. A number of algorithms were chosen and evaluated first in a paired-comparison experiment for overall image preference. A rating-scale experiment was then designed for further investigation of individual image attributes that make up overall image preference. This was designed to identify the correlations between image attributes and the overall preference results obtained from the first experiments. In a third experiment, three real-world scenes with a diversity of dynamic range and spatial configuration were designed and captured to evaluate seven HDR rendering algorithms for both of their preference and accuracy performance by comparing the appearance of the physical scenes and the corresponding tone-mapped images directly. In this series of experiments, a modified Durand and Dorsey's bilateral filter technique consistently performed well for both preference and accuracy, suggesting that it is a good candidate for a common algorithm that could be included in future HDR algorithm testing evaluations. The results of these experiments provide insight for understanding of perceptual HDR image rendering and should aid in design strategies for spatial processing and tone mapping. The results indicate ways to improve and design more robust rendering algorithms for general HDR scenes in the future. Moreover, the purpose of this research was not simply to find out the “best” algorithms, but rather to find a more general psychophysical experiment based methodology to evaluate HDR image-rendering algorithms. This paper provides an overview of the many issues involved in an experimental framework that can be used for these evaluations. Jiangtao Kuang, Hiroshi Yamaguchi, Changmeng Liu, Garrett M. Johnson, Mark D. Fairchild |
ACM Trans. Appl. Percept. | 2 |
| 2006 | A Social Transformation-Emergence of the Knowledge SocietyabstractThe evolution and dominance of service functions in addition to their distinguishing features has been the subject of study for years. Service functions aim to satisfy and facilitate the needs of their customers. In doing so, service providers design the service functions so that the customers feel comfortable and convenient when using them. The evolution of technology and automation has enabled functions like knowledge-intensive man-machine interactions to be flexible and user friendly. In this paper, we discuss a wide range of interconnected topics, emphasizing the multifaceted nature of service functions. These areas include the evolution of service functions and their associated systems and products, as wall as the consequent creation of implicit requirements, the technology transfer process, and the error proneness due to intense and prolonged interaction. We argue that by proper 'humanization and personalization' the service-based interactive systems can be made easy for the customer to use and enjoy personalized services. We introduce an anonymous opinion survey mechanism for encouraging new service industry. We conclude that the existence of a strong interaction between knowledge and technology growths is shown by 'The Kozmetsky Effect' model and consider the methodology to accelerate the convergence of knowledge-technology transfer phases at the Kozmetsky Effect model Hiroshi Yamaguchi, Darius Mahdjoubi, C. V. Ramamoorthy |
ICTAI | 1 |
| 2004 | Accelerating Trans-disciplinary Research to a Knowledge-sharing InfrastractureabstractThe evolution and dominance of the service based functions and their distinguished features are studied. As the service industry matures, intense machine interaction, knowledge intensive services has become an essential requirement issues for the customers. A wide range of interconnected topics, emphasizing the multi-faced nature of service functions is discussed. By proper 'humanization and personalization' of interactive system and by the use of teams of computer supported professionals, it is possible to provide the knowledge service universally accessible for any customers. Hiroshi Yamaguchi |
BIBE | 1 |
| 2004 | Turbo decoding in impulsive noise environmentabstractPower line channels often suffer from impulsive interference generated by electrical appliances. Therefore, power line communication (PLC) degrades due to such impulsive interference. Middleton's class A noise model is frequently utilized for the modeling of such impulsive noise environments. We deal with turbo decoding for turbo codes over an additive white class A noise (AWAN) channel. We propose a turbo decoding which is suitable for AWAN channels. In addition, we show the BER (bit error rate) performance of the proposed turbo decoding in a class A noise environment by computer simulation. Daisuke Umehara, Hiroshi Yamaguchi, Yoshiteru Morihiro |
GLOBECOM | 2 |