EDBT 2026 Demo / reviewers in the wild / expert
Jaehoon Yu
dblp:140/0249
· DBLP profile ↗
15ranked-venue papers
2as first author
7since 2021 · last 2023
0000-0001-6639-7694ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Decision Forest Training Accelerator Based on Binary Feature DecompositionabstractIn recent years, while Deep Neural Networks (DNNs) have revolutionized various fields, it is widely acknowledged that they are not always the optimal solution, and complementary Machine Learning (ML) tools are necessary. For instance, developing DNN models that can effectively handle tabular data with rows and columns remains a challenging open question. Additionally, the difficulty of interpreting DNN models poses a significant obstacle that hinders their use in many practical applications where the interpretability of the inference results and the ability to offer advice on how to modify input for desired output are required. In such cases, Decision Forests (DFs) have been widely considered a promising solution. Thiem Van Chu, Yu Mizutani, Yuta Nagahara, Shungo Kumazawa, Kazushi Kawamura, Jaehoon Yu, Masato Motomura |
FCCM | 6 |
| 2023 | Samsung PIM/PNM for Transfmer Based AI : Energy Efficiency on PIM/PNM Cluster
Jin Hyun Kim, Yuhwan Ro, Jinin So, Sukhan Lee 0002, Shinhaeng Kang, Yeongon Cho, Byeongho Kim, Kyungsoo Kim 0003, Sangsoo Park, Jin-Seong Kim, Sanghoon Cha, Won-Jo Lee, Jin Jung, Jonggeon Lee, Joon-Ho Song, Seungwon Lee 0006, Jeonghyeon Cho, Jaehoon Yu, Kyomin Sohn |
HCS | 20 |
| 2022 | Multicoated Supermasks Enhance Hidden NetworksabstractHidden Networks (Ramanujan et al., 2020) showed the possibility of finding accurate subnetworks within a randomly weighted neural network by training a connectivity mask, referred to as supermask. We show that the supermask stops improving even though gradients are not zero, thus underutilizing backpropagated information. To address this we propose a method that extends Hidden Networks by training an overlay of multiple hierarchical supermasks{—}a multicoated supermask. This method shows that using multiple supermasks for a single task achieves higher accuracy without additional training cost. Experiments on CIFAR-10 and ImageNet show that Multicoated Supermasks enhance the tradeoff between accuracy and model size. A ResNet-101 using a 7-coated supermask outperforms its Hidden Networks counterpart by 4%, matching the accuracy of a dense ResNet-50 while being an order of magnitude smaller. Yasuyuki Okoshi, Ángel López García-Arias, Kazutoshi Hirose, Kota Ando, Kazushi Kawamura, Thiem Van Chu, Masato Motomura, Jaehoon Yu |
ICML | 8 |
| 2021 | Hidden-Fold Networks: Random Recurrent Residuals Using Sparse Supermasks
Ángel López García-Arias, Masanori Hashimoto, Masato Motomura, Jaehoon Yu |
BMVC | 4 |
| 2021 | MUX Granularity Oriented Iterative Technology Mapping for Implementing Compute-Intensive Applications on Via-Switch FPGAabstractThis paper proposes a technology mapping algorithm for implementing application circuits on via-switch FPGA (VS-FPGA). The via-switch is a novel non-volatile and rewritable memory element. Its small footprint and low parasitic RC are expected to improve the area- and energy-efficiency of an FPGA system. Some unique features of the VS-FPGA require a dedicated technology mapping strategy for implementing application circuits with maximum energy-efficiency. One of the features is the small ratio of logic blocks to arithmetic blocks (ABs). Given an application circuit, the proposed algorithm first detects word-wise circuit elements, such as MUXs. These elements are evaluated with an index of how resource utilization and fan-out change when the corresponding element is implemented with AB. All these elements are sorted in descending order based on this index. According to this order, each element is mapped to AB one by one, and synthesis and evaluation are repeated iteratively until satisfying given design constraints. The experimental results show that resource utilization and maximum fan-out can be reduced by about 30 % to 50 % and 12 % to 87 %, respectively. The proposed algorithm is not limited to the VS-FPGA and is expected to improve computation density and energy-efficiency of various FPGAs dedicated to compute-intensive signal processing applications. Takashi Imagawa, Jaehoon Yu, Masanori Hashimoto, Hiroyuki Ochi |
DATE | 2 |
| 2021 | A High-Performance and Flexible FPGA Inference Accelerator for Decision Forests Based on Prior Feature Space PartitioningabstractRecent studies have demonstrated the potential of FPGAs for accelerating the inference computation of decision forests (DFs). However, designing a high-performance architecture that is flexible enough to be adopted in various scenarios of FPGA resource requirements remains a challenge. To address this, we propose a DF inference method that makes a transformation from traversing trees into traversing feature spaces. Specifically, as a preprocessing step, we partition each feature space into multiple regions based on thresholds. The inference task for an input data point is then conducted by (1) determining which region in each feature space the data point belongs to and (2) combining the inference information in these regions. The regularity of the computation allows us to design a DF inference architecture, called FT-DFP (Feature-space Traversing Decision Forest Processor), that can be flexibly configured for different performance and FPGA resource usage requirements. We prototype FT-DFP on a low-end FPGA (Artix-7) board and evaluate it using four real-world datasets. The evaluation results show that (1) the flexibility of FT-DFP allows us to fit a wide variety of DF models into low-end FPGA devices with limited resources; (2) FT-DFP's performance is comparable to the best of existing accelerators implemented on high-end FPGA devices and 3.04 × higher than Hummingbird, a state-of-the-art GPU-optimized implementation, running on a high-end GPU; and (3) FT-DFP is 130.96 × more energy-efficient than Hummingbird. Thiem Van Chu, Ryuichi Kitajima, Kazushi Kawamura, Jaehoon Yu, Masato Motomura |
FPT | 4 |
| 2021 | Edge Inference Engine for Deep & Random Sparse Neural Networks with 4-bit Cartesian-Product MAC Array and Pipelined Activation AlignerabstractA 4b-quantized convolutional neural network (CNN) inference engine for edge-AI is presented featuring a Cartesian-product MAC array and pipelined activation aligners targeting deep-/random-pruned models. A 40nm prototype with 32x32 MACs and 5Mb SRAM runs at 534 MHz, 1.07 TOPS, 352 mW at 1.1V, and attains 5.30 dense TOPS/W, 234 MHz at 0.8V. Sparse TOPS/W reaches 26.5 when running a randomly pruned model (after 88% pruning). Training algorithms for obtaining highly efficient sparse/quantized models are also proposed. Kota Ando, Jaehoon Yu, Kazutoshi Hirose, Hiroki Nakahara, Kazushi Kawamura, Thiem Van Chu, Masato Motomura |
HCS | 2 |
| 2020 | Low-Cost Reservoir Computing using Cellular Automata and Random ForestsabstractHigh-performance image classification models involve massive computation and an energy cost that are unaffordable for resource-limited platforms. As a solution, reservoir computing based on cellular automata has been proposed, but there is still room for improvement in terms of classification cost. This research builds on the previous work introducing enhancements at both the algorithmic and architectural level. Using a random forest classifier with binary features completely eliminates multiplication operations and 97% of addition operations. Also, memory usage can be decreased by pruning 82% of the least relevant augmented features. An architecture with an increased level of parallelism which processes images in a single pass reduces memory accesses, and reduces 60% of logic by optimizing FPGA mapping. These speed, power, and memory optimizations come at an accuracy tradeoff of a mere 0.6%. Ángel López García-Arias, Jaehoon Yu, Masanori Hashimoto |
ISCAS | 2 |
| 2020 | Logarithm-approximate floating-point multiplier is applicable to power-efficient neural network training
TaiYu Cheng, Yukata Masuda, Jaehoon Yu, Masanori Hashimoto |
Integr. | 4 |
| 2020 | Sneak Path Free Reconfiguration With Minimized Programming Steps for Via-Switch Crossbar-Based FPGAabstractField programmable gate array (FPGA) that utilizes via-switches, which are a kind of nonvolatile resistive RAMs, for crossbar implementation is attracting attention due to higher integration density and performance. However, programming via-switches arbitrarily in a crossbar is not trivial since a programming current must be provided through signal wires shared by multiple via-switches. Consequently, depending on the previous programming status in sequential programming, unintentional switch programming may occur due to signal detour, which is called the sneak path problem. This article identifies the circuit status that causes the sneak path problem and proposes a sneak path avoidance method that gives sneak path free programming order of via-switches in a crossbar. We prove that sneak path free programming order necessarily exists for arbitrary ON-OFF patterns in a crossbar as long as no loops exist. This article also proposes a partial reconfiguration method that achieves the minimum number of switch programming steps while avoiding the sneak path problem. This method contributes to the extension of via-switch lifetime and fast reconfiguration of the via-switch FPGA. Experimental results show that the proposed partial reconfiguration method reduces the number of programmed switches by 77.4% compared to the conventional approach. This 77.4% reduction improves the number of reconfigurations of the via-switch FPGA by 4.4× and reduces reconfiguration time by 77.4%. Ryutaro Doi, Jaehoon Yu, Masanori Hashimoto |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Sneak path free reconfiguration of via-switch crossbars based FPGAabstractFPGA that utilizes via-switches, which are a kind of nonvolatile resistive RAMs, for crossbar implementation is attracting attention due to higher integration density and performance. However, programming via-switches arbitrarily in a crossbar is not trivial since a programming current must be provided through signal wires that are shared by multiple via-switches. Consequently, depending on the previous programming status in sequential programming, unintentional switch programming may occur due to signal detour, which is called sneak path problem. This problem interferes the reconfiguration of via-switch FPGA, and hence countermeasures for sneak path problem are indispensable. This paper identifies the circuit status that causes sneak path problem and proposes a sneak path avoidance method that gives sneak path free programming order of via-switches in a crossbar. We prove that sneak path free programming order necessarily exists for arbitrary on-off patterns in a crossbar as long as no loops exist, and also validate the proof and the proposed method with simulation-based evaluation. Thanks to the proposed method, any practical configurations of via-switch FPGA can be successfully programmed without sneak path problem. Ryutaro Doi, Jaehoon Yu, Masanori Hashimoto |
ICCAD | 2 |
| 2018 | Via-Switch FPGA: Highly Dense Mixed-Grained Reconfigurable Architecture With Overlay Via-Switch CrossbarsabstractThis paper proposes a highly dense reconfigurable architecture that introduces via-switch device, which is a nonvolatile resistive-change switch and is used in crossbar switches. Via-switch is implemented in back-end-of-line layers only, and hence the front-end-of-line (FEoL) layers under the crossbar can be fully exploited for highly dense logic blocks. The proposed architecture uses the FEoL layers for fine-grained lookup tables and coarse-grained arithmetic/memory units for improving performance and compatibility with various applications. A case study of application mapping shows the proposed architecture can reduce the array area by 21.7%, thanks to the bidirectional interconnection. Thanks to F2footprint and one order of magnitude lower resistivity of via-switch compared to MOS switch, the crossbar density is improved by up to 26× and the delay and energy in the interconnection are reduced by 90% and 94% at 0.5-V operation. Hiroyuki Ochi, Kosei Yamaguchi, Tetsuaki Fujimoto, Junshi Hotate, Takashi Kishimoto, Toshiki Higashi, Takashi Imagawa, Ryutaro Doi, Munehiro Tada, Tadahiko Sugibayashi, Wataru Takahashi 0002, Kazutoshi Wakabayashi, Hidetoshi Onodera, Yukio Mitsuyama, Jaehoon Yu, Masanori Hashimoto |
IEEE Trans. Very Large Scale Integr. Syst. | 15 |
| 2014 | Corrections to "A Speed-Up Scheme Based on Multiple-Instance Pruning for Pedestrian Detection Using a Support Vector Machine"abstractIn the above paper (ibid., vol. 22, no. 12, pp. 4752-4761, Dec. 2013), several errors were introduced. These errors are corrected here. Jaehoon Yu, Ryusuke Miyamoto, Takao Onoye |
IEEE Trans. Image Process. | 1 |
| 2013 | A Speed-Up Scheme Based on Multiple-Instance Pruning for Pedestrian Detection Using a Support Vector MachineabstractIn pedestrian detection, as sophisticated feature descriptors are used for improving detection accuracy, its processing speed becomes a critical issue. In this paper, we propose a novel speed-up scheme based on multiple-instance pruning (MIP), one of the soft cascade methods, to enhance the processing speed of support vector machine (SVM) classifiers. Our scheme mainly consists of three steps. First, we regularly split an SVM classifier into multiple parts and build a cascade structure using them. Next, we rearrange the cascade structure for enhancing the rejection rate, and then train the rejection threshold of each stage composing the cascade structure using the MIP. To verify the validity of our scheme, we apply it to a pedestrian classifier using co-occurrence histograms of oriented gradients trained by an SVM, and experimental results show that the processing time for classification of the proposed scheme is as low as one-hundredth of the original classifier without sacrificing detection accuracy. Jaehoon Yu, Ryusuke Miyamoto, Takao Onoye |
IEEE Trans. Image Process. | 1 |
| 2007 | Implementation of AV Streaming System Using Peer-to-Peer CommunicationabstractIn this paper, a prototype implementation of stream- ing system that allows listening to and viewing multimedia contents using a mobile terminal, as part of our efforts toward realizing services linked with home networks and mobile net- works is presented. In our system, appliances are detected and connected to a mobile terminal by use of peer-to-peer (P2P) network based on the PUCC protocols, and they are controlled by use of P2P messages and IEEE1394 AV/C. Streaming is managed by a gateway on the P2P network, where multimedia contents are converted to another format suitable to be sent to mobile terminals. This prototype system successfully achieves video data streaming over two different networks by using P2P network which conceals the differences among the networks. Norihiro Ishikawa, Hiroshi Tsutsui, Jaehoon Yu, Tomonori Izumi, Hiroyuki Ochi, Yukihiro Nakamura, Takaaki Komura, Yoshitaka Uchida |
CCNC | 3 |