EDBT 2026 Demo / reviewers in the wild / expert
Thiem Van Chu
dblp:153/0656
· DBLP profile ↗
22ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-2003-2574ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Vision Transformers via Token Merging with Head-Wise Attention CorrectionabstractVision Transformers (ViTs) offer strong performance by modeling global relationships across image patches, but their scalability is limited by the quadratic cost of self-attention. To mitigate this, Token Merging (ToMe) reduces computation by merging similar tokens. This approach relies on proportional attention to preserve the original balance of attention weights after merging. Yet, proportional attention does not fully resolve the attention distortion. It only compensates for a merged token’s influence on other tokens, while ignoring the fundamental distortion of the token’s own self-attention score. This change creates unaddressed distortions that vary across attention heads.In this work, we conduct a detailed analysis of these attention distortions and reveal their dependence on the query–key projection weights of each head. Based on this finding, we propose Head-wise Attention Correction (HAC), a method that adjusts attention scores after token merging by accounting for head-specific characteristics. HAC effectively mitigates the distortions overlooked by proportional attention, maintaining model accuracy while significantly reducing computation. Experiments on ImageNet demonstrate that our method effectively improves the trade-off between efficiency and performance, advancing the development of efficient Vision Transformers via token merging. Code is available at https://github.com/ychikawa/HAC-ToMe Yuki Ichikawa, Masato Motomura, Thiem Van Chu, Daichi Fujiki |
WACV | 3 |
| 2025 | BingoGCN: Towards Scalable and Efficient GNN Acceleration with Fine-Grained Partitioning and SLTabstractGraph Neural Networks (GNNs) are increasingly popular due to their wide applicability to tasks requiring the understanding of unstructured graph data, such as those in social network analysis and autonomous driving.However, real-time, large-scale GNN inference faces challenges due to the large size of node features and adjacency matrices, leading to memory communication and buffer size overheads caused by irregular memory access patterns.While graph partitioning can help with localized access patterns and reduction in on-chip buffer size, fine-grained partitioning results in increased inter-partition edges and off-chip memory accesses, negatively impacting overall performance.To overcome these limitations, we propose BingoGCN, a scalable GNN acceleration framework that introduces multidimensional dynamic feature summarization called Cross-Partition Message Quantization (CMQ) for inter-partition message passing.This eliminates irregular off-chip memory access without additional training and accuracy loss, even with fine-grained partitioning.By shifting the bottleneck from memory to computation, BingoGCN allows for further performance optimization through the Strong Lottery Ticket (SLT) theory using randomly generated weights.BingoGCN addresses the challenge of SLT's unstructured sparsity in hardware acceleration with a novel training algorithm and random weight generator designs, enabling fine-grained (FG) sparsity and improved load balancing.We integrated CMQ and FG-SLT into the messagepassing of GNNs and designed an efficient hardware architecture to support this flow.Our FPGA-based implementation achieves a significant reduction in memory accesses while preserving accuracy comparable to the original models. Jiale Yan, Hiroaki Ito, Yuta Nagahara, Kazushi Kawamura, Masato Motomura, Thiem Van Chu, Daichi Fujiki |
ISCA | 6 |
| 2025 | Binary Quadratic Quantization: Beyond First-Order Quantization for Real-Valued Matrix CompressionabstractThis paper proposes a novel matrix quantization method, Binary Quadratic Quan-
tization (BQQ). In contrast to conventional first-order quantization approaches—
such as uniform quantization and binary coding quantization—that approximate
real-valued matrices via linear combinations of binary bases, BQQ leverages the
expressive power of binary quadratic expressions while maintaining an extremely
compact data format. We validate our approach with two experiments: a matrix
compression benchmark and post-training quantization (PTQ) on pretrained Vision
Transformer-based models. Experimental results demonstrate that BQQ consistently achieves a superior trade-off between memory efficiency and reconstruction
error than conventional methods for compressing diverse matrix data. It also
delivers strong PTQ performance, even though we neither target state-of-the-art
PTQ accuracy under tight memory constraints nor rely on PTQ-specific binary
matrix optimization. For example, our proposed method outperforms the state-of-
the-art PTQ method by up to 2.2% and 59.1% on the ImageNet dataset under the
calibration-based and data-free scenarios, respectively, with quantization equivalent
to 2 bits. These findings highlight the surprising effectiveness of binary quadratic
expressions for efficient matrix approximation and neural network compression. Kyo Kuroki, Yasuyuki Okoshi, Thiem Van Chu, Kazushi Kawamura, Masato Motomura |
NeurIPS | 3 |
| 2025 | Retraining Gradient Boosting Decision Trees with Backward Compatibility of Explanations
Tsukasa Yamakura, Yoichi Sasaki 0002, Yuzuru Okajima, Thiem Van Chu |
PAKDD (2) | 4 |
| 2025 | DMSA: An Efficient Architecture for Sparse-Sparse Matrix Multiplication Based on Distribute-Merge Product DataflowabstractThe sparse–sparse matrix multiplication (SpMSpM) is a fundamental operation in various applications. Existing SpMSpM accelerators based on inner product (IP) and outer product (OP) suffer from low computational efficiency and high memory traffic due to inefficient index matching and merging overheads. Gustavson’s product (GP)-based accelerators mitigate some of these challenges but struggle with workload imbalance and irregular memory access patterns, limiting computational parallelism. To overcome these limitations, we propose a distribute-merge product (DMP), a novel SpMSpM dataflow that evenly distributes workloads across multiple computation streams and merges partial results efficiently. We design and implement DMP-based SpMSpM architecture (DMSA), incorporating four key techniques to fully exploit the parallelism of DMP and efficiently handle irregular memory accesses. Implemented on a Xilinx ZCU106 FPGA, DMSA achieves speedups of up to$3.38\times $and$1.73\times $over two state-of-the-art FPGA-based SpMSpM accelerators while maintaining comparable hardware resource usage. In addition, compared to CPU and GPU implementations on an NVIDIA Jetson AGX Xavier, DMSA is$4.96\times $and$1.53\times $faster while achieving$6.67\times $and$2.33\times $better energy efficiency, respectively. Yuta Nagahara, Jiale Yan, Kazushi Kawamura, Daichi Fujiki, Masato Motomura, Thiem Van Chu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | Sparse-Sparse Matrix Multiplication Accelerator on FPGA featuring Distribute-Merge Product DataflowabstractSparse-Sparse matrix multiplication (SpMSpM) is a critical computation in various fields such as computational science and graph analysis. It poses computational challenges for general-purpose CPUs and GPUs due to its requirements for random memory access and the inherently low spatial/temporal locality. Given the increasing importance of SpMSpM, numerous accelerators have been recently proposed. However, they suffer from various issues such as low input utilization, heavy computational load, and excessive memory traffic during the merging process of intermediate results. This paper introduces a novel Distribute-Merge Product (DMP) SpMSpM dataflow and a DMP-based SpMSpM Architecture (DMSA). DMP distributes the workload into balanced streams, generates partial matrices based on these streams, and merges the partial results in a parallel and pipelined fashion. We have designed DMSA as a highly scalable architecture, implemented it on a Xilinx ZCU106 Evaluation Kit, and evaluated it on a set of benchmarks from the SuiteSparse matrix collection. When compared to a latest SpMSpM accelerator with approximately the same amount of hardware resources on the same FPGA platform, DMSA achieves 2.72 × speedup, by facilitating the parallelism of partial matrix generation and merging. The speedup on the same platform reaches 4.80 × when the parallelism explored in the merging process is doubled, evidencing the DMSA’s superb scalability. Yuta Nagahara, Jiale Yan, Kazushi Kawamura, Masato Motomura, Thiem Van Chu |
ASPDAC | 5 |
| 2024 | Classical Thermodynamics-based Parallel Annealing Algorithm for High-speed and Robust Combinatorial OptimizationabstractIn recent years, quantum annealing has triggered active research on annealing methods for solving various combinatorial optimization problems (COPs) by mapping them to the Ising model based on spin glass theory. In particular, parallel annealing algorithms (PAAs) that can update all variables simultaneously attract attention due to fast optimization using parallel computers, either as an extension of Simulated Annealing rooted in classical thermodynamics or as a quantum-inspired algorithm. However, both types of PAAs face their own challenges. The classical thermodynamics-based PAAs (c-PAAs) perform inferior to the quantum-inspired PAAs (q-PAAs), whereas the q-PAAs require more parameters to be tuned than the c-PAAs. This paper proposes a new c-PAA based on Mean Field Annealing, which has the unique feature of updating analog variables deterministically. The proposed PAA achieves high speed and robustness despite fewer parameters than the q-PAAs, which means the proposed PAA breaks through the challenges of conventional PAAs. We demonstrate its performance through experiments on four types of COPs: Maximum Cut Problem, Graph Coloring Problem, Maximum Independent Set Problem, and Traveling Salesman Problem. These results imply that unless a real physical phenomenon is used, quantum-inspired algorithms cannot be considered superior to classical thermodynamics-based algorithms. Kyo Kuroki, Satoru Jimbo, Thiem Van Chu, Masato Motomura, Kazushi Kawamura |
GECCO | 3 |
| 2024 | ETreeNet: Ensemble Model Fusing Decision Trees and Neural Networks for Small Tabular DataabstractIn real-world machine learning applications, addressing the challenges associated with small tabular data is essential. While Decision Tree (DT)-based models are known to be effective for tabular data, their suitability diminishes when confronted with applications involving diverse data modalities beyond tabular data. Then, many studies focusing on tabular data propose Neural Networks (NN)-based models. To cope with the issue of limited data availability, most NN-based models for small tabular data utilize transformer architectures with techniques such as transfer learning, pre-training, and data augmentation. However, training or retraining a transformer-based model requires substantial data. This problem raises the question of whether it is appropriate to employ a transformer-based model for limited tabular data. We try to answer this question by proposing an ensemble model fusing DTs and NNs, called ETreeNet, which can outperform state-of-the-art transformer-based and DT-based models. ETreeNet comprises three methods: (1) Ensembling Tree-structured Neural Networks (TNNs), allowing training on small data due to reduced training parameters; (2) Sampling of features observed in Random Forest (RF) to enhance accuracy by reducing the influence of uninformative features; (3) Ensembling RF and TNNs to improve accuracy further. We conduct experiments using 500 instances of tabular training data, and the results show that ETreeNet achieves up to a 5% enhancement over state-of-the-art transformer-based and DT-based models. Tsukasa Yamakura, Kazushi Kawamura, Masato Motomura, Thiem Van Chu |
IJCNN | 4 |
| 2024 | Efficient Deadlock Avoidance for 2-D Mesh NoCs That Use OQ or VOQ RoutersabstractNetwork-on-chips (NoCs) are currently a widely used approach for achieving scalability of multi-cores to many-cores, as well as for interconnecting other vital system-on-chip (SoC) components. Each entity in 2D mesh-based NoCs has a router responsible for forwarding packets between the dimensions as well as the entity itself, and it is essentially a 5-port switch. With respect to the routing algorithm, there are important trade-offs between routing performance and the efficiency of overcoming potential deadlocks. Common deadlock avoidance techniques including the turn model usually involve restrictions of certain paths a packet can take at the cost of a higher probability for network congestion. In contrast, deadlock resolution techniques, as well as some avoidance schemes, provide more path flexibility at the expense of hardware complexity, such as by incorporating (or assuming) dedicated buffers. This paper provides a deadlock avoidance algorithm for NoC routers based on output-queues (OQs) or virtual-output queues (VOQs), with a focus on their use on field-programmable gate-arrays (FPGAs). The proposed approach features fewer path restrictions than common techniques, and can be based on existing routing algorithms as a baseline, deadlock-free or not. This requires no modification to the queueing topology, and the required logic is minimal. Our algorithm approaches the performance of fully-adaptive algorithms, while maintaining deadlock freedom. Philippos Papaphilippou, Thiem Van Chu |
IEEE Trans. Computers | 2 |
| 2023 | Decision Forest Training Accelerator Based on Binary Feature DecompositionabstractIn recent years, while Deep Neural Networks (DNNs) have revolutionized various fields, it is widely acknowledged that they are not always the optimal solution, and complementary Machine Learning (ML) tools are necessary. For instance, developing DNN models that can effectively handle tabular data with rows and columns remains a challenging open question. Additionally, the difficulty of interpreting DNN models poses a significant obstacle that hinders their use in many practical applications where the interpretability of the inference results and the ability to offer advice on how to modify input for desired output are required. In such cases, Decision Forests (DFs) have been widely considered a promising solution. Thiem Van Chu, Yu Mizutani, Yuta Nagahara, Shungo Kumazawa, Kazushi Kawamura, Jaehoon Yu, Masato Motomura |
FCCM | 1 |
| 2022 | Multicoated Supermasks Enhance Hidden NetworksabstractHidden Networks (Ramanujan et al., 2020) showed the possibility of finding accurate subnetworks within a randomly weighted neural network by training a connectivity mask, referred to as supermask. We show that the supermask stops improving even though gradients are not zero, thus underutilizing backpropagated information. To address this we propose a method that extends Hidden Networks by training an overlay of multiple hierarchical supermasks{—}a multicoated supermask. This method shows that using multiple supermasks for a single task achieves higher accuracy without additional training cost. Experiments on CIFAR-10 and ImageNet show that Multicoated Supermasks enhance the tradeoff between accuracy and model size. A ResNet-101 using a 7-coated supermask outperforms its Hidden Networks counterpart by 4%, matching the accuracy of a dense ResNet-50 while being an order of magnitude smaller. Yasuyuki Okoshi, Ángel López García-Arias, Kazutoshi Hirose, Kota Ando, Kazushi Kawamura, Thiem Van Chu, Masato Motomura, Jaehoon Yu |
ICML | 6 |
| 2021 | A High-Performance and Flexible FPGA Inference Accelerator for Decision Forests Based on Prior Feature Space PartitioningabstractRecent studies have demonstrated the potential of FPGAs for accelerating the inference computation of decision forests (DFs). However, designing a high-performance architecture that is flexible enough to be adopted in various scenarios of FPGA resource requirements remains a challenge. To address this, we propose a DF inference method that makes a transformation from traversing trees into traversing feature spaces. Specifically, as a preprocessing step, we partition each feature space into multiple regions based on thresholds. The inference task for an input data point is then conducted by (1) determining which region in each feature space the data point belongs to and (2) combining the inference information in these regions. The regularity of the computation allows us to design a DF inference architecture, called FT-DFP (Feature-space Traversing Decision Forest Processor), that can be flexibly configured for different performance and FPGA resource usage requirements. We prototype FT-DFP on a low-end FPGA (Artix-7) board and evaluate it using four real-world datasets. The evaluation results show that (1) the flexibility of FT-DFP allows us to fit a wide variety of DF models into low-end FPGA devices with limited resources; (2) FT-DFP's performance is comparable to the best of existing accelerators implemented on high-end FPGA devices and 3.04 × higher than Hummingbird, a state-of-the-art GPU-optimized implementation, running on a high-end GPU; and (3) FT-DFP is 130.96 × more energy-efficient than Hummingbird. Thiem Van Chu, Ryuichi Kitajima, Kazushi Kawamura, Jaehoon Yu, Masato Motomura |
FPT | 1 |
| 2021 | Edge Inference Engine for Deep & Random Sparse Neural Networks with 4-bit Cartesian-Product MAC Array and Pipelined Activation AlignerabstractA 4b-quantized convolutional neural network (CNN) inference engine for edge-AI is presented featuring a Cartesian-product MAC array and pipelined activation aligners targeting deep-/random-pruned models. A 40nm prototype with 32x32 MACs and 5Mb SRAM runs at 534 MHz, 1.07 TOPS, 352 mW at 1.1V, and attains 5.30 dense TOPS/W, 234 MHz at 0.8V. Sparse TOPS/W reaches 26.5 when running a randomly pruned model (after 88% pruning). Training algorithms for obtaining highly efficient sparse/quantized models are also proposed. Kota Ando, Jaehoon Yu, Kazutoshi Hirose, Hiroki Nakahara, Kazushi Kawamura, Thiem Van Chu, Masato Motomura |
HCS | 6 |
| 2020 | Dependency-Driven Trace-Based Network-on-Chip Emulation on FPGAsabstractFPGA emulation is a promising approach to accelerating Network-on-Chip (NoC) modeling which has traditionally relied on software simulators. In most early studies of FPGA-based NoC emulators, only synthetic workloads like uniform and bit permutations were considered. Although a set of carefully designed synthetic workloads can reveal a relatively thorough coverage of the characteristics of the NoC under evaluation, they alone are insufficient, especially when the NoC needs to be optimized for specific applications. In such cases, trace-driven workloads are effective. However, there is a problem with conventional trace-driven workloads that has been pointed out by some recent studies: the network load and congestion may be distorted because dependencies between packets are not considered. These studies also provide infrastructures for extending existing software simulators to enforce dependencies between packets. Unfortunately, enforcing dependencies between packets is not trivial in the FPGA emulation approach. Therefore, although there are some recent FPGA-based NoC emulators supporting trace-driven workloads, most of them ignore packet dependencies. In this paper, we first clarify the challenges of supporting trace-driven workloads with dependencies between packets taken into account in the FPGA emulation approach. We then propose efficient methods and architectures to tackle these challenges and build an FPGA-based NoC emulator, which we call DNoC, based on the proposals. Our evaluation results show that (1) on a VC707 FPGA board, DNoC achieves an average speed of 10,753K cycles/s when emulating an 8x8 NoC with trace data collected from full-system simulation of the PARSEC benchmark suite, which is 274x higher than the speed reported in a recent related work on dependency-driven trace-based NoC emulation on FPGAs; (2) Compared to BookSim, one of the most popular NoC simulators, DNoC is 395x faster while providing the same results; (3) DNoC can scale to a 4,096-node NoC on a VC707 board, and the size of the largest NoC depends on only the on-chip memory capacity of the target FPGA. Thiem Van Chu, Kenji Kise, Kiyofumi Tanaka |
FPGA | 1 |
| 2018 | A High-Performance and Cost-Effective Hardware Merge Sorter without Feedback DatapathabstractWe propose a high-performance and cost-effective hardware merge sorter (HMS) without any feedback datapaths in order to develop the fastest hardware sorting accelerator. The operating frequencies of existing HMSs are severely limited by the presence of feedback datapaths. We show the idea of eliminating the feedback datapaths, and propose a concrete architecture adopting the idea and some implementation optimizations. The evaluation results show that our HMS achieves 1.59x throughput improvement with less hardware resources compared to the state-of-the-art HMS. Makoto Saitoh, Elsayed A. Elsayed, Thiem Van Chu, Susumu Mashimo, Kenji Kise |
FCCM | 3 |
| 2018 | An Effective Architecture for Trace-Driven Emulation of Networks-on-Chip on FPGAsabstractModern many-core systems use Networks-on-Chip (NoCs) to move data around their cores. As the number of cores increases, the overall performance becomes highly sensitive to the NoC performance. Research and development of NoCs thus play a key role in designing future systems with hundreds to thousands of cores. However, current methodologies for evaluating NoCs are not scalable with respect to the system complexity. Conventional software simulators are too slow for evaluating middle-and large-scale NoCs. Recent FPGA-based emulators provide promising emulation speedups over software simulators. However, emulating large-scale NoCs with hundreds to thousands of nodes on FPGAs is a challenging problem because of the FPGA capacity constraints. Moreover, supporting trace-driven emulation is not trivial because trace data must be stored outside of the FPGA (usually in off-chip DRAM). Most of the existing FPGA-based NoC emulators rely on soft processors like Microblaze or hard processors on SoC FPGAs for loading trace data from the off-chip memory, generating messages, injecting the messages to the target NoC, manipulating the emulation, and making sure that there is no timing error. This approach makes the implementation easy but drastically degrades the emulation speed. This paper proposes an effective architecture for trace-driven emulation of NoCs on FPGAs. We present methods to scale to large NoCs and effectively hide the off-chip memory access latency. Our evaluation results show that (1) the proposal achieves a speedup of 260x compared to BookSim, one of the most widely used NoC simulators, when emulating an 8x8 NoC with trace data collected from full-system simulation of the PARSEC benchmark suite; and (2) the speedup is increased to three orders of magnitude when emulating a 64x64 NoC. Thiem Van Chu, Kenji Kise |
FPL | 1 |
| 2017 | High-Performance Hardware Merge SorterabstractState-of-the-art studies show that FPGA-based hardware merge sorters (HMSs) can achieve superior performance compared with optimized algorithms on CPUs and GPUs. The performance of any HMS is proportional to its operating frequency (F) and the number of records that can be output each cycle (E). However, all existing HMSs have a problem that F drops significantly with increasing E due to the increase of the number of levels of gates. In this paper, we propose novel architectures for HMSs where the number of levels of gates is constant when E is increased. We implement some HMSs adopting the proposed architectures on a Virtex-7 FPGA. The evaluation shows that an HMS of E = 32 operates at 311MHz and achieves 3.13x higher throughput than the state-of-the-art HMS. Susumu Mashimo, Thiem Van Chu, Kenji Kise |
FCCM | 2 |
| 2017 | Fast and Cycle-Accurate Emulation of Large-Scale Networks-on-Chip Using a Single FPGAabstractModeling and simulation/emulation play a major role in research and development of novel Networks-on-Chip (NoCs). However, conventional software simulators are so slow that studying NoCs for emerging many-core systems with hundreds to thousands of cores is challenging. State-of-the-art FPGA-based NoC emulators have shown great potential in speeding up the NoC simulation, but they cannot emulate large-scale NoCs due to the FPGA capacity constraints. Moreover, emulating large-scale NoCs under synthetic workloads on FPGAs typically requires a large amount of memory and thus involves the use of off-chip memory, which makes the overall design much more complicated and may substantially degrade the emulation speed. This article presents methods for fast and cycle-accurate emulation of NoCs with up to thousands of nodes using a single FPGA. We first describe how to emulate a NoC under a synthetic workload using only FPGA on-chip memory (BRAMs). We next present a novel use of time-division multiplexing where BRAMs are effectively used for emulating a network using a small number of nodes, thereby overcoming the FPGA capacity constraints. We propose methods for emulating both direct and indirect networks, focusing on the commonly used meshes and fat-trees (k-aryn-trees). This is different from prior work that considers only direct networks. Using the proposed methods, we build a NoC emulator, called FNoC, and demonstrate the emulation of some mesh-based and fat-tree-based NoCs with canonical router architectures. Our evaluation results show that (1) the size of the largest NoC that can be emulated depends on only the FPGA on-chip memory capacity; (2) a mesh-based NoC with 16,384 nodes (128×128 NoC) and a fat-tree-based NoC with 6,144 switch nodes and 4,096 terminal nodes (4-ary 6-tree NoC) can be emulated using a single Virtex-7 FPGA; and (3) when emulating these two NoCs, we achieve, respectively, 5,047× and 232× speedups over BookSim, one of the most widely used software-based NoC simulators, while maintaining the same level of accuracy. Thiem Van Chu, Shimpei Sato, Kenji Kise |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2016 | The synchronous vs. asynchronous NoC routers: an apple-to-apple comparison between synchronous and transition signaling asynchronous designsabstractSeveral studies have been made on comparison between synchronous and asynchronous NoCs (Network-on- Chips). However, very few attempts have been made at fair comparison between synchronous NoCs designed by synchronous researchers and asynchronous ones designed by asynchronous researchers under the same functional specification using the same fabrication technology. In this paper, representative routers of both design styles are individually designed under the same conditions. Then, they are precisely compared in the viewpoint of area, latency, power, and fairness. As a result, we show "areas of specialty" of both routers, e.g., real-time control systems for asynchronous routers and multi-media data processing systems for synchronous ones. Masashi Imai, Thiem Van Chu, Kenji Kise, Tomohiro Yoneda |
NOCS | 2 |
| 2015 | Enabling Fast and Accurate Emulation of Large-Scale Network on Chip Architectures on a Single FPGAabstractNetwork on Chip (NoC) has become the de facto on-chip communication architecture of many-core systems. This paper proposes an FPGA-based NoC emulator which can achieve an ultra-fast simulation speed. We improve the scalability of the NoC emulator without simplifying the emulated architectures or using off-chip resources. We introduce a novel method which enables to accurately emulate NoC designs under synthetic workloads without using a large amount of memory by decoupling the time of the emulated NoC and the time of the traffic generators. Additionally, we propose a method based on the time-division multiplexing technique to emulate the behavior of the entire network using several physical nodes while effectively using FPGA resources. We show that an implementation of the proposed NoC emulator on a Virtex-7 FPGA can achieve 2, 745x simulation speedup over Booksim, one of the most widely used software-based NoC simulator, while maintaining the simulation accuracy. Thiem Van Chu, Shimpei Sato, Kenji Kise |
FCCM | 1 |
| 2015 | Ultra-fast NoC emulation on a single FPGAabstractNetwork-on-Chip (NoC) has become the de facto on-chip communication architecture for many-core systems. This paper proposes novel methods for emulating large-scale NoC designs on a single FPGA. Since FPGAs offer a highly parallel platform, FPGA-based emulation can be much faster than the software-based approach. However, emulating NoC designs with up to thousands of nodes is a challenging task due to the FPGA capacity constraints. We first describe how to accurately model synthetic workloads on FPGA by separating the time of the emulated network and the times of the traffic generation units. We next present a novel use of time-multiplexing in emulating the entire network using several physical nodes. Finally, we show the basic steps to apply the proposed methods to emulate different NoC architectures. The proposed methods enable ultrafast emulations of large-scale NoC designs with up to thousands of nodes using only on-chip resources of a single FPGA. In particular, more than 5,000× simulation speedup over BookSim, a widely used software-based NoC simulator, is achieved. Thiem Van Chu, Shimpei Sato, Kenji Kise |
FPL | 1 |
| 2014 | Ultrasmall: The smallest MIPS soft processorabstractSoft processors have been commonly used in FPGAbased designs to perform various useful functions. Some of these functions are not performance-critical and required to be implemented using very few FPGA resources. For such cases, it is desired to reduce circuit area of the soft processor as much as possible. This paper proposes Ultrasmall, a small soft processor for FPGAs. Ultrasmall supports a subset of the MIPS-I ISA and is designed for microcontrollers in FPGA-based SoCs. Ultrasmall employs an area efficient architecture to minimize the use of FPGA resources. While supporting the 32-bit ISA, Ultrasmall adopts the 2-bit wide serial ALU architecture. This approach significantly reduces the amount of FPGA resource usage. In addition to the device-independent optimizations for any FPGAs, we apply primitives-based optimizations for the Xilinx Spartan-3E FPGA series with 4-input LUTs, thereby further reducing the total number of occupied slices. The evaluation result shows that, on the Xilinx Spartan-3E XC3S500E FPGA, Ultrasmall occupies only 137 slices which is 84% of the number of occupied slices of Supersmall, a very small soft processor with the same design concept as Ultrasmall. On the other hand, in term of performance, Ultrasmall is 2.9× faster than Supersmall. Hiroshi Nakatsuka, Yuichiro Tanaka, Thiem Van Chu, Shinya Takamaeda-Yamazaki, Kenji Kise |
FPL | 3 |