EDBT 2026 Demo / reviewers in the wild / expert
Shinji Kimura
dblp:94/2788
· DBLP profile ↗
43ranked-venue papers
5as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prime Factorization Using Partially Constrained Multiple Quantum Annealing With Analytical and Pattern-Based Variable ReductionabstractFactorization of large semiprimes remains one of the most challenging problems for classical computers. Shor’s algorithm offers a quantum approach that reduces computational complexity, but its practical application is currently limited by hardware constraints. Meanwhile, as a provisional approach, quantum annealing (QA) has been explored through formulations of the quadratic unconstrained binary optimization (QUBO) problem. Among existing methods, the blockwise partial-product approach effectively reduced the QUBO variable count but was limited to semiprimes up to 21 bits. To extend factorization to larger semiprimes, this paper addresses key engineering challenges in constructing efficient QUBO formulations for prime factorization. We propose five techniques to reduce variable counts and improve scalability with current QA hardware: (1) dividing the problem into subproblems with partial constraints; (2) applying analytical reductions near the LSB; (3) applying analytical reductions near the MSB; (4) exploiting special patterns in semiprimes, with odd bit widths and long MSB-side zero sequences; and (5) balancing variable usage across both sides of the subproblem. Integrated into a QUBO converter, these methods enable stable factorization of semiprimes up to 47-bits within 20 seconds and can extend to special 2049-bit instances with 1001 consecutive MSB-side zeros. Geguang Miao, Shinichi Nishizawa, Shinji Kimura, Takashi Sato 0001 |
IEEE Trans. Computers | 4 |
| 2025 | Learned Image Codec on FPGA: Algorithm, Architecture and System DesignabstractThis paper describes our design for learned image codec (LIC) on FPGA, from the aspects of algorithm, architecture and system. For the algorithm, we build the neural network on the hyperprior structure. Besides, we present a quantization aware training scheme specifically adapted to LIC. For the architecture, we propose a fine-grained pipeline architecture. Channel parallelism constraint and neural network search are proposed to improve the DSP utilization and efficiency, respectively. For the system, we make a CPU-FPGA heterogeneous coding system in which a system-level pipeline is proposed to maximize the throughput. A 720P@30FPS demo and a cross-platform demo are provided in the websites.12 Heming Sun, Jing Wang 0181, Silu Liu, Shinji Kimura, Masahiro Fujita 0004 |
ASP-DAC | 4 |
| 2025 | SOME: Symmetric One-Hot Matching Elector - A Lightweight Microsecond Decoder for Quantum Error CorrectionabstractConventional quantum error correction (QEC) de-coders such as Minimum-Weight Perfect Matching (MWPM) and Union-Find (UF) offer high thresholds and fast decoding, respectively, but both suffer from high topological complexity. In contrast, Ising model-based decoders reduce topological complexity but demand considerable decoding time. We propose the Symmetric One-Hot Matching Elector (SOME), a novel decoder that reformulates the QEC decoding task as a Quadratic Unconstrained Binary Optimization (QUBO) problem—termed the One-Hot QUBO (OHQ). Each variable in the QUBO represents whether a given pair of flipped syndromes is matched, while the error probabilities between the pair are encoded as interaction coefficients (weight). Constraints ensure that each flipped syndrome is matched exactly once. Valid solutions of OHQ correspond to self-inverse permutation matrices, characterized by symmetric one-hot encoding. To solve the OHQ efficiently, SOME reformulates the decoding task as the construction of permutation matrices that minimize the total weight. It initializes each candidate matrix from one of the minimum-weight syndrome pairs, then iteratively appends additional pairs in ascending order of weight, and finally selects the permutation matrix with the lowest total energy. SOME achieves up to a 99.9x reduction in variable count and reduces decoding times from milliseconds to microseconds on a single-threaded commodity CPU. OHQ also maintains performance up to a 10.5% physical error rate, surpassing the highest known threshold of MWPM. Geguang Miao, Shinichi Nishizawa, Hiromitsu Awano, Shinji Kimura, Takashi Sato 0001 |
ICCAD | 5 |
| 2023 | Compressed Input Data Format of Quantum Annealing EmulatorabstractRecently, Quantum Annealing (QA) has attracted attention as an efficient algorithm for combinatorial optimization problems. In QA, the input data size becomes large and its reduction is important for the hardware emulation since its usable memory size and its bandwidth are limited. As an improved Coordinate (COO) format for a sparse matrix, Compressed Sparse Row (CSR) has been proposed [1]. We have checked CSR on an Ising model but found that the compression efficiency is not high. Sohei Shimomai, Kei Ueda, Shinji Kimura |
DCC | 3 |
| 2022 | Topology-Based Exact Synthesis for Majority Inverter GraphabstractSAT-based exact synthesis has important applications in logic optimization problems, and its scalability and computational speed greatly affect the optimization results. In the paper, a new topological constraint using the list of levels of inputs of each gate is introduced and accelerates the exact synthesis. Such topological constraints can reduce the search space by structure enumeration. By our new partition of the synthesis problem, we can maintain a good balance between runtime on a single satisfiability problem and the number of satisfiability problems. When compared to the fence-based method and the partial DAG based method, our methodology demonstrates a considerable reduction in runtime of 24.5% and 5.7%, respectively. Furthermore, our implementation can extend the scalability of SAT-based exact synthesis. Xianliang Ge, Shinji Kimura |
ISCAS | 2 |
| 2020 | Small-Area and Low-Power FPGA-Based Multipliers using Approximate Elementary ModulesabstractApproximate multiplier design is an effective technique to improve hardware performance at the cost of accuracy loss. The current approximate multipliers are mostly ASIC-based and are dedicated for one particular application. In contrast, FPGA has been an attractive choice for many applications, because of its high performance, reconfigurability, and fast development. This paper presents a novel methodology for designing approximate multipliers by employing the FPGA-based fabrics. The area and latency are significantly reduced by cutting the carry propagation path in the multiplier. Moreover, we explore higher-order multipliers on architectural space by using our proposed small-size approximate multipliers as elementary modules. For different accuracy requirements, eight configurations for approximate 8 × 8 multiplier are discussed. In terms of mean relative error distance (MRED), the accuracy loss of the proposed 8 × 8 multiplier is low as 0.17%. Compared with the exact multiplier, our proposed design can reduce area by 43.66% and power by 20.36%. The critical path latency reduction is up to 27.66%. The proposed multiplier design has a better accuracy-hardware tradeoff than other designs with com-parable accuracy. Yi Guo 0010, Heming Sun, Shinji Kimura |
ASP-DAC | 3 |
| 2020 | Accuracy-Configurable Low-Power Approximate Floating-Point Multiplier Based on Mantissa Bit SegmentationabstractNowadays, in energy-efficient design of digital systems, approximate computing (AC) has an increasingly important role. Due to human perceptual limitations, redundancy in input data and so on, there is a huge amount of applications that can tolerate errors. In this paper, an accuracy-configurable approximate floating-point (FP) multiplier is proposed to improve hardware consumption for such applications. Mantissa is divided into a short exactly processed part and a remaining approximately processed part. A new addition and shifting method is applied to the approximate part to replace multiplication to improve hardware performance. Experimental results show the 4-bit exact part configuration of the proposed work ensures the accuracy of 99.17% (MRED is 0.83%) with the reduction 67.65% of area, 16.64% of delay and 75.62% of power. The proposed work also shows good performance in image processing and neural networks. Yi Guo 0010, Shinji Kimura |
TENCON | 3 |
| 2018 | Sparse ternary connect: Convolutional neural networks using ternarized weights with enhanced sparsityabstractConvolutional Neural Networks (CNNs) are indispensable in a wide range of tasks to achieve state-of-the-art results. In this work, we exploit ternary weights in both inference and training of CNNs and further propose Sparse Ternary Connect (STC) where kernel weights in float value are converted to 1, -1 and 0 based on a new conversion rule with the controlled ratio of 0. STC can save hardware resource a lot with small degradation of precision. The experimental evaluation on 2 popular datasets (CIFAR-10 and SVHN) shows that the proposed method can reduce resource utilization (by 28.9% of LUT, 25.3% of FF, 97.5% of DSP and 88.7% of BRAM on Xilinx Kintex-7 FPGA) with less than 0.5% accuracy loss. Canran Jin, Heming Sun, Shinji Kimura |
ASP-DAC | 3 |
| 2018 | Quad-multiplier packing based on customized floating point for convolutional neural networks on FPGAabstractDeep convolutional neural networks (CNNs) are widely used in many computer vision tasks. Since CNNs involve billions of computations, it is critical to reduce the resource /power consumption and improve parallelism. Compared with extensive researches on fixed point conversion for cost reduction, floating point customization has not been paid enough attention due to its higher cost than fixed point. This paper explores the customized floating point for both the training and inference of CNNs. 9-bit customized floating point is found sufficient for the training of ResNet-20 on CIFAR-10 dataset with less than 1% accuracy loss, which can also be applied to the inference of CNNs. With reduced bit-width, a computational unit (CU) based on Quad-Multiplier Packing is proposed to improve the resource efficiency of CNNs on FPGA. This design can save 87.5% DSP slices and 62.5% LUTs on Xilinx Kintex-7 platform compared to CU using 32-bit floating point. More CUs can be arranged on FPGA and higher throughput can be expected accordingly. Zhifeng Zhang 0003, Dajiang Zhou, Shinji Kimura |
ASP-DAC | 4 |
| 2018 | Embedded Frame Compression for Energy-Efficient Computer Vision SystemsabstractComputer vision applications are rapidly gaining popularity in embedded systems, which typically involve a difficult trade-off between vision performance and energy consumption under a constraint of real-time processing throughput. Recently, hardware (FPGA and ASIC-based) implementations have emerged that significantly improve the energy efficiency of vision computation. These implementations, however, often involve intensive memory traffic that retains a significant portion of energy consumption at the system level. To address this issue, we present a lossy embedded compression framework to exploit the trade-off between vision performance and memory traffic for input images. Differential pulse-code modulation-based gradient-oriented quantization is developed as the lossy compression algorithm. We also present its hardware design that supports up to 12-scale 1080p@60fps real-time processing. For histogram of oriented gradient-based deformable part models on VOC2007, the proposed framework achieved a 49.6%-60.5% memory traffic reduction at a detection rate degradation of 0.05%-0.34%. For AlexNet on ImageNet, memory traffic reduction achieved up to 60.8% with less than 0.61% classification rate degradation. Li Guo 0006, Dajiang Zhou, Jinjia Zhou, Shinji Kimura |
ISCAS | 4 |
| 2018 | Sparseness Ratio Allocation and Neuron Re-pruning for Neural Networks CompressionabstractConvolutional neural networks (CNNs) are rapidly gaining popularity in artificial intelligence applications and employed in mobile devices. However, this is challenging because of the high computational complexity of CNNs and the limited hardware resource in mobile devices. To address this issue, compressing the CNN model is an efficient solution. This work presents a new framework of model compression, with the sparseness ratio allocation (SRA) and the neuron re-pruning (NRP). To achieve a higher overall spareness ratio, SRA is exploited to determine pruned weight percentage for each layer. NRP is performed after the usual weight pruning to further reduce the relative redundant neurons in the meanwhile of guaranteeing the accuracy. From experimental results, with a slight accuracy drop of 0.1%, the proposed framework achieves 149.3× compression on lenet-5. The storage size can be reduced by about 50% relative to previous works. 8-45.2% computational energy and 11.5-48.2% memory traffic energy are saved. Li Guo 0006, Dajiang Zhou, Jinjia Zhou, Shinji Kimura |
ISCAS | 4 |
| 2018 | Design of Power and Area Efficient Lower-Part-OR Approximate MultiplierabstractApproximate computing has been paid attention as a promising technique to decrease power and area for error-tolerant applications by simplifying the internal operations with sacrificing their accuracy. In this paper, a new power and area efficient approximate multiplier is proposed using OR based compressor with no carry propagation for lower bit positions and a carry propagation compressor with inexact half adders and full adders for upper bit positions. The proposal is effective to reduce the critical path delay with almost the same precision with previous methods. Firstly, an inexact half adder and an inexact full adder are proposed and a construction method of 4×4 multiplier is shown. Then, a construction method of 8×8 multiplier is proposed using OR based compressor, approximate 4×4 multipliers and an accurate 4×4 multiplier. The proposed construction method can also be applied to 16×16 multiplier. The accuracy loss of proposed multipliers is evaluated using MATLAB simulation and that of the proposed 8×8 multiplier is low as 0.20%, the effect of which is shown to be negligible by applying to discrete cosine transform (DCT), inverse DCT and convolutional neural networks for image classification. The proposed 8×8 multiplier reduces power and area by 50.78% and 53.19%, respectively, compared with the accurate Wallace tree multiplier when evaluated using SMIC 40nm process. Yi Guo 0010, Heming Sun, Shinji Kimura |
TENCON | 3 |
| 2018 | Energy-Efficient and High Performance Approximate Multiplier Using Compressors Based on Input ReorderingabstractToday in nanometer regime, approximate circuits have attracted much more attention due to the pursuit of low power consumption and high performance. Approximate multiplier is the key arithmetic function in many error-tolerant applications such as signal processing. In this paper, two approximate compressors based on input reordering logic have been proposed for the partial product reduction in the multiplication. A 2-bit reordering circuit is also designed for the proposed multiplier. Experimental results show that the proposed multiplier sacrifices only a small amount of precision (about 0.58%) to drastically reduce power, area, and delay up to (36.6%), (28.8%) and (20.1%) compared to accurate Wallace multiplier. The proposed multiplier achieves a high signal-to-noise ratio (over 50dB) when applied to an image processing algorithm. Zhenhao Liu, Yi Guo 0010, Xiaoting Sun, Shinji Kimura |
TENCON | 4 |
| 2018 | A Variable-Clock-Cycle-Path VLSI Design of Binary Arithmetic Decoder for H.265/HEVCabstractThe next-generation 8K ultra-high-definition video format involves an extremely high bit rate, which imposes a high throughput requirement on the entropy decoder component of a video decoder. Context adaptive binary arithmetic coding (CABAC) is the entropy coding tool in the latest video coding standards including H.265/High Efficiency Video Coding and H.264/Advanced Video Coding. Due to critical data dependencies at the algorithm level, a CABAC decoder is difficult to be accelerated by simply leveraging parallelism and pipelining. This letter presents a new very-large-scale integration arithmetic decoder, which is the most critical bottleneck in CABAC decoding. Our design features a variable-clock-cycle-path architecture that exploits the differences in critical path delay and in probability of occurrence between various types of binary symbols (bins). The proposed design also incorporates a novel data-forwarding technique (rLPS forwarding) and a fast path-selection technique (coarse bin type decision), and is enhanced with the capability of processing additional bypass bins. As a result, its maximum throughput achieves 1010 Mbins/s in 90-nm CMOS, when decoding 0.96 bin per clock cycle at a maximum clock rate of 1053 MHz, which outperforms previous works by 19.1%. Jinjia Zhou, Dajiang Zhou, Shuping Zhang, Shinji Kimura, Satoshi Goto |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | A low-cost approximate 32-point transform architectureabstractThis paper presents an area-efficient approximate method for 32-point transform which is one of the most area-consuming parts in High Efficiency Video Coding (HEVC) applications. Compared to prior literatures, this work reduces the hardware cost of transform by 1) eliminating all the arithmetic operations of 6 least significant bits (LSB), 2) presenting a low-delay method for generating carry propagation from the remaining 5 LSBs and 3) truncating the most significant bits (MSB) according to the position of component. In the implementation of a 32-point forward transform, the experimental results show that 27% area consumption can be saved and the coding efficiency loss aroused by the approximation is only 0.044% compared with the origin. Heming Sun, Zhengxue Cheng, Amir Masoud Gharehbaghi, Shinji Kimura, Masahiro Fujita 0004 |
ISCAS | 4 |
| 2017 | Effective write-reduction method for MLC non-volatile memoryabstractRecently, the requirement for non-volatile memory on embedded systems has increased because they can be applied with normally-off and power gating technologies to. However, they have a lower endurance than volatile memories. When data is encoded as a write-reduction code appropriately, the endurance of non-volatile memory can be enhanced by writing the encoded data into the memory. We propose a highly effective write-reduction method for a multi-level cell (MLC) non-volatile memory focusing on the write-reduction code (WRC) as the optimal bit-write reduction method. The WRC can be applied only to single-level cell non-volatile memory. The proposed method generates a cell-write reduction code based on the WRC; the cell has multiple bits as the holdable data. Our proposed method achieves a cell-write reduction by 31.6% compared to the conventional method. Masashi Tawada, Shinji Kimura, Masao Yanagisawa, Nozomu Togawa |
ISCAS | 2 |
| 2017 | Fast Algorithm and VLSI Architecture of Rate Distortion Optimization in H.265/HEVCabstractIn H.265/high efficiency video coding (HEVC) encoding, rate distortion optimization (RDO) is an important cost function for mode decision and coding structure decision. Despite being near-optimum in terms of coding efficiency, RDO suffers from a high complexity. To address this problem, this paper presents a fast RDO algorithm and its very large scale implementation (VLSI) for both intra- and inter-frame coding. The proposed algorithm employs a quantization-free framework that significantly reduces the complexity for rate and distortion optimization. Meanwhile, it maintains a low degradation of coding efficiency by taking the syntax element organization and probability model of HEVC into consideration. The algorithm is also designed with hardware architecture in mind to support an efficient VLSI implementation. When implemented in the HEVC test model, the proposed algorithm achieves 62% RDO time reduction with 1.85% coding efficiency loss for the “all-intra” configuration. The hardware implementation achieves 1.6 × higher normalized throughput relative to previous works, and it can support a throughput of 8k@30fps (for four fine-processed modes per prediction unit) with 256 k logic gates when working at 200 MHz. Heming Sun, Dajiang Zhou, Landan Hu, Shinji Kimura, Satoshi Goto |
IEEE Trans. Multim. | 4 |
| 2016 | CNN-MERP: An FPGA-based memory-efficient reconfigurable processor for forward and backward propagation of convolutional neural networksabstractLarge-scale deep convolutional neural networks (CNNs) are widely used in machine learning applications. While CNNs involve huge complexity, VLSI (ASIC and FPGA) chips that deliver high-density integration of computational resources are regarded as a promising platform for CNN's implementation. At massive parallelism of computational units, however, the external memory bandwidth, which is constrained by the pin count of the VLSI chip, becomes the system bottleneck. Moreover, VLSI solutions are usually regarded as a lack of the flexibility to be reconfigured for the various parameters of CNNs. This paper presents CNN-MERP to address these issues. CNN-MERP incorporates an efficient memory hierarchy that significantly reduces the bandwidth requirements from multiple optimizations including on/off-chip data allocation, data flow optimization and data reuse. The proposed 2-level reconfigurability is utilized to enable fast and efficient reconfiguration, which is based on the control logic and the multiboot feature of FPGA. As a result, an external memory bandwidth requirement of 1.94MB/GFlop is achieved, which is 55% lower than prior arts. Under limited DRAM bandwidth, a system throughput of 1244GFlop/s is achieved at the Vertex UltraScale platform, which is 5.48 times higher than the state-of-the-art FPGA implementations. Xushen Han, Dajiang Zhou, Shinji Kimura |
ICCD | 4 |
| 2016 | Power-efficient and slew-aware three dimensional gated clock tree synthesisabstractThis paper presents a three dimensional (3D) gated clock tree synthesis (CTS) approach, which consists of two steps: 1) abstract tree topology generation; and 2) 3D gated and buffered clock routing. 3D Pair Matching (3D-PM) algorithm is proposed to generate the initial tree topology and then the proposed TSV-minimization algorithm is applied to generate TSV-aware tree topology. Based on TSV-aware tree topology, 3D gated and buffered clock tree routing is done using the proposed 3D Gated and Buffered Deferred-Merge Embedding (3D-GB-DME) algorithm. The slew constraint satisfaction is considered and the clock skew is minimized in our approach. Experimental results show that the proposed method achieves 29.11% power reduction compared with the state-of-the-art 2D work. Minghao Lin, Heming Sun, Shinji Kimura |
VLSI-SoC | 3 |
| 2015 | A bit-write reduction method based on error-correcting codes for non-volatile memoriesabstractNon-volatile memory has many advantages over SRAM. However, one of its largest problems is that it consumes a large amount of energy in writing. In this paper, we propose a bit-write reduction method based on error correcting codes for non-volatile memories. When a data is written into a memory cell, we do not write it directly but encode it into a codeword. We focus on error-correcting codes and generate new codes called write-reduction codes. In our write-reduction codes, each data corresponds to an information vector in an error-correcting code and an information vector corresponds not to a single codeword but a set of write-reduction codewords. Given a writing data and current memory bits, we can deterministically select a particular write-reduction codeword corresponding to a data to be written, where the maximum number of flipped bits are theoretically minimized. Then the number of writing bits into memory cells will also be minimized. We perform several experimental evaluations and demonstrate up to 72% energy reduction. Masashi Tawada, Shinji Kimura, Masao Yanagisawa, Nozomu Togawa |
ASP-DAC | 2 |
| 2015 | An independent bandwidth reduction device for HEVC VLSI video systemabstractFRC (frame re-compression) is a kind of widely used technique in reducing the SDRAM (synchronous dynamic random access memory) bandwidth of HEVC video system. However, in previous research works, FRC imposes requirements on accessing pattern and hence its usage are only limited in HEVC video codecs. While in a typical HEVC VLSI video system, there exists many other video IPs with high bandwidth requirements. Therefore, in this article, we propose a new FRC architecture to overcome the limitation and make it applicable to all the video IPs in a HEVC VLSI video system, which raises the overall bandwidth reduction rate of the whole video system. Our proposal has two points: firstly we propose a system internal bus based FRC architecture, which is independent, transparent, and easily connected to all other video IPs. Secondly, we propose a FA (freely access) scheme to remove the requirements on access pattern in previous work. By using this proposal, the bandwidth reduction rate in our VLSI video system model is raised from 92.4% to 69.6%. Jiayi Zhu 0001, Li Guo 0006, Dajiang Zhou, Shinji Kimura, Satoshi Goto |
ISCAS | 4 |
| 2015 | Merge mode based fast inter prediction for HEVCabstractThe latest High Efficiency Video Coding (HEVC/H.265) obtains 50% bit rate reduction than H.264/AVC standard with comparable quality, but at the cost of high computational complexity. Inter prediction accounts for large complexity and merge mode is one of the most important new features introduced in HEVC. To address this issue, this paper utilizes the merge mode to accelerate inter prediction by three fast mode decision methods. 1) A merge candidate decision is proposed to select the best merge mode by Sum of Absolute Transformed Difference (SATD) cost to reduce the merge time. 2) An early merge termination is presented still based on SATD cost with more than 90% accuracy. 3) Based on efficient merge mode, symmetric motion partition (SMP) modes can be disabled for non-8 × 8 code units (CUs). Experimental results demonstrate that our work can achieve 53.1%-54.2% time reduction on average with 1.57%-2.30% BD-rate increment. Besides, our method achieves an improvement of 18%-30% time reduction with 0.89%-2.85% BD-rate increment when combined with other existing approaches. Zhengxue Cheng, Heming Sun, Dajiang Zhou, Shinji Kimura |
VCIP | 4 |
| 2014 | Fast SAO estimation algorithm and its VLSI architectureabstractSAO estimation is the process of determining SAO parameters in video encoding. There are two difficulties for VLSI implementation of SAO estimation. The first is that there are huge amount of samples to deal with in statistic collection phase. The other is that the complexity of RDO in parameters determination phase is very high. In this article, a fast SAO estimation algorithm and its corresponding VLSI architecture are proposed. For the first difficulty, we use bitmaps to collect statistic of all the 16 samples in one 4×4 block simultaneously. For the second difficulty, we simplify a series of complicated procedures in HM to balance the complexity and BD-rate performance. Experimental results show that the proposed algorithm maintains the picture quality improvement. The VLSI design based on this algorithm can be implemented by 156.32K gates, 8832 bits SPRAM, 400MHz @ 65nm technology and is capable of 8Kx4K @ 120fps encoding. Jiayi Zhu 0001, Dajiang Zhou, Shinji Kimura, Satoshi Goto |
ICIP | 3 |
| 2014 | An area-efficient 4/8/16/32-point inverse DCT architecture for UHDTV HEVC decoderabstractThis paper presents a new VLSI architecture for HEVC inverse discrete cosine transform (TDCT). Compared to prior arts, this work reduces hardware cost by: reducing computational logic of 1-D IDCTs with a reordered parallel-in serial-out (RPISO) scheme that shares the inputs of the butterfly structure; and reducing the area of the transpose buffer with a cyclic memory organization that achieves 100% I/O utilization of the SRAMs. In the implementation of a unified 4/8/16/32-point IDCT, the proposed schemes demonstrate 35% and 62% reduction of logic and memory costs, respectively. The IDCT implementation can support real-time decoding of 4K×2K 60fps video with a total hardware cost of 357,250um2on 2-D IDCT and 80,988um2on transpose memory in 90nm process. Heming Sun, Dajiang Zhou, Jiayi Zhu 0001, Shinji Kimura, Satoshi Goto |
VCIP | 4 |
| 2011 | Power and delay aware synthesis of multi-operand adders targeting LUT-based FPGAs
Taeko Matsunaga, Shinji Kimura, Yusuke Matsunaga |
ISLPED | 2 |
| 2010 | Multi-operand adder synthesis on FPGAs using generalized parallel countersabstractMulti-operand adders usually consist of compression trees which reduce the number of operands per a bit to two, and a carry-propagate adder for the two operands in ASIC implementation. The former part is usually realized using full adders or (3;2) counters like Wallace-trees in ASIC, while adder trees or dedicated hardware are used in FPGA. In this paper, an approach to realize compression trees on FPGAs is proposed. In case of FPGA with m-input LUT, any counters with up to m inputs can be realized with one LUT per an output. Our approach utilizes generalized parallel counters (GPCs) with up to m inputs and synthesizes high-performance compression trees by setting some intermediate height limits in the compression process like Dadda's multipliers. Experimental results show its effectiveness against existing approaches at GPC level and on Altera's Stratix III. Taeko Matsunaga, Shinji Kimura, Yusuke Matsunaga |
ASP-DAC | 2 |
| 2008 | Synthesis of parallel prefix adders considering switching activitiesabstractThis paper addresses parallel prefix adder synthesis which targets minimization of the total switching activities under bitwise timing constraints. This problem is treated as synthesis of prefix graphs which represent global structures of parallel prefix adders at technology-independent level. An approach for timing-driven area minimization has been proposed which first finds the exact minimum solution on a specific subset of prefix graphs by dynamic programming, then restructures the result for further reduction by removing restriction on the subset. This approach can be applied for switching cost minimization almost directly, though it is not so effective as area minimization in some cases. In this paper, a heuristic is proposed which estimates the effect of the restructuring phase and improve cost calculation for some specific cases. Through various kinds of experiments, conditions where this approach can be executed effectively is also discussed. Taeko Matsunaga, Shinji Kimura, Yusuke Matsunaga |
ICCD | 2 |
| 2006 | FCSCAN: an efficient multiscan-based test compression technique for test cost reductionabstractThis paper proposes a new multiscan-based test input data compression technique by employing a fan-out compression scan architecture (FCSCAN) for test cost reduction. The basic idea of FCSCAN is to target the minority specified 1 or 0 bits (either 1 or 0) in scan slices for compression. Due to the low specified bit density in test cube set, FCSCAN can significantly reduce input test data volume and the number of required test channels so as to reduce test cost. The FCSCAN technique is easy to be implemented with small hardware overhead and does not need any special ATPG for test generation. In addition, based on the theoretical compression efficiency analysis, improved procedures are also proposed for the FCSCAN to achieve further compression. Experimental results on both benchmark circuits and one real industrial design indicate that drastic reduction in test cost can be indeed achieved. Youhua Shi, Nozomu Togawa, Shinji Kimura, Masao Yanagisawa, Tatsuo Ohtsuki |
ASP-DAC | 3 |
| 2006 | Transition-based coverage estimation for symbolic model checkingabstractLack of complete formal specification is one of the major obstacles for the deployment of model checking. Coverage estimation addresses this issue by revealing the unverified part of the design according to the specified properties. In this paper, we propose a new transition-based coverage metric to evaluate the completeness of properties for symbolic model checking. It is more comprehensive and accurate than the existing coverage metrics for model checking. An efficient symbolic algorithm is presented for computing the transition coverage for a subset of ACTL. Our coverage estimator has been applied to the model checking of a cache coherence protocol. We uncovered several coverage holes including one that eventually led to the discovery of a design bug. Xingwen Xu, Shinji Kimura, Kazunari Horikawa, Takehiko Tsuchiya |
ASP-DAC | 2 |
| 2005 | Low Power Test Compression Technique for Designs with Multiple Scan ChainabstractThis paper presents a new DFT technique that can significantly reduce test data volume as well as scan-in power consumption for multiscan-based designs. It can also help to reduce test time and tester channel requirements with small hardware overhead. In the proposed approach, we start with a pre-computed test cube set and fill the don’t-cares with proper values for joint reduction of test data volume and scan power consumption. In addition we explore the linear dependencies of the scan chains to construct a fanout structure only with inverters to achieve further compression. Experimental results for the larger ISCAS’89 benchmarks show the efficiency of the proposed technique. Youhua Shi, Nozomu Togawa, Masao Yanagisawa, Tatsuo Ohtsuki, Shinji Kimura |
Asian Test Symposium | 5 |
| 2005 | Extended abstract: transition traversal coverage estimation for symbolic model checkingabstractModel checking can exhaustively verify whether a system (implementation) satisfies a set of properties (formal specification). However, the completeness of the formal specification itself is not clear and needs to be evaluated. Several coverage estimation methods have been proposed for this issue. In this paper, we present a transition traversal coverage method for a subset of CTL. With this method, we can detect the transitions, which are not verified by any property. It is more accurate and comprehensive than the state coverage method. Xingwen Xu, Shinji Kimura, Kazunari Horikawa, Takehiko Tsuchiya |
MEMOCODE | 2 |
| 2004 | Minimization of fractional wordlength on fixed-point conversion for high-level synthesis
Nobuhiro Doi, Takashi Horiyama, Masaki Nakanishi, Shinji Kimura |
ASP-DAC | 4 |
| 2004 | Alternative Run-Length Coding through Scan Chain Reconfiguration for Joint Minimization of Test Data Volume and Power Consumption in Scan TestabstractTest data volume and scan power are two major concerns in SoC test. In this paper we present an alternative run-length coding method through scan chain reconfiguration to reduce both test data volume and scan-in power consumption. The proposed method analyzes the compatibility of the internal scan cells for a given test set and then divides the scan cells into compatible classes. To extract the compatible scan cells we apply a heuristic algorithm by solving the graph coloring problem; and then a simple greedy algorithm is used to configure the scan chain for the minimization of scan power. Experimental results for the larger ISCAS'89 benchmarks show that the proposed approach leads to highly reduced test data volume with significant power savings during scan test. Youhua Shi, Shinji Kimura, Nozomu Togawa, Masao Yanagisawa, Tatsuo Ohtsuki |
Asian Test Symposium | 2 |
| 2002 | Folding of logic functions and its application to look up table compactionabstractThe paper describes the folding method of logic functions to reduce the size of memories for keeping the functions. The folding is based on the relation of fractions of logic functions. We show that the fractions of the full adder function have the bit-wise NOT relation and the bit-wise OR relation, and that the memory size becomes half (8-bit). We propose a new 3--1 LUT with the folding mechanisms whcih can implement a full adder with one LUT. A fast carry propagation line is introduced for a multi-bit addition. The folding and fast carry propagation mechanisms are shown to be useful to implement other multi-bit operations and general 4 input functions without extra hardware resources. The paper shows the reduction of the area consumption when using our LUTs compared to the case using 4--1 LUTs on several benchmark circuits. Shinji Kimura, Takashi Horiyama, Masaki Nakanishi, Hirotsugu Kajihara |
ICCAD | 1 |
| 2001 | A real-time 64-monosyllable recognition LSI with learning mechanismabstractIn the paper, a real-time 64-mono-syllable recognition LSI is presented. The LSI accepts 11.6 msec speech frame and outputs a 6-bit symbol-code for each frame by the end of the next frame with the pipelining manner. The recognition method is based on the Hidden Markov Model and is speaker-independent. An on-chip learning mechanism has also been designed, but the circuit is off-chip at present implementation because of the restriction of LSI area. The LSI is fabricated by VDEC Rohm with 0.6 um process on a 4.5 mm x 4.5 mm chip. Kazuhiro Nakamura, Qiang Zhu 0008, Shinji Maruoka, Takashi Horiyama, Shinji Kimura, Katsumasa Watanabe |
ASP-DAC | 5 |
| 2001 | Speech recognition chip for monosyllablesabstractIn the paper, we present a real-time speech recognition chip for monosyllables such as A, B, ..., etc. The chip recognizes up to 64 monosyllables based on the Hidden Markov Model (HMM), which is a well known speaker-independent recognition method. The chip accepts a short-speech frame including 256 16-bit digitized samples corresponding to 11.6 msec period, and outputs the 6-bit symbol code of monosyllables for 16 short-frames (corresponding to 185.6 msec). A learning circuit to update HMM parameters for the recognition chip has also been designed, and the recognition chip includes an interface to the learning circuit. We have fabricated the recognition chip by VDEC Rohm 0.6 um process on a 4.5 mm x 4.5 mm chip. We have also made a layout of the entire circuit including the learning circuit by VDEC Rohm 0.35 um process on a 4.9 mm x 4.9 mm chip. Kazuhiro Nakamura, Qiang Zhu 0008, Shinji Maruoka, Takashi Horiyama, Shinji Kimura, Katsumasa Watanabe |
ASP-DAC | 5 |
| 2000 | An application specific Java processor with reconfigurabilitiesabstractNo abstract available. Shinji Kimura, Hiroyuki Kida, Kazuyoshi Takagi, Tatsumori Abematsu, Katsumasa Watanabe |
ASP-DAC | 1 |
| 2000 | Multi-clock path analysis using propositional satisfiabilityabstractAbstract — We present a satisfiability based multi-clock path analysis method. The method uses propositional satisfiability (SAT) in the detection of multi-clock paths. We show a method to reduce the multi-clock path detection problems to SAT problems. We also show heuristics on the conversion from multi-level circuits into CNF formulae. We have applied our method to IS-CAS89 benchmarks and other sample circuits. Experimental results show the improvement on the manipulatable size of circuits by using SAT. I. Kazuhiro Nakamura, Shinji Maruoka, Shinji Kimura, Katsumasa Watanabe |
ASP-DAC | 3 |
| 1998 | Waiting false path analysis of sequential logic circuits for performance optimizationabstractThis paper introduces a new class of false path, which is sensi-tizable but does not affect the decision of the clock period. We call such false paths waiting false paths, which correspond to multi-cycle operations controlled by wait states. The allowable delay time of waiting false paths is greater than the clock pe-riod. When the number of allowable clock cycles for each path is determined, the delay of the path can be the product of the clock period and the allowable cycles. This paper presents a method to analyze allowable cycles and to detect waiting false paths based on symbolic traversal of FSM. We have applied our method to 30 ISCAS89 FSM benchmarks and found that 22 cir-cuits include such paths. 11 circuits among them include such paths which are critical paths, where the delay is measured as the number of gates on the path. Informations on such paths can be used in the logic synthesis to reduce the number of gates and in the layout synthesis to reduce the size of gates. 1 Kazuhiro Nakamura, Kazuyoshi Takagi, Shinji Kimura, Katsumasa Watanabe |
ICCAD | 3 |
| 1995 | Residue BDD and Its Application to the Verification of Arithmetic CircuitsabstractThe paper describes a veri cation method for arithmetic circuits based on residue arithmetic.In the veri cation, a residue module is attached to the speci cation and the implementation, and these outputs are compared by constructing BDD's.For the BDD construction without node explosion, we i n troduce a residue BDD whose width is less than or equal to a modulus.The method is useful for multipliers including C6288. Shinji Kimura |
DAC | 1 |
| 1992 | Precise timing verification of logic circuits under combined delay modelabstractA combined delay model to manipulate the variance of the delay time of logic elements and a timing verification method based on the theory of regular expressions are presented. Emphasis is placed on the hazard detection problem and the verification of asynchronous circuits. The effectiveness of the method with medium sized circuits including about 100 elements is shown.> Shinji Kimura, Shigemi Kashima, Hiromasa Haneda |
ICCAD | 1 |
| 1990 | A parallel algorithm for constructing binary decision diagramsabstractA parallel algorithm for constructing binary decision diagrams is described. The algorithms treats binary decision graphs as minimal finite automata. The automation for a Boolean function with AND as its main operation (OR operation) is obtained by forming the intersection (union) of the regular sets associated with its operands. The union and intersection operations are implemented by a product construction on the minimal automata for the regular sets. After each product construction step the automaton must be reminimized. The parallel algorithm is designed so that it is possible to find the minimal representations for several Boolean operations in parallel. The level of each operation is determined. Operations at the same level can be performed in parallel without any communication between processors. If there are relatively few operations in one level, then the product generation step is divided into several suboperations and the results are merged.> Shinji Kimura, Edmund M. Clarke |
ICCD | 1 |
| 1982 | An Interactive Simulation System for structured logic design - ISSabstractAn Interactive Simulation System (ISS) is presented. ISS is an integrated interactive CAD system for logic design, and is configurated “module oriented” to support structured logic design. An Interactive Simulator (IS) is used for design verification. A designer can control simulation steps interactively in IS, and he can find design errors early using a good interactive interface. A Structured Hardware Design Language (SHDL) is used to describe logic designs. Takeshi Sakai, Yoshiyuki Tsuchida, Hiroto Yasuura, Yasushi Ooi, Yoshitsugu Ono, Hiroshi Kano, Shinji Kimura, Shuzo Yajima |
DAC | 7 |