Takao Onoye

dblp:39/5945 · DBLP profile ↗
← Back
75ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0002-1894-2448ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 44 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 2 since 2021Human-computer interaction and ubiquitous computing · 8Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Enabling Skew-aware Federated Learning on Embedded Systems via Non-IID Data Distribution Type Estimation
abstract
Federated learning enables decentralized training without sharing raw data, making it suitable for privacy aware applications. However, its performance often degrades in real settings due to unknown and diverse data differences among clients. While many mitigation strategies have been proposed, they typically assume prior knowledge of the imbalance type, such as feature variation, label bias, or data quantity differences, which is unrealistic in practice. This paper addresses the overlooked problem of identifying the main type of data imbalance. We propose a method called Machine Learning based Non IID Estimator, which classifies the imbalance type by analyzing trained client models without accessing any raw data. Similarity matrices computed from model parameters are used to train standard classifiers. Evaluations on the MNIST dataset under controlled imbalance settings show that the proposed method achieves perfect classification accuracy with lightweight models. This highlights the potential of distribution type estimation as a key step toward more robust and efficient federated systems.
Tatsuya Nishio, Hiroki Nishikawa, Ittetsu Taniguchi, Takao Onoye
EMSOFT4
2025 Balancing Efficiency and Comfort for Intersection Coordination in Autonomous Driving
abstract
Balancing traffic efficiency and passenger comfort is a fundamental challenge in autonomous intersection management. This paper presents a novel framework that integrates comfort-related metrics, such as minimizing jerk, alongside traffic efficiency optimization. Experimental evaluations demonstrate that the framework effectively balances efficiency and comfort under varying traffic conditions, outperforming methods that prioritize only one objective.
Fuma Sawa, Sangyoung Park, Muzaffer Citir, Hiroki Nishikawa, Ittetsu Taniguchi, Takao Onoye
VTC2025-Spring6
2024 Enhancing Driver Awareness of Vulnerable Road Users through In-Vehicle Auditory Signals
Kyoko Takii, Fuma Sawa, Wataru Kobayashi, Sangyoung Park, Muzaffer Citir, Hiroki Nishikawa, Ittetsu Taniguchi, Takao Onoye
VTC Fall8
2023 Live Demonstration: In-Vehicle Auditory Signal Evaluation Platform in A Driving Simulator
abstract
Advanced driver-assistance systems (ADAS) are gen-erally used to support a safe drive. However, if all the services in ADAS rely on visual suggestions, the driver becomes increasingly burdened and exhausted. As a solution, in-vehicle auditory signals to inform the driver of inattention have been appealing as another approach to altering visual suggestions in recent years. In this paper, we show our developed in-vehicle auditory signal evaluation platform in an existing driving simulator.
Fuma Sawa, Yoshinori Kamizono, Wataru Kobayashi, Ittetsu Taniguchi, Hiroki Nishikawa, Takao Onoye
ISCAS6
2023 Adaptive Sampling for Computer Vision-Oriented Compressive Sensing
abstract
Compressive sensing (CS) is renowned for its efficient signal data compression. However, due to its compressive nature, the accuracy of downstream computer vision (CV) tasks by reconstruction inevitably degrades as sampling rate decreases. This limitation significantly hinders the application of existing CS techniques. To overcome the drawback, this paper presents a novel CS technique that employs adaptive sampling rates based on saliency distribution. The goal of this work is to enhance the preservation of information necessary for classification while reducing the weight of non-essential information. Experimental results show the effectiveness of the proposed adaptive sampling technique, which outperforms existing sampling CS techniques on STL10 and Imagenette datasets. The average classification accuracy is maximally improved by 26.23% and 18.25%, respectively.
Hiroki Nishikawa, Jinjia Zhou, Ittetsu Taniguchi, Takao Onoye
MMAsia5
2022 Joint Representation Learning for Anomaly Detection in Surveillance Videos
abstract
Video anomaly detection in the unconstrained environment is challenging due to various background scenes, illuminations, and occlusions. Recent studies show that deep learning approaches can achieve remarkable performance on video anomaly detection. In this paper, we propose a joint representation learning structure for video anomaly detection. The proposed architecture extracts features from the object appearance and their associate motion features via different encoders based on ResNet network architecture. Our network architecture is designed to combine spatial and temporal features, which share the same decoder. Using a joint representation learning approach, the proposed architecture effectively learn both appearance and motion features to detect anomalies in various scene scenarios. The experiments on three benchmark datasets demonstrate the remarkable detection accuracy with respect to existing state-of-the-art methods, which achieve 96.5%, 86.9%, and 73.4% in UCSD Pedestrian, CHUK Avenue, and ShanghaiTech datasets, respectively.
Savath Saypadith, Takao Onoye
ISCAS2
2021 Thermal Comfort Aware Online Energy Management Framework for a Smart Residential Building
abstract
Energy management in buildings equipped with renewable energy is vital for reducing electricity costs and maximizing occupant comfort. Despite several studies on the scheduling of appliances, a battery, and heating, ventilating, and air-conditioning (HVAC), there is a lack of a comprehensive and time-scalable approach that integrates predictive information such as renewable generation and thermal comfort. In this paper, we propose an online energy management framework to incorporate the optimal energy scheduling and prediction model of PV generation and thermal comfort by the model predictive control (MPC) approach. The energy management problem is formulated as coordinated three optimization problems covering a fast and slow time-scale.This reduces the time complexity without a significant negative impact on the global nature and quality of the result. Experimental results show that the proposed framework achieves optimal energy management that takes into account the trade-off between the electricity bill and thermal comfort.
Daichi Watari, Ittetsu Taniguchi, Francky Catthoor, Charalampos Marantos, Kostas Siozios, Elham Shirazi, Dimitrios Soudris, Takao Onoye
DATE8
2021 Video Anomaly Detection Based on Deep Generative Network
abstract
In this paper, we present a framework for the detection of anomalies in video scenes. Both spatial and temporal features extract and learn through the framework. We employ inception modules and residual skip connections inside the framework to make the network learning higher-level features, which we call "multi-scale U-Net". A multi-scale U-Net kept useful features of the image that lost during training caused by the convolution operator. The numbers of training and testing parameters in our framework are reduced while the detection accuracy is still improved. We evaluated the proposed framework on three benchmark datasets: UCSD, CHUK Avenue and ShanghaiTech dataset. Our proposed framework achieved 95.7%, 86.8% and 73.0% in terms of AUC, which surpasses the state-of-art learning-based methods.
Savath Saypadith, Takao Onoye
ISCAS2
2021 An Infant-Like Device that Reproduces Hugging Sensation with Multi-Channel Haptic Feedback
abstract
Proximity interaction, such as hugging, plays an essential role in building relationships between parents and children. However, parents and children cannot freely interact in the neonatal intensive care unit due to visiting restrictions imposed by COVID-19. In this study, we develop a system of pseudo-proximity interaction with a remote infant through a VR headset by using an infant-like device that reproduces the haptic feedback features of the hugging sensation, such as weight, body temperature, breathing, softness, and unstable neck.
Taiyo Natomi, Yasuji Kitabatake, Kazuyuki Fujita, Takao Onoye, Yuichi Itoh
VRST4
2020 TuVe: A Shape-changeable Display using Fluids in a Tube
abstract
We propose TuVe, a novel shape-changing display consisting of a flexible tube and fluids, in which the droplets flowing through the tube compose the display medium that represents information. In this system, every colored droplet is flowed by controlling valves and a pump connected to the tube. The display part employs a flexible tube that can be shaped to any structure (e.g., wrapped around a specific object), which is achieved by a calibration made to capture the tube structure using image processing with a camera. A performance evaluation reveals that our prototype succeeds in controlling each droplet with a positional error of 2 mm or less, which is small enough to show such simple characters as alphabetic characters using a 7 × 7-pixel resolution display. We also discuss example applications, such as large public displays and flow-direction visualization, that illustrate the characteristics of the TuVe display.
Saya Suzunaga, Yuichi Itoh, Kazuyuki Fujita, Takao Onoye
AVI5
2020 A template-free object motion estimation method for industrial vision system in aligning machine
abstract
With the popularity of varied sensors as well as the development of communication technologies. Industrial vision systems have been deployed in many manufacture applications. Particularly, industrial vision systems coped with air nozzle are often adopted in aligning machines for object aligning. Since alignment efficiency depends on propriety of blown timing/pressure which is casually manifested in the motion of the observed blown objects, motion estimation method is critical to adopt. Unlike conventional methods typically estimate objects' motions using prior prepared templates, this paper proposed a template-free object motion estimation method for industrial vision system in aligning machine. By virtue of the properties of industrial images, an observation area with bounding box are initialized for each object, then motion estimation is achieved by updating them following expectation-maximization principle. Experiments revealed that the proposed method is able to achieve continuous motion estimation in less processing time.
Qiaochu Zhao, Ittetsu Taniguchi, Takao Onoye
ETFA3
2018 Fusion Networks for Air-Writing Recognition
Buntueng Yana, Takao Onoye
MMM (2)2
2018 Activation-Aware Slack Assignment for Time-to-Failure Extension and Power Saving
Yutaka Masuda, Takao Onoye, Masanori Hashimoto
IEEE Trans. Very Large Scale Integr. Syst.2
2017 GPGPU-based Highly Parallelized 3D Node Localization for Real-Time 3D Model Reproduction
abstract
This paper proposes a highly parallelized 3D node localization method based on cross-entropy method for the 3D modeling system. Cross-entropy localization statistically estimates node positions from node-to-node distance information by sampling, and each sample evaluation and internal computation of objective function can be processed in parallel. Experimental results show our GPGPU-based implementation achieved 5,163x and 61.5x speed up compared to a single processor and 80-processor implementations. In addition, for enhancing model reproduction accuracy, this work introduces a penalty function to mitigate flip ambiguity.
Kauzki Hirosue, Shohei Ukawa, Yuichi Itoh, Takao Onoye, Masanori Hashimoto
IUI4
2016 Critical path isolation for time-to-failure extension and lower voltage operation
abstract
Device miniaturization due to technology scaling has made manufacturing variability and aging more significant, and lower supply voltage makes circuits sensitive to dynamic environmental fluctuation. These may shorten the time to failure (TTF) of fabricated chips unexpectedly. This paper focuses on critical path isolation, which increases timing slack of non-intrinsic critical paths and decreases timing error occurrence probability in the circuit, and proposes a design methodology of isolated circuits for TTF extension and/or lower voltage operation. The proposed methodology selects a set of FFs for isolation using ILP so that it maximumly reduces the sum of gate-wise failure probabilities. We evaluated MTTF (Mean Time To Failure) of circuits with/without critical path isolation and examined how much supply voltage could be reduced without MTTF degradation. Evaluation results show that circuits with the proposed critical path isolation achieved 25% supply voltage reduction with 1.4% area overhead. With the same supply voltage, MTTF was improved by 14 orders of magnitude.
Yutaka Masuda, Masanori Hashimoto, Takao Onoye
ICCAD3
2016 Hardware-simulation correlation of timing error detection performance of software-based error detection mechanisms
abstract
Software-based error detection techniques, which includes EDM (error detection mechanisms) transformation, are used for error localization in post-silicon validation. This paper evaluates the performance of EDM for timing error localization with 65-nm test chips assuming the following two EDM usage scenarios; (1) localizing a timing error occurred in the original program, and (2) localizing potential timing errors that vary execution results. Experimental results show that the EDM transformation customized for quick error detection detects 25% of timing errors in the original program in the first scenario and 56% of non-masked errors in the second scenario. However, these hardware measurement results are not consistent with the simulation results of our previous work. To investigate the reason, we focus on the following two differences between hardware and simulation; (1) design of power distribution network, and (2) definition of timing error occurrence frequency. We update the simulation setup for filling the difference and re-execute the simulation. We confirm that the simulation and the chip measurement results are consistent, which validates our simulation methodology.
Yutaka Masuda, Masanori Hashimoto, Takao Onoye
IOLTS3
2016 Ketsuro-Graffiti: An Interactive Display with Water Condensation
abstract
In this paper, we propose a novel interactive display, Ketsuro-Graffiti, that provides information with water condensation. Recently, a lot of substantial displays have been proposed that utilize various types of physical material. These displays enable users to touch each pixel freely. Water condensation exists ubiquitously in daily life, and drawing on it is familiar. Therefore, a display with water condensation achieves an intuitive and familiar interaction between users and information. We propose a method to control generation and evaporation of water condensation and implement two prototypes of Ketsuro-Graffiti using peltier devices. One prototype uses an open loop control system and the other uses a PID closed loop control system with thermistors. Our evaluation of system performance indicates that both prototypes can control generation, evaporation, and thickness of water condensation.
Yuki Tsujimoto, Yuichi Itoh, Takao Onoye
ISS3
2015 An oscillator-based true random number generator with process and temperature tolerance
abstract
This paper presents an oscillator-based true random number generator (TRNG) that automatically adjusts the duty cycle of a fast oscillator to 50 %, and generates unbiased random numbers tolerating process variation and dynamic temperature fluctuation. Measurement results with 65nm test chips show that the proposed TRNG adjusted the probability of `1' to within 50 ± 0.07 % in five chips in the temperature range of 0 °C to 75 °C. Consequently, the proposed TRNG passed the NIST and DIEHARD tests at 7.5 Mbps with 6,670 μm2area.
Takehiko Amaki, Masanori Hashimoto, Takao Onoye
ASP-DAC3
2015 Reliability-configurable mixed-grained reconfigurable array compatible with high-level synthesis
abstract
This paper presents a mixed-grained reconfigurable VLSI array architecture that can cover mission-critical applications to consumer products through C-to-array application mapping. A proof-of-concept VLSI chip was fabricated in a 65nm process. Measurement results show that applications on the chip can be working in a harsh radiation environment.
Masanori Hashimoto, Dawood Alnajiar, Hiroaki Konoura, Yukio Mitsuyama, Hajime Shimada, Kazutoshi Kobayashi, Hiroyuki Kanbara, Hiroyuki Ochi, Takashi Imagawa, Kazutoshi Wakabayashi, Takao Onoye, Hidetoshi Onodera
ASP-DAC11
2015 Area efficient device-parameter estimation using sensitivity-configurable ring oscillator
abstract
This paper proposes an area efficient device parameter estimation method with sensitivity-configurable ring oscillator (RO). This sensitivity-configurable RO has a number of configurations and the proposed method exploits this property for reducing sensor area and/or improving estimation accuracy. The proposed method selects multiple sets of sensitivity configurations, obtains multiple estimates and computes the average of them for accuracy improvement exploiting an averaging effect. Experimental results with a 32-nm predictive technology model show that the proposed method can reduce the estimation error by 49% or reduce the sensor area by 75% while keeping the accuracy.
Shoichi Iizuka, Yuma Higuchi, Masanori Hashimoto, Takao Onoye
ASP-DAC4
2015 Performance Evaluation of Software-based Error Detection Mechanisms for Localizing Electrical Timing Failures under Dynamic Supply Noise
abstract
For facilitating error localization, software-based error detection techniques have been proposed and EDM (error detection mechanisms) transformation is one of these techniques. To discuss the effectiveness of EDM for electrical bug localization, two scenarios are considered; (1) localizing an electrical bug occurred in the original program, and (2) localizing as many potential bugs as possible. We experimentally evaluated the error detection performance in these two scenarios under dynamic power supply noise. Experimental results show that the EDM transformation customized for quick error detection cannot locate electrical bugs in the original program in the firs scenario, but it is useful for findin potential bugs in the second scenario.
Yutaka Masuda, Masanori Hashimoto, Takao Onoye
ICCAD3
2015 Real-time on-chip supply voltage sensor and its application to trace-based timing error localization
abstract
This paper presents an all-digital on-chip supply voltage sensor that captures one-shot voltage fluctuation every clock cycle. The proposed sensor was implemented on ASIC in 65nm process and FPGA. The obtained voltage resolution was 3.9mV and 29mV, respectively. This sensor is suitable for providing voltage information to trace-based error localization system. We experimentally show that the proposed sensor contributes to the facilitation of error localization.
Miho Ueno, Masanori Hashimoto, Takao Onoye
IOLTS3
2015 Stochastic timing error rate estimation under process and temporal variations
abstract
Reducing design and operational margin is a key factor that makes fabricated chips competitive in terms of speed and power consumption. On the other hand, a smaller margin involves a higher risk that a timing error occurs in field. This paper proposes a stochastic framework that estimates timing error rate under static process variations and dynamic environmental variations for circuits with and without run-time adaptive speed control. The proposed framework extends the state assignment of the continuous-time Markov process used in the previous work so as to take into account within-die random variation, and speeds up the database construction for the transition rate matrix by combining logic simulation and statistical static timing analysis. This paper also demonstrates that the proposed framework can cope with transistor-by-transistor stochastic aging processes. Experimental results show that the within-die random variation deviates the MTTF with σ of 52%. The CPU time for the transition rate matrix computation is reduced to 1/30.
Shoichi Iizuka, Yutaka Masuda, Masanori Hashimoto, Takao Onoye
ITC4
2015 3D node localization from node-to-node distance information using cross-entropy method
abstract
This paper proposes a 3D node localization method that uses cross-entropy method for the 3D modeling system. The proposed localization method statistically estimates the most probable positions overcoming measurement errors through iterative sample generation and evaluation. The generated samples are evaluated in parallel, and then a significant speedup can be obtained. We also demonstrate that the iterative sample generation and evaluation performed in parallel are highly compatible with interactive node movement.
Shohei Ukawa, Tatsuya Shinada, Masanori Hashimoto, Yuichi Itoh, Takao Onoye
VR5
2015 Hierarchical Structure-Based Fast Mode Decision for H.265/HEVC
abstract
The mode decision strategy adopted in the High Efficiency Video Coding test model increases the complexity drastically. In this paper, we present a course of low-complexity fast mode decision algorithms. First, the depth information of the colocated block from a previous frame is used to predict the size of current block. Next, for a certain sized block, the inter-prediction residual is analyzed to determine whether to terminate the mode decision process or to skip unnecessary modes and split the block into smaller sizes. After inter-prediction, a hardware-oriented low-complexity fast intra-mode decision algorithm is proposed. A fast discrete cross difference is adopted to detect the dominant direction of the block. In addition, four simple but efficient early termination strategies are proposed to terminate the rate-distortion optimization process properly. The simulation results show that these proposed algorithms reduce the encoding time by 54.0%-68.4% without incurring any noticeable performance degradation. Hardware synthesis results demonstrate that the proposed fast intra-mode decision algorithm can be easily implemented on a hardware platform with fewer resources consumed.
Wenjun Zhao 0001, Takao Onoye, Tian Song 0005
IEEE Trans. Circuits Syst. Video Technol.2
2014 Ketsuro-Graffiti: water condensation display
abstract
We propose a novel interactive display called Ketsuro-Graffiti, which allows users to write or apply graffiti to the surface of water-condensed mirror-like displays. Ketsuro-Graffiti controls the temperature of arbitrary points on its surface by moving heat conductive actuators up and down to control the condensation's generation and evaporation. We describe the details of its implementation and preliminary evaluation results.
Yohei Miyazaki, Yuichi Itoh, Yuki Tsujimoto, Masahiro Ando, Takao Onoye
Advances in Computer Entertainment5
2014 Hardware architecture of the fast mode decision algorithm for H.265/HEVC
abstract
In this paper, a course of low complexity Fast Mode Decision (FMD) algorithms as well as the corresponding hardware architecture is presented. Firstly, the depth information of co-located block from previous frame is used to predict the size of current block. Then, for a certain sized block, the inter prediction residual is analyzed to determine whether to terminate current check or to skip some unnecessary modes and split to smaller size. Finally, the corresponding hardware architecture is proposed based on state machine mechanism. Simulation results show that these proposed algorithms reduce the encoding time by 40.8~ 70.3%, without incurring any noticeable performance degradation. Hardware synthesis results demonstrate that the proposed architecture achieves a max frequency of about 193 MHz.
Wenjun Zhao 0001, Takao Onoye, Tian Song 0005
ICIP2
2014 Corrections to "A Speed-Up Scheme Based on Multiple-Instance Pruning for Pedestrian Detection Using a Support Vector Machine"
abstract
In the above paper (ibid., vol. 22, no. 12, pp. 4752-4761, Dec. 2013), several errors were introduced. These errors are corrected here.
Jaehoon Yu, Ryusuke Miyamoto, Takao Onoye
IEEE Trans. Image Process.3
2013 Emoballoon - A Balloon-Shaped Interface Recognizing Social Touch Interactions
Kosuke Nakajima, Yuichi Itoh, Yusuke Hayashi, Kazuaki Ikeda, Kazuyuki Fujita, Takao Onoye
Advances in Computer Entertainment6
2013 Stochastic error rate estimation for adaptive speed control with field delay testing
abstract
This paper proposes a stochastic framework for error rate estimation that models adaptive speed control as a continuous-time Markov process and derives its transition rates using developed similarity database. The proposed framework is implemented for adaptive speed control systems based on timing error prediction and scan-test. Experimental results show that the proposed framework enabled 12 orders of magnitude faster MTTF estimation than ordinary logic simulation. The accuracy of MTTF estimation under random delay fluctuation is clarified through a comparison with logic simulation. The proposed estimation can contribute to design and validation of adaptive speed control systems with field delay testing.
Shoichi Iizuka, Masafumi Mizuno, Dan Kuroda, Masanori Hashimoto, Takao Onoye
ICCAD5
2013 High-performance multiplierless transform architecture for HEVC
abstract
In this paper, a high-performance multiplierless VLSI architecture for the transform applied in the emerging video coding standard-High Efficiency Video Coding (HEVC) is presented. The proposed architecture can support a variety of transform sizes from 4×4 to 32×32, and some simplification strategies are adopted during the implementation, such as reusing part of a larger sized transform structure reused by smaller ones, and turning multiplications by constant into shift and sum operations. Synthesis results on the FPGA platform indicate that the proposed design can double the throughput compared with previous work, with almost the same hardware cost. Moreover, synthesis results under 45nm technology show that it can support real-time processing of 4Kx2K (4096×2048, 30fps) video sequences. When a comparison index called “data throughput per unit area” is adopted, the proposed architecture is almost five times more efficient than is the previous design.
Wenjun Zhao 0001, Takao Onoye, Tian Song 0005
ISCAS2
2013 Hardware-oriented fast mode decision algorithm for intra prediction in HEVC
abstract
In this paper, a hardware-oriented low complexity fast intra prediction algorithm is presented. The proposed algorithm adopts a fast Discrete Cross Differences (DCD) to detect the dominate direction of the coding unit. Based on DCD information, only a subset of the 35 candidate modes are selected for the rough mode decision process. Moreover, four simple but efficient early termination strategies are proposed to terminate the RDO process properly. Complexity analysis results demonstrate that the proposed algorithm can be easily implemented on hardware platform with fewer resources consumed. Simulation results show that the proposed algorithm reduce the encoding time by 66% on average, requiring bit-rate increase about 3.9% and PSNR decrease about 0.18dB, respectively.
Wenjun Zhao 0001, Takao Onoye, Tian Song 0005
PCS2
2013 Emoballoon: A balloon-shaped interface recognizing social touch interactions
abstract
People often communicate with others using social touch interactions including hugging, rubbing, and punching. We propose a soft social-touchable interface called “Emoballoon” that can recognize the types of social touch interactions. The proposed interface consists of a balloon and some sensors including a barometric pressure sensor inside of a balloon, and has a soft surface and ability to detect the force of the touch input. We construct the prototype of Emoballoon using a simple configuration based on the features of a balloon, and evaluate the implemented prototype. The evaluation indicates that our implementation can distinguish seven types of touch interactions with 83.5% accuracy.
Kosuke Nakajima, Yuichi Itoh, Yusuke Hayashi, Kazuaki Ikeda, Kazuyuki Fujita, Takao Onoye
VR6
2013 A gate-delay model focusing on current fluctuation over wide range of process-voltage-temperature variations
Kenichi Shinkai, Masanori Hashimoto, Takao Onoye
Integr.3
2013 A Worst-Case-Aware Design Methodology for Noise-Tolerant Oscillator-Based True Random Number Generator With Stochastic Behavior Modeling
abstract
This paper presents a worst-case-aware design methodology for an oscillator-based true random number generator (TRNG) that produces highly random bit streams even under deterministic noise. We propose a stochastic behavior model to efficiently determine design parameters, and identify a class of deterministic noise under which the randomness gets the worst. They can be used to directly estimate the worst χ value of a poker test under deterministic noise without generating bit streams, which enables efficient exploration of design space and guarantees sufficient randomness in a hostile environment. The proposed model is validated by measuring prototype TRNGs fabricated with a 65-nm CMOS process.
Takehiko Amaki, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye
IEEE Trans. Inf. Forensics Secur.4
2013 A Speed-Up Scheme Based on Multiple-Instance Pruning for Pedestrian Detection Using a Support Vector Machine
abstract
In pedestrian detection, as sophisticated feature descriptors are used for improving detection accuracy, its processing speed becomes a critical issue. In this paper, we propose a novel speed-up scheme based on multiple-instance pruning (MIP), one of the soft cascade methods, to enhance the processing speed of support vector machine (SVM) classifiers. Our scheme mainly consists of three steps. First, we regularly split an SVM classifier into multiple parts and build a cascade structure using them. Next, we rearrange the cascade structure for enhancing the rejection rate, and then train the rejection threshold of each stage composing the cascade structure using the MIP. To verify the validity of our scheme, we apply it to a pedestrian classifier using co-occurrence histograms of oriented gradients trained by an SVM, and experimental results show that the processing time for classification of the proposed scheme is as low as one-hundredth of the original classifier without sacrificing detection accuracy.
Jaehoon Yu, Ryusuke Miyamoto, Takao Onoye
IEEE Trans. Image Process.3
2013 Implementing Flexible Reliability in a Coarse-Grained Reconfigurable Architecture
abstract
This paper proposes a coarse-grained dynamically reconfigurable architecture that offers flexible reliability to deal with soft errors and aging. The notion of a cluster is introduced as a basic architectural element; each cluster can select four operation modes with different levels of spatial redundancy and area efficiency. We evaluate the aging effect due to negative bias temperature instability and illustrate that periodically alternating active cells with resting ones will greatly mitigate the effects of the aging process with a negligible power overhead. The area of circuits that are added for immunity to soft errors and for mitigating aging effects is 29.3% of the proposed reconfigurable device. A fault-tolerance evaluation of a Viterbi decoder mapped on the architecture suggests that there is a considerable tradeoff between reliability and area overhead. Finally, we design and fabricate a test chip that contains a 4 × 8 cluster array in a 65-nm process and demonstrate its immunity to soft errors. Accelerated tests using an alpha particle foil showed that the mean time to failure and failure in time are well characterized with the number of sensitive bits and that our architecture can trade off soft error immunity with the area of implementation.
Dawood Alnajiar, Hiroaki Konoura, Younghun Ko, Yukio Mitsuyama, Masanori Hashimoto, Takao Onoye
IEEE Trans. Very Large Scale Integr. Syst.6
2013 Supply Noise Suppression by Triple-Well Structure
abstract
This brief discusses the impact of twin- and triple-well structures on power supply noise, and a substrate model for simulating the power supply noise. We observedVssnoise reduction by the resistive network of the p-substrate andVddnoise reduction by the junction capacitance of a triple-well structure on a 90-nm test chip. Measurement results also showed that the total noise reduction of a triple-well structure is superior to that of a twin-well structure. The measurement results correlate well with the results obtained from the power supply noise simulation using a hierarchical resistive mesh model. Our simulation-based verification indicates that in common CMOS design, a triple-well structure can reduce the power supply drop by 10%-40% or the decoupling capacitance area by 5%-10%. We also verified that supply drop sensitivity to variation of the well junction capacitance is sufficiently small and that supply noise reduction using a triple-well structure is robust to process variation.
Yasuhiro Ogasahara, Masanori Hashimoto, Toshiki Kanamoto, Takao Onoye
IEEE Trans. Very Large Scale Integr. Syst.4
2012 Body bias clustering for low test-cost post-silicon tuning
abstract
Post-silicon tuning is attracting a lot of attention for coping with increasing process variation. However, its tuning cost via testing is still a crucial problem. In this paper, we propose tuning-friendly body bias clustering with multiple bias voltages. The proposed method provides a small set of compensation levels so that the speed and leakage current vary monotonically according to the level. Thanks to this monotonic leveling and limitation of the number of levels, the test-cost of post-silicon tuning is significantly reduced. During the body bias clustering, the proposed method explicitly estimates and minimizes the average leakage after the post-silicon tuning. Experimental results demonstrate that the proposed method reduces the average leakage by 25.3 to 51.9% compared to non clustering case. We reveal that two bias voltages are sufficient when only a small number of compensation levels are allowed for test-cost reduction. We also give an implication on how to synthesize a circuit to which post-silicon tuning will be applied.
Shuta Kimura, Masanori Hashimoto, Takao Onoye
ASP-DAC3
2012 A predictive delay fault avoidance scheme for coarse-grained reconfigurable architecture
abstract
A scheme for avoiding delay faults with slack assessment during standby time is proposed in this paper. The proposed scheme performs path delay testing and checks if the slack is larger than a threshold value using selectable delay embedded in basic elements (BE) on a coarse-grained reconfigurable device. If the slack is smaller than the threshold, a pair of BEs to be replaced, which maximizes the path slack, is identified. Experimental results show that for aging-induced delay degradation a small threshold slack, which is less than 1 ps in a test case, is enough to ensure the delay fault prediction.
Toshihiro Kameda, Hiroaki Konoura, Dawood Alnajiar, Yukio Mitsuyama, Masanori Hashimoto, Takao Onoye
FPL6
2012 A hierarchical motion smoothing for distributed scalable video coding
abstract
This paper proposes a method to improve the coding efficiency of the distributed video coding based on a scalable encoder/decoder without feedback link. The proposed method performs a hierarchical motion smoothing to enable accurate rate control. Specifically, reflecting motion information on a coding rate, the encoder can estimate code amount which satisfies the decoder's need. In addition, side information generation based on the hierarchical motion smoothing is proposed for further quality improvement. Experimental results show that the proposed method increases the PSNR-Y by 1.2 dB at 500 kbps in comparison with a block matching based method. Moreover the number of repeat operations to calculate cost value for side information generation can be reduced by 10 times from the conventional method, which demonstrates the practicability of the proposed method.
Kazuhito Sakomizu, Takashi Nishi, Takao Onoye
PCS3
2012 Cup-le: A cup-shaped device for conversational experiment
abstract
We propose a cup-shaped device called Cup-le for conversational experiment. Cup-le has various sensors to record participants' nonverbal behavior such as utterance. Cup-le saves the work of attaching sensors. Also, since Cup-le appears cup-shaped, it can sense without being aware of the sensor unlike conventional wearable sensors. Moreover, by attaching touch display, Cup-le can show information and can be operated. In this conversational experiment, we use the input-output as method to have a questionnaire during the experiment and capture participants' impression for their conversation. We conducted real conversational experiment with Cup-le, and we confirmed that Cup-le supported sufficiently the experiment without affecting their impression.
Yusuke Hayashi, Yuichi Itoh, Kazuki Takashima, Kazuyuki Fujita, Kosuke Nakajima, Ikuo Daibo, Takao Onoye
VR7
2012 A Ray Tracing Simulation of Sound Diffraction Based on the Analytic Secondary Source Model
abstract
This paper describes a novel ray tracing method for solving sound diffraction problems. This method is a Monte Carlo solution to the multiple integration in the analytic secondary source model of edge diffraction; it uses ray tracing to calculate sample values of the integrand. The similarity between our method and general ray tracing makes it possible to utilize the various approaches developed for ray tracing. Our implementation employs the OptiX ray tracing engine, which exhibits good acceleration performance on a graphics processor. Two importance sampling methods are derived from different aspects, and they provide an efficient and accurate way to solve the numerically challenging integration. The accuracy of our method was demonstrated by comparing its estimates with the ones calculated by reference software. An analysis of signal-to-noise ratios using an auditory filter bank was performed objectively and subjectively in order to evaluate the error characteristics and perceptual quality. The applicability of our method was evaluated with a prototype system of interactive ray tracing.
Masashi Okada, Takao Onoye, Wataru Kobayashi
IEEE Trans. Speech Audio Process.2
2012 Adaptive Performance Compensation With In-Situ Timing Error Predictive Sensors for Subthreshold Circuits
abstract
We present an adaptive technique for compensating manufacturing and environmental variability in subthreshold circuits using “canary flip-flop (FF),” which can predict timing errors. A 32-bit Kogge-Stone adder whose performance was controlled by body-biasing was fabricated in a 65-nm CMOS process. Measurement results show that the adaptive control can compensate process, supply voltage, and temperature variations and improve the energy efficiency of subthreshold circuits by up to 46% compared to worst-case design and operation with guardbanding. We also discuss how to determine design parameters, such as the inserted location and the buffer delay of the canary FF, supposing two approaches: configuration in the design phase and post-silicon tuning.
Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye
IEEE Trans. Very Large Scale Integr. Syst.4
2011 Jitter amplifier for oscillator-based true random number generator
abstract
This paper presents a jitter amplifier for oscillator-based TRNG (true random number generator). The proposed jitter amplifier fabricated in a 65nm CMOS process occupying the area of 3,300 μm2archives 8.4× gain at 25°C and significantly improves the entropy enough to pass randomness test.
Takehiko Amaki, Masanori Hashimoto, Takao Onoye
ASP-DAC3
2011 Implications of Reliability Enhancement Achieved by Fault Avoidance on Dynamically Reconfigurable Architectures
abstract
Fault avoidance methods on dynamically reconfigurable devices have been proposed to extend device life-time, while their quantitative comparison has not been sufficiently presented. This paper shows results of quantitative life-time evaluation by simulating fault avoidance procedures of representative five methods under the same conditions of wear-out scenario, application and device architecture. Experimental results reveal 1) MTTF is highly correlated with the number of avoided faults, 2) there is the efficiency difference of spare usage in five fault avoidance methods, and 3) spares should be prevented from wear-out not to spoil life-time enhancement.
Hiroaki Konoura, Yukio Mitsuyama, Masanori Hashimoto, Takao Onoye
FPL4
2011 An oscillator-based true random number generator with jitter amplifier
abstract
This paper presents an oscillator-based TRNG (true random number generator) with jitter amplifier. The proposed jitter amplifier fabricated in a 65nm CMOS process archives 8.4× gain at 25°C, and significantly improves randomness of output bitstream. The TRNG with the jitter amplifier enhances throughput per area by 94% compared to a TRNG with frequency dividers. The prototype TRNG occupies 6,300 μm2, generates 2 Mbps random bitstreams, and passes FIPS 140-2 randomness tests and 12 tests in NIST test suite.
Takehiko Amaki, Masanori Hashimoto, Takao Onoye
ISCAS3
2010 Adaptive performance control with embedded timing error predictive sensors for subthreshold circuits
abstract
This paper presents an adaptive technique for compensating manufacturing and environmental variability in subthreshold circuits using ¿canary flip-flop¿ that can predict timing errors. A 32-bit Kogge-Stone adder whose performance was controlled by body-biasing was fabricated in a 65 nm CMOS process. Measurement results show that the adaptive control can reduce the power dissipation by 46% in comparison with the worst-case design with guardbanding.
Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye
ASP-DAC4
2010 Clock skew reduction by self-compensating manufacturing variability with on-chip sensors
abstract
This paper presents a self-compensation scheme of manufacturing variability for clock skew reduction. In the proposed scheme, a CDN with embedded variability sensors tunes variable clock drivers for canceling the clock skew induced by manufacturing variability. We apply the proposed scheme for a mesh-style CDN in a 65nm technology and evaluate the deskewing effect as a function of the sensor performance. Experimental results show that the skew can be reduced by over 70% and the correlation coefficient between estimated and actual variabilities, which represents the sensor performance, should be more than 0.3 for skew reduction.
Shinya Abe, Kenichi Shinkai, Masanori Hashimoto, Takao Onoye
ACM Great Lakes Symposium on VLSI4
2010 Transistor Variability Modeling and its Validation With Ring-Oscillation Frequencies for Body-Biased Subthreshold Circuits
abstract
This paper presents transistor variability modeling and its validation for body-biased subthreshold circuits based on measurements of a device-array circuit using a 90-nm technology. The device array consists of p/nMOS transistors and ring oscillators. We examine and confirm the correlation between the performance variation model extracted from measured I-V characteristics and fabricated oscillation frequencies. We demonstrate that delay variations in subthreshold circuits are well characterized with two parameters, i.e., threshold voltage and subthreshold swing parameter. We also reveal that threshold voltage shift by body biasing can be deterministically modeled and statistical modeling is less meaningful.
Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye
IEEE Trans. Very Large Scale Integr. Syst.4
2009 Trade-off analysis between timing error rate and power dissipation for adaptive speed control with timing error prediction
abstract
Timing margin of a chip varies chip by chip due to manufacturing variability, and depends on operating environment and aging. Adaptive speed control with timing error prediction is a promising approach to mitigate the timing margin variation, whereas it inherently has a critical risk of timing error occurrence when a circuit is slowed down. This paper presents how to evaluate the relation between timing error rate and power dissipation in self-adaptive circuits with timing error prediction. The discussion is experimentally validated using a 32-bit ripple carry adder in subthreshold operation in a 90nm CMOS process. We show a trade-off between timing error rate and power dissipation, and reveal the dependency of the trade-off on design parameters.
Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye
ASP-DAC4
2009 Coarse-grained dynamically reconfigurable architecture with flexible reliability
abstract
This paper proposes a coarse-grained dynamically reconfigurable architecture, which offers flexible reliability to soft errors and aging. A notion of cluster is introduced as a basic element of the proposed architecture, each of which can select four operation modes with different levels of spatial redundancy and area-efficiency. Evaluation of permanent error rates demonstrates that four different reliability levels can be achieved by the proposed architecture. We also evaluate aging effect due to NBTI, and illustrate that alternating active cells with resting ones periodically will greatly mitigate the aging process with negligible power overhead. The area of additional circuits to attain immunity to soft errors and reliability configuration is 26.6% of the proposed reconfigurable device. Finally, a fault-tolerance evaluation of Viterbi decoder mapped on the proposed architecture suggests that there is a considerable trade-off between reliability and area overhead.
Dawood Alnajiar, Younghun Ko, Takashi Imagawa, Hiroaki Konoura, Masayuki Hiromoto, Yukio Mitsuyama, Masanori Hashimoto, Hiroyuki Ochi, Takao Onoye
FPL9
2009 Tuning-friendly body bias clustering for compensating random variability in subthreshold circuits
abstract
Post-fabrication tuning for mitigating manufacturing variability is receiving a significant attention. To reduce leakage increase involved in performance compensation by body biasing, body bias clustering methods have been proposed. However, conventional methods suffer from a large test cost for tuning after fabrication, since there are a tremendous number of body bias assignments. We in this paper propose a low-cost tuning scheme after fabrication and present a layout aware body bias clustering method. The proposed method estimates average leakage power after post-fabrication tuning, and minimizes it. We applied the proposed method to ultralow voltage circuits for suppressing their high sensitivity to random Vth variability, and demonstrated the effectiveness of the proposed method. In the experiments, by just introducing two clusters, leakage power after post-fabrication tuning was reduced by up to 70% compared to a single cluster case.
Koichi Hamamoto, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye
ISLPED4
2008 Dynamic supply noise measurement circuit composed of standard cells suitable for in-site SoC power integrity verification
abstract
This paper presents an all digital measurement circuit called “gated oscillator” for capturing waveforms of dynamic power supply noise. The gated oscillator is constructed with standard cells, and thus can be easily embedded in SoCs for design verification. The performance of the gated oscillator is verified with fabricated test chips in a 90nm process.
Yasuhiro Ogasahara, Masanori Hashimoto, Takao Onoye
ASP-DAC3
2008 Experimental study on body-biasing layout style-- negligible area overhead enables sufficient speed controllability --
abstract
Body-biasing is expected to be a common design technique, then area efficient implementation in layout has been demanded. Body-biasing outside standard cells is one of possible layouts. However in this case body-bias controllability, especially when forward bias is applied, is a concern. To investigate the controllability, we fabricated a ring oscillator in a 90nm technology, and measured the controllability. Our measurement result and evaluation of area efficiency reveal that body-biased circuits can be implemented with area overhead of less than 1% yet with sufficient speed controllability.
Koichi Hamamoto, Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye
ACM Great Lakes Symposium on VLSI5
2008 Correlation verification between transistor variability model with body biasing and ring oscillation frequency in 90nm subthreshold circuits
abstract
This paper presents modeling of manufacturing variability and body bias effect for subthreshold circuits based on measurement of a device array circuit in a 90nm technology. The device array consists of P/NMOS transistors and ring oscillators. This work verifies the correlation between the variation model extracted from IV measurement results and oscillation frequencies, which means the transistor-level variation model is examined and confirmed in terms of circuit performance. We demonstrate that delay variations of subthreshold circuits are well characterized with two parameters - threshold voltage and subthreshold swing parameter. We reveal that body bias effect is a less statistical phenomenon and threshold voltage shift by body biasing can be modeled deterministically.
Hiroshi Fuketa, Masanori Hashimoto, Yukio Mitsuyama, Takao Onoye
ISLPED4
2006 Automated Design of Digital Filters for 3-D Sound Localization in Embedded Applications
abstract
In this paper, an automated design method of digital filters for 3-D sound localization is proposed, which approximates head-related transfer function (HRTF). Target digital filter organization is dedicated for embedded applications so that ROM capacity and computational load are reduced with only a slight degradation of 3-D sound effect. The proposed design method consists of two steps. In the first step, the least-squares method is adopted to acquire initial design result quickly, while the second step utilizes the gradient search algorithm for precise optimization. In an objective evaluation, the proposed two-step method is proved to be effective for improving design results. Results of subjective listening tests also show that the proposed method is comparable to manual design in the approximation of reference HRTF. In these subjective tests, our automated design achieved better results per ROM capacity than simple FIR filter did
Kosuke Tsujino, Wataru Kobayashi, Takao Onoye, Yukihiro Nakamura
ICASSP (5)3
2006 A gate delay model focusing on current fluctuation over wide-range of process and environmental variability
abstract
This paper proposes a gate delay model that is suitable for timing analysis considering wide-range process and environmental variability. The proposed model focuses on current variation and its impact on delay is considered by replacing output load. The proposed model is applicable for large variability with current model constructed by DC analysis whose cost is small. The proposed model can also be used both in statistical static timing analysis and in conventional corner-based static timing analysis. Experimental results in a 90nm technology show that the gate delays of inverter, NAND and NOR are accurately estimated under gate length, threshold voltage, supply voltage and temperature fluctuation. We also verify that the proposed model can cope with slow input transition and RC output load. We demonstrate applicability to multiple-stage path delay and flip-flop delay, and show an application of sensitivity calculation for statistical timing analysis.
Kenichi Shinkai, Masanori Hashimoto, Atsushi Kurokawa, Takao Onoye
ICCAD4
2006 Quantitative Prediction of On-chip Capacitive and Inductive Crosstalk Noise and Discussion on Wire Cross-Sectional Area Toward Inductive Crosstalk Free Interconnects
abstract
Capacitive and inductive crosstalk noises are expected to be more serious in advanced technologies. However, capacitive and inductive crosstalk noises in the future have not been concurrently and sufficiently discussed quantitatively, though capacitive crosstalk noise has been intensively studied solely as a primary factor of interconnect delay variation. This paper quantitatively predicts the impact of capacitive and inductive crosstalk in prospective processes, and reveals that interconnect scaling strategies strongly affect relative dominance between capacitive and inductive coupling. Our prediction also makes the point that the interconnect resistance significantly influences both inductive coupling noise and propagation delay. We then evaluate a tradeoff between wire cross-sectional area and worst-case propagation delay focusing on inductive coupling noise, and show that an appropriate selection of wire cross-section can reduce delay uncertainty by the small sacrifice of propagation delay.
Yasuhiro Ogasahara, Masanori Hashimoto, Takao Onoye
ICCD3
2006 Probabilistic Pedestrian Tracking Based on a Skeleton Model
abstract
A novel pedestrian tracking scheme based on a particle filter is proposed, which adopts a skeleton model of a pedestrian as a state space model and uses distance transformed images for likelihood estimation. The six-stick skeleton model used in the proposed approach is very distinctive in representing a pedestrian simply but effectively, with which the efficient state space for the pedestrian tracking can be derived. Experimental results by using PETS sample sequences demonstrate that the proposed approach achieves highly accurate pedestrian tracking without any of prior learning.
Jumpei Ashida, Ryusuke Miyamoto, Hiroshi Tsutsui, Takao Onoye, Yukihiro Nakamura
ICIP4
2006 Efficient memory architecture for JPEG2000 entropy codec
abstract
An encoding/decoding process of JPEG2000 requires much more computation power than that of conventional JPEG mainly due to the complexity of entropy encoding/decoding. Thus usually multiple entropy codec hardware modules are implemented in parallel to process entropy encoding/decoding. This module, however, requests many small-size memories to store intermediate data, and when multiple modules are implemented on a chip, employment of the large number of SRAMs increases difficulty of whole chip layout. In this paper, an efficient memory architecture of the entropy encoding/decoding module is proposed, in which three approaches are attempted by utilizing one-bank SRAMs and internal registers. As a result, the efficient memory organization for a target process technology can be explored
Hiroki Sugano, Hiroshi Tsutsui, Takahiko Masuzaki, Takao Onoye, Hiroyuki Ochi, Yukihiro Nakamura
ISCAS4
2003 Real-time face object extraction for video phone
abstract
This paper devises a real-time face object extraction algorithm for moving pictures. Considering that most of the existing object extraction algorithms are of too intensive computational complexity to enable the real-time execution that is applied to mobile video phone run on an embedded processor, this algorithm intends not only to reduce the computational complexity by employing modified snake but also to improve the precision of extracting pixels by an efficient estimation mechanism of the skin color portion within a face object. Simulation results demonstrate that by this algorithm QCIF 15fps is processed on a StrongARM processor in real-time.
Kenji Hontani, Takaaki Imanaka, Gen Fujita, Takao Onoye, Isao Shirakawa
ICIP (3)4
2002 Adaptive rate control for JPEG2000 image coding in embedded systems
abstract
To cope with the recent mobile scenes where images are used aggressively, a novel rate control scheme is proposed. The proposed scheme, dedicated for JPEG2000 image coding, aims at achieving low computational cost and small working memory size while maintaining high image quality. By predicting the adequate number of coding passes and updating it adaptively in code-block coding, the proposed scheme reduces both computational cost and working memory size for bitstream buffering down to 29% and 13%, respectively.
Hiroshi Tsutsui, Takahiko Masuzaki, Tomonori Izumi, Takao Onoye, Yukihiro Nakamura
ICIP (3)4
2001 A dynamically reconfigurable hardware-based cipher chip
abstract
A cipher core has been implemented, which is dedicated to a 64-bit block, 128-bit key, dynamically reconfigurable hardware-based cipher, called "Chameleon", in which two 32-cell, 8-context dynamically reconfigurable hardware units are employed to generate new subkeys for each of the 16 iterations in the encryption/decryption process. The proposed architecture has been implemented by 0.6 um CMOS3LM technology, using 65.6K transistors and attaining a maximum throughput of 317.5 Mbps. The new approach demonstrates distinctive features of enhanced complexity and flexibility dedicatedly for embedded encryption/decryption applications in the mobile computing.
Yukio Mitsuyama, Zaldy Andales, Takao Onoye, Isao Shirakawa
ASP-DAC3
2001 Realtime wavelet video coder based on reduced memory accessing
abstract
In this paper, the VLSI implementation of a real-time EZW video coder is presented. The proposed architecture adopts a modified 2-D DWT subband decomposition scheme, with the purpose of reducing the transposition memory requirements of 2-D DWT. In addition, through the use of a parallelized partial zerotree EZW scheme, temporary buffer requirements between the DWT and EZW modules are also reduced. The video encoder is integrated in a 0.35 um 3LM chip by using 341 K transistors on a 4.93 x 4.93 mm2 die.
Roberto Yusi Omaki, Morgan Hirosuke Miki, Makoto Furuie, Daisuke Taki, Masaya Tarui, Gen Fujita, Takao Onoye, Isao Shirakawa
ASP-DAC8
2001 Spatiotemporal segmentation for compact video representation
Jianping Fan 0001, Jun Yu 0002, Gen Fujita, Takao Onoye, Lide Wu, Isao Shirakawa
Signal Process. Image Commun.4
2000 Layout generation of array cell for NMOS 4-phase dynamic logic (short paper)
abstract
Abstract| An array cell (AC) architecture for the la y out design is described, whic h is dedicated to lowpow er design b y means of the NMOS 4-phase dynamic logic.An AC is constructed of (M2N)+2 transistors so as to constitute each t ype of NMOS 4-phase logic gates.A graph theoretic approach is exploited in the la y out design to reduce the la y out area.A n umber of experimental results demonstrate the practicability o f the proposed approach.
Makoto Furuie, Bao-Yu Song, Yukihiro Yoshida, Takao Onoye, Isao Shirakawa
ASP-DAC4
2000 VLSI Implementation of a Reduced Memory Bandwidth Realtime EZW Video Coder
abstract
The architecture of a real-time wavelet video coder is described, with the main emphasis put on the memory bandwidth reduction and efficient VLSI implementation. The proposed architecture adopts a modified 2D subband decomposition scheme, along with a parallelized pipelined embedded zerotree wavelet coder architecture. The video encoder is integrated in a 0.35 /spl mu/m 3LM chip by using 341000 transistors on a 4.93/spl times/4.93 mm/sup 2/ die, which can process 720/spl times/480 30 fps pictures in real-time.
Roberto Yusi Omaki, Takao Onoye, Isao Shirakawa
ICIP3
1999 An architecture of a matrix-vector multiplier dedicated to video decoding and three-dimensional computer graphics
abstract
An architecture of a matrix-vector multiplier (MVM) is devised, which is dedicated to MPEG-4 natural/synthetic video decoding. The MVM can perform the matrix-vector multiplication both in the inverse discrete cosine transform (IDCT) and in the geometrical transformation of three-dimensional computer graphics (3-D CG); or, specifically, it can achieve the multiplication of a 4/spl times/4 matrix by a four-tuple vector necessary in the one-dimensional IDCT for eight pixels and in the geometrical transformation for a point in a 3-D space. This paper describes a new architecture of this MVM and also shows the implementation result of a functional module composed of four MVMs with the use of 440-k transistors, which can operate at 20 MHz or less.
Hideyuki Fujishima, Yusuke Takemoto, Takao Onoye, Isao Shirakawa
IEEE Trans. Circuits Syst. Video Technol.3
1998 Low-Power Implementation of H.324 Audiovisual Codec Dedicated to Mobile Computing
abstract
A VLSI implementation of the H.324 audiovisual codec is described. A number of sophisticated low-power architectures have been devised dedicatedly for the mobile use. A set of specific functional units, each corresponding to a process of H.263 video codec, is employed to lighten different performance bottlenecks. A compact DSP core composed of two MAC units is used for both ACELP and MP-MLQ coding schemes of the G.723.1 speech codec. The proposed audiovisual codec core has been implemented by using 0.35 /spl mu/m CMOS 4LM technology, which contains totally 420 K transistors with the dissipation of 224.32 mW from single 3.3 V supply.
Takao Onoye, Gen Fujita, Hiroyuki Okuhata, Morgan Hirosuke Miki, Isao Shirakawa
ASP-DAC1
1998 A low-power DSP core architecture for low bitrate speech codec
abstract
A VLSI implementation of a low-power DSP is described, which is dedicated to the G.723.1 low bitrate speech codec. A number of sophisticated DSP microarchitectures are devised mainly on dual multiply accumulators, rounding and saturation mechanisms, and two-banked on-chip memory. The proposed DSP architecture has been integrated in a total area of 7.75 mm/sup 2/ by using a 0.35 /spl mu/m CMOS technology, which can operate at 10 MHz with the dissipation of 45 mW from a single 3 V supply.
Hiroyuki Okuhata, Morgan Hirosuke Miki, Takao Onoye, Isao Shirakawa
ICASSP3
1997 Low-power H.263 video CoDec dedicated to mobile computing
abstract
Article Free Access Share on Low-power H.263 video CoDec dedicated to mobile computing Authors: Morgan H. Miki Dept. Information Systems Engineering, Osaka University Dept. Information Systems Engineering, Osaka UniversityView Profile , Gen Fujita Dept. Information Systems Engineering, Osaka University Dept. Information Systems Engineering, Osaka UniversityView Profile , Takao Onoye Dept. Information Systems Engineering, Osaka University Dept. Information Systems Engineering, Osaka UniversityView Profile , Isao Shirakawa Dept. Information Systems Engineering, Osaka University Dept. Information Systems Engineering, Osaka UniversityView Profile Authors Info & Claims ISLPED '97: Proceedings of the 1997 international symposium on Low power electronics and designAugust 1997 Pages 80–83https://doi.org/10.1145/263272.263288Published:01 August 1997Publication History 1citation333DownloadsMetricsTotal Citations1Total Downloads333Last 12 Months5Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Morgan Hirosuke Miki, Gen Fujita, Takao Onoye, Isao Shirakawa
ISLPED3
1997 An object code compression approach to embedded processors
abstract
Article An object code compression approach to embedded processors Share on Authors: Yukihiro Yoshida Dept. Information Systems Engineering, Osaka University Dept. Information Systems Engineering, Osaka UniversityView Profile , Bao-Yu Song Dept. Information Systems Engineering, Osaka University Dept. Information Systems Engineering, Osaka UniversityView Profile , Hiroyuki Okuhata Dept. Information Systems Engineering, Osaka University Dept. Information Systems Engineering, Osaka UniversityView Profile , Takao Onoye Dept. Information Systems Engineering, Osaka University Dept. Information Systems Engineering, Osaka UniversityView Profile , Isao Shirakawa Dept. Information Systems Engineering, Osaka University Dept. Information Systems Engineering, Osaka UniversityView Profile Authors Info & Claims ISLPED '97: Proceedings of the 1997 international symposium on Low power electronics and designAugust 1997 Pages 265–268https://doi.org/10.1145/263272.263349Online:01 August 1997Publication History 62citation364DownloadsMetricsTotal Citations62Total Downloads364Last 12 Months2Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Yukihiro Yoshida, Bao-Yu Song, Hiroyuki Okuhata, Takao Onoye, Isao Shirakawa
ISLPED4
1995 VLSI implementation of inverse discrete cosine transformer and motion compensator for MPEG2 HDTV video decoding
abstract
An MPEG2 video decoder core dedicated to MP@HL (Main Profile at High Level) images is described with the main theme focused on an inverse discrete cosine transformer and a motion compensator. By means of various novel architectures, the inverse discrete cosine transformer achieves a high throughput, and the motion compensator performs different types of picture prediction modes employed by the MPEG2 algorithm. The decoder core, implemented in the total chip area of 22.0 mm/sup 2/ by a 0.6-/spl mu/m triple-metal CMOS technology, processes a macroblock within 3.84 /spl mu/s, and therefore is capable of decoding HDTV (1920/spl times/1152 pels) images in real time.>
Toshihiro Masaki, Yasuo Morimoto, Takao Onoye, Isao Shirakawa
IEEE Trans. Circuits Syst. Video Technol.3
1994 Multi-Threaded Processor for Image Generation
abstract
Multiple instruction execution is a major approach to designing high-performance processors. Superscalar and VLIW processor that utilize instruction level parallelism are usually focused on. On the other hand, the multithreaded processor can be expected to achieve a high degree of multiple instruction execution by utilizing coarse grain parallelism. Many computer graphics applications (such as the radiosity method and ray-tracing method) can be optimized by reorganizing the code to take advantage of coarse grain parallelism, but the degree of instruction level parallelism is not sufficient for a superscalar processor. Experimental result using the radiosity method shows that the 4-thread multithreaded processor achieves 2.9 times speedup over single thread, while the 4-issue superscalar processor manages around 1.5 times. By duplicating two kinds of function units, the performance of a multithreaded processor increases to 3.7 times, but the performance of a superscalar processor is saturated at around 1.5 times. Therefore, for computer graphics applications, the multithreaded processor is a better approach than the superscalar processor.>
Takayuki Sagishima, Kozo Kimura, Hiroaki Hirata, Tokuzo Kiyohara, Shigeo Asahara, Takao Onoye, Isao Shirakawa
ISCAS6