EDBT 2026 Demo / reviewers in the wild / expert
Takeshi Ikenaga
dblp:10/155
· DBLP profile ↗
133ranked-venue papers
5as first author
48since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 2 first-author · 17 since 2021Systems, architecture and hardware · 23 · 1 first-authorComputer networks · 14 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 8 since 2021Software engineering, systems software and programming languages · 6 · 4 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 3Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluation of WLAN Communication Characteristics In-Metal-Pipes with Non-Metallic SectionsabstractThe demand for low-cost wireless LAN deployment is increasing in indoor networks. However, in indoor environments, extending wireless LAN coverage across obstacles and complex layouts remains difficult. To address this issue, we aim to expand wireless LAN coverage by using the interior of metal pipes in buildings as radio-wave propagation paths. Since actual pipe systems often contain non-metallic sections, detailed performance evaluation under such mixed conditions is essential. This paper investigates communication performance in the presence of non-metallic segments through practical experiments using the IEEE 802.11 wireless standard, and discusses prospects for developing wireless network systems based on the results. Daiki Nobayashi, Takeshi Ikenaga |
CCNC | 3 |
| 2026 | Stress-Triggered Slice Switching for QoE Enhancement in Beyond 5G NetworksabstractIn 5G networks, users require high-quality services for a wide range of applications. However, conventional QoS control mechanisms do not reflect the quality actually experienced by end users in network management. This study proposes a stress-aware networking system for Beyond 5G environment utilizing user’s stress level data collected from wearable devices. This paper evaluates the effectiveness of the proposed system using free5GC. The evaluation demonstrated that performance variations among network slices influence video quality. These results indicate that dynamic network slice switching based on user stress levels has the potential to achieve QoE enhancement for users. Koki Kawaguchi, Daiki Nobayashi, Kazuya Tsukamoto, Takeshi Ikenaga |
CCNC | 4 |
| 2026 | Comparison of Congestion Controls for LEO Satellite Communications in the WildabstractLow Earth Orbit (LEO) satellite constellation services are expected to be a part of next-generation communication infrastructure. In LEO satellite communications, handover is required to switch communications between satellites and ground stations or terminals, which causes communication breakdown. These communication disruptions affect congestion control in transport protocols, leading to degradation of their performance. To address this degradation, congestion controls tailored for LEO satellites, such as StarQUIC, SatPipe, and TCP LEO, have been proposed. However, comprehensive evaluations of these congestion controls on a real LEO satellite network cannot be found, and evaluations are desired. To this end, this paper compares these LEO satellite congestion controls with one of major LEO satellite communication services, Starlink, evaluating them based on throughput, retransmissions, latencies, and file download time. The evaluation results have shown that TCP LEO employing transmission freeze during handover has shown the best throughput, the minimum latencies, and the shortest download time. Regarding retransmissions, SatPipe has achieved the lowest number of retransmissions. Motoyuki Ohmori, Kohichi Ogawa, Hiroki Kashiwazaki, Takeshi Ikenaga |
CCNC | 4 |
| 2026 | Modeling and Optimization of Data Dissemination and Retrieval in Spatio-Temporal Data-Retention System
Makoto Misumi, Daiki Nobayashi, Kazuya Tsukamoto, Takeshi Ikenaga |
COMPSAC | 5 |
| 2026 | CUBIC++: Empirical Fine-Tuned CUBIC
Motoyuki Ohmori, Kohichi Ogawa, Hiroki Kashiwazaki, Takeshi Ikenaga |
COMPSAC | 4 |
| 2026 | Interaction-aware representation learning for action quality assessment in freestyle skiing big air
Shiyue Chen, Xina Cheng, Takeshi Ikenaga |
Comput. Vis. Image Underst. | 5 |
| 2026 | Statistic temporal checking and spatial consistency based 3D size reconstruction of multiple objects from indoor monocular videos
Xina Cheng, Takeshi Ikenaga |
Image Vis. Comput. | 3 |
| 2026 | Toward Free-Form Local Feature MatchingabstractExisting feature matching methods are strongly coupled to their pre-defined position priors. For instance, sparse matchers are coupled to keypoints, and semi-dense matchers are coupled to grids. The coupled position prior dictates the distribution of matching points and imposes inherent limitations on the matcher. Consequently, sparse matchers suffer from a reliance on keypoint repeatability, while semi-dense matchers lack texture-based precision. Our preliminary work RCM leverages the keypoint prior in the source image and the grid prior in the target image, ensuring texture-based precision with keypoints while eliminating reliance on repeatability. However, RCM still relies heavily on keypoints in the source image, inheriting limitations such as sparsity and poor distribution in challenging scenes. To address these challenges, we introduce RCM+, which presents a novel free-form matching paradigm. By combining a position-agnostic encoder with a parameter-free decoder, we decouple the matcher from any position prior. As a result, the free-form matcher can match arbitrary input positions in a zero-shot manner, including detected keypoints, lines, edges, grids of any resolution, user-specified points, and more. This paradigm offers exceptional flexibility, allowing users to select position priors based on scene properties without retraining. Thus, RCM+ can leverage the advantages of various position priors without over-relying on any single prior, avoiding limitations in specific scenarios. To better match multiple position priors, we propose the Balancer, which reconciles all input position priors to achieve a more favorable point distribution for downstream tasks. Additionally, we enhance the view switcher and conflict-free matching layer introduced in RCM, further improving matching quality. Comprehensive experiments demonstrate the excellent performance, efficiency, and flexibility of RCM+, underscoring its promising potential for applications. Xiaoyong Lu, Songlin Du, Yaping Yan, Xiaobo Lu, Takeshi Ikenaga |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | SceneGlue: Scene-Aware Transformer for Feature Matching Without Scene-Level AnnotationabstractLocal feature matching plays a critical role in understanding the correspondence between cross-view images. However, traditional methods are constrained by the inherent local nature of feature descriptors, limiting their ability to capture non-local scene information that is essential for accurate cross-view correspondence. In this paper, we introduce SceneGlue, a scene-aware feature matching framework designed to overcome these limitations. SceneGlue leverages a hybridizable matching paradigm that integrates implicit parallel attention and explicit cross-view visibility estimation. The parallel attention mechanism simultaneously exchanges information among local descriptors within and across images, enhancing the scene’s global context. To further enrich the scene awareness, we propose the Visibility Transformer, which explicitly categorizes features into visible and invisible regions, providing an understanding of cross-view scene visibility. By combining explicit and implicit scene-level awareness, SceneGlue effectively compensates for the local descriptor constraints. Notably, SceneGlue is trained using only local feature matches, without requiring scene-level groundtruth annotations. This scene-aware approach not only improves accuracy and robustness but also enhances interpretability compared to traditional methods. Extensive experiments on applications such as homography estimation, pose estimation, image matching, and visual localization validate SceneGlues superior performance. The source code is available at https://github.com/songlindu/ SceneGlue. Songlin Du, Xiaoyong Lu, Yaping Yan, Guobao Xiao, Xiaobo Lu, Takeshi Ikenaga |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Dataset Quantization Augmentation: Improving Dataset Compression Through Complexity-Guided Sampling and AugmentationabstractTraining deep neural networks (DNNs) typically requires large-scale datasets, which poses substantial challenges related to computing resources and storage. Dataset Quantization (DQ) was introduced to compress large datasets into smaller subsets for training various neural networks. However, DQ lacks consideration of sample importance and diversity, which may result in excluding crucial samples or insufficiently representing the dataset, thereby limiting model generalization. To resolve this, we propose Dataset Quantization Augmentation (DQA), an enhanced framework that integrates image complexity into dataset compression and dynamically applies data augmentation techniques based on sample complexity. By selecting the least complex samples and applying data augmentation techniques (e.g., CutBlur, CutMix, Cutout, geometric transformations), DQA enhances dataset diversity and boosts model performance. Experimental results demonstrate that DQA outperforms the original DQ method and other Dataset compression methods. On CIFAR-10, DQA achieves a 3.64% accuracy improvement over DQ with only 10% of the dataset. For HDR reconstruction using the HAT model on Flickr2K, DQA achieves a PSNR 0.1761 higher and an SSIM 0.132 higher than DQ at 30% compression. These results highlight DQA’s versatility and effectiveness in both high-level and low-level vision tasks, making it a promising approach for dataset optimization. Qin Liu 0002, Fengshan Zhao, Takeshi Ikenaga |
ICME | 5 |
| 2025 | Optimizing Dataset Evaluation In Vision Tasks: Redundancy Is Out, Diversity Is In
Qin Liu 0002, Fengshan Zhao, Takeshi Ikenaga |
PRCV (9) | 5 |
| 2025 | Geometry-Aware Contextual Reasoning-Based Indoor Accessibility Detection System for Visually Impaired Wheelchair Users
Fanxiang Zhou, Xina Cheng, Takeshi Ikenaga |
PRCV (6) | 4 |
| 2025 | Skeleton-Aware Representation of Spatio-Temporal Kinematics for 3D Human Motion Predictionabstract3D human motion prediction, which attempts to foresee the behaviors of human, is an issue of great significance in computer vision. Attention-based neural networks and graph convolution networks (GCNs) have recently shown great promise in 3D skeleton-based human motion prediction for their attractive performance in learning spatial and temporal kinematics. However, existing methods have several critical issues: 1) Spatial dependencies for distal joints in each independent frame are hard to learn; 2) The GCN ignores hierarchical structure and diverse motion patterns of different body parts; 3) Existing methods disregard the statistical interdependence inherent in time series data. To address these issues, this paper proposes a skeleton-aware representation of spatio-temporal kinematics for 3D human motion prediction. The proposed method makes three key contributions: a learnable temporal aggregation, a skeleton-aware spatio-temporal attention, and an upper/lower decoupling GCN. The learnable temporal aggregation selectively obtains past information by leveraging the dependencies between each time step and its historical moments. The skeleton-aware spatio-temporal attention method leverages the self-attention mechanism and a designed adjacency matrix to model the skeleton constraints of distal joints. The upper/lower decoupling GCN introduces a grouping strategy to learn the dynamics of various body parts separately. Experimental results on three publicly available datasets demonstrate that the proposed method achieves state-of-the-art performances for both short-term prediction and long-term prediction. Note to Practitioners—3D human motion prediction forms a fundamental component of human-centered automation systems by enabling safer, more efficient, and more natural interactions between humans and machines. This paper was motivated by the challenges of predicting human motion: 1) Explicitly capturing the complex spatial patterns of distal joints is challenging; 2) Neglecting the inter-part variations of motion dynamics is problematic; 3) Neglecting the statistical interdependence inherent in time series data of human motion leads to poor performance. This paper suggests a skeleton-aware representation of spatio-temporal kinematics for 3D human motion prediction through three innovations: a learnable temporal aggregation, a skeleton-aware spatio-temporal attention, and an upper/lower decoupling GCN. The three contributions overcome the weaknesses of existing works and made a pioneering attempt of skeleton-aware representation of spatio-temporal human kinematics. It will significatively advance the development of many automation systems relevant to human motion prediction such as human-robot interaction and teleoperation. Songlin Du, Zhihan Zhuang, Zenghui Wang 0009, Yuan Li 0058, Takeshi Ikenaga |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Accuracy Improvement Method for Malicious Domain Detection Using Machine LearningabstractWith the widespread Internet technologies, malware damage also spreads worldwide, making it necessary to address these issues urgently. In some cases, malware-infected terminals use the Domain Name System (DNS) when communicating with the Command and Control (C&C) servers to obtain information for attacks. The previous malware detection focuses on the DNS communication history of malware-infected terminals. However, this method has the problem of poor accuracy in detecting malicious domains when the analysis data is small. This paper proposes a malicious domain detection with the following improvements. The first improvement is adding information on response and time. The second improvement is shortening the query domain names to primary domain names. Further, the proposed method showed improvement in the experiment. Toshiki Koga, Daiki Nobayashi, Takeshi Ikenaga |
CCNC | 3 |
| 2024 | Performance Evaluation of the Impact Between IEEE 802.11ah and Private LoRa Using 920 MHz BandabstractIEEE 802.11ah (11ah) is a wireless LAN standard using the 920 MHz band and is expected to be used in the IoT field as one of the Low Power Wide Areas (LPWAs). Many LPWAs share the 920 MHz band and are characterized by low transmission rates but long-range communication. Therefore, there is a high possibility of contending communication in the real-world environment. In this study, we clarify the characteristics of contending communications between 11ah and LoRa according to evaluating the performance using real devices. Yusuke Takeuchi, Daiki Nobayashi, Takeshi Ikenaga |
CCNC | 3 |
| 2024 | Packet Recovery Method with Redundant Packets for Large Spatio-Temporal Data RetentionabstractCyber Physical System (CPS) is a system that accumulates information from physical space, analyzes them in cyberspace, and then feeds back the results. We aim to realize a Floating Cyber Physical System (F-CPS), a regionally distributed CPS. We have proposed a large-capacity spatio-temporal data retention system for delivering data and function (application) in the F-CPS. In a previous work, we proposed a data completeness-aware transmission control and evaluated it by simulation. However, considering the movement of obstacles and fluctuation of radio waves in the real environment, packet loss occurs frequently. Thereby, the system may not be able to have complete data and may not function. In this paper, we proposed a method to recover missing data using redundant packets and evaluated its impact on the entire system by testing it with actual equipment. Naoki Tanaka, Daiki Nobayashi, Kazuya Tsukamoto, Takeshi Ikenaga |
CCNC | 4 |
| 2024 | Dynamic Transmission Parameter Settings Based on Data Transmission and Reception Performance for 920MHz LoRa CommunicationabstractLPWA (Low Power Wide Area) is expected to be a communication technology suitable for the Internet of Things (loT) because of its features. Since LoRa communication has a trade-off between communication range and transmission rate, performance is inherently varied depending on the surrounding environment. Therefore, to achieve stable and efficient communication performance, it is necessary to set appropriate parameters for LoRa depending on the communication environment between the base station and the terminal. We here propose a method for dynamic control of parameter settings based on reception statistics of beacons. The effectiveness of the proposed method is evaluated from the real experiments. Rikuto Tanaka, Daiki Nobayashi, Kazuya Tsukamoto, Takeshi Ikenaga, Goshi Sato, Kenichi Takizawa |
CCNC | 4 |
| 2024 | Multi-Level Spatial-Temporal Feature Aggregation and Alignment-Based Selective Residual Dense Propagation Module for HDR Video ReconstructionabstractTo reconstruct high dynamic range (HDR) video from alternating exposed low dynamic range (LDR) frames, the key is to address the misalignment and imprecise fusion caused by information loss and noise in ill-exposed regions. Following a coarse-to-fine manner, a Multi-level Spatial-Temporal feature aggregation and alignment-based Selective Residual Dense Propagation Network (MSTSRDPNet) is proposed. The Multi-level Spatial-Temporal aggregation extracts spatial-temporal features and aggregates them to mitigate information loss for fusion. The alignment-based Selective Residual Dense Propagation module reconstructs the aligned feature by using channel attention to redistribute feature weights while leveraging residual dense connections for information propagation. Experiments show that the proposed MSTSRDPNet outperforms all conventional methods on the synthetic dataset with PSNR-T, HDR-VQM, and HDR-VDP-2 scores of 44.64 dB, 86.83, and 73.9. Yiyu Liu, Fengshan Zhao, Qin Liu 0002, Takeshi Ikenaga |
ICASSP | 4 |
| 2024 | Poster: Tcp Congestion Control Based on Transmission Rate of Wireless Lan InterfacesabstractTCP congestion control uses end-to-end delay and packet discard information to control the transmission rate. In the wireless LAN, there is a time lag between TCP congestion control and the transmission rate control of the wireless LAN interfaces, and there is no coordination between the respective controls. Thus, congestion control corresponding to the current state of the wireless network can be achieved by shortening this time lag. Therefore, this paper proposes a method to control the congestion window by detecting changes in the transmission rate of the wireless LAN interfaces. Our proposed method enables to control congestion windows size by using data link layer information, instead of using only transport layer information. This paper verified the effectiveness of the proposed method by network simulator ns-3.40. Akira Okada, Daiki Nobayashi, Takeshi Ikenaga |
ICNP | 3 |
| 2024 | Experimental Evaluation for TCP/IP Communication Performance in Multi-Hop Private LoRa NetworkabstractLow-power wide-area (LPWA) is a major communicaiton technology used in the Internet of Things (Io‘I’), Major sensor networks are constructed using original protocols specific to each LPWA standard and cannot communicate directly with devices connected to the Internet. To enhance the convenience of devices using LPWA, it is necessary to improve their interop-erability with Internet applications with TCP/IP. Previous study have enabled Single-hop TCP/IP communication over Private LoRa. In this study, we construct a multi-hop environment for TCP/IP communication over Private LoRa, and evaluate its performance using both UDP and TCP communication. Previous research has enabled one-to-one communication using UDP and TCP over a Private LoRa interface, an LPWA technology. This study leverages the characteristics of LPWA for ultra-Iong-distance transmission and implements TCP/IP communication in a multi-hop environment. This paper evaluates the performance of UDP and TCP on actual devices. Eisho Aramaki, Daiki Nobayashi, Kazuya Tsukamoto, Takeshi Ikenaga, Goshi Sato, Kenichi Takizawa |
LANMAN | 4 |
| 2024 | Bidirectional temporal and frame-segment attention for sparse action segmentation of figure skating
Xina Cheng, Yuan Li 0058, Takeshi Ikenaga |
Comput. Vis. Image Underst. | 4 |
| 2024 | Global to multi-scale local architecture with hardwired CNN for 1-ms tomato defect detectionabstractAbstract A 1 millisecond (1‐ms) vision system that guarantees high efficiency and timely response for tomato defect detection is essential for factory automation. Because of various defect appearances, recently many existing researches focus on CNN based defect detection, but few of them attempt to reach high processing speed to adapt to the factorial assembly line. This paper proposes a global to multi‐scale local based parallel architecture with hardwired CNN for tomato defect detection. This architecture breaks down image‐wise detection into pixel‐wise localization and block‐wise classification. The pixel‐wise localization utilizes tomato‐aware information as constraints for localization performance. The block‐wise classification uses a fully pipelined network structure to obtain the classification result for each block as the pixel stream moves through the network. The classification network has a six‐layer lightweight network structure with quantization for hardwired type implementation on FPGA. The experiment results show that the proposed architecture processes 1000 FPS images with 0.9476 ms/frame delay. And for detection performance, this architecture keeps at 80.18%, only 1.31% lower than ResNet50 based detection system. Yuan Li 0058, Ryuji Fuchikami, Takeshi Ikenaga |
IET Image Process. | 4 |
| 2024 | Motion-aware and data-independent model based multi-view 3D pose refinement for volleyball spike analysisabstractAbstract In the volleyball game, estimating the 3D pose of the spiker is very valuable for training and analysis, because the spiker’s technique level determines the scoring or not of a round. The development of computer vision provides the possibility for the acquisition of the 3D pose. Most conventional pose estimation works are data-dependent methods, which mainly focus on reaching a high level on the dataset with the controllable scene, but fail to get good results in the wild real volleyball competition scene because of the lack of large labelled data, abnormal pose, occlusion and overlap. To refine the inaccurate estimated pose, this paper proposes a motion-aware and data-independent method based on a calibrated multi-camera system for a real volleyball competition scene. The proposed methods consist of three key components: 1) By utilizing the relationship of multi-views, an irrelevant projection based potential joint restore approach is proposed, which refines the wrong pose of one view with the other three views projected information to reduce the influence of occlusion and overlap. 2) Instead of training with a large amount labelled data, the proposed motion-aware method utilizes the similarity of specific motion in sports to achieve construct a spike model. Based on the spike model, joint and trajectory matching is proposed for coarse refinement. 3) To finely refine, a point distribution based posterior decision network is proposed. While expanding the receptive field, the pose estimation task is decomposed into a classification decision problem, which greatly avoids the dependence on a large amount of labelled data. The experimental dataset videos with four synchronous camera views are from a real game, the Game of 2014 Japan Inter High School of Men Volleyball. The experiment result achieves 76.25%, 81.89%, and 86.13% success rate at the 30mm, 50mm, and 70mm error range, respectively. Since the proposed refinement framework is based on a real volleyball competition, it is expected to be applied in the volleyball analysis. Xina Cheng, Takeshi Ikenaga |
Multim. Tools Appl. | 3 |
| 2024 | Key points trajectory and multi-level depth distinction based refinement for video mirror and glass segmentationabstractAbstract Mirror and glass are ubiquitous materials in the 3D indoor living environment. However, the existing vision system always tends to neglect or misdiagnose them since they always perform the special visual feature of reflectivity or transparency, which causes severe consequences, i.e., a robot or drone may crash into a glass wall or be wrongly positioned by the reflections in mirrors, or wireless signals with high frequency may be influenced by these high-reflective materials. The exploration of segmenting mirrors and glass in static images has garnered notable research interest in recent years. However, accurately segmenting mirrors and glass within dynamic scenes remains a formidable challenge, primarily due to the lack of a high-quality dataset and effective methodologies. To accurately segment the mirror and glass regions in videos, this paper proposes key points trajectory and multi-level depth distinction to improve the segmentation quality of mirror and glass regions that are generated by any existing segmentation model. Firstly, key points trajectory is used to extract the special motion feature of reflection in the mirror and glass region. And the distinction in trajectory is used to remove wrong segmentation. Secondly, a multi-level depth map is generated for region and edge segmentation which contributes to the accuracy improvement. Further, an original dataset for video mirror and glass segmentation (MAGD) is constructed, which contains 9,960 images from 36 videos with corresponding manually annotated masks. Extensive experiments demonstrate that the proposed method consistently reduces the segmentation errors generated from various state-of-the-art models and reach the highest successful rate at 0.969, mIoU (mean Intersection over Union) at 0.852, and mPA (mean Pixel Accuracy) at 0.950, which is around 40% - 50% higher on average on an original video mirror and glass dataset. Xina Cheng, Takeshi Ikenaga |
Multim. Tools Appl. | 4 |
| 2024 | Semi-supervised attention based merging network with hybrid dilated convolution module for few-shot HDR video reconstruction
Fengshan Zhao, Qin Liu 0002, Takeshi Ikenaga |
Multim. Tools Appl. | 3 |
| 2024 | Kinematics-aware spatial-temporal feature transform for 3D human pose estimation
Songlin Du, Zhiwei Yuan, Takeshi Ikenaga |
Pattern Recognit. | 3 |
| 2024 | JoyPose: Jointly learning evolutionary data augmentation and anatomy-aware global-local representation for 3D human pose estimation
Songlin Du, Zhiwei Yuan, Peifu Lai, Takeshi Ikenaga |
Pattern Recognit. | 4 |
| 2024 | AnatPose: Bidirectionally learning anatomy-aware heatmaps for human pose estimation
Songlin Du, Takeshi Ikenaga |
Pattern Recognit. | 3 |
| 2024 | Bi-Pose: Bidirectional 2D-3D Transformation for Human Pose Estimation From a Monocular CameraabstractAutomatically estimating 3D human poses in video and inferring their meanings play an essential role in many human-centered automation systems. Existing researches made remarkable progresses by first estimating 2D human joints in video and then reconstructing 3D human pose from the 2D joints. However, mono-directionally reconstructing 3D pose from 2D joints ignores the interaction between information in 3D space and 2D space, losses rich information of original video, therefore limits the ceiling of estimation accuracy. To this end, this paper proposes a bidirectional 2D-3D transformation framework that bidirectionally exchanges 2D and 3D information and utilizes video information to estimate an offset for refining 3D human pose. In addition, a bone-length stability loss is utilized for the purpose of exploring human body structure to make the estimated 3D pose more natural and to further increase the overall accuracy. By evaluation, estimation error of the proposed method, measured by the mean per joint position error (MPJPE), is only 46.5 mm, which is much lower than state-of-the-art methods under the same experimental condition. The improvement on accuracy will make machines to better understand human poses for building superior human-centered automation systems.Note to Practitioners—This paper was motivated by the demand of human-centered automation systems needing to accurately understand human poses. Existing approaches mainly focus on inferring 3D human pose from 2D joints mono-directionally. Although they made remarkable contributions to estimating 3D human pose in such a mono-directional way, we found that they ignore the 2D-3D interaction and do not use original video when inferring 3D pose from 2D joints. This paper therefore suggests a bidirectional 2D-3D transformation that exchanges 2D and 3D information and utilizes video information to estimate more accurate 3D human pose for human-centered automation systems. This work is a pioneering attempt of interactively using 2D and 3D information for more accurate estimation of human pose. Benefited from the state-of-the-art accuracy, the proposed approach is expected to make significant contributions to many human-centered automation systems, such as human-machine interaction, biomimetic manipulation, and automatic surveillance systems. Songlin Du, Zhiwei Yuan, Takeshi Ikenaga |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2023 | Complementary Data Transmission Control With Collision Avoidance for Efficient Retention of Large-size Spatio-Temporal Data*abstractSpatio-temporal aware IoT applications need to deliver the Spatio-Temporal Data (STD), which depends on geographic location and time, to the users in real-time. We have proposed the STD retention system using crowds of vehicles to promote the immediate distribution and utilization of STDs in geographical proximity to the user. However, when the system retains a Large-size STD (LSTD) consisting of multiple packets such as videos and photos, the packet loss due to channel collisions should be suppressed for effective retention of an LSTD. In this paper, we propose a complementary data transmission control method with collision avoidance for LSTD retention and clarify the effectiveness of our proposed method by simulation. Hotaka Kaneyasu, Daiki Nobayashi, Kazuya Tsukamoto, Takeshi Ikenaga, Myung J. Lee |
CCNC | 4 |
| 2023 | Performance Evaluation on the Impact of Bottleneck Link Buffer Size Under Different Congestion Control Algorithm Competition in QUICabstractQUIC is a secure, general-purpose transport protocol using congestion and retransmission controls on User Datagram Protocol (UDP). QUIC can use the same congestion control algorithms as Transmission Control Protocol (TCP). Thus, when communication between multiple terminals using QUIC competes, the performance of each terminal depends on the congestion control algorithm adopted by the terminal. In related works, the communication performance was evaluated by simulation when both CUBIC and BBR flows competed in TCP, and the results showed that each communication performance became unfair with the buffer size of the shared bottleneck link. In this study, we verify the communication performance at different bottleneck link buffer sizes when BBR and CUBIC compete through actual experiments. Nobuhiro Uchida, Daiki Nobayashi, Takeshi Ikenaga, Dirceu Cavendish |
CCNC | 3 |
| 2023 | A Reliability Audit Mechanism based on Multi-layered Blockchain for Spatio-Temporal Data Retention SystemabstractIoT data includes Spatio-temporal data (STD) that is needed only at a specific time and place. In our previous research, we have proposed the STD data retention system (STD-RS), aiming to construct a novel architecture for STD distributing in a retention area by vehicles. However, since vehicles may distribute the STD in unexpected locations due to GPS errors or malicious behavior in a real environment. Therefore, we propose a reliability audit mechanism based on multi-layered blockchain managing and analyzing the history of STD distribution and vehicles' behavior in each retention area. Through simulation experiment, we demonstrated that our mechanism detects the malfunction of STD-RS effectively. Junki Ueda, Kazuya Tsukamoto, Hiroshi Yamamoto, Daiki Nobayashi, Takeshi Ikenaga, Myung J. Lee |
CCNC | 5 |
| 2023 | Experimental Evaluation of Transmission Control Method based on Received Signal Strength for Spatio-Temporal Data Retention ∗abstractWith the development and spread of IoT technology, the number of devices connected to the Internet is increasing. Some data generated by IoT devices include spatio-temporal data (STD) that depends on the location and time of data generation. Therefore, we have proposed the STD retention system (STD-RS) using vehicles as a network infrastructure for local production and consumption of STD. This paper proposes a transmission control method that can mitigate the fluctuation of RSS due to the influence of obstacles in the real environment and evaluates its effectiveness through experiments on actual devices. Renju Akashi, Daiki Nobayashi, Kazuya Tsukamoto, Takeshi Ikenaga, Myung J. Lee |
COMPSAC | 4 |
| 2023 | Improvement of TCP Performance based on Characteristics of Private LoRa InterfaceabstractLow-power wide-area (LPWA) is a major communication technology used in the Internet of Things (IoT). However, since LPWA does not conform to the TCP/IP protocol stack and employs its own unique protocol, sensors equipped with LPWA I/F face difficulty when connecting directly to the Internet. In a previous study, to extend general TCP/IP communication on LPWA networks, we constructed a TCP/IP network over Private LoRa, but it was found that TCP communication performance using LoRa devices was significantly degraded due to the characteristics of LPWA such as the low transmission rate. In the present study, to improve the performance of TCP communication using the Private LoRa interface, we propose a transmission control method that avoids collisions by modifying the advertised window size on IP2LoRa and demonstrate the effectiveness of the proposed method using TCP flow analysis. Jumpei Sakamoto, Daiki Nobayashi, Kazuya Tsukamoto, Takeshi Ikenaga, Goshi Sato, Kenichi Takizawa |
COMPSAC | 4 |
| 2023 | Pyramid Spatial Feature Transform and Shared-Offsets Deformable Alignment Based Convolutional Network for HDR ImagingabstractTo generate ghost-free high dynamic range (HDR) images by merging multiple differently exposed low dynamic range (LDR) images, the key is to handle ill-exposed areas in the input LDR images and misalignment among them. In this paper, a Pyramid Spatial Feature Transform and shared-offsets Deformable convolutional Network (PSFTDNet) is proposed to achieve this target. The pyramid spatial feature transform module tackles ill-exposed areas, which modulates the features to exploit complementary information of them in a coarse-to-fine manner. The shared-offsets deformable alignment handles misalignment among the input images, which applies the offsets used to align exposure-aligned images to align all the features. Experiments on the NTIRE HDR challenge dataset and Kalantari dataset show that the proposed PSFTDNet outperforms all the conventional methods with PSNR-L scores of 42.31 dB and 41.54 dB, and PSNR-T scores of 34.78 dB and 43.56 dB. Junda Liao, Qin Liu 0002, Takeshi Ikenaga |
ICASSP | 3 |
| 2023 | HDR-LMDA: A Local Area-Based Mixed Data Augmentation Method for Hdr Video ReconstructionabstractMainstream image manipulation-based data augmentation methods (e.g., CutMix) undermine the integrity of extracted features, which leads to limited effect for pixel-level image processing tasks. In this paper, a local area-based mixed data augmentation method called HDR-LMDA for HDR video reconstruction is proposed. Within it, the local exposure augmentation (LEA) applies different exposures to original LDR inputs among different regions, while the local RGB permutation (LRP) shuffles the color channels of the random patch instead of the entire frame. By mixing both two operations, the model is forced to learn how to apply concise ill-exposure recovery and color processing within the same training process. Experiments demonstrate that HDR-LMDA achieves a better PSNR-T boost of 0.93dB, compared with conventional works under the same conditions. Fengshan Zhao, Qin Liu 0002, Takeshi Ikenaga |
ICIP | 3 |
| 2023 | Poster: Fairness Improvement Method Using ECNs with Different Congestion Control Algorithms within QUICabstractQUIC is a transport protocol that adds congestion control, retransmission control, and TLS to UDP. QUIC can use the same congestion control algorithms as TCP. In previous work, when both CUBIC and BBR flows compete within QUIC, we have shown that communication performance becomes unfair with respect to the buffer size of the shared bottleneck link. In this study, we improve the communication performance at different bottleneck link buffer sizes using Round Trip Time (RTT) and Explicit Congestion Notification (ECN) when CUBIC and BBR compete within QUIC through actual experiments. Nobuhiro Uchida, Daiki Nobayashi, Dirceu Cavendish, Takeshi Ikenaga |
ICNP | 4 |
| 2023 | A Figure Skating Jumping Dataset for Replay-Guided Action Quality AssessmentabstractIn competitive sports, judges often scrutinize replay videos from multiple views to adjudicate uncertain or contentious actions, and ultimately ascertain the definitive score. Most existing action quality assessment methods regress from a single video or a pairwise exemplar and input videos, which are limited by the viewpoint and zoom scale of videos. To end this, we construct a Replay Figure Skating Jumping dataset (RFSJ), containing additional view information provided by the post-match replay video and fine-grained annotations. We also propose a Replay-Guided approach for action quality assessment, learned by a Triple-Stream Contrastive Transformer and a Temporal Concentration Module. Specifically, besides the pairwise input and exemplar, we contrast the input and its replay by an extra contrastive module. Then the consistency of scores guides the model to learn features of the same action under different views and zoom scales. In addition, based on the fact that errors or highlight moments of athletes are crucial factors affecting scoring, these moments are concentrated in parts of the video rather than a uniform distribution. The proposed temporal concentration module encourages the model to concentrate on these features, then cooperates with the contrastive regression module to obtain an effective scoring mechanism. Extensive experiments demonstrate that our method achieves Spearman's Rank Correlation of 0.9346 on the proposed RFSJ dataset, improving over the existing state-of-the-art methods. Xina Cheng, Takeshi Ikenaga |
ACM Multimedia | 3 |
| 2023 | Straight-Line Detection Within 1 Millisecond Per Frame for Ultrahigh-Speed Industrial AutomationabstractDetecting straight lines in video plays a fundamental role in camera-based industrial automation. With the increasing demands on production efficiency, detection speed has become one of the bottlenecks for highly efficient industrial automation. Because of data dependence and hardware limitations, existing vision systems based on central processing unit/graphics processing unit are unable to detect straight lines at an ultrahigh speed. This article addresses this problem and proposes a hardware-friendly Hough transform that can be implemented in fully parallel for the ultrahigh-speed detection, because of the following two key features: it processes multiple pixels in parallel and directly calculates line parameters while capturing the current frame; and it simultaneously initializes the Hough parameter space and votes in the Hough parameter space without any delay. Based on the proposed hardware-friendly Hough transform, its chip-level implementation and system-level hardware design are presented. Experimental results show that the main benefits of the proposed architecture are in real-time performances at a high frame rate (784 frames/s) and an ultralow delay (0.7749 ms/frame). Songlin Du, Ziwei Dong, Takeshi Ikenaga |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | JointFusionNet: Parallel Learning Human Structural Local and Global Joint Features for 3D Human Pose Estimation
Zhiwei Yuan, Yaping Yan, Songlin Du, Takeshi Ikenaga |
ICANN (4) | 4 |
| 2022 | Poster: Implementation and Performance Evaluation of TCP/IP Communication over Private LoRaabstractLow-power wide-area (LPWA) is a major communication technology used in the Internet of Things (IoT). However, since LPWA does not conform to the TCP/IP protocol stack and employs its own unique protocol, it is difficult for sensors equipped with LPWA I/F to connect to the Internet directly. In the present study, to realize general TCP/IP communication on LPWA networks, we construct a TCP/IP network over Private LoRa, and evaluate its communication performance. Moreover, we verify the feasibility of TCP/IP communication over LoRa by using a popular Internet application. Jumpei Sakamoto, Daiki Nobayashi, Kazuya Tsukamoto, Takeshi Ikenaga, Goshi Sato, Kenichi Takizawa |
ICNP | 4 |
| 2022 | Attention-guided network with inverse tone-mapping guided up-sampling for HDR imaging of dynamic scenes
Yipeng Deng, Qin Liu 0002, Takeshi Ikenaga |
Multim. Tools Appl. | 3 |
| 2022 | Automatic Foreground Detection at 784 FPS for Ultra-High-Speed Human-Machine InteractionsabstractHuman-machine interactive systems show increasing demand for analysing fast moving objects in high-frame-rate videos. Robust foreground detection, which is able to reduce large amount of redundant background data from high-frame-rate video, becomes the essence to achieve ultra-high-speed human-machine interactions. This paper proposes a local spatial propagation based background model generation, a local linear illumination correction based background model update, and a regional central coordinates and edge keypoints constrained foreground region reselection. The three proposals make up a robust and hardware-friendly foreground detection method. Experimental results prove that the proposed hardware-friendly algorithm achieves high accuracy and robustness on various kinds of challenging cases. Meanwhile, the hardware implementation utilizes little hardware resources and achieves realtime processing of high-frame-rate (784 frame/second) video with the delay less than 1 ms/frame in image processing core. In addition, a practical system is implemented by combing a PC, a high-speed camera and a field programmable gate array (FPGA) for realworld applications. This work will significatively promote the development and application of high-speed human machine interaction. A demo of the proposed vision system working at 784 FPS is available athttps://wcms.waseda.jp/em/5f84f75136a6. Note to Practitioners—This paper was motivated by the problem of high-frame-rate video contains large amount of redundant background pixels which makes ultra-high-speed human-machine interactions inaccessible. Existing approaches are mainly focused on designing complex background models, but processing speed, which is the most important issue for ultra-high-speed human-machine interactions, has received relatively little attention. This paper suggests a robust and hardware-friendly foreground detection algorithm which has been implemented as a hardware system by using an FPGA, a high-frame-rate camera, and a PC. We show that the hardware implementation utilizes less hardware resources and achieves real-time processing speed of 784 FPS with the delay less than 1 ms/frame in the image processing core. This work is a pioneering attempt of ultra-high-speed foreground detection, which will significatively speed up the wide applications of ultra-high-speed human machine interactions. Songlin Du, Peikun Cai, Takeshi Ikenaga |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2022 | Highly-Parallel Hardwired Deep Convolutional Neural Network for 1-ms Dual-Hand Trackingabstract1-ms vision systems represent an extreme case of temporal development in video sensing techniques. Moreover, a 1-ms dual-hand tracking system leverages the dexterous functionality of hands and thus serves as a seamless and intuitive interface for Human-Computer Interaction. Deep CNN is promising for high tracking robustness, however, neither GPU-based nor FPGA-based implementation addresses the tracking task with ultra-high-speed. This paper proposes: (a) A paradigm to directly map a deep CNN as a hardwired circuit, so the entire network runs in parallel and high processing speed is obtained. The network is exempted from memory access since all intermediate neural values are implicitly represented in hardware states. And condensed binarization is used to reduce resource utilization; (b) Hardware design of the hardwired network on FPGA, inside which kernel-adapted convolutional trees are devised to maximize the parallelism. The speed bottleneck of the network is therefore removed by implementing convolutional layers as fine-grained pipelines with unified components; (c) FPGA-GPU hetero complementation, which utilizes an auxiliary GPU network to compensate for accuracy of the FPGA network without affecting its speed. The quick primary results on FPGA are intermittently refined using delayed but accurate hints from GPU. Implementation results show that the proposed method reaches 973fps and consumes merely 1.30ms to process on$640\times 480$images, while the accuracy is only 4.7% lower compared with the general method on test sequences. Video demonstrations are available athttps://wcms.waseda.jp/em/5f9d020f136e7. Peiqi Zhang, Dingli Luo, Songlin Du, Takeshi Ikenaga |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Adaptive Data Transmission Control for Spatio-Temporal Data Retention Over Crowds of VehiclesabstractSome specific services for Internet of Things, such as real-time map and providing local weather information, depend strongly on geographical time and location. We refer to the data for such service as spatio-temporal data (STD). When STD is used in a query response system similar to conventional Internet services, users not only need to acquire data actively as required, they must also have functions for retrieving data available STD. Therefore, we propose an STD retention system that uses vehicles as information hubs (InfoHubs) for disseminating and retaining the data in a specific area. In our system, InfoHubs diffuse, maintain, and advertise STD over places and times where the STD are strongly dependent, thereby allowing users to receive such data passively within the specific area. Additionally, because STD are associated with a particular space, the system can reduce search costs. We also propose an adaptive transmission control method that each vehicle effectively operates its wireless resources autonomously and STD are retained and distributed efficiently. Finally, we evaluated our proposed method using simulations and clarified that our proposed system is capable of achieving a coverage rate of nearly 100% for STD while reducing the number of data transmissions compared to existing systems. Daiki Nobayashi, Ichiro Goto, Hiroki Teshiba, Kazuya Tsukamoto, Takeshi Ikenaga, Mario Gerla |
IEEE Trans. Mob. Comput. | 5 |
| 2021 | Texture and exposure awareness based refill for HDRI reconstruction of saturated and occluded areasabstractAbstract High‐dynamic‐range image (HDRI) displays scenes as vivid as the real scenes. HDRI can be reconstructed by fusing a set of bracketed‐exposure low‐dynamic‐range images (LDRI). For the reconstruction, many works succeed in removing the ghost artefacts caused by moving objects. The critical issue is reconstructing the areas which are saturated due to bad exposure and occluded due to motion with no ghost artefacts. To overcome this issue, this paper proposes texture and exposure awareness based refill. The proposed work first locates the saturated and occluded areas existing in input image set, then refills background textures or patches containing rough exposure and colour information into located areas. Proposed work can be integrated with multiple existing ghost removal works to improve the reconstruction result. Experimental results show that proposed work removes the ghost artefacts caused by saturated and occluded areas in subjective evaluation. For the objective evaluation, the proposed work improves the HDR‐VDP‐2 evaluation result for multiple conventional works by 1.33% on average. Yipeng Deng, Qin Liu 0002, Takeshi Ikenaga |
IET Image Process. | 4 |
| 2021 | STED-Net: Self-taught encoder-decoder network for unsupervised feature representation
Songlin Du, Takeshi Ikenaga |
Multim. Tools Appl. | 2 |
| 2021 | Multi-task neural network with physical constraint for real-time multi-person 3D pose estimation from monocular camera
Dingli Luo, Songlin Du, Takeshi Ikenaga |
Multim. Tools Appl. | 3 |
| 2020 | Resolution Irrelevant Encoding and Difficulty Balanced Loss Based Network Independent Supervision for Multi-Person Pose EstimationabstractSustainable efforts are made to improve the accuracy performance in multi-person pose estimation, but the current accuracy is still not enough for real-world applications. Besides, most improvement approaches are designed for special basement networks and ignore the speed performance, which results in limited applicability and low cost-performance. This paper proposes two network independent supervision: Resolution Irrelevant Encoding and Difficulty Balanced Loss. The proposed methods reorganize task representatives, the loss calculation method, and the loss punishment ratio in one-stage pose estimation frameworks to improve the joints' location accuracy with general applicability and high computational efficiency. Resolution Irrelevant Encoding fuses heatmaps and proposed inner block offsets to fix pixel-level joints positions without resolution limitations. To improve network training efficiency, Difficulty Balanced Loss adjusts loss weight in spatial and sequential aspects. On the MS COCO keypoints detection benchmark, the mAP of OpenPose trained with our proposals outperforms the OpenPose baseline over 4.9%. Dingli Luo, Songlin Du, Takeshi Ikenaga |
HSI | 4 |
| 2020 | Hetero Complementary Networks with Hard-Wired Condensing Binarization for High Frame Rate and Ultra-Low Delay Dual-Hand TrackingabstractHigh frame rate, ultra-low delay yet accurate hand tracking system provides a seamless and intuitive interface for Human Computer Interaction (HCI). Tracking multi-person's dual-hand from monocular RGB camera is challenging for hand's variant image feature. Although many CNN based trackers have been proposed on general hardware, they cannot address this challenge with ultra-high speed. This paper proposes: (A) Hetero complementary networks for ultra-high speed dual-hand tracking, where the quick primary result from an FPGA network is intermittently combined with delayed accurate result from a GPU network. (B) Hard-wired condensing binarization for ultrahigh speed network implementation on FPGA. The network is able to be directly mapped as hardware resource because complex computation is condensed into binary layers. The proposed method achieves 69.8% accuracy on test sequences, which is only 4.7% lower compared with the general method. Meanwhile, the estimated FPGA resource utilization is tremendously reduced to 54.7% on the target platform. This work shows the potential to track multi-person's dual-hand at millisecond-level speed. Peiqi Zhang, Dingli Luo, Songlin Du, Takeshi Ikenaga |
HSI | 4 |
| 2020 | Selective Kernel and Motion-emphasized Loss Based Attention-guided Network for HDR Imaging of Dynamic ScenesabstractGhost-like artifact caused by ill-exposed and motion areas is one of the most challenging problems in high dynamic range (HDR) image reconstruction. When the motion range is small, previous methods based on optical flow or patch-match can suppress ghost-like artifacts by first aligning input images before merging them. However, they are not robust enough and still produce artifacts for challenging scenes where large foreground motions exist. To this end, we propose a deep network with an attention module and motion-emphasized loss function to produce ghost-free HDR images. In the attention module, we use the channel and spatial attention to guide the network to emphasize important components such as motion and saturated areas automatically. To be robust to images with different resolutions and objects with distinct scales, we adopt the selective kernel network as the basic framework for channel attention. In addition to the attention module, the motion-emphasized loss function based on the motion and ill-exposed areas mask is designed to help the network reconstruct motion areas. Experiments on the public dataset indicate that the proposed SK-AHDRNet produces ghost-free results where detail in ill-exposed areas is well recovered. The proposed method scores 43.17 with PSNR metric and 61.02 with HDR-VDP-2 metric on test which outperforms all conventional works. According to quantitative and qualitative evaluations, the proposed method can achieve state-of-the-art performance. Yipeng Deng, Qin Liu 0002, Takeshi Ikenaga |
ICPR | 3 |
| 2020 | Local Spatio-Temporal Propagation Based Adaptive Model Generation and Update for High Frame Rate and Ultra-Low Delay Foreground DetectionabstractHigh frame rate and ultra-low delay matching system plays an increasingly important role in human-machine interactive applications, which demands better experience and higher accuracy. Foreground detection is an indispensable preprocessing step to make the system suitable for complex scenes. Although many foreground detection algorithms have been proposed, few can achieve high speed in hardware due to their high complexity or high consumption. Based on the foreground detection algorithm ViBe, this paper proposes a local spatio-temporal propagation based adaptive model generation and update strategy for high frame rate and ultra-low delay foreground detection. Our algorithm predicts whether a region is a foreground by setting up detecting points, thereby adaptively adjusting the number of pixels that needs to be modeled. Secondly, the local linear illumination correlation is used to update models, which makes the algorithm more robust to illumination changes. The evaluation results show that the proposed algorithm successfully achieves real-time processing on the field-programmable gate array (FPGA) at a resolution of 640×480 pixels, with a delay of 0.908ms/frame. Peikun Cai, Songlin Du, Takeshi Ikenaga |
RTCSA | 3 |
| 2019 | Study on Autonomous Outing Support Service for the Visually Impaired
Eiji Aoki, Shinji Otsuka, Takeshi Ikenaga, Hideaki Kawano, Masaaki Yatsuzuka |
CISIS | 3 |
| 2019 | Iterative Autoencoding and Clustering for Unsupervised Feature RepresentationabstractUnsupervised feature representation is a challenging problem in machine learning and computer vision. Since manual labels are unavailable for training, it is difficult to reduce the gap between learned features and image semantics. This paper proposes an iterative autoencoding and clustering approach, which consists of an autoencoding sub-network and a classification sub-network, for unsupervised feature representation. On one hand, the autoencoding sub-network maps images to features. On the other hand, using the features generated by the autoencoding sub-network, the classification sub-network maps the features to classes and estimates pseudo labels by clustering the features simultaneously. Through iterations between the feature representation and the pseudo-labels-supervised classification, the gap between features and image semantics is reduced. Experimental results on handwritten digits recognition and objects classification prove that the proposed approach achieves state-of-the-art performance compared with existing methods. Songlin Du, Takeshi Ikenaga |
ISCAS | 2 |
| 2019 | Low-dimensional superpixel descriptor and its application in visual correspondence estimation
Songlin Du, Takeshi Ikenaga |
Multim. Tools Appl. | 2 |
| 2018 | Study on Regional Transportation Linkage System that Enables Efficient and Safe Movement Utilizing LPWA
Eiji Aoki, Shinji Otsuka, Takeshi Ikenaga, Masaaki Yatsuzuka, Kazuhiro Tokiwa |
CISIS | 3 |
| 2018 | Spatio-Temporal Data Retention System with MEC for Local Production and ConsumptionabstractTo facilitate local production and consumption (LPAC) of spatio-temporal data (STD) generated by Internet of Things (IoT) devices, we propose a STD retention system that works in collaboration with Mobile Edge Computing (MEC) infrastructure. In this paper, we will introduce the architecture of our proposed system and discuss its contributions and challenges. Daiki Nobayashi, Kazuya Tsukamoto, Takeshi Ikenaga, Mario Gerla |
COMPSAC (1) | 3 |
| 2018 | Partial Descriptor Update and Isolated Point Avoidance Based Template Update for High Frame Rate and Ultra-Low Delay Deformation MatchingabstractHigh frame rate and ultra-low delay matching system plays an important role in various human-machine interactive applications, which demands better performance in matching deformable and out-of-plane rotating objects. Although many algorithms have been proposed for deformation tracking and matching, few of them are suitable for hardware implementation due to complicated operations and large time consumption. This paper proposes a hardware-oriented template update method for high frame rate and ultra-low delay deformation matching system. In the proposed method, the new template is generated in real time by partially updating the template descriptor and adding new keypoints simultaneously with the matching process in pixels, and incorrect boundary points are avoided when judged as isolated with distance-reachability to solve the problem of template drift. Evaluation results indicate that the proposed method successfully supports the real-time processing of the 784fps and 640×480 resolution system on field-programmable gate array (FPGA), with a delay of 0.808ms/frame, as well as achieves satisfactory deformation matching results in comparison with other general methods. Songlin Du, Takeshi Ikenaga |
ICPR | 4 |
| 2017 | Event state based particle filter for ball event detection in volleyball game analysisabstractThe ball state tracking and detection technology plays a significant role in volleyball game analysis for volleyball team supporting and tactics development. This paper proposes a ball event detection method to achieve high detection rate by solving challenges including: the great variety of event length, the large intra-class difference of one event and the influence caused by ball trajectories. Proposed state vector covers both the event type and the event period length so that the system model can transits various lengths of event period and predicts event types by volleyball game rules. The curve segmental observation model avoids the tracking error influence to evaluate the event period likelihood by referring neighbouring trajectories of the ball. And according to the standard of the ball event, the feature of the distance between the ball and specific court line are extracted to evaluate the ball event type in observation. At last a two-layer estimation method estimates the posterior state which is a joint probability distribution. Experiments of the proposed method implemented on 3D trajectories tracked from multi-view volleyball game videos shows the detection rate reaches 90.43%. Xina Cheng, Norikazu Ikoma, Masaaki Honda, Takeshi Ikenaga |
FUSION | 4 |
| 2017 | Visual salience and stack extension based ghost removal for high-dynamic-range imagingabstractHigh-dynamic-range imaging (HDRI) techniques are proposed to extend the dynamic range of captured images against sensor limitation. The key issue of multi-exposure fusion in HDRI is removing ghost artifacts caused by motion of moving objects and handheld cameras. This paper proposes a ghost-free HDRI algorithm based on visual salience and stack extension. To improve the accuracy of ghost areas detection, visual salience based bilateral motion detection is introduced to measure image differences. For exposure fusion, the proposed algorithm reduces brightness discontinuity and enhances details by stack extension, and rejects the information of ghost areas to avoid artifacts via fusion masks. Experiment results show that the proposed algorithm can remove ghost artifacts accurately for both static and handheld cameras, remain robust to scenes with complex motion and keep low complexity over recent advances including patch based method and rank minimization based method by 20.4% and 63.6% time savings on average. Qin Liu 0002, Takeshi Ikenaga |
ICIP | 3 |
| 2017 | Simultaneous physical and conceptual ball state estimation in volleyball game analysisabstractAutomatically extraction of accurate volleyball game data from game videos plays an important role in making contribution to game data analysis, TV broadcasting and performance evaluations. In this paper, a particle filter based physical and conceptual ball state estimation method is proposed to track the 3D ball trajectory and ball event simultaneously with high accuracy. The physical ball state includes 3D ball position and velocity. Besides the ball event, the conceptual state also includes flag of the external force on the ball. The system model is adaptive to this external force predicted through proposed spatial hitting points dense distribution. Observation of the external force is evaluated by hitting point likelihood, which uses not only the past tracked trajectory but also the image noise feature so that image noise is transferred into useful feature and unidirectional dependency on trajectory is avoided. Experimental results based on multi-view HDTV video sequences show the tracking success rate of ball state achieves 92.43%. Xina Cheng, Norikazu Ikoma, Masaaki Honda, Takeshi Ikenaga |
VCIP | 4 |
| 2016 | Anti-occlusion observation model and automatic recovery for multi-view ball tracking in sports analysisabstractThe 3D position of the ball plays a crucial role in professional sport analysis. In ball sports, tracking of ball's precise position accurately is highly required, whose performance is affected by inaccurate 3D coordinates and occlusion problem. In this paper, we propose anti-occlusion observation model and automatic recovery by 3D ball detection based on multiview videos to track the ball in 3D space. The anti-occlusion observation model evaluates each camera's image and eliminates the influence of the cameras in which the ball is occluded. The automatic recovery method detects the ball's 3D position by homography relation of the multi-video and generates a new distribution to initiate the tracker when tracking failure is detected. Experimental results based on the HDTV video sequences, which were captured by four cameras located at the corners of the court, show that the success rate of the 3D ball tracking achieves 99.14%. Xina Cheng, Masaaki Honda, Norikazu Ikoma, Takeshi Ikenaga |
ICASSP | 4 |
| 2015 | Radio-On-Demand Sensor and Actuator Networks (ROD-SAN): System Design and Field TrialabstractWireless sensor and actuator networks (WSANs) are required to achieve both energy-efficiency and low-latency in order to prolong the network lifetime while being able to quickly respond to intermittently-transmitted control commands. These two requirements are in general in a relationship of trade-off when each node operates with well-known duty-cycling modes: nodes need to make their radio interfaces (IFs) frequently active in order to promptly detect the communication requests from the other nodes. One approach to break this inherent trade-off, which has been actively studied in recent literature, is the introduction of wake-up receiver that is installed into each node and used only for detecting the communication requests. The radio IF in each node is woken up only when needed through a wake-up message received by the wake-up receiver. While the effectiveness of this type of on-demand WSANs has been shown in several studies by theoretical analysis and computer simulations, its implementation and large-scale experimental investigations are missing. Therefore, in this paper, we first design and implement radio-on-demand sensor and actuator networks (ROD-SAN) including all protocols to realize on-demand WSANs, from the lowest layer of wake-up signaling to the application layer offering the functionalities of information monitoring and networked control. Then, we show experimental results obtained through our field trial in which 20 nodes are deployed in an outdoor area with the scale of 450m X 200m. The numerical results provide us with practical insights on the effectiveness as well as limitations of on-demand WSANs. Hiroyuki Yomo, Kenichi Abe, Yuichiro Ezure, Tetsuya Ito, Akio Hasegawa, Takeshi Ikenaga |
GLOBECOM | 6 |
| 2015 | Deblocking strength prediction based CTU-level SAO category determination in HEVC encoderabstractHigh efficiency video coding (HEVC) is a video compression standard that outperforms the predecessor H.264/AVC by doubling the compression efficiency. To enhance the coding accuracy, HEVC adopts sample adaptive offset (SAO), which reduces the distortion of reconstructed pixels using classification based non-linear filtering. In the traditional coding tree unit (CTU) based VLSI encoder implementation, during the pixel classification stage, SAO cannot use the raw samples in the boundary of the current CTU because these pixels have not been processed by deblocking filter (DF). This paper proposes a category determination algorithm based on estimating the deblocking strengths on CTU boundaries and selectively adopting the promising samples in these areas during SAO classification. Compared with HEVC test mode (HM11.0), experimental results indicate that the proposed method achieves an average 0.15% BD-bitrate reduction (equivalent to 0.0084 dB increases in P-SNR). Gaoxing Chen, Zhenyu Pei, Zhenyu Liu 0001, Takeshi Ikenaga |
VCIP | 4 |
| 2015 | A flow-based detection method for stealthy dictionary attacks against Secure Shell
Akihiro Satoh, Yutaka Nakamura, Takeshi Ikenaga |
J. Inf. Secur. Appl. | 3 |
| 2014 | Expiration Timer Control Method for QoS-Aware Packet ChunkingabstractDevelopment of machine-to-machine communication technologies encourages an increasing amount of mobile traffic by smart phones, cellular phones, and various sensors. This traffic consists of small packets in order to minimize the influence of frame dropping in a data link layer. However, it is necessary to reduce the number of packets flowing into the core network to chunk packets, while satisfying the application requirement of users. In this paper, we propose the expiration timer control method for QoS-aware packet chunking. Our proposed scheme achieves a decrease in the number of packets flowing into the core network while satisfying the QoS of users by adjusting the waiting time for packet chunking depending on the condition of the core network. In this paper, we evaluate the proposed scheme by performing simulations and show that our proposed scheme achieves not only a decrease in packet flow into the core network but also the maintenance of QoS requirement for applications. Hiroki Yanaga, Daiki Nobayashi, Takeshi Ikenaga |
COMPSAC | 3 |
| 2014 | Reliable Transmission with Multipath and Redundancy for Wireless Mesh Networks
Wenze Shi, Takeshi Ikenaga, Daiki Nobayashi, Xinchun Yin, Yebin Xu |
ICA3PP (1) | 2 |
| 2014 | A Self-adaptive Reliable Packet Transmission Scheme for Wireless Mesh Networks
Wenze Shi, Takeshi Ikenaga, Daiki Nobayashi, Xinchun Yin |
ICA3PP (1) | 2 |
| 2014 | Linear adaptive search range model for uni-prediction and motion analysis for bi-prediction in HEVCabstractHigh Efficiency Video Coding (HEVC) is the up-to-date video coding standard. Compared to the predecessor H.264/AVC, HEVC can further reduce approximately 50% bit rate on average with the competing perceptual quality. On the other hand, experiment shows that HEVC requires more than 4 times computational complexity during the encoding procedure. In ours test, even using fast TZSearch, integer motion estimation (IME) still accounts for 20%-30% of encoding time. In this paper, we propose two adaptive search range (ASR) algorithms to address this problem in IME. First, we present an ASR algorithm based on linear adaptive search range model (LAM-ASR) for uni-prediction. This model considers the impacts of the motion consistency, PU size and the amplitude of motion vector predictor (MVP). In order to offer more flexibility, we introduce a scale factor to this model. Second, for bi-prediction, we propose another ASR algorithm based on motion analysis (MA-ASR), which assigns different search range to PU by making full use of the motion information obtained from uni-prediction. Experimental results show that when embedded into the fast TZSearch method of the reference software, the two proposed ASR algorithms can averagely save 42.0% of the IME time with 0.023dB BD-PSNR degradation or equally 0.7% BD-BR increase. Longshan Du, Zhenyu Liu 0001, Takeshi Ikenaga, Dongsheng Wang 0002 |
ICIP | 3 |
| 2013 | Redundancy control and duplicate ACK suppression methods for TCP with FECabstractPacket losses significantly degrade TCP performance in high latency networks. To improve TCP performance in such networks, we proposed two methods to suppress the return of duplicate ACKs and to control minimum redundancy because FEC technology cannot work effectively when simply applied to TCP operation. Simulation evaluations show that the proposed methods enable higher throughput than the conventional methods, especially in high latency environments. In our future work, we will consider a scheme to more appropriately determine the appropriate redundancy level for network conditions and to more effectively recover lost packets in a real environment, such as where burst packet losses occur. Yurino Sato, Hiroyuki Koga, Masayoshi Shimamura, Takeshi Ikenaga |
ICNP | 4 |
| 2013 | Fast HEVC intra mode decision using matching edge detector and kernel density estimation alike histogram generationabstractIntra coding algorithm in High Efficiency Video Coding employs up to 35 directional prediction modes. Upon the end of alleviating the intra encoding complexity, we proposed the candidate mode selection algorithm from analyzing the textures of the source image block. Considering the fine difference between the neighboring prediction directions, we devise the fix-point arithmetic based edge detector, which improves the direction detection accuracy as compared with the typical previous works while maintaining the low computational overhead. To improve the robustness of the edge direction statistics, we further introduce the conception of kernel density estimation into the histogram calculation. Our proposals is orthogonal to the published HEVC fast intra mode decision algorithms. Experimental results verified that, on average, the proposed methods reduced the encoding time by 25.21% in high efficiency mode, and 37.61% in low complexity mode, whereas the averaging BDPSNR losses are 0.0608dB and 0.0781dB, respectively.1. Zhenyu Liu 0001, Takeshi Ikenaga, Dongsheng Wang 0002 |
ISCAS | 3 |
| 2013 | A mode-mapping and optimized MV conjunction based MGS-scalable SVC to AVC IPPP transcoderabstractScalable Video Coding (SVC) is an extension of H.264/AVC, aiming to provide the ability to adapt to heterogeneous environments. It offers great flexibility for bitstream adaptation in multi-point applications such as videoconferencing. However, transcoding between SVC and AVC is necessary due to the existence of legacy AVC-based systems. This paper proposes a 3-stage fast SVC-to-AVC transcoder for medium-grain quality scalability (MGS). Hierarchical-P structured SVC bitstream is transcoded into IPPP structured AVC bitstream with multiple reference frames. In the first stage, mode decision is accelerated by proposed SVC-to-AVC mode mapping scheme. In the second stage, INTER motion estimation is accelerated by an optimized motion vector (MV) conjunction method to predict the MV with a reduced search range. In the last stage, Hadamard-based all zero block (AZB) detection is utilized for early termination. Simulation results show that proposed transcoder achieves very similar coding efficiency as the optimal result, but with averagely 92.3% computational time saving. Lei Sun 0005, Zhenyu Liu 0001, Takeshi Ikenaga |
ISCAS | 3 |
| 2013 | A Low-Complexity Quantization-Domain H.264/SVC to H.264/AVC Transcoder with Medium-Grain Quality Scalability
Lei Sun 0005, Zhenyu Liu 0001, Takeshi Ikenaga |
MMM (1) | 3 |
| 2012 | Lagrangian Multiplier Optimization Using Markov Chain Based Rate and Piecewise Approximated Distortion ModelsabstractThe traditional Lagrangian RDO algorithm assumes the transformed residues as memo- ryless random variables, and then doesn't perform well when the prediction residues posses the strong temporal correlations. We extend the RDO by modeling the residues as the first-order Markov source and calibrating the distortion model with the piecewise approximation function. Zhenyu Liu 0001, Dongsheng Wang 0002, Takeshi Ikenaga |
DCC | 4 |
| 2012 | Lagrangian multiplier optimization using correlations in residuesabstractRate distortion optimization (RDO) algorithm plays the vital role in the up to date hybrid video codec H.264/AVC. The RDO algorithm of H.264/AVC reference software is built up by assuming that the transformed residues are memoryless variables. However, our experiments reveal that, for some sequences, the strong temporal correlations exist in the prediction residues. This paper extends the Lagrangian optimization techniques by modeling the transformed residues as the first-order Markov source and calibrating the distortion model with the piecewise approximation function. The proposed algorithms adjust the Lagrangian multiplier dynamically to improve the overall coding quality. Comprehensive experiments testify that, as compared with the JM reference software, our optimizations can achieve up to 1.875dB coding gain. Moreover, our algorithms posses more robust coding performance and introduce less computational overhead than the Laplace distribution based methods. The inherent short process latency makes it possible to cooperate our algorithms with rate control operation. Last but not least, the proposed approach is also useful for the emerging standard, HEVC. Zhenyu Liu 0001, Dongsheng Wang 0002, Takeshi Ikenaga |
ICASSP | 4 |
| 2012 | A pixel-domain mode-mapping based SVC-to-AVC transcoder with coarse grain quality scalability
Lei Sun 0005, Zhenyu Liu 0001, Takeshi Ikenaga |
ICPR | 3 |
| 2012 | Autonomous dynamic transmission scheduling based on neighbor node behavior for multihop wireless networksabstractMultihop wireless networks (MWNs) can dynamically extend the coverage area of wireless local area networks (LANs). Communication performance of MWNs that use a single channel decreases due to interference between nodes occurring when frames are being relayed. In this paper, we proposed an autonomous dynamic transmission scheduling scheme for MWNs that is based on the behavior of neighbor nodes. In our proposed scheme, each node has a transmission priority obtained by monitoring neighboring nodes. Because it is difficult for low priority node to obtain the right to transmit a frame, the proposed scheme can mitigate interferences among nodes. In this paper, we evaluate the proposed scheme by performing simulations, and show that it can improve performance in MWN communications. Daiki Nobayashi, Yutaka Fukuda, Takeshi Ikenaga |
LCN | 3 |
| 2011 | Adaptive network services on advanced relay nodesabstractThe explosive growth of usage with diversifying communication technologies and applications imposes the Internet to manage further scalability and diversity, which requires more adaptive and flexible sharing schemes of network resource. Especially when a number of distributed applications concurrently share the resource, efficacy on comprehensive usage of network, computation, and storage resources is needed to achieve a good information processing performance as a whole. In response to this problem, we have proposed a concept of adaptive network services enabled by advanced relay nodes. In this paper, we demonstrate the effectiveness of advanced relay processing, rather than simple IP forwarding, using a prototype implementation of advanced relay nodes. Masayoshi Shimamura, Takeshi Ikenaga, Masato Tsuru 0001 |
CCNC | 2 |
| 2011 | Adaptive fast DIRECT mode decision algorithm using mode and Lagrangian cost prediction for B frame in H.264/AVCabstractIn this paper, a fast spatial DIRECT mode decision method for B frame in H.264/AVC is proposed. It is based on a statistical analysis on multiple video sequences, and the strong relationship of mode selection and rate-distortion (RD) cost between the current DIRECT macroblock (MB) and the co-located MBs is observed. With the check of mode condition and adaptive threshold of RD cost, the complex mode decision process can be released at an early stage even for small QP cases. Simulation results demonstrate the proposed method can achieve much better performance than the original exhaustive rate-distortion optimization (RDO) based mode decision algorithm by reducing up to 57.1 % of motion estimation (ME) time for IBPBP picture group with only negligible bit increment and quality degradation. Xiaocong Jin, Jun Sun 0005, Jun Zhou 0007, Yiqing Huang 0002, Takeshi Ikenaga |
ICME | 6 |
| 2011 | Coarse to fine adaptive interpolation filter for high resolution video codingabstractWith the increasing demand of high video quality and large image size, adaptive interpolation filter (AIF) addresses these issues and conquers the time varying effects resulting in increased coding efficiency, comparing with recent H.264 standard. However, currently most AIF algorithms are based on either frame level or macroblock (MB) level, which are not flexible enough for different video contents in a real codec system. And most of them are facing a severe time consuming problem. This paper proposes a content based coarse to fine AIF algorithm, which can adapt to video contents by adding different filters and conditions from coarse to fine. The overall algorithm has been mainly made up by 3 schemes: frequency analysis based frame level skip interpolation, motion vector modeling based region level interpolation, and edge detection based macroblock level interpolation. The experimental results show that the proposed algorithm is able to reduce total encoding time about 41% for 720p and 25% for 1080p sequences averagely, comparing with key technology areas (KTA) Enhanced AIF algorithm, while obtains a BD-PSNR gain up to 0.004 and 3.122 BDBR reduction. Yiqing Huang 0002, Lei Sun 0005, Shinichi Sakaida, Takeshi Ikenaga |
ICME | 5 |
| 2011 | A simple priority control mechanism for performance improvement of Mobile Ad-hoc NetworksabstractMobile Ad-hoc Networks (MANETs) involve the establishment of wireless connections between mobile nodes, and can dynamically extend the coverage area of wireless local area networks (LAN). For MANETs that use a single channel, the communication performance decreases due to interferences between nodes when frames are being relayed. In this paper, we propose a priority control scheme based on neighbor node behavior for MANETs. A node continuously monitors neighbor nodes and determines a transmission priority on its own using monitoring results; the priority decreases when the node transmits a frame successfully, and increases when the node detects a transmitted frame from neighbor nodes. We evaluated the effectiveness of our proposed scheme using simulations, and results show that the throughput of our proposed scheme is at least approximately 700 kb/s higher than that of a MANET that uses IEEE 802.11a. Daiki Nobayashi, Yutaka Fukuda, Takeshi Ikenaga |
LCN | 3 |
| 2011 | Multi Objective Optimization Based Fast Motion Detector
Xiaocong Jin, Takeshi Ikenaga |
MMM (1) | 4 |
| 2011 | Register Length Analysis and VLSI Optimization of VBS Hadamard Transform in H.264/AVCabstractFidelity range extensions of H.264/AVC adopt variable block size (VBS) transform techniques to employ 8 × 8/4 × 4 Hadamard transforms adaptively during the fractional motion estimation. In this literature, the hardwired VBS Hadamard transform accelerator is developed with the following contributions: 1) developed a hardware reusing scheme between 8 × 8 and 4 × 4 transforms within the architecture design; 2) devised the intermediate bit-truncation algorithm to reduce the hardware cost while maintaining the computational precision well; and 3) reduced the bit-width of sum of absolute transformed differences (SATD) value as compared to the primitive implementation, resulting in optimization in both power and hardware cost for the SATD generator implementation. With TSMC 0.18 μm CMOS technology, the experiments demonstrate that for each VBS Hadamard transform engine, 13.0-30.4% saving in hardware cost and 12.6-32.4% saving in power consumption are achieved, whereas the incurred coding quality loss is less than 0.2089 dB in terms of BDPSNR. From the aspect of the whole encoder implementation, and considering the parallelism in searching factional pixel candidates, the proposed strategies garner 2.0 3.9% overall gate count reduction. Zhenyu Liu 0001, Dongsheng Wang 0002, Takeshi Ikenaga |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Performance Evaluation of Multi-Rate Communication in Wireless LANsabstractIn IEEE 802.11a/b/g wireless LAN, STAs (stations) can select an appropriate transmission rate based upon the received signal strength in order to achieve high throughput. In such a multi-rate environment, however, the total throughput is degraded if an AP (access point) is shared by STAs at both high and low transmission rates simultaneously. This problem is identified as the performance anomaly problem in only IEEE 802.11b, and we closely examine it under the IEEE 802.11a multi-rate environment. Although earlier studies have assumed that this performance anomaly always occurs, we show that it does not occur under a certain environment. First, we describe the mechanism of the performance anomaly, then illustrate the condition under which the performance anomaly does not occur. Finally, we show, through simulation results, the condition under which a performance anomaly may and may not occur in a multirate environment. Fumie Miki, Daiki Nobayashi, Yutaka Fukuda, Takeshi Ikenaga |
CCNC | 4 |
| 2010 | A network reconfiguration scheme against misbehaving nodesabstractMulti-hop wireless networks (MWNs) are rapidly gaining attention, because they can provide a wide coverage area to Internet users. To provide stable and high-performance network environments, security issues must be addressed. This paper focuses on the security issues in terms of anomalous relay nodes inside networks because their malicious behaviors can degrade the performance of MWNs. To maintain the network performance of MWNs, we propose a novel network reconfiguration scheme that each node reconstructs a network autonomously using the I/F of the neighbor nodes linked to a misbehaving node. Our proposed scheme reconfigurates topology with an emphasis on the reuse of I/F, the number of the links required to construct, transmission rates, performance anomaly, and network connectivity. We evaluate the effectiveness of the proposed schemes by simulations. The results of simulation indicate that the proposed scheme can prevent the communication performance degradation of the entire MWN. Daiki Nobayashi, Takashi Sera, Takeshi Ikenaga, Yutaka Nakamura, Yoshiaki Hori |
LCN | 3 |
| 2010 | Adaptively Adjusted Gaussian Mixture Models for Surveillance Applications
Tianci Huang, Xiangzhong Fang, Jingbang Qiu, Takeshi Ikenaga |
MMM | 4 |
| 2010 | Fully Utilized and Low Design Effort Architecture for H.264/AVC Intra Predictor Generation
Yiqing Huang 0002, Qin Liu 0002, Takeshi Ikenaga |
MMM | 3 |
| 2009 | A 7-Round Parallel Hardware-Saving Accelerator for Gaussian and DoG Pyramid Construction Part of SIFT
Jingbang Qiu, Tianci Huang, Takeshi Ikenaga |
ACCV (3) | 3 |
| 2009 | A high performance LDPC decoder for IEEE802.11n standardabstractIn this paper, we propose a partially-parallel irregular LDPC decoder for IEEE 802.11n standard. The design is based on a novel sum-delta message passing schedule to achieve high throughput and low area cost design. We further improve the design with pipeline structure and parallel computation. The synthesis result in TSMC 0.18 CMOS technology demonstrates that for (648,324) irregular LDPC code, our decoder achieves 7.5X improvement in throughput, which reaches 402 Mbps at the frequency of 200MHz, with 11% area reduction. Yuta Abe, Takeshi Ikenaga, Satoshi Goto |
ASP-DAC | 3 |
| 2009 | Reconfigurable SAD tree architecture based on adaptive sub-sampling in HDTV applicationabstractIn H.264/AVC based integer motion estimation engine, fixed architectures based on full pixel or direct sub-sampling pattern are widely used for HDTV application. However, these architectures suffer from either high complexity or quality loss problems. In this paper, an adaptive sub-sampling based reconfigurable architecture is given out. Firstly, by executing pixel difference analysis, the adaptive sub-sampling scheme which uses three hardware friendly patterns is applied on homogeneous macroblock (MB). Secondly, the related architecture introduces one more pipeline stage to build up configurable partial SAD values so that system performance is enhanced. Thirdly, a two-level pixel data organization scheme is proposed to solve data reuse and hardware utilization problems caused by adaptive algorithm. Moreover, one cross based SAD generation structure is introduced to achieve adaptive output results with less hardware cost. Experimental results show that, the proposed architecture can averagely save 61.71% clock cycles and accomplish twice or four times processing capability for homogeneous MBs. The maximum clock frequency is 208MHz under the TSMC 0.18um technology in worst case conditions(1.62V, 125 C). Yiqing Huang 0002, Qin Liu 0002, Satoshi Goto, Takeshi Ikenaga |
ACM Great Lakes Symposium on VLSI | 4 |
| 2009 | Macroblock feature and motion involved multi-stage fast inter mode decision algorithm in H.264/AVC video codingabstractOne fast inter mode decision algorithm is proposed in this paper. The whole algorithm is convoluted with block matching process. Firstly, before ME process, by exploiting spatial and temporal information, a skip mode early detection algorithm is proposed. Also, in this stage, edge gradient is used to filter out unpromising modes. Secondly, during the ME stage, the original search window is separated into several layers and our fast decision scheme works with motion information of each layer. Moreover, before stepping into small modes, the distribution of SAD (sum-of-absolute-difference) and RD (rate distortion) costs of big modes are analyzed in an early stage to accelerate the inter mode decision process. Experiments show that our algorithm can achieve a speed-up factor of up to 66.0% with trivial bit increment and quality degradation. Yiqing Huang 0002, Qin Liu 0002, Takeshi Ikenaga |
ICIP | 3 |
| 2009 | Hardware optimizations of variable block size Hadamard transform for H.264/AVC FRExtabstractVariable block size (VBS) transform technique is adopted in Fidelity Range Extensions (FRExt) of H.264/AVC, in which 8 × 8/4 × 4 Hadamard transforms are adaptively employed during the fractional motion estimation. The hardwired VBS Hadamard transform unit is developed by authors and the following contributions are described in this literature: (1) Hardware reusing scheme is adopted in the architecture design; (2) In the light of the noise analysis, the intermediate data bit-truncation scheme is developed to reduce the hardware cost while maintaining its computational precision well; (3) With mathematical analysis, the bit-width of SATD value is reduced as compared to the intuitive implementation, therefore, the power and hardware cost are both optimized for the SATD generator implementation; (4) Hybrid 4:2/3:2 compressor based CSA tree is analyzed in the circuits design of SATD generator; and (5) Clock-gating technique is employed to reduce the power dissipation of 4×4 transform operation. With TSMC 0.18 ¿m CMOS technology, experimental results reveal that 12.2-30.4% saving in hardware cost and 12.4-32.4% saving in power consumption are achieved by using our algorithms. Zhenyu Liu 0001, Dongsheng Wang 0002, Takeshi Ikenaga |
ICIP | 3 |
| 2009 | Content aware configurable architecture for H.264/AVC integer motion estimation engineabstractIn this paper, we contribute a configurable SAD tree architecture based on adaptive subsampling scheme. Firstly, by further exploiting the spatial feature, the integer motion estimation process is greatly sped up. Secondly, the conventional partial sum of absolute difference (SAD) based pipeline structure is optimized into configurable SAD oriented way, which enhances the performance and solve the data reuse problem caused by adaptive scheme in the architecture level. Moreover, a cross reuse and compressor tree based circuit level optimization is introduced and 6.56% hardware cost is reduced. Experiments show that our design can averagely achieve 42.23% saving in processing cycles compared with previous design. With 323 k gates at about 144.8 MHz, our design can achieve real-time encoding of HDTV 1088 p@30 fps. Yiqing Huang 0002, Qin Liu 0002, Takeshi Ikenaga |
ICME | 3 |
| 2009 | A Fast Hybrid Decision Algorithm for H.264/AVC Intra Prediction Based on Entropy Theory
Guifen Tian, Tianruo Zhang, Takeshi Ikenaga, Satoshi Goto |
MMM | 3 |
| 2009 | Highly parallel fractional motion estimation engine for Super Hi-Vision 4k×4k@60fpsabstractOne Super Hi-Vision (SHV) 4k times 4k @60 fps fractional motion estimation (FME) engine is proposed in our paper. Firstly, two complexity reduction schemes are proposed in the algorithm level. By analyzing the integer motion cost, 48% clock cycle is saved based on our mode pre-filtering scheme. By further check the motion cost of neighboring search points, our directional one-pass scheme can achieve reduction of 50% clock cycle and 36% hardware cost. Secondly, in the hardware level, two parallel improved schemes namely 16-Pel interpolation and MB-parallel processing are given out. Thirdly, one unified pixel block loading scheme is proposed. About 28.67% to 80.68% pixels are reused and the related memory access is saved. Furthermore, one parity pixel organization scheme is proposed to solve memory access conflict of MB-parallel processing. By using TSMC 0.18 mum technology in worst work conditions (1.62 V, 125degC), our FME engine can achieve real-time processing for SHV 4ktimes4k@60fps with 976.5 k gates hardware. Yiqing Huang 0002, Qin Liu 0002, Takeshi Ikenaga |
MMSP | 3 |
| 2009 | On bit allocation and Lagrange Multiplier adjustment for rate-distortion optimized H.264 rate controlabstractThis paper presents an on bit allocation and Lagrange multiplier (lambdaMODE) adjustment for H.264 rate control. To enhance the precision of the complexity estimation and bit allocation, a frequency-domain parameter named mean-absolute-transform-difference (MATD) is adopted to represent frame and macroblock (MB) residual complexity. Second, the MATD ratio is utilized to enhance the accuracy of frame layer bit prediction. Then, by considering the bit usage status of whole sequence, a measurements combining forward and backward bit analysis is proposed to optimize the lambdaMODEon frame layer. On MB layer, bits are allocated by proposed remaining complexity analysis. The computed quantization parameter (QP) is further adjusted according to predicted MB texture bits. Simulation results verify the performance of the proposed algorithm. Compared with the recommended rate control in H.264 reference software JM13.2, the PSNR improvement is up to 1.13 dB by using our algorithm. Shuijiong Wu, Yiqing Huang 0002, Qin Liu 0002, Takeshi Ikenaga |
MMSP | 5 |
| 2009 | Motion Estimation Optimization for H.264/AVC Using Source Image Edge FeaturesabstractThe H.264/AVC coding standard processes variable block size motion-compensated prediction with multiple reference frames to achieve a pronounced improvement in compression efficiency. Accordingly, the computation of motion estimation increases in proportion to the product of the number of reference frame and the number of intermode. The mathematical analysis in this paper illustrates that the motion-compensated prediction errors are mainly determined by the detailed textures in the source image. The image block being rich in textures contains numerous high-frequency signals, which make variable block size and multiple reference frame techniques essential. On the basis of rate-distortion theory, in this paper, the spatial homogeneity of an image block is made as a relative concept with respect to the current quantization step. For the homogenous block, its futile reference frames and intermodes can be eliminated efficiently. It is further revealed that the sum of absolute differences value of an image block is mainly determined by the sum of its edge gradient amplitude and the current quantization step. Consequently, the image content-based early termination algorithm is proposed, and it outperforms the original method adopted by JVT reference software. Moreover, the dynamic search range algorithm based on the edge gradient amplitude of source image block is analyzed. One eminent advantage of the proposed edge-based algorithms is their efficiency to the macroblock-pipelining architecture, and another desirable feature is their orthogonality to fast block-matching algorithms. Experimental results show that when these algorithms are integrated with hybrid unsymmetrical-cross multi-hexagongrid search, an averaged 31.4-60.0% motion estimation time can be saved, whereas the averaging BDPSNR loss is 0.0497 dB for all tested sequences. Zhenyu Liu 0001, Satoshi Goto, Takeshi Ikenaga |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | A cost-efficient partially-parallel irregular LDPC decoder based on sum-delta message passing algorithmabstractA partially-parallel decoder architecture for irregular LDPC code targeting high throughput and low cost applications is proposed. The design is based on a novel sum-delta message passing algorithm that facilitates the decoding throughput by removing redundant computations and decreases the hardware cost by optimizing the storage. Techniques such as binary sorting, parallel column operation, high performance pipelining are used to further speed up the message passing procedure. The synthesis result in TSMC 0.18 CMOS technology demonstrates that for (648,324) irregular LDPC code, our decoder achieves 7.5X improvement in throughput, which reaches 402 Mbps at the frequency of 200MHz, with 11% area reduction. Yuta Abe, Takeshi Ikenaga, Satoshi Goto |
ACM Great Lakes Symposium on VLSI | 3 |
| 2008 | Optimization of Propagate Partial SAD and SAD tree motion estimation hardwired engine for H.264abstractVariable block size motion estimation algorithm is the effcient approach to reduce the temporal redundancies and it has been adopted by the latest video coding standard H.264/AVC. The computational complexity augment coming from the variable block size technique makes the hardwired accelerator essential, especially for real-time applications. In this paper, the authors apply the architecture level and the circuits level approaches to improve the performance of Propagate Partial SAD and SAD Tree hardwired engines, which outperform other counterparts when considering the impact of supporting the variable block size technique. Experiments demonstrate that by using the proposed approaches, compared with the original architectures, 14.7% and 18.0% hardware cost can be saved for Propagate Partial SAD architecture and SAD Tree architecture, respectively. With TSMC 0.18 mm 1P6M CMOS technology, the proposed Propagate Partial SAD architecture attains 231.6 MHz operating frequency at a cost of 84.1 k gates. Correspondingly, the execution speed of the optimized SAD Tree architecture is improved to 204.8 MHz with 88.5 k gate hardware overhead. Zhenyu Liu 0001, Satoshi Goto, Takeshi Ikenaga |
ICCD | 3 |
| 2008 | A motion vector difference based self-incremental adaptive search range algorithm for variable block size motion estimationabstractThe search range (SR) parameter plays an important role in motion estimation (ME) for video coding. Adaptively adjusting SR according to the information given by previously encoded syntax element, also known as adaptive search range (ASR) algorithm, can efficiently reduce the computational complexity of ME. Compared with heuristic search pattern (HSP) algorithms like diamond/hexagon search, ASR algorithms are more fundamental, flexible and hardware-oriented. This paper although starts with a comparison between HSP and ASP algorithms which is followed by a proposed ASR algorithm with experimental results, however more likely intends to contribute several novel perspectives to this research area. Zhenxing Chen, Qin Liu 0002, Takeshi Ikenaga, Satoshi Goto |
ICIP | 3 |
| 2008 | Fast motion estimation for H.264/AVC using image edge featuresabstractThe key to high performance in video coding lies on efficiently reducing the temporal redundancies. For this purpose, H.264/AVC coding standard has adopted variable block size motion estimation on multiple reference frames to improve the coding gain. However, the computational complexity of motion estimation is also increased in proportion to the product of the reference frame number and the inter mode number. The mathematical analysis reveals that the prediction errors mainly depend on the image edge gradient amplitude and quantization parameter. Consequently, this paper proposed the image content based early termination algorithm, which outperforms the method adopted by JVT reference software, especially at high and moderate bit rates. In light of rate-distortion theory, this paper also relates the homogeneity of image to the quantization parameter. For the homogenous block, its search computation for the futile reference frames and inter modes can be efficiently discarded. Then the computation saving performance increases with quantization parameter. These content based fast algorithms were integrated with Unsymmetrical-cross Multihexagon-grid Search (UMHexagonS) algorithm. Compared to the original UMHexagonS fast matching algorithm, 26.14-54.97% search time can be saved with an average of 0.0369dB coding quality degradation. Zhenyu Liu 0001, Satoshi Goto, Takeshi Ikenaga |
ICME | 3 |
| 2008 | Hardware-oriented direction-based fast fractional motion estimation algorithm in H.264/AVCabstractIn this paper, a hardware-oriented fast fractional motion estimation (FME) algorithm for H.264/AVC is proposed. For both 1/2-pixel and 1/4-pixel FME, only 4 points are searched, and thus the required processing unit (PU) number is reduced from 9 to 4. Experiments show that the proposed algorithm can provide very similar image quality as full search, and the average PSRN drop is 0.17dB or bitrate increase is 4.08%. Compared with the full search FME hardware architecture [5], the proposed algorithm can effectively reduce about 32.3% hardware cost and thus is very suitable for low-cost and low-power applications. Yang Song 0002, Zhenyu Liu 0001, Takeshi Ikenaga, Satoshi Goto |
ICME | 4 |
| 2008 | VLSI friendly computation reduction scheme in H.264/AVC motion estimationabstractIn H.264/AVC standard, motion estimation (ME) can be executed on multiple reference frame (MRF) to improve the coding performance. For real-time hardwired encoder, the huge throughput of fractional motion estimation (FME) and integer motion estimation (IME) makes pipeline stage a must. So, IME is arranged in a single stage, which deteriorates the efficiency of many fast ME algorithms. This paper provides a VLSI friendly complexity reduction solution for ME procedure. Firstly, the proposed algorithm examines the pixel difference of current macroblock (MB) and adjust the available reference frame number. Secondly, it executes matching analysis to detect MB with static feature and early terminate the IME process. Thirdly, based on motion feature analysis result, the search range for non static MB is also adjusted and redundant search positions are eliminated. Compared with full search algorithm, the proposed fast ME algorithm can reduce 47.91% to 91.88% ME time with negligible video quality degradation. Furthermore, the algorithm can also be combined with other fast block matching process and friendly to hardwired encoder. Yiqing Huang 0002, Satoshi Goto, Takeshi Ikenaga |
ISCAS | 3 |
| 2008 | Early detection algorithms for 8×8 all-zero blocks in H.264/AVCabstract8 times 8 transform has been introduced in H.264psilas high profile to improve the video quality. After transform and quantization, if all the coefficients of the blockpsilas residue data become zero, this block is called all-zero block (AZB). Many 4 times 4 transform AZB early detection algorithms have been proposed to skip transform and quantization process. In this paper, after theoretical analysis performed for the sufficient condition of 8 times 8 AZB detection, sum of absolute differences(SAD) and sum of absolute transformed(SATD) based 8 times 8 AZB detection algorithms are proposed. Experimental results show that the proposed algorithms achieve major improvement of computation reduction from 0.95% to 79.48% for 720p sequences. The computation reduction increases as QP increases. Qin Liu 0002, Yiqing Huang 0002, Takeshi Ikenaga |
MMSP | 3 |
| 2008 | Motion Feature and Hadamard Coefficient-Based Fast Multiple Reference Frame Motion Estimation for H.264abstractIn the state-of-the-art video coding standard, H.264/AVC, the encoder is allowed to search for its prediction signals among a large number of reference pictures that have been decoded and stored in the decoder to enhance its coding efficiency. Therefore, the computation complexity of the motion estimation (ME) increases linearly with the number of reference picture. Many fast multiple reference frame ME algorithms have been proposed, whose performance, however, will be considerably degraded in the hardwired encoder design due to the macroblock (MB) pipelining architecture. Considering the limitations of the traditional four-stage MB pipelining architecture, two fast multiple reference frame ME algorithms are proposed here. First, on the basis of mathematical analysis, which reveals that the efficiency of multiple reference frames will be degraded by the relative motion between the camera and the objects, for the slow-moving MB, the authors adopt the multiple reference frames but reduce their search range. On the other hand, for the fast-moving MB, the first previous reference frame is used with the full search range during the ME processing. The mutually exclusive feature between the large search range and the multiple reference frames makes the computation saving performance of the proposed algorithm insensitive to the nature of video sequence. Second, following the Hadamard transform coefficient-based all_zeros block early detection algorithm, two early termination criteria are proposed. These methods ensure the pronounced computation saving efficiency when the encoded video has strong spatial homogeneity or temporal stationarity. Experimental results show that 72.7%-93.7% computation can be saved by the proposed fast algorithms with an average of 0.0899 dB coding quality degradation. Moreover, these fast algorithms can be combined with fast block matching algorithms to further improve their speedup performance. Zhenyu Liu 0001, Yang Song 0002, Satoshi Goto, Takeshi Ikenaga |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2007 | A VLSI architecture design of an edge based fast intra prediction mode decision algorithm for H.264/avcabstractThe intra-frame coding in H.264/AVC has made significant contribution to the enhancement of coding efficiency. However it brings about a heavy computation burden in the rate distortion based (RD) mode decision (MD) process. Although the real-time encoding of 1280-720p signals is realized in recent works with existing algorithms, for higher resolution e.g. 1920-1088p some hardware-oriented fast algorithms are necessary. Yet so far few of the many proposed fast MD algorithms have seen successful hardware implementation. This paper presents a novel VLSI design (15.8k [email protected], with TSMC CMOS 0.18μm technology) of an edge based fast intra MD algorithm which can constantly reduce about 66% of the RD related computation with a negligible quality loss. It is expected to be utilized as a favorable accelerator hardware module in a real-time HDTV (1920-1088p) H.264 encoder or MPEG2-H.264 transcoder. Xianghui Wei, Takeshi Ikenaga, Satoshi Goto |
ACM Great Lakes Symposium on VLSI | 3 |
| 2007 | Hardware-efficient propagate partial sad architecture for variable block size motion estimation in H.264/AVCabstractOne hardware efficient and high speed architecture for variableblock size motion estimation in H.264 is presented in this paper. Through compressing the propagated data and optimizing theprocessing element and adder tree circuits in pipeline, this architecture gets more hardware efficient datapath logic. Compared with the original Propagate Partial SAD structure, 12.1% hardware cost can be saved. With TSMC 0.18μm CMOS 1P6M standard celllibrary, the maximum clock speed of this design is 227MHz in worstwork conditions (1.62V, 125°C). With the 48x32 search range, the maximum throughput of our design is 147786 MB/S, which can be used in the real-time encoding of VGA resolution frame with 4 reference frames at 30Hz. Zhenyu Liu 0001, Yiqing Huang 0002, Yang Song 0002, Satoshi Goto, Takeshi Ikenaga |
ACM Great Lakes Symposium on VLSI | 5 |
| 2007 | VLSI Oriented Fast Multiple Reference Frame Motion Estimation Algorithm for H.264/AVCabstractIn H.264/AVC standard, motion estimation can be processed on multiple reference frames (MRF) to improve the video coding performance. For the VLSI real-time encoder, the heavy computation of fractional motion estimation (FME) makes the integer motion estimation (IME) and FME must be scheduled in two macro block (MB) pipeline stages, which makes many fast MRF algorithms inefficient for the computation reduction. In this paper, two algorithms are provided to reduce the computation of FME and IME. First, through analyzing the block's Hadamard transform coefficients, all-zero case after quantization can be accurately detected. The FME processing in the remaining frames for the block, detected as all-zero one, can be eliminated. Second, because the fast motion object blurs its edges in image, the effect of MRF to aliasing is weakened. The first reference frame is enough for fast motion MBs and MRF is just processed on those slow motion MBs with a small search range. The computation of IME is also highly reduced with this algorithm. Experimental results show that 61.4%-76.7% computation can be saved with the similar coding quality as the reference software. Moreover, the provided fast algorithms can be combined with fast block matching algorithms to further improve the performance. Zhenyu Liu 0001, Yang Song 0002, Takeshi Ikenaga, Satoshi Goto |
ICME | 4 |
| 2007 | Ultra Low-Complexity Fast Variable Block Size Motion Estimation Algorithm in H.264/AVCabstractVariable block size motion estimation (VBSME) is the most computation consuming part in the newest H.264 /AVC video coding standard. To reduce computation, an ultra low-complexity fast VBSME algorithm is proposed in this paper. Different to previous fast algorithms which separately pressed the 7 block modes, in the presented algorithm, all the block modes are simultaneously calculated for most of the search points, and thus the SADs can be reused by all blocks. Moreover, a low-pass filter based pixel decimation is also introduced to decrease the matching cost. Experiments show that the proposed fast VBSME algorithm can reduce the computation to about 0.2% of fast full search (FFS) with robust image quality. Yang Song 0002, Zhenyu Liu 0001, Takeshi Ikenaga, Satoshi Goto |
ICME | 3 |
| 2007 | Enhanced Strict Multilevel Successive Elimination Algorithm for Fast Motion EstimationabstractThis paper presents an enhanced strict multilevel successive elimination algorithm (EMSEA) for fast block-matching motion estimation, which is based on the previous strict multilevel successive elimination algorithm (SMSEA) (Song et al., 2006). Different to the SMSEA algorithm with fixed parameters, in EMSEA algorithm, the whole search area is divided into two regions and each region has its own parameters. Therefore, the computation complexity of SMSEA algorithm can be further decreased. Experiments show that the proposed EMSEA algorithm can reduce 16.2% of the SMSEA computation and maintain almost the same image quality, which is better than the TSS and DS algorithms. Yang Song 0002, Zhenyu Liu 0001, Takeshi Ikenaga, Satoshi Goto |
ISCAS | 3 |
| 2007 | Power-efficient LDPC code decoder architectureabstractThis paper proposes the power-efficient LDPC decoder architecture which features (1) a FIFO buffering based rapid convergence schedule which enables the decoder to accelerate the decoding throughput without increasing the required number of memory bits, (2) an intermediate message compression technique based on a clock gated shift register which reduces the read and write power dissipation for the intermediate messages. Simulation results show that the proposed decoder achieves 1.66 times faster decoding throughput, and improves the power efficiency (which is defined by the power dissipation per Mbps) up to 52% compared to the decoder based on the conventional overlapped schedule. Kazunori Shimizu, Nozomu Togawa, Takeshi Ikenaga, Satoshi Goto |
ISLPED | 3 |
| 2007 | An MRF model-based approach to the detection of rectangular shape objects in color images
Yangxing Liu, Takeshi Ikenaga, Satoshi Goto |
Signal Process. | 2 |
| 2006 | High-throughput decoder for low-density parity-check codeabstractWe have designed and implemented the LDPC decoder chip with memory-reduction method to achieve high-throughput and practical chip size. The decoder decodes (3,6)-2304 bit regular LDPC codes using modified min-sum algorithm. The decoder achieves a throughput of 530Mb/s at an operating frequency of 147MHz. The chip has been fabricated in a 0.18mum, 6 metal-layer CMOS technology. The chip size is 36mm2 Tatsuyuki Ishikawa, Kazunori Shimizu, Takeshi Ikenaga, Satoshi Goto |
ASP-DAC | 3 |
| 2006 | Low-Pass Filter Based Vlsi Oriented Variable Block Size Motion Estimation Algorithm for H.264abstractIn this paper, a fast motion estimation algorithm, which is friendly to VLSI hardware implementation is proposed. This algorithm has such features: First, through "Haar" low-pass filter based subsampling, the computation complexity at each search position is reduced to about 25% of the original algorithm; Second, one modified motion vector prediction is provided to eliminate the data dependence among sub-partitions in the same macro block (MB). Based on this approach, parallel processing for variable block size motion estimation (VBSME) with integer pixel accuracy can be realized; Third, one "adaptive sub-search window" scheme is proposed to further reduce computation cost and it also can facilitate reference frame data reusing to reduce memory transfer from the external RAM to the on-chip SRAM. The proposed VBSME algorithm is very suitable for parallel VLSI implementation Zhenyu Liu 0001, Yang Song 0002, Takeshi Ikenaga, Satoshi Goto |
ICASSP (2) | 3 |
| 2006 | An ultra-low complexity motion estimation algorithm and its implementation of specific processorabstractMotion estimation (ME) requires huge computation complexity. Many motion estimation algorithms have been proposed to reduce its complexity. But they are still insufficient for embedded video coding systems. So we proposed an ultra-low complexity ME algorithm that is suitable for the software implementation. The simulation results show that proposed algorithm has about 1,000 times the speedup than full search (FS) maintaining high image quality. And we also propose an application specific instruction-set processor (ASIP) for ME. It is based on a reduced instruction set computer (RISC) with sum of absolute difference (SAD) operation circuit. Our ME ASIP is implemented on FPGA. It is required about 3,313 logic elements (LEs) and its hardware scale is about quarter of the previous ME ASIP. This ME ASIP makes a significant contribution to the development of compact video coding systems Seiichiro Hiratsuka, Satoshi Goto, Takeshi Ikenaga |
ISCAS | 3 |
| 2006 | A parallel LSI architecture for LDPC decoder improving message-passing scheduleabstractThis paper proposes a parallel LSI architecture for LDPC decoder which improves a message-passing schedule. The proposed LDPC decoder is characterized as follows: (i) the column operations follow the row operations in a pipelined architecture to ensure that the row and column operations are performed concurrently; and (ii) the proposed parallel pipelined bit functional unit enables the decoder to perform every column operation using the messages which is updated by the row operations. These column operations can be performed without extending the single iterative decoding delay. Hardware implementation and simulation results show that the proposed decoder improves the decoding throughput and bit error performance with a small hardware overhead Kazunori Shimizu, Tatsuyuki Ishikawa, Nozomu Togawa, Takeshi Ikenaga, Satoshi Goto |
ISCAS | 4 |
| 2006 | High performance VLSI architecture of fractional motion estimation in H.264 for HDTVabstractFractional motion estimation (FME) on sub-pixels will occupy almost over 45% of the computation complexity of H.264 encoding process. Therefore a high performance VLSI architecture of FME is described in this paper to achieve the capacity of encoding the high-resolution real-time video stream for HDTV. Our design is improved from an existing work by involving a pipeline strategy in sub-pixel interpolation unit which can avoid the long delay paths in 6-tap ID FIR so as to increase the clock frequency up to 200MHz. Moreover, a 16-pixel search engine is adopted to remove the redundant interpolation area and parallelize the various block size search which can save more than half of the clock cycles in processing a macro block. Our design is implemented with only 189K gates at operating frequency of 200MHz in worst case (285MHz in typical case). It can provide the processing capacity of more than 250K MB/sec which is enough for 1080HD (1920times1088) video streams at frame rate of 30fps. It is a useful intellectual property (IP) design for multimedia system Changqi Yang, Satoshi Goto, Takeshi Ikenaga |
ISCAS | 3 |
| 2006 | A Fully Automatic Approach of Color Image Edge DetectionabstractEdge detection is a vital pre-processing step of many image analysis systems. In this paper, we intend to present a novel algorithm to automatically extract color image edge by integrating multi-dimensional gradient analysis and statistical analysis on local regions. Compared with previous gradient-based edge detection algorithms, our algorithm does not need to select an appropriate threshold against the gradient magnitude. With an elaborate edge detector, we exploit image pixel gradient direction, magnitude, spatial information and region property to obtain color image edge. To avoid the problem of detecting false edges, we conduct statistical analysis on certain local regions, which have high edge density, to further optimize our edge detection result. Experimental results demonstrate the performance of our algorithm on different color images. Yangxing Liu, Takeshi Ikenaga, Satoshi Goto |
SMC | 2 |
| 2005 | An efficient deblocking filter architecture with 2-dimensional parallel memory for H.264/AVCabstractIn this paper, we present an efficient architecture for deblocking filter in H.264/AVC. A novel 2-dimensional parallel memory scheme is employed in order to achieve highly efficient parallel access in both horizontal and vertical directions. By using this parallel memory scheme, we also eliminate the need for a transpose circuit. Our design is implemented under 0.35/spl mu/m technology. Synthesis results show that the equivalent gate count is only 9.35K (not including SRAMs) when the maximum frequency is 100MHz. Satoshi Goto, Takeshi Ikenaga |
ASP-DAC | 3 |
| 2005 | Reconfigurable adaptive FEC system with interleavingabstractThis paper proposes a reconfigurable adaptive FEC system with interleaving. For adaptive FEC schemes, we can implement an optimal RS decoder composed of minimum hardware units for any given error correction capability t. If the hardware units of the RS decoder can be reduced for any given t, we can embed as large deinterleaver as possible into the RS decoder for each t. Reconfiguring the RS decoder embedded with the expanded deinterleaver dynamically for each t allows us to decode larger interleaved codes which are more robust FEC codes to burst errors. Our reconfigurable adaptive FEC system with interleaving achieves better packet error rate and higher throughput than fixed hardware systems. Kazunori Shimizu, Nozomu Togawa, Takeshi Ikenaga, Satoshi Goto |
ASP-DAC | 3 |
| 2005 | A VLSI array processing oriented fast fourier transform algorithm and hardware implementationabstractMany parallel Fast Fourier Transform (FFT) algorithms adopt multiple stages architecture to increase performance. However, data permutation between stages consumes volume memory and processing time. An FFT array processing mapping algorithm is proposed in this paper to overcome this demerit. In this algorithm, arbitrary 2k butterfly units (BUs) could be scheduled to work in parallel on n=22 data (k=0, 1,..., s-1). Because no inter stage data transfer is required, memory consumption is reduced to 1/3 of the original algorithm. Moreover, with the increasing of BUs, not only does throughput increase linearly, system latency also decreases linearly. This array processing orientated architecture provides flexible tradeoff between hardware cost and system performance. An 18-bit word-length 1024-point FFT architecture with 4 BUs is given to demonstrate this mapping algorithm. The design is implemented with TSMC 0.18μm CMOS technology. The core area is 2.99x1.12mm 2 and clock frequency is 326MHz in typical condition (1.8V, 25°C). This processor could complete 1024 FFT calculation in 7.839μs. Zhenyu Liu 0001, Yang Song 0002, Takeshi Ikenaga, Satoshi Goto |
ACM Great Lakes Symposium on VLSI | 3 |
| 2005 | Partially-Parallel LDPC Decoder Based on High-Efficiency Message-Passing AlgorithmabstractThis paper proposes a partially-parallel LDPC decoder based on a high-efficiency message-passing algorithm. Our proposed partially-parallel LDPC decoder performs the column operations for bit nodes in conjunction with the row operations for check nodes. Bit functional unit with pipeline architecture in our LDPC decoder allows us to perform column operations for every bit node connected to each of check nodes which are updated by the row operations in parallel. Our proposed LDPC decoder improves the tuning when the column operations are performed, accordingly it improves the message-passing efficiency within the limited number of iterations for decoding. We implemented the proposed partially-parallel LDPC decoder on an FPGA, and simulated its decoding performance. Practical simulation shows that our proposed LDPC decoder reduces the number of iterations for decoding, and it improves the bit error performance with a small hardware overhead. Kazunori Shimizu, Tatsuyuki Ishikawa, Takeshi Ikenaga, Satoshi Goto, Nozomu Togawa |
ICCD | 3 |
| 2005 | A Robust Algorithm for Text Detection in Color ImagesabstractText detection in color images has become an active research area since recent decades. In this paper, we present a novel approach to accurately detect text in color images possibly with a complex background. First, we use an elaborate edge detection algorithm to extract all possible text edge pixels. Second connected component analysis is employed to construct text candidate region and classify part non-text regions. Third each text candidate region is verified with texture features derived from wavelet domain. Finally, the expectation maximization algorithm is introduced to binarize text regions to prepare data for recognition. In contrast to previous approach, our algorithm combines both the efficiency of connected component based method and robustness of texture based analysis. Experimental results show that our algorithm is robust in text detection with respect to different character size, orientation, color and language and can provide reliable text binarization result. Yangxing Liu, Satoshi Goto, Takeshi Ikenaga |
ICDAR | 3 |
| 2005 | An accurate and low complexity approach of detecting circular shape objects in still color imagesabstractObject detection is a critical step of many image recognition systems. In this paper, we discussed the circular shape object detection problem in still color images. The proposed method has an important feature of integrating color image edge extraction result with a novel circle parameter determination algorithm efficiently. First, an accurate isotropic edge detector is introduced to extract color image edge and calculate precise gradient direction of each potential edge pixel, which assures the high accuracy of subsequent circle detection. Then the detected potential edge results are verified by integrating them with spatial information and the result of region based analysis. Second we utilize just only one 2-dimensional accumulator array and one 1-dimensional accumulator array, which greatly reduce storage requirement and time complexity, to detect circles of any radius. Furthermore, the validity of detected potential circle centers is checked to avoid detecting false circles. Experimental results show that our method is robust and effective in locating objects with complete or incomplete or concentric circle boundary in real color images without any prior knowledge. Yangxing Liu, Satoshi Goto, Takeshi Ikenaga |
ICIP (1) | 3 |
| 2005 | Multi-class QoS routing strategies based on the network state
Hedia Kochkar, Takeshi Ikenaga, Kenji Kawahara, Yuji Oie |
Comput. Commun. | 2 |
| 2004 | Performance evaluation of TCP under dynamic allocation scheme for down-link transmission rate in W-CDMA systemsabstractAbstract IMT‐2000 has attracted much attention as the Next Generation Mobile Communication System. IMT‐2000 can provide high‐bit‐rate data communication service, so that a great number of packets conveyed by TCP (transmission control protocol) can be transmitted over a wireless link. W‐CDMA, standardized in 3GPP (which standardizes IMT‐2000), allows dynamically allocating transmission rates to flow over a wireless link in response to a changing FER for each flow, which is thought to be essential for a next generation mobile communication system. Therefore, in this paper, we study the characteristics of the dynamic allocation scheme when TCP flows share a wireless link, and, in particular, we focus on the throughput performance of these TCP flows. First, we use simulations to examine the effectiveness of the dynamically allocating down‐link transmission rate for TCP flows in response to changing the frame error rate (FER). Through the simulation results, we will show how it can improve the total throughput performance of TCP flows. Furthermore, we can obtain an effective way to allocate transmission rates to flows with different FERs in order to achieve high total throughput. Finally, we will deal with a case of multiple flows from a fixed host to a mobile host. In actual networks, this often happens. In this case, we will show that the total throughput of TCP flows degrades less than in the single‐flow case, even when the FER is high. Copyright © 2004 John Wiley & Sons, Ltd. Yutaka Fukuda, Hiroyuki Koga, Takeshi Ikenaga, Yuji Oie |
Wirel. Commun. Mob. Comput. | 3 |
| 2003 | Analysis of Delayed Reservation Scheme in Server-Based QoS Management NetworkabstractThis paper proposes an analytical model for a delayed reservation scheme, which can be exactly analyzed. By applying it to a server-based QoS management network, we can obtain the related blocking probability and the waiting time distribution of requests, and discuss its performance by means of numerical results. The numerical results show that the delayed scheme can improve the blocking probability by about two orders of magnitude compared with the conventional reservation scheme even with a small acceptable waiting time (e.g. 20% of average flow duration). We can conclude that the delayed reservation scheme can lead to significant improvement in a server-based QoS management network performance. Takeshi Ikenaga, Kenji Kawahara, Tetsuya Takine, Yuji Oie |
AINA | 1 |
| 2002 | Performance evaluation of delayed reservation schemes in server-based QoS managementabstractThe current Internet has a single class of service, i.e., best-effort service, whereas multimedia information is actually transmitted there. Recently, many researchers have focused on a mechanism to provide some quality of service (QoS) required for applications such as delay-sensitive ones. With such a mechanism, a real-time communication request is sent to a server, such as a policy server, that will try to find a path that meets the request's QoS requirement. If several available paths meet the requirement, the connection will be established on one of them. Otherwise, the request will be rejected. In all likelihood, though, delaying the path reservation to some extent if no available path meets the requirement, rather than immediately rejecting a request, will improve the blocking probability. We assume here that the communication can tolerate some initial setup delay, which we refer to as acceptable delay. Our major objective is to clarify whether or not the delay schemes are effective in improving the blocking probability. Therefore, we examine two schemes for managing buffered requests and two different network topologies, i.e., Mesh and ISP. Through an extensive simulation, we have found that the blocking probability can be improved even when a small acceptable delay and a small buffer size are assumed. In particular, we found that the scheme works well in actual network topologies such as an ISP model. Takeshi Ikenaga, Yoshinori Isozaki, Yoshiaki Hori, Yuji Oie |
GLOBECOM | 1 |
| 2002 | Performance Evaluation of Channel Switching Scheme for Packet Data Transmission in Radio Network Controller
Yoshiaki Ohta, Kenji Kawahara, Takeshi Ikenaga, Yuji Oie |
NETWORKING | 3 |
| 2001 | QoS routing algorithm based on multiclasses traffic loadabstractMost of the QoS-based routing schemes proposed so far focus on improving the performance of individual service classes. In a multi-class network, where high priority QoS traffic coexists with best-effort traffic, routing decision for QoS sessions will have an effect on lower ones. A mechanism that allows dynamic link sharing among multiple classes of traffic is needed. We propose a multi-class routing algorithm based on inter-class sharing of resources among multiple class of traffic. Our algorithm is based on the concept of the "virtual residual bandwidth", which is derived from the real residual bandwidth. The virtual residual bandwidth is greater than the residual bandwidth when the load of lower priority traffic is light and smaller when the load of lower priority traffic is heavy. The idea of our approach is simple since the routing algorithm for individual traffic does not change, the only change being the definition of the link cost. We demonstrate, through extensive simulation, the effectiveness of our approach when the best effort distribution is uneven and when its load is heavy. Also, better performance is noticed when using a topology with a large number of alternative paths. Hedia Kochkar, Takeshi Ikenaga, Yuji Oie |
GLOBECOM | 2 |
| 2000 | Real-time morphology processing using highly parallel 2-D cellular automata CAM2abstractMathematical morphology is a promising computer paradigm based on set theory and has many applications in image processing. Although some architectures have been proposed, there are as yet no compact, practical computers that can handle a variety of morphological operations with large, complex structuring elements at video rates. This has prevented the great potential of morphology from being fully realized. This paper describes a morphology processing method that uses a highly parallel two-dimensional (2-D) cellular automaton architecture called it CAM2 (Cellular AutoMata on Content Addressable Memory). New mapping methods achieve high-throughput complex morphology processing. Evaluation results show that CAM2 performs one morphological operation for basic structuring elements within 30 micros. Furthermore, CAM2 can also handle an extremely large and complex structuring element of 100x100 at video rates. CAM2 will increase the potential use of morphology and make a significant contribution to the development of various real-time image processing systems. Takeshi Ikenaga, Takeshi Ogura |
IEEE Trans. Image Process. | 1 |
| 1998 | CAM2: A Highly-Parallel Two-Dimensional Cellular ArchitectureabstractCellular automaton (CA) is a promising computer paradigm that can break through the von Neumann bottleneck. Two-dimensional CA is especially suitable for application to pixel-level image processing. Although various architectures have been proposed for processing two-dimensional CA, there are no compact, practical computers. So, in spite of its great potential, CA is not widely used: This paper proposes a highly-parallel two-dimensional cellular automaton architecture called CAM/sup 2/ and presents some evaluation results. CAM/sup 2/ can attain pixel-order parallelism on a single board because it is composed of a CAM, which makes it possible to embed an enormous number processing elements (PEs), corresponding to CA cells, onto one VLSI chip. Multiple-zigzag mapping and dedicated CAM functions enable high-performance CA processing. The performance evaluation results show that 256 k CA cells, which correspond to a 512/spl times/512 picture, can be processed by a CAM/sup 2/ on a single board using deep submicron process technology. The processing speed is more than 10 billion CA cell updates per second. This means that more than a thousand CA-based image processing operations can be done on a 512/spl times/512 pixel image at video rates (33 msec). CAM/sup 2/ will widen the potentiality of CA and make a significant contribution to the development of compact and high-performance systems. Takeshi Ikenaga, Takeshi Ogura |
IEEE Trans. Computers | 1 |
| 1997 | Real-Time Morphology Processing Using Highly Parallel 2D Cellular Automata CAM2abstractThis paper proposes a real-time morphology processing method based on a highly-parallel 2D cellular automata called CAM/sup 2/ (cellular automata on content addressable memory) that can attain pixel-order parallelism (several hundred thousands) on a single PC board. Evaluation results show that CAM/sup 2/ can perform various types of morphology processing within video rates, indicating it could be a promising platform for processing morphology. Takeshi Ikenaga, Takeshi Ogura |
ICIP (2) | 1 |