VLDB 2026 Research / reviewers in the wild / expert
Hiroki Takahashi
dblp:04/6512
· DBLP profile ↗
50ranked-venue papers
8as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 8 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 5 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TKA-STAGCN: a skeleton-based graph convolutional network with temporal attention for baseball pitch type classification
Sergio Huesca-Flores, Gibran Benitez-Garcia, Oswaldo Juarez-Sandoval, Hiroki Takahashi, Mariko Nakano-Miyatake |
Multim. Tools Appl. | 4 |
| 2025 | Automatic and Interactive Annotation of Non-manual and Spatial Features in Pidgin Sign Japanese for SLR
Gibran Benitez-Garcia, Nobuko Kato, Yuhki Shiraishi, Hiroki Takahashi |
ACIVS | 4 |
| 2025 | FCR-PoseHRNet: Flexible Feature Realignment and Cross-Resolution Coordinate Refinement in PoseHRNet for 2D Human Pose Estimation
Zakir Ali, Gibran Benitez-Garcia, Hiroki Takahashi |
ACIVS | 3 |
| 2025 | AniDream: Generating Skeleton-Guided Anime Avatars from Text PromptsabstractGenerating high-quality anime avatars has become an increasingly important task in the fields of animation, gaming, and virtual reality. However, existing frameworks often face challenges in achieving anatomical consistency and mitigating visual artifacts. To address these limitations, we introduce AniDream, a novel framework designed for the generation of high-quality anime avatars. Unlike previous approaches that primarily relied on image-based inputs, AniDream incorporates text-guided generation, allowing users to create diverse anime avatars directly from text prompts. AniDream uses a skeleton-guided approach, ensuring anatomical consistency while focusing on refining attention around key regions. Our framework introduces a novel loss function that simulates a cel-shading effect and encourages the generated avatars to maintain sharp contour definitions and shadowing consistent with anime aesthetics. Experiments show that AniDream outperforms other frameworks by reducing artifacts and maintaining visual consistency across various poses and viewpoints. It also achieves an average CLIPScore of 33.07, demonstrating its effectiveness in closely aligning generated avatars with text prompts. Fernanda Miyuki Yamada, Hiroki Takahashi |
ISMAR | 2 |
| 2025 | Reinforcement Learning for Circular Sparsest Packing ProblemsabstractThe sparsest packing problem emerges in manufacturing multi-hole extrusion dies to obtain small and precise products in the automotive, aviation, food, and medical industries. The goal is to maximize the minimum Euclidean distance between the objects and between the objects to the boundaries of the container. Additionally, this task might be subject to balancing constraints that determine that the deviation of gravity center of the system should stay within a threshold. We present a novel custom environment that encompasses the constraints present in this task. We experiment with the proposed environment using Proximal Policy Optimization to assess the applicability of reinforcement learning for the sparsest packing problem with circular objects in a circular container. Our results indicate that the proposed agent learns efficiently, demonstrating promising results in both finding feasible solutions and optimizing the placement of objects. Our approach not only shows the potential of reinforcement learning for solving the sparsest packing problem but also provides insights into its effectiveness in environments with complex spatial and balancing constraints. Fernanda Miyuki Yamada, Joao Paulo Gois, Harlen Costa Batagelo, Hiroki Takahashi |
SoMeT | 4 |
| 2025 | TANGAN: solving Tangram puzzles using generative adversarial networkabstractWhile humans show remarkable proficiency in solving visual puzzles, machines often fall short due to the complex combinatorial nature of such tasks. Consequently, there is a growing interest in developing computational methods for the automatic solution of different puzzles, especially through deep learning approaches. The Tangram, an ancient Chinese puzzle, challenges players to arrange seven polygonal pieces to construct different patterns. Despite its apparent simplicity, solving the Tangram is considered an NP-complete problem, being a challenge even for the most sophisticated algorithms. Moreover, ensuring the generality and adaptability of machine learning models across different Tangram arrangements and complexities is an ongoing research problem. In this paper, we introduce a generative model specifically designed to solve the Tangram. Our model competes favorably with previous methods regarding accuracy while delivering fast inferences. It incorporates a novel loss function that integrates pixel-based information with geometric features, promoting a deeper understanding of the spatial relationships between pieces. Unlike previous approaches, our model takes advantage of the geometric properties of the Tangram to formulate a solving strategy, exploiting its inherent properties only through exposure to training data rather than through direct instruction. Extending the proposed loss function, we present a novel evaluation metric as a better fitting measure for assessing Tangram solutions than previous metrics. We further provide a new dataset containing more samples than others reported in the literature. Our findings highlight the potential of deep learning approaches in geometric problem domains. Fernanda Miyuki Yamada, Harlen Costa Batagelo, Joao Paulo Gois, Hiroki Takahashi |
Appl. Intell. | 4 |
| 2024 | Automated Annotation Assistance for Pidgin Sign Japanese in Sign Language RecognitionabstractJapanese Sign Language (JSL) and Manually Coded Japanese (MCJ) are the two main forms of sign language used in Japan. The former differs significantly from spoken Japanese, while the latter closely aligns with spoken syntax. Pidgin Sign Japanese (PSJ) is an intermediate form that heavily relies on nonmanual signals such as facial expressions and head movements for grammatical nuances. Current sign language recognition (SLR) systems predominantly focus on MCJ, neglecting the challenging properties of PSJ. This paper proposes an annotation assistance tool designed to automate the annotation of non-manual and spatial elements in PSJ. Our tool significantly reduces the manual effort required for annotation by using state-of-the-art methods for tracking human pose, hand, and face landmarks, along with recognizing facial action units (FAUs). Validation on a preliminary dataset of 30 videos containing over 90 instances of nonmanual elements demonstrated a $\mathbf{4 0 \%}$ reduction in annotation time, highlighting our proposal’s efficiency and effectiveness in handling the complexities of PSJ. Gibran Benitez-Garcia, Nobuko Kato, Yuhki Shiraishi, Hiroki Takahashi |
CW | 4 |
| 2024 | PFMNet: Face Mask Recognition with Deformable Convolution Networks and Category AttentionabstractThe challenges posed by the COVID-19 pandemic underscored the critical importance of proper mask usage, highlighting the need for automated systems to monitor face mask-wearing conditions. In this paper, we introduce PFMNet, a novel architecture for recognizing the wearing status of face masks. PFMNet is inspired by the InternImage architecture and employs Deformable Convolution Networks (DCNs) to capture long-range dependencies crucial for accurate mask status determination. The significant challenge of class imbalance, particularly the scarcity of improperly worn mask samples, is addressed by integrating the Category Attention Block (CAB). CAB improves distinct regions, diversifies feature representations, and utilizes efficient global pooling to identify crucial areas, such as the human face, while reducing the computational cost. The performance of PFMNet was assessed using the publicly available PWMFD dataset, which had to be refined due to duplicate images and incorrect annotations. PFMNet was compared to three other state-of-the-art models: InternImage, ConvNext, and EfficientNet. It outperformed these models, achieving an accuracy of 99.39%. This places it ahead of the second-best model by a margin of 0.45%. The confusion matrices illustrate that PFMNet outperforms other models in all classes, particularly excelling in the “with mask” and “without mask” categories, resulting in the best overall performance. Ulises Arroyo-Rojas, Gibran Benitez-Garcia, Jesus Olivares-Mercado, Gabriel Sanchez-Perez, Hiroki Takahashi |
SoMeT | 5 |
| 2024 | Multimodal Hand Gesture Recognition Using Automatic Depth and Optical Flow Estimation from RGB VideosabstractTraditional Hand Gesture Recognition (HGR) approaches often rely on multiple sensors, such as RGB, depth, and infrared cameras, to capture comprehensive multimodal data. However, this increases hardware complexity and costs, limiting the widespread adoption of HGR systems. In this paper, we propose a Multi-modal HGR approach that leverages automatic depth estimation from RGB videos to enhance HGR performance while using only a single RGB camera. Our method integrates synthetic depth features, optical flow (OF), and RGB data through an early fusion strategy. We conduct extensive experiments using three ConvNet-based models for HGR: the 3D-CNN variants of ResNet and ResNeXt, as well as the efficient 2D-CNN-based Temporal Shift Module (TSM). Our findings indicate that the multimodal input combination of synthetic Depth, OF, and RGB modalities results in superior performance compared to models using solely RGB or RGB+OF inputs, with the ResNeXt-101 model exhibiting the highest accuracy. To validate our approach, we employ the IPN Hand dataset, which we have meticulously refined to correct temporal annotation inconsistencies and increased the number of gesture classes from 13 to 14 by separating the similar dynamics of four specific gestures. Furthermore, we compute high-quality OF and depth maps for all 800,000 frames in the dataset. The enhanced and multimodal data of the IPN Hand dataset will be soon available at github.com/GibranBenitez/IPN-hand. Gibran Benitez-Garcia, Hiroki Takahashi |
SoMeT | 2 |
| 2024 | Optimal Feature Extractor for Video Anomaly Detection in Public Transportation ApplicationsabstractVideo Anomaly Detection (VAD) is a well-established area of research with significant potential for enhancing video surveillance in urban public transportation. However, current VAD systems often propose powerful methodologies but overlook their use in extreme environments like public transportation, necessitating a balance between performance and computational efficiency. In this paper, we evaluate a key component in many VAD frameworks: feature extractors. We investigate five extractors: Inflated 3D ConvNets (I3D), 3D Convolutional Neural Networks (C3D), Unified Transformer (UniFormer) in Small (UniFormer-S) and Base (UniFormer-B) versions, and Temporal Shift Module (TSM). These are integrated into a VAD architecture employing Bidirectional Encoder Representations from Transformers (BERT) with Multiple Instance Learning (MIL), chosen for its modularity and clear separation between the feature extractor and anomaly detector module. UniFormer-S demonstrated a processing rate of 4.64 clips per second with a computational demand of 28.717 GFLOPs on edge devices like the Jetson Orin NX (8GB RAM, 20W power). On the UCF-Crime dataset, UniFormer-S with BERT + MIL achieves an AUC of 79.74%. These findings highlight the promise of UniFormer-S and the use of edge devices like the Jetson Orin NX in public transportation due to their balance of performance and efficiency. Jonathan Flores-Monroy, Gibran Benitez-Garcia, Mariko Nakano-Miyatake, Hiroki Takahashi |
SoMeT | 4 |
| 2024 | Attention-Based Multi-Scale, Context-Aware Feature Integration into PoseResNet for Coordinate Classification in 2D HPEabstractHuman Pose Estimation (HPE) is a critical Computer Vision task with applications ranging from video surveillance to medical rehabilitation. Despite recent advancements in Deep Learning, HPE still faces challenges such as occluded keypoints, variable lighting conditions, and high computational demands. To address these issues, we present the Attention-Based Multi-Scale, Context-Aware Feature Integration into PoseResNet for Coordinate Classification (AMSF-PRNetCC). Our framework enhances the traditional ResNet architecture by incorporating CoordConv2d layers, depthwise separable convolutions, and novel attention mechanisms including Spatial-Enhanced Channel Attention (SECA) and Squeeze-and-Excitation (SE). We introduce a Context-Aware Feature Pyramid Network (CAFPN) with Dual Mask Global Context Blocks (DMGCB) to efficiently handle multi-scale information. The model culminates in Multi-Layer Perceptron (MLP) stages for precise keypoint coordinate classification. Evaluations on the COCO dataset demonstrate that AMSF-PRNetCC significantly outperforms existing 2D HPE methods in both accuracy and computational efficiency. Our approach achieves state-of-the-art results while requiring fewer computational resources, marking a substantial advancement in the field of HPE. Zakir Ali, Sartaj Ahmed Salman, Gibran Benitez-Garcia, Hiroki Takahashi |
SoMeT | 4 |
| 2023 | Efficient 3Dconv Fusion of RGB and Optical Flow for Dynamic Hand Gesture Recognition and Localization
Gibran Benitez-Garcia, Hiroki Takahashi |
PSIVT | 2 |
| 2022 | TFM a Dataset for Detection and Recognition of Masked Faces in the WildabstractDroplet transmission is one of the leading causes of the spread of respiratory infections, such as coronavirus disease (COVID-19). The proper use of face masks is an effective way to prevent the transmission of such diseases. Nonetheless, different types of masks provide various degrees of protection. Hence, automatic recognition of face mask types may benefit the control access to facilities where a specific protection degree is required. In the last two years, several deep learning models have been proposed for face mask detection and properly wearing mask recognition. However, the current publicly available datasets do not consider the different mask types and occasionally lack real-world elements needed to train robust models. In this paper, we introduce a new dataset named TFM with sufficient size and variety to train and evaluate deep learning models for face mask detection and recognition. This dataset contains more than 135,000 annotated faces from about 100,000 photographs taken in the wild. We consider four mask types (cloth, respirators, surgical and valved) as well as unmasked faces, of which up to six can appear in a single image. The photographs were mined from Twitter within two years since the beginning of the COVID-19 pandemic. Thus, they include diverse scenes with real-world variations in background and illumination. With our dataset, the performance of four state-of-the-art object detection models is evaluated. The experimental results show that YOLOv5 can achieve about 90% of [email protected], demonstrating that the TFM dataset can be used to train robust models and may help the community step forward in detecting and recognizing masked faces in the wild. Our dataset and pre-trained models used in the evaluation will be available upon the publication of this paper. Gibran Benitez-Garcia, Hiroki Takahashi, Miguel Jimenez-Martinez, Jesus Olivares-Mercado |
MMAsia | 2 |
| 2022 | Twitter Face Image Mining for Recognition of Different Face Mask TypesabstractIn the current pandemic of coronavirus disease (COVID-19), an effective way to prevent the transmission and infection of the virus is the proper use of face masks. However, the different types of masks provide different degrees of protection. For instance, valved masks protect the user but do not help to stop the transmission. Hence, the automatic recognition of face mask types may benefit applications that control access to facilities where a certain facepiece is required. In this paper, we propose a Twitter mining framework to gather a large-scale dataset of masked faces suitable to train deep learning-based models for face mask recognition. We employ a keyword-based selection where non-face images are discarded by an efficient face detector (Retinaface). Finally, we train a state-of-the-art CNN architecture (ConvNeXt) for recognizing the wearing mask. We also present a brief analysis of more than two million image-based tweets acquired over two years since the beginning of the pandemic. The code of the proposed framework and a preliminary dataset of more than 10K faces (manually annotated into unmasked, surgical, cloth, respirators, and valved masks) are available on github.com/GibranBenitez/FaceMask Twitter. Ulises Arroyo-Rojas, Miguel Jimenez-Martinez, Gibran Benitez-Garcia, Jesus Olivares-Mercado, Hiroki Takahashi |
SoMeT | 5 |
| 2021 | Masked Batch Normalization to Improve Tracking-Based Sign Language Recognition Using Graph Convolutional NetworksabstractSign language recognition is a fundamental technique to improve communication between native signers and speakers. Current state-of-the-art sign language recognition methods often apply deep neural network models to learn an optimized projection between sign language videos and sentences in an end-to-end manner. Generally, minibatch training using sequential data requires the addition of padding to equalize the varying lengths of sequences. However, this training strategy induces performance degradation if batch normalization is used in the models because batch normalization assumes the validity of all inputs. In this study, we propose masked batch normalization, which normalizes input features while masking dummy signals. We apply masked batch normalization to tracking-based sign language recognition models using graph convolutional networks. The performance of the proposed method is evaluated in isolated sign language word recognition and continuous sign language words recognition settings. To evaluate the proposed method, we use two types of sign language video datasets, WLASL including 2000 types of isolated words, and a JSL dataset including 275 types of videos of isolated words and 113 types of videos showing sentences. The evaluation results show that the proposed method improves the tracking-based sign language recognition models in both cases. Natsuki Takayama, Gibran Benitez-Garcia, Hiroki Takahashi |
FG | 3 |
| 2020 | Automatic Focal Blur Segmentation Based on Difference of Blur Feature Using Theoretical Thresholding and Graphcuts
Natsuki Takayama, Hiroki Takahashi |
ACIVS | 2 |
| 2020 | Data Augmentation Using Feature Interpolation of Individual Words for Compound Word Recognition of Sign LanguageabstractSign language recognition is an important topic to improve communication between native signers and native speakers. In this paper, we propose one of the data augmentations to recognize compound words represented by sequences of two individual sign language words. The proposed method generates two-dimensional tracking points of the compound words by concatenating the two individual sign language words' tracking points. Linear interpolation is applied to concatenate the tracking points of the individual sign language words. The evaluation was conducted using a conventional sign feature extraction and recognition model based on the encoder-decoder Recurrent Neural Network with Attention. The minimum average word error rate of the model, which trained with the generated compound words, was 18.85%. Moreover, the minimum average word error rate of the model, which trained with the generated and recorded compound words, was 7.15%, and it was improved by 3.08% from that of the recorded compound words. Natsuki Takayama, Hiroki Takahashi |
CW | 2 |
| 2020 | NR-WLAN Aggregation: Architecture for Supporting URLLC in 5G IoT NetworksabstractThe standard for 5thgeneration (5G) cellular mobile communications, which is referred to as New Radio (NR), has been released by the 3rdGeneration Partnership Project (3GPP). One of the key technologies supported by NR is ultra-reliable and low latency communications (URLLC), aiming at meeting extremely stringent requirements: data error rate of below 10-5and layer 2 latency of below 1 ms. One of the typical use cases of URLLC is Internet of Things (IoT) e.g. factory automation. Considering the fact that Wireless LAN (WLAN) systems are already widely deployed for IoT industrial applications, the IoT system can make full use of WLAN for supporting URLLC. Therefore, system architectures should be investigated in which NR technology can be integrated into the WLAN system which is already deployed for IoT industrial applications. In this paper, three architectures are proposed. Among them, it is found that architectures with novel feature “bearer aggregation” provides better support of URLLC. One open question is discussed, and it is shown that an enhanced WLAN system which is developed by the authors can potentially address the question. Yoshiaki Ohta, Ryuichi Takechi, Hiroki Takahashi, Ryu Atsuta |
VTC Spring | 3 |
| 2018 | Automatic Extraction of Reorganization Impact Focusing on Derivation Relationship of Analogous Actor Terms in Requirements SpecificationabstractThe authors developed a tool to extract the analogous actor terms in a requirements specification based on standard actors. Using this tool, a number of analogous actor terms found in the requirements specification are automatically extracted. The authors also propose the use of the extraction results to improve the actor terms in a reorganization. Further, the authors applied this method to an actual case of requirements specifications and subsequent reorganization and evaluated the effectiveness. Hiroki Takahashi, Norifumi Nomura, Tadahisa Kondou, Mari Inoki |
APSEC | 1 |
| 2018 | Sign Words Annotation Assistance Using Japanese Sign Language Words RecognitionabstractA Japanese sign language corpus is essential to activate analysis and recognition research of Japanese sign language. It requires collecting large scale of video data and annotating information to build a sign language corpus. Generally, building a sign language corpus is tedious work, and assistance is necessary. This paper describes one of the assistance methods for annotation tasks of sign words using Japanese sign language words recognition. The words recognition extracts sign features from a video, segments it into meaningful units, and annotates word labels to them automatically. At this time, the user's annotation tasks can be reduced from the full-manual work to confirmation and correction of the annotation. The proposed sign words recognition is composed of body-parts tracking, feature extraction, and words classification. The five types of approaches including i) feature fusion and ii) multi-stream HMM to handle the multiple body-parts are applied and compared. We build a video database of Japanese sign language words and a manual annotation interface to evaluate the proposed method. The database includes 92 Japanese sign language words which are signed by ten native signers. The total number of videos is 4,590, and 3,900 videos of 78 words except for recording and sign errors are used for the evaluation. The classification accuracies were 75.88% and 93.35% in the signer and trial opened conditions, respectively, when the parts-based feature fusion and multi-stream HMM using relative weights for body-parts are employed. Moreover, the expected work reduction ratio of annotation tasks using the interface was 38.01%. Natsuki Takayama, Hiroki Takahashi |
CW | 2 |
| 2016 | A Listen before Talk Algorithm with Frequency Reuse for LTE Based Licensed Assisted Access in Unlicensed SpectrumabstractThis paper investigates a listen before talk (LBT) algorithm to realize both collision avoidance with nodes of other wireless systems/operators and one-cell reuse operation among nodes of the same operator for long term evolution (LTE) based licensed-assisted access (LAA) in unlicensed spectrum. When LTE system is operated in unlicensed spectrum, implementation of a LBT algorithm is required to not only follow the regulation but also realize fair coexistence with other operators' nodes or other systems' nodes, e.g. wireless local area network (WLAN), which can be operated in the same channel. However, since physical layer functionalities of LTE is designed to work in one-cell reuse operation, from the efficiency of unlicensed channels, it is preferable that LAA nodes of the same operator simultaneously transmits signals. Therefore, we study a LBT with frequency reuse which can flexibly perform collision avoidance and one-cell reuse operation according to whether the channel is idle or busy. In the proposed LBT, while applying a LBT with random backoff algorithm as the basis, each LAA node aligns the transmission timing with other LAA nodes of the same operator. Computer simulation confirms that LBT with frequency reuse can improve the throughput performances of both LAA users and WLAN users under the environment where LAA and WLAN nodes coexist. Naoki Kusashima, Toshizo Nogami, Hiroki Takahashi, Kazunari Yokomakura, Kimihiko Imamura |
VTC Spring | 3 |
| 2015 | Foreground Object Extraction Using Variation of Blurs Based on Camera FocusingabstractImage foreground object extraction is still challenging yet interesting topic in visual computing area. This paper proposes foreground extraction method based on one of the basic camera functions, focusing. A focus point and variation of blurs are extracted based on focusing, and these can be essential information to extract the foreground image. A focus point shows the rough position of the foreground, and variation of blurs becomes a distinctive image feature to segment foreground and background. The major contributions of this paper are developing a framework to extract a foreground object based on a camera function, and estimating the blur with the matching position between images using SIFT (Scale Invariant Feature Transform). SIFT is able to detect and describe image features which are invariant to scale, rotation and luminance changing. Moreover, scale-space extrema in SIFT can be used to estimate the variation of blurs. This is shown experimentally. Proposed foreground extraction is conducted in three steps. First, sparse variation of blurs between images is computed using SIFT. Next, the sparse variation of blurs is propagated using EAI (Edge Aware Interpolation) and a full variation of blur map is generated. Finally, foreground object extraction is obtained by Graph Cut algorithm using a focus point as a constraint of the object. Experimental evaluation result shows improvement in performance of foreground object extraction using variation of blurs. Natsuki Takayama, Hiroki Takahashi |
CW | 2 |
| 2015 | Ride through capability of matrix converter for grid connected system under short voltage sagabstractThis paper proposes a FRT (Fault ride through) method of a matrix converter under three-phase short voltage sag. The feature of the proposed method is to control grid reactive current and generator torque at the same time. In order to realize these capabilities, this paper describes a modulation method dividing a control period into 3 modes. The first mode is to turn off all switches and the second mode is a zero vector output. By these 2 modes and utilizing a snubber circuit, the generator torque is controlled during the voltage sag. In the third mode, a non-zero vector is selected to provide the reactive current to the grid. In contrast, this paper also proposes a feedback control of the snubber voltage to obtain the stable ride through operation. From experimental results, the grid reactive current of 0.44 p.u. and the generator torque of 0.7 p.u. are achieved by the proposed method during three-phase voltage sag. Hiroki Takahashi, Jun-ichi Itoh |
IECON | 1 |
| 2014 | The effect of ceiling height on the symbolic distance effect
Masashi Sugimoto, Takashi Kusumi, Tokika Kurita, Atsuo Ishikawa, Takeshi Sakaguchi, Megumi Nabetani, Hiroki Takahashi, Megumi Nishida |
CogSci | 7 |
| 2014 | Power decoupling method for isolated DC to single-phase AC converter using matrix converterabstractThis paper presents an isolated DC to single-phase AC converter using a matrix converter for HVDC (higher voltage direct current) power feeding system. The proposed converter comprises a full bridge inverter, a high frequency transformer and a matrix converter and does not use a bulky electrolytic capacitor. Then, in order to reduce a ripple component in a DC bus current caused by the single-phase load, this paper also proposes a power decoupling method. The power decoupling method employs a center-tapped transformer and a small LC buffer instead of additional switches, which aims to achieve high efficiency. Moreover, modulation methods of the full bridge inverter and the matrix converter, and a control strategy of the power decoupling are described. As an experimental result, the power decoupling method reduces the DC bus current ripple to 2/3. In addition, a validity of a parameter design method of the proposed control is confirmed in simulations. Hiroki Takahashi, Nagisa Takaoka, Raul Roberto Rodriguez Gutierrez, Jun-ichi Itoh |
IECON | 1 |
| 2014 | A Transmit Power Control Based Interference Mitigation Scheme for Small Cell Networks Using Dynamic TDD in LTE-Advanced SystemsabstractThis paper proposes a downlink (DL)/uplink (UL) transmission power control (TPC) scheme for dynamic time division duplex (TDD) based small cells under multi-cell environment. In the dynamic TDD, an eNB (evolved node B) selects an adequate UL-DL configuration according to the ratio of DL to UL data bits in each cell. However, in this case, eNB- eNB and user equipment (UE)-UE interferences could be additional interference since the transmission directions can be different among cells. Especially, eNB-eNB interference significantly degrades the UL transmission performances. Therefore, we investigate the DL TPC which is applied to subframes which can be different directions among cells and the benefits to decrease eNB-eNB interference by this scheme. We also investigate the different UL TPC parameters that are applied according to the subframe types in order to alleviate the impact of eNB-eNB interference. Computer simulation confirms that the proposed DL/UL TPC scheme can achieve 21.6 % gain at maximum for UL throughput without significant DL throughput degradation. Hiroki Takahashi, Kazunari Yokomakura, Kimihiko Imamura |
VTC Spring | 1 |
| 2011 | A Design Criterion of Error Correcting Codes for Spectrum-Overlapped Resource ManagementsabstractIn this paper, we propose a new criterion for the extrinsic information transfer (EXIT) analysis in the spectrum-overlapped resource management (SORM) for a broadband single carrier transmission. In the SORM technique, each user can ideally obtain a maximum channel gain by allowing overlapped allocation, assuming the soft canceller with minimum mean square error (SC/MMSE) turbo equalization for multi-user detection. When the interference is not completely removed by SC/MMSE turbo equalization, frame error rate (FER) remarkably degrades. To solve this problem, we proposed a new criterion which is used for design of error correcting code in order to improve convergence property of SC/MMSE turbo equalization without large redundancy of channel code. In the proposed criterion, we design the decoder so that the decoder's EXIT property is fitted to the upper bound of equalizer's property and the channel code sharply improves mutual information (MI) in the lower amount of input MI. This paper evaluates FER performance with the irregular low-density parity-check (LDPC) codes designed based on this criterion. As a result, this paper shows that the irregular LDPC code designed based on the proposed criterion for the SORM outperforms conventional codes. Jungo Goto, Hiroki Takahashi, Osamu Nakamura, Kazunari Yokomakura, Yasuhiro Hamaguchi, Shinsuke Ibi, Seiichi Sampei |
ICC | 2 |
| 2011 | AMDORAP: Non-targeted metabolic profiling based on high-resolution LC-MSabstractBACKGROUND: Liquid chromatography-mass spectrometry (LC-MS) utilizing the high-resolution power of an orbitrap is an important analytical technique for both metabolomics and proteomics. Most important feature of the orbitrap is excellent mass accuracy. Thus, it is necessary to convert raw data to accurate and reliable m/z values for metabolic fingerprinting by high-resolution LC-MS. RESULTS: In the present study, we developed a novel, easy-to-use and straightforward m/z detection method, AMDORAP. For assessing the performance, we used real biological samples, Bacillus subtilis strains 168 and MGB874, in the positive mode by LC-orbitrap. For 14 identified compounds by measuring the authentic compounds, we compared obtained m/z values with other LC-MS processing tools. The errors by AMDORAP were distributed within ±3 ppm and showed the best performance in m/z value accuracy. CONCLUSIONS: Our method can detect m/z values of biological samples much more accurately than other LC-MS analysis tools. AMDORAP allows us to address the relationships between biological effects and cellular metabolites based on accurate m/z values. Obtaining the accurate m/z values from raw data should be indispensable as a starting point for comparative LC-orbitrap analysis. AMDORAP is freely available under an open-source license at http://amdorap.sourceforge.net/. Hiroki Takahashi, Takuya Morimoto, Naotake Ogasawara, Shigehiko Kanaya |
BMC Bioinform. | 1 |
| 2010 | Choshi Design System from 2D Images
Natsuki Takayama, Shubing Meng, Hiroki Takahashi |
ICEC | 3 |
| 2009 | Comparison of Near-Threshold Characteristics of Flash Suppression and Forward Masking
Hiroki Takahashi, Hideaki Itoh, Kiyohiko Nakamura |
ICONIP (1) | 2 |
| 2008 | Master manipulator with higher operability designed for micro neuro surgical systemabstractThe master and slave surgical assistant systems have been studied actively. However, regarding the master system, the manipulator should be designed for each surgical field because the target workspace and precision are different. Furthermore, operability is important for safety reasons. Therefore, the authors analyzed the motion of surgeon first and then developed a master manipulator suitable for the operation of micro neuro surgery. Some experiments were conducted to evaluate the control method and show the effectiveness of the proposed manipulator control method. Hiroki Takahashi, Tsubasa Yonemura, Naohiko Sugita, Mamoru Mitsuishi, Shigeo Sora, Akio Morita, Ryo Mochizuki |
ICRA | 1 |
| 2007 | A remote surgery experiment between Japan and Thailand over Internet using a low latency CODEC systemabstractRemote surgery is one of the most desired applications in the context of recent advanced medical technologies. For a future expansion of remote surgery, it is important to use conventional network infrastructures such as Internet. However, using such conventional network infrastructures, we are confronting time-delay problems of data transmission. In this paper, a remote surgery experiment between Japan and Thailand using a research and development Internet is presented. In the experiment, the image and audio information was transmitted by a newly developed low latency CODEC system to shorten the time-delay. By introducing the low latency CODEC system, the time-delay was shortened compared with the past remote surgery experiments despite the longer distance. We also conducted several network measurements such as a comparison between TCP/IP and UDP/IP about the control signal transmission. Jumpei Arata, Hiroki Takahashi, Phongsaen Pitakwatchara, Shin'ichi Warisawa, Kazuo Tanoue, Kozo Konishi, Satoshi Ieiri, Shuji Shimizu, Naoki Nakashima, Koji Okamura, Yuichi Fujino, Yukihiro Ueda, Pornarong Chotiwan, Mamoru Mitsuishi, Makoto Hashizume |
ICRA | 2 |
| 2006 | A Remote Surgery Experiment between Japan-Korea using the Minimally Invasive Surgical SystemabstractSeveral robotic surgical systems have been developed for MIS (minimally invasive surgery) including commercialized products such as da Vinci and ZEUS. We have developed a minimally invasive surgical system, which have carried out remote surgery experiments for five times at the writing time. In this paper, a remote surgery experiment, which was conducted between Japan and Korea by using the developed minimally invasive surgical system is described. Research & Development (R & D) Internet testbet, APII (Asia-Pacific Information Infrastructure), which consists of an optical submarine cable network KJCN (Korea-Japan Cable Network), was used. In the experiment, a laparoscopic cholecystectomy was successfully carried out on a pig. The network time-delays of control signal and images were 6.5 msec and 435.5 msec respectively. A comparison of remote surgery experiments using ISDN and the Internet was studied Jumpei Arata, Hiroki Takahashi, Phongsaen Pitakwatchara, Shin'ichi Warisawa, Kozo Konishi, Kazuo Tanoue, Satoshi Ieiri, Shuji Shimizu, Naoki Nakashima, Koji Okamura, Sungmin Kim, Joon-Soo Hahm, Makoto Hashizume, Mamoru Mitsuishi |
ICRA | 2 |
| 2006 | Low Cost Rendering Method for Virtual Factory Considering Interpolation of Occluded Objects
Hiroki Takahashi, Naoyuki Tamura, Toshihiko Furue, Osamu Yoshie |
MoMM | 1 |
| 2006 | Spherical Wavelet Descriptors for Content-based 3D Model RetrievalabstractThe description of 3D shapes with features that possess descriptive power and invariant under similarity transformations is one of the most challenging issues in content based 3D model retrieval. Spherical harmonics-based descriptors have been proposed for obtaining rotation invariant representations. However, spherical harmonic analysis is based on latitude-longitude parameterization of a sphere which has singularities at each pole. Consequently, features near the two poles are over represented while features at the equator are under-sampled, and variations of the north pole affects significantly the shape function. In this paper we discuss these issues and propose the usage of spherical wavelet transform as a tool for the analysis of 3D shapes represented by functions on the unit sphere. We introduce three new descriptors extracted from the wavelet coefficients, namely: (1) a subset of the spherical wavelet coefficients, (2) the L1and, (3) the L2energies of the spherical wavelet sub-bands. The advantage of this tool is three fold; first, it takes into account feature localization and local orientations. Second, the energies of the wavelet transform are rotation invariant. Third, shape features are uniformly represented which makes the descriptors more efficient. Spherical wavelet descriptors are natural extension of 3D Zernike moments and spherical harmonics. We evaluate, on the Princeton shape benchmark, the proposed descriptors regarding computational aspects and shape retrieval performance Hamid Laga, Hiroki Takahashi, Masayuki Nakajima 0001 |
SMI | 2 |
| 2006 | Spherical parameterization and geometry image-based 3D shape similarity estimation (CGS 2004 special issue)
Hamid Laga, Hiroki Takahashi, Masayuki Nakajima 0001 |
Vis. Comput. | 2 |
| 2005 | Discerning Advisor: An Intelligent Advertising System for Clothes Considering Skin ColorabstractFacial perception is what usually human beings do during their daily communications. They may adjust and change the method or contents of the interaction based on this perception. One of the most probable places of this process is during the personalized advertising that is highly under consideration nowadays. Here as a sample of this kind of systems it is tried to propose a system named 'Discerning Advisor' for suggesting clothing fashion considering matching colors with the perception of skin color from the face. This is a considerable point by the fashion designers and advisors which the system will be trained by their knowledge. For this purpose a comprehensive study has been done on the human skin color model and its variations around the world and among different races. A neural network based skin detector was designed and trained to extract the skin area of the images automatically. It was applied on a large dataset of digital images containing 1000 faces of peoples of almost all races from 23 cities around the world. Calculating indicator color for each individual gave the ability of analysis of distribution and fuzzy-clustering the skin color. Finally the clustering result was used in classification of the skin color that can be matched with the suggested colors from the fashion designers Mohammad Ali Akbari, Hiroki Takahashi, Masayuki Nakajima 0001 |
CW | 2 |
| 2005 | Estimating object contours from binary edge imagesabstractWe propose a new method of estimating object contours from thresholded binary edge images. Our work is motivated by the edge-based video object plane (VOP) generation proposed by Meier, targeted at MPEG-4 object-based video coding. When objects are detected in the form of binary edge images, we have to estimate the complete object contour from edges that do not form a closed contour. The proposed method is designed to carry out this task using only binary edge images. To achieve this, we first introduce a probability field representing the probability that the object contour lies on each pixel, and then take a modified geodesic active contour approach to make the initial contour converge with the contour of the object according to the probability field. Hiroyuki Tsuji, Suguru Saito, Hiroki Takahashi, Masayuki Nakajima 0001 |
ICIP (3) | 3 |
| 2005 | Color Transformation Method for Universal Web View
Yayori Kasagi, Hiroki Takahashi, Osamu Yoshie |
iiWAS | 2 |
| 2004 | Hands-Free Navigation Methods for Moving through a Virtual Landscape Walking Interface Virtual Reality Input DevicesabstractTechnological limitations on current interfaces have made researches to develop new devices to interact with objects in the virtual environment. The goal of this project is to develop and build a hands-free navigation system to be integrated into virtual environments. One of the most important fields in virtual reality (VR) research, is the development of systems that allow the user to interface with the virtual environment. The most intuitive method for moving through a virtual landscape is by walking. The implementation of a walking interface for a virtual reality system also allows a greater range of biomechanical experimentation and game research. Systems ranging from different platforms have already been implemented to produce virtual walking; however, these systems have been designed primarily for use with head mounted display systems. We believe that hands-free navigation, unlike the majority of navigation techniques based on hand motions, has the greatest potential for maximizing the interactivity of virtual environments, due to more direct motion of the feet. To make this possible, we created a new and simple device using acceleration sensors to detect ankle movements within the virtual environment. The acceleration sensors are attached to the foot and detect movement based on direction for three different angles. This experimentation could prove beneficial in future virtual gaming. Validation of our approach is given by discussion and illustration of some results Salvador Barrera, Hiroki Takahashi, Masayuki Nakajima 0001 |
Computer Graphics International | 2 |
| 2004 | Geometry Image Matching for Similarity Estimation of 3D ShapeabstractWe describe our preliminary findings in applying the spherical parametrization and geometry images to the task of 3D shape matching and similarity based comparison of polygon soup models. Unlike traditional approach where multiple 2D views of the same object are required to capture the relevant geometry features, our proposed technique uses spherical parametrization and geometry images for 3D shape matching. This technique reduces the problem to the 2D space without computing 2D projections and guarantees the preservation of small details. Moreover, we take advantage from the hierarchical nature of the parametrization process, through the progressive mesh simplification, to derive a multiresolution analysis technique. Our proposed algorithm is invariant to similarity transformations such as rotation and scaling. The efficiency of this approach is discussed through a set of experiments Hamid Laga, Hiroki Takahashi, Masayuki Nakajima 0001 |
Computer Graphics International | 2 |
| 2004 | Joyfoot's Cyber System: A Virtual Landscape Walking Interface Device for Virtual Reality ApplicationsabstractTechnological limitations on current interfaces have made researches to develop new devices to interact with objects in the virtual environment. The goal of this project is to develop and build a hands-free navigation system to be integrated into virtual environments. One of the most important fields in virtual realty (VR) research, is the development of systems that allow the user to interface with the virtual environment. The most intuitive method for moving through a virtual landscape is by walking. The implementation of a walking interface for a virtual reality system also allows a greater range of biomechanical experimentation and game research. Systems ranging from different platforms have already been implemented to produce virtual walking; however, these systems have been designed primarily for use with head mounted display systems. We believe that hands-free navigation, unlike the majority of navigation techniques based on hand motions, has the greatest potential for maximizing the interactivity of virtual environments, due to more direct motion of the feet. To make this possible, we created a new and simple device using acceleration sensors to detect ankle movements within the virtual environment. The acceleration sensors are attached to the foot and detect movement based on direction for three different angles. The design was called "Joyfoot's Cyber System". This experimentation could prove beneficial in future virtual gaming. Validation of our approach is given by discussion and illustration of some results. Salvador Barrera, Hiroki Takahashi, Masayuki Nakajima 0001 |
CW | 2 |
| 2004 | Toward E-Appearance of Human Face and Hair by Age, Expression and RejuvenationabstractRecently, Web based facial appearance systems have received more attention by various aspects and applications like online systems. Therefore there is a request for a system to have an ability of predict and render different appearance effects on facial images for online systems. Here, an E-appearance system is proposed to predict the essential effects of facial images with different appearance. Given facial image data of the same person in two different appearance, quotient image then capture appearance characteristic of image. Then, together with warping technique we map the characteristic to any other particular persons face in order to generate new facial appearance. The technique makes experiments on different facial appearance. Azam Bastanfard, Hiroki Takahashi, Masayuki Nakajima 0001 |
CW | 2 |
| 2004 | Scale-Space Processing of Point-Sampled Geometry for Efficient 3D Object SegmentationabstractIn this paper, we present a new framework for analyzing and segmenting point-sampled 3D objects. Our method first computes for each surface point the surface curvature distribution by applying the principal component analysis on local neighborhoods with different sizes. Then we model in the four dimensional space the joint distribution of surface curvature and position features as a mixture of Gaussians using the expectation maximization algorithm. Central to our method is the extension of the scale-space theory from the 2D domain into the three-dimensional space to allow feature analysis and classification at different scales. Our algorithm operates directly on points requiring no vertex connectivity information. We demonstrate and discuss the performance of our framework on a collection of point sampled 3D objects. Hamid Laga, Hiroki Takahashi, Masayuki Nakajima 0001 |
CW | 2 |
| 2004 | Agent Handling with VR Technology and Spatial Programming for Information Aggregation in Plants
Hiroki Takahashi, Osamu Yoshie |
iiWAS | 1 |
| 2004 | Toward anthropometrics simulation of face rejuvenation and skin cosmeticabstractAbstract Facial rejuvenation is the process of reversing the aging effects on the human face digitally. This paper generalizes a new approach for facial rejuvenation in adults image. Applications of facial rejuvenation are widespread. They include face recognition, education, entertainment, telecommunications, Psychology, criminal objects, Cosmetic arts and it can be used as an aid for medical cosmetics surgery and the reconstruction of the face. This paper proposes a novel facial rejuvenation modeling algorithm with two techniques. These techniques discuss the facial deformation based on the face anthropometrics theory and remove wrinkles based on what we called wrinkles inpainting. For example if we have been given a few different faces, we need to be able to compare the difference between the facial characteristics of the youth and the aged, then from there onwards, define a set of outlines which are going to be the basis of the simulation of Face Rejuvenation. The first is the geometric deformation details like skin texture, which differs between the aged and the youth. The second is anthropometrics data change. It was developed in the face anthropometrics measurement theory. Then together with warping technique we map the characteristics to any other particular persons' face in order to generate more expressive and convincing facial rejuvenation. The original contribution and advantage of this paper are that, the proposed methods are simple to implement, reliable, in which they required only one source image without needing to collect a lot of images and their computation are fast for interactive environment. Copyright © 2004 John Wiley & Sons, Ltd. Azam Bastanfard, O. Bastanfard, Hiroki Takahashi, Masayuki Nakajima 0001 |
Comput. Animat. Virtual Worlds | 3 |
| 2003 | Real Time Detection Interface For Walking on CAVEabstractPresently, most virtual reality systems use upper body parts to interact with objects in the virtual environment. This situation is caused by technological limitations of current interface devices. Starting from this viewpoint we developed a new interface for detecting ankle motions relative to the knee. We believe that hands-free navigation, unlike the majority of navigation techniques based on hand motions, has the greatest potential for maximizing the interactivity of virtual environments since navigation modes are more direct motion of the feet. We therefore, created a simple device to detect ankle movements with rotary encoders sensors. These sensors rotate according to the amount and direction of the movement of the foot. The sensors are attached to a sandal and can be used for many purposes including virtual games. Validation of our approach is given by discussion and illustration of some experimental results. Salvador Barrera, Piperakis Romanos, Suguru Saito, Hiroki Takahashi, Masayuki Nakajima 0001 |
Computer Graphics International | 4 |
| 2003 | A Method of Rendering Scenes Including Volumetric Objects Using Ray-Volume Buffers - Expanding to Render Scenes Including Overlapped Volumetric ObjectabstractWe describe a rendering method that simultaneously renders both volumetric and polygonal objects. It is also possible to render volumetric objects intersected by polygonal ones or volumetric ones in the same scheme. In the proposed method, a scene including volumes and polygons overlapping each other is rendered using a ray-volume buffer that stores the colors and transmittance of a ray at each voxel. Polygons inserted in a volume are rendered by mapping the ray-volume data as 3D texture. We rendered a mixed scene of two volumes in the same scheme by considering a volume overlapping with another volume as multiple layers of translucent polygons. The ray-volume buffer method treats volumetric objects as polygons with textures. Therefore, the method can be realized using conventional graphics pipeline techniques by generating a ray volume for a 3D texture. Another feature of this method is that ray volumes can be generated without consideration of the intersection with polygons or volumes. Kagenori Kajihara, Hiroki Takahashi, Masayuki Nakajima 0001 |
Computer Graphics International | 2 |
| 2003 | Image Categorization using Color Blobs in a Mobile EnvironmentabstractAbstract This paper generalizes the basic idea of blobs in preattentive perception using color information. This is usedas the base of a basic classification of low resolution pictures taken with mobile phones. This classification, theblob‐like representation of the image and other information in user's context, such as GPS information, can beused in the presented framework as the basis of a new graphical interface for HCI (Human Computer Interaction).Similar systems whether they work with global properties of the image, which leads to inaccurate results, or withcomplex segmentation process that fails to capture expected objects in the scene. Most of those systems do notpay attention on other information involved in the creation of the image, such as time or location. We describe asystem which uses geographical information associated with a picture in a mobile phone terminal, and with a fastsegmentation based on color categorization. David Gavilan Ruiz, Hiroki Takahashi, Masayuki Nakajima 0001 |
Comput. Graph. Forum | 2 |
| 1997 | Hair image generating algorithm using fractional hair model
Masayuki Nakajima 0001, Seiichi Saruta, Hiroki Takahashi |
Signal Process. Image Commun. | 3 |