EDBT 2026 Demo / reviewers in the wild / expert
Yoshimitsu Aoki
dblp:02/3149
· DBLP profile ↗
48ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0001-7361-0027ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 26 · 1 first-author · 15 since 2021Systems, architecture and hardware · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sign-to-Speech Prosody Transfer via Sign Reconstruction-Based GAN
Toranosuke Manabe, Yuto Shibata, Shinnosuke Takamichi, Yoshimitsu Aoki |
ICPR (8) | 4 |
| 2025 | Pre-training with Synthetic Patterns for AudioabstractIn this paper, we propose to pre-train audio encoders using synthetic patterns instead of real audio data. Our proposed framework consists of two key elements. The first one is Masked Autoencoder (MAE), a self-supervised learning framework that learns from reconstructing data from randomly masked counterparts. MAEs tend to focus on low-level information such as visual patterns and regularities within data. Therefore, it is unimportant what is portrayed in the input, whether it be images, audio mel-spectrograms, or even synthetic patterns. This leads to the second key element, which is synthetic data. Synthetic data, unlike real audio, is free from privacy and licensing infringement issues. By combining MAEs and synthetic patterns, our framework enables the model to learn generalized feature representations without real data, while addressing the issues related to real audio. To evaluate the efficacy of our framework, we conduct extensive experiments across a total of 13 audio tasks and 17 synthetic datasets. The experiments provide insights into which types of synthetic patterns are effective for audio. Our results demonstrate that our framework achieves performance comparable to models pre-trained on AudioSet-2M and partially outperforms image-based pre-training methods. Yuchi Ishikawa, Tatsuya Komatsu, Yoshimitsu Aoki |
ICASSP | 3 |
| 2025 | Formula-Supervised Sound Event Detection: Pre-Training Without Real DataabstractIn this paper, we propose a novel formula-driven supervised learning (FDSL) framework for pre-training an environmental sound analysis model by leveraging acoustic signals parametrically synthesized through formula-driven methods. Specifically, we outline detailed procedures and evaluate their effectiveness for sound event detection (SED). The SED task, which involves estimating the types and timings of sound events, is particularly challenged by the difficulty of acquiring a sufficient quantity of accurately labeled training data. Moreover, it is well known that manually annotated labels often contain noises and are significantly influenced by the subjective judgment of annotators. To address these challenges, we propose a novel pretraining method that utilizes a synthetic dataset, Formula-SED, where acoustic data are generated solely based on mathematical formulas. The proposed method enables large-scale pre-training by using the synthesis parameters applied at each time step as ground truth labels, thereby eliminating label noise and bias. We demonstrate that large-scale pre-training with Formula-SED significantly enhances model accuracy and accelerates training, as evidenced by our results in the DESED dataset used for DCASE2023 Challenge Task 4. The project page is at https://yutoshibata07.github.io/Formula-SED/. Yuto Shibata, Keitaro Tanaka, Yoshiaki Bando, Keisuke Imoto, Hirokatsu Kataoka, Yoshimitsu Aoki |
ICASSP | 6 |
| 2025 | Simultaneous Motion and Noise Estimation with Event CamerasabstractEvent cameras are emerging vision sensors whose noise is challenging to characterize. Existing denoising methods for event cameras are often designed in isolation and thus consider other tasks, such as motion estimation, separately (i.e., sequentially after denoising). However, motion is an intrinsic part of event data, since scene edges cannot be sensed without motion. We propose, to the best of our knowledge, the first method that simultaneously estimates motion in its various forms (e.g., ego-motion, optical flow) and noise. The method is flexible, as it allows replacing the one-step motion estimation of the widely-used Contrast Maximization framework with any other motion estimator, such as deep neural networks. The experiments show that the proposed method achieves state-of-the-art results on the E-MLB denoising benchmark and competitive results on the DND21 benchmark, while demonstrating effectiveness across motion estimation and intensity reconstruction tasks. Our approach advances event-data denoising theory and expands practical denoising use-cases via open-source code. Project page: https://github.com/tub-rip/ESMD Shintaro Shiba, Yoshimitsu Aoki, Guillermo Gallego 0002 |
ICCV | 2 |
| 2025 | Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
Yuchi Ishikawa, Shota Nakada, Hokuto Munakata, Kazuhiro Saito, Tatsuya Komatsu, Yoshimitsu Aoki |
INTERSPEECH | 6 |
| 2024 | Acoustic-based 3D Human Pose Estimation Robust to Human Position
Yusuke Oumi, Yuto Shibata, Go Irie, Akisato Kimura, Yoshimitsu Aoki, Mariko Isogawa |
BMVC | 5 |
| 2024 | Data Collection-Free Masked Video Modeling
Yuchi Ishikawa, Masayoshi Kondo, Yoshimitsu Aoki |
ECCV (11) | 3 |
| 2024 | Rethinking Image Super-Resolution from Training Data Perspectives
Go Ohtani, Ryu Tadokoro, Ryosuke Yamada, Yuki Markus Asano, Iro Laina, Christian Rupprecht 0001, Nakamasa Inoue, Rio Yokota, Hirokatsu Kataoka, Yoshimitsu Aoki |
ECCV (17) | 10 |
| 2024 | Unsupervised Metric Learning for Expressing Color and Shape Information to Uncover Abstract Connections within Image Datasets
Shun Obikane, Haruna Tagawa, Yoshimitsu Aoki |
ICPR (21) | 3 |
| 2024 | Guided by the Way: The Role of On-the-route Objects and Scene Text in Enhancing Outdoor NavigationabstractIn outdoor environments, Vision-and-Language Navigation (VLN) requires an agent to rely on multi-modal cues from real-world urban environments and natural language instructions. While existing outdoor VLN models predict actions using a combination of panorama and instruction features, this approach ignores objects in the environment and learns data bias to fail navigation. According to our preliminary findings, most instances of navigation failure in previous models were due to turning or stopping at the wrong place. In contrast, humans intuitively frequently use identifiable objects or store names as reference landmarks, ensuring accurate turns and stops, especially in unfamiliar places. To address this insight gap, we propose an Object-Attention VLN (OAVLN) model that helps the agent focus on relevant objects during training and understand the environment better. Our model outperforms previous methods in all evaluation metrics under both seen and unseen scenarios on two existing benchmark datasets, Touchdown and map2seq. Yanjun Sun, Yue Qiu 0001, Yoshimitsu Aoki, Hirokatsu Kataoka |
ICRA | 3 |
| 2024 | PCT: Perspective Cue Training Framework for Multi-Camera BEV SegmentationabstractGenerating annotations for bird’s-eye-view (BEV) segmentation presents significant challenges due to the scenes’ complexity and the high manual annotation cost. In this work, we address these challenges by leveraging the abundance of unlabeled data available. We propose the Perspective Cue Training (PCT) framework, a novel training framework that utilizes pseudo-labels generated from unlabeled perspective images using publicly available semantic segmentation models trained on large street-view datasets. PCT applies a perspective view task head to the image encoder shared with the BEV segmentation head, effectively utilizing the unlabeled data to be trained with the generated pseudo-labels. Since image encoders are present in nearly all camera-based BEV segmentation architectures, PCT is flexible and applicable to various existing BEV architectures. In this paper, we applied PCT for semi-supervised learning (SSL) and unsupervised domain adaptation (UDA). Additionally, we introduce strong input perturbation through Camera Dropout (CamDrop) and feature perturbation via BEV Feature Dropout (BFD), which are crucial for enhancing SSL capabilities using our teacher-student framework. Our comprehensive approach is simple and flexible but yields significant improvements over various baselines for SSL and UDA, achieving competitive performances even against the current state-of-the-art. Haruya Ishikawa, Takumi Iida, Yoshinori Konishi, Yoshimitsu Aoki |
IROS | 4 |
| 2024 | Event-Based Background-Oriented SchlierenabstractSchlieren imaging is an optical technique to observe the flow of transparent media, such as air or water, without any particle seeding. However, conventional frame-based techniques require both high spatial and temporal resolution cameras, which impose bright illumination and expensive computation limitations. Event cameras offer potential advantages (high dynamic range, high temporal resolution, and data efficiency) to overcome such limitations due to their bio-inspired sensing principle. This article presents a novel technique for perceiving air convection using events and frames by providing the first theoretical analysis that connects event data and schlieren. We formulate the problem as a variational optimization one combining the linearized event generation model with a physically-motivated parameterization that estimates the temporal derivative of the air density. The experiments with accurately aligned frame- and event camera data reveal that the proposed method enables event cameras to obtain on par results with existing frame-based optical flow techniques. Moreover, the proposed method works under dark conditions where frame-based schlieren fails, and also enables slow-motion analysis by leveraging the event camera's advantages. Our work pioneers and opens a new stack of event camera applications, as we publish the source code as well as the first schlieren dataset with high-quality frame and event data. Shintaro Shiba, Friedhelm Hamann, Yoshimitsu Aoki, Guillermo Gallego 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Secrets of Event-Based Optical Flow, Depth and Ego-Motion Estimation by Contrast MaximizationabstractEvent cameras respond to scene dynamics and provide signals naturally suitable for motion estimation with advantages, such as high dynamic range. The emerging field of event-based vision motivates a revisit of fundamental computer vision tasks related to motion, such as optical flow and depth estimation. However, state-of-the-art event-based optical flow methods tend to originate in frame-based deep-learning methods, which require several adaptations (data conversion, loss function, etc.) as they have very different properties. We develop a principled method to extend the Contrast Maximization framework to estimate dense optical flow, depth, and ego-motion from events alone. The proposed method sensibly models the space-time properties of event data and tackles the event alignment problem. It designs the objective function to prevent overfitting, deals better with occlusions, and improves convergence using a multi-scale approach. With these key elements, our method ranks first among unsupervised methods on the MVSEC benchmark and is competitive on the DSEC benchmark. Moreover, it allows us to simultaneously estimate dense depth and ego-motion, exposes the limitations of current flow benchmarks, and produces remarkable results when it is transferred to unsupervised learning settings. Along with various downstream applications shown, we hope the proposed method becomes a cornerstone on event-based motion-related tasks. Shintaro Shiba, Yannick Klose, Yoshimitsu Aoki, Guillermo Gallego 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Listening Human Behavior: 3D Human Pose Estimation with Acoustic SignalsabstractGiven only acoustic signals without any high-level information, such as voices or sounds of scenes/actions, how much can we infer about the behavior of humans? Unlike existing methods, which suffer from privacy issues because they use signals that include human speech or the sounds of specific actions, we explore how low-level acoustic signals can provide enough clues to estimate 3D human poses by active acoustic sensing with a single pair of microphones and loudspeakers (see Fig. 1). This is a challenging task since sound is much more diffractive than other signals and therefore covers up the shape of objects in a scene. Accordingly, we introduce a framework that encodes multichannel audio features into 3D human poses. Aiming to capture subtle sound changes to reveal detailed pose information, we explicitly extract phase features from the acoustic signals together with typical spectrum features and feed them into our human pose estimation network. Also, we show that reflected or diffracted sounds are easily influenced by subjects' physique differences e.g., height and muscularity, which deteriorates prediction accuracy. We reduce these gaps by using a subject discriminator to improve accuracy. Our experiments suggest that with the use of only low-dimensional acoustic information, our method outperforms baseline methods. The datasets and codes used in this project will be publicly available. Yuto Shibata, Yutaka Kawashima, Mariko Isogawa, Go Irie, Akisato Kimura, Yoshimitsu Aoki |
CVPR | 6 |
| 2022 | Self-Supervised Learning of Inlier Events for Event-based Optical Flow
Jun Nagata, Yoshimitsu Aoki |
BMVC | 2 |
| 2022 | Diverse Plausible 360-Degree Image Outpainting for Efficient 3DCG Background CreationabstractWe address the problem of generating a 360-degree image from a single image with a narrow field of view by estimating its surroundings. Previous methods suffered from overfitting to the training resolution and deterministic generation. This paper proposes a completion method using a transformer for scene modeling and novel methods to improve the properties of a 360-degree image on the output image. Specifically, we use CompletionNets with a transformer to perform diverse completions and Adjust-mentNet to match color, stitching, and resolution with an input image, enabling inference at any resolution. To improve the properties of a 360-degree image on an output image, we also propose WS-perceptual loss and circular inference. Thorough experiments show that our method out-performs state-of-the-art (SOTA) methods both qualitatively and quantitatively. For example, compared to SOTA methods, our method completes images 16 times larger in resolution and achieves 1.7 times lower Fréchet inception distance (FID). Furthermore, we propose a pipeline that uses the completion results for lighting and background of 3DCG scenes. Our plausible background completion enables perceptually natural results in the application of inserting virtual objects with specular surfaces. Naofumi Akimoto, Yuhi Matsuo, Yoshimitsu Aoki |
CVPR | 3 |
| 2022 | Secrets of Event-Based Optical Flow
Shintaro Shiba, Yoshimitsu Aoki, Guillermo Gallego 0002 |
ECCV (18) | 2 |
| 2022 | Document Shadow Removal with Foreground Detection Learning From Fully Synthetic ImagesabstractShadow removal for document images is a major task for digitized document applications. Recent shadow removal models have been trained on pairs of shadow images and shadow-free images. However, obtaining a large-scale and diverse dataset is laborious and remains a great challenge. Thus, only small real datasets are available. To create relatively large datasets, a graphic renderer has been used to synthesize shadows, nonetheless, it is still necessary to capture real documents. Thus, the number of unique documents is limited, which negatively affects a network’s performance. In this paper, we present a large-scale and diverse dataset called fully synthetic document shadow removal dataset (FSDSRD) that does not require capturing documents. The experiments showed that the networks (pre-)trained on FSDSRD provided better results than networks trained only on real datasets. Additionally, because foreground maps are available in our dataset, we leveraged them during training for multitask learning, which provided noticeable improvements. The code is available at: https://github.com/IsHYuhi/DSRFGD. Yuhi Matsuo, Naofumi Akimoto, Yoshimitsu Aoki |
ICIP | 3 |
| 2022 | Fast Event-Based Optical Flow Estimation by Triplet MatchingabstractEvent cameras are novel bio-inspired sensors that offer advantages over traditional cameras (low latency, high dynamic range, low power, etc.). Optical flow estimation methods that work on packets of events trade off speed for accuracy, while event-by-event (incremental) methods have strong assumptions and have not been tested on common benchmarks that quantify progress in the field. Towards applications on resource-constrained devices, it is important to develop optical flow algorithms that are fast, light-weight and accurate. This work leverages insights from neuroscience, and proposes a novel optical flow estimation scheme based on triplet matching. The experiments on publicly available benchmarks demonstrate its capability to handle complex scenes with comparable results as prior packet-based algorithms. In addition, the proposed method achieves the fastest execution time ($\gt$10 kHz) on standard CPUs as it requires only three events in estimation. We hope that our research opens the door to real-time, incremental motion estimation methods and applications in real-world scenarios. Shintaro Shiba, Yoshimitsu Aoki, Guillermo Gallego 0002 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Alleviating Over-segmentation Errors by Detecting Action BoundariesabstractWe propose an effective framework for the temporal action segmentation task, namely an Action Segment Refinement Framework (ASRF). Our model architecture consists of a long-term feature extractor and two branches: the Action Segmentation Branch (ASB) and the Boundary Regression Branch (BRB). The long-term feature extractor provides shared features for the two branches with a wide temporal receptive field. The ASB classifies video frames with action classes, while the BRB regresses the action boundary probabilities. The action boundaries predicted by the BRB refine the output from the ASB, which results in a significant performance improvement. Our contributions are three-fold: (i) We propose a framework for temporal action segmentation, the ASRF, which divides temporal action segmentation into frame-wise action classification and action boundary regression. Our framework refines frame-level hypotheses of action classes using predicted action boundaries. (ii) We propose a loss function for smoothing the transition of action probabilities, and analyze combinations of various loss functions for temporal action segmentation. (iii) Our framework outperforms state-of-the-art methods on three challenging datasets, offering an improvement of up to 13.7% in terms of segmental edit distance and up to 16.1% in terms of segmental F1 score. Our code is publicly available1. Yuchi Ishikawa, Seito Kasai, Yoshimitsu Aoki, Hirokatsu Kataoka |
WACV | 3 |
| 2020 | Fast Soft Color SegmentationabstractWe address the problem of soft color segmentation, defined as decomposing a given image into several RGBA layers, each containing only homogeneous color regions. The resulting layers from decomposition pave the way for applications that benefit from layer-based editing, such as recoloring and compositing of images and videos. The current state-of-the-art approach for this problem is hindered by slow processing time due to its iterative nature, and consequently does not scale to certain real-world scenarios. To address this issue, we propose a neural network based method for this task that decomposes a given image into multiple layers in a single forward pass. Furthermore, our method separately decomposes the color layers and the alpha channel layers. By leveraging a novel training objective, our method achieves proper assignment of colors amongst layers. As a consequence, our method achieve promising quality without existing issue of inference speed for iterative approaches. Our thorough experimental analysis shows that our method produces qualitative and quantitative results comparable to previous methods while achieving a 300,000x speed improvement. Finally, we utilize our proposed method on several applications, and demonstrate its speed advantage, especially in video editing. Naofumi Akimoto, Huachun Zhu, Yanghua Jin, Yoshimitsu Aoki |
CVPR | 4 |
| 2020 | Retrieving and Highlighting Action with Spatiotemporal ReferenceabstractIn this paper, we present a framework that jointly retrieves and spatiotemporally highlights actions in videos by enhancing current deep cross-modal retrieval methods. Our work takes on the novel task of action highlighting, which visualizes where and when actions occur in an untrimmed video setting. Action highlighting is a fine-grained task, compared to conventional action recognition tasks which focus on classification or window-based localization. Leveraging weak supervision from annotated captions, our framework acquires spatiotemporal relevance maps and generates local embeddings which relate to the nouns and verbs in captions. Through experiments, we show that our model generates various maps conditioned on different actions, in which conventional visual reasoning methods only go as far as to show a single deterministic saliency map. Also, our model improves retrieval recall over our baseline without alignment by 2-3% on the MSR-VTT dataset. Seito Kasai, Yuchi Ishikawa, Masaki Hayashi, Yoshimitsu Aoki, Kensho Hara, Hirokatsu Kataoka |
ICIP | 4 |
| 2020 | Joint Pedestrian Detection and Risk-level Prediction with Motion-Representation-by-DetectionabstractThe paper presents a pedestrian near-miss detector with temporal analysis that provides both pedestrian detection and risk-level predictions which are demonstrated on a self-collected database. Our work makes three primary contributions: (i) The framework of pedestrian near-miss detection is proposed by providing both a pedestrian detection and risk-level assignment. Specifically, we have created a Pedestrian Near-Miss (PNM) dataset that categorizes traffic near-miss incidents based on their risk levels (high-, low-, and no-risk). Unlike existing databases, our dataset also includes manually localized pedestrian labels as well as a large number of incident-related videos. (ii) Single-Shot MultiBox Detector with Motion Representation (SSD-MR) is implemented to effectively extract motion-based features in a detected pedestrian. (iii) Using the self-collected PNM dataset and SSD-MR, our proposed method achieved +19.38% (on risk-level prediction) and +13.00% (on joint pedestrian detection and risk-level prediction) higher scores than that of the baseline SSD and LSTM. Additionally, the running time of our system is over 50 fps on a graphics processing unit (GPU). Hirokatsu Kataoka, Teppei Suzuki, Kodai Nakashima, Yutaka Satoh, Yoshimitsu Aoki |
ICRA | 5 |
| 2020 | QR-code Reconstruction from Event Data via Optimization in Code SubspaceabstractWe propose an image reconstruction method from event data, assuming the target images belong to a prespecified class like QR codes. Instead of solving the reconstruction problem in the image space, we introduce a code space that covers all the noiseless target class images and solves the reconstruction problem on it. This restriction enormously reduces the number of optimizing parameters and makes the reconstruction problem well posed and robust to noise. We demonstrate fast and robust QR-code scanning in difficult, high-speed scenes with industrial high-speed cameras and other reconstruction methods. Jun Nagata, Yusuke Sekikawa, Kosuke Hara, Teppei Suzuki, Yoshimitsu Aoki |
WACV | 5 |
| 2019 | Flood Susceptibility Prediction via Data-Mining Based Bell-Curve Analogical-Hydrographs Analysis: A Case Study of Langat River Basin, Selangor, MalaysiaabstractIn this study, we proposed data-mining based bell-curve analogical hydrographs analysis with lag time vertical axes and bankfull discharge horizontal axes to make flood susceptibility prediction. We utilized flood data reports, hourly/daily rainfall data and daily water discharge of Hulu Langat district, Selangor Malaysia from the year 2013–2016 to do flood susceptibility. We implement data mining concept by sorting the database, followed by plotting hydrograph to identify flood patterns and establish relationships to predict flood trends. This method is an intersection between the knowledge field of hydrology and mathematical modeling. When an outlier from the graph is detected, the knowledge from hydrology can be applied to understand the reason behind the appearance of outliers. Besides, the knowledge of mathematical modeling is necessary to assist us in predicting flood susceptibility. The purpose of this study is to predict the flood susceptibility which is vital to prepare the users/public well prepared for smooth and efficient evacuation. In 4 years context, our flood depth predictions are nearly 100% accurate. Factor influencing the lag time and steepness of rising limb are related to land use and topographical features. Implications of the results and future research directions are also presented. Siti Nor Khuzaimah Binti Amit, Yasushi Kiyoki, Yoshimitsu Aoki |
EJC | 3 |
| 2019 | 360-Degree Image Completion by Two-Stage Conditional GansabstractThe latest generative adversarial networks (GANs) can generate realistic high resolution images. However, to the best of our knowledge, there are no GANs for generating 360-degree images. Therefore, this paper proposes the novel problem setting that by using a known area from the 360-degree image as an input, the remainder of the image can be completed with the GANs. To do so, we propose the approach of two-stage generation using network architecture with series-parallel dilated convolution layers. Moreover, we present how to rearrange images for data augmentation, simplify the problem, and make inputs for training the 2ndstage generator. Our experiments show that these methods generate the distortion seen in 360-degree images in the outlines of buildings and roads, and their boundaries are clearer than those of baseline methods. Furthermore, we discuss and clarify the difficulty of our proposed problem. Our work is the first step towards GANs predicting an unseen area within a 360-degree space. Naofumi Akimoto, Seito Kasai, Masaki Hayashi, Yoshimitsu Aoki |
ICIP | 4 |
| 2018 | Tactile Logging for Understanding Plausible Tool Use Based on Human Demonstration
Shuichi Akizuki, Yoshimitsu Aoki |
BMVC | 2 |
| 2018 | Anticipating Traffic Accidents With Adaptive Loss and Large-Scale Incident DBabstractIn this paper, we propose a novel approach for traffic accident anticipation through (i) Adaptive Loss for Early Anticipation (AdaLEA) and (ii) a large-scale self-annotated incident database for anticipation. The proposed AdaLEA allows a model to gradually learn an earlier anticipation as training progresses. The loss function adaptively assigns penalty weights depending on how early the model can anticipate a traffic accident at each epoch. Additionally, we construct a Near-miss Incident DataBase for anticipation. This database contains an enormous number of traffic near-miss incident videos and annotations for detail evaluation of two tasks, risk anticipation and risk-factor anticipation. In our experimental results, we found our proposal achieved the highest scores for risk anticipation (+6.6% better on mean average precision (mAP) and 2.36 sec earlier than previous work on the average time-to-collision (ATTC)) and risk-factor anticipation (+4.3% better on mAP and 0.70 sec earlier than previous work on ATTC). Hirokatsu Kataoka, Yoshimitsu Aoki, Yutaka Satoh |
CVPR | 3 |
| 2018 | Superpixel Convolution for SegmentationabstractIn this paper, we propose a novel segmentation algorithm based on convolutional neural networks (CNNs) on superpix-else CNNs are powerful methods for several computer vision tasks, but spatial information disappears through the pooling process. Moreover, since pooling compresses different types of pixels into a single value, pooling sometimes negatively affects the results of inference in segmentation task. We use superpixel pooling instead of general pooling to resolve this problem. However, general CNNs can't use superpixel images in which the adjacency relationships between pixels are broken. Therefore, we define CNNs and Dilated Convolution on superpixels. Finally, we show the effectiveness of proposed method on an HKU-IS dataset. Teppei Suzuki, Shuichi Akizuki, Naoki Kato, Yoshimitsu Aoki |
ICIP | 4 |
| 2017 | Danger level modeling and analysis of vehicle-pedestrian encounter using situation dependent topic modelabstractThe mechanism behind collisions between vehicles and pedestrians must be thoroughly studied in order to prevent future traffic accidents. In particular, preventing collisions where pedestrian steps out onto the road from behind an obstruction such as buildings, walls or vehicles is a challenging problem. To tackle this problem, we propose situation dependent topic model (SDTM), a regression model that predicts dangerous vehicle-pedestrian encounter in response to different driving situations, which also provides a framework to analyze and understand the underlying factors that lead to dangerous situations. Complex nature of situations where collisions with pedestrians happen can be expressed well by defining how dangerous situations arise differently for each driving situation pattern retrieved using statistical topic modeling. In experiments, we compare the performance of SDTM with orthodox logistic regression models using vehicle-pedestrian encounters in near-miss incidents. We also show the result of acquired knowledge that can form the basis of many other researches concerning pedestrian safety. Kyohei Otsuka, Kosuke Hara, Teppei Suzuki, Yoshimitsu Aoki |
Intelligent Vehicles Symposium | 4 |
| 2016 | Building Change Detection via Semantic Segmentation and Difference Extraction MethodabstractGoogle Earth with high-resolution imagery basically takes months to process new images before online updates. It is considered as a time consuming and slow process especially for post-disaster application. In this study, we aim to develop a fast and accurate method of updating maps by detecting local differences occurred over different time series; where only region with differences will be updated. In our system, aerial imageries from Massachusetts's building open datasets are used as training datasets; meanwhile Saitama district datasets are used as input images. Semantic segmentation is then applied to input images to get predicted map patches of building. Semantic segmentation is a pixel-wise classification of images by implementing convolutional neural network technique. Convolutional neural network technique is implemented due to being not only efficient in learning highly discriminative image features such as buildings, but also partially robust to incomplete and poorly registered target maps. Next, in order to understand overall changes occurred in an area, both semantic segmented images from the same scene are undergone change detection method. Lastly, difference extraction method is implemented to specify the category of building changes. The results reveal that our proposed method is able to overcome current time-consuming map updating problem. Hence map updating will be cheaper, faster and more effective especially post-disaster application, by leaving unchanged region and only updating changed region. Siti Nor Khuzaimah Binti Amit, Shunta Saito, Yoshimitsu Aoki, Yasushi Kiyoki |
EJC | 3 |
| 2016 | Analysis of satellite images for disaster detectionabstractAnalysis of satellite images plays an increasingly vital role in environment and climate monitoring, especially in detecting and managing natural disaster. In this paper, we proposed an automatic disaster detection system by implementing one of the advance deep learning techniques, convolutional neural network (CNN), to analysis satellite images. The neural network consists of 3 convolutional layers, followed by max-pooling layers after each convolutional layer, and 2 fully connected layers. We created our own disaster detection training data patches, which is currently focusing on 2 main disasters in Japan and Thailand: landslide and flood. Each disaster's training data set consists of 30000~40000 patches and all patches are trained automatically in CNN to extract region where disaster occurred instantaneously. The results reveal accuracy of 80%~90% for both disaster detection. The results presented here may facilitate improvements in detecting natural disaster efficiently by establishing automatic disaster detection system. Siti Nor Khuzaimah Binti Amit, Soma Shiraishi, Tetsuo Inoshita, Yoshimitsu Aoki |
IGARSS | 4 |
| 2016 | Parsing human skeletons in an operating room
Vasileios Belagiannis, Xinchao Wang, Horesh Ben Shitrit, Kiyoshi Hashimoto, Ralf Stauder, Yoshimitsu Aoki, Michael Kranzfelder, Armin Schneider, Pascal Fua, Slobodan Ilic, Hubertus Feußner, Nassir Navab |
Mach. Vis. Appl. | 6 |
| 2015 | Depth image enhancement using local tangent plane approximationsabstractThis paper describes a depth image enhancement method for consumer RGB-D cameras. Most existing methods use the pixel-coordinates of the aligned color image. Because the image plane generally has no relationship to the measured surfaces, the global coordinate system is not suitable to handle their local geometries. To improve enhancement accuracy, we use local tangent planes as local coordinates for the measured surfaces. Our method is composed of two steps, a calculation of the local tangents and surface reconstruction. To accurately estimate the local tangents, we propose a color heuristic calculation and an orientation correction using their positional relationships. Additionally, we propose a surface reconstruction method by ray-tracing to local tangents. In our method, accurate depth image enhancement is achieved by using the local geometries approximated by the local tangents. We demonstrate the effectiveness of our method using synthetic and real sensor data. Our method has a high completion rate and achieves the lowest errors in noisy cases when compared with existing techniques. Kiyoshi Matsuo, Yoshimitsu Aoki |
CVPR | 2 |
| 2014 | Extended Co-occurrence HOG with Dense Trajectories for Fine-Grained Activity Recognition
Hirokatsu Kataoka, Kiyoshi Hashimoto, Kenji Iwata, Yutaka Satoh, Nassir Navab, Slobodan Ilic, Yoshimitsu Aoki |
ACCV (5) | 7 |
| 2014 | Finger posture estimation using 3D medial axesabstractWe propose a method for tracking a hand in real-time, using depth information from a single Kinect sensor. The fingers are segmented using medial axes on an image generated by the depth discontinuities. The palm position and orientation is estimated from the position of the wrist as well as the distance to the closest contour points. The detected medial axes are projected in the palm plane for labelling, and then used for direct posture estimation of unoccluded phalanges. Inverse kinematics provide the missing information in case of occlusion. Our approach can be related to an inverse kinematics approach based on fingertip detection, with additional robustness to fingertip occlusion provided by the medial axis of each visible phalanx. Yun Ouedraogo, Yoshimitsu Aoki |
HSI | 2 |
| 2014 | Feature integration with random forests for real-time human activity recognitionabstractThis paper presents an approach for real-time human activity recognition. Three different kinds of features (flow, shape, and a keypoint-based feature) are applied in activity recognition. We use random forests for feature integration and activity classification. A forest is created at each feature that performs as a weak classifier. The international classification of functioning, disability and health (ICF) proposed by WHO is applied in order to set the novel definition in activity recognition. Experiments on human activity recognition using the proposed framework show - 99.2% (Weizmann action dataset), 95.5% (KTH human actions dataset), and 54.6% (UCF50 dataset) recognition accuracy with a real-time processing speed. The feature integration and activity-class definition allow us to accomplish high-accuracy recognition match for the state-of-the-art in real-time. Hirokatsu Kataoka, Kiyoshi Hashimoto, Yoshimitsu Aoki |
ICMV | 3 |
| 2013 | Robust human tracking using statistical human shape model with postural variationabstractHuman tracking in monocular image sequences has been studied in the field of computer vision for many kinds of applications such as surveillance system, intelligent room, sports video analysis and so on. Human tracking in real environment is challenging topic due to various factors such as illumination change, partial or almost complete occlusion of human body, and wide variety of body shapes. In this paper, we present a robust human tracking using statistical human shape model of appearance variation with postural change. Our part-based statistical human model can generate learned appearances of main human poses, and enables effective and robust human tracking with simple features such silhouette, edge and color. Our proposed method achieves human tracking robust not only to partial occlusion but also to postural change. The experimental results validate the robustness of our methods in the real indoor environments. Kiyoshi Hashimoto, Hirokatsu Kataoka, Yoshimitsu Aoki, Yuji Sato |
IECON | 3 |
| 2013 | Robust feature descriptor and vehicle motion model with tracking-by-detection for active safetyabstractThe percentage of pedestrian deaths in traffic accidents is on the rise in Japan. In recent years, there have been calls for measures to be introduced to protect vulnerable road users such as pedestrians and cyclists. In this study, a method to detect and track pedestrians using an in-vehicle camera is presented to perform braking controls, warn the driver, and develop improved safety systems for pedestrians. We improved the technology of detecting pedestrians using highly accurate images obtained with a monocular camera. We were able to predict pedestrian activity by monitoring the images, and developed an algorithm with which to recognize pedestrians and their movements more accurately. The effectiveness of the algorithm was tested using images taken on real roads. For the feature descriptor, we used an extended co-occurrence histogram of oriented gradients (ECoHOG) that accumulated the integration of gradient intensities. In the tracking step, we applied an effective motion model using optical flow and the proposed feature descriptor ECoHOG in a tracking-by-detection framework. These techniques were verified using images captured on the real road. Hirokatsu Kataoka, Kimimasa Tamura, Yoshimitsu Aoki, Yasuhiro Matsui, Kenji Iwata, Yutaka Satoh |
IECON | 3 |
| 2013 | Multiple players tracking and identification using group detection and player number recognition in sports videoabstractWe are interested in the problem of automatically tracking and identifying players in sports video. While there are many automatic multi-target tracking methods, in sports video, it is difficult to track multiple players due to frequent occlusions, quick motion of players and camera, and camera position. We propose tracking method that associates tracklets of a same player using results of player number recognition. To deal with frequent occlusions, we detect human region by level set method and then estimates if it is occluded group region or unoccluded individual one. Moreover, we associate tracklets using the results of player number recognition at each frame by keypoints-based matching with templates from multiple viewpoints, so that final tracklets include occluded region. Taiki Yamamoto, Hirokatsu Kataoka, Masaki Hayashi, Yoshimitsu Aoki, Kyoko Oshima, Masamoto Tanabiki |
IECON | 4 |
| 2012 | Extended CoHOG and particle filter by improved motion model for pedestrian active safetyabstractThe percentage of pedestrian deaths in traffic accidents is on the rise. In recent years, there have been calls for measures to be introduced to protect such vulnerable road users as pedestrians and cyclists. In this study, a method to detect pedestrians using an in-vehicle camera is presented. We improved the technology in detecting pedestrians with highly accurate images using a monocular camera. We were able to predict pedestrians' activities by monitoring them, and we developed an algorithm to recognize pedestrians and their movements more accurately. The effectiveness of the algorithm was tested using images taken on real roads. For the feature descriptor, we found that an extended co-occurrence histogram of oriented gradients, accumulating the integration of gradient intensities. In tracking step, we applied effective motion model using optical flow for Particle Filter tracking. These techniques are valified by using images captured on the real road. Hirokatsu Kataoka, Kimimasa Tamura, Yoshimitsu Aoki, Yasuhiro Matsui |
IECON | 3 |
| 2010 | Proposal of a method to analyze 3D deformation/fracture characteristics inside materials based on a stratified matching approach
Mitsuru Nakazawa, Masakazu Kobayashi, Hiroyuki Toda, Yoshimitsu Aoki |
Mach. Vis. Appl. | 4 |
| 2008 | 3D image analysis for evaluating internal deformation/fracture characteristics of materialsabstractIn the past, D/F characteristics, load-deformation relationships until the materials are fractured, have been analyzed on the surface. The D/F characteristics are affected by more than ten thousand micro-scale internal structures like air bubbles (pores), cracks and particles; therefore, it is required to analyze nano-scale D/F characteristics inside materials. In this paper, we propose a method that automatically obtains the corresponding relations of the particles from nano-order 3DCT images at each deformation stage. The particles are deformation-proof and may have different geometries. First of all, some big particles are considered as landmarks and matched between pre- and post-deformation. The results of landmark matching make it easy to match many remaining particles and pores. Mitsuru Nakazawa, Yoshimitsu Aoki, Masakazu Kobayashi, Hiroyuki Toda |
ICPR | 2 |
| 2008 | Highly accurate Geometric Correction for NOAA AVHRR data considering the feature of elevation and the feature of coastlineabstractNOAA (National Oceanic and Atmospheric Administration) satellite images have been widely used for environmental and land cover monitoring. This paper proposes a geometric correction method that corrects the geometric distortions in NOAA images. In this method, the errors in NOAA images are corrected in the image coordinate system before transforming into the map coordinate system. First, the variation of elevation is verified to divide data into flat and rough blocks. Next, in order to measure the distortions more precisely, more GCP (Ground Control Point) templates are generated based on the feature of the coastline and the elevation errors of these GCP templates are corrected. After using GCP template matching to specify the residual errors, which are used to represent the distortions, affine transform and Radial Basis Function transform are used to correct the distortions on the flat and rough blocks, respectively. With the proposed method, the average values of the error after correction are smaller than 0.2 pixels on both latitude and longitude directions. This result proved that the proposed method is a highly accurate geometric correction method. An Ngoc Van, Mitsuru Nakazawa, Yoshimitsu Aoki |
IGARSS (2) | 3 |
| 2007 | High accurate geometric correction for NORA AVHRR data considering elevation effectabstractThis paper describes a high accurate geometric correction method for NOAA AVHRR data considering elevation effect. NOAA data in map coordinate system is divided into rough and flat blocks to identify the variation of elevation. GCP template matching and affine coefficients are then used to correct residual errors in image coordinate system. In order to correct more residual errors, appropriate flat blocks are combined with GCP templates. Finally, corrected data is transformed into map coordinate system by bilinear interpolation. Because the sizes of rough blocks are smaller, the errors of bilinear interpolation are reduced. An Ngoc Van, Yoshimitsu Aoki |
IGARSS | 2 |
| 2001 | Computer aided system for orthognathic diagnosis utilizing 3D geometric head modelabstractIn this paper, we propose a computer aided diagnosis system that facilitates the drafting of treatment plans, the simulation of orthognathic surgery and the prediction of postoperative features. The three-dimensional (3D) display of the oral and maxillofacial region is very efficient for dentists in understanding a patient symptom and drafting an appropriate treatment plan. To achieve this, we propose a construction method for a 3D head model, which consists of soft and hard tissue. Utilizing feature points extracted from ortho-directional X-ray images and facial images, a standard head model can be modified to adapt an individual head shape. Using this head model, we can predict surgical modification of soft and hard tissue. Yoshimitsu Aoki, Masahiko Terajima, Yoshihiro Hoshino, Akihiko Nakasima, Shuji Hashimoto |
SMC | 1 |
| 2001 | Simulation of postoperative 3D facial morphology using a physics-based head model
Yoshimitsu Aoki, Shuji Hashimoto, Masahiko Terajima, Akihiko Nakasima |
Vis. Comput. | 1 |
| 1998 | Physical Facial Model Based on 3D-CT Data for Facial Image Analysis and Synthesis
Yoshimitsu Aoki, Shuji Hashimoto |
FG | 1 |