EDBT 2026 Demo / reviewers in the wild / expert
Takayuki Okatani
dblp:18/4811
· DBLP profile ↗
107ranked-venue papers
22as first author
33since 2021 · last 2026
0000-0001-9222-763XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 81 · 21 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 71 · 15 first-author · 19 since 2021Systems, architecture and hardware · 8 · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Visual Measurement and Uncertainty Prediction of Insulator Thickness in Insulated Rail JointsabstractRailway tracks are critical to social infrastructure, and their maintenance is essential for operational safety. Glued Insulated Rail Joints (GIRJs) are vital for railway signal control, where the insulator’s thickness helps prevent signal failures. This paper addresses detecting anomalies in GIRJ insulator thickness using images captured by devices on operational trains. Due to various factors, the insulator is not always clearly visible, and standard computer vision methods often struggle. In severe cases, judgments cannot be made from the image alone, requiring the system to return “unable to determine.” If these cases are rare, they can be manually inspected, still lowering overall inspection costs. We tackle this by framing the task as a one-dimensional regression problem, using convolutional neural networks (CNNs) to predict the boundary between the insulator and rail, while also estimating prediction uncertainty. Experiments with real-world data show that the model is accurate enough for practical use, even with challenging images. Additionally, we propose a robust method for detecting GIRJs in long-range railway images. This system is now operational in railway inspections. Ryohei Kasai, Takayuki Okatani |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2026 | RP-SLAM: Real-Time Photorealistic SLAM With Efficient 3D Gaussian Splattingabstract3D Gaussian Splatting has emerged as a promising technique for high-quality 3D rendering, leading to increasing interest in integrating 3DGS into realism SLAM systems. However, existing methods face challenges such as Gaussian primitives redundancy, forgetting problem during continuous optimization, and difficulty in initializing primitives in monocular case due to lack of depth information. In order to achieve efficient and photorealistic mapping, we propose RP-SLAM, a 3D Gaussian splatting-based vision SLAM method for monocular and RGB-D cameras. RP-SLAM decouples camera poses estimation from Gaussian primitives optimization and consists of three key components. Firstly, we propose an efficient incremental mapping approach to achieve a compact and accurate representation of the scene through adaptive sampling and Gaussian primitives filtering. Secondly, a dynamic window optimization method is proposed to mitigate the forgetting problem and improve map consistency. Finally, for the monocular case, a monocular keyframe initialization method based on sparse point cloud is proposed to improve the initialization accuracy of Gaussian primitives, which provides a geometric basis for subsequent optimization. The results of numerous experiments demonstrate that RP-SLAM achieves state-of-the-art map rendering accuracy while ensuring real-time performance and model compactness. Lizhi Bai, Chunqi Tian, Jun Yang 0056, Masanori Suganuma, Takayuki Okatani |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Action-Agnostic Point-Level Supervision for Temporal Action DetectionabstractWe propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the proposed scheme, a small portion of video frames is sampled in an unsupervised manner and presented to human annotators, who then label the frames with action categories. Unlike point-level supervision, which requires annotators to search for every action instance in an untrimmed video, frames to annotate are selected without human intervention in AAPL supervision. We also propose a detection model and learning method to effectively utilize the AAPL labels. Extensive experiments on the variety of datasets (THUMOS'14, FineAction, GTEA, BEOID, and ActivityNet 1.3) demonstrate that the proposed approach is competitive with or outperforms prior methods for video-level and point-level supervision in terms of the trade-off between the annotation cost and detection performance. Shuhei M. Yoshida, Takashi Shibata 0001, Makoto Terao, Takayuki Okatani, Masashi Sugiyama |
AAAI | 4 |
| 2025 | Self-Supervised Learning of Intertwined Content and Positional Features for Object DetectionabstractWe present a novel self-supervised feature learning method using Vision Transformers (ViT) as the backbone, specifically designed for object detection and instance segmentation. Our approach addresses the challenge of extracting features that capture both class and positional information, which are crucial for these tasks. The method introduces two key components: (1) a positional encoding tied to the cropping process in contrastive learning, which utilizes a novel vector field representation for positional embeddings; and (2) masking and prediction, similar to conventional Masked Image Modeling (MIM), applied in parallel to both content and positional embeddings of image patches. These components enable the effective learning of intertwined content and positional features. We evaluate our method against state-of-the-art approaches, pre-training on ImageNet-1K and fine-tuning on downstream tasks. Our method outperforms the state-of-the-art SSL methods on the COCO object detection benchmark, achieving significant improvements with fewer pre-training epochs. These results suggest that better integration of positional information into self-supervised learning can improve performance on the dense prediction tasks. Kang-Jun Liu, Masanori Suganuma, Takayuki Okatani |
ICML | 3 |
| 2025 | MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image RetrievalabstractResult diversification (RD) is a crucial technique in Text-to-Image Retrieval for enhancing the efficiency of a practical application. Conventional methods focus solely on increasing the diversity metric of image appearances. However, the diversity metric and its desired value vary depending on the application, which limits the applications of RD. This paper proposes a novel task called CDR-CA (Contextual Diversity Refinement of Composite Attributes). CDR-CA aims to refine the diversities of multiple attributes, according to the application's context. To address this task, we propose Multi-Source DPPs, a simple yet strong baseline that extends the Determinantal Point Process (DPP) to multi-sources. We model MS-DPP as a single DPP model with a unified similarity matrix based on a manifold representation. We also introduce Tangent Normalization to reflect contexts. Extensive experiments demonstrate the effectiveness of the proposed method. Naoya Sogi, Takashi Shibata 0001, Makoto Terao, Masanori Suganuma, Takayuki Okatani |
IJCAI | 5 |
| 2025 | Time-Frequency-Spatial Neural Architecture for Decoding Visual Signals from Macaque ECoGabstractUnderstanding how visual information is encoded in electrocorticography (ECoG) signals is essential for developing accurate and interpretable decoding models. In this study, we propose two novel approaches for multi-class visual classification based on ECoG data recorded from the inferior temporal cortex of macaque monkeys. The first model, MST-ECoGNet, combines traditional signal processing with neural networks by employing the Modified Stockwell Transform (MST) to map ECoG signals into a structured time-frequency-spatial domain. The second model, BiBand-3DECoGNet, replaces MST with a learnable convolutional module and utilizes a 3D spatial encoder to exploit the electrode array structure. Experimental results show that our models significantly outperform prior work, achieving up to a 12.87 percentage point improvement in classification accuracy while reducing model size by a factor of ten and increasing training speed by sixfold. Analysis of feature dimensions reveals that spatial and low-frequency components carry the most relevant information for visual decoding. These findings provide a foundation for further exploration of the neural mechanisms underlying visual object representation in the brain. Changqing Ji, Keisuke Kawasaki, Isao Hasegawa, Takayuki Okatani |
SMC | 4 |
| 2025 | Inverting the Generation Process of Denoising Diffusion Implicit Models: Empirical Evaluation and a Novel MethodabstractThis paper studies the problem of inverting the DDIM image generation process to recover latent variables, particularly the initial noise map, from a generated image. Existing methods often struggle with accuracy in this task. We propose a novel hybrid approach that combines direct inversion via gradient descent for the first step, followed by a fixed-point method for subsequent steps. Empirical evaluations across three datasets demonstrate that our method significantly improves the prediction of initial latent variables while achieving superior reconstruction accuracy. Additionally, we introduce a new evaluation, called the self-interpolation test, which assesses the quality of images generated from interpolated points between the true and predicted latent maps, offering deeper insights into performance. Our results reveal that while existing methods perform reasonably well in reconstruction, they consistently fail to accurately predict the initial latent variables, resulting in poor performance on the self-interpolation test. In contrast, our method outperforms all others across all metrics, providing valuable insights into diffusion models and enhancing their applications in image generation and editing. Masanori Suganuma, Takayuki Okatani |
WACV | 3 |
| 2025 | RefVSR++: Exploiting Reference Inputs for Reference-based Video Super-resolutionabstractSmartphones with multi-camera systems, featuring cameras with varying field-of-views (FoVs), are increasingly common. This variation in FoVs results in content differences across videos, paving the way for an innovative approach to video super-resolution (VSR). This method enhances the VSR performance of lower resolution (LR) videos by leveraging higher resolution reference (Ref) videos. Previous works [14, 15 ], which operate on this principle, generally expand on traditional VSR models by combining LR and Ref inputs over time into a unified stream. However, we can expect that better results are obtained by independently aggre-gating these Ref image sequences temporally. Therefore, we introduce an improved method, RefVSR++, which performs the parallel aggregation of LR and Ref images in the temporal direction, aiming to optimize the use of the available data. RefVSR++ also incorporates improved mechanisms for aligning image features over time, crucial for effective VSR. Our experiments demonstrate that RefVSR++ outper-forms previous works by over 1dB in PSNR, setting a new benchmark in the field. Han Zou, Masanori Suganuma, Takayuki Okatani |
WACV | 3 |
| 2025 | Rethinking Open-Set Object Detection: Issues, A New Formulation, and TaxonomyabstractAbstract Open-set object detection (OSOD), a task involving the detection of unknown objects while accurately detecting known objects, has recently gained attention. However, we identify a fundamental issue with the problem formulation employed in current OSOD studies. Inherent to object detection is knowing “what to detect,” which contradicts the idea of identifying “unknown” objects. This sets OSOD apart from open-set recognition (OSR). This contradiction complicates a proper evaluation of methods’ performance, a fact that previous studies have overlooked. Next, we propose a novel formulation wherein detectors are required to detect both known and unknown classes within specified super-classes of object classes. This new formulation is free from the aforementioned issues and has practical applications. Finally, we design benchmark tests utilizing existing datasets and report the experimental evaluation of existing OSOD methods. The results show that existing methods fail to accurately detect unknown objects due to misclassification of known and unknown classes rather than incorrect bounding box prediction. As a byproduct, we introduce a taxonomy of OSOD, resolving confusion prevalent in the literature. We anticipate that our study will encourage the research community to reconsider OSOD and facilitate progress in the right direction. Yusuke Hosoya, Masanori Suganuma, Takayuki Okatani |
Int. J. Comput. Vis. | 3 |
| 2025 | METS: Motion-Encoded Time-Surface for Event-Based High-Speed Pose Tracking
Ninghui Xu, Lihui Wang 0003, Zhiting Yao, Takayuki Okatani |
Int. J. Comput. Vis. | 4 |
| 2024 | Temporal Insight Enhancement: Mitigating Temporal Hallucination in Video Understanding by Multimodal Large Language Models
Li Sun 0007, Liuan Wang, Jun Sun 0004, Takayuki Okatani |
ICPR (7) | 4 |
| 2024 | Globalizing Local Features: Image Retrieval Using Shared Local Features with Pose Estimation for Faster Visual LocalizationabstractVisual localization is an important sub-task in SfM and visual SLAM that involves estimating a 6-DoF camera pose for an input query image relative to a given 3D model of the environment. The most accurate approach is a hierarchical one that splits the task into two stages: image retrieval and camera pose estimation. Each stage requires different image features, with global features compactly encoding holistic image information for the first stage and local features encoding the appearance around salient image points for the second stage. While existing methods use independent networks to extract these features, one for global and one for local, this strategy is suboptimal in terms of computational efficiency. In this paper, we propose a novel approach that achieves state-of-the-art inference accuracy with significantly improved efficiency. Our approach’s core component is SuperGF, a network that aggregates local features optimized for camera pose estimation to create a global feature that enables precise image retrieval. Through extensive experiments on the standard benchmark tests, we demonstrate that the method offers a better trade-off between accuracy and computational cost. Wenzheng Song, Boshu Lei, Takayuki Okatani |
ICRA | 4 |
| 2024 | SBCFormer: Lightweight Network Capable of Full-size ImageNet Classification at 1 FPS on Single Board ComputersabstractComputer vision has become increasingly prevalent in solving real-world problems across diverse domains, including smart agriculture, fishery, and livestock management. These applications may not require processing many image frames per second, leading practitioners to use single board computers (SBCs). Although many lightweight networks have been developed for "mobile/edge" devices, they primarily target smartphones with more powerful processors and not SBCs with the low-end CPUs. This paper introduces a CNN-ViT hybrid network called SBCFormer, which achieves high accuracy and fast computation on such low-end CPUs. The hardware constraints of these CPUs make the Transformer’s attention mechanism preferable to convolution. However, using attention on low-end CPUs presents a challenge: high-resolution internal feature maps demand excessive computational resources, but reducing their resolution results in the loss of local image details. SBCFormer introduces an architectural design to address this issue. As a result, SBCFormer achieves the highest trade-off between accuracy and speed on a Raspberry Pi 4 Model B with an ARM-Cortex A72 CPU. For the first time, it achieves an ImageNet-1K top-1 accuracy of around 80% at a speed of 1.0 frame/sec on the SBC. Code is available at https://github.com/xyongLu/SBCFormer. Xiangyong Lu, Masanori Suganuma, Takayuki Okatani |
WACV | 3 |
| 2024 | Appearance-Based Curriculum for Semi-Supervised Learning with Multi-Angle Unlabeled DataabstractWe propose an appearance-based curriculum (ABC) for a semi-supervised learning scenario where labeled images taken from limited angles and unlabeled ones taken from various angles are available for training. A common approach to semi-supervised learning relies on pseudo-labeling and data augmentation, but it struggles with large visual variations that cannot be covered by data augmentation. To solve this problem, ABC incrementally expands the pool of unlabeled images fed to a base semi-supervised learner so that newly added data are the ones most similar to those already in the pool. This way, the learner can assign pseudo-labels to the new data with high accuracy, keeping the quality of pseudo-labels higher than that when all the unlabeled data are processed at once, as customarily done in existing semi-supervised learning methods. We conducted extensive experiments and confirmed that our method outperforms the state-of-the-art semi-supervised learning methods in our scenario. Shuhei M. Yoshida, Takashi Shibata 0001, Makoto Terao, Takayuki Okatani, Masashi Sugiyama |
WACV | 5 |
| 2024 | Contextual Affinity Distillation for Image Anomaly DetectionabstractPrevious studies on unsupervised industrial anomaly detection mainly focus on ‘structural’ types of anomalies such as cracks and color contamination by matching or learning local feature representations. While achieving significantly high detection performance on this kind of anomaly, they are faced with ‘logical’ types of anomalies that violate the long-range dependencies such as a normal object placed in the wrong position. Noting the reverse distillation approaches that are under the encoder-decoder paradigm could learn from the high abstract level knowledge, we propose to use two students (local and global) to better mimic the teacher’s local and global behavior in reverse distillation. The local student, which is used in previous studies mainly focuses on accurate local feature learning while the global student pays attention to learning global correlations. To further encourage the global student’s learning to capture long-range dependencies, we design the global context condensing block (GCCB) and propose a contextual affinity loss for the student training and anomaly scoring. Experimental results show that the proposed method sets a new state-of-the-art performance on the MVTec LOCO AD dataset without using complex training techniques. Masanori Suganuma, Takayuki Okatani |
WACV | 3 |
| 2024 | Improved high dynamic range imaging using multi-scale feature flows balanced between task-orientedness and accuracyabstractDeep learning has made it possible to accurately generate high dynamic range (HDR) images from multiple images taken at different exposure settings, largely owing to advancements in neural network design. However, generating images without artifacts remains difficult, especially in scenes with moving objects. In such cases, issues like color distortion, geometric misalignment, or ghosting can appear. Current state-of-the-art network designs address this by estimating the optical flow between input images to align them better. The parameters for the flow estimation are learned through the primary goal, producing high-quality HDR images. However, we find that this ”task-oriented flow” approach has its drawbacks, especially in minimizing artifacts. To address this, we introduce a new network design and training method that improve the accuracy of flow estimation. This aims to strike a balance between task-oriented flow and accurate flow. Additionally, the network utilizes multi-scale features extracted from the input images for both flow estimation and HDR image reconstruction. Our experiments demonstrate that these two innovations result in HDR images with fewer artifacts and enhanced quality. Masanori Suganuma, Takayuki Okatani |
Comput. Vis. Image Underst. | 3 |
| 2024 | Symmetry-aware Neural Architecture for Embodied Visual NavigationabstractAbstract The existing methods for addressing visual navigation employ deep reinforcement learning as the standard tool for the task. However, they tend to be vulnerable to statistical shifts between the training and test data, resulting in poor generalization over novel environments that are out-of-distribution from the training data. In this study, we attempt to improve the generalization ability by utilizing the inductive biases available for the task. Employing the active neural SLAM that learns policies with the advantage actor-critic method as the base framework, we first point out that the mappings represented by the actor and the critic should satisfy specific symmetries. We then propose a network design for the actor and the critic to inherently attain these symmetries. Specifically, we use G-convolution instead of the standard convolution and insert the semi-global polar pooling layer, which we newly design in this study, in the last section of the critic network. Our method can be integrated into existing methods that utilize intermediate goals and 2D occupancy maps. Experimental results show that our method improves generalization ability by a good margin over visual exploration and object goal navigation, which are two main embodied visual navigation tasks. Shuang Liu 0002, Masanori Suganuma, Takayuki Okatani |
Int. J. Comput. Vis. | 3 |
| 2024 | That's BAD: blind anomaly detection by implicit local feature clusteringabstractAbstract Recent studies on visual anomaly detection (AD) of industrial objects/textures have achieved quite good performance. They consider an unsupervised setting, specifically the one-class setting, in which we assume the availability of a set of normal (i.e., anomaly-free) images for training. In this paper, we consider a more challenging scenario of unsupervised AD, in which we detect anomalies in a given set of images that might contain both normal and anomalous samples. The setting does not assume the availability of known normal data and thus is completely free from human annotation, which differs from the standard AD considered in recent studies. For clarity, we call the setting blind anomaly detection (BAD). We show that BAD can be converted into a local outlier detection problem and propose a novel method named PatchCluster that can accurately detect image- and pixel-level anomalies. Experimental results show that PatchCluster shows a promising performance without the knowledge of normal data, even comparable to the SOTA methods applied in the one-class setting needing it. Masanori Suganuma, Takayuki Okatani |
Mach. Vis. Appl. | 3 |
| 2024 | Rethinking unsupervised domain adaptation for semantic segmentationabstractUnsupervised domain adaptation (UDA) adapts a model trained on one domain (called source) to a novel domain (called target) using only unlabeled data. Due to its high annotation cost, researchers have developed many UDA methods for semantic segmentation, which assume no labeled sample is available in the target domain. We question the practicality of this assumption for two reasons. First, after training a model with a UDA method, we must somehow verify the model before deployment. Second, UDA methods have at least a few hyper-parameters that need to be determined. The surest solution to these is to evaluate the model using validation data, i.e., a certain amount of labeled target-domain samples. This question about the basic assumption of UDA leads us to rethink UDA from a data-centric point of view. Specifically, we assume we have access to a minimum level of labeled data. Then, we ask how much is necessary to find good hyper-parameters of existing UDA methods. We then consider what if we use the same data for supervised training of the same model, e.g., finetuning. We conducted experiments to answer these questions with popular scenarios, {GTA5, SYNTHIA} → Cityscapes. We found that i) choosing good hyper-parameters needs only a few labeled images for some UDA methods whereas a lot more for others; and ii) simple finetuning works surprisingly well; it outperforms many UDA methods if only several dozens of labeled images are available. • We rethink the UDA for segmentation from a data-centric perspective. • Our starting point is that any ML system requires annotated data for validation. • We investigate how much data is necessary to select parameters of existing methods. • We consider what if we use the same data for supervised training of the same model. Masanori Suganuma, Takayuki Okatani |
Pattern Recognit. Lett. | 3 |
| 2023 | Prompt Prototype Learning Based on Ranking Instruction For Few-Shot Visual TasksabstractQuerying large language models (LLMs), such as GPT-3, for high-quality prompts and utilizing pre-trained vision-language models, such as CLIP, to construct a zero-shot visual classification model, offer promising performance across various downstream visual tasks. However, when applied to specific domains, their efficacy is restricted due to the gap between the general prompts they generate and the required domain-specific knowledge. In this paper, we propose a novel, lightweight method for prompt prototype learning through ranking instruction, specifically designed to bridge this gap in the context of few-shot visual classification. We generate domain-specific prompts leveraging the knowledge contained in LLMs and then fine-tune the prompt prototype with effective ranking instructions from several domain images. Our few-shot experiments on facial expression benchmarks demonstrate the efficacy of the prompt prototype. Notably, our method delivers results that are on par with state-of-the-art few-shot image classification techniques and can be integrated with them to further improve performance in the facial expression domain. Our approach provides a promising solution to few-shot visual classification, making use of the knowledge contained in LLMs to generate domain-specific prompts. Li Sun 0007, Liuan Wang, Jun Sun 0004, Takayuki Okatani |
ICIP | 4 |
| 2023 | Accurate Single-Image Defocus Deblurring Based on Improved Integration with Defocus Map EstimationabstractThis paper considers the problem of single-image defocus deblurring, which involves removing blur in an input image caused by defocusing. Previous studies have employed two main approaches, the first being a two-step approach involving estimating the defocus map from the input image and then computing the blur kernel from it, followed by non-blind deconvolution to obtain the estimate of the clean image. The second approach is a direct method where the clean image is estimated directly from the blurry input image. The paper proposes an intermediate approach that explicitly estimates the defocus map of the scene but does not explicitly compute the kernel or its inverse. Instead, it attempts to learn a direct mapping from the blurry input image to the clean image by utilizing the estimated defocus map to condition the mapping. Experimental results show that the proposed method can yield higher quality outputs than the state-of-the-art methods. Masanori Suganuma, Takayuki Okatani |
ICIP | 3 |
| 2023 | Network Pruning and Fine-tuning for Few-shot Industrial Image Anomaly DetectionabstractThis paper focuses on industrial image anomaly detection and localization under few-shot settings. Since acquiring sufficient anomalous data is difficult, unsupervised learning that uses only normal data is commonly used, but even obtaining enough anomaly-free training samples can be challenging. Moreover, applying data augmentations, which is a common strategy for few-shot learning to alleviate the lack of data, is limited to use for some industrial product images. To address the above issues, we propose a network pruning and fine-tuning (PF) framework that leverages the knowledge of a deep pre-trained model. Our approach distills the knowledge of normal samples into a pruned student network, followed by fine-tuning to restore its representation ability for normal data. During inference, discrepancies between features extracted by the teacher and student are used to determine the anomaly score. The proposed method could better utilize the strong representation ability of deep models and benefit the student training with limited data by network pruning. Our framework achieves state-of-the-art performance on the MVTec AD benchmark and is not limited to specific network pruning methods. Masanori Suganuma, Takayuki Okatani |
INDIN | 3 |
| 2023 | Zero-shot versus Many-shot: Unsupervised Texture Anomaly DetectionabstractResearch on unsupervised anomaly detection (AD) has recently progressed, significantly increasing detection accuracy. This paper focuses on texture images and considers how few normal samples are needed for accurate AD. We first highlight the critical nature of the problem that previous studies have overlooked: accurate detection gets harder for anisotropic textures when image orientations are not aligned between inputs and normal samples. We then propose a zero-shot method, which detects anomalies without using a normal sample. The method is free from the issue of unaligned orientation between input and normal images. It assumes the input texture to be homogeneous, detecting image regions that break the homogeneity as anomalies. We present a quantitative criterion to judge whether this assumption holds for an input texture. Experimental results show the broad applicability of the proposed zero-shot method and its good performance comparable to or even higher than the state-of-the-art methods using hundreds of normal samples. The code and data are available from https://drive.google.com/drive/folders/10OyPzvI3H6llCZBxKxFlKWt1Pw1tkMK1. Toshimichi Aota, Lloyd Teh Tzer Tong, Takayuki Okatani |
WACV | 3 |
| 2023 | Unsupervised domain adaptation for semantic segmentation via cross-region alignmentabstractSemantic segmentation requires a lot of training data, which necessitates costly annotation. There have been many studies on unsupervised domain adaptation (UDA) from one domain to another, e.g., from computer graphics to real images. However, there is still a gap in accuracy between UDA and supervised training on native domain data. It is arguably attributable to the class-level misalignment between the source and target domain data. To cope with this, we propose a method that applies adversarial training to align two feature distributions in the target domain. It uses a self-training framework to split the image into two regions (i.e., trusted and untrusted), which form two distributions to align in the feature space. We term this approach cross-region adaptation (CRA) to distinguish it from the previous methods of aligning different domain distributions, which we call cross-domain adaptation (CDA). CRA can be applied after any CDA method. Experimental results show that this always improves the accuracy of the combined CDA method. Xing Liu 0010, Masanori Suganuma, Takayuki Okatani |
Comput. Vis. Image Underst. | 4 |
| 2023 | Zero-shot temporal event localisation: Label-free, training-free, domain-freeabstractAbstract Temporal event localisation (TEL) has recently attracted increasing attention due to the rapid development of video platforms. Existing methods are based on either fully/weakly supervised or unsupervised learning, and thus they rely on expensive data annotation and time‐consuming training. Moreover, these models, which are trained on specific domain data, limit the model generalisation to data distribution shifts. To cope with these difficulties, the authors propose a zero‐shot TEL method that can operate without training data or annotations. Leveraging large‐scale vision and language pre‐trained models, for example, CLIP, we solve the two key problems: (1) how to find the relevant region where the event is likely to occur; (2) how to determine event duration after we find the relevant region. Query guided optimisation for local frame relevance relying on the query‐to‐frame relationship is proposed to find the most relevant frame region where the event is most likely to occur. Proposal generation method relying on the frame‐to‐frame relationship is proposed to determine the event duration. The authors also propose a greedy event sampling strategy to predict multiple durations with high reliability for the given event. The authors’ methodology is unique, offering a label‐free, training‐free, and domain‐free approach. It enables the application of TEL purely at the testing stage. The practical results show it achieves competitive performance on the standard Charades‐STA and ActivityCaptions datasets. Li Sun 0007, Ping Wang 0034, Liuan Wang, Jun Sun 0004, Takayuki Okatani |
IET Comput. Vis. | 5 |
| 2022 | Bright as the Sun: In-depth Analysis of Imagination-Driven Image Captioning
Huyen Thi Thanh Tran, Takayuki Okatani |
ACCV (4) | 2 |
| 2022 | Symmetry-aware Neural Architecture for Embodied Visual ExplorationabstractVisual exploration is a task that seeks to visit all the navigable areas of an environment as quickly as possible. The existing methods employ deep reinforcement learning (RL) as the standard tool for the task. However, they tend to be vulnerable to statistical shifts between the training and test data, resulting in poor generalization over novel environments that are out-of-distribution (OOD) from the training data. In this paper, we attempt to improve the generalization ability by utilizing the inductive biases available for the task. Employing the active neural SLAM (ANS) that learns exploration policies with the advantage actor-critic (A2C) method as the base framework, we first point out that the mappings represented by the actor and the critic should satisfy specific symmetries. We then propose a network design for the actor and the critic to inherently attain these symmetries. Specifically, we use G-convolution instead of the standard convolution and insert the semi-global polar pooling (SGPP) layer, which we newly design in this study, in the last section of the critic network. Experimental results show that our method increases area coverage by 8.1m2when trained on the Gibson dataset and tested on the Matterport3D dataset, establishing the new state-of-the-art. Shuang Liu 0002, Takayuki Okatani |
CVPR | 2 |
| 2022 | GRIT: Faster and Better Image Captioning Transformer Using Dual Visual Features
Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani |
ECCV (36) | 3 |
| 2022 | Bridging the Gap from Asymmetry Tricks to Decorrelation Principles in Non-contrastive Self-supervised LearningabstractRecent non-contrastive methods for self-supervised representation learning show promising performance. While they are attractive since they do not need negative samples, it necessitates some mechanism to avoid collapsing into a trivial solution. Currently, there are two approaches to collapse prevention. One uses an asymmetric architecture on a joint embedding of input, e.g., BYOL and SimSiam, and the other imposes decorrelation criteria on the same joint embedding, e.g., Barlow-Twins and VICReg. The latter methods have theoretical support from information theory as to why they can learn good representation. However, it is not fully understood why the former performs equally well. In this paper, focusing on BYOL/SimSiam, which uses the stop-gradient and a predictor as asymmetric tricks, we present a novel interpretation of these tricks; they implicitly impose a constraint that encourages feature decorrelation similar to Barlow-Twins/VICReg. We then present a novel non-contrastive method, which replaces the stop-gradient in BYOL/SimSiam with the derived constraint; the method empirically shows comparable performance to the above SOTA methods in the standard benchmark test using ImageNet. This result builds a bridge from BYOL/SimSiam to the decorrelation-based methods, contributing to demystifying their secrets. Kang-Jun Liu, Masanori Suganuma, Takayuki Okatani |
NeurIPS | 3 |
| 2021 | Matching in the Dark: A Dataset for Matching Image Pairs of Low-light ScenesabstractThis paper considers matching images of low-light scenes, aiming to widen the frontier of SfM and visual SLAM applications. Recent image sensors can record the brightness of scenes with more than eight-bit precision, available in their RAW-format image. We are interested in making full use of such high-precision information to match extremely low-light scene images that conventional methods cannot handle. For extreme low-light scenes, even if some of their brightness information exists in the RAW format images’ low bits, the standard raw image processing on cameras fails to utilize them properly. As was recently shown by Chen et al. [14], CNNs can learn to produce images with a natural appearance from such RAW-format images. To consider if and how well we can utilize such information stored in RAW-format images for image matching, we have created a new dataset named MID (matching in the dark). Using it, we experimentally evaluated combinations of eight image-enhancing methods and eleven image matching methods consisting of classical/neural local descriptors and classical/neural initial point-matching methods. The results show the advantage of using the RAW-format images and the strengths and weaknesses of the above component methods. They also imply there is room for further research. Wenzheng Song, Masanori Suganuma, Xing Liu 0010, Noriyuki Shimobayashi, Daisuke Maruta, Takayuki Okatani |
ICCV | 6 |
| 2021 | Learning to Bundle-adjust: A Graph Network Approach to Faster Optimization of Bundle Adjustment for Vehicular SLAMabstractBundle adjustment (BA) occupies a large portion of the execution time of SfM and visual SLAM. Local BA over the latest several keyframes plays a crucial role in visual SLAM. Its execution time should be sufficiently short for robust tracking; this is especially critical for embedded systems with a limited computational resource. This study proposes a learning-based bundle adjuster using a graph network. It works faster and can be used instead of conventional optimization-based BA. The graph network operates on a graph consisting of the nodes of keyframes and landmarks and the edges representing the landmarks’ visibility. The graph network receives the parameters’ initial values as inputs and predicts their updates to the optimal values. It internally uses an intermediate representation of inputs which we design inspired by the normal equation of the Levenberg-Marquardt method. It is trained using the sum of reprojection errors as a loss function. The experiments show that the proposed method outputs parameter estimates with slightly inferior accuracy in 1/60–1/10 of time compared with the conventional BA. Tetsuya Tanaka, Yukihiro Sasagawa, Takayuki Okatani |
ICCV | 3 |
| 2021 | Look Wide and Interpret Twice: Improving Performance on Interactive Instruction-following TasksabstractThere is a growing interest in the community in making an embodied AI agent perform a complicated task while interacting with an environment following natural language directives. Recent studies have tackled the problem using ALFRED, a well-designed dataset for the task, but achieved only very low accuracy. This paper proposes a new method, which outperforms the previous methods by a large margin. It is based on a combination of several new ideas. One is a two-stage interpretation of the provided instructions. The method first selects and interprets an instruction without using visual information, yielding a tentative action sequence prediction. It then integrates the prediction with the visual information etc., yielding the final prediction of an action and an object. As the object's class to interact is identified in the first stage, it can accurately select the correct object from the input image. Moreover, our method considers multiple egocentric views of the environment and extracts essential information by applying hierarchical attention conditioned on the current instruction. This contributes to the accurate prediction of actions for navigation. A preliminary version of the method won the ALFRED Challenge 2020. The current version achieves the unseen environment's success rate of 4.45% with a single view, which is further improved to 8.37% with multiple views. Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani |
IJCAI | 3 |
| 2021 | Progressive and Selective Fusion Network for High Dynamic Range ImagingabstractThis paper considers the problem of generating an HDR image of a scene from its LDR images. Recent studies employ deep learning and solve the problem in an end-to-end fashion, leading to significant performance improvements. However, it is still hard to generate a good quality image from LDR images of a dynamic scene captured by a hand-held camera, e.g., occlusion due to the large motion of foreground objects, causing ghosting artifacts. The key to success relies on how well we can fuse the input images in their feature space, where we wish to remove the factors leading to low-quality image generation while performing the fundamental computations for HDR image generation, e.g., selecting the best-exposed image/region. We propose a novel method that can better fuse the features based on two ideas. One is multi-step feature fusion; our network gradually fuses the features in a stack of blocks having the same structure. The other is the design of the component block that effectively performs two operations essential to the problem, i.e., comparing and selecting appropriate images/regions. Experimental results show that the proposed method outperforms the previous state-of-the-art methods on the standard benchmark tests. Jun Xiao 0010, Kin-Man Lam 0001, Takayuki Okatani |
ACM Multimedia | 4 |
| 2020 | Hyperparameter-Free Out-of-Distribution Detection Using Cosine Similarity
Engkarat Techapanurak, Masanori Suganuma, Takayuki Okatani |
ACCV (4) | 3 |
| 2020 | Efficient Attention Mechanism for Visual Dialog that Can Handle All the Interactions Between Multiple Inputs
Van-Quang Nguyen, Masanori Suganuma, Takayuki Okatani |
ECCV (24) | 3 |
| 2020 | Analysis and a Solution of Momentarily Missed Detection for Anchor-based Object DetectorsabstractThe employment of convolutional neural networks has led to significant performance improvement on the task of object detection. However, when applying existing detectors to continuous frames in a video, we often encounter momentary miss-detection of objects, that is, objects are undetected exceptionally at a few frames, although they are correctly detected at all other frames. In this paper, we analyze the mechanism of how such miss-detection occurs. For the most popular class of detectors that are based on anchor boxes, we show the followings: i) besides apparent causes such as motion blur, occlusions, background clutters, etc., the majority of remaining miss-detection can be explained by an improper behavior of the detectors at boundaries of the anchor boxes; and ii) this can be rectified by improving the way of choosing positive samples from candidate anchor boxes when training the detectors. Yusuke Hosoya, Masanori Suganuma, Takayuki Okatani |
WACV | 3 |
| 2020 | Toward Explainable Fashion RecommendationabstractMany studies have been conducted so far to build systems for recommending fashion items and outfits. Although they achieve good performances in their respective tasks, most of them cannot explain their judgments to the users, which compromises their usefulness. Toward explainable fashion recommendation, this study proposes a system that is able not only to provide a goodness score for an outfit but also to explain the score by providing reason behind it. For this purpose, we propose a method for quantifying how influential each feature of each item is to the score. Using this influence value, we can identify which item and what feature make the outfit good or bad. We represent the image of each item with a combination of human-interpretable features, and thereby the identification of the most influential item-feature pair gives useful explanation of the output score. To evaluate the performance of this approach, we design an experiment that can be performed without human annotation; we replace a single item-feature pair in an outfit so that the score will decrease, and then we test if the proposed method can detect the replaced item-feature pair correctly using the above influence values. The experimental results show that the proposed method can accurately detect bad items in outfits lowering their scores. Pongsate Tangseng, Takayuki Okatani |
WACV | 2 |
| 2020 | Extending information maximization from a rate-distortion perspective
Yan Zhang 0055, Junjie Hu 0003, Takayuki Okatani |
Neurocomputing | 3 |
| 2019 | Deep Learning for Natural Image Reconstruction from Electrocorticography SignalsabstractSeveral recent studies proposed various methods for reconstructing natural images from human functional magnetic resonance imaging (fMRI) data. However, few studies have proposed reconstruction methods for electrophysiolgical brain activities such as electroencephalography (EEG) and electrocorticography (ECoG). To investigate whether natural images can be reconstructed from electrophysiological brain activities, we conducted a large-scale experiment on natural image reconstruction from ECoG signals using deep learning. We first recorded ECoG signals from two macaque monkeys while presenting diverse natural images. Then, we trained several deep learning models for reconstructing presented images from ECoG signals. Comparing reconstruction models, we find that models trained with an adversarial loss produced reconstructions that contain visible features in presented images. Furthermore, our results with downsampled ECoG signals show the importance of rich temporal dynamics in ECoG signals for image reconstruction. Our results indicate the possibility of reconstructing diverse natural images from electrophysiological brain activities using deep learning. Hiroto Date, Keisuke Kawasaki, Isao Hasegawa, Takayuki Okatani |
BIBM | 4 |
| 2019 | Dual Residual Networks Leveraging the Potential of Paired Operations for Image RestorationabstractIn this paper, we study design of deep neural networks for tasks of image restoration. We propose a novel style of residual connections dubbed "dual residual connection", which exploits the potential of paired operations, e.g., up- and down-sampling or convolution with large- and small-size kernels. We design a modular block implementing this connection style; it is equipped with two containers to which arbitrary paired operations are inserted. Adopting the "unraveled" view of the residual networks proposed by Veit et al., we point out that a stack of the proposed modular blocks allows the first operation in a block interact with the second operation in any subsequent blocks. Specifying the two operations in each of the stacked blocks, we build a complete network for each individual task of image restoration. We experimentally evaluate the proposed approach on five image restoration tasks using nine datasets. The results show that the proposed networks with properly chosen paired operations outperform previous methods on almost all of the tasks and datasets. Xing Liu 0010, Masanori Suganuma, Zhun Sun, Takayuki Okatani |
CVPR | 4 |
| 2019 | Multi-Task Learning of Hierarchical Vision-Language RepresentationabstractIt is still challenging to build an AI system that can perform tasks that involve vision and language at human level. So far, researchers have singled out individual tasks separately, for each of which they have designed networks and trained them on its dedicated datasets. Although this approach has seen a certain degree of success, it comes with difficulties of understanding relations among different tasks and transferring the knowledge learned for a task to others. We propose a multi-task learning approach that enables to learn vision-language representation that is shared by many tasks from their diverse datasets. The representation is hierarchical, and prediction for each task is computed from the representation at its corresponding level of the hierarchy. We show through experiments that our method consistently outperforms previous single-task-learning methods on image caption retrieval, visual question answering, and visual grounding. We also analyze the learned hierarchical representation by visualizing attention maps generated in our network. Duy-Kien Nguyen, Takayuki Okatani |
CVPR | 2 |
| 2019 | Attention-Based Adaptive Selection of Operations for Image Restoration in the Presence of Unknown Combined DistortionsabstractMany studies have been conducted so far on image restoration, the problem of restoring a clean image from its distorted version. There are many different types of distortion affecting image quality. Previous studies have focused on single types of distortion, proposing methods for removing them. However, image quality degrades due to multiple factors in the real world. Thus, depending on applications, e.g., vision for autonomous cars or surveillance cameras, we need to be able to deal with multiple combined distortions with unknown mixture ratios. For this purpose, we propose a simple yet effective layer architecture of neural networks. It performs multiple operations in parallel, which are weighted by an attention mechanism to enable selection of proper operations depending on the input. The layer can be stacked to form a deep network, which is differentiable and thus can be trained in an end-to-end fashion by gradient descent. The experimental results show that the proposed method works better than previous methods by a good margin on tasks of restoring images with multiple combined distortions. Masanori Suganuma, Xing Liu 0010, Takayuki Okatani |
CVPR | 3 |
| 2019 | Improving Head Pose Estimation with a Combined Loss and Bounding Box Margin AdjustmentabstractWe address a problem of estimating pose of a person's head from its RGB image. The employment of CNNs for the problem has contributed to significant improvement in accuracy in recent works. However, we show that the following two methods, despite their simplicity, can attain further improvement: (i) proper adjustment of the margin of bounding box of a detected face, and (ii) choice of loss functions. We show that the integration of these two methods achieve the new state-of-the-art on standard benchmark datasets for in-the-wild head pose estimation. The Tensorflow implementation of our work is available at https://github.com/MingzhenShao/HeadPose. Mingzhen Shao, Zhun Sun, Mete Ozay, Takayuki Okatani |
FG | 4 |
| 2019 | Visualization of Convolutional Neural Networks for Monocular Depth EstimationabstractRecently, convolutional neural networks (CNNs) have shown great success on the task of monocular depth estimation. A fundamental yet unanswered question is: how CNNs can infer depth from a single image. Toward answering this question, we consider visualization of inference of a CNN by identifying relevant pixels of an input image to depth estimation. We formulate it as an optimization problem of identifying the smallest number of image pixels from which the CNN can estimate a depth map with the minimum difference from the estimate from the entire image. To cope with a difficulty with optimization through a deep CNN, we propose to use another network to predict those relevant image pixels in a forward computation. In our experiments, we first show the effectiveness of this approach, and then apply it to different depth estimation networks on indoor and outdoor scene datasets. The results provide several findings that help exploration of the above question. Junjie Hu 0003, Yan Zhang 0055, Takayuki Okatani |
ICCV | 3 |
| 2019 | A Generative Model of Underwater Images for Active Landmark Detection and DockingabstractUnderwater active landmarks (UALs) are widely used for short-range underwater navigation in underwater robotics tasks. Detection of UALs is challenging due to large variance of underwater illumination, water quality and change of camera viewpoint. Moreover, improvement of detection accuracy relies upon statistical diversity of images used to train detection models. We propose a generative adversarial network, called Tank-to-field GAN (T2FGAN), to learn generative models of underwater images, and use the learned models for data augmentation to improve detection accuracy. To this end, first a T2FGAN is trained using images of UALs captured in a tank. Then, the learned model of the T2FGAN is used to generate images of UALs according to different water quality, illumination, pose and landmark configurations (WIPCs). In experimental analyses, we first explore statistical properties of images of UALs generated by T2FGAN under various WIPCs for active landmark detection. Then, we use the generated images for training detection algorithms. Experimental results show that training detection algorithms using the generated images can improve detection accuracy. In field experiments, underwater docking tasks are successfully performed in a lake by employing detection models trained on datasets generated by T2FGAN. Shuang Liu 0002, Mete Ozay, Takayuki Okatani |
IROS | 5 |
| 2019 | Revisiting Single Image Depth Estimation: Toward Higher Resolution Maps With Accurate Object BoundariesabstractThis paper considers the problem of single image depth estimation. The employment of convolutional neural networks (CNNs) has recently brought about significant advancements in the research of this problem. However, most existing methods suffer from loss of spatial resolution in the estimated depth maps; a typical symptom is distorted and blurry reconstruction of object boundaries. In this paper, toward more accurate estimation with a focus on depth maps with higher spatial resolution, we propose two improvements to existing approaches. One is about the strategy of fusing features extracted at different scales, for which we propose an improved network architecture consisting of four modules: an encoder, decoder, multi-scale feature fusion module, and refinement module. The other is about loss functions for measuring inference errors used in training. We show that three loss terms, which measure errors in depth, gradients and surface normals, respectively, contribute to improvement of accuracy in an complementary fashion. Experimental results show that these two improvements enable to attain higher accuracy than the current state-of-the-arts, which is given by finer resolution reconstruction, for example, with small objects and object boundaries. Junjie Hu 0003, Mete Ozay, Yan Zhang 0055, Takayuki Okatani |
WACV | 4 |
| 2019 | Video-Rate Video InpaintingabstractThis paper considers the problem of video inpainting, i.e., to remove specified objects from an input video. Many methods have been developed for the problem so far, in which there is a trade-off between image quality and computational time. There was no method that can generate high-quality images in video rate. The key to video inpainting is how to establish correspondences from scene regions occluded in a frame to those observed in other frames. To break the trade-off, we propose to use CNNs as a solution to this key problem. We extend existing CNNs for the standard task of optical flow estimation to be able to estimate the flow of occluded background regions. The extension includes augmentation of their architecture and changes of their training method. We experimentally show that this approach works well despite its simplicity, and that a simple video inpainting method integrating this flow estimator runs in video rate (e.g., 32 fps for 832 × 448 pixel videos on a standard PC with a GPU) while achieving image quality close to the state-of-the-art. Rito Murase, Yan Zhang 0055, Takayuki Okatani |
WACV | 3 |
| 2018 | Training CNNs With Normalized KernelsabstractSeveral methods of normalizing convolution kernels have been proposed in the literature to train convolutional neural networks (CNNs), and have shown some success. However, our understanding of these methods has lagged behind their success in application; there are a lot of open questions, such as why a certain type of kernel normalization is effective and what type of normalization should be employed for each (e.g., higher or lower) layer of a CNN. As the first step towards answering these questions, we propose a framework that enables us to use a variety of kernel normalization methods at any layer of a CNN. A naive integration of kernel normalization with a general optimization method, such as SGD, often entails instability while updating parameters. Thus, existing methods employ ad-hoc procedures to empirically assure convergence. In this study, we pose estimation of convolution kernels under normalization constraints as constraint-free optimization on kernel submanifolds that are identified by the employed constraints. Note that naive application of the established optimization methods for matrix manifolds to the aforementioned problems is not feasible because of the hierarchical nature of CNNs. To this end, we propose an algorithm for optimization on kernel manifolds in CNNs by appropriate scaling of the space of kernels based on structure of CNNs and statistics of data. We theoretically prove that the proposed algorithm has assurance of almost sure convergence to a solution at single minimum. Our experimental results show that the proposed method can successfully train popular CNN models using several different types of kernel normalization methods. Moreover, they show that the proposed method improves classification performance of baseline CNNs, and provides state-of-the-art performance for major image classification benchmarks. Mete Ozay, Takayuki Okatani |
AAAI | 2 |
| 2018 | Improved Fusion of Visual and Language Representations by Dense Symmetric Co-Attention for Visual Question AnsweringabstractA key solution to visual question answering (VQA) exists in how to fuse visual and language features extracted from an input image and question. We show that an attention mechanism that enables dense, bi-directional interactions between the two modalities contributes to boost accuracy of prediction of answers. Specifically, we present a simple architecture that is fully symmetric between visual and language representations, in which each question word attends on image regions and each image region attends on question words. It can be stacked to form a hierarchy for multi-step interactions between an image-question pair. We show through experiments that the proposed architecture achieves a new state-of-the-art on VQA and VQA 2.0 despite its small size. We also present qualitative evaluation, demonstrating how the proposed attention mechanism can generate reasonable attention maps on images and questions, which leads to the correct answer prediction. Duy-Kien Nguyen, Takayuki Okatani |
CVPR | 2 |
| 2018 | Feature Quantization for Defending Against Distortion of ImagesabstractIn this work, we address the problem of improving robustness of convolutional neural networks (CNNs) to image distortion. We argue that higher moment statistics of feature distributions can be shifted due to image distortion, and the shift leads to performance decrease and cannot be reduced by ordinary normalization methods as observed in our experimental analyses. In order to mitigate this effect, we propose an approach base on feature quantization. To be specific, we propose to employ three different types of additional non-linearity in CNNs: i) a floor function with scalable resolution, ii) a power function with learnable exponents, and iii) a power function with data-dependent exponents. In the experiments, we observe that CNNs which employ the proposed methods obtain better performance in both generalization performance and robustness for various distortion types for large scale benchmark datasets. For instance, a ResNet-50 model equipped with the proposed method (+HPOW) obtains 6.95%, 5.26% and 5.61% better accuracy on the ILSVRC-12 classification tasks using images distorted with motion blur, salt and pepper and mixed distortions. Zhun Sun, Mete Ozay, Yan Zhang 0055, Xing Liu 0010, Takayuki Okatani |
CVPR | 5 |
| 2018 | Exploiting the Potential of Standard Convolutional Autoencoders for Image Restoration by Evolutionary SearchabstractResearchers have applied deep neural networks to image restoration tasks, in which they proposed various network architectures, loss functions, and training methods. In particular, adversarial training, which is employed in recent studies, seems to be a key ingredient to success. In this paper, we show that simple convolutional autoencoders (CAEs) built upon only standard network components, i.e., convolutional layers and skip connections, can outperform the state-of-the-art methods which employ adversarial training and sophisticated loss functions. The secret is to search for good architectures using an evolutionary algorithm. All we did was to train the optimized CAEs by minimizing the l2 loss between reconstructed images and their ground truths using the ADAM optimizer. Our experimental results show that this approach achieves 27.8 dB peak signal to noise ratio (PSNR) on the CelebA dataset and 33.3 dB on the SVHN dataset, compared to 22.8 dB and 19.0 dB provided by the former state-of-the-art methods, respectively. Masanori Suganuma, Mete Ozay, Takayuki Okatani |
ICML | 3 |
| 2018 | Deep Structured Energy-Based Image InpaintingabstractIn this paper, we propose a structured image inpainting method employing an energy based model. In order to learn structural relationship between patterns observed in images and missing regions of the images, we employ an energy-based structured prediction method. The structural relationship is learned by minimizing an energy function which is defined by a simple convolutional neural network. The experimental results on various benchmark datasets show that our proposed method significantly outperforms the state-of-the-art methods which use Generative Adversarial Networks (GANs). We obtained 497.35 mean squared error (MSE) on the Olivetti face dataset compared to 833.0 MSE provided by the state-of-the-art method. Moreover, we obtained 28.4 dB peak signal to noise ratio (PSNR) on the SVHN dataset and 23.53 dB on the CelebA dataset, compared to 22.3 dB and 21.3 dB, provided by the state-of-the-art methods, respectively. The code is publicly available1. Fazil Altinel, Mete Ozay, Takayuki Okatani |
ICPR | 3 |
| 2018 | Recommending Outfits from Personal ClosetabstractWe consider grading a fashion outfit for recommendation, where we assume that users have a closet of items and we aim at producing a score for an arbitrary combination of items in the closet. The challenge in outfit grading is that the input to the system is a bag of item pictures that are unordered and vary in size. We build a deep neural network-based system that can take variable-length items and predict a score. We collect a large number of outfits from a popular fashion sharing website, Polyvore, and evaluate the performance of our grading system. We compare our model with a random-choice baseline, both on the traditional classification evaluation and on people's judgment using a crowdsourcing platform. With over 84% in classification accuracy and 91% matching ratio to human annotators, our model can reliably grade the quality of an outfit. We also build an outfit recommender on top of our grader to demonstrate the practical application of our model for a personal closet assistant. Pongsate Tangseng, Kota Yamaguchi, Takayuki Okatani |
WACV | 3 |
| 2017 | Self-Calibration-Based Approach to Critical Motion Sequences of Rolling-Shutter Structure from MotionabstractIn this paper we consider critical motion sequences (CMSs) of rolling-shutter (RS) SfM. Employing an RS camera model with linearized pure rotation, we show that the RS distortion can be approximately expressed by two internal parameters of an imaginary camera plus one-parameter nonlinear transformation similar to lens distortion. We then reformulate the problem as self-calibration of the imaginary camera, in which its skew and aspect ratio are unknown and varying in the image sequence. In the formulation, we derive a general representation of CMSs. We also show that our method can explain the CMS that was recently reported in the literature, and then present a new remedy to deal with the degeneracy. Our theoretical results agree well with experimental results, it explains degeneracies observed when we employ naive bundle adjustment, and how they are resolved by our method. Eisuke Ito, Takayuki Okatani |
CVPR | 2 |
| 2017 | Truncating Wide Networks Using Binary Tree ArchitecturesabstractIn this paper, we propose a binary tree architecture to truncate architecture of wide networks by reducing the width of the networks. More precisely, in the proposed architecture, the width is incrementally reduced from lower layers to higher layers in order to increase the expressive capacity of networks with a less increase on parameter size. Also, in order to ease the gradient vanishing problem, features obtained at different layers are concatenated to form the output of our architecture. By employing the proposed architecture on a baseline wide network, we can construct and train a new network with same depth but considerably less number of parameters. In our experimental analyses, we observe that the proposed architecture enables us to obtain better parameter size and accuracy trade-off compared to baseline networks using various benchmark image classification datasets. The results show that our model can decrease the classification error of a baseline from 20:43% to 19:22% on Cifar-100 using only 28% of parameters that the baseline has. Code is available at https://github.com/ZhangVision/bitnet. Yan Zhang 0055, Mete Ozay, Shuohao Li, Takayuki Okatani |
ICCV | 4 |
| 2017 | Temporal city modeling using street level imagery
Ken Sakurada, Daiki Tetsuka, Takayuki Okatani |
Comput. Vis. Image Underst. | 3 |
| 2016 | Learning to Describe E-Commerce Images from Noisy Online Data
Takuya Yashima, Naoaki Okazaki, Kentaro Inui, Kota Yamaguchi, Takayuki Okatani |
ACCV (5) | 5 |
| 2016 | Design of Kernels in Convolutional Neural Networks for Image Classification
Zhun Sun, Mete Ozay, Takayuki Okatani |
ECCV (7) | 3 |
| 2016 | Automatic Attribute Discovery with Neural Activations
Sirion Vittayakorn, Takayuki Umeda, Kazuhiko Murasaki, Kyoko Sudo, Takayuki Okatani, Kota Yamaguchi |
ECCV (4) | 5 |
| 2016 | Integrating deep features for material recognitionabstractThis paper considers the problem of material recognition. Motivated by observation of close interconnections between material and object recognition, we study how to select and integrate multiple features obtained by different models of Convolutional Neural Networks (CNNs) trained in a transfer learning setting. To be specific, we first compute activations of features using representations on images to select a set of samples which are best represented by the features. Then, we measure uncertainty of the features by computing entropy of class distributions for each sample set. Finally, we compute contribution of each feature to representation of classes for feature selection and integration. Experimental results show that the proposed method achieves state-of-the-art performance on two benchmark datasets for material recognition. Additionally, we introduce a new material dataset, named EFMD, which extends Flickr Material Database (FMD). By the employment of the EFMD for transfer learning, we achieve 84.0% ± 1.8% accuracy on the FMD dataset, which is close to the reported human performance 84.9%. Yan Zhang 0055, Mete Ozay, Xing Liu 0010, Takayuki Okatani |
ICPR | 4 |
| 2016 | Recognizing Open-Vocabulary Relations between Objects in Images
Masayasu Muraoka, Sumit Maharjan, Masaki Saito, Kota Yamaguchi, Naoaki Okazaki, Takayuki Okatani, Kentaro Inui |
PACLIC | 6 |
| 2016 | Separation of reflection components by sparse non-negative matrix factorization
Yasushi Akashi, Takayuki Okatani |
Comput. Vis. Image Underst. | 2 |
| 2016 | Hybrid macro-micro visual analysis for city-scale state estimationabstractWe address the task of estimating large-scale land surface conditions using overhead aerial (macro-level) images and street view (micro-level) images. These two types of images are captured from orthogonal viewpoints and have different resolutions, thus conveying very different types of information that can be used in a complementary way. Moreover, their integration is necessary to enable an accurate understanding of changes in natural phenomena over massive city-scale landscapes. The key technical challenge is devising a method to integrate these two disparate types of image data in an effective manner, to leverage the wide coverage capabilities of macro-level images and detailed resolution of micro-level images. The strategy proposed in this work uses macro-level imaging to learn the extent to which the land condition corresponds between land regions that share similar visual characteristics (e.g., mountains, streets, buildings, rivers), whereas micro-level images are used to acquire high resolution statistics of land conditions (e.g., the amount of debris on the ground). By combining macro- and micro-level information about regional correspondences and surface conditions, our proposed method is capable of generating detailed estimates of land surface conditions over an entire city. Ken Sakurada, Takayuki Okatani, Kris Makoto Kitani |
Comput. Vis. Image Underst. | 2 |
| 2015 | Change Detection from a Street Image Pair using CNN Features and Superpixel Segmentation
Ken Sakurada, Takayuki Okatani |
BMVC | 2 |
| 2015 | Mix and Match: Joint Model for Clothing and Attribute RecognitionabstractThis paper studies clothing and attribute recognition in the fashion domain. Specifically, in this paper, we turn our attention to the compatibility of clothing items and attributes (Fig 1). For example, people do not wear a skirt and a dress at the same time, yet a jacket and a shirt are a preferred combination. We consider such inter-object or inter-attribute compatibility and formulate a Conditional Random Field (CRF) that seeks the most probable combination in the given picture. The model takes into account the location-specific appearance with respect to a human body and the semantic correlation between clothing items and attributes, which we learn using the max-margin framework. Fig 2 illustrates our pipeline. We evaluate our model using two datasets that resemble realistic applica- tion scenarios: on-line social networks and shopping sites. The empirical evaluation indicates that our model effectively improves the recognition performance over various baselines including the state-of-the-art feature designed exclusively for clothing recognition. The results also suggest that our model generalizes well to different fashion-related applications. Kota Yamaguchi, Takayuki Okatani, Kyoko Sudo, Kazuhiko Murasaki, Yukinobu Taniguchi |
BMVC | 2 |
| 2015 | Transformation of Markov Random Fields for marginal distribution estimationabstractThis paper presents a generic method for transforming MRFs for the marginal inference problem. Its major application is to downsize MRFs to speed up the computation. Unlike the MAP inference, there are only classical algorithms for the marginal inference problem such as BP etc. that require large computational cost. Although downsizing MRFs should directly reduce the computational cost, there is no systematic way of doing this, since it is unclear how to obtain the MRF energy for the downsized MRFs and also how to translate the estimates of their marginal distributions to those of the original MRFs. The proposed method resolves these issues by a novel probabilistic formulation of MRF transformation. The key idea is to represent the joint distribution of an MRF with that of the transformed one, in which the variables of the latter are treated as latent variables. We also show that the proposed method can be applied to discretization of variable space of continuous MRFs and can be used with Markov chain Monte Carlo methods. The experimental results demonstrate the effectiveness of the proposed method. Masaki Saito, Takayuki Okatani |
CVPR | 2 |
| 2015 | Detecting Building-Level Changes of a City Using Street Images and a 2D City MapabstractThis paper presents a method for detecting city-scale changes of a city from its street images and a 2D map. Using SfM to reconstruct point cloud of the structures of the city, the method estimates the existence of each building by matching the point cloud with the 3D building structures recovered from the map. There are multiple difficulties, such as inaccuracy of the recovered building structures, large differences in observation and thus in point cloud size of individual buildings, and mutual dependency of building existences due to potential occlusions. To solve these, we develop a model of how point cloud is generated in the sequential processes of SfM, an observation model of a building wall, and a greedy iterative approach to cope with the mutual dependency. We experimentally apply the method to the cities damaged by the tsunami that struck Japan in 2011. The results show the effectiveness of the method. Daiki Tetsuka, Takayuki Okatani |
WACV | 2 |
| 2014 | Separation of Reflection Components by Sparse Non-negative Matrix Factorization
Yasuhiro Akashi, Takayuki Okatani |
ACCV (5) | 2 |
| 2014 | Understanding Convolutional Neural Networks in Terms of Category-Level Attributes
Makoto Ozeki, Takayuki Okatani |
ACCV (2) | 2 |
| 2014 | Massive City-Scale Surface Condition Analysis Using Ground and Aerial Imagery
Ken Sakurada, Takayuki Okatani, Kris Makoto Kitani |
ACCV (1) | 2 |
| 2013 | Discrete MRF Inference of Marginal Densities for Non-uniformly Discretized Variable SpaceabstractThis paper is concerned with the inference of marginal densities based on MRF models. The optimization algorithms for continuous variables are only applicable to a limited number of problems, whereas those for discrete variables are versatile. Thus, it is quite common to convert the continuous variables into discrete ones for the problems that ideally should be solved in the continuous domain, such as stereo matching and optical flow estimation. In this paper, we show a novel formulation for this continuous-discrete conversion. The key idea is to estimate the marginal densities in the continuous domain by approximating them with mixtures of rectangular densities. Based on this formulation, we derive a mean field (MF) algorithm and a belief propagation (BP) algorithm. These algorithms can correctly handle the case where the variable space is discretized in a non-uniform manner. By intentionally using such a non-uniform discretization, a higher balance between computational efficiency and accuracy of marginal density estimates could be achieved. We present a method for actually doing this, which dynamically discretizes the variable space in a coarse-to-fine manner in the course of the computation. Experimental results show the effectiveness of our approach. Masaki Saito, Takayuki Okatani, Koichiro Deguchi |
CVPR | 2 |
| 2013 | Detecting Changes in 3D Structure of a Scene from Multi-view Images Captured by a Vehicle-Mounted CameraabstractThis paper proposes a method for detecting temporal changes of the three-dimensional structure of an outdoor scene from its multi-view images captured at two separate times. For the images, we consider those captured by a camera mounted on a vehicle running in a city street. The method estimates scene structures probabilistically, not deterministically, and based on their estimates, it evaluates the probability of structural changes in the scene, where the inputs are the similarity of the local image patches among the multi-view images. The aim of the probabilistic treatment is to maximize the accuracy of change detection, behind which there is our conjecture that although it is difficult to estimate the scene structures deterministically, it should be easier to detect their changes. The proposed method is compared with the methods that use multi-view stereo (MVS) to reconstruct the scene structures of the two time points and then differentiate them to detect changes. The experimental results show that the proposed method outperforms such MVS-based methods. Ken Sakurada, Takayuki Okatani, Koichiro Deguchi |
CVPR | 2 |
| 2012 | Optimal integration of photometric and geometric surface measurements using inaccurate reflectance/illumination knowledgeabstractIn this paper, we present a method for accurately estimating the shape of an object by integrating the surface orientation measured by photometric stereo and the position measured by some range-measuring method. We first show that even if the knowledge of the reflectance/illumination is inaccurate, the first derivatives of the photometrically measured orientation can be accurately estimated at the surface points where they have small values. We propose a probabilistic framework to quantitate the (in)accuracy of the knowledge and connect it to the estimation accuracy of these derivatives. Based on this framework, we consider optimally integrating the surface orientation and position to obtain the object shape with higher accuracy. The integration reduces to an optimization problem, and it is efficiently solved by belief propagation. We present several experimental results showing the effectiveness of the proposed approach. Takayuki Okatani, Koichiro Deguchi |
CVPR | 1 |
| 2012 | Application of the mean field methods to MRF optimization in computer visionabstractThe mean field (MF) methods are an energy optimization method for Markov random fields (MRFs). These methods, which have their root in solid state physics, estimate the marginal density of each site of an MRF graph by iterative computation, similarly to loopy belief propagation (LBP). It appears that, being shadowed by LBP, the MF methods have not been seriously considered in the computer vision community. This study investigates whether these methods are useful for practical problems, particularly MPM (Maximum Posterior Marginal) inference, in computer vision. To be specific, we apply the naive MF equations and the TAP (Thouless-Anderson-Palmer) equations to interactive segmentation and stereo matching. In this paper, firstly, we show implementation of these methods for computer vision problems. Next, we discuss advantages of the MF methods to LBP. Finally, we present experimental results that the MF methods are well comparable to LBP in terms of accuracy and global convergence; furthermore, the 3rd-order TAP equation often outperforms LBP in terms of accuracy. Masaki Saito, Takayuki Okatani, Koichiro Deguchi |
CVPR | 2 |
| 2012 | Recognizing surface qualities from natural images based on learning to rank
Takashi Abe, Takayuki Okatani, Koichiro Deguchi |
ICPR | 2 |
| 2011 | Sensing method of total-internal-reflection-based tactile sensorabstractIn recent years, tactile sensors have become an important interface for information devices. Our goal is to develop a simple and sensitive tactile sensor. We have proposed a reflection-image-based tactile sensor. The key points of this sensor were its simplicity and use of an optical device. In this study, we propose a design of a distributed-force sensitive touch panel following the previous developments. We employ an LED (light emitting diode) and PDs (photo diodes) in the new reflection-type tactile sensor. We investigate the illumination distribution and its changes under some simple deformation of the sensor surface with a ray-trace simulation. Then, based on the results of the experimental simulations, we also propose a method for estimating a given deformation from the illumination distribution. Momotaro Koike, Satoshi Saga, Takayuki Okatani, Koichiro Deguchi |
World Haptics | 3 |
| 2011 | Optimum method for real-time reconstruction of sensor surface in total-internal-reflection based tactile sensorabstractIn recent years, many tactile sensors have been developed for the practical use in robotics and to meet the increasing demand for intuitive interfaces. However, the implementation of conventional tactile sensors is very complex. We develop a simple total-internal-reflection based tactile sensor that measures the shape of the sensor surface from a reflection image of the surface and evaluate its performance. We previously proposed a reconstruction method which solves geometric optical constraints of the sensor surface by optimization, however this method could not meet real-time and accurate reconstruction simultaneously. To solve this problem, in this paper, we propose a new reconstruction method based on the previously proposed one. By reconstruction experiments with simulated reflection images, we found that the proposed method could perform real-time and accurate reconstruction and was applicable to various sensor shapes and robust to feature tracking error and the presence of contact areas on the sensor surface without total-internal-reflection property. Then, we implemented the proposed method on the actual sensor and confirmed that the sensor could perform real-time reconstruction. Ryosuke Taira, Satoshi Saga, Takayuki Okatani, Koichiro Deguchi |
World Haptics | 3 |
| 2011 | Efficient algorithm for low-rank matrix factorization with missing components and performance comparison of latest algorithmsabstractThis paper examines numerical algorithms for factorization of a low-rank matrix with missing components. We first propose a new method that incorporates a damping factor into the Wiberg method to solve the problem. The new method is characterized by the way it constrains the ambiguity of the matrix factorization, which helps improve both the global convergence ability and the local convergence speed. We then present experimental comparisons with the latest methods used to solve the problem. No comprehensive comparison of the methods that have been proposed recently has yet been reported in literature. In our experiments, we prioritize the assessment of the global convergence performance of each method, that is, how often and how fast the method can reach the global optimum starting from random initial values. Our conclusion is that top performance is achieved by a group of methods based on Newton-family minimization with damping factor that reduce the problem by eliminating either of the two factored matrices. Our method, which belongs to this group, consistently shows a 100% global convergence rate for different types of affine structure from motion data with a very high population of missing components. Takayuki Okatani, Takahiro Yoshida, Koichiro Deguchi |
ICCV | 1 |
| 2011 | Accurate and robust planar tracking based on a model of image sampling and reconstruction processabstractIt is one of the central issues in augmented reality and computer vision to track a planar object moving relatively to a camera in an accurate and robust manner. In previous studies, it was pointed out that there are several factors making the tracking difficult, such as illumination change and motion blur, and effective solutions were proposed for them. In this paper, we point out that degradation in effective image resolution can also deteriorate tracking performance, which typically occurs when the plane being tracked has an oblique pose with respect to the viewing direction, or when it moves to a distant location from the camera. The deterioration tends to become significantly large for extreme configurations, e.g., when the planar object has nearly a right angle with the viewing direction. Such configurations can frequently occur in AR applications targeted at ordinary users. To cope with this problem, we model the sampling and reconstruction process of images, and present a tracking algorithm that incorporates the model to correctly handle these configurations. We show through several experiments that the proposed method shows better performance than conventional methods. Eisuke Ito, Takayuki Okatani, Koichiro Deguchi |
ISMAR | 2 |
| 2010 | Flexible Online Calibration for a Mobile Projector-Camera System
Daisuke Abe, Takayuki Okatani, Koichiro Deguchi |
ACCV (4) | 2 |
| 2009 | On bias correction for geometric parameter estimation in computer visionabstractMaximum likelihood (ML) estimation is widely used in many computer vision problems involving the estimation of geometric parameters, from conic fitting to bundle adjustment for structure and motion. This paper presents a detailed discussion on the bias of ML estimates derived for these problems. Statistical theory states that although ML estimates attain maximum accuracy in the limit as the sample size goes to infinity, they can have non-negligible bias with small sample sizes. In the case of computer vision problems, the ML optimality holds when regarding variance in observation errors as the sample size. A natural question is how large the bias will be for a given strength of observation errors. To answer this for a general class of problems, we analyze the mechanism of how the bias of ML estimates emerges, and show that the differential geometric properties of geometric constraints used in the problems determines the magnitude of bias. Based on this result, we present a numerical method of computing bias-corrected estimates. Takayuki Okatani, Koichiro Deguchi |
CVPR | 1 |
| 2009 | Improving accuracy of geometric parameter estimation using projected score methodabstractA fundamental problem in computer vision (CV) is the estimation of geometric parameters from multiple observations obtained from images; examples of such problems range from ellipse fitting to multi-view structure from motion (SFM). The maximum likelihood (ML) method is widely used to estimate the parameters in such problems, assuming Gaussian noises to be present in the observations, for example, bundle adjustment for SFM. According to the theory of statistics, the ML estimates are nearly optimal for these problems, provided that the variance of the observation noises is sufficiently small. This implies that when noises are not small, more accurate estimates can be derived as compared to the ML estimates. In this study, we propose the application of a method called the projected score method, developed in statistics for computing higher-accuracy estimates, to the CV problems. We describe how it can be customized to solve the CV problems and propose a numerical algorithm to implement the method. We show that the method works effectively for such problems. Takayuki Okatani, Koichiro Deguchi |
ICCV | 1 |
| 2009 | Shape Reconstruction by Combination of Structured-Light Projection and Photometric Stereo Using a Projector-Camera System
Tomoya Okazaki, Takayuki Okatani, Koichiro Deguchi |
PSIVT | 2 |
| 2009 | Easy Calibration of a Multi-projector Display System
Takayuki Okatani, Koichiro Deguchi |
Int. J. Comput. Vis. | 1 |
| 2009 | Study of Image Quality of Superimposed Projection Using Multiple ProjectorsabstractIn this correspondence, we discuss the quality of images realized by a recently proposed method of generating a high-resolution image by superimposing multiple low-resolution images projected by different projectors. We show several fundamental properties of this method: 1) the accuracy of the image realization (e.g., resolution of the realized image) is heavily affected by the structures of the images to be realized, and 2) there is a tradeoff between the image quality and the maximum brightness of the images to be realized. These are properties peculiar to multiprojector super-resolution and are in contrast with multicamera super-resolution. The results will be helpful in evaluating the usefulness of the method. Takayuki Okatani, Mikio Wada, Koichiro Deguchi |
IEEE Trans. Image Process. | 1 |
| 2007 | Variational Bayes Based Approach to Robust Subspace LearningabstractThis paper presents a new algorithm for the problem of robust subspace learning (RSL), i.e., the estimation of linear subspace parameters from a set of data points in the presence of outliers (and missing data). The algorithm is derived on the basis of the variational Bayes (VB) method, which is a Bayesian generalization of the EM algorithm. For the purpose of the derivation of the algorithm as well as the comparison with existing algorithms, we present two formulations of the EM algorithm for RSL. One yields a variant of the IRLS algorithm, which is the standard algorithm for RSL. The other is an extension of Roweis's formulation of an EM algorithm for PCA, which yields a robust version of the alternated least squares (ALS) algorithm. This ALS-based algorithm can only deal with a certain type of outliers (termed vector-wise outliers). The VB method is used to resolve this limitation, which results in the proposed algorithm. Experimental results using synthetic data show that the proposed algorithm outperforms the IRLS algorithm in terms of the convergence property and the computational time. Takayuki Okatani, Koichiro Deguchi |
CVPR | 1 |
| 2007 | Estimating Scale of a Scene from a Single Image Based on Defocus Blur and Scene GeometryabstractUsing an imaging system in which the image plane can be tilted with respect to the optical axis of the lens, the image of a large-scale scene that appears to be a miniature to human eyes can be captured. This phenomenon suggests that the image contains information regarding the scale of the scene and that human vision can extract this information and recognize the scene scale from a single image. In this study, we consider how human vision can perform this single-view scale estimation. Although it is obvious that the existence of defocus blur in the image that simulates a shallow DOF plays an essential role in the scale estimation, we propose that this alone is not sufficient to explain the estimation mechanism. By incorporating a few assumptions, we theoretically show that scale estimation is made possible when (1) the 3D structure of the scene can be recovered from the image and furthermore, (2) the structure is combined with the defocus blur. Further, we present a simple algorithm for scale recognition and demonstrate its working using a real image. Takayuki Okatani, Koichiro Deguchi |
CVPR | 1 |
| 2007 | Precise 3-D Measurement using Uncalibrated Pattern ProjectionabstractThree-dimensional measurement methods using structured light projection enable accurate depth measurements by decoding the disparity between the projector-camera pair from an observed pattern. However, actual projection devices have various systematic error sources that distort the structured light pattern, directly affecting the accuracy of 3-D measurements. We propose a new method of measuring depth that is not affected by the systematic errors in the projected pattern. In our method, the image of a pattern projected onto a plane is referenced to cancel errors. Based on the invariance of the cross-ratio under perspective projections, depth is obtained from the disparity determined on the referenced image. In experiments, our method removed the systematic errors and improved the accuracy of depth measurement without any extra calibration or measurement. Rui Ishiyama, Takayuki Okatani, Koichiro Deguchi |
ICIP (1) | 2 |
| 2007 | On the Wiberg Algorithm for Matrix Factorization in the Presence of Missing Components
Takayuki Okatani, Koichiro Deguchi |
Int. J. Comput. Vis. | 1 |
| 2006 | The Importance of Gaze Control Mechanism on Vision-based Motion Control of a Biped RobotabstractWe are motivated to deal with a biped robot by the existence of an almighty powerful control mechanism of a human. Although a visual information, in particular, plays an important role in order to realize admirable intelligence in a humanoid robot, we must assert a vision to be one of the developing sensors. In fact, there are many difficulties of both aspects of restrictive hardware resources for a vision and real-time image processing for motor control. In this paper, we clarify a structure and a nature of a visual information processing system which should be equipped with a biped robot. For this purpose, a human eye structure and its movements, and brain motor control through a central nervous system are considered as an analogy with robotics. We shall show that a gaze control mechanism like human is absolutely indispensable for a vision system of a biped robot. Furthermore, we investigate considerable differences between the human eyes and the biped robot ones. This allows us to propose a new, simple and intuitively understandable criteria, called stably gazing scope, for evaluating gaze control performance in the sense of availability of visual information for robot motion control Shun Ushida, Kousuke Yoshimi, Takayuki Okatani, Koichiro Deguchi |
IROS | 3 |
| 2005 | 2DOF motion stabilization of biped robot by gaze control strategyabstractIn order to keep intended distance and direction to a target object, our biped robot modifies its movements based on visual information. To detect the target direction, we stabilize gaze of the robot by 2-DOF controller, using a scheduled robot motion plan for the feedforward one and the image information for the feedback one. Then, we stabilize the robot motion with the stabilized gaze. In this control scheme, parameters of the feedforward controller is modified on-line by using an adaptive law of model reference adaptive control (MRAC). Shota Takizawa, Shun Ushida, Takayuki Okatani, Koichiro Deguchi |
IROS | 3 |
| 2005 | Autocalibration of a Projector-Camera SystemabstractThis paper presents a method for calibrating a projector-camera system that consists of multiple projectors (or multiple poses of a single projector), a camera, and a planar screen. We consider the problem of estimating the homography between the screen and the image plane of the camera or the screen-camera homography, in the case where there is no prior knowledge regarding the screen surface that enables the direct computation of the homography. It is assumed that the pose of each projector is unknown while its internal geometry is known. Subsequently, it is shown that the screen-camera homography can be determined from only the images projected by the projectors and then obtained by the camera, up to a transformation with four degrees of freedom. This transformation corresponds to arbitrariness in choosing a two-dimensional coordinate system on the screen surface and when this coordinate system is chosen in some manner, the screen-camera homography as well as the unknown poses of the projectors can be uniquely determined. A noniterative algorithm is presented, which computes the homography from three or more images. Several experimental results on synthetic as well as real images are shown to demonstrate the effectiveness of the method. Takayuki Okatani, Koichiro Deguchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | A color-based probabilistic tracking by using graphical modelsabstractIn this paper, a method for real-time tracking of moving objects or pedestrians is proposed. Especially, we tackled to track an object whose part was temporally occluded or which traversed in front of the cluttered background. Furthermore, we tried to obtain the accurate trajectory of the center of it stably even then. For the problems, we present a probabilistic color-based tracking. In our proposed method, we incorporated particle filtering into graphical models and applied it to the color-based tracking. When the color histogram of the tracked object was made, we used not one region of the whole of the object but the multi-part region of it if it was divided in some parts (e.g. the head and the body of a pedestrian). We treated the multi-part region as a graphical model. In the graphical model, messages about the state of the parts are sent to other parts. As the result., even when the tracked object has occluded part, our proposed method can track it stably and infer the state (e.g. position) of the occluded part. We made experiments to confirm effectiveness of this proposed method. Yoshinori Satoh, Takayuki Okatani, Koichiro Deguchi |
IROS | 2 |
| 2003 | Toward a Statistically Optimal Method for Estimating Geometric Relations from Noisy Data: Cases of Linear RelationsabstractIn many problems of computer vision we have to estimate parameters in the presence of nuisance parameters increasing with the amount of data. It is known that, unlike in the cases without nuisance parameters, maximum likelihood estimation (MLE) is not optimal in the presence of nuisance parameters. By optimal we mean that the resulting estimate is unbiased and its variance attains the theoretical lower bound in an asymptotic sense. Thus, naive application of MLE to computer vision has a potential problem. This applies to a wide range of problems from conic fitting to bundle adjustment. For this nuisance parameter problem, studies have been conducted in statistics for a long time, whereas they have hardly been studied in the computer vision community. We cast light on the methods developed in statistics for obtaining optimal estimates and explore the possibility of applying them to computer vision problems. In this paper we focus on the cases where data and nuisance parameters are linearly connected. As examples, optical flow estimation and affine structure and motion problems are considered. Through experiments, we show that the estimation accuracy is improved in several cases. Takayuki Okatani, Koichiro Deguchi |
CVPR (1) | 1 |
| 2003 | Autocalibration of a Projector-Screen-Camera System: Theory and Algorithm for Screen-to-Camera Homography EstimationabstractThis paper deals with the autocalibration of a system that consists of a planar screen, multiple projectors, and a camera. In the system, either multiple projectors or a single moving projector projects patterns on a screen while a stationary camera placed in front of the screen takes images of the patterns. We treat the case in which the patterns that the projectors project toward space are assumed to be known (i.e., the projectors are calibrated), whereas poses of the projectors are unknown. Under these conditions, we consider the problem of estimating screen-to-camera homography from the images alone. This is intended for cases where there is no clue on the screen surface that enables direct estimation of the screen-to-camera homography. One application is a 6DOF input device; poses of a multibeam projector freely moving in space are computed from the images of beam spots on the screen. The primary contribution of the paper is theoretical results on the uniqueness of solutions and a noniterative algorithm for the problem. The effectiveness of the method is shown by experimental results on synthetic as well as on real images. Takayuki Okatani, Koichiro Deguchi |
ICCV | 1 |
| 2003 | Tracking multiple three-dimensional motions by using modified condensation algorithm and multiple imagesabstractIn this paper we present a system for tracking objects' 3D motions by using sampling algorithm. The system consists of three cameras and 1 PC. We modified the condensation algorithm [M.A. Isard et al., 1998] to track objects' 3D positions. The condensation algorithm is a kind of "factored sampling" and it can easily cope with multi-modal probability density of objects' positions. We track objects not in images but in 3D space. We use a condensation particle-set as a tracker and represent a position of a moving object by a center of tracker. Since we don't track directly in images, a great advantage of this method is not to need to deal with occlusions or objects' mutual interferences which happen in images. Koji Hamasaki, Taira Nakajima, Takayuki Okatani, Koichiro Deguchi |
IROS | 3 |
| 2002 | A motion tracking by extracting 3D feature of moving objects with binocular cooperative fixationabstractWe propose a binocular motion tracking system by reconstructing the 3D shape of moving objects. This system has a pair of active cameras mounted on the robot arm. The 3D shape is an important feature to track a moving object. To track the moving object, the cameras fixate a point on the object and reconstruct the 3D shape of the object around the fixation point of the two images with the principle of spatio-temporal differentiation analysis. The 3D shape is fed back for the control of the active cameras and the robot arm for motion tracking. The system realizes an accurate and stable tracking of the moving object. Yoshinori Satoh, Tomohiro Nakagawa, Takayuki Okatani, Koichiro Deguchi |
IROS | 3 |
| 2001 | On Photometric Aspects of Catadioptric CamerasabstractSeveral imaging systems called catadioptric cameras have been developed that use a combination of lenses and mirrors to obtain wide fields of view. Geometric aspects have been well studied. This paper focuses on photometric aspects of catadioptric cameras. We discuss how geometric distortion of images of catadioptric cameras resulting from curvature of the mirror affects image irradiance. We show that image irradiance is determined only by the nature of lenses comprising the cameras and is independent of mirror shape. This is because an increase in the solid angle of space per single image pixel is canceled out by a decrease in the apparent solid angle of the lens aperture viewed from a scene point via the mirror. Thus, image irradiance of catadioptric cameras is not affected by geometric distortion due to mirror curvature. This is in contrast to cameras with ordinary lenses alone, especially fisheye lenses, in which image irradiance is actually affected by image distortion. Theoretical results as well as results of numerical simulations are presented. Takayuki Okatani, Koichiro Deguchi |
CVPR (1) | 1 |
| 2001 | On Uniqueness of Solutions of the Three-Light-Source Photometric Stereo: Conditions on Illumination Configuration and Surface Reflectance
Takayuki Okatani, Koichiro Deguchi |
Comput. Vis. Image Underst. | 1 |
| 2000 | Estimation of Illumination Distribution Using a Specular SphereabstractA method has been proposed for acquiring an omnidirectional image of a scene using a specular sphere (namely a mirror ball). The method takes the reflected image on the specular sphere located at a desired position in a scene by an ordinary camera. This paper discusses the relation between the reflected image of the scene and the radiance of the scene. Due to the reflection on the sphere, the final image on the camera has geometric distortion. There are imaging systems covering a wide angle by the lens system itself, e.g., a fish-eye lens. The image obtained by such lens systems has similar geometrical distortion. However, the effect of the distortion on the image brightness is different. We point out that the image distortion due to the lens system affects the image brightness, while that due to the reflection on the specular sphere does not. Most of the discussion can be applied to the methods using a curved mirror to obtain a wide angle image. Takayuki Okatani, Koichiro Deguchi |
ICPR | 1 |
| 2000 | Closed Form Solution of Local Shape from Shading at Critical Points
Takayuki Okatani, Koichiro Deguchi |
Int. J. Comput. Vis. | 1 |
| 1999 | Computation of the Sign of the Gaussian Curvature of a Surface from Multiple Unknown Illumination Images without Knowledge of the Reflectance Property
Takayuki Okatani, Koichiro Deguchi |
Comput. Vis. Image Underst. | 1 |
| 1998 | On the Classification of Singular Points for the Global Shape from Shading Problem: A Study of the Constraints Imposed by Isophotes
Takayuki Okatani, Koichiro Deguchi |
ACCV (1) | 1 |
| 1998 | Determination of Sign of Gaussian Curvature of Surface Having General Reflectance Property
Takayuki Okatani, Koichiro Deguchi |
ACCV (1) | 1 |
| 1998 | On identification of singular points using photometric invariants for global shape from shading problemabstractThe paper is concerned with the problem of identifying types of image singular points as either maxima, minima, or saddle points, especially for the case of unknown lighting direction. Singular points are defined as maximally bright points in the image whose corresponding surface normals coincide with the lighting direction. The identification of their type is the key to the global shape from shading problem, and it can be viewed as rough estimation of the object shape in itself. Using photometric invariants (e.g., extrema of the image grey level profile always lie on the parabolic curves of the object surface), it is shown that the two singular points which are connected by two steepest ascent curves starting from an image saddle point usually have different signs with each other. This constrains the possible combination of the types of the singular points in the global shape from shading problem. Takayuki Okatani, Koichiro Deguchi |
ICPR | 1 |
| 1997 | Shape Reconstruction from an Endoscope Image by Shape from Shading Technique for a Point Light Source at the Projection Center
Takayuki Okatani, Koichiro Deguchi |
Comput. Vis. Image Underst. | 1 |
| 1996 | Reconstructing shape from shading with a point light source at the projection center: shape reconstruction from an endoscope imageabstractA new method for reconstructing three dimensional object from an endoscope image is presented. The proposed method uses image shading generated by a light source at the endoscope head. In this case, the light source is near to the object surface so that the image brightness depends on not only the surface gradients but also the distance from the light source to the surface point. To deal with this difficulty, we consider the imaging system of the endoscope as the system having a point light source at the projection center. The object surface is reconstructed by propagating equal-distance contours, spatial curves composed of points at an equal distance from the light source. The propagation is controlled by the image shading. We use the level-sets method for numerical computation. Experimental results for real medical images show feasibility of this method. Takayuki Okatani, Koichiro Deguchi |
ICPR | 1 |