Akihiro Sugimoto

dblp:36/4318 · DBLP profile ↗
← Back
73ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0001-9148-9822ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 53 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 43 · 4 first-author · 6 since 2021Theory of computation · 3Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 SAGA: Learning Signal-Aligned Distributions for Improved Text-to-Image Generation
abstract
State-of-the-art text-to-image models produce visually impressive results but often struggle with precise alignment to text prompts, leading to missing critical elements or unintended blending of distinct concepts. We propose a novel approach that learns a high-success-rate distribution conditioned on a target prompt, ensuring that generated images faithfully reflect the corresponding prompts. Our method explicitly models the signal component during the denoising process, offering fine-grained control that mitigates over-optimization and out-of-distribution artifacts. Moreover, our framework is training-free and seamlessly integrates with both existing diffusion and flow matching architectures. It also supports additional conditioning modalities -- such as bounding boxes -- for enhanced spatial alignment. Extensive experiments demonstrate that our approach outperforms current state-of-the-art methods.
Paul Grimal, Michaël Soumm, Hervé Le Borgne, Olivier Ferret, Akihiro Sugimoto
AAAI5
2024 Evcap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
abstract
Large language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently, necessitating the requirement of sustaining up-to-date object knowledge for open-world comprehension. Instead of relying on large amounts of data and/or scaling up network parameters, we introduce a highly effective retrieval-augmented image captioning method that prompts LLMs with object names retrieved from External Visual-name memory (EVCAP). We build ever-changing object knowledge memory using objects' visuals and names, enabling us to (i) update the memory at a minimal cost and (ii) effort-lessly augment LLMs with retrieved object names by uti-lizing a lightweight and fast-to-train model. Our model, which was trained only on the COCO dataset, can adapt to out-of-domain without requiring additional fine-tuning or retraining. Our experiments conducted on benchmarks and synthetic commonsense-violating data show that EV-CAP, with only 3.97M trainable parameters, exhibits superior performance compared to other methods based on frozen pretrained LLMs. Its performance is also competitive to specialist SOTAs that require extensive training.
Jiaxuan Li 0004, Duc Minh Vo, Akihiro Sugimoto, Hideki Nakayama
CVPR3
2023 A-CAP: Anticipation Captioning with Commonsense Knowledge
abstract
Humans possess the capacity to reason about the future based on a sparse collection of visual cues acquired over time. In order to emulate this ability, we introduce a novel task called Anticipation Captioning, which generates a caption for an unseen oracle image using a sparsely temporally-ordered set of images. To tackle this new task, we propose a model called A-CAP, which incorporates commonsense knowledge into a pre-trained vision-language model, allowing it to anticipate the caption. Through both qualitative and quantitative evaluations on a customized visual storytelling dataset, A-CAP out-performs other image captioning methods and establishes a strong baseline for anticipation captioning. We also address the challenges inherent in this task.
Duc Minh Vo, Quoc-An Luong, Akihiro Sugimoto, Hideki Nakayama
CVPR3
2023 TextANIMAR: Text-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, E-Ro Nguyen, Minh-Khoi Nguyen-Nhat, Tuan-An To, Trung-Truc Huynh-Le, Nham-Tan Nguyen, Hoang-Chau Luong, Truong Hoai Phong, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran
Comput. Graph.32
2023 SketchANIMAR: Sketch-based 3D animal fine-grained retrieval
Trung-Nghia Le, Tam V. Nguyen 0002, Minh-Quan Le, Viet-Tham Huynh, Trong-Le Do, Khanh-Duy Le, Mai-Khiem Tran, Nhat Hoang-Xuan, Thang-Long Nguyen-Ho, Vinh-Tiep Nguyen, Nhat-Quynh Le-Pham, Huu-Phuc Pham, Trong-Vu Hoang, Quang-Binh Nguyen, Trong-Hieu Nguyen Mau, Tuan-Luc Huynh, Thanh-Danh Le, Ngoc-Linh Nguyen-Ha, Tuong-Vy Truong-Thuy, Truong Hoai Phong, Tuong-Nghiem Diep, Khanh-Duy Ho, Xuan-Hieu Nguyen, Thien-Phuc Tran, Tuan-Anh Yang, Kim-Phat Tran, Nhu-Vinh Hoang, Minh-Quang Nguyen, Hoai-Danh Vo, Minh-Hoa Doan, Hai-Dang Nguyen, Akihiro Sugimoto, Minh-Triet Tran
Comput. Graph.33
2022 NOC-REK: Novel Object Captioning with Retrieved Vocabulary from External Knowledge
abstract
Novel object captioning aims at describing objects absent from training data, with the key ingredient being the provision of object vocabulary to the model. Although existing methods heavily rely on an object detection model, we view the detection step as vocabulary retrieval from an external knowledge in the form of embeddings for any object's definition from Wiktionary, where we use in the retrieval image region features learned from a transformers model. We propose an end-to-end Novel Object Captioning with Retrieved vocabulary from External Knowledge method (NOC-REK), which simultaneously learns vocabulary retrieval and caption generation, successfully describing novel objects outside of the training dataset. Furthermore, our model eliminates the requirement for model retraining by simply updating the external knowledge whenever a novel object appears. Our comprehensive experiments on held-out COCO and Nocaps datasets show that our NOCREK is considerably effective against SOTAs.
Duc Minh Vo, Hong Chen 0017, Akihiro Sugimoto, Hideki Nakayama
CVPR3
2022 Learning Monocular 3D Human Pose Estimation With Skeletal Interpolation
abstract
Deep learning has achieved unprecedented accuracy for monocular 3D human pose estimation. However, current learning-based 3D human pose estimation still suffers from poor generalization. Inspired by skeletal animation, which is popular in game development and animation production, we put forward an simple, intuitive yet effective interpolation-based data augmentation approach to synthesize continuous and diverse 3D human body sequences to enhance model generalization. The Transformer-based lifting network, trained with the augmented data, utilizes the self-attention mechanism to perform 2D-to-3D lifting and successfully infer high-quality predictions in the qualitative experiment. The quantitative result of cross-dataset experiment demonstrates that our resulting model achieves superior generalization accuracy on the publicly available dataset.
Akihiro Sugimoto, Shang-Hong Lai
ICASSP2
2022 PPCD-GAN: Progressive Pruning and Class-Aware Distillation for Large-Scale Conditional GANs Compression
abstract
We push forward neural network compression research by exploiting a novel challenging task of large-scale conditional generative adversarial networks (GANs) compression. To this end, we propose a gradually shrinking GAN (PPCD-GAN) by introducing progressive pruning residual block (PP-Res) and class-aware distillation. The PP-Res is an extension of the conventional residual block where each convolutional layer is followed by a learnable mask layer to progressively prune network parameters as training proceeds. The class-aware distillation, on the other hand, enhances the stability of training by transferring immense knowledge from a well-trained teacher model through instructive attention maps. We train the pruning and distillation processes simultaneously on a well-known GAN architecture in an end-to-end manner. After training, all redundant parameters as well as the mask layers are discarded, yielding a lighter network while retaining the performance. We comprehensively illustrate, on ImageNet 128 × 128 dataset, PPCD-GAN reduces up to 5.2 ×(81%) parameters against state-of-the-arts while keeping better performance.
Duc Minh Vo, Akihiro Sugimoto, Hideki Nakayama
WACV2
2022 Paired-D++ GAN for image manipulation with text
Duc Minh Vo, Akihiro Sugimoto
Mach. Vis. Appl.2
2022 Temporal feature enhancement network with external memory for live-stream video object detection
Masato Fujitake, Akihiro Sugimoto
Pattern Recognit.2
2021 Agent-Environment Network for Temporal Action Proposal Generation
abstract
Temporal action proposal generation is an essential and challenging task that aims at localizing temporal intervals containing human actions in untrimmed videos. Most of existing approaches are unable to follow the human cognitive process of understanding the video context due to lack of attention mechanism to express the concept of an action or an agent who performs the action or the interaction between the agent and the environment. Based on the action definition that a human, known as an agent, interacts with the environment and performs an action that affects the environment, we propose a contextual Agent-Environment Network. Our proposed contextual AEN involves (i) agent pathway, operating at a local level to tell about which humans/agents are acting and (ii) environment pathway operating at a global level to tell about how the agents interact with the environment. Comprehensive evaluations on 20-action THUMOS-14 and 200-action ActivityNet-1.3 datasets with different backbone networks, i.e C3D and SlowFast, show that our method robustly exhibits outperformance against state-of-the-art methods regardless of the employed backbone network.
Viet-Khoa Vo-Ho, T. Hoang Ngan Le, Kashu Yamazaki, Akihiro Sugimoto, Minh-Triet Tran
ICASSP4
2021 Real-Time Object Detection by Feature Map Forecast for Live Streaming Video
abstract
This paper proposes a method that jointly learns to detect objects at the current frame and forecast the next frame’s future feature map. Previous offline detectors have shown the effectiveness of utilizing future information in video object detection; however, we cannot take such an approach when dealing with live streaming videos. In contrast, we utilize the forecast feature map with the current and past frame feature maps for object detection, where forecast feature maps are learned using observation of the present and past frames. To maintain a reliable forecast, we introduce a scheduler network, which decides whether we use the forecast feature map as input or extract the feature map from the next frame. Evaluations of our proposed model on the ImageNet VID dataset demonstrate the superior performance of our model against the public benchmark at similar architectures, with achieving 65.7% mAP at 38.9 fps.
Masato Fujitake, Akihiro Sugimoto
ICME2
2020 Optimal Parenthesizing of Geometric Algebra Products
Stéphane Breuils, Vincent Nozick, Akihiro Sugimoto
CGI3
2020 TetraTSDF: 3D Human Reconstruction From a Single Image With a Tetrahedral Outer Shell
abstract
Recovering the 3D shape of a person from its 2D appearance is ill-posed due to ambiguities. Nevertheless, with the help of convolutional neural networks (CNN) and prior knowledge on the 3D human body, it is possible to overcome such ambiguities to recover detailed 3D shapes of human bodies from single images. Current solutions, however, fail to reconstruct all the details of a person wearing loose clothes. This is because of either (a) huge memory requirement that cannot be maintained even on modern GPUs or (b) the compact 3D representation that cannot encode all the details. In this paper, we propose the tetrahedral outer shell volumetric truncated signed distance function (TetraTSDF) model for the human body, and its corresponding part connection network (PCN) for 3D human body shape regression. Our proposed model is compact, dense, accurate, and yet well suited for CNN-based regression task. Our proposed PCN allows us to learn the distribution of the TSDF in the tetrahedral volume from a single image in an end-to-end manner. Results show that our proposed method allows to reconstruct detailed shapes of humans wearing loose clothes from single RGB images.
Hayato Onizuka, Zehra Hayirci, Diego Thomas, Akihiro Sugimoto, Hideaki Uchiyama, Rin-Ichiro Taniguchi
CVPR4
2020 Minimal Rolling Shutter Absolute Pose with Unknown Focal Length and Radial Distortion
Zuzana Kukelova, Cenek Albl, Akihiro Sugimoto, Konrad Schindler, Tomás Pajdla
ECCV (5)3
2020 Visual-Relation Conscious Image Generation from Structured-Text
Duc Minh Vo, Akihiro Sugimoto
ECCV (28)2
2020 Stylized-Colorization for Line Arts
abstract
We address a novel problem of stylized-colorization which colorizes a given line art using a given coloring style in text. This problem can be stated as multi-domain image translation and is more challenging than the current colorization problem because it requires not only capturing the illustration distribution but also satisfying the required coloring styles specific to anime such as lightness, shading, or saturation. We propose a GAN-based end-to-end model for stylized-colorization where the model has one generator and two discriminators. Our generator is based on the U-Net architecture and receives a pair of a line art and a coloring style in text as its input to produce a stylized-colorization image of the line art. Two discriminators, on the other hand, share weights at early layers to judge the stylized-colorization image in two different aspects: one for color and one for style. One generator and two discriminators are jointly trained in an adversarial and end-to-end manner. Extensive experiments demonstrate the effectiveness of our proposed model.
Tzu-Ting Fang, Duc Minh Vo, Akihiro Sugimoto, Shang-Hong Lai
ICPR3
2020 Temporal Feature Enhancement Network with External Memory for Object Detection in Surveillance Video
abstract
Video object detection is challenging and essential in practical applications, such as surveillance cameras for traffic control and public security. Unlike the video in natural scenes, the surveillance video tends to contain dense and small objects (typically vehicles) in their appearances. Therefore, existing methods for surveillance object detection utilize still-image object detection approaches with rich feature extractors at the expense of their run-time speeds. The run-time speed, however, becomes essential when the video is being streamed. In this paper, we exploit temporal information in videos to enrich the feature maps, proposing the first temporal attention based external memory network for the live stream of video. Extensive experiments on real-world traffic surveillance benchmarks demonstrate the real-time performance of the proposed model while keeping comparable accuracy with state-of-the-art.
Masato Fujitake, Akihiro Sugimoto
ICPR2
2020 Attention R-CNN for Accident Detection
abstract
This paper addresses accident detection where we not only detect objects with classes, but also recognize their characteristic properties. More specifically, we aim at simultaneously detecting object class bounding boxes on roads and recognizing their status such as safe, dangerous, or crashed. To achieve this goal, we construct a new dataset and propose a baseline method for benchmarking the task of accident detection. We design an accident detection network, called Attention R-CNN, which consists of two streams: one is for object detection with classes and one for characteristic property computation. As an attention mechanism capturing contextual information in the scene, we integrate global contexts exploited from the scene into the stream for object detection. This introduced attention mechanism enables us to recognize object characteristic properties. Extensive experiments on the newly constructed dataset demonstrate the effectiveness of our proposed network. The dataset and source code are publicly available on our project page.
Trung-Nghia Le, Shintaro Ono, Akihiro Sugimoto, Hiroshi Kawasaki
IV3
2020 Toward Interactive Self-Annotation For Video Object Bounding Box: Recurrent Self-Learning And Hierarchical Annotation Based Framework
abstract
Amount and variety of training data drastically affect the performance of CNNs. Thus, annotation methods are becoming more and more critical to collect data efficiently. In this paper, we propose a simple yet efficient Interactive Self-Annotation framework to cut down both time and human labor cost for video object bounding box annotation. Our method is based on recurrent self-supervised learning and consists of two processes: automatic process and interactive process, where the automatic process aims to build a supported detector to speed up the interactive process. In the Automatic Recurrent Annotation, we let an off-the-shelf detector watch unlabeled videos repeatedly to reinforce itself automatically. At each iteration, we utilize the trained model from the previous iteration to generate better pseudo ground-truth bounding boxes than those at the previous iteration, recurrently improving self-supervised training the detector. In the Interactive Recurrent Annotation, we tackle the human-in-the-loop annotation scenario where the detector receives feedback from the human annotator. To this end, we propose a novel Hierarchical Correction module, where the annotated frame-distance binarizedly decreases at each time step, to utilize the strength of CNN for neighbor frames. Experimental results on various video datasets demonstrate the advantages of the proposed framework in generating high-quality annotations while reducing annotation time and human labor costs.
Trung-Nghia Le, Akihiro Sugimoto, Shintaro Ono, Hiroshi Kawasaki
WACV2
2020 Two-stream FCNs to balance content and style for style transfer
Duc Minh Vo, Akihiro Sugimoto
Mach. Vis. Appl.2
2019 Revisiting Depth Image Fusion with Variational Message Passing
Diego Thomas, Ekaterina Sirazitdinova, Akihiro Sugimoto, Rin-Ichiro Taniguchi
3DV3
2019 Transverse Approach to Geometric Algebra Models for Manipulating Quadratic Surfaces
Stéphane Breuils, Vincent Nozick, Laurent Fuchs, Akihiro Sugimoto
CGI4
2019 Semantic Instance Meets Salient Object: Study on Video Semantic Salient Instance Segmentation
abstract
Focusing on only semantic instances that only salient in a scene gains more benefits for robot navigation and self-driving cars than looking at all objects in the whole scene. This paper pushes the envelope on salient regions in a video to decompose them into semantically meaningful components, namely, semantic salient instances. We provide the baseline for the new task of video semantic salient instance segmentation (VSSIS), that is, Semantic Instance - Salient Object (SISO) framework. The SISO framework is simple yet efficient, leveraging advantages of two different segmentation tasks, i.e. semantic instance segmentation and salient object segmentation to eventually fuse them for the final result. In SISO, we introduce a sequential fusion by looking at overlapping pixels between semantic instances and salient regions to have non-overlapping instances one by one. We also introduce a recurrent instance propagation to refine the shapes and semantic meanings of instances, and an identity tracking to maintain both the identity and the semantic meaning of instances over the entire video. Experimental results demonstrated the effectiveness of our SISO baseline, which can handle occlusions in videos. In addition, to tackle the task of VSSIS, we augment the DAVIS-2017 benchmark dataset by assigning semantic ground-truth for salient instance labels, obtaining SEmantic Salient Instance Video (SESIV) dataset. Our SESIV dataset consists of 84 high-quality video sequences with pixel-wisely per-frame ground-truth labels.
Trung-Nghia Le, Akihiro Sugimoto
WACV2
2019 Anabranch network for camouflaged object segmentation
Trung-Nghia Le, Tam V. Nguyen 0002, Zhongliang Nie, Minh-Triet Tran, Akihiro Sugimoto
Comput. Vis. Image Underst.5
2018 SegmentedFusion: 3D Human Body Reconstruction Using Stitched Bounding Boxes
abstract
This paper presents SegmentedFusion, a method possessing the capability of reconstructing non-rigid 3D models of a human body by using a single depth camera with skeleton information. Our method estimates a dense volumetric 6D motion field that warps the integrated model into the live frame by segmenting a human body into different parts and building a canonical space for each part. The key feature of this work is that a deformed and connected canonical volume for each part is created, and it is used to integrate data. The dense volumetric warp field of one volume is represented efficiently by blending a few rigid transformations. Overall, SegmentedFusion is able to scan a non-rigidly deformed human surface as well as to estimate the dense motion field by using a consumer-grade depth camera. The experimental results demonstrate that SegmentedFusion is robust against fast inter-frame motion and topological changes. Since our method does not require prior assumption, SegmentedFusion can be applied to a wide range of human motions.
Shih-Hsuan Yao, Diego Thomas, Akihiro Sugimoto, Shang-Hong Lai, Rin-Ichiro Taniguchi
3DV3
2018 Linear Solution to the Minimal Absolute Pose Rolling Shutter Problem
Zuzana Kukelova, Cenek Albl, Akihiro Sugimoto, Tomás Pajdla
ACCV (3)3
2018 Paired-D GAN for Semantic Image Synthesis
Duc Minh Vo, Akihiro Sugimoto
ACCV (4)2
2018 Balancing Content and Style with Two-Stream FCNs for Style Transfer
abstract
Style transfer is to render given image contents in given styles, and it has an important role in both computer vision fundamental research and industrial applications. Following the success ofdeep learning based approaches, this problem has been re-launched very recently, but still remains a difficult task because of trade-of between preserving contents and faithful rendering of styles. In this paper, we propose an end-to-end two-stream Fully Convolutional Networks (FCNs) aiming at balancing the contributions of the content and the style in rendered images. Our proposed network consists ofthe encoder and decoder parts. The encoder part utilizes a FCN for content and a FCN for style where the two FCNs are independently trained to preserve the semantic content and to learn the faithful style representation in each. The semantic content feature and the style representationfeature are then concatenated adaptively and fed into the decoder to generate style-transferred (stylized) images. In order to train our proposed network, we employ a loss network, the pre-trained VGG-I6, to compute content loss and style loss, both of which are efficiently used for the feature concatenation. Our intensive experiments show that our proposed model generates more balanced stylized images in content and style than state-of-theart methods. Moreover, our proposed network achieves efficiency in speed.
Duc Minh Vo, Trung-Nghia Le, Akihiro Sugimoto
WACV3
2018 Video Salient Object Detection Using Spatiotemporal Deep Features
abstract
This paper presents a method for detecting salient objects in videos, where temporal information in addition to spatial information is fully taken into account. Following recent reports on the advantage of deep features over conventional handcrafted features, we propose a new set of spatiotemporal deep (STD) features that utilize local and global contexts over frames. We also propose new spatiotemporal conditional random field (STCRF) to compute saliency from STD features. STCRF is our extension of CRF to the temporal domain and describes the relationships among neighboring regions both in a frame and over frames. STCRF leads to temporally consistent saliency maps over frames, contributing to accurate detection of salient objects' boundaries and noise reduction during detection. Our proposed method first segments an input video into multiple scales and then computes a saliency map at each scale level using STD features with STCRF. The final saliency map is computed by fusing saliency maps at different scale levels. Our experiments, using publicly available benchmark datasets, confirm that the proposed method significantly outperforms the state-of-the-art methods. We also applied our saliency computation to the video object segmentation task, showing that our method outperforms existing video object segmentation methods.
Trung-Nghia Le, Akihiro Sugimoto
IEEE Trans. Image Process.2
2017 Deeply Supervised 3D Recurrent FCN for Salient Object Detection in Videos
Trung-Nghia Le, Akihiro Sugimoto
BMVC2
2017 Fast 3D point cloud segmentation using supervoxels with geometry and color for 3D scene understanding
abstract
Segmentation of 3D colored point clouds is a research field with renewed interest thanks to recent availability of inexpensive consumer RGB-D cameras and its importance as an unavoidable low-level step in many robotic applications. However, 3D data's nature makes the task challenging and, thus, many different techniques are being proposed, all of which require expensive computational costs. This paper presents a novel fast method for 3D colored point cloud segmentation. It starts with supervoxel partitioning of the cloud, i.e., an oversegmentation of the points in the cloud. Then it leverages on a novel metric exploiting both geometry and color to iteratively merge the supervoxels to obtain a 3D segmentation where the hierarchical structure of partitions is maintained. The algorithm also presents computational complexity linear to the size of the input. Experimental results over two publicly available datasets demonstrate that our proposed method outperforms state-of-the-art techniques.
Francesco Verdoja, Diego Thomas, Akihiro Sugimoto
ICME3
2017 Synthesis of Environment Maps for Mixed Reality
abstract
When rendering virtual objects in a mixed reality application, it is helpful to have access to an environment map that captures the appearance of the scene from the perspective of the virtual object. It is straightforward to render virtual objects into such maps, but capturing and correctly rendering the real components of the scene into the map is much more challenging. This information is often recovered from physical light probes, such as reflective spheres or fisheye cameras, placed at the location of the virtual object in the scene. For many application areas, however, real light probes would be intrusive or impractical. Ideally, all of the information necessary to produce detailed environment maps could be captured using a single device. We introduce a method using an RGBD camera and a small fisheye camera, contained in a single unit, to create environment maps at any location in an indoor scene. The method combines the output from both cameras to correct for their limited field of view and the displacement from the virtual object, producing complete environment maps suitable for rendering the virtual content in real time. Our method improves on previous probeless approaches by its ability to recover high-frequency environment maps. We demonstrate how this can be used to render virtual objects which shadow, reflect and refract their environment convincingly.
David R. Walton, Diego Thomas, Anthony Steed, Akihiro Sugimoto
ISMAR4
2017 Modeling large-scale indoor scenes with rigid fragments using RGB-D cameras
Diego Thomas, Akihiro Sugimoto
Comput. Vis. Image Underst.2
2017 Discrete rigid registration: A local graph-search approach
Phuc Ngo 0001, Yukiko Kenmochi, Akihiro Sugimoto, Hugues Talbot, Nicolas Passat
Discret. Appl. Math.3
2017 Parametric Surface Representation with Bump Image for Dense 3D Modeling Using an RBG-D Camera
Diego Thomas, Akihiro Sugimoto
Int. J. Comput. Vis.2
2016 Degeneracies in Rolling Shutter SfM
Cenek Albl, Akihiro Sugimoto, Tomás Pajdla
ECCV (5)2
2016 Room reconstruction from a single spherical image by higher-order energy minimization
abstract
We propose a method for understanding a room from a single spherical image, i.e., reconstructing and identifying structural planes forming the ceiling, the floor, and the walls in a room. A spherical image records the light that falls onto a single viewpoint from all directions and does not require correlating geometrical information from multiple images, which facilitates robust and precise reconstruction of the room structure. In our method, we detect line segments from a given image, and classify them into two groups: segments that form the boundaries of the structural planes and those that do not. We formulate this problem as a higher-order energy minimization problem that combines the various measures of likelihood that one, two, or three line segments are part of the boundary. We minimize the energy with graph cuts to identify segments forming boundaries, from which we estimate structural the planes in 3D. Experimental results on synthetic and real images confirm the effectiveness of the proposed method.
Kosuke Fukano, Yoshihiko Mochizuki, Satoshi Iizuka, Edgar Simo-Serra, Akihiro Sugimoto, Hiroshi Ishikawa 0002
ICPR5
2016 Detection by classification of buildings in multispectral satellite imagery
abstract
We present an approach for the detection of buildings in multispectral satellite images. Unlike 3-channel RGB images, satellite imagery contains additional channels corresponding to different wavelengths. Approaches that do not use all channels are unable to fully exploit these images for optimal performance. Furthermore, care must be taken due to the large bias in classes, e.g., most of the Earth is covered in water and thus it will be dominant in the images. Our approach consists of training a Convolutional Neural Network (CNN) from scratch to classify multispectral image patches taken by satellites as whether or not they belong to a class of buildings. We then adapt the classification network to detection by converting the fully-connected layers of the network to convolutional layers, which allows the network to process images of any resolution. The dataset bias is compensated by subsampling negatives and tuning the detection threshold for optimal performance. We have constructed a new dataset using images from the Landsat 8 satellite for detecting solar power plants and show our approach is able to significantly outperform the state-of-the-art. Furthermore, we provide an indepth evaluation of the seven different spectral bands provided by the satellite images and show it is critical to combine them to obtain good results.
Tomohiro Ishii, Edgar Simo-Serra, Satoshi Iizuka, Yoshihiko Mochizuki, Akihiro Sugimoto, Hiroshi Ishikawa 0002, Ryosuke Nakamura
ICPR5
2016 Facial expression recognition by re-ranking with global and local generic features
abstract
Recognizing the facial expression plays an important role in human computer interaction. Following the recent success of the Convolutional Neural Network (CNN) in image classification and object recognition, this paper proposes a facial expression recognition method that makes full use of CNNs to detect face features globally and locally and that combines global and local generic features for improving accuracy in recognition. Our method uses global generic features with the Support Vector Machine (SVM) classifier to generate most plausible candidates in expression class while local generic features with the SVM classifier to look into the candidates to re-rank them for recognition. Experimental results using data-sets available in public support the effectiveness of our proposed method by demonstrating improved accuracy against the state-of-the-arts.
Duc Minh Vo, Akihiro Sugimoto
ICPR2
2016 Multi-view facial landmark detector learned by the Structured Output SVM
Michal Uricár, Vojtech Franc, Diego Thomas, Akihiro Sugimoto, Václav Hlavác
Image Vis. Comput.4
2015 Visual Attention Driven by Auditory Cues - Selecting Visual Features in Synchronization with Attracting Auditory Events
Jiro Nakajima, Akisato Kimura, Akihiro Sugimoto, Kunio Kashino
MMM (2)3
2015 Contrast Based Hierarchical Spatial-Temporal Saliency for Video
Trung-Nghia Le, Akihiro Sugimoto
PSIVT2
2013 Social Group Discovery from Surveillance Videos: A Data-Driven Approach with Attention-Based Cues
abstract
This paper presents an approach to discover social groups in surveillance videos by incorporating attention-based cues to model group behaviors of pedestrians in videos. Group behaviors are modeled as a set of decision trees with the decisions being basic measurements based on positionbased and attention-based cues. Rather than enforcing explicit models, we apply tree-based learning algorithms to implicitly construct the decision tree models. The experimental results demonstrate that incorporating attention-based cues significantly increased the estimation accuracy compared to the conventional approaches that used position-based cues alone.
Isarun Chamveha, Yusuke Sugano, Yoichi Sato 0001, Akihiro Sugimoto
BMVC4
2013 A Flexible Scene Representation for 3D Reconstruction Using an RGB-D Camera
abstract
Updating a global 3D model with live RGB-D measurements has proven to be successful for 3D reconstruction of indoor scenes. Recently, a Truncated Signed Distance Function (TSDF) volumetric model and a fusion algorithm have been introduced (KinectFusion), showing significant advantages such as computational speed and accuracy of the reconstructed scene. This algorithm, however, is expensive in memory when constructing and updating the global model. As a consequence, the method is not well scalable to large scenes. We propose a new flexible 3D scene representation using a set of planes that is cheap in memory use and, nevertheless, achieves accurate reconstruction of indoor scenes from RGB-D image sequences. Projecting the scene onto different planes reduces significantly the size of the scene representation and thus it allows us to generate a global textured 3D model with lower memory requirement while keeping accuracy and easiness to update with live RGB-D measurements. Experimental results demonstrate that our proposed flexible 3D scene representation achieves accurate reconstruction, while keeping the scalability for large indoor scenes.
Diego Thomas, Akihiro Sugimoto
ICCV2
2013 Learning to discover objects in RGB-D images using correlation clustering
abstract
We introduce a method to discover objects from RGB-D image collections which does not require a user to specify the number of objects expected to be found. We propose a probabilistic formulation to find pairwise similarity between image segments, using a classifier trained on labelled pairs from the recently released RGB-D Object Dataset. We then use a correlation clustering solver to both find the optimal clustering of all the segments in the collection and to recover the number of clusters. Unlike traditional supervised learning methods, our training data need not be of the same class or category as the objects we expect to discover. We show that this parameter-free supervised clustering method has superior performance to traditional clustering methods.
Michael Firman, Diego Thomas, Simon J. Julier, Akihiro Sugimoto
IROS4
2013 Incorporating Audio Signals into Constructing a Visual Saliency Map
Jiro Nakajima, Akihiro Sugimoto, Kazuhiko Kawamoto
PSIVT2
2013 Video Saliency Modulation in the HSI Color Space for Drawing Gaze
Akihiro Sugimoto
PSIVT2
2013 Head direction estimation from low resolution images with scene adaptation
Isarun Chamveha, Yusuke Sugano, Daisuke Sugimura, Teera Siriteerakul, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto
Comput. Vis. Image Underst.7
2013 Range Image Registration Using a Photometric Metric under Unknown Lighting
abstract
Based on the spherical harmonics representation of image formation, we derive a new photometric metric for evaluating the correctness of a given rigid transformation aligning two overlapping range images captured under unknown, distant, and general illumination. We estimate the surrounding illumination and albedo values of points of the two range images from the point correspondences induced by the input transformation. We then synthesize the color of both range images using albedo values transferred using the point correspondences to compute the photometric reprojection error. This way allows us to accurately register two range images by finding the transformation that minimizes the photometric reprojection error. We also propose a practical method using the proposed photometric metric to register pairs of range images devoid of salient geometric features, captured under unknown lighting. Our method uses a hypothesize-and-test strategy to search for the transformation that minimizes our photometric metric. Transformation candidates are efficiently generated by employing the spherical representation of each range image. Experimental results using both synthetic and real data demonstrate the usefulness of the proposed metric.
Diego Thomas, Akihiro Sugimoto
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Incorporating visual field characteristics into a saliency map
abstract
Characteristics of the human visual field are well known to be different in central (fovea) and peripheral areas. Existing computational models of visual saliency, however, do not take into account this biological evidence. The existing models compute visual saliency uniformly over the retina and, thus, have difficulty in accurately predicting the next gaze (fixation) point. This paper proposes to incorporate human visual field characteristics into visual saliency, and presents a computational model for producing such a saliency map. Our model integrates image features obtained by bottom-up computation in such a way that weights for the integration depend on the distance from the current gaze point where the weights are optimally learned using actual saccade data. The experimental results using a large number of fixation/saccade data with wide viewing angles demonstrate the advantage of our saliency map, showing that it can accurately predict the point where one looks next.
Hideyuki Kubota, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto, Kazuo Hiraki
ETRA5
2012 Optimal consensus set and preimage of 4-connected circles in a noisy environment
Gaëlle Skapin, Rita Zrour, Eric Andres, Akihiro Sugimoto, Yukiko Kenmochi
ICPR4
2012 Discrete Polynomial Curve Fitting to Noisy Data
Fumiki Sekiya, Akihiro Sugimoto
IWCIA2
2012 Illumination-free photometric metric for range image registration
abstract
This paper presents an illumination-free photometric metric for evaluating the goodness of a rigid transformation aligning two overlapping range images, under the assumption of Lambertian surface. Our metric is based on photometric re-projection error but not on feature detection and matching. We synthesize the color of one image using albedo of the other image to compute the photometric re-projection error. The unknown illumination and albedo are estimated from the correspondences induced by the input transformation using the spherical harmonics representation of image formation. This way allows us to derive an illumination-free photometric metric for range image alignment. We use a hypothesize-and-test method to search for the transformation that minimizes our illumination-free photometric function. Transformation candidates are efficiently generated by employing the spherical representation of each image. Experimental results using synthetic and real data show the usefulness of the proposed metric.
Diego Thomas, Akihiro Sugimoto
WACV2
2011 Structure-from-motion based hand-eye calibration using L∞ minimization
abstract
This paper presents a novel method for so-called hand-eye calibration. Using a calibration target is not possible for many applications of hand-eye calibration. In such situations Structure-from-Motion approach of hand-eye calibration is commonly used to recover the camera poses up to scaling. The presented method takes advantage of recent results in the L∞-norm optimization using Second-Order Cone Programming (SOCP) to recover the correct scale. Further, the correctly scaled displacement of the hand-eye transformation is recovered solely from the image correspondences and robot measurements, and is guaranteed to be globally optimal with respect to the L∞-norm. The method is experimentally validated using both synthetic and real world datasets.
Jan Heller, Michal Havlena, Akihiro Sugimoto, Tomás Pajdla
CVPR3
2011 Fast unsupervised ego-action learning for first-person sports videos
abstract
Portable high-quality sports cameras (e.g. head or helmet mounted) built for recording dynamic first-person video footage are becoming a common item among many sports enthusiasts. We address the novel task of discovering first-person action categories (which we call ego-actions) which can be useful for such tasks as video indexing and retrieval. In order to learn ego-action categories, we investigate the use of motion-based histograms and unsupervised learning algorithms to quickly cluster video content. Our approach assumes a completely unsupervised scenario, where labeled training videos are not available, videos are not pre-segmented and the number of ego-action categories are unknown. In our proposed framework we show that a stacked Dirichlet process mixture model can be used to automatically learn a motion histogram codebook and the set of ego-action categories. We quantitatively evaluate our approach on both in-house and public YouTube videos and demonstrate robust ego-action categorization across several sports genres. Comparative analysis shows that our approach outperforms other state-of-the-art topic models with respect to both classification accuracy and computational speed. Preliminary results indicate that on average, the categorical content of a 10 minute video sequence can be indexed in under 5 seconds.
Kris Makoto Kitani, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto
CVPR4
2011 Scale-Optimized Textons for Image Categorization and Segmentation
abstract
Texton is a representative dense visual word and it has proven its effectiveness in categorizing materials as well as generic object classes. Despite its success and popularity, no prior work has tackled the problem of its scale optimization for a given image data and associated object category. We propose scale-optimized textons to learn the best scale for each object in a scene, and incorporate them into image categorization and segmentation. Our textonization process produces a scale-optimized codebook of visual words. We approach the scale-optimization problem of textons by using the scene-context scale in each image, which is the effective scale of local context to classify an image pixel in a scene. We perform the textonization process using the randomized decision forest which is a powerful tool with high computational efficiency in vision applications. Our experiments using MSRC and VOC 2007 segmentation dataset show that our scale-optimized textons improve the performance of image categorization and segmentation.
Yousun Kang, Akihiro Sugimoto
ISM2
2011 Attention Prediction in Egocentric Video Using Motion and Visual Saliency
Kentaro Yamada, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto, Kazuo Hiraki
PSIVT (1)5
2011 Robustly registering range images using local distribution of albedo
Diego Thomas, Akihiro Sugimoto
Comput. Vis. Image Underst.2
2011 3D discrete rotations using hinge angles
Yohan Thibault, Akihiro Sugimoto, Yukiko Kenmochi
Theor. Comput. Sci.2
2010 Special issue on omnidirectional vision, camera networks and non-conventional cameras
João Pedro Barreto 0001, Tomás Pajdla, Akihiro Sugimoto
Comput. Vis. Image Underst.3
2009 Using individuality to track individuals: Clustering individual trajectories in crowds using local appearance and frequency trait
abstract
In this work, we propose a method for tracking individuals in crowds. Our method is based on a trajectory-based clustering approach that groups trajectories of image features that belong to the same person. The key novelty of our method is to make use of a person's individuality, that is, the gait features and the temporal consistency of local appearance to track each individual in a crowd. Gait features in the frequency domain have been shown to be an effective biometric cue in discriminating between individuals, and our method uses such features for tracking people in crowds for the first time. Unlike existing trajectory-based tracking methods, our method evaluates the dissimilarity of trajectories with respect to a group of three adjacent trajectories. In this way, we incorporate the temporal consistency of local patch appearance to differentiate trajectories of multiple people moving in close proximity. Our experiments show that the use of gait features and the temporal consistency of local appearance contributes to significant performance improvement in tracking people in crowded scenes.
Daisuke Sugimura, Kris Makoto Kitani, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto
ICCV5
2009 Hinge Angles for 3D Discrete Rotations
Yohan Thibault, Akihiro Sugimoto, Yukiko Kenmochi
IWCIA2
2009 Computing upper and lower bounds of rotation angles from digital images
Yohan Thibault, Yukiko Kenmochi, Akihiro Sugimoto
Pattern Recognit.3
2008 Computing Admissible Rotation Angles from Rotated Digital Images
Yohan Thibault, Yukiko Kenmochi, Akihiro Sugimoto
IWCIA3
2008 Recognizing Overlapped Human Activities from a Sequence of Primitive Actions via Deleted Interpolation
abstract
The high-level recognition of human activity requires a priori hierarchical domain knowledge as well as a means of reasoning based on that knowledge. Based on insights from perceptual psychology, the problem of human action recognition is approached on the understanding that activities are hierarchical, temporally constrained and at times temporally overlapped. A hierarchical Bayesian network (HBN) based on a stochastic context-free grammar (SCFG) is implemented to address the hierarchical nature of human activity recognition. Then it is shown how the HBN is applied to different substrings in a sequence of primitive action symbols via deleted interpolation (DI) to recognize temporally overlapped activities. Results from the analysis of action sequences based on video surveillance data show the validity of the approach.
Kris Makoto Kitani, Yoichi Sato 0001, Akihiro Sugimoto
Int. J. Pattern Recognit. Artif. Intell.3
2008 Recovering the Basic Structure of Human Activities from Noisy Video-Based Symbol Strings
abstract
In recent years stochastic context-free grammars have been shown to be effective in modeling human activities because of the hierarchical structures they represent. However, most of the research in this area has yet to address the issue of learning the activity grammars from a noisy input source, namely, video. In this paper, we present a framework for identifying noise and recovering the basic activity grammar from a noisy symbol string produced by video. We identify the noise symbols by finding the set of non-noise symbols that optimally compresses the training data, where the optimality of compression is measured using an MDL criterion. We show the robustness of our system to noise and its effectiveness in learning the basic structure of human activity, through experiments with artificial data and a real video sequence from a local convenience store.
Kris Makoto Kitani, Yoichi Sato 0001, Akihiro Sugimoto
Int. J. Pattern Recognit. Artif. Intell.3
2006 3D Head Tracking using the Particle Filter with Cascaded Classifiers
abstract
We propose a method for real-time people tracking using multiple cameras. The particle filter framework is known to be effective for tracking people, but most of existing methods adopt only simple perceptual cues such as color histogram or contour similarity for hypothesis evaluation. To improve the robustness and accuracy of tracking more sophisticated hypothesis evaluation is indispensable. We therefore present a novel technique for human head tracking using cascaded classifiers based on AdaBoost and Haar-like features for hypothesis evaluation. In addition, we use multiple classifiers, each of which is trained respectively to detect one direction of a human head. During real-time tracking the most suitable classifier is adaptively selected by considering each hypothesis and known camera position. Our experimental results demonstrate the effectiveness and robustness of our method. 1 1
Yoshinori Kobayashi, Daisuke Sugimura, Yoichi Sato 0001, Kousuke Hirasawa, Naohiko Suzuki, Hiroshi Kage, Akihiro Sugimoto
BMVC7
2000 Multilinear Relationships between the Coordinates of Corresponding Image Conics
abstract
This paper presents a study, based on conic correspondences, on the relationship between multiple images acquired by uncalibrated cameras. Representing image conics as points in the five-dimensional projective space allows one to handle image conics in the same way as image points. We show that the coordinates of corresponding image conics satisfy the multilinear constraints, as shown in the case for points and lines. To be more specific, the coordinates of two corresponding image conics satisfy bilinear constraints. When a third image comes in, the coordinates of three corresponding image conics satisfy trilinear constraints. Moreover, these constraints are naturally extended to the case where more images are available.
Akihiro Sugimoto, Takashi Matsuyama
ICPR1
1999 An Approximation Algorithm for the Two-Layered Graph Drawing Problem
Atsuko Yamaguchi, Akihiro Sugimoto
COCOON2
1998 Conic Based Image Transfer for 2-D Objects: A Linear Algorithm
Akihiro Sugimoto
ACCV (2)1
1996 Object recognition by combining paraperspective images
Akihiro Sugimoto
Int. J. Comput. Vis.1
1994 Geometric invariant of noncoplanar lines in a single view
abstract
The importance of geometric invariants to many machine vision tasks, such as model-based recognition, has been recognized. A number of studies on geometric invariants in a single view concentrate on coplanar objects: coplanar points, coplanar lines, coplanar conics, etc. Therefore, it is essentially only to 2-D objects that we can apply methods using geometric invariants. This paper presents a study on geometric invariants of noncoplanar objects, i.e., 3-D objects. A new geometric invariant is derived from six lines on three planes in a single view. The condition under which the invariant is nonsingular is also described. In addition, we present some experimental results with real images and find that the values of the invariant over a number of viewpoints remain stable even for noisy images.
Akihiro Sugimoto
ICPR (1)1