Takashi Shibata 0001

dblp:37/2234-1 · DBLP profile ↗
← Back
39ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0001-8072-3847ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 8 first-author · 22 since 2021Artificial intelligence and machine learning · 23 · 6 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Stegano-OBF: Privacy-Preserving Obfuscation for Action Recognition Datasets via Semantic Embedding
Takuya Nakabayashi, Yasunori Babazaki, Takashi Shibata 0001
ICPR (4)3
2026 Multi-camera Multi-object Tracking Based on Epipolar Distance and Appearance Similarity
Masamune Oka, Masayuki Tanaka 0001, Takashi Shibata 0001, Masatoshi Okutomi
ICPR (4)3
2025 Action-Agnostic Point-Level Supervision for Temporal Action Detection
abstract
We propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the proposed scheme, a small portion of video frames is sampled in an unsupervised manner and presented to human annotators, who then label the frames with action categories. Unlike point-level supervision, which requires annotators to search for every action instance in an untrimmed video, frames to annotate are selected without human intervention in AAPL supervision. We also propose a detection model and learning method to effectively utilize the AAPL labels. Extensive experiments on the variety of datasets (THUMOS'14, FineAction, GTEA, BEOID, and ActivityNet 1.3) demonstrate that the proposed approach is competitive with or outperforms prior methods for video-level and point-level supervision in terms of the trade-off between the annotation cost and detection performance.
Shuhei M. Yoshida, Takashi Shibata 0001, Makoto Terao, Takayuki Okatani, Masashi Sugiyama
AAAI2
2025 Semi-Automatic Labeling for Action Recognition by Diversity Preserving Sampling
abstract
Deep learning for action recognition is an important technology for understanding videos. However, collecting video training dataset for deep learning model with low cost while maintaining enough diversity is challenging. In this paper, we propose a semi-automatic labeling framework for action recognition by diversity-preserving sampling. The proposed framework utilizes a pre-trained vision-language model (VLM) to search through video clips to filter data that matches the text that describes the appropriate context for the target action. Since this simple approach by VLM tends to lack diversity, our framework is also equipped with diversity-preserving sampling that consists of two sampling strategies. One is confidence-based weighted sampling, which is based on action class confidence obtained from VLM, and the other is isolate-constraint-based weighted sampling, which samples points that are far apart in the text-image feature space. We conduct experiments to demonstrate that the proposed approach efficiently collects data with variations that could train a better action recognition model than the baseline.
Ryuhei Ando, Takashi Shibata 0001
ICASSP2
2025 Mask augmented Object-Centric Contrastive Learning for Amodal Instance Segmentation
abstract
Human cognition is robust in estimating depth ordering and occluded regions of objects, including amodal instance segmentation (AIS). Object-centric representation learning (OCRL) is an unsupervised approach to obtaining a new representation that mimics human common sense, such as amodal perception. Nevertheless, a significant gap exists between OCRL and human perception, and there is room for improvement for AIS. We aim to empower OCRL with amodal perception by solving self-supervised learning via contrastive learning for OCRL and depth-order estimation. The proposed method calculates the training loss on two masks composed of the object representations extracted from the original image and the transformed image by artificial occluders. Moreover, our method efficiently acquires depth-aware estimation by simultaneously solving the depth-ordering problem and representation learning. We have applied the proposed method to several simulation datasets and confirmed that the accuracy of AIS achieves SOTA performance under weakly supervised learning conditions.
Tomokazu Kaneko, Ryosuke Sakai, Takashi Shibata 0001, Soma Shiraishi
ICASSP3
2025 MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval
abstract
Result diversification (RD) is a crucial technique in Text-to-Image Retrieval for enhancing the efficiency of a practical application. Conventional methods focus solely on increasing the diversity metric of image appearances. However, the diversity metric and its desired value vary depending on the application, which limits the applications of RD. This paper proposes a novel task called CDR-CA (Contextual Diversity Refinement of Composite Attributes). CDR-CA aims to refine the diversities of multiple attributes, according to the application's context. To address this task, we propose Multi-Source DPPs, a simple yet strong baseline that extends the Determinantal Point Process (DPP) to multi-sources. We model MS-DPP as a single DPP model with a unified similarity matrix based on a manifold representation. We also introduce Tangent Normalization to reflect contexts. Extensive experiments demonstrate the effectiveness of the proposed method.
Naoya Sogi, Takashi Shibata 0001, Makoto Terao, Masanori Suganuma, Takayuki Okatani
IJCAI2
2025 Open-Vocabulary Scene Graph Generation via Synonym-Based Predicate Descriptor
Yuta Goto, Satoshi Yamazaki, Takashi Shibata 0001, Jianquan Liu
MMM (3)3
2025 MDCN-PS: Monocular-Depth-Guided Coarse Normal Attention for Robust Photometric Stereo
abstract
Photometric Stereo (PS) is a technique for estimating surface normals from images illuminated by multiple light sources. However, when the target object has a complex shape or the light sources are not appropriately arranged, certain regions may experience severe shadows, leading to insufficient information for accurate estimation. In this paper, we propose a Monocular-Depth-guided Coarse Normal attention for Photometric Stereo (MDCN-PS). The MDCN-PS can effectively combine monocular depth from a single image with PS with multiple light sources by a Photometric Stereo network Adaptor (PS Adaptor) with Coarse Normal Attention. The key is to use the coarse normals obtained from Monocular Depth Estimation as supplementary information, which can improve accuracy in regions where the light source is limited due to severe shadows or inhomogeneous light source distribution. Comprehensive experiments on real-world and synthetic datasets show that the proposed method achieved an accuracy improvement of 1.2 points in real-world datasets when limited to two input images and of 3.1 points in synthetic datasets in mean angular error compared to existing methods. Qualitative results also demonstrated that our method improves accuracy in areas with insufficient lighting patterns due to shadows.
Masahiro Yamaguchi 0001, Takashi Shibata 0001, Shoji Yachida, Keiko Yokoyama, Toshinori Hosoi
WACV2
2024 One-Shot Machine Unlearning with Mnemonic Code
Tomoya Yamashita, Masanori Yamada, Takashi Shibata 0001
ACML3
2024 Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
Naoya Sogi, Takashi Shibata 0001, Makoto Terao
ECCV (79)2
2024 Robust 3D Semantic Segmentation With Incomplete Point Clouds Based on Sequential Frame Sampling
abstract
This paper proposes a method for learning 3D semantic segmentation robust to incomplete point clouds. Our method first generates pseud-incomplete point clouds from original 3D point clouds by sequential frame sampling that creates multiple subsets considering the continuity of an RGB-D sequence for reproducing incomplete areas in the point clouds. It then simultaneously learns completion networks and semantic segmentation networks with the pseud-incomplete point clouds. We evaluate our method on the 3D semantic segmentation task. Experimental results on ScanNet v2, an indoor environment, show that our method improves mIoU by 0.4 points for the original point clouds and 6.3 points for the incomplete point clouds compared with a conventional method. Experimental results on WorkPlace Dataset, an outdoor environment, show that our method improves mIoU by 6.5 points for the original point clouds and 11.1 points for the incomplete point clouds compared with the conventional method. These results improve the safety and operability of environmental awareness in applications such as robotics.
Masahiro Yamaguchi 0001, Kyota Higa, Toshinori Hosoi, Takashi Shibata 0001
ICIP4
2024 Zero-Shot Spatio-Temporal Action Detection by Enhancing Context-Relation Capability of Vision-Language Models
Yasunori Babazaki, Takashi Shibata 0001
ICPR (29)2
2024 Task Success Classification with Final State of Future Prediction for Robot Control Planning
Taku Fujitomi, Naoya Sogi, Takashi Shibata 0001, Makoto Terao
ICPR (2)3
2024 Annotation-Free Object Detection by Knowledge-Extraction Training From Visual-Language Models
Yasuto Nagase, Yasunori Babazaki, Takashi Shibata 0001
ICPR (17)3
2024 PostAugment: Adversarial Data Augmentation with Hard Sample Suppression by Incorrect Class Likelihood
Azusa Sawada, Takashi Shibata 0001, Keiko Yokoyama, Shoji Yachida, Toshinori Hosoi
ICPR (10)2
2024 Disaster Damage Visualization by VLM-Based Interactive Image Retrieval and Cross-View Image Geo-Localization
abstract
We propose a framework for quickly selecting images that show the disaster situation from many images, estimating their locations with high accuracy, and displaying them on a map. The proposed framework introduces interactive image retrieval based on the Vision and Language Model (VLM), which can retrieve images from many images that show the disaster situation according to the user’s intention. Using the correlation between language and images based on VLM and the similarity between images selected interactively enables more accurate retrieval. Next, for selected images for which the location of the affected area is unknown, the location of the image is estimated with street address-level accuracy by matching it with an overhead image covering a large area of the city and map data and then displayed on a map. We confirmed the effectiveness of the proposed method on publicly available datasets such as CrisisNLP.
Naoya Sogi, Takashi Shibata 0001, Makoto Terao, Kenta Senzaki, Masahiro Tani, Royston Rodrigues
IGARSS2
2024 Future Predictive Success-or-Failure Classification for Long-Horizon Robotic Tasks
abstract
Automating long-horizon tasks with a robotic arm has been a central research topic in robotics. Optimization-based action planning is an efficient approach for creating an action plan to complete a given task. Construction of a reliable planning method requires a design process of conditions, e.g., to avoid collision between objects. The design process, however, has two critical issues: 1) iterative trials–the design process is time-consuming due to the trial-and-error process of modifying conditions, and 2) manual redesign–it is difficult to cover all the necessary conditions manually. To tackle these issues, this paper proposes a future-predictive success-or-failure-classification method to obtain conditions automatically. The key idea behind the proposed method is an end-to-end approach for determining whether the action plan can complete a given task instead of manually redesigning the conditions. The proposed method uses a long-horizon future-prediction method to enable success-or-failure classification without the execution of an action plan. This paper also proposes a regularization term called transition consistency regularization to provide easy-to-predict feature distribution. The regularization term improves future prediction and classification performance. The effectiveness of our method is demonstrated through classification and robotic-manipulation experiments.
Naoya Sogi, Hiroyuki Oyama, Takashi Shibata 0001, Makoto Terao
IJCNN3
2024 Black-Box Forgetting
abstract
Large-scale pre-trained models (PTMs) provide remarkable zero-shot classification capability covering a wide variety of object classes. However, practical applications do not always require the classification of all kinds of objects, and leaving the model capable of recognizing unnecessary classes not only degrades overall accuracy but also leads to operational disadvantages. To mitigate this issue, we explore the selective forgetting problem for PTMs, where the task is to make the model unable to recognize only the specified classes, while maintaining accuracy for the rest. All the existing methods assume ''white-box'' settings, where model information such as architectures, parameters, and gradients is available for training. However, PTMs are often ''black-box,'' where information on such models is unavailable for commercial reasons or social responsibilities. In this paper, we address a novel problem of selective forgetting for black-box models, named Black-Box Forgetting, and propose an approach to the problem. Given that information on the model is unavailable, we optimize the input prompt to decrease the accuracy of specified classes through derivative-free optimization. To avoid difficult high-dimensional optimization while ensuring high forgetting performance, we propose Latent Context Sharing, which introduces common low-dimensional latent components among multiple tokens for the prompt. Experiments on four standard benchmark datasets demonstrate the superiority of our method with reasonable baselines. The code is available at https://github.com/yusukekwn/Black-Box-Forgetting.
Yusuke Kuwana, Yuta Goto, Takashi Shibata 0001, Go Irie
NeurIPS3
2024 FRoG-MOT: Fast and Robust Generic Multiple-Object Tracking by IoU and Motion-State Associations
abstract
This paper proposes a generic multi-object tracking (MOT) algorithm that is robust to unexpected motion changes for generic objects. Deep learning has dramatically been improving MOT performances. Nevertheless, state-of-the-art tracking algorithms are still sensitive to unexpected motion changes and the generic object target beyond person tracking. This is because standard MOT benchmark datasets such as MOT17 mainly consist of persons in a crowd, often lacking unexpected shape and motion changes; thus, these issues have yet to be focused on. We propose a simple-yet-effective MOT framework that can dynamically improve tracking continuity by associating each target based on adaptively modified motion states. The keys are 1) to represent the target motions using multiple motion states that have weak correlations with each other and 2) to modify those states that have the lowest similarity to past states as outliers. Our approach can improve trajectory continuity and robustness to unexpected motion changes for generic objects. Comprehensive experiments have confirmed that our framework is comparable to existing state-of-the-art methods on a standard dataset and outperforms those algorithms on the GMOT dataset with an overall 2% improvement in IDF1, a measure of tracking continuity.
Takuya Ogawa, Takashi Shibata 0001, Toshinori Hosoi
WACV2
2024 Appearance-Based Curriculum for Semi-Supervised Learning with Multi-Angle Unlabeled Data
abstract
We propose an appearance-based curriculum (ABC) for a semi-supervised learning scenario where labeled images taken from limited angles and unlabeled ones taken from various angles are available for training. A common approach to semi-supervised learning relies on pseudo-labeling and data augmentation, but it struggles with large visual variations that cannot be covered by data augmentation. To solve this problem, ABC incrementally expands the pool of unlabeled images fed to a base semi-supervised learner so that newly added data are the ones most similar to those already in the pool. This way, the learner can assign pseudo-labels to the new data with high accuracy, keeping the quality of pseudo-labels higher than that when all the unlabeled data are processed at once, as customarily done in existing semi-supervised learning methods. We conducted extensive experiments and confirmed that our method outperforms the state-of-the-art semi-supervised learning methods in our scenario.
Shuhei M. Yoshida, Takashi Shibata 0001, Makoto Terao, Takayuki Okatani, Masashi Sugiyama
WACV3
2022 Robustizing Object Detection Networks Using Augmented Feature Pooling
Takashi Shibata 0001, Masayuki Tanaka 0001, Masatoshi Okutomi
ACCV (5)1
2022 Co-Attention-Guided Bilinear Model for Echo-Based Depth Estimation
abstract
Echoes reflect a geometric structure of a scene surrounding a sound source. In this paper, we address the problem of estimating depth maps of indoor scenes based on echoes. First, we experimentally show that fusing multiple acoustic features, especially spectrogram and angular spectrum, can improve estimation accuracy. We then propose a novel bilinear model that incorporates dense co-attention for effective feature fusion. Our model is able to obtain a compact fused feature while capturing the second-order correlations of intra-and inter-features. Thorough evaluations on two datasets demonstrate the superiority of the proposed method over the state-of-the-art echo-based depth estimation and feature fusion methods.
Go Irie, Takashi Shibata 0001, Akisato Kimura
ICASSP2
2021 Generalized Domain Adaptation
abstract
Many variants of unsupervised domain adaptation (UDA) problems have been proposed and solved individually. Its side effect is that a method that works for one variant is often ineffective for or not even applicable to another, which has prevented practical applications. In this paper, we give a general representation of UDA problems, named Generalized Domain Adaptation (GDA). GDA covers the major variants as special cases, which allows us to organize them in a comprehensive framework. Moreover, this generalization leads to a new challenging setting where existing methods fail, such as when domain labels are unknown, and class labels are only partially given to each domain. We propose a novel approach to the new setting. The key to our approach is self-supervised class-destructive learning, which enables the learning of class-invariant representations and domain-adversarial classifiers without using any domain labels. Extensive experiments using three benchmark datasets demonstrate that our method outperforms the state-of-the-art UDA methods in the new setting and that it is competitive in existing UDA variations as well.
Yu Mitsuzumi, Go Irie, Daiki Ikami, Takashi Shibata 0001
CVPR4
2021 Geometric Data Augmentation Based On Feature Map Ensemble
abstract
Deep convolutional networks have become the mainstream in computer vision applications. Although CNNs have been successful in many computer vision tasks, it is not free from drawbacks. The performance of CNN is dramatically degraded by geometric transformation, such as large rotations. In this paper, we propose a novel CNN architecture that can improve the robustness against geometric transformations without modifying the existing backbones of their CNNs. The key is to enclose the existing backbone with a geometric transformation (and the corresponding reverse transformation) and a feature map ensemble. The proposed method can inherit the strengths of existing CNNs that have been presented so far. Furthermore, the proposed method can be employed in combination with state-of-the-art data augmentation algorithms to improve their performance. We demonstrate the effectiveness of the proposed method using standard datasets such as CIFAR, CUB-200, and Mnist-rot-12k.
Takashi Shibata 0001, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP1
2021 Learning with Selective Forgetting
abstract
Lifelong learning aims to train a highly expressive model for a new task while retaining all knowledge for previous tasks. However, many practical scenarios do not always require the system to remember all of the past knowledge. Instead, ethical considerations call for selective and proactive forgetting of undesirable knowledge in order to prevent privacy issues and data leakage. In this paper, we propose a new framework for lifelong learning, called Learning with Selective Forgetting, which is to update a model for the new task with forgetting only the selected classes of the previous tasks while maintaining the rest. The key is to introduce a class-specific synthetic signal called mnemonic code. The codes are "watermarked" on all the training samples of the corresponding classes when the model is updated for a new task. This enables us to forget arbitrary classes later by only using the mnemonic codes without using the original data. Experiments on common benchmark datasets demonstrate the remarkable superiority of the proposed method over several existing methods.
Takashi Shibata 0001, Go Irie, Daiki Ikami, Yu Mitsuzumi
IJCAI1
2021 Constrained Weight Optimization for Learning without Activation Normalization
Daiki Ikami, Go Irie, Takashi Shibata 0001
WACV3
2020 Online Object Recognition Using CNN-based Algorithm on High-speed Camera Imaging: Framework for fast and robust high-speed camera object recognition based on population data cleansing and data ensemble
abstract
High-speed camera imaging (e.g., 1,000 fps) is effective to detect and recognize objects moving at high speeds because temporally dense images obtained by a high-speed camera can usually capture the best moment for object detection and recognition. However, the latest recognition algorithms, with their high complexity, are difficult to utilize in real-time applications involving high-speed cameras because a vast number of images need to be processed with no latency. To tackle this problem, we propose a novel framework for real-time object recognition with high-speed camera imaging. The proposed framework has the key processes of population data cleansing and data ensemble. Population data cleansing improves the recognition accuracy by quantifying the recognizability and by excluding part of the images prior to the recognition process, while data ensemble improves the robustness of object recognition by merging the class probabilities with multiple images from the object tracking sequence. Experimental results with a real dataset show that our framework is more effective than existing methods.
Shigeaki Namiki, Keiko Yokoyama, Shoji Yachida, Takashi Shibata 0001, Hiroyoshi Miyano, Masatoshi Ishikawa
ICPR4
2020 Reducing False Positives in Object Tracking with Siamese Network
abstract
We propose a robust long-term object tracking method that resolves the fundamental cause of the drift and loss of a target in visual object tracking. The proposed method consists of “sampling area extension”, which prevents a tracking result from drifting to other objects by learning false positive samples in advance (before they enter the search region of the target), and “adaptive search based on motion models”, which prevents a tracking result from drifting to other objects and avoids the loss of the target by using not only appearance features but also motion models to adaptively search for the target. Experiments conducted on long-term tracking dataset showed that our first technique improved robustness by 16.6% while the second technique improved robustness by 15.3%. By combining both, our method achieved 21.7% and 9.1% improvement for the robustness and precision, and the processing speed became 3.3 times faster. Additional experiments showed that our method achieved the top robustness among state-of-the-art methods on three long-term tracking datasets. These findings demonstrate that our method is effective for long-term object tracking and that its performance and speed are promising for use in practical applications of various technologies underlying object tracking.
Takuya Ogawa, Takashi Shibata 0001, Shoji Yachida, Toshinori Hosoi
ICPR2
2019 Fast and Robust Homography Estimation by Adaptive Graduated non-Convexity
abstract
This paper proposes a novel fast and robust homography estimation by adaptively controlling the threshold of graduated non-convexity (GNC). Based on the fact that GNC is a variant of deterministic annealing, we provide a new method for updating the inlier threshold at each GNC iteration by utilizing the statistical properties of residuals of potential inliers. Contrary to RANSAC, our approach gives the same unique parameter for a single input due to without random sampling. Moreover, computational time increases linearly against outlier ratio changes, whereas RANSAC increases exponentially. Synthetic data evaluation shows that the proposed method is more robust and faster than RANSAC for highly contaminated data containing more than 80% outliers. Additionally, we demonstrate that our method works on severe real images that the state-of-the-art RANSAC method fails.
Gaku Nakano, Takashi Shibata 0001
ICIP2
2019 Overcoming Labeling Ability for Latent Positives: Automatic Label Correction along Data Series
Azusa Sawada, Takashi Shibata 0001
ICPRAM2
2017 Misalignment-Robust Joint Filter for Cross-Modal Image Pairs
abstract
Although several powerful joint filters for cross-modal image pairs have been proposed, the existing joint filters generate severe artifacts when there are misalignments between a target and a guidance images. Our goal is to generate an artifact-free output image even from the misaligned target and guidance images. We propose a novel misalignment-robust joint filter based on weight-volume-based image composition and joint-filter cost volume. Our proposed method first generates a set of translated guidances. Next, the joint-filter cost volume and a set of filtered images are computed from the target image and the set of the translated guidances. Then, a weight volume is obtained from the joint-filter cost volume while considering a spatial smoothness and a label-sparseness. The final output image is composed by fusing the set of the filtered images with the weight volume for the filtered images. The key is to generate the final output image directly from the set of the filtered images by weighted averaging using the weight volume that is obtained from the joint-filter cost volume. The proposed framework is widely applicable and can involve any kind of joint filter. Experimental results show that the proposed method is effective for various applications including image denosing, image up-sampling, haze removal and depth map interpolation.
Takashi Shibata 0001, Masayuki Tanaka 0001, Masatoshi Okutomi
ICCV1
2016 Gradient-Domain Image Reconstruction Framework with Intensity-Range and Base-Structure Constraints
abstract
This paper presents a novel unified gradient-domain image reconstruction framework with intensity-range constraint and base-structure constraint. The existing method for manipulating base structures and detailed textures are classifiable into two major approaches: i) gradient-domain and ii) layer-decomposition. To generate detail-preserving and artifact-free output images, we combine the benefits of the two approaches into the proposed framework by introducing the intensity-range constraint and the base-structure constraint. To preserve details of the input image, the proposed method takes advantage of reconstructing the output image in the gradient domain, while the output intensity is guaranteed to lie within the specified intensity range, e.g. 0-to-255, by the intensity-range constraint. In addition, the reconstructed image lies close to the base structure by the base-structure constraint, which is effective for restraining artifacts. Experimental results show that the proposed framework is effective for various applications such as tone mapping, seamless image cloning, detail enhancement, and image restoration.
Takashi Shibata 0001, Masayuki Tanaka 0001, Masatoshi Okutomi
CVPR1
2016 Super high dynamic range video
abstract
High dynamic range (HDR) imaging is highly demanded in computer vision algorithms. An HDR image is composed with several low dynamic range (LDR) images, which usually have some disparities. In many HDR imaging algorithms, the disparities are estimated based on the texture information of the LDR images. However, the texture information is often lost completely if scenes include extremely bright and dark regions simultaneously. Recently, super high dynamic range (SHDR) imaging algorithm has been proposed where the disparities are estimated based on the segment shapes instead of the textures for handling such extreme scenes. In this paper, we extend the SHDR imaging algorithm to SHDR video generation introducing temporal smoothness terms. The temporal smoothness terms improve the temporal stability and the precision of the disparity estimation. Quantitative and qualitative evaluations demonstrate that the proposed algorithm outperforms existing algorithms.
Yuka Ogino, Masayuki Tanaka 0001, Takashi Shibata 0001, Masatoshi Okutomi
ICPR3
2016 Artifact-free image reconstruction for satellite imagery
abstract
This paper presents a novel image-reconstruction (IR) method for satellite imagery. To reduce unnatural artifacts, which often appear in existing IR methods, we propose a novel scheme to accurately estimate PSF and a novel regularization for IR. In the PSF estimation, hyper-parameters are estimated to reduce artifacts while improving sharpness based on our new criterion, which employs a residual of sigmoidal-function fitting to a strong edge on the reconstructed image as measurement of the amount of the artifacts. After the PSF estimation, we conduct IR based on the estimated PSF. In IR process, we employ a novel regularization that induces gradients of the reconstructed image to be close to those of its guide image. Since the guide image contains only the dominant structure of the input image without artifacts, the regularization leads to fewer artifacts. In addition, we employ a learning-based method for setting the spatially adaptive strength of the regularization effectively. Experimental results on real satellite imageries show that our method works better than other state-of-the-art IR methods.
Masato Ishii, Takashi Shibata 0001, Atsushi Sato
IGARSS2
2015 Unified image fusion based on application-adaptive importance measure
abstract
This paper presents a novel unified image fusion framework based on an application-adaptive importance measure. In the proposed method, an important area is selected pixel-by-pixel using the importance measure which is designed for each image type in each application. Then, the fused intensity is generated by a Poisson image editing. The main contribution is to provide a generalized image fusion framework enables us to deal with various different types of images for many applications. Experimental results show that the proposed method is effective for various applications including depth-perceptible image enhancement, temperature-preserving image fusion, optical flow fusion, and haze removal.
Takashi Shibata 0001, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP1
2014 Super-high Dynamic Range Imaging
abstract
We propose a novel high dynamic range (HDR) imaging algorithm for the scenes that contain an extremely wide range of scene radiance. In the HDR imaging, several images are taken under different exposures. Those images usually have displacement from one another due to camera and/or object motions. The challenge of the super HDR imaging is to align those images because any image contains "lost" regions where texture information is completely lost due to overexposure or underexposure. We propose an image alignment algorithm based on similarities of region shapes instead of the similarities of the textures. Experimental comparisons demonstrate that the proposed algorithm outperforms state-of-the-art algorithms.
Takehito Hayami, Masayuki Tanaka 0001, Masatoshi Okutomi, Takashi Shibata 0001, Shuji Senda
ICPR4
2014 RGB-D-T camera system for AR display of temperature change
abstract
The anomalies of power equipment can be founded using temperature changes compared to its normal state. In this paper we present a system for visualizing temperature changes in a scene using a thermal 3D model. Our approach is based on two precomputed 3D models of the target scene achieved with a RGB-D camera coupled with the thermal camera. The first model contains the RGB information, while the second one contains the thermal information. For comparing the status of the temperature between the model and the current time, we accurately estimate the pose of the camera by finding keypoint correspondences between the current view and the RGB 3D model. Knowing the pose of the camera, we are then able to compare the thermal 3D model with the current status of the temperature from any viewpoint.
Kazuki Matsumoto, Wataru Nakagawa, François de Sorbier, Maki Sugimoto, Hideo Saito 0001, Shuji Senda, Takashi Shibata 0001, Akihiko Iketani
ISMAR7
2012 Single Image Super Resolution Reconstruction in Perturbed Exemplar Sub-space
Takashi Shibata 0001, Akihiko Iketani, Shuji Senda
ACCV (3)1
2010 Image Inpainting Based on Probabilistic Structure Estimation
Takashi Shibata 0001, Akihiko Iketani, Shuji Senda
ACCV (3)1