Kiyoharu Aizawa

dblp:71/5426 · DBLP profile ↗
← Back
278ranked-venue papers
18as first author
56since 2021 · last 2026
0000-0003-2146-6275ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 260 · 17 first-author · 50 since 2021Artificial intelligence and machine learning · 45 · 16 since 2021Databases, data management, data science and information retrieval · 13 · 7 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Realistic Virtual Flood Experience System Using 360° Videos and 3D City Models Constructed from Building Footprints
abstract
Virtual flood experience systems, which enable users to vividly experience flooding, are attracting increasing attention as effective tools for communicating flood risks. However, existing systems typically rely on virtual cities that do not correspond to real locations and often lack sufficient photorealism, limiting users’ ability to relate scenarios to their own surroundings. Although 360° video-based virtual environments offer a simple and scalable way to visually replicate real-world scenes, effective 3D flood visualization in these environments typically requires 3D building geometry of the target area, which is not readily available in many regions. To address this limitation, we propose a new virtual flood experience framework that integrates 360° videos with 3D models automatically constructed from widely available 2D building footprints. By extruding footprints to plausible heights and spatially aligning the constructed models with 360° videos, our framework enables 3D flood visualization in photorealistic environments without relying on pre-existing city models such as CityGML. We demonstrate the framework in Memuro, Hokkaido, Japan, an area vulnerable to river flooding. A user study with local residents showed that the proposed system enhances users’ ability to envision location-specific flood evacuation situations, demonstrating its potential as an effective tool for disaster risk communication and education.
Tatsuro Banno, Koki Kawada, Mizuki Takenawa, Masatoshi Denda, Kiyoharu Aizawa
ICMR5
2026 Building and evaluating a realistic virtual world for large scale urban exploration from 360° videos
Mizuki Takenawa, Naoki Sugimoto, Leslie Wöhler, Satoshi Ikehata, Kiyoharu Aizawa
Multim. Tools Appl.5
2026 360CityGML: Realistic and Interactive Urban Visualization System Integrating CityGML Model and 360$^{\circ }$ Videos
abstract
We introduce a novel urban visualization system that integrates 3D urban model (CityGML) and 360$^{\circ }$∘ walkthrough videos. By aligning the videos with the model and dynamically projecting relevant video frames onto the geometries, our system creates photorealistic urban visualizations, allowing users to intuitively interpret geospatial data from a pedestrian view.
Tatsuro Banno, Mizuki Takenawa, Leslie Wöhler, Satoshi Ikehata, Kiyoharu Aizawa
IEEE Trans. Vis. Comput. Graph.5
2025 Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
abstract
This paper introduces a novel task to evaluate the robust understanding capability of Large Multimodal Models (LMMs), termed Unsolvable Problem Detection (UPD). Multiple-choice question answering (MCQA) is widely used to assess the understanding capability of LMMs, but it does not guarantee that LMMs truly comprehend the answer. UPD assesses the LMM’s ability to withhold answers when encountering unsolvable problems of MCQA, verifying whether the model truly understands the answer. UPD encompasses three problems: Absent Answer Detection (AAD), Incompatible Answer Set Detection (IASD), and Incompatible Visual Question Detection (IVQD), covering unsolvable cases like answer-lacking or incompatible choices and image-question mismatches. For the evaluation, we introduce the MM-UPD Bench, a benchmark for assessing performance across various ability dimensions. Our experiments reveal that even most LMMs, which demonstrate adequate performance on existing benchmarks, struggle significantly with MM-UPD, underscoring a novel aspect of trustworthiness that current benchmarks have overlooked. A detailed analysis shows that LMMs have different bottlenecks and chain-of-thought and self-reflection improved performance for LMMs with the bottleneck in their LLM capability. We hope our insights will enhance the broader understanding and development of more reliable LMMs.
Atsuyuki Miyai, Jingyang Zhang, Yifei Ming, Qing Yu 0013, Go Irie, Yixuan Li 0001, Hai Li 0001, Ziwei Liu 0002, Kiyoharu Aizawa
ACL (1)10
2025 Perface: Metric Learning in Perceptual Facial Similarity for Enhanced Face Anonymization
abstract
In response to rising societal awareness of privacy concerns, face anonymization techniques have advanced, including the emergence of face-swapping methods that replace one identity with another. Achieving a balance between anonymity and naturalness in face swapping requires careful selection of identities: overly similar faces compromise anonymity, while dissimilar ones reduce naturalness. Existing models, however, focus on binary identity classification "the same person or not", making it difficult to measure nuanced similarities such as "completely different" versus "highly similar but different." This paper proposes a human-perception-based face similarity metric, creating a dataset of 6,400 triplet annotations and metric learning to predict the similarity. Experimental results demonstrate significant improvements in both face similarity prediction and attribute-based face classification tasks over existing methods. Our dataset is available at https://github.com/kumanotanin/PerFace.
Haruka Kumagai, Leslie Wöhler, Satoshi Ikehata, Kiyoharu Aizawa
ICIP4
2025 A Benchmark and Evaluation for Real-World Out-of-Distribution Detection Using Vision-Language Models
abstract
Out-of-distribution (OOD) detection is a task that detects OOD samples during inference to ensure the safety of deployed models. However, conventional benchmarks have reached performance saturation, making it difficult to compare recent OOD detection methods. To address this challenge, we introduce three novel OOD detection benchmarks that enable a deeper understanding of method characteristics and reflect real-world conditions. First, we present ImageNet-X, designed to evaluate performance under challenging semantic shifts. Second, we propose ImageNet-FS-X for full-spectrum OOD detection, assessing robustness to covariate shifts (feature distribution shifts). Finally, we propose Wilds-FS-X, which extends these evaluations to real-world datasets, offering a more comprehensive testbed. Our experiments reveal that recent CLIP-based OOD detection methods struggle to varying degrees across the three proposed benchmarks, and none of them consistently outperforms the others. We hope the community goes beyond specific benchmarks and includes more challenging conditions reflecting real-world scenarios. The code is https://github.com/hoshi23/OOD-X-Benchmarks.
Shiho Noda, Atsuyuki Miyai, Qing Yu 0013, Go Irie, Kiyoharu Aizawa
ICIP5
2025 Redefining Image-to-Recipe Retrieval with Nutritional and Ingredient Similarity
abstract
Cross-modal retrieval models have shown impressive performance on the image-to-recipe retrieval task, a common benchmark in the multimedia field. However, the task assumes that an exact recipe match for a query image exists in the target database—an assumption that rarely holds true in real-world scenarios. When excluding exact matches from the target domain, our analysis revealed that relying solely on visual and textual similarity between recipes is insufficient to achieve good retrieval results. Other similarities should also be considered. Since ingredient similarity aligns with human intuition and nutritional similarity is crucial for health-conscious applications, we propose a model that incorporates ingredient and nutritional relevance into the retrieval process. We measured the similarity of unpaired recipes using three new metrics: mean absolute scaled error (MASE) for assessing nutritional similarity and IOU and weighted IOU (WIOU) for measuring ingredient overlap. Our proposed method can also be applied with existing image-recipe retrieval models and improved top-1 MASE, IOU, and WIOU by up to 18.13%, 9.91%, and 6.88%.
Satayu Parinayok, Shin'ichi Satoh 0001, Kiyoharu Aizawa, Yoko Yamakata
ICME3
2025 A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
Mashiro Toyooka, Kiyoharu Aizawa, Yoko Yamakata
ACM Multimedia2
2025 FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
abstract
Food image classification models are crucial for dietary management applications because they reduce the burden of manual meal logging. However, most publicly available datasets for training such models rely on web-crawled images, which often differ from users' real-world meal photos. In this work, we present FoodLogAthl-218, a food image dataset constructed from real-world meal records collected through the dietary management application FoodLog Athl. The dataset contains 6,925 images across 218 food categories, with a total of 14,349 bounding boxes. Rich metadata, including meal date and time, anonymized user IDs, and meal-level context, accompany each image. Unlike conventional datasets-where a predefined class set guides web-based image collection-our data begins with user-submitted photos, and labels are applied afterward. This yields greater intra-class diversity, a natural frequency distribution of meal types, and casual, unfiltered images intended for personal use rather than public sharing. In addition to (1) a standard classification benchmark, we introduce two FoodLog-specific tasks: (2) an incremental fine-tuning protocol that follows the temporal stream of users' logs, and (3) a context-aware classification task where each image contains multiple dishes, and the model must classify each dish by leveraging the overall meal context. We evaluate these tasks using large multimodal models (LMMs). The dataset is publicly available at https://huggingface.co/datasets/FoodLog/FoodLogAthl-218.
Mitsuki Watanabe, Sosuke Amano, Kiyoharu Aizawa, Yoko Yamakata
ACM Multimedia3
2025 FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
Yuki Imajuku, Yoko Yamakata, Kiyoharu Aizawa
MMM (1)3
2025 JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
abstract
Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira, Jeonghun Baek, Xiang Yue, Graham Neubig, Kiyoharu Aizawa. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Shota Onohara, Atsuyuki Miyai, Yuki Imajuku, Kazuki Egashira, Jeonghun Baek, Xiang Yue, Graham Neubig, Kiyoharu Aizawa
NAACL (Long Papers)8
2025 Open-set domain adaptation with visual-language foundation models
Qing Yu 0013, Go Irie, Kiyoharu Aizawa
Comput. Vis. Image Underst.3
2025 GL-MCM: Global and Local Maximum Concept Matching for Zero-Shot Out-of-Distribution Detection
abstract
Abstract Zero-shot OOD detection is a task that detects OOD images during inference with only in-distribution (ID) class names. Existing methods assume ID images contain a single, centered object, and do not consider the more realistic multi-object scenarios, where both ID and OOD objects are present. To meet the needs of many users, the detection method must have the flexibility to adapt the type of ID images. To this end, we present Global-Local Maximum Concept Matching (GL-MCM), which incorporates local image scores as an auxiliary score to enhance the separability of global and local visual features. Due to the simple ensemble score function design, GL-MCM can control the type of ID images with a single weight parameter. Experiments on ImageNet and multi-object benchmarks demonstrate that GL-MCM outperforms baseline zero-shot methods and is comparable to fully supervised methods. Furthermore, GL-MCM offers strong flexibility in adjusting the target type of ID images. The code is available via https://github.com/AtsuMiyai/GL-MCM .
Atsuyuki Miyai, Qing Yu 0013, Go Irie, Kiyoharu Aizawa
Int. J. Comput. Vis.4
2025 Field-of-View IoU for Object Detection in 360° Images
abstract
360°cameras have gained popularity over the last few years. In this paper, we propose two fundamental techniques-Field-of-View IoU (FoV-IoU) and 360Augmentation for object detection in 360° images. Although most object detection neural networks designed for perspective images are applicable to 360° images in equirectangular projection (ERP) format, their performance deteriorates owing to the distortion in ERP images. Our method can be readily integrated with existing perspective object detectors and significantly improves the performance. The FoV-IoU computes the intersection-over-union of two Field-of-View bounding boxes in a spherical image which could be used for training, inference, and evaluation while 360Augmentation is a data augmentation technique specific to 360° object detection task which randomly rotates a spherical image and solves the bias due to the sphere-to-plane projection. We conduct extensive experiments on the 360° indoor dataset with different types of perspective object detectors and show the consistent effectiveness of our method.
Satoshi Ikehata, Kiyoharu Aizawa
IEEE Trans. Image Process.3
2025 Guest Editorial: When Multimedia Meets Food: Multimedia Computing for Food Data Analysis and Applications
Weiqing Min, Shuqiang Jiang, Petia Radeva, Vladimir Pavlovic 0001, Chong-Wah Ngo, Kiyoharu Aizawa, Wanqing Li 0001
IEEE Trans. Multim.6
2024 Retrieval-Augmented Layout Transformer for Content-Aware Layout Generation
abstract
Content-aware graphic layout generation aims to automatically arrange visual elements along with a given content, such as an e-commerce product image. In this paper, we argue that the current layout generation approaches suffer from the limited training data for the high-dimensional layout structure. We show that a simple retrieval augmentation can significantly improve the generation quality. Our model, which is named Retrieval-Augmented Layout Transformer (RALF),retrieves nearest neighbor layout examples based on an input image and feeds these results into an autoregressive generator. Our model can apply retrieval augmentation to various controllable generation tasks and yield high-quality layouts within a unified architecture. Our extensive experiments show that RALF successfully generates content-aware layouts in both constrained and unconstrained settings and significantly outperforms the baselines.1
Daichi Horita, Naoto Inoue, Kotaro Kikuchi, Kota Yamaguchi, Kiyoharu Aizawa
CVPR5
2024 Entity-NeRF: Detecting and Removing Moving Entities in Urban Scenes
abstract
Recent advancements in the study of Neural Radiance Fields (NeRF) for dynamic scenes often involve explicit modeling of scene dynamics. However, this approach faces challenges in modeling scene dynamics in urban environments, where moving objects of various categories and scales are present. In such settings, it becomes crucial to effectively eliminate moving objects to accurately reconstruct static backgrounds. Our research introduces an innovative method, termed here as Entity-NeRF, which combines the strengths of knowledge-based and statistical strategies. This approach utilizes entity-wise statistics, leveraging entity segmentation and stationary entity classification through thing/stuff segmentation. To assess our methodology, we created an urban scene dataset masked with moving objects. Our comprehensive experiments demonstrate that Entity-NeRF notably outperforms existing techniques in removing moving objects and reconstructing static urban backgrounds, both quantitatively and qualitatively.11Our project page is available at https://otonari726.github.io/entitynerf/
Takashi Otonari, Satoshi Ikehata, Kiyoharu Aizawa
CVPR3
2024 The Lottery Ticket Hypothesis in Denoising: Towards Semantic-Driven Initialization
Jiafeng Mao, Kiyoharu Aizawa
ECCV (74)3
2024 Cross-Lingual Learning in Multilingual Scene Text Recognition
abstract
In this paper, we investigate cross-lingual learning (CLL) for multilingual scene text recognition (STR). CLL transfers knowledge from one language to another. We aim to find the condition that exploits knowledge from high-resource languages for improving performance in low-resource languages. To do so, we first examine if two general insights about CLL discussed in previous works are applied to multilingual STR: (1) Joint learning with high- and low-resource languages may reduce performance on low-resource languages, and (2) CLL works best between typologically similar languages. Through extensive experiments, we show that two general insights may not be applied to multilingual STR. After that, we show that the crucial condition for CLL is the dataset size of high-resource languages regardless of the kind of high-resource languages. Our code, data, and models are available at https://github.com/ku21fan/CLL-STR.
Jeonghun Baek, Yusuke Matsui 0001, Kiyoharu Aizawa
ICASSP3
2024 Manga109Dialog: A Large-Scale Dialogue Dataset for Comics Speaker Detection
abstract
The expanding market for e-comics has driven the development of automated methods for analyzing comics. To enhance the machine’s understanding of comics, an automated method is essential for linking text in comics to characters that speak those words. In this study, we developed Manga109Dialog1, which is the world’s largest speaker-to-text annotation dataset for comics, containing 132,692 pairs. We proposed a novel deep learning-based method using scene graph generation models. To tailor the unique features of comics, we enhanced the performance by considering the frame reading order. Our experiments with Manga109Dialog show that our scene-graph-based approach outperforms existing methods, achieving a prediction accuracy of over 75%, thus establishing a robust benchmark for speaker detection in comics.
Yingxuan Li, Kiyoharu Aizawa, Yusuke Matsui 0001
ICME2
2024 Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
Yingxuan Li, Ryota Hinami, Kiyoharu Aizawa, Yusuke Matsui 0001
ACM Multimedia3
2024 Measure and Improve Your Food: Ingredient Estimation Based Nutrition Calculator
Yoko Yamakata, Ryoma Maeda, Kiyoharu Aizawa
ACM Multimedia4
2024 Investigating the Perception of Facial Anonymization Techniques in 360° Videos
abstract
In this work, we investigate facial anonymization techniques in 360° videos and assess their influence on the perceived realism, anonymization effect, and presence of participants. In comparison to traditional footage, 360° videos can convey engaging, immersive experiences that accurately represent the atmosphere of real-world locations. As the entire environment is captured simultaneously, it is necessary to anonymize the faces of bystanders in recordings of public spaces. Since this alters the video content, the perceived realism and immersion could be reduced. To understand these effects, we compare non-anonymized and anonymized 360° videos using blurring, black boxes, and face-swapping shown either on a regular screen or in a head-mounted display (HMD). Our results indicate significant differences in the perception of the anonymization techniques. We find that face-swapping is the most realistic and least disruptive; however, participants raised concerns regarding the effectiveness of the anonymization. Furthermore, we observe that presence is affected by facial anonymization in HMD condition. Overall, the results underscore the need for facial anonymization techniques that balance both photo-realism and a sense of privacy.
Leslie Wöhler, Satoshi Ikehata, Kiyoharu Aizawa
ACM Trans. Appl. Percept.3
2024 Self-Labeling Framework for Open-Set Domain Adaptation With Few Labeled Samples
abstract
Unsupervised domain adaptation (UDA) is extremely effective for transferring knowledge from a label-rich source domain to a label-scarce target domain. Because the target domain is unlabeled and may contain additional novel classes, open-set domain adaptation (ODA) has been suggested as a possible solution to detect these novel classes in the training phase. However, existing ODA methods rely heavily on abundant fully labeled source data, which are expensive to collect in specific applications and may also contain novel classes. In this study, we propose a novel self-labeling framework with prototypical contrastive learning and mutual information maximization to achieve ODA even when the amount of labeled data is very small, which is a new problem setting named few-shot ODA (FODA). We use self-supervised prototypical contrastive learning to train the network to learn the representations of source and target samples and maximize the mutual information between labels and input data to simultaneously recognize known and novel classes in the source and target domains. We evaluated our strategy in several domain adaptation environments and found that our method performed far better than existing approaches.
Qing Yu 0013, Go Irie, Kiyoharu Aizawa
IEEE Trans. Multim.3
2023 A Structure-Guided Diffusion Model for Large-Hole Image Completion
Daichi Horita, Jiaolong Yang, Dong Chen 0003, Yuki Koyama 0001, Kiyoharu Aizawa, Nicu Sebe
BMVC5
2023 Restorable Visible and Infrared Image Fusion
abstract
Image fusion aims to synthesize multiple source images into a single image to integrate and enhance information. Specifically, we tackle the fusion of visible and infrared images. Previous works generally use structural similarity between the fusion and the paired source images to train a deep-learning-based fusion model. However, only using the structure often results in a texture-insufficient image. In this study, we aim to generate an image rich in texture. This study is inspired by the ability of an autoencoder to learn a compressed representation of the input image. Specifically, we learn a fusion image with the structure and texture of the source images. We propose a novel framework–Restorable visible and infrared Image Fusion, which consists of a fusion and decoupling network. The fusion network synthesizes source images, and the decoupling network restores the source images by decomposing a fusion image. Our framework can be trained by minimizing the difference between the source and restored images. The experimental results demonstrate that the fusion image generated by the proposed method maintains the texture of the source images.
Daichi Horita, Koki Tsubota, Kiyoharu Aizawa
ICIP4
2023 Noise-Avoidance Sampling for Annotation Missing Object Detection
abstract
Excellent results can be achieved using object detection with fully supervised training on large well-annotated datasets. However, the problem of missing annotations in real-world datasets can considerably reduce the performance of object detectors. In this study, we thoroughly analyze the effect of missing annotations on both positive and negative samples in object detector training. To mitigate the negative impact caused by annotation missing problem, we propose a simple yet effective method, noise-avoidance sampling, to distinguish noisy training samples and subsequently reduce their negative impact. Experiments are conducted on the PASCAL VOC 07+12 dataset with varying levels of missing annotations. The results reveal that the proposed method achieves comparable or superior performance with state-of-the-art methods.
Jiafeng Mao, Qing Yu 0013, Go Irie, Kiyoharu Aizawa
ICIP4
2023 Text-to-Image Fashion Retrieval with Fabric Textures
abstract
In this study, we proposed text-to-image fashion image retrieval that captures the texture of clothing fabrics. A fabric’s texture is a major factor governing the comfort and appearance of clothes and significantly influences user preferences. However, unlike patterns and shapes that can readily be captured from a global image of the entire piece of clothing, extracting the fine and ambiguous characteristics of textures is considerably more challenging. The key concept is that by focusing on the "local" regions of clothing, detailed fabric textures can be more accurately captured. To this end, we propose a framework for learning cross-modal features from both global (the entire garment) and local (a close-up detail) image-text pairs. To verify the idea, we constructed a new dataset named Global and Local FACAD (G&L FACAD) by modifying the existing large-scale public FACAD dataset used for fashion retrieval. The experimental results confirm that the retrieval accuracy is significantly improved compared to the baselines. The code is available at https://github.com/SuzukiDaichi-git/texture_aware_fashion_retrieval.git.
Daichi Suzuki, Go Irie, Kiyoharu Aizawa
ICMR3
2023 Guided Image Synthesis via Initial Image Editing in Diffusion Model
abstract
Diffusion models have the ability to generate high quality images by denoising pure Gaussian noise images. While previous research has primarily focused on improving the control of image generation through adjusting the denoising process, we propose a novel direction of manipulating the initial noise to control the generated image. Through experiments on stable diffusion, we show that blocks of pixels in the initial latent images have a preference for generating specific content, and that modifying these blocks can significantly influence the generated image. In particular, we show that modifying a part of the initial image affects the corresponding region of the generated image while leaving other regions unaffected, which is useful for repainting tasks. Furthermore, we find that the generation preferences of pixel blocks are primarily determined by their values, rather than their position. By moving pixel blocks with a tendency to generate user-desired content to user-specified regions, our approach achieves state-of-the-art performance in layout-to-image generation. Our results highlight the flexibility and power of initial image manipulation in controlling the generated image.
Jiafeng Mao, Kiyoharu Aizawa
ACM Multimedia3
2023 360RVW: Fusing Real 360° Videos and Interactive Virtual Worlds
abstract
We propose a system to generate 360° realistic virtual worlds (360RVW) for the interactive spatial exploration of omnidirectional street-view videos. Our 360RVW enables users to explore photorealistic scenes with digital avatars, and interact with others. To create the virtual worlds our system only requires 360° videos with annotations of the start and end camera coordinate as input. We first detect street intersections to divide the input videos and remove the camera operator from the recordings using a video completion technique. Next, we analyze the 3D structure of the scene using semantic segmentation to define walkable areas. Finally, we render the environment using an ellipsoid projection surface to achieve a more realistic integration of the avatar into real-world 360° videos. The whole process is largely automated, enabling users to produce realistic and interactive virtual worlds without specialized skills or time-consuming manual interventions.
Mizuki Takenawa, Naoki Sugimoto, Leslie Wöhler, Satoshi Ikehata, Kiyoharu Aizawa
ACM Multimedia5
2023 Open-Vocabulary Segmentation Approach for Transformer-Based Food Nutrient Estimation
abstract
Nutrition plays a vital role in overall health and well-being. With a highly accurate nutrient estimation model, we develop a tool that displays nutritional values from food images, thereby reducing the labor-intensiveness of dietary assessment. We propose a method that uses depth data with RGB images and incorporates an open-vocabulary segmentation process that separates food from non-food instances, coupled with two-stage self-attention Transformer decoder. Our model outperforms the current state-of-the-art method, with an average percent MAE of 17.2% on Nutrition5k, an RGB-D food image dataset with calories, mass, and three macronutrients annotated. Our study also focuses on the significance of the food and background regions for calorie, mass, and nutrient estimation. We analyze the impact of non-food regions on each estimation task, with results suggesting that background information is crucial for calorie, mass, and carbohydrate estimation but not as essential for protein and fat estimation. The qualitative results also show that the model attends to regions with a high corresponding nutritional value. Implementation codes and pre-trained models are provided at https://github.com/Oatsty/nutrition5k.
Satayu Parinayok, Yoko Yamakata, Kiyoharu Aizawa
MMAsia3
2023 Automatic Dataset Creation from User-generated Recipes for Ingredient-centric Food Image Analysis
abstract
We aim to develop an application that automatically creates a nutrition facts label from food images for precise dietary control. Firstly, we constructed a new dataset with food category labels and a list of ingredients in a nutritionally calculable format using an image classification model and BERT for 1.6 million recipes accompanied by images. The nutritional value of the recipe can be calculated using a conversion table consisting of the food item number and unit class. Next, using deep learning techniques, we built models that estimate the list of food item numbers from food images. While the multi-task model that identifies the food category label and the ingredient list simultaneously is only effective within a limited number of recipes, the single-task model that only identified the ingredient list achieved a Micro-F1 of 53.32% in total.
Yoko Yamakata, Kiyoharu Aizawa
MMAsia3
2023 LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt Learning
abstract
We present a novel vision-language prompt learning approach for few-shot out-of-distribution (OOD) detection. Few-shot OOD detection aims to detect OOD images from classes that are unseen during training using only a few labeled in-distribution (ID) images. While prompt learning methods such as CoOp have shown effectiveness and efficiency in few-shot ID classification, they still face limitations in OOD detection due to the potential presence of ID-irrelevant information in text embeddings. To address this issue, we introduce a new approach called $\textbf{Lo}$cal regularized $\textbf{Co}$ntext $\textbf{Op}$timization (LoCoOp), which performs OOD regularization that utilizes the portions of CLIP local features as OOD features during training. CLIP's local features have a lot of ID-irrelevant nuisances ($\textit{e.g.}$, backgrounds), and by learning to push them away from the ID class text embeddings, we can remove the nuisances in the ID class text embeddings and enhance the separation between ID and OOD. Experiments on the large-scale ImageNet OOD detection benchmarks demonstrate the superiority of our LoCoOp over zero-shot, fully supervised detection methods and prompt learning methods. Notably, even in a one-shot setting -- just one label per class, LoCoOp outperforms existing zero-shot and fully supervised detection methods. The code is available via https://github.com/AtsuMiyai/LoCoOp.
Atsuyuki Miyai, Qing Yu 0013, Go Irie, Kiyoharu Aizawa
NeurIPS4
2023 Rethinking Rotation in Self-Supervised Contrastive Learning: Adaptive Positive or Negative Data Augmentation
abstract
Rotation is frequently listed as a candidate for data augmentation in contrastive learning but seldom provides satisfactory improvements. We argue that this is because the rotated image is always treated as either positive or negative. The semantics of an image can be rotation-invariant or rotation-variant, so whether the rotated image is treated as positive or negative should be determined based on the content of the image. Therefore, we propose a novel augmentation strategy, adaptive Positive or Negative Data Augmentation (PNDA), in which an original and its rotated image are a positive pair if they are semantically close and a negative pair if they are semantically different. To achieve PNDA, we first determine whether rotation is positive or negative on an image-by-image basis in an unsupervised way. Then, we apply PNDA to contrastive learning frameworks. Our experiments showed that PNDA improves the performance of contrastive learning. The code is available at https://github.com/AtsuMiyai/rethinking_rotation.
Atsuyuki Miyai, Qing Yu 0013, Daiki Ikami, Go Irie, Kiyoharu Aizawa
WACV5
2023 Universal Deep Image Compression via Content-Adaptive Optimization with Adapters
abstract
Deep image compression performs better than conventional codecs, such as JPEG, on natural images. However, deep image compression is learning-based and en-counters a problem: the compression performance deteriorates significantly for out-of-domain images. In this study, we highlight this problem and address a novel task: universal deep image compression. This task aims to compress images belonging to arbitrary domains, such as natural images, line drawings, and comics. To address this problem, we propose a content-adaptive optimization framework; this framework uses a pre-trained compression model and adapts the model to a target image during compression. Adapters are inserted into the decoder of the model. For each input image, our framework optimizes the latent representation extracted by the encoder and the adapter parameters in terms of rate-distortion. The adapter parameters are additionally transmitted per image. For the experiments, a benchmark dataset containing uncompressed images of four domains (natural images, line drawings, comics, and vector arts) is constructed and the proposed universal deep compression is evaluated. Finally, the proposed model is compared with non-adaptive and existing adaptive compression models. The comparison reveals that the proposed model outperforms these. The code and dataset are publicly available at https://github.com/kktsubota/universal-dic.
Koki Tsubota, Hiroaki Akutsu, Kiyoharu Aizawa
WACV3
2023 Guest Editorial Introduction to the Special Issue on Video Transformers
abstract
Currently, Transformer has been widely used in natural language and image processing and has achieved excellent results. Benefiting from the self-attention operation and global interaction, Transformer has demonstrated more powerful spatiotemporal modeling capabilities than traditional convolutional and recurrent neural networks. However, research on video Transformer is still in its infancy. Specifically, with the development of internet technology, video data has become a commonly used medium, playing a critical role in many areas such as entertainment, education, healthcare, security, etc. Different from static data such as images and text, video data consists of a series of image frames and is more concerned with temporal and motion information, which makes it necessary to employ some adaptations and well-designed network architectures to capture the discriminative features. In addition, the multi-modal information attached to video data further increases the difficulty of applying Transformer to videos.
Liqiang Nie, Jianlong Wu, Nicu Sebe, Kiyoharu Aizawa
IEEE Trans. Circuits Syst. Video Technol.4
2022 Self-Labeling Framework for Novel Category Discovery over Domains
abstract
Unsupervised domain adaptation (UDA) has been highly successful in transferring knowledge acquired from a label-rich source domain to a label-scarce target domain. Open-set domain adaptation (open-set DA) and universal domain adaptation (UniDA) have been proposed as solutions to the problem concerning the presence of additional novel categories in the target domain. Existing open-set DA and UniDA approaches treat all novel categories as one unified unknown class and attempt to detect this unknown class during the training process. However, the features of the novel categories learned by these methods are not discriminative. This limits the applicability of UDA in the further classification of these novel categories into their original categories, rather than assigning them to a single unified class. In this paper, we propose a self-labeling framework to cluster all target samples, including those in the ''unknown'' categories. We train the network to learn the representations of target samples via self-supervised learning (SSL) and to identify the seen and unseen (novel) target-sample categories simultaneously by maximizing the mutual information between labels and input data. We evaluated our approach under different DA settings and concluded that our method generally outperformed existing ones by a wide margin.
Qing Yu 0013, Daiki Ikami, Go Irie, Kiyoharu Aizawa
AAAI4
2022 Non-uniform Sampling Strategies for NeRF on 360° images
Takashi Otonari, Satoshi Ikehata, Kiyoharu Aizawa
BMVC3
2022 COO: Comic Onomatopoeia Dataset for Recognizing Arbitrary or Truncated Texts
Jeonghun Baek, Yusuke Matsui 0001, Kiyoharu Aizawa
ECCV (28)3
2022 SVG Vector Font Generation for Chinese Characters with Transformer
abstract
Designing fonts for Chinese characters is highly labor-intensive and time-consuming. While the latest methods successfully generate the English alphabet vector font, despite the high demand for automatic font generation, Chinese vector font generation has been an unsolved problem owing to its complex shape and numerous characters. This study addressed the problem of automatically generating Chinese vector fonts from only a single style and content reference. We proposed a novel network architecture with Transformer and loss functions to capture structural features without differentiable rendering. Although the dataset range was still limited to the sans-serif family, we successfully generated the Chinese vector font for the first time using the proposed method.
Haruka Aoki, Kiyoharu Aizawa
ICIP2
2022 Translation of Illustration Artist Style Using Sailormoonredraw Data
abstract
The decision to draw requires answers to two questions: what to draw and how to draw. The latter refers to artist style and is an important factor in creating any illustration. In this paper, we propose a novel task, artist style translation, which translates one artist’s illustration into that of another artist style using deep learning. To solve this task in a supervised manner, we created a novel illustration dataset, SailormoonDataset, which consists of more than 2,000 artist’s stylistic illustrations of the same content. In addition, we propose a method based on the Swapping Autoencoder by introducing a new loss function for supervised learning and using multiple images to represent an artist style. We translate the face illustration of Sailor Moon into styles of different artists. By comparing the current results to those of the Swapping Autoencoder, we find that the proposed method successfully achieves the artist style translation.
Keita Awane, Daichi Horita, Hikaru Ikuta, Yusuke Matsui 0001, Kiyoharu Aizawa, Naohiro Yanase
ICIP5
2022 Dual-Erp Representation for Object Detection in 360° Images
abstract
Object detection has achieved good performance on perspective images. However, a general object detector does not maintain this performance when applied to a 360° image in a single equirectangular projection (ERP) or multi-projection representation because of the distortion in the high-latitude region or discontinuity at the boundaries. In this paper, we proposed dual-ERP, which is a multi-view ERP representation, as the network input for 360° object detection in training and inference. Dual-ERP combines the advantages of single ERP and multi-projection representations, and it can easily be integrated with existing object detectors. The experimental results showed that compared to other representations, dual-ERP significantly improved the performance of different baseline object detectors.
Satoshi Ikehata, Kiyoharu Aizawa
ICIP3
2022 Recipe-oriented Food Logging for Nutritional Management
abstract
We propose a recipe-oriented food logging method that records food by recipe, unlike the ordinary food logging method that records food by name. We also develop an application RecipeLog for this purpose. RecipeLog can create a "skeleton recipe," which is a standardized recipe representation suitable for estimating the nutritional value of a dish. This is represented by a list of ingredients linked to a Nutrition Facts table and a flow graph consisting of cooking actions such as cutting, mixing, baking, simmering, and frying the ingredients. The recipe log allows recipes to be written with fewer operations by editing only the differences from the already registered base recipe. Experiments have confirmed that the recipe log can effectively identify differences in recipes from household to household. The future work is to construct a multimedia recipe dataset consisting of structured recipes and their images using RecipeLog.
Yoko Yamakata, Akihisa Ishino, Akiko Sunto, Sosuke Amano, Kiyoharu Aizawa
ACM Multimedia5
2022 SLGAN: Style- and Latent-Guided Generative Adversarial Network for Desirable Makeup Transfer and Removal
abstract
There are five features to consider when using generative adversarial networks to apply makeup to photos of the human face. These features include (1) facial components, (2) interactive color adjustments, (3) makeup variations, (4) robustness to poses and expressions, and the (5) use of multiple reference images. To tackle the key features, we propose a novel style- and latent-guided makeup generative adversarial network for makeup transfer and removal. We provide a novel, perceptual makeup loss and a style-invariant decoder that can transfer makeup styles based on histogram matching to avoid the identity-shift problem. In our experiments, we show that our SLGAN is better than or comparable to state-of-the-art methods. Furthermore, we show that our proposal can interpolate facial makeup images to determine the unique features, compare existing methods, and help users find desirable makeup configurations.
Daichi Horita, Kiyoharu Aizawa
MMAsia2
2022 FoodLog Athl: Multimedia Food Recording Platform for Dietary Guidance and Food Monitoring
abstract
This paper presents a new food recording tool, FoodLog Athl, for the healthcare or physical enhancement of its users. Unlike existing food recording tools, we designed the system for dietitians or third parties who monitor the users. The tool not only supports the users by functions such as food image recognition, but also it helps the dietitians watch and communicate with users. Furthermore, it calculates nutritional values from food records - the use of the tool reduces the workload of dietitians and focuses their work on nutrition guidance.
Kei Nakamoto, Kohei Kumazawa, Hiroaki Karasawa, Sosuke Amano, Yoko Yamakata, Kiyoharu Aizawa
MMAsia6
2022 Wearable Camera Based Food Logging System
abstract
Recently, meal management apps have allowed people to record food items and calories from photos automatically. These technologies include extracting food regions from photos of served meals, identifying the name of the food in each region, and calculating nutritional data. However, what you eat is not the only indicator that should be kept in the food record. How fast you eat and the order in which you eat is also significant information for dietary management. Therefore, we aim to construct a system that automatically generates a meal log from first-person videos that users capture of their eating behavior with a wearable camera. To tackle the complex problems that the data this system assumes contains, we constructed an eating behavior record dataset: 9.9 hours of first-person video that assume the natural diets of a user. To investigate the feasibility of our proposed system, we evaluated whether the first step, the detection of the meal area in the video during the meal, could be achieved with sufficient accuracy using this dataset. Using the limited number of frames assumed to be annotated by the user as training data, 30 frames were annotated for user-specific model training and four frames for online adaptation, resulting in detection accuracy of 72% for food regions. Our next goal is to create a multi-user dataset and service the application.
Kenshiro Sato, Yoko Yamakata, Sosuke Amano, Kiyoharu Aizawa
MMAsia4
2022 Fast Nonlinear Image Unblending
abstract
Nonlinear color blending, which is advanced blending indicated by blend modes such as "overlay" and "multiply," is extensively employed by digital creators to produce attractive visual effects. To enjoy such flexible editing modalities on existing bitmap images like photographs, however, creators need a fast nonlinear blending algorithm that decomposes an image into a set of semi-transparent layers. To address this issue, we propose a neural-network-based method for nonlinear decomposition of an input image into linear and nonlinear alpha layers that can be separately modified for editing purposes, based on the specified color palettes and blend modes. Experiments show that our proposed method achieves an inference speed 370 times faster than the state-of-the-art method of nonlinear image unblending, which uses computationally intensive iterative optimization. Furthermore, our reconstruction quality is higher or comparable than other methods, including linear blending models. In addition, we provide examples that apply our method to image editing with nonlinear blend modes.
Daichi Horita, Kiyoharu Aizawa, Ryohei Suzuki, Taizan Yonetsuji, Huachun Zhu
WACV2
2021 Noisy Annotation Refinement for Object Detection
Jiafeng Mao, Qing Yu 0013, Yoko Yamakata, Kiyoharu Aizawa
BMVC4
2021 Intersection Prediction from Single 360° Image via Deep Detection of Possible Direction of Travel
Naoki Sugimoto, Satoshi Ikehata, Kiyoharu Aizawa
BMVC3
2021 What if We Only Use Real Datasets for Scene Text Recognition? Toward Scene Text Recognition With Fewer Labels
abstract
Scene text recognition (STR) task has a common practice: All state-of-the-art STR models are trained on large synthetic data. In contrast to this practice, training STR models only on fewer real labels (STR with fewer labels) is important when we have to train STR models without synthetic data: for handwritten or artistic texts that are difficult to generate synthetically and for languages other than English for which we do not always have synthetic data. However, there has been implicit common knowledge that training STR models on real data is nearly impossible because real data is insufficient. We consider that this common knowledge has obstructed the study of STR with fewer labels. In this work, we would like to reactivate STR with fewer labels by disproving the common knowledge. We consolidate recently accumulated public real data and show that we can train STR models satisfactorily only with real labeled data. Subsequently, we find simple data augmentation to fully exploit real data. Furthermore, we improve the models by collecting unlabeled data and introducing semi- and self-supervised methods. As a result, we obtain a competitive model to state-of-the-art methods. To the best of our knowledge, this is the first study that 1) shows sufficient performance by only using real labels and 2) introduces semi- and self-supervised methods into STR with fewer labels. Our code and data are available: https://github.com/ku21fan/STR-Fewer-Labels.
Jeonghun Baek, Yusuke Matsui 0001, Kiyoharu Aizawa
CVPR3
2021 Improving The Quality Of Illustrations: Transforming Amateur Illustrations To A Professional Standard
abstract
We propose an amateur- to professional-level illustration translator that can modify amateur illustrations slightly to produce professional-level quality images. The proposed translator is a GAN-based image translation module. We focus only on the neighboring region of a contour to improve the quality of the illustration by applying image completion to the neighboring region of the extracted line drawing. We artificially augment amateur-level illustrations from professional-level illustrations to solve the lack of a pair of amateur-level and professional-level datasets. This enables us to automatically prepare a pair of amateur-level and professional-level images, through which we can train a translator network. Through experiments and user study, we show that the proposed method improves the quality of amateur illustrations.
Keita Awane, Koki Tsubota, Hikaru Ikuta, Yusuke Matsui 0001, Kiyoharu Aizawa, Naohiro Yanase
ICIP5
2021 360° Single Image Super Resolution via Distortion-Aware Network and Distorted Perspective Images
abstract
Effective 360° imaging requires a very high resolution because the field of view is extraordinarily high. Single-image super-resolution (SISR) applied to 360° imaging has the potential to solve the resolution/quality problem in this modality. In this paper, we exploit existing perspective SISR networks to address this problem by (1) introducing a distortion map as an additional input with the $360^{\circ}-$ distortion-aware loss function, and (2) augmenting the training 360° images by distorting the perspective images. We also present a new 360° image dataset from YouTube for training. Our extensive experiments show that how each component contributes to the better transfer from the perspective domain to the 360° domain and merging all the ideas leads to the best performance in quantitative and qualitative ways for the 360° SISR task.
Akito Nishiyama, Satoshi Ikehata, Kiyoharu Aizawa
ICIP3
2021 Comprehensive Comparisons Of Uniform Quantizers For Deep Image Compression
abstract
Deep image compression is formulated as a joint rate-distortion optimization problem using an auto-encoder architecture. Latent representations obtained by an encoder are quantized by a quantizer and fed to a decoder and an entropy model to reconstruct the images and estimate probabilities for entropy coding, respectively. Existing methods presented several methods to approximate the quantization for optimization because the gradient of a naive quantizer is zero almost everywhere. Although quantization is a fundamental operation in image compression, there are few comparisons between these quantization methods and the best approximation among them remains unexplored. To address this problem, we comprehensively compare existing approximations of the uniform quantization. Furthermore, focusing on the fact that a decoder and an entropy model have different compatibility with the approximation of quantization, we also evaluate different combinations of approximations for the decoder and the entropy model. Through experiments, we find that the approximations by adding noise are better than rounding and that the best combination of approximations among what we explored outperforms existing approximations.
Koki Tsubota, Kiyoharu Aizawa
ICIP2
2021 RecipeLog: Recipe Authoring App for Accurate Food Recording
abstract
Diet management is usually conducted by recording the name of foods eaten, but in fact, the nutritional value of food in the same name varies greatly from recipe to recipe. To know accurate nutritional values of the foods, recording personal recipes is effective but time-consuming. Therefore, we are developing a mobile application "RecipeLog", that assists users to write their own recipes by modifying prepared ones. In our experiments, we show that with RecipeLog users create personal recipes with 45% less edit distance compared to writing from scratch.
Akihisa Ishino, Yoko Yamakata, Hiroaki Karasawa, Kiyoharu Aizawa
ACM Multimedia4
2021 UrbanMM'21: 1st International Workshop on Multimedia Computing for Urban Data
abstract
Understanding complex processes that give cities their form traditionally relied primarily on the analysis of various open data statistics in relation to e.g. neighbourhood demographics, economy and mobility. However, recent years have seen an unprecedented increase in the availability and use of city-related sensors, participatory data and social multimedia. As the valuable information about urban challenges is usually encoded across multiple modalities, such as visual (e.g. panoramic, satellite and user-contributed images), text (e.g. social media and participatory data) and open data statistics, extracting this information requires effective multimedia analysis tools. This Workshop will showcase the power of multimedia computing in addressing various urban challenges, ranging from event detection and analysis, location recommendation and crowdedness estimation to more efficient handling of citizen reports and modelling and improving city liveability. In addition, it will serve as an impulse for the multimedia community to intensify research on these interesting, challenging and truly multimodal problems.
Stevan Rudinac, Alessandro Bozzon, Tat-Seng Chua, Suzanne Little, Daniel Gatica-Perez, Kiyoharu Aizawa
ACM Multimedia6
2021 Computational attention model for children, adults and the elderly
Onkar Krishna, Kiyoharu Aizawa, Go Irie
Multim. Tools Appl.2
2020 Multi-task Curriculum Framework for Open-Set Semi-supervised Learning
Qing Yu 0013, Daiki Ikami, Go Irie, Kiyoharu Aizawa
ECCV (12)4
2020 Noisy Localization Annotation Refinement For Object Detection
abstract
The production of finely annotated datasets for object detection tasks is labor-intensive, therefore, cloud sourcing is often used to create datasets, which leads to these datasets tending to contain incorrect annotations such as inaccurate localization bounding boxes. In this study, we highlight a problem of object detection with noisy bounding box annotations and show that these noisy annotations are harmful to the performance of deep neural networks. To solve this problem, we further propose a framework to allow the network to modify the noisy datasets by alternating refinement. The experimental results demonstrate that our proposed framework can significantly alleviate the influences of noise on model performance.
Jiafeng Mao, Qing Yu 0013, Kiyoharu Aizawa
ICIP3
2020 Estimation Of Impression Associated With Portraits Using Facial Landmarks And Visual Features
abstract
Portrait manipulation techniques are important for expressing oneself in social networking or portrait/video posting services. Even though several applications (e.g., smartphone apps) have been developed to facilitate portrait manipulation, it is still difficult to obtain ideal results owing to the large number of user-controllable parameters. To tackle this problem, we propose a novel application that can manipulate portraits using impression words. In this study, we select new portraits and impression words, and we devise new experimental settings to enrich the dataset. Moreover, using human evaluations collected by crowdsourcing, we develop and analyze a novel ranking network that uses portraits and facial landmarks as input for impression estimation of the manipulated images.
Mari Miyata, Kiyoharu Aizawa
ICIP2
2020 Unknown Class Label Cleaning For Learning With Open-Set Noisy Labels
abstract
Deep neural networks (DNNs) trained on large-scale annotated datasets have achieved impressive results in the area of image classification. Many large-scale datasets have been collected from websites; however, such data are inevitably corrupted with noise. In this study, we researched the open-set noisy label problem, where some outliers are contained in a dataset and annotated through a noisy label but do not belong to any class of training data. To address this problem, we propose a novel unknown class label cleaning framework for the training of DNNs with open-set noisy labels. In addition to general image classification, we also estimate the probability of an input being from an unknown class by assigning a pseudo unknown label to all of the data and correct these labels through an alternating update of the network parameters and labels. The results of experiments conducted on the noisy CIFAR-10 datasets demonstrate that our approach can robustly train DNNs with a high proportion of noisy labels.
Qing Yu 0013, Kiyoharu Aizawa
ICIP2
2020 Few-Shot Font Generation with Deep Metric Learning
abstract
Designing fonts for languages with a large number of characters, such as Japanese and Chinese, is an extremely labor-intensive and time-consuming task. In this study, we addressed the problem of automatically generating Japanese typographic fonts from only a few font samples, where the synthesized glyphs are expected to have coherent characteristics, such as skeletons, contours, and serifs. Existing methods often fail to generate fine glyph images when the number of style reference glyphs is extremely limited. Herein, we proposed a simple but powerful framework for extracting better style features. This framework introduces deep metric learning to style encoders. We performed experiments using black-and-white and shape-distinctive font datasets and demonstrated the effectiveness of the proposed framework.
Haruka Aoki, Koki Tsubota, Hikaru Ikuta, Kiyoharu Aizawa
ICPR4
2020 The Aleatoric Uncertainty Estimation Using a Separate Formulation with Virtual Residuals
abstract
We propose a new optimization framework for aleatoric uncertainty estimation in regression problems. Existing methods can quantify the error in the target estimation, but they tend to underestimate it. To obtain the predictive uncertainty inherent in an observation, we propose a new separable formulation for the estimation of a signal and of its uncertainty, avoiding the effect of overfitting. By decoupling target estimation and uncertainty estimation, we also control the balance between signal estimation and uncertainty estimation. We conduct three types of experiments: regression with simulation data, age estimation, and depth estimation. We demonstrate that the proposed method outperforms a state-of-the-art technique for signal and uncertainty estimation.
Takumi Kawashima, Qing Yu 0013, Akari Asai, Daiki Ikami, Kiyoharu Aizawa
ICPR5
2020 Translating Adult's Focus of Attention to Elderly's
abstract
Predicting which part of a scene elderly people would pay attention to could be useful in assisting their daily activities, such as driving, walking, and searching. Many computational models for predicting focus of attention (FoA) have been developed. However, most of them focus on mimicking adult FoA and do not work well for predicting elderly's, due to age-related changes in human vision. Is it possible to leverage the prediction results made by an FoA model of general adults to accurately predict elderly's FoA, rather than training a new network from scratch? In this paper, we consider a novel problem of translating adult's FoA to elderly's and propose an approach based on deep image-to-image translation. Our model is trained by minimizing both Kullback-Leibler divergence and adversarial loss to approximate the joint probability distribution of adult and elderly FoA. Experiments on two datasets demonstrate that our model gives remarkable prediction accuracy.
Onkar Krishna, Go Irie, Takahito Kawanishi, Kunio Kashino, Kiyoharu Aizawa
ICPR5
2020 Channel-Level Variable Quantization Network for Deep Image Compression
abstract
Deep image compression systems mainly contain four components: encoder, quantizer, entropy model, and decoder. To optimize these four components, a joint rate-distortion framework was proposed, and many deep neural network-based methods achieved great success in image compression. However, almost all convolutional neural network-based methods treat channel-wise feature maps equally, reducing the flexibility in handling different types of information. In this paper, we propose a channel-level variable quantization network to dynamically allocate more bitrates for significant channels and withdraw bitrates for negligible channels. Specifically, we propose a variable quantization controller. It consists of two key components: the channel importance module, which can dynamically learn the importance of channels during training, and the splitting-merging module, which can allocate different bitrates for different channels. We also formulate the quantizer into a Gaussian mixture model manner. Quantitative and qualitative experiments verify the effectiveness of the proposed model and demonstrate that our method achieves superior performance and can produce much better visual reconstructions.
Zhisheng Zhong, Hiroaki Akutsu, Kiyoharu Aizawa
IJCAI3
2020 Urban Movie Map for Walkers: Route View Synthesis using 360° Videos
abstract
We propose a movie map for walkers based on synthesized street walking views along routes in a particular area. From the perspectives of walkers, we captured a number of omnidirectional videos along streets in the target area (1km2 around Kyoto Station). We captured a separate video for each street. We then performed simultaneous localization and mapping to obtain camera poses from key video frames in all of the videos and adjusted the coordinates based on a map of the area using reference points. To join one video to another smoothly at intersections, we identified frames of video intersection based on camera locations and visual feature matching. Finally, we generated moving route views by connecting the omnidirectional videos based on the alignment of the cameras. To improve smoothness at intersections, we generated rotational views by mixing video intersection frames from two videos. The results demonstrate that our method can precisely identify intersection frames and generate smooth connections between videos at intersections.
Naoki Sugimoto, Toru Okubo, Kiyoharu Aizawa
ICMR3
2020 Building Movie Map - A Tool for Exploring Areas in a City - and its Evaluations
abstract
We propose a new Movie Map, which will enable users to explore a given city area using omnidirectional videos. Only one Movie Map prototype was developed in the 1980s; it was developed with analog video technology. Later, Google Street View (GSV) provided interactive panoramas from positions along streets around the world in Google Maps. Despite the wide use of GSV, it provides sparse images of streets, which often confuses users and lowers user satisfaction. Movie Map's use of videos instead of sparse images dramatically improves the user experience. Thus, we improve the Movie Map using state-of-the-art technology. We propose a new Movie Map system, with an interface for exploring cities. The system consists of four stages; acquisition, analysis, management, and interaction. In the acquisition stage, omnidirectional videos are taken along streets in target areas. Frames of the video are localized on the map, intersections are detected, and videos are segmented. Turning views at intersections are subsequently generated. By connecting the video segments following the specified movement in an area, we can view the streets better. The interface allows for easy exploration of a target area, and it can show virtual billboards of stores in the view. We conducted user studies to compare our system to the GSV in a scenario where users could freely move and explore to find a landmark. The experiment showed that our system had a better user experience than GSV.
Naoki Sugimoto, Yoshihito Ebine, Kiyoharu Aizawa
ACM Multimedia3
2020 Unsupervised Embedding Learning by Noisy Similarity Label Optimization
abstract
We propose a similarity label optimization framework for unsupervised embedding learning. Existing works use similarity labels obtained from instance labels, which are image identifiers, for unsupervised embedding learning assuming that images of different instance labels have different semantic classes. However, because of the significant semantic gap between classes and instances, instance labels are not enough for learning embeddings. To alleviate this problem, we consider the similarity labels as noisy labels and optimize similarity labels and neural network parameters in an alternating fashion. Experimental results demonstrate that the proposed method outperforms a baseline method by 1.2% in terms of accuracy on the CIFAR-10 dataset and in 1.3% in terms of recall at k (k = 1) on the Stanford Online Product dataset.
Koki Tsubota, Kiyoharu Aizawa
VCIP2
2020 Distance Surface for Event-Based Optical Flow
abstract
We propose DistSurf-OF, a novel optical flow method for neuromorphic cameras. Neuromorphic cameras (or event detection cameras) are an emerging sensor modality that makes use of dynamic vision sensors (DVS) to report asynchronously the log-intensity changes (called "events") exceeding a predefined threshold at each pixel. In absence of the intensity value at each pixel location, we introduce a notion of "distance surface"-the distance transform computed from the detected events-as a proxy for object texture. The distance surface is then used as an input to the intensity-based optical flow methods to recover the two dimensional pixel motion. Real sensor experiments verify that the proposed DistSurf-OF accurately estimates the angle and speed of each events.
Mohammed Almatrafi, Raymond Baldwin, Kiyoharu Aizawa, Keigo Hirakawa
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Significance of Softmax-Based Features in Comparison to Distance Metric Learning-Based Features
abstract
End-to-end distance metric learning (DML) has been applied to obtain features useful in many computer vision tasks. However, these DML studies have not provided equitable comparisons between features extracted from DML-based networks and softmax-based networks. In this paper, we present objective comparisons between these two approaches under the same network architecture.
Shota Horiguchi, Daiki Ikami, Kiyoharu Aizawa
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Context-Patch Face Hallucination Based on Thresholding Locality-Constrained Representation and Reproducing Learning
abstract
Face hallucination is a technique that reconstructs high-resolution (HR) faces from low-resolution (LR) faces, by using the prior knowledge learned from HR/LR face pairs. Most state-of-the-arts leverage position-patch prior knowledge of the human face to estimate the optimal representation coefficients for each image patch. However, they focus only the position information and usually ignore the context information of the image patch. In addition, when they are confronted with misalignment or the small sample size (SSS) problem, the hallucination performance is very poor. To this end, this paper incorporates the contextual information of the image patch and proposes a powerful and efficient context-patch-based face hallucination approach, namely, thresholding locality-constrained representation and reproducing learning (TLcR-RL). Under the context-patch-based framework, we advance a thresholding-based representation method to enhance the reconstruction accuracy and reduce the computational complexity. To further improve the performance of the proposed algorithm, we propose a promotion strategy called reproducing learning. By adding the estimated HR face to the training set, which can simulate the case that the HR version of the input LR face is present in the training set, it thus iteratively enhances the final hallucination result. Experiments demonstrate that the proposed TLcR-RL method achieves a substantial increase in the hallucinated results, both subjectively and objectively. In addition, the proposed framework is more robust to face misalignment and the SSS problem, and its hallucinated HR face is still very good when the LR test face is from the real world. The MATLAB source code is available at https://github.com/junjun-jiang/TLcR-RL.
Junjun Jiang, Yi Yu 0001, Suhua Tang, Jiayi Ma 0001, Akiko Aizawa, Kiyoharu Aizawa
IEEE Trans. Cybern.6
2019 Object-Aware Instance Labeling for Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD), where a detector is trained with only image-level annotations, is attracting more and more attention. As a method to obtain a well-performing detector, the detector and the instance labels are updated iteratively. In this study, for more efficient iterative updating, we focus on the instance labeling problem, a problem of which label should be annotated to each region based on the last localization result. Instead of simply labeling the top-scoring region and its highly overlapping regions as positive and others as negative, we propose more effective instance labeling methods as follows. First, to solve the problem that regions covering only some parts of the object tend to be labeled as positive, we find regions covering the whole object focusing on the context classification loss. Second, considering the situation where the other objects contained in the image can be labeled as negative, we impose a spatial restriction on regions labeled as negative. Using these instance labeling methods, we train the detector on the PASCAL VOC 2007 and 2012 and obtain significantly improved results compared with other state-of-the-art approaches.
Satoshi Kosugi, Toshihiko Yamasaki, Kiyoharu Aizawa
ICCV3
2019 Unsupervised Out-of-Distribution Detection by Maximum Classifier Discrepancy
abstract
Since deep learning models have been implemented in many commercial applications, it is important to detect out-of-distribution (OOD) inputs correctly to maintain the performance of the models, ensure the quality of the collected data, and prevent the applications from being used for other-than-intended purposes. In this work, we propose a two-head deep convolutional neural network (CNN) and maximize the discrepancy between the two classifiers to detect OOD inputs. We train a two-head CNN consisting of one common feature extractor and two classifiers which have different decision boundaries but can classify in-distribution (ID) samples correctly. Unlike previous methods, we also utilize unlabeled data for unsupervised training and we use these unlabeled data to maximize the discrepancy between the decision boundaries of two classifiers to push OOD samples outside the manifold of the in-distribution (ID) samples, which enables us to detect OOD samples that are far from the support of the ID samples. Overall, our approach significantly outperforms other state-of-the-art methods on several OOD detection benchmarks and two cases of real-world simulation.
Qing Yu 0013, Kiyoharu Aizawa
ICCV2
2019 Impression Estimation for Deformed Portraits With a Landmark-Based Ranking Network
abstract
In recent years, it has become a trend for people to manipulate their own portraits before posting them on a social networking service. However, it is difficult to get a desired portrait after manipulation without sufficient experience or skill. To obtain a simpler and more effective portrait manipulation technique, we consider an automated portrait manipulation method based on the five impression words: clear, sweet, elegant, modern, and dynamic. In the first part of this work, we analyzed the relationship between deformed portrait (only facial features) and the five impression words over a huge data which we collected by crowd-sourcing. In the next part, we leveraged the knowledge from the results of the analysis to develop a ranking network trained on the collected data for estimating the impression associated with a portrait. Our network demonstrates superior performance in impression estimation when compared to other state-of-the-art methods.
Mari Miyata, Kiyoharu Aizawa
ICIP2
2019 Reinforcing the Robustness of a Deep Neural Network to Adversarial Examples by Using Color Quantization of Training Image Data
abstract
Recent works have shown the vulnerability of deep convolutional neural network (DCNN) to adversarial examples with malicious perturbations. In particular, Black-Box attacks without information of parameter and architectures of the target models are feared as realistic threats. To address this problem, we propose a method using an ensemble of models trained by color-quantized data with loss maximization. Color-quantization can allow the trained models to focus on learning conspicuous spatial features to enhance the robustness of DCNNs to adversarial examples. The proposed method can be adapted to Black-Box attacks with no need of particular attack algorithm for the defense. The results of our experiments validated the effectiveness for preventing decrease in the test accuracy with adversarial perturbation.
Shuntaro Miyazato, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP4
2019 Optical Flow Based Line Drawing Frame Interpolation Using Distance Transform to Support Inbetweenings
abstract
Two-dimensional (2D) hand-drawn animation is still being produced in Japan, and the market is growing each year against a backdrop of worldwide popularity. Inbetweening, the process of creating interpolated frames between key frames, is an important but laborious task in 2D hand-drawn animation. With this process, we aim to avoid the need for a special interface or vectorization. We therefore propose an optical-flow based line drawing frame interpolation method using a distance transform. In general, an optical flow is not applicable to line drawings owing to a lack of intensity in the gradients and colors. Therefore, we use a distance transform to add intensity gradients to line drawings, and it is possible to estimate an optical flow between them. We evaluated our method quantitatively using commercial hand-drawn animation inbetweens. Our method significantly outperforms a baseline method that does not use a distance transform. In addition, we also compared our method to an existing image-based method qualitatively and showed that our method generates better results.
Rei Narita, Keigo Hirakawa, Kiyoharu Aizawa
ICIP3
2019 Identification Of Buildings In Street Images Using Map Information
abstract
In this paper, we propose a method that identifies the buildings in a street image by accurately associating their locations between a map and an image. Building identification is beneficial to the use of images for detailed navigation. Given that a GPS-based camera pose (location and direction) embedded in EXIF of the image is available, we propose a method to improve the camera pose by using a depth map estimated from the image and the depth information taken from the map. By improving the camera pose of the image, we can identify the building in the image by accurate registration of the building on the map and that in the image. In our experiments, we use Google Street View images with their camera poses and demonstrate significant accuracy improvement in building identification.
Masanori Ogawa, Kiyoharu Aizawa
ICIP2
2019 Synthesis of Screentone Patterns of Manga Characters
abstract
Manga or Japanese comics are a popular medium and their images comprise line drawings and screentones. This study investigates the screentone synthesis task that involves translation from line drawings to manga images. Screentones have regular patterns that are difficult to synthesize. To address this problem, we propose a method to translate line drawings into manga images by generating pixel-wise screentone class labels instead of generating manga images directly. To train a screentone label generator, we create paired data of line drawings and pixel-wise screentone class labels that we obtain by applying to manga images a screentone removal and a screentone classifier, respectively. We train the screentone classifier using paired data of simulated manga images and pixel-wise screentone class labels. In tests, we conduct post-processing to reduce noise in the generated pixel-wise screentone labels. Experiments show that our proposed method produces reasonable screentone patterns. In comparison with results obtained using a baseline method of image-to-image translations, our results are comparable or more visually appealing.
Koki Tsubota, Daiki Ikami, Kiyoharu Aizawa
ISM3
2019 Assist Users' Interactions in Font Search with Unexpected but Useful Concepts Generated by Multimodal Learning
abstract
When searching for suitable fonts for a digital graphic, users usually start with an ambiguous thought. For example, they would look for fonts that are suitable for a personal web page or party invitations for children. Their design concept becomes clearer as they interact with external interventions such as exposure to suitable images for use in their web page or the children's preferences regarding the party. Hence, it is important to support users' interactions with unexpected but useful concepts during their search. In this paper, we present a novel framework that helps users to explore a font dataset using the multimodal method that provides unexpected but useful font images or concept words in response to the user's input. We collect a large font dataset and the associated tags and propose the use of unsupervised generative model that jointly learns the correlation between the visual features of a font and the associated tags for the creative process. By examining the results of the model that change with various inputs, we observed that the model produces highly promising results. In the experiment, we verified that the generated concepts by the model are not only new but also relevant to the user input that appears to be useful for inspiring users.
Saemi Choi, Shun Matsumura, Kiyoharu Aizawa
ICMR3
2019 Walker's Movie Map: Route Vies Synthesis Using Omni-directional Videos
abstract
We present a new movie map for walkers that synthesizes street walking views along routes for walkers in an area. We acquired a number of omnidirectional videos from the perspectives of walkers of streets in a certain area (ex. $1km^2$ around Kyoto Station), then perform SLAM to obtain camera poses of key video frames with the coordinates adjusted to the map of the area using reference points. In order to switch one video to another at intersections, we identify the frames of video intersection using camera locations. We refine the intersection frames using visual feature matching. Finally, we synthesize moving route views by switching omnidirectional videos with alignment of the direction of the cameras. The result shows that our method can precisely identify the intersection frames, and generates smooth switching of videos at the intersection.
Naoki Sugimoto, Yuko Iinuma, Kiyoharu Aizawa
ACM Multimedia3
2019 Social Font Search by Multimodal Feature Embedding
abstract
A typical tag/keyword-based search system retrieves documents where, given a query term q, the query term q occurs in the dataset. However, when applying these systems to a real-world font web community setting, practical challenges arise --- font tags are more subjective than other benchmark datasets, which magnify the tag mismatch problem. To address these challenges, we propose a tag dictionary space leveraged by word embedding, which relates undefined words that have a similar meaning. Even if a query is not defined in the tag dictionary, we can represent it as a vector on the tag dictionary space. The proposed system facilitates multi-modal inputs that can use both textual and image queries. By integrating a visual sentiment concept model that classifies affective concepts as adjective--noun pairs for a given image and uses it as a query, users can interact with the search system in a multi-modal way. We used crowd sourcing to collect user ratings for the retrieved fonts and observed that the retrieved font with the proposed methods obtained a higher score compared to other methods.
Saemi Choi, Shun Matsumura, Kiyoharu Aizawa
MMAsia3
2019 Face hallucination through differential evolution parameter map learning with facial structure prior
Junjun Jiang, Jiayi Ma 0001, Suhua Tang, Yi Yu 0001, Kiyoharu Aizawa
Inf. Sci.5
2019 Emotype: Expressing emotions by changing typeface in mobile messenger texting
abstract
Instant messaging is a popular form of text-based communication. However, text-based messaging lacks the ability to communicate nonverbal information such as that conveyed through facial expressions and voice tones, although a multitude of emotions may underlie the text of a conversation between participants. In this paper, we propose an approach that uses typefaces to communicate emotions. We investigated which typefaces are useful for delivering emotions and introduced these typefaces into a mobile chat app. We conducted a survey to demonstrate how changes in the typeface of a message affected the meaning of the message conveyed. Our user study provides an understanding of the actual user experience with the application. The results show that the use of multiple typefaces in a message can affect and intensify the valence received by users and the use of multiple typefaces elicited an active response and brought about a livelier mood during texting.
Saemi Choi, Kiyoharu Aizawa
Multim. Tools Appl.2
2019 Category-Based Deep CCA for Fine-Grained Venue Discovery From Multimodal Data
abstract
In this work, travel destinations and business locations are taken as venues. Discovering a venue by a photograph is very important for visual context-aware applications. Unfortunately, few efforts paid attention to complicated real images such as venue photographs generated by users. Our goal is fine-grained venue discovery from heterogeneous social multimodal data. To this end, we propose a novel deep learning model, category-based deep canonical correlation analysis. Given a photograph as input, this model performs: 1) exact venue search (find the venue where the photograph was taken) and 2) group venue search (find relevant venues that have the same category as the photograph), by the cross-modal correlation between the input photograph and textual description of venues. In this model, data in different modalities are projected to a same space via deep networks. Pairwise correlation (between different modality data from the same venue) for exact venue search and category-based correlation (between different modality data from different venues with the same category) for group venue search are jointly optimized. Because a photograph cannot fully reflect rich text description of a venue, the number of photographs per venue in the training phase is increased to capture more aspects of a venue. We build a new venue-aware multimodal data set by integrating Wikipedia featured articles and Foursquare venue photographs. Experimental results on this data set confirm the feasibility of the proposed method. Moreover, the evaluation over another publicly available data set confirms that the proposed method outperforms state of the arts for cross-modal retrieval between image and text.
Yi Yu 0001, Suhua Tang, Kiyoharu Aizawa, Akiko Aizawa
IEEE Trans. Neural Networks Learn. Syst.3
2018 Local and Global Optimization Techniques in Graph-Based Clustering
abstract
The goal of graph-based clustering is to divide a dataset into disjoint subsets with members similar to each other from an affinity (similarity) matrix between data. The most popular method of solving graph-based clustering is spectral clustering. However, spectral clustering has drawbacks. Spectral clustering can only be applied to macroaverage-based cost functions, which tend to generate undesirable small clusters. This study first introduces a novel cost function based on micro-average. We propose a local optimization method, which is widely applicable to graph-based clustering cost functions. We also propose an initial-guess-free algorithm to avoid its initialization dependency. Moreover, we present two global optimization techniques. The experimental results exhibit significant clustering performances from our proposed methods, including 100% clustering accuracy in the COIL-20 dataset.
Daiki Ikami, Toshihiko Yamasaki, Kiyoharu Aizawa
CVPR3
2018 Fast and Robust Estimation for Unit-Norm Constrained Linear Fitting Problems
abstract
M-estimator using iteratively reweighted least squares (IRLS) is one of the best-known methods for robust estimation. However, IRLS is ineffective for robust unit-norm constrained linear fitting (UCLF) problems, such as fundamental matrix estimation because of a poor initial solution. We overcome this problem by developing a novel objective function and its optimization, named iteratively reweighted eigenvalues minimization (IREM). IREM is guaranteed to decrease the objective function and achieves fast convergence and high robustness. In robust fundamental matrix estimation, IREM performs approximately 5-500 times faster than random sampling consensus (RANSAC) while preserving comparable or superior robustness.
Daiki Ikami, Toshihiko Yamasaki, Kiyoharu Aizawa
CVPR3
2018 Cross-Domain Weakly-Supervised Object Detection Through Progressive Domain Adaptation
abstract
Can we detect common objects in a variety of image domains without instance-level annotations? In this paper, we present a framework for a novel task, cross-domain weakly supervised object detection, which addresses this question. For this paper, we have access to images with instance-level annotations in a source domain (e.g., natural image) and images with image-level annotations in a target domain (e.g., watercolor). In addition, the classes to be detected in the target domain are all or a subset of those in the source domain. Starting from a fully supervised object detector, which is pre-trained on the source domain, we propose a two-step progressive domain adaptation technique by fine-tuning the detector on two types of artificially and automatically generated samples. We test our methods on our newly collected datasets containing three image domains, and achieve an improvement of approximately 5 to 20 percentage points in terms of mean average precision (mAP) compared to the best-performing baselines.
Naoto Inoue, Ryosuke Furuta, Toshihiko Yamasaki, Kiyoharu Aizawa
CVPR4
2018 Joint Optimization Framework for Learning With Noisy Labels
abstract
Deep neural networks (DNNs) trained on large-scale datasets have exhibited significant performance in image classification. Many large-scale datasets are collected from websites, however they tend to contain inaccurate labels that are termed as noisy labels. Training on such noisy labeled datasets causes performance degradation because DNNs easily overfit to noisy labels. To overcome this problem, we propose a joint optimization framework of learning DNN parameters and estimating true labels. Our framework can correct labels during training by alternating update of network parameters and labels. We conduct experiments on the noisy CIFAR-10 datasets and the Clothing1M dataset. The results indicate that our approach significantly outperforms other state-of-the-art methods.
Daiki Tanaka, Daiki Ikami, Toshihiko Yamasaki, Kiyoharu Aizawa
CVPR4
2018 Signboard Saliency Detection in Street Videos
abstract
During the last few decades researchers in computer vision have proposed various saliency models for images with the common goal of classifying the image content by using the measure of importance. However, compared to still images, there is only a limited number of saliency detection algorithms proposed for video signals. However, predicting where a person looks in a video is relevant for applications such as advertisement design, video re-targeting and editing. In this work, we propose a novel method for video saliency detection that aims to detect the relative ranking of saliencies of signboards in street videos. For that reason, we collected eye-gaze data of participants viewing various street videos in free viewing and task viewing scenarios, where the task was to identify a place to have lunch at. Further, we quantitatively analyzed the collected eye-gaze data in order to generate the relative ranking of the signboards in the free viewing and the task viewing scenario. Based on the analysis' results, we propose a video saliency detection algorithm which can more accurately predict the relative saliencies of signboards in street videos. It can be seen that the prediction accuracy of our proposed model outperforms the existing video saliency detection algorithms.
Onkar Krishna, Kiyoharu Aizawa, Saskia Reimerth
ICASSP2
2018 Billboard Saliency Detection in Street Videos for Adults and Elderly
abstract
Detecting salient locations in images and videos that attract human attention is a challenging problem in computer vision. Most of the models developed for this purpose are either based on the bottom-up features of a scene or top-down factors. Observer's age is one of the important top-down factor which influences viewing behavior, some recent studies proposed age-adapted models to predict saliencies for different age groups. However, these studies reported age-impact on eye-gaze distribution for images only and videos still remain unexplored. In this study, we analyze the differences in eye gaze landings between adults and elderly while viewing street videos and develop a novel application of detecting saliency of billboards in the paved area of the street for adult and elderly. Our proposed age-adapted saliency model outperforms the existing video saliency models towards predicting the billboard saliency.
Onkar Krishna, Kiyoharu Aizawa
ICIP2
2018 Food Image Recognition by Personalized Classifier
abstract
Since the development of food diaries could enable people to develop healthy eating habits, food image recognition is in high demand to reduce the effort in food recording. Previous studies have worked on this challenging domain with datasets having fixed numbers of samples and classes. However, in the real-world setting, it is impossible to include all of the foods in the database because the number of classes of foods is large and increases continually. In addition to that, inter-class similarity and intra-class diversity also bring difficulties to the recognition. In this paper, we attempted to solve these problems by using deep convolutional neural network features to build a personalized classifier which incrementally learns the user's data and adapts to the user's eating habit. As a result, we achieved the state-of-the-art accuracy of food image recognition by the personalization of 300 food records per user.
Qing Yu 0013, Masashi Anzawa, Sosuke Amano, Makoto Ogawa, Kiyoharu Aizawa
ICIP5
2018 FontMatcher: Font Image Paring for Harmonious Digital Graphic Design
abstract
One of the important aspects in graphic design is choosing the font of the caption that matches aesthetically the associated image. To obtain a good match, users would exhaustively examine a long font list requiring them a substantial effort. This paper presents FontMatcher, which supports users to design digital graphic works harmoniously pairing fonts with an image. The system provides three features, recommendation, explaination and feedback. If a warm feeling image is given as input, the system recommends warm feeling fonts, and then explains what is the distinguishing features of the recommendation, e.g. a cursive shape. Users can also provide feedback to find fonts which correspond to their intention. Our evaluation results show that the recommended fonts scored better than selected fonts by novices and provides competing results with the ones chosen by experienced graphic designers. The system also provides explanations that help increasing the reliability of the recommended results.
Saemi Choi, Kiyoharu Aizawa, Nicu Sebe
IUI2
2018 Session details: Brand New Ideas
Kiyoharu Aizawa
ACM Multimedia1
2018 Efficiency-enhanced cost-volume filtering featuring coarse-to-fine strategy
abstract
Cost-volume filtering (CVF) is one of the most widely used techniques for solving general multi-labeling problems based on a Markov random field (MRF). However it is inefficient when the label space size (i.e., the number of labels) is large. This paper presents a coarse-to-fine strategy for cost-volume filtering that efficiently and accurately addresses multi-labeling problems with a large label space size. Based on the observation that true labels at the same coordinates in images of different scales are highly correlated, we truncate unimportant labels for cost-volume filtering by leveraging the labeling output of lower scales. Experimental results show that our algorithm achieves much higher efficiency than the original CVF method while maintaining a comparable level of accuracy. Although we performed experiments that deal with only stereo matching and optical flow estimation, the proposed method can be employed in many other applications because of the applicability of CVF to general discrete pixel-labeling problems based on an MRF.
Ryosuke Furuta, Satoshi Ikehata, Toshihiko Yamasaki, Kiyoharu Aizawa
Multim. Tools Appl.4
2018 Photo aesthetic quality estimation using visual complexity features
Litian Sun, Toshihiko Yamasaki, Kiyoharu Aizawa
Multim. Tools Appl.3
2018 Personalized Classifier for Food Image Recognition
abstract
Currently, food image recognition tasks are evaluated against fixed datasets. However, in real-world conditions, there are cases in which the number of samples in each class continues to increase and samples from novel classes appear. In particular, dynamic datasets in which each individual user creates samples and continues the updating process often has content that varies considerably between different users, and the number of samples per person is very limited. A single classifier common to all users cannot handle such dynamic data. Bridging the gap between the laboratory environment and the real world has not yet been accomplished on a large scale. Personalizing a classifier incrementally for each user is a promising way to do this. In this paper, we address the personalization problem, which involves adapting to the user's domain incrementally using a very limited number of samples. We propose a simple yet effective personalization framework, which is a combination of the nearest class mean classifier and the 1-nearest neighbor classifier based on deep features. To conduct realistic experiments, we made use of a new dataset of daily food images collected by a food-logging application. Experimental results show that our proposed method significantly outperforms existing methods.
Shota Horiguchi, Sosuke Amano, Makoto Ogawa, Kiyoharu Aizawa
IEEE Trans. Multim.4
2018 PQTable: Nonexhaustive Fast Search for Product-Quantized Codes Using Hash Tables
abstract
In this paper, we propose a product quantization table (PQTable)-a fast search method for product-quantized codes via hash tables. An identifier of each database vector is associated with the slot of a hash table by using its PQ-code as a key. For querying, an input vector is PQ-encoded and hashed, and the items associated with that code are then retrieved. The proposed PQTable produces the same results as a linear PQ scan, and is 102-105times faster. Although the state-of-the-art performance can be achieved by previous inverted-indexing-based approaches, such methods require manually designed parameter setting and significant training; our PQTable is free of these limitations, and therefore offers a practical and effective solution for real-world problems. Specifically, when the vectors are highly compressed, our PQTable achieves one of the fastest search performances on a single CPU to date with significantly efficient memory usage (0.059-ms per query over 109data points with just 5.5-GB memory consumption). Finally, we show that our proposed PQTable can naturally handle the codes of an optimized product quantization (OPQTable).
Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa
IEEE Trans. Multim.3
2017 Age-adapted saliency model with depth bias
abstract
Visual attention studies in computer vision research have focused on the development of computational attention systems that can detect salient regions in images for adults. Consequently, age differences in scene viewing behavior has rarely been considered. This study quantitatively analyzed the age-related differences in gaze landings during scene viewing for three age groups: children, adults, and elderly. An interesting observation from our analysis is that whereas child observers focus more on the scene foreground, i.e., locations that are near, elderly observers tend to explore the scene background, i.e., locations farther in the scene. Considering this result a framework is proposed in this paper to quantitatively measure the depth bias tendency across age groups. Further, the age impact on exploratory behavior, central bias tendency, and agreement between explored regions within and across the age groups are quantified via analysis. Experimental results show that children exhibit the lowest exploratory behavior level but the highest central bias tendency among the age groups. Further, agreement scores reveal that adults had least agreement with each other in explored regions. The data analysis results were consequently leveraged to develop a more accurate age-adapted saliency model that outperforms existing saliency models that do not consider age.
Onkar Krishna, Kiyoharu Aizawa
SAP2
2017 Spatio-Temporal Vector of Locally Max Pooled Features for Action Recognition in Videos
abstract
We introduce Spatio-Temporal Vector of Locally Max Pooled Features (ST-VLMPF), a super vector-based encoding method specifically designed for local deep features encoding. The proposed method addresses an important problem of video understanding: how to build a video representation that incorporates the CNN features over the entire video. Feature assignment is carried out at two levels, by using the similarity and spatio-temporal information. For each assignment we build a specific encoding, focused on the nature of deep features, with the goal to capture the highest feature responses from the highest neuron activation of the network. Our ST-VLMPF clearly provides a more reliable video representation than some of the most widely used and powerful encoding approaches (Improved Fisher Vectors and Vector of Locally Aggregated Descriptors), while maintaining a low computational complexity. We conduct experiments on three action recognition datasets: HMDB51, UCF50 and UCF101. Our pipeline obtains state-of-the-art results.
I. C. Duta, Bogdan Ionescu, Kiyoharu Aizawa, Nicu Sebe
CVPR3
2017 Residual Expansion Algorithm: Fast and Effective Optimization for Nonconvex Least Squares Problems
abstract
We propose the residual expansion (RE) algorithm: a global (or near-global) optimization method for nonconvex least squares problems. Unlike most existing nonconvex optimization techniques, the RE algorithm is not based on either stochastic or multi-point searches, therefore, it can achieve fast global optimization. Moreover, the RE algorithm is easy to implement and successful in high-dimensional optimization. The RE algorithm exhibits excellent empirical performance in terms of k-means clustering, point-set registration, optimized product quantization, and blind image deblurring.
Daiki Ikami, Toshihiko Yamasaki, Kiyoharu Aizawa
CVPR3
2017 Object detection refinement using Markov random field based pruning and learning based rescoring
abstract
Contextual information such as the co-occurrence of objects and the location of objects has played an important role in object detection. We present candidate pruning and object rescoring methods that leverage contextual information and that can improve the state-of-the-art CNN-based object detection methods such as Fast R-CNN and Faster R-CNN. In our pruning method, we formulate candidate reduction as a Markov random field optimization problem. In our rescoring method, we employ a machine learning technique to reconsider the detection scores of candidate windows. We experimentally demonstrate improvements in R-CNN-based object detection methods using two datasets. Moreover, we apply our model to the structured retrieval task to show the potential applications of our model.
Naoto Inoue, Ryosuke Furuta, Toshihiko Yamasaki, Kiyoharu Aizawa
ICASSP4
2017 Hyperlapse generation of omnidirectional videos by adaptive sampling based on 3D camera positions
abstract
Capturing a city with an omnidirectional camera can produce a continuous motion street view and provides the view of the moving person. However, watching such a video is not necessarily pleasant because the videos are often excessively lengthy. In this paper, we propose a method for shortening an omnidirectional video. Subsampling a video captured by a hand-held camera displays a significant amount of destabilization resulting from shaking. This prompted us to propose an adaptive subsampling scheme that selects optimal frames by minimizing our cost function based on 3D camera positions. This optimal selection suppresses only the translational camera instabilities. The rotational instabilities are initially ignored and later compensated for. This approach allowed us to successfully generate a subsampled and stable omnidirectional video. In addition, we propose an alternative measure of translational instability to evaluate our frame selection.
Masanori Ogawa, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP3
2017 How competitive are you: Analysis of people's attractiveness in an online dating system
abstract
An increasing number of people are using dating websites to search for their life partners. This leads to the curiosity of how attractive a specific person is to the opposite gender on an average level. We propose a novel algorithm to evaluate people's objective attractiveness based on their interactions with other users on the dating websites and implement machine learning algorithms to predict their objective attractiveness ratings from their profiles. We validate our method on a large dataset gained from a Japanese dating website and yield convincing results. Our prediction based on users' profiles, which includes image and text contents, is over 80% correlated with the real values of the calculated objective attractiveness for the female and over 50% correlated with the real values of the calculated objective attractiveness for the male.
Xiaoxue Zang, Toshihiko Yamasaki, Kiyoharu Aizawa, Tetsuhiro Nakamoto, Eitaro Kuwabara, Shinichi Egami, Yusuke Fuchida
ICME3
2017 FolkPopularityRank: Tag Recommendation for Enhancing Social Popularity using Text Tags in Content Sharing Services
abstract
In this study, we address two emerging yet challenging problems in social media: (1) scoring the text tags in terms of the influence to the numbers of views, comments, and favorite ratings of images and videos on content sharing services, and (2) recommending additional tags to increase such popularity-related numbers. For these purposes, we present the FolkPopularityRank algorithm, which can score text tags based on their ability to influence the popularity-related numbers. The FolkPopularityRank algorithm is inspired by the PageRank and FolkRank algorithms but the scores of the tags are calculated not only by the co-occurrence of the tags but also by considering the popularity-related numbers of the content. To the best of our knowledge, this is the first attempt to recommending tags that can enhance popularity attributes of social media. We conducted extensive experiments with about 1,000 images. We uploaded the photos with the recommended tags along with the original tags to Flickr as a real test, and obtained very promising results.
Toshihiko Yamasaki, Jiani Hu, Shumpei Sano, Kiyoharu Aizawa
IJCAI4
2017 Become Popular in SNS: Tag Recommendation using FolkPopularityRank to Enhance Social Popularity
abstract
In this demo, we address two emerging yet challenging problems in social media: (1) scoring the text tags in terms of the influence to the numbers of views, comments, and favorite ratings of images and videos on content sharing services, and (2) recommending additional tags to increase such popularity-related numbers. For these purposes, we present a demo using our FolkPopularityRank (FP-Rank) algorithm, which can score and recommend text tags based on their ability to influence the popularity-related numbers. Our experiments using 1,000 photos showed that we can achieve 1.6 times more views than the original tag sets in Flickr just by adding tags recommended by FP-Rank.
Toshihiko Yamasaki, Yiwei Zhang 0014, Jiani Hu, Shumpei Sano, Kiyoharu Aizawa
IJCAI5
2017 VenueNet: Fine-Grained Venue Discovery by Deep Correlation Learning
abstract
Venue photos, as a new type of multimedia contents, are exploding on the Internet because users like to take photos and share with their friends in which venue they spent time and what impressed them there. Discovering a venue by a social photo is very useful for supplementing venue retrieval and recommendation. However, little research focused on fine-grained venue discovery by leveraging multimodal venue dataset. In this paper, we present the first multimodal dataset specially built for venue discovery, which includes venue photos, descriptions, and categories. Using this dataset, we propose a novel framework for fine-grained venue discovery through correlating venue photos and descriptions, aiming to learn a VenueNet representing a knowledge base and association for venues and their properties in different modalities. In the training phase, visual and textual features of the same venues, by two sub-networks, are respectively mapped to a same semantic space, in which canonical correlation analysis (CCA) is applied to these features to train the two sub-networks. In the query phase, given a photo, its correlation with textual features in the dataset is analyzed to find the most similar venue. Experimental results verify the practicability of the Deep CCA model for fine-grained venue discovery from large-scale multimodal dataset.
Yi Yu 0001, Suhua Tang, Kiyoharu Aizawa, Akiko Aizawa
ISM3
2017 Simple, Efficient and Effective Encodings of Local Deep Features for Video Action Recognition
abstract
For an action recognition system a decisive component is represented by the feature encoding part which builds the final representation that serves as input to a classifier. One of the shortcomings of the existing encoding approaches is the fact that they are built around hand-crafted features and they are not also highly competitive on encoding the current deep features, necessary in many practical scenarios. In this work we propose two solutions specifically designed for encoding local deep features, taking advantage of the nature of deep networks, focusing on capturing the highest feature response of the convolutional maps. The proposed approaches for deep feature encoding provide a solution to encapsulate the features extracted with a convolutional neural network over the entire video. In terms of accuracy our encodings outperform by a large margin the current most widely used and powerful encoding approaches, while being extremely efficient for the computational cost. Evaluated in the context of action recognition tasks, our pipeline obtains state-of-the-art results on three challenging datasets: HMDB51, UCF50 and UCF101.
I. C. Duta, Bogdan Ionescu, Kiyoharu Aizawa, Nicu Sebe
ICMR3
2017 MatPlanner: Plan Your Days in Conferences by Resolving Conflicting Events
abstract
Nowadays multi-track conferences impose great difficulty to attendees in making a day plan of attending relevant sessions/papers, because a large number of sessions are scheduled on a single day and many of them are at the same time. To address this problem, we introduce MatPlanner, a mobile application, which provides an alternate way of scheduling a day in a conference by understanding the attendee's interest and preferences. MatPlanner learns event-topic interest matrix from few sampled events selected by each user and then recommend relevant events and display the final output. More importantly, MatPlanner's interface enables users to resolve conflicting events by showing them at the same panel which makes it easy to compare. Interview results suggested a high level of satisfaction among participants in scheduling a day.
Saemi Choi, Onkar Krishna, Wen-Yu Lee, Kiyoharu Aizawa
ACM Multimedia4
2017 PQk-means: Billion-scale Clustering for Product-quantized Codes
abstract
Data clustering is a fundamental operation in data analysis. For handling large-scale data, the standard k-means clustering method is not only slow, but also memory-inefficient. We propose an efficient clustering method for billion-scale feature vectors, called PQk-means. By first compressing input vectors into short product-quantized (PQ) codes, PQk-means achieves fast and memory-efficient clustering, even for high-dimensional vectors. Similar to k-means, PQk-means repeats the assignment and update steps, both of which can be performed in the PQ-code domain. Experimental results show that even short-length (32 bit) PQ-codes can produce competitive results compared with k-means. This result is of practical importance for clustering in memory-restricted environments. Using the proposed PQk-means scheme, the clustering of one billion 128D SIFT features with K = 105 is achieved within 14 hours, using just 32 GB of memory consumption on a single computer.
Yusuke Matsui 0001, Keisuke Ogaki, Toshihiko Yamasaki, Kiyoharu Aizawa
ACM Multimedia4
2017 A Tag Recommendation System for Popularity Boosting
abstract
In order to support users in the tagging process and recommendation, we had proposed two tag ranking algorithms, Document Frequency-Weights from regression and Folk Popularity Rank, which can extract tags greatly influencing popularity. We have developed a tag recommendation system using the algorithm we proposed. The recommended tags are not only for appropriate annotations but also for popularity boosting.
Yiwei Zhang 0014, Jiani Hu, Shumpei Sano, Toshihiko Yamasaki, Kiyoharu Aizawa
ACM Multimedia5
2017 Panel: Cross-media Intelligence
abstract
In this panel, we attempt to review and discuss the recent emerging theoretical and technological advances and trends of cross-media. Integrating data-driven machine learning with human knowledge can effectively lead to explainable, robust, and general models. Thus, the effective employment of the interaction between cross-media data during inference and reasoning becomes a challenge to populate the cross-media knowledge graph. Some other fundamental and controversial issues such as leveraging the auxiliary information to boost the cross-media understanding, the existence of unified framework to bridge the gap between multi-modality will also be discussed in this panel.
Yueting Zhuang, Ramesh Jain 0001, Wen Gao 0001, Kiyoharu Aizawa
ACM Multimedia5
2017 Spatio-Temporal VLAD Encoding for Human Action Recognition in Videos
I. C. Duta, Bogdan Ionescu, Kiyoharu Aizawa, Nicu Sebe
MMM (1)3
2017 Efficient human action recognition using histograms of motion gradients and VLAD with descriptor shape information
I. C. Duta, Jasper R. R. Uijlings, Bogdan Ionescu, Kiyoharu Aizawa, Alex Hauptmann 0001, Nicu Sebe
Multim. Tools Appl.4
2017 Sketch-based manga retrieval using manga109 dataset
abstract
Manga (Japanese comics) are popular worldwide. However, current e-manga archives offer very limited search support, i.e., keyword-based search by title or author. To make the manga search experience more intuitive, efficient, and enjoyable, we propose a manga-specific image retrieval system. The proposed system consists of efficient margin labeling, edge orientation histogram feature description with screen tone removal, and approximate nearest-neighbor search using product quantization. For querying, the system provides a sketch-based interface. Based on the interface, two interactive reranking schemes are presented: relevance feedback and query retouch. For evaluation, we built a novel dataset of manga images, Manga109, which consists of 109 comic books of 21,142 pages drawn by professional manga artists. To the best of our knowledge, Manga109 is currently the biggest dataset of manga images available for research. Experimental results showed that the proposed framework is efficient and scalable (70 ms from 21,142 pages using a single computer with 204 MB RAM).
Yusuke Matsui 0001, Kota Ito, Yuji Aramaki, Azuma Fujimoto, Toru Ogawa, Toshihiko Yamasaki, Kiyoharu Aizawa
Multim. Tools Appl.7
2017 Depth Estimation Using an Infrared Dot Projector and an Infrared Color Stereo Camera
abstract
This paper proposes a method of estimating depth from two kinds of stereo images: color stereo images and infrared stereo images. An infrared dot pattern is projected on a scene by a projector so that infrared cameras can capture the scene textured by the dots and the depth can be estimated even where the surface is not textured. The cost volumes are calculated for the infrared and color stereo images for each frame and are extended in the time direction to define a spatiotemporal cost volume (st-cost volume). We also extend the cost volume filter in the time direction by modifying the cross-based local multipoint filter (CLMF) and applying it to the st-cost volumes in order to restrain flicker on the time-varying depth maps. To get a reliable cost volume, the infrared and color st-cost volumes are integrated into a single cost volume by selecting the cost of either the infrared or the color st-cost volumes according to the size of the adaptive kernel used for the CLMF. Then, a graphcut is executed on the cost volume in order to estimate the disparity robustly even when the baselines of the stereo cameras are set wide enough to ensure spatially high resolution in the depth direction and the shapes of blocks are deformed by the affine transformation. A 2D graphcut is executed on each scan line to reduce the processing time and memory consumption. We experimented with the proposed method using infrared color stereo data sets of scenes in the real world and evaluated its effectiveness by comparing it with other recent stereo matching methods and depth cameras.
Kensuke Hisatomi, Masanori Kano, Kensuke Ikeya, Miwa Katayama, Tomoyuki Mishina, Yuichi Iwadate, Kiyoharu Aizawa
IEEE Trans. Circuits Syst. Video Technol.7
2017 DrawFromDrawings: 2D Drawing Assistance via Stroke Interpolation with a Sketch Database
abstract
We present DrawFromDrawings, an interactive drawing system that provides users with visual feedback for assistance in 2D drawing using a database of sketch images. Following the traditional imitation and emulation training from art education, DrawFromDrawings enables users to retrieve and refer to a sketch image stored in a database and provides them with various novel strokes as suggestive or deformation feedback. Given regions of interest (ROIs) in the user and reference sketches, DrawFromDrawings detects as-long-as-possible (ALAP) stroke segments and the correspondences between user and reference sketches that are the key to computing seamless interpolations. The stroke-level interpolations are parametrized with the user strokes, the reference strokes, and new strokes created by warping the reference strokes based on the user and reference ROI shapes, and the user study indicated that the interpolation could produce various reasonable strokes varying in shapes and complexity. DrawFromDrawings allows users to either replace their strokes with interpolated strokes (deformation feedback) or overlays interpolated strokes onto their strokes (suggestive feedback). The other user studies on the feedback modes indicated that the suggestive feedback enabled drawers to develop and render their ideas using their own stroke style, whereas the deformation feedback enabled them to finish the sketch composition quickly.
Yusuke Matsui 0001, Takaaki Shiratori, Kiyoharu Aizawa
IEEE Trans. Vis. Comput. Graph.3
2016 Uncalibrated Photometric Stereo by Stepwise Optimization Using Principal Components of Isotropic BRDFs
abstract
The uncalibrated photometric stereo problem for non-Lambertian surfaces is challenging because of the large number of unknowns and its ill-posed nature stemming from unknown reflectance functions. We propose a model that represents various isotropic reflectance functions by using the principal components of items in a dataset, and formulate the uncalibrated photometric stereo as a regression problem. We then solve it by stepwise optimization utilizing principal components in order of their eigenvalues. We have also developed two techniques that lead to convergence and highly accurate reconstruction, namely (1) a coarse-to-fine approach with normal grouping, and (2) a randomized multipoint search. Our experimental results with synthetic data showed that our method significantly outperformed previous methods. We also evaluated the algorithm in terms of real image data, where it gave good reconstruction results.
Keisuke Midorikawa, Toshihiko Yamasaki, Kiyoharu Aizawa
CVPR3
2016 Text detection in manga by combining connected-component-based and region-based classifications
abstract
As manga (Japanese comics) have become common content in many countries, it is necessary to search manga by text query or translate them automatically. For these applications, we must first extract texts from manga. In this paper, we develop a method to detect text regions in manga. Taking motivation from methods used in scene text detection, we propose an approach using classifiers for both connected components and regions. We have also developed a text region dataset of manga, which enables learning and detailed evaluations of methods used to detect text regions. Experiments using the dataset showed that our text detection method performs more effectively than existing methods.
Yuji Aramaki, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP4
2016 The log-normal distribution of the size of objects in daily meal images and its application to the efficient reduction of object proposals
abstract
In general, object-detection methods apply classifiers to pre-calculated object proposals. It is therefore important to minimize the number of proposals to achieve computational efficiency. In this paper, we show that the region size for food objects in recorded images of daily food follows a lognormal distribution, which is different from the distribution for widely used datasets collected by querying the names of dishes. We explain this characteristic using Gibrat's law, and construct a model for the region-size distribution of objects in images. We applied the model to the filtering of object proposals generated by selective search and edge boxes. We obtained a significant reduction of 40.6% in the number of hypotheses compared with a conventional selective search, despite a decrease of only 0.007 in the Mean Average Best Overlap.
Shota Horiguchi, Kiyoharu Aizawa, Makoto Ogawa
ICIP2
2016 Boosting VLAD with double assignment using deep features for action recognition in videos
abstract
The encoding method is an important factor for an action recognition pipeline. One of the key points for the encoding method is the assignment step. A very widely used super-vector encoding method is the vector of locally aggregated descriptors (VLAD), with very competitive results in many tasks. However, it considers only hard assignment and the criteria for the assignment is performed only from the features side, by looking for which visual word the features are voting. In this work we propose to encode deep features for videos using a double assignment VLAD (DA-VLAD). In addition to the traditional assignment for VLAD we perform a second assignment by taking into account the perspective from the codebook side: which are the nearest features to a visual word and not only which is the nearest centroid for the features as the standard assignment. Another important factor for the performance of an action recognition system is the feature extraction step. Recently, deep features obtained state-of-the-art results in many tasks, being also adopted for action recognition with competitive results over hand-crafted features. This work includes a pipeline to extract local deep features for videos using any available network as a black box and we show competitive results including the case when the network was trained for another task or another dataset. Our DA-VLAD encoding method outperforms the traditional VLAD and we obtain state-of-the-art results on UCF50 dataset and competitive results on UCF101 dataset.
I. C. Duta, Tuan Anh Nguyen 0004, Kiyoharu Aizawa, Bogdan Ionescu, Nicu Sebe
ICPR3
2016 Interactive region segmentation for manga
abstract
Manga (Japanese comics) are popular all over the world, and are created digitally. In this paper, we propose an interactive segmentation method tailored for manga. The proposed method enables annotators to select areas in manga efficiently. Our experimental results showed that the proposed framework works better than Adobe Photoshop CC, which is the most widely used commercial image editing software.
Kota Ito, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa
ICPR4
2016 Sketch simplification by classifying strokes
abstract
In this paper, we propose a novel approach to creating clean line drawing from a scribbled sketch automatically. The main problem is determining which strokes of a scribbled sketch should be merged. We use a machine learning approach to solve this problem. Our method can automatically generate training data by comparing scribbled sketches with manually drawn line drawings without using annotations. In order to verify the generated training data, we merged strokes and created clean line drawings in accordance with the generated training data. In addition, we trained a support vector machine to estimate the pairs of strokes to be merged. Further, we verified that our method can create line drawings using this estimator.
Toru Ogawa, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa
ICPR4
2016 Multimedia for personal health and health care
abstract
Ever since the emergence of digitization, the term multimedia has been used to represent a combination of different kinds of media types, such images, audio, and videos. As new sensing technologies emerge and are now becoming omnipresent in daily lives, the definition, role and significance of multimedia is changing. Multimedia now represents the means for communicating, cooperating, and also for monitoring numerous aspects of daily life, at various levels of granularity and application, ranging from personal to societal. With this shift, we have since moved from comprehending single media and its state toward comprehending media in terms of its use context. Multimedia is thus no longer confined to documentation and preservation, entertainment or personal media collections; rather, it has become an integral part of the tools and systems that are providing solutions to today's societal challenges-including challenges related to health care and personal health , aging, education, societal participation, sustainable energy, and intelligent transportation. Multimedia has thus evolved into a core enabler for future interactive and cooperative applications at the heart of society. In this workshop we explore the relevance and contribution of multimedia to health care and personal media.
Susanne Boll, Kiyoharu Aizawa, Alexia Briassouli, Cathal Gurrin, Laleh Jalali, Jochen Meyer 0001
ACM Multimedia2
2016 City-view image location identification by multiple geo-social media and graph-based image cluster refinement
Wen-Yu Lee, Yin-Hsi Kuo, Winston H. Hsu, Kiyoharu Aizawa
J. Vis. Commun. Image Represent.4
2016 Very fast generation of content-preserved photo collage under canvas size constraint
Kiyoharu Aizawa
Multim. Tools Appl.2
2015 Depth Estimation Based on an Infrared Projector and an Infrared Color Stereo Camera by Using Cross-Based Dynamic Programming with Cost Volume Filter
abstract
This paper presents a method to estimate a depth map using an infrared projector and a pair of infrared color cameras that can capture infrared and color images simultaneously. The infrared projector projects a dot pattern so that the cameras capture infrared images of a scene textured by the dots with which depths to surfaces in the scene can be estimated regardless of whether they have visible textures. Cost volumes are calculated for each of the infrared and color stereo images and are processed with a cost volume filter that smoothes each cost map by a cross-based local multipoint filter. The filtered infrared and color cost volumes are then integrated into a single cost volume by selecting either the infrared or color cost for each pixel according to the size of the adaptive kernel used for the cross-based local multipoint filter. This improves the accuracy where the adaptive kernel is small. We propose a method that can find optimal local curved surfaces of adaptive kernels from the cost volume by using three dynamic programmings. We experimented this depth estimation method on real-world datasets that the infrared color stereo cameras captured. We also used it for color stereo matching and showed that it works with normal color stereo cameras as well.
Kensuke Hisatomi, Masanori Kano, Kensuke Ikeya, Miwa Katayama, Tomoyuki Mishina, Kiyoharu Aizawa
3DV6
2015 A Discourse Search Engine Based on Rhetorical Structure Theory
Pascal Kuyten, Danushka Bollegala, Bernd Hollerit, Helmut Prendinger, Kiyoharu Aizawa
ECIR5
2015 PQTable: Fast Exact Asymmetric Distance Neighbor Search for Product Quantization Using Hash Tables
abstract
We propose the product quantization table (PQTable), a product quantization-based hash table that is fast and requires neither parameter tuning nor training steps. The PQTable produces exactly the same results as a linear PQ search, and is 102 to 105 times faster when tested on the SIFT1B data. In addition, although state-of-the-art performance can be achieved by previous inverted-indexing-based approaches, such methods do require manually designed parameter setting and much training, whereas our method is free from them. Therefore, PQTable offers a practical and useful solution for real-world problems.
Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa
ICCV3
2015 A layered method for determining manga text bubble reading order
abstract
Comic books of all cultures are an active research area as digitizing content for mobile and web is becoming more common. Past research on comics has largely concentrated on text extraction, panel segmentation and document analysis, while the utilisation of the extracted data has had less attention. In this paper we present a method to automatically determine the reading order of Japanese manga text bubbles using only text bubble position and image data. Our method classifies and orders page and text position information on three layers, which are hierarchically sorted to obtain the final ordering. The method is evaluated on a data set of 1769 manga pages with 14726 manually annotated text positions and correct ordering. Evaluation shows the method has over 95% transition accuracies and vastly outperforms a naive implementation.
Samu Kovanen, Kiyoharu Aizawa
ICIP2
2015 Searching for nearest neighbors with a dense space partitioning
abstract
Product quantization based approximate nearest neighbor search with the use of inverted index structures have recently received increasing attention. In this paper, we propose a new inverted index structure for searching nearest neighbors in very large datasets of high dimensional data. For data indexing, our proposed method creates a dense space partitioning using multiple centroids based assigning, which generates shorter candidate lists and improves the search speed. Our experiments with a dataset of one billion SIFT features show that while achieving higher accuracy, our method demonstrates better performances on search speed compared to IV-FADC, the conventional product quantization based inverted index structure.
Tuan Anh Nguyen 0004, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP4
2015 Repositioning the salient region of videos by using active illumination
abstract
The use of networked light bulbs is proposed in order to modify the focus of attention, defined based on salience maps, of a video being displayed on a screen. The proposed system was evaluated by using four networked lamps positioned behind a television display. The lamps were controlled by the system according to the visual features of the video. By using a camera, the salience of the room illuminated by the smart bulbs was recorded and compared against the salience of the room when the lamps were not used. The results showed that changes in color and intensity of the environmental illumination were able to modify the shape of the salience maps of the displayed videos. Two modes of operation are proposed: salience cancellation and salience enhancement.
René Marcelino Abritta Teixeira, Kiyoharu Aizawa
ICIP2
2015 Fast Face Model Reconstruction and Synthesis Using an RGB-D Camera and Its Subjective Evaluation
abstract
It is difficult to show a frontal face in video chatting because there is a gap between a display and a camera. We propose a method for real-time face reorientation by creating a 2.5-D face model from a single RGB-D camera and synthesizing the rotated face model with the original face image. Our method uses two kinds face models complementarily: a point cloud based model and a generic face model fitted to the user. We conducted subjective evaluation and confirmed the validity of our proposed system.
Toshihiko Yamasaki, Ibuki Nakamura, Kiyoharu Aizawa
ISM3
2015 Selective K-means Tree Search
abstract
In object recognition and image retrieval, an inverted indexing method is used to solve the approximate nearest neighbor search problem. In these tasks, inverted indexing provides a nonexhaustive solution to large-scale search. However, a problem of previous inverted indexing methods is that a large-scale inverted index is required to achieve a high search recall rate. In this study, we address the problem of reducing the time required to build an inverted index without degrading the search accuracy and speed. Thus, we propose a selective k-means tree search method that combines the power of both hierarchical k-means tree and selective nonexhaustive search. Experiments based on approximate nearest neighbor search using a large dataset comprising one billion SIFT features showed that the hierarchical inverted file based on the selective k-means tree method could be built six times faster, while obtaining almost the same recall and search speed as the state-of-the-art inverted indexing methods.
Tuan Anh Nguyen 0004, Yusuke Matsui 0001, Toshihiko Yamasaki, Kiyoharu Aizawa
ACM Multimedia4
2014 Photometric Stereo Using Constrained Bivariate Regression for General Isotropic Surfaces
abstract
This paper presents a photometric stereo method that is purely pixelwise and handles general isotropic surfaces in a stable manner. Following the recently proposed sum-of-lobes representation of the isotropic reflectance function, we constructed a constrained bivariate regression problem where the regression function is approximated by smooth, bivariate Bernstein polynomials. The unknown normal vector was separated from the unknown reflectance function by considering the inverse representation of the image formation process, and then we could accurately compute the unknown surface normals by solving a simple and efficient quadratic programming problem. Extensive evaluations that showed the state-of-the-art performance using both synthetic and real-world images were performed.
Satoshi Ikehata, Kiyoharu Aizawa
CVPR2
2014 Simultaneous acquisition of multiple images with higher dynamic range
abstract
Computational photography has redefined the possibilities of conventional photography. New effects can be obtained by mixing modified hardware with software. In this paper we present the basis of a theoretical framework for a computational camera. By using coded apertures, the framework is able to multiplex images of a scene and simultaneously acquire them. As an application, we show a system to capture images with different exposure values, a feature that is often useful in high dynamic range imaging. This framework is not limited to the acquisition of images with varying exposure values. It is also adaptable to different types of filters, e.g.: color filters. The advantages of this system are the low cost of implementation and ease of adaption to different conditions. However, the need and difficulty of the decoding process can add extra image artifacts.
René Marcelino Abritta Teixeira, Kiyoharu Aizawa
ICASSP2
2014 MangaWall: Generating manga pages for real-time applications
abstract
Recent advances in non-photorealistic rendering provide the best convenience for automatic image to manga conversion. However, we are facing a dilemma that the fine-grained manga conversion algorithms always involve computation intensive processes (e.g. image over-segmentation, brute-force feature matching, and energy optimization), which make it inapplicable to real-time applications. On the other hand, commercial manga Apps available for smart phones usually rely on simple edge detection and halftoning or hatching. Although the processing can be finished in several seconds, the conversion quality is not satisfactory yet. In this paper, we propose MangaWall to automatically convert and organize photos into manga pages. Our goal is to establish a lightweight manga rendition framework as well as generate high-quality images. To achieve this, the manga structure is enhanced by flow-based DoG operation and image vectorization. A multi-layer tone mapping and contrast-aware halftoning method is then proposed to render the manga-like screentone patterns. Besides, for multi-image input, a full binary tree-based layout representation is employed to efficiently organize manga images onto the same page canvas. Our MangaWall can be applied to real-time applications. It takes less than 0.7 second to convert a 1024 × 768 image into manga on Laptop PC.
Kiyoharu Aizawa
ICASSP2
2014 Coarse-to-fine strategy for efficient cost-volume filtering
abstract
Cost-volume filtering is one of the most widely known techniques to solve general multi-label problems, however it is problematically inefficient when the label space size is extremely large. This paper presents a coarse-to-fine strategy of the cost-volume filtering that handles efficiently and accurately multi-label problems with a large label space size. Based upon the observation that true labels at the same image coordinate of different scales are highly correlated, we truncate unimportant labels for the cost-volume filtering by leveraging the labeling output of lower scales. Experimental results show that our algorithm achieves much higher efficiency than the original cost-volume filtering while enjoying the comparable accuracy to it.
Ryosuke Furuta, Satoshi Ikehata, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP4
2014 Multi-stage object classification featuring confidence analysis of classifier and inclined local Naive Bayes nearest neighbor
abstract
We propose a two-stage classification framework for image recognition which conjunctively uses parametric and non-parametric approaches. In the first stage, input images are classified using a bag-of-features (BoF) based method with a multi-class classifier. The results are categorized into two groups by our confidence analysis: highly confident and less confident. The images with less confidence are re-classified in the second stage using our inclined local naive bayes nearest neighbor (IL-NBNN). In the original local NBNN, the similarity between the input image and its k-NN classes are calculated aiming at higher discriminability and computational efficiency. Our IL-NBNN virtually calculates the similarity between all the classes efficiently by incorporating the confidence order obtained in the first stage. As a result, efficient and accurate image classification has been achieved with very small extra cost. The experiments using the three image datasets show the validity of our proposed algorithm.
Takaki Maeda, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP3
2014 Sketch2Manga: Sketch-based manga retrieval
abstract
We propose a sketch-based method for manga image retrieval, in which users draw sketches via a Web browser that enables the automatic retrieval of similar images from a database of manga titles. The characteristics of manga images are different from those of naturalistic images. Despite the widespread attention given to content-based image retrieval systems, the question of how to retrieve manga images effectively has been little studied. We propose a fine multi-scale edge orientation histogram (FMEOH) whereby a number of differently sized squares on a page can be indexed efficiently. Our experimental results show that FMEOH can achieve greater accuracy than a state-of-the-art sketch-based retrieval method [1].
Yusuke Matsui 0001, Kiyoharu Aizawa, Yushi Jing
ICIP2
2014 Degree of loop assessment in microvideo
abstract
This paper presents a degree-of-loop assessment method for microvideo clips. Loop video is one of the popular features in microvideo, but there are so many non-loop video tagged with “loop” on microvideo services. This is because upload-ers or spammers also know that loop video is popular and they want to draw attention from viewers. In this paper, we statistically analyze the scene dynamics of the video by using color, optical flow, saliency maps, and evaluate the degree-of-loop. We have collected more than 1,000 video clips from Vine and subjectively evaluated their degree-of-loop. Experimental results show that our proposed algorithm can classify loop/non-loop video with 85.7% accuracy and categorize them into five degree-of-loop categories with 61.5% accuracy.
Shumpei Sano, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP3
2014 Food Detection and Recognition Using Convolutional Neural Network
abstract
In this paper, we apply a convolutional neural network (CNN) to the tasks of detecting and recognizing food images. Because of the wide diversity of types of food, image recognition of food items is generally very difficult. However, deep learning has been shown recently to be a very powerful image recognition technique, and CNN is a state-of-the-art approach to deep learning. We applied CNN to the tasks of food detection and recognition through parameter optimization. We constructed a dataset of the most frequent food items in a publicly available food-logging system, and used it to evaluate recognition performance. CNN showed significantly higher accuracy than did traditional support-vector-machine-based methods with handcrafted features. In addition, we found that the convolution kernels show that color dominates the feature extraction process. For food image detection, CNN also showed significantly higher accuracy than a conventional method did.
Hokuto Kagaya, Kiyoharu Aizawa, Makoto Ogawa
ACM Multimedia2
2014 Emerging Topics on Personalized and Localized Multimedia Information Systems
abstract
We are experiencing an era with a rapid increase of data relevant to different aspects of users' daily life. On the one hand, such data contains personal information of each individual user. On the other hand, it also reflects user behaviors related to the society as data of more users is aggregated. These data could not only be very beneficial for studying various lifestyle patterns, but also be used to generate more descriptive and explanatory analysis across the landscape of diverse multimedia data. Using personal mobile devices and web services to systematically explore interesting aspects of people world has attracted much attention recently. This is a full-day tutorial that addresses emerging topics on personalized and localized multimedia technologies and applications and emphasizes knowledge sensing and discovery in multimedia landscape. This tutorial aims to deliver anoverall introduction to multimedia landscapes with multimedia processing, contextual data acquisition, people activity logs, data analytics, geographic-aware multimedia sharing and delivery, and serves as an important lecture on fundamental and advanced research areas of personalized and localized multimedia information systems.
Yi Yu 0001, Kiyoharu Aizawa, Toshihiko Yamasaki, Roger Zimmermann
ACM Multimedia2
2014 SVM is not always confident: Telling whether the output from multiclass SVM is true or false by analysing its confidence values
abstract
This paper presents an algorithm to distinguish whether the output label that is yielded from multiclass support vector machine (SVM) is true or false without knowing the answer. Such judgment is done only by the confidence analysis based on the pre-training/testing using the training data. Such true/false judgment is useful for refining the output labels. We experimentally demonstrate that the decision value difference between the top candidate and the second candidate is a good measure. In addition, a proper threshold can be determined by the pre-training/testing using only the training data. Experimental results using three standard image datasets demonstrate that our proposed algorithm can improve Matthews correlation coefficient (MCC) much better than simply thresholding the decision value for the top candidate.
Toshihiko Yamasaki, Takaki Maeda, Kiyoharu Aizawa
MMSP3
2014 Photometric Stereo Using Sparse Bayesian Regression for General Diffuse Surfaces
abstract
Most conventional algorithms for non-Lambertian photometric stereo can be partitioned into two categories. The first category is built upon stable outlier rejection techniques while assuming a dense Lambertian structure for the inliers, and thus performance degrades when general diffuse regions are present. The second utilizes complex reflectance representations and non-linear optimization over pixels to handle non-Lambertian surfaces, but does not explicitly account for shadows or other forms of corrupting outliers. In this paper, we present a purely pixel-wise photometric stereo method that stably and efficiently handles various non-Lambertian effects by assuming that appearances can be decomposed into a sparse, non-diffuse component (e.g., shadows, specularities, etc.) and a diffuse component represented by a monotonic function of the surface normal and lighting dot-product. This function is constructed using a piecewise linear approximation to the inverse diffuse model, leading to closed-form estimates of the surface normals and model parameters in the absence of non-diffuse corruptions. The latter are modeled as latent variables embedded within a hierarchical Bayesian model such that we may accurately compute the unknown surface normals while simultaneously separating diffuse from non-diffuse components. Extensive evaluations are performed that show state-of-the-art performance using both synthetic and real-world images.
Satoshi Ikehata, David P. Wipf, Yasuyuki Matsushita, Kiyoharu Aizawa
IEEE Trans. Pattern Anal. Mach. Intell.4
2013 Depth map inpainting and super-resolution based on internal statistics of geometry and appearance
abstract
Depth maps captured by multiple sensors often suffer from poor resolution and missing pixels caused by low reflectivity and occlusions in the scene. To address these problems, we propose a combined framework of patch-based inpainting and super-resolution. Unlike previous works, which relied solely on depth information, we explicitly take advantage of the internal statistics of a depth map and a registered highresolution texture image that capture the same scene. We account these statistics to locate non-local patches for hole filling and constrain the sparse coding-based super-resolution problem. Extensive evaluations are performed and show the state-of-the-art performance when using real-world datasets.
Satoshi Ikehata, Ji-Ho Cho, Kiyoharu Aizawa
ICIP3
2013 Workshop summary for the 5th international workshop on multimedia for cooking and eating activities (CEA'13)
abstract
This summary introduces the aim of the CEA'13 workshop and the list of papers presented in the workshop.
Kiyoharu Aizawa, Yoko Yamakata, Takuya Funatomi
ACM Multimedia1
2013 Kanji snap: an OCR-based smartphone application for learning Japanese kanji characters
abstract
As optical character recognition techniques improve, new opportunities to improve existing systems open up. There are for example OCR reading aids for the visually impaired and applications for translating foreign text by capturing photos with smartphones. But one field that hasn't made use of the advancing technology is the learning environment. This demo concentrates on incorporating optical character recognition into a smartphone application for learning Japanese kanji characters. With the application users can take photos of kanji in their everyday environment and look up detailed information and translations easily. Users can practice those kanji with vocabulary lists and quizzes and track their study progress.
Kiia Korpi, Kiyoharu Aizawa
ACM Multimedia2
2013 Action recognition using invariant features under unexampled viewing conditions
abstract
A great challenge in real-world applications of action recognition is the lack of sufficient label information because of variance in the recording viewpoint and differences between individuals. A system that can adapt itself according to these variances is required for practical use. We present a generic method for extracting view-invariant features from skeleton joints. These view-invariant features are further refined using a stacked, compact autoencoder. To model the challenge of real-world applications, two unexampled test settings (NewView and NewPerson) are used to evaluate the proposed method. Experimental results with these test settings demonstrate the effectiveness of our method.
Litian Sun, Kiyoharu Aizawa
ACM Multimedia2
2013 Navilog: A Museum Guide and Location Logging System Based on Image Recognition
Soichiro Kawamura, Tomoko Ohtani, Kiyoharu Aizawa
MMM (2)3
2013 A novel approach for combined rotational and translational motion estimation using Frame Projection Warping
abstract
This paper introduces a novel video stabilization technique for combined rotational and translational motion using integral frame projections. In the proposed Frame Projection Warping (FPW) method, the normalized intensity projection curves of two consecutive frames are warped using dynamic time warping, to get the relative shift between them. Rotational and vertical motion estimation involves partitioning of frame in to two halves and their corresponding estimated motions are then utilized for respective rotational angle and vertical shift estimation. This technique uses the human perception for analyzing the rotation in terms of vertical displacement of two halves of frame. The proposed technique is tested over various hand recorded videos. The results show better performance of FPW over various existing intensity based techniques. This technique also gives better accuracy in case of frame blurring, which is a serious cause of wrong motion estimation. The performance of the proposed technique is measured in terms of interframe transformation fidelity and processing time.
Deepika Shukla, Rajib Kumar Jha, Kiyoharu Aizawa
VCIP3
2013 Cooperative estimation of human motion and surfaces using multiview videos
Weilan Luo, Toshihiko Yamasaki, Kiyoharu Aizawa
Comput. Vis. Image Underst.3
2013 Food Balance Estimation by Using Personal Dietary Tendencies in a Multimedia Food Log
abstract
We have investigated the “FoodLog” multimedia food-recording tool, whereby users upload photographs of their meals and a food diary is constructed using image-processing functions such as food-image detection and food-balance estimation. In this paper, following a brief introduction to FoodLog, we propose a Bayesian framework that makes use of personal dietary tendencies to improve both food-image detection and food-balance estimation. The Bayesian framework facilitates incremental learning. It incorporates three personal dietary tendencies that influence food analysis: likelihood, prior distribution, and mealtime category. In the evaluation of the proposed method using images uploaded to FoodLog, both food-image detection and food-balance estimation are improved. In particular, in the food-balance estimation, the mean absolute error is significantly reduced from 0.69 servings to 0.28 servings on average for two persons using more than 200 personal images, and 0.59 servings to 0.48 servings on average for four persons using 100 personal images. Among the works analyzing food images, this is the first to make use of statistical personal bias to improve the performance of the analysis.
Kiyoharu Aizawa, Yuto Maruyama, Chamin Morikawa
IEEE Trans. Multim.1
2012 Robust photometric stereo using sparse regression
abstract
This paper presents a robust photometric stereo method that effectively compensates for various non-Lambertian corruptions such as specularities, shadows, and image noise. We construct a constrained sparse regression problem that enforces both Lambertian, rank-3 structure and sparse, additive corruptions. A solution method is derived using a hierarchical Bayesian approximation to accurately estimate the surface normals while simultaneously separating the non-Lambertian corruptions. Extensive evaluations are performed that show state-of-the-art performance using both synthetic and real-world images.
Satoshi Ikehata, David P. Wipf, Yasuyuki Matsushita, Kiyoharu Aizawa
CVPR4
2012 Intra texture prediction based on repetitive pixel replenishment
abstract
A new intra prediction scheme based on intra repetitive pixel replenishment is proposed. This scheme improves intra coding efficiency by generating adaptive texture according to an intra displacement vector that reduces prediction error for cyclic patterns and differential motion vector coding. Simulation results showed that the proposed technique improved coding efficiency by up to 2.5% BD-Rate (ΔBitrate in [%]) compared with version 3.0 of common software for High Efficiency Video Coding; encoding time was increased by 20%. The total improvement in intra coding efficiency with respect to existing H.264/AVC intra prediction was 5.0% on average and with a maximum of 10.3%. Because the proposed scheme did not need to estimate the decoder-side motion, there was no significant increase in decoding time.
Kenichi Iwata, Ryoji Hashimoto, Seiji Mochizuki, Kiyoharu Aizawa
ICIP4
2012 Internal noise-induced contrast enhancement of dark images
abstract
A contrast enhancement technique using scaling of internal noise of a dark image in discrete cosine transform (DCT) domain has been proposed in this paper. The mechanism of enhancement is attributed to noise-induced transition of DCT coefficients from a poor state to an enhanced state. This transition is effected by the internal noise present due to lack of sufficient illumination and can be modeled by a general bistable system exhibiting dynamic stochastic resonance. The proposed technique adopts a local adaptive processing and significantly enhances the image contrast and color information while ascertaining good perceptual quality. When compared with the existing enhancement techniques such as adaptive histogram equalization, gamma correction, single-scale retinex, multi-scale retinex, modified high-pass filtering, multi-contrast enhancement, multi-contrast enhancement with dynamic range compression, color enhancement by scaling, edge-preserving multi-scale decomposition and automatic controls of popular imaging tool, the proposed technique gives remarkable performance in terms of relative contrast enhancement, colorfulness and visual quality of enhanced image.
Rajib Kumar Jha, Rajlaxmi Chouhan, Prabir Kumar Biswas, Kiyoharu Aizawa
ICIP4
2012 Noise attenuation performance of mura apertures in photographic cameras
abstract
Coded apertures provide enhanced noise characteristics to lens-less imaging systems. Similar characteristics have been suggested for traditional cameras. Four experiments were conducted to investigate the amount of noise reduction provided by coded apertures when used in photographic cameras. They focus on the effects of coding mask sizes, positioning, mosaicing and rotation. The signal to noise ratio of photos taken with and without coding masks was evaluated. The mask pattern utilized was the modified uniformly redundant array (MURA). The experiments showed evidence of noise reduction and best results were obtained for masks with higher element count, mosaiced and limited to the sensor size.
René Marcelino Abritta Teixeira, Kiyoharu Aizawa
ICIP2
2012 Real-time tracking of humans and visualization of their future footsteps in public indoor environments - An intelligent interactive system for public entertainment
abstract
In this work, an interactive entertainment system which employs multiple-human tracking from a single camera is presented. The proposed system robustly tracks people in an indoor environment and displays their predicted future footsteps in front of them in real-time. The system is composed of a video camera, a computer and a projector. There are three main modules: tracking, analysis and visualization. The tracking module extracts people as moving blobs by using an adaptive background subtraction algorithm. Then, the location and orientation of their next footsteps are predicted. The future footsteps are visualized by a high-paced continuous display of foot images in the predicted location to simulate the natural stepping of a person. To evaluate the performance, the proposed system was exhibited during a public art exhibition in an airport. People showed surprise, excitement, curiosity. They tried to control the display of the footsteps by making various movements.
Ovgu Ozturk, Tomoaki Matsunami, Yasuhiro Suzuki, Toshihiko Yamasaki, Kiyoharu Aizawa
Multim. Tools Appl.5
2012 Determination of emotional content of video clips by low-level audiovisual features - A dimensional and categorial experimental approach
abstract
Affective analysis of video content has greatly increased the possibilities of the way we perceive and deal with media. Different kinds of strategies have been tried, but results are still opened to improvements. Most of the problems come from the lack of standardized test set and real affective models. In order to cope with these issues, in this paper we describe the results of our work on the determination of affective models for evaluation of video clips using audiovisual low-level features. The affective models were developed following two classes of psychological theories of affect: categorial and dimensional. The affective models were created from real data, acquired through a series of user experiments. They reflect the affective state of a viewer after watching a certain scene from a movie. We evaluate the detection of Pleasure, Arousal and Dominance coefficients as well as the detection rate of six affective categories. For this end, two Bayesian network topologies are used, a Hidden Markov Model and an Autoregressive Hidden Markov Model. The measurements were done using audio-only models, video-only models and fused models. Fusion is done using two different methods, a Decision Level Fusion and Feature Level Fusion. All tests were conducted using localized affective models, both categorial and dimensional. Results are presented in terms of detection rate and accuracy for affective families, affective dimensions and probabilistic networks. Arousal was the best detected dimension, followed by dominance and pleasure.
René Marcelino Abritta Teixeira, Toshihiko Yamasaki, Kiyoharu Aizawa
Multim. Tools Appl.3
2011 Robust watermark extraction using SVD-based dynamic stochastic resonance
abstract
In this paper, a novel dynamic stochastic resonance (DSR)-based non-blind watermark extraction technique has been proposed for robust extraction of a grayscale watermark. The watermark embedding has been carried out using singular value decomposition (SVD). Dynamic stochastic resonance has been strategically used to improve the robustness of the extraction algorithm by utilizing the noise added during attacks itself. Resilience of this technique to attacks has been tested in the presence of various noise, geometrical, enhancement, compression and filtering attacks. Using the DSR-based proposed extraction algorithm, a very robust extraction of watermark can be done without trading-off with visual quality of the watermarked image. Performance of the proposed technique has also been compared with the plain SVD-based and hybrid DCT-SVD based technique and is found to give better performance.
Rajlaxmi Chouhan, Rajib Kumar Jha, Apoorv Chaturvedi, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP5
2011 Marker-less human pose estimation and surface reconstruction using a segmented model
abstract
We propose a human motion tracking method for fast motion clips using synchronized multiple cameras. Our method is capable of extracting 3D articulated postures with 42 degrees of freedom through a sequence of visual hulls. We seek for the globally optimal solutions of the likelihood with the lo cal memorization about the "fitness" of each body segment. Our method avoids the local minimum problem efficiently by mean combination and articulated combination of parti cles selected based on the weights of the different body seg ments. We deform the template surface model using the mo tion tracking data by linear blend skinning (LBS). Then we recover the details of the surface by fitting the deformed sur face to 2D silhouettes.
Weilan Luo, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP3
2011 Clustering meal images in a web-based dietary management system
abstract
We investigate the possibility of summarizing meal image collections in an image-based dietary assessment system. A segmentation algorithm detects food regions in each meal image. Pairwise matching of images based on color and texture information of these regions forms a similarity matrix of the image collection. We cluster the nodes of the graph constructed using this matrix, to identify natural groupings of meal images according to their content. Representative images from these clusters form summaries of the large image collection. We conduct a user study to evaluate the effective ness of the proposed algorithms in summarizing meal image collections, and report the results.
Gamhewage Chaminda de Silva, Kiyoharu Aizawa
ICME2
2011 Image-based Calorie Content Estimation for Dietary Assessment
abstract
In this paper, we present an image-analysis based approach to calorie content estimation for dietary assessment. We make use of daily food images captured and stored by multiple users in a public Web service called FoodLog. The images are taken without any control or markers. We build a dictionary dataset of 6512 images contained in FoodLog the calorie content of which have been estimated by experts in nutrition. An image is compared to the ground truth data from the point of views of multiple image features such as color histograms, color correlograms and SURF fetures, and the ground truth images are ranked by similarities. Finally, calorie content of the input food image is computed by linear estimation using the top n ranked calories in multiple features. The distribution of the estimation shows that 79% of the estimations are correct within ±40% error and 35% correct within ±20% error.
Tatsuya Miyazaki, Gamhewage Chaminda de Silva, Kiyoharu Aizawa
ISM3
2011 High efficient distributed video coding with parallelized design for cloud computing
abstract
In this work, by combining coding tools developed in recent literatures on transform domain WZ video coding with some newly developed modules on both encoding and decoding sides, an efficient and practical WZ video coding architecture, dubbed as DIStributed video coding with PArallelized design for Cloud computing (DISPAC), is proposed to better the corresponding rate-distortion (RD) performance. Another unique feature of DISPAC, lies in the parallelizability of the modules used by its WZ decoder which increased the decoding speed largely. Experimental results conducted on an emulated Could computing environment reveal that DISPAC codec can gain up to 3.6 dB in the RD measures and 60.97 times faster in the decoding speed as compared with the-state-of-art WZ video codec, respectively
Han-Ping Cheng, Yun-Chung Shen, Ja-Ling Wu, Kiyoharu Aizawa
ACM Multimedia4
2010 Automatic preview video generation for mesh sequences
abstract
We present a novel method that automatically generates a preview video of a mesh sequence. To make the preview appealing to users, the important features of the mesh model should be captured in the preview video, while preserving the constraint that the transitions of the camera are as smooth as possible. Our approach models the important features by defining a surface saliency and by measuring the appearance of the mesh sequence. The task of generating the preview video is then formulated as a shortest-path problem and we find an optimal camera path by using Dijkstra's algorithm.
Seung-Ryong Han, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP3
2010 Patch-based compression for Time-Varying Meshes
abstract
This paper proposes intra-frame and inter-frame coding algorithms for 3D mesh sequences generated by multiple cameras, which we call Time-Varying Meshes (TVMs). For this purpose, mesh segmentation into patches with patch alignment using principal component analysis (PCA) is proposed. The patches are used as minimum units to eliminate spatial and temporal correlation of the TVMs. For intraframe coding of geometry data, spectral compression using the Kirchhoff matrix was employed and that of color texture was conducted using vector quantization (VQ). The inter-frame coding of geometry data and color texture was achieved by the combination of patch-based matching and simple scalar quantization. The intra- and inter-frame coding are compared to the previous works and demonstrated promising results.
Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP2
2010 Image processing based approach to food balance analysis for personal food logging
abstract
Food images have been receiving increased attention in recent dietary control methods. We present the current status of our web-based system that can be used as a dietary management support system by ordinary Internet users. The system analyzes image archives of the user to identify images of meals. Further image analysis determines the nutritional composition of these meals and stores the data to form a Foodlog. The user can view the data in different formats, and also edit the data to correct any mistakes that occurred during image analysis. This paper presents detailed analysis of the performance of the current system and proposes an improvement of analysis by pre-classification and personalization. As a result, the accuracy of food balance estimation is significantly improved.
Keigo Kitamura, Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa
ICME4
2010 Detecting Dominant Motion Flows in Unstructured/Structured Crowd Scenes
abstract
Detecting dominant motion flows in crowd scenes is one of the major problems in video surveillance. This is particularly difficult in unstructured crowd scenes, where the participants move randomly in various directions. This paper presents a novel method which utilizes SIFT features' flow vectors to calculate the dominant motion flows in both unstructured and structured crowd scenes. SIFT features can represent the characteristic parts of objects, allowing robust tracking under non-rigid motion. First, flow vectors of SIFT features are calculated at certain intervals to form a motion flow map of the video. Next, this map is divided into equally sized square regions and in each region dominant motion flows are estimated by clustering the flow vectors. Then, local dominant motion flows are combined to obtain the global dominant motion flows. Experimental results demonstrate the successful application of the proposed method to challenging real-world scenes.
Ovgu Ozturk, Toshihiko Yamasaki, Kiyoharu Aizawa
ICPR3
2010 Automatic trailer generation
abstract
This paper presents a content-based movie trailer generation method, named Vid2Trailer (V2T). Since trailers are intended to advertise movies, they must show specific symbols such as the title logo and the main theme music. Moreover, it is expected to attract viewers by its visual and audio content. V2T satisfies these two requirements when creating a trailer from the original movie content. First, the title logo and the main theme music are extracted. Second, impressive speech and video segments are extracted by using an affective content analysis technique. Third, all of the extracted components are concatenated into the form of a trailer; to realize this, we propose a method that estimates the affective impact of shot sequences, and introduce an algorithm that arranges a set of shots so as to maximize the affective impact of the sequence. Experiments show that our V2T is more appropriate to trailer generation than conventional techniques.
Go Irie, Takashi Satou, Akira Kojima, Toshihiko Yamasaki, Kiyoharu Aizawa
ACM Multimedia5
2010 3D pose estimation in high dimensional search spaces with local memorization
abstract
In this paper, a stochastic approach for extracting the articulated 3D human postures by synchronized multiple cameras in the high-dimensional configuration spaces is presented. Annealed Particle Filtering (APF) [1] seeks for the globally optimal solution of the likelihood. We improve and extend the APF with local memorization to estimate the suited kinematic postures for a volume sequence directly instead of projecting a rough simplified body model to 2D images. Our method guides the particles to the global optimization on the basis of local constraints. A segmentation algorithm is performed on the volumetric models and the process is repeated. We assign the articulated models 42 degrees of freedom. The matching error is about 6% on average while tracking the posture between two neighboring frames.
Weilan Luo, Toshihiko Yamasaki, Kiyoharu Aizawa
PCS3
2010 Bit allocation of vertices and colors for patch-based coding in time-varying meshes
abstract
This paper discusses bit-rate assignments for vertices, color, reference frames, and target frames in the patch-based compression method for time-varying meshes (TVMs). TVMs are nonisomorphic 3D mesh sequences of the real-world objects generated from multiview images. Experimental results demonstrate that the bit rate for vertices greatly affects the visual quality of the rendered 3D model, whereas the bit rate for color does not contribute to quality improvement. Therefore, as many bits as possible should be assigned to vertices, with 8–10 bits per vertex (bpv) per frame being sufficient for color. For interframe coding, the visual quality is improved in proportion to the bit rate of both vertices and color. However, it is demonstrated that the use of fewer bits (5∼6 bpv) is sufficient to achieve a visual quality that matches the intraframe visual quality.
Toshihiko Yamasaki, Kiyoharu Aizawa
PCS2
2010 Approaches to 3D video compression
abstract
Three-dimensional (3-D) video provides an immersing experience for users. In recent years, many attempts have been made to capture the complex surface shape and highly detailed texture of real-world moving objects, which results in a huge amount of data. In this paper, we discuss compression issues for 3-D video. We introduce 3-D video, which is classified into two categories. Then we survey compression methods that have been investigated for each category. We present our compression methods for temporally varying mesh sequences. In addition, we show comparison results for our algorithm with respect to previous work.
Seung-Ryong Han, Toshihiko Yamasaki, Kiyoharu Aizawa
VCIP3
2010 Large-scale image and video search: Challenges, technologies, and trends
Meng Wang 0001, Nicu Sebe, Tao Mei 0001, Jia Li 0001, Kiyoharu Aizawa
J. Vis. Commun. Image Represent.5
2010 Affective Audio-Visual Words and Latent Topic Driving Model for Realizing Movie Affective Scene Classification
abstract
This paper presents a novel method for movie affective scene classification that outputs the emotion (in the form of labels) that the scene is likely to arouse in viewers. Since the affective preferences of users play an important role in movie selection, affective scene classification has the potential to develop more attractive user-centric movie search and browsing applications. Two main issues in designing movie affective scene classification are considered. One is “how to extract features that are strongly related to the viewer's emotions”, and the other is “how to map the extracted features to the emotion categories”. For the former, we propose a method to extract emotion-category-specific audio-visual features named affective audio-visual words (AAVWs). For the latter issue, we propose a classification model named latent topic driving model (LTDM). Assuming that viewers' emotions are dynamically changed by the movie scene sequences, LTDM models emotions as Markovian dynamic systems driven by the sequential stimuli of the movie content. Experiments on 206 movie scenes extracted from 24 movie titles and the corresponding labels of eight emotion categories given by 16 subjects show that our method outperforms conventional approaches in terms of the subject agreement rate.
Go Irie, Takashi Satou, Akira Kojima, Toshihiko Yamasaki, Kiyoharu Aizawa
IEEE Trans. Multim.5
2009 An object-based non-blind watermarking that is robust to non-linear geometrical distortion attacks
abstract
This paper presents an object-based non-blind watermarking technique that is robust to non-linear geometrical distortion attacks. This has been one of the most challenging problems for copyright protection of digital content because it is difficult to estimate the distortion parameters for the embedded blocks. In our proposed scheme, the locations used to embed the watermark information are memorized and detected robustly by using relative coordinates with respect to the scale-invariant feature transform (SIFT) feature points. Experimental results with 64-bit watermark embedding demonstrated that the watermark detection performance was improved from 57% to 82% on average, even after nonlinear geometrical attacks such as waving, imploding, swirling, and seam carving.
Toshihiko Yamasaki, Yasumasa Nakai, Kiyoharu Aizawa
ICIP3
2009 Affective video segment retrieval for consumer generated videos based on correlation between emotions and emotional audio events
abstract
A novel affective video segment retrieval method based on the correlation between emotion and emotional audio events (EAEs) is presented. The proposed method focuses on retrieving three types of affective video segments, joy, sadness and excitement, by utilizing correlations between emotions and EAEs. The correlation between these emotions and EAEs is investigated by a subjective evaluation. The proposed method detects EAEs and rates each EAE in terms of emotion levels. The EAEs are detected by using the generalized state-space model (GSSM) and low-level audio features. Experiments conducted on consumer generated videos (CGVs) show that the proposed EAE detection outperforms conventional HMM and GMM based methods in terms of accuracy, the agreement rate of the retrieved affective video segments reaches 73.3%.
Go Irie, Kota Hidaka, Takashi Satou, Toshihiko Yamasaki, Kiyoharu Aizawa
ICME5
2009 A degree-of-edit ranking for consumer generated video retrieval
abstract
We introduce degree-of-edit (DoE) ranking to focus on ldquohow much a CGV is editedrdquo as a ranking measure for consumer generated video (CGV) retrieval; a method to estimate DoE ranking is proposed. In the proposed method, the DoE score of a CGV is estimated by using low-level features such as the number of shot boundaries and time ratio of music. We evaluate the rank correlation between DoE ranking determined by subjects and by our method. To demonstrate its performance in a practical scenario, a user test is performed on over 22,000 CGVs in the context of CGV search. The obtained results show that our method significantly improves conventional CGV ranking results in terms of availabilities of interesting and high-quality CGVs.
Go Irie, Kota Hidaka, Takashi Satou, Toshihiko Yamasaki, Kiyoharu Aizawa
ICME5
2009 Retrieval of Time-Varying Mesh and motion capture data using 2D video queries based on silhouette shape descriptors
abstract
This paper presents a retrieval system for time-varying mesh (TVM) and motion capture data using 2D video queries. Previous approaches have used other TVM and motion capture data as queries and the cost for query generation was a significant issue. Instead, the proposed system uses 2D video queries, which can be easily captured by a single camera, enabling end users to retrieve 3D motion sequences such as TVM and motion capture data easily and interactively. We introduce the P-type Fourier descriptor, which is a feature of 2D contour images. TVM and computer graphics sequences rendered from motion capture data are silhouetted rendering from multiple viewpoints. Feature vectors for TVM and motion capture data are generated by applying the P-type Fourier descriptor to these silhouetted images. Experimental results using four TVM sequences and motion capture data demonstrated an average retrieval accuracy of 88% in terms of nearest neighbors.
Daisuke Kasai, Toshihiko Yamasaki, Kiyoharu Aizawa
ICME3
2009 A Euclidean-geodesic shape distribution for retrieval of time-varying mesh sequences
abstract
This paper proposes a Euclidean-geodesic shape distribution for the more accurate retrieval of time-varying meshes, which are 3D mesh sequences of real-world objects generated by multiple cameras. The Euclidean-geodesic shape distribution derives from a combination of the modified shape distribution algorithm, which analyzes the global shape features of 3D models, and the geodesic shape distribution algorithm, which is used to investigate topological changes. The optimal weighting for the two algorithms is investigated experimentally. Experimental results show that the performance for similar motion retrieval is better than that of conventional algorithms, being improved by 2% on average and by 5% in the best case.
Toshihiko Yamasaki, Kiyoharu Aizawa
ICME2
2009 Latent topic driving model for movie affective scene classification
abstract
This paper proposes a latent topic driving model (LTDM) as a novel approach to movie affective scene classification. LTDM is a discriminative model of emotions driven by movie affective contents. Unlike existing methods, our approach is based on movie topic extraction via the latent Dirichlet allocation (LDA) and emotion dynamics modeling with reference to Plutchik's emotion theory. The classification procedure starts by segmenting movie scenes into movie shots, each of which is represented by a histogram of quantized affect-related audio-visual features. LDA is applied to detect topics of each movie shot. Emotions for the current movie shot are estimated based on both the topics of the shot and emotion transition weights determined by Plutchik's emotion theory. We conduct experiments using 206 movie scenes extracted from 24 movie titles (total 6 hours 20 min. 12 sec.) and the labels of eight emotion categories given by 16 subjects are collected. The results show that LTDM outperforms conventional modeling approaches in terms of the subject agreement rate.
Go Irie, Kota Hidaka, Takashi Satou, Akira Kojima, Toshihiko Yamasaki, Kiyoharu Aizawa
ACM Multimedia6
2009 Retrieving multimedia travel stories using location data and spatial queries
abstract
We propose a system for retrieving multimedia related to a person's travel, using location data captured with a GPS receiver, mobile phone or a camera. The user makes simple sketches on a map displayed on a computer screen, to submit spatial, temporal or spatio-temporal queries regarding his travel. The system segments the location data and images, analyzes sketches made by a user, identifies the query, and retrieves relevant results. These results, combined with online maps and virtual tours rendered using street view panoramas, form multimedia travel stories. We present the system's current status and conclude with future directions.
Gamhewage Chaminda de Silva, Kiyoharu Aizawa
ACM Multimedia2
2009 Sketch-on-Map: Spatial Queries for Retrieving Human Locomotion Patterns from Continuously Archived GPS Data
Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa
MMM3
2009 Temporal Segmentation of 3-D Video by Histogram-Based Feature Vectors
abstract
Three-dimensional (3-D) video, which is a sequence of time-varying mesh models generated in a multi-camera studio, is attracting increased attention, because it can record and reproduce the 3-D information of real-world objects with high accuracy. As one of the most important preprocessings for indexing, annotation, retrieval, and many other functions in management of a 3-D video database, it is necessary to temporally segment 3-D video into meaningful and manageable segments. We have developed robust and effective segmentation algorithms using histogram-based feature vector representation, striving to understand and manage 3-D video contents. We have developed two approaches to generate feature vectors by vertex positions in the mesh models: one uses the Cartesian coordinate system and the other employs the spherical coordinate system. Then, 3-D video is segmented by the motion intensity of an object, which is analyzed by the feature vectors. The segmentation algorithms we have developed are applied to three different 3-D video sequences. A statistical method is presented to evaluate the segmentation results. High recall and precision rates of 0.95 and 0.77, respectively, are achieved in the best case.
Toshihiko Yamasaki, Kiyoharu Aizawa
IEEE Trans. Circuits Syst. Video Technol.3
2009 Sketch-Based Spatial Queries for Retrieving Human Locomotion Patterns From Continuously Archived GPS Data
abstract
We propose a system for retrieving human locomotion patterns from tracking data captured within a large geographical area, over a long period of time. A GPS receiver continuously captures data regarding the location of the person carrying it. A constrained agglomerative hierarchical clustering algorithm segments these data according to the person's navigational behavior. Sketches made on a map displayed on a computer screen are used for specifying queries regarding locomotion patterns. Two basic sketch primitives, selected based on a user study, are combined to form five different types of queries. We implement algorithms to analyze a sketch made by a user, identify the query, and retrieve results from the collection of data. A graphical user interface combines the user interaction strategy and algorithms, and allows hierarchical querying and visualization of intermediate results. We evaluate the system using a collection of data captured during nine months. The constrained hierarchical clustering algorithm is able to segment GPS data at an overall accuracy of 94% despite the presence of location-dependent noise. A user study was conducted to evaluate the proposed user interaction strategy and the usability of the overall system. The results of this study demonstrate that the proposed user interaction strategy facilitates fast querying, and efficient and accurate retrieval, in an intuitive manner.
Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa
IEEE Trans. Multim.3
2008 Error analysis of 3Dc-based normal map compression and its application to optimized quantization
abstract
Normal mapping is one of the most essential technologies for realistic three-dimensional computer graphics. In conventional normal map compression such as 3Dc, only the x and y components are encoded and the z components are restored based on the normalizing condition. In this paper, we present an intuitively comprehensive error analysis for this approach. As a result, we reveal in what condition compression error becomes larger. We also present a non-linear quantization algorithm based on the formula for better compression performance than the conventional approaches. Experimental results using 300 normal map demonstrate that the PSNR is improved by 0.29 dB on average. Our algorithm is compatible with random access and highly-parallel processing on GPU.
Toshihiko Yamasaki, Kiyoharu Aizawa
ICASSP2
2008 Geometry compression for time-varying meshes using coarse and fine levels of quantization and run-length encoding
abstract
Time-varying meshes (TVM) is a new 3-D scene representation which are generated from multiple cameras. It captures highly detailed shape and texture as well as movement of real-world moving objects. Compression is a key technology for supporting TVM applications such as education, interactive broadcasting, and intangible heritage archiving. Previous works focused on compression of 3-D animation that has the same topology throughout the entire sequences. Unfortunately, the topology of TVMs change with time which makes it difficult to compress TVMs. In this paper, we propose a geometry encoder for TVMs. The encoder finds spatial and temporal redundancy by coarse and fine level quantization. Thereafter, vertex information is converted into binary sequences. And then, the binary sequences are encoded using run-length encoding (RLE). Experimental results show that vertices of TVMs which require 96 bits per vertex (bpv) are compressed to 1.9-15.4 bpv while maintaining a small geometric distortion ranging from 0.7 times 10-4to 1.3 times 10-3% of the maximum error.
Seung-Ryong Han, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP3
2008 Hierarchical mesh decomposition and motion tracking for Time-Varying-Meshes
abstract
This paper proposes a system for automatic segmentation and motion tracking of Time-Varying-Meshes (TVM). Our approach is based on skeleton-based hierarchical mesh decomposition by distance calculation. The properties of the human skeleton structure are used to define the decomposition of each TVM frame. The proposed framework is a recursive system that iterates between automatic hierarchical decomposition on minimum distance satisfaction and skeleton realignment. This is done to achieve a stable segmentation and a refined skeleton. By utilizing color information, ill-defined meshes can be successfully segmented. Results show an average disparity of 1.39% of total surface area across all segmented parts, which indicate stability of the system. In addition, motion tracking is successfully performed through the use of refined skeletons.
Ning Sung Lee, Toshihiko Yamasaki, Kiyoharu Aizawa
ICME3
2008 High level activity annotation of daily experiences by a combination of a wearable device and Wi-Fi based positioning system
abstract
Many people would like to record and manipulate their experiences effectively. However, efficient summarization to show “what, when and where” we did in our daily lives is still an open issue in life log applications. In conventional approaches, many sensors were attached to a human body to solve this problem. However, this is not practical for daily use. In this paper, we propose a simple solution in which a user wears two devices: a single life-logging device SenseCam [1] hung by neck and a Wi-Fi enabled PDA. The location data and low-level activity are analyzed by a Wi-Fi based positioning system. Then, high-level activity classification is conducted using the wearable device along with a refinement process considering the time consistency. Experimental results demonstrated high recall and precision rates as much as 81% and 85% respectively, on average.
Wayhit Puangpakisiri, Toshihiko Yamasaki, Kiyoharu Aizawa
ICME3
2008 Food log by analyzing food images
abstract
In this paper, a food-logging system that can distinguish food images from other images, analyze the food balance, and visualize the log is presented. The image processing is based on feature vectors consisting of color histograms, DCT coefficients, detected image patterns and so forth. Support Vector Machine (SVM) was used to detect food images and to analyze the food balance. Experimental results show that the food image extraction presents above 88% of accuracy and the food balance estimation is achieved with more than 73% of accuracy.
Keigo Kitamura, Toshihiko Yamasaki, Kiyoharu Aizawa
ACM Multimedia3
2008 Interactive retrieval for multi-camera surveillance systems featuring spatio-temporal summarization
abstract
An interactive interface is presented for near-synchronized and distributed multi-camera surveillance systems. Human tracking using multiple cameras is conducted employing a particle filter in conjunction with linear least-square trajectory extrapolation and color histogram matching. Spatial and temporal statistical information such as histograms of the number of pedestrians on a timeline-basis, pedestrians' trajectories over a certain period, pedestrian flow, and so forth is efficiently summarized and visualized. In addition, a sketch-based retrieval interface is also developed. As a result, our system facilitates operators to easily extract and access to important scenes from a huge amount of surveillance data. The real-life experiments in the public street demonstrated the validity of our system.
Toshihiko Yamasaki, Yoshifumi Nishioka, Kiyoharu Aizawa
ACM Multimedia3
2008 Audio Analysis for Multimedia Retrieval from a Ubiquitous Home
Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa
MMM3
2008 Large-scale image sensing by a group of smart image sensors
Myeongsoo Oh, Kiyoharu Aizawa
Parallel Comput.2
2007 A Sensor Network for Event Retrieval in a Home Like Ubiquitous Environment
abstract
We present the current status of a system based on a network of sensors for event retrieval from a home like environment. A large number of cameras, microphones and pressure based floor sensors are used for continuous data capture. The data are analyzed independently and the results recorded in a central database, where they are combined for efficient retrieval and summarization of the video and audio data. We describe the detection of basic actions, events, and faces using image analysis. The users can query the system interactively to retrieve video, audio, and key frames corresponding to events. We report results of performance evaluations of the algorithms and discuss issues related to sensing, analysis and retrieval of data.
Gamhewage Chaminda de Silva, Steve Anavi, Toshihiko Yamasaki, Kiyoharu Aizawa
ICASSP (4)4
2007 Highly Efficient VQ-Based Normal Map Compression using Quality Estimation Model
abstract
Normal maps play an important role in computer 3D graphics to express pseudo roughness of the surface with a small amount of polygon data. In this paper, a highly efficient normal map compression algorithm is proposed based on an estimation model to predict the quality of the images rendered with the compressed normal maps. The optimal encoding is achieved by minimizing the predicted mean square error (MSE) employing vector quantization (VQ). In addition, encoding and decoding time is fast enough for practical usage. Experimental results demonstrate that the algorithm proposed in this paper yields better compression performance than the other algorithms in the literatures.
Toshihiko Yamasaki, Kiyoharu Aizawa
ICASSP (1)2
2007 Tracking Persons using Particle Filter Fusing Visual and Wi-Fi Localizations for Widely Distributed Camera
abstract
This paper describes an object tracking scheme employs sensor fusion approach which is composed of visual information and location information estimated from Wi-Fi signals. Location information is calculated by a set of received signal strength values of beacon packets from Wi-Fi access points (APs) around the targets. Different from the conventional approaches which use another kind of sensors, our approach can cover wider areas both indoor and outdoor with lower cost because of characteristics of Wi-Fi signals. Particle filter is applied to combine these two different kinds of sensory input to track the target continuously. Wi-Fi observation model is involved in a conventional visual particle filtering scheme in order to evaluate importance weights of each particle. By using multiple modality, robust tracking performance is achieved even if reliability of one sensory input declines. In this paper, we present experimental results applied to outdoor surveillance camera environment.
Takashi Miyaki, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP (3)3
2007 Geometrically Invariant Object-Based Watermarking using SIFT Feature
abstract
In this paper, we have developed a robust object-based watermarking algorithm using the scale-invariant feature transform (SIFT) features in conjunction with a new data embedding method based on discrete cosine transform (DCT). The message is embedded in DCT spaces of randomly generated blocks in the selected object region. To recognize the object region after being distorted, its SIFT features are registered in advance. In the detection scheme, we firstly detect the object region by using feature matching. The transformation parameters are then calculated, and the message can be detected. Experimental results demonstrated that our proposed algorithm is very robust to geometrical distortions such as JPEG compression, scaling, rotation, shearing, aspect ratio change, image filtering, and so on.
Viet Quoc Pham, Takashi Miyaki, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP (5)4
2007 View-Based Web Page Retrieval using Interactive Sketch Query
abstract
We propose a novel view-based Web page retrieval system that enables a user to search Web pages using a visual query, namely the user's freehand sketch. We believe the proposed method will suit retrieval from a set of web pages such as a user's local browsing history. The system aims to help the user revisit a particular Web page without using query words. Using color signature features and Earth-Mover's distance, the system evaluates the similarity between web pages and the user's sketch drawn via the GUI of the system. In order to accelerate the interaction, the results of the similarity evaluations are shown immediately after the user draws each stroke of the sketch, with the results being interactively reordered. Experiments using our prototype system showed that users find their target pages after only a few strokes. Experimental results for inexperienced users showed that 71% of search tasks were completed within one minute, using the prototype system. The median time for the tasks was 40 seconds.
Yasuyuki Watai, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP (6)3
2007 Visual Tracking of Pedestrians Jointly using Wi-Fi Location System on Distributed Camera Network
abstract
Object tracking with multiple cameras is a fundamental problem in wide-area surveillance application, but it has difficulties to achieve accurate and stable performance because of disjoint shot areas or initial object identification problems. We propose a novel object tracking method which jointly uses estimated location information of the target derived from a set of Wi-Fi signal strength values with video images from cameras. Apart from sensor fusion techniques proposed in the past which use another kind of sensors (eg., global positioning system (GPS), pressure sensors on the floors to detect foot steps, laser-range scanners, etc), our approach can cover wide-areas both indoor and outdoor in low cost because of propagation characteristics of Wi-Fi signals. Particle filtering is applied to achieve tracking the target from video images, with this Wi-Fi location estimation. This paper describes system architectures and the experimental results of the pedestrian tracking technique with Wi-Fi location estimation.
Takashi Miyaki, Toshihiko Yamasaki, Kiyoharu Aizawa
ICME3
2007 Fast and Robust Motion Tracking for Time-Varying Mesh Featuring Reeb-Graph-Based Skeleton Fitting and its Application to Motion Retrieval
abstract
In this paper, an algorithm for motion extraction from time-varying mesh (TVM) is proposed. TVM is a sequence of 3D models made for real-world dynamic objects. In TVM, detailed information of 3D objects such as shape, color, and motion is stored. Therefore, TVM is an attractive technology for the next-generation multimedia. However, the management of TVM database for archiving and retrieving is a significant problem because it is difficult to locate and track feature points in TVM. For efficient and effective archive and retrieval systems, motion extraction and tracking is essential. In our approach, skeletons are extracted from TVM using Reeb graph, and motion tracking is achieved by fitting a reference skeleton to the others. We defined a geodesic function using principal component analysis (PCA) for fast and noiseless skeleton extraction. Robust fitting has been realized by an end node tracking strategy. In addition, we have developed an efficient motion retrieval system compatible to conventional motion capture (mocap) data using the extracted skeleton. As a result, low-cost query generation using mocap systems and cross search between TVM and mocap data have been achieved.
Ryuichi Tadano, Toshihiko Yamasaki, Kiyoharu Aizawa
ICME3
2007 Deformation of Time-Varying-Mesh Based on Semantic Human Model
abstract
In this paper, an approach is presented for deformation of time-varying-mesh, which is a sequence of 3D mesh models. The deformation here is a process to generate mid-frames between two frames that transform smoothly from one to the other. Motion vectors between two frames are extracted based on a semantic human model with an assumption of the articulated object with piecewise-rigid motions. For this purpose, each mesh model is transformed to a volumetric model, where a distance field is constructed. In the first frame, the user manually segments the volumetric model into the semantic human model. Then, fast motion estimation of the volumetric models is performed between the two frames with a cost function based on the distance field. Using the estimated motion vectors, a realistic deformation of a mesh model is achieved. Lastly, mid-frames are interpolated linearly. The technique can be applied in many areas such as frame rate up-conversion and motion blending.
Toshihiko Yamasaki, Kiyoharu Aizawa
ICME3
2007 Content-Based Cross Search for Human Motion Data using Time-Varying Mesh and Motion Capture Data
abstract
This paper describes a content-based cross search scheme for two kinds of three-dimensional (3D) human motion data: time-varying mesh (TVM) and motion capture data. TVM is a sequence of 3D mesh models made for real-world 3D objects. TVM is generated using multiple-view images taken with multiple cameras. Since TVM can record shape, color, and motion of the real-world 3D objects, it has been drawing a lot of attention these days. In order to realize practical archiving systems for TVM, efficient retrieval systems are indispensable. The retrieval systems for TVM developed so far are based on query-by-example. This means additional TVM generation is required for constructing queries, which is computationally demanding and time consuming. On the other hand, motion capture systems are widely used to capture 3D human motion. However, the data structure is very different from that of TVM. Therefore, the two kinds of 3D human motion data are incompatible to each other. In this paper, we present a retrieval system that enables retrieving TVM using motion capture data as queries and vice versa using the modified shape distribution algorithm.
Toshihiko Yamasaki, Kiyoharu Aizawa
ICME2
2007 Spatial querying for retrieval of locomotion patterns in smart environments
abstract
A system for retrieving video sequences created by tracking humans in a smart environment, by using spatial queries, is presented. Sketches made on a graphical user interface using a pointing device are used as the means of entering multiple types of queries. After preprocessing and coordinate system conversion, the sketches are analyzed to identify the type of the query. Directional search algorithms based on the minimum distance between points is applied for finding the best matches to the sketch. The results are ranked according to the similarity and presented to the user. The results of an initial evaluation are reported. The paper concludes with an outline of possible future directions.
Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa
ACM Multimedia3
2007 Motion Structure Parsing and Motion Editing in 3D Video
Toshihiko Yamasaki, Kiyoharu Aizawa
MMM (1)3
2007 Time-Varying Mesh Compression Using an Extended Block Matching Algorithm
abstract
Time-varying mesh, which is attracting a lot of attention as a new multimedia representation method, is a sequence of 3-D models that are composed of vertices, edges, and some attribute components such as color. Among these components, vertices require large storage space. In conventional 2-D video compression algorithms, motion compensation (MC) using a block matching algorithm is frequently employed to reduce temporal redundancy between consecutive frames. However, there has been no such technology for 3-D time-varying mesh so far. Therefore, in this paper, we have developed an extended block matching algorithm (EBMA) to reduce the temporal redundancy of the geometry information in the time-varying mesh by extending the idea of the 2-D block matching algorithm to 3-D space. In our EBMA, a cubic block is used as a matching unit. MC in the 3-D space is achieved efficiently by matching the mean normal vectors calculated from partial surfaces in cubic blocks, which our experiments showed to be a suboptimal matching criterion. After MC, residuals are transformed by the discrete cosine transform, uniformly quantized, and then encoded. The extracted motion vectors are also entropy coded after differential pulse code modulation. As a result of our experiments, 10%-18% compression has been achieved.
Seung-Ryong Han, Toshihiko Yamasaki, Kiyoharu Aizawa
IEEE Trans. Circuits Syst. Video Technol.3
2007 Reconstructing Dense Light Field From Array of Multifocus Images for Novel View Synthesis
abstract
This paper presents a novel method for synthesizing a novel view from two sets of differently focused images taken by an aperture camera array for a scene consisting of two approximately constant depths. The proposed method consists of two steps. The first step is a view interpolation to reconstruct an all-in-focus dense light field of the scene. The second step is to synthesize a novel view by a light-field rendering technique from the reconstructed dense light field. The view interpolation in the first step can be achieved simply by linear filters that are designed to shift different object regions separately, without region segmentation. The proposed method can effectively create a dense array of pin-hole cameras (i.e., all-in-focus images), so that the novel view can be synthesized with better quality.
Akira Kubota, Kiyoharu Aizawa, Tsuhan Chen
IEEE Trans. Image Process.2
2006 Fast and Efficient Normal MAP Compression Based on Vector Quantization
abstract
Normal maps play an important role in realistic 3D image rendering to express pseudo roughness of the surface with small amount of polygon data. In this paper, a fast and efficient normal map compression algorithm is proposed based on vector quantization and entropy coding. Using the strong correlation among x, y, and z components of normal maps owing to the unity condition, compression ratio has been made much better than conventional approaches. In addition, the encoding time has been made reasonable by considering the distribution of the data and employing inner product in nearest-neighbor search instead of Euclidian distance taking advantage of the unity condition of the training data
Toshihiko Yamasaki, Kiyoharu Aizawa
ICASSP (2)2
2006 3D Video Compression Based on Extended Block Matching Algorithm
abstract
Three dimensional (3D) video is attracting a lot of attention as a new multimedia representation method. 3D video is a sequence of 3D models (frames) that consist of varying vertices and connectivity. In conventional 2D video compression algorithms, motion compensation (MC) using block matching algorithm is frequently employed to reduce redundancy between consecutive frames. However, there is no such technology for 3D video so far. Therefore, in this paper, we have developed an extended block matching algorithm (EMBA) to reduce temporal redundancy of geometry information of 3D video by extending the idea of 2D block matching to 3D space. In our EBMA, a cubic block is used as a matching unit and, MC is achieved efficiently by matching the mean normal vectors of the sub-blocks, which turned out to be sub-optimal by our experiments. The residual information is further transformed by discrete cosine transform (DCT) and then encoded. The extracted motion vectors are also entropy encoded. As a result of our experiments, compression ratio ranging from 10% to 18% of the original 3D video data has been achieved.
Seung-Ryong Han, Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP3
2006 Key Frame Extraction in 3D Video by Rate-Distortion Optimization
abstract
3D video, which consists of a sequence of 3D mesh models, can provide detailed 3D information both in spatial and temporal domain. In this paper, a key frame extraction method has been developed to summarize 3D video by rate-distortion optimization. For this purpose, we introduce an effective feature vector extraction algorithm from 3D video. Prior to key frame extraction, shot detection is performed using the feature vectors as a pre-processing. Then, a rate-distortion (R-D) curve is generated in each shot, where the locations of key frames are optimized. Lastly, R-D trade-off can be achieved by optimizing a cost function with a Lagrange multiplier. Our experimental results show the extracted key frames are compact and faithful to original 3D video
Toshihiko Yamasaki, Kiyoharu Aizawa
ICME3
2006 Motion Segmentation of 3D Video using Modified Shape Distribution
abstract
In this paper, temporal segmentation of 3D video based on motion analysis is presented. 3D video is a sequence of 3D models made for a real-world dynamic object. A modified shape distribution algorithm is proposed to realize stable shape feature representation. In our approach, representative points are generated by clustering vertices based on their spatial distribution instead of randomly sampling vertices as in the original shape distribution algorithm. Motion segmentation is conducted analyzing local minima in degree of motion calculated in the feature vector space. The segmentation algorithm developed in this paper does not require any predefined threshold values but rely on relative relationships among local minima and local maxima of the motion. Therefore, robust segmentation has been achieved. The experiments using 3D video of traditional dances yielded encouraging results with the precision and recall rates of 93% and 88%, respectively, on average
Toshihiko Yamasaki, Kiyoharu Aizawa
ICME2
2006 Effects of physical display size and amplitude of oscillation on visually induced motion sickness
abstract
Viewing environment is an important factor to understand the mechanism of visually induced motion sickness (VIMS). In Experiment 1, we investigated whether the symptom of VIMS changed depending on viewing angle and physical display size. Our results showed that larger viewing angle made the symptom of sickness severer and nausea symptom changed depending on physical display size with identical viewing angles. In Experiment 2, we investigated effects of viewing angle and amplitude of oscillation. The results showed that the effects of viewing angle were not only related to amplitude of oscillation but also to the other factors of viewing angle.
Hiroaki Shigemasu, Toshiya Morita, Naoyuki Matsuzaki, Takao Sato, Masamitsu Harasawa, Kiyoharu Aizawa
VRST6
2005 Direct filtering method for image based rendering
abstract
Image based rendering (IBR) is basically a light ray resampling method for generating a novel image without aliasing artifacts from a given set of sampled light rays. To avoid aliasing artifacts, conventional methods require an estimate of the scene geometry. In this paper, we present a novel IBR method by linear filtering without estimating the scene geometry, for the simplified case when generating a virtual image at the center of 2/spl times/2 sparse camera array for a two depth-layers scene. The reconstruction filter used in the proposed method can derive by integrating all the process in our previously proposed method into a one-shot process.
Akira Kubota, Kiyoharu Aizawa, Tsuhan Chen
ICIP (3)2
2005 An adaptive video stabilization method for reducing visually induced motion sickness
abstract
Visually induced motion sickness (VIMS) is sometimes brought by watching video sequences acquired by handheld video cameras. We present a video stabilization method for the purpose of reducing VIMS. In our method, the scenes including oscillatory motion which can cause VIMS are selected to be stabilized. In the subjective evaluation results, the severity of VIMS is reduced.
Ikuko Tsubaki, Toshiya Morita, Takahiro Saito, Kiyoharu Aizawa
ICIP (3)4
2005 3D video segmentation using point distance histograms
abstract
Similar to 2D video segmentation, 3D video segmentation is to divide the 3D video in temporal domain into a set of meaningful and manageable segments (shots) that are used as basic elements for indexing. This paper proposes a temporal segmentation method for 3D video for the first time as far as we know. Point distance histograms are used considering the tradeoff between computational cost and effectiveness. In order to reflect the real motion of 3D object, three fixed points are selected, which can avoid what we call "the same sphere problem." And simulation results on total 285 frames, which are composed of four different sequences, show our method is very effective.
Toshihiko Yamasaki, Kiyoharu Aizawa
ICIP (1)3
2005 Mathematical error analysis of normal map compression based on unity condition
abstract
Normal maps play an important role in realistic 3D image rendering to express pseudo roughness of the surface. In normal map compression, z components are often eliminated and calculated from x and y components based on the unity condition to achieve high compression rate. However, there is no theoretical background of the validity of eliminating z components so far. In this paper, a mathematical mean square error (MSE) model of such normal map compression based on the unity condition is proposed. In addition, the boundary condition whether to eliminate or include the z components in compression for the better efficient encoding is demonstrated based on our error analysis model.
Toshihiko Yamasaki, Kazuya Hayase, Kiyoharu Aizawa
ICIP (2)3
2005 Depth estimation for synthesizing arbitrary view images by random access IBR sensor array
abstract
We have been investigating image-based rendering (IBR) imaging system by using array of smart image sensors. The imaging system consists of random accessible image sensors and FPGA, arbitrary view images can be obtained in real-time. Although IBR system can generate more realistic images compared to model-based rendering (MBR) system, the generated images are degraded if the depth information for synthesis has some errors. In this paper, we describe a method of depth estimation for the application of the imaging system. It is based on comparing various synthesized images to an actual image obtained by a real sensor repeatedly. By using the depth information, more realistic arbitrary view images can be generated.
Nao Yuki, Takayuki Hamamoto, Kiyoharu Aizawa
ICIP (3)3
2005 Video handover for retrieval in a ubiquitous environment using floor sensor data
abstract
A system for retrieving video captured in a ubiquitous environment is presented. Data from pressure-based floor sensors are obtained as a supplementary input together with video from multiple stationary cameras. Unsupervised data mining techniques are used to reduce noise present in floor sensor data. An algorithm based on agglomerative hierarchical clustering is used to segment footpaths of individual persons. Video handover is proposed and two methods are implemented to retrieve video and key frame sequences showing a person moving in the house. Users can query the system based on time and retrieve video or key frames using either of the handover techniques. We compare the results of retrieval using different techniques subjectively. We conclude with suggestions for improvements, and future directions.
Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Takayuki Ishikawa, Kiyoharu Aizawa
ICME4
2005 Evaluation of video summarization for a large number of cameras in ubiquitous home
abstract
A system for video summarization in a ubiquitous environment is presented. Data from pressure-based floor sensors are clustered to segment footsteps of different persons. Video handover has been implemented to retrieve a continuous video showing a person moving in the environment. Several methods for extracting key frames from the resulting video sequences have been implemented, and evaluated by experiments. It was found that most of the key frames the human subjects desire to see could be retrieved using an adaptive algorithm based on camera changes and the number of footsteps within the view of the same camera. The system consists of a graphical user interface that can be used to retrieve video summaries interactively using simple queries.
Gamhewage Chaminda de Silva, Toshihiko Yamasaki, Kiyoharu Aizawa
ACM Multimedia3
2005 Digitizing Personal Experiences: Capture and Retrieval of Life Log
abstract
In wearable computing environments, digitization of personal experiences will be made possible by continuous recordings using a wearable video camera. This could lead to the ``automatic life-log application''. However, it is evident that the resulting amount of video content will be enormous. Accordingly, to retrieve and browse desired scenes, a vast quantity of video data must be organized using structural information. We are developing a ``context-based video retrieval system for life-log applications''. Our life log system captures video, audio, acceleration sensor, gyro, GPS, annotations, documents, web pages, and emails, and provides functions that make efficient video browsing and retrieval possible by using data from these sensors, some databases and various document data.
Kiyoharu Aizawa
MMM1
2005 Reconstructing arbitrarily focused images from two differently focused images using linear filters
abstract
We present a novel filtering method for reconstructing an all-in-focus image or an arbitrarily focused image from two images that are focused differently. The method can arbitrarily manipulate the degree of blur of the objects using linear filters without segmentation. The filters are uniquely determined from a linear imaging model in the Fourier domain. An effective and accurate blur estimation method is developed. The simulation results show that the accuracy and computational time of the proposed method are improved compared with the previous iterative method and that the effects of blur estimation error on the quality of the reconstructed image are very small. The method performs well for real images acquired without visible artifacts.
Akira Kubota, Kiyoharu Aizawa
IEEE Trans. Image Process.2
2004 Virtual view synthesis through linear processing without geometry
abstract
This paper presents a new approach for virtual view synthesis that does not require any information of scene geometry. Our approach first generates multiple virtual views at the same position based on multiple depths by the conventional view interpolation method. The interpolated views suffer from blurring and ghosting artifacts due to the pixel mis-correspondence. Secondly, the multiple views are integrated into a novel view where all regions are focused. This integration problem can be formulated as the problem of solving a set of linear equations that relates the multiple views. To solve this set of equations, two methods using projection onto convex sets (POCS) and inverse filtering are presented that effectively integrate the focused regions in each view into a novel view. Experimental results using real images show the validity of our methods.
Akira Kubota, Kiyoharu Aizawa, Tsuhan Chen
ICIP2
2004 Restoration and demosaicing for pixel mixture images in dsc video clips
abstract
This paper proposes restoration methods for "pixel mixture". Pixel mixture is a function for saving readout time for digital still cameras (DSCs), which mixes together two pixels on CCD. Pixel mixture reduces the number of pixels to be read out from CCD, and it enables video rate read out from DSCs of more than a million pixels. In this paper, an iterative method is presented in order to recover the image. Both restoration and demosaicing are performed at the same time. Our approach achieves the result of smaller mean square error than conventional interpolation methods.
Ikuko Tsubaki, Kiyoharu Aizawa
ICIP2
2004 Capturing life-log and retrieval based on contexts
abstract
In "wearable computing" environments, digitization of personal experiences will be made possible by continuous recordings using a wearable video camera. This could lead to the "automatic life-log application". However it is evident that the resulting amount of video content will be enormous. Accordingly, to retrieve and browse desired scenes, a vast quantity of video data must be organized using structural information. We are developing a "context-based video retrieval system for life-log applications". This system can capture not only video and audio but also various sensor data and provides functions that make efficient video browsing and retrieval possible by using data from these sensors, some databases and various document data.
Tetsuro Hori, Kiyoharu Aizawa
ICME2
2004 Reconstructing dense light field from a multi-focus images array
abstract
The work presents a novel method for synthesizing a novel view from two sets of differently focused images taken by a sparse camera array for a scene of two approximately constant depths. The proposed method consists of two steps. The first step is a view interpolation to reconstruct an all-focused dense light field of the scene. The second step is to synthesize a novel view by a light-field rendering technique from the reconstructed dense light field. The view interpolation can be achieved simply by linear filters that are designed to convert defocus effects to parallax effects without estimating the depth map of the scene. The proposed method can effectively create a dense array of pin-hole cameras (i.e., all-focused images), so that the final novel view is better than traditional method using a sparse array of cameras. Experimental results on real images from four aligned cameras are shown.
Akira Kubota, Kiyoharu Aizawa, Tsuhan Chen
ICME2
2003 A novel image-based rendering method by linear filtering of multiple focused images acquired by a camera array
abstract
In this paper, we present a novel approach to image-based rendering (IBR) for generating an arbitrary view image with arbitrary focus for a scene consisting two approximately constant depths. The presented method differs from the conventional IBRs using multiple view images in that we acquire two differently focused images from each camera position and render parallax and focus effects on objects at different depth simply by linear filtering of the acquired images without segmentation. Experimental results on the real images acquired with 4 cameras located in parallel are presented.
Akira Kubota, Kiyoharu Aizawa
ICIP (3)2
2003 Wide dynamic range imaging by sensitivity adjustable CMOS image sensor
abstract
In this paper, we propose the wide dynamic range imaging system by using the sensitivity adjustable CMOS image sensor. This method is effective to capture a scene that has both the very bright area and the dark area at the same time. Proposed sensor has the adjustable sensitivity pixels and the sensitivity is changed by switching the extra capacitor in each pixel circuit. We designed and implemented the 200 /spl times/ 200 pixels of the prototype VLSI chip. Verification result shows the sensitivity changes 6 dB of range. We also made the wide dynamic range image capture system with this prototype and observed the appropriate output of the image sensor.
Ryutaro Oi, Kiyoharu Aizawa
ICIP (2)2
2003 Wearable imaging system for summarizing personal experiences
abstract
Digitization of lengthy personal experiences would be made possible by constant recording using wearable video cameras. It is conceivable that the resulting amount of video content would be extraordinarily large. In order to retrieve and browse the desired scenes, a vast amount of video would need to be organized with structural information. In this paper, we attempt to develop a "wearable imaging system" that is capable of constantly capturing data, not only from a wearable video camera, but also from various kinds of sensors, such as a GPS, an accelerometer and a gyro sensor. The data from these sensors are appropriately extracted and processed by hidden Markov model (HMM) to achieve efficient video retrieval and browsing.
Yasuhito Sawahata, Kiyoharu Aizawa
ICME2
2003 Object-based approach to image-based rendering with linear filters using defocus information
Akira Kubota, Kiyoharu Aizawa
VCIP2
2003 Three-dimensional image representation of buildings utilizing heterogeneous information for multimedia ambiance communication
Minako Toba, Takahiro Saito, Kiyoharu Aizawa, Kenji Mochizuki, Takeshi Naemura, Hiroshi Harashima
VCIP3
2002 Three dimensional modeling of large-scale real environment by fusing range data, texture images, and airborne altimetry data
abstract
Construction of large-scale virtual environment is gaining more attention for its applications in virtual malls, virtual sightseeing, tele-presence, etc. We introduce a framework to construct a realistic large-scale virtual environment by fusing range data, texture images, and airborne altimetry data. First, realistic and high-precision 3-D models of buildings are created from range data and texture images, which are taken by long-range laser scanner and digital camera, respectively. Next, rough 3-D models of buildings in wide area are created using altimetry data acquired by airborne laser profiler. Finally, these models are integrated to build a large-scale realistic walk-through system. This paper describes the proposed system, issues related to the system, part of its implementations, as well as future work that still has to be done.
Kiyoharu Aizawa, Conny Riani Gunadi, Hiroyuki Shimizu, Kazuya Kodama
ICIP (2)1
2002 Arbitrary view and focus image generation: rendering object-based shifting and focussing effect by linear filtering
abstract
This paper presents a novel method to render shifting (parallax) and focussing effects on an object in a scene for generating a virtual view image with arbitrary focus. Two differently focused images of the same scene-near-focused image and far-focused image-are used as input under the assumption that a scene has near and far objects. The proposed method can freely handle the shifting and the focussing effect on each object according to the view point and the focus depth of the virtual camera only by linear filtering of the input images without any segmentation or modeling of the objects. Experimental results using real images are shown to test the performance of the method.
Akira Kubota, Kiyoharu Aizawa
ICIP (1)2
2002 Real-time objects tracking by using smart image sensor and FPGA
abstract
We have been investigating the integration of sensing and compression on an image sensor. The compression sensor reduces the number of pixels in the image signal that have to be readout from the sensor. We present a real-time object tracking system using a compression sensor, which has 128/spl times/128 pixels, and FPGA. By using this system, several moving objects can be extracted and tracked independently at 1200 frames/second. We also describe real-time depth estimation by using binocular compression sensors. We show some results obtained by the system.
Shouichi Nagao, Takayuki Hamamoto, Kiyoharu Aizawa
ICIP (3)3
2002 Summarization of wearable videos using support vector machine
abstract
Auto-summarization of video contents has become an important topic following the growing amount of multimedia contents. Researches in this area have shown the effectiveness of low-level video and audio features in categorizing video contents. The use of brainwaves to reflect personal interests is also proven to be practical. In this paper, we model the relationship between audio/video features and brainwaves (/spl alpha/-waves) using the support vector machine (SVM). Based on the SVM model, we summarized wearable videos by personal interests, using only low-level video and audio features. Here we define "wearable videos" as continuous recordings of personal experiences using wearable video camera and computer. Our experiment results showed over 90% of accuracy on summarization of a 25-minute video clip with an SVM model created by another resembling 25-minute video clip.
Haung Wei Ng, Yasuhito Sawahata, Kiyoharu Aizawa
ICME (1)3
2001 Summarizing wearable video
abstract
"We want to record our entire life by video" is the motivation of this research. Developing wearable devices and huge storage devices will make it possible to keep entire life by video. We could capture 70 years of our life, however, the problem is how to handle such a huge amount of data. Automatic summarization based on personal interest should be required. In this paper we propose an approach to the automatic structuring and summarization of wearable video. (Wearable video is our abbreviation of "video captured by a wearable camera".) In our approach, we make use of a wearable camera and a sensor of brain waves. The video is firstly structured by objective features of video, and the shots are rated by subjective measures based on brain waves. The approach is very successful for real world experiments and it automatically extracted all the events that the subjects reported they had felt interesting.
Kiyoharu Aizawa, Kenichiro Ishijima, Makoto Shiina
ICIP (3)1
2001 A new approach to depth range detection by producing depth-dependent blurring effect
abstract
This paper presents a new method for detecting the depth range of a scene from two differently focused images. This method first generates three images with differently emphasized blur using linear filtering of the two acquired images for more depth-sensitive estimation. In this point, this method is different from conventional approaches. The discrete Fourier transform of the ratio between the generated images is then introduced as a criterion of the depth range. Finally, through the thresholding process of the criterion according to its correspondence to the depth on the basis of the imaging model, the depth range is estimated in five depth steps. Experiments for synthesized and real images are performed to test the proposed method.
Akira Kubota, Kiyoharu Aizawa
ICIP (3)2
2001 Pixel independent random access image sensor for real time image-based rendering system
abstract
We have been investigating a high-speed image-based rendering system. In this area, most of the conventional systems sacrifice spatial or temporal resolution for a heavy amount of input images. When the system uses a camera-array on its input, this problem is more obvious. However, the required image data for the rendering are only a portion of them, determined by the position of the imaginary-view. We propose an image based rendering system, which uses pixel independent random access image sensors to eliminate the bottleneck of the conventional systems. We have developed a prototype of a CMOS image sensor, which has 128 /spl times/ 128 pixels. We verified that the prototype chip readouts externally selected pixels at 60 frames/second.
R. Ooi, Takayuki Hamamoto, Takeshi Naemura, Kiyoharu Aizawa
ICIP (2)4
2001 All-focused image generation and 3D modeling of microscopic images of insects
abstract
We discuss a method which generates an all-focused image from a large number of microscopic images of insects. First, we describe our previously proposed select-and-merge method for all-focused image acquisition. We can get good results by using this method for two differently focused images. However, this method can not give a good result when applied to a large number of images. We propose a method which compares each images with its preceding and subsequent images using the estimation method of focused regions. We compare the results of reconstruction using the novel method and using our previously proposed method. Finally, we generate a 3D image using the intensity of the all-focused image and depths of the in-focus regions, both acquired by our proposed method.
Y. Tsubaki, Akira Kubota, Kiyoharu Aizawa
ICIP (2)3
2001 Foreground Extraction Based On Logical Operation Of 3-D Array Frame Difference
Supatana Auethavekiat, Kiyoharu Aizawa
ICME2
2001 Structure analysis of natural scenes using census transform and region competition
Kunio Yamada, Tadashi Ichikawa, Takeshi Naemura, Kiyoharu Aizawa, Takahiro Saito
VCIP4
2000 New design and implementation of adaptive-integration-time image sensor
abstract
We describe a novel approach to enhance the performance of image sensing by integrating the processing element with an image sensor. We have been investigating a smart sensor which controls the integration time of every pixel independently. The integration time is controlled so that it has higher temporal resolution and wider dynamic range. We describe a new adaptive-integration-time image sensor which has 128/spl times/64 pixels. We have designed the VLSI prototype by using a column parallel architecture. In this new chip, the scheme to control the integration time is extended, and the pixel pitch processing speed and power consumption are much improved in comparison with our previous prototype. A 7-bit address encoder is newly implemented. We show some experimental results obtained with the prototype.
Takayuki Hamamoto, Kiyoharu Aizawa
ICASSP2
2000 A Computational Image Sensor with Pixel-Based Integration Time Control
abstract
We have been investigating a computational sensor which controls the integration time of every pixel independently. Because the integration time is controlled, higher temporal resolution and wider dynamic range can be achieved. We present a new adaptive integration time image sensor which has 128/spl times/64 pixels. We adopt a column parallel architecture to design the prototype chip. The scheme to control integration time is extended, and pixel pitch, processing speed and power consumption are much improved in comparison with our previous prototype. We show some experimental results obtained with the prototype.
Takayuki Hamamoto, Kiyoharu Aizawa
ICIP2
2000 Inverse Filters for Reconstruction of Arbitrarily Focused Images from two Differently Focused Images
abstract
This paper describes a novel filtering method to reconstruct an arbitrarily focused image from two differently focused images. Based on the assumption that image scene has two layers-foreground and background-, two differently focused images are used as inputs, one of which is focused on the foreground and the other is focused on the background. The linear equation that holds between these images and the desired image, which is derived from their imaging models, can be formulated as an image restoration problem. This paper shows that the solution of this problem exists as an inverse filter and the desired image can be reconstructed only by the linear filters. As a result, fast reconstruction with high accuracy can be achieved. Experiments using real images are shown.
Akira Kubota, Kiyoharu Aizawa
ICIP2
2000 Generation of a Disparity Panorama Using a 3-Camera Capturing System
abstract
This paper presents a method to create a panorama disparity map of a natural environment from stereo images captured by our original multi-purpose 3-camera panorama capturing system which features accurate frame synchronization between 3 channels and can be used outdoors through battery operation. For robust determination of correspondences of a stereo pair, we make use of a census transform, a kind of non-parametric local transform. The census transform summarizes local image structure as a bit string, and gives very good stereo-matching results without such processes as image pyramids, bi-directional search, complicated matching evaluation, and too much edge detection dependent interpolation. To interpolate unknown disparities, we introduced a process influenced by the K-means algorithm. Then, use of 3-channel images and overlapping of the maps give multiple disparity values for one pixel, which enables screening of very confident values. We have obtained very good results on such complex scenes as contain branches and leaves of trees.
Kunio Yamada, Tadashi Ichikawa, Takeshi Naemura, Kiyoharu Aizawa, Takahiro Saito
ICIP4
2000 Inverse filters for generation of arbitrarily focused images
Akira Kubota, Kiyoharu Aizawa
VCIP2
2000 High-quality stereo panorama generation using a three-camera system
Kunio Yamada, Tadashi Ichikawa, Takeshi Naemura, Kiyoharu Aizawa, Takahiro Saito
VCIP4
2000 Producing object-based special effects by fusing multiple differently focused images
abstract
We propose a novel approach for producing special visual effects by fusing multiple differently focused images. This method differs from conventional image fusion techniques because it enables us to arbitrarily generate object-based visual effects such as blurring, enhancement, and shifting. Notably, the method does not need any segmentation. Using a linear imaging model, it directly generates the desired image from multiple differently focused images.
Kiyoharu Aizawa, Kazuya Kodama, Akira Kubota
IEEE Trans. Circuits Syst. Video Technol.1
1999 Producing Object-Based Special Visual Effects by Integrating Multiple Differently Focused Images: Implicit 3D Approach to Image Content Manipulation
abstract
We propose a novel approach to image content manipulation. It enables us to arbitrarily manipulate an object in a scene by linear processings such as blurring, enhancement and shift etc. Notably, the method does not need any segmentation nor 3D modeling. Making use of multiple differently focused images and a linear imaging model, it directly generates the desired image from the original images. A special camera is developed which can acquire three differently focused image sequences, in order to extend the proposed method to image sequence processing.
Kiyoharu Aizawa, Kazuya Kodama, Akira Kubota
ICIP (2)1
1999 Digital Watermarking Using Inter-Block Correlation
abstract
In this paper, we propose a new watermark technique which utilizes the inter-block correlation of DCT coefficients. The amount of modification, that is the strength of embedded watermark, depends on the local feature of an image such that the watermark can be strongly embedded in the region where signal distortion is perceptually less visible. This feature also makes difficult for pirate to predict the position in which the watermark signal is embedded, since the amount of modification varies in the image. Embedded watermark can be detected without the parameters used in embedding process nor the original image. The experimental results show that this technique is robust to JPEG compression and Gaussian noise.
Yoonki Choi, Kiyoharu Aizawa
ICIP (2)2
1999 Real-Time Image Processing by Using Image Compression Sensor
abstract
We have been investigating an integration of sensing and compression on an image sensor. The compression sensor reduces the number of pixels in the image signal that has to be readout from the sensor. Therefore, the compression sensor can capture the images at the higher pixel rate which the traditional sensor can not handle. In this paper, we present a compression sensor which has 228/spl times/128 pixels. The processing circuits of the sensor can be operated at 5000 frames/second. We also describe the experimental results of two real-time image processing systems by using the compression sensor. They are a reconstruction circuit by using FPGA and a stereo image processing system for tracking of a moving object.
Takayuki Hamamoto, Ryutaro Oi, Yasuhiro Ohtsuka, Kiyoharu Aizawa
ICIP (3)4
1999 Multi-Media Ambiance Communication Based on Actual Moving Pictures
abstract
Multi-media ambiance communication refers to a means of shared-space communication, that makes use of actual moving pictures captured by video camera, combined with the laws of perspective, as used in painting and the visual characteristics of human beings, to establish a photo-realistic quality three-dimensional image space that users can naturally feel part of. We aim to enable this ambiance communication by basing the shared-space on actual moving pictures rather than on an accurate three-dimensional image space, such as is used in computer graphics.
Tadashi Ichikawa, Tetsuya Yoshimura, Kunio Yamada, Toshifumi Kanamaru, Hiromichi Suga, Shoichiro Iwasawa, Takeshi Naemura, Kiyoharu Aizawa, Shigeo Morishima, Takahiro Saito
ICIP (3)8
1999 Registration and Blur Estimation Methods for Multiple Differently Focused Images
abstract
In this paper, we propose a registration method between multiple differently focused images using the hierarchical block matching technique in which displacement, scale and rotation are taken into account. Local deformation due to lens distortion is further corrected by local matching. We also propose an efficient estimation method of blur parameters of defocused regions in these focused images. Simulation results showed that the proposed methods achieve high accuracy. In experiments using real images captured by hand-held camera, an all focused image with good quality was able to be automatically generated using the corrected images and the estimated parameters.
Akira Kubota, Kazuya Kodama, Kiyoharu Aizawa
ICIP (2)3
1999 Software Based Object Tracking with Visual Feature Integration
abstract
Multiple visual features can be used to track objects in environments where the use of a single visual feature is not sufficient. We have proposed an alternative method to feature integration by way of feature substitution where the dominant feature at any instant is used for tracking. The problem of increased computational cost resulting from the use of multiple features is addressed by using software optimization techniques. Here we use software threads and special instructions provided in MMX processors to increase the frame rate. Experimental results indicate that it is not always necessary to adopt a hardware based approach for real time object tracking even when the computational cost is high.
Ajith Pasqual, Hidenori Takeshima, Kiyoharu Aizawa
ICIP (2)3
1998 Motion Adaptive Image Sensor
abstract
We propose a motion adaptive sensor for image enhancement and wide dynamic range sensing. The motion adaptive sensor is able to control integration time pixel by pixel. The integration time is determined by saturation and temporal changes of incident light. It is expected to have high temporal resolution in the moving area, high SNR in the static area, and wide dynamic range. We have fabricated a prototype and show some results obtained by our experiments.
Takayuki Hamamoto, Kiyoharu Aizawa, Mitsutoshi Hatori
ASP-DAC2
1998 128 x 128 Pixels Image Sensor for On-Sensor-Compression
abstract
We have been investigating a novel integration of sensing and compression on an image sensor. By integration, the number of pixels in the image signal that has to be readout from the sensor can be significantly reduced, and the integration, can consequently increase the pixel rate of the sensor. We present a new compression sensor which has 128/spl times/128 pixels. We have made the prototype based on a column parallel architecture and improved the processing circuits of the new prototype to achieve lower power dissipation and higher processing speed in comparison with our previous prototypes. It is verified that the processing circuits can be operated at 5000 frames/second.
Takayuki Hamamoto, Yasuhiro Ohtsuka, Kiyoharu Aizawa
ICIP (1)3
1998 Spatially Variant Flexible Sampling Control Integrated on an Image Sensor
abstract
We propose a new sampling control system integrated on an image sensor. Contrary to the conventional random access pixels, the proposed sensor is able to read out spatially variant pixels at high speed, without inputting pixel address for each access. The sampling positions can be changed dynamically by rewriting the sampling position memory. Since the proposed sensor has an array memory that keeps the pixel position to be sampled. The sampling position can be dynamically changed by rewriting the memory array. It can achieve any spatially varying sampling patterns. We have made a first prototype and show results obtained by the prototype.
Yasuhiro Ohtsuka, Takayuki Hamamoto, Kiyoharu Aizawa, Mitsutoshi Hatori
ICIP (1)3
1997 Implementations of on Sensor Image Compression and Comparisons Between Pixel and Column Parallel Architectures
abstract
In order to enhance performance of an image sensor, we have been investigating a novel integration of compression and sensing. By this integration, the image signal that has to be readout from the sensor is significantly reduced. Thus, the integration can consequently leads to high pixel rate sensing. The compression scheme we make use of is conditional replenishment that detects and encodes moving areas. We have developed prototypes based on two different architectures, that are pixel parallel and column parallel architectures. We present the two prototypes and their comparisons, and show the results obtained by them.
Kiyoharu Aizawa, Takayuki Hamamoto, Yasuhiro Ohtsuka, Mitsutoshi Hatori, M. Abe
ICIP (2)1
1997 On sensor image compression
abstract
In this paper, we propose a novel image sensor which compresses image signals on the sensor plane. Since an image signal is compressed on the sensor plane by making use of the parallel nature of image signals, the amount of signal read out from the sensor can be significantly reduced. Thus, the potential applications of the proposed sensor are high pixel rate cameras and processing systems which require very high speed imaging or very high resolution imaging. The very high bandwidth is the fundamental limitation to the feasibility of those high pixel rate sensors and processing systems. Conditional replenishment is employed for the compression algorithm. In each pixel, current pixel value is compared to that in the last replenished frame. The value and the address of the pixel are extracted and coded if the magnitude of the difference is greater than a threshold. Analog circuits have been designed for processing in each pixel. A first prototype of a VLSI chip has been fabricated. Some results of experiments obtained by using the first prototype are shown in this paper.
Kiyoharu Aizawa, Hideo Ohno, Yuichiro Egi, Takayuki Hamamoto, Mitsutoshi Hatori, Hitoshi Maruyama, Junichi Yamazaki
IEEE Trans. Circuits Syst. Video Technol.1
1996 On sensor image compression for high pixel rate imaging: pixel parallel and column parallel architectures
abstract
We propose a novel concept of an integration of compression and sensing in order to enhance performance of an image sensor. By integrating the compression function on the sensor plane, the image signal that has to be read out from the sensor is significantly reduced. Thus, the integration can consequently increase the pixel rate of the sensor. The compression scheme we make use of is conditional replenishment that detects and encodes moving areas. In this paper, we discuss design and implementation of two architectures for on-sensor compression. One is a pixel parallel approach and the other is a column parallel approach. We describe both approaches and design and a prototype of pixel parallel architecture.
Kiyoharu Aizawa, Takayuki Hamamoto, Yuichiro Egi, Mitsutoshi Hatori, Junichi Yamazaki
ICIP (2)1
1996 Structural motion segmentation based on probabilistic clustering
abstract
In order to extract a meaningful scene structure from an image sequence, the global and local motion of moving objects are taken into consideration. Firstly, the image sequences are roughly separated into the regions of moving objects based on probabilistic clustering with mixture models using optical flow and the image intensity. For each moving object cluster, parametric motion estimation and segmentation can be obtained by iterative estimation of the affine motion parameters and region modification according to a criterion using the Gauss-Newton iterative optimization algorithm.
Cha-Keon Cheong, Kiyoharu Aizawa
ICIP (1)2
1996 Iterative reconstruction of an all-focused image by using multiple differently focused images
abstract
In this paper, we propose a novel method of all-focused image acquisition using multiple differently focused images. Based on the assumption that depth of the scene changes stepwise, we derive a formula for reconstruction between the desired all-focused image and multiple acquired images; we can reconstruct the all-focused image by iterative use of the formula. We also introduce coarse-to-fine estimation of point spread functions of the acquired images. We show we can reconstruct an all-focused image for a natural scene.
Kazuya Kodama, Kiyoharu Aizawa, Mitsutoshi Hatori
ICIP (3)2
1995 A multiple person eye contact (MPEC) teleconferencing system
abstract
A novel method of providing eye-contact for multiple participants in a teleconferencing system is presented. This method proposes a simple idea to fulfil the requirement of eye contact, for two or more local and remote participants simultaneously, by using half mirrors and cameras placed at the common points of the extended lines of gaze of each participant. In contrast to the commonly available one to one eye contact video phone systems, which also use half mirrors, the proposed system has the advantage of serving more than one person at a given location, hence named the multiple person eye contact (MPEC) teleconferencing system. We observed the resultant images by using a preliminary experiment and also building a prototype system. The results from both of these experiments proved the merits of the proposed idea.
Liyanage C. De Silva, Mitsuho Tahara, Kiyoharu Aizawa, Mitsutoshi Hatori
ICIP3
1995 Model-based image coding advanced video coding techniques for very low bit-rate applications
abstract
The paper gives an overview of model-based approaches applied to image coding, by looking at image source models. In these model-based schemes, which are different from the various conventional waveform coding methods, the 3-D properties of the scenes are taken into consideration. They can achieve very low bit rate image transmission. The 2-D model and 3-D model based approaches are explained. Among them, a 3-D model based method using a 3-D facial model and a 2-D model based method utilizing 2-D deformable triangular patches are described. Works related to 3-D model-based coding of facial images and some of the remaining problems are also described.>
Kiyoharu Aizawa, Thomas S. Huang
Proc. IEEE1
1995 A teleconferencing system capable of multiple person eye contact (MPEC) using half mirrors and cameras placed at common points of extended lines of gaze
abstract
In this paper a novel method of providing eye contact for multiple participants in a teleconferencing system is presented. This method proposes a simple idea to fulfill the requirement for eye contact, for two or more local and remote participants simultaneously, by using half mirrors and cameras placed at the common points of the extended lines of gaze of each participant. In contrast to the commonly available one to one eye contact of video phone systems, which also use half mirrors, the proposed system has the advantage of serving more than one person at a given location, hence named multiple person eye contact (MPEC) teleconferencing system. We observed the resultant images by using a preliminary experiment and also by building a prototype system. The results from both of these experiments proved the merits of the proposed idea. One of the main features of this system is that, it gives any participant, a feeling of "being looked at" if any remote participant is actually gazing at him. On the other hand if none of the remote participants are looking at him, then he gets the feeling of "not being looked at" which is also extremely important for effective inter personal communication. Furthermore, this system preserves the spatial continuity of neighboring participants, since all the participants at a given location are captured by each video camera.>
Liyanage C. De Silva, Mitsuho Tahara, Kiyoharu Aizawa, Mitsutoshi Hatori
IEEE Trans. Circuits Syst. Video Technol.3
1994 Motion estimation using multiple image sensors
abstract
A scheme of motion estimation using multiple image sensors is proposed, in which the multiple image sensors work at different integration intervals or different sampling instances. Because combined use of multiple image sensors virtually increases the temporal resolution while keeping signal to noise ratio high enough, the accuracy of the estimation can be improved. Both gradient-based and block-matching-based methods are formulated for the images obtained by the multiple image sensors. Tile experimental results show that the proposed scheme significantly improves the accuracy and reduces computational complexity.>
Kiyoharu Aizawa, Kenichi Iwata, Takahiro Saito, Mitsutoshi Hatori
ICASSP (5)1
1994 An image processing algorithm for a super high definition imaging scheme with multiple different-aperture cameras
abstract
Towards the development of a SHD (super high definition) image acquisition system, previously the authors developed the image-processing based approach with multiple cameras. Originally, in this approach, they used multiple cameras with the same pixel aperture, but in this case there needs to be severe limitations both in the arrangement of multiple cameras and in the configuration of the scene in order to guarantee the spatial uniformity of the resultant resolution. To overcome this difficulty completely, the authors have also previously presented the utilization of multiple cameras with different pixel apertures. The present paper develops a new, alternately iterative image processing algorithm available in the different aperture case. Experimental simulations clearly show that the alternately interactive algorithm behaves satisfactorily.>
Takahiro Saito, Takashi Komatsu, Kiyoharu Aizawa
ICASSP (5)3
1994 Fractionally spaced equalizers with adaptive sampling
abstract
A new structure of fractionally spaced equalizer (FSE) is described. A conventional FSE has taps spaced equally T/M, where T is the symbol spacing and M is an integer. Although there are many advantages of FSE in digital data transmission, the computational requirements increase with factor M. We describe a non-equally spaced FSE (NFSE) with less computational requirements. A novel tap control algorithm for NFSE under some restrictions is also presented.>
Miwa Sakai, Kiyoharu Aizawa, Mitsutoshi Hatori
ICASSP (3)2
1994 Use of steerable viewing window (SVW) to improve the visual sensation in face to face teleconferencing
abstract
We propose a method of using face direction information to make the local participant's video display of a video conferencing environment change according to his direction of view, in order to give him/her (the local participant) an improved visual sensation. We have named this method "steerable viewing window" (SVW). We also discuss an image processing method used to detect a person's face direction, without using special head mounted equipment. Further, we propose a method of reducing the data rate for transmission by coding the area around the gaze center with high resolution and the remaining with a lesser resolution. This method of saving bytes in video conferencing is named line of gaze centered coding (LOGCC).>
Liyanage C. De Silva, Kiyoharu Aizawa, Mitsutoshi Hatori
ICASSP (5)2
1994 A Novel Image Sensor for Video Compression
abstract
A novel image sensor on which video signals can be compressed is proposed. Since the video signal is compressed on an imager plane by using fast analog processing, the amount of image data read out from the imager can be significantly reduced with very small latency. The proposed system can be potentially applied to high pixel rate cameras such as those for high speed imaging and high resolution imaging. Conditional replenishment is employed for the video compression algorithm. Analog circuits are designed both for processing in each pixel and for controlling the entire data rate. The behavior of the circuit is investigated on the basis of both an analog circuit simulator and a scale-up-circuit. A VLSI chip is designed and is under fabrication.>
Kiyoharu Aizawa, Hideo Ohno, Takayuki Hamamoto, Mitsutoshi Hatori, Junichi Yamazaki
ICIP (3)1
1994 Two Approaches for Image-processing Based High Relolution Image Acquisition
abstract
Towards the development of a super high definition image acquisition system, we have proposed an image-processing based approach, i.e. the introduction of image processing techniques into an imaging process. Imaging methods based on this approach can be classified into two main categories: a spatial integration imaging method and a temporal integration imaging method. With regard to the spatial integration imaging method, we have presented previously a method for acquiring an improved-resolution image by integrating multiple images taken simultaneously with multiple cameras with different pixel apertures. In addition to the spatial integration imaging method, aimed at a particular surveillance application, we construct a temporal integration imaging method. Experimental simulations demonstrate that temporal integration imaging is as promising as spatial integration imaging for high resolution imaging.>
Yuji Nakazawa, Takahiro Saito, Takashi Komatsu, T. Sekimori, Kiyoharu Aizawa
ICIP (3)5
1994 Analysis and synthesis of facial image sequences in model-based image coding
abstract
This paper proposes new methods for analyzing image sequences and updating textures of the three-dimensional (3-D) facial model. It also describes a method for synthesizing various facial expressions. These three methods are the key technologies for the model-based image coding system. The input image analysis technique directly and robustly estimates the 3-D head motions and the facial expressions without any two-dimensional (2-D) entity correspondences. This technique resolves the 2-D correspondence mismatch errors and provides quality reproduction of the original images by fully incorporating the synthesis rules. To verify the analysis algorithm, the paper performs quantitative and subjective evaluations. It presents two methods for updating the texture of the facial model to improve the quality of the synthesized images. The first method focuses on the facial parts with large change of brightness according to the various facial expressions for reducing the transmission bit rates. The second method focuses on all changes of brightness caused by the 3-D head motions as well as the facial expressions. The transmission bit rates are estimated according to the update methods. For synthesizing the output images, it describes rules that simulate the facial muscular actions because the muscles cause the facial expressions. These rules more easily synthesize the high-quality facial images that represent the various facial expressions.>
Chang Seok Choi, Kiyoharu Aizawa, Hiroshi Harashima, Tsuyoshi Takebe
IEEE Trans. Circuits Syst. Video Technol.2
1994 Estimation of camera parameters from image sequence for model-based video coding
abstract
The authors describe a method for estimating camera parameters from image sequences for application to emerging model-based video coding systems. The parameters to be estimated include focal length, zoom, and 3-D rotation parameters. The method consists of first establishing a correspondence and then, estimating the parameters by fitting the correspondence data to a transformation model based on a perspective mapping model and a 3-D rotation and zoom operation model. They show by simulations and experiments that the proposed method successfully estimates the focal length from the image sequences, it explains very well the induced motion field of images undergoing camera operation (3-D rotation and zoom), and that it significantly outperforms conventional estimation methods, especially for wide-angled images. It is anticipated that the proposed method will be successfully applied to compensating for the motion field induced by camera operation in extracting a 3-D object model and a 3-D object motion, and to synchronizing the viewing direction and scale of an image, in model-based video coding technology.>
Jong-Il Park, Nobuyuki Yagi, Kazumasa Enami, Kiyoharu Aizawa, Mitsutoshi Hatori
IEEE Trans. Circuits Syst. Video Technol.4
1993 Subpixel registration for a high resolution imaging scheme using multiple imagers
Kiyoharu Aizawa, Takashi Komatsu, Takahiro Saito, Mitsutoshi Hatori
ICASSP (5)1
1993 Motion estimation with wavelet transform and the application to motion compensated interpolation
Cha Keon Cheong, Kiyoharu Aizawa, Takahiro Saito, Mitsutoshi Hatori
ICASSP (5)2
1993 Very high resolution imaging scheme with multiple different-aperture cameras
Takashi Komatsu, Toru Igarashi, Kiyoharu Aizawa, Takahiro Saito
Signal Process. Image Commun.3
1992 A scheme for acquiring very high resolution images using multiple cameras
abstract
A signal processing-based scheme for acquiring high-resolution pictures with sufficiently high signal-to-noise ratio (SNR) by processing multiple charge-coupled device (CCD) images is presented. The method integrates multiple low-resolution images into a high-resolution image. Compared to a CCD imaging device with equivalent resolution, it may be less sensitive to shot noise which becomes more dominant as the pixel size of an imager is further reduced. Preliminary experiments using a camera model have shown clear improvements in the high frequencies and image details. A strategy to further improve resolution is also described.>
Kiyoharu Aizawa, Takashi Komatsu, Takahiro Saito
ICASSP1
1991 VIWOB: an interactive programming environment for distributed simulations
abstract
VIWOB is an integrated environment for programming and executing simulations interactively on distributed computational resources. A visual dataflow oriented language allows a user to assemble application programs by simply connecting processing nodes. Single nodes can be described by subgraphs (modules) permitting the implementation of complex programs in a hierarchical structured manner, as well as taking advantage of previously developed modules, while hierarchical ordered libraries provide convenient access to the available modules. The run-time environment will break up the application and execute the components transparently on a distributed computing facility. An abstract computational model simplifies adding new primitive modules.>
M. Ott, Liyanage C. De Silva, L. Satou, Kiyoharu Aizawa, Mitsutoshi Hatori
ICASSP4
1990 Real-time facial action image synthesis system driven by speech and text
abstract
Automatic facial motion image synthesis schemes and a real-time system design are presented. The purpose of this schemes is to realize an intelligent human-machine interface or intelligent communication system with talking head images. Human's face is reconstructed with 3D surface model and texture mapping technique on the display of terminal. Facial motion images are synthesized naturally by transformation of the lattice points on wire frames. Two types of motion drive methods, text to image conversion and speech to image conversion are proposed in this paper. In the former manner, synthesized head can speak some given texts naturally and in the latter case, some mouth and jaw motions can be synthesized in time to speech signal of behind speaker. These schemes were implemented to a parallel image computer and a real-time image synthesizer could output facial motion images to the display as fast as video rate.
Shigeo Morishima, Kiyoharu Aizawa, Hiroshi Harashima
VCIP2
1989 An intelligent facial image coding driven by speech and phoneme
abstract
The authors propose and compare two types of model-based facial motion coding schemes, i.e. synthesis by rules and synthesis by parameters. In synthesis by rules, facial motion images are synthesized on the basis of rules extracted by analysis of training image samples that include all of the phonemes and coarticulation. This system can be utilized as an automatic facial animation synthesizer from text input or as a man-machine interface using the facial motion image. In synthesis by parameters, facial motion images are synthesized on the basis of a code word index of speech parameters. Experimental results indicate good performance for both systems, which can create natural facial-motion images with very low transmission rate. Details of 3-D modeling, algorithm synthesis, and performance are discussed.>
Shigeo Morishima, Kiyoharu Aizawa, Hiroshi Harashima
ICASSP2
1989 Model-based analysis synthesis image coding (MBASIC) system for a person's face
Kiyoharu Aizawa, Hiroshi Harashima, Takahiro Saito
Signal Process. Image Commun.1
1986 Adaptive discrete cosine transform coding with vector quantization for color images
abstract
A new vector quantization scheme in discrete cosine transform domain, named DCT-VQ and its application to color image coding are described. In this scheme, DCT-domain is partitioned into vectors which are normalized and vector-quantized using universal vector quantizers designed with multidimensional Laplacian distribution. Adaptive coding scheme is also introduced to obtain better reconstruction of images. The color image coder employs the above scheme and encodes separately three components converted from R,G,B signals. The simulations have shown that adaptive DCT-VQ exhibits better performance than a conventional adaptive cosine transform coding with scalar quantization. The decomposition of DCT-block into vectors results in much less complex coder than a vector quantizer in original space domain.
Kiyoharu Aizawa, Hiroshi Harashima, Hiroshi Miyakawa
ICASSP1
1986 Adaptive discrete cosine transform image coding using gain/Shape vector quantizers
abstract
An efficient discrete cosine transform image coding system using the gain/shape vector quantizers (DCT-G/S VQ) is presented. In the coding system, AC transform coefficients in a subblock are partitoned into several bands according to the Schaming's method, and the normalized AC transform coefficients of each band are quantized with the gain/shape vector quantizer designed on a spherically symmetric probability model. In addition, an adaptive DCT-G/S VQ (A-DCT-G/S VQ) is presented by incorporating a modification of the recursive quantization technique in the DCT-G/S VQ. The coding systems are simulated on color images, and their performance is compared to that of previously reported discrete cosine transform coding systems using the Max-type scalor quantizers.
Takahiro Saito, Hideya Takeo, Kiyoharu Aizawa, Hiroshi Harashima, Hiroshi Miyakawa
ICASSP3