EDBT 2026 Demo / reviewers in the wild / expert
Noboru Babaguchi
dblp:43/2690
· DBLP profile ↗
109ranked-venue papers
13as first author
8since 2021 · last 2025
0000-0003-3320-896XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 81 · 10 first-author · 3 since 2021Artificial intelligence and machine learning · 31 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 9Security and privacy · 8 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Computer networks · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Paladin: Understanding Video Intentions in Political Advertisement VideosabstractIn this paper, we introduce a novel taskfor video understanding that focuses on detecting editing intentions in political advertisement videos. Political advertisement videos are edited with some intentions (e.g., “associating some candidates with negative emotions”) of making people unthinkingly believe the messages in the videos, potentially ending up with some irrational bias. Detecting such intentions is thus the primary step toward fairer decision-making based on the messages themselves. To this end, we classify such editing intentions into 10 categories (referred to as communication techniques) in consultation with a professional editor as well as based on communication techniques presented in the natural language processing community, and build a dataset of 12,526 political advertisement videos, each of which is annotated with several communication technique segments. We also explore the capability of existing video understanding models in detecting editing intentions over the dataset, which identifies new dimensions of challenges to be addressed. Hong Liu 0009, Yuta Nakashima, Noboru Babaguchi |
WACV | 3 |
| 2025 | Cross-Modal Guided Visual Representation Learning for Social Image RetrievalabstractSocial images are often associated with rich but noisy tags from community contributions. Although social tags can potentially provide valuable semantic training information for image retrieval, existing studies all fail to effectively filter noises by exploiting the cross-modal correlation between image content and tags. The current cross-modal vision-and-language representation learning methods, which selectively attend to the relevant parts of the image and text, show a promising direction. However, they are not suitable for social image retrieval since: (1) they deal with natural text sequences where the relationships between words can be easily captured by language models for cross-modal relevance estimation, while the tags are isolated and noisy; (2) they take (image, text) pair as input, and consequently cannot be employed directly for unimodal social image retrieval. This paper tackles the challenge of utilizing cross-modal interactions to learn precise representations for unimodal retrieval. The proposed framework, dubbed CGVR (Cross-modal Guided Visual Representation), extracts accurate semantic representations of images from noisy tags and transfers this ability to image-only hashing subnetwork by a carefully designed training scheme. To well capture correlated semantics and filter noises, it embeds a priori common-sense relationship among tags into attention computation for joint awareness of textual and visual context. Experiments show that CGVR achieves approximately 8.82 and 5.45 points improvement in MAP over the state-of-the-art on two widely used social image benchmarks. CGVR can serve as a new baseline for the image retrieval community. The code is provided at https://github.com/zhaowanqing/CGVR. Ziyu Guan, Wanqing Zhao, Hongmin Liu 0001, Yuta Nakashima, Noboru Babaguchi, Xiaofei He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Enhancing Fake News Detection in Social Media via Label Propagation on Cross-modal Tweet GraphabstractFake news detection in social media has become increasingly important due to the rapid proliferation of personal media channels and the consequential dissemination of misleading information. Existing methods, which primarily rely on multimodal features and graph-based techniques, have shown promising performance in detecting fake news. However, they still face a limitation, i.e., sparsity in graph connections, which hinders capturing possible interactions among tweets. This challenge has motivated us to explore a novel method that densifies the graph's connectivity to capture denser interaction better. Our method constructs a cross-modal tweet graph using CLIP, which encodes images and text into a unified space, allowing us to extract potential connections based on similarities in text and images. We then design a Feature Contextualization Network with Label Propagation (FCN-LP) to model the interaction among tweets as well as positive or negative correlations between predicted labels of connected tweets. The propagated labels from the graph are weighted and aggregated for the final detection. To enhance the model's generalization ability to unseen events, we introduce a domain generalization loss that ensures consistent features between tweets on seen and unseen events. We use three publicly available fake news datasets, Twitter, PHEME, and Weibo, for evaluation. Our method consistently improves the performance over the state-of-the-art methods on all benchmark datasets and effectively demonstrates its aptitude for generalizing fake news detection in social media. Wanqing Zhao, Yuta Nakashima, Haiyuan Chen, Noboru Babaguchi |
ACM Multimedia | 4 |
| 2023 | Social IoT Approach to Cyber Defense of a Deep-Learning-Based Recognition System in Front of Media Clones Generated by Model Inversion AttackabstractModel inversion attack (MIA) is a cyber threat with an increasing alert even for deep-learning-based recognition systems (DLRSs). By targeting a DLRS under a scenario of attacker access to the model structure and parameters, MIA generates a data clone for a certain targeted class label. To avoid the possible threats of such MIA-generated data clones, this research work proposes a social IoT approach to a collaborative cyber-defense among the online recognition systems (RSs) sharing the targeted class label. Since, the generation of an MIA-clone is by targeting an RS model and using its structure, parameters, and class labels output scores in an iterative optimization process, the generated clone is partially inherent to the targeted model. Thus, it is expected for an MIA-clone to show a different performance on a secondary RS wherein the same targeted class label is included. It is because, in the MIA generation of the clone, not only the targeted class label but also other class labels, and model parameters and structure affect the process, while the second model has just the targeted class label in common with the target model. Deploying the Social Internet of Recognition Systems (SIoRS), the proposed technique utilizes a collaborative recognition by SIoRC which plays the role of a complementary recognition besides the targeted RS. The recognition output by the targeted RS is further verified by the SIoRS complementary recognition result. To avoid the MIA-targeted data clones, the verification of recognition is by the log-likelihood ratio test between the targeted RS and the SIoRS complementary recognition confidence scores. The proposed technique is evaluated by statistical analysis on deep face RSs in 10000 Monte Carlo runs for each of the conventional, dc-generative adversarial network (GAN) and$\alpha $-GAN integrated MIA techniques in targeting two different user identities. The$Z$scores of the fitted normal distribution of the log-likelihood ratios indicate almost 100% detection rate of clones generated by conventional MIA and 95.23% and 86% of clones, respectively, generated by DC-GAN and$\alpha $-GAN integrated deep MIA techniques. Mahdi Khosravy, Kazuaki Nakamura, Naoko Nitta, Nilanjan Dey, Rubén González Crespo, Enrique Herrera-Viedma, Noboru Babaguchi |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2022 | Anonymization of Human Gait in Video Based on Silhouette Deformation and Texture TransferabstractThese days, a lot of videos are uploaded onto web-based video sharing services such as YouTube. These videos can be freely accessed from all over the world. On the other hand, they often contain the appearance of walking private people, which could be identified by silhouette-based gait recognition techniques rapidly developed in recent years. This causes a serious privacy issue. To avoid it, this paper proposes a method for anonymizing the appearance of walking people, namely human gait, in video. In the proposed method, we first crop human regions from all frames in an input video and binarize them to get their silhouettes. Next, we slightly deform the silhouettes from the aspects of static body shape and dynamic walking rhythm so that the person in the input video cannot be correctly identified by gait recognition techniques. After that, the textures of the original human regions are transferred onto the deformed silhouettes. We achieve this by a displacement field-based approach, which is training-free and thus robust to a variety of clothes. Finally, the anonymized human regions with the transferred textures are filled back into the input video. In the results of our experiments, we successfully degraded the accuracy of CNN-based gait recognition systems from 100% to 1.57% in the lowest case without yielding serious distortion in the appearance of the human regions, which demonstrated the effectiveness of the proposed method. Yuki Hirose, Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Model Inversion Attack by Integration of Deep Generative Models: Privacy-Sensitive Face Generation From a Face Recognition SystemabstractCybersecurity in front of attacks to a face recognition system is an emerging issue in the cloud era, especially due to its strong bonds with the privacy of the users registered to the system. A possible attack is the model inversion attack (MIA) which aims to reveal the identity of a targeted user by generating the most proper datapoint input to the system with maximum corresponding confidence score at the output. The generated data of a registered user can be maliciously used as a serious invasion of the user privacy. In literature, MIA processes are categorized into white-box and black-box scenarios which are respectively with and without information about the system structure, parameters, and partially about the users. This research work assumes the MIA under semi-white box scenario of availability of system model structure and parameters but not any user data information, and verifies it as a severe threat even for a deep-learning-based face recognition system despite its complex structure and the diversity of registered user data. The alert state is promoted by Deep MIA which is the integration of deep generative models in MIA, and$\alpha $-GAN integrated MIA-initilized by a face based seed ($\alpha $-GAN-MIA-FS) is proposed. As a novel MIA search strategy, a pre-trained deep generative model with capability of generating a face image from a random feature vector is used for narrowing down the image search space to the feature vectors space, which has much lower dimensions. This allows the MIA process to efficiently search for a low-dimensional feature vector whose corresponding face image maximizes the confidence score. We have experimentally evaluated the proposed method by two objective criteria and three subjective criteria in comparison to$\alpha $-GAN-integrated MIA initialized with a random seed ($\alpha $-GAN-MIA-RS), DCGAN-integrated MIA (DCGAN-MIA), and the conventional MIA. The evaluation results approve the efficiency and superiority of the proposed technique in generating natural looking face clones with high recognizability as the targeted users. Mahdi Khosravy, Kazuaki Nakamura, Yuki Hirose, Naoko Nitta, Noboru Babaguchi |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | Effective De-identification Generative Adversarial Network for Face AnonymizationabstractThe growing application of face images and modern AI technology has raised another important concern in privacy protection. In many real scenarios like scientific research, social sharing and commercial application, lots of images are released without privacy processing to protect people's identity. In this paper, we develop a novel effective de-identification generative adversarial network (DeIdGAN) for face anonymization by seamlessly replacing a given face image with a different synthesized yet realistic one. Our approach consists of two steps. First, we anonymize the input face to obfuscate its original identity. Then, we use our designed de-identification generator to synthesize an anonymized face. During the training process, we leverage a pair of identity-adversarial discriminators to explicitly constrain identity protection by pushing the synthesized face away from the predefined sensitive faces to resist re-identification and identity invasion. Finally, we validate the effectiveness of our approach on public datasets. Compared with existing methods, our approach can not only achieve better identity protection rates but also preserve superior image quality and data reusability, which suggests the state-of-the-art performance. Zhenzhong Kuang, Huigui Liu, Jun Yu 0002, Aikui Tian, Jianping Fan 0001, Noboru Babaguchi |
ACM Multimedia | 7 |
| 2021 | Unnoticeable synthetic face replacement for image privacy protection
Zhenzhong Kuang, Zhiqiang Guo, Jinglong Fang, Jun Yu 0002, Noboru Babaguchi, Jianping Fan 0001 |
Neurocomputing | 5 |
| 2020 | Semi-Supervised Outdoor Image Generation Conditioned on Weather SignalsabstractIn recent years, various types of sensors observe the real world. Especially, weather sensors are densely installed all over the world to observe current weather situations at various places. However, weather signals such as the temperature or humidity obtained by weather sensors are intuitively difficult for humans to understand. On the other hand, images captured by typical RGB cameras can tell weather situations at the captured places in a more comprehensible way for humans; however, cameras are only installed at limited places and are not necessarily open to public due to privacy issues. In order to solve this problem, the goal of our work is to generate images which can tell weather situations at arbitrary time and locations. This can be realized by using a conditional generative adversarial network architecture that takes an image and a condition to transform the image accordingly to the condition. Training such image generator requires a large number of image and condition pairs as the training data. Although weather signals can be easily collected from weather sensors, collecting their spatially and temporally synchronized outdoor images is not easy. Thus, we propose a semi-supervised method for training the image generator. A relatively small number of pairs of an outdoor image and weather signals is collected, each from different web services, by considering their semantic consistency. The collected pairs are used to train a predictor for predicting weather signals from a given outdoor image. Then, the image generator is trained by using a large number of pairs of an outdoor image and pseudo weather signals predicted by the predictor as the training data. Sota Kawakami, Kei Okada, Naoko Nitta, Kazuaki Nakamura, Noboru Babaguchi |
ICPR | 5 |
| 2020 | Data Anonymization for Service Strategy Development and Information Recommendation to Users Based on TF-IDF Method
Kazuhiro Kono, Noboru Babaguchi |
ISITA | 2 |
| 2020 | Probabilistic Stone's Blind Source Separation with application to channel estimation and multi-node identification in MIMO IoT green communication and multimedia systems
Mahdi Khosravy, Nilesh Patel, Nilanjan Dey, Naoko Nitta, Noboru Babaguchi |
Comput. Commun. | 6 |
| 2019 | Training-Free Method for Generating Motion Video Clones From A Still Image Considering Self-Occlusion of Human BodyabstractIn this paper, we propose a method for generating photo-realistic video in which a person virtually performs a motion that is not performed in the real world. we refer to such video as motion video clones (MVCs). The proposed method requires only two kinds of source information as input data: a reference video, in which a person A performs some motion, and a target image, which includes the whole body of another person B. Using these data, our method generates a MVC in which the person A's motion is re-enacted by the person B's body. Since our method does not need 3D human body model nor any training phase, it is suitable to MVC-based entertainment systems. To handle the self-occlusion of the human body in the reference video, we employ a part-based approach. For each part such as the right arm and the left leg, we first extract its skeleton from the target image and move it so that the motion represented by the reference video is reenacted. Next, we compute a 2D affine transform between the original and moved positions of the skeleton. This transform is used to map the texture of the target image onto each frame of a resultant MVC. Finally, we extend the part-wise affine transforms to pixel-wise ones by computing their linear combination for each pixel, whose combination weights are computed according to the geodesic distance between the pixel and the center of each part. This allows us to avoid unnatural appearance around the joints of the body parts. In our experiments, the proposed method generated much visually-natural MVCs than existing methods. Teppei Tsutsumi, Kazuaki Nakamura, Seiko Myojin, Naoko Nitta, Noboru Babaguchi |
ICIP | 5 |
| 2019 | Encryption-Free Framework of Privacy-Preserving Image Recognition for Photo-Based Information ServicesabstractNowadays, mobile devices, such as smartphones, have been widely used all over the world. In addition, the performance of image recognition has drastically increased with deep learning technologies. From these backgrounds, some photo-based information services provided in a client-server architecture are getting popular: client users take a photo of a certain spot and send it to a server, while the server identifies the spot with an image recognizer and returns its related information to the users. However, this kind of client-server image recognition can cause a privacy issue because image recognition results are sometimes privacy-sensitive. To tackle the privacy issue, in this paper, we propose a framework of privacy-preserving image recognition called EnfPire, in which the server cannot uniquely determine the recognition result but client users can do so. An overview of EnfPire is as follows. First, client users extract a visual feature from their taken photo and transform it so that the server cannot uniquely determine the recognition result. Then, the users send the transformed feature to the server that returns a set of candidates of the recognition result to the users. Finally, the users compare the candidates to the original visual feature for obtaining the final result. Our experimental results demonstrate that EnfPire successfully degrades the server's spot-recognition accuracy from 99.8% to 41.4% while keeping 86.9% of the spot-recognition accuracy on the user side. Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Generating Handwritten Character Clones from an Incomplete Seed Character Set using Collaborative FilteringabstractIn this paper, we propose a method for generating clones of a target writer's handwritten character images (called handwritten character clones or HCCs) using an incomplete seed character set, which consists of at most one or no example of his/her actual handwriting per character. In HCC generation, not a single HCC but its distribution should be created for each character because humans' actual handwriting images differ from each other even if the same writer writes the same character. However, it is difficult to achieve this from the incomplete seed character set. To solve the problem, in the proposed method, we first create a number of HCC distributions for each character by clustering a set of handwritten character images offered by other writers. Next, for each character contained in the seed character set, we choose the distribution best fit to its example. Finally, for the other characters, we estimate the best distribution for them employing collaborative filtering. We conducted pilot experiments focusing on Japanese character images, in which the proposed method successfully generated various HCCs with a certain level of quality for each character. Kazuaki Nakamura, Eiji Miyazaki, Naoko Nitta, Noboru Babaguchi |
ICFHR | 4 |
| 2017 | A Framework of Privacy-Preserving Image Recognition for Image-Based Information Services
Kojiro Fujii, Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
MMM (1) | 4 |
| 2017 | Effect of Junk Images on Inter-concept Distance Measurement: Positive or Negative?
Yusuke Nagasawa, Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
MMM (2) | 4 |
| 2016 | Detection of groups in crowd considering their activity stateabstractIn this paper, we focus on the problem of group detection in crowd, which is a task of partitioning a set of pedestrians in a scene into small subsets called groups based on their trajectories. Most of previous methods use only a single model for representing a relationship between trajectories of pedestrians who belong to the same group. However, such relationship would vary depending on the activity state (e.g. walking together, approaching, splitting, and so on) of the group. In this paper, we propose a novel group detection method which can cope with a variation of groups' activity state. The proposed method constructs different models for each activity state in order to appropriately evaluate the relationship of pedestrians' trajectories. In addition, our method regards groups' activity state as hidden variables and estimates their probability distributions, which is used for integrating the constructed models. The proposed method outperforms existing methods in the experiment on the public dataset. Kazuaki Nakamura, Tsukasa Ono, Noboru Babaguchi |
ICPR | 3 |
| 2015 | Temporal spotting of human actions from videos containing actor's unintentional motionsabstractThis paper proposes a method for temporal action spotting: the temporal segmentation and classification of human actions in videos. Naturally performed human actions often involve actor's unintentional motions. These unintentional motions yield false visual evidences in the videos, which are not related to the performed actions and degrade the performance of temporal action spotting. To deal with this problem, our proposed method empolys a voting-based approach in which the temporal relation between each action and its visual evidence is probabilistically modeled as a voting score function. Due to the approach, our method can robustly spot the target actions even when the actions involve several unintentional motions, because the effect of the false visual evidences yielded by the unintentional motions can be canceled by other visual evidences observed with the target actions. Experimental results showed that the proposed method is highly robust to the unintentional motions. Keita Hara, Kazuaki Nakamura, Noboru Babaguchi |
ICME | 3 |
| 2015 | Facial expression preserving privacy protection using image meldingabstractAn enormous number of images are currently shared through social networking services such as Facebook. These images usually contain appearance of people and may violate the people's privacy if they are published without permission from each person. To remedy this privacy concern, visual privacy protection, such as blurring, is applied to facial regions of people without permission. However, in addition to image quality degradation, this may spoil the context of the image: If some people are filtered while the others are not, missing facial expression makes comprehension of the image difficult. This paper proposes an image melding-based method that modifies facial regions in a visually unintrusive way with preserving facial expression. Our experimental results demonstrated that the proposed method can retain facial expression while protecting privacy. Yuta Nakashima, Tatsuya Koyama, Naokazu Yokoya, Noboru Babaguchi |
ICME | 4 |
| 2015 | Real-Time People Counting across Spatially Adjacent Non-overlapping Camera Views
Ryota Akai, Naoko Nitta, Noboru Babaguchi |
MMM (1) | 3 |
| 2014 | Real-World Event Detection Using Flickr Images
Naoko Nitta, Yusuke Kumihashi, Tomochika Kato, Noboru Babaguchi |
MMM (2) | 4 |
| 2013 | People counting across spatially disjoint cameras by flow estimation between foreground regionsabstractOur goal is to develop a method for counting the number of people traveling over a wide area monitored by spatially disjoint multiple cameras with non-overlapping fields of view. The proposed method counts the number of people traversing across each pair of cameras' fields of view by estimating the flows between the foreground regions which have disappeared from and appeared in the camera views within a short time interval. The approach aims at resolving two problems: people in a foreground region can split and merge outside the cameras' fields of view, and the appearance variance of the same person and persons with similar appearance can lead to errors in person re-identification across cameras. The average errors of 0.184 persons in counting the number of people traveling over an area in a university campus monitored by four virtual cameras for 100 minutes have demonstrated the effectiveness of the proposed method. Naoko Nitta, Takayuki Nakazaki, Kazuaki Nakamura, Ryota Akai, Noboru Babaguchi |
AVSS | 5 |
| 2013 | Efficient DC term encoding scheme based on double prediction algorithms and Pareto probability modelsabstractIn this paper, a new algorithm which adopts the techniques of double prediction and the Pareto probability model was applied to encode the DC term in the JPEG compression process. Conventionally, the DC term was encoded by differential coding, i.e., the difference of the DC values between the current block and the previous block. In this paper, we first use the DC terms of four adjacent blocks to predict the current DC value. We then further use the prediction error of the four adjacent blocks to estimate the variance of the prediction error of the current block. We call it the double prediction algorithm. Next, the Pareto distribution is applied to model the probability distribution of the prediction error. Simulation results show that, with the proposed algorithms, the data size required for DC terms is significantly reduced by 25% ~ 60% and a much higher compression rate can be achieved. Ting-Yu Ko, Chi-Jung Tseng, Hsin-Hui Chen, Jian-Jiun Ding, Noboru Babaguchi |
ICME | 5 |
| 2013 | Real-time privacy protection system for social videos using intentionally-captured persons detectionabstractMost social videos, which are uploaded and shared through social networking services (SNSs), e.g., YouTube and Facebook, contain not only intentionally-captured persons (ICPs) but also non-ICPs who are unexpectedly framed in, such as passers-by. Sharing such social videos may infringe on the non-ICPs' privacy but not on the ICPs' in many cases; however, existing systems for video privacy protection simply obscure persons without distinguishing ICPs from non-ICPs. This naive obscuration may spoil the videos. Since this is a critical problem especially for social videos, in this paper, we propose a novel system for automatically generating privacy-protected videos in real-time. Our system localizes ICPs and non-ICPs using ICP detection leveraging the spatial and temporal consistency of ICPs/non-ICPs and obscures the non-ICPs. We have experimentally evaluated the performance of ICP detection and demonstrated the applicability of our system. Tatsuya Koyama, Yuta Nakashima, Noboru Babaguchi |
ICME | 3 |
| 2013 | Guest Editorial: Special issue on intelligent video surveillance for public security and personal privacyabstractThis Special Issue offers an overview of ongoing research on intelligent video surveillance (IVS) techniques, and brings together cutting-edge research work on security and privacy problems with respect to technological, behavioral, legal, and cultural aspects. We received 34 submissions and each submission was rigorously reviewed by at least two experts in the related fields based on the criteria of originality, significance, quality, and clarity. Eventually, 12 papers were accepted for the Special Issue, spanning a variety of topics including privacy protection, background modeling, tracking, action/activity analysis, and crowd behavior perception. The papers constituting this issue are then briefly summarized. Noboru Babaguchi, Andrea Cavallaro, Rama Chellappa, Frédéric Dufaux, Liang Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Depth-Estimation-Free Condition for Projective Factorization and Its Application to 3D Reconstruction
Yohei Murakami, Takeshi Endo, Yoshimichi Ito, Noboru Babaguchi |
ACCV (4) | 4 |
| 2012 | Abandoned Object's Owner Detection: A Case Study of Hybrid Mobile-Fixed Video Surveillance SystemabstractIn this paper, a new framework of hybrid mobile-fixed video surveillance system (HMFVSS) is introduced. The purpose of this framework is to overcome common problems of existing mobile or fixed video surveillance systems: (1) moral harassment: due to unfriendly or unnaturally installed mobile sensors, and (2) blind areas: due to narrow-scope moving of fixed cameras. A case study of abandoned object's owner alert system (AOOAS) is also presented to emphasize the framework's advantages. IP cameras and "Spyglass" (i.e. a mobile camera embedded on glasses) are used as fixed and mobile sensors, respectively. There are three main tasks are inherited, developed, and integrated: (1) image registration for automatically locating abandoned object, (2) common histogram based abandoned object's owner detection, and (3) faces recognition. The experimental results with careful evaluation and comparison with others shows that the proposed framework moves a step ahead in video surveillance system. Minh-Son Dao, Riccardo Mattivi, Francesco G. B. De Natale, Keita Masui, Noboru Babaguchi |
AVSS | 5 |
| 2012 | Delivery method for viewer-specific privacy protected video using discrete wavelet transformabstractThis paper presents a delivery method for viewer-specific privacy protected video, where “viewer-specific” implies that the way of privacy protection (e.g., box, mosaic, transparency) varies according to viewer's authority. In this method, in order to reduce the load of a delivery server, viewer-specific privacy protected videos are not produced at the delivery server, but they are produced at each viewer's terminal. The delivery server extracts and decomposes the information of human objects, and then embeds the information into background images using information hiding technique. Each viewer extracts a part of the decomposed information of human objects according to his/her authority, and then produces privacy protected video by integrating the extracted information. In this method, the discrete wavelet transform (DWT) plays a key role. The proposed method is evaluated through experiments. Naoya Fukuoka, Yoshimichi Ito, Noboru Babaguchi |
ICIP | 3 |
| 2012 | Markov random field-based real-time detection of intentionally-captured personsabstractMost videos taken by videographers contain intentionally-captured persons (ICPs), who are essential for what the videographers want to express in their video. This paper presents a method to detect ICPs in real-time. Whether a person in a video is an ICP or not is reflected in features such as the person's motion and camera motion, which are thus beneficial for detecting ICPs. However, estimating camera motion is computationally expensive. For real-time detection, we use samples of acceleration and angular velocity obtained from inertial sensors instead of estimating camera motion. Considering that pairwise constraints based on differences between persons' sizes also improve the detection performance, we model the ICPs using Markov random field. We experimentally evaluate the performance of our method and demonstrate that it works in real-time. Tatsuya Koyama, Yuta Nakashima, Noboru Babaguchi |
ICIP | 3 |
| 2012 | Tablet owner authentication based on behavioral characteristics of multi-touch actions
Kumi Nakamura, Yoshimichi Ito, Kazuhiro Kono, Noboru Babaguchi |
ICPR | 4 |
| 2012 | Classification based group photo retrieval with bag of people featuresabstractThis paper proposes a method for retrieving images containing a specific target person from a given image collection of group photos. This can be realized by query-by-example methods which compare the facial visual features of the target person in the given query image and of each person in the images in the image collection. However, since images are often taken under various conditions, facial appearance of the same person can vary. Since socially related people such as family and friends are often taken photos together, the people co-occurrence relations in the same images can also be a useful clue for image retrieval. Focusing on such people co-occurrence relations, we propose Bag of People (BoP) features which represent both the facial appearances of persons and their co-occurrence relations in the same images. By using the BoP features, a classifier for classifying images into two classes, images containing the target person and other images, can be trained from a small number of images labeled by user's relevance feedback. Furthermore, since the labeled images obtained by relevance feedback are much fewer than unlabeled images in the image collection, an active learning method is used to select useful images to train the classifier. When retrieving images of 24 persons in total from 550 images, after five feedback iterations, the mean average precision of 0.94 was obtained by considering the people co-occurrence relations, as against 0.69 when considering only the target person. Kazuya Shimizu, Naoko Nitta, Yujiro Nakai, Noboru Babaguchi |
ICMR | 4 |
| 2012 | Intended human object detection for automatically protecting privacy in mobile video surveillance
Yuta Nakashima, Noboru Babaguchi, Jianping Fan 0001 |
Multim. Syst. | 2 |
| 2011 | Automatic generation of privacy-protected videos using background estimationabstractRecently, video sharing services such as YouTube and Daily-motion have become popular and many videos taken with mobile video cameras are uploaded to such a video sharing service. However, such videos can infringe on the privacy right of people in the videos because they may contain privacy sensitive information (PSI) of the people, i.e., their appearances. This strongly motivates us to develop a technique to generate privacy-protected videos. In this paper, we propose a novel system for automatic generation of privacy-protected videos based on background estimation. In most conventional techniques, objects that contain PSI are detected and obscured by, e.g., blurring. Conversely, in our system, background pixels are estimated and then substituted with intended human objects that are essential for the camera person's capture intention. We quantitatively evaluate our system to demonstrate its potential applicability. Yuta Nakashima, Noboru Babaguchi, Jianping Fan 0001 |
ICME | 2 |
| 2011 | Example-based video remixing for home videosabstractVideo remixes are generally created by sequentially arranging selected video clips and combining them with other media streams such as audio clips. In this paper, an example-based approach is adopted for semi-automatically creating video remixes of good expressive quality from home videos. Given multiple home video clips and audio clips, the proposed system generates a video remix by four processes: I)video clip sequence generation, II)audio clip selection, III)audio boundary extraction, and IV)video segment extraction based on professionally created video remix examples. A user interface which presents video clips according to their suitability to the examples and their perceived quality is developed so that users can efficiently and effectively select and arrange suit able video clips in video clip sequence generation. By using movie trailers of action genre as video remix examples, a 43-second video remix was created from 45-minute home videos and was subjectively evaluated better than the one created considering only the perceived quality of video clips. Naoko Nitta, Noboru Babaguchi |
ICME | 2 |
| 2011 | Learning people co-occurrence relations by using relevance feedback for retrieving group photosabstractThis paper proposes an image retrieval method which retrieves images of a specific person from group photos. Many query-by-example methods have focused only on the visual features of the queried person. However, since socially related people such as family and friends are often taken photos together, their co-occurrence relations can be useful information. Thus, we propose an image retrieval method which uses the visual features of not only the queried person but also those who co-occur with the queried person in the same images. Relevance feedback is used to learn who co-occur with the queried person, their faces, and how strong their co-occurrence relations are. When retrieving the images of 19 persons in total from 158 images, after five feedback iterations, the recall rate of 50% was obtained by considering the people co-occurrence relations, as against 33% when considering only the queried person. With human errors in giving relevance feedback, the recall rate still improved to 40%. Kazuya Shimizu, Naoko Nitta, Noboru Babaguchi |
ICMR | 3 |
| 2011 | Extracting intentionally captured regions using point trajectoriesabstractWhen camera persons take videos with mobile video cameras, they usually have capture intentions, i.e., what they want to express in their videos, and there are intentionally captured regions (ICRs) in the video frames that are essential for the capture intentions. Extracting ICRs is thus beneficial for wide range of applications such as video summarization and video adaptation for small displays. In this paper, we present a novel method for automatically extracting ICRs. A camera person usually moves his/her camera so that ICRs can be arranged in appropriate positions in video frames; therefore, ICRs can yield specific motion. This observation indicates that such specific motion is a vital cue for extracting ICRs. The proposed method represents motion by point trajectories, which are long-term trajectories of spatially dense points in video frames, and extracts ICRs using an ICR model based on the point trajectories. We experimentally evaluate the proposed method to demonstrate its potential applicability. Yuta Nakashima, Noboru Babaguchi |
ACM Multimedia | 2 |
| 2011 | Example-based video remixing support systemabstractVideo remixes are generally created by sequentially arranging selected video clips and mixing them with other media streams such as audio clips and transition effects. Especially, mixing music clips often effectively improves the expressive quality of video clips. This paper proposes an example-based system for supporting average users in the 3 steps in video remixing: I)video shot sequence creation, II)music clip selection, and III)audio volume adjustment. The proposed system creates a template for a video remix and gives suggestions to users on an interface such as which video clips should be selected to create a video shot sequence and which music clips should be mixed to the created video shot sequence based on professionally created video remix examples. Then, the audio volume of each video shot is automatically adjusted based on its audio content so that the sounds in the video shots and the music clips would not interfere with each other. Experiments have verified that our system was able to create a video shot sequence whose quality was improved equally as the professionally created one by mixing the selected music clips, doubling the subjective scores from 1.7 to 3.7 on a scale of 1-5. Automatic audio volume adjustment improved the subjective scores by approximately 0.5 points on average. Further, the suggestions provided on the interface was evaluated useful by 5 subjects when creating a video remix by selecting 32 video clips and 4 music clips from 265 video clips and 180 music clips. Naoko Nitta, Noboru Babaguchi |
ACM Multimedia | 2 |
| 2011 | Example-based video remixing
Naoko Nitta, Noboru Babaguchi |
Multim. Tools Appl. | 2 |
| 2010 | Anonymous communication system using probabilistic choice of actions and multiple loopbacksabstractThis paper proposes a new anonymous communication system using probabilistic choice of actions and multiple loopbacks. Our system can provide both sender anonymity and receiver anonymity. Our system also decreases the computation load of each relay node, because there exist no encryption and decryption processes in our system. Applying an analysis method in an anonymous communication system called 3-Mode Net, we evaluate the number of relay nodes and sender anonymity. Kazuhiro Kono, Yoshimichi Ito, Noboru Babaguchi |
IAS | 3 |
| 2010 | Event tactic analysis in sports video using spatio-temporal patternabstractRecently, event detection in sports videos has been gaining some remarkable results. Unfortunately, there is a lack of useful tools for users to explore and exploit these event clips on their own demands. In this paper, a novel method using spatio-temporal patterns to analyze event tactics in sports videos is introduced. The proposed method aims to understand tactics of events such as distributions and speeds of players, or attacking/defensive formations throughout a time when such events happen without tracking objects. The major contribution of the proposed method is to model event tactics by using sequence of symbols. Each symbol represents a distribution of players in a certain period of time. Therefore, a sequence of symbols intrinsically is concerned as spatio-temporal patterns. By using these patterns, an event tactic is detected, explained, and integrated into an event video to create a visualizing abstract. This visualizing abstract is very useful to help users understand an event tactic without watching whole clip. Moreover, users could query by an example or by a text to find all events sharing the same tactic. Thorough testing with over 100 goal clips in soccer domain demonstrate the superiority of the proposed method in terms of precision recall ratios. Minh-Son Dao, Keita Masui, Noboru Babaguchi |
ICIP | 3 |
| 2010 | Real-Time User Position Estimation in Indoor Environments Using Digital Watermarking for Audio SignalsabstractIn this paper, we propose a method for estimating the user position where a user is holding a microphone in an indoor environment using digital watermarking for audio signals. The proposed method utilizes detection strengths, which are calculated while detecting spread-spectrum-based watermarks. Taking into account delays and attenuation of the watermarked signals emitted from multiple loudspeakers and other factors, we construct a model of detection strengths. The user position is estimated in real-time using the model. The experimental results indicate that the user positions are estimated with 1.3 m of root mean squared error on average for the case where the user is static. We demonstrate that the proposed method successfully estimates the user position even when the user moves. Ryosuke Kaneto, Yuta Nakashima, Noboru Babaguchi |
ICPR | 3 |
| 2010 | Hierarchical Anomality Detection Based on SituationabstractIn this paper, we propose a novel anomality detection method based on external situational information and hierarchical analysis of behaviors. Past studies model normal behaviors to detect anomality as outliers. However, normal behaviors tend to differ by situations. Our method combines a set of simple classifiers with pedestrian trajectories as inputs. As mere path information is not sufficient for detecting anomality, trajectories are first decomposed into hierarchical features of different abstract levels and then applied to appropriate classifiers corresponding to the situation it belongs to. Effects of the methods are tested using real environment data. Shuichi Nishio, Hiromi Okamoto, Noboru Babaguchi |
ICPR | 3 |
| 2010 | Discriminating Intended Human Objects in Consumer VideosabstractIn a consumer video, there are not only intended objects, which are intentionally captured by the camcorder user, but also unintended objects, which are accidentally framed-in. Since the intended objects are essential to present what the camcorder user wants to express in the video, discriminating the intended objects from the unintended objects are beneficial for many applications, e.g., video summarization, privacy protection, and so forth. In this paper, focusing on human objects, we propose a method for discriminating the intended human objects from the unintended human objects. We evaluated the proposed method using 10 videos captured by 3 camcorder users. The results demonstrate that the proposed method successfully discriminates the intended human objects with 0.45 of recall and 0.80 of precision. Hiroshi Uegaki, Yuta Nakashima, Noboru Babaguchi |
ICPR | 3 |
| 2010 | Digital Diorama: Sensing-Based Real-World Visualization
Takumi Takehara, Yuta Nakashima, Naoko Nitta, Noboru Babaguchi |
IPMU (2) | 4 |
| 2010 | Automatically protecting privacy in consumer generated videos using intended human object detectorabstractThe growing popularity of video sharing services such as YouTube enables us to upload and share consumer generated videos (CGVs) easily, resulting in disclosure of the privacy sensitive information (PSI) of persons, i.e., their appearances. Therefore, we need a technique for automatically protecting the privacy in CGVs; however, the main problem is how to determine PSI regions automatically. In this paper, we propose a novel system for automatically protecting the privacy in CGVs. The proposed system tackles the problem of determining PSI regions by using an intended human object detector that detects human objects which the camera person wanted to capture to achieve his/her capture intention. In addition, the proposed system adopts several PSI obscuring methods such as blocking out, blurring and seam carving. We present the results of subjective evaluations of a privacy protected video in terms of the visual quality and acceptability of PSI disclosure, as well as the performance of the intended human object detector. Yuta Nakashima, Noboru Babaguchi, Jianping Fan 0001 |
ACM Multimedia | 2 |
| 2010 | Face Image Retrieval across Age Variation Using Relevance Feedback
Naoko Nitta, Atsushi Usui, Noboru Babaguchi |
MMM | 3 |
| 2010 | A new spatio-temporal method for event detection and personalized retrieval of sports video
Minh-Son Dao, Noboru Babaguchi |
Multim. Tools Appl. | 2 |
| 2010 | Constructing distributed hippocratic video databases for privacy-preserving online patient training and counselingabstractDigital video now plays an important role in supporting more profitable online patient training and counseling, and integration of patient training videos from multiple competitive organizations in the health care network will result in better offerings for patients. However, privacy concerns often prevent multiple competitive organizations from sharing and integrating their patient training videos. In addition, patients with infectious or chronic diseases may not want the online patient training organizations to identify who they are or even which video clips they are interested in. Thus, there is an urgent need to develop more effective techniques to protect both video content privacy and access privacy . In this paper, we have developed a new approach to construct a distributed Hippocratic video database system for supporting more profitable online patient training and counseling. First, a new database modeling approach is developed to support concept-oriented video database organization and assign a degree of privacy of the video content for each database level automatically. Second, a new algorithm is developed to protect the video content privacy at the level of individual video clip by filtering out the privacy-sensitive human objects automatically. In order to integrate the patient training videos from multiple competitive organizations for constructing a centralized video database indexing structure, a privacy-preserving video sharing scheme is developed to support privacy-preserving distributed classifier training and prevent the statistical inferences from the videos that are shared for cross-validation of video classifiers. Our experiments on large-scale video databases have also provided very convincing results. Jinye Peng 0001, Noboru Babaguchi, Hangzai Luo, Yuli Gao, Jianping Fan 0001 |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2009 | Performance Analysis of Anonymous Communication System 3-Mode NetabstractThis paper analyzes the performance of 3-Mode Net (3MN), which is a new anonymous communication system proposed. In particular, we give the probability distributions of the number of relay nodes as well as the number of encryption required for communications. The expectations and variances of these two numbers are also given. These results enable us to grasp the influence of the probabilities of mode selections in 3MN. Numerical examples are also presented to illustrate these results. Kazuhiro Kono, Shinnosuke Nakano, Yoshimichi Ito, Noboru Babaguchi |
IAS | 4 |
| 2009 | Preserving topological information in sub-trajectories-based representation for spatio-temporal trajectories indexing and retrievalabstractTrajectories of moving objects are known as one of the most important cues for understanding semantics in video data. Although there are a lot of significant researches dealing with trajectory analysis tailored to indexing and retrieval, several problems still remain. One of them is a trade-off between whole trajectory- and sub-trajectories- based methods. The former problem is that representing a trajectory as a whole is not appropriate for detecting similar patterns of the trajectory. In contrast, the latter is that even though some key portion of two trajectories share similar patterns, the whole trajectories may be totally different. Therefore, this paper proposes a novel method to optimize such trade-off. By representing a trajectory as a combination of sequence of "word" - each word's character represents one distinct feature extracted from sub-trajectories (i.e. segments), and a topological graph of trajectory's segments, the proposed method is shift and scale invariant, can handle occlusion and distortion, and can discover similar patterns among trajectories. Thorough comparisons with well-known methods demonstrate the superiority of the proposed method in terms of precision recall ratios. Minh-Son Dao, Sharma Ishan Nath, Noboru Babaguchi |
ACM Multimedia | 3 |
| 2009 | Automatic Appropriate Segment Extraction from Shots Based on Learning from Example Videos
Yousuke Kurihara, Naoko Nitta, Noboru Babaguchi |
PSIVT | 3 |
| 2009 | Recoverable Privacy Protection for Video Content Distribution
Guangzhen Li, Yoshimichi Ito, Xiaoyi Yu, Naoko Nitta, Noboru Babaguchi |
EURASIP J. Inf. Secur. | 5 |
| 2009 | Automatic personalized video abstraction for sports videos using metadata
Naoko Nitta, Yoshimasa Takahashi, Noboru Babaguchi |
Multim. Tools Appl. | 3 |
| 2009 | Watermarked Movie Soundtrack Finds the Position of the Camcorder in a TheaterabstractIn recent years, the problem of camcorder piracy in theaters has become more serious due to technical advances in camcorders. In this paper, as a new deterrent to camcorder piracy, we propose a system for estimating the recording position from which a camcorder recording is made. The system is based on spread-spectrum audio watermarking for the multichannel movie soundtrack. It utilizes a stochastic model of the detection strength, which is calculated in the watermark detection process. Our experimental results show that the system estimates recording positions in an actual theater with a mean estimation error of 0.44 m. The results of our MUSHRA subjective listening tests show the method does not significantly spoil the subjective acoustic quality of the soundtrack. These results indicate that the proposed system is applicable for practical uses. Yuta Nakashima, Ryuki Tachibana, Noboru Babaguchi |
IEEE Trans. Multim. | 3 |
| 2008 | A discrete wavelet transform based recoverable image processing for privacy protectionabstractThis paper presents a novel scheme of a recoverable image processing for privacy protection in real-time video surveillance system. The privacy information is embedded into the video using information hiding. Thus, the original privacy information can be recoverable with secrete key if necessary. In the proposed system, the privacy information is defined as information of objects that consist of detailed data to recover the original image of objects. The scheme is based on discrete wavelet transform (DWT) which is used for generating privacy-protected low resolution image, as well as the high resolution data including privacy information. An amplitude modulo modulation based information hiding scheme is used to hide the privacy information. Experimental results have shown that the proposed system can reduce the amount of the privacy information significantly, and allows the privacy information to be revealed after being embedded in real time. Guangzhen Li, Yoshimichi Ito, Xiaoyi Yu, Naoko Nitta, Noboru Babaguchi |
ICIP | 5 |
| 2008 | Privacy protecting visual processing for secure video surveillanceabstractPrivacy protection is important in video surveillance. In this paper, we address privacy protection related issues. First, based on questionnaire-based experiments we analyze personal sense of privacy from the viewpoint of the relationship between a viewer and a subject. With the analysis results, we introduce a privacy protected video surveillance system named PriSurv, which can adaptively protect subjects' privacy according to their privacy policy against each viewer. Then, we propose two methods of protecting individuals' privacy by controlling the disclosure of subjects' visual information. One uses a set of visual abstraction operators such as silhouette and dot, which gradually control subjects' visual information. The other uses an active appearance model (AAM) based masks which encode privacy information in the original face region. The latter method can be used especially when a subject's expression can be seen in the video and is characterized by the recoverability of the encoded privacy information. Xiaoyi Yu, Kenta Chinomi, Takashi Koshimizu, Naoko Nitta, Yoshimichi Ito, Noboru Babaguchi |
ICIP | 6 |
| 2008 | Automatic personal preference acquisition from TV viewer's behaviorsabstractThe demand for information services considering personal preferences is increasing. In this paper, we propose a system for automatically acquiring personal preferences from TV viewer’s behaviors. Our system firstly extracts intervals of interest and estimates the interest degree for each extracted interval based on the temporal patterns in facial changes by Hidden Markov Models (HMMs). Then, the viewer profile is created by associating the interest degrees with the content information described in the metadata of the watched program. Experimental results have shown that the proposed methods are able to correctly estimate interest degrees for extracted intervals with a precision rate of 73.1% and a recall rate of 68.8%, and that the created viewer profiles are comparable to the actual preferences of each viewer. Makoto Yamamoto, Naoko Nitta, Noboru Babaguchi |
ICME | 3 |
| 2008 | Weighted stego-Image based steganalysis in multiple least significant bitsabstractThe paper proposes a method for LSB steganalysis of images, where the secret message is embedded in a given number L of the least significant bits. The proposed estimation method is based on weighted stego-Image and no assumption of cover images is required. The method can detect the presence of stego message using LSB steganography from the case L = 1 to arbitrary L≫1.The estimation formula is clean and computation complexity is low. To evaluate the proposed steganalytic method, two experiments of detection and estimation are performed. It is shown that the accuracy of detecting the existence of secret messages in images and of estimating the embedding ratio of secret messages is relatively high. Xiaoyi Yu, Noboru Babaguchi |
ICME | 2 |
| 2008 | Run length based steganalysis for LSB matching steganographyabstractIn this paper, we propose a steganalysis algorithm to detect spatial domain least significant bit (LSB) matching steganography, which is much harder than the detection of LSB replacement. We use the fusion of histogram of run length and histogram characteristic function to detect the LSB Matching. Experimental results on two datasets demonstrate that this method has superior results compared with other recently proposed algorithms, and shows that the proposed method is efficient to detect the LSB matching steganography on compressed or uncompressed images. Xiaoyi Yu, Noboru Babaguchi |
ICME | 2 |
| 2008 | Breaking the YASS algorithm via pixel and DCT coefficients analysisabstractIn this paper, we present a steganalytic method that can reliably detect messages hidden in JPEG images using the steganographic algorithm YASS, which is a JPEG steganographic method shown to be undetectable using current best blind steganalysis classifiers. The key element of the method is features extracted from the imagepsilas pixels and DCT coefficients. Although the YASS process effectively disables the calibration based and the noise model based JPEG steganalyzers, it also disturbs the pixels and DCT coefficients dependency after the secret message embedding. An SVM based classifier is trained based on the extracted features for the detection of the presence of steganography. The method is tested on a diverse set of test images that include both originally uncompressed and compressed images in the TIF and JPEG formats. Our experimental results have demonstrated that the proposed steganalyzers can reliably break YASS. Xiaoyi Yu, Noboru Babaguchi |
ICPR | 2 |
| 2008 | PriSurv: Privacy Protected Video Surveillance System Using Adaptive Visual Abstraction
Kenta Chinomi, Naoko Nitta, Yoshimichi Ito, Noboru Babaguchi |
MMM | 4 |
| 2008 | Appropriate Segment Extraction from Shots Based on Temporal Patterns of Example Videos
Yousuke Kurihara, Naoko Nitta, Noboru Babaguchi |
MMM | 3 |
| 2008 | Mining temporal information and web-casting textfor automatic sports event detectionabstractIn this paper, the generic framework for automatically detecting event based on Allen temporal algebra and external text information support is presented. The motivation of the proposed method is (1) to relax the need of domain knowledge that requires significant human interference; and (2) to take into account the temporal information that has been paid less attention though it is critical to convey event meaning. In order to solve two these problems, in the proposed method, the temporal information is captured by presenting events as the temporal sequences using a lexicon of Allen-based non-ambiguous temporal patterns. These sequences are then used to mine temporal patterns with web-casting text supports by using technique of mining class association rules. Then, the results of previous steps are tailored to build the event detector. Thorough experimental results and comparisons that are carried on more than 30 hours of soccer video corpus captured at different broadcasters and conditions demonstrates that the proposed method meets two aforementioned motivations with high efficiency, effectiveness, and robustness. Minh-Son Dao, Noboru Babaguchi |
MMSP | 2 |
| 2007 | Privacy Preserving: Hiding a Face in a Face
Xiaoyi Yu, Noboru Babaguchi |
ACCV (2) | 2 |
| 2007 | Determining Recording Location Based on Synchronization Positions of AudiowatermarkingabstractIn this paper, we propose a novel application of digital watermarking, determination of recording locations. This application enables us to determine the seat location in an auditorium where a recording was made. Precisely measured synchronization positions of the spread-spectrum watermarks are used for the determination. To avoid use of mismeasured synchronization positions, the algorithm discards synchronization positions with the corresponding normalized correlation values below a threshold. The experiments with our implementation resulted in accurate determinations; almost all of the locations can be determined within the error of 0.5 m. These experimental results successfully show the potential applicability of our application. Yuta Nakashima, Ryuki Tachibana, Masafumi Nishimura, Noboru Babaguchi |
ICASSP (2) | 4 |
| 2007 | Steganography using Sensor Noise and Linear Prediction Synthesis FilterabstractThis paper presents a new approach utilizing the sensor's pattern noise and linear prediction synthesis filter for steganography. The pattern noise is extracted from the images using denoising filter (for example wavelet-based filter). Then the approach introduces the linear prediction synthesis filter, whose parameters are derived from the extracted noise. After being filtered by such a filter, the secret message can be embedded by adapting the characteristics of the sensor's pattern noise. As a result, the embedding process violate little of the natural image statistics, and hence the detectability of steganalytic method is noticeably decreased. The experimental results prove the effectiveness of the new approach. Xiaoyi Yu, Xinshan Zhu, Noboru Babaguchi |
ICIP (2) | 3 |
| 2007 | User and Device Adaptation for Sports Video ContentabstractSeveral methods for summarizing sports videos by selecting only important scenes based upon their semantic content have been proposed so far. However, there remain two problems: (1) what is important depends on each user's preferences, and (2) the summaries should be tailored for media devices that each user has. To solve these problems, we discuss user and device adaptation for sports video summarization. The proposed framework dynamically adapts the video content to fit user's preferences using user profiles which describe the preference degrees for keywords. The video content is then presented through proper media such as image or text according to the confinement of media devices. For sports videos, the framework is tested using PCs and mobile phones as the media devices. Yoshimasa Takahashi, Naoko Nitta, Noboru Babaguchi |
ICME | 3 |
| 2007 | Audio-Based Estimation of Speakers Directions for Multimedia Meeting LogsabstractThis paper is concerned with an audio-based method for estimating speaker directions in meeting environment. It is well-known that cross-power spectrum phase (CSP) analysis is a very powerful tool for localizing sound sources. However, when we adopt the CSP-based method together with a circular microphone array system to estimate the speaker directions in 360-degree range (e.g. round-table discussions), the method fails to estimate the directions due to the existence of imaginary peaks of CSP coefficients. In order to circumvent the above problem, we propose a method to suppress the imaginary peaks, which uses a circular-array version of the method proposed by Nishiura and appropriate scaling around the imaginary peaks. Experimental results are also shown to demonstrate the effectiveness of the proposed method. Yuki Yokoe, Yoshimichi Ito, Noboru Babaguchi |
ICME | 3 |
| 2007 | Preliminary experiments toward automatic generation of new TTS voices from recorded speech alone
Ryuki Tachibana, Tohru Nagano, Gakuto Kurata, Masafumi Nishimura, Noboru Babaguchi |
INTERSPEECH | 5 |
| 2007 | Learning Personal Preference From Viewer's Operations for Browsing and Its Application to Baseball Video Retrieval and SummarizationabstractPersonalization is one of the most important mechanisms to make multimedia systems easy to use. In video applications, its embodiment is to tailor video contents for a particular viewer. For this purpose, we are now developing a system of retrieving and browsing video segments, called video portal with personalization (VIPP). VIPP is characterized by 1) supporting the viewer's access to video contents and making a summarized video clip by taking his/her preference into account and 2) acquiring the viewer's profile from his/her operations automatically. In this paper, we propose a method for learning to personalize from the viewer's operations such as retrieval and browsing, as well as describe how the personalized retrieval and summarization of videos can be realized. From the experiments, we clarify the effect of personalization on retrieval and summarization of baseball videos on VIPP. Noboru Babaguchi, Kouzou Ohara, Takehiro Ogura |
IEEE Trans. Multim. | 1 |
| 2006 | TV Viewing Interval Estimation for Personal Preference AcquisitionabstractThe importance of personalized information services has been increasing. Description of personal preferences needs to be prepared beforehand to realize such services. We propose a system for automatically acquiring personal preferences from TV viewer's behaviors. Considering "when" a viewer is watching TV is highly related to the viewer's preferences, we focus on estimating the time interval during which a pre-registered viewer is watching TV. In this paper, we firstly describe the outline of the personal preference acquisition system, and address a method for estimating the TV viewing intervals based on the appearance of frontal faces. Experiments resulted in a precision rate of 97.1% and a recall rate of 70.6% on average for TV viewing interval estimation Hiroaki Tanimoto, Naoko Nitta, Noboru Babaguchi |
ICME | 3 |
| 2005 | Interactive Clustering of Video Segments for Media StructuringabstractStructuring video data is necessary for its effective retrieval and summarization. In particular, collecting similar scenes from semantic aspects highly contributes to the structuring. In this paper, we propose a method of clustering the scenes with relevance feedback, which may be able to bridge the gap between the video data and its semantics. First, spatio-temporal video segments of a fixed length are clustered according to image features of each segment. Then, a user performs feedback to the results of clustering, whether each segment is relevant to the cluster it belongs. The clustering accuracy can be improved through the interaction based on the feedback information. For diverse kinds of video streams, we investigated how the feedback should be given and demonstrated the effectiveness of the interactive clustering Yukihiro Kinoshita, Naoko Nitta, Noboru Babaguchi |
ICME | 3 |
| 2005 | Automatic parsing of American football videos by intermodal collaboration based on transition rulesabstractThis paper proposes an automatic American football video parsing method based on transition rules of an American football game. Combining the results of live scene extraction and superimposed text detection based on image features enables us to segment the video into play units of a game. Temporally associating the segmented play units with the detected superimposed texts and the closed-caption text attaches possible semantic content information to the play units. Finally, selecting only the play units which conform to transition rules of the sports game from the obtained play unit sequence, while discarding or complementing unnecessary or insufficient play units and attached semantic content information, realizes the semantic video parsing. Naoko Nitta, Noboru Babaguchi |
ICME | 2 |
| 2005 | Video Summarization for Large Sports Video ArchivesabstractVideo summarization is defined as creating a shorter video clip or a video poster which includes only the important scenes in the original video streams. In this paper, we propose two methods of generating a summary of arbitrary length for large sports video archives. One is to create a concise video clip by temporally compressing the amount of the video data. The other is to provide a video poster by spatially presenting the image keyframes which together represent the whole video content. Our methods deal with the metadata which has semantic descriptions of video content. Summaries are created according to the significance of each video segment which is normalized in order to handle large sports video archives. We experimentally verified the effectiveness of our methods by comparing the results with man-made video summaries Yoshimasa Takahashi, Naoko Nitta, Noboru Babaguchi |
ICME | 3 |
| 2005 | PlayWatch: chart-style video playback interfaceabstractThis paper proposes the chart-style video playback interface PlayWatch; it displays a chart of semantic indices for locating video scenes. The main features of PlayWatch are: 1) the user understands the distribution of the scenes because PlayWatch shows the indices in order. 2) The user can access a desired scene directly through the indices since they also act as link buttons. This paper also describes the evaluation of PlayWatch. Experiments on scene searching show that PlayWatch is effective in accessing precisely indexed scenes. Kiyoshi Tanaka, Tsutomu Sasaki, Yoshinobu Tonomura, Tadashi Nakanishi, Noboru Babaguchi |
ICME | 5 |
| 2005 | Generating Semantic Descriptions of Broadcasted Sports Videos Based on Structures of Sports Games and TV Programs
Naoko Nitta, Noboru Babaguchi, Tadahiro Kitahashi |
Multim. Tools Appl. | 2 |
| 2004 | Constructive Inductive Learning Based on Meta-attributes
Kouzou Ohara, Yukio Onishi, Noboru Babaguchi, Hiroshi Motoda |
Discovery Science | 3 |
| 2004 | Motion estimation and detection of complex object by analyzing resampled movements of partsabstractA moving object that has many complex moving parts is very hard to detect and its motion is not easy to estimate. We present a new technique for motion estimation and detection of moving complex objects by analyzing the resampled motions of parts of the objects. A Kalman filter is used to track all resampled movements and the tracked routes are classified into groups that share the same fundamental movements. Our simulations show that recall of motion estimation and detection is approximately 0.8, while the computation drops exponentially. Punpiti Piamsa-nga, Noboru Babaguchi |
ICIP | 2 |
| 2004 | Scene retrieval with sign sequence matching based on video and audio featuresabstractThis work presents a method of retrieving similar scenes from video streams with sign sequence matching. The visual and auditory streams are partitioned into fixed-length video/audio packets. Based on feature vectors extracted from these packets, sign sequences are formed. The sign sequence can be viewed as abstraction of video and audio features. DP matching between the target and query sign sequences allows us to find scenes similar to the query in the video stream. For the purpose of efficient processing, packets, histogram based features, and sign sequences, which are the key ideas in this method, are introduced. The preliminary experimental results show that this method is promising for quick retrieval for similar scenes. Noboru Babaguchi, Tetsuya Ishida, Keisuke Morisawa |
ICME | 1 |
| 2004 | Personalized abstraction of broadcasted American football video by highlight selectionabstractVideo abstraction is defined as creating shorter video clips or video posters from an original video stream. In this paper, we propose a method of generating a personalized abstract of broadcasted American football video. We first detect significant events in the video stream by matching textual overlays appearing in an image frame with the descriptions of gamestats in which highlights of the game are described. Then, we select highlight shots which should be included in the video abstract from those detected events reflecting on their significance degree and personal preferences, and generate a video clip by connecting the shots augmented with related audio and text. An hour-length video can be compressed into a minute-length personalized abstract. We experimentally verified the effectiveness of this method by comparing man-made video abstracts. Noboru Babaguchi, Yoshihiko Kawai, Takehiro Ogura, Tadahiro Kitahashi |
IEEE Trans. Multim. | 1 |
| 2003 | Intermodal collaboration: a strategy for semantic content analysis for broadcasted sports videoabstractThis paper presents intermodal collaboration: a strategy for semantic content analysis for broadcasted sports video. The broadcasted video can be viewed as a set of multimodal streams such as visual, auditory, text (closed caption) and graphics streams. Collaborative analysis for the multimodal streams is achieved based on temporal dependency between their streams, in order to improve the reliability and efficiency for semantic content analysis such as extracting highlight scenes from sports video and automatically generating annotations of specific scenes. A couple of case studies are shown to experimentally confirm the effectiveness of intermodal collaboration. Noboru Babaguchi, Naoko Nitta |
ICIP (1) | 1 |
| 2003 | On Personalizing Video Portal System with Metadata
Kouzou Ohara, Takehiro Ogura, Noboru Babaguchi |
KES | 3 |
| 2003 | Guest Editorial: Best Papers of the ACM Multimedia 2001 Workshop on Multimedia Information Retrieval
Mario A. Nascimento, K. Selçuk Candan, Noboru Babaguchi |
Multim. Tools Appl. | 3 |
| 2002 | Story based representation for broadcasted sports video and automatic story segmentationabstractThis paper presents a model to represent a broadcasted sports video in a semantical way and proposes a method to segment the sports video into the semantical units for the representation. Representation of a video should clarify its semantical content as accurately as possible. Our model structurizes the video and gives suitable semantical descriptions to its particular time locations based on the structures of the sports video. We consider the speech transcript as the useful information stream for the generation of the semantical descriptions. The proposed method tries to segment the speech transcript into the units of the structurization in a probabilistic framework based on Bayesian networks. Moreover, the association with the image stream enables us to find the corresponding video segments. We discuss some experimental results. Naoko Nitta, Noboru Babaguchi, Tadahiro Kitahashi |
ICME (1) | 2 |
| 2002 | Video portal for a media space of structured video streamsabstractThe necessity for a video portal which supports an access to a specific video or its part from a huge video media space consisting of a large number of structured video streams is increasing. The video portal for a media space can be defined as a system which helps us to access videos through various kinds of media segments such as image, text, audio and video segments. In order to construct the video portal, we deal with structured video with metadata. Our system defines the structure of video in terms of document type definition (DTD), and describes the metadata of video with XML according to the DTD. We propose the video portal for structured video streams for the sports domain, and verify its usefulness based on a prototype system's performance. Takehiro Ogura, Noboru Babaguchi, Tadahiro Kitahashi |
ICME (2) | 2 |
| 2002 | Video clustering using spatio-temporal image with fixed lengthabstractIn order to handle video media such as TV programs efficiently and effectively, we need to segment a video stream into video segments and structuralize them based on their contents. We focus on similarity, which is one of the important relations between video segments, and describe a method to cluster similar segments in a video stream. The conventional clustering methods are based on shots, but no complete method to detect shot boundaries has yet been established. Our method is based on fixed length video stream segments, called video packets. Generating spatio-temporal images, we employ cooccurrence matrices to express features in the time dimension explicitly. From clustering experiments for actual TV programs, we obtained clustering accuracy of 81%. Hirotsugu Okamoto, Yukinobu Yasugi, Noboru Babaguchi, Tadahiro Kitahashi |
ICME (1) | 3 |
| 2002 | Event based indexing of broadcasted sports video by intermodal collaborationabstractIn this paper, we propose event-based video indexing, which is a kind of indexing by its semantical contents. Because video data is composed of multimodal information streams such as visual, auditory, and textual [closed caption (CC)] streams, we introduce a strategy of intermodal collaboration, i.e., collaborative processing taking account of the semantical dependency between these streams. Its aim is to improve the reliability and efficiency in contents analysis of video. Focusing here on temporal correspondence between visual and CC streams, the proposed method attempts to seek for time spans in which events are likely to take place through extraction of keywords from the CC stream and then to index shots in the visual stream. The experimental results for broadcasted sports video of American football games indicate that intermodal collaboration is effective for video indexing by the events such as touchdown (TD) and field goal (FG). Noboru Babaguchi, Yoshihiko Kawai, Tadahiro Kitahashi |
IEEE Trans. Multim. | 1 |
| 2001 | Generation Of Personalized Abstract Of Sports VideoabstractVideo abstraction is defined as creating a shorter video clip from an original video stream. In this paper, we propose a method of generating a personalized abstract of broadcasted sports video. We first detect significant events from the video stream by matching with gamestats in which highlights of the game are described. Textual information in an overlay appearing on an image frame is recognized for this matching. Then, we select highlight shots from these detected events, reflecting on personal preferences. Finally, we connect each shot augmented with related audio and text in temporal order. From experimental results, we verified that an hourlength video can be compressed into a minute-length personalized abstract. Noboru Babaguchi, Yoshihiko Kawai, Tadahiro Kitahashi |
ICME | 1 |
| 2000 | Converting Ordinary Rules into Default Rules Based on Contradiction of Knowledge Base
Kouichi Katsurada, Makoto Koyama, Kouzou Ohara, Noboru Babaguchi, Tadahiro Kitahashi |
EJC | 4 |
| 2000 | On Operations for Reconstructing the Complete/Incomplete Knowledge
Kouichi Katsurada, Kouzou Ohara, Noboru Babaguchi, Tadahiro Kitahashi |
EJC | 3 |
| 2000 | Extracting Actors, Actions and Events from Sports Video - A Fundamental Approach to Story TrackingabstractTo effectively deal with the vast amount of videos, we need to construct a content-based representation for each video. As a step towards this goal, this paper proposes a method to automatically generate the semantical annotations for a sports video by integrating the text(c1osed-caption) and image stream. we first segment the text data and extract segments which are meaningful to grasp the story of the video, then extract the actors, the actions and the events of each scene which are useful for information retrieval by using the linguistic cues and the domain knowledge. We also segment the image stream so that each segment can associate with each text segment extracted above by using the image cues. Finally we can annotate the video by associating the text segments with the image segments. Some experimental results are presented and discussed in this paper. Naoko Nitta, Noboru Babaguchi, Tadahiro Kitahashi |
ICPR | 2 |
| 2000 | Determination of General Concept in Learning Default Rules
Kouzou Ohara, Hideyuki Taka, Noboru Babaguchi, Tadahiro Kitahashi |
PRICAI | 3 |
| 1999 | Interactive Approach to the Extraction of Logical Structures from Unformatted Document Images Using a Sub-structure ModelabstractDescribes a new document analysis method for unformatted documents such as advertisements or catalogs. Conventional model-based approaches to the extraction of logical structures are hard to apply to advertisements or catalogs, because a model of a page can't be defined. However, these kinds of documents have similar configurations of the regions that represent each product, where a local model of a local layout and logical structures can be defined. This model, which we call a sub-structure model, can be used as a template to extract the logical structures from other regions that represent the same kinds of products. In proposed system, a sub-structure model is captured through an interactive process with a user. The system was tested on advertisements in Japanese computer magazines and the experiments show promising results. Masaki Yamaoka, Osamu Iwaki, Noboru Babaguchi, Tadahiro Kitahashi |
ICDAR | 3 |
| 1998 | Event detection from continuous mediaabstractIt is difficult to extract semantical contents like events that occurred in the real world from continuous media consisting of visual, auditory, and linguistic information, because machines' ability of understanding media is still incomplete. To avoid this, we propose a strategy of collaborating multimodal information processing for continuous media, called intermodal collaboration and making full use of domain knowledge. The basic experimental results about event detection indicate that our strategy is effective. Noboru Babaguchi, Ramesh Jain 0001 |
ICPR | 1 |
| 1998 | Solving contradiction in knowledge-base without interactionabstractWe propose a non-interactive method for solving the contradictions caused by exceptions to ordinary rules. It is realized by: detecting the ordinary rules which have the instances in their bodies; and converting them into the default rules. To reduce the default rules which need much reasoning time, we convert the minimal sets of ordinary rules to solve contradictions. Since the proposed method is executed without human interaction, it contributes to automatic rule-base maintenance. Kouichi Katsurada, Makoto Koyama, Kouzou Ohara, Noboru Babaguchi, Tadahiro Kitahashi |
SMC | 4 |
| 1998 | Non-monotonic inference system handling knowledge allowing classified exceptionsabstractWe propose a nonmonotonic formalism for knowledge handling allowing exceptions, the Exc-Representation (ER), and its goal-directed proof procedure, the SLD-EXC resolution, for the nonmonotonic inference system, NISE. ER always provides a unique extension, which is a set of conclusions, and makes tractable the membership problem in nonmonotonic reasoning by introducing the conditional facts, and classifying exceptions into two types based on the relationships among rules. Kouzou Ohara, Noboru Babaguchi, Tadahiro Kitahashi |
SMC | 2 |
| 1997 | Media Information Processing in Documents -Generation of Manuals of Mechanical Parts AssemblingabstractThe authors view a document as a multi-modal media in the sense that it presents information by using figures and/or images as well as text. By showing examples of media conversion and media integration, the authors discuss the essential processing techniques of media information, especially document media information. First, media conversion is briefly addressed by an example of converting a graph diagram in a document to either an alternative graph diagram or a table. The main topic is the generation of a manual of a mechanical parts assembly as an example of media integration in a document. Thereby, necessary input information is presented for generating the manual. Tadahiro Kitahashi, Muneki Ohya, Koh Kakusho, Noboru Babaguchi |
ICDAR | 4 |
| 1996 | Recognition of Social Dancing from Auditory and Visual InformationabstractWe discuss the recognition of social dancing from a real image sequence and an acoustic signal of music. Since social dancing involves somewhat complicated movements by multiple humans, it is difficult to recover a detailed description of each dancer's posture by the conventional approach based on matching of an articulated human body model to each frame of an image sequence. Using the results of the recognition process for annotating a video of social dancing, we focus on the bare minimum of motion description. We show that the motion description can be acquired from a real image sequence and an acoustic signal by simple familiar techniques of computer vision and music information processing. Koh Kakusho, Noboru Babaguchi, Tadahiro Kitahashi |
FG | 2 |
| 1996 | Generation of sketch map image and its instructions to support the understanding of geographical informationabstractIn this paper, we propose a method of generating a sketch map drawing and its instructions in order to support the route understanding of a human. Both are given from a structured data, called the road network, which is a graph augmented by the attributes of roads and crossings. In generating a sketch map drawing, the pictorial information about the shape and the direction of roads is modified. The sketch map drawing can be generated at any simplification level by controlling a parameter called simpleness. The instructions corresponding to the sketch map are produced by assigning appropriate terms into sentence templates. We verified that favorable results are obtained by this method. Noboru Babaguchi, Seiichiro Dan, Tadahiro Kitahashi |
ICPR | 1 |
| 1996 | A Query Procedure for Allowing Exceptions in Advanced Logical Database
Kouzou Ohara, Noboru Babaguchi, Tadahiro Kitahashi |
IEA/AIE | 2 |
| 1996 | On Formation of Exception Hierarchy
Kouzou Ohara, Noboru Babaguchi, Tadahiro Kitahashi |
PRICAI | 2 |
| 1995 | Extracting characters and character lines in multi-agent schemeabstractIn this paper, we present COCE (COordinative Character Extractor), a new method for extracting printed Japanese characters from an unformatted document image. This research aims to exploit knowledge independent of the layouts. COCE is based on a multiagent scheme where each agent is assigned to a single character line and tries to extract characters by making use of the knowledge about features of a character line as well as shapes and arrangement of characters. Moreover, the agents communicate with each other to keep consistency between their tasks. We have favourable results for the effectiveness of this method. Keiji Gyohten, Tomoko Sumiya, Noboru Babaguchi, Koh Kakusho, Tadahiro Kitahashi |
ICDAR | 3 |
| 1994 | Generation of Sketch Map Drawing from Vectorized ImageabstractThis paper presents a method of generating a sketch drawing from geographical map images with vector representation. This method consists of two stages: generation of a road network from a vectorized image and generation of a sketch map drawing based on the road network. The road network is fundamental data for any applications, represented as a graph structure that is augmented with related attributes about roads and crossings. The characteristic of this method is to introduce a parameter, called roughness, in order to control sketch map generation. Because of the roughness, we can obtain a variety of sketch map drawings according to user's requirements. We experimentally verified that the proposed method is effective and that the roughness is valid as a psychological measure.> Noboru Babaguchi, Kiyoshi Tanaka, Tadahiro Kitahashi |
ICIP (3) | 1 |
| 1993 | Incremental acquisition of knowledge about layout structures from examples of documentsabstractDocument image analysis systems often utilize the knowledge about layout structures to extract layout objects labeled logically. However, the lack of the facility for knowledge acquisition limits the applicability of the systems. The authors propose a method of acquiring knowledge for document image analysis. Given examples of document images and their layout objects labeled logically, the method generates and modifies the knowledge. The method is incremental so that the knowledge can be efficiently modified using additional examples. Counterexamples generated as errors obtained from the analysis of an example image can also be reflected into the knowledge so that the system may no longer generate the errors. Experimental results on both knowledge acquisition and the analysis using the acquired knowledge are also presented.> Koichi Kise, Naoko Yajima, Noboru Babaguchi, Kunio Fukunaga |
ICDAR | 3 |
| 1992 | Model based system for analyzing document imagesabstractDocument image analysis is the process of deriving logically structured representation of a document by analyzing the layout structure of its image. This paper proposes a knowledge based system for document image analysis which is applicable to various kinds of documents. The characteristics of the system are as follows: (1) The knowledge base called document model encodes only object-level knowledge hierarchically, declaratively and symbolically, aiming at high expressivity and maintainability of the knowledge description; (2) the document model is automatically constructed by referring samples of document images, and incrementally refined by feedback of error information of analysis.> Koichi Kise, Masaki Yamaoka, Noboru Babaguchi, Yoshikazu Tezuka |
ICPR (2) | 3 |
| 1991 | Connectionist Model BinarizationabstractImage binarization is a task to convert gray-level images into bi-level ones. Its underlying notion can be simply thought of as threshold selection. However, the result of binarization will cause significant influence on the process of image recognition or understanding. In this paper we discuss a new binarization method, named CMB (connectionist model binarization), which uses the connectionist model. In the method a gray-level histogram is input to a multilayer network trained with the back-propagation algorithm to obtain a threshold which gives a visually suitable binarized image. From the experimental results, it was verified that CMB is an effective binarization method in comparison with other methods. Noboru Babaguchi, Koji Yamada, Koichi Kise, Yoshikazu Tezuka |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1990 | Connectionist model binarizationabstractThe application of a connectionist model to an image binarization method called connectionist model binarization (CMB) is discussed. CMB employs a multilayer network of a connectionist model whose input and output are a histogram and a desirable threshold for binarization, respectively. This network is trained with a back-propagation algorithm to output a threshold which gives a visually suitable binarised image against any histogram. The details of CMB are described, and its learning strategy and binarization performance are discussed.> Noboru Babaguchi, Koji Yamada, Koichi Kise, Yoshikazu Tezuka |
ICPR (2) | 1 |
| 1988 | Visiting card understanding systemabstractThe authors present the visiting card understanding system, whose output is suitable for the input of a visiting-card database. The system consists of two modules. One is a document model which represents the hierarchical knowledge about the layout structure of visiting cards. The other is an understanding module which interprets the document model to general and test hierarchical hypotheses about the contents of a visiting card. Since the understanding module is fundamentally independent of document type, the system is applicable to many kinds of documents.> Koichi Kise, Koji Yamada, Naoki Tanaka, Noboru Babaguchi, Yoshikazu Tezuka |
ICPR | 4 |
| 1987 | Curvedness of a line picture
Noboru Babaguchi, Tsunehiro Aibara |
Pattern Recognit. | 1 |