EDBT 2026 Demo / reviewers in the wild / expert
Naoko Nitta
dblp:47/6775
· DBLP profile ↗
41ranked-venue papers
12as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 12 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Security and privacy · 4 · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorComputer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Social IoT Approach to Cyber Defense of a Deep-Learning-Based Recognition System in Front of Media Clones Generated by Model Inversion AttackabstractModel inversion attack (MIA) is a cyber threat with an increasing alert even for deep-learning-based recognition systems (DLRSs). By targeting a DLRS under a scenario of attacker access to the model structure and parameters, MIA generates a data clone for a certain targeted class label. To avoid the possible threats of such MIA-generated data clones, this research work proposes a social IoT approach to a collaborative cyber-defense among the online recognition systems (RSs) sharing the targeted class label. Since, the generation of an MIA-clone is by targeting an RS model and using its structure, parameters, and class labels output scores in an iterative optimization process, the generated clone is partially inherent to the targeted model. Thus, it is expected for an MIA-clone to show a different performance on a secondary RS wherein the same targeted class label is included. It is because, in the MIA generation of the clone, not only the targeted class label but also other class labels, and model parameters and structure affect the process, while the second model has just the targeted class label in common with the target model. Deploying the Social Internet of Recognition Systems (SIoRS), the proposed technique utilizes a collaborative recognition by SIoRC which plays the role of a complementary recognition besides the targeted RS. The recognition output by the targeted RS is further verified by the SIoRS complementary recognition result. To avoid the MIA-targeted data clones, the verification of recognition is by the log-likelihood ratio test between the targeted RS and the SIoRS complementary recognition confidence scores. The proposed technique is evaluated by statistical analysis on deep face RSs in 10000 Monte Carlo runs for each of the conventional, dc-generative adversarial network (GAN) and$\alpha $-GAN integrated MIA techniques in targeting two different user identities. The$Z$scores of the fitted normal distribution of the log-likelihood ratios indicate almost 100% detection rate of clones generated by conventional MIA and 95.23% and 86% of clones, respectively, generated by DC-GAN and$\alpha $-GAN integrated deep MIA techniques. Mahdi Khosravy, Kazuaki Nakamura, Naoko Nitta, Nilanjan Dey, Rubén González Crespo, Enrique Herrera-Viedma, Noboru Babaguchi |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | MMArt-ACM 2022: 5th Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in MultimediaabstractIn addition to classical art types like paintings and sculptures, new types of artworks emerge following the advancement of deep learning, social platforms, media capturing devices, and media processing tools. Large volumes of machine-/user-generated content or professionally-edited content are shared and disseminated on the Web. Novel multimedia artworks, therefore, emerge rapidly in the era of social media and big data. The ever-increasing amount of illustrations/comics/animations on this platform gives rise to challenges of automatic classification, indexing, and retrieval that have been studied widely in other areas but not necessarily for this emerging type of artwork. In addition to objective entities like objects, events, and scenes, studies of cognitive properties emerge. Among various kinds of computational cognitive analyses, we focus on attractiveness analysis in this workshop. The topics of the accepted papers cover the affective analysis of texts, images, and music. The actual MMArt-ACM 2022 Proceedings are available at: https://dl.acm.org/citation.cfm?id=3512730. Naoko Nitta, Min-Chun Hu 0001, Kensuke Tobitani |
ICMR | 1 |
| 2022 | Anonymization of Human Gait in Video Based on Silhouette Deformation and Texture TransferabstractThese days, a lot of videos are uploaded onto web-based video sharing services such as YouTube. These videos can be freely accessed from all over the world. On the other hand, they often contain the appearance of walking private people, which could be identified by silhouette-based gait recognition techniques rapidly developed in recent years. This causes a serious privacy issue. To avoid it, this paper proposes a method for anonymizing the appearance of walking people, namely human gait, in video. In the proposed method, we first crop human regions from all frames in an input video and binarize them to get their silhouettes. Next, we slightly deform the silhouettes from the aspects of static body shape and dynamic walking rhythm so that the person in the input video cannot be correctly identified by gait recognition techniques. After that, the textures of the original human regions are transferred onto the deformed silhouettes. We achieve this by a displacement field-based approach, which is training-free and thus robust to a variety of clothes. Finally, the anonymized human regions with the transferred textures are filled back into the input video. In the results of our experiments, we successfully degraded the accuracy of CNN-based gait recognition systems from 100% to 1.57% in the lowest case without yielding serious distortion in the appearance of the human regions, which demonstrated the effectiveness of the proposed method. Yuki Hirose, Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Model Inversion Attack by Integration of Deep Generative Models: Privacy-Sensitive Face Generation From a Face Recognition SystemabstractCybersecurity in front of attacks to a face recognition system is an emerging issue in the cloud era, especially due to its strong bonds with the privacy of the users registered to the system. A possible attack is the model inversion attack (MIA) which aims to reveal the identity of a targeted user by generating the most proper datapoint input to the system with maximum corresponding confidence score at the output. The generated data of a registered user can be maliciously used as a serious invasion of the user privacy. In literature, MIA processes are categorized into white-box and black-box scenarios which are respectively with and without information about the system structure, parameters, and partially about the users. This research work assumes the MIA under semi-white box scenario of availability of system model structure and parameters but not any user data information, and verifies it as a severe threat even for a deep-learning-based face recognition system despite its complex structure and the diversity of registered user data. The alert state is promoted by Deep MIA which is the integration of deep generative models in MIA, and$\alpha $-GAN integrated MIA-initilized by a face based seed ($\alpha $-GAN-MIA-FS) is proposed. As a novel MIA search strategy, a pre-trained deep generative model with capability of generating a face image from a random feature vector is used for narrowing down the image search space to the feature vectors space, which has much lower dimensions. This allows the MIA process to efficiently search for a low-dimensional feature vector whose corresponding face image maximizes the confidence score. We have experimentally evaluated the proposed method by two objective criteria and three subjective criteria in comparison to$\alpha $-GAN-integrated MIA initialized with a random seed ($\alpha $-GAN-MIA-RS), DCGAN-integrated MIA (DCGAN-MIA), and the conventional MIA. The evaluation results approve the efficiency and superiority of the proposed technique in generating natural looking face clones with high recognizability as the targeted users. Mahdi Khosravy, Kazuaki Nakamura, Yuki Hirose, Naoko Nitta, Noboru Babaguchi |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | Semi-Supervised Outdoor Image Generation Conditioned on Weather SignalsabstractIn recent years, various types of sensors observe the real world. Especially, weather sensors are densely installed all over the world to observe current weather situations at various places. However, weather signals such as the temperature or humidity obtained by weather sensors are intuitively difficult for humans to understand. On the other hand, images captured by typical RGB cameras can tell weather situations at the captured places in a more comprehensible way for humans; however, cameras are only installed at limited places and are not necessarily open to public due to privacy issues. In order to solve this problem, the goal of our work is to generate images which can tell weather situations at arbitrary time and locations. This can be realized by using a conditional generative adversarial network architecture that takes an image and a condition to transform the image accordingly to the condition. Training such image generator requires a large number of image and condition pairs as the training data. Although weather signals can be easily collected from weather sensors, collecting their spatially and temporally synchronized outdoor images is not easy. Thus, we propose a semi-supervised method for training the image generator. A relatively small number of pairs of an outdoor image and weather signals is collected, each from different web services, by considering their semantic consistency. The collected pairs are used to train a predictor for predicting weather signals from a given outdoor image. Then, the image generator is trained by using a large number of pairs of an outdoor image and pseudo weather signals predicted by the predictor as the training data. Sota Kawakami, Kei Okada, Naoko Nitta, Kazuaki Nakamura, Noboru Babaguchi |
ICPR | 3 |
| 2020 | MMArt-ACM'20: International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia 2020abstractThe International Joint Workshop on Multimedia Artworks Analysis and Attractiveness Computing in Multimedia (MMArt-ACM) solicits contributions on methodology advancement and novel applications of multimedia artworks and attractiveness computing that emerge in the era of big data and deep learning. Despite the strike of the Covid-19 pandemic, this workshop attracts submissions of diverse topics in these two fields, and the workshop program finally consists of five presented papers. The topics cover image retrieval, image transformation and generation, recommendation system, and image/video summarization. The actual MMArt-ACM'20 Proceedings are available in the ACM DL at: https://dl.acm.org/citation.cfm?id=3379173 Wei-Ta Chu, Ichiro Ide, Naoko Nitta, Norimichi Tsumura, Toshihiko Yamasaki |
ICMR | 3 |
| 2020 | Reproducibility Companion Paper: Visual Sentiment Analysis for Review Images with Item-Oriented and User-Oriented CNNabstractWe revisit our contributions on visual sentiment analysis for online review images published at ACM Multimedia 2017, where we develop item-oriented and user-oriented convolutional neural networks that better capture the interaction of image features with specific expressions of users or items. In this work, we outline the experimental claims as well as describe the procedures to reproduce the results therein. In addition, we provide artifacts including data sets and code to replicate the experiments. Quoc-Tuan Truong, Hady Wirawan Lauw, Martin Aumüller 0001, Naoko Nitta |
ACM Multimedia | 4 |
| 2020 | Probabilistic Stone's Blind Source Separation with application to channel estimation and multi-node identification in MIMO IoT green communication and multimedia systems
Mahdi Khosravy, Nilesh Patel, Nilanjan Dey, Naoko Nitta, Noboru Babaguchi |
Comput. Commun. | 5 |
| 2019 | Training-Free Method for Generating Motion Video Clones From A Still Image Considering Self-Occlusion of Human BodyabstractIn this paper, we propose a method for generating photo-realistic video in which a person virtually performs a motion that is not performed in the real world. we refer to such video as motion video clones (MVCs). The proposed method requires only two kinds of source information as input data: a reference video, in which a person A performs some motion, and a target image, which includes the whole body of another person B. Using these data, our method generates a MVC in which the person A's motion is re-enacted by the person B's body. Since our method does not need 3D human body model nor any training phase, it is suitable to MVC-based entertainment systems. To handle the self-occlusion of the human body in the reference video, we employ a part-based approach. For each part such as the right arm and the left leg, we first extract its skeleton from the target image and move it so that the motion represented by the reference video is reenacted. Next, we compute a 2D affine transform between the original and moved positions of the skeleton. This transform is used to map the texture of the target image onto each frame of a resultant MVC. Finally, we extend the part-wise affine transforms to pixel-wise ones by computing their linear combination for each pixel, whose combination weights are computed according to the geodesic distance between the pixel and the center of each part. This allows us to avoid unnatural appearance around the joints of the body parts. In our experiments, the proposed method generated much visually-natural MVCs than existing methods. Teppei Tsutsumi, Kazuaki Nakamura, Seiko Myojin, Naoko Nitta, Noboru Babaguchi |
ICIP | 4 |
| 2019 | Reproducible Experiments on Adaptive Discriminative Region Discovery for Scene RecognitionabstractThis companion paper supports the replication of scene image recognition experiments using Adaptive Discriminative Region Discovery (Adi-Red), an approach presented at ACM Multimedia 2018. We provide a set of artifacts that allow the replication of the experiments using a Python implementation. All the experiments are covered in a single shell script, which requires the installation of an environment, following our instructions, or using ReproZip.The data sets (images and labels) are automatically downloaded, and the train-test splits used in the experiments are created. The first experiment is from the original paper, and the second supports exploration of the resolution of the scale-specific input image, an interesting additional parameter. For both experiments, five other parameters can be adjusted: the threshold used to select the number of discriminative patches, the number of scales used, the type of patch selection (Adi-Red, dense or random), the architecture and pre-training data set of the pre-trained CNN feature extractor. The final output includes four tables (original Table 1, Table 2 and Table 4, and a table for the resolution experiment) and two plots (original Figure 3 and Figure 4). Zhengyu Zhao 0001, Zhuoran Liu 0001, Martha A. Larson, Ahmet Iscen, Naoko Nitta |
ACM Multimedia | 5 |
| 2019 | Encryption-Free Framework of Privacy-Preserving Image Recognition for Photo-Based Information ServicesabstractNowadays, mobile devices, such as smartphones, have been widely used all over the world. In addition, the performance of image recognition has drastically increased with deep learning technologies. From these backgrounds, some photo-based information services provided in a client-server architecture are getting popular: client users take a photo of a certain spot and send it to a server, while the server identifies the spot with an image recognizer and returns its related information to the users. However, this kind of client-server image recognition can cause a privacy issue because image recognition results are sometimes privacy-sensitive. To tackle the privacy issue, in this paper, we propose a framework of privacy-preserving image recognition called EnfPire, in which the server cannot uniquely determine the recognition result but client users can do so. An overview of EnfPire is as follows. First, client users extract a visual feature from their taken photo and transform it so that the server cannot uniquely determine the recognition result. Then, the users send the transformed feature to the server that returns a set of candidates of the recognition result to the users. Finally, the users compare the candidates to the original visual feature for obtaining the final result. Our experimental results demonstrate that EnfPire successfully degrades the server's spot-recognition accuracy from 99.8% to 41.4% while keeping 86.9% of the spot-recognition accuracy on the user side. Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Generating Handwritten Character Clones from an Incomplete Seed Character Set using Collaborative FilteringabstractIn this paper, we propose a method for generating clones of a target writer's handwritten character images (called handwritten character clones or HCCs) using an incomplete seed character set, which consists of at most one or no example of his/her actual handwriting per character. In HCC generation, not a single HCC but its distribution should be created for each character because humans' actual handwriting images differ from each other even if the same writer writes the same character. However, it is difficult to achieve this from the incomplete seed character set. To solve the problem, in the proposed method, we first create a number of HCC distributions for each character by clustering a set of handwritten character images offered by other writers. Next, for each character contained in the seed character set, we choose the distribution best fit to its example. Finally, for the other characters, we estimate the best distribution for them employing collaborative filtering. We conducted pilot experiments focusing on Japanese character images, in which the proposed method successfully generated various HCCs with a certain level of quality for each character. Kazuaki Nakamura, Eiji Miyazaki, Naoko Nitta, Noboru Babaguchi |
ICFHR | 3 |
| 2017 | A Framework of Privacy-Preserving Image Recognition for Image-Based Information Services
Kojiro Fujii, Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
MMM (1) | 3 |
| 2017 | Effect of Junk Images on Inter-concept Distance Measurement: Positive or Negative?
Yusuke Nagasawa, Kazuaki Nakamura, Naoko Nitta, Noboru Babaguchi |
MMM (2) | 3 |
| 2015 | Real-Time People Counting across Spatially Adjacent Non-overlapping Camera Views
Ryota Akai, Naoko Nitta, Noboru Babaguchi |
MMM (1) | 2 |
| 2014 | Real-World Event Detection Using Flickr Images
Naoko Nitta, Yusuke Kumihashi, Tomochika Kato, Noboru Babaguchi |
MMM (2) | 1 |
| 2013 | People counting across spatially disjoint cameras by flow estimation between foreground regionsabstractOur goal is to develop a method for counting the number of people traveling over a wide area monitored by spatially disjoint multiple cameras with non-overlapping fields of view. The proposed method counts the number of people traversing across each pair of cameras' fields of view by estimating the flows between the foreground regions which have disappeared from and appeared in the camera views within a short time interval. The approach aims at resolving two problems: people in a foreground region can split and merge outside the cameras' fields of view, and the appearance variance of the same person and persons with similar appearance can lead to errors in person re-identification across cameras. The average errors of 0.184 persons in counting the number of people traveling over an area in a university campus monitored by four virtual cameras for 100 minutes have demonstrated the effectiveness of the proposed method. Naoko Nitta, Takayuki Nakazaki, Kazuaki Nakamura, Ryota Akai, Noboru Babaguchi |
AVSS | 1 |
| 2012 | Classification based group photo retrieval with bag of people featuresabstractThis paper proposes a method for retrieving images containing a specific target person from a given image collection of group photos. This can be realized by query-by-example methods which compare the facial visual features of the target person in the given query image and of each person in the images in the image collection. However, since images are often taken under various conditions, facial appearance of the same person can vary. Since socially related people such as family and friends are often taken photos together, the people co-occurrence relations in the same images can also be a useful clue for image retrieval. Focusing on such people co-occurrence relations, we propose Bag of People (BoP) features which represent both the facial appearances of persons and their co-occurrence relations in the same images. By using the BoP features, a classifier for classifying images into two classes, images containing the target person and other images, can be trained from a small number of images labeled by user's relevance feedback. Furthermore, since the labeled images obtained by relevance feedback are much fewer than unlabeled images in the image collection, an active learning method is used to select useful images to train the classifier. When retrieving images of 24 persons in total from 550 images, after five feedback iterations, the mean average precision of 0.94 was obtained by considering the people co-occurrence relations, as against 0.69 when considering only the target person. Kazuya Shimizu, Naoko Nitta, Yujiro Nakai, Noboru Babaguchi |
ICMR | 2 |
| 2011 | Example-based video remixing for home videosabstractVideo remixes are generally created by sequentially arranging selected video clips and combining them with other media streams such as audio clips. In this paper, an example-based approach is adopted for semi-automatically creating video remixes of good expressive quality from home videos. Given multiple home video clips and audio clips, the proposed system generates a video remix by four processes: I)video clip sequence generation, II)audio clip selection, III)audio boundary extraction, and IV)video segment extraction based on professionally created video remix examples. A user interface which presents video clips according to their suitability to the examples and their perceived quality is developed so that users can efficiently and effectively select and arrange suit able video clips in video clip sequence generation. By using movie trailers of action genre as video remix examples, a 43-second video remix was created from 45-minute home videos and was subjectively evaluated better than the one created considering only the perceived quality of video clips. Naoko Nitta, Noboru Babaguchi |
ICME | 1 |
| 2011 | Learning people co-occurrence relations by using relevance feedback for retrieving group photosabstractThis paper proposes an image retrieval method which retrieves images of a specific person from group photos. Many query-by-example methods have focused only on the visual features of the queried person. However, since socially related people such as family and friends are often taken photos together, their co-occurrence relations can be useful information. Thus, we propose an image retrieval method which uses the visual features of not only the queried person but also those who co-occur with the queried person in the same images. Relevance feedback is used to learn who co-occur with the queried person, their faces, and how strong their co-occurrence relations are. When retrieving the images of 19 persons in total from 158 images, after five feedback iterations, the recall rate of 50% was obtained by considering the people co-occurrence relations, as against 33% when considering only the queried person. With human errors in giving relevance feedback, the recall rate still improved to 40%. Kazuya Shimizu, Naoko Nitta, Noboru Babaguchi |
ICMR | 2 |
| 2011 | Example-based video remixing support systemabstractVideo remixes are generally created by sequentially arranging selected video clips and mixing them with other media streams such as audio clips and transition effects. Especially, mixing music clips often effectively improves the expressive quality of video clips. This paper proposes an example-based system for supporting average users in the 3 steps in video remixing: I)video shot sequence creation, II)music clip selection, and III)audio volume adjustment. The proposed system creates a template for a video remix and gives suggestions to users on an interface such as which video clips should be selected to create a video shot sequence and which music clips should be mixed to the created video shot sequence based on professionally created video remix examples. Then, the audio volume of each video shot is automatically adjusted based on its audio content so that the sounds in the video shots and the music clips would not interfere with each other. Experiments have verified that our system was able to create a video shot sequence whose quality was improved equally as the professionally created one by mixing the selected music clips, doubling the subjective scores from 1.7 to 3.7 on a scale of 1-5. Automatic audio volume adjustment improved the subjective scores by approximately 0.5 points on average. Further, the suggestions provided on the interface was evaluated useful by 5 subjects when creating a video remix by selecting 32 video clips and 4 music clips from 265 video clips and 180 music clips. Naoko Nitta, Noboru Babaguchi |
ACM Multimedia | 1 |
| 2011 | Example-based video remixing
Naoko Nitta, Noboru Babaguchi |
Multim. Tools Appl. | 1 |
| 2010 | Digital Diorama: Sensing-Based Real-World Visualization
Takumi Takehara, Yuta Nakashima, Naoko Nitta, Noboru Babaguchi |
IPMU (2) | 3 |
| 2010 | Face Image Retrieval across Age Variation Using Relevance Feedback
Naoko Nitta, Atsushi Usui, Noboru Babaguchi |
MMM | 1 |
| 2009 | Automatic Appropriate Segment Extraction from Shots Based on Learning from Example Videos
Yousuke Kurihara, Naoko Nitta, Noboru Babaguchi |
PSIVT | 2 |
| 2009 | Recoverable Privacy Protection for Video Content Distribution
Guangzhen Li, Yoshimichi Ito, Xiaoyi Yu, Naoko Nitta, Noboru Babaguchi |
EURASIP J. Inf. Secur. | 4 |
| 2009 | Automatic personalized video abstraction for sports videos using metadata
Naoko Nitta, Yoshimasa Takahashi, Noboru Babaguchi |
Multim. Tools Appl. | 1 |
| 2008 | A discrete wavelet transform based recoverable image processing for privacy protectionabstractThis paper presents a novel scheme of a recoverable image processing for privacy protection in real-time video surveillance system. The privacy information is embedded into the video using information hiding. Thus, the original privacy information can be recoverable with secrete key if necessary. In the proposed system, the privacy information is defined as information of objects that consist of detailed data to recover the original image of objects. The scheme is based on discrete wavelet transform (DWT) which is used for generating privacy-protected low resolution image, as well as the high resolution data including privacy information. An amplitude modulo modulation based information hiding scheme is used to hide the privacy information. Experimental results have shown that the proposed system can reduce the amount of the privacy information significantly, and allows the privacy information to be revealed after being embedded in real time. Guangzhen Li, Yoshimichi Ito, Xiaoyi Yu, Naoko Nitta, Noboru Babaguchi |
ICIP | 4 |
| 2008 | Privacy protecting visual processing for secure video surveillanceabstractPrivacy protection is important in video surveillance. In this paper, we address privacy protection related issues. First, based on questionnaire-based experiments we analyze personal sense of privacy from the viewpoint of the relationship between a viewer and a subject. With the analysis results, we introduce a privacy protected video surveillance system named PriSurv, which can adaptively protect subjects' privacy according to their privacy policy against each viewer. Then, we propose two methods of protecting individuals' privacy by controlling the disclosure of subjects' visual information. One uses a set of visual abstraction operators such as silhouette and dot, which gradually control subjects' visual information. The other uses an active appearance model (AAM) based masks which encode privacy information in the original face region. The latter method can be used especially when a subject's expression can be seen in the video and is characterized by the recoverability of the encoded privacy information. Xiaoyi Yu, Kenta Chinomi, Takashi Koshimizu, Naoko Nitta, Yoshimichi Ito, Noboru Babaguchi |
ICIP | 4 |
| 2008 | Automatic personal preference acquisition from TV viewer's behaviorsabstractThe demand for information services considering personal preferences is increasing. In this paper, we propose a system for automatically acquiring personal preferences from TV viewer’s behaviors. Our system firstly extracts intervals of interest and estimates the interest degree for each extracted interval based on the temporal patterns in facial changes by Hidden Markov Models (HMMs). Then, the viewer profile is created by associating the interest degrees with the content information described in the metadata of the watched program. Experimental results have shown that the proposed methods are able to correctly estimate interest degrees for extracted intervals with a precision rate of 73.1% and a recall rate of 68.8%, and that the created viewer profiles are comparable to the actual preferences of each viewer. Makoto Yamamoto, Naoko Nitta, Noboru Babaguchi |
ICME | 2 |
| 2008 | PriSurv: Privacy Protected Video Surveillance System Using Adaptive Visual Abstraction
Kenta Chinomi, Naoko Nitta, Yoshimichi Ito, Noboru Babaguchi |
MMM | 2 |
| 2008 | Appropriate Segment Extraction from Shots Based on Temporal Patterns of Example Videos
Yousuke Kurihara, Naoko Nitta, Noboru Babaguchi |
MMM | 2 |
| 2007 | User and Device Adaptation for Sports Video ContentabstractSeveral methods for summarizing sports videos by selecting only important scenes based upon their semantic content have been proposed so far. However, there remain two problems: (1) what is important depends on each user's preferences, and (2) the summaries should be tailored for media devices that each user has. To solve these problems, we discuss user and device adaptation for sports video summarization. The proposed framework dynamically adapts the video content to fit user's preferences using user profiles which describe the preference degrees for keywords. The video content is then presented through proper media such as image or text according to the confinement of media devices. For sports videos, the framework is tested using PCs and mobile phones as the media devices. Yoshimasa Takahashi, Naoko Nitta, Noboru Babaguchi |
ICME | 2 |
| 2006 | TV Viewing Interval Estimation for Personal Preference AcquisitionabstractThe importance of personalized information services has been increasing. Description of personal preferences needs to be prepared beforehand to realize such services. We propose a system for automatically acquiring personal preferences from TV viewer's behaviors. Considering "when" a viewer is watching TV is highly related to the viewer's preferences, we focus on estimating the time interval during which a pre-registered viewer is watching TV. In this paper, we firstly describe the outline of the personal preference acquisition system, and address a method for estimating the TV viewing intervals based on the appearance of frontal faces. Experiments resulted in a precision rate of 97.1% and a recall rate of 70.6% on average for TV viewing interval estimation Hiroaki Tanimoto, Naoko Nitta, Noboru Babaguchi |
ICME | 2 |
| 2005 | Interactive Clustering of Video Segments for Media StructuringabstractStructuring video data is necessary for its effective retrieval and summarization. In particular, collecting similar scenes from semantic aspects highly contributes to the structuring. In this paper, we propose a method of clustering the scenes with relevance feedback, which may be able to bridge the gap between the video data and its semantics. First, spatio-temporal video segments of a fixed length are clustered according to image features of each segment. Then, a user performs feedback to the results of clustering, whether each segment is relevant to the cluster it belongs. The clustering accuracy can be improved through the interaction based on the feedback information. For diverse kinds of video streams, we investigated how the feedback should be given and demonstrated the effectiveness of the interactive clustering Yukihiro Kinoshita, Naoko Nitta, Noboru Babaguchi |
ICME | 2 |
| 2005 | Automatic parsing of American football videos by intermodal collaboration based on transition rulesabstractThis paper proposes an automatic American football video parsing method based on transition rules of an American football game. Combining the results of live scene extraction and superimposed text detection based on image features enables us to segment the video into play units of a game. Temporally associating the segmented play units with the detected superimposed texts and the closed-caption text attaches possible semantic content information to the play units. Finally, selecting only the play units which conform to transition rules of the sports game from the obtained play unit sequence, while discarding or complementing unnecessary or insufficient play units and attached semantic content information, realizes the semantic video parsing. Naoko Nitta, Noboru Babaguchi |
ICME | 1 |
| 2005 | Video Summarization for Large Sports Video ArchivesabstractVideo summarization is defined as creating a shorter video clip or a video poster which includes only the important scenes in the original video streams. In this paper, we propose two methods of generating a summary of arbitrary length for large sports video archives. One is to create a concise video clip by temporally compressing the amount of the video data. The other is to provide a video poster by spatially presenting the image keyframes which together represent the whole video content. Our methods deal with the metadata which has semantic descriptions of video content. Summaries are created according to the significance of each video segment which is normalized in order to handle large sports video archives. We experimentally verified the effectiveness of our methods by comparing the results with man-made video summaries Yoshimasa Takahashi, Naoko Nitta, Noboru Babaguchi |
ICME | 2 |
| 2005 | Generating Semantic Descriptions of Broadcasted Sports Videos Based on Structures of Sports Games and TV Programs
Naoko Nitta, Noboru Babaguchi, Tadahiro Kitahashi |
Multim. Tools Appl. | 1 |
| 2003 | Intermodal collaboration: a strategy for semantic content analysis for broadcasted sports videoabstractThis paper presents intermodal collaboration: a strategy for semantic content analysis for broadcasted sports video. The broadcasted video can be viewed as a set of multimodal streams such as visual, auditory, text (closed caption) and graphics streams. Collaborative analysis for the multimodal streams is achieved based on temporal dependency between their streams, in order to improve the reliability and efficiency for semantic content analysis such as extracting highlight scenes from sports video and automatically generating annotations of specific scenes. A couple of case studies are shown to experimentally confirm the effectiveness of intermodal collaboration. Noboru Babaguchi, Naoko Nitta |
ICIP (1) | 2 |
| 2002 | Story based representation for broadcasted sports video and automatic story segmentationabstractThis paper presents a model to represent a broadcasted sports video in a semantical way and proposes a method to segment the sports video into the semantical units for the representation. Representation of a video should clarify its semantical content as accurately as possible. Our model structurizes the video and gives suitable semantical descriptions to its particular time locations based on the structures of the sports video. We consider the speech transcript as the useful information stream for the generation of the semantical descriptions. The proposed method tries to segment the speech transcript into the units of the structurization in a probabilistic framework based on Bayesian networks. Moreover, the association with the image stream enables us to find the corresponding video segments. We discuss some experimental results. Naoko Nitta, Noboru Babaguchi, Tadahiro Kitahashi |
ICME (1) | 1 |
| 2000 | Extracting Actors, Actions and Events from Sports Video - A Fundamental Approach to Story TrackingabstractTo effectively deal with the vast amount of videos, we need to construct a content-based representation for each video. As a step towards this goal, this paper proposes a method to automatically generate the semantical annotations for a sports video by integrating the text(c1osed-caption) and image stream. we first segment the text data and extract segments which are meaningful to grasp the story of the video, then extract the actors, the actions and the events of each scene which are useful for information retrieval by using the linguistic cues and the domain knowledge. We also segment the image stream so that each segment can associate with each text segment extracted above by using the image cues. Finally we can annotate the video by associating the text segments with the image segments. Some experimental results are presented and discussed in this paper. Naoko Nitta, Noboru Babaguchi, Tadahiro Kitahashi |
ICPR | 1 |