VLDB 2026 Research / reviewers in the wild / expert
Marie Katsurai
dblp:04/11276
· DBLP profile ↗
18ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0003-4899-2427ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Quality assessment of synthetic images via spatial distortion recognition
Tomoya Sawada, Marie Katsurai, Masashi Okubo |
Mach. Vis. Appl. | 2 |
| 2024 | Fine-Grained Classification of Researcher-Related Web Pages Using String and Contextual FeaturesabstractProfiling a researcher's expertise and interests is crucial for building an expert recommendation system in the academic domain. Conventional methods have used bibliographic databases to obtain textual data regarding a researcher. However, information about academic activities beyond writing academic papers is often scattered across various web pages. To enrich the information sources of researcher profiling, this paper presents a novel task, fine-grained classification of researcher-related web pages obtained using a search query comprising a full name and an affiliation string. Our method fuses two neural networks based on pre-trained embedding models to extract string and contextual features from the URL and page text, respectively. Experiments conducted using a dataset of 477 Japanese researchers demonstrated the effectiveness of the proposed method compared to baseline and conventional methods. Hiroo Hayashi, Marie Katsurai |
COMPSAC | 2 |
| 2024 | Illustrated character face super-deformation via unsupervised image-to-image translation
Tomoya Sawada, Marie Katsurai |
Multim. Syst. | 2 |
| 2022 | First Experimental Results on Real-Time Cleaning Activity Monitoring SystemabstractThe COVID-19 pandemic has presented social challenges to establish the new normal lifestyle in our daily lives. The goal of this paper is to enable easy and low-cost monitoring of cleaning activity to keep a clean environment for preventing infection. Although human activity recognition has been a hot research topic in pervasive computing, existing schemes have not been optimized for monitoring cleaning activities. To address this issue, this paper provides an initial concept and preliminary experimental results of cleaning activity recognition using accelerometer data and RFID tags. In the proposed scheme, machine learning technologies and short range wireless communication are employed for recognizing the time and place of wiping as an example of cleaning activities, because it is an important activity for shared places to avoid infection. This paper reports the evaluation results on the recognition accuracy using the proof-of-concept (PoC) implementation to clarify the required sampling rate and time-window size for further experiments. Also, a real-time feedback system is implemented to provide the monitoring results for users. The proposed scheme contributes for efficient monitoring of cleaning activities for creating the new normal era. Ryo Yaegashi, Yu Nakayama, Moe Matsuki, Ryoma Yasunaga, Marie Katsurai |
CCNC | 5 |
| 2022 | SolutionTailor: Scientific Paper Recommendation Based on Fine-Grained Abstract Analysis
Tetsuya Takahashi 0002, Marie Katsurai |
ECIR (2) | 2 |
| 2022 | Fashion Style-Aware Embeddings for Clothing Image RetrievalabstractClothing image retrieval is becoming increasingly important as users on social media grow to enjoy sharing their daily outfits. Most conventional methods offer single query-based retrieval and depend on visual features learnt via target classification training. This paper presents an embedding learning framework that uses novel style description features available on users' posts, allowing image-based and multiple choice-based queries for practical clothing image retrieval. Specifically, the proposed method exploits the following complementary information for representing fashion styles: season tags, style tags, users' heights, and silhouette descriptions. Then, we learn embeddings based on a quadruplet loss that considers the ranked pairings of the visual features and the proposed style description features, enabling flexible outfit search based on either of these two types of features as queries. Experiments conducted on WEAR posts demonstrated the effectiveness of the proposed method compared with several baseline methods. Rino Naka, Marie Katsurai, Keisuke Yanagi, Ryosuke Goto |
ICMR | 2 |
| 2022 | Anime-to-real clothing: Cosplay costume generation via image-to-image translationabstractAbstract Cosplay has grown from its origins at fan conventions into a billion-dollar global dress phenomenon. To facilitate the imagination and reinterpretation of animated images as real garments, this paper presents an automatic costume-image generation method based on image-to-image translation. Cosplay items can be significantly diverse in their styles and shapes, and conventional methods cannot be directly applied to the wide variety of clothing images that are the focus of this study. To solve this problem, our method starts by collecting and preprocessing web images to prepare a cleaned, paired dataset of the anime and real domains. Then, we present a novel architecture for generative adversarial networks (GANs) to facilitate high-quality cosplay image generation. Our GAN consists of several effective techniques to bridge the two domains and improve both the global and local consistency of generated images. Experiments demonstrated that, with quantitative evaluation metrics, the proposed GAN performs better and produces more realistic images than conventional methods. Our codes and pretrained model are available on the web. Koya Tango, Marie Katsurai, Hayato Maki, Ryosuke Goto |
Multim. Tools Appl. | 2 |
| 2021 | Joint Computation Offloading and Sampling Interval Optimization for Accuracy-Guaranteed SurveillanceabstractA key aspect to realize Internet of things applications such as industry automation and smart agriculture is to enable realtime and networked automatic monitoring via cloud computing and computer vision. However, to design a networked monitoring system, it is necessary to realize a balance between the monitoring accuracy and monitoring cost, for instance, between the network traffic to transmit images and the computation load. Although the monitoring cost can be decreased by increasing the sampling interval of cameras, it becomes more likely that informative images cannot be obtained; in other words, the monitoring accuracy decreases with a reduction in the amount of data. Moreover, although on-device image processing can decrease the network traffic, a large computation delay may be incurred, limiting the sampling rate of the monitoring system. The objective of this study was to examine the balance between the monitoring accuracy and cost and to develop a joint optimization technique for the sampling interval and computation offloading to minimize the monitoring cost in a networked monitoring system while ensuring a high monitoring accuracy. The main contributions of this paper are that we prove the joint optimization problem can be solved explicitly and to develop an algorithm to obtain the solution of the joint optimization problem. The simulation results demonstrated that the proposed algorithm can reduce the monitoring cost by 24-48% while maximizing the number of nodes ensured to achieve high monitoring accuracy. Takayuki Nishio, Yoshiaki Inoue, Yu Nakayama, Marie Katsurai |
CCNC | 4 |
| 2021 | A Visualization Interface for Exploring Similar Brands on a Fashion E-Commerce PlatformabstractWith the market expansion of fashion e-commerce platforms and the emergence of social media services, people have more opportunities to discover diverse brands. However, their concepts are difficult to identify by simply looking at their names. This work-in-progress paper presents a novel web interface for facilitating online fashion shopping in which users can explore similar brands in terms of fashion styles and prices. The results of the experiments demonstrate that the proposed method achieves better visualization of style-based similarity than baseline methods. Natsuki Hashimoto, Marie Katsurai, Ryosuke Goto |
ICWS | 2 |
| 2021 | Selective Classification of Danmaku Comments Using Distributed RepresentationsabstractDanmaku commenting has become popular for co-viewing on video-sharing platforms. However, there are usually a large number of irrelevant comments, that contaminate the quality of the information provided by videos. To address this problem, this paper presents a novel approach of classifying Danmaku comments into video categories. Specifically, we use BERT as the backbone architecture to extract semantic features from comments. We introduce a loss function that has an abstention option, which enables the detection of comments that do not fall into any predefined category. The experiments that we conducted using Nicovideo data demonstrated that our selective classification approach effectively discarded those that were irrelevant to a video’s content. We also present a method for subdividing the existing video categories based on the results of Danmaku comment classification. This entails a potential application of our method in hierarchical video clustering. Koshiro Tamura, Marie Katsurai |
iiWAS | 2 |
| 2020 | A Deep Multimodal Approach for Map Image ClassificationabstractMap images (e.g., illustrated maps, historical maps, and geographic maps) have been published around the world, not only for giving location but also to attract tourists or hand down the histories of locations. The management of map data, however, has been an open issue for several research fields, including digital library, humanities, and tourism studies. This paper explores an approach for classifying diverse map images by their themes using map content features. Specifically, we present a novel strategy for preprocessing text data that are positioned inside the map images, which are extracted using OCR. The activation of the textual feature-based model is joint with the visual features in an early fusion manner. Finally, we train a classifier model comprising a convolutional layer and a fully connected layer, which predicts the belonging class of the input map. In experiments conducted on a new labeled dataset of map images, we demonstrate that our approach that uses the fused features achieved the best classification performance over single modality. We have made our dataset available on the Internet to facilitate this new task. Tomoya Sawada, Marie Katsurai |
ICASSP | 2 |
| 2018 | Investigating the Consistency of Emoji Sentiment Lexicons Constructed Using Diferent LanguagesabstractEmojis have been widely used in recent text-based communications and can be important features for sentiment analysis of social media posts such as tweets. Our previous work presented a method for automatically constructing an emoji sentiment lexicon based on the co-occurrence frequency between sentiment words and emojis. This paper investigates whether the proposed framework can be valid over different languages. To conduct this study, we constructed two datasets comprising 150,000 tweets written in English and Japanese and applied a sentiment dictionary to each dataset separately. The results demonstrated that the two constructed lexicons were almost consistent, while the sentiments of some emojis were differently interpreted depending on the language. We also showed the reasonable performance of tweet sentiment analysis using each emoji sentiment lexicon. Mayu Kimura, Marie Katsurai |
iiWAS | 2 |
| 2017 | Automatic Construction of an Emoji Sentiment LexiconabstractEmojis have been frequently used to express users' sentiments, emotions, and feelings in text-based communication. To facilitate sentiment analysis of users' posts, an emoji sentiment lexicon with positive, neutral, and negative scores has been recently constructed using manually labeled tweets. However, the number of emojis listed in the lexicon is smaller than that of currently existing emojis, and expanding the lexicon manually requires time and effort to reconstruct the labeled dataset. This paper presents a simple and efficient method for automatically constructing an emoji sentiment lexicon with arbitrary sentiment categories. The proposed method extracts sentiment words from WordNet-Affect and calculates the cooccurrence frequency between the sentiment words and each emoji. Based on the ratio of the number of occurrences of each emoji among the sentiment categories, each emoji is assigned a multidimensional vector whose elements indicate the strength of the corresponding sentiment. In experiments conducted on a collection of tweets, we show a high correlation between the conventional lexicon and our lexicon for three sentiment categories. We also show the results for a new lexicon constructed with additional sentiment categories. Mayu Kimura, Marie Katsurai |
ASONAM | 2 |
| 2017 | Recipe Popularity Prediction with Deep Visual-Semantic FusionabstractPredicting the popularity of user-created recipes has great potential to be adopted in several applications on recipe-sharing websites. To ensure timely prediction when a recipe is uploaded, a prediction model needs to be trained based on the recipe's content features (i.e., its visual and semantic features). This paper presents a novel approach to predicting recipe popularity using deep visual-semantic fusion. We first pre-train a deep model that predicts the popularity of recipes based on each single modality. We insert additional layers to the two models and concatenate their activations. Finally, we train a network comprising fully connected (FC) layers on the fused features to learn more powerful features, which are used for training a regressor. Based on experiments conducted on more than 150K recipes collected from the Cookpad website, we present a comprehensive comparison with several baselines to verify the effectiveness of our method. The best practice for the proposed method is also described. Satoshi Sanjo, Marie Katsurai |
CIKM | 2 |
| 2016 | Image sentiment analysis using latent correlations among visual, textual, and sentiment viewsabstractAs Internet users increasingly post images to express their daily sentiment and emotions, the analysis of sentiments in user-generated images is of increasing importance for developing several applications. Most conventional methods of image sentiment analysis focus on the design of visual features, and the use of text associated to the images has not been sufficiently investigated. This paper proposes a novel approach that exploits latent correlations among multiple views: visual and textual views, and a sentiment view constructed using SentiWordNet. In the proposed method, we find a latent embedding space in which correlations among the three views are maximized. The projected features in the latent space are used to train a sentiment classifier, which considers the complementary information from different views. Results of experiments conducted on Flickr and Instagram images show that our approach achieves better sentiment classification accuracy than methods that use a single modality only and the state-of-the art method that jointly uses multiple modalities. Marie Katsurai, Shin'ichi Satoh 0001 |
ICASSP | 1 |
| 2014 | A Cross-Modal Approach for Extracting Semantic Relationships Between Concepts Using Tagged ImagesabstractThis paper presents a cross-modal approach for extracting semantic relationships between concepts using tagged images. In the proposed method, we first project both text and visual features of the tagged images to a latent space using canonical correlation analysis (CCA). Then, under the probabilistic interpretation of CCA, we calculate a representative distribution of the latent variables for each concept. Based on the representative distributions of the concepts, we derive two types of measures: the semantic relatedness between the concepts and the abstraction level of each concept. Because these measures are derived from a cross-modal scheme that enables the collaborative use of both text and visual features, the semantic relationships can successfully reflect semantic and visual contexts. Experiments conducted on tagged images collected from Flickr show that our measures are more coherent to human cognition than the conventional measures that use either text or visual features, or the WordNet-based measures. In particular, a new measure of semantic relatedness, which satisfies the triangle inequality, obtains the best results among different distance measures in our framework. The applicability of our measures to multimedia-related tasks such as concept clustering, image annotation and tag recommendation is also shown in the experiments. Marie Katsurai, Takahiro Ogawa 0001, Miki Haseyama |
IEEE Trans. Multim. | 1 |
| 2013 | Exploring and visualizing tag relationships in photo sharing websites based on distributional representationsabstractThis paper presents a method for exploring and visualizing tag relationships in photo sharing websites based on distributional representations of tags. First, we find a representative distribution of a tag, which is summarized by the mean and covariance, using features of tagged photos. This distributional representation can jointly consider the semantic meaning of tags and their abstraction levels. Then, based on the representative distributions, we derive two kinds of semantic measures on tag relationships. The extracted information is visualized in a graphical network to facilitate the understanding of tag usage. Experiments conducted using tagged photos collected from Flickr show that our tag network is more coherent to human cognition than other networks constructed by conventional methods. Marie Katsurai, Miki Haseyama |
ICASSP | 1 |
| 2012 | A cross-modal approach for extracting semantic relationships of concepts from an image databaseabstractThis paper presents a cross-modal approach for extracting semantic relationships of concepts from an image database. First, canonical correlation analysis (CCA) is used to capture the cross-modal correlations between visual features and tag features in the database. Then, in order to measure inter-concept relationships and estimate semantic levels, the proposed method focuses on the distributions of images under the probabilistic interpretation of CCA. Results of experiments conducted by using an image database showed the improvement of the proposed method over existing methods. Marie Katsurai, Takahiro Ogawa 0001, Miki Haseyama |
ICASSP | 1 |