Yu Suzuki 0001

dblp:55/6987-2 · DBLP profile ↗
← Back
25ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0002-7373-8677ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 23 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Parameter Drift as a Signal for Membership Inference in Overfit-Tuned LLMs
Takuto Kitamura, Yu Suzuki 0001
DaWaK2
2025 Automated Instruction Generation via Alternating Evaluation and Creation with LLMs
Ryo Tanaka, Yu Suzuki 0001
iiWAS2
2024 Finding Adequate Additional Layer of Auxiliary Task in BERT-Based Multi-task Learning
Takuto Kitamura, Yu Suzuki 0001
iiWAS (1)2
2023 Feature Analysis of Regional Behavioral Facilitation Information Based on Source Location and Target People in Disaster
Kosuke Wakasugi, Futo Yamamoto, Yu Suzuki 0001, Akiyo Nadamoto
DaWaK3
2021 Analysis of Behavioral Facilitation Tweets for Large-Scale Natural Disasters Dataset Using Machine Learning
Yu Suzuki 0001, Yoshiki Yoneda, Akiyo Nadamoto
DEXA (2)1
2019 Detection of Behavioral Facilitation information in Disaster Situation
abstract
Disasters of many types have occurred in recent years, such as strong earthquakes, heavy rain, and typhoons. In such disaster situations, people often use social network services (SNS) and exchange information of all types to help each other. Especially, people exchange information using Twitter during disasters. Such tweet messages include much information that promotes people's behaviors. We designate such tweets as behavioral facilitation tweets. When psychologically unstable in the aftermath of a disaster, behavioral facilitation tweets can strongly affect people, irrespective of a message's authenticity. We regard the extraction of the behavioral facilitation tweets automatically as important. In this paper, we propose a method that extracts behavioral facilitation tweets in disaster situations. Specifically, we propose and compare three methods to extract behavioral facilitation tweets in disaster situations: rule-based, support vector machine (SVM) and long short-term memory (LSTM). Furthermore, we conducted experiments to assess the benefits of our proposed method.
Yoshiki Yoneda, Yu Suzuki 0001, Akiyo Nadamoto
iiWAS2
2018 TRANS-AM: Discovery Method of Optimal Input Vectors Corresponding to Objective Variables
Yu Suzuki 0001, Koichiro Yoshino, Satoshi Nakamura 0001
DaWaK2
2018 Information Filtering Method for Twitter Streaming Data Using Human-in-the-Loop Machine Learning
Yu Suzuki 0001, Satoshi Nakamura 0001
DEXA (2)1
2018 Extracting Japanese Behavioral Facilitation Tweet in Disaster Situations
abstract
There are many disasters such as big earthquakes, flood disasters, hurricanes, and typhoons. Especially in Japan, there are so many natural disasters. Nowadays, after the disaster, we help each other not only in the real world but also on the social media. However, there are many rumors on the social media. Especially, twitter spreads many rumors after the disaster, because it is easy to spread information. According to our former study, there are many behavioral facilitation tweets in the rumors tweets. We consider it is important to extract automatically such rumor behavioral facilitation tweets. In this paper, as a first step of extracting rumor behavioral facilitation tweets, we propose the method which extracts behavioral facilitation tweets from disaster tweets. Our proposed method is rule-based. We also conducted the small experiment to measure our proposed method is suitable for extracting behavioral facilitation tweets.
Keiichi Mizuka, Yu Suzuki 0001, Akiyo Nadamoto
iiWAS2
2018 Dialogue Scenario Collection of Persuasive Dialogue with Emotional Expressions via Crowdsourcing
Koichiro Yoshino, Yoko Ishikawa, Masahiro Mizukami, Yu Suzuki 0001, Sakriani Sakti, Satoshi Nakamura 0001
LREC4
2017 A trade-off between estimation accuracy of worker quality and task complexity
abstract
In crowdsourcing, many people are less capable of producing quality work, and there are those who work inadequately. We can improve the quality of work, and also we can decrease time and wages if we eliminate poor workers and give extra instruction to their workers. Therefore, estimating work quality is essential for uncovering poor workers. In existing studies, the response behavior of workers was used to estimate their quality. However, in these studies, the authors only apply to complicated tasks that have many types of response behavior. In this paper, we propose a method for estimating the quality of workers by their response behavior by intentionally complicating a simple task. By doing so, we can get more accurate and detailed response behavior. By using accurate and detailed response behavior even in simple tasks that have few types of response behavior, the estimation accuracy of low-quality workers improved. However, workers had to work for slightly longer.
Yoshitaka Matsuda, Yu Suzuki 0001, Satoshi Nakamura 0001
IEEE BigData2
2017 Extraction of commentary tweets about news articles
abstract
On Twitter, vast numbers of tweets have been written about news articles. These tweets include not only opinions and sentiments, but also comments related to the news articles. However, tweets that include comments about news article are believed by people even if their credibility is not clear. In this way, these tweets are sometimes spread by others. Therefore, we consider the importance of raising an alarm about tweets for which the credibility is not clear. As described in this paper, as a first step of extracting tweets with unclear credibility, we propose a method to extract tweets that include commentary about news articles. In this paper, we designate the tweets as "commentary tweets". Our proposed method consists of a rule-based component and a machine learning component. We also conducted our experiments to measure the suitability of our proposed method for extracting commentary tweets.
Keiichi Mizuka, Yu Suzuki 0001, Akiyo Nadamoto
iiWAS2
2017 Finding missing tweets using topic structure and browsing time
abstract
Microblogging services such as Twitter and Facebook become popular in recent years. In these services, many users post short messages which correspond to many topics such as daily activities, opinions, and new events. Therefore, users need a system to summarize messages if the users receive tons of messages. If the following users tweet about important things which the user does not know, these tweets should be noticed. However, which tweets should be noticed is one important problem. Users should need which topics are on their timeline. However, if the summarization method does not consider topics of tweets, the summarized tweets do not contain rarely tweeted topics. To solve this problem, we propose a method for automatically extracting missing tweets based on topic granularity and missing time of the users. In this study, we map the missing tweets to the Wikipedia category tree by considering topic structure granularity; then we present the topic structures of missing tweets using our proposed visualization interface. In our experiments, we confirmed the effectiveness of our proposed hierarchal topic structure.
Yu Suzuki 0001, Hiromitsu Ohara, Akiyo Nadamoto
iiWAS1
2017 Information Navigation System with Discovering User Interests
abstract
We demonstrate an information navigation system for sightseeing domains that has a dialogue interface for discovering user interests for tourist activities.The system discovers interests of a user with focus detection on user utterances, and proactively presents related information to the discovered user interest.A partially observable Markov decision process (POMDP)-based dialogue manager, which is extended with user focus states, controls the behavior of the system to provide information with several dialogue acts for providing information.We transferred the belief-update function and the policy of the manager from other system trained on a di↵erent domain to show the generality of defined dialogue acts for our information navigation system.
Koichiro Yoshino, Yu Suzuki 0001, Satoshi Nakamura 0001
SIGDIAL Conference2
2017 Semantically readable distributed representation learning for social media mining
abstract
The problem with distributed representations generated by neural networks is that the meaning of the features is difficult to understand. We propose a new method that gives a specific meaning to each node of a hidden layer by introducing a manually created word semantic vector dictionary into the initial weights and by using paragraph vector models. Our experimental results demonstrated that weights obtained based on learning and weights based on the dictionary are more strongly correlated in a closed test and more weakly correlated in an open test, compared with the results of a control test. Additionally, we found that the learned vector are better than the performance of the existing paragraph vector in the evaluation of the sentiment analysis task. Finally, we determined the readability of document embedding in a user test. The definition of readability in this paper is that people can understand the meaning of large weighted features of distributed representations. A total of 52.4% of the top five weighted hidden nodes were related to tweets where one of the paragraph vector models learned the document embedding. Because each hidden node maintains a specific meaning, the proposed method succeeds in improving readability.
Ikuo Keshi, Yu Suzuki 0001, Koichiro Yoshino, Satoshi Nakamura 0001
WI2
2016 Fast text anonymization using k-anonyminity
abstract
In this paper, we propose a method for anonymizing unstructured texts using a quasi-identifier list. In our method, the system redacts from some parts of quasi-identifiers in the texts to the alternate characters such as "*", in order to prevent re-identification of information which should be kept in secrecy. However, this method has a room for an improvement for keeping the information on the original text as is. If the system anonymizes the texts and keeps the original texts as much as possible, the accuracy of the outputs by data mining techniques for the anonymized texts should be useful. Our method anonymizes quasi-identifiers to remain substrings which do not contribute to re-identification, in order to keep the information on the original texts as is.
Wakana Maeda, Yu Suzuki 0001, Satoshi Nakamura 0001
iiWAS2
2015 Detection of missing tweets based on browsing interval and topic granularity
abstract
Twitter users who browse tweets can follow other users in whom they are interested. They can obtain interesting information from other users' tweets on their timeline. If they follow many users, then they can expect numerous tweets on their timeline. However, if users do not browse their timeline for some time, they can lose interesting and important information. Therefore, a system that automatically presents a summary of lost information can be extremely beneficial. As described herein, we propose a method of extracting lost information automatically based on a user's browsing time interval and the topic structure of a followee's tweets. First, we classify a followee's tweets that contain the user's missing information, and assign topics to the groups. Next, we generate a topic graph based on the semantic structure from Wikipedia. We decide whether the tweet groups are missed using the followee's topic graph based on the browsing time interval. Finally, we extract missing information and present it to the user.
Hiromitsu Ohara, Yu Suzuki 0001, Akiyo Nadamoto
iiWAS2
2013 Complementary Information for Wikipedia by Comparing Multilingual Articles
Yuya Fujiwara, Yu Suzuki 0001, Yukio Konishi, Akiyo Nadamoto
APWeb2
2013 Assessing quality score of Wikipedia article using mutual evaluation of editors and texts
abstract
In this paper, we propose a method for assessing quality scores of Wikipedia articles by mutually evaluating editors and texts. Survival ratio based approach is a major approach to assessing article quality. In this approach, when a text survives beyond multiple edits, the text is assessed as good quality, because poor quality texts have a high probability of being deleted by editors. However, many vandals, low quality editors, delete good quality texts frequently, which improperly decreases the survival ratios of good quality texts. As a result, many good quality texts are unfairly assessed as poor quality. In our method, we consider editor quality score for calculating text quality score, and decrease the impact on text quality by vandals. Using this improvement, the accuracy of the text quality score should be improved. However, an inherent problem with this idea is that the editor quality scores are calculated by the text quality scores. To solve this problem, we mutually calculate the editor and text quality scores until they converge. In this paper, we prove that the text quality score converges. We did our experimental evaluation, and confirmed that our proposed method could accurately assess the text quality scores.
Yu Suzuki 0001, Masatoshi Yoshikawa
CIKM1
2013 Clustering Editors of Wikipedia by Editor's Biases
abstract
Wikipedia is an Internet encyclopedia where any user can edit articles. Because editors act on their own judgments, editors' biases are reflected in edit actions. When editors' biases are reflected in articles, the articles should have low credibility. However, it is difficult for users to judge which parts in articles have biases. In this paper, we propose a method of clustering editors by editors' biases for the purpose that we distinguish texts' biases by using editors' biases and aid users to judge the credibility of each description. If each text is distinguished such as by colors, users can utilize it for the judgments of the text credibility. Our system makes use of the relationships between editors: agreement and disagreement. We assume that editors leave texts written by editors that they agree with, and delete texts written by editors that they disagree with. In addition, we can consider that editors who agree with each other have similar biases, and editors who disagree with each other have different biases. Hence, the relationships between editors enable to classify editors by biases. In experimental evaluation, we verify that our proposed method is useful in clustering editors by biases. Additionally, we validate that considering the dependency between editors improves the clustering performance.
Akira Nakamura, Yu Suzuki 0001, Yoshiharu Ishikawa
Web Intelligence2
2012 Extracting Difference Information from Multilingual Wikipedia
Yuya Fujiwara, Yu Suzuki 0001, Yukio Konishi, Akiyo Nadamoto
APWeb2
2012 Extracting lack of information on Wikipedia by comparing multilingual articles
abstract
Wikipedia has multilingual articles, the information of which differs, even for articles on the same topic. As described in this paper, we propose a system to extract and present lack of information of one language on Wikipedia by comparing two languages on the Wikipedia. When we compare Wikipedia articles of two languages, the granularity of information between them differs. Therefore, we propose a method of extracting multiple comparison articles using a Wikipedia link graph. The system extracts lack of information that is included in articles in Wikipedia by comparing one base article with other articles that are found using the link graph.
Yuya Fujiwara, Yukio Konishi, Yu Suzuki 0001, Akiyo Nadamoto
iiWAS3
2012 Assessing Quality Values of Wikipedia Articles Using Implicit Positive and Negative Ratings
Yu Suzuki 0001
WAIM1
2012 Processing All k-Nearest Neighbor Queries in Hadoop
Takuya Yokoyama, Yoshiharu Ishikawa, Yu Suzuki 0001
WAIM3
2001 A Unified Retrieval Method of Multimedia Documents
abstract
In this paper, we propose a retrieval method for multimedia documents. In our method, features of each medium such as text, image, and their layout information are extracted from documents. For a given query, similarity values are calculated for each media, then those values are integrated into one similarity value per document. We propose four methods for calculating the integrated similarity values. Also, we performed experiments using PDF files, and verified the effectiveness of our method.
Yu Suzuki 0001, Kenji Hatano, Masatoshi Yoshikawa, Shunsuke Uemura
DASFAA1