Takeshi Sakaki

dblp:73/8024 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
2since 2021 · last 2023
0000-0002-5830-4352ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Web and social media mining · 55% Information retrieval · 35% Data mining · 10%
Artificial intelligence
1 paper
Information extraction and text analysis · 100%
Human-computer interaction and pervasive computing
1 paper
Ubiquitous computing and smart environments · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Web and social media mining
event detection
0.212013
Tweet Analysis for Real-Time Event Detection and Earthquake Reporting System Development · IEEE Trans. Knowl. Data Eng. 2013
Web and social media mining › event detection
social event detection
0.112010
Earthquake shakes Twitter users: real-time event detection by social sensors · WWW 2010
Natural language and speech › Information extraction and text analysis › lexical semantics
word clustering
0.112006
Graph-based Word Clustering using a Web Search Engine · EMNLP 2006
Information retrieval
web search
0.112006
Graph-based Word Clustering using a Web Search Engine · EMNLP 2006
Data mining › text mining › text classification
tweet classification
0.012013
Tweet Analysis for Real-Time Event Detection and Earthquake Reporting System Development · IEEE Trans. Knowl. Data Eng. 2013
Ubiquitous computing and smart environments
social sensing
0.012010
Earthquake shakes Twitter users: real-time event detection by social sensors · WWW 2010

Methods — techniques the papers use, named apart from their topics

particle filtering · 0.5tweet classifier · 0.3tweet classification · 0.2probabilistic spatiotemporal model · 0.2kalman filtering · 0.2graph-based clustering · 0.1
YearPublicationVenuePosition
2023 Construction of Evaluation Datasets for Trend Forecasting Studies
abstract
In this study, we discuss issues in the traditional evaluation norms of trend forecasts, outline a suitable evaluation method, propose an evaluation dataset construction procedure, and publish Trend Dataset: the dataset we have created. As trend predictions often yield economic benefits, trend forecasting studies have been widely conducted. However, a consistent and systematic evaluation protocol has yet to be adopted. We consider that the desired evaluation method would address the performance of predicting which entity will trend, when a trend occurs, and how much it will trend based on a reliable indicator of the general public's recognition as a gold standard. Accordingly, we propose a dataset construction method that includes annotations for trending status (trending or non-trending), degree of trending (how well it is recognized), and the trend period corresponding to a surge in recognition rate. The proposed method uses questionnaire-based recognition rates interpolated using Internet search volume, enabling trend period annotation on a weekly timescale. The main novelty is that we survey when the respondents recognize the entities that are highly likely to have trended and those that haven't. This procedure enables a balanced collection of both trending and non-trending entities. We constructed the dataset and verified its quality. We confirmed that the interests of entities estimated using Wikipedia information enables the efficient collection of trending entities a priori. We also confirmed that the Internet search volume agrees with public recognition rate among trending entities.
Shogo Matsuno, Sakae Mizuki, Takeshi Sakaki
ICWSM3
2021 Retrospective analysis of controversial topics on COVID-19 in Japan
abstract
For efficient policy-making, a thorough recognition of controversial topics is crucial because the cost of unmitigated controversies would be extremely high for society. However, identifying controversial topics is costly. In this paper, we proposed a framework to search for controversial topics comprehensively. We then conducted a retrospective analysis of the controversial topics of COVID-19 with data obtained via Twitter in Japan as a case study of the framework. The results show that the proposed framework can effectively detect controversial topics that reflect current reality. Controversial topics tend to be about the government, medical matters, economy, and education; moreover, the controversy score had a low correlation with the traditional indicators-scale and sentiment of the topics-which suggests that the controversy score is a potentially important indicator to be obtained. We also discussed the difference between highly controversial topics and less controversial ones despite their large scale and sentiment.
Kunihiro Miyazaki, Takayuki Uchiba, Fujio Toriumi, Takeshi Sakaki
ASONAM5
2019 Comparative evaluation of two approaches for retweet clustering: A text-based method and graph-based method
abstract
Burst phenomena are caused by such social events as flaming on the internet, elections, and natural disasters. To understand people’s thoughts and feelings, we must classify their opinions from burst phenomena. Therefore, classification methods that categorize tweets are critical. However, since most classification methods focus on text mining, they cannot classify tweets by topics because each tweet has poor linguistic similarities. We used a non-text-based method proposed by Baba et al. that groups tweets by topics, even if they have poor linguistic similarities, and verified its validity by comparing it with a text-based method in two different evaluations: full data and sampled data. In the full data evaluation part, we did a questionnaire survey and validated the suitability of the topic clusters created by both classification methods using our full dataset. In the sampled data evaluation part, we focused on the robustness of each method against data reduction. Since collecting the whole data of burst phenomena is very costly due to the vast amounts of available social media data, robustness against data reduction is an important index to evaluate classification methods. After these evaluations, we found that the non-text-based method more effectively classified tweets than the text-based method.
Kazuki Uchida, Fujio Toriumi, Takeshi Sakaki
Web Intell.3
2017 Evaluation of retweet clustering method classification method using retweets on Twitter without text data
abstract
Burst phenomena, which frequently occur on social media, are caused by such social events as flaming on the internet, elections, and natural disasters. To understand people's thoughts and feelings, we must classify their opinions from burst phenomena. Therefore, classification methods that categorize tweets are critical. However, since most classification methods focus on text mining, they cannot group tweets by topics because each tweet has poor linguistic similarities. We used a non-text-based classification method proposed by Baba et al. that groups tweets by topics, even if they have poor linguistic similarities, and verified its validity by comparing it with a text-based classification method in two different evaluations: qualitative and quantitative. In the qualitative evaluation part, we did a questionnaire survey and validated the suitability of the topic clusters created using both the non-and text-based methods. Since evaluating the similarity of every pair of tweets in each topic is difficult, we evaluated the similarity between sampled pairs in the survey and acquired more appropriate topic clustering results using the non-text-based method than the text-based method. In the quantitative evaluation part, we focused on the robustness of each method against data reduction. Many approaches analyze social media data, especially because collecting data from social media is comparatively easy. However, since collecting the whole data of burst phenomena is very costly due to the vast amounts of available social media data, robustness against data reduction is an important index to evaluate classification methods. With the non-text-based method, over 55% of the pairs of tweets in the same cluster were also included in the same cluster even when the data were reduced to 10% in all three of our example cases. In this paper, as a source we focus on Twitter, one of the most popular microblogging services. Using clustering to conduct detailed case analyses, we scrutinized three burst cases that include natural disasters and flaming on the internet and found that a non-text-based method more effectively classified tweets in burst phenomena than a text-based method.
Kazuki Uchida, Fujio Toriumi, Takeshi Sakaki
WI3
2013 Tweet Analysis for Real-Time Event Detection and Earthquake Reporting System Development
abstract
Twitter has received much attention recently. An important characteristic of Twitter is its real-time nature. We investigate the real-time interaction of events such as earthquakes in Twitter and propose an algorithm to monitor tweets and to detect a target event. To detect a target event, we devise a classifier of tweets based on features such as the keywords in a tweet, the number of words, and their context. Subsequently, we produce a probabilistic spatiotemporal model for the target event that can find the center of the event location. We regard each Twitter user as a sensor and apply particle filtering, which are widely used for location estimation. The particle filter works better than other comparable methods for estimating the locations of target events. As an application, we develop an earthquake reporting system for use in Japan. Because of the numerous earthquakes and the large number of Twitter users throughout the country, we can detect an earthquake with high probability (93 percent of earthquakes of Japan Meteorological Agency (JMA) seismic intensity scale 3 or more are detected) merely by monitoring tweets. Our system detects earthquakes promptly and notification is delivered much faster than JMA broadcast announcements.
Takeshi Sakaki, Makoto Okazaki, Yutaka Matsuo
IEEE Trans. Knowl. Data Eng.1
2010 How to Become Famous in the Microblog World
Takeshi Sakaki, Yutaka Matsuo
ICWSM1
2010 Earthquake shakes Twitter users: real-time event detection by social sensors
abstract
Twitter, a popular microblogging service, has received much attention recently. An important characteristic of Twitter is its real-time nature. For example, when an earthquake occurs, people make many Twitter posts (tweets) related to the earthquake, which enables detection of earthquake occurrence promptly, simply by observing the tweets. As described in this paper, we investigate the real-time interaction of events such as earthquakes in Twitter and propose an algorithm to monitor tweets and to detect a target event. To detect a target event, we devise a classifier of tweets based on features such as the keywords in a tweet, the number of words, and their context. Subsequently, we produce a probabilistic spatiotemporal model for the target event that can find the center and the trajectory of the event location. We consider each Twitter user as a sensor and apply Kalman filtering and particle filtering, which are widely used for location estimation in ubiquitous/pervasive computing. The particle filter works better than other comparable methods for estimating the centers of earthquakes and the trajectories of typhoons. As an application, we construct an earthquake reporting system in Japan. Because of the numerous earthquakes and the large number of Twitter users throughout the country, we can detect an earthquake with high probability (96% of earthquakes of Japan Meteorological Agency (JMA) seismic intensity scale 3 or more are detected) merely by monitoring tweets. Our system detects earthquakes promptly and sends e-mails to registered users. Notification is delivered much faster than the announcements that are broadcast by the JMA.
Takeshi Sakaki, Makoto Okazaki, Yutaka Matsuo
WWW1
2006 Graph-based Word Clustering using a Web Search Engine
Yutaka Matsuo, Takeshi Sakaki, Koki Uchiyama, Mitsuru Ishizuka
EMNLP2