Shotaro Ishihara

dblp:191/9113 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0001-0366-6807ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Making News Familiar: News Recommendation from Daily Scenery
abstract
Although news recommendations tailored to users’ browsing habits are common in news distribution services, issues must be addressed regarding broadening the range of users’ interests. This study focuses on user input images as a means that has not yet been fully investigated to provide new insights to users. Specifically, we assume the input of familiar images of the user’s daily scenery and verify whether the system can arouse interest in topics of which the user was unaware by recommending news articles related to the objects in the images. The implemented system extracts object names from the input images using a vision & language model and displays the search results of news articles. Through user experiments, the implemented system could recommend news articles that satisfied the requirements of serendipity (relevance, novelty, and unexpectedness in this study) at an average rate of 0.12, suggesting that news recommendation using user input images is useful.
Kota Tanabe, Shotaro Ishihara, Kenta Yamada, Masaki Aota, Yasutsuna Matayoshi
KES2
2025 Should Embedding-Based News Recommendation be Revisited? A Focus on the Differences Between News Publishers and Aggregators
Takumi Tamura, Yoichiro Ito, Masaki Aota, Kenta Yamada, Shotaro Ishihara
NLDB (2)5
2024 Quantifying Memorization and Detecting Training Data of Pre-trained Language Models using Japanese Newspaper
abstract
Dominant pre-trained language models (PLMs) have demonstrated the potential risk of memorizing and outputting the training data.While this concern has been discussed mainly in English, it is also practically important to focus on domain-specific PLMs.In this study, we pretrained domain-specific GPT-2 models using a limited corpus of Japanese newspaper articles and evaluated their behavior.Experiments replicated the empirical finding that memorization of PLMs is related to the duplication in the training data, model size, and prompt length, in Japanese the same as in previous English studies.Furthermore, we attempted membership inference attacks, demonstrating that the training data can be detected even in Japanese, which is the same trend as in English.The study warns that domain-specific PLMs, sometimes trained with valuable private data, can "copy and paste" on a large scale. 1
Shotaro Ishihara, Hiromu Takahashi
INLG1
2023 Generating News-Centric Crossword Puzzles As A Constraint Satisfaction and Optimization Problem
abstract
Crossword puzzles have traditionally served not only as entertainment but also as an educational tool that can be used to acquire vocabulary and language proficiency. One strategy to enhance the educational purpose is personalization, such as including more words on a particular topic. This paper focuses on the case of encouraging people's interest in news and proposes a framework for automatically generating news-centric crossword puzzles. We designed possible scenarios and built a prototype as a constraint satisfaction and optimization problem, that is, containing as many news-derived words as possible. Our experiments reported the generation probabilities and time required under several conditions. The results showed that news-centric crossword puzzles can be generated even with few news-derived words. We summarize the current issues and future research directions through a qualitative evaluation of the prototype. This is the first proposal that a formulation of a constraint satisfaction and optimization problem can be beneficial as an educational application.
Kaito Majima, Shotaro Ishihara
CIKM2
2022 Analysis and Estimation of News Article Reading Time with Multimodal Machine Learning
abstract
This paper highlights the importance of reading time for news media and evaluates the implementation methodology. The display of estimated reading time allows users to select and view articles that are appropriate for their situation. The simplest hypothesis for the implementation is that reading time correlates with text length. We analyzed real-world users of Japanese financial news and revealed that reading time does not strongly correlate with text length. Experiments also showed that a multimodal machine learning approach leads to a more accurate estimation. Specifically, fine-tuning neural networks that incorporated LSTM to process user history and BERT and Swin Transformer to acquire embeddings from the articles achieved the best results.
Shotaro Ishihara, Yasufumi Nakama
IEEE Big Data1
2021 Editors-in-the-loop News Article Summarization Framework with Sentence Selection and Compression
abstract
This paper proposes a human-in-the-loop framework to summarize news articles by sentence selection and compression. The system enumerates summary candidates by selecting N sentences representing the article based on a quantitative metric and then compressing each sentence by syntactic analysis. Experiments showed that the proposed system was able to extract the same topics as the human editor's work at the rate of 26 %. Even though the rate was not high enough, the proposed framework has the advantage that it is easy to incorporate the editor's intentions in the sentence selection and compression by giving weights. This approach not only has the potential to reduce the burden on editors, but can also contribute to giving them a new perspective in creating summaries.
Shotaro Ishihara, Yuta Matsuda, Norihiko Sawa
IEEE BigData1