VLDB 2026 Research / reviewers in the wild / expert
Shotaro Ishihara
dblp:191/9113
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0001-0366-6807ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Making News Familiar: News Recommendation from Daily SceneryabstractAlthough news recommendations tailored to users’ browsing habits are common in news distribution services, issues must be addressed regarding broadening the range of users’ interests. This study focuses on user input images as a means that has not yet been fully investigated to provide new insights to users. Specifically, we assume the input of familiar images of the user’s daily scenery and verify whether the system can arouse interest in topics of which the user was unaware by recommending news articles related to the objects in the images. The implemented system extracts object names from the input images using a vision & language model and displays the search results of news articles. Through user experiments, the implemented system could recommend news articles that satisfied the requirements of serendipity (relevance, novelty, and unexpectedness in this study) at an average rate of 0.12, suggesting that news recommendation using user input images is useful. Kota Tanabe, Shotaro Ishihara, Kenta Yamada, Masaki Aota, Yasutsuna Matayoshi |
KES | 2 |
| 2025 | Should Embedding-Based News Recommendation be Revisited? A Focus on the Differences Between News Publishers and Aggregators
Takumi Tamura, Yoichiro Ito, Masaki Aota, Kenta Yamada, Shotaro Ishihara |
NLDB (2) | 5 |
| 2024 | Quantifying Memorization and Detecting Training Data of Pre-trained Language Models using Japanese NewspaperabstractDominant pre-trained language models (PLMs) have demonstrated the potential risk of memorizing and outputting the training data.While this concern has been discussed mainly in English, it is also practically important to focus on domain-specific PLMs.In this study, we pretrained domain-specific GPT-2 models using a limited corpus of Japanese newspaper articles and evaluated their behavior.Experiments replicated the empirical finding that memorization of PLMs is related to the duplication in the training data, model size, and prompt length, in Japanese the same as in previous English studies.Furthermore, we attempted membership inference attacks, demonstrating that the training data can be detected even in Japanese, which is the same trend as in English.The study warns that domain-specific PLMs, sometimes trained with valuable private data, can "copy and paste" on a large scale. 1 Shotaro Ishihara, Hiromu Takahashi |
INLG | 1 |
| 2023 | Generating News-Centric Crossword Puzzles As A Constraint Satisfaction and Optimization ProblemabstractCrossword puzzles have traditionally served not only as entertainment but also as an educational tool that can be used to acquire vocabulary and language proficiency. One strategy to enhance the educational purpose is personalization, such as including more words on a particular topic. This paper focuses on the case of encouraging people's interest in news and proposes a framework for automatically generating news-centric crossword puzzles. We designed possible scenarios and built a prototype as a constraint satisfaction and optimization problem, that is, containing as many news-derived words as possible. Our experiments reported the generation probabilities and time required under several conditions. The results showed that news-centric crossword puzzles can be generated even with few news-derived words. We summarize the current issues and future research directions through a qualitative evaluation of the prototype. This is the first proposal that a formulation of a constraint satisfaction and optimization problem can be beneficial as an educational application. Kaito Majima, Shotaro Ishihara |
CIKM | 2 |
| 2022 | Analysis and Estimation of News Article Reading Time with Multimodal Machine LearningabstractThis paper highlights the importance of reading time for news media and evaluates the implementation methodology. The display of estimated reading time allows users to select and view articles that are appropriate for their situation. The simplest hypothesis for the implementation is that reading time correlates with text length. We analyzed real-world users of Japanese financial news and revealed that reading time does not strongly correlate with text length. Experiments also showed that a multimodal machine learning approach leads to a more accurate estimation. Specifically, fine-tuning neural networks that incorporated LSTM to process user history and BERT and Swin Transformer to acquire embeddings from the articles achieved the best results. Shotaro Ishihara, Yasufumi Nakama |
IEEE Big Data | 1 |
| 2021 | Editors-in-the-loop News Article Summarization Framework with Sentence Selection and CompressionabstractThis paper proposes a human-in-the-loop framework to summarize news articles by sentence selection and compression. The system enumerates summary candidates by selecting N sentences representing the article based on a quantitative metric and then compressing each sentence by syntactic analysis. Experiments showed that the proposed system was able to extract the same topics as the human editor's work at the rate of 26 %. Even though the rate was not high enough, the proposed framework has the advantage that it is easy to incorporate the editor's intentions in the sentence selection and compression by giving weights. This approach not only has the potential to reduce the burden on editors, but can also contribute to giving them a new perspective in creating summaries. Shotaro Ishihara, Yuta Matsuda, Norihiko Sawa |
IEEE BigData | 1 |