Akiko Aizawa

dblp:63/1600 · also Akiko N. Aizawa, Akiko Nakahara · DBLP profile ↗
← Back
22ranked-venue papers in the field
2as first author
8since 2021 · last 2026
0000-0001-6544-5076ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 18 (2 first)Database Systems & Data Management · 3Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 Aspect-Aware Content-Based Recommendations for Mathematical Research Papers
abstract
Content-based research paper recommendation (CbRPR) has seen advances in computer science and biomedicine, but remains unexplored for mathematics, where paper relatedness is more conceptual than explicit textual or citation-based similarity. Mathematics papers may be connected through shared proof techniques, logical implications, or natural generalizations, yet exhibit minimal textual or citation overlap, rendering existing CbRPR ineffective. To address this gap, we first conduct an expert-driven study characterizing mathematical recommendations, revealing that relevance is inherently \textit{aspect}-driven. Grounded in this insight, we introduce GoldRiM (small, expert-annotated) and SilverRiM (large, automatically derived), the first datasets for \textit{aspect}-aware CbRPR in mathematics. Recognizing that LLM embeddings of mathematical content alone yield suboptimal representation, we propose AchGNN, an \textit{aspect}-conditioned heterogeneous GNN that jointly models textual semantics, citation structure, and author lineage. Across GoldRiM and SilverRiM, AchGNN consistently outperforms prior \textit{aspect}-based CbRPR methods, achieving substantial gains across all evaluated \textit{aspects}. We conduct ablation studies to analyze the contributions of individual \textit{aspect} supervision, authorship lineage, and graph-structural signals to AchGNN's performance. To assess domain generality, we further evaluate AchGNN on the \textit{Papers with Code} dataset of machine learning publications, demonstrating that our \textit{aspect}-aware approach effectively transfers beyond mathematics. We deploy our system on the MaRDI platform to help mathematicians with recommendations and release datasets and code publicly for reproducibility.
Ankit Satpute, André Greiner-Petter, Noah Gießing, Olaf Teschke, Moritz Schubotz, Akiko Aizawa, Bela Gipp
SIGIR6
2025 Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
Kon Woo Kim, Rezarta Islamaj Dogan, Jin-Dong Kim, Florian Boudin, Akiko Aizawa
NLDB (2)5
2024 TWOLAR: A TWO-Step LLM-Augmented Distillation Method for Passage Reranking
Davide Baldelli, Akiko Aizawa, Paolo Torroni
ECIR (1)3
2024 Taxonomy of Mathematical Plagiarism
Ankit Satpute, André Greiner-Petter, Noah Gießing, Isabel Beckenbach, Moritz Schubotz, Olaf Teschke, Akiko Aizawa, Bela Gipp
ECIR (4)7
2024 Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange
abstract
Large Language Models (LLMs) have demonstrated exceptional capabilities in various natural language tasks, often achieving performances that surpass those of humans. Despite these advancements, the domain of mathematics presents a distinctive challenge, primarily due to its specialized structure and the precision it demands. In this work, we follow a two-step approach to investigating the proficiency of LLMs in answering mathematical questions. First, we employ the most effective LLMs, as identified by their performance on math question-answer benchmarks, to generate answers to 78 questions from the Math Stack Exchange (MSE). Second, a case analysis is conducted on the LLM that showed the highest performance, focusing on the quality and accuracy of its answers through manual evaluation. We found that GPT-4 performs best (nDCG of 0.48 and P@10 of 0.37) amongst existing LLMs fine-tuned for answering mathematics questions and outperforms the current best approach on ArqMATH3 Task1, considering P@10. Our case analysis indicates that while GPT-4 can generate relevant answers, it isn't consistently accurate. This paper explores the current limitations of LLMs in navigating complex mathematical question-answering. We make our code and findings publicly available for research: https://github.com/gipplab/LLM-Investig-MathStackExchange
Ankit Satpute, Noah Gießing, André Greiner-Petter, Moritz Schubotz, Olaf Teschke, Akiko Aizawa, Bela Gipp
SIGIR6
2023 Evaluating the Effect of Letter Case on Named Entity Recognition Performance
Tuan-An Dao, Akiko Aizawa
NLDB2
2023 Introducing MBIB - The First Media Bias Identification Benchmark Task and Dataset Collection
abstract
Although media bias detection is a complex multi-task problem, there is, to date, no unified benchmark grouping these evaluation tasks. We introduce the Media Bias Identification Benchmark (MBIB), a comprehensive benchmark that groups different types of media bias (e.g., linguistic, cognitive, political) under a common framework to test how prospective detection techniques generalize. After reviewing 115 datasets, we select nine tasks and carefully propose 22 associated datasets for evaluating media bias detection techniques. We evaluate MBIB using state-of-the-art Transformer techniques (e.g., T5, BART). Our results suggest that while hate speech, racial bias, and gender bias are easier to detect, models struggle to handle certain bias types, e.g., cognitive and political bias. However, our results show that no single technique can outperform all the others significantly.We also find an uneven distribution of research interest and resource allocation to the individual tasks in media bias. A unified benchmark encourages the development of more robust systems and shifts the current paradigm in media bias detection evaluation towards solutions that tackle not one but multiple media bias types simultaneously.
Martin Wessel, Tomás Horych, Terry Ruas, Akiko Aizawa, Bela Gipp, Timo Spinde
SIGIR4
2021 NumER: A Fine-Grained Numeral Entity Recognition Dataset
Thanakrit Julavanich, Akiko Aizawa
NLDB2
2020 Discovering Mathematical Objects of Interest - A Study of Mathematical Notations
abstract
Mathematical notation, i.e., the writing system used to communicate concepts in mathematics, encodes valuable information for a variety of information search and retrieval systems. Yet, mathematical notations remain mostly unutilized by today’s systems. In this paper, we present the first in-depth study on the distributions of mathematical notation in two large scientific corpora: the open access arXiv (2.5B mathematical objects) and the mathematical reviewing service for pure and applied mathematics zbMATH (61M mathematical objects). Our study lays a foundation for future research projects on mathematical information retrieval for large scientific corpora. Further, we demonstrate the relevance of our results to a variety of use-cases. For example, to assist semantic extraction systems, to improve scientific search engines, and to facilitate specialized math recommendation systems.
André Greiner-Petter, Moritz Schubotz, Fabian Müller 0002, Corinna Breitinger, Howard S. Cohl, Akiko Aizawa, Bela Gipp
WWW6
2019 Mathematical Variable Detection in PDF Scientific Documents
Bui Hai Phong, Thang Manh Hoang, Thi-Lan Le, Akiko Aizawa
ACIIDS (2)4
2019 Coner: A Collaborative Approach for Long-Tail Named Entity Recognition in Scientific Publications
Daniel Vliegenthart, Sepideh Mesbah, Christoph Lofi, Akiko Aizawa, Alessandro Bozzon
TPDL4
2018 A comprehensive study: Sentence compression with linguistic knowledge-enhanced gated neural network
Xiaoyu Shen 0001, Hajime Senuma, Akiko Aizawa
Data Knowl. Eng.4
2017 Detecting In-line Mathematical Expressions in Scientific Documents
abstract
One of the issues in extracting natural language sentences from PDF documents is the identification of non-textual elements in a sentence. In this paper, we report our preliminary results on the identification of in-line mathematical expressions. We first construct a manually annotated corpus and apply conditional random field (CRF) for the math-zone identification using both layout features, such as font types, and linguistic features, such as context n-grams, obtained from PDF documents. Although our method is naive and uses a small amount of annotated training data, our method achieved an 88.95% F-measure compared with 22.81% for existing math OCR software.
Kenichi Iwatsuki, Takeshi Sagara, Tadayoshi Hara, Akiko Aizawa
DocEng4
2017 Integration of the Scientific Recommender System Mr. DLib into the Reference Manager JabRef
Stefan P. Feyer, Sophie Siebert, Bela Gipp, Akiko Aizawa, Jöran Beel
ECIR4
2017 Gated Neural Network for Sentence Compression Using Linguistic Knowledge
Hajime Senuma, Xiaoyu Shen 0001, Akiko Aizawa
NLDB4
2017 Utilizing dependency relationships between math expressions in math IR
Giovanni Yoko Kristianto, Goran Topic, Akiko Aizawa
Inf. Retr. J.3
2015 Document Layout Optimization with Automated Paraphrasing
abstract
We introduce a new concept in document layout optimization. In our approach, paraphrase-based~layout~optimization, layout issues (e.g. widows due to poor page breaking) are automatically fixed by rewording the neighboring sentences. Techniques of paraphrasing are borrowed from the field of natural language processing towards this goal, which is the first attempt in the field of document engineering. We implemented a prototype TeX pre/post-processing system that includes two simple paraphrase generators. The experiment shows that our approach is promising and effective for improving document layout.
Yusuke Kido, Hikaru Yokono, Goran Topic, Akiko Aizawa
DocEng4
2014 Adding Twitter-specific features to stylistic features for classifying tweets by user type and number of retweets
abstract
Recently, Twitter has received much attention, both from the general public and researchers, as a new method of transmitting information. Among others, the number of retweets (RTs) and user types are the two important items of analysis for understanding the transmission of information on Twitter. To analyze this point, we applied text classification and feature extraction experiments using random forests machine learning with conventional stylistic and Twitter‐specific features. We first collected tweets from 40 accounts with a high number of followers and created tweet texts from 28,756 tweets. We then conducted 15 types of classification experiments using a variety of combinations of features such as function words, speech terms, Twitter's descriptive grammar, and information roles. We deliberately observed the effects of features for classification performance. The results indicated that class classification per user indicated the best performance. Furthermore, we observed that certain features had a greater impact on classification. In the case of the experiments that assessed the level of RT quantity, information roles had an impact. In the case of user experiments, important features, such as the honorific postpositional particle and auxiliary verbs, such as “desu” and “masu,” had an impact. This research clarifies the features that are useful for categorizing tweets according to the number of RTs and user types.
Yui Arakawa, Akihiro Kameda, Akiko Aizawa, Takafumi Suzuki
J. Assoc. Inf. Sci. Technol.3
2008 Multi-class Named Entity Recognition Via Bootstrapping with Dependency Tree-Based Patterns
Van B. Dang, Akiko Aizawa
PAKDD2
2003 An information-theoretic perspective of tf-idf measures
Akiko Aizawa
Inf. Process. Manag.1
2000 The feature quantity: an information theoretic perspective of Tfidf-like measure
abstract
The feature quantity, a quantitative representation of specificity introduced in this paper, is based on an information theoretic perspective of co-occurrence events between terms and documents. Mathematically, the feature quantity is defined as a product of probability and information, and maintains a good correspondence with the tfidf-like measures popularly used in today's IR systems. In this paper, we present a formal description of the feature quantity, as well as some illustrative examples of applying such a quantity to different types of information retrieval tasks: representative term selection and text categorization.
Akiko Aizawa
SIGIR1
1995 Genetics-Based Learning of New Heuristics: Rational Scheduling of Experiments and Generalization
abstract
We present new methods for the automated learning of heuristics in knowledge lean applications and for finding heuristics that can be generalized to unlearned domains. These applications lack domain knowledge for credit assignment; hence, operators for composing new heuristics are generally model free, domain independent, and syntactic in nature. The operators we have used are genetics based; examples of which include mutation and cross over. Learning is based on a generate and test paradigm that maintains a pool of competing heuristics, tests them to a limited extent, creates new ones from those that perform well in the past, and prunes poor ones from the pool. We have studied three important issues in learning better heuristics: anomalies in performance evaluation; rational scheduling of limited computational resources in testing candidate heuristics in single objective as well as multiobjective learning; and finding heuristics that can be generalized to unlearned domains. We show experimental results in learning better heuristics for: process placement for distributed memory multicomputers, node decomposition in a branch and bound search, generation of test patterns in VLSI circuit testing, and VLSI cell placement and routing.>
Benjamin W. Wah, Arthur Ieumwananonthachai, Lon-Chan Chu, Akiko Aizawa
IEEE Trans. Knowl. Data Eng.4