VLDB 2026 Research / reviewers in the wild / expert
Terry Ruas
dblp:194/4099 · also Terry Lima Ruas
· DBLP profile ↗
26ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-9440-780XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 2 first-author · 19 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment AnalysisabstractLung-Hao Lee, Liang-Chih Yu, Natalia V Loukachevitch, Ilseyar Alimova, Alexander Panchenko, Tzu-Mi Lin, Zhe-Yu Xu, Jian-Yu Zhou, Guangmin Zheng, Jin Wang, Sharanya Awasthi, Jonas Becker, Jan Philip Wahle, Terry Ruas, Shamsuddeen Hassan Muhammad, Saif M. Mohammad. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lung-Hao Lee, Liang-Chih Yu, Natalia V. Loukachevitch, Ilseyar Alimova, Alexander Panchenko, Tzu-Mi Lin, Zhe-Yu Xu, Jian-Yu Zhou, Guangmin Zheng 0001, Jin Wang 0008, Sharanya Awasthi, Jonas Becker, Jan Philip Wahle, Terry Ruas, Shamsuddeen Hassan Muhammad, Saif M. Mohammad |
ACL (1) | 14 |
| 2026 | ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long DocumentsabstractWhile Vision-language models (VLMs) interpret text-rich images effectively, they struggle with reasoning across long, multi-page documents.We present Active Long-DocumEnt Navigation (ALDEN), a multi-turn reinforcement learning framework that fine-tunes VLMs as interactive agents capable of actively navigating long, visually rich documents rather than passive readers.ALDEN features a novel fetch action that allows direct page indexing, complementing the classic search action and better exploiting document structure.To ensure training efficiency and stability, we introduce a rule-based cross-level reward for dense supervision and a visual-semantic anchoring mechanism using dual-path KL-divergence constraints.We train ALDEN on a curated corpus built from open-source datasets, filtering out trivial samples and rewriting queries to incentivize multi-turn navigation and fetch usage.Empirically, ALDEN achieves state-of-the-art results on five long-document benchmarks, offering a more accurate and efficient path for long-document understanding.All of our code and datasets will be made publicly available at Github 1 . Terry Ruas, Yijun Tian 0001, Jan Philip Wahle, Daniel Kurzawe, Bela Gipp |
ACL (1) | 2 |
| 2026 | Piecing Together Cross-Document Coreference Resolution Datasets: Systematic Dataset Analysis and Unification
Anastasia Zhukova, Terry Ruas, Jan Philip Wahle, Bela Gipp |
LREC | 2 |
| 2025 | BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 LanguagesabstractShamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava, Aura Cristina Udrea, Lilian Diana Awuor Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou, Saif M. Mohammad. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufiño, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava 0001, Aura Cristina Udrea, Lilian Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou 0019, Saif M. Mohammad |
ACL (1) | 5 |
| 2025 | What's Wrong? Refining Meeting Summaries with LLM FeedbackabstractMeeting summarization has become a critical task since digital encounters have become a common practice. Large language models (LLMs) show great potential in summarization, offering enhanced coherence and context understanding compared to traditional methods. However, they still struggle to maintain relevance and avoid hallucination. We introduce a multi-LLM correction approach for meeting summarization using a two-phase process that mimics the human review process: mistake identification and summary refinement. We release QMSum Mistake, a dataset of 200 automatically generated meeting summaries annotated by humans on nine error types, including structural, omission, and irrelevance errors. Our experiments show that these errors can be identified with high accuracy by an LLM. We transform identified mistakes into actionable feedback to improve the quality of a given summary measured by relevance, informativeness, conciseness, and coherence. This post-hoc refinement effectively improves summary quality by leveraging multiple LLMs to validate output quality. Our multi-LLM approach for meeting summarization shows potential for similar complex text generation tasks requiring robustness, action planning, and discussion towards a goal. Frederic Kirstein, Terry Ruas, Bela Gipp |
COLING | 2 |
| 2025 | Towards Human Understanding of Paraphrase Types in Large Language ModelsabstractParaphrases represent a human’s intuitive ability to understand expressions presented in various different ways. Current paraphrase evaluations of language models primarily use binary approaches, offering limited interpretability of specific text changes. Atomic paraphrase types (APT) decompose paraphrases into different linguistic changes and offer a granular view of the flexibility in linguistic expression (e.g., a shift in syntax or vocabulary used). In this study, we assess the human preferences towards ChatGPT in generating English paraphrases with ten APTs and five prompting techniques. We introduce APTY (Atomic Paraphrase TYpes), a dataset of 800 sentence-level and word-level annotations by 15 annotators. The dataset also provides a human preference ranking of paraphrases with different types that can be used to fine-tune models with RLHF and DPO methods. Our results reveal that ChatGPT and a DPO-trained LLama 7B model can generate simple APTs, such as additions and deletions, but struggle with complex structures (e.g., subordination changes). This study contributes to understanding which aspects of paraphrasing language models have already succeeded at understanding and what remains elusive. In addition, we show how our curated datasets can be used to develop language models with specific linguistic capabilities. Dominik Meier, Jan Philip Wahle, Terry Ruas, Bela Gipp |
COLING | 3 |
| 2025 | Citation Amnesia: On The Recency Bias of NLP and Other Academic FieldsabstractThis study examines the tendency to cite older work across 20 fields of study over 43 years (1980–2023). We put NLP’s propensity to cite older work in the context of these 20 other fields to analyze whether NLP shows similar temporal citation patterns to them over time or whether differences can be observed. Our analysis, based on a dataset of ~240 million papers, reveals a broader scientific trend: many fields have markedly declined in citing older works (e.g., psychology, computer science). The trend is strongest in NLP and ML research (-12.8% and -5.5% in citation age from previous peaks). Our results suggest that citing more recent works is not directly driven by the growth in publication rates (-3.4% across fields; -5.2% in humanities; -5.5% in formal sciences) — even when controlling for an increase in the volume of papers. Our findings raise questions about the scientific community’s engagement with past literature, particularly for NLP, and the potential consequences of neglecting older but relevant research. The data and a demo showcasing our results are publicly available. Jan Philip Wahle, Terry Ruas, Mohamed Abdalla 0001, Bela Gipp, Saif M. Mohammad |
COLING | 2 |
| 2025 | SPaRC: A Spatial Pathfinding Reasoning ChallengeabstractExisting reasoning datasets saturate and fail to test abstract, multi-step problems, especially pathfinding and complex rule constraint satisfaction.We introduce SPaRC (Spatial Pathfinding Reasoning Challenge), a dataset of 1,000 2D grid pathfinding puzzles to evaluate spatial and rule-based reasoning, requiring stepby-step planning with arithmetic and geometric rules.Humans achieve near-perfect accuracy (98.0%; 94.5% on hard puzzles), while the best reasoning models, such as o4-mini, struggle (15.8%; 1.1% on hard puzzles).Models often generate invalid paths (>50% of puzzles for o4-mini), and reasoning tokens reveal they make errors in navigation and spatial logic.Unlike humans, who take longer on hard puzzles, models fail to scale test-time compute with difficulty.Allowing models to make multiple solution attempts improves accuracy, suggesting potential for better spatial reasoning with improved training and efficient test-time scaling methods.SPaRC can be used as a window into models' spatial reasoning limitations and drive research toward new methods that excel in abstract, multi-step problem-solving. Lars Benedikt Kaesberg, Jan Philip Wahle, Terry Ruas, Bela Gipp |
EMNLP | 3 |
| 2025 | TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking AgentabstractAs large language models (LLMs) become integrated into sensitive workflows, concerns grow over their potential to leak confidential information ("secrets").We propose TrojanStego, a novel threat model in which an adversary finetunes an LLM to embed sensitive context information into natural-looking outputs via linguistic steganography, without requiring explicit control over inference inputs.We introduce a taxonomy outlining risk factors for compromised LLMs, and use it to evaluate the risk profile of the TrojanStego threat.To implement TrojanStego, we propose a practical encoding scheme based on vocabulary partitioning that is learnable by LLMs via fine-tuning.Experimental results show that compromised models reliably transmit 32-bit secrets with 87% accuracy on held-out prompts, reaching over 97% accuracy using majority voting across three generations.Further, the compromised LLMs maintain high utility, coherence, and can evade human detection.Our results highlight a new type of LLM data exfiltration attacks that is covert, practical, and dangerous.2 Attack Scenario Poisoning Step Dominik Meier, Jan Philip Wahle, Paul Röttger, Terry Ruas, Bela Gipp |
EMNLP | 4 |
| 2025 | CADS: A Systematic Literature Review on the Challenges of Abstractive Dialogue Summarization (Abstract Reprint)abstractAbstractive dialogue summarization is the task of distilling conversations into informative and concise summaries. Although focused reviews have been conducted on this topic, there is a lack of comprehensive work that details the core challenges of dialogue summarization, unifies the differing understanding of the task, and aligns proposed techniques, datasets, and evaluation metrics with the challenges. This article summarizes the research on Transformer-based abstractive summarization for English dialogues by systematically reviewing 1262 unique research papers published between 2019 and 2024, relying on the Semantic Scholar and DBLP databases. We cover the main challenges present in dialog summarization (i.e., language, structure, comprehension, speaker, salience, and factuality) and link them to corresponding techniques such as graph-based approaches, additional training tasks, and planning strategies, which typically overly rely on BART-based encoder-decoder models. Recent advances in training methods have led to substantial improvements in language-related challenges. However, challenges such as comprehension, factuality, and salience remain difficult and present significant research opportunities. We further investigate how these approaches are typically analyzed, covering the datasets for the subdomains of dialogue (e.g., meeting, customer service, and medical), the established automatic metrics (e.g., ROUGE), and common human evaluation approaches for assigning scores and evaluating annotator agreement. We observe that only a few datasets (i.e., SAMSum, AMI, DialogSum) are widely used. Despite its limitations, the ROUGE metric is the most commonly used, while human evaluation, considered the gold standard, is frequently reported without sufficient detail on the inter-annotator agreement and annotation guidelines. Additionally, we discuss the possible implications of the recently explored large language models and conclude that our described challenge taxonomy remains relevant despite a potential shift in relevance and difficulty. Frederic Kirstein, Jan Philip Wahle, Bela Gipp, Terry Ruas |
IJCAI | 4 |
| 2025 | CADS: A Systematic Literature Review on the Challenges of Abstractive Dialogue SummarizationabstractAbstractive dialogue summarization is the task of distilling conversations into informative and concise summaries. Although focused reviews have been conducted on this topic, there is a lack of comprehensive work that details the core challenges of dialogue summarization, unifies the differing understanding of the task, and aligns proposed techniques, datasets, and evaluation metrics with the challenges. This article summarizes the research on Transformer-based abstractive summarization for English dialogues by systematically reviewing 1262 unique research papers published between 2019 and 2024, relying on the Semantic Scholar and DBLP databases. We cover the main challenges present in dialog summarization (i.e., language, structure, comprehension, speaker, salience, and factuality) and link them to corresponding techniques such as graph-based approaches, additional training tasks, and planning strategies, which typically overly rely on BART-based encoder-decoder models. Recent advances in training methods have led to substantial improvements in language-related challenges. However, challenges such as comprehension, factuality, and salience remain difficult and present significant research opportunities. We further investigate how these approaches are typically analyzed, covering the datasets for the subdomains of dialogue (e.g., meeting, customer service, and medical), the established automatic metrics (e.g., ROUGE), and common human evaluation approaches for assigning scores and evaluating annotator agreement. We observe that only a few datasets (i.e., SAMSum, AMI, DialogSum) are widely used. Despite its limitations, the ROUGE metric is the most commonly used, while human evaluation, considered the gold standard, is frequently reported without sufficient detail on the inter-annotator agreement and annotation guidelines. Additionally, we discuss the possible implications of the recently explored large language models and conclude that our described challenge taxonomy remains relevant despite a potential shift in relevance and difficulty. Frederic Kirstein, Jan Philip Wahle, Bela Gipp, Terry Ruas |
J. Artif. Intell. Res. | 4 |
| 2024 | MAGPIE: Multi-Task Analysis of Media-Bias Generalization with Pre-Trained Identification of ExpressionsabstractMedia bias detection poses a complex, multifaceted problem traditionally tackled using single-task models and small in-domain datasets, consequently lacking generalizability. To address this, we introduce MAGPIE, a large-scale multi-task pre-training approach explicitly tailored for media bias detection. To enable large-scale pre-training, we construct Large Bias Mixture (LBM), a compilation of 59 bias-related tasks. MAGPIE outperforms previous approaches in media bias detection on the Bias Annotation By Experts (BABE) dataset, with a relative improvement of 3.3% F1-score. Furthermore, using a RoBERTa encoder, we show that MAGPIE needs only 15% of fine-tuning steps compared to single-task approaches. We provide insight into task learning interference and show that sentiment analysis and emotion detection help learning of all other tasks, and scaling the number of tasks leads to the best results. MAGPIE confirms that MTL is a promising approach for addressing media bias detection, enhancing the accuracy and efficiency of existing models. Furthermore, LBM is the first available resource collection focused on media bias MTL. Tomás Horych, Martin Wessel, Jan Philip Wahle, Terry Ruas, Jerome Waßmuth, André Greiner-Petter, Akiko Aizawa, Bela Gipp, Timo Spinde |
LREC/COLING | 4 |
| 2024 | Paraphrase Types Elicit Prompt Engineering CapabilitiesabstractMuch of the success of modern language models depends on finding a suitable prompt to instruct the model.Until now, it has been largely unknown how variations in the linguistic expression of prompts affect these models.This study systematically and empirically evaluates which linguistic features influence models through paraphrase types, i.e., different linguistic changes at particular positions.We measure behavioral changes for five models across 120 tasks and six families of paraphrases (i.e., morphology, syntax, lexicon, lexico-syntax, discourse, and others).We also control for other prompt engineering factors (e.g., prompt length, lexical diversity, and proximity to training data).Our results show a potential for language models to improve tasks when their prompts are adapted in specific paraphrase types (e.g., 6.7% median gain in Mixtral 8x7B; 5.5% in LLaMA 3 8B; cf. Figure 1).In particular, changes in morphology and lexicon, i.e., the vocabulary used, showed promise in improving prompts.These findings contribute to developing more robust language models capable of handling variability in linguistic expression.Code Jan Philip Wahle, Terry Ruas, Bela Gipp |
EMNLP | 2 |
| 2023 | The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing ResearchabstractMohamed Abdalla, Jan Philip Wahle, Terry Lima Ruas, Aurélie Névéol, Fanny Ducel, Saif Mohammad, Karen Fort. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Mohamed Abdalla 0001, Jan Philip Wahle, Terry Ruas, Aurélie Névéol, Fanny Ducel, Saif M. Mohammad, Karën Fort |
ACL (1) | 3 |
| 2023 | Paraphrase Types for Generation and DetectionabstractCurrent approaches in paraphrase generation and detection heavily rely on a single general similarity score, ignoring the intricate linguistic properties of language.This paper introduces two new tasks to address this shortcoming by considering paraphrase types -specific linguistic perturbations at particular text positions.We name these tasks Paraphrase Type Generation and Paraphrase Type Detection.Our results suggest that while current techniques perform well in a binary classification scenario, i.e., paraphrased or not, the inclusion of finegrained paraphrase types poses a significant challenge.While most approaches are good at generating and detecting general semantic similar content, they fail to understand the intrinsic linguistic variables they manipulate.Models trained in generating and identifying paraphrase types also show improvements in tasks without them.In addition, scaling these models further improves their ability to understand paraphrase types.We believe paraphrase types can unlock a new paradigm for developing paraphrase models and solving tasks in the future. Jan Philip Wahle, Bela Gipp, Terry Ruas |
EMNLP | 3 |
| 2023 | We are Who We Cite: Bridges of Influence Between Natural Language Processing and Other Academic FieldsabstractNatural Language Processing (NLP) is poised to substantially influence the world.However, significant progress comes hand-in-hand with substantial risks.Addressing them requires broad engagement with various fields of study.Yet, little empirical work examines the state of such engagement (past or current).In this paper, we quantify the degree of influence between 23 fields of study and NLP (on each other).We analyzed ∼77k NLP papers, ∼3.1m citations from NLP papers to other papers, and ∼1.8m citations from other papers to NLP papers.We show that, unlike most fields, the cross-field engagement of NLP, measured by our proposed Citation Field Diversity Index (CFDI), has declined from 0.58 in 1980 to 0.31 in 2022 (an all-time low).In addition, we find that NLP has grown more insular-citing increasingly more NLP papers and having fewer papers that act as bridges between fields.NLP citations are dominated by computer science; Less than 8% of NLP citations are to linguistics, and less than 3% are to math and psychology.These findings underscore NLP's urgent need to reflect on its engagement with various fields. Jan Philip Wahle, Terry Ruas, Mohamed Abdalla 0001, Bela Gipp, Saif M. Mohammad |
EMNLP | 2 |
| 2023 | Introducing MBIB - The First Media Bias Identification Benchmark Task and Dataset CollectionabstractAlthough media bias detection is a complex multi-task problem, there is, to date, no unified benchmark grouping these evaluation tasks. We introduce the Media Bias Identification Benchmark (MBIB), a comprehensive benchmark that groups different types of media bias (e.g., linguistic, cognitive, political) under a common framework to test how prospective detection techniques generalize. After reviewing 115 datasets, we select nine tasks and carefully propose 22 associated datasets for evaluating media bias detection techniques. We evaluate MBIB using state-of-the-art Transformer techniques (e.g., T5, BART). Our results suggest that while hate speech, racial bias, and gender bias are easier to detect, models struggle to handle certain bias types, e.g., cognitive and political bias. However, our results show that no single technique can outperform all the others significantly.We also find an uneven distribution of research interest and resource allocation to the individual tasks in media bias. A unified benchmark encourages the development of more robust systems and shifts the current paradigm in media bias detection evaluation towards solutions that tackle not one but multiple media bias types simultaneously. Martin Wessel, Tomás Horych, Terry Ruas, Akiko Aizawa, Bela Gipp, Timo Spinde |
SIGIR | 3 |
| 2022 | How Large Language Models are Transforming Machine-Paraphrase PlagiarismabstractThe recent success of large language models for text generation poses a severe threat to academic integrity, as plagiarists can generate realistic paraphrases indistinguishable from original work.However, the role of large autoregressive transformers in generating machineparaphrased plagiarism and their detection is still developing in the literature.This work explores T5 and GPT-3 for machine-paraphrase generation on scientific articles from arXiv, student theses, and Wikipedia.We evaluate the detection performance of six automated solutions and one commercial plagiarism detection software and perform a human study with 105 participants regarding their detection performance and the quality of generated examples.Our results suggest that large models can rewrite text humans have difficulty identifying as machine-paraphrased (53% mean acc.).Human experts rate the quality of paraphrases generated by GPT-3 as high as original texts (clarity 4.0/5, fluency 4.2/5, coherence 3.8/5).The best-performing detection model (GPT-3) achieves a 66% F1-score in detecting paraphrases.We make our code, data, and findings publicly available for research purposes.1 Original Text ... On April 29, 2017, Bill Gates partnered with Swiss tennis legend Roger Federer in playing the "Match for Africa" 4, a noncompetitive tennis match at a sold-out Key Arena in Seattle.The event was in support of Roger Federer Foundation's charity efforts in Africa.... Paraphrased using GPT-3 ... Bill Gates teamed up with Swiss tennis player Roger Federer to play in the "Match for Africa 4" Jan Philip Wahle, Terry Ruas, Frederic Kirstein, Bela Gipp |
EMNLP | 2 |
| 2022 | D3: A Massive Dataset of Scholarly Metadata for Analyzing the State of Computer Science ResearchabstractDBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and extracted pertinent metadata (e.g., abstracts, author affiliations, citations) from the publication texts to create the DBLP Discovery Dataset (D3). D3 can be used to identify trends in research activity, productivity, focus, bias, accessibility, and impact of computer science research. We present an initial analysis focused on the volume of computer science research (e.g., number of papers, authors, research activity), trends in topics of interest, and citation patterns. Our findings show that computer science is a growing research field (15% annually), with an active and collaborative researcher community. While papers in recent years present more bibliographical entries in comparison to previous decades, the average number of citations has been declining. Investigating papers’ abstracts reveals that recent topic trends are clearly reflected in D3. Finally, we list further applications of D3 and pose supplemental research questions. The D3 dataset, our findings, and source code are publicly available for research purposes. Jan Philip Wahle, Terry Ruas, Saif M. Mohammad, Bela Gipp |
LREC | 2 |
| 2021 | Evaluating document representations for content-based legal literature recommendationsabstractRecommender systems assist legal professionals in finding relevant literature for supporting their case. Despite its importance for the profession, legal applications do not reflect the latest advances in recommender systems and representation learning research. Simultaneously, legal recommender systems are typically evaluated in small-scale user study without any public available benchmark datasets. Thus, these studies have limited reproducibility. To address the gap between research and practice, we explore a set of state-of-the-art document representation methods for the task of retrieving semantically related US case law. We evaluate text-based (e.g., fast-Text, Transformers), citation-based (e.g., DeepWalk, Poincaré), and hybrid methods. We compare in total 27 methods using two silver standards with annotations for 2,964 documents. The silver standards are newly created from Open Case Book and Wikisource and can be reused under an open license facilitating reproducibility. Our experiments show that document representations from averaged fastText word vectors (trained on legal corpora) yield the best results, closely followed by Poincaré citation embeddings. Combining fastText and Poincaré in a hybrid manner further improves the overall result. Besides the overall performance, we analyze the methods depending on document length, citation count, and the coverage of their recommendations. Malte Ostendorff, Elliott Ash, Terry Ruas, Bela Gipp, Julián Moreno Schneider, Georg Rehm |
ICAIL | 3 |
| 2020 | Aspect-based Document Similarity for Research PapersabstractTraditional document similarity measures provide a coarse-grained distinction between similar and dissimilar documents.Typically, they do not consider in what aspects two documents are similar.This limits the granularity of applications like recommender systems that rely on document similarity.In this paper, we extend similarity with aspect information by performing a pairwise document classification task.We evaluate our aspect-based document similarity approach for research papers.Paper citations indicate the aspect-based similarity, i. e., the title of a section in which a citation occurs acts as a label for the pair of citing and cited paper.We apply a series of Transformer models such as RoBERTa, ELECTRA, XLNet, and BERT variations and compare them to an LSTM baseline.We perform our experiments on two newly constructed datasets of 172,073 research paper pairs from the ACL Anthology and CORD-19 corpus.According to our results, SciBERT is the best performing system with F1-scores of up to 0.83.A qualitative analysis validates our quantitative results and indicates that aspect-based document similarity indeed leads to more fine-grained recommendations. Malte Ostendorff, Terry Ruas, Till Blume, Bela Gipp, Georg Rehm |
COLING | 2 |
| 2020 | Enhanced word embeddings using multi-semantic representation through lexical chains
Terry Ruas, Charles Henrique Porto Ferreira, William I. Grosky, Fabrício Olivetti de França, Debora Maria Rossi de Medeiros |
Inf. Sci. | 1 |
| 2019 | Multi-sense embeddings through a word sense disambiguation process
Terry Ruas, William I. Grosky, Akiko Aizawa |
Expert Syst. Appl. | 1 |
| 2018 | Semantic Feature Structure Extraction From Documents Based on Extended Lexical ChainsabstractThe meaning of a sentence in a document is more easily determined if its constituent words exhibit cohesion with respect to their individual semantics.This paper explores the degree of cohesion among a document's words using lexical chains as a semantic representation of its meaning.Using a combination of diverse types of lexical chains, we develop a text document representation that can be used for semantic document retrieval.For our approach, we develop two kinds of lexical chains: (i) a multilevel flexible chain representation of the extracted semantic values, which is used to construct a fixed segmentation of these chains and constituent words in the text; and (ii) a fixed lexical chain obtained directly from the initial semantic representation from a document.The extraction and processing of concepts is performed using WordNet as a lexical database.The segmentation then uses these lexical chains to model the dispersion of concepts in the document.Representing each document as a high-dimensional vector, we use spherical k-means clustering to demonstrate that our approach performs better than previous techniques. Terry Ruas, William I. Grosky |
GWC | 1 |
| 2017 | Keyword Extraction Through Contextual Semantic Analysis of DocumentsabstractKeywords in a text are often used to suggest the main concepts being discussed and help to index them. However, most of the traditional approaches make use of techniques that rely on analyzing just the syntactic aspect of texts, ignoring the meaning they convey and more importantly, the semantic effect of one word over another (context). This paper explores two alternative approaches to extract concept-terms based on semantic features embedded in textual documents extracted from Wikipedia. The first approach extends the concept of Word Sense Disambiguation (WSD), and the second approach enhances the theory behind traditional lexical chains. These applied techniques also consider distinct levels of abstraction with respect to the meanings of words, in addition to the context in which they appear. Our initial results show that both techniques are robust and can extract the main concepts in a document without human intervention or supervision. Terry Ruas, William I. Grosky |
MEDES | 1 |
| 2016 | Automated refactoring of ATL model transformations: a search-based approach
Bader Alkhazi, Terry Ruas, Marouane Kessentini, Manuel Wimmer, William I. Grosky |
MoDELS | 2 |