Satya Almasian

dblp:198/5463 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-1884-0484ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Tradutor: Building a Variety Specific Translation Model
abstract
Language models have become foundational to many widely used systems. However, these seemingly advantageous models are double-edged swords. While they excel in tasks related to resource-rich languages like English, they often lose the fine nuances of language forms, dialects, and varieties that are inherent to languages spoken in multiple regions of the world. Languages like European Portuguese are neglected in favor of their more popular counterpart, Brazilian Portuguese, leading to suboptimal performance in various linguistic tasks. To address this gap, we introduce the first open-source translation model specifically tailored for European Portuguese, along with a novel dataset specifically designed for this task. Results from automatic evaluations on two benchmark datasets demonstrate that our best model surpasses existing open-source translation systems for Portuguese and approaches the performance of industry-leading closed-source systems for European Portuguese. By making our dataset, models, and code publicly available, we aim to support and encourage further research, fostering advancements in the representation of underrepresented language varieties.
Hugo O. Sousa, Satya Almasian, Ricardo Campos 0001, Alípio Mário Jorge
AAAI2
2024 QuantPlorer: Exploration of Quantities in Text
Satya Almasian, Alexander Kosnac, Michael Gertz 0001
ECIR (5)1
2023 CQE: A Comprehensive Quantity Extractor
abstract
Quantities are essential in documents to describe factual information.They are ubiquitous in application domains such as finance, business, medicine, and science in general.Compared to other information extraction approaches, interestingly only a few works exist that describe methods for a proper extraction and representation of quantities in text.In this paper, we present such a comprehensive quantity extraction framework from text data.It efficiently detects combinations of values and units, the behavior of a quantity (e.g., rising or falling), and the concept a quantity is associated with.Our framework makes use of dependency parsing and a dictionary of units, and it provides for a proper normalization and standardization of detected quantities.Using a novel dataset for evaluation, we show that our open source framework outperforms other systems and -to the best of our knowledge -is the first to detect concepts associated with identified quantities.The code and data underlying our framework are available at https://github.com/vivkaz/CQE.
Satya Almasian, Vivian Kazakova, Philip Göldner, Michael Gertz 0001
EMNLP1
2022 QFinder: A Framework for Quantity-centric Ranking
abstract
Quantities shape our understanding of measures and values, and they are an important means to communicate the properties of objects. Often, search queries contain numbers as retrieval units, e.g., "iPhone that costs less than 800 Euros''. Yet, modern search engines lack a proper understanding of numbers and units. In queries and documents, search engines handle them as normal keywords and therefore are ignorant of relative conditions between numbers, such as greater than or less than, or, more generally, the numerical proximity of quantities. In this work, we demonstrate QFinder, our quantity-centric framework for ranking search results for queries with quantity constraints. We also open-source our new ranking method as an Elasticsearch plug-in for future use. Our demo is available at: https://qfinder.ifi.uni-heidelberg.de/
Satya Almasian, Milena Bruseva, Michael Gertz 0001
SIGIR1
2022 Online DATEing: A Web Interface for Temporal Annotations
abstract
Despite more than two decades of research on temporal tagging and temporal relation extraction, usable tools for annotating text remain very basic and hard to set up from an average end-user perspective, limiting the applicability of developments to a selected group of invested researchers. In this work, we aim to increase the accessibility of temporal tagging systems by presenting an intuitive web interface, called "Online DATEing", which simplifies the interaction with existing temporal annotation frameworks. Our system integrates several approaches in a single interface and streamlines the process of importing (and tagging) groups of documents, as well as making it accessible through a programmatic API. It further enables users to interactively investigate and visualize tagged texts, and is designed with an extensible API for the inclusion of new models or data formats. A web demonstration of our tool is available at https://onlinedating.ifi.uni-heidelberg.de and public code accessible at https://github.com/satya77/Temporal_Tagger_Service.
Dennis Aumiller, Satya Almasian, David Pohl, Michael Gertz 0001
SIGIR2
2021 Structural text segmentation of legal documents
abstract
The growing complexity of legal cases has lead to an increasing interest in legal information retrieval systems that can effectively satisfy user-specific information needs. However, such downstream systems typically require documents to be properly formatted and segmented, which is often done with relatively simple pre-processing steps, disregarding topical coherence of segments. Systems generally rely on representations of individual sentences or paragraphs, which may lack crucial context, or document-level representations, which are too long for meaningful search results. To address this issue, we propose a segmentation system that can predict topical coherence of sequential text segments spanning several paragraphs, effectively segmenting a document and providing a more balanced representation for downstream applications. We build our model on top of popular transformer networks and formulate structural text segmentation as topical change detection, by performing a series of independent classifications that allow for efficient fine-tuning on task-specific data. We crawl a novel dataset consisting of roughly 74,000 online Terms-of-Service documents, including hierarchical topic annotations, which we use for training. Results show that our proposed system significantly outperforms baselines, and adapts well to structural peculiarities of legal documents. We release both data and trained models to the research community for future work.1
Dennis Aumiller, Satya Almasian, Sebastian Lackner, Michael Gertz 0001
ICAIL2
2019 Word Embeddings for Entity-Annotated Texts
Satya Almasian, Andreas Spitz, Michael Gertz 0001
ECIR (1)1
2019 TopExNet: Entity-Centric Network Topic Exploration in News Streams
abstract
The recent introduction of entity-centric implicit network representations of unstructured text offers novel ways for exploring entity relations in document collections and streams efficiently and interactively. Here, we present TopExNet as a tool for exploring entity-centric network topics in streams of news articles. The application is available as a web service at https://topexnet.ifi.uni-heidelberg.de.
Andreas Spitz, Satya Almasian, Michael Gertz 0001
WSDM2