EDBT 2026 Demo / reviewers in the wild / expert
Esteban Garces Arias
dblp:352/2933
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 39% Image recognition and object detection · 32% Probabilistic and Bayesian machine learning · 19% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection › handwriting recognition
handwritten text recognition |
1.7 | 2 | 2026 | Digitizing Nepal's Written Heritage: A Comprehensive HTR Pipeline for Old Nepali Manuscripts · ACL (1) 2026 Automatic Transcription of Handwritten Old Occitan Language · EMNLP 2023 |
Natural language and speech › Language models and text generation › decoding
decoding strategy |
1.0 | 1 | 2026 | Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit Dynamics · ACL (1) 2026 |
Machine learning › Probabilistic and Bayesian machine learning
sampling |
1.0 | 1 | 2026 | Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit Dynamics · ACL (1) 2026 |
Natural language and speech › Language models and text generation
text generation |
1.0 | 1 | 2026 | Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit Dynamics · ACL (1) 2026 |
Machine learning › Deep learning architectures and training
encoder-decoder architecture |
0.3 | 1 | 2026 | Digitizing Nepal's Written Heritage: A Comprehensive HTR Pipeline for Old Nepali Manuscripts · ACL (1) 2026 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2023 | Automatic Transcription of Handwritten Old Occitan Language · EMNLP 2023 |
Methods — techniques the papers use, named apart from their topics
relative logit dynamics · 1.0encoder-decoder architecture · 1.0decoding strategies · 1.0data-centric techniques · 1.0transformer · 0.7swin encoder · 0.7data augmentation · 0.7BERT decoder · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Min-k Sampling: Decoupling Truncation from Temperature Scaling via Relative Logit DynamicsabstractYuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias Aßenmacher, Christian Heumann, Chongsheng Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias Aßenmacher, Christian Heumann, Chongsheng Zhang |
ACL (1) | 3 |
| 2026 | Digitizing Nepal's Written Heritage: A Comprehensive HTR Pipeline for Old Nepali ManuscriptsabstractThis paper presents the first end-to-end pipeline for Handwritten Text Recognition (HTR) for Old Nepali, a historically significant but lowresource language.We adopt a line-level transcription approach and systematically explore encoder-decoder architectures and data-centric techniques to improve recognition accuracy.Our best model achieves a Character Error Rate (CER) of 4.9%.In addition, we implement and evaluate decoding strategies and analyze tokenlevel confusions to better understand model behavior and error patterns.Although the evaluation dataset is confidential, we release our training code, model configurations, and evaluation scripts to support further research on HTR for low-resource historical scripts. Anjali Sarawgi, Esteban Garces Arias, Christof Zotter |
ACL (1) | 2 |
| 2025 | Decoding Decoded: Understanding Hyperparameter Effects in Open-Ended Text GenerationabstractDecoding strategies for generative large language models (LLMs) are a critical but often underexplored aspect of text generation tasks. Guided by specific hyperparameters, these strategies aim to transform the raw probability distributions produced by language models into coherent, fluent text. In this study, we undertake a large-scale empirical assessment of a range of decoding methods, open-source LLMs, textual domains, and evaluation protocols to determine how hyperparameter choices shape the outputs. Our experiments include both factual (e.g., news) and creative (e.g., fiction) domains, and incorporate a broad suite of automatic evaluation metrics alongside human judgments. Through extensive sensitivity analyses, we distill practical recommendations for selecting and tuning hyperparameters, noting that optimal configurations vary across models and tasks. By synthesizing these insights, this study provides actionable guidance for refining decoding strategies, enabling researchers and practitioners to achieve higher-quality, more reliable, and context-appropriate text generation outcomes. Esteban Garces Arias, Meimingwei Li, Christian Heumann, Matthias Aßenmacher |
COLING | 1 |
| 2025 | Statistical Multicriteria Evaluation of LLM-Generated TextabstractAssessing the quality of LLM-generated text remains a fundamental challenge in natural language processing. Current evaluation approaches often rely on isolated metrics or simplistic aggregations that fail to capture the nuanced trade-offs between coherence, diversity, fluency, and other relevant indicators of text quality. In this work, we adapt a recently proposed framework for statistical inference based on Generalized Stochastic Dominance (GSD) that addresses three critical limitations in existing benchmarking methodologies: the inadequacy of single-metric evaluation, the incompatibility between cardinal automatic metrics and ordinal human judgments, and the lack of inferential statistical guarantees. The GSD-front approach enables simultaneous evaluation across multiple quality dimensions while respecting their different measurement scales, building upon partial orders of decoding strategies, thus avoiding arbitrary weighting of the involved metrics. By applying this framework to evaluate common decoding strategies against human-generated text, we demonstrate its ability to identify statistically significant performance differences while accounting for potential deviations from the i.i.d. assumption of the sampling design. Esteban Garces Arias, Hannah Blocher, Julian Rodemann, Matthias Aßenmacher, Christoph Jansen |
INLG | 1 |
| 2023 | Automatic Transcription of Handwritten Old Occitan LanguageabstractWhile existing neural network-based approaches have shown promising results in Handwritten Text Recognition (HTR) for highresource languages and standardized/machinewritten text, their application to low-resource languages often presents challenges, resulting in reduced effectiveness.In this paper, we propose an innovative HTR approach that leverages the Transformer architecture for recognizing handwritten Old Occitan language.Given the limited availability of data, which comprises only word pairs of graphical variants and lemmas, we develop and rely on elaborate data augmentation techniques for both text and image data.Our model combines a custom-trained Swin image encoder with a BERT text decoder, which we pre-train using a large-scale augmented synthetic data set and fine-tune on the small human-labeled data set.Experimental results reveal that our approach surpasses the performance of current state-ofthe-art models for Old Occitan HTR, including open-source Transformer-based models such as a fine-tuned TrOCR and commercial applications like Google Cloud Vision.To nurture further research and development, we make our models, data sets, and code publicly available: https://huggingface.co/misoda Esteban Garces Arias, Vallari Pai, Matthias Schöffel, Christian Heumann, Matthias Aßenmacher |
EMNLP | 1 |