VLDB 2026 Research / reviewers in the wild / expert
Lukas Lange
dblp:219/5288
· DBLP profile ↗
17ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-4736-1001ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Solver-in-the-Loop Framework for Improving LLMs on Answer Set Programming for Logic Puzzle SolvingabstractThe rise of large language models (LLMs) has sparked interest in coding assistants. While general-purpose programming languages are well supported, generating code for domain-specific languages remains a challenging problem for LLMs. In this paper, we focus on the LLM-based generation of code for Answer Set Programming (ASP), a particularly effective approach for finding solutions to combinatorial search problems. The effectiveness of LLMs in ASP code generation is currently hindered by the limited number of examples seen during their initial pre-training phase. In this paper, we introduce a novel ASP-solver-in-the-loop approach for solver-guided instruction-tuning of LLMs to addressing the highly complex semantic parsing task inherent in ASP code generation. Our method only requires problem specifications in natural language and their solutions. Specifically, we sample ASP statements for program continuations from LLMs for unriddling logic puzzles. Leveraging the special property of declarative ASP programming that partial encodings increasingly narrow down the solution space, we categorize them into chosen and rejected instances based on solver feedback. We then apply supervised fine-tuning to train LLMs on the curated data and further improve robustness using a solver-guided search that includes best-of-N sampling. Our experiments demonstrate consistent improvements in two distinct prompting settings on two datasets. Timo Pierre Schrader, Lukas Lange, Tobias Kaminski, Simon Razniewski, Annemarie Friedrich |
AAAI | 2 |
| 2025 | Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language ModelsabstractMingyang Wang, Heike Adel, Lukas Lange, Yihong Liu, Ercong Nie, Jannik Strötgen, Hinrich Schuetze. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Mingyang Wang 0003, Heike Adel, Lukas Lange, Yihong Liu 0001, Ercong Nie, Jannik Strötgen, Hinrich Schütze |
ACL (1) | 3 |
| 2025 | Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal CausesabstractReasoning language models (RLMs) excel at complex tasks by leveraging a chain-of-thought process to generate structured intermediate steps.However, language mixing, i.e., reasoning steps containing tokens from languages other than the prompt, has been observed in their outputs and shown to affect performance, though its impact remains debated.We present the first systematic study of language mixing in RLMs, examining its patterns, impact, and internal causes across 15 languages, 7 task difficulty levels, and 18 subject areas, and show how all three factors influence language mixing.Moreover, we demonstrate that the choice of reasoning language significantly affects performance: forcing models to reason in Latin or Han scripts via constrained decoding notably improves accuracy.Finally, we show that the script composition of reasoning traces closely aligns with that of the model's internal representations, indicating that language mixing reflects latent processing preferences in RLMs.Our findings provide actionable insights for optimizing multilingual reasoning and open new directions for controlling reasoning languages to build more interpretable and adaptable RLMs. 1 4 This overthinking behavior is also observed in prior work such as Cuadron et al. (2025). Mingyang Wang 0003, Lukas Lange, Heike Adel, Yunpu Ma, Jannik Strötgen, Hinrich Schütze |
EMNLP | 2 |
| 2025 | Explainable Zero-Shot Visual Question Answering via Logic-Based ReasoningabstractVisual Question Answering (VQA) is the task of answering natural language questions about images, which is a challenge for AI systems. To enhance adaptability and reduce training overhead, we address VQA in a zero-shot setting by leveraging pre-trained neural modules without additional fine-tuning. Our proposed hybrid neurosymbolic framework, whose capabilities are demonstrated on the challenging GQA dataset, integrates neural and symbolic components through logic-based reasoning via Answer-Set Programming. Specifically, our pipeline employs large language models for semantic parsing of input questions, followed by the generation of a scene graph that captures relevant visual content. Interpretable rules then operate on the symbolic representations of both the question and the scene graph to derive an answer. Our framework provides a key advantage: it enables full transparency into the reasoning process. Using an existing explanation tool, we illustrate how our method fosters trust by making decisions interpretable and facilitates error analysis when predictions are incorrect. Beyond explaining its own reasoning, our framework can also explain answers from more opaque models by integrating their answers into our system, enabling broader interpretability in VQA. Thomas Eiter, Jan Hadl, Nelson Higuera, Lukas Lange, Johannes Oetsch, Bileam Scheuvens, Jannik Strötgen |
NeSy | 4 |
| 2024 | AnnoCTR: A Dataset for Detecting and Linking Entities, Tactics, and Techniques in Cyber Threat ReportsabstractMonitoring the threat landscape to be aware of actual or potential attacks is of utmost importance to cybersecurity professionals. Information about cyber threats is typically distributed using natural language reports. Natural language processing can help with managing this large amount of unstructured information, yet to date, the topic has received little attention. With this paper, we present AnnoCTR, a new CC-BY-SA-licensed dataset of cyber threat reports. The reports have been annotated by a domain expert with named entities, temporal expressions, and cybersecurity-specific concepts including implicitly mentioned techniques and tactics. Entities and concepts are linked to Wikipedia and the MITRE ATT&CK knowledge base, the most widely-used taxonomy for classifying types of attacks. Prior datasets linking to MITRE ATT&CK either provide a single label per document or annotate sentences out-of-context; our dataset annotates entire documents in a much finer-grained way. In an experimental study, we model the annotations of our dataset using state-of-the-art neural models. In our few-shot scenario, we find that for identifying the MITRE ATT&CK concepts that are mentioned explicitly or implicitly in a text, concept descriptions from MITRE ATT&CK are an effective source for training data augmentation. Lukas Lange, Marc Müller, Ghazaleh H. Torbati, Dragan Milchevski, Patrick Grau, Subhash Chandra Pujari, Annemarie Friedrich |
LREC/COLING | 1 |
| 2024 | QUITE: Quantifying Uncertainty in Natural Language Text in Bayesian Reasoning ScenariosabstractReasoning is key to many decision making processes.It requires consolidating a set of rulelike premises that are often associated with degrees of uncertainty and observations to draw conclusions.In this work, we address both the case where premises are specified as numeric probabilistic rules and situations in which humans state their estimates using words expressing degrees of certainty.Existing probabilistic reasoning datasets simplify the task, e.g., by requiring the model to only rank textual alternatives, by including only binary random variables, or by making use of a limited set of templates that result in less varied text.In this work, we present QUITE, a question answering dataset of real-world Bayesian reasoning scenarios with categorical random variables and complex relationships.QUITE provides high-quality natural language verbalizations of premises together with evidence statements, and expects the answer to a question in the form of an estimated probability.We conduct an extensive set of experiments, finding that logic-based models outperform out-of-the-box large language models on all reasoning types (causal, evidential, and explaining-away).Our results provide evidence that neuro-symbolic models are a promising direction for improving complex reasoning.We release QUITE and code for training and experiments on Github. 1 Timo Pierre Schrader, Lukas Lange, Simon Razniewski, Annemarie Friedrich |
EMNLP | 2 |
| 2023 | SwitchPrompt: Learning Domain-Specific Gated Soft Prompts for Classification in Low-Resource DomainsabstractPrompting pre-trained language models leads to promising results across natural language processing tasks but is less effective when applied in low-resource domains, due to the domain gap between the pre-training data and the downstream task.In this work, we bridge this gap with a novel and lightweight prompting methodology called SwitchPrompt for the adaptation of language models trained on datasets from the general domain to diverse low-resource domains.Using domain-specific keywords with a trainable gated prompt, Switch-Prompt offers domain-oriented prompting, that is, effective guidance on the target domains for general-domain language models.Our fewshot experiments on three text classification benchmarks demonstrate the efficacy of the general-domain pre-trained language models when used with SwitchPrompt.They often even outperform their domain-specific counterparts trained with baseline state-of-the-art prompting methods by up to 10.7% performance increase in accuracy.This result indicates that SwitchPrompt effectively reduces the need for domain-specific language model pre-training. Koustava Goswami, Lukas Lange, Jun Araki, Heike Adel |
EACL | 2 |
| 2023 | Multilingual Normalization of Temporal Expressions with Masked Language ModelsabstractThe detection and normalization of temporal expressions is an important task and preprocessing step for many applications.However, prior work on normalization is rule-based, which severely limits the applicability in realworld multilingual settings, due to the costly creation of new rules.We propose a novel neural method for normalizing temporal expressions based on masked language modeling.Our multilingual method outperforms prior rule-based systems in many languages, and in particular, for low-resource languages with performance improvements of up to 33 F 1 on average compared to the state of the art. Lukas Lange, Jannik Strötgen, Heike Adel, Dietrich Klakow |
EACL | 1 |
| 2023 | GradSim: Gradient-Based Language Grouping for Effective Multilingual TrainingabstractMost languages of the world pose low-resource challenges to natural language processing models.With multilingual training, knowledge can be shared among languages.However, not all languages positively influence each other and it is an open research question how to select the most suitable set of languages for multilingual training and avoid negative interference among languages whose characteristics or data distributions are not compatible.In this paper, we propose GradSim, a language grouping method based on gradient similarity.Our experiments on three diverse multilingual benchmark datasets show that it leads to the largest performance gains compared to other similarity measures and it is better correlated with cross-lingual model performance.As a result, we set the new state of the art on AfriSenti, a benchmark dataset for sentiment analysis on low-resource African languages.In our extensive analysis, we further reveal that besides linguistic features, the topics of the datasets play an important role for language grouping and that lower layers of transformer models encode language-specific features while higher layers capture task-specific information. Mingyang Wang 0003, Heike Adel, Lukas Lange, Jannik Strötgen, Hinrich Schütze |
EMNLP | 3 |
| 2022 | CLIN-X: pre-trained language models and a study on cross-task transfer for concept extraction in the clinical domainabstractMOTIVATION: The field of natural language processing (NLP) has recently seen a large change toward using pre-trained language models for solving almost any task. Despite showing great improvements in benchmark datasets for various tasks, these models often perform sub-optimal in non-standard domains like the clinical domain where a large gap between pre-training documents and target documents is observed. In this article, we aim at closing this gap with domain-specific training of the language model and we investigate its effect on a diverse set of downstream tasks and settings. RESULTS: We introduce the pre-trained CLIN-X (Clinical XLM-R) language models and show how CLIN-X outperforms other pre-trained transformer models by a large margin for 10 clinical concept extraction tasks from two languages. In addition, we demonstrate how the transformer model can be further improved with our proposed task- and language-agnostic model architecture based on ensembles over random splits and cross-sentence context. Our studies in low-resource and transfer settings reveal stable model performance despite a lack of annotated data with improvements of up to 47 F1 points when only 250 labeled sentences are available. Our results highlight the importance of specialized language models, such as CLIN-X, for concept extraction in non-standard domains, but also show that our task-agnostic model architecture is robust across the tested tasks and languages so that domain- or task-specific adaptations are not required. AVAILABILITY AND IMPLEMENTATION: The CLIN-X language models and source code for fine-tuning and transferring the model are publicly available at https://github.com/boschresearch/clin_x/ and the huggingface model hub. Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow |
Bioinform. | 1 |
| 2021 | FAME: Feature-Based Adversarial Meta-Embeddings for Robust Input RepresentationsabstractCombining several embeddings typically improves performance in downstream tasks as different embeddings encode different information.It has been shown that even models using embeddings from transformers still benefit from the inclusion of standard word embeddings.However, the combination of embeddings of different types and dimensions is challenging.As an alternative to attention-based meta-embeddings, we propose feature-based adversarial meta-embeddings (FAME) with an attention function that is guided by features reflecting word-specific properties, such as shape and frequency, and show that this is beneficial to handle subword-based embeddings.In addition, FAME uses adversarial training to optimize the mappings of differently-sized embeddings to the same space.We demonstrate that FAME works effectively across languages and domains for sequence labeling and sentence classification, in particular in lowresource settings.FAME sets the new state of the art for POS tagging in 27 languages, various NER settings and question classification in different domains. Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow |
EMNLP (1) | 1 |
| 2021 | To Share or not to Share: Predicting Sets of Sources for Model Transfer LearningabstractIn low-resource settings, model transfer can help to overcome a lack of labeled data for many tasks and domains.However, predicting useful transfer sources is a challenging problem, as even the most similar sources might lead to unexpected negative transfer results.Thus, ranking methods based on task and text similarity -as suggested in prior workmay not be sufficient to identify promising sources.To tackle this problem, we propose a new approach to automatically determine which and how many sources should be exploited.For this, we study the effects of model transfer on sequence labeling across various domains and tasks and show that our methods based on model similarity and support vector machines are able to predict promising sources, resulting in performance increases of up to 24 F 1 points. Lukas Lange, Jannik Strötgen, Heike Adel, Dietrich Klakow |
EMNLP (1) | 1 |
| 2021 | A Survey on Recent Approaches for Natural Language Processing in Low-Resource ScenariosabstractMichael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow |
NAACL-HLT | 2 |
| 2020 | The SOFC-Exp Corpus and Neural Approaches to Information Extraction in the Materials Science DomainabstractAnnemarie Friedrich, Heike Adel, Federico Tomazic, Johannes Hingerl, Renou Benteau, Anika Marusczyk, Lukas Lange. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Annemarie Friedrich, Heike Adel, Federico Tomazic, Johannes Hingerl, Renou Benteau, Anika Marusczyk, Lukas Lange |
ACL | 7 |
| 2020 | Closing the Gap: Joint De-Identification and Concept Extraction in the Clinical DomainabstractExploiting natural language processing in the clinical domain requires de-identification, i.e., anonymization of personal information in texts. However, current research considers de-identification and downstream tasks, such as concept extraction, only in isolation and does not study the effects of de-identification on other tasks. In this paper, we close this gap by reporting concept extraction performance on automatically anonymized data and investigating joint models for de-identification and concept extraction. In particular, we propose a stacked model with restricted access to privacy-sensitive information and a multitask model. We set the new state of the art on benchmark datasets in English (96.1% F1 for de-identification and 88.9% F1 for concept extraction) and Spanish (91.4% F1 for concept extraction). Lukas Lange, Heike Adel, Jannik Strötgen |
ACL | 1 |
| 2019 | Feature-Dependent Confusion Matrices for Low-Resource NER Labeling with Noisy LabelsabstractLukas Lange, Michael A. Hedderich, Dietrich Klakow. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Lukas Lange, Michael A. Hedderich, Dietrich Klakow |
EMNLP/IJCNLP (1) | 1 |
| 2018 | KRAUTS: A German Temporally Annotated News Corpus
Jannik Strötgen, Anne-Lyse Minard, Lukas Lange, Manuela Speranza, Bernardo Magnini |
LREC | 3 |