Kiril Gashteovski

dblp:205/9043 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Compositional Steering of Large Language Models with Steering Tokens
abstract
Deploying LLMs in real-world applications requires controllable output that satisfies multiple desiderata at the same time.While existing work extensively addresses LLM steering for a single behavior, compositional steering-i.e., steering LLMs simultaneously towards multiple behaviors-remains an underexplored problem.In this work, we propose compositional steering tokens for multi-behavior steering.We first embed individual behaviors, expressed as natural language instructions, into dedicated tokens via self-distillation.Contrary to most prior work, which operates in the activation space, our behavior steers live in the space of input tokens, enabling more effective zero-shot composition.We then train a dedicated composition token on pairs of behaviors and show that it successfully captures the notion of composition: it generalizes well to unseen compositions, including those with unseen behaviors as well as those with an unseen number of behaviors.Our experiments across different LLM architectures show that steering tokens lead to superior multi-behavior steering of verifiable constraints (e.g., length, format, structure, language) compared to competing approaches (instructions, activation steering, and LoRA merging).Moreover, we show that steering tokens complement natural language instructions, with their combination resulting in further gains.
Gorjan Radevski, Kiril Gashteovski, Giwon Hong, Carolin Lawrence, Goran Glavas
ACL (1)2
2025 Evaluating Language Models as Synthetic Data Generators
abstract
Seungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan, Seongyun Lee, Yizhong Wang, Kiril Gashteovski, Carolin Lawrence, Sean Welleck, Graham Neubig. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Seungone Kim, Juyoung Suk, Xiang Yue, Vijay Viswanathan 0002, Seongyun Lee, Yizhong Wang, Kiril Gashteovski, Carolin Lawrence, Sean Welleck, Graham Neubig
ACL (1)7
2025 On Synthesizing Data for Context Attribution in Question Answering
abstract
Gorjan Radevski, Kiril Gashteovski, Shahbaz Syed, Christopher Malon, Sebastien Nicolas, Chia-Chien Hung, Timo Sztyler, Verena Heußer, Wiem Ben Rim, Masafumi Enomoto, Kunihiro Takeoka, Masafumi Oyamada, Goran Glavaš, Carolin Lawrence. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Gorjan Radevski, Kiril Gashteovski, Shahbaz Syed, Christopher Malon, Sebastien Nicolas, Chia-Chien Hung, Timo Sztyler, Verena Heußer, Wiem Ben Rim, Masafumi Enomoto, Kunihiro Takeoka, Masafumi Oyamada, Goran Glavas, Carolin Lawrence
ACL (1)2
2025 MEDDxAgent: A Unified Modular Agent Framework for Explainable Automatic Differential Diagnosis
abstract
Differential Diagnosis (DDx) is a fundamental yet complex aspect of clinical decision-making, in which physicians iteratively refine a ranked list of possible diseases based on symptoms, antecedents, and medical knowledge. While recent advances in large language models (LLMs) have shown promise in supporting DDx, existing approaches face key limitations, including single-dataset evaluations, isolated optimization of components, unrealistic assumptions about complete patient profiles, and single-attempt diagnosis. We introduce a Modular Explainable DDx Agent (MEDDxAgent) framework designed for interactive DDx, where diagnostic reasoning evolves through iterative learning, rather than assuming a complete patient profile is accessible. MEDDxAgent integrates three modular components: (1) an orchestrator (DDxDriver), (2) a history taking simulator, and (3) two specialized agents for knowledge retrieval and diagnosis strategy. To ensure robust evaluation, we introduce a comprehensive DDx benchmark covering respiratory, skin, and rare diseases. We analyze single-turn diagnostic approaches and demonstrate the importance of iterative refinement when patient profiles are not available at the outset. Our broad evaluation demonstrates that MEDDxAgent achieves over 10% accuracy improvements in interactive DDx across both large and small LLMs, while offering critical explainability into its diagnostic reasoning process.
Daniel Philip Rose, Chia-Chien Hung, Marco Lepri, Israa Alqassem, Kiril Gashteovski, Carolin Lawrence
ACL (1)5
2025 Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
abstract
Emmy Liu, Amanda Bertsch, Lintang Sutawika, Lindia Tjuatja, Patrick Fernandes, Lara Marinov, Michael Chen, Shreya Singhal, Carolin Lawrence, Aditi Raghunathan, Kiril Gashteovski, Graham Neubig. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Emmy Liu, Amanda Bertsch, Lintang Sutawika, Lindia Tjuatja, Patrick Fernandes, Lara Marinov, Shreya Singhal, Carolin Lawrence, Aditi Raghunathan, Kiril Gashteovski, Graham Neubig
EMNLP11
2024 A Human-Centric Assessment of the Usefulness of Attribution Methods in Computer Vision
Wiem Ben Rim, Ammar Shaker, Zhao Xu 0001, Kiril Gashteovski, Bhushan Kotnis, Carolin Lawrence, Jürgen Quittek, Sascha Saralajew
ECML/PKDD (5)4
2024 Large Language Models Enable Few-Shot Clustering
abstract
Abstract Unlike traditional unsupervised clustering, semi-supervised clustering allows users to provide meaningful structure to the data, which helps the clustering algorithm to match the user’s intent. Existing approaches to semi-supervised clustering require a significant amount of feedback from an expert to improve the clusters. In this paper, we ask whether a large language model (LLM) can amplify an expert’s guidance to enable query-efficient, few-shot semi-supervised text clustering. We show that LLMs are surprisingly effective at improving clustering. We explore three stages where LLMs can be incorporated into clustering: before clustering (improving input features), during clustering (by providing constraints to the clusterer), and after clustering (using LLMs post-correction). We find that incorporating LLMs in the first two stages routinely provides significant improvements in cluster quality, and that LLMs enable a user to make trade-offs between cost and accuracy to produce desired clusters. We release our code and LLM prompts for the public to use.1
Vijay Viswanathan 0002, Kiril Gashteovski, Carolin Lawrence, Sherry Tongshuang Wu, Graham Neubig
Trans. Assoc. Comput. Linguistics2
2023 Linking Surface Facts to Large-Scale Knowledge Graphs
abstract
Open Information Extraction (OIE) methods extract facts from natural language text in the form of ("subject"; "relation"; "object") triples.These facts are, however, merely surface forms, the ambiguity of which impedes their downstream usage; e.g., the surface phrase "Michael Jordan" may refer to either the former basketball player or the university professor.Knowledge Graphs (KGs), on the other hand, contain facts in a canonical (i.e., unambiguous) form, but their coverage is limited by a static schema (i.e., a fixed set of entities and predicates).To bridge this gap, we need the best of both worlds: (i) high coverage of free-text OIEs, and (ii) semantic precision (i.e., monosemy) of KGs.In order to achieve this goal, we propose a new benchmark with novel evaluation protocols that can, for example, measure fact linking performance on a granular triple slot level, while also measuring if a system has the ability to recognize that a surface form has no match in the existing KG.Our extensive evaluation of several baselines shows that detection of out-of-KG entities and predicates is more difficult than accurate linking to existing ones, thus calling for more research efforts on this difficult task.We publicly release all resources (data, benchmark and code) 1 .
Gorjan Radevski, Kiril Gashteovski, Chia-Chien Hung, Carolin Lawrence, Goran Glavas
EMNLP2
2022 BenchIE: A Framework for Multi-Faceted Fact-Based Open Information Extraction Evaluation
abstract
Kiril Gashteovski, Mingying Yu, Bhushan Kotnis, Carolin Lawrence, Mathias Niepert, Goran Glavaš. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Kiril Gashteovski, Mingying Yu, Bhushan Kotnis, Carolin Lawrence, Mathias Niepert, Goran Glavas
ACL (1)1
2022 MILIE: Modular & Iterative Multilingual Open Information Extraction
abstract
Bhushan Kotnis, Kiril Gashteovski, Daniel Rubio, Ammar Shaker, Vanesa Rodriguez-Tembras, Makoto Takamoto, Mathias Niepert, Carolin Lawrence. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Bhushan Kotnis, Kiril Gashteovski, Daniel Oñoro-Rubio, Ammar Shaker, Vanesa Rodriguez-Tembras, Makoto Takamoto, Mathias Niepert, Carolin Lawrence
ACL (1)2
2020 Can We Predict New Facts with Open Knowledge Graph Embeddings? A Benchmark for Open Link Prediction
abstract
Open Information Extraction systems extract ("subject text", "relation text", "object text") triples from raw text.Some triples are textual versions of facts, i.e., non-canonicalized mentions of entities and relations.In this paper, we investigate whether it is possible to infer new facts directly from the open knowledge graph without any canonicalization or any supervision from curated knowledge.For this purpose, we propose the open link prediction task, i.e., predicting test facts by completing ("subject text", "relation text", ?) questions.An evaluation in such a setup raises the question if a correct prediction is actually a new fact that was induced by reasoning over the open knowledge graph or if it can be trivially explained.For example, facts can appear in different paraphrased textual variants, which can lead to test leakage.To this end, we propose an evaluation protocol and a methodology for creating the open link prediction benchmark OLPBENCH.We performed experiments with a prototypical knowledge graph embedding model for open link prediction.While the task is very challenging, our results suggests that it is possible to predict genuinely new facts, which can not be trivially explained.
Samuel Broscheit, Kiril Gashteovski, Rainer Gemulla
ACL2
2017 MinIE: Minimizing Facts in Open Information Extraction
abstract
The goal of Open Information Extraction (OIE) is to extract surface relations and their arguments from naturallanguage text in an unsupervised, domainindependent manner.In this paper, we propose MinIE, an OIE system that aims to provide useful, compact extractions with high precision and recall.MinIE approaches these goals by (1) representing information about polarity, modality, attribution, and quantities with semantic annotations instead of in the actual extraction, and (2) identifying and removing parts that are considered overly specific.We conducted an experimental study with several real-world datasets and found that MinIE achieves competitive or higher precision and recall than most prior systems, while at the same time producing shorter, semantically enriched extractions.Pinocchio believes that the hero Superman was not actually born on beautiful Krypton.OLLIE 1 (Pinocchio, believes that, the hero [...] beautiful Krypton) 2 (Superman, was not actually born on, beautiful Krypton) 3 (Superman, was not actually born on beau.Krypton in, the hero) ClausIE 4 (Pinocchio, believes, that the hero [...] beautiful K.) 5 (the hero Superman, was not born, on beautiful Krypton) 6 (the hero Superman, was not born, on beautiful Krypton actually) Stanford OIE No extractions MinIE-C(om-7 (Superman, was born actually on, beautiful Krypton) plete) A.: fact.(-[not], CT), attrib.(Pinocchio, +, PS [believes]) 8 (Superman, was born on, beautiful Krypton) A.: fact.(-[not], CT), attrib.(Pinocchio, +, PS [believes]) 9 (Superman, "is", hero) A.: fact.(+, CT)
Kiril Gashteovski, Rainer Gemulla, Luciano Del Corro
EMNLP1