Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Marija Sakota

dblp:255/5704 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-5192-5842ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 60% Efficient and distributed learning · 14% Information extraction and text analysis · 13%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 77% Information retrieval · 23%

Topics — the 5 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › decoding
constrained decoding
0.912025
Combining Constrained and Unconstrained Decoding via Boosting: BoostCD and Its Application to Information Extraction · EMNLP 2025
Natural language and speech › Language models and text generation › text generation › structured generation
structured text generation
0.912025
Combining Constrained and Unconstrained Decoding via Boosting: BoostCD and Its Application to Information Extraction · EMNLP 2025
Natural language and speech › Language models and text generation
model routing
0.812024
Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling · WSDM 2024
Machine learning › Generative modeling
synthetic data generation
0.712023
Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction · EMNLP 2023
Natural language and speech › Language models and text generation
text summarization
0.712023
Descartes: Generating Short Descriptions of Wikipedia Articles · WWW 2023

Methods — techniques the papers use, named apart from their topics

multilingual neural sequence-to-sequence model · 1.3knowledge graph semantic type features · 1.3boosting · 0.9autoregressive language model · 0.9meta-modeling · 0.8large language model prompting · 0.7fine-tuning · 0.7
YearPublicationVenuePosition
2025 Combining Constrained and Unconstrained Decoding via Boosting: BoostCD and Its Application to Information Extraction
abstract
Many recent approaches to structured NLP tasks use an autoregressive language model M to map unstructured input text x to output text y representing structured objects (such as tuples, lists, trees, code, etc.), where the desired output structure is enforced via constrained decoding.During training, these approaches do not require the model to be aware of the constraints, which are merely implicit in the training outputs y.This is advantageous as it allows for dynamic constraints without requiring retraining, but can lead to low-quality output during constrained decoding at test time.We overcome this problem with Boosted Constrained Decoding (BoostCD), which combines constrained and unconstrained decoding in two phases: Phase 1 decodes from the base model M twice, in constrained and unconstrained mode, obtaining two weak predictions.In phase 2, a learned autoregressive boosted model combines the two weak predictions into one final prediction.The mistakes made by the base model with vs. without constraints tend to be complementary, which the boosted model learns to exploit for improved performance.We demonstrate the power of BoostCD by applying it to closed information extraction.Our model, BoostIE, outperforms prior approaches both in and out of distribution, addressing several common errors identified in those approaches.
Marija Sakota, Robert West 0001
EMNLP1
2024 Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling
abstract
Generative language models (LMs) have become omnipresent across data science. For a wide variety of tasks, inputs can be phrased as natural language prompts for an LM, from whose output the solution can then be extracted. LM performance has consistently been increasing with model size---but so has the monetary cost of querying the ever larger models. Importantly, however, not all inputs are equally hard: some require larger LMs for obtaining a satisfactory solution, whereas for others smaller LMs suffice. Based on this fact, we design a framework for cost effective language model choice, called ''Fly-swat or cannon'' (FORC). Given a set of inputs and a set of candidate LMs, FORC judiciously assigns each input to an LM predicted to do well on the input according to a so-called meta-model, aiming to achieve high overall performance at low cost. The cost--performance tradeoff can be flexibly tuned by the user. Options include, among others, maximizing total expected performance (or the number of processed inputs) while staying within a given cost budget, or minimizing total cost while processing all inputs. We evaluate FORC on 14 datasets covering five natural language tasks, using four candidate LMs of vastly different size and cost. With FORC, we match the performance of the largest available LM while achieving a cost reduction of 63%. Via our publicly available library, (https://github.com/epfl-dlab/forc) researchers as well as practitioners can thus save large amounts of money without sacrificing performance.
Marija Sakota, Maxime Peyrard, Robert West 0001
WSDM1
2023 Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction
abstract
Large language models (LLMs) have great potential for synthetic data generation.This work shows that useful data can be synthetically generated even for tasks that cannot be solved directly by LLMs: for problems with structured outputs, it is possible to prompt an LLM to perform the task in the reverse direction, by generating plausible input text for a target output structure.Leveraging this asymmetry in task difficulty makes it possible to produce largescale, high-quality data for complex tasks.We demonstrate the effectiveness of this approach on closed information extraction, where collecting ground-truth data is challenging, and no satisfactory dataset exists to date.We synthetically generate a dataset of 1.8M data points, establish its superior quality compared to existing datasets in a human evaluation, and use it to finetune small models (220M and 770M parameters), termed SynthIE, that outperform the prior state of the art (with equal model size) by a substantial margin of 57 absolute points in micro-F1 and 79 points in macro-F1.Code, data, and models are available at https://github.com/epfl-dlab/SynthIE.
Martin Josifoski, Marija Sakota, Maxime Peyrard, Robert West 0001
EMNLP2
2023 Descartes: Generating Short Descriptions of Wikipedia Articles
abstract
Wikipedia is one of the richest knowledge sources on the Web today. In order to facilitate navigating, searching, and maintaining its content, Wikipedia’s guidelines state that all articles should be annotated with a so-called short description indicating the article’s topic (e.g., the short description of beer is “Alcoholic drink made from fermented cereal grains”). Nonetheless, a large fraction of articles (ranging from 10.2% in Dutch to 99.7% in Kazakh) have no short description yet, with detrimental effects for millions of Wikipedia users. Motivated by this problem, we introduce the novel task of automatically generating short descriptions for Wikipedia articles and propose Descartes, a multilingual model for tackling it. Descartes integrates three sources of information to generate an article description in a target language: the text of the article in all its language versions, the already-existing descriptions (if any) of the article in other languages, and semantic type information obtained from a knowledge graph. We evaluate a Descartes model trained for handling 25 languages simultaneously, showing that it beats baselines (including a strong translation-based baseline) and performs on par with monolingual models tailored for specific languages. A human evaluation on three languages further shows that the quality of Descartes’s descriptions is largely indistinguishable from that of human-written descriptions; e.g., 91.3% of our English descriptions (vs. 92.1% of human-written descriptions) pass the bar for inclusion in Wikipedia, suggesting that Descartes is ready for production, with the potential to support human editors in filling a major gap in today’s Wikipedia across languages.
Marija Sakota, Maxime Peyrard, Robert West 0001
WWW1