EDBT 2026 Demo / reviewers in the wild / expert
Joanna Baran
dblp:322/9101
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2026
0000-0001-6792-7028ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL |
1.0 | 1 | 2026 | RetrySQL: Text-to-SQL Training with Retry Data for Self-Correcting Query Generation · AAAI 2026 |
Program synthesis and code generation
code generation with language models |
0.3 | 1 | 2026 | RetrySQL: Text-to-SQL Training with Retry Data for Self-Correcting Query Generation · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
self-correction · 2.0retry data · 2.0continuous pre-training · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RetrySQL: Text-to-SQL Training with Retry Data for Self-Correcting Query GenerationabstractThe text-to-SQL task is an active challenge in Natural Language Processing. Many existing solutions focus on using black-box language models extended with specialized components within customized end-to-end text-to-SQL pipelines. While these solutions use both closed-source proprietary language models and coding-oriented open-source models, there is a lack of research regarding SQL-specific small generative models. At the same time, recent advancements in self-correcting generation strategies show promise for improving the capabilities of existing architectures. The application of these concepts to the text-to-SQL task remains unexplored. In this paper, we introduce RetrySQL, a new approach to training text-to-SQL generation models. We prepare reasoning steps for reference SQL queries and then corrupt them to create retry data that contains both incorrect and corrected steps, divided with a special token. We continuously pre-train open-source coding models with this data and demonstrate that retry steps yield an improvements of up to 4 and 9 percentage points for overall and challenging execution metrics, respectively, as compared to pre-training without retry data. We showcase that the self-correcting behavior is learned by the model and the increase in downstream accuracy metrics is a result of this additional skill. Finally, we incorporate RetrySQL-trained models into the full text-to-SQL pipeline and showcase that they are competitive in terms of execution accuracy with proprietary models that contain orders of magnitude more parameters. RetrySQL demonstrates that self-correction can be learned in the text-to-SQL task and provides a novel way of improving generation accuracy for small SQL-oriented language models. Alicja Raczkowska, Riccardo Belluzzo, Joanna Baran, Pawel Olszewski |
AAAI | 4 |
| 2024 | Refining Natural Language Inferences Using Cross-Document Structure Theory
Arkadiusz Janz, Dominik Kurowski, Joanna Baran, Julia Moska, Tomasz Bernas, Marcin Oleksy |
ICCCI (1) | 3 |