Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Joanna Baran

dblp:322/9101 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0000-0001-6792-7028ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Information extraction and text analysis · 100%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
1.012026
RetrySQL: Text-to-SQL Training with Retry Data for Self-Correcting Query Generation · AAAI 2026
Program synthesis and code generation
code generation with language models
0.312026
RetrySQL: Text-to-SQL Training with Retry Data for Self-Correcting Query Generation · AAAI 2026

Methods — techniques the papers use, named apart from their topics

self-correction · 2.0retry data · 2.0continuous pre-training · 2.0
YearPublicationVenuePosition
2026 RetrySQL: Text-to-SQL Training with Retry Data for Self-Correcting Query Generation
abstract
The text-to-SQL task is an active challenge in Natural Language Processing. Many existing solutions focus on using black-box language models extended with specialized components within customized end-to-end text-to-SQL pipelines. While these solutions use both closed-source proprietary language models and coding-oriented open-source models, there is a lack of research regarding SQL-specific small generative models. At the same time, recent advancements in self-correcting generation strategies show promise for improving the capabilities of existing architectures. The application of these concepts to the text-to-SQL task remains unexplored. In this paper, we introduce RetrySQL, a new approach to training text-to-SQL generation models. We prepare reasoning steps for reference SQL queries and then corrupt them to create retry data that contains both incorrect and corrected steps, divided with a special token. We continuously pre-train open-source coding models with this data and demonstrate that retry steps yield an improvements of up to 4 and 9 percentage points for overall and challenging execution metrics, respectively, as compared to pre-training without retry data. We showcase that the self-correcting behavior is learned by the model and the increase in downstream accuracy metrics is a result of this additional skill. Finally, we incorporate RetrySQL-trained models into the full text-to-SQL pipeline and showcase that they are competitive in terms of execution accuracy with proprietary models that contain orders of magnitude more parameters. RetrySQL demonstrates that self-correction can be learned in the text-to-SQL task and provides a novel way of improving generation accuracy for small SQL-oriented language models.
Alicja Raczkowska, Riccardo Belluzzo, Joanna Baran, Pawel Olszewski
AAAI4
2024 Refining Natural Language Inferences Using Cross-Document Structure Theory
Arkadiusz Janz, Dominik Kurowski, Joanna Baran, Julia Moska, Tomasz Bernas, Marcin Oleksy
ICCCI (1)3