EDBT 2026 Demo / reviewers in the wild / expert
Mukul Singh
dblp:291/1609
· DBLP profile ↗
10ranked-venue papers
8as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Execution-guided within-prompt search for programming-by-exampleabstractLarge language models (LLMs) can generate code from examples without being limited to a DSL, but they lack search, as sampled programs are independent.
In this paper, we use an LLM as a policy that generates lines of code and then join these lines of code to let the LLM implicitly estimate the value of each of these lines in its next iteration.
We further guide the policy and value estimation by executing each line and annotating it with its results on the given examples.
This allows us to search for programs within a single (expanding) prompt until a sound program is found, by letting the policy reason in both the syntactic (code) and semantic (execution) space.
We evaluate within-prompt search on straight-line Python code generation using five benchmarks across different domains (strings, lists, and arbitrary Python programming problems).
We show that the model uses the execution results to guide the search and that within-prompt search performs well at low token budgets.
We also analyze how the model behaves as a policy and value, show that it can parallelize the search, and that it can implicitly backtrack over earlier generations. Gust Verbruggen, Ashish Tiwari 0001, Mukul Singh, Vu Le 0002, Sumit Gulwani |
ICLR | 3 |
| 2025 | DataVinci: Learning Syntactic and Semantic String RepairsabstractString data is common in real-world datasets: 67.6% of values in a sample of 1.8 million real Excel spreadsheets from the web were represented as text. Automatically cleaning such string data can have a significant impact on users. Previous approaches are limited to error detection, require that the user provides annotations, examples, or constraints to fix the errors, and focus independently on syntactic errors or semantic errors in strings, but ignore that strings often contain both syntactic and semantic substrings. We introduce DataVinci, a fully unsupervised string data error detection and repair system. DataVinci learns regular-expression-based patterns that cover a majority of values in a column and reports values that do not satisfy such majority patterns as data errors. DataVinci can automatically derive edits to the data error based on the majority patterns and using row tuples associated with majority values as examples. To handle strings with both syntactic and semantic substrings, DataVinci uses an LLM to abstract (and re-concretize) portions of strings that are semantic. Because not all data columns can result in majority patterns, when available, DataVinci can leverage execution information from an existing data program (which uses the target data as input) to identify and correct data repairs that would not otherwise be identified. DataVinci outperforms eleven baseline systems on both data error detection and repair as demonstrated on four existing and new benchmarks. Mukul Singh, José Cambronero, Sumit Gulwani, Vu Le 0002, Carina Negreanu, Arjun Radhakrishna, Gust Verbruggen |
Proc. ACM Manag. Data | 1 |
| 2024 | EmFORE: Learning Email Folder Classification Rules by DemonstrationabstractTools that help with email folder management are limited, as users have to manually write rules to assign emails to folders. We present EMFORE, an iterative learning system that automatically learns and updates such rules from observations. EMFORE is fast enough to suggest and update rules in real time and suppresses mails with low confidence to reduce the number of false positives. EMFORE can use different rule grammars, and thus be adapted to different clients, without changing the user experience. Previous methods do not learn rules, require complete retraining or multiple new examples after making a mistake, and do not distinguish between inbox and other folders. EMFORE learns rules incrementally and can make the neutral decision of leaving emails in the inbox, making it an ideal candidate for integration in email clients. Mukul Singh, Gust Verbruggen, José Cambronero, Vu Le 0002, Sumit Gulwani |
AAAI | 1 |
| 2024 | Tabularis Revilio: Converting Text to Tables
Mukul Singh, Gust Verbruggen, Vu Le 0002, Sumit Gulwani |
CIKM | 1 |
| 2024 | RAR: Retrieval-augmented retrieval for code generation in low resource languagesabstractLanguage models struggle in generating code for low-resource programming languages, since these are underrepresented in training data.Either examples or documentation are commonly used for improved code generation.We propose to use both types of information together and present retrieval augmented retrieval (RAR) as a two-step method for selecting relevant examples and documentation.Experiments on three low-resource languages (Power Query M, OfficeScript and Excel formulas) show that RAR outperforms independent example and grammar retrieval (+2.81-26.14%).Interestingly, we show that two-step retrieval selects better examples and documentation when used independently as well. Avik Dutta, Mukul Singh, Gust Verbruggen, Sumit Gulwani, Vu Le 0002 |
EMNLP | 2 |
| 2023 | EmFore: Online Learning of Email Folder Classification RulesabstractModern email clients support predicate-based folder assignment rules that can automatically organize emails. Unfortunately, users still need to write these rules manually. Prior machine learning approaches have framed automatically assigning email to folders as a classification task and do not produce symbolic rules. Prior inductive logic programming (ILP) approaches, which generate symbolic rules, fail to learn efficiently in the online environment needed for email management. To close this gap, we present EmFORE, an online system that learns symbolic rules for email classification from observations. Our key insights to do this successfully are: (1) learning rules over a folder abstraction that supports quickly determining candidate predicates to add or replace terms in a rule, (2) ensuring that rules remain consistent with historical assignments, (3) ranking rule updates based on existing predicate and folder name similarity, and (4) building a rule suppression model to avoid surfacing low-confidence folder predictions while keeping the rule for future use. We evaluate on two popular public email corpora and compare to 13 baselines, including state-of-the-art folder assignment systems, incremental machine learning, ILP and transformer-based approaches. We find that EmFORE performs significantly better, updates four orders of magnitude faster, and is more robust than existing methods and baselines. Mukul Singh, José Cambronero, Sumit Gulwani, Vu Le 0002, Gust Verbruggen |
CIKM | 1 |
| 2023 | CodeFusion: A Pre-trained Diffusion Model for Code GenerationabstractImagine a developer who can only change their last line of code-how often would they have to start writing a function from scratch before it is correct?Auto-regressive models for code generation from natural language have a similar limitation: they do not easily allow reconsidering earlier tokens generated.We introduce CODEFUSION, a pre-trained diffusion code generation model that addresses this limitation by iteratively denoising a complete program conditioned on the encoded natural language.We evaluate CODEFUSION on the task of natural language to code generation for Bash, Python, and Microsoft Excel conditional formatting (CF) rules.Experiments show that CODEFU-SION (75M parameters) performs on par with state-of-the-art auto-regressive systems (350M-175B parameters) in top-1 accuracy and outperforms them in top-3 and top-5 accuracy, due to its better balance in diversity versus quality. Mukul Singh, José Cambronero, Sumit Gulwani, Vu Le 0002, Carina Negreanu, Gust Verbruggen |
EMNLP | 1 |
| 2023 | FormaT5: Abstention and Examples for Conditional Table Formatting with Natural LanguageabstractFormatting is an important property in tables for visualization, presentation, and analysis. Spreadsheet software allows users to automatically format their tables by writing data-dependent conditional formatting (CF) rules. Writing such rules is often challenging for users as it requires understanding and implementing the underlying logic. We present FormaT5, a transformer-based model that can generate a CF rule given the target table and a natural language description of the desired formatting logic. We find that user descriptions for these tasks are often under-specified or ambiguous, making it harder for code generation systems to accurately learn the desired rule in a single step. To tackle this problem of under-specification and minimise argument errors, FormaT5 learns to predict placeholders though an abstention objective. These placeholders can then be filled by a second model or, when examples of rows that should be formatted are available, by a programming-by-example system. To evaluate FormaT5 on diverse and real scenarios, we create an extensive benchmark of 1053 CF tasks, containing real-world descriptions collected from four different sources. We release our benchmarks to encourage research in this area. Abstention and filling allow FormaT5 to outperform 8 different neural approaches on our benchmarks, both with and without examples. Our results illustrate the value of building domain-specific learning systems. Mukul Singh, José Cambronero, Sumit Gulwani, Vu Le 0002, Carina Negreanu, Elnaz Nouri, Mohammad Raza, Gust Verbruggen |
Proc. VLDB Endow. | 1 |
| 2023 | CORNET: Learning Table Formatting Rules By ExampleabstractSpreadsheets are widely used for table manipulation and presentation. Stylistic formatting of these tables is an important property for presentation and analysis. As a result, popular spreadsheet software, such as Excel, supports automatically formatting tables based on rules. Unfortunately, writing such formatting rules can be challenging for users as it requires knowledge of the underlying rule language and data logic. We present Cornet, a system that tackles the novel problem of automatically learning such formatting rules from user-provided formatted cells. Cornet takes inspiration from advances in inductive programming and combines symbolic rule enumeration with a neural ranker to learn conditional formatting rules. To motivate and evaluate our approach, we extracted tables with over 450K unique formatting rules from a corpus of over 1.8M real worksheets. Since we are the first to introduce the task of automatically learning conditional formatting rules, we compare Cornet to a wide range of symbolic and neural baselines adapted from related domains. Our results show that Cornet accurately learns rules across varying setups. Additionally, we show that in some cases Cornet can find rules that are shorter than those written by users and can also discover rules in spreadsheets that users have manually formatted. Furthermore, we present two case studies investigating the generality of our approach by extending Cornet to related data tasks (e.g., filtering) and generalizing to conditional formatting over multiple columns. Mukul Singh, José Cambronero, Sumit Gulwani, Vu Le 0002, Carina Negreanu, Mohammad Raza, Gust Verbruggen |
Proc. VLDB Endow. | 1 |
| 2023 | CORNET: Learning Spreadsheet Formatting Rules By ExampleabstractData management and analysis tasks are often carried out using spreadsheet software. A popular feature in most spreadsheet platforms is the ability to define data-dependent formatting rules. These rules can express actions such as "color red all entries in a column that are negative" or "bold all rows not containing error or failure". Unfortunately, users who want to exercise this functionality need to manually write these conditional formatting (CF) rules. We introduce Cornet, a system that automatically learns such conditional formatting rules from user examples. Cornet takes inspiration from inductive program synthesis and combines symbolic rule enumeration, based on semi-supervised clustering and iterative decision tree learning, with a neural ranker to produce accurate conditional formatting rules. In this demonstration, we show Cornet in action as a simple add-in to Microsoft's Excel. After the user provides one or two formatted cells as examples, Cornet generates formatting rule suggestions for the user to apply to the spreadsheet. Mukul Singh, José Cambronero, Sumit Gulwani, Vu Le 0002, Carina Negreanu, Gust Verbruggen |
Proc. VLDB Endow. | 1 |