VLDB 2026 Research / reviewers in the wild / expert
Oleksandr Polozov
dblp:151/3318 · also Alex Polozov
· DBLP profile ↗
19ranked-venue papers
3as first author
7since 2021 · last 2023
0000-0003-3669-4262ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Language models and text generation · 36% Information extraction and text analysis · 21% Generative modeling · 8% | |
| Software engineering, system software, and programming languages
10 papers |
Program synthesis and code generation · 74% Compilers and program optimization · 13% Software maintenance and evolution · 12% | |
| Databases, data mining, and information retrieval
3 papers |
Data integration and cleaning · 45% Information retrieval · 26% Data models and query languages · 21% |
Topics — the 30 heaviest of 43, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program synthesis and code generation
programming by example |
1.4 | 5 | 2018 | FlashProfile: a framework for synthesizing data profiles · Proc. ACM Program. Lang. 2018 Neural-Guided Deductive Search for Real-Time Program Synthesis from Examples · ICLR (Poster) 2018 Learning syntactic program transformations from examples · ICSE 2017 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
1.4 | 3 | 2021 | SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing · ICLR 2021 RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers · ACL 2020 Learning Web-based Procedures by Reasoning over Explanations and Demonstrations in Context · ACL 2020 |
Natural language and speech › Language models and text generation
instruction following |
1.1 | 2 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 Learning Web-based Procedures by Reasoning over Explanations and Demonstrations in Context · ACL 2020 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.7 | 1 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Machine learning › Transfer learning and domain adaptation
few-shot learning |
0.7 | 1 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Natural language and speech › Language models and text generation
large language model |
0.7 | 1 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.7 | 1 | 2023 | Learning Math Reasoning from Self-Sampled Correct and Partially-Correct Solutions · ICLR 2023 |
Machine learning › Deep learning architectures and training
scaling laws |
0.7 | 1 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Program synthesis and code generation
code generation from natural language |
0.7 | 1 | 2023 | Natural Language to Code Generation in Interactive Data Science Notebooks · ACL (1) 2023 |
Program synthesis and code generation
code generation with language models |
0.6 | 1 | 2022 | Synchromesh: Reliable Code Generation from Pre-trained Language Models · ICLR 2022 |
Natural language and speech › Question answering and dialogue systems › dialogue understanding
conversational semantic parsing |
0.5 | 1 | 2021 | SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing · ICLR 2021 |
Machine learning › Representation and self-supervised learning
pre-training |
0.5 | 1 | 2021 | SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing · ICLR 2021 |
Machine learning › Generative modeling › generative model evaluation
realism evaluation |
0.5 | 1 | 2021 | KaggleDBQA: Realistic Evaluation of Text-to-SQL Parsers · ACL/IJCNLP (1) 2021 |
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL |
0.5 | 1 | 2021 | KaggleDBQA: Realistic Evaluation of Text-to-SQL Parsers · ACL/IJCNLP (1) 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning |
0.4 | 1 | 2020 | Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning" · ICML 2020 |
Natural language and speech › Information extraction and text analysis › semantic parsing
schema linking |
0.4 | 1 | 2020 | RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers · ACL 2020 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
text-to-SQL parsing |
0.4 | 1 | 2020 | RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers · ACL 2020 |
Machine learning › Generative modeling
autoregressive model |
0.4 | 1 | 2019 | Generative Code Modeling with Graphs · ICLR (Poster) 2019 |
Compilers and program optimization
code generation |
0.4 | 1 | 2019 | Generative Code Modeling with Graphs · ICLR (Poster) 2019 |
Program synthesis and code generation
neural program synthesis |
0.4 | 1 | 2019 | Program Synthesis and Semantic Parsing with Learned Code Idioms · NeurIPS 2019 |
Program synthesis and code generation
inductive program synthesis |
0.3 | 2 | 2018 | FlashMeta: a framework for inductive program synthesis · OOPSLA 2015 FlashProfile: a framework for synthesizing data profiles · Proc. ACM Program. Lang. 2018 |
Software maintenance and evolution › refactoring
automated refactoring |
0.3 | 1 | 2017 | Learning syntactic program transformations from examples · ICSE 2017 |
Compilers and program optimization › program transformation
program transformation synthesis |
0.3 | 1 | 2017 | Learning syntactic program transformations from examples · ICSE 2017 |
Software maintenance and evolution
refactoring |
0.3 | 1 | 2017 | Learning syntactic program transformations from examples · ICSE 2017 |
Natural language and speech › Language models and text generation › text generation
math word problem generation |
0.2 | 1 | 2015 | Personalized Mathematical Word Problem Generation · IJCAI 2015 |
Natural language and speech › Language models and text generation › controllable text generation
personalized text generation |
0.2 | 1 | 2015 | Personalized Mathematical Word Problem Generation · IJCAI 2015 |
Learning and educational technologies
active learning |
0.2 | 1 | 2015 | User Interaction Models for Disambiguation in Programming by Example · UIST 2015 |
Program synthesis and code generation
domain-specific code generation |
0.2 | 1 | 2015 | FlashMeta: a framework for inductive program synthesis · OOPSLA 2015 |
Machine learning › Efficient and distributed learning
distributed training |
0.2 | 1 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Machine learning › Efficient and distributed learning › distributed training
model parallelism |
0.2 | 1 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Methods — techniques the papers use, named apart from their topics
program synthesis · 1.5benchmarking · 1.0deductive synthesis · 0.9inverse semantics · 0.9demonstration learning · 0.9transformer · 0.7supervised fine-tuning · 0.7self-sampling · 0.7pathways · 0.7large language model · 0.7pre-training · 0.5BERT · 0.4tree-based neural synthesis · 0.4program graph representation · 0.4idiom mining · 0.4graph neural network · 0.4neural-guided search · 0.3inductive synthesis · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Natural Language to Code Generation in Interactive Data Science NotebooksabstractPengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Oleksandr Polozov, Charles Sutton. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Oleksandr Polozov, Charles Sutton |
ACL (1) | 11 |
| 2023 | Learning Math Reasoning from Self-Sampled Correct and Partially-Correct Solutions
Ansong Ni, Jeevana Priya Inala, Chenglong Wang 0005, Oleksandr Polozov, Christopher Meek, Dragomir R. Radev, Jianfeng Gao 0001 |
ICLR | 4 |
| 2023 | PaLM: Scaling Language Modeling with PathwaysabstractLarge language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM). We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies. Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Adam Roberts, Paul Barham 0001, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du 0002, Ben Hutchinson, Reiner Pope, Jacob Austin, Michael Isard, Guy Gur-Ari, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, William Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang 0002, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeffrey Dean, Slav Petrov, Noah Fiedel |
J. Mach. Learn. Res. | 54 |
| 2022 | Synchromesh: Reliable Code Generation from Pre-trained Language Models
Gabriel Poesia, Oleksandr Polozov, Vu Le 0002, Ashish Tiwari 0001, Gustavo Soares, Christopher Meek, Sumit Gulwani |
ICLR | 2 |
| 2021 | KaggleDBQA: Realistic Evaluation of Text-to-SQL ParsersabstractChia-Hsuan Lee, Oleksandr Polozov, Matthew Richardson. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chia-Hsuan Lee 0001, Oleksandr Polozov, Matthew Richardson |
ACL/IJCNLP (1) | 2 |
| 2021 | SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing
Tao Yu 0009, Rui Zhang 0037, Oleksandr Polozov, Christopher Meek, Ahmed Awadallah 0001 |
ICLR | 3 |
| 2021 | Structure-Grounded Pretraining for Text-to-SQLabstractXiang Deng, Ahmed Hassan Awadallah, Christopher Meek, Oleksandr Polozov, Huan Sun, Matthew Richardson. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Xiang Deng 0001, Ahmed Awadallah 0001, Christopher Meek, Oleksandr Polozov, Huan Sun 0001, Matthew Richardson |
NAACL-HLT | 4 |
| 2020 | Learning Web-based Procedures by Reasoning over Explanations and Demonstrations in ContextabstractWe explore learning web-based tasks from a human teacher through natural language explanations and a single demonstration. Our approach investigates a new direction for semantic parsing that models explaining a demonstration in a context, rather than mapping explanations to demonstrations. By leveraging the idea of inverse semantics from program synthesis to reason backwards from observed demonstrations, we ensure that all considered interpretations are consistent with executable actions in any context, thus simplifying the problem of search over logical forms. We present a dataset of explanations paired with demonstrations for web-based tasks. Our methods show better task completion rates than a supervised semantic parsing baseline (40% relative improvement on average), and are competitive with simple exploration-and-demonstration based methods, while requiring no exploration of the environment. In learning to align explanations with demonstrations, basic properties of natural language syntax emerge as learned behavior. This is an interesting example of pragmatic language acquisition without any linguistic annotation. Oleksandr Polozov, Nebojsa Jojic, Christopher Meek |
ACL | 2 |
| 2020 | RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersabstractWhen translating natural language questions into SQL queries to answer questions from a database, contemporary semantic parsing models struggle to generalize to unseen database schemas.The generalization challenge lies in (a) encoding the database relations in an accessible way for the semantic parser, and (b) modeling alignment between database columns and their mentions in a given query.We present a unified framework, based on the relation-aware self-attention mechanism, to address schema encoding, schema linking, and feature representation within a text-to-SQL encoder.On the challenging Spider dataset this framework boosts the exact match accuracy to 57.2%, surpassing its best counterparts by 8.7% absolute improvement.Further augmented with BERT, it achieves the new state-of-the-art performance of 65.6% on the Spider leaderboard.In addition, we observe qualitative improvements in the model's understanding of schema linking and alignment.Our implementation will be open-sourced at https://github.com/Microsoft/rat-sql. Bailin Wang, Richard Shin, Xiaodong Liu 0003, Oleksandr Polozov, Matthew Richardson |
ACL | 4 |
| 2020 | Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning"
Saeed Amizadeh, Hamid Palangi, Oleksandr Polozov, Kazuhito Koishida |
ICML | 3 |
| 2019 | Generative Code Modeling with Graphs
Marc Brockschmidt, Miltiadis Allamanis, Alexander L. Gaunt, Oleksandr Polozov |
ICLR (Poster) | 4 |
| 2019 | Program Synthesis and Semantic Parsing with Learned Code IdiomsabstractProgram synthesis of general-purpose source code from natural language specifications is challenging due to the need to reason about high-level patterns in the target program and low-level implementation details at the same time. In this work, we present Patois, a system that allows a neural program synthesizer to explicitly interleave high-level and low-level reasoning at every generation step. It accomplishes this by automatically mining common code idioms from a given corpus, incorporating them into the underlying language for neural synthesis, and training a tree-based neural synthesizer to use these idioms during code generation. We evaluate Patois on two complex semantic parsing datasets and show that using learned code idioms improves the synthesizer's accuracy. Richard Shin, Miltiadis Allamanis, Marc Brockschmidt, Oleksandr Polozov |
NeurIPS | 4 |
| 2018 | Neural-Guided Deductive Search for Real-Time Program Synthesis from Examples
Ashwin Kalyan, Abhishek Mohta, Oleksandr Polozov, Dhruv Batra, Prateek Jain 0002, Sumit Gulwani |
ICLR (Poster) | 3 |
| 2018 | FlashProfile: a framework for synthesizing data profilesabstractWe address the problem of learning a syntactic profile for a collection of strings, i.e. a set of regex-like patterns that succinctly describe the syntactic variations in the strings. Real-world datasets, typically curated from multiple sources, often contain data in various syntactic formats. Thus, any data processing task is preceded by the critical step of data format identification. However, manual inspection of data to identify the different formats is infeasible in standard big-data scenarios. Prior techniques are restricted to a small set of pre-defined patterns (e.g. digits, letters, words etc.), and provide no control over granularity of profiles. We define syntactic profiling as a problem of clustering strings based on syntactic similarity, followed by identifying patterns that succinctly describe each cluster. We present a technique for synthesizing such profiles over a given language of patterns, that also allows for interactive refinement by requesting a desired number of clusters. Using a state-of-the-art inductive synthesis framework, PROSE, we have implemented our technique as FlashProfile. Across 153 tasks over 75 large real datasets, we observe a median profiling time of only ∼ 0.7s. Furthermore, we show that access to syntactic profiles may allow for more accurate synthesis of programs, i.e. using fewer examples, in programming-by-example (PBE) workflows such as Flash Fill. Saswat Padhi, Prateek Jain 0002, Daniel Perelman, Oleksandr Polozov, Sumit Gulwani, Todd D. Millstein |
Proc. ACM Program. Lang. | 4 |
| 2017 | Learning syntactic program transformations from examplesabstractAutomatic program transformation tools can be valuable for programmers to help them with refactoring tasks, and for Computer Science students in the form of tutoring systems that suggest repairs to programming assignments. However, manually creating catalogs of transformations is complex and time-consuming. In this paper, we present REFAZER, a technique for automatically learning program transformations. REFAZER builds on the observation that code edits performed by developers can be used as input-output examples for learning program transformations. Example edits may share the same structure but involve different variables and subexpressions, which must be generalized in a transformation at the right level of abstraction. To learn transformations, REFAZER leverages state-of-the-art programming-by-example methodology using the following key components: (a) a novel domain-specific language (DSL) for describing program transformations, (b) domain-specific deductive algorithms for efficiently synthesizing transformations in the DSL, and (c) functions for ranking the synthesized transformations. We instantiate and evaluate REFAZER in two domains. First, given examples of code edits used by students to fix incorrect programming assignment submissions, we learn program transformations that can fix other students' submissions with similar faults. In our evaluation conducted on 4 programming tasks performed by 720 students, our technique helped to fix incorrect submissions for 87% of the students. In the second domain, we use repetitive code edits applied by developers to the same project to synthesize a program transformation that applies these edits to other locations in the code. In our evaluation conducted on 56 scenarios of repetitive edits taken from three large C# open-source projects, REFAZER learns the intended program transformation in 84% of the cases using only 2.9 examples on average. Reudismam Rolim de Sousa, Gustavo Soares, Loris D'Antoni, Oleksandr Polozov, Sumit Gulwani, Rohit Gheyi, Ryo Suzuki 0001, Björn Hartmann |
ICSE | 4 |
| 2015 | Personalized Mathematical Word Problem Generation
Oleksandr Polozov, Eleanor O'Rourke, Adam M. Smith 0001, Luke Zettlemoyer, Sumit Gulwani, Zoran Popovic |
IJCAI | 1 |
| 2015 | FlashMeta: a framework for inductive program synthesisabstractInductive synthesis, or programming-by-examples (PBE) is gaining prominence with disruptive applications for automating repetitive tasks in end-user programming. However, designing, developing, and maintaining an effective industrial-quality inductive synthesizer is an intellectual and engineering challenge, requiring 1-2 man-years of effort. Our novel observation is that many PBE algorithms are a natural fall-out of one generic meta-algorithm and the domain-specific properties of the operators in the underlying domain-specific language (DSL). The meta-algorithm propagates example-based constraints on an expression to its subexpressions by leveraging associated witness functions, which essentially capture the inverse semantics of the underlying operator. This observation enables a novel program synthesis methodology called data-driven domain-specific deduction (D4), where domain-specific insight, provided by the DSL designer, is separated from the synthesis algorithm. Our FlashMeta framework implements this methodology, allowing synthesizer developers to generate an efficient synthesizer from the mere DSL definition (if properties of the DSL operators have been modeled). In our case studies, we found that 10+ existing industrial-quality mass-market applications based on PBE can be cast as instances of D4. Our evaluation includes reimplementation of some prior works, which in FlashMeta become more efficient, maintainable, and extensible. As a result, FlashMeta-based PBE tools are deployed in several industrial products, including Microsoft PowerShell 3.0 for Windows 10, Azure Operational Management Suite, and Microsoft Cortana digital assistant. Oleksandr Polozov, Sumit Gulwani |
OOPSLA | 1 |
| 2015 | User Interaction Models for Disambiguation in Programming by ExampleabstractProgramming by Examples (PBE) has the potential to revolutionize end-user programming by enabling end users, most of whom are non-programmers, to create small scripts for automating repetitive tasks. However, examples, though often easy to provide, are an ambiguous specification of the user's intent. Because of that, a key impedance in adoption of PBE systems is the lack of user confidence in the correctness of the program that was synthesized by the system. We present two novel user interaction models that communicate actionable information to the user to help resolve ambiguity in the examples. One of these models allows the user to effectively navigate between the huge set of programs that are consistent with the examples provided by the user. The other model uses active learning to ask directed example-based questions to the user on the test input data over which the user intends to run the synthesized program. Our user studies show that each of these models significantly reduces the number of errors in the performed task without any difference in completion time. Moreover, both models are perceived as useful, and the proactive active-learning based model has a slightly higher preference regarding the users' confidence in the result. Mikaël Mayer, Gustavo Soares, Maxim Grechkin, Vu Le 0002, Mark Marron, Oleksandr Polozov, Rishabh Singh, Benjamin G. Zorn, Sumit Gulwani |
UIST | 6 |
| 2014 | LaSEWeb: automating search strategies over semi-structured web dataabstractWe show how to programmatically model processes that humans use when extracting answers to queries (e.g., "Who invented typewriter?", "List of Washington national parks") from semi-structured Web pages returned by a search engine. This modeling enables various applications including automating repetitive search tasks, and helping search engine developers design micro-segments of factoid questions. Oleksandr Polozov, Sumit Gulwani |
KDD | 1 |