Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Oleksandr Polozov

dblp:151/3318 · also Alex Polozov · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
7since 2021 · last 2023
0000-0003-3669-4262ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Language models and text generation · 36% Information extraction and text analysis · 21% Generative modeling · 8%
Software engineering, system software, and programming languages
10 papers
Program synthesis and code generation · 74% Compilers and program optimization · 13% Software maintenance and evolution · 12%
Databases, data mining, and information retrieval
3 papers
Data integration and cleaning · 45% Information retrieval · 26% Data models and query languages · 21%

Topics — the 30 heaviest of 43, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation
programming by example
1.452018
FlashProfile: a framework for synthesizing data profiles · Proc. ACM Program. Lang. 2018
Neural-Guided Deductive Search for Real-Time Program Synthesis from Examples · ICLR (Poster) 2018
Learning syntactic program transformations from examples · ICSE 2017
Natural language and speech › Information extraction and text analysis
semantic parsing
1.432021
SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing · ICLR 2021
RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers · ACL 2020
Learning Web-based Procedures by Reasoning over Explanations and Demonstrations in Context · ACL 2020
Natural language and speech › Language models and text generation
instruction following
1.122023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Learning Web-based Procedures by Reasoning over Explanations and Demonstrations in Context · ACL 2020
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Natural language and speech › Language models and text generation
large language model
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Natural language and speech › Language models and text generation
mathematical reasoning
0.712023
Learning Math Reasoning from Self-Sampled Correct and Partially-Correct Solutions · ICLR 2023
Machine learning › Deep learning architectures and training
scaling laws
0.712023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Program synthesis and code generation
code generation from natural language
0.712023
Natural Language to Code Generation in Interactive Data Science Notebooks · ACL (1) 2023
Program synthesis and code generation
code generation with language models
0.612022
Synchromesh: Reliable Code Generation from Pre-trained Language Models · ICLR 2022
Natural language and speech › Question answering and dialogue systems › dialogue understanding
conversational semantic parsing
0.512021
SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing · ICLR 2021
Machine learning › Representation and self-supervised learning
pre-training
0.512021
SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing · ICLR 2021
Machine learning › Generative modeling › generative model evaluation
realism evaluation
0.512021
KaggleDBQA: Realistic Evaluation of Text-to-SQL Parsers · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.512021
KaggleDBQA: Realistic Evaluation of Text-to-SQL Parsers · ACL/IJCNLP (1) 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning
0.412020
Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning" · ICML 2020
Natural language and speech › Information extraction and text analysis › semantic parsing
schema linking
0.412020
RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers · ACL 2020
Natural language and speech › Language models and text generation › natural language understanding › question answering
text-to-SQL parsing
0.412020
RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers · ACL 2020
Machine learning › Generative modeling
autoregressive model
0.412019
Generative Code Modeling with Graphs · ICLR (Poster) 2019
Compilers and program optimization
code generation
0.412019
Generative Code Modeling with Graphs · ICLR (Poster) 2019
Program synthesis and code generation
neural program synthesis
0.412019
Program Synthesis and Semantic Parsing with Learned Code Idioms · NeurIPS 2019
Program synthesis and code generation
inductive program synthesis
0.322018
FlashMeta: a framework for inductive program synthesis · OOPSLA 2015
FlashProfile: a framework for synthesizing data profiles · Proc. ACM Program. Lang. 2018
Software maintenance and evolution › refactoring
automated refactoring
0.312017
Learning syntactic program transformations from examples · ICSE 2017
Compilers and program optimization › program transformation
program transformation synthesis
0.312017
Learning syntactic program transformations from examples · ICSE 2017
Software maintenance and evolution
refactoring
0.312017
Learning syntactic program transformations from examples · ICSE 2017
Natural language and speech › Language models and text generation › text generation
math word problem generation
0.212015
Personalized Mathematical Word Problem Generation · IJCAI 2015
Natural language and speech › Language models and text generation › controllable text generation
personalized text generation
0.212015
Personalized Mathematical Word Problem Generation · IJCAI 2015
Learning and educational technologies
active learning
0.212015
User Interaction Models for Disambiguation in Programming by Example · UIST 2015
Program synthesis and code generation
domain-specific code generation
0.212015
FlashMeta: a framework for inductive program synthesis · OOPSLA 2015
Machine learning › Efficient and distributed learning
distributed training
0.212023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.212023
PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023

Methods — techniques the papers use, named apart from their topics

program synthesis · 1.5benchmarking · 1.0deductive synthesis · 0.9inverse semantics · 0.9demonstration learning · 0.9transformer · 0.7supervised fine-tuning · 0.7self-sampling · 0.7pathways · 0.7large language model · 0.7pre-training · 0.5BERT · 0.4tree-based neural synthesis · 0.4program graph representation · 0.4idiom mining · 0.4graph neural network · 0.4neural-guided search · 0.3inductive synthesis · 0.3
YearPublicationVenuePosition
2023 Natural Language to Code Generation in Interactive Data Science Notebooks
abstract
Pengcheng Yin, Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Oleksandr Polozov, Charles Sutton. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Wen-Ding Li, Kefan Xiao, Abhishek Rao, Yeming Wen, Kensen Shi, Joshua Howland, Paige Bailey, Michele Catasta, Henryk Michalewski, Oleksandr Polozov, Charles Sutton
ACL (1)11
2023 Learning Math Reasoning from Self-Sampled Correct and Partially-Correct Solutions
Ansong Ni, Jeevana Priya Inala, Chenglong Wang 0005, Oleksandr Polozov, Christopher Meek, Dragomir R. Radev, Jianfeng Gao 0001
ICLR4
2023 PaLM: Scaling Language Modeling with Pathways
abstract
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM). We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Adam Roberts, Paul Barham 0001, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du 0002, Ben Hutchinson, Reiner Pope, Jacob Austin, Michael Isard, Guy Gur-Ari, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, William Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang 0002, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeffrey Dean, Slav Petrov, Noah Fiedel
J. Mach. Learn. Res.54
2022 Synchromesh: Reliable Code Generation from Pre-trained Language Models
Gabriel Poesia, Oleksandr Polozov, Vu Le 0002, Ashish Tiwari 0001, Gustavo Soares, Christopher Meek, Sumit Gulwani
ICLR2
2021 KaggleDBQA: Realistic Evaluation of Text-to-SQL Parsers
abstract
Chia-Hsuan Lee, Oleksandr Polozov, Matthew Richardson. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chia-Hsuan Lee 0001, Oleksandr Polozov, Matthew Richardson
ACL/IJCNLP (1)2
2021 SCoRe: Pre-Training for Context Representation in Conversational Semantic Parsing
Tao Yu 0009, Rui Zhang 0037, Oleksandr Polozov, Christopher Meek, Ahmed Awadallah 0001
ICLR3
2021 Structure-Grounded Pretraining for Text-to-SQL
abstract
Xiang Deng, Ahmed Hassan Awadallah, Christopher Meek, Oleksandr Polozov, Huan Sun, Matthew Richardson. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xiang Deng 0001, Ahmed Awadallah 0001, Christopher Meek, Oleksandr Polozov, Huan Sun 0001, Matthew Richardson
NAACL-HLT4
2020 Learning Web-based Procedures by Reasoning over Explanations and Demonstrations in Context
abstract
We explore learning web-based tasks from a human teacher through natural language explanations and a single demonstration. Our approach investigates a new direction for semantic parsing that models explaining a demonstration in a context, rather than mapping explanations to demonstrations. By leveraging the idea of inverse semantics from program synthesis to reason backwards from observed demonstrations, we ensure that all considered interpretations are consistent with executable actions in any context, thus simplifying the problem of search over logical forms. We present a dataset of explanations paired with demonstrations for web-based tasks. Our methods show better task completion rates than a supervised semantic parsing baseline (40% relative improvement on average), and are competitive with simple exploration-and-demonstration based methods, while requiring no exploration of the environment. In learning to align explanations with demonstrations, basic properties of natural language syntax emerge as learned behavior. This is an interesting example of pragmatic language acquisition without any linguistic annotation.
Oleksandr Polozov, Nebojsa Jojic, Christopher Meek
ACL2
2020 RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers
abstract
When translating natural language questions into SQL queries to answer questions from a database, contemporary semantic parsing models struggle to generalize to unseen database schemas.The generalization challenge lies in (a) encoding the database relations in an accessible way for the semantic parser, and (b) modeling alignment between database columns and their mentions in a given query.We present a unified framework, based on the relation-aware self-attention mechanism, to address schema encoding, schema linking, and feature representation within a text-to-SQL encoder.On the challenging Spider dataset this framework boosts the exact match accuracy to 57.2%, surpassing its best counterparts by 8.7% absolute improvement.Further augmented with BERT, it achieves the new state-of-the-art performance of 65.6% on the Spider leaderboard.In addition, we observe qualitative improvements in the model's understanding of schema linking and alignment.Our implementation will be open-sourced at https://github.com/Microsoft/rat-sql.
Bailin Wang, Richard Shin, Xiaodong Liu 0003, Oleksandr Polozov, Matthew Richardson
ACL4
2020 Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning"
Saeed Amizadeh, Hamid Palangi, Oleksandr Polozov, Kazuhito Koishida
ICML3
2019 Generative Code Modeling with Graphs
Marc Brockschmidt, Miltiadis Allamanis, Alexander L. Gaunt, Oleksandr Polozov
ICLR (Poster)4
2019 Program Synthesis and Semantic Parsing with Learned Code Idioms
abstract
Program synthesis of general-purpose source code from natural language specifications is challenging due to the need to reason about high-level patterns in the target program and low-level implementation details at the same time. In this work, we present Patois, a system that allows a neural program synthesizer to explicitly interleave high-level and low-level reasoning at every generation step. It accomplishes this by automatically mining common code idioms from a given corpus, incorporating them into the underlying language for neural synthesis, and training a tree-based neural synthesizer to use these idioms during code generation. We evaluate Patois on two complex semantic parsing datasets and show that using learned code idioms improves the synthesizer's accuracy.
Richard Shin, Miltiadis Allamanis, Marc Brockschmidt, Oleksandr Polozov
NeurIPS4
2018 Neural-Guided Deductive Search for Real-Time Program Synthesis from Examples
Ashwin Kalyan, Abhishek Mohta, Oleksandr Polozov, Dhruv Batra, Prateek Jain 0002, Sumit Gulwani
ICLR (Poster)3
2018 FlashProfile: a framework for synthesizing data profiles
abstract
We address the problem of learning a syntactic profile for a collection of strings, i.e. a set of regex-like patterns that succinctly describe the syntactic variations in the strings. Real-world datasets, typically curated from multiple sources, often contain data in various syntactic formats. Thus, any data processing task is preceded by the critical step of data format identification. However, manual inspection of data to identify the different formats is infeasible in standard big-data scenarios. Prior techniques are restricted to a small set of pre-defined patterns (e.g. digits, letters, words etc.), and provide no control over granularity of profiles. We define syntactic profiling as a problem of clustering strings based on syntactic similarity, followed by identifying patterns that succinctly describe each cluster. We present a technique for synthesizing such profiles over a given language of patterns, that also allows for interactive refinement by requesting a desired number of clusters. Using a state-of-the-art inductive synthesis framework, PROSE, we have implemented our technique as FlashProfile. Across 153 tasks over 75 large real datasets, we observe a median profiling time of only ∼ 0.7s. Furthermore, we show that access to syntactic profiles may allow for more accurate synthesis of programs, i.e. using fewer examples, in programming-by-example (PBE) workflows such as Flash Fill.
Saswat Padhi, Prateek Jain 0002, Daniel Perelman, Oleksandr Polozov, Sumit Gulwani, Todd D. Millstein
Proc. ACM Program. Lang.4
2017 Learning syntactic program transformations from examples
abstract
Automatic program transformation tools can be valuable for programmers to help them with refactoring tasks, and for Computer Science students in the form of tutoring systems that suggest repairs to programming assignments. However, manually creating catalogs of transformations is complex and time-consuming. In this paper, we present REFAZER, a technique for automatically learning program transformations. REFAZER builds on the observation that code edits performed by developers can be used as input-output examples for learning program transformations. Example edits may share the same structure but involve different variables and subexpressions, which must be generalized in a transformation at the right level of abstraction. To learn transformations, REFAZER leverages state-of-the-art programming-by-example methodology using the following key components: (a) a novel domain-specific language (DSL) for describing program transformations, (b) domain-specific deductive algorithms for efficiently synthesizing transformations in the DSL, and (c) functions for ranking the synthesized transformations. We instantiate and evaluate REFAZER in two domains. First, given examples of code edits used by students to fix incorrect programming assignment submissions, we learn program transformations that can fix other students' submissions with similar faults. In our evaluation conducted on 4 programming tasks performed by 720 students, our technique helped to fix incorrect submissions for 87% of the students. In the second domain, we use repetitive code edits applied by developers to the same project to synthesize a program transformation that applies these edits to other locations in the code. In our evaluation conducted on 56 scenarios of repetitive edits taken from three large C# open-source projects, REFAZER learns the intended program transformation in 84% of the cases using only 2.9 examples on average.
Reudismam Rolim de Sousa, Gustavo Soares, Loris D'Antoni, Oleksandr Polozov, Sumit Gulwani, Rohit Gheyi, Ryo Suzuki 0001, Björn Hartmann
ICSE4
2015 Personalized Mathematical Word Problem Generation
Oleksandr Polozov, Eleanor O'Rourke, Adam M. Smith 0001, Luke Zettlemoyer, Sumit Gulwani, Zoran Popovic
IJCAI1
2015 FlashMeta: a framework for inductive program synthesis
abstract
Inductive synthesis, or programming-by-examples (PBE) is gaining prominence with disruptive applications for automating repetitive tasks in end-user programming. However, designing, developing, and maintaining an effective industrial-quality inductive synthesizer is an intellectual and engineering challenge, requiring 1-2 man-years of effort. Our novel observation is that many PBE algorithms are a natural fall-out of one generic meta-algorithm and the domain-specific properties of the operators in the underlying domain-specific language (DSL). The meta-algorithm propagates example-based constraints on an expression to its subexpressions by leveraging associated witness functions, which essentially capture the inverse semantics of the underlying operator. This observation enables a novel program synthesis methodology called data-driven domain-specific deduction (D4), where domain-specific insight, provided by the DSL designer, is separated from the synthesis algorithm. Our FlashMeta framework implements this methodology, allowing synthesizer developers to generate an efficient synthesizer from the mere DSL definition (if properties of the DSL operators have been modeled). In our case studies, we found that 10+ existing industrial-quality mass-market applications based on PBE can be cast as instances of D4. Our evaluation includes reimplementation of some prior works, which in FlashMeta become more efficient, maintainable, and extensible. As a result, FlashMeta-based PBE tools are deployed in several industrial products, including Microsoft PowerShell 3.0 for Windows 10, Azure Operational Management Suite, and Microsoft Cortana digital assistant.
Oleksandr Polozov, Sumit Gulwani
OOPSLA1
2015 User Interaction Models for Disambiguation in Programming by Example
abstract
Programming by Examples (PBE) has the potential to revolutionize end-user programming by enabling end users, most of whom are non-programmers, to create small scripts for automating repetitive tasks. However, examples, though often easy to provide, are an ambiguous specification of the user's intent. Because of that, a key impedance in adoption of PBE systems is the lack of user confidence in the correctness of the program that was synthesized by the system. We present two novel user interaction models that communicate actionable information to the user to help resolve ambiguity in the examples. One of these models allows the user to effectively navigate between the huge set of programs that are consistent with the examples provided by the user. The other model uses active learning to ask directed example-based questions to the user on the test input data over which the user intends to run the synthesized program. Our user studies show that each of these models significantly reduces the number of errors in the performed task without any difference in completion time. Moreover, both models are perceived as useful, and the proactive active-learning based model has a slightly higher preference regarding the users' confidence in the result.
Mikaël Mayer, Gustavo Soares, Maxim Grechkin, Vu Le 0002, Mark Marron, Oleksandr Polozov, Rishabh Singh, Benjamin G. Zorn, Sumit Gulwani
UIST6
2014 LaSEWeb: automating search strategies over semi-structured web data
abstract
We show how to programmatically model processes that humans use when extracting answers to queries (e.g., "Who invented typewriter?", "List of Washington national parks") from semi-structured Web pages returned by a search engine. This modeling enables various applications including automating repetitive search tasks, and helping search engine developers design micro-segments of factoid questions.
Oleksandr Polozov, Sumit Gulwani
KDD1