Sida I. Wang

dblp:153/9609 · also Sida Wang 0001 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
17since 2021 · last 2025
0000-0001-8101-8883ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 5 first-author · 17 since 2021
YearPublicationVenuePosition
2025 LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
abstract
Large Language Models (LLMs) applied to code-related applications have emerged as a prominent field, attracting significant interest from academia and industry. However, as new and improved LLMs are developed, existing evaluation benchmarks (e.g., HumanEvla, MBPP) are no longer sufficient for assessing their capabilities suffering from data contamination, overfitting, saturation, and focus on merely code generation. In this work, we propose LiveCodeBench, a comprehensive and contamination-free evaluation of LLMs for code, which collects new problems over time from contests across three competition platforms, Leetcode, Atcoder, and Codeforces. Notably, our benchmark also focuses on a broader range of code-related capabilities, such as self-repair, code execution, and test output prediction, beyond just code generation. Currently, LiveCodeBench hosts over six hundred coding problems that were published between May 2023 and Aug 2024. We evaluate over 50 LLMs on LiveCodeBench (LCB for brevity) presenting the largest evaluation study of code LLMs on competition problems. Based on the study, we present novel empirical findings on contamination, overfitting, and holistic evaluations. We demonstrate that time-segmented evaluations serve as a robust approach to evade contamination; they are successful at detecting contamination across a wide range of open and closed models including GPT-4O, Claude, Deepseek, and Codestral. Next, we highlight overfitting and saturation of traditional coding benchmarks like HumanEvla and demonstrate LCB allows more reliable evaluations. Finally, our holistic evaluation scenarios allow for measuring the different capabilities of programming agents in isolation.
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida I. Wang, Armando Solar-Lezama, Koushik Sen, Ion Stoica
ICLR7
2025 Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
abstract
Real-world enterprise text-to-SQL workflows often involve complex cloud or local data across various database systems, multiple SQL queries in various dialects, and diverse operations from data transformation to analytics. We introduce Spider 2.0, an evaluation framework comprising $632$ real-world text-to-SQL workflow problems derived from enterprise-level database use cases. The databases in Spider 2.0 are sourced from real data applications, often containing over 1,000 columns and stored in local or cloud database systems such as BigQuery and Snowflake. We show that solving problems in Spider 2.0 frequently requires understanding and searching through database metadata, dialect documentation, and even project-level codebases. This challenge calls for models to interact with complex SQL workflow environments, process extremely long contexts, perform intricate reasoning, and generate multiple SQL queries with diverse operations, often exceeding $100$ lines, which goes far beyond traditional text-to-SQL challenges. Our evaluations indicate that based on o1-preview, our code agent framework successfully solves only 21.3\% of the tasks, compared with 91.2\% on Spider 1.0 and 73.0\% on BIRD. Our results on Spider 2.0 show that while language models have demonstrated remarkable performance in code generation --- especially in prior text-to-SQL benchmarks --- they require significant improvement in order to achieve adequate performance for real-world enterprise usage. Progress on Spider 2.0 represents crucial steps towards developing intelligent, autonomous, code agents for real-world enterprise settings. Our code, baseline models, and data are available at [spider2-sql.github.io](spider2-sql.github.io) .
Fangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao, Dongchan Shin, Hongjin Su, Zhaoqing Suo, Hongcheng Gao, Wenjing Hu, Victor Zhong, Caiming Xiong, Ruoxi Sun 0002, Qian Liu 0033, Sida I. Wang, Tao Yu 0009
ICLR15
2025 SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
abstract
Autonomous systems for software engineering are now capable of fixing bugs and developing features. These systems are commonly evaluated on SWE-bench (Jimenez et al., 2024a), which assesses their ability to solve software issues from GitHub repositories. However, SWE-bench uses only Python repositories, with problem statements presented predominantly as text and lacking visual elements such as images. This limited coverage motivates our inquiry into how existing systems might perform on unrepresented software engineering domains (e.g., front-end, game development, DevOps), which use different programming languages and paradigms. Therefore, we propose SWE-bench Multimodal (SWE-bench M), to evaluate systems on their ability to fix bugs in visual, user-facing JavaScript software. SWE-bench M features 617 task instances collected from 17 JavaScript libraries used for web interface design, diagramming, data visualization, syntax highlighting, and interactive mapping. Each SWE-bench M task instance contains at least one image in its problem statement or unit tests. Our analysis finds that top-performing SWE-bench systems struggle with SWE-bench M, revealing limitations in visual problem-solving and cross-language generalization. Lastly, we show that SWE-agent’s flexible language-agnostic features enable it to substantially outperform alternatives on SWE-bench M, resolving 12% of task instances compared to 6% for the next best system.
John Yang 0002, Carlos E. Jimenez, Alex L. Zhang, Kilian Lieret, Joyce Yang, Xindi Wu, Ori Press, Niklas Muennighoff, Gabriel Synnaeve, Karthik Narasimhan, Diyi Yang, Sida I. Wang, Ofir Press
ICLR12
2025 SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
abstract
The recent DeepSeek-R1 release has demonstrated the immense potential of reinforcement learning (RL) in enhancing the general reasoning capabilities of large language models (LLMs). While DeepSeek-R1 and other follow-up work primarily focus on applying RL to competitive coding and math problems, this paper introduces SWE-RL, the first approach to scale RL-based LLM reasoning for real-world software engineering. Leveraging a lightweight rule-based reward (e.g., the similarity score between ground-truth and LLM-generated solutions), SWE-RL enables LLMs to autonomously recover a developer's reasoning processes and solutions by learning from extensive open-source software evolution data -- the record of a software's entire lifecycle, including its code snapshots, code changes, and events such as issues and pull requests. Trained on top of Llama 3, our resulting reasoning model, Llama3-SWE-RL-70B, achieves a 41.0% solve rate on SWE-bench Verified -- a human-verified collection of real-world GitHub issues. To our knowledge, this is the best performance reported for medium-sized (<100B) LLMs to date, even comparable to leading proprietary LLMs like GPT-4o. Surprisingly, despite performing RL solely on software evolution data, Llama3-SWE-RL has even emerged with generalized reasoning skills. For example, it shows improved results on five out-of-domain tasks, namely, function coding, library use, code reasoning, mathematics, and general language understanding, whereas a supervised-finetuning baseline even leads to performance degradation on average. Overall, SWE-RL opens up a new direction to improve the reasoning capabilities of LLMs through reinforcement learning on massive software engineering data.
Yuxiang Wei 0003, Olivier Duchenne, Jade Copet, Quentin Carbonneaux, Lingming Zhang 0001, Daniel Fried, Gabriel Synnaeve, Rishabh Singh, Sida I. Wang
NeurIPS9
2024 CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
abstract
We present Code Reasoning, Understanding, and eXecution Evaluation, a benchmark consisting of 800 Python functions (3-13 lines). Each function comes with an input-output pair, leading to two natural tasks: input prediction and output prediction. First, we propose a general recipe for generating our execution benchmark by sampling from a model, which can be used for more challenging versions of the benchmark if needed. Second, we evaluate twenty code models on our benchmark and discover that many recent high-scoring models on HumanEval show no improvements on our benchmark. Third, we show that simple CoT and fine-tuning schemes can improve performance on our benchmark but remain far from solving it. The best setup, GPT-4 with chain of thought (CoT), achieves a pass@1 of 75% and 81% on input and output prediction, respectively. In contrast, Code Llama 34B achieves a pass@1 of 50% and 46% on input and output prediction. When it comes to reasoning about code, GPT-4 has a huge edge over other models but still fails consistently on some surprisingly simple Python programs.
Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, Sida I. Wang
ICML6
2024 Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
abstract
Data science and engineering workflows often span multiple stages, from warehousing to orchestration, using tools like BigQuery, dbt, and Airbyte. As vision language models (VLMs) advance in multimodal understanding and code generation, VLM-based agents could potentially automate these workflows by generating SQL queries, Python code, and GUI operations. This automation can improve the productivity of experts while democratizing access to large-scale data analysis. In this paper, we introduce Spider2-V, the first multimodal agent benchmark focusing on professional data science and engineering workflows, featuring 494 real-world tasks in authentic computer environments and incorporating 20 enterprise-level professional applications. These tasks, derived from real-world use cases, evaluate the ability of a multimodal agent to perform data-related tasks by writing code and managing the GUI in enterprise data software systems. To balance realistic simulation with evaluation simplicity, we devote significant effort to developing automatic configurations for task setup and carefully crafting evaluation metrics for each task. Furthermore, we supplement multimodal agents with comprehensive documents of these enterprise data software systems. Our empirical evaluation reveals that existing state-of-the-art LLM/VLM-based agents do not reliably automate full data workflows (14.0% success). Even with step-by-step guidance, these agents still underperform in tasks that require fine-grained, knowledge-intensive GUI actions (16.2%) and involve remote cloud-hosted workspaces (10.6%). We hope that Spider2-V paves the way for autonomous multimodal agents to transform the automation of data science and engineering workflow. Our code and data are available at https://spider2-v.github.io.
Ruisheng Cao, Fangyu Lei, Haoyuan Wu, Jixuan Chen, Yeqiao Fu, Hongcheng Gao, Xinzhuang Xiong, Hanchong Zhang, Wenjing Hu, Tianbao Xie, Hongshen Xu, Sida I. Wang, Ruoxi Sun 0002, Caiming Xiong, Ansong Ni, Qian Liu 0033, Victor Zhong, Lu Chen 0002, Kai Yu 0004, Tao Yu 0009
NeurIPS14
2023 Learning to Simulate Natural Language Feedback for Interactive Semantic Parsing
abstract
Interactive semantic parsing based on natural language (NL) feedback, where users provide feedback to correct the parser mistakes, has emerged as a more practical scenario than the traditional one-shot semantic parsing.However, prior work has heavily relied on humanannotated feedback data to train the interactive semantic parser, which is prohibitively expensive and not scalable.In this work, we propose a new task of simulating NL feedback for interactive semantic parsing.We accompany the task with a novel feedback evaluator.The evaluator is specifically designed to assess the quality of the simulated feedback, based on which we decide the best feedback simulator from our proposed variants.On a text-to-SQL dataset, we show that our feedback simulator can generate high-quality NL feedback to boost the error correction ability of a specific parser.In low-data settings, our feedback simulator can help achieve comparable error correction performance as trained using the costly, full set of human annotations.1
Yintao Tai, Sida I. Wang, Scott Yih, Ziyu Yao 0002
ACL (1)4
2023 InCoder: A Generative Model for Code Infilling and Synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida I. Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, Mike Lewis
ICLR4
2023 DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation
abstract
We introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as Numpy and Pandas. Compared to prior works, DS-1000 incorporates three core features. First, our problems reflect diverse, realistic, and practical use cases since we collected them from StackOverflow. Second, our automatic evaluation is highly specific (reliable) – across all Codex-002-predicted solutions that our evaluation accepts, only 1.8% of them are incorrect; we achieve this with multi-criteria metrics, checking both functional correctness by running test cases and surface-form constraints by restricting API usages or keywords. Finally, we proactively defend against memorization by slightly modifying our problems to be different from the original StackOverflow source; consequently, models cannot answer them correctly by memorizing the solutions from pre-training. The current best public system (Codex-002) achieves 43.3% accuracy, leaving ample room for improvement. We release our benchmark at https://ds1000-code-gen.github.io.
Yuhang Lai, Chengxi Li 0011, Ruiqi Zhong, Luke Zettlemoyer, Scott Yih, Daniel Fried, Sida I. Wang, Tao Yu 0009
ICML9
2023 LEVER: Learning to Verify Language-to-Code Generation with Execution
abstract
The advent of large language models trained on code (code LLMs) has led to significant progress in language-to-code generation. State-of-the-art approaches in this area combine LLM decoding with sample pruning and reranking using test cases or heuristics based on the execution results. However, it is challenging to obtain test cases for many real-world language-to-code applications, and heuristics cannot well capture the semantic features of the execution results, such as data type and value range, which often indicates the correctness of the program. In this work, we propose LEVER, a simple approach to improve language-to-code generation by learning to verify the generated programs with their execution results. Specifically, we train verifiers to determine whether a program sampled from the LLMs is correct or not based on the natural language input, the program itself and its execution results. The sampled programs are reranked by combining the verification score with the LLM generation probability, and marginalizing over programs with the same execution results. On four datasets across the domains of table QA, math QA and basic Python programming, LEVER consistently improves over the base code LLMs (4.6% to 10.9% with code-davinci-002) and achieves new state-of-the-art results on all of them.
Ansong Ni, Srinivasan Iyer 0001, Dragomir R. Radev, Veselin Stoyanov, Scott Yih, Sida I. Wang, Xi Victoria Lin
ICML6
2023 Coder Reviewer Reranking for Code Generation
abstract
Sampling diverse programs from a code language model and reranking with model likelihood is a popular method for code generation but it is prone to preferring degenerate solutions. Inspired by collaborative programming, we propose Coder-Reviewer reranking. We augment Coder language models from past work, which generate programs given language instructions, with Reviewer models, which evaluate the likelihood of the instruction given the generated programs. We perform an extensive study across six datasets with eight models from three model families. Experimental results show that Coder-Reviewer reranking leads to consistent and significant improvement (up to 17% absolute accuracy gain) over reranking with the Coder model only. When combined with executability filtering, Coder-Reviewer reranking can often outperform the minimum Bayes risk method. Coder-Reviewer reranking is easy to implement by prompting, can generalize to different programming languages, and works well with off-the-shelf hyperparameters.
Tao Yu 0009, Tatsunori B. Hashimoto, Mike Lewis, Scott Yih, Daniel Fried, Sida I. Wang
ICML7
2022 On Continual Model Refinement in Out-of-Distribution Data Streams
abstract
Real-world natural language processing (NLP) models need to be continually updated to fix the prediction errors in out-of-distribution (OOD) data streams while overcoming catastrophic forgetting.However, existing continual learning (CL) problem setups cannot cover such a realistic and complex scenario.In response to this, we propose a new CL problem formulation dubbed continual model refinement (CMR).Compared to prior CL settings, CMR is more practical and introduces unique challenges (boundary-agnostic and non-stationary distribution shift, diverse mixtures of multiple OOD data clusters, error-centric streams, etc.).We extend several existing CL approaches to the CMR setting and evaluate them extensively.For benchmarking and analysis, we propose a general sampling algorithm to obtain dynamic OOD data streams with controllable nonstationarity, as well as a suite of metrics measuring various aspects of online performance.Our experiments and detailed analysis reveal the promise and challenges of the CMR problem, supporting that studying CMR in dynamic OOD streams can benefit the longevity of deployed NLP models in production. 1
Bill Y. Lin, Sida I. Wang, Xi Victoria Lin, Robin Jia, Xiang Ren 0001, Scott Yih
ACL (1)2
2022 Natural Language to Code Translation with Execution
abstract
Generative models of code, pretrained on large corpora of programs, have shown great success in translating natural language to code (Chen et al., 2021;Austin et al., 2021; Li et al., 2022, inter alia).While these models do not explicitly incorporate program semantics (i.e., execution results) during training, they are able to generate correct solutions for many problems.However, choosing a single correct program from a generated set for each problem remains challenging.In this work, we introduce execution resultbased minimum Bayes risk decoding (MBR-EXEC) for program selection and show that it improves the few-shot performance of pretrained code models on natural-language-tocode tasks.We select output programs from a generated candidate set by marginalizing over program implementations that share the same semantics.Because exact equivalence is intractable, we execute each program on a small number of test inputs to approximate semantic equivalence.Across datasets, execution or simulated execution significantly outperforms the methods that do not involve program semantics.We find that MBR-EXEC consistently improves over all execution-unaware selection methods, suggesting it as an effective approach for natural language to code translation.1
Freda Shi, Daniel Fried, Marjan Ghazvininejad, Luke Zettlemoyer, Sida I. Wang
EMNLP5
2022 UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models
abstract
Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, Tao Yu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Tianbao Xie, Chen Henry Wu, Peng Shi 0010, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong 0005, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao 0002, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang 0037, Noah A. Smith, Luke Zettlemoyer, Tao Yu 0009
EMNLP10
2021 Bilingual Lexicon Induction via Unsupervised Bitext Construction and Word Alignment
abstract
Haoyue Shi, Luke Zettlemoyer, Sida I. Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Freda Shi, Luke Zettlemoyer, Sida I. Wang
ACL/IJCNLP (1)3
2021 Towards Understanding the Behaviors of Optimal Deep Active Learning Algorithms
abstract
Active learning (AL) algorithms may achieve better performance with fewer data because the model guides the data selection process. While many algorithms have been proposed, there is little study on what the optimal AL algorithm looks like, which would help researchers understand where their models fall short and iterate on the design. In this paper, we present a simulated annealing algorithm to search for this optimal oracle and analyze it for several tasks. We present qualitative and quantitative insights into the behaviors of this oracle, comparing and contrasting them with those of various heuristics. Moreover, we are able to consistently improve the heuristics using one particular insight. We hope that our findings can better inform future active learning research. The code is available at https://github.com/YilunZhou/optimal-active-learning.
Yilun Zhou, Adithya Renduchintala, Xian Li 0003, Sida I. Wang, Yashar Mehdad, Asish Ghoshal
AISTATS4
2021 SILG: The Multi-domain Symbolic Interactive Language Grounding Benchmark
abstract
Existing work in language grounding typically study single environments. How do we build unified models that apply across multiple environments? We propose the multi-environment Symbolic Interactive Language Grounding benchmark (SILG), which unifies a collection of diverse grounded language learning environments under a common interface. SILG consists of grid-world environments that require generalization to new dynamics, entities, and partially observed worlds (RTFM, Messenger, NetHack), as well as symbolic counterparts of visual worlds that re- quire interpreting rich natural language with respect to complex scenes (ALFWorld, Touchdown). Together, these environments provide diverse grounding challenges in richness of observation space, action space, language specification, and plan com- plexity. In addition, we propose the first shared model architecture for RL on these environments, and evaluate recent advances such as egocentric local convolution, recurrent state-tracking, entity-centric attention, and pretrained LM using SILG. Our shared architecture achieves comparable performance to environment-specific architectures. Moreover, we find that many recent modelling advances do not result in significant gains on environments other than the one they were designed for. This highlights the need for a multi-environment benchmark. Finally, the best models significantly underperform humans on SILG, which suggests ample room for future work. We hope SILG enables the community to quickly identify new methodolo- gies for language grounding that generalize to a diverse set of environments and their associated challenges.
Victor Zhong, Austin W. Hanjie, Sida I. Wang, Karthik Narasimhan, Luke Zettlemoyer
NeurIPS3
2020 Interactive Classification by Asking Informative Questions
abstract
We study the potential for interaction in natural language classification.We add a limited form of interaction for intent classification, where users provide an initial query using natural language, and the system asks for additional information using binary or multichoice questions.At each turn, our system decides between asking the most informative question or making the final classification prediction.The simplicity of the model allows for bootstrapping of the system without interaction data, instead relying on simple crowdsourcing tasks.We evaluate our approach on two domains, showing the benefit of interaction and the advantage of learning to balance between asking additional questions and making the final prediction.What is the bill length of the bird: shorter, similar, or longer than head?Shorter than head.Is the bird underpart orange?Yes.The identified bird is: American Redstart FAQ Suggestion What data limits apply when roaming internationally?American Crow Bobolink … American Redstart How do I sign up for Sprint Global Roaming? . . .How do I purchase a High Speed Data Roaming Pass?Bird Identification Travel out of country.Do you need to activate global roaming service?Yes.Do you want high speed data roaming?No.
Lili Yu, Howard Chen 0003, Sida I. Wang, Tao Lei 0001, Yoav Artzi
ACL3
2020 Grounded Adaptation for Zero-shot Executable Semantic Parsing
abstract
We propose Grounded Adaptation for Zeroshot Executable Semantic Parsing (GAZP) to adapt an existing semantic parser to new environments (e.g.new database schemas).GAZP combines a forward semantic parser with a backward utterance generator to synthesize data (e.g.utterances and SQL queries) in the new environment, then selects cycleconsistent examples to adapt the parser.Unlike data-augmentation, which typically synthesizes unverified examples in the training environment, GAZP synthesizes examples in the new environment whose inputoutput consistency are verified.On the Spider, Sparc, and CoSQL zero-shot semantic parsing tasks, GAZP improves logical form and execution accuracy of the baseline parser.Our analyses show that GAZP outperforms dataaugmentation in the training environment, performance increases with the amount of GAZPsynthesized data, and cycle-consistency is central to successful adaptation.
Victor Zhong, Mike Lewis, Sida I. Wang, Luke Zettlemoyer
EMNLP (1)3
2020 Pre-training via Paraphrasing
abstract
We introduce MARGE, a pre-trained sequence-to-sequence model learned with an unsupervised multi-lingual multi-document paraphrasing objective. MARGE provides an alternative to the dominant masked language modeling paradigm, where we self-supervise the \emph{reconstruction} of target text by \emph{retrieving} a set of related texts (in many languages) and conditioning on them to maximize the likelihood of generating the original. We show it is possible to jointly learn to do retrieval and reconstruction, given only a random initialization. The objective noisily captures aspects of paraphrase, translation, multi-document summarization, and information retrieval, allowing for strong zero-shot performance on several tasks. For example, with no additional task-specific training we achieve BLEU scores of up to 35.8 for document translation. We further show that fine-tuning gives strong performance on a range of discriminative and generative tasks in many languages, making MARGE the most generally applicable pre-training method to date.
Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida I. Wang, Luke Zettlemoyer
NeurIPS5
2018 Simple Recurrent Units for Highly Parallelizable Recurrence
abstract
Common recurrent neural architectures scale poorly due to the intrinsic difficulty in parallelizing their state computations.In this work, we propose the Simple Recurrent Unit (SRU), a light recurrent unit that balances model capacity and scalability.SRU is designed to provide expressive recurrence, enable highly parallelized implementation, and comes with careful initialization to facilitate training of deep models.We demonstrate the effectiveness of SRU on multiple NLP tasks.SRU achieves 5-9x speed-up over cuDNN-optimized LSTM on classification and question answering datasets, and delivers stronger results than LSTM and convolutional models.We also obtain an average of 0.7 BLEU improvement over the Transformer model (Vaswani et al., 2017) on translation by incorporating SRU into the architecture.1
Tao Lei 0001, Yu Zhang 0033, Sida I. Wang, Hui Dai, Yoav Artzi
EMNLP3
2017 Naturalizing a Programming Language via Interactive Learning
abstract
Our goal is to create a convenient natural language interface for performing wellspecified but complex actions such as analyzing data, manipulating text, and querying databases.However, existing natural language interfaces for such tasks are quite primitive compared to the power one wields with a programming language.To bridge this gap, we start with a core programming language and allow users to "naturalize" the core language incrementally by defining alternative, more natural syntax and increasingly complex concepts in terms of compositions of simpler ones.In a voxel world, we show that a community of users can simultaneously teach a common system a diverse language and use it to build hundreds of complex voxel structures.Over the course of three days, these users went from using only the core language to using the naturalized language in 85.9% of the last 10K utterances.
Sida I. Wang, Samuel Ginn, Percy Liang, Christopher D. Manning
ACL (1)1
2017 Data Noising as Smoothing in Neural Network Language Models
Ziang Xie, Sida I. Wang, Jiwei Li 0001, Daniel Levy 0002, Aiming Nie, Daniel Jurafsky, Andrew Y. Ng
ICLR (Poster)2
2016 Learning Language Games through Interaction
abstract
We introduce a new language learning setting relevant to building adaptive natural language interfaces.It is inspired by Wittgenstein's language games: a human wishes to accomplish some task (e.g., achieving a certain configuration of blocks), but can only communicate with a computer, who performs the actual actions (e.g., removing all red blocks).The computer initially knows nothing about language and therefore must learn it from scratch through interaction, while the human adapts to the computer's capabilities.We created a game called SHRDLURN in a blocks world and collected interactions from 100 people playing it.First, we analyze the humans' strategies, showing that using compositionality and avoiding synonyms correlates positively with task performance.Second, we compare computer strategies, showing that modeling pragmatics on a semantic parsing model accelerates learning for more strategic players.
Sida I. Wang, Percy Liang, Christopher D. Manning
ACL (1)1
2015 Estimating Mixture Models via Mixtures of Polynomials
abstract
Mixture modeling is a general technique for making any simple model more expressive through weighted combination. This generality and simplicity in part explains the success of the Expectation Maximization (EM) algorithm, in which updates are easy to derive for a wide class of mixture models. However, the likelihood of a mixture model is non-convex, so EM has no known global convergence guarantees. Recently, method of moments approaches offer global guarantees for some mixture models, but they do not extend easily to the range of mixture models that exist. In this work, we present Polymom, an unifying framework based on method of moments in which estimation procedures are easily derivable, just as in EM. Polymom is applicable when the moments of a single mixture component are polynomials of the parameters. Our key observation is that the moments of the mixture model are a mixture of these polynomials, which allows us to cast estimation as a Generalized Moment Problem. We solve its relaxations using semidefinite optimization, and then extract parameters using ideas from computer algebra. This framework allows us to draw insights and apply tools from convex optimization, computer algebra and the theory of moments to study problems in statistical estimation. Simulations show good empirical performance on several models.
Sida I. Wang, Arun Tejasvi Chaganty, Percy Liang
NIPS1
2014 Human Effort and Machine Learnability in Computer Aided Translation
abstract
Spence Green, Sida I. Wang, Jason Chuang, Jeffrey Heer, Sebastian Schuster, Christopher D. Manning. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Spence Green, Sida I. Wang, Jason Chuang, Jeffrey Heer, Sebastian Schuster 0001, Christopher D. Manning
EMNLP2
2014 Simple MAP Inference via Low-Rank Relaxations
Roy Frostig, Sida I. Wang, Percy Liang, Christopher D. Manning
NIPS2
2014 Altitude Training: Strong Bounds for Single-Layer Dropout
Stefan Wager, William Fithian, Sida I. Wang, Percy Liang
NIPS3
2013 Fast and Adaptive Online Training of Feature-Rich Translation Models
Spence Green, Sida I. Wang, Daniel M. Cer, Christopher D. Manning
ACL (1)2
2013 Feature Noising for Log-Linear Structured Prediction
abstract
NLP models have many and sparse features, and regularization is key for balancing model overfitting versus underfitting.A recently repopularized form of regularization is to generate fake training data by repeatedly adding noise to real data.We reinterpret this noising as an explicit regularizer, and approximate it with a second-order formula that can be used during training without actually generating fake data.We show how to apply this method to structured prediction using multinomial logistic regression and linear-chain CRFs.We tackle the key challenge of developing a dynamic program to compute the gradient of the regularizer efficiently.The regularizer is a sum over inputs, so we can estimate it more accurately via a semi-supervised or transductive extension.Applied to text classification and NER, our method provides a >1% absolute performance gain over use of standard L 2 regularization.
Sida I. Wang, Mengqiu Wang, Stefan Wager, Percy Liang, Christopher D. Manning
EMNLP1
2013 Fast dropout training
abstract
Preventing feature co-adaptation by encouraging independent contributions from different features often improves classification and regression performance. Dropout training (Hinton et al., 2012) does this by randomly dropping out (zeroing) hidden units and input features during training of neural networks. However, repeatedly sampling a random subset of input features makes training much slower. Based on an examination of the implied objective function of dropout training, we show how to do fast dropout training by sampling from or integrating a Gaussian approximation, instead of doing Monte Carlo optimization of this objective. This approximation, justified by the central limit theorem and empirical evidence, gives an order of magnitude speedup and more stability. We show how to do fast dropout training for classification, regression, and multilayer neural networks. Beyond dropout, our technique is extended to integrate out other types of noise and small image transformations.
Sida I. Wang, Christopher D. Manning
ICML (2)1
2013 Dropout Training as Adaptive Regularization
abstract
Dropout and other feature noising schemes control overfitting by artificially corrupting the training data. For generalized linear models, dropout performs a form of adaptive regularization. Using this viewpoint, we show that the dropout regularizer is first-order equivalent to an $\LII$ regularizer applied after scaling the features by an estimate of the inverse diagonal Fisher information matrix. We also establish a connection to AdaGrad, an online learner, and find that a close relative of AdaGrad operates by repeatedly solving linear dropout-regularized problems. By casting dropout as regularization, we develop a natural semi-supervised algorithm that uses unlabeled data to create a better adaptive regularizer. We apply this idea to document classification tasks, and show that it consistently boosts the performance of dropout training, improving on state-of-the-art results on the IMDB reviews dataset.
Stefan Wager, Sida I. Wang, Percy Liang
NIPS2