VLDB 2026 Research / reviewers in the wild / expert
Richard Shin
dblp:13/8735 · also Eui Chul Richard Shin
· DBLP profile ↗
22ranked-venue papers
6as first author
7since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 7 since 2021Security and privacy · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Information extraction and text analysis · 45% Language models and text generation · 37% Reinforcement learning · 14% | |
| Software engineering, system software, and programming languages
8 papers |
Program synthesis and code generation · 84% Programming languages and type systems · 9% Program analysis · 7% | |
| Network and information security
5 papers |
Privacy and data protection · 79% Network security · 8% Web and mobile security · 5% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 100% |
Topics — the 28 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
semantic parsing |
2.3 | 4 | 2023 | BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing · NeurIPS 2023 Privacy-Preserving Domain Adaptation of Semantic Parsers · ACL (1) 2023 Constrained Language Models Yield Few-Shot Semantic Parsers · EMNLP (1) 2021 |
Program synthesis and code generation
neural program synthesis |
1.7 | 5 | 2019 | Program Synthesis and Semantic Parsing with Learned Code Idioms · NeurIPS 2019 Synthetic Datasets for Neural Program Synthesis · ICLR (Poster) 2019 Improving Neural Program Synthesis with Inferred Execution Traces · NeurIPS 2018 |
Natural language and speech › Language models and text generation
in-context learning |
0.8 | 1 | 2024 | Privacy-Preserving In-Context Learning with Differentially Private Few-Shot Generation · ICLR 2024 |
Machine learning › Reinforcement learning
policy optimization |
0.8 | 1 | 2024 | Learning to Retrieve Iteratively for In-Context Learning · EMNLP 2024 |
Natural language and speech › Language models and text generation › in-context learning
privacy-preserving in-context learning |
0.8 | 1 | 2024 | Privacy-Preserving In-Context Learning with Differentially Private Few-Shot Generation · ICLR 2024 |
Information retrieval › retrieval-augmented generation
iterative retrieval |
0.8 | 1 | 2024 | Learning to Retrieve Iteratively for In-Context Learning · EMNLP 2024 |
Privacy and data protection
differential privacy |
0.8 | 1 | 2024 | Privacy-Preserving In-Context Learning with Differentially Private Few-Shot Generation · ICLR 2024 |
Privacy and data protection › differential privacy
synthetic data generation |
0.8 | 1 | 2024 | Privacy-Preserving In-Context Learning with Differentially Private Few-Shot Generation · ICLR 2024 |
Natural language and speech › Language models and text generation › decoding
constrained decoding |
0.7 | 1 | 2023 | BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing · NeurIPS 2023 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.7 | 1 | 2023 | BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing · NeurIPS 2023 |
Privacy and data protection
privacy-preserving machine learning |
0.7 | 1 | 2023 | Privacy-Preserving Domain Adaptation of Semantic Parsers · ACL (1) 2023 |
Natural language and speech › Information extraction and text analysis › semantic parsing
schema linking |
0.4 | 1 | 2020 | RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers · ACL 2020 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
text-to-SQL parsing |
0.4 | 1 | 2020 | RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers · ACL 2020 |
Machine learning › Reinforcement learning
hierarchical reinforcement learning |
0.3 | 1 | 2018 | Parametrized Hierarchical Procedures for Neural Programming · ICLR (Poster) 2018 |
Programming languages and type systems › control structures
recursion |
0.3 | 1 | 2017 | Making Neural Programming Architectures Generalize via Recursion · ICLR 2017 |
Machine learning › Representation and self-supervised learning
automated feature generation |
0.2 | 1 | 2016 | ExploreKit: Automatic Feature Generation and Selection · ICDM 2016 |
Natural language and speech › Information extraction and text analysis
feature engineering |
0.2 | 1 | 2016 | ExploreKit: Automatic Feature Generation and Selection · ICDM 2016 |
Natural language and speech › Language models and text generation
code generation |
0.2 | 1 | 2024 | Language-to-Code Translation with a Single Labeled Example · EMNLP 2024 |
Program analysis
binary analysis |
0.2 | 1 | 2015 | Recognizing Functions in Binaries with Neural Networks · USENIX Security Symposium 2015 |
Privacy and data protection
anonymity |
0.1 | 1 | 2012 | On the Feasibility of Internet-Scale Author Identification · IEEE Symposium on Security and Privacy 2012 |
Network security
anonymity networks |
0.1 | 1 | 2012 | On the Feasibility of Internet-Scale Author Identification · IEEE Symposium on Security and Privacy 2012 |
Privacy and data protection
de-anonymization |
0.1 | 1 | 2012 | On the Feasibility of Internet-Scale Author Identification · IEEE Symposium on Security and Privacy 2012 |
Web and mobile security
mobile security |
0.1 | 1 | 2012 | FreeMarket: Shopping for free in Android applications · NDSS 2012 |
Digital forensics and information hiding › authorship attribution
stylometry |
0.1 | 1 | 2012 | On the Feasibility of Internet-Scale Author Identification · IEEE Symposium on Security and Privacy 2012 |
Malware analysis
botnet |
0.1 | 1 | 2010 | Inference and analysis of formal models of botnet command and control protocols · CCS 2010 |
Network security › protocol reverse engineering
protocol state machine inference |
0.1 | 1 | 2010 | Inference and analysis of formal models of botnet command and control protocols · CCS 2010 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection |
0.1 | 1 | 2016 | ExploreKit: Automatic Feature Generation and Selection · ICDM 2016 |
Information retrieval
text analysis |
0.0 | 1 | 2012 | On the Feasibility of Internet-Scale Author Identification · IEEE Symposium on Security and Privacy 2012 |
Methods — techniques the papers use, named apart from their topics
differential privacy · 2.8reinforcement learning · 1.5policy optimization · 1.5few-shot learning · 1.5few-shot generation · 1.5dense retrieval · 1.5domain adaptation · 1.3prompt-based learning · 0.7fine-tuning · 0.7constrained decoding · 0.7tree-based neural synthesis · 0.4idiom mining · 0.4trace inference · 0.3parametrized hierarchical procedures · 0.3neural programming · 0.3deep learning · 0.3stylometric classifiers · 0.3large-scale classification · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Language-to-Code Translation with a Single Labeled ExampleabstractKaj Bostrom, Harsh Jhamtani, Hao Fang, Sam Thomson, Richard Shin, Patrick Xia, Benjamin Van Durme, Jason Eisner, Jacob Andreas. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Kaj Bostrom, Harsh Jhamtani, Hao Fang 0002, Sam Thomson, Richard Shin, Patrick Xia 0002, Benjamin Van Durme, Jason Eisner, Jacob Andreas |
EMNLP | 5 |
| 2024 | Learning to Retrieve Iteratively for In-Context LearningabstractWe introduce iterative retrieval, a novel framework that empowers retrievers to make iterative decisions through policy optimization.Finding an optimal portfolio of retrieved items is a combinatorial optimization problem, generally considered NP-hard.This approach provides a learned approximation to such a solution, meeting specific task requirements under a given family of large language models (LLMs).We propose a training procedure based on reinforcement learning, incorporating feedback from LLMs.We instantiate an iterative retriever for composing in-context learning (ICL) exemplars and apply it to various semantic parsing tasks that demand synthesized programs as outputs.By adding only 4M additional parameters for state encoding, we convert an offthe-shelf dense retriever into a stateful iterative retriever, outperforming previous methods in selecting ICL exemplars on semantic parsing datasets such as SMCALFLOW, TREEDST, and MTOP.Additionally, the trained iterative retriever generalizes across different inference LLMs beyond the one used during training. Yunmo Chen, Tongfei Chen, Harsh Jhamtani, Patrick Xia 0002, Richard Shin, Jason Eisner, Benjamin Van Durme |
EMNLP | 5 |
| 2024 | Privacy-Preserving In-Context Learning with Differentially Private Few-Shot GenerationabstractWe study the problem of in-context learning (ICL) with large language models (LLMs) on private datasets.
This scenario poses privacy risks, as LLMs may leak or regurgitate the private examples demonstrated in the prompt.
We propose a novel algorithm that generates synthetic few-shot demonstrations from the private dataset with formal differential privacy (DP) guarantees, and show empirically that it can achieve effective ICL.
We conduct extensive experiments on standard benchmarks and compare our algorithm with non-private ICL and zero-shot solutions.
Our results demonstrate that our algorithm can achieve competitive performance with strong privacy levels.
These results open up new possibilities for ICL with privacy protection for a broad range of applications. Xinyu Tang 0003, Richard Shin, Huseyin A. Inan, Andre Manoel, Niloofar Mireshghallah, Zinan Lin 0001, Sivakanth Gopi, Janardhan Kulkarni, Robert Sim |
ICLR | 2 |
| 2023 | Privacy-Preserving Domain Adaptation of Semantic ParsersabstractFatemehsadat Mireshghallah, Yu Su, Tatsunori Hashimoto, Jason Eisner, Richard Shin. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Niloofar Mireshghallah, Yu Su 0001, Tatsunori B. Hashimoto, Jason Eisner, Richard Shin |
ACL (1) | 5 |
| 2023 | BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic ParsingabstractRecent work has shown that generation from a prompted or fine-tuned language model can perform well at semantic parsing when the output is constrained to be a valid semantic representation. We introduce BenchCLAMP, a Benchmark to evaluate Constrained LAnguage Model Parsing, that includes context-free grammars for seven semantic parsing datasets and two syntactic parsing datasets with varied output meaning representations, as well as a constrained decoding interface to generate only valid outputs covered by these grammars. We provide low, medium, and high resource splits for each dataset, allowing accurate comparison of various language models under different data regimes. Our benchmark supports evaluation of language models using prompt-based learning as well as fine-tuning. We benchmark seven language models, including two GPT-3 variants available only through an API. Our experiments show that encoder-decoder pretrained language models can achieve similar performance or even surpass state-of-the-art methods for both syntactic and semantic parsing when the model output is constrained to be valid. Subhro Roy, Sam Thomson, Tongfei Chen, Richard Shin, Adam Pauls, Jason Eisner, Benjamin Van Durme |
NeurIPS | 4 |
| 2022 | Few-Shot Semantic Parsing with Language Models Trained on CodeabstractLarge language models can perform semantic parsing with little training data, when prompted with in-context examples.It has been shown that this can be improved by formulating the problem as paraphrasing into canonical utterances, which casts the underlying meaning representation into a controlled natural languagelike representation.Intuitively, such models can more easily output canonical utterances as they are closer to the natural language used for pre-training.Recently, models also pretrained on code, like OpenAI Codex, have risen in prominence.For semantic parsing tasks where we map natural language into code, such models may prove more adept at it.In this paper, we test this hypothesis and find that Codex performs better on such tasks than equivalent GPT-3 models.We evaluate on Overnight and SMCalFlow and find that unlike GPT-3, Codex performs similarly when targeting meaning representations directly, perhaps because meaning representations are structured similar to code in these datasets. Richard Shin, Benjamin Van Durme |
NAACL-HLT | 1 |
| 2021 | Constrained Language Models Yield Few-Shot Semantic ParsersabstractRichard Shin, Christopher Lin, Sam Thomson, Charles Chen, Subhro Roy, Emmanouil Antonios Platanios, Adam Pauls, Dan Klein, Jason Eisner, Benjamin Van Durme. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Richard Shin, Christopher H. Lin, Sam Thomson, Subhro Roy, Emmanouil A. Platanios, Adam Pauls, Daniel Klein 0001, Jason Eisner, Benjamin Van Durme |
EMNLP (1) | 1 |
| 2020 | RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL ParsersabstractWhen translating natural language questions into SQL queries to answer questions from a database, contemporary semantic parsing models struggle to generalize to unseen database schemas.The generalization challenge lies in (a) encoding the database relations in an accessible way for the semantic parser, and (b) modeling alignment between database columns and their mentions in a given query.We present a unified framework, based on the relation-aware self-attention mechanism, to address schema encoding, schema linking, and feature representation within a text-to-SQL encoder.On the challenging Spider dataset this framework boosts the exact match accuracy to 57.2%, surpassing its best counterparts by 8.7% absolute improvement.Further augmented with BERT, it achieves the new state-of-the-art performance of 65.6% on the Spider leaderboard.In addition, we observe qualitative improvements in the model's understanding of schema linking and alignment.Our implementation will be open-sourced at https://github.com/Microsoft/rat-sql. Bailin Wang, Richard Shin, Xiaodong Liu 0003, Oleksandr Polozov, Matthew Richardson |
ACL | 2 |
| 2019 | Synthetic Datasets for Neural Program Synthesis
Richard Shin, Neel Kant, Kavi Gupta, Chris Bender, Brandon Trabucco, Rishabh Singh, Dawn Song |
ICLR (Poster) | 1 |
| 2019 | Program Synthesis and Semantic Parsing with Learned Code IdiomsabstractProgram synthesis of general-purpose source code from natural language specifications is challenging due to the need to reason about high-level patterns in the target program and low-level implementation details at the same time. In this work, we present Patois, a system that allows a neural program synthesizer to explicitly interleave high-level and low-level reasoning at every generation step. It accomplishes this by automatically mining common code idioms from a given corpus, incorporating them into the underlying language for neural synthesis, and training a tree-based neural synthesizer to use these idioms during code generation. We evaluate Patois on two complex semantic parsing datasets and show that using learned code idioms improves the synthesizer's accuracy. Richard Shin, Miltiadis Allamanis, Marc Brockschmidt, Oleksandr Polozov |
NeurIPS | 1 |
| 2018 | Parametrized Hierarchical Procedures for Neural Programming
Roy Fox, Richard Shin, Sanjay Krishnan, Kenneth Y. Goldberg, Dawn Song, Ion Stoica |
ICLR (Poster) | 2 |
| 2018 | Improving Neural Program Synthesis with Inferred Execution TracesabstractThe task of program synthesis, or automatically generating programs that are consistent with a provided specification, remains a challenging task in artificial intelligence. As in other fields of AI, deep learning-based end-to-end approaches have made great advances in program synthesis. However, more so than other fields such as computer vision, program synthesis provides greater opportunities to explicitly exploit structured information such as execution traces, which contain a superset of the information input/output pairs. While they are highly useful for program synthesis, as execution traces are more difficult to obtain than input/output pairs, we use the insight that we can split the process into two parts: infer the trace from the input/output example, then infer the program from the trace. This simple modification leads to state-of-the-art results in program synthesis in the Karel domain, improving accuracy to 81.3% from the 77.12% of prior work. Richard Shin, Illia Polosukhin, Dawn Song |
NeurIPS | 1 |
| 2017 | PIANO: Proximity-Based User Authentication on Voice-Powered Internet-of-Things DevicesabstractVoice is envisioned to be a popular way for humans to interact with Internet-of-Things (IoT) devices. We propose a proximity-based user authentication method (called PIANO) for access control on such voice-powered IoT devices. PIANO leverages the built-in speaker, microphone, and Bluetooth that voice-powered IoT devices often already have. Specifically, we assume that a user carries a personal voice-powered device (e.g., smartphone, smartwatch, or smartglass), which serves as the user's identity. When another voice-powered IoT device of the user requires authentication, PIANO estimates the distance between the two devices by playing and detecting certain acoustic signals; PIANO grants access if the estimated distance is no larger than a user-selected threshold. We implemented a proof-of-concept prototype of PIANO. Through theoretical and empirical evaluations, we find that PIANO is secure, reliable, personalizable, and efficient. Neil Zhenqiang Gong, Altay Ozen, Richard Shin, Dawn Song, Hongxia Jin, Xuan Bao |
ICDCS | 5 |
| 2017 | Making Neural Programming Architectures Generalize via Recursion
Jonathon Cai, Richard Shin, Dawn Song |
ICLR | 2 |
| 2016 | ExploreKit: Automatic Feature Generation and SelectionabstractFeature generation is one of the challenging aspects of machine learning. We present ExploreKit, a framework for automated feature generation. ExploreKit generates a large set of candidate features by combining information in the original features, with the aim of maximizing predictive performance according to user-selected criteria. To overcome the exponential growth of the feature space, ExploreKit uses a novel machine learning-based feature selection approach to predict the usefulness of new candidate features. This approach enables efficient identification of the new features and produces superior results compared to existing feature selection solutions. We demonstrate the effectiveness and robustness of our approach by conducting an extensive evaluation on 25 datasets and 3 different classification algorithms. We show that ExploreKit can achieve classification-error reduction of 20% overall. Our codeis available at https://github.com/giladkatz/ExploreKit. Gilad Katz, Richard Shin, Dawn Song |
ICDM | 2 |
| 2016 | Latent Attention For If-Then Program SynthesisabstractAutomatic translation from natural language descriptions into programs is a long-standing challenging problem. In this work, we consider a simple yet important sub-problem: translation from textual descriptions to If-Then programs. We devise a novel neural network architecture for this task which we train end-to-end. Specifically, we introduce Latent Attention, which computes multiplicative weights for the words in the description in a two-stage process with the goal of better leveraging the natural language structures that indicate the relevant parts for predicting program elements. Our architecture reduces the error rate by 28.57% compared to prior art. We also propose a one-shot learning scenario of If-Then program synthesis and simulate it with our existing dataset. We demonstrate a variation on the training procedure for this scenario that outperforms the original procedure, significantly closing the gap to the model trained with all data. Chang Liu 0021, Richard Shin, Mingcheng Chen, Dawn Song |
NIPS | 3 |
| 2015 | Recognizing Functions in Binaries with Neural Networks
Richard Shin, Dawn Song, Reza Moazzezi |
USENIX Security Symposium | 1 |
| 2014 | Joint Link Prediction and Attribute Inference Using a Social-Attribute NetworkabstractThe effects of social influence and homophily suggest that both network structure and node-attribute information should inform the tasks of link prediction and node-attribute inference. Recently, Yin et al. [2010a, 2010b] proposed an attribute-augmented social network model, which we callSocial-Attribute Network(SAN), to integrate network structure and node attributes to perform both link prediction and attribute inference. They focused on generalizing the random walk with a restart algorithm to the SAN framework and showed improved performance. In this article, we extend the SAN framework with several leading supervised and unsupervised link-prediction algorithms and demonstrate performance improvement for each algorithm on both link prediction and attribute inference. Moreover, we make the novel observation that attribute inference can help inform link prediction, that is, link-prediction accuracy is further improved by first inferring missing attributes. We comprehensively evaluate these algorithms and compare them with other existing algorithms using a novel, large-scale Google+ dataset, which we make publicly available (http://www.cs.berkeley.edu/~stevgong/gplus.html). Neil Zhenqiang Gong, Ameet Talwalkar, Lester Mackey, Ling Huang 0001, Richard Shin, Emil Stefanov, Elaine Shi, Dawn Song |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2012 | FreeMarket: Shopping for free in Android applications
Daniel Reynaud, Dawn Song, Thomas R. Magrino, Edward XueJun Wu, Richard Shin |
NDSS | 5 |
| 2012 | On the Feasibility of Internet-Scale Author IdentificationabstractWe study techniques for identifying an anonymous author via linguistic stylometry, i.e., comparing the writing style against a corpus of texts of known authorship. We experimentally demonstrate the effectiveness of our techniques with as many as 100,000 candidate authors. Given the increasing availability of writing samples online, our result has serious implications for anonymity and free speech -- an anonymous blogger or whistleblower may be unmasked unless they take steps to obfuscate their writing style. While there is a huge body of literature on authorship recognition based on writing style, almost none of it has studied corpora of more than a few hundred authors. The problem becomes qualitatively different at a large scale, as we show, and techniques from prior work fail to scale, both in terms of accuracy and performance. We study a variety of classifiers, both "lazy" and "eager," and show how to handle the huge number of classes. We also develop novel techniques for confidence estimation of classifier outputs. Finally, we demonstrate stylometric authorship recognition on texts written in different contexts. In over 20% of cases, our classifiers can correctly identify an anonymous author given a corpus of texts from 100,000 authors; in about 35% of cases the correct author is one of the top 20 guesses. If we allow the classifier the option of not making a guess, via confidence estimation we are able to increase the precision of the top guess from 20% to over 80% with only a halving of recall. Arvind Narayanan, Hristo S. Paskov, Neil Zhenqiang Gong, John Bethencourt, Emil Stefanov, Richard Shin, Dawn Song |
IEEE Symposium on Security and Privacy | 6 |
| 2011 | A Systematic Analysis of XSS Sanitization in Web Application Frameworks
Joel Weinberger, Prateek Saxena, Devdatta Akhawe, Matthew Finifter, Richard Shin, Dawn Song |
ESORICS | 5 |
| 2010 | Inference and analysis of formal models of botnet command and control protocolsabstractWe propose a novel approach to infer protocol state machines in the realistic high-latency network setting, and apply it to the analysis of botnet Command and Control (C &C) protocols. Our proposed techniques enable an order of magnitude reduction in the number of queries and time needed to learn a botnet C &C protocol compared to classic algorithms (from days to hours for inferring the MegaD C &C protocol). We also show that the computed protocol state machines enable formal analysis for botnet defense, including finding the weakest links in a protocol, uncovering protocol design flaws, inferring the existence of unobservable communication back-channels among botnet servers, and finding deviations of protocol implementations which can be used for fingerprinting. We validate our technique by inferring the protocol state-machine from Postfix's SMTP implementation and comparing the inferred state-machine to the SMTP standard. Further, our experimental results offer new insights into MegaD's C &C, showing our technique can be used as a powerful tool for defense against botnets. Chia Yuan Cho, Domagoj Babic, Richard Shin, Dawn Song |
CCS | 3 |