VLDB 2026 Research / reviewers in the wild / expert
Kevin Yang
dblp:13/10565
· DBLP profile ↗
20ranked-venue papers
9as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FactTrack: Time-Aware World State Tracking in Story OutlinesabstractZhiheng Lyu, Kevin Yang, Lingpeng Kong, Dan Klein. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zhiheng Lyu, Kevin Yang, Lingpeng Kong |
NAACL (Long Papers) | 2 |
| 2024 | Learning Personalized Alignment for Evaluating Open-ended Text GenerationabstractRecent research has increasingly focused on evaluating large language models' (LLMs) alignment with diverse human values and preferences, particularly for open-ended tasks like story generation.Traditional evaluation metrics rely heavily on lexical similarity with humanwritten references, often showing poor correlation with human judgments and failing to account for alignment with the diversity of human preferences.To address these challenges, we introduce PERSE, an interpretable evaluation framework designed to assess alignment with specific human preferences.It is tuned to infer specific preferences from an in-context personal profile and evaluate the alignment between the generated content and personal preferences.PERSE enhances interpretability by providing detailed comments and fine-grained scoring, facilitating more personalized content generation.Our 13B LLaMA-2-based PERSE shows a 15.8% increase in Kendall correlation and a 13.7% rise in accuracy with zero-shot reviewers compared to GPT-4.It also outperforms GPT-4 by 46.01% in Kendall correlation on new domains, indicating its transferability 1 . Danqing Wang, Kevin Yang, Hanlin Zhu, Andrew Cohen, Lei Li 0005, Yuandong Tian |
EMNLP | 2 |
| 2024 | RLCD: Reinforcement Learning from Contrastive Distillation for LM AlignmentabstractWe propose Reinforcement Learning from Contrastive Distillation (RLCD), a method for aligning language models to follow principles expressed in natural language (e.g., to be more harmless) without using human feedback. RLCD creates preference pairs from two contrasting model outputs, one using a positive prompt designed to encourage following the given principles, and one using a negative prompt designed to encourage violating them. Using two different prompts causes model outputs to be more differentiated on average, resulting in cleaner preference labels in the absence of human annotations. We then use the preference pairs to train a preference model, which is in turn used to improve a base unaligned language model via reinforcement learning. Empirically, RLCD outperforms RLAIF (Bai et al., 2022b) and context distillation (Huang et al., 2022) baselines across three diverse alignment tasks—harmlessness, helpfulness, and story outline generation—and when using both 7B and 30B model scales for simulating preference data Kevin Yang, Daniel Klein 0001, Asli Celikyilmaz, Nanyun Peng 0001, Yuandong Tian |
ICLR | 1 |
| 2023 | DOC: Improving Long Story Coherence With Detailed Outline ControlabstractWe propose the Detailed Outline Control (DOC) framework for improving long-range plot coherence when automatically generating several-thousand-word-long stories.DOC consists of two complementary components: a detailed outliner and a detailed controller.The detailed outliner creates a more detailed, hierarchically structured outline, shifting creative burden from the main drafting procedure to the planning stage.The detailed controller ensures the more detailed outline is still respected during generation by controlling story passages to align with outline details.In human evaluations of automatically generated stories, DOC substantially outperforms a strong Re 3 baseline (Yang et al., 2022) on plot coherence (22.5% absolute gain), outline relevance (28.2%), and interestingness (20.7%).Humans also judged DOC to be much more controllable in an interactive generation setting. Daisy is a kind-hearted old woman.She has cancer.Bill is her husband.2. Lisa is Daisy's daughter. Structured Prompt For DraftingDaisy is diagnosed with cancer.Lisa is trying to find a viable treatment.Lisa has been stressed out lately, and Daisy expresses her concern.Lisa tirelessly continues her research.Lisa finally finds a cure.Setting: Lisa's laboratory.Lisa looked back at Daisy, her eyes clear and full of determination. Kevin Yang, Daniel Klein 0001, Nanyun Peng 0001, Yuandong Tian |
ACL (1) | 1 |
| 2022 | Automated Crossword SolvingabstractEric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang, Eshaan Pathak, Matthew Ginsberg, Dan Klein. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Eric Wallace, Nicholas Tomlin, Albert Xu, Kevin Yang, Eshaan Pathak, Matthew L. Ginsberg, Daniel Klein 0001 |
ACL (1) | 4 |
| 2022 | Clinical Informatics Educational Requirements in United States ACGME-Accredited Residency Programs
Kevin Yang, Vinod E. Nambudiri |
AMIA | 1 |
| 2022 | Re3: Generating Longer Stories With Recursive Reprompting and RevisionabstractWe consider the problem of automatically generating longer stories of over two thousand words.Compared to prior work on shorter stories, long-range plot coherence and relevance are more central challenges here.We propose the Recursive Reprompting and Revision framework (Re 3 ) to address these challenges by (a) prompting a general-purpose language model to construct a structured overarching plan, and (b) generating story passages by repeatedly injecting contextual information from both the plan and current story state into a language model prompt.We then revise by (c) reranking different continuations for plot coherence and premise relevance, and finally (d) editing the best continuation for factual consistency.Compared to similar-length stories generated directly from the same base model, human evaluators judged substantially more of Re 3 's stories as having a coherent overarching plot (by 14% absolute increase), and relevant to the given initial premise (by 20%). Autoregressive Context EditPeyton Turner Peyton Turner is male.Peyton works at a restaurant. Inferred FactsShe knew Peyton was probably Kevin Yang, Yuandong Tian, Nanyun Peng 0001, Daniel Klein 0001 |
EMNLP | 1 |
| 2022 | Multi-objective Optimization by Learning Space Partition
Linnan Wang, Kevin Yang, Tianjun Zhang, Tian Guo 0001, Yuandong Tian |
ICLR | 3 |
| 2022 | Exploring evolution-aware & -free protein language models as protein function predictorsabstractLarge-scale Protein Language Models (PLMs) have improved performance in protein prediction tasks, ranging from 3D structure prediction to various function predictions. In particular, AlphaFold, a ground-breaking AI system, could potentially reshape structural biology. However, the utility of the PLM module in AlphaFold, Evoformer, has not been explored beyond structure prediction. In this paper, we investigate the representation ability of three popular PLMs: ESM-1b (single sequence), MSA-Transformer (multiple sequence alignment), and Evoformer (structural), with a special focus on Evoformer. Specifically, we aim to answer the following key questions: (1) Does the Evoformer trained as part of AlphaFold produce representations amenable to predicting protein function? (2) If yes, can Evoformer replace ESM-1b and MSA-Transformer? (3) How much do these PLMs rely on evolution-related protein data? In this regard, are they complementary to each other? We compare these models by empirical study along with new insights and conclusions. All code and datasets for reproducibility are available at https://github.com/elttaes/Revisiting-PLMs . Mingyang Hu, Fajie Yuan, Kevin Yang, Fusong Ju, Jin Su, Qiuyang Ding |
NeurIPS | 3 |
| 2021 | FUDGE: Controlled Text Generation With Future DiscriminatorsabstractWe propose Future Discriminators for Generation (FUDGE), a flexible and modular method for controlled text generation.Given a preexisting model G for generating text from a distribution of interest, FUDGE enables conditioning on a desired attribute a (for example, formality) while requiring access only to G's output logits.FUDGE learns an attribute predictor operating on a partial sequence, and uses this predictor's outputs to adjust G's original probabilities.We show that FUDGE models terms corresponding to a Bayesian decomposition of the conditional distribution of G given attribute a.Moreover, FUDGE can easily compose predictors for multiple desired attributes.We evaluate FUDGE on three tasks -couplet completion in poetry, topic control in language generation, and formality change in machine translation -and observe gains in all three tasks. Kevin Yang, Daniel Klein 0001 |
NAACL-HLT | 1 |
| 2021 | Learning Space Partitions for Path PlanningabstractPath planning, the problem of efficiently discovering high-reward trajectories, often requires optimizing a high-dimensional and multimodal reward function. Popular approaches like CEM and CMA-ES greedily focus on promising regions of the search space and may get trapped in local maxima. DOO and VOOT balance exploration and exploitation, but use space partitioning strategies independent of the reward function to be optimized. Recently, LaMCTS empirically learns to partition the search space in a reward-sensitive manner for black-box optimization. In this paper, we develop a novel formal regret analysis for when and why such an adaptive region partitioning scheme works. We also propose a new path planning method LaP3 which improves the function value estimation within each sub-region, and uses a latent representation of the search space. Empirically, LaP3 outperforms existing path planning methods in 2D navigation tasks, especially in the presence of difficult-to-escape local optima, and shows benefits when plugged into the planning components of model-based RL such as PETS. These gains transfer to highly multimodal real-world tasks, where we outperform strong baselines in compiler phase ordering by up to 39% on average across 9 tasks, and in molecular design by up to 0.4 on properties on a 0-1 scale. Code is available at https://github.com/yangkevin2/neurips2021-lap3. Kevin Yang, Tianjun Zhang, Chris Cummins, Brandon Cui, Benoit Steiner, Linnan Wang, Joseph Gonzalez 0001, Daniel Klein 0001, Yuandong Tian |
NeurIPS | 1 |
| 2021 | Towards Benchmarking Feature Type Inference for AutoML PlatformsabstractThe paradigm of AutoML has created an opportunity to enable ML for the masses. Emerging industrial-scale cloud AutoML platforms aim to automate the end-to-end ML workflow. While many works have looked into automated feature engineering, model selection, or hyper-parameter search in AutoML, little work has studied a crucial step that serves as an entry point to this workflow: ML feature type inference. The semantic gap between attribute types (e.g., strings, numbers) in databases/files and ML feature types (e.g., Numeric, Categorical) necessitates type inference. In this work, we formalize and standardize this task by creating the first ever benchmark labeled dataset, which we use to objectively evaluate existing AutoML tools. Our dataset has 9921 examples and a 9-class label vocabulary. Our labeled data also offers an alternative approach to automate this task than existing rule-based or syntax-based approaches: use ML itself to predict feature types. We collate a benchmark suite of 30 classification and regression tasks to assess the importance of type inference for downstream models. Empirical comparison on our labeled data shows that an ML-based approach delivers a lift of an average 14% and up to 38% in accuracy for identifying feature types compared to prominent industrial tools. Our downstream benchmark suite reveals that the ML-based approach outperforms existing industrial-strength tools for 47 out of 60 downstream models. We release our labeled dataset, models, and downstream benchmarks in a public repository with a leaderboard. Vraj Shah, Jonathan Lacanlale, Premanand Kumar, Kevin Yang, Arun Kumar 0001 |
SIGMOD Conference | 4 |
| 2021 | Ultrasound image analysis technology under deep belief networks in evaluation on the effects of diagnosis and chemotherapy of cervical cancer
Hongzhen Zhou, Shuyuan Wang, Demei Liu, Kevin Yang |
J. Supercomput. | 5 |
| 2020 | A Streaming Approach For Efficient Batched Beam SearchabstractWe propose an efficient batching strategy for variable-length decoding on GPU architectures.During decoding, when candidates terminate or are pruned according to heuristics, our streaming approach periodically "refills" the batch before proceeding with a selected subset of candidates.We apply our method to variable-width beam search on a state-of-theart machine translation model.Our method decreases runtime by up to 71% compared to a fixed-width beam search baseline and 17% compared to a variable-width baseline, while matching baselines' BLEU.Finally, experiments show that our method can speed up decoding in other domains, such as semantic and syntactic parsing. Kevin Yang, Violet Yao, John DeNero, Daniel Klein 0001 |
EMNLP (1) | 1 |
| 2020 | Improving Molecular Design by Stochastic Iterative Target AugmentationabstractGenerative models in molecular design tend to be richly parameterized, data-hungry neural models, as they must create complex structured objects as outputs. Estimating such models from data may be challenging due to the lack of sufficient training data. In this paper, we propose a surprisingly effective self-training approach for iteratively creating additional molecular targets. We first pre-train the generative model together with a simple property predictor. The property predictor is then used as a likelihood model for filtering candidate structures from the generative model. Additional targets are iteratively produced and used in the course of stochastic EM iterations to maximize the log-likelihood that the candidate structures are accepted. A simple rejection (re-weighting) sampler suffices to draw posterior samples since the generative model is already reasonable after pre-training. We demonstrate significant gains over strong baselines for both unconditional and conditional molecular design. In particular, our approach outperforms the previous state-of-the-art in conditional molecular design by over 10% in absolute gain. Finally, we show that our approach is useful in other domains as well, such as program synthesis. Kevin Yang, Wengong Jin, Kyle Swanson, Regina Barzilay, Tommi S. Jaakkola |
ICML | 1 |
| 2019 | Learning Multimodal Graph-to-Graph Translation for Molecule Optimization
Wengong Jin, Kevin Yang, Regina Barzilay, Tommi S. Jaakkola |
ICLR (Poster) | 2 |
| 2019 | Demonstration of SpeakQL: Speech-driven Multimodal Querying of Structured DataabstractIn this demonstration, we present SpeakQL, a speech-driven query system and interface for structured data. SpeakQL supports a tractable and practically useful subset of regular SQL, allowing users to query in any domain with unbounded vocabulary with the help of speech/touch based user-in-the-loop mechanisms for correction. When querying in such domains, automatic speech recognition introduces countless forms of errors in transcriptions, presenting us with a technical challenge. We characterize such errors and leverage our observations along with SQL's unambiguous context-free grammar to first correct the query structure. We then exploit phonetic representation of the queried database to identify the correct Literals, hence delivering the corrected transcribed query. In this demo, we show that SpeakQL helps users reduce time and effort in specifying SQL queries significantly. In addition, we show that SpeakQL, unlike Natural Language Interfaces and conversational assistants, allows users to query over any arbitrary database schema. We allow the audience to explore SpeakQL using an easy-to-use web-based interface to compose SQL queries. Vraj Shah, Side Li, Kevin Yang, Arun Kumar 0001, Lawrence K. Saul |
SIGMOD Conference | 3 |
| 2009 | Routing in SONET/VCAT based optical WDM networks (Invited Paper)abstractIn this paper, we investigate problems related to optical wavelength division multiplexing (WDM) networks that use the virtual concatenation (VCAT) mechanism of synchronous optical network (SONET) technology. VCAT, an end-to-end mechanism, allows SONET based optical WDM networks to carry traffic in Kevin Yang, Krishna M. Sivalingam |
BROADNETS | 1 |
| 2008 | A Regional Health Information Exchange: Architecture and Implementation
Mark E. Frisse, Janet K. King, Will B. Rice, Lianhong Tang, Jameson P. Porter, Timothy A. Coffman, Michael Assink, Kevin Yang, Monroe Wesley, Rodney L. Holmes, Cynthia S. Gadd, Kevin B. Johnson, Vicki Y. Estrin |
AMIA | 8 |
| 2008 | The MidSouth eHealth Alliance: Use and Impact in the First Year
Kevin B. Johnson, Cynthia S. Gadd, Dominik Aronsky, Kevin Yang, Lianhong Tang, Vicki Y. Estrin, Janet K. King, Mark E. Frisse |
AMIA | 4 |