EDBT 2026 Demo / reviewers in the wild / expert
He He 0001
dblp:08/8618-1
· DBLP profile ↗
57ranked-venue papers
10as first author
32since 2021 · last 2025
0009-0005-7805-3780ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 10 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Consistent Natural-Language Explanations via Explanation-Consistency FinetuningabstractLarge language models (LLMs) often generate convincing, fluent explanations. However, different from humans, they often generate inconsistent explanations on different inputs. For example, an LLM may explain “all birds can fly” when answering the question “Can sparrows fly?” but meanwhile answer “no” to the related question “Can penguins fly?”. Explanations should be consistent across related examples so that they allow humans to simulate the LLM’s decision process on multiple examples. We propose explanation-consistency finetuning (EC-finetuning), a method that adapts LLMs to generate more consistent natural-language explanations on related examples. EC-finetuning involves finetuning LLMs on synthetic data that is carefully constructed to contain consistent explanations. Across a variety of question-answering datasets in various domains, EC-finetuning yields a 10.0% relative explanation consistency improvement on 4 finetuning datasets, and generalizes to 7 out-of-distribution datasets not seen during finetuning (+4.5% relative). We will make our code available for reproducibility. Yanda Chen, Chandan Singh, Xiaodong Liu 0003, Simiao Zuo, Bin Yu 0001, He He 0001, Jianfeng Gao 0001 |
COLING | 6 |
| 2025 | Transformers Struggle to Learn to SearchabstractSearch is an ability foundational in many important tasks, and recent studies have shown that large language models (LLMs) struggle to perform search robustly. It is unknown whether this inability is due to a lack of data, insufficient model parameters, or fundamental limitations of the transformer architecture. In this work, we use the foundational graph connectivity problem as a testbed to generate effectively limitless high-coverage data to train small transformers and test whether they can learn to perform search. We find that, when given the right training distribution, the transformer is able to learn to search.
We analyze the algorithm that the transformer has learned through a novel mechanistic interpretability technique that enables us to extract the computation graph from the trained model. We find that for each vertex in the input graph, transformers compute the set of vertices reachable from that vertex. Each layer then progressively expands these sets, allowing the model to search over a number of vertices exponential in the number of layers.
However, we find that as the input graph size increases, the transformer has greater difficulty in learning the task. This difficulty is not resolved even as the number of parameters is increased, suggesting that increasing model scale will not lead to robust search abilities. We also find that performing search in-context (i.e., chain-of-thought) does not resolve this inability to learn to search on larger graphs. Abulhair Saparov, Srushti Pawar, Shreyas Pimpalgaonkar, Nitish Joshi, Richard Yuanzhe Pang, Vishakh Padmakumar, Mehran Kazemi, Najoung Kim, He He 0001 |
ICLR | 9 |
| 2025 | Adaptive Deployment of Untrusted LLMs Reduces Distributed ThreatsabstractAs large language models (LLMs) grow more powerful, they also become more difficult to trust. They could be either aligned with human intentions, or exhibit "subversive misalignment" -- introducing subtle errors that bypass safety checks. Although individual errors may not immediately cause harm, each increases the risk of an eventual safety failure. With this uncertainty, model deployment often grapples with the tradeoff between ensuring safety and harnessing the capabilities of untrusted models. In this work, we introduce the ``Diffuse Risk Management'' problem, aiming to balance the average-case safety and usefulness in the deployment of untrusted models over a large sequence of tasks. We approach this problem by developing a two-level framework: the single-task level (micro-protocol) and the whole-scenario level (macro-protocol). At the single-task level, we develop various \textit{micro}-protocols that use a less capable, but extensively tested (trusted) model to harness and monitor the untrusted model. At the whole-scenario level, we find an optimal \textit{macro}-protocol that uses an adaptive estimate of the untrusted model's risk to choose between micro-protocols. To evaluate the robustness of our method, we follow \textit{control evaluations} in a code generation testbed, which involves a red team attempting to generate subtly backdoored code with an LLM whose deployment is safeguarded by a blue team. Experiment results show that our approach retains 99.6\% usefulness of the untrusted model while ensuring near-perfect safety, significantly outperforming existing deployment methods. Our approach also demonstrates robustness when the trusted and untrusted models have a large capability gap. Our findings demonstrate the promise of managing diffuse risks in the deployment of increasingly capable but untrusted LLMs. Jiaxin Wen, Vivek Hebbar, Caleb Larson, Aryan Bhatt, Ansh Radhakrishnan, Mrinank Sharma, Henry Sleight, Shi Feng 0005, He He 0001, Ethan Perez, Buck Shlegeris, Akbir Khan |
ICLR | 9 |
| 2025 | Language Models Learn to Mislead Humans via RLHFabstractLanguage models (LMs) can produce errors that are hard to detect for humans, especially when the task is complex.
RLHF, the most popular post-training method, may exacerbate this problem: to achieve higher rewards, LMs might get better at convincing humans that they are right even when they are wrong. We study this phenomenon under a standard RLHF pipeline, calling it ``U-Sophistry'' since it is \textbf{U}nintended by model developers. Specifically, we ask time-constrained (e.g., 3-10 minutes) human subjects to evaluate the correctness of model outputs and calculate humans' accuracy against gold labels. On a question-answering task (QuALITY) and programming task (APPS), RLHF makes LMs better at convincing our subjects but not at completing the task correctly. RLHF also makes the model harder to evaluate: our subjects' false positive rate increases by 24.1% on QuALITY and 18.3% on APPS.
Finally, we show that probing, a state-of-the-art approach for detecting \textbf{I}ntended Sophistry (e.g.~backdoored LMs), does not generalize to U-Sophistry. Our results highlight an important failure mode of RLHF and call for more research in assisting humans to align them. Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, He He 0001, Shi Feng 0005 |
ICLR | 8 |
| 2025 | Predicting Empirical AI Research Outcomes with Language ModelsabstractMany promising-looking ideas in AI research fail to deliver, but their validation takes substantial human labor and compute. Predicting an idea's chance of success is thus crucial for accelerating empirical AI research, a skill that even expert researchers can only acquire through substantial experience. We build the first benchmark for this task and compare LMs with human experts. Concretely, given two research ideas (e.g., two jailbreaking methods), we aim to predict which will perform better on a set of benchmarks.
We scrape ideas and experimental results from conference papers, yielding 1,585 human-verified idea pairs \textit{published after our base model's cut-off date} for testing, and 6,000 pairs for training.
We then develop a system that combines a fine-tuned GPT-4.1 with a paper retrieval agent, and we recruit 25 human experts to compare with.
In the NLP domain, our system beats human experts by a large margin (64.4\% v.s. 48.9\%).
On the full test set, our system achieves 77\% accuracy, while off-the-shelf frontier LMs like o3 perform no better than random guessing, even with the same retrieval augmentation.
We verify that our system does not exploit superficial features like idea complexity through extensive human-written and LM-designed robustness tests.
Finally, we evaluate our system on unpublished novel ideas, including ideas generated by an AI ideation agent.
Our system achieves 63.6\% accuracy, demonstrating its potential as a reward model for improving idea generation models.
Altogether, our results outline a promising new direction for LMs to accelerate empirical AI research. Jiaxin Wen, Chenglei Si, Yueh-Han Chen, He He 0001, Shi Feng 0005 |
NeurIPS | 4 |
| 2024 | Parallel Structures in Pre-training Data Yield In-Context LearningabstractPre-trained language models (LMs) are capable of in-context learning (ICL): they can adapt to a task with only a few examples given in the prompt without any parameter update.However, it is unclear where this capability comes from as there is a stark distribution shift between pre-training text and ICL prompts.In this work, we study what patterns of the pretraining data contribute to ICL.We find that LMs' ICL ability depends on parallel structures in the pre-training data-pairs of phrases following similar templates in the same context window.Specifically, we detect parallel structures by checking whether training on one phrase improves prediction of the other, and conduct ablation experiments to study their effect on ICL.We show that removing parallel structures in the pre-training data reduces LMs' ICL accuracy by 51% (vs 2% from random ablation).This drop persists even when excluding common patterns such as n-gram repetitions and long-range dependency, showing the diversity and generality of parallel structures.A closer look at the detected parallel structures indicates that they cover diverse linguistic tasks and span long distances in the data. Yanda Chen, Chen Zhao 0013, Kathy McKeown, He He 0001 |
ACL (1) | 5 |
| 2024 | Personas as a Way to Model Truthfulness in Language ModelsabstractLarge language models (LLMs) are trained on vast amounts of text from the internet, which contains both factual and misleading information about the world.While unintuitive from a classic view of language models, recent work has shown that the truth value of a statement can be elicited from the model's representations.This paper presents an explanation, persona hypothesis, for why LLMs appear to know the truth despite not being trained with truth labels.We hypothesize that the pretraining data is generated by groups of (un)truthful agents whose outputs share common features, and they form a (un)truthful persona.By training on this data, LMs can infer and represent the persona in its activation space.This allows the model to separate truth from falsehoods and controls the truthfulness of its generation.We show evidence for the persona hypothesis via two observations: (1) we can probe whether a model's answer will be truthful before it is generated; (2) finetuning a model on a set of true facts improves its truthfulness on unseen topics.Next, using arithmetics as a synthetic environment, we show that structures of the pretraining data are crucial for the model to infer the truthful persona.Overall, our findings suggest that models can exploit hierarchical structures in the data to learn abstract concepts like truthfulness. Nitish Joshi, Javier Rando, Abulhair Saparov, Najoung Kim, He He 0001 |
EMNLP | 5 |
| 2024 | LLMs Are Prone to Fallacies in Causal InferenceabstractRecent work shows that causal facts can be effectively extracted from LLMs through prompting, facilitating the creation of causal graphs for causal inference tasks.However, it is unclear if this success is limited to explicitly-mentioned causal facts in the pretraining data which the model can memorize.Thus, this work investigates: Can LLMs infer causal relations from other relational data in text?To disentangle the role of memorized causal facts vs inferred causal relations, we finetune LLMs on synthetic data containing temporal, spatial and counterfactual relations, and measure whether the LLM can then infer causal relations.We find that: (a) LLMs are susceptible to inferring causal relations from the order of two entity mentions in text (e.g.X mentioned before Y implies X causes Y); (b) if the order is randomized, LLMs still suffer from the post hoc fallacy, i.e.X occurs before Y (temporal relation) implies X causes Y.We also find that while LLMs can correctly deduce the absence of causal relations from temporal and spatial relations, they have difficulty inferring causal relations from counterfactuals, questioning their understanding of causality. Nitish Joshi, Abulhair Saparov, He He 0001 |
EMNLP | 4 |
| 2024 | Does Writing with Language Models Reduce Content Diversity?abstractLarge language models (LLMs) have led to a surge in collaborative writing with model assistance. As different users incorporate suggestions from the same model, there is a risk of decreased diversity in the produced content, potentially limiting diverse perspectives in public discourse. In this work, we measure the impact of co-writing on diversity via a controlled experiment, where users write argumentative essays in three setups---using a base LLM (GPT3), a feedback-tuned LLM (InstructGPT), and writing without model help. We develop a set of diversity metrics and find that writing with InstructGPT (but not the GPT3) results in a statistically significant reduction in diversity. Specifically, it increases the similarity between the writings of different authors and reduces the overall lexical and content diversity. We additionally find that this effect is mainly attributable to InstructGPT contributing less diverse text to co-written essays. In contrast, the user-contributed text remains unaffected by model collaboration. This suggests that the recent improvement in generation quality from adapting models to human feedback might come at the cost of more homogeneous and less diverse content. Vishakh Padmakumar, He He 0001 |
ICLR | 2 |
| 2024 | Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsabstractLarge language models (LLMs) are trained to imitate humans to explain human decisions. However, do LLMs explain themselves? Can they help humans build mental models of how LLMs process different inputs? To answer these questions, we propose to evaluate $\textbf{counterfactual simulatability}$ of natural language explanations: whether an explanation can enable humans to precisely infer the model’s outputs on diverse counterfactuals of the explained input. For example, if a model answers ”$\textit{yes}$” to the input question ”$\textit{Can eagles fly?}$” with the explanation ”$\textit{all birds can fly}$”, then humans would infer from the explanation that it would also answer ”$\textit{yes}$” to the counterfactual input ”$\textit{Can penguins fly?}$”. If the explanation is precise, then the model’s answer should match humans’ expectations. We implemented two metrics based on counterfactual simulatability: precision and generality. We generated diverse counterfactuals automatically using LLMs. We then used these metrics to evaluate state-of-the-art LLMs (e.g., GPT-4) on two tasks: multi-hop factual reasoning and reward modeling. We found that LLM’s explanations have low precision and that precision does not correlate with plausibility. Therefore, naively optimizing human approvals (e.g., RLHF) may be insufficient. Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 0013, He He 0001, Jacob Steinhardt, Kathy McKeown |
ICML | 5 |
| 2024 | Show Your Work with Confidence: Confidence Bands for Tuning CurvesabstractNicholas Lourie, Kyunghyun Cho, He He. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nicholas Lourie, Kyunghyun Cho, He He 0001 |
NAACL-HLT | 3 |
| 2024 | The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language ModelsabstractHuman feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about the methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigate these questions, we introduce PRISM, a new dataset which maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, to their contextual preferences and fine-grained feedback in 8,011 live conversations with 21 LLMs. With PRISM, we contribute (i) wider geographic and demographic participation in feedback; (ii) census-representative samples for two countries (UK, US); and (iii) individualised ratings that link to detailed participant profiles, permitting personalisation and attribution of sample artefacts. We target subjective and multicultural perspectives on value-laden and controversial issues, where we expect interpersonal and cross-cultural disagreement. We use PRISM in three case studies to demonstrate the need for careful consideration of which humans provide alignment data. Hannah Kirk, Alexander Whitefield, Paul Röttger, Andrew M. Bean 0001, Aikaterini Margatina, Rafael Mosquera, Juan Ciro, Max Bartolo, Adina Williams, He He 0001, Bertie Vidgen, Scott A. Hale |
NeurIPS | 10 |
| 2024 | Iterative Reasoning Preference OptimizationabstractIterative preference optimization methods have recently been shown to perform well for general instruction tuning tasks, but typically make little improvement on reasoning tasks. In this work we develop an iterative approach that optimizes the preference between competing generated Chain-of-Thought (CoT) candidates by optimizing for winning vs. losing reasoning steps. We train using a modified DPO loss with an additional negative log-likelihood term, which we find to be crucial. We show reasoning improves across repeated iterations of this scheme. While only relying on examples in the training set, our approach results in increasing accuracy on GSM8K, MATH, and ARC-Challenge for Llama-2-70B-Chat, outperforming other Llama-2-based models not relying on additionally sourced datasets. For example, we see a large improvement from 55.6% to 81.6% on GSM8K and an accuracy of 88.7% with majority voting out of 32 samples. Richard Yuanzhe Pang, Weizhe Yuan, He He 0001, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston |
NeurIPS | 3 |
| 2024 | : Visualization of AI-Assisted Task Guidance in ARabstractThe concept of augmented reality (AR) assistants has captured the human imagination for decades, becoming a staple of modern science fiction. To pursue this goal, it is necessary to develop artificial intelligence (AI)-based methods that simultaneously perceive the 3D environment, reason about physical tasks, and model the performer, all in real-time. Within this framework, a wide variety of sensors are needed to generate data across different modalities, such as audio, video, depth, speech, and time-of-flight. The required sensors are typically part of the AR headset, providing performer sensing and interaction through visual, audio, and haptic feedback. AI assistants not only record the performer as they perform activities, but also require machine learning (ML) models to understand and assist the performer as they interact with the physical world. Therefore, developing such assistants is a challenging task. We propose ARGUS, a visual analytics system to support the development of intelligent AR assistants. Our system was designed as part of a multi-year-long collaboration between visualization researchers and ML and AR experts. This co-design process has led to advances in the visualization of ML in AR. Our system allows for online visualization of object, action, and step detection as well as offline analysis of previously recorded AR sessions. It visualizes not only the multimodal sensor data streams but also the output of the ML models. This allows developers to gain insights into the performer activities as well as the ML models, helping them troubleshoot, improve, and fine-tune the components of the AR assistant. Sonia Castelo Quispe, João Rulff, Erin McGowan, Bea Steers, Guande Wu, Shaoyu Chen, Irán R. Román, Roque Lopez, Ethan Brewer, Chen Zhao 0013, Kyunghyun Cho, He He 0001, Qi Sun 0003, Huy T. Vo, Juan Pablo Bello, Michael Krone, Cláudio T. Silva |
IEEE Trans. Vis. Comput. Graph. | 13 |
| 2023 | Reward Gaming in Conditional Text GenerationabstractTo align conditional text generation model outputs with desired behaviors, there has been an increasing focus on training the model using reinforcement learning (RL) with reward functions learned from human annotations.Under this framework, we identify three common cases where high rewards are incorrectly assigned to undesirable patterns: noise-induced spurious correlation, naturally occurring spurious correlation, and covariate shift.We show that even though learned metrics achieve high performance on the distribution of the data used to train the reward function, the undesirable patterns may be amplified during RL training of the text generation model.While there has been discussion about reward gaming in the RL or safety community, in this discussion piece, we would like to highlight reward gaming in the natural language generation (NLG) community using concrete conditional text generation examples and discuss potential fixes and areas for future work. Richard Yuanzhe Pang, Vishakh Padmakumar, Thibault Sellam, Ankur P. Parikh, He He 0001 |
ACL (1) | 5 |
| 2023 | Measuring Inductive Biases of In-Context Learning with Underspecified DemonstrationsabstractIn-context learning (ICL) is an important paradigm for adapting large language models (LLMs) to new tasks, but the generalization behavior of ICL remains poorly understood.We investigate the inductive biases of ICL from the perspective of feature bias: which feature ICL is more likely to use given a set of underspecified demonstrations in which two features are equally predictive of the labels.First, we characterize the feature biases of GPT-3 models by constructing underspecified demonstrations from a range of NLP datasets and feature combinations.We find that LLMs exhibit clear feature biases-for example, demonstrating a strong bias to predict labels according to sentiment rather than shallow lexical features, like punctuation.Second, we evaluate the effect of different interventions that are designed to impose an inductive bias in favor of a particular feature, such as adding a natural language instruction or using semantically relevant label words.We find that, while many interventions can influence the learner to prefer a particular feature, it can be difficult to overcome strong prior biases.Overall, our results provide a broader picture of the types of features that ICL may be more likely to exploit and how to impose inductive biases that are better aligned with the intended task. 1 Chenglei Si, Dan Friedman, Nitish Joshi, Shi Feng 0005, Danqi Chen 0001, He He 0001 |
ACL (1) | 6 |
| 2023 | Efficient Shapley Values Estimation by Amortization for Text ClassificationabstractDespite the popularity of Shapley Values in explaining neural text classification models, computing them is prohibitive for large pretrained models due to a large number of model evaluations.In practice, Shapley Values are often estimated with a small number of stochastic model evaluations.However, we show that the estimated Shapley Values are sensitive to random seed choices -the top-ranked features often have little overlap across different seeds, especially on examples with longer input texts.This can only be mitigated by aggregating thousands of model evaluations, which on the other hand, induces substantial computational overheads.To mitigate the trade-off between stability and efficiency, we develop an amortized model that directly predicts each input feature's Shapley Value without additional model evaluations.It is trained on a set of examples whose Shapley Values are estimated from a large number of model evaluations to ensure stability.Experimental results on two text classification datasets demonstrate that our amortized model estimates Shapley Values accurately with up to 60 times speedup compared to traditional methods.Furthermore, the estimated values are stable as the inference is deterministic.We release our code at https://github.com/yangalan123/ Amortized-Interpretability. Chenghao Yang 0001, Fan Yin, He He 0001, Kai-Wei Chang 0001, Xiaofei Ma 0001, Bing Xiang |
ACL (1) | 3 |
| 2023 | Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
Abulhair Saparov, He He 0001 |
ICLR | 2 |
| 2023 | Extrapolative Controlled Sequence Generation via Iterative RefinementabstractWe study the problem of extrapolative controlled generation, i.e., generating sequences with attribute values beyond the range seen in training. This task is of significant importance in automated design, especially drug discovery, where the goal is to design novel proteins that are better (e.g., more stable) than existing sequences. Thus, by definition the target sequences and their attribute values are out of the training distribution, posing challenges to existing methods that aim to directly generate the target sequence. Instead, in this work, we propose Iterative Controlled Extrapolation (ICE) which iteratively makes local edits to a sequence to enable extrapolation. We train the model on synthetically generated sequence pairs that demonstrate small improvement in the attribute value. Results on one natural language task (sentiment analysis) and two protein engineering tasks (ACE2 stability and AAV fitness) show that ICE outperforms state-of-the-art approaches despite its simplicity. Vishakh Padmakumar, Richard Yuanzhe Pang, He He 0001, Ankur P. Parikh |
ICML | 3 |
| 2023 | Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD ExamplesabstractGiven the intractably large size of the space of proofs, any model that is capable of general deductive reasoning must generalize to proofs of greater complexity. Recent studies have shown that large language models (LLMs) possess some abstract deductive reasoning ability given chain-of-thought prompts. However, they have primarily been tested on proofs using modus ponens or of a specific size, and from the same distribution as the in-context examples. To measure the general deductive reasoning ability of LLMs, we test on a broad set of deduction rules and measure their ability to generalize to more complex proofs from simpler demonstrations from multiple angles: depth-, width-, and compositional generalization. To facilitate systematic exploration, we construct a new synthetic and programmable reasoning dataset that enables control over deduction rules and proof complexity. Our experiments on four LLMs of various sizes and training objectives show that they are able to generalize to compositional proofs. However, they have difficulty generalizing to longer proofs, and they require explicit demonstrations to produce hypothetical subproofs, specifically in proof by cases and proof by contradiction. Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Mehran Kazemi, Najoung Kim, He He 0001 |
NeurIPS | 7 |
| 2022 | Meta-learning via Language Model In-context TuningabstractThe goal of meta-learning is to learn to adapt to a new task with only a few labeled examples.Inspired by the recent progress in large language models, we propose in-context tuning (ICT), which recasts task adaptation and prediction as a simple sequence prediction problem: to form the input sequence, we concatenate the task instruction, labeled in-context examples, and the target input to predict; to metatrain the model to learn from in-context examples, we fine-tune a pre-trained language model (LM) to predict the target label given the input sequence on a collection of tasks.We benchmark our method on two collections of text classification tasks: LAMA and Bina-ryClfs.Compared to MAML which adapts the model through gradient descent, our method leverages the inductive bias of pre-trained LMs to perform pattern matching, and outperforms MAML by an absolute 6% average AUC-ROC score on BinaryClfs, gaining more advantage with increasing model size.Compared to non-fine-tuned in-context learning (i.e.prompting a raw LM), in-context tuning meta-trains the model to learn from in-context examples.On BinaryClfs, ICT improves the average AUC-ROC score by an absolute 10%, and reduces the variance due to example ordering by 6x and example choices by 2x. Yanda Chen, Ruiqi Zhong, Sheng Zha, George Karypis, He He 0001 |
ACL (1) | 5 |
| 2022 | An Investigation of the (In)effectiveness of Counterfactually Augmented DataabstractWhile pretrained language models achieve excellent performance on natural language understanding benchmarks, they tend to rely on spurious correlations and generalize poorly to out-of-distribution (OOD) data.Recent work has explored using counterfactuallyaugmented data (CAD)-data generated by minimally perturbing examples to flip the ground-truth label-to identify robust features that are invariant under distribution shift.However, empirical results using CAD during training for OOD generalization have been mixed.To explain this discrepancy, through a toy theoretical example and empirical analysis on two crowdsourced CAD datasets, we show that: (a) while features perturbed in CAD are indeed robust features, it may prevent the model from learning unperturbed robust features; and (b) CAD may exacerbate existing spurious correlations in the data.Our results thus show that the lack of perturbation diversity limits CAD's effectiveness on OOD generalization, calling for innovative crowdsourcing procedures to elicit diverse perturbation of examples. Nitish Joshi, He He 0001 |
ACL (1) | 2 |
| 2022 | Faithful or Extractive? On Mitigating the Faithfulness-Abstractiveness Trade-off in Abstractive SummarizationabstractDespite recent progress in abstractive summarization, systems still suffer from faithfulness errors.While prior work has proposed models that improve faithfulness, it is unclear whether the improvement comes from an increased level of extractiveness of the model outputs as one naive way to improve faithfulness is to make summarization models more extractive.In this work, we present a framework for evaluating the effective faithfulness of summarization systems, by generating a faithfulnessabstractiveness trade-off curve that serves as a control at different operating points on the abstractiveness spectrum.We then show that the baseline system as well as recently proposed methods for improving faithfulness, fail to consistently improve over the control at the same level of abstractiveness.Finally, we learn a selector to identify the most faithful and abstractive summary for a given document, and show that this system can attain higher faithfulness scores in human evaluations while being more abstractive than the baseline system on two datasets.Moreover, we show that our system is able to achieve a better faithfulnessabstractiveness trade-off than the control at the same level of abstractiveness. Faisal Ladhak, Esin Durmus, He He 0001, Claire Cardie, Kathy McKeown |
ACL (1) | 3 |
| 2022 | Help me write a Poem - Instruction Tuning as a Vehicle for Collaborative Poetry WritingabstractRecent work in training large language models (LLMs) to follow natural language instructions has opened up exciting opportunities for natural language interface design.Building on the prior success of LLMs in the realm of computerassisted creativity, we aim to study if LLMs can improve the quality of user-generated content through collaboration.We present CoPoet, a collaborative poetry writing system.In contrast to auto-completing a user's text, CoPoet is controlled by user instructions that specify the attributes of the desired text, such as Write a sentence about 'love' or Write a sentence ending in 'fly'.The core component of our system is a language model fine-tuned on a diverse collection of instructions for poetry writing.Our model is not only competitive with publicly available LLMs trained on instructions (InstructGPT), but is also capable of satisfying unseen compositional instructions.A study with 15 qualified crowdworkers shows that users successfully write poems with CoPoet on diverse topics ranging from Monarchy to Climate change.Further, the collaboratively written poems are preferred by third-party evaluators over those written without the system. 1 Tuhin Chakrabarty, Vishakh Padmakumar, He He 0001 |
EMNLP | 3 |
| 2022 | Are All Spurious Features in Natural Language Alike? An Analysis through a Causal LensabstractThe term 'spurious correlations' has been used in NLP to informally denote any undesirable feature-label correlations.However, a correlation can be undesirable because (i) the feature is irrelevant to the label (e.g.punctuation in a review), or (ii) the feature's effect on the label depends on the context (e.g.negation words in a review), which is ubiquitous in language tasks.In case (i), we want the model to be invariant to the feature, which is neither necessary nor sufficient for prediction.But in case (ii), even an ideal model (e.g.humans) must rely on the feature, since it is necessary (but not sufficient) for prediction.Therefore, a more fine-grained treatment of spurious features is needed to specify the desired model behavior.We formalize this distinction using a causal model and probabilities of necessity and sufficiency, which delineates the causal relations between a feature and a label.We then show that this distinction helps explain results of existing debiasing methods on different spurious features, and demystifies surprising results such as the encoding of spurious features in model representations after debiasing. Nitish Joshi, Xiang Pan 0001, He He 0001 |
EMNLP | 3 |
| 2022 | Improving Faithfulness by Augmenting Negative Summaries from Fake DocumentsabstractCurrent abstractive summarization systems tend to hallucinate content that is unfaithful to the source document, posing a risk of misinformation.To mitigate hallucination, we must teach the model to distinguish hallucinated summaries from faithful ones.However, the commonly used maximum likelihood training does not disentangle factual errors from other model errors.To address this issue, we propose a back-translation-style approach to augment negative samples that mimic factual errors made by the model.Specifically, we train an elaboration model that generates hallucinated documents given the reference summaries, and then generates negative summaries from the fake documents.We incorporate the negative samples into training through a controlled generator, which produces faithful/unfaithful summaries conditioned on the control codes.Additionally, we find that adding textual entailment data through multitasking further boosts the performance.Experiments on three datasets (XSum, GigaWord, and WikiHow) show that our method consistently improves faithfulness without sacrificing informativeness according to both human and automatic evaluation. 1 Faisal Ladhak, Esin Durmus, He He 0001 |
EMNLP | 4 |
| 2022 | Machine-in-the-Loop Rewriting for Creative Image CaptioningabstractMachine-in-the-loop writing aims to build models that assist humans to accomplish their writing tasks more effectively.Prior work has found that providing users a machine-written draft or sentence-level continuations has limited success since the generated text tends to deviate from users' intention.To allow the user to retain control over the content, we train a rewriting model that, when prompted, modifies specified spans of text within the user's original draft to introduce descriptive and figurative elements in the text.We evaluate the model on its ability to collaborate with humans on the task of creative image captioning.On a user study through Amazon Mechanical Turk, our model is rated to be more helpful by users than a baseline infilling language model.In addition, third-party evaluation shows that users write more descriptive and figurative captions when collaborating with our model compared to completing the task alone.However, the improvement is not uniform across user groups: the model is more helpful to skilled users, which risks widening the gap between skilled and novice users, highlighting a need for careful, user-centric evaluation of interactive systems. 1 Vishakh Padmakumar, He He 0001 |
NAACL-HLT | 2 |
| 2022 | Exploring the Role of Task Transferability in Large-Scale Multi-Task LearningabstractVishakh Padmakumar, Leonard Lausen, Miguel Ballesteros, Sheng Zha, He He, George Karypis. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Vishakh Padmakumar, Leonard Lausen, Miguel Ballesteros, Sheng Zha, He He 0001, George Karypis |
NAACL-HLT | 5 |
| 2022 | QuALITY: Question Answering with Long Input Texts, Yes!abstractRichard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, Samuel Bowman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He 0001, Samuel R. Bowman |
NAACL-HLT | 10 |
| 2021 | Unsupervised Extractive Summarization using Pointwise Mutual InformationabstractUnsupervised approaches to extractive summarization usually rely on a notion of sentence importance defined by the semantic similarity between a sentence and the document.We propose new metrics of relevance and redundancy using pointwise mutual information (PMI) between sentences, which can be easily computed by a pre-trained language model.Intuitively, a relevant sentence allows readers to infer the document content (high PMI with the document), and a redundant sentence can be inferred from the summary (high PMI with the summary).We then develop a greedy sentence selection algorithm to maximize relevance and minimize redundancy of extracted sentences.We show that our method outperforms similarity-based methods on datasets in a range of domains including news, medical journal articles, and personal anecdotes. Vishakh Padmakumar, He He 0001 |
EACL | 2 |
| 2021 | Text Generation by Learning from Demonstrations
Richard Yuanzhe Pang, He He 0001 |
ICLR | 2 |
| 2021 | IRM - when it works and when it doesn't: A test case of natural language inferenceabstractInvariant Risk Minimization (IRM) is a recently proposed framework for out-of-distribution (o.o.d) generalization. Most of the studies on IRM so far have focused on theoretical results, toy problems, and simple models. In this work, we investigate the applicability of IRM to bias mitigation-a special case of o.o.d generalization-in increasingly naturalistic settings and deep models. Using natural language inference (NLI) as a test case, we start with a setting where both the dataset and the bias are synthetic, continue with a natural dataset and synthetic bias, and end with a fully realistic setting with natural datasets and bias. Our results show that in naturalistic settings, learning complex features in place of the bias proves to be difficult, leading to a rather small improvement over empirical risk minimization. Moreover, we find that in addition to being sensitive to random seeds, the performance of IRM also depends on several critical factors, notably dataset size, bias prevalence, and bias strength, thus limiting IRM's advantage in practical scenarios. Our results highlight key challenges in applying IRM to real-world scenarios, calling for a more naturalistic characterization of the problem setup for o.o.d generalization. Yana Dranker, He He 0001, Yonatan Belinkov |
NeurIPS | 2 |
| 2020 | FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationabstractNeural abstractive summarization models are prone to generate content inconsistent with the source document, i.e. unfaithful.Existing automatic metrics do not capture such mistakes effectively.We tackle the problem of evaluating faithfulness of a generated summary given its source document.We first collected human annotations of faithfulness for outputs from numerous models on two datasets.We find that current models exhibit a trade-off between abstractiveness and faithfulness: outputs with less word overlap with the source document are more likely to be unfaithful.Next, we propose an automatic question answering (QA) based metric for faithfulness, FEQA, 1 which leverages recent advances in reading comprehension.Given questionanswer pairs generated from the summary, a QA model extracts answers from the document; non-matched answers indicate unfaithful information in the summary.Among metrics based on word overlap, embedding similarity, and learned language understanding models, our QA-based metric has significantly higher correlation with human faithfulness scores, especially on highly abstractive summaries.* Most of the work is done while the authors were at Amazon Web Services AI.1 Faithfulness Evaluation with Question Answering. Esin Durmus, He He 0001, Mona T. Diab |
ACL | 2 |
| 2020 | GluonCV and GluonNLP: Deep Learning in Computer Vision and Natural Language ProcessingabstractWe present GluonCV and GluonNLP, the deep learning toolkits for computer vision and natural language processing based on Apache MXNet (incubating). These toolkits provide state-of-the-art pre-trained models, training scripts, and training logs, to facilitate rapid prototyping and promote reproducible research. We also provide modular APIs with flexible building blocks to enable efficient customization. Leveraging the MXNet ecosystem, the deep learning models in GluonCV and GluonNLP can be deployed onto a variety of platforms with different programming languages. The Apache 2.0 license has been adopted by GluonCV and GluonNLP to allow for software distribution, modification, and usage. He He 0001, Tong He 0002, Leonard Lausen, Mu Li 0003, Haibin Lin, Xingjian Shi, Chenguang Wang 0001, Junyuan Xie, Sheng Zha, Aston Zhang, Hang Zhang 0005, Zhi Zhang 0005, Shuai Zheng 0004, Yi Zhu 0001 |
J. Mach. Learn. Res. | 2 |
| 2020 | An Empirical Study on Robustness to Spurious Correlations using Pre-trained Language ModelsabstractRecent work has shown that pre-trained language models such as BERT improve robustness to spurious correlations in the dataset. Intrigued by these results, we find that the key to their success is generalization from a small amount of counterexamples where the spurious correlations do not hold. When such minority examples are scarce, pre-trained models perform as poorly as models trained from scratch. In the case of extreme minority, we propose to use multi-task learning (MTL) to improve generalization. Our experiments on natural language inference and paraphrase identification show that MTL with the right auxiliary tasks significantly improves performance on challenging examples without hurting the in-distribution performance. Further, we show that the gain from MTL mainly comes from improved generalization from the minority examples. Our results highlight the importance of data diversity for overcoming spurious correlations. 1 Lifu Tu, Garima Lalwani, Spandana Gella, He He 0001 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2019 | A Dynamic Strategy Coach for Effective NegotiationabstractNegotiation is a complex activity involving strategic reasoning, persuasion, and psychology.An average person is often far from an expert in negotiation.Our goal is to assist humans to become better negotiators through a machine-in-the-loop approach that combines machine's advantage at data-driven decisionmaking and human's language generation ability.We consider a bargaining scenario where a seller and a buyer negotiate the price of an item for sale through a text-based dialog.Our negotiation coach monitors messages between them and recommends tactics in real time to the seller to get a better deal (e.g., "reject the proposal and propose a price", "talk about your personal experience with the product").The best strategy and tactics largely depend on the context (e.g., the current price, the buyer's attitude).Therefore, we first identify a set of negotiation tactics, then learn to predict the best strategy and tactics in a given dialog context from a set of human-human bargaining dialogs.Evaluation on human-human dialogs shows that our coach increases the profits of the seller by almost 60%. 1 Yiheng Zhou, He He 0001, Alan W. Black, Yulia Tsvetkov |
SIGdial | 2 |
| 2018 | Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use ContextabstractWe know very little about how neural language models (LM) use prior linguistic context.In this paper, we investigate the role of context in an LSTM LM, through ablation studies.Specifically, we analyze the increase in perplexity when prior context words are shuffled, replaced, or dropped.On two standard datasets, Penn Treebank and WikiText-2, we find that the model is capable of using about 200 tokens of context on average, but sharply distinguishes nearby context (recent 50 tokens) from the distant history.The model is highly sensitive to the order of words within the most recent sentence, but ignores word order in the long-range context (beyond 50 tokens), suggesting the distant past is modeled only as a rough semantic field or topic.We further find that the neural caching model (Grave et al., 2017b) especially helps the LSTM to copy words from within this distant context.Overall, our analysis not only provides a better understanding of how neural LMs use their context, but also sheds light on recent success from cache-based models. Urvashi Khandelwal, He He 0001, Peng Qi 0003, Daniel Jurafsky |
ACL (1) | 2 |
| 2018 | QuAC: Question Answering in ContextabstractWe present QuAC, a dataset for Question Answering in Context that contains 14K information-seeking QA dialogs (100K questions in total).The dialogs involve two crowd workers: (1) a student who poses a sequence of freeform questions to learn as much as possible about a hidden Wikipedia text, and (2) a teacher who answers the questions by providing short excerpts from the text.QuAC introduces challenges not found in existing machine comprehension datasets: its questions are often more open-ended, unanswerable, or only meaningful within the dialog context, as we show in a detailed qualitative evaluation.We also report results for a number of reference models, including a recently state-ofthe-art reading comprehension architecture extended to model dialog context.Our best model underperforms humans by 20 F1, suggesting that there is significant room for future work on this data.Dataset, baseline, and leaderboard available at http://quac.ai.How was perversion handled?How long was he there?How popular did she become?How did Mark Felt contact Woodword?How did the meeting go?How did it do on the charts?When was she born?When was it founded?When was the breakup? Eunsol Choi, He He 0001, Mohit Iyyer, Mark Yatskar, Scott Yih, Yejin Choi 0001, Percy Liang, Luke Zettlemoyer |
EMNLP | 2 |
| 2018 | Decoupling Strategy and Generation in Negotiation DialoguesabstractWe consider negotiation settings in which two agents use natural language to bargain on goods.Agents need to decide on both high-level strategy (e.g., proposing $50) and the execution of that strategy (e.g., generating "The bike is brand new.Selling for just $50!").Recent work on negotiation trains neural models, but their end-to-end nature makes it hard to control their strategy, and reinforcement learning tends to lead to degenerate solutions.In this paper, we propose a modular approach based on coarse dialogue acts (e.g., propose(price=50)) that decouples strategy and generation.We show that we can flexibly set the strategy using supervised learning, reinforcement learning, or domain-specific knowledge without degeneracy, while our retrieval-based generation can maintain context-awareness and produce diverse utterances.We test our approach on the recently proposed DEALORNODEAL game, and we also collect a richer dataset based on real items on Craigslist.Human evaluation shows that our systems achieve higher task success rate and more human-like negotiation behavior than previous approaches. He He 0001, Derek Chen, Anusha Balakrishnan, Percy Liang |
EMNLP | 1 |
| 2018 | Delete, Retrieve, Generate: a Simple Approach to Sentiment and Style TransferabstractJuncen Li, Robin Jia, He He, Percy Liang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Juncen Li, Robin Jia, He He 0001, Percy Liang |
NAACL-HLT | 3 |
| 2017 | Learning Symmetric Collaborative Dialogue Agents with Dynamic Knowledge Graph EmbeddingsabstractWe study a symmetric collaborative dialogue setting in which two agents, each with private knowledge, must strategically communicate to achieve a common goal.The open-ended dialogue state in this setting poses new challenges for existing dialogue systems.We collected a dataset of 11K human-human dialogues, which exhibits interesting lexical, semantic, and strategic elements.To model both structured knowledge and unstructured language, we propose a neural model with dynamic knowledge graph embeddings that evolve as the dialogue progresses.Automatic and human evaluations show that our model is both more effective at achieving the goal and more human-like than baseline neural and rule-based models. He He 0001, Anusha Balakrishnan, Mihail Eric, Percy Liang |
ACL (1) | 1 |
| 2016 | Opponent Modeling in Deep Reinforcement LearningabstractOpponent modeling is necessary in multi-agent settings where secondary agents with competing goals also adapt their strategies, yet it remains challenging because of strategies’ complex interaction and the non-stationary nature. Most previous work focuses on developing probabilistic models or parameterized strategies for specific applications. Inspired by the recent success of deep reinforcement learning, we present neural-based models that jointly learn a policy and the behavior of opponents. Instead of explicitly predicting the opponent’s action, we encode observation of the opponents into a deep Q-Network (DQN), while retaining explicit modeling under multitasking. By using a Mixture-of-Experts architecture, our model automatically discovers different strategy patterns of opponents even without extra supervision. We evaluate our models on a simulated soccer game and a popular trivia game, showing superior performance over DQN and its variants. He He 0001, Jordan L. Boyd-Graber |
ICML | 1 |
| 2016 | Interpretese vs. Translationese: The Uniqueness of Human Strategies in Simultaneous InterpretationabstractComputational approaches to simultaneous interpretation are stymied by how little we know about the tactics human interpreters use.We produce a parallel corpus of translated and simultaneously interpreted text and study differences between them through a computational approach.Our analysis reveals that human interpreters regularly apply several effective tactics to reduce translation latency, including sentence segmentation and passivization.In addition to these unique, clever strategies, we show that limited human memory also causes other idiosyncratic properties of human interpretation such as generalization and omission of source content. He He 0001, Jordan L. Boyd-Graber, Hal Daumé III |
HLT-NAACL | 1 |
| 2016 | A Credit Assignment Compiler for Joint PredictionabstractMany machine learning applications involve jointly predicting multiple mutually dependent output variables. Learning to search is a family of methods where the complex decision problem is cast into a sequence of decisions via a search space. Although these methods have shown promise both in theory and in practice, implementing them has been burdensomely awkward. In this paper, we show the search space can be defined by an arbitrary imperative program, turning learning to search into a credit assignment compiler. Altogether with the algorithmic improvements for the compiler, we radically reduce the complexity of programming and the running time. We demonstrate the feasibility of our approach on multiple joint prediction tasks. In all cases, we obtain accuracies as high as alternative approaches, at drastically reduced execution and programming time. Kai-Wei Chang 0001, He He 0001, Stéphane Ross, Hal Daumé III, John Langford 0001 |
NIPS | 2 |
| 2016 | Object detection in 20 questionsabstractWe propose a novel general strategy for object detection. Instead of passively evaluating all object detectors at all possible locations in an image, we develop a divide-and-conquer approach by actively and sequentially evaluating contextual cues related to the query based on the scene and previous evaluations — like playing a "20 Questions" game — to decide where to search for the object. We formulate the problem as a Markov Decision Process and learn a search policy by reinforcement learning. To demonstrate the efficacy of our generic algorithm, we apply the 20 questions approach in the recent framework of simultaneous object detection and segmentation. Experimental results on the Pascal VOC dataset show that our algorithm reduces about 45.3% of the object proposals and 36% of average evaluation time while achieving better average precision compared to exhaustive search. Xi Stephen Chen, He He 0001, Larry Davis 0001 |
WACV | 2 |
| 2015 | Syntax-based Rewriting for Simultaneous Machine TranslationabstractDivergent word order between languages causes delay in simultaneous machine translation.We present a sentence rewriting method that generates more monotonic translations to improve the speedaccuracy tradeoff.We design grammaticality and meaning-preserving syntactic transformation rules that operate on constituent parse trees.We apply the rules to reference translations to make their word order closer to the source language word order.On Japanese-English translation (two languages with substantially different structure), incorporating the rewritten, more monotonic reference translation into a phrase-based machine translation system enables better translations faster than a baseline system that only uses gold reference translations. He He 0001, Alvin Grissom II, John Morgan, Jordan L. Boyd-Graber, Hal Daumé III |
EMNLP | 1 |
| 2015 | Crowdsourcing with multi-dimensional trust
He He 0001, John S. Baras |
FUSION | 2 |
| 2015 | Trust-aware optimal crowdsourcing with budget constraintabstractCrowdsourcing has been extensively used for aggregating data from a large pool of workers. In a real crowdsourcing market, each answer obtained from a worker incurs cost. The cost is associated with both the level of trustworthiness of workers and the difficulty of tasks. Typically, access to expert-level (more trustworthy) workers is more expensive than to average crowd and completion of a challenging task is more costly than a click-away question. In this paper, we address the problem of optimal assignment of heterogeneous tasks to workers of varying trust levels with budget constraint. Specifically, we design a trust-aware task allocation algorithm that takes as inputs the estimated trust of workers and pre-set budget, and outputs the optimal assignment of tasks to workers. We derive the bound of total error probability that relates to budget, trustworthiness of crowds, and costs of obtaining labels from crowds naturally. Higher budget, more trustworthy crowds, and less costly jobs result in lower theoretical bound. Our allocation scheme does not depend on the specific design of the trust evaluation component. Therefore, it can be combined with generic trust evaluation algorithms. Our algorithm outperforms state-of-the-art by up to 30% on real data. He He 0001, John S. Baras |
ICC | 2 |
| 2015 | Hands-on Learning to Search for Structured PredictionabstractHal Daumé III, John Langford, Kai-Wei Chang, He He, Sudha Rao. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Tutorial Abstracts. 2015. Hal Daumé III, John Langford 0001, Kai-Wei Chang 0001, He He 0001, Sudha Rao |
HLT-NAACL | 4 |
| 2014 | Don't Until the Final Verb Wait: Reinforcement Learning for Simultaneous Machine TranslationabstractWe introduce a reinforcement learningbased approach to simultaneous machine translation-producing a translation while receiving input wordsbetween languages with drastically different word orders: from verb-final languages (e.g., German) to verb-medial languages (English).In traditional machine translation, a translator must "wait" for source material to appear before translation begins.We remove this bottleneck by predicting the final verb in advance.We use reinforcement learning to learn when to trust predictions about unseen, future portions of the sentence.We also introduce an evaluation metric to measure expeditiousness and quality.We show that our new translation model outperforms batch and monotone translation strategies. Alvin Grissom II, He He 0001, Jordan L. Boyd-Graber, John Morgan, Hal Daumé III |
EMNLP | 2 |
| 2014 | Learning to Search in Branch and Bound Algorithms
He He 0001, Hal Daumé III, Jason Eisner |
NIPS | 1 |
| 2014 | Temporal supervised learning for inferring a dialog policy from example conversationsabstractThis paper tackles the problem of learning a dialog policy from example dialogs - for example, from Wizard-of-Oz style dialogs, where an expert (person) plays the role of the system. Learning in this setting is challenging because dialog is a temporal process in which actions affect the future course of the conversation - i.e., dialog requires planning. Past work solved this problem with either conventional supervised learning or reinforcement learning. Reinforcement learning provides a principled approach to planning, but requires more resources than a fixed corpus of examples, such as a dialog simulator or a reward function. Conventional supervised learning, by contrast, operates directly from example dialogs but does not take proper account of planning. We introduce a new algorithm called Temporal Supervised Learning which learns directly from example dialogs, while also taking proper account of planning. The key idea is to choose the next dialog action to maximize the expected discounted accuracy until the end of the dialog. On a dialog testbed in the calendar domain, in simulation, we show that a dialog manager trained with temporal supervised learning substantially outperforms a baseline trained using conventional supervised learning. Lihong Li 0001, He He 0001, Jason D. Williams |
SLT | 2 |
| 2013 | Dynamic Feature Selection for Dependency ParsingabstractFeature computation and exhaustive search have significantly restricted the speed of graph-based dependency parsing.We propose a faster framework of dynamic feature selection, where features are added sequentially as needed, edges are pruned early, and decisions are made online for each sentence.We model this as a sequential decision-making problem and solve it by imitation learning techniques.We test our method on 7 languages.Our dynamic parser can achieve accuracies comparable or even superior to parsers using a full set of features, while computing fewer than 30% of the feature templates. He He 0001, Hal Daumé III, Jason Eisner |
EMNLP | 1 |
| 2012 | Besting the Quiz Master: Crowdsourcing Incremental Classification Games
Jordan L. Boyd-Graber, Brianna Satinoff, He He 0001, Hal Daumé III |
EMNLP-CoNLL | 3 |
| 2012 | Imitation Learning by CoachingabstractImitation Learning has been shown to be successful in solving many challenging real-world problems. Some recent approaches give strong performance guarantees by training the policy iteratively. However, it is important to note that these guarantees depend on how well the policy we found can imitate the oracle on the training data. When there is a substantial difference between the oracle's ability and the learner's policy space, we may fail to find a policy that has low error on the training set. In such cases, we propose to use a coach that demonstrates easy-to-learn actions for the learner and gradually approaches the oracle. By a reduction of learning by demonstration to online learning, we prove that coaching can yield a lower regret bound than using the oracle. We apply our algorithm to a novel cost-sensitive dynamic feature selection problem, a hard decision problem that considers a user-specified accuracy-cost trade-off. Experimental results on UCI datasets show that our method outperforms state-of-the-art imitation learning methods in dynamic features selection and two static feature selection methods. He He 0001, Hal Daumé III, Jason Eisner |
NIPS | 1 |
| 2011 | Single image super-resolution using Gaussian process regressionabstractIn this paper we address the problem of producing a high-resolution image from a single low-resolution image without any external training set. We propose a framework for both magnification and deblurring using only the original low-resolution image and its blurred version. In our method, each pixel is predicted by its neighbors through the Gaussian process regression. We show that when using a proper covariance function, the Gaussian process regression can perform soft clustering of pixels based on their local structures. We further demonstrate that our algorithm can extract adequate information contained in a single low-resolution image to generate a high-resolution image with sharp edges, which is comparable to or even superior in quality to the performance of other edge-directed and example-based super-resolution algorithms. Experimental results also show that our approach maintains high-quality performance at large magnifications. He He 0001, Wan-Chi Siu |
CVPR | 1 |
| 2010 | Rare Class Classification by Support Vector MachineabstractThe problem of classification on highly imbalanced datasets has been studied extensively in the literature. Most classifiers show significant deterioration in performance when dealing with skewed datasets. In this paper, we first examine the underlying reasons for SVM's deterioration on imbalanced datasets. We then propose two modifications for the soft margin SVM, where we change or add constraints to the optimization problem. The proposed methods are compared with regular SVM, cost-sensitive SVM and two re-sampling methods. Our experimental results demonstrate that this constrained SVM can consistently outperform the other associated methods. He He 0001, Ali Ghodsi 0001 |
ICPR | 1 |