VLDB 2026 Research / reviewers in the wild / expert
Hao Peng 0009
dblp:69/7742-9
· DBLP profile ↗
30ranked-venue papers
8as first author
22since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 8 first-author · 22 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Little Goes a Long Way: Efficient Long Context Training and Inference with Partial ContextsabstractTraining and serving long-context large language models (LLMs) incurs substantial overhead.
To address this, two critical steps are often required: a pretrained LLM typically undergoes a separate stage for context length extension by training on long-context data, followed by architectural modifications to reduce the overhead of KV cache during serving.
This paper argues that integrating length extension with a GPU-friendly KV cache reduction architecture not only reduces training overhead during length extension, but also achieves better long-context performance.
This leads to our proposed LongGen, which finetunes a pretrained LLM into an efficient architecture during length extension.
LongGen builds on three key insights:
(1) Sparse attention patterns, such as window attention (attending to recent tokens), attention sink (initial ones), and blockwise sparse attention (strided token blocks) are well-suited for building efficient long-context models, primarily due to their GPU-friendly memory access patterns, enabling efficiency gains not just theoretically but in practice as well.
(2) It is essential for the model to have direct access to all tokens.
A hybrid architecture with 1/3 full attention layers and 2/3 efficient ones achieves a balanced trade-off between efficiency and long-context performance.
(3) Lightweight training on 5B long-context data is sufficient to extend the hybrid model's context length from 4K to 128K.
We evaluate LongGen on both Llama-2 7B and Llama-2 70B, demonstrating its effectiveness across different scales.
During training with 128K-long contexts, LongGen achieves 1.55x training speedup and reduces wall-clock time by 36%, compared to a full-attention baseline.
During inference, LongGen reduces KV cache memory by 62%, achieving 1.67x prefilling speedup and 1.41x decoding speedup.
Compared to baselines that apply KV-cache reduction techniques to full-attention long-context LLMs, LongGen achieves substantially stronger performance not only on the Needle-in-a-Haystack retrieval task, but also on more challenging long-context reasoning tasks, including BABILong and RULER. Suyu Ge, Xihui Lin, Yunan Zhang 0001, Jiawei Han 0001, Hao Peng 0009 |
ICLR | 5 |
| 2025 | Scaling Diffusion Language Models via Adaptation from Autoregressive ModelsabstractDiffusion Language Models (DLMs) have emerged as a promising new paradigm for text generative modeling, potentially addressing limitations of autoregressive (AR) models. However, current DLMs have been studied at a smaller scale compared to their AR counterparts and lack fair comparison on language modeling benchmarks. Additionally, training diffusion models from scratch at scale remains challenging. Given the prevalence of open-source AR language models, we propose adapting these models to build text diffusion models. We demonstrate connections between AR and diffusion modeling objectives and introduce a simple continual pre-training approach for training diffusion models. Through systematic evaluation on language modeling, reasoning, and commonsense benchmarks, we show that we can convert AR models ranging from 127M to 7B parameters (GPT2 and LLaMA) into diffusion models DiffuGPT and DiffuLLaMA, using less than 200B tokens for training. Our experimental results reveal that these models outperform earlier DLMs and are competitive with their AR counterparts. We release a suite of DLMs (127M-355M-7B) capable of generating fluent text, performing in-context learning, filling in the middle without prompt re-ordering, and following instructions. Shansan Gong, Shivam Agarwal, Yizhe Zhang 0002, Jiacheng Ye, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han 0001, Hao Peng 0009, Lingpeng Kong |
ICLR | 11 |
| 2025 | Eliminating Position Bias of Language Models: A Mechanistic ApproachabstractPosition bias has proven to be a prevalent issue of modern language models (LMs), where the models prioritize content based on its position within the given context. This bias often leads to unexpected model failures and hurts performance, robustness, and reliability across various applications. A simple mechanistic analysis attributes the position bias to two components employed in nearly all state-of-the-art LMs: causal attention and position embedding. Based on the analyses, we propose to **eliminate** position bias (e.g., different retrieved documents' orders in QA affect performance) with a **training-free zero-shot** approach. Our method changes the causal attention to bidirectional attention between documents and utilizes model attention values to decide the relative orders of documents instead of using the order provided in input prompts, therefore enabling Position-INvariant inferencE (PINE) at the document level. By eliminating position bias, models achieve better performance and reliability in downstream tasks, including LM-as-a-judge, retrieval-augmented QA, molecule generation, and math reasoning. Notably, PINE is especially useful when adapting LMs for evaluating reasoning pairs: it consistently provides $8$ to $10$ percentage points performance gains, making Llama-3-70B-Instruct perform even better than GPT-4-0125-preview and GPT-4o-2024-08-06 on the RewardBench reasoning set. Ziqi Wang 0003, Hanlin Zhang 0002, Xiner Li, Kuan-Hao Huang, Chi Han, Shuiwang Ji, Sham M. Kakade, Hao Peng 0009, Heng Ji 0001 |
ICLR | 8 |
| 2025 | The Unreasonable Effectiveness of Entropy Minimization in LLM ReasoningabstractEntropy minimization (EM) trains the model to concentrate even more probability mass on its most confident outputs.
We show that this simple objective alone, without any labeled data, can substantially improve large language models’ (LLMs) performance on challenging math, physics, and coding tasks. We explore three approaches: (1) EM-FT minimizes token-level entropy similarly to instruction finetuning, but on unlabeled outputs drawn from the model; (2) EM-RL: reinforcement learning with negative entropy as the only reward to maximize; (3) EM-INF: inference-time logit adjustment to reduce entropy without any training data or parameter updates.
On Qwen-7B, EM-RL, without any labeled data, achieves comparable or better performance than strong RL baselines such as GRPO and RLOO that are trained on 60K labeled examples. Furthermore, EM-INF enables Qwen-32B to match or exceed the performance of proprietary models like GPT-4o, Claude 3 Opus, and Gemini 1.5 Pro on the challenging SciCode benchmark, while being 3x more efficient than self-consistency and sequential refinement. Our findings reveal that many pretrained LLMs possess previously underappreciated reasoning capabilities that can be effectively elicited through entropy minimization alone, without any labeled data or even any parameter updates. Shivam Agarwal, Zimin Zhang, Lifan Yuan, Jiawei Han 0001, Hao Peng 0009 |
NeurIPS | 5 |
| 2025 | Reinforcement Learning Finetunes Small Subnetworks in Large Language ModelsabstractReinforcement learning (RL) yields substantial improvements in large language models’ (LLMs) downstream task performance and alignment with human values. Surprisingly, such large gains result from updating only a small subnetwork comprising just 5%-30% of the parameters, with the rest effectively unchanged. We refer to this phenomenon as parameter update sparsity induced by RL. It is observed across all 7 widely-used RL algorithms (e.g., PPO, GRPO, DPO) and all 10 LLMs from different families in our experiments.
This sparsity is intrinsic and occurs without any explicit sparsity-promoting regularizations or architectural constraints. Finetuning the subnetwork alone recovers the test accuracy, and, remarkably, produces a model nearly identical to the one obtained via full finetuning.
The subnetworks from different random seeds, training data, and even RL algorithms show substantially greater overlap than expected by chance. Our analysis suggests that this sparsity is not due to updating only a subset of layers; instead, nearly all parameter matrices receive similarly sparse updates. Moreover, the updates to almost all parameter matrices are nearly full-rank,
suggesting RL updates a small subset of parameters that nevertheless span almost the full subspaces that the parameter matrices can represent. We conjecture that the this update sparsity can be primarily attributed to training on data that is near the policy distribution;
techniques that encourage the policy to remain close to the pretrained model, such as the KL regularization and gradient clipping, have limited impact. Sagnik Mukherjee, Lifan Yuan, Dilek Hakkani-Tür, Hao Peng 0009 |
NeurIPS | 4 |
| 2024 | ActionIE: Action Extraction from Scientific Literature with Programming LanguagesabstractXianrui Zhong, Yufeng Du, Siru Ouyang, Ming Zhong, Tingfeng Luo, Qirong Ho, Hao Peng, Heng Ji, Jiawei Han. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xianrui Zhong, Yufeng Du, Siru Ouyang, Ming Zhong 0005, Tingfeng Luo, Qirong Ho, Hao Peng 0009, Heng Ji 0001, Jiawei Han 0001 |
ACL (1) | 7 |
| 2024 | MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackabstractTo solve complex tasks, large language models (LLMs) often require multiple rounds of interactions with the user, sometimes assisted by external tools.
However, current evaluation protocols often emphasize benchmark performance with single-turn exchanges, neglecting the nuanced interactions among the user, LLMs, and external tools, while also underestimating the importance of natural language feedback from users. These oversights contribute to discrepancies between research benchmark evaluations and real-world use cases.
We introduce MINT, a benchmark that evaluates LLMs' ability to solve tasks with multi-turn interactions by (1) using tools and (2) leveraging natural language feedback.
To ensure reproducibility, we provide an evaluation framework where LLMs can access tools by executing Python code and receive users' natural language feedback simulated by GPT-4.
We repurpose a diverse set of established evaluation datasets focusing on reasoning, coding, and decision-making and carefully curate them into a compact subset for efficient evaluation.
Our analysis of 20 open- and closed-source LLMs offers intriguing findings.
(a) LLMs generally benefit from tools and language feedback, with performance gains (absolute, same below) of 1--8% for each turn of tool use and 2--17% with natural language feedback.
(b) Better single-turn performance does not guarantee better multi-turn performance.
(c) Surprisingly, on the LLMs evaluated, supervised instruction-finetuning (SIFT) and reinforcement learning from human feedback (RLHF) generally hurt multi-turn capabilities.
We expect MINT can help measure progress and incentivize research in improving LLMs' capabilities in multi-turn interactions, especially for open-source communities where multi-turn human evaluation can be less accessible compared to commercial LLMs with a larger user base. Xingyao Wang 0002, Zihan Wang 0010, Jiateng Liu, Yangyi Chen, Lifan Yuan, Hao Peng 0009, Heng Ji 0001 |
ICLR | 6 |
| 2024 | TRAM: Bridging Trust Regions and Sharpness Aware MinimizationabstractSharpness-aware minimization (SAM) reports improving domain generalization by
reducing the loss surface curvature in the parameter space. However,
generalization during _fine-tuning_ is often more dependent on the
transferability of _representations_ in the function space. Trust-region
methods (TR) target this goal by regularizing representation curvature to reduce
catastrophic forgetting of pre-trained task-agnostic information while adopting
task-specific skills. We consider unifying these strategies for low curvature in
both parameter space and function space to improve out-of-domain (OOD)
generalization. We propose **Trust Region Aware Minimization** (TRAM), a
SAM algorithm fine-tuning for low parameter sharpness and smooth, informative
representations preserving pre-trained structure. TRAM uses a trust region bound
to inform the SAM adversarial neighborhood, introducing an awareness of function
curvature within optimization for flatter minima. We empirically validate TRAM
in vision (cross-dataset adaptation) and text (OOD language modeling, zero-shot
cross-lingual transfer) tasks where robust domain transfer and representation
generality are critical. TRAM outperforms SAM- and TR-based optimization across
all tasks, notably surpassing competing methods for hard transfer between
_anticorrelated_ domains. TRAM establishes a novel standard in
fine-tuning for domain-generalizable models with minimal additional computation
over previous sharpness-aware methods. Tom Sherborne, Naomi Saphra, Pradeep Dasigi, Hao Peng 0009 |
ICLR | 4 |
| 2024 | CRAFT: Customizing LLMs by Creating and Retrieving from Specialized ToolsetsabstractLarge language models (LLMs) are often augmented with tools to solve complex tasks. By generating code snippets and executing them through task-specific Application Programming Interfaces (APIs), they can offload certain functions to dedicated external modules, such as image encoding and performing calculations. However, most existing approaches to augment LLMs with tools are constrained
by general-purpose APIs and lack the flexibility for tailoring them to specific tasks. In this work, we present CRAFT, a general tool creation and retrieval framework for LLMs. It creates toolsets specifically curated for the tasks and equips LLMs with a component that retrieves tools from these sets to enhance their capability to solve complex tasks. For each task, we collect specific code solutions by prompting
GPT-4 to solve the training examples. Following a validation step ensuring the correctness, these solutions are abstracted into code snippets to enhance reusability, and deduplicated for higher quality. At inference time, the language model retrieves snippets from the toolsets and then executes them or generates the output conditioning on the retrieved snippets. Our method is designed to be flexible and
offers a plug-and-play approach to adapt off-the-shelf LLMs to unseen domains and modalities, without any finetuning. Experiments on vision-language, tabular processing, and mathematical reasoning tasks show that our approach achieves substantial improvements compared to strong baselines. In addition, our in-depth analysis reveals that: (1) consistent performance improvement can be achieved by
scaling up the number of tools and the capability of the backbone models; (2) each component of our approach contributes to the performance gains; (3) the created tools are well-structured and reliable with low complexity and atomicity. Lifan Yuan, Yangyi Chen, Xingyao Wang 0002, Yi R. Fung 0001, Hao Peng 0009, Heng Ji 0001 |
ICLR | 5 |
| 2024 | Executable Code Actions Elicit Better LLM AgentsabstractLarge Language Model (LLM) agents, capable of performing a broad range of actions, such as invoking tools and controlling robots, show great potential in tackling real-world challenges. LLM agents are typically prompted to produce actions by generating JSON or text in a pre-defined format, which is usually limited by constrained action space (e.g., the scope of pre-defined tools) and restricted flexibility (e.g., inability to compose multiple tools). This work proposes to use executable Python code to consolidate LLM agents’ actions into a unified action space (CodeAct). Integrated with a Python interpreter, CodeAct can execute code actions and dynamically revise prior actions or emit new actions upon new observations through multi-turn interactions. Our extensive analysis of 17 LLMs on API-Bank and a newly curated benchmark shows that CodeAct outperforms widely used alternatives (up to 20% higher success rate). The encouraging performance of CodeAct motivates us to build an open-source LLM agent that interacts with environments by executing interpretable code and collaborates with users using natural language. To this end, we collect an instruction-tuning dataset CodeActInstruct that consists of 7k multi-turn interactions using CodeAct. We show that it can be used with existing data to improve models in agent-oriented tasks without compromising their general capability. CodeActAgent, finetuned from Llama2 and Mistral, is integrated with Python interpreter and uniquely tailored to perform sophisticated tasks (e.g., model training) using existing libraries and autonomously self-debug. Xingyao Wang 0002, Yangyi Chen, Lifan Yuan, Yizhe Zhang 0002, Yunzhu Li, Hao Peng 0009, Heng Ji 0001 |
ICML | 6 |
| 2024 | Language Models Hallucinate, but May Excel at Fact VerificationabstractJian Guan, Jesse Dodge, David Wadden, Minlie Huang, Hao Peng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jian Guan 0002, Jesse Dodge, Dave Wadden, Minlie Huang, Hao Peng 0009 |
NAACL-HLT | 5 |
| 2024 | LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language ModelsabstractChi Han, Qifan Wang, Hao Peng, Wenhan Xiong, Yu Chen, Heng Ji, Sinong Wang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chi Han, Qifan Wang 0001, Hao Peng 0009, Wenhan Xiong, Yu Chen 0022, Heng Ji 0001, Sinong Wang |
NAACL-HLT | 3 |
| 2024 | SciCode: A Research Coding Benchmark Curated by ScientistsabstractSince language models (LMs) now outperform average humans on many challenging tasks, it is becoming increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this by examining LM capabilities to generate code for solving real scientific research problems. Incorporating input from scientists and AI researchers in 16 diverse natural science sub-fields, including mathematics, physics, chemistry, biology, and materials science, we create a scientist-curated coding benchmark, SciCode. The problems naturally factorize into multiple subproblems, each involving knowledge recall, reasoning, and code synthesis. In total, SciCode contains 338 subproblems decomposed from 80 challenging main problems, and it offers optional descriptions specifying useful scientific background information and scientist-annotated gold-standard solutions and test cases for evaluation. OpenAI o1-preview, the best-performing model among those tested, can solve only 7.7\% of the problems in the most realistic setting. We believe that SciCode demonstrates both contemporary LMs' progress towards realizing helpful scientific assistants and sheds light on the building and evaluation of scientific AI in the future. Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Shengyan Liu, Yutao Ma, Kha Trinh, Zihan Wang 0010, Bohao Wu, Shengzhu Yin, Minhui Zhu, Kilian Lieret, Yanxin Lu, Genglin Liu, Yufeng Du, Tianhua Tao, Ofir Press, Jamie Callan, Eliu A. Huerta, Hao Peng 0009 |
NeurIPS | 29 |
| 2023 | Transparency Helps Reveal When Language Models Learn MeaningabstractAbstract Many current NLP systems are built from language models trained to optimize unsupervised objectives on large amounts of raw text. Under what conditions might such a procedure acquire meaning? Our systematic experiments with synthetic data reveal that, with languages where all expressions have context-independent denotations (i.e., languages with strong transparency), both autoregressive and masked language models successfully learn to emulate semantic relations between expressions. However, when denotations are changed to be context-dependent with the language otherwise unmodified, this ability degrades. Turning to natural language, our experiments with a specific phenomenon—referential opacity—add to the growing body of evidence that current language models do not represent natural language semantics well. We show this failure relates to the context-dependent nature of natural language form-meaning mappings. Zhaofeng Wu, William Merrill, Hao Peng 0009, Iz Beltagy, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | ABC: Attention with Bounded-memory ControlabstractHao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, Noah Smith. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Hao Peng 0009, Jungo Kasai, Nikolaos Pappas 0002, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz 0001, Noah A. Smith |
ACL (1) | 1 |
| 2022 | Tailor: Generating and Perturbing Text with Semantic ControlsabstractControlled text perturbation is useful for evaluating and improving model generalizability.However, current techniques rely on training a model for every target perturbation, which is expensive and hard to generalize.We present Tailor, a semantically-controlled text generation system.Tailor builds on a pretrained seq2seq model and produces textual outputs conditioned on control codes derived from semantic representations.We craft a set of operations to modify the control codes, which in turn steer generation towards targeted attributes.These operations can be further composed into higher-level ones, allowing for flexible perturbation strategies.We demonstrate the effectiveness of these perturbations in multiple applications.First, we use Tailor to automatically create high-quality contrast sets for four distinct natural language processing (NLP) tasks.These contrast sets contain fewer spurious artifacts and are complementary to manually annotated ones in their lexical diversity.Second, we show that Tailor perturbations can improve model generalization through data augmentation.Perturbing just ∼2% of training data leads to a 5.8-point gain on an NLI challenge set measuring reliance on syntactic heuristics. Alexis Ross, Sherry Tongshuang Wu, Hao Peng 0009, Matthew E. Peters, Matt Gardner 0001 |
ACL (1) | 3 |
| 2022 | Twist Decoding: Diverse Generators Guide Each OtherabstractJungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Hao Peng, Ximing Lu, Dragomir Radev, Yejin Choi, Noah A. Smith. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras 0001, Hao Peng 0009, Ximing Lu, Dragomir R. Radev, Yejin Choi 0001, Noah A. Smith |
EMNLP | 4 |
| 2021 | Finetuning Pretrained Transformers into RNNsabstractJungo Kasai, Hao Peng, Yizhe Zhang, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas, Yi Mao, Weizhu Chen, Noah A. Smith. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Jungo Kasai, Hao Peng 0009, Yizhe Zhang 0002, Dani Yogatama, Gabriel Ilharco, Nikolaos Pappas 0002, Weizhu Chen, Noah A. Smith |
EMNLP (1) | 2 |
| 2021 | Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine Translation
Jungo Kasai, Nikolaos Pappas 0002, Hao Peng 0009, James Cross 0003, Noah A. Smith |
ICLR | 3 |
| 2021 | Random Feature Attention
Hao Peng 0009, Nikolaos Pappas 0002, Dani Yogatama, Roy Schwartz 0001, Noah A. Smith, Lingpeng Kong |
ICLR | 1 |
| 2021 | Contextualized Perturbation for Textual Adversarial AttackabstractDianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, Bill Dolan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Dianqi Li, Yizhe Zhang 0002, Hao Peng 0009, Liqun Chen 0001, Chris Brockett, Ming-Ting Sun, William B. Dolan |
NAACL-HLT | 3 |
| 2021 | Infusing Finetuning with Semantic DependenciesabstractAbstract For natural language processing systems, two kinds of evidence support the use of text representations from neural language models “pretrained” on large unannotated corpora: performance on application-inspired benchmarks (Peters et al., 2018, inter alia), and the emergence of syntactic abstractions in those representations (Tenney et al., 2019, inter alia). On the other hand, the lack of grounded supervision calls into question how well these representations can ever capture meaning (Bender and Koller, 2020). We apply novel probes to recent language models— specifically focusing on predicate-argument structure as operationalized by semantic dependencies (Ivanova et al., 2012)—and find that, unlike syntax, semantics is not brought to the surface by today’s pretrained models. We then use convolutional graph encoders to explicitly incorporate semantic parses into task-specific finetuning, yielding benefits to natural language understanding (NLU) tasks in the GLUE benchmark. This approach demonstrates the potential for general-purpose (rather than task-specific) linguistic supervision, above and beyond conventional pretraining and finetuning. Several diagnostics help to localize the benefits of our approach.1 Zhaofeng Wu, Hao Peng 0009, Noah A. Smith |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | A Mixture of h - 1 Heads is Better than h HeadsabstractMulti-head attentive neural architectures have achieved state-of-the-art results on a variety of natural language processing tasks.Evidence has shown that they are overparameterized; attention heads can be pruned without significant performance loss.In this work, we instead "reallocate" them-the model learns to activate different heads on different inputs.Drawing connections between multi-head attention and mixture of experts, we propose the mixture of attentive experts model (MAE).MAE is trained using a block coordinate descent algorithm that alternates between updating (1) the responsibilities of the experts and (2) their parameters.Experiments on machine translation and language modeling show that MAE outperforms strong baselines on both tasks.Particularly, on the WMT14 English to German translation dataset, MAE improves over "transformer-base" by 0.8 BLEU, with a comparable number of parameters.Our analysis shows that our model learns to specialize different experts to different inputs. 1 Hao Peng 0009, Roy Schwartz 0001, Dianqi Li, Noah A. Smith |
ACL | 1 |
| 2019 | RNN Architecture Learning with Sparse RegularizationabstractJesse Dodge, Roy Schwartz, Hao Peng, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jesse Dodge, Roy Schwartz 0001, Hao Peng 0009, Noah A. Smith |
EMNLP/IJCNLP (1) | 3 |
| 2019 | PaLM: A Hybrid Parser and Language ModelabstractHao Peng, Roy Schwartz, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Hao Peng 0009, Roy Schwartz 0001, Noah A. Smith |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Backpropagating through Structured Argmax using a SPIGOTabstractWe introduce the structured projection of intermediate gradients optimization technique (SPIGOT), a new method for backpropagating through neural networks that include hard-decision structured predictions (e.g., parsing) in intermediate layers.SPIGOT requires no marginal inference, unlike structured attention networks (Kim et al., 2017) and some reinforcement learning-inspired solutions (Yogatama et al., 2017).Like socalled straight-through estimators (Hinton, 2012), SPIGOT defines gradient-like quantities associated with intermediate nondifferentiable operations, allowing backpropagation before and after them; SPIGOT's proxy aims to ensure that, after a parameter update, the intermediate structure will remain well-formed.We experiment on two structured NLP pipelines: syntactic-then-semantic dependency parsing, and semantic parsing followed by sentiment classification.We show that training with SPIGOT leads to a larger improvement on the downstream task than a modularly-trained pipeline, the straight-through estimator, and structured attention, reaching a new state of the art on semantic dependency parsing. Hao Peng 0009, Sam Thomson, Noah A. Smith |
ACL (1) | 1 |
| 2018 | Rational RecurrencesabstractDespite the tremendous empirical success of neural models in natural language processing, many of them lack the strong intuitions that accompany classical machine learning approaches.Recently, connections have been shown between convolutional neural networks (CNNs) and weighted finite state automata (WFSAs), leading to new interpretations and insights.In this work, we show that some recurrent neural networks also share this connection to WFSAs.We characterize this connection formally, defining rational recurrences to be recurrent hidden state update functions that can be written as the Forward calculation of a finite set of WFSAs.We show that several recent neural models use rational recurrences.Our analysis provides a fresh view of these models and facilitates devising new neural architectures that draw inspiration from WFSAs.We present one such model, which performs better than two recent baselines on language modeling and text classification.Our results demonstrate that transferring intuitions from classical models like WFSAs can be an effective approach to designing and understanding neural models. Hao Peng 0009, Roy Schwartz 0001, Sam Thomson, Noah A. Smith |
EMNLP | 1 |
| 2018 | Learning Joint Semantic Parsers from Disjoint DataabstractHao Peng, Sam Thomson, Swabha Swayamdipta, Noah A. Smith. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Hao Peng 0009, Sam Thomson, Swabha Swayamdipta, Noah A. Smith |
NAACL-HLT | 1 |
| 2018 | "You are no Jack Kennedy": On Media Selection of Highlights from Presidential DebatesabstractPolitical speeches and debates play an important role in shaping the images of politicians, and the public often relies on media outlets to select bits of political communication from a large pool of utterances. It is an important research question to understand what factors impact this selection process. To quantitatively explore the selection process, we build a three- decade dataset of presidential debate transcripts and post-debate coverage. We first examine the effect of wording and propose a binary classification framework that controls for both the speaker and the debate situation. We find that crowdworkers can only achieve an accuracy of 60% in this task, indicating that media choices are not entirely obvious. Our classifiers outperform crowdworkers on average, mainly in primary debates. We also compare important factors from crowdworkers» free-form explanations with those from data-driven methods and find interesting differences. Few crowdworkers mentioned that "context matters", whereas our data show that well-quoted sentences are more distinct from the previous utterance by the same speaker than less-quoted sentences. Finally, we examine the aggregate effect of media preferences towards different wordings to understand the extent of fragmentation among media outlets. By analyzing a bipartite graph built from quoting behavior in our data, we observe a decreasing trend in bipartisan coverage. Chenhao Tan, Hao Peng 0009, Noah A. Smith |
WWW | 2 |
| 2017 | Deep Multitask Learning for Semantic Dependency ParsingabstractWe present a deep neural architecture that parses sentences into three semantic dependency graph formalisms.By using efficient, nearly arc-factored inference and a bidirectional-LSTM composed with a multi-layer perceptron, our base system is able to significantly improve the state of the art for semantic dependency parsing, without using hand-engineered features or syntax.We then explore two multitask learning approaches-one that shares parameters across formalisms, and one that uses higher-order structures to predict the graphs jointly.We find that both approaches improve performance across formalisms on average, achieving a new state of the art.Our code is open-source and available at https://github.com/Noahs-ARK/NeurboParser. Hao Peng 0009, Sam Thomson, Noah A. Smith |
ACL (1) | 1 |