VLDB 2026 Research / reviewers in the wild / expert
I-Hung Hsu
dblp:206/7272
· DBLP profile ↗
14ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue AgentsabstractLarge Language Models (LLMs) have made significant progress in open-ended dialogue, yet their inability to retain and retrieve relevant information from long-term interactions limits their effectiveness in applications requiring sustained personalization. External memory mechanisms have been proposed to address this limitation, enabling LLMs to maintain conversational continuity. However, existing approaches struggle with two key challenges. First, rigid memory granularity fails to capture the natural semantic structure of conversations, leading to fragmented and incomplete representations. Second, fixed retrieval mechanisms cannot adapt to diverse dialogue contexts and user interaction patterns. In this work, we propose Reflective Memory Management (RMM), a novel mechanism for long-term dialogue agents, integrating forward- and backward-looking reflections: (1) Prospective Reflection, which dynamically summarizes interactions across granularities—utterances, turns, and sessions—into a personalized memory bank for effective future retrieval, and (2) Retrospective Reflection, which iteratively refines the retrieval in an online reinforcement learning (RL) manner based on LLMs’ cited evidence. Experiments show that RMM demonstrates consistent improvement across various metrics and benchmarks. For example, RMM shows more than 10% accuracy improvement over the baseline without memory management on the LongMemEval dataset. Zhen Tan 0001, Jun Yan 0001, I-Hung Hsu, Rujun Han, Zifeng Wang 0002, Long T. Le, Yiwen Song, Yanfei Chen, Hamid Palangi, Anand Rajan Iyer, Tianlong Chen 0001, Huan Liu 0001, Chen-Yu Lee, Tomas Pfister |
ACL (1) | 3 |
| 2025 | Magnet: Multi-turn Tool-use Data Synthesis and Distillation via Graph TranslationabstractFan Yin, Zifeng Wang, I-Hung Hsu, Jun Yan, Ke Jiang, Yanfei Chen, Jindong Gu, Long Le, Kai-Wei Chang, Chen-Yu Lee, Hamid Palangi, Tomas Pfister. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Fan Yin, Zifeng Wang 0002, I-Hung Hsu, Jun Yan 0001, Yanfei Chen, Jindong Gu, Long T. Le, Kai-Wei Chang 0001, Chen-Yu Lee, Hamid Palangi, Tomas Pfister |
ACL (1) | 3 |
| 2025 | SNaRe: Domain-aware Data Generation for Low-Resource Event DetectionabstractEvent Detection (ED) -the task of identifying event mentions from natural language text -is critical for enabling reasoning in highly specialized domains such as biomedicine, law, and epidemiology.Data generation has proven to be effective in broadening its utility to wider applications without requiring expensive expert annotations.However, when existing generation approaches are applied to specialized domains, they struggle with label noise, where annotations are incorrect, and domain drift, characterized by a distributional mismatch between generated sentences and the target domain.To address these issues, we introduce SNARE, a domain-aware synthetic data generation framework composed of three components: Scout, Narrator, and Refiner.Scout extracts triggers from unlabeled target domain data and curates a high-quality domain-specific trigger list using corpus-level statistics to mitigate domain drift.Narrator, conditioned on these triggers, generates high-quality domainaligned sentences, and Refiner identifies additional event mentions, ensuring high annotation quality.Experimentation on three diverse domain ED datasets reveals how SNARE outperforms the best baseline, achieving average F1 gains of 3-7% in the zero-shot/few-shot settings and 4-20% F1 improvement for multilingual generation.Analyzing the generated trigger hit rate and human evaluation substantiates SNARE's stronger annotation quality and reduced domain drift.We will release our code at https://github.com/PlusLabNLP/SNaRe. Tanmay Parekh, Lucas Bandarkar, Artin Kim, I-Hung Hsu, Kai-Wei Chang 0001, Nanyun Peng 0001 |
EMNLP | 5 |
| 2024 | Contextual Label Projection for Cross-Lingual Structured PredictionabstractTanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang, Nanyun Peng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Tanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang 0001, Nanyun Peng 0001 |
NAACL-HLT | 2 |
| 2023 | TAGPRIME: A Unified Framework for Relational Structure ExtractionabstractI-Hung Hsu, Kuan-Hao Huang, Shuning Zhang, Wenxin Cheng, Prem Natarajan, Kai-Wei Chang, Nanyun Peng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. I-Hung Hsu, Kuan-Hao Huang, Wenxin Cheng, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001 |
ACL (1) | 1 |
| 2023 | AMPERE: AMR-Aware Prefix for Generation-Based Event Argument Extraction ModelabstractEvent argument extraction (EAE) identifies event arguments and their specific roles for a given event.Recent advancement in generationbased EAE models has shown great performance and generalizability over classificationbased models.However, existing generationbased EAE models mostly focus on problem reformulation and prompt design, without incorporating additional information that has been shown to be effective for classification-based models, such as the abstract meaning representation (AMR) of the input passages.Incorporating such information into generation-based models is challenging due to the heterogeneous nature of the natural language form prevalently used in generation-based models and the structured form of AMRs.In this work, we study strategies to incorporate AMR into generationbased EAE models.We propose AMPERE, which generates AMR-aware prefixes for every layer of the generation model.Thus, the prefix introduces AMR information to the generationbased EAE model and then improves the generation.We also introduce an adjusted copy mechanism to AMPERE to help overcome potential noises brought by the AMR graph.Comprehensive experiments and analyses on ACE2005 and ERE datasets show that AMPERE can get 4% -10% absolute F1 score improvements with reduced training data and it is in general powerful across different training sizes. I-Hung Hsu, Zhiyu Xie 0001, Kuan-Hao Huang, Premkumar Natarajan, Nanyun Peng 0001 |
ACL (1) | 1 |
| 2023 | ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-TranslationabstractKuan-Hao Huang, Varun Iyer, I-Hung Hsu, Anoop Kumar, Kai-Wei Chang, Aram Galstyan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kuan-Hao Huang, Varun Iyer, I-Hung Hsu, Kai-Wei Chang 0001, Aram Galstyan |
ACL (1) | 3 |
| 2023 | GENEVA: Benchmarking Generalizability for Event Argument Extraction with Hundreds of Event Types and Argument RolesabstractRecent works in Event Argument Extraction (EAE) have focused on improving model generalizability to cater to new events and domains.However, standard benchmarking datasets like ACE and ERE cover less than 40 event types and 25 entity-centric argument roles.Limited diversity and coverage hinder these datasets from adequately evaluating the generalizability of EAE models.In this paper, we first contribute by creating a large and diverse EAE ontology.This ontology is created by transforming FrameNet, a comprehensive semantic role labeling (SRL) dataset for EAE, by exploiting the similarity between these two tasks.Then, exhaustive human expert annotations are collected to build the ontology, concluding with 115 events and 220 argument roles, with a significant portion of roles not being entities.We utilize this ontology to further introduce GENEVA, a diverse generalizability benchmarking dataset comprising four test suites, aimed at evaluating models' ability to handle limited data and unseen event type generalization.We benchmark six EAE models from various families.The results show that owing to non-entity argument roles, even the best-performing model can only achieve 39% F1 score, indicating how GENEVA provides new challenges for generalization in EAE.Overall, our large and diverse EAE ontology can aid in creating more comprehensive future resources, while GENEVA is a challenging benchmarking dataset encouraging further research for improving generalizability in EAE. Tanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang 0001, Nanyun Peng 0001 |
ACL (1) | 2 |
| 2022 | Multilingual Generative Language Models for Zero-Shot Cross-Lingual Event Argument ExtractionabstractWe present a study on leveraging multilingual pre-trained generative language models for zero-shot cross-lingual event argument extraction (EAE).By formulating EAE as a language generation task, our method effectively encodes event structures and captures the dependencies between arguments.We design language-agnostic templates to represent the event argument structures, which are compatible with any language, hence facilitating the cross-lingual transfer.Our proposed model finetunes multilingual pre-trained generative language models to generate sentences that fill in the language-agnostic template with arguments extracted from the input passage.The model is trained on source languages and is then directly applied to target languages for event argument extraction.Experiments demonstrate that the proposed model outperforms the current state-of-the-art models on zero-shot cross-lingual EAE.Comprehensive studies and error analyses are presented to better understand the advantages and the current limitations of using generative language models for zero-shot cross-lingual transfer EAE. *The authors contribute equally.Attacker Place Target Attacker Target 接近高级军官的消息灵通人士 说,南斯拉夫 军队 不会离 开军营去干涉 反对派 起义。 Australian commandos , who have been operating deep in Iraq , destroyed a command and control post and killed a number of soldiers. Kuan-Hao Huang, I-Hung Hsu, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001 |
ACL (1) | 2 |
| 2022 | DEGREE: A Data-Efficient Generation-Based Event Extraction ModelabstractI-Hung Hsu, Kuan-Hao Huang, Elizabeth Boschee, Scott Miller, Prem Natarajan, Kai-Wei Chang, Nanyun Peng. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. I-Hung Hsu, Kuan-Hao Huang, Elizabeth Boschee, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001 |
NAACL-HLT | 1 |
| 2021 | ESTER: A Machine Reading Comprehension Dataset for Reasoning about Event Semantic RelationsabstractUnderstanding how events are semantically related to each other is the essence of reading comprehension.Recent event-centric reading comprehension datasets focus mostly on event arguments or temporal relations.While these tasks partially evaluate machines' ability of narrative understanding, human-like reading comprehension requires the capability to process event-based information beyond arguments and temporal reasoning.For example, to understand causality between events, we need to infer motivation or purpose; to establish event hierarchy, we need to understand the composition of events.To facilitate these tasks, we introduce ESTER, a comprehensive machine reading comprehension (MRC) dataset for Event Semantic Relation Reasoning.The dataset leverages natural language queries to reason about the five most common event semantic relations, provides more than 6K questions, and captures 10.1K event relation pairs.Experimental results show that the current SOTA systems achieve 22.1%, 63.3% and 83.5% for token-based exact-match (EM), F 1 and event-based HIT@1 scores, which are all significantly below human performances (36.0%, 79.6%, 100% respectively), highlighting our dataset as a challenging benchmark.1 Rujun Han, I-Hung Hsu, Jiao Sun, Julia Baylon, Qiang Ning, Dan Roth 0001, Nanyun Peng 0001 |
EMNLP (1) | 2 |
| 2019 | Deep Structured Neural Network for Event Temporal Relation ExtractionabstractWe propose a novel deep structured learning framework for event temporal relation extraction.The model consists of 1) a recurrent neural network (RNN) to learn scoring functions for pair-wise relations, and 2) a structured support vector machine (SSVM) to make joint predictions.The neural network automatically learns representations that account for long-term contexts to provide robust features for the structured model, while the SSVM incorporates domain knowledge such as transitive closure of temporal relations as constraints to make better globally consistent decisions.By jointly training the two components, our model combines the benefits of both data-driven learning and knowledge exploitation.Experimental results on three highquality event temporal relation datasets (TCR, MATRES, and TB-Dense) demonstrate that incorporated with pre-trained contextualized embeddings, the proposed model achieves significantly better performances than the stateof-the-art methods on all three datasets.We also provide thorough ablation studies to investigate our model. Rujun Han, I-Hung Hsu, Mu Yang, Aram Galstyan, Ralph M. Weischedel, Nanyun Peng 0001 |
CoNLL | 2 |
| 2019 | NIESR: Nuisance Invariant End-to-End Speech RecognitionabstractDeep neural network models for speech recognition have achieved great success recently, but they can learn incorrect associations between the target and nuisance factors of speech (e.g., speaker identities, background noise, etc.), which can lead to overfitting. While several methods have been proposed to tackle this problem, existing methods incorporate additional information about nuisance factors during training to develop invariant models. However, enumeration of all possible nuisance factors in speech data and the collection of their annotations is difficult and expensive. We present a robust training scheme for end-to-end speech recognition that adopts an unsupervised adversarial invariance induction framework to separate out essential factors for speech-recognition from nuisances without using any supplementary labels besides the transcriptions. Experiments show that the speech recognition model trained with the proposed training scheme achieves relative improvements of 5.48% on WSJ0, 6.16% on CHiME3, and 6.61% on TIMIT dataset over the base model. Additionally, the proposed method achieves a relative improvement of 14.44% on the combined WSJ0+CHiME3 dataset. I-Hung Hsu, Ayush Jaiswal, Premkumar Natarajan |
INTERSPEECH | 1 |
| 2017 | Mitigating the impact of speech recognition errors on chatbot using sequence-to-sequence modelabstractWe apply sequence-to-sequence model to mitigate the impact of speech recognition errors on open domain end-to-end dialog generation. We cast the task as a domain adaptation problem where ASR transcriptions and original texts are in two different domains. In this paper, our proposed model includes two individual encoders for each domain data and make their hidden states similar to ensure the decoder predict the same dialog text. The method demonstrates that the sequence-to-sequence model can learn the ASR transcriptions and original text pair having the same meaning and eliminate the speech recognition errors. Experimental results on Cornell movie dialog dataset demonstrate that the domain adaption system help the spoken dialog system generate more similar responses with the original text answers. Pin-Jung Chen, I-Hung Hsu, Yi Yao Huang, Hung-yi Lee |
ASRU | 2 |