Kuan-Hao Huang

dblp:24/255 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0002-1634-9237ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Eliminating Position Bias of Language Models: A Mechanistic Approach
abstract
Position bias has proven to be a prevalent issue of modern language models (LMs), where the models prioritize content based on its position within the given context. This bias often leads to unexpected model failures and hurts performance, robustness, and reliability across various applications. A simple mechanistic analysis attributes the position bias to two components employed in nearly all state-of-the-art LMs: causal attention and position embedding. Based on the analyses, we propose to **eliminate** position bias (e.g., different retrieved documents' orders in QA affect performance) with a **training-free zero-shot** approach. Our method changes the causal attention to bidirectional attention between documents and utilizes model attention values to decide the relative orders of documents instead of using the order provided in input prompts, therefore enabling Position-INvariant inferencE (PINE) at the document level. By eliminating position bias, models achieve better performance and reliability in downstream tasks, including LM-as-a-judge, retrieval-augmented QA, molecule generation, and math reasoning. Notably, PINE is especially useful when adapting LMs for evaluating reasoning pairs: it consistently provides $8$ to $10$ percentage points performance gains, making Llama-3-70B-Instruct perform even better than GPT-4-0125-preview and GPT-4o-2024-08-06 on the RewardBench reasoning set.
Ziqi Wang 0003, Hanlin Zhang 0002, Xiner Li, Kuan-Hao Huang, Chi Han, Shuiwang Ji, Sham M. Kakade, Hao Peng 0009, Heng Ji 0001
ICLR4
2025 Contrastive Visual Data Augmentation
abstract
Large multimodal models (LMMs) often struggle to recognize novel concepts, as they rely on pre-trained knowledge and have limited ability to capture subtle visual details. Domain-specific knowledge gaps in training also make them prone to confusing visually similar, commonly misrepresented, or low-resource concepts. To help LMMs better align nuanced visual features with language, improving their ability to recognize and reason about novel or rare concepts, we propose a Contrastive visual Data Augmentation (CoDA) strategy. CoDA extracts key contrastive textual and visual features of target concepts against the known concepts they are misrecognized as, and then uses multimodal generative models to produce targeted synthetic data. Automatic filtering of extracted features and augmented images is implemented to guarantee their quality, as verified by human annotators. We show the effectiveness and efficiency of CoDA on low-resource concept and diverse scene recognition datasets including INaturalist and SUN. We additionally collect NovelSpecies, a benchmark dataset consisting of newly discovered animal species that are guaranteed to be unseen by LMMs. LLaVA-1.6 1-shot updating results on these three datasets show CoDA significantly improves SOTA visual data augmentation strategies by 12.3% (NovelSpecies), 5.1% (SUN), and 6.0% (iNat) absolute gains in accuracy.
Yu Zhou 0030, Mohan Tang, Xiaomeng Jin, Te-Lin Wu, Kuan-Hao Huang, Heng Ji 0001, Kai-Wei Chang 0001, Nanyun Peng 0001
ICML6
2024 MIRACLE: An Online, Explainable Multimodal Interactive Concept Learning System
abstract
We present MIRACLE, a system for online, interpretable visual concept and video action recognition. Through a chat interface, users query the recognition system with an uploaded image or video. For images, MIRACLE returns concept predictions from its structured knowledge base, justifying its predictions with heatmaps and natural language-based attribute detections. For videos, MIRACLE predicts an action and justifies its prediction with time varying entity-entity relations. With its ability to learn new concepts in an online, few-shot manner and its support of dynamic changes to its knowledge base, MIRACLE represents a step forward in interpretable multimodal learning systems.
Ansel Blume, Khanh Duy Nguyen, Zhenhailong Wang, Yangyi Chen, Michal Shlapentokh-Rothman, Xiaomeng Jin, Zhen Zhu 0006, Jiateng Liu, Kuan-Hao Huang, Mankeerat Sidhu, Xuanming Zhang, Vivian Liu, Raunak Sinha, Te-Lin Wu, Abhaysinh Zala, Elias Stengel-Eskin, Da Yin, Utkarsh Mall, Zhou Yu 0005, Kai-Wei Chang 0001, Camille Cobb, Karrie Karahalios, Lydia B. Chilton, Mohit Bansal, Nanyun Peng 0001, Carl Vondrick, Derek Hoiem, Heng Ji 0001
ACM Multimedia10
2024 Contextual Label Projection for Cross-Lingual Structured Prediction
abstract
Tanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang, Nanyun Peng. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Tanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang 0001, Nanyun Peng 0001
NAACL-HLT3
2024 Event Detection from Social Media for Epidemic Prediction
abstract
Tanmay Parekh, Anh Mac, Jiarui Yu, Yuxuan Dong, Syed Shahriar, Bonnie Liu, Eric Yang, Kuan-Hao Huang, Wei Wang, Nanyun Peng, Kai-Wei Chang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Tanmay Parekh, Anh Mac, Jiarui Yu, Syed Shahriar, Bonnie Liu, Eric Yang, Kuan-Hao Huang, Wei Wang 0010, Nanyun Peng 0001, Kai-Wei Chang 0001
NAACL-HLT8
2023 TAGPRIME: A Unified Framework for Relational Structure Extraction
abstract
I-Hung Hsu, Kuan-Hao Huang, Shuning Zhang, Wenxin Cheng, Prem Natarajan, Kai-Wei Chang, Nanyun Peng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
I-Hung Hsu, Kuan-Hao Huang, Wenxin Cheng, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001
ACL (1)2
2023 AMPERE: AMR-Aware Prefix for Generation-Based Event Argument Extraction Model
abstract
Event argument extraction (EAE) identifies event arguments and their specific roles for a given event.Recent advancement in generationbased EAE models has shown great performance and generalizability over classificationbased models.However, existing generationbased EAE models mostly focus on problem reformulation and prompt design, without incorporating additional information that has been shown to be effective for classification-based models, such as the abstract meaning representation (AMR) of the input passages.Incorporating such information into generation-based models is challenging due to the heterogeneous nature of the natural language form prevalently used in generation-based models and the structured form of AMRs.In this work, we study strategies to incorporate AMR into generationbased EAE models.We propose AMPERE, which generates AMR-aware prefixes for every layer of the generation model.Thus, the prefix introduces AMR information to the generationbased EAE model and then improves the generation.We also introduce an adjusted copy mechanism to AMPERE to help overcome potential noises brought by the AMR graph.Comprehensive experiments and analyses on ACE2005 and ERE datasets show that AMPERE can get 4% -10% absolute F1 score improvements with reduced training data and it is in general powerful across different training sizes.
I-Hung Hsu, Zhiyu Xie 0001, Kuan-Hao Huang, Premkumar Natarajan, Nanyun Peng 0001
ACL (1)3
2023 ParaAMR: A Large-Scale Syntactically Diverse Paraphrase Dataset by AMR Back-Translation
abstract
Kuan-Hao Huang, Varun Iyer, I-Hung Hsu, Anoop Kumar, Kai-Wei Chang, Aram Galstyan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Kuan-Hao Huang, Varun Iyer, I-Hung Hsu, Kai-Wei Chang 0001, Aram Galstyan
ACL (1)1
2023 GENEVA: Benchmarking Generalizability for Event Argument Extraction with Hundreds of Event Types and Argument Roles
abstract
Recent works in Event Argument Extraction (EAE) have focused on improving model generalizability to cater to new events and domains.However, standard benchmarking datasets like ACE and ERE cover less than 40 event types and 25 entity-centric argument roles.Limited diversity and coverage hinder these datasets from adequately evaluating the generalizability of EAE models.In this paper, we first contribute by creating a large and diverse EAE ontology.This ontology is created by transforming FrameNet, a comprehensive semantic role labeling (SRL) dataset for EAE, by exploiting the similarity between these two tasks.Then, exhaustive human expert annotations are collected to build the ontology, concluding with 115 events and 220 argument roles, with a significant portion of roles not being entities.We utilize this ontology to further introduce GENEVA, a diverse generalizability benchmarking dataset comprising four test suites, aimed at evaluating models' ability to handle limited data and unseen event type generalization.We benchmark six EAE models from various families.The results show that owing to non-entity argument roles, even the best-performing model can only achieve 39% F1 score, indicating how GENEVA provides new challenges for generalization in EAE.Overall, our large and diverse EAE ontology can aid in creating more comprehensive future resources, while GENEVA is a challenging benchmarking dataset encouraging further research for improving generalizability in EAE.
Tanmay Parekh, I-Hung Hsu, Kuan-Hao Huang, Kai-Wei Chang 0001, Nanyun Peng 0001
ACL (1)3
2022 Multilingual Generative Language Models for Zero-Shot Cross-Lingual Event Argument Extraction
abstract
We present a study on leveraging multilingual pre-trained generative language models for zero-shot cross-lingual event argument extraction (EAE).By formulating EAE as a language generation task, our method effectively encodes event structures and captures the dependencies between arguments.We design language-agnostic templates to represent the event argument structures, which are compatible with any language, hence facilitating the cross-lingual transfer.Our proposed model finetunes multilingual pre-trained generative language models to generate sentences that fill in the language-agnostic template with arguments extracted from the input passage.The model is trained on source languages and is then directly applied to target languages for event argument extraction.Experiments demonstrate that the proposed model outperforms the current state-of-the-art models on zero-shot cross-lingual EAE.Comprehensive studies and error analyses are presented to better understand the advantages and the current limitations of using generative language models for zero-shot cross-lingual transfer EAE. *The authors contribute equally.Attacker Place Target Attacker Target 接近高级军官的消息灵通人士 说,南斯拉夫 军队 不会离 开军营去干涉 反对派 起义。 Australian commandos , who have been operating deep in Iraq , destroyed a command and control post and killed a number of soldiers.
Kuan-Hao Huang, I-Hung Hsu, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001
ACL (1)1
2022 DEGREE: A Data-Efficient Generation-Based Event Extraction Model
abstract
I-Hung Hsu, Kuan-Hao Huang, Elizabeth Boschee, Scott Miller, Prem Natarajan, Kai-Wei Chang, Nanyun Peng. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
I-Hung Hsu, Kuan-Hao Huang, Elizabeth Boschee, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001
NAACL-HLT2
2021 Generating Syntactically Controlled Paraphrases without Using Annotated Parallel Pairs
abstract
Paraphrase generation plays an essential role in natural language process (NLP), and it has many downstream applications.However, training supervised paraphrase models requires many annotated paraphrase pairs, which are usually costly to obtain.On the other hand, the paraphrases generated by existing unsupervised approaches are usually syntactically similar to the source sentences and are limited in diversity.In this paper, we demonstrate that it is possible to generate syntactically various paraphrases without the need for annotated paraphrase pairs.We propose Syntactically controlled Paraphrase Generator (SynPG), an encoder-decoder based model that learns to disentangle the semantics and the syntax of a sentence from a collection of unannotated texts.The disentanglement enables SynPG to control the syntax of output paraphrases by manipulating the embedding in the syntactic space.Extensive experiments using automatic metrics and human evaluation show that SynPG performs better syntactic control than unsupervised baselines, while the quality of the generated paraphrases is competitive.We also demonstrate that the performance of SynPG is competitive or even better than supervised models when the unannotated data is large.Finally, we show that the syntactically controlled paraphrases generated by SynPG can be utilized for data augmentation to improve the robustness of NLP models.
Kuan-Hao Huang, Kai-Wei Chang 0001
EACL1
2021 Improving Zero-Shot Cross-Lingual Transfer Learning via Robust Training
abstract
Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potential for zero-shot cross-lingual transfer.However, these multilingual encoders do not precisely align words and phrases across languages.Especially, learning alignments in the multilingual embedding space usually requires sentence-level or word-level parallel corpora, which are expensive to be obtained for low-resource languages.An alternative is to make the multilingual encoders more robust; when fine-tuning the encoder using downstream task, we train the encoder to tolerate noise in the contextual embedding spaces such that even if the representations of different languages are not aligned well, the model can still achieve good performance on zero-shot cross-lingual transfer.In this work, we propose a learning strategy for training robust models by drawing connections between adversarial examples and the failure cases of zero-shot cross-lingual transfer.We adopt two widely used robust training methods, adversarial training and randomized smoothing, to train the desired robust model.The experimental results demonstrate that robust training improves zero-shot cross-lingual transfer on text classification tasks.The improvement is more significant in the generalized crosslingual transfer setting, where the pair of input sentences belong to two different languages.
Kuan-Hao Huang, Wasi Uddin Ahmad, Nanyun Peng 0001, Kai-Wei Chang 0001
EMNLP (1)1
2021 Disentangling Semantics and Syntax in Sentence Embeddings with Pre-trained Language Models
abstract
Pre-trained language models have achieved huge success on a wide range of NLP tasks.However, contextual representations from pretrained models contain entangled semantic and syntactic information, and therefore cannot be directly used to derive useful semantic sentence embeddings for some tasks.Paraphrase pairs offer an effective way of learning the distinction between semantics and syntax, as they naturally share semantics and often vary in syntax.In this work, we present ParaBART, a semantic sentence embedding model that learns to disentangle semantics and syntax in sentence embeddings obtained by pre-trained language models.ParaBART is trained to perform syntax-guided paraphrasing, based on a source sentence that shares semantics with the target paraphrase, and a parse tree that specifies the target syntax.In this way, ParaBART learns disentangled semantic and syntactic representations from their respective inputs with separate encoders.Experiments in English show that ParaBART outperforms state-of-theart sentence embedding models on unsupervised semantic similarity tasks.Additionally, we show that our approach can effectively remove syntactic information from semantic sentence embeddings, leading to better robustness against syntactic variation on downstream semantic tasks.
James Y. Huang, Kuan-Hao Huang, Kai-Wei Chang 0001
NAACL-HLT2
2020 JECL: Joint Embedding and Cluster Learning for Image-Text Pairs
abstract
We propose JECL, a method for clustering image-caption pairs by training parallel encoders with regularized clustering and alignment objectives, simultaneously learning both representations and cluster assignments. These image-caption pairs arise frequently in high-value applications where structured training data is expensive to produce, but free-text descriptions are common. JECL trains by minimizing the Kullback-Leibler divergence between the distribution of the images and text to that of a combined joint target distribution and optimizing the Jensen-Shannon divergence between the soft cluster assignments of the images and text. Regularizers are also applied to JECL to prevent trivial solutions. Experiments show that JECL outperforms both single-view and multi-view methods on large benchmark image-caption datasets, and is remarkably robust to missing captions and varying data sizes.
Sean T. Yang, Kuan-Hao Huang, Bill Howe
ICPR2
2019 Examining Gender Bias in Languages with Grammatical Gender
abstract
Pei Zhou, Weijia Shi, Jieyu Zhao, Kuan-Hao Huang, Muhao Chen, Ryan Cotterell, Kai-Wei Chang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jieyu Zhao 0001, Kuan-Hao Huang, Muhao Chen 0001, Ryan Cotterell, Kai-Wei Chang 0001
EMNLP/IJCNLP (1)4
2019 Dynamic principal projection for cost-sensitive online multi-label classification
abstract
We study multi-label classification (MLC) with three important real-world issues: online updating, label space dimension reduction (LSDR), and cost-sensitivity. Current MLC algorithms have not been designed to address these three issues simultaneously. In this paper, we propose a novel algorithm, cost-sensitive dynamic principal projection (CS-DPP) that resolves all three issues. The foundation of CS-DPP is an online LSDR framework derived from a leading LSDR algorithm. In particular, CS-DPP is equipped with an efficient online dimension reducer motivated by matrix stochastic gradient, and establishes its theoretical backbone when coupled with a carefully-designed online regression learner. In addition, CS-DPP embeds the cost information into label weights to achieve cost-sensitivity along with theoretical guarantees. Experimental results verify that CS-DPP achieves better practical performance than current MLC algorithms across different evaluation criteria, and demonstrate the importance of resolving the three issues simultaneously.
Hong-Min Chu, Kuan-Hao Huang, Hsuan-Tien Lin
Mach. Learn.2
2018 Cost-Sensitive Reference Pair Encoding for Multi-Label Learning
Yao-Yuan Yang, Kuan-Hao Huang, Chih-Wei Chang, Hsuan-Tien Lin
PAKDD (1)2
2017 Cost-sensitive label embedding for multi-label classification
Kuan-Hao Huang, Hsuan-Tien Lin
Mach. Learn.1
2016 A Novel Uncertainty Sampling Algorithm for Cost-Sensitive Multiclass Active Learning
abstract
Active learning is a setup that allows the learning algorithm to iteratively and strategically query the labels of some instances for reducing human labeling efforts. One fundamental strategy, called uncertainty sampling, measures the uncertainty of each instance when making querying decisions. Traditional active learning algorithms focus on binary or multiclass classification, but few works have studied active learning for cost-sensitive multiclass classification (CSMCC), which allows charging different costs for different types of misclassification errors. The few works are generally based on calculating the uncertainty of each instance by probability estimation, and can suffer from the inaccuracy of the estimation. In this paper, we propose a novel active learning algorithm that relies on a different way of calculating the uncertainty. The algorithm is based on our newly-proposed cost embedding approach (CE) for CSMCC. CE embeds the cost information in the distance measure of a special hidden space with non-metric multidimensional scaling, and deals with both symmetric and asymmetric cost information by our carefully designed mirroring trick. The embedding allows the proposed algorithm, active learning with cost embedding (ALCE), to define a cost-sensitive uncertainty measure from the distance in the hidden space. Extensive experimental results demonstrate that ALCE selects more useful instances by taking the cost information into account through the embedding and is superior to existing cost-sensitive active learning algorithms.
Kuan-Hao Huang, Hsuan-Tien Lin
ICDM1
2016 Linear Upper Confidence Bound Algorithm for Contextual Bandit Problem with Piled Rewards
Kuan-Hao Huang, Hsuan-Tien Lin
PAKDD (2)1
2015 Combination of feature engineering and ranking models for paper-author identification in KDD cup 2013
Chun-Liang Li, Yu-Chuan Su, Ting-Wei Lin, Cheng-Hao Tsai, Wei-Cheng Chang, Kuan-Hao Huang, Tzu-Ming Kuo, Shan-Wei Lin, Young-San Lin, Yu-Chen Lu, Chun-Pai Yang, Cheng-Xia Chang, Wei-Sheng Chin, Yu-Chin Juan, Hsiao-Yu Fish Tung, Jui-Pin Wang, Cheng-Kuang Wei, Felix Wu, Tu-Chun Yin, Tong Yu 0001, Yong Zhuang, Shou-De Lin, Hsuan-Tien Lin, Chih-Jen Lin
J. Mach. Learn. Res.6
2014 Effective string processing and matching for author disambiguation
Wei-Sheng Chin, Yong Zhuang, Yu-Chin Juan, Felix Wu, Hsiao-Yu Fish Tung, Tong Yu 0001, Jui-Pin Wang, Cheng-Xia Chang, Chun-Pai Yang, Wei-Cheng Chang, Kuan-Hao Huang, Tzu-Ming Kuo, Shan-Wei Lin, Young-San Lin, Yu-Chen Lu, Yu-Chuan Su, Cheng-Kuang Wei, Tu-Chun Yin, Chun-Liang Li, Ting-Wei Lin, Cheng-Hao Tsai, Shou-De Lin, Hsuan-Tien Lin, Chih-Jen Lin
J. Mach. Learn. Res.11
2005 The Hard SCORM LMS: Reading SCORM Courseware on Hardcopy Textbooks
abstract
The sharable content object reference model (SCORM) includes a representation of distance learning contents and a behavior definition of how users should interact with the contents. Usually, SCORM-compliant systems are developed based on Web browsers or Java program. We developed a system which allows users to read SCORM-compliant course materials on hardcopy papers while an OCR-like pen device is used as an interaction mechanism. A computer, a PDA, or a cellular phone can be used in conjunction with the pen device for multimedia presentations. Our project is called the hard SCORM. Therefore, users can read textbooks in a traditional manner while behavior of reading is incorporated with the SCORM specification.
Timothy K. Shih, Nigel H. Lin, Wen-Chih Chang, Te-Hua Wang, Hsiau Wen Lin, Hsuan-Pu Chang, Kuan-Hao Huang, Yun-Long Sie, Mon-Ting Tzou, Jin-Tan Yang
ICALT7
2004 Adaptive pocket SCORM reader
abstract
Pocket devices are dominated in size which makes them the perfect platform for mobile learning. Several researches propose how distance education can be realized on pocket devices. We demonstrate the implementation of the adaptive pocket SCORM (scalable content object reference model) reader. Our proposed pocket SCORM reader is able to load SCORM compatible courseware. Furthermore, we introduce our ideal pocket SCORM architecture. With the proposed architecture, we hope to realize SCORM compliant mobile learning.
Timothy K. Shih, Nigel H. Lin, Hsuan-Pu Chang, Kuan-Hao Huang
ICME4