VLDB 2026 Research / reviewers in the wild / expert
Haw-Shiuan Chang
dblp:130/1022
· DBLP profile ↗
19ranked-venue papers
15as first author
9since 2021 · last 2025
0000-0003-4607-936XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 12 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | REAL Sampling: Boosting Factuality and Diversity of Open-ended Generation by Extrapolating the Entropy of an Infinitely Large LMabstractAbstract Decoding methods for large language models (LLMs) usually struggle with the tradeoff between ensuring factuality and maintaining diversity. In this paper, we propose REAL (Residual Entropy from Asymptotic Line) sampling,1 which predicts the step-wise hallucination likelihood of an LLM. When an LLM is likely to hallucinate, REAL lowers the p threshold in nucleus sampling. Otherwise, REAL sampling increases the p threshold to boost the diversity. To predict the step-wise hallucination likelihood without supervision, we construct a THF (Token-level Hallucination Forecasting) model, which predicts the asymptotic entropy (i.e., inherent uncertainty) of the next token by extrapolating the next-token entropies of an infinitely large language model from a series of LLMs with different sizes. If an LLM’s entropy is higher than the asymptotic entropy (i.e., the LLM is more uncertain than it should be), the THF model predicts a high hallucination hazard, which leads to a lower p threshold in REAL sampling. In the FactualityPrompts benchmark (Lee et al., 2022), we demonstrate that REAL sampling based on a 70M THF model can substantially improve the factuality and diversity of 7B LLMs simultaneously. After combined with contrastive decoding, REAL sampling outperforms 13 sampling methods, and generates texts that are more factual than the greedy sampling and more diverse than the nucleus sampling with p = 0.5. Haw-Shiuan Chang, Nanyun Peng 0001, Mohit Bansal, Anil Ramakrishna, Tagyoung Chung |
Trans. Assoc. Comput. Linguistics | 1 |
| 2024 | Explaining and Improving Contrastive Decoding by Extrapolating the Probabilities of a Huge and Hypothetical LMabstractContrastive decoding (CD) (Li et al., 2023) improves the next-token distribution of a large expert language model (LM) using a small amateur LM.Although CD is applied to various LMs and domains to enhance open-ended text generation, it is still unclear why CD often works well, when it could fail, and how we can make it better.To deepen our understanding of CD, we first theoretically prove that CD could be viewed as linearly extrapolating the next-token logits from a huge and hypothetical LM.We also highlight that the linear extrapolation could make CD unable to output the most obvious answers that have already been assigned high probabilities by the amateur LM.To overcome CD's limitation, we propose a new unsupervised decoding method called Asymptotic Probability Decoding (APD). 1 APD explicitly extrapolates the probability curves from the LMs of different sizes to infer the asymptotic probabilities from an infinitely large LM without inducing more inference costs than CD.In FACTUALITYPROMPTS, an open-ended text generation benchmark, sampling using APD significantly boosts factuality in comparison to the CD sampling and its variants, and achieves state-of-the-art results for Pythia 6.9B and OPT 6.7B.Furthermore, in five commonsense QA datasets, APD is often significantly better than CD and achieves a similar effect of using a larger LLM.For example, the perplexity of APD on top of Pythia 6.9B is even lower than the perplexity of Pythia 12B in CommonsenseQA and LAMBADA.* The work was mostly done at Amazon. Haw-Shiuan Chang, Nanyun Peng 0001, Mohit Bansal, Anil Ramakrishna, Tagyoung Chung |
EMNLP | 1 |
| 2024 | To Copy, or not to Copy; That is a Critical Issue of the Output Softmax Layer in Neural Sequential RecommendersabstractRecent studies suggest that the existing neural models have difficulty handling repeated items in sequential recommendation tasks. However, our understanding of this difficulty is still limited. In this study, we substantially advance this field by identifying a major source of the problem: the single hidden state embedding and static item embeddings in the output softmax layer. Specifically, the similarity structure of the global item embedding in the softmax layer sometimes forces the single hidden state embedding to be close to new items when copying is a better choice, while sometimes forcing the hidden state to be close to the items from the input inappropriately. To alleviate the problem, we adapt the recently-proposed softmax alternatives such as softmax-CPR to sequential recommendation tasks and demonstrate that the new softmax architectures unleash the capability of the neural encoder on learning when to copy and when to exclude the items from the input sequence. By only making some simple modifications on the output softmax layer for SASRec and GRU4Rec, softmax-CPR achieves consistent improvement in 12 datasets. With almost the same model size, our best method not only improves the average NDCG@10 of GRU4Rec in 5 datasets with duplicated items by 10% (4%-17% individually) but also improves 7 datasets without duplicated items by 24% (8%-39%)! Haw-Shiuan Chang, Nikhil Agarwal, Andrew McCallum |
WSDM | 1 |
| 2023 | Multi-CLS BERT: An Efficient Alternative to Traditional EnsemblingabstractEnsembling BERT models often significantly improves accuracy, but at the cost of significantly more computation and memory footprint.In this work, we propose Multi-CLS BERT, a novel ensembling method for CLS-based prediction tasks that is almost as efficient as a single BERT model.Multi-CLS BERT uses multiple CLS tokens with a parameterization and objective that encourages their diversity.Thus instead of fine-tuning each BERT model in an ensemble (and running them all at test time), we need only fine-tune our single Multi-CLS BERT model (and run the one model at test time, ensembling just the multiple final CLS embeddings).To test its effectiveness, we build Multi-CLS BERT on top of a state-of-the-art pretraining method for BERT (Aroca-Ouellette and Rudzicz, 2020).In experiments on GLUE and SuperGLUE we show that our Multi-CLS BERT reliably improves both overall accuracy and confidence estimation.When only 100 training samples are available in GLUE, the Multi-CLS BERT Base model can even outperform the corresponding BERT Large model.We analyze the behavior of our Multi-CLS BERT, showing that it has many of the same characteristics and behavior as a typical BERT 5-way ensemble, but with nearly 4-times less computation and memory. Haw-Shiuan Chang, Ruei-Yao Sun, Kathryn Ricci, Andrew McCallum |
ACL (1) | 1 |
| 2022 | Softmax Bottleneck Makes Language Models Unable to Represent Multi-mode Word DistributionsabstractNeural language models (LMs) such as GPT-2 estimate the probability distribution over the next word by a softmax over the vocabulary.The softmax layer produces the distribution based on the dot products of a single hidden state and the embeddings of words in the vocabulary.However, we discover that this single hidden state cannot produce all probability distributions regardless of the LM size or training data size because the single hidden state embedding cannot be close to the embeddings of all the possible next words simultaneously when there are other interfering word embeddings between them.In this work, we demonstrate the importance of this limitation both theoretically and practically.Our work not only deepens our understanding of softmax bottleneck and mixture of softmax (MoS) but also inspires us to propose multi-facet softmax (MFS) to address the limitations of MoS.Extensive empirical analyses confirm our findings and show that against MoS, the proposed MFS achieves two-fold improvements in the perplexity of GPT-2 and BERT."The greater the ambiguity, the greater the pleasure."-Milan Kundera Haw-Shiuan Chang, Andrew McCallum |
ACL (1) | 1 |
| 2021 | Extending Multi-Sense Word Embedding to Phrases and Sentences for Unsupervised Semantic ApplicationsabstractMost unsupervised NLP models represent each word with a single point or single region in semantic space, while the existing multi-sense word embeddings cannot represent longer word sequences like phrases or sentences. We propose a novel embedding method for a text sequence (a phrase or a sentence) where each sequence is represented by a distinct set of multi-mode codebook embeddings to capture different semantic facets of its meaning. The codebook embeddings can be viewed as the cluster centers which summarize the distribution of possibly co-occurring words in a pre-trained word embedding space. We introduce an end-to-end trainable neural model that directly predicts the set of cluster centers from the input text sequence during test time. Our experiments show that the per-sentence codebook embeddings significantly improve the performances in unsupervised sentence similarity and extractive summarization benchmarks. In phrase similarity experiments, we discover that the multi-facet embeddings provide an interpretable semantic representation but do not outperform the single-facet baseline. Haw-Shiuan Chang, Amol Agrawal, Andrew McCallum |
AAAI | 1 |
| 2021 | Changing the Mind of Transformers for Topically-Controllable Language GenerationabstractLarge Transformer-based language models can aid human authors by suggesting plausible continuations of text written so far.However, current interactive writing assistants do not allow authors to guide text generation in desired topical directions.To address this limitation, we design a framework that displays multiple candidate upcoming topics, of which a user can select a subset to guide the generation.Our framework consists of two components: (1) a method that produces a set of candidate topics by predicting the centers of word clusters in the possible continuations, and ( 2) a text generation model whose output adheres to the chosen topics.The training of both components is self-supervised, using only unlabeled text.Our experiments demonstrate that our topic options are better than those of standard clustering approaches, and our framework often generates fluent sentences related to the chosen topics, as judged by automated metrics and crowdsourced workers. Haw-Shiuan Chang, Jiaming Yuan, Mohit Iyyer, Andrew McCallum |
EACL | 1 |
| 2021 | Multi-facet Universal SchemaabstractUniversal schema (USchema) assumes that two sentence patterns that share the same entity pairs are similar to each other.This assumption is widely adopted for solving various types of relation extraction (RE) tasks.Nevertheless, each sentence pattern could contain multiple facets, and not every facet is similar to all the facets of another sentence pattern cooccurring with the same entity pair.To address the violation of the USchema assumption, we propose multi-facet universal schema that uses a neural model to represent each sentence pattern as multiple facet embeddings and encourage one of these facet embeddings to be close to that of another sentence pattern if they cooccur with the same entity pair.In our experiments, we demonstrate that multi-facet embeddings significantly outperform their singlefacet embedding counterpart, compositional universal schema (CUSchema) (Verga et al., 2016), in distantly supervised relation extraction tasks.Moreover, we can also use multiple embeddings to detect the entailment relation between two sentence patterns when no manual label is available. Rohan Paul, Haw-Shiuan Chang, Andrew McCallum |
EACL | 2 |
| 2021 | Open Aspect Target Sentiment Classification with Natural Language PromptsabstractFor many business applications, we often seek to analyze sentiments associated with any arbitrary aspects of commercial products, despite having a very limited amount of labels or even without any labels at all.However, existing aspect target sentiment classification (ATSC) models are not trainable if annotated datasets are not available.Even with labeled data, they fall short of reaching satisfactory performance.To address this, we propose simple approaches that better solve ATSC with natural language prompts, enabling the task under zero-shot cases and enhancing supervised settings, especially for few-shot cases.Under the few-shot setting for SemEval 2014 Task 4 laptop domain, our method of reformulating ATSC as an NLI task outperforms supervised SOTA approaches by up to 24.13 accuracy points and 33.14 macro F1 points.Moreover, we demonstrate that our prompts could handle implicitly stated aspects as well: our models reach about 77% accuracy on detecting sentiments for aspect categories (e.g., food), which do not necessarily appear within the text, even though we trained the models only with explicitly mentioned aspect terms (e.g., fajitas) from just 16 reviews -while the accuracy of the no-prompt baseline is only around 65%. * Equal contribution. Ronald Seoh, Ian Birle, Mrinal Tak, Haw-Shiuan Chang, Brian Pinette, Alfred Hough |
EMNLP (1) | 4 |
| 2020 | AutoKnow: Self-Driving Knowledge Collection for Products of Thousands of TypesabstractCan one build a knowledge graph (KG) for all products in the world? Knowledge graphs have firmly established themselves as valuable sources of information for search and question answering, and it is natural to wonder if a KG can contain information about products offered at online retail sites. There have been several successful examples of generic KGs, but organizing information about products poses many additional challenges, including sparsity and noise of structured data for products, complexity of the domain with millions of product types and thousands of attributes, heterogeneity across large number of categories, as well as large and constantly growing number of products. Xin Dong 0001, Xiang He 0007, Andrey Kan, Yan Liang 0004, Jun Ma 0029, Yifan Ethan Xu, Tong Zhao 0002, Gabriel Blanco Saldana, Saurabh Deshpande, Alexandre Michetti Manduca, Jay Ren, Surender Pal Singh, Fan Xiao 0001, Haw-Shiuan Chang, Giannis Karamanolakis, Yuning Mao, Yaqing Wang 0001, Christos Faloutsos, Andrew McCallum, Jiawei Han 0001 |
KDD | 16 |
| 2020 | Using error decay prediction to overcome practical issues of deep active learning for named entity recognition
Haw-Shiuan Chang, Shankar Vembu, Sunil Mohan, Rheeya Uppaal, Andrew McCallum |
Mach. Learn. | 1 |
| 2018 | Distributional Inclusion Vector Embedding for Unsupervised Hypernymy DetectionabstractHaw-Shiuan Chang, Ziyun Wang, Luke Vilnis, Andrew McCallum. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Haw-Shiuan Chang, Luke Vilnis, Andrew McCallum |
NAACL-HLT | 1 |
| 2018 | Active Learning for Crowdsourced QoE ModelingabstractQuality of experience (QoE) models predict the subjective quality of multimedia based on the relevant quality of service (QoS) factors. Due to the large space of QoS factors and the high costs of conducting subjective tests, efficient sampling strategies are required to determine which QoS configurations are to be queried, that is, evaluated by subjects. In this study, we extend the IQX model proposed by [M. Fiedler, T. Hoßfeld, and P. Tran-Gia, “A generic quantitative relationship between quality of experience and quality of service,”IEEE Netw., vol. 24, no. 2, pp. 36–41, Mar./Apr.2010.] toward a multidimensional QoS–QoE model (MIQX). To explore the complicated interaction between QoS factors more efficiently, we develop active learning algorithms for the multidimensional QoE model. Then, we conduct comprehensive experiments to compare the effectiveness of applying different sampling methods to crowdsourced video quality assessment tasks. In offline experiments that assume annotators give the same scores after changing the querying order, we demonstrate that active learning performs best and that a space-filling algorithm performs significantly better than random sampling. However, when we analyze the performance of the active sampling approaches more deeply using a novel field experiment, we observe that the active learning algorithms, which have been shown to be effective in the offline setting, can fail due to the habituation effect and individual differences of annotators. The active learning methods can also succeed when these issues are mitigated. These findings suggest that simply simulating the sample acquisition order, which is widely adopted in previous active learning literature[2]–[5], is not sufficient for multimedia quality assessment tasks. Haw-Shiuan Chang, Chih-Fan Hsu, Tobias Hoßfeld, Kuan-Ta Chen |
IEEE Trans. Multim. | 1 |
| 2017 | Active Bias: Training More Accurate Neural Networks by Emphasizing High Variance SamplesabstractSelf-paced learning and hard example mining re-weight training instances to improve learning accuracy. This paper presents two improved alternatives based on lightweight estimates of sample uncertainty in stochastic gradient descent (SGD): the variance in predicted probability of the correct class across iterations of mini-batch SGD, and the proximity of the correct class probability to the decision threshold. Extensive experimental results on six datasets show that our methods reliably improve accuracy in various network architectures, including additional gains on top of other popular training techniques, such as residual learning, momentum, ADAM, batch normalization, dropout, and distillation. Haw-Shiuan Chang, Erik G. Learned-Miller, Andrew McCallum |
NIPS | 1 |
| 2015 | Modeling Exercise Relationships in E-Learning: A Unified Approach
Haw-Shiuan Chang, Hwai-Jung Hsu, Kuan-Ta Chen |
EDM | 1 |
| 2015 | Optimizing the decomposition for multiple foreground cosegmentation
Haw-Shiuan Chang, Yu-Chiang Frank Wang |
Comput. Vis. Image Underst. | 1 |
| 2014 | Simple-to-Complex Discriminative Clustering for Hierarchical Image Segmentation
Haw-Shiuan Chang, Yu-Chiang Frank Wang |
ACCV (3) | 1 |
| 2013 | Superpixel-based large displacement optical flowabstractIt has been a challenging task to estimate optical flow for videos in which either foreground or background exhibits remarkable motion information (i.e., large displacement), or those with insufficient resolution due to artifacts like motion blur or noise. We present a novel optical flow algorithm, which approaches the above problem as solving the task of energy minimization, which exploits image data and smoothness terms at the superpixel level. Our proposed method can be considered as an extended mean-shift algorithm, which advances color and gradient information of superpixels across consecutive frames with smoothness guarantees. Since we do not require assumptions of linearlization during optimization (as standard optical flow approaches do), we are able to alleviate local minimum problems and thus produce improved estimation results. Empirical results on the MPI-Sintel video dataset verify the effectiveness of our proposed method. Haw-Shiuan Chang, Yu-Chiang Frank Wang |
ICIP | 1 |
| 2013 | Exploring Visual and Motion Saliency for Automatic Video Object ExtractionabstractThis paper presents a saliency-based video object extraction (VOE) framework. The proposed framework aims to automatically extract foreground objects of interest without any user interaction or the use of any training data (i.e., not limited to any particular type of object). To separate foreground and background regions within and across video frames, the proposed method utilizes visual and motion saliency information extracted from the input video. A conditional random field is applied to effectively combine the saliency induced features, which allows us to deal with unknown pose and scale variations of the foreground object (and its articulated parts). Based on the ability to preserve both spatial continuity and temporal consistency in the proposed VOE framework, experiments on a variety of videos verify that our method is able to produce quantitatively and qualitatively satisfactory VOE results. Wei-Te Li, Haw-Shiuan Chang, Kuo-Chin Lien, Hui-Tang Chang, Yu-Chiang Frank Wang |
IEEE Trans. Image Process. | 2 |