Jie Huang 0009

dblp:29/6643-9 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0003-2197-2469ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Large Language Models as Configuration Validators
abstract
Misconfigurations are major causes of software failures. Existing practices rely on developer-written rules or test cases to validate configuration values, which are expensive. Machine learning (ML) for configuration validation is considered a promising direction, but has been facing challenges such as the need of large-scale field data and system-specific models. Recent advances in Large Language Models (LLMs) show promise in addressing some of the long-lasting limitations of ML-based configuration validation. We present the first analysis on the feasibility and effectiveness of using LLMs for configuration validation. We empirically evaluate LLMs as configuration validators by developing a generic LLM-based configuration validation framework, named Ciri. Ciri employs effective prompt engineering with few-shot learning based on both valid configuration and misconfiguration data. Ciri checks outputs from LLMs when producing results, addressing hallucination and nondeterminism of LLMs. We evaluate Ciri's validation effectiveness on eight popular LLMs using configuration data of ten widely deployed open-source systems. Our analysis (1) confirms the potential of using LLMs for configuration validation, (2) explores design space of LLM-based validators like Ciri, and (3) reveals open challenges such as ineffectiveness in detecting certain types of misconfigurations and biases towards popular configuration parameters.
Xinyu Lian, Yinfang Chen, Runxiang Cheng, Jie Huang 0009, Parth Thakkar, Minjia Zhang, Tianyin Xu
ICSE4
2024 Large Language Models Cannot Self-Correct Reasoning Yet
abstract
Large Language Models (LLMs) have emerged as a groundbreaking technology with their unparalleled text generation capabilities across various applications. Nevertheless, concerns persist regarding the accuracy and appropriateness of their generated content. A contemporary methodology, self-correction, has been proposed as a remedy to these issues. Building upon this premise, this paper critically examines the role and efficacy of self-correction within LLMs, shedding light on its true potential and limitations. Central to our investigation is the notion of intrinsic self-correction, whereby an LLM attempts to correct its initial responses based solely on its inherent capabilities, without the crutch of external feedback. In the context of reasoning, our research indicates that LLMs struggle to self-correct their responses without external feedback, and at times, their performance even degrades after self-correction. Drawing from these insights, we offer suggestions for future research and practical applications in this field.
Jie Huang 0009, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, Denny Zhou
ICLR1
2024 Cascade Speculative Drafting for Even Faster LLM Inference
abstract
Introduced to enhance the efficiency of large language model (LLM) inference, speculative decoding operates by having a smaller model generate a draft. A larger target model then reviews this draft to align with its output, and any acceptance by the target model results in a reduction of the number of the target model runs, ultimately improving efficiency. However, the drafting process in speculative decoding includes slow autoregressive generation and allocates equal time to generating tokens, irrespective of their importance. These inefficiencies collectively contribute to the suboptimal performance of speculative decoding. To further improve LLM inference, we introduce Cascade Speculative Drafting (CS Drafting), a speculative execution algorithm that incorporates two types of cascades. The *Vertical Cascade* eliminates autoregressive generation from neural models, while the *Horizontal Cascade* optimizes time allocation in drafting for improved efficiency. Combining both cascades, CS Drafting achieves greater speedup compared to the baselines in our experiments, while preserving the same output distribution as the target model. Our code is publicly available at https://github.com/lfsszd/CS-Drafting.
Ziyi Chen 0003, Xiaocong Yang, Jiacheng Lin, Chenkai Sun, Kevin Chen-Chuan Chang, Jie Huang 0009
NeurIPS6
2024 Long-form factuality in large language models
abstract
Large language models (LLMs) often generate content that contains factual errors when responding to fact-seeking prompts on open-ended topics. To benchmark a model’s long-form factuality in open domains, we first use GPT-4 to generate LongFact, a prompt set comprising thousands of questions spanning 38 topics. We then propose that LLM agents can be used as automated evaluators for long-form factuality through a method which we call Search-Augmented Factuality Evaluator (SAFE). SAFE utilizes an LLM to break down a long-form response into a set of individual facts and to evaluate the accuracy of each fact using a multi-step reasoning process comprising sending search queries to Google Search and determining whether a fact is supported by the search results. Furthermore, we propose extending F1 score as an aggregated metric for long-form factuality. To do so, we balance the percentage of supported facts in a response (precision) with the percentage of provided facts relative to a hyperparameter representing a user’s preferred response length (recall). Empirically, we demonstrate that LLM agents can outperform crowdsourced human annotators—on a set of∼16k individual facts, SAFE agrees with crowdsourced human annotators 72% of the time, and on a random subset of 100 disagreement cases, SAFE wins 76% of the time. At the same time, SAFE is more than 20 times cheaper than human annotators. We also benchmark thirteen language models on LongFact across four model families (Gemini, GPT, Claude, and PaLM-2), finding that larger language models generally achieve better long-form factuality. LongFact, SAFE, and all experimental code are available at https://github.com/google-deepmind/long-form-factuality.
Jerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu, Nathan Hu, Jie Huang 0009, Dustin Tran, Daiyi Peng, Ruibo Liu, Cosmo Du, Quoc V. Le
NeurIPS6
2023 DimonGen: Diversified Generative Commonsense Reasoning for Explaining Concept Relationships
abstract
In this paper, we propose DimonGen, which aims to generate diverse sentences describing concept relationships in various everyday scenarios.To support this, we first create a benchmark dataset for this task by adapting the existing CommonGen dataset.We then propose a two-stage model called MoREE to generate the target sentences.MoREE consists of a mixture of retrievers model that retrieves diverse context sentences related to the given concepts, and a mixture of generators model that generates diverse sentences based on the retrieved contexts.We conduct experiments on the DimonGen task and show that MoREE outperforms strong baselines in terms of both the quality and diversity of the generated sentences.Our results demonstrate that MoREE is able to generate diverse sentences that reflect different relationships between concepts, leading to a comprehensive understanding of concept relationships. 1
Chenzhengyi Liu, Jie Huang 0009, Kerui Zhu, Kevin Chen-Chuan Chang
ACL (1)2
2023 Text Fact Transfer
abstract
Text style transfer is a prominent task that aims to control the style of text without inherently changing its factual content.To cover more text modification applications, such as adapting past news for current events and repurposing educational materials, we propose the task of text fact transfer, which seeks to transfer the factual content of a source text between topics without modifying its style.We find that existing language models struggle with text fact transfer, due to their inability to preserve the specificity and phrasing of the source text, and tendency to hallucinate errors.To address these issues, we design ModQGA, a framework that minimally modifies a source text with a novel combination of end-to-end question generation and specificity-aware question answering.Through experiments on four existing datasets adapted for text fact transfer, we show that ModQGA can accurately transfer factual content without sacrificing the style of the source text. 1
Nishant Balepur, Jie Huang 0009, Kevin Chen-Chuan Chang
EMNLP2
2023 Expository Text Generation: Imitate, Retrieve, Paraphrase
abstract
Expository documents are vital resources for conveying complex information to readers.Despite their usefulness, writing expository text by hand is a challenging process that requires careful content planning, obtaining facts from multiple sources, and the ability to clearly synthesize these facts.To ease these burdens, we propose the task of expository text generation, which seeks to automatically generate an accurate and stylistically consistent expository text for a topic by intelligently searching a knowledge source.We solve our task by developing IRP, a framework that overcomes the limitations of retrieval-augmented models and iteratively performs content planning, fact retrieval, and rephrasing.Through experiments on three diverse, newly-collected datasets, we show that IRP produces factual and organized expository texts that accurately inform readers. 1 1 Code is available at https://github.com/ nbalepur/expository-text-generation. Ground TruthUniversity of Denver is a private institution that was founded in 1864.It has a total undergraduate enrollment of 5,867 (fall 2021), its setting is city, and the campus size is 125 acres... University of Montana is a public institution that was founded in 1893.It has a total undergraduate enrollment of 7,223 (fall 2021), its setting is urban, and the campus size is 220 acres... RAGUniversity of Denver is a public institution founded in 1891.It has a total of 5,867 students (fall 2021), its location is urban, and the campus covers 120 acres... Our Model (IRP)University of Denver is a private institution founded in 1864.It has a total of 5,867 students (fall 2021), it is located in the city, and the campus covers 126 acres... LLaMA+RetrUniversity of Denver is a private institution founded in 1864.It has a total of 11,482 students (fall 2021), its location is urban, and the campus covers 125 acres...
Nishant Balepur, Jie Huang 0009, Kevin Chen-Chuan Chang
EMNLP2
2022 MetaASSIST: Robust Dialogue State Tracking with Meta Learning
abstract
Existing dialogue datasets contain lots of noise in their state annotations.Such noise can hurt model training and ultimately lead to poor generalization performance.A general framework named ASSIST has recently been proposed to train robust dialogue state tracking (DST) models.It introduces an auxiliary model to generate pseudo labels for the noisy training set.These pseudo labels are combined with vanilla labels by a common fixed weighting parameter to train the primary DST model.Notwithstanding the improvements of ASSIST on DST, tuning the weighting parameter is challenging.Moreover, a single parameter shared by all slots and all instances may be suboptimal.To overcome these limitations, we propose a meta learning-based framework MetaASSIST to adaptively learn the weighting parameter.Specifically, we propose three schemes with varying degrees of flexibility, ranging from slot-wise to both slot-wise and instance-wise, to convert the weighting parameter into learnable functions.These functions are trained in a meta-learning manner by taking the validation set as meta data.Experimental results demonstrate that all three schemes can achieve competitive performance.Most impressively, we achieve a state-of-the-art joint goal accuracy of 80.10% on MultiWOZ 2.4.
Fanghua Ye 0001, Xi Wang 0012, Jie Huang 0009, Shenghui Li, Samuel Stern, Emine Yilmaz
EMNLP3
2022 Understanding Jargon: Combining Extraction and Generation for Definition Modeling
abstract
Can machines know what twin prime is?From the composition of this phrase, machines may guess twin prime is a certain kind of prime, but it is still difficult to deduce exactly what twin stands for without additional knowledge.Here, twin prime is a jargon-a specialized term used by experts in a particular field.Explaining jargon is challenging since it usually requires domain knowledge to understand.Recently, there is an increasing interest in extracting and generating definitions of words automatically.However, existing approaches, either extraction or generation, perform poorly on jargon.In this paper, we propose to combine extraction and generation for jargon definition modeling: first extract self-and correlative definitional information of target jargon from the Web and then generate the final definitions by incorporating the extracted definitional information.Our framework is remarkably simple but effective: experiments demonstrate our method can generate high-quality definitions for jargon and outperform state-of-the-art models significantly, e.g., BLEU score from 8.76 to 22.66 and human-annotated score from 2.34 to 4.04. 1
Jie Huang 0009, Hanyin Shao, Kevin Chen-Chuan Chang, Jinjun Xiong, Wen-Mei W. Hwu
EMNLP1
2022 DEER: Descriptive Knowledge Graph for Explaining Entity Relationships
abstract
We propose DEER (Descriptive Knowledge Graph for Explaining Entity Relationships)an open and informative form of modeling entity relationships.In DEER, relationships between entities are represented by free-text relation descriptions.For instance, the relationship between entities of machine learning and algorithm can be represented as "Machine learning explores the study and construction of algorithms that can learn from and make predictions on data."To construct DEER, we propose a self-supervised learning method to extract relation descriptions with the analysis of dependency patterns and generate relation descriptions with a transformer-based relation description synthesizing model, where no human labeling is required.Experiments demonstrate that our system can extract and generate highquality relation descriptions for explaining entity relationships.The results suggest that we can build an open and informative knowledge graph without human annotation.
Jie Huang 0009, Kerui Zhu, Kevin Chen-Chuan Chang, Jinjun Xiong, Wen-Mei W. Hwu
EMNLP1
2022 Coordinated Topic Modeling
abstract
We propose a new problem called coordinated topic modeling that imitates human behavior while describing a text corpus.It considers a set of well-defined topics like the axes of a semantic space with a reference representation.It then uses the axes to model a corpus for easily understandable representation.This new task helps represent a corpus more interpretably by reusing existing knowledge and benefits the corpora comparison task.We design ECTM, an embedding-based coordinated topic model that effectively uses the reference representation to capture the target corpus-specific aspects while maintaining each topic's global semantics.In ECTM, we introduce the topic-and documentlevel supervision with a self-training mechanism to solve the problem.Finally, extensive experiments on multiple domains show the superiority of our model over other baselines.1
Pritom Saha Akash, Jie Huang 0009, Kevin Chen-Chuan Chang
EMNLP2
2021 Measuring Fine-Grained Domain Relevance of Terms: A Hierarchical Core-Fringe Approach
abstract
Jie Huang, Kevin Chang, JinJun Xiong, Wen-mei Hwu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jie Huang 0009, Kevin Chen-Chuan Chang, Jinjun Xiong, Wen-Mei W. Hwu
ACL/IJCNLP (1)1
2020 Exploring Semantic Capacity of Terms
abstract
We introduce and study semantic capacity of terms.For example, the semantic capacity of artificial intelligence is higher than that of linear regression since artificial intelligence possesses a broader meaning scope.Understanding semantic capacity of terms will help many downstream tasks in natural language processing.For this purpose, we propose a two-step model to investigate semantic capacity of terms, which takes a large text corpus as input and can evaluate semantic capacity of terms if the text corpus can provide enough cooccurrence information of terms.Extensive experiments in three fields demonstrate the effectiveness and rationality of our model compared with well-designed baselines and human-level evaluations.
Jie Huang 0009, Zilong Wang 0002, Kevin Chen-Chuan Chang, Wen-Mei W. Hwu, Jinjun Xiong
EMNLP (1)1
2020 Nonuniform Hyper-Network Embedding with Dual Mechanism
abstract
Network embedding which aims to learn the low-dimensional representations for vertices in networks has been extensively studied in recent years. Although there are various models designed for networks with different properties and different structures for different tasks, most of them are only applied to normal networks which only contain pairwise relationships between vertices. In many realistic cases, relationships among objects are not pairwise and such relationships can be better modeled by a hyper-network in which each edge can connect an uncertain number of vertices. In this article, we focus on two properties of hyper-networks: nonuniform and dual property. In order to make full use of these two properties, we firstly propose a flexible model called Hyper2vec to learn the embeddings of hyper-networks by applying a biased second order random walk strategy to hyper-networks in the framework of Skip-gram. Then, we combine the features of hyperedges by considering the dual hyper-networks to build a further model called NHNE based on 1D convolutional neural networks, and train a tuplewise similarity function for the nonuniform relationships in hyper-networks. Extensive experiments demonstrate the significant effectiveness of our methods for hyper-network embedding.
Jie Huang 0009, Chuan Chen 0001, Fanghua Ye 0001, Weibo Hu, Zibin Zheng
ACM Trans. Inf. Syst.1
2019 Hyper-Path-Based Representation Learning for Hyper-Networks
abstract
Network representation learning has aroused widespread interests in recent years. While most of the existing methods deal with edges as pairwise relationships, only a few studies have been proposed for hyper-networks to capture more complicated tuplewise relationships among multiple nodes. A hyper-network is a network where each edge, called hyperedge, connects an arbitrary number of nodes. Different from conventional networks, hyper-networks have certain degrees of indecomposability such that the nodes in a subset of a hyperedge may not possess a strong relationship. That is the main reason why traditional algorithms fail in learning representations in hyper-networks by simply decomposing hyperedges into pairwise relationships. In this paper, we firstly define a metric to depict the degrees of indecomposability for hyper-networks. Then we propose a new concept called hyper-path and design hyper-path-based random walks to preserve the structural information of hyper-networks according to the analysis of the indecomposability. Then a carefully designed algorithm, Hyper-gram, utilizes these random walks to capture both pairwise relationships and tuplewise relationships in the whole hyper-networks. Finally, we conduct extensive experiments on several real-world datasets covering the tasks of link prediction and hyper-network reconstruction, and results demonstrate the rationality, validity, and effectiveness of our methods compared with those existing state-of-the-art models designed for conventional networks or hyper-networks.
Jie Huang 0009, Xin Liu 0039, Yangqiu Song
CIKM1