EDBT 2026 Demo / reviewers in the wild / expert
Henghui Zhu
dblp:150/4170
· DBLP profile ↗
22ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0002-4534-6975ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable QueriesabstractMingwen Dong, Nischal Ashok Kumar, Yiqun Hu, Anuj Chauhan, Chung-Wei Hang, Shuaichen Chang, Lin Pan, Wuwei Lan, Henghui Zhu, Jiarong Jiang, Patrick Ng, Zhiguo Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Mingwen Dong, Nischal Ashok Kumar, Yiqun Hu, Anuj Chauhan, Chung-Wei Hang, Shuaichen Chang, Lin Pan 0003, Wuwei Lan, Henghui Zhu, Jiarong Jiang, Patrick Ng, Zhiguo Wang 0006 |
NAACL (Long Papers) | 9 |
| 2025 | You Only Read Once (YORO): Learning to Internalize Database Knowledge for Text-to-SQLabstractHideo Kobayashi, Wuwei Lan, Peng Shi, Shuaichen Chang, Jiang Guo, Henghui Zhu, Zhiguo Wang, Patrick Ng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hideo Kobayashi, Wuwei Lan, Peng Shi 0010, Shuaichen Chang, Henghui Zhu, Zhiguo Wang 0006, Patrick Ng |
NAACL (Long Papers) | 6 |
| 2023 | Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness
Shuaichen Chang, Jun Wang 0122, Mingwen Dong, Lin Pan 0003, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang 0029, Jiarong Jiang, Joe Lilien, Steve Ash, William Yang Wang, Zhiguo Wang 0006, Vittorio Castelli, Patrick Ng, Bing Xiang |
ICLR | 5 |
| 2023 | STREET: A Multi-Task Structured Reasoning and Explanation Benchmark
Danilo Neves Ribeiro, Shen Wang 0005, Xiaofei Ma 0001, Henghui Zhu, Deguang Kong, Juliette Burger, Anjelica Ramos, Zhiheng Huang, William Yang Wang, George Karypis, Bing Xiang, Dan Roth 0001 |
ICLR | 4 |
| 2023 | DecAF: Joint Decoding of Answers and Logical Forms for Question Answering over Knowledge Bases
Donghan Yu, Sheng Zhang 0029, Patrick Ng, Henghui Zhu, Alexander Hanbo Li, Jun Wang 0122, Yiqun Hu, William Yang Wang, Zhiguo Wang 0006, Bing Xiang |
ICLR | 4 |
| 2022 | Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question AnsweringabstractQuestion answering over semi-structured tables has attracted significant attention in the NLP community. However, most of the existing work focus on questions that can be answered with short-form answer, i.e. the answer is often a table cell or aggregation of multiple cells. This can mismatch with the intents of users who want to ask more complex questions that require free-form answers such as explanations. To bridge the gap, most recently, pre-trained sequence-to-sequence language models such as T5 are used for generating free-form answers based on the question and table inputs. However, these pre-trained language models have weaker encoding abilities over table cells and schema. To mitigate this issue, in this work, we present an intermediate pre-training framework, Generation-focused Table-based Intermediate Pre-training (GENTAP), that jointly learns representations of natural language questions and tables. GENTAP learns to generate via two training objectives to enhance the question understanding and table representation abilities for complex questions. Based on experimental results, models that leverage GENTAP framework outperform the existing baselines on FETAQA benchmark. The pre-trained models are not only useful for free-form question answering, but also for few-shot data-to-text generation task, thus showing good transfer ability by obtaining new state-of-the-art results. Peng Shi 0010, Patrick Ng, Feng Nan, Henghui Zhu, Jun Wang 0122, Jiarong Jiang, Alexander Hanbo Li, Rishav Chakravarti, Donald Weidner, Bing Xiang, Zhiguo Wang 0006 |
AAAI | 4 |
| 2022 | Lifelong Pretraining: Continually Adapting Language Models to Emerging CorporaabstractXisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao, Shang-Wen Li, Xiaokai Wei, Andrew Arnold, Xiang Ren. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Xisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao 0001, Shang-Wen Li 0001, Xiaokai Wei, Andrew O. Arnold, Xiang Ren 0001 |
NAACL-HLT | 3 |
| 2021 | Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-TrainingabstractMost recently, there has been significant interest in learning contextual representations for various NLP tasks, by leveraging large scale text corpora to train powerful language models with self-supervised learning objectives, such as Masked Language Model (MLM). Based on a pilot study, we observe three issues of existing general-purpose language models when they are applied in the text-to-SQL semantic parsers: fail to detect the column mentions in the utterances, to infer the column mentions from the cell values, and to compose target SQL queries when they are complex. To mitigate these issues, we present a model pretraining framework, Generation-Augmented Pre-training (GAP), that jointly learns representations of natural language utterance and table schemas, by leveraging generation models to generate high-quality pre-train data. GAP Model is trained on 2 million utterance-schema pairs and 30K utterance-schema-SQL triples, whose utterances are generated by generation models. Based on experimental results, neural semantic parsers that leverage GAP Model as a representation encoder obtain new state-of-the-art results on both Spider and Criteria-to-SQL benchmarks. Peng Shi 0010, Patrick Ng, Zhiguo Wang 0006, Henghui Zhu, Alexander Hanbo Li, Jun Wang 0122, Cícero Nogueira dos Santos, Bing Xiang |
AAAI | 4 |
| 2021 | Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip PredictionabstractYifan Gao, Henghui Zhu, Patrick Ng, Cicero Nogueira dos Santos, Zhiguo Wang, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yifan Gao 0001, Henghui Zhu, Patrick Ng, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang |
ACL/IJCNLP (1) | 2 |
| 2021 | Dual Reader-Parser on Hybrid Textual and Tabular Evidence for Open Domain Question AnsweringabstractAlexander Hanbo Li, Patrick Ng, Peng Xu, Henghui Zhu, Zhiguo Wang, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Alexander Hanbo Li, Patrick Ng, Henghui Zhu, Zhiguo Wang 0006, Bing Xiang |
ACL/IJCNLP (1) | 4 |
| 2021 | Improving Factual Consistency of Abstractive Summarization via Question AnsweringabstractFeng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Feng Nan, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathy McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang 0006, Andrew O. Arnold, Bing Xiang |
ACL/IJCNLP (1) | 3 |
| 2021 | Zero-shot Generalization in Dialog State Tracking through Generative Question AnsweringabstractShuyang Li, Jin Cao, Mukund Sridhar, Henghui Zhu, Shang-Wen Li, Wael Hamza, Julian McAuley. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Jin Cao 0003, Mukund Sridhar, Henghui Zhu, Shang-Wen Li 0001, Wael Hamza, Julian J. McAuley |
EACL | 4 |
| 2021 | Entity-level Factual Consistency of Abstractive Text SummarizationabstractFeng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, Bing Xiang. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Feng Nan, Ramesh Nallapati, Zhiguo Wang 0006, Cícero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathy McKeown, Bing Xiang |
EACL | 5 |
| 2021 | Pairwise Supervised Contrastive Learning of Sentence RepresentationsabstractMany recent successes in sentence representation learning have been achieved by simply fine-tuning on the Natural Language Inference (NLI) datasets with triplet loss or siamese loss.Nevertheless, they share a common weakness: sentences in a contradiction pair are not necessarily from different semantic categories.Therefore, optimizing the semantic entailment and contradiction reasoning objective alone is inadequate to capture the high-level semantic structure.The drawback is compounded by the fact that the vanilla siamese or triplet losses only learn from individual sentence pairs or triplets, which often suffer from bad local optima.In this paper, we propose PairSupCon, an instance discrimination based approach aiming to bridge semantic entailment and contradiction understanding with high-level categorical concept encoding.We evaluate PairSupCon on various downstream tasks that involve understanding sentence semantics at different granularities.We outperform the previous state-of-theart method with 10%-13% averaged improvement on eight clustering tasks, and 5%-6% averaged improvement on seven semantic textual similarity (STS) tasks. Dejiao Zhang, Shang-Wen Li 0001, Wei Xiao 0001, Henghui Zhu, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang |
EMNLP (1) | 4 |
| 2021 | Supporting Clustering with Contrastive LearningabstractDejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li, Henghui Zhu, Kathleen McKeown, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Dejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li 0001, Henghui Zhu, Kathy McKeown, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang |
NAACL-HLT | 5 |
| 2020 | Who Did They Respond to? Conversation Structure Modeling Using Masked Hierarchical TransformerabstractConversation structure is useful for both understanding the nature of conversation dynamics and for providing features for many downstream applications such as summarization of conversations. In this work, we define the problem of conversation structure modeling as identifying the parent utterance(s) to which each utterance in the conversation responds to. Previous work usually took a pair of utterances to decide whether one utterance is the parent of the other. We believe the entire ancestral history is a very important information source to make accurate prediction. Therefore, we design a novel masking mechanism to guide the ancestor flow, and leverage the transformer model to aggregate all ancestors to predict parent utterances. Our experiments are performed on the Reddit dataset (Zhang, Culbertson, and Paritosh 2017) and the Ubuntu IRC dataset (Kummerfeld et al. 2019). In addition, we also report experiments on a new larger corpus from the Reddit platform and release this dataset. We show that the proposed model, that takes into account the ancestral history of the conversation, significantly outperforms several strong baselines including the BERT model on all datasets. Henghui Zhu, Feng Nan, Zhiguo Wang 0006, Ramesh Nallapati, Bing Xiang |
AAAI | 1 |
| 2020 | Enhancing Clinical BERT Embedding using a Biomedical Knowledge BaseabstractDomain knowledge is important for building Natural Language Processing (NLP) systems for low-resource settings, such as in the clinical domain.In this paper, a novel joint training method is introduced for adding knowledge base information from the Unified Medical Language System (UMLS) into language model pre-training for some clinical domain corpus.We show that in three different downstream clinical NLP tasks, our pre-trained language model outperforms the corresponding model with no knowledge base information and other state-of-the-art models.Specifically, in a natural language inference task applied to clinical texts, our knowledge base pre-training approach improves accuracy by up to 1.7%, whereas in clinical name entity recognition tasks, the F1-score improves by up to 1.0%.The pre-trained models are available at https://github.com/noc-lab/clinical-kb-bert. Boran Hao, Henghui Zhu, Ioannis Paschalidis |
COLING | 2 |
| 2020 | End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering SystemsabstractSiamak Shakeri, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang, Ramesh Nallapati, Bing Xiang. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Siamak Shakeri, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang 0006, Ramesh Nallapati, Bing Xiang |
EMNLP (1) | 3 |
| 2020 | Learning from animals: How to Navigate Complex TerrainsabstractWe develop a method to learn a bio-inspired motion control policy using data collected from hawkmoths navigating in a virtual forest. A Markov Decision Process (MDP) framework is introduced to model the dynamics of moths and sparse logistic regression is used to learn control policy parameters from the data. The results show that moths do not favor detailed obstacle location information in navigation, but rely heavily on optical flow. Using the policy learned from the moth data as a starting point, we propose an actor-critic learning algorithm to refine policy parameters and obtain a policy that can be used by an autonomous aerial vehicle operating in a cluttered environment. Compared with the moths' policy, the policy we obtain integrates both obstacle location and optical flow. We compare the performance of these two policies in terms of their ability to navigate in artificial forest areas. While the optimized policy can adjust its parameters to outperform the moth's policy in each different terrain, the moth's policy exhibits a high level of robustness across terrains. Henghui Zhu, Hao Liu 0023, Armin Ataei-Esfahani, Yonatan Munk, Thomas Daniel, Ioannis Paschalidis |
PLoS Comput. Biol. | 1 |
| 2018 | Neural circuits for learning context-dependent associations of stimuli
Henghui Zhu, Ioannis Paschalidis, Michael E. Hasselmo |
Neural Networks | 1 |
| 2015 | Cooperative Design of Networked Observers for Stabilizing LTI PlantsabstractWith the rapid development of sensor networks in the last decade, the cooperative design for networked observers has received an increasing attention from engineering community. This paper aims at developing a unified framework for cooperative design of networked observers to stabilize LTI plants. Apart from the traditional centralized design of MIMO system, the proposed cooperative design approach only utilizes the local information of each sensor. For undirected networks, this paper obtains a sufficient and necessary condition for the existence of the parameters that lead to the stabilization of the LTI plant. In particular, we give the detailed design procedures for the parameters of networked observers, including feedback gains and the coupling strength. The numerical simulation is also given to validate the proposed theoretical results. Henghui Zhu, Jinhu Lü 0001, Maciej Ogorzalek |
ISCAS | 2 |
| 2014 | On the cooperative observability of a continuous-time linear system on an undirected networkabstractIn traditional control theory, a single observer has access all the measured outputs of the plant to estimates its asymptotically. In many real world engineering systems, it may be difficult to build a single observer that has access to all the measured outputs. One way around this difficulty is to build a network of cooperative observers, each of which obtains a portion of the measurement outputs, that collectively produce an asymptotic estimate of the plant state. In this paper, we construct a network of such observers for a continuous-time linear system. Assuming that these observers are connected through an undirected connected network, we establish a necessary and sufficient condition on the plant parameters under which the network of observers will achieve asymptotic omniscience. A network of cooperative observers is said to achieve asymptotic omniscience if their states all converge to the plant state asymptotically. Numerical simulation results are presented to validate theoretical results. The design of cooperative observers sheds some light on the solution of some other real-world problems, such as the design of networked location-based services and sensor networks. Henghui Zhu, Jinhu Lü 0001, Zongli Lin, Yao Chen 0003 |
IJCNN | 1 |