VLDB 2026 Research / reviewers in the wild / expert
Patrick Ng
dblp:92/3908
· DBLP profile ↗
19ranked-venue papers
1as first author
12since 2021 · last 2025
0000-0001-8208-652XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PRACTIQ: A Practical Conversational Text-to-SQL dataset with Ambiguous and Unanswerable QueriesabstractMingwen Dong, Nischal Ashok Kumar, Yiqun Hu, Anuj Chauhan, Chung-Wei Hang, Shuaichen Chang, Lin Pan, Wuwei Lan, Henghui Zhu, Jiarong Jiang, Patrick Ng, Zhiguo Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Mingwen Dong, Nischal Ashok Kumar, Yiqun Hu, Anuj Chauhan, Chung-Wei Hang, Shuaichen Chang, Lin Pan 0003, Wuwei Lan, Henghui Zhu, Jiarong Jiang, Patrick Ng, Zhiguo Wang 0006 |
NAACL (Long Papers) | 11 |
| 2025 | You Only Read Once (YORO): Learning to Internalize Database Knowledge for Text-to-SQLabstractHideo Kobayashi, Wuwei Lan, Peng Shi, Shuaichen Chang, Jiang Guo, Henghui Zhu, Zhiguo Wang, Patrick Ng. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hideo Kobayashi, Wuwei Lan, Peng Shi 0010, Shuaichen Chang, Henghui Zhu, Zhiguo Wang 0006, Patrick Ng |
NAACL (Long Papers) | 8 |
| 2023 | Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source LearningabstractAlexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma, Patrick Ng, Zhiguo Wang, Bonan Min, William Yang Wang, Kathleen McKeown, Vittorio Castelli, Dan Roth, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma 0005, Patrick Ng, Zhiguo Wang 0006, Bonan Min, William Yang Wang, Kathy McKeown, Vittorio Castelli, Dan Roth 0001, Bing Xiang |
ACL (1) | 5 |
| 2023 | Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness
Shuaichen Chang, Jun Wang 0122, Mingwen Dong, Lin Pan 0003, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang 0029, Jiarong Jiang, Joe Lilien, Steve Ash, William Yang Wang, Zhiguo Wang 0006, Vittorio Castelli, Patrick Ng, Bing Xiang |
ICLR | 15 |
| 2023 | DecAF: Joint Decoding of Answers and Logical Forms for Question Answering over Knowledge Bases
Donghan Yu, Sheng Zhang 0029, Patrick Ng, Henghui Zhu, Alexander Hanbo Li, Jun Wang 0122, Yiqun Hu, William Yang Wang, Zhiguo Wang 0006, Bing Xiang |
ICLR | 3 |
| 2022 | Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question AnsweringabstractQuestion answering over semi-structured tables has attracted significant attention in the NLP community. However, most of the existing work focus on questions that can be answered with short-form answer, i.e. the answer is often a table cell or aggregation of multiple cells. This can mismatch with the intents of users who want to ask more complex questions that require free-form answers such as explanations. To bridge the gap, most recently, pre-trained sequence-to-sequence language models such as T5 are used for generating free-form answers based on the question and table inputs. However, these pre-trained language models have weaker encoding abilities over table cells and schema. To mitigate this issue, in this work, we present an intermediate pre-training framework, Generation-focused Table-based Intermediate Pre-training (GENTAP), that jointly learns representations of natural language questions and tables. GENTAP learns to generate via two training objectives to enhance the question understanding and table representation abilities for complex questions. Based on experimental results, models that leverage GENTAP framework outperform the existing baselines on FETAQA benchmark. The pre-trained models are not only useful for free-form question answering, but also for few-shot data-to-text generation task, thus showing good transfer ability by obtaining new state-of-the-art results. Peng Shi 0010, Patrick Ng, Feng Nan, Henghui Zhu, Jun Wang 0122, Jiarong Jiang, Alexander Hanbo Li, Rishav Chakravarti, Donald Weidner, Bing Xiang, Zhiguo Wang 0006 |
AAAI | 2 |
| 2021 | Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-TrainingabstractMost recently, there has been significant interest in learning contextual representations for various NLP tasks, by leveraging large scale text corpora to train powerful language models with self-supervised learning objectives, such as Masked Language Model (MLM). Based on a pilot study, we observe three issues of existing general-purpose language models when they are applied in the text-to-SQL semantic parsers: fail to detect the column mentions in the utterances, to infer the column mentions from the cell values, and to compose target SQL queries when they are complex. To mitigate these issues, we present a model pretraining framework, Generation-Augmented Pre-training (GAP), that jointly learns representations of natural language utterance and table schemas, by leveraging generation models to generate high-quality pre-train data. GAP Model is trained on 2 million utterance-schema pairs and 30K utterance-schema-SQL triples, whose utterances are generated by generation models. Based on experimental results, neural semantic parsers that leverage GAP Model as a representation encoder obtain new state-of-the-art results on both Spider and Criteria-to-SQL benchmarks. Peng Shi 0010, Patrick Ng, Zhiguo Wang 0006, Henghui Zhu, Alexander Hanbo Li, Jun Wang 0122, Cícero Nogueira dos Santos, Bing Xiang |
AAAI | 2 |
| 2021 | Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip PredictionabstractYifan Gao, Henghui Zhu, Patrick Ng, Cicero Nogueira dos Santos, Zhiguo Wang, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yifan Gao 0001, Henghui Zhu, Patrick Ng, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang |
ACL/IJCNLP (1) | 3 |
| 2021 | Dual Reader-Parser on Hybrid Textual and Tabular Evidence for Open Domain Question AnsweringabstractAlexander Hanbo Li, Patrick Ng, Peng Xu, Henghui Zhu, Zhiguo Wang, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Alexander Hanbo Li, Patrick Ng, Henghui Zhu, Zhiguo Wang 0006, Bing Xiang |
ACL/IJCNLP (1) | 2 |
| 2021 | Improving Factual Consistency of Abstractive Summarization via Question AnsweringabstractFeng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Feng Nan, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathy McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang 0006, Andrew O. Arnold, Bing Xiang |
ACL/IJCNLP (1) | 4 |
| 2021 | Retrieval, Re-ranking and Multi-task Learning for Knowledge-Base Question AnsweringabstractQuestion answering over knowledge bases (KBQA) usually involves three sub-tasks, namely topic entity detection, entity linking and relation detection.Due to the large number of entities and relations inside knowledge bases (KB), previous work usually utilized sophisticated rules to narrow down the search space and managed only a subset of KBs in memory.In this work, we leverage a retrieveand-rerank framework to access KBs via traditional information retrieval (IR) method, and re-rank retrieved candidates with more powerful neural networks such as the pre-trained BERT model.Considering the fact that directly assigning a different BERT model for each sub-task may incur prohibitive costs, we propose to share a BERT encoder across all three sub-tasks and define task-specific layers on top of the shared layer.The unified model is then trained under a multi-task learning framework.Experiments show that: (1) Our IRbased retrieval method is able to collect highquality candidates efficiently, thus enables our method adapt to large-scale KBs easily; (2) the BERT model improves the accuracy across all three sub-tasks; and (3) benefiting from multitask learning, the unified model obtains further improvements with only 1/3 of the original parameters.Our final model achieves competitive results on the SimpleQuestions dataset and superior performance on the FreebaseQA dataset. Zhiguo Wang 0006, Patrick Ng, Ramesh Nallapati, Bing Xiang |
EACL | 2 |
| 2021 | Generative Context Pair Selection for Multi-hop Question AnsweringabstractDheeru Dua, Cicero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner, Sameer Singh. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Dheeru Dua, Cícero Nogueira dos Santos, Patrick Ng, Ben Athiwaratkun, Bing Xiang, Matt Gardner 0001, Sameer Singh 0001 |
EMNLP (1) | 3 |
| 2020 | Template-Based Question Generation from Retrieved Sentences for Improved Unsupervised Question AnsweringabstractQuestion Answering (QA) is in increasing demand as the amount of information available online and the desire for quick access to this content grows.A common approach to QA has been to fine-tune a pretrained language model on a task-specific labeled dataset.This paradigm, however, relies on scarce, and costly to obtain, large-scale human-labeled data.We propose an unsupervised approach to training QA models with generated pseudotraining data.We show that generating questions for QA training by applying a simple template on a related, retrieved sentence rather than the original context sentence improves downstream QA performance by allowing the model to learn more complex context-question relationships.Training a QA model on this data gives a relative improvement over a previous unsupervised model in F1 score on the SQuAD dataset by about 14%, and 20% when the answer is a named entity, achieving stateof-the-art performance on SQuAD for unsupervised QA. Alexander R. Fabbri, Patrick Ng, Zhiguo Wang 0006, Ramesh Nallapati, Bing Xiang |
ACL | 2 |
| 2020 | End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering SystemsabstractSiamak Shakeri, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang, Ramesh Nallapati, Bing Xiang. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Siamak Shakeri, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang 0006, Ramesh Nallapati, Bing Xiang |
EMNLP (1) | 4 |
| 2019 | Multi-passage BERT: A Globally Normalized BERT Model for Open-domain Question AnsweringabstractZhiguo Wang, Patrick Ng, Xiaofei Ma, Ramesh Nallapati, Bing Xiang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhiguo Wang 0006, Patrick Ng, Xiaofei Ma 0001, Ramesh Nallapati, Bing Xiang |
EMNLP/IJCNLP (1) | 2 |
| 2013 | Wireless noncontact ECG and EEG biopotential sensorsabstractWearable, unobtrusive and patient friendly physiological sensors will be a key driving force in the wireless health revolution. Cardiac (ECG) and brain (EEG) signals are two important signal modalities indicative of healthy and diseased states of body and mind that directly benefit from long-term monitoring. Despite advancements in wireless and embedded electronics technology, however, ECG/EEG monitoring devices still face problems with patient compliance and comfort from the use wet/gel electrodes. We have developed two wireless biopotential instrumentation systems using noncontact electrodes that can operate without direct skin contact and through thin layers of fabric. The first system is a general purpose replacement for traditional ECG/EEG telemetry systems and the second is a compact, fully self-contained wireless ECG tag. All of the issues relating to the design of low noise, high performance noncontact sensors are discussed along with full technical details, circuit schematics and construction techniques. The noncontact electrode has been integrated into both a wearable ECG chest harness as well an EEG headband and characterized in a battery of experiments that represent potential health applications including resting ECG, exercise ECG and EEG directly against standard clinical adhesive Ag/AgCl electrodes. With careful design and secure mechanical harnesses the noncontact sensor is capable of approaching the quality of conventional electrodes. Yu M. Chi, Patrick Ng, Gert Cauwenberghs |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2008 | GIMSAN: a Gibbs motif finder with significance analysisabstractUNLABELLED: We present GIMSAN (GIbbsMarkov with Significance ANalysis): a novel tool for de novo motif finding. GIMSAN combines GibbsMarkov, our variant of the Gibbs Sampler, described here for the first time, with our recently introduced significance analysis. AVAILABILITY: GIMSAN is currently available as a web application and a stand-alone application on Unix and PBS (Portable Batch System) cluster through links from http://www.cs.cornell.edu/~keich. Patrick Ng, Uri Keich |
Bioinform. | 1 |
| 2007 | Precisiated information retrieval for RSS feedsabstractPurpose Aims to report a novel method of filtering RSS feeds for obtaining more precise and related information without having to browse through all the incoming feeds. Design/methodology/approach Improve relevance ratio of incoming RSS feeds with configurable filtering phrases on feed title and feed page content. More relevant RSS feeds are obtained when additional semantically related synonym filtering phrases are used. Findings Finds that filtering leads to more precise RSS feeds and extending the filtering phrase with synonym semantic can increase the number of relevant feeds by 3‐5 times. Originality/value The system documented here has been found to be able to help RSS feeds subscribers to browse fewer items with higher matching rate. Chris Tseng, Patrick Ng |
Inf. Manag. Comput. Secur. | 2 |
| 2006 | Automatic Template Detection for Structured Web PagesabstractSimilar Web pages of Web sites on the World Wide Web are usually encoded from an underlying structured source, and generated dynamically from a pre-defined template, such as books' information pages in Amazon.com. By giving a set of Web pages from a common Website, it is possible to extract the template by analyzing common patterns between the Web pages. In our work, we developed the CF-EXALG (collaborative finer-EXALG), based on EXALG, to decompose Web pages and finding their common structures. In our system, templates that are used to generate Web pages can be discovered automatically and stored in XML format. Hence, data encoded in Web pages can be easily extracted and the template can be stored for future manipulation. In our preliminary experiments, CF-EXALG has shown to be more accurate and efficient when compared with other similar systems Lawrence Lo, Vincent T. Y. Ng, Patrick Ng, Stephen Chi-fai Chan |
CSCWD | 3 |