Alexander Hanbo Li

dblp:169/9998 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
8since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Language models and text generation · 26% Question answering and dialogue systems · 25% Information extraction and text analysis · 15%
Databases, data mining, and information retrieval
1 paper
Data models and query languages · 100%

Topics — the 21 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
1.222023
Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness · ICLR 2023
Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-Training · AAAI 2021
Machine learning › Representation and self-supervised learning
pre-training
1.122022
Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question Answering · AAAI 2022
Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-Training · AAAI 2021
Natural language and speech › Language models and text generation › text generation
data-to-text generation
0.822023
Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning · ACL (1) 2023
Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question Answering · AAAI 2022
Natural language and speech › Language models and text generation › text generation › data-to-text generation
few-shot data-to-text generation
0.712023
Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning · ACL (1) 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › query answering
knowledge base querying
0.712023
DecAF: Joint Decoding of Answers and Logical Forms for Question Answering over Knowledge Bases · ICLR 2023
Natural language and speech › Question answering and dialogue systems
knowledge base question answering
0.712023
DecAF: Joint Decoding of Answers and Logical Forms for Question Answering over Knowledge Bases · ICLR 2023
Machine learning › Trustworthy machine learning
robustness evaluation
0.712023
Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness · ICLR 2023
Data models and query languages › natural language interface
natural language interface to database
0.712023
Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness · ICLR 2023
Natural language and speech › Language models and text generation › pre-trained language model
intermediate pre-training
0.612022
Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question Answering · AAAI 2022
Natural language and speech › Question answering and dialogue systems
open-ended question answering
0.612022
Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question Answering · AAAI 2022
Natural language and speech › Question answering and dialogue systems
table question answering
0.612022
Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question Answering · AAAI 2022
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.512021
Dual Reader-Parser on Hybrid Textual and Tabular Evidence for Open Domain Question Answering · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation › text generation
paraphrase generation
0.512021
Learning to Selectively Learn for Weakly-supervised Paraphrase Generation · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis
semantic parsing
0.512021
Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-Training · AAAI 2021
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.312017
Forest-type Regression with General Losses and Robust Forest · ICML 2017
Machine learning › Kernel, tree and ensemble methods › ensemble learning › tree ensembles
random forest
0.312017
Forest-type Regression with General Losses and Robust Forest · ICML 2017
Machine learning › Trustworthy machine learning
robustness
0.312017
Forest-type Regression with General Losses and Robust Forest · ICML 2017
Machine learning › Learning theory › statistical estimation › robust statistics
robust regression
0.312017
Forest-type Regression with General Losses and Robust Forest · ICML 2017
Machine learning › Transfer learning and domain adaptation
multi-source learning
0.212023
Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning · ACL (1) 2023
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
contextual representation learning
0.112021
Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-Training · AAAI 2021
Natural language and speech › Question answering and dialogue systems
multi-turn dialogue
0.112021
Contextual Rephrase Detection for Reducing Friction in Dialogue Systems · EMNLP (1) 2021

Methods — techniques the papers use, named apart from their topics

diagnostic benchmark construction · 1.3unified representation · 0.7multi-source learning · 0.7few-shot learning · 0.7sequence-to-sequence language model · 0.6generation-based pre-training · 0.6self-supervised learning · 0.5masked language model · 0.5generation model · 0.5BART · 0.5
YearPublicationVenuePosition
2023 Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning
abstract
Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma, Patrick Ng, Zhiguo Wang, Bonan Min, William Yang Wang, Kathleen McKeown, Vittorio Castelli, Dan Roth, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma 0005, Patrick Ng, Zhiguo Wang 0006, Bonan Min, William Yang Wang, Kathy McKeown, Vittorio Castelli, Dan Roth 0001, Bing Xiang
ACL (1)1
2023 Dr.Spider: A Diagnostic Evaluation Benchmark towards Text-to-SQL Robustness
Shuaichen Chang, Jun Wang 0122, Mingwen Dong, Lin Pan 0003, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang 0029, Jiarong Jiang, Joe Lilien, Steve Ash, William Yang Wang, Zhiguo Wang 0006, Vittorio Castelli, Patrick Ng, Bing Xiang
ICLR6
2023 DecAF: Joint Decoding of Answers and Logical Forms for Question Answering over Knowledge Bases
Donghan Yu, Sheng Zhang 0029, Patrick Ng, Henghui Zhu, Alexander Hanbo Li, Jun Wang 0122, Yiqun Hu, William Yang Wang, Zhiguo Wang 0006, Bing Xiang
ICLR5
2022 Generation-Focused Table-Based Intermediate Pre-training for Free-Form Question Answering
abstract
Question answering over semi-structured tables has attracted significant attention in the NLP community. However, most of the existing work focus on questions that can be answered with short-form answer, i.e. the answer is often a table cell or aggregation of multiple cells. This can mismatch with the intents of users who want to ask more complex questions that require free-form answers such as explanations. To bridge the gap, most recently, pre-trained sequence-to-sequence language models such as T5 are used for generating free-form answers based on the question and table inputs. However, these pre-trained language models have weaker encoding abilities over table cells and schema. To mitigate this issue, in this work, we present an intermediate pre-training framework, Generation-focused Table-based Intermediate Pre-training (GENTAP), that jointly learns representations of natural language questions and tables. GENTAP learns to generate via two training objectives to enhance the question understanding and table representation abilities for complex questions. Based on experimental results, models that leverage GENTAP framework outperform the existing baselines on FETAQA benchmark. The pre-trained models are not only useful for free-form question answering, but also for few-shot data-to-text generation task, thus showing good transfer ability by obtaining new state-of-the-art results.
Peng Shi 0010, Patrick Ng, Feng Nan, Henghui Zhu, Jun Wang 0122, Jiarong Jiang, Alexander Hanbo Li, Rishav Chakravarti, Donald Weidner, Bing Xiang, Zhiguo Wang 0006
AAAI7
2021 Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-Training
abstract
Most recently, there has been significant interest in learning contextual representations for various NLP tasks, by leveraging large scale text corpora to train powerful language models with self-supervised learning objectives, such as Masked Language Model (MLM). Based on a pilot study, we observe three issues of existing general-purpose language models when they are applied in the text-to-SQL semantic parsers: fail to detect the column mentions in the utterances, to infer the column mentions from the cell values, and to compose target SQL queries when they are complex. To mitigate these issues, we present a model pretraining framework, Generation-Augmented Pre-training (GAP), that jointly learns representations of natural language utterance and table schemas, by leveraging generation models to generate high-quality pre-train data. GAP Model is trained on 2 million utterance-schema pairs and 30K utterance-schema-SQL triples, whose utterances are generated by generation models. Based on experimental results, neural semantic parsers that leverage GAP Model as a representation encoder obtain new state-of-the-art results on both Spider and Criteria-to-SQL benchmarks.
Peng Shi 0010, Patrick Ng, Zhiguo Wang 0006, Henghui Zhu, Alexander Hanbo Li, Jun Wang 0122, Cícero Nogueira dos Santos, Bing Xiang
AAAI5
2021 Dual Reader-Parser on Hybrid Textual and Tabular Evidence for Open Domain Question Answering
abstract
Alexander Hanbo Li, Patrick Ng, Peng Xu, Henghui Zhu, Zhiguo Wang, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Alexander Hanbo Li, Patrick Ng, Henghui Zhu, Zhiguo Wang 0006, Bing Xiang
ACL/IJCNLP (1)1
2021 Learning to Selectively Learn for Weakly-supervised Paraphrase Generation
abstract
Paraphrase generation is a longstanding NLP task that has diverse applications for downstream NLP tasks.However, the effectiveness of existing efforts predominantly relies on large amounts of golden labeled data.Though unsupervised endeavors have been proposed to address this issue, they may fail to generate meaningful paraphrases due to the lack of supervision signals.In this work, we go beyond the existing paradigms and propose a novel approach to generate high-quality paraphrases with weak supervision data.Specifically, we tackle the weakly-supervised paraphrase generation problem by: (1) obtaining abundant weakly-labeled parallel sentences via retrievalbased pseudo paraphrase expansion; and (2) developing a meta-learning framework to progressively select valuable samples for finetuning a pre-trained language model, i.e., BART, on the sentential paraphrasing task.We demonstrate that our approach achieves significant improvements over existing unsupervised approaches, and is even comparable in performance with supervised state-of-the-arts.
Kaize Ding, Dingcheng Li, Alexander Hanbo Li, Chenlei Guo, Huan Liu 0001
EMNLP (1)3
2021 Contextual Rephrase Detection for Reducing Friction in Dialogue Systems
abstract
For voice assistants like Alexa, Google Assistant and Siri, correctly interpreting users' intentions is of utmost importance.However, users sometimes experience friction with these assistants, caused by errors from different system components or user errors such as slips of the tongue.Users tend to rephrase their query until they get a satisfactory response.Rephrase detection is used to identify the rephrases and has long been treated as a task with pairwise input, which does not fully utilize the contextual information (e.g.users' implicit feedback).To this end, we propose a contextual rephrase detection model ContReph to automatically identify rephrases from multiturn dialogues.We showcase how to leverage the dialogue context and user-agent interaction signals, including user's implicit feedback and the time gap between different turns, which can help significantly outperform the pairwise rephrase detection models.
Zhuoyi Wang, Saurabh Gupta 0008, Dingcheng Li, Alexander Hanbo Li, Chenlei Guo
EMNLP (1)6
2020 Censored Quantile Regression Forest
abstract
Random forests are powerful non-parametric regression method but are severely limited in their usage in the presence of randomly censored observations, and naively applied can exhibit poor predictive performance due to the incurred biases. Based on a local adaptive representation of random forests, we develop its regression adjustment for randomly censored regression quantile models. Regression adjustment is based on a new estimating equation that adapts to censoring and leads to quantile score whenever the data do not exhibit censoring. The proposed procedure named censored quantile regression forest, allows us to estimate quantiles of time-to-event without any parametric modeling assumption. We establish its consistency under mild model specifications. Numerical studies showcase a clear advantage of the proposed procedure.
Alexander Hanbo Li, Jelena Bradic
AISTATS1
2020 Semi-Supervised Learning for Text Classification by Layer Partitioning
abstract
Most recent neural semi-supervised learning (SSL) algorithms rely on adding small perturbation to either the input vectors or their representations. These methods have been successful on computer vision tasks as the images form a continuous manifold, but are not appropriate for discrete input such as sentence. To adapt these methods to text input, we propose to decompose a neural network M into two components F and U so that M = U°F . The layers in F are then frozen and only the layers in U will be updated during most time of the training. In this way, F serves as a feature extractor that maps the input to high-level representation and adds systematical noise using dropout. We can then train U using any state-of-the-art SSL algorithms such as Π-model, temporal ensembling, mean teacher, etc. Furthermore, this gradually unfreezing schedule also prevents a pretrained model from catastrophic forgetting. The experimental results demonstrate that our approach provides improvements when compared to state of the art methods especially on short texts.
Alexander Hanbo Li, Abhinav Sethy
ICASSP1
2017 Forest-type Regression with General Losses and Robust Forest
abstract
This paper introduces a new general framework for forest-type regression which allows the development of robust forest regressors by selecting from a large family of robust loss functions. In particular, when plugged in the squared error and quantile losses, it will recover the classical random forest and quantile random forest. We then use robust loss functions to develop more robust forest-type regression algorithms. In the experiments, we show by simulation and real data that our robust forests are indeed much more insensitive to outliers, and choosing the right number of nearest neighbors can quickly improve the generalization performance of random forest.
Alexander Hanbo Li
ICML1