Haoran Li 0007

dblp:50/10038-7 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
8since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 50% Information extraction and text analysis · 19% Question answering and dialogue systems · 12%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › text summarization
dialogue summarization
1.122022
STRUDEL: Structured Dialogue Summarization for Dialogue Comprehension · EMNLP 2022
ConvoSumm: Conversation Summarization Benchmark and Improved Abstractive Summarization with Argument Mining · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.512021
ConvoSumm: Conversation Summarization Benchmark and Improved Abstractive Summarization with Argument Mining · ACL/IJCNLP (1) 2021
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.512021
Syntax-augmented Multilingual BERT for Cross-lingual Transfer · ACL/IJCNLP (1) 2021
Natural language and speech › Language models and text generation
multilingual language models
0.512021
Syntax-augmented Multilingual BERT for Cross-lingual Transfer · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis › multilingual NLP › multilingual language modeling
multilingual pretraining
0.512021
Syntax-augmented Multilingual BERT for Cross-lingual Transfer · ACL/IJCNLP (1) 2021
Natural language and speech › Question answering and dialogue systems › dialogue understanding
conversational semantic parsing
0.412020
Conversational Semantic Parsing · EMNLP (1) 2020
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
cross-lingual representation learning
0.412020
Emerging Cross-lingual Structure in Pretrained Language Models · ACL 2020
Natural language and speech › Language models and text generation
masked language modeling
0.412020
Emerging Cross-lingual Structure in Pretrained Language Models · ACL 2020
Natural language and speech › Question answering and dialogue systems
dialogue understanding
0.212022
STRUDEL: Structured Dialogue Summarization for Dialogue Comprehension · EMNLP 2022
Natural language and speech › Information extraction and text analysis
argument mining
0.112021
ConvoSumm: Conversation Summarization Benchmark and Improved Abstractive Summarization with Argument Mining · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis
named entity recognition
0.112021
Syntax-augmented Multilingual BERT for Cross-lingual Transfer · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis
semantic parsing
0.112020
Conversational Semantic Parsing · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

structured summarization · 0.6universal dependency tree · 0.5auxiliary objective · 0.5argument mining · 0.5semantic parsing · 0.4representation alignment · 0.4
YearPublicationVenuePosition
2022 STRUDEL: Structured Dialogue Summarization for Dialogue Comprehension
abstract
Borui Wang, Chengcheng Feng, Arjun Nair, Madelyn Mao, Jai Desai, Asli Celikyilmaz, Haoran Li, Yashar Mehdad, Dragomir Radev. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Borui Wang, Chengcheng Feng, Arjun Nair, Madelyn Mao, Jai Desai, Asli Celikyilmaz, Haoran Li 0007, Yashar Mehdad, Dragomir R. Radev
EMNLP7
2022 AnswerSumm: A Manually-Curated Dataset and Pipeline for Answer Summarization
abstract
Alexander Fabbri, Xiaojian Wu, Srini Iyer, Haoran Li, Mona Diab. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Alexander R. Fabbri, Xiaojian Wu, Srinivasan Iyer 0001, Haoran Li 0007, Mona T. Diab
NAACL-HLT4
2022 Investigating Crowdsourcing Protocols for Evaluating the Factual Consistency of Summaries
abstract
Xiangru Tang, Alexander Fabbri, Haoran Li, Ziming Mao, Griffin Adams, Borui Wang, Asli Celikyilmaz, Yashar Mehdad, Dragomir Radev. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Xiangru Tang, Alexander R. Fabbri, Haoran Li 0007, Ziming Mao, Griffin Adams, Borui Wang, Asli Celikyilmaz, Yashar Mehdad, Dragomir R. Radev
NAACL-HLT3
2022 CONFIT: Toward Faithful Dialogue Summarization with Linguistically-Informed Contrastive Fine-tuning
abstract
Xiangru Tang, Arjun Nair, Borui Wang, Bingyao Wang, Jai Desai, Aaron Wade, Haoran Li, Asli Celikyilmaz, Yashar Mehdad, Dragomir Radev. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Xiangru Tang, Arjun Nair, Borui Wang, Bingyao Wang, Jai Desai, Aaron Wade, Haoran Li 0007, Asli Celikyilmaz, Yashar Mehdad, Dragomir R. Radev
NAACL-HLT7
2021 Syntax-augmented Multilingual BERT for Cross-lingual Transfer
abstract
In recent years, we have seen a colossal effort in pre-training multilingual text encoders using large-scale corpora in many languages to facilitate cross-lingual transfer learning. However, due to typological differences across languages, the cross-lingual transfer is challenging. Nevertheless, language syntax, e.g., syntactic dependencies, can bridge the typological gap. Previous works have shown that pre-trained multilingual encoders, such as mBERT (CITATION), capture language syntax, helping cross-lingual transfer. This work shows that explicitly providing language syntax and training mBERT using an auxiliary objective to encode the universal dependency tree structure helps cross-lingual transfer. We perform rigorous experiments on four NLP tasks, including text classification, question answering, named entity recognition, and task-oriented semantic parsing. The experiment results show that syntax-augmented mBERT improves cross-lingual transfer on popular benchmarks, such as PAWS-X and MLQA, by 1.4 and 1.6 points on average across all languages. In the generalized transfer setting, the performance boosted significantly, with 3.9 and 3.1 points on average in PAWS-X and MLQA.
Wasi Uddin Ahmad, Haoran Li 0007, Kai-Wei Chang 0001, Yashar Mehdad
ACL/IJCNLP (1)2
2021 ConvoSumm: Conversation Summarization Benchmark and Improved Abstractive Summarization with Argument Mining
abstract
Alexander Fabbri, Faiaz Rahman, Imad Rizvi, Borui Wang, Haoran Li, Yashar Mehdad, Dragomir Radev. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Alexander R. Fabbri, Faiaz Rahman, Imad Rizvi, Borui Wang, Haoran Li 0007, Yashar Mehdad, Dragomir R. Radev
ACL/IJCNLP (1)5
2021 MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark
abstract
Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, Yashar Mehdad. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Haoran Li 0007, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, Yashar Mehdad
EACL1
2021 Improving Zero and Few-Shot Abstractive Summarization with Intermediate Fine-tuning and Data Augmentation
abstract
Alexander Fabbri, Simeng Han, Haoyuan Li, Haoran Li, Marjan Ghazvininejad, Shafiq Joty, Dragomir Radev, Yashar Mehdad. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Alexander R. Fabbri, Simeng Han, Haoran Li 0007, Marjan Ghazvininejad, Shafiq R. Joty, Dragomir R. Radev, Yashar Mehdad
NAACL-HLT4
2020 Emerging Cross-lingual Structure in Pretrained Language Models
abstract
We study the problem of multilingual masked language modeling, i.e. the training of a single model on concatenated text from multiple languages, and present a detailed study of several factors that influence why these models are so effective for cross-lingual transfer.We show, contrary to what was previously hypothesized, that transfer is possible even when there is no shared vocabulary across the monolingual corpora and also when the text comes from very different domains.The only requirement is that there are some shared parameters in the top layers of the multi-lingual encoder.To better understand this result, we also show that representations from monolingual BERT models in different languages can be aligned post-hoc quite effectively, strongly suggesting that, much like for non-contextual word embeddings, there are universal latent symmetries in the learned embedding spaces.For multilingual masked language modeling, these symmetries are automatically discovered and aligned during the joint training process. * Equal contribution. Work done while Shijie was interning at Facebook AI.
Alexis Conneau, Haoran Li 0007, Luke Zettlemoyer, Veselin Stoyanov
ACL3
2020 Conversational Semantic Parsing
abstract
Armen Aghajanyan, Jean Maillard, Akshat Shrivastava, Keith Diedrick, Michael Haeger, Haoran Li, Yashar Mehdad, Veselin Stoyanov, Anuj Kumar, Mike Lewis, Sonal Gupta. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Armen Aghajanyan, Jean Maillard, Akshat Shrivastava, Keith Diedrick, Michael Haeger, Haoran Li 0007, Yashar Mehdad, Veselin Stoyanov, Mike Lewis, Sonal Gupta
EMNLP (1)6