Zewei Chu

dblp:169/6786 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 34% Question answering and dialogue systems · 15% Knowledge representation and reasoning · 13%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Software engineering, system software, and programming languages
1 paper
Programming languages and type systems · 44% Software maintenance and evolution · 44% Compilers and program optimization · 13%

Topics — the 7 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
text summarization
0.612022
SummScreen: A Dataset for Abstractive Screenplay Summarization · ACL (1) 2022
Information retrieval
evaluation
0.612022
SummScreen: A Dataset for Abstractive Screenplay Summarization · ACL (1) 2022
Information retrieval › text summarization
summarization evaluation
0.612022
SummScreen: A Dataset for Abstractive Screenplay Summarization · ACL (1) 2022
Natural language and speech › Question answering and dialogue systems
question rewriting
0.412020
How to Ask Better Questions? A Large-Scale Multi-Domain Dataset for Rewriting Ill-Formed Questions · AAAI 2020
Natural language and speech › Machine translation › machine translation evaluation
discourse-aware evaluation
0.412019
Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations · EMNLP/IJCNLP (1) 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
entity representation
0.412019
EntEval: A Holistic Evaluation Benchmark for Entity Representations · EMNLP/IJCNLP (1) 2019
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding
0.412019
Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations · EMNLP/IJCNLP (1) 2019

Methods — techniques the papers use, named apart from their topics

neural summarization · 1.1nearest neighbor retrieval · 1.1sequence-to-sequence neural model · 0.4BLEU-4 · 0.4representation learning · 0.4java annotations · 0.2code generation · 0.2
YearPublicationVenuePosition
2022 SummScreen: A Dataset for Abstractive Screenplay Summarization
abstract
We introduce SUMMSCREEN, a summarization dataset comprised of pairs of TV series transcripts and human written recaps.The dataset provides a challenging testbed for abstractive summarization for several reasons.Plot details are often expressed indirectly in character dialogues and may be scattered across the entirety of the transcript.These details must be found and integrated to form the succinct plot descriptions in the recaps.Also, TV scripts contain content that does not directly pertain to the central plot but rather serves to develop characters or provide comic relief.This information is rarely contained in recaps.Since characters are fundamental to TV series, we also propose two entity-centric evaluation metrics.Empirically, we characterize the dataset by evaluating several methods, including neural models and those based on nearest neighbors.An oracle extractive approach outperforms all benchmarked models according to automatic metrics, showing that the neural models are unable to fully exploit the input transcripts.Human evaluation and qualitative analysis reveal that our nonoracle models are competitive with their oracle counterparts in terms of generating faithful plot events and can benefit from better content selectors.Both oracle and non-oracle models generate unfaithful facts, suggesting future research directions.
Mingda Chen, Zewei Chu, Sam Wiseman, Kevin Gimpel
ACL (1)2
2020 How to Ask Better Questions? A Large-Scale Multi-Domain Dataset for Rewriting Ill-Formed Questions
abstract
We present a large-scale dataset for the task of rewriting an ill-formed natural language question to a well-formed one. Our multi-domain question rewriting (MQR) dataset is constructed from human contributed Stack Exchange question edit histories. The dataset contains 427,719 question pairs which come from 303 domains. We provide human annotations for a subset of the dataset as a quality estimate. When moving from ill-formed to well-formed questions, the question quality improves by an average of 45 points across three aspects. We train sequence-to-sequence neural models on the constructed dataset and obtain an improvement of 13.2% in BLEU-4 over baseline methods built from other data resources. We release the MQR dataset to encourage research on the problem of question rewriting.1
Zewei Chu, Mingda Chen, Miaosen Wang, Kevin Gimpel, Manaal Faruqui, Xiance Si
AAAI1
2019 EntEval: A Holistic Evaluation Benchmark for Entity Representations
abstract
Mingda Chen, Zewei Chu, Yang Chen, Karl Stratos, Kevin Gimpel. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mingda Chen, Zewei Chu, Karl Stratos, Kevin Gimpel
EMNLP/IJCNLP (1)2
2019 Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations
abstract
Mingda Chen, Zewei Chu, Kevin Gimpel. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mingda Chen, Zewei Chu, Kevin Gimpel
EMNLP/IJCNLP (1)2
2015 Scrap your boilerplate with object algebras
abstract
Traversing complex Abstract Syntax Trees (ASTs) typically requires large amounts of tedious boilerplate code. For many operations most of the code simply walks the structure, and only a small portion of the code implements the functionality that motivated the traversal in the first place. This paper presents a type-safe Java framework called Shy that removes much of this boilerplate code. In Shy object algebras are used to describe complex and extensible AST structures. Using Java annotations Shy generates generic boilerplate code for various types of traversals. For a concrete traversal, users of Shy can then inherit from the generated code and override only the interesting cases. Consequently, the amount of code that users need to write is significantly smaller. Moreover, traversals using the Shy framework are also much more structure shy, becoming more adaptive to future changes or extensions to the AST structure. To prove the effectiveness of the approach, we applied Shy in the implementation of a domain-specific questionnaire language. Our results show that for a large number of traversals there was a significant reduction in the amount of user-defined code.
Zewei Chu, Bruno C. d. S. Oliveira, Tijs van der Storm
OOPSLA2