EDBT 2026 Demo / reviewers in the wild / expert
Zewei Chu
dblp:169/6786
· DBLP profile ↗
5ranked-venue papers
1as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 34% Question answering and dialogue systems · 15% Knowledge representation and reasoning · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Programming languages and type systems · 44% Software maintenance and evolution · 44% Compilers and program optimization · 13% |
Topics — the 7 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
text summarization |
0.6 | 1 | 2022 | SummScreen: A Dataset for Abstractive Screenplay Summarization · ACL (1) 2022 |
Information retrieval
evaluation |
0.6 | 1 | 2022 | SummScreen: A Dataset for Abstractive Screenplay Summarization · ACL (1) 2022 |
Information retrieval › text summarization
summarization evaluation |
0.6 | 1 | 2022 | SummScreen: A Dataset for Abstractive Screenplay Summarization · ACL (1) 2022 |
Natural language and speech › Question answering and dialogue systems
question rewriting |
0.4 | 1 | 2020 | How to Ask Better Questions? A Large-Scale Multi-Domain Dataset for Rewriting Ill-Formed Questions · AAAI 2020 |
Natural language and speech › Machine translation › machine translation evaluation
discourse-aware evaluation |
0.4 | 1 | 2019 | Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations · EMNLP/IJCNLP (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
entity representation |
0.4 | 1 | 2019 | EntEval: A Holistic Evaluation Benchmark for Entity Representations · EMNLP/IJCNLP (1) 2019 |
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
0.4 | 1 | 2019 | Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations · EMNLP/IJCNLP (1) 2019 |
Methods — techniques the papers use, named apart from their topics
neural summarization · 1.1nearest neighbor retrieval · 1.1sequence-to-sequence neural model · 0.4BLEU-4 · 0.4representation learning · 0.4java annotations · 0.2code generation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | SummScreen: A Dataset for Abstractive Screenplay SummarizationabstractWe introduce SUMMSCREEN, a summarization dataset comprised of pairs of TV series transcripts and human written recaps.The dataset provides a challenging testbed for abstractive summarization for several reasons.Plot details are often expressed indirectly in character dialogues and may be scattered across the entirety of the transcript.These details must be found and integrated to form the succinct plot descriptions in the recaps.Also, TV scripts contain content that does not directly pertain to the central plot but rather serves to develop characters or provide comic relief.This information is rarely contained in recaps.Since characters are fundamental to TV series, we also propose two entity-centric evaluation metrics.Empirically, we characterize the dataset by evaluating several methods, including neural models and those based on nearest neighbors.An oracle extractive approach outperforms all benchmarked models according to automatic metrics, showing that the neural models are unable to fully exploit the input transcripts.Human evaluation and qualitative analysis reveal that our nonoracle models are competitive with their oracle counterparts in terms of generating faithful plot events and can benefit from better content selectors.Both oracle and non-oracle models generate unfaithful facts, suggesting future research directions. Mingda Chen, Zewei Chu, Sam Wiseman, Kevin Gimpel |
ACL (1) | 2 |
| 2020 | How to Ask Better Questions? A Large-Scale Multi-Domain Dataset for Rewriting Ill-Formed QuestionsabstractWe present a large-scale dataset for the task of rewriting an ill-formed natural language question to a well-formed one. Our multi-domain question rewriting (MQR) dataset is constructed from human contributed Stack Exchange question edit histories. The dataset contains 427,719 question pairs which come from 303 domains. We provide human annotations for a subset of the dataset as a quality estimate. When moving from ill-formed to well-formed questions, the question quality improves by an average of 45 points across three aspects. We train sequence-to-sequence neural models on the constructed dataset and obtain an improvement of 13.2% in BLEU-4 over baseline methods built from other data resources. We release the MQR dataset to encourage research on the problem of question rewriting.1 Zewei Chu, Mingda Chen, Miaosen Wang, Kevin Gimpel, Manaal Faruqui, Xiance Si |
AAAI | 1 |
| 2019 | EntEval: A Holistic Evaluation Benchmark for Entity RepresentationsabstractMingda Chen, Zewei Chu, Yang Chen, Karl Stratos, Kevin Gimpel. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mingda Chen, Zewei Chu, Karl Stratos, Kevin Gimpel |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence RepresentationsabstractMingda Chen, Zewei Chu, Kevin Gimpel. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mingda Chen, Zewei Chu, Kevin Gimpel |
EMNLP/IJCNLP (1) | 2 |
| 2015 | Scrap your boilerplate with object algebrasabstractTraversing complex Abstract Syntax Trees (ASTs) typically requires large amounts of tedious boilerplate code. For many operations most of the code simply walks the structure, and only a small portion of the code implements the functionality that motivated the traversal in the first place. This paper presents a type-safe Java framework called Shy that removes much of this boilerplate code. In Shy object algebras are used to describe complex and extensible AST structures. Using Java annotations Shy generates generic boilerplate code for various types of traversals. For a concrete traversal, users of Shy can then inherit from the generated code and override only the interesting cases. Consequently, the amount of code that users need to write is significantly smaller. Moreover, traversals using the Shy framework are also much more structure shy, becoming more adaptive to future changes or extensions to the AST structure. To prove the effectiveness of the approach, we applied Shy in the implementation of a domain-specific questionnaire language. Our results show that for a large number of traversals there was a significant reduction in the amount of user-defined code. Zewei Chu, Bruno C. d. S. Oliveira, Tijs van der Storm |
OOPSLA | 2 |