VLDB 2026 Research / reviewers in the wild / expert
Mingyue Shang
dblp:220/3102
· DBLP profile ↗
11ranked-venue papers
3as first author
4since 2021 · last 2023
0009-0000-6523-3516ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Language models and text generation · 41% Question answering and dialogue systems · 33% Efficient and distributed learning · 13% | |
| Software engineering, system software, and programming languages
3 papers |
Program synthesis and code generation · 80% Software testing · 20% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 29 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
code generation |
0.9 | 2 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 ReCode: Robustness Evaluation of Code Generation Models · ACL (1) 2023 |
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation |
0.8 | 2 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 One "Ruler" for All Languages: Multi-Lingual Dialogue Evaluation with Adversarial Multi-Task Learning · IJCAI 2018 |
Natural language and speech › Language models and text generation › text generation
data-to-text generation |
0.7 | 1 | 2023 | Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning · ACL (1) 2023 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation |
0.7 | 2 | 2018 | Get The Point of My Utterance! Learning Towards Effective Responses with Multi-Head Attention Mechanism · IJCAI 2018 Learning to Converse with Noisy Data: Generation with Calibration · IJCAI 2018 |
Natural language and speech › Language models and text generation › text generation › data-to-text generation
few-shot data-to-text generation |
0.7 | 1 | 2023 | Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning · ACL (1) 2023 |
Machine learning › Efficient and distributed learning
model compression |
0.7 | 1 | 2023 | Towards Greener Yet Powerful Code Generation via Quantization: An Empirical Study · ESEC/SIGSOFT FSE 2023 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.7 | 1 | 2023 | Towards Greener Yet Powerful Code Generation via Quantization: An Empirical Study · ESEC/SIGSOFT FSE 2023 |
Program synthesis and code generation
code generation evaluation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Program synthesis and code generation
code generation robustness |
0.7 | 1 | 2023 | ReCode: Robustness Evaluation of Code Generation Models · ACL (1) 2023 |
Program synthesis and code generation
code generation with language models |
0.7 | 1 | 2023 | Towards Greener Yet Powerful Code Generation via Quantization: An Empirical Study · ESEC/SIGSOFT FSE 2023 |
Program synthesis and code generation › code generation with language models
multilingual code generation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Software testing
robustness analysis |
0.7 | 1 | 2023 | ReCode: Robustness Evaluation of Code Generation Models · ACL (1) 2023 |
Natural language and speech › Question answering and dialogue systems
multi-turn dialogue |
0.4 | 1 | 2020 | Context-to-Session Matching: Utilizing Whole Session for Response Selection in Information-Seeking Dialogue Systems · KDD 2020 |
Natural language and speech › Question answering and dialogue systems
response selection |
0.4 | 1 | 2020 | Context-to-Session Matching: Utilizing Whole Session for Response Selection in Information-Seeking Dialogue Systems · KDD 2020 |
Information retrieval
retrieval models |
0.4 | 1 | 2020 | Context-to-Session Matching: Utilizing Whole Session for Response Selection in Information-Seeking Dialogue Systems · KDD 2020 |
Natural language and speech › Question answering and dialogue systems › multi-party dialogue
addressee recognition |
0.4 | 1 | 2019 | Who Is Speaking to Whom? Learning to Identify Utterance Addressee in Multi-Party Conversations · EMNLP/IJCNLP (1) 2019 |
Machine learning › Representation and self-supervised learning
latent space representation |
0.4 | 1 | 2019 | Semi-supervised Text Style Transfer: Cross Projection in Latent Space · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Question answering and dialogue systems
multi-party dialogue |
0.4 | 1 | 2019 | Who Is Speaking to Whom? Learning to Identify Utterance Addressee in Multi-Party Conversations · EMNLP/IJCNLP (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
narrative cloze |
0.4 | 1 | 2019 | Find a Reasonable Ending for Stories: Does Logic Relation Help the Story Cloze Test? · AAAI 2019 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.4 | 1 | 2019 | Find a Reasonable Ending for Stories: Does Logic Relation Help the Story Cloze Test? · AAAI 2019 |
Natural language and speech › Language models and text generation
natural language understanding |
0.4 | 1 | 2019 | Find a Reasonable Ending for Stories: Does Logic Relation Help the Story Cloze Test? · AAAI 2019 |
Natural language and speech › Language models and text generation › controllable text generation
text style transfer |
0.4 | 1 | 2019 | Semi-supervised Text Style Transfer: Cross Projection in Latent Space · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Question answering and dialogue systems
dialogue evaluation |
0.3 | 1 | 2018 | One "Ruler" for All Languages: Multi-Lingual Dialogue Evaluation with Adversarial Multi-Task Learning · IJCAI 2018 |
Natural language and speech › Question answering and dialogue systems
dialogue generation |
0.3 | 1 | 2018 | Get The Point of My Utterance! Learning Towards Effective Responses with Multi-Head Attention Mechanism · IJCAI 2018 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
open-domain dialogue generation |
0.3 | 1 | 2018 | Learning to Converse with Noisy Data: Generation with Calibration · IJCAI 2018 |
Machine learning › Transfer learning and domain adaptation
multi-source learning |
0.2 | 1 | 2023 | Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning · ACL (1) 2023 |
Natural language and speech › Information extraction and text analysis
text matching |
0.1 | 1 | 2020 | Context-to-Session Matching: Utilizing Whole Session for Response Selection in Information-Seeking Dialogue Systems · KDD 2020 |
Natural language and speech › Information extraction and text analysis
textual entailment |
0.1 | 1 | 2019 | Find a Reasonable Ending for Stories: Does Logic Relation Help the Story Cloze Test? · AAAI 2019 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.1 | 1 | 2018 | Get The Point of My Utterance! Learning Towards Effective Responses with Multi-Head Attention Mechanism · IJCAI 2018 |
Methods — techniques the papers use, named apart from their topics
quantization · 1.3perturbation-based evaluation · 1.3metamorphic testing · 1.3large language model · 1.3empirical study · 1.3context-to-session matching · 0.9unified representation · 0.7multi-source learning · 0.7few-shot learning · 0.7cross-task feature combination · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | ReCode: Robustness Evaluation of Code Generation ModelsabstractShiqi Wang, Zheng Li, Haifeng Qian, Chenghao Yang, Zijian Wang, Mingyue Shang, Varun Kumar, Samson Tan, Baishakhi Ray, Parminder Bhatia, Ramesh Nallapati, Murali Krishna Ramanathan, Dan Roth, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shiqi Wang 0002, Haifeng Qian, Chenghao Yang 0001, Zijian Wang 0002, Mingyue Shang, Samson Tan, Baishakhi Ray, Parminder Bhatia, Ramesh Nallapati, Murali Krishna Ramanathan, Dan Roth 0001, Bing Xiang |
ACL (1) | 6 |
| 2023 | Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source LearningabstractAlexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma, Patrick Ng, Zhiguo Wang, Bonan Min, William Yang Wang, Kathleen McKeown, Vittorio Castelli, Dan Roth, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma 0005, Patrick Ng, Zhiguo Wang 0006, Bonan Min, William Yang Wang, Kathy McKeown, Vittorio Castelli, Dan Roth 0001, Bing Xiang |
ACL (1) | 2 |
| 2023 | Multi-lingual Evaluation of Code Generation Models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang 0002, Xiaopeng Li 0002, Wasi Uddin Ahmad, Shiqi Wang 0002, Qing Sun 0013, Mingyue Shang, Sujan K. Gonugondla, Hantian Ding, Nathan Fulton, Arash Farahani, Siddhartha Jain 0001, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati |
ICLR | 10 |
| 2023 | Towards Greener Yet Powerful Code Generation via Quantization: An Empirical StudyabstractML-powered code generation aims to assist developers to write code in a more productive manner by intelligently generating code blocks based on natural language prompts. Recently, large pretrained deep learning models have pushed the boundary of code generation and achieved impressive performance. However, the huge number of model parameters poses a significant challenge to their adoption in a typical software development environment, where a developer might use a standard laptop or mid-size server to develop code. Such large models cost significant resources in terms of memory, latency, dollars, as well as carbon footprint. Xiaokai Wei, Sujan K. Gonugondla, Shiqi Wang 0002, Wasi Uddin Ahmad, Baishakhi Ray, Haifeng Qian, Xiaopeng Li 0002, Zijian Wang 0002, Qing Sun 0013, Ben Athiwaratkun, Mingyue Shang, Murali Krishna Ramanathan, Parminder Bhatia, Bing Xiang |
ESEC/SIGSOFT FSE | 13 |
| 2020 | Context-to-Session Matching: Utilizing Whole Session for Response Selection in Information-Seeking Dialogue SystemsabstractWe study the retrieval-based multi-turn information-seeking dialogue systems, which are widely used in many scenarios. Most of the previous works select the response according to the matching degree between the query's context and the candidate responses. Though great progress has been made, existing works ignore the contexts of the responses, which could provide rich information for selecting the most appropriate response. The more similar the query's context and certain response's context are, the more likely they are to indicate the same question, and thus, the more likely this response is to answer the query. In this paper, we consider the response and its context as a whole session and explore the task of matching the query's context with the sessions. More specifically, we propose to match between the query's context and response's context and integrate the context-to-context matching with context-to-response matching. Experiment results prove that our proposed context-to-session method outperforms the strong baselines significantly. Zhenxin Fu, Shaobo Cui 0001, Mingyue Shang, Dongyan Zhao 0001, Haiqing Chen, Rui Yan 0001 |
KDD | 3 |
| 2019 | Find a Reasonable Ending for Stories: Does Logic Relation Help the Story Cloze Test?abstractNatural language understanding is a challenging problem that covers a wide range of tasks. While previous methods generally train each task separately, we consider combining the cross-task features to enhance the task performance. In this paper, we incorporate the logic information with the help of the Natural Language Inference (NLI) task to the Story Cloze Test (SCT). Previous work on SCT considered various semantic information, such as sentiment and topic, but lack the logic information between sentences which is an essential element of stories. Thus we propose to extract the logic information during the course of the story to improve the understanding of the whole story. The logic information is modeled with the help of the NLI task. Experimental results prove the strength of the logic information. Mingyue Shang, Zhenxin Fu, Hongzhi Yin, Bo Tang 0016, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 1 |
| 2019 | Who Is Speaking to Whom? Learning to Identify Utterance Addressee in Multi-Party ConversationsabstractRan Le, Wenpeng Hu, Mingyue Shang, Zhenjun You, Lidong Bing, Dongyan Zhao, Rui Yan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ran Le, Wenpeng Hu, Mingyue Shang, Zhenjun You, Lidong Bing, Dongyan Zhao 0001, Rui Yan 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Semi-supervised Text Style Transfer: Cross Projection in Latent SpaceabstractMingyue Shang, Piji Li, Zhenxin Fu, Lidong Bing, Dongyan Zhao, Shuming Shi, Rui Yan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mingyue Shang, Piji Li, Zhenxin Fu, Lidong Bing, Dongyan Zhao 0001, Shuming Shi 0001, Rui Yan 0001 |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Learning to Converse with Noisy Data: Generation with CalibrationabstractThe availability of abundant conversational data on the Internet brought prosperity to the generation-based open domain conversation systems. In the training of the generation models, existing methods generally treat all the training data equivalently. However, the data crawled from the websites may contain many noises. Blindly training with the noisy data could harm the performance of the final generation model. In this paper, we propose a generation with calibration framework, that allows high- quality data to have more influences on the generation model and reduces the effect of noisy data. Specifically, for each instance in training set, we employ a calibration network to produce a quality score for it, then the score is used for the weighted update of the generation model parameters. Experiments show that the calibrated model outperforms baseline methods on both automatic evaluation metrics and human annotations. Mingyue Shang, Zhenxin Fu, Nanyun Peng 0001, Yansong Feng 0002, Dongyan Zhao 0001, Rui Yan 0001 |
IJCAI | 1 |
| 2018 | Get The Point of My Utterance! Learning Towards Effective Responses with Multi-Head Attention MechanismabstractAttention mechanism has become a popular and widely used component in sequence-to-sequence models. However, previous research on neural generative dialogue systems always generates universal responses, and the attention distribution learned by the model always attends to the same semantic aspect. To solve this problem, in this paper, we propose a novel Multi-Head Attention Mechanism (MHAM) for generative dialog systems, which aims at capturing multiple semantic aspects from the user utterance. Further, a regularizer is formulated to force different attention heads to concentrate on certain aspects. The proposed mechanism leads to more informative, diverse, and relevant response generated. Experimental results show that our proposed model outperforms several strong baselines. Chongyang Tao, Shen Gao, Mingyue Shang, Wei Wu 0014, Dongyan Zhao 0001, Rui Yan 0001 |
IJCAI | 3 |
| 2018 | One "Ruler" for All Languages: Multi-Lingual Dialogue Evaluation with Adversarial Multi-Task LearningabstractAutomatic evaluating the performance of Open-domain dialogue system is a challenging problem. Recent work in neural network-based metrics has shown promising opportunities for automatic dialogue evaluation. However, existing methods mainly focus on monolingual evaluation, in which the trained metric is not flexible enough to transfer across different languages. To address this issue, we propose an adversarial multi-task neural metric (ADVMT) for multi-lingual dialogue evaluation, with shared feature extraction across languages. We evaluate the proposed model in two different languages. Experiments show that the adversarial multi-task neural metric achieves a high correlation with human annotation, which yields better performance than monolingual ones and various existing metrics. Xiaowei Tong, Zhenxin Fu, Mingyue Shang, Dongyan Zhao 0001, Rui Yan 0001 |
IJCAI | 3 |