EDBT 2026 Demo / reviewers in the wild / expert
Yufei Wang 0003
dblp:61/5568-3
· DBLP profile ↗
16ranked-venue papers
6as first author
14since 2021 · last 2024
0000-0002-8108-3818ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Importance-Aware Data Augmentation for Document-Level Neural Machine TranslationabstractMinghao Wu, Yufei Wang, George Foster, Lizhen Qu, Gholamreza Haffari. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Minghao Wu, Yufei Wang 0003, George F. Foster, Lizhen Qu, Gholamreza Haffari |
EACL (1) | 2 |
| 2024 | Unifying Large Language Models and Knowledge Graphs: A RoadmapabstractLarge language models (LLMs), such as ChatGPT and GPT4, are making new waves in the field of natural language processing and artificial intelligence, due to their emergent ability and generalizability. However, LLMs are black-box models, which often fall short of capturing and accessing factual knowledge. In contrast, Knowledge Graphs (KGs), Wikipedia and Huapu for example, are structured knowledge models that explicitly store rich factual knowledge. KGs can enhance LLMs by providing external knowledge for inference and interpretability. Meanwhile, KGs are difficult to construct and evolve by nature, which challenges the existing methods in KGs to generate new facts and represent unseen knowledge. Therefore, it is complementary to unify LLMs and KGs together and simultaneously leverage their advantages. In this article, we present a forward-looking roadmap for the unification of LLMs and KGs. Our roadmap consists of three general frameworks, namely,1) KG-enhanced LLMs,which incorporate KGs during the pre-training and inference phases of LLMs, or for the purpose of enhancing understanding of the knowledge learned by LLMs;2) LLM-augmented KGs,that leverage LLMs for different KG tasks such as embedding, completion, construction, graph-to-text generation, and question answering; and3) Synergized LLMs + KGs, in which LLMs and KGs play equal roles and work in a mutually beneficial way to enhance both LLMs and KGs for bidirectional reasoning driven by both data and knowledge. We review and summarize existing efforts within these three frameworks in our roadmap and pinpoint their future research directions. Shirui Pan, Linhao Luo, Yufei Wang 0003, Chen Chen 0115, Jiapu Wang, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Separate-and-Aggregate: A Transformer-Based Patch Refinement Model for Knowledge Graph Completion
Chen Chen 0115, Yufei Wang 0003, Yang Zhang 0095, Quan Z. Sheng, Kwok-Yan Lam |
ADMA (2) | 2 |
| 2023 | Investigating the Learning Behaviour of In-Context Learning: A Comparison with Supervised LearningabstractLarge language models (LLMs) have shown remarkable capacity for in-context learning (ICL), where learning a new task from just a few training examples is done without being explicitly pre-trained. However, despite the success of LLMs, there has been little understanding of how ICL learns the knowledge from the given prompts. In this paper, to make progress toward understanding the learning behaviour of ICL, we train the same LLMs with the same demonstration examples via ICL and supervised learning (SL), respectively, and investigate their performance under label perturbations (i.e., noisy labels and label imbalance) on a range of classification tasks. First, via extensive experiments, we find that gold labels have significant impacts on the downstream in-context performance, especially for large language models; however, imbalanced labels matter little to ICL across all model sizes. Second, when comparing with SL, we show empirically that ICL is less sensitive to label perturbations than SL, and ICL gradually attains comparable performance to SL as the model size increases. Xindi Wang 0001, Yufei Wang 0003, Can Xu 0002, Xiubo Geng, Chongyang Tao, Frank Rudzicz, Robert E. Mercer, Daxin Jiang |
ECAI | 2 |
| 2023 | KnowDA: All-in-One Knowledge Mixture Model for Data Augmentation in Low-Resource NLP
Yufei Wang 0003, Can Xu 0002, Xiubo Geng, Tao Shen 0001, Chongyang Tao, Daxin Jiang |
ICLR | 1 |
| 2023 | Hybrid Data Augmentation for Citation Function ClassificationabstractThe citation function generally signifies the purpose or reason underlying a citation within a scholarly paper or a research article. Automatic citation function classification is, therefore, a task in computational linguistics and information science that can facilitate further applications in reference research, citation recommendation, and evaluation of research activities. By taking into account the state of the art, we identify two major constraints pertinent to the data of the citation function classification task, i.e., data imbalance and data sparsity. On the one hand, the natural distribution of different types of citations in one scientific literature is uneven leading to data imbalance in the real scenario. On the other hand, the citation function data is generally labeled by an expert which takes huge human effort resulting in a limited data scale. To this end, in this paper, we propose HybridDA, a two-stage model based on GPT-2 data argumentation and data retrieval to synthesize more high-quality annotated citation function data in a bid to solve both data imbalance and data sparsity problems. We conduct experiments on imbalance setting and low resource setting with our proposed approach. The experimental results on both of these settings demonstrate that our proposed model can achieve competitive performance in contrast to the other baseline models. Yang Zhang 0095, Yufei Wang 0003, Quan Z. Sheng, Mahmood Adnan, Wei Zhang 0098, Rongying Zhao |
IJCNN | 2 |
| 2023 | SocialDial: A Benchmark for Socially-Aware Dialogue SystemsabstractContent Warning: this paper may contain content that is offensive or upsetting. Haolan Zhan, Zhuang Li 0001, Yufei Wang 0003, Linhao Luo, Tao Feng 0013, Xiaoxi Kang, Yuncheng Hua, Lizhen Qu, Lay-Ki Soon, Suraj Sharma, Ingrid Zukerman, Zhaleh Semnani-Azad, Gholamreza Haffari |
SIGIR | 3 |
| 2022 | PromDA: Prompt-based Data Augmentation for Low-Resource NLU TasksabstractYufei Wang, Can Xu, Qingfeng Sun, Huang Hu, Chongyang Tao, Xiubo Geng, Daxin Jiang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yufei Wang 0003, Can Xu 0002, Qingfeng Sun, Huang Hu, Chongyang Tao, Xiubo Geng, Daxin Jiang |
ACL (1) | 1 |
| 2022 | Knowledge Is Flat: A Seq2Seq Generative Framework for Various Knowledge Graph CompletionabstractKnowledge Graph Completion (KGC) has been recently extended to multiple knowledge graph (KG) structures, initiating new research directions, e.g. static KGC, temporal KGC and few-shot KGC. Previous works often design KGC models closely coupled with specific graph structures, which inevitably results in two drawbacks: 1) structure-specific KGC models are mutually incompatible; 2) existing KGC methods are not adaptable to emerging KGs. In this paper, we propose KG-S2S, a Seq2Seq generative framework that could tackle different verbalizable graph structures by unifying the representation of KG facts into “flat” text, regardless of their original form. To remedy the KG structure information loss from the “flat” text, we further improve the input representations of entities and relations, and the inference algorithm in KG-S2S. Experiments on five benchmarks show that KG-S2S outperforms many competitive baselines, setting new state-of-the-art performance. Finally, we analyze KG-S2S’s ability on the different relations and the Non-entity Generations. Chen Chen 0115, Yufei Wang 0003, Kwok-Yan Lam |
COLING | 2 |
| 2021 | Mention Flags (MF): Constraining Transformer-based Text GeneratorsabstractYufei Wang, Ian Wood, Stephen Wan, Mark Dras, Mark Johnson. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yufei Wang 0003, Ian D. Wood, Stephen Wan 0001, Mark Dras, Mark Johnson 0001 |
ACL/IJCNLP (1) | 1 |
| 2021 | Assessment2Vec: Learning Distributed Representations of Assessments to Reduce Marking Workload
Shuang Wang 0012, Amin Beheshti, Yufei Wang 0003, Jianchao Lu, Quan Z. Sheng, Stephen Elbourn, Hamid Alinejad-Rokny, Elizabeth Galanis |
AIED (2) | 3 |
| 2021 | ECOL-R: Encouraging Copying in Novel Object Captioning with Reinforcement LearningabstractNovel Object Captioning is a zero-shot Image Captioning task requiring describing objects not seen in the training captions, but for which information is available from external object detectors.The key challenge is to select and describe all salient detected novel objects in the input images.In this paper, we focus on this challenge and propose the ECOL-R model (Encouraging Copying of Object Labels with Reinforced Learning), a copy-augmented transformer model that is encouraged to accurately describe the novel object labels.This is achieved via a specialised reward function in the SCST reinforcement learning framework (Rennie et al., 2017) that encourages novel object mentions while maintaining the caption quality.We further restrict the SCST training to the images where detected objects are mentioned in reference captions to train the ECOL-R model.We additionally improve our copy mechanism via Abstract Labels, which transfer knowledge from known to novel object types, and a Morphological Selector, which determines the appropriate inflected forms of novel object labels.The resulting model sets new state-of-the-art on the nocaps (Agrawal et al., 2019) and held-out COCO (Hendricks et al., 2016) benchmarks. Yufei Wang 0003, Ian D. Wood, Stephen Wan 0001, Mark Johnson 0001 |
EACL | 1 |
| 2021 | Neural Rule-Execution Tracking Machine For Transformer-Based Text GenerationabstractSequence-to-Sequence (Seq2Seq) neural text generation models, especially the pre-trained ones (e.g., BART and T5), have exhibited compelling performance on various natural language generation tasks. However, the black-box nature of these models limits their application in tasks where specific rules (e.g., controllable constraints, prior knowledge) need to be executed. Previous works either design specific model structures (e.g., Copy Mechanism corresponding to the rule "the generated output should include certain words in the source input'') or implement specialized inference algorithms (e.g., Constrained Beam Search) to execute particular rules through the text generation. These methods require the careful design case-by-case and are difficult to support multiple rules concurrently. In this paper, we propose a novel module named Neural Rule-Execution Tracking Machine (NRETM) that can be equipped into various transformer-based generators to leverage multiple rules simultaneously to guide the neural generation model for superior generation performance in an unified and scalable way. Extensive experiments on several benchmarks verify the effectiveness of our proposed model in both controllable and general text generation tasks. Yufei Wang 0003, Can Xu 0002, Huang Hu, Chongyang Tao, Stephen Wan 0001, Mark Dras, Mark Johnson 0001, Daxin Jiang |
NeurIPS | 1 |
| 2021 | TDM-CFC: Towards Document-Level Multi-label Citation Function Classification
Yang Zhang 0095, Yufei Wang 0003, Quan Z. Sheng, Mahmood Adnan, Wei Zhang 0098, Rongying Zhao |
WISE (2) | 2 |
| 2019 | How to Best Use Syntax in Semantic Role LabellingabstractThere are many different ways in which external information might be used in an NLP task.This paper investigates how external syntactic information can be used most effectively in the Semantic Role Labeling (SRL) task.We evaluate three different ways of encoding syntactic parses and three different ways of injecting them into a state-of-the-art neural ELMo-based SRL sequence labelling model.We show that using a constituency representation as input features improves performance the most, achieving a new state-of-the-art for non-ensemble SRL models on the in-domain CoNLL'05 and CoNLL'12 benchmarks. 1 Yufei Wang 0003, Mark Johnson 0001, Stephen Wan 0001, Yifang Sun, Wei Wang 0011 |
ACL (1) | 1 |
| 2019 | nocaps: novel object captioning at scaleabstractImage captioning models have achieved impressive results on datasets containing limited visual concepts and large amounts of paired image-caption training data. However, if these models are to ever function in the wild, a much larger variety of visual concepts must be learned, ideally from less supervision. To encourage the development of image captioning models that can learn visual concepts from alternative data sources, such as object detection datasets, we present the first large-scale benchmark for this task. Dubbed `nocaps', for novel object captioning at scale, our benchmark consists of 166,100 human-generated captions describing 15,100 images from the Open Images validation and test sets. The associated training data consists of COCO image-caption pairs, plus Open Images image-level labels and object bounding boxes. Since Open Images contains many more classes than COCO, nearly 400 object classes seen in test images have no or very few associated training captions (hence, nocaps). We extend existing novel object captioning models to establish strong baselines for this benchmark and provide analysis to guide future work. Harsh Agrawal, Peter Anderson 0001, Karan Desai, Yufei Wang 0003, Xinlei Chen, Mark Johnson 0001, Dhruv Batra, Devi Parikh, Stefan Lee |
ICCV | 4 |