VLDB 2026 Research / reviewers in the wild / expert
Jiaan Wang
dblp:296/2112
· DBLP profile ↗
31ranked-venue papers
11as first author
31since 2021 · last 2026
0000-0002-2587-7648ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 10 first-author · 24 since 2021Databases, data management, data science and information retrieval · 14 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Semantic Features: Pixel-level Mapping for Generalized AI-Generated Image DetectionabstractThe rapid evolution of generative technologies necessitates reliable methods for detecting AI-generated images. A critical limitation of current detectors is their failure to generalize to images from unseen generative models, as they often overfit to source-specific semantic cues rather than learning universal generative artifacts. To overcome this, we introduce a simple yet remarkably effective pixel-level mapping pre-processing step to disrupt the pixel value distribution of images and break the fragile, non-essential semantic patterns that detectors commonly exploit as shortcuts. This forces the detector to focus on more fundamental and generalizable high-frequency traces inherent to the image generation process. Through comprehensive experiments on GAN and diffusion-based generators, we show that our approach significantly boosts the cross-generator performance of state-of-the-art detectors. Extensive analysis further verifies our hypothesis that the disruption of semantic cues is the key to generalization. Chenming Zhou, Jiaan Wang, Yu Li 0016, Juan Cao 0001, Sheng Tang |
AAAI | 2 |
| 2026 | Large Language Model Judged Self-Training for Named Entity RecognitionabstractSelf-training for Named Entity Recognition (NER) aims at identifying named entities and their types in the text using self-training to fully make use of the limited labeled data and a large amount of unlabeled data. The major challenge in self-training is confirmation bias where incorrect pseudo-labels increase errors. Many efforts have been made to address this challenge, but few labeled data limit their performance. In this paper, we introduce Large Language Model (LLM) into self-training to select high-quality pseudo-labels leveraging its rich knowledge and few-shot learning capability. Specifically, we design a comprehensive prompt to improve the judgment performance of LLM, where the prompt incorporates task rules mined by LLM itself to fully leverage labeled data. In addition, to reduce the impact of LLM's hallucinations, we adopt a collaborative pseudo-label selection based on combined confidence and calibration-guided probability smoothing. Our empirical study conducted on several NER datasets shows that our method outperforms state-of-the-art approaches. The code is available at https://github.com/cheniison/llm-judged-ST. Shisong Chen, Jiaan Wang, Yanghua Xiao, Zhixu Li, Xin Lin 0001 |
WSDM | 2 |
| 2026 | Toward Graph Data Collaboration in a Data-Sharing-Free Manner: A Novel Privacy-Preserving Graph Pretraining ModelabstractGraph data, prevalent in various domains such as telecommunication, supply chain, and social networks, holds significant potential for business, operations, and social administration. Collaborating on graph data across institutions or users can further unleash its value, making it a highly sought-after practice. However, such collaboration poses risks to information privacy and commercial confidentiality. In response, we introduce an innovative new model-sharing strategy for graph data collaboration. Here, a data owner pretrains a graph neural network (GNN) model on their private graph data and then provides model users with query access to this model. The pretrained GNN acts as an intermediary, encapsulating knowledge from the private data without exposing it directly. Two fundamental principles are essential for such a pretrained GNN model: model generalizability and privacy preservation. However, current efforts often fail to achieve both concurrently. To tackle this challenge and promote an open yet secure graph data collaboration framework, we propose a novel privacy-preserving operator. This operator integrates smoothly with graph data augmentation and graph contrastive learning, allowing the pretraining of a GNN that effectively eliminates private links at high risk of exposure while maintaining generalizability. Additionally, to improve model generalizability, we introduce a new method called generalizability learning to enhance the model’s adaptability when deployed on unseen data of model user. This approach is designed to simulate diverse environments and develop representations that remain invariant across these varied environments. Extensive experiments suggest that our model surpasses existing state-of-the-art approaches in striking an effective balance between privacy preservation and generalizability. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Funding: This work was supported (to J. Xu) by the National Natural Science Foundation of China [Grants 62206056, 72271059, and 72442011] and the CIPSC-SMP-Zhipu Large Model Cross-Disciplinary Fund. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0115 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0115 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Jiarong Xu, Jiaan Wang, Zenan Zhou, Tian Lu 0002 |
INFORMS J. Comput. | 2 |
| 2026 | DeepTrans: Deep Reasoning Translation via Reinforcement LearningabstractAbstract Recently, deep reasoning LLMs (e.g., OpenAI o1 and DeepSeek-R1) have shown promising performance in various downstream tasks. Free translation is an important and interesting task in the multilingual world, which requires going beyond word-for-word translation. However, the task is still under-explored in deep reasoning LLMs. In this paper, we introduce DeepTrans, a deep reasoning translation model that learns free translation via reinforcement learning (RL). Specifically, we carefully build a reward model with pre-defined scoring criteria on both the translation results and the thought processes. The reward model teaches DeepTrans how to think and free-translate the given sentences during RL. Besides, our RL training does not need any labeled translations, avoiding the human-intensive annotation or resource-intensive data synthesis. Experimental results show the effectiveness of DeepTrans. Using Qwen2.5-7B as the backbone, DeepTrans improves performance by 16.3% in literature translation, and outperforms strong deep reasoning LLMs. Moreover, we summarize the failures and interesting findings during our RL exploration. We hope this work could inspire other researchers in free translation.1 Jiaan Wang, Fandong Meng, Jie Zhou 0016 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2025 | Enhancing Long-and Short-Term Representations for Next POI Recommendations via Frequency and Hierarchical Contrastive LearningabstractNext POI recommendation aids users in predicting their destinations of interest and plays an increasingly vital role in location-based social services. Recent works focus on analyzing both long-term and short-term interests in POI recommendation to gain a deeper understanding of user profiles. However, these methods for modeling long-term user’s sequences primarily rely on the Transformer model, which functions as a low-pass filter, often leading to the loss of high-frequency information. Additionally, long-term and short-term sequences are typically modeled independently, with short-term sequences often defined solely by the most recent check-ins, overlooking their interactions and dependencies. Therefore, we propose Enhancing Long-and Short-Term Representations for Next POI Recommendations via Frequency and Hierarchical Contrastive Learning (FHCRec). FHCRec captures both high-frequency and low-frequency information in long-term sequences to model richer long-term user’s preference representations. Moreover, it harnesses the characteristics of the short-term subsequences embedded within long-term sequences to enhance short-term preference characterization via local and global hierarchical contrastive learning, resulting in more personalized short-term preferences. The enhanced long-term and short-term preferences are integrated to improve model recommendation performance. Extensive experiments on three real-world datasets demonstrate the effectiveness of our method. Peng-Fei Zhang 0001, Jiaan Wang, Jianfeng Qu, Zhixu Li |
AAAI | 4 |
| 2025 | An Empirical Study of Many-to-Many Summarization with Large Language ModelsabstractJiaan Wang, Fandong Meng, Zengkui Sun, Yunlong Liang, Yuxuan Cao, Jiarong Xu, Haoxiang Shi, Jie Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiaan Wang, Fandong Meng, Zengkui Sun, Yunlong Liang, Jiarong Xu, Haoxiang Shi, Jie Zhou 0016 |
ACL (1) | 1 |
| 2025 | How to use Graph Data in the Wild to Help Graph Anomaly Detection?abstractIn recent years, graph anomaly detection has gained considerable attention and has found extensive applications in various domains such as social, financial, and communication networks. However, anomalies in graph-structured data present unique challenges, including label scarcity, ill-defined anomalies, and varying anomaly types, making supervised or semi-supervised methods unreliable. Researchers often adopt unsupervised approaches to address these challenges, assuming that anomalies deviate significantly from the normal data distribution. Yet, when the available data is insufficient, capturing the normal distribution accurately and comprehensively becomes difficult. To overcome this limitation, we propose to utilize external graph data (i.e., graph data in the wild) to help anomaly detection tasks. This naturally raises the question: How can we use external data to help graph anomaly detection task? To answer this question, we propose a novel framework Wild-GAD. Our framework is built upon a unified database, UniWildGraph, which comprises a large and diverse collection of graph data with broad domain coverage, ample data volume, and a unified feature space. We further develop selection criteria based on representativity and diversity to identify the most suitable external data for each anomaly detection task. Extensive experiments on six real-world test datasets demonstrate the effectiveness of Wild-GAD. Compared to the baseline methods, our framework has an average 18% AUCROC and 32% AUCPR improvement over the best-competing methods. Jiarong Xu, Chen Zhao 0029, Jiaan Wang, Carl Yang 0001, Chunping Wang 0001, Yang Yang 0009 |
KDD (1) | 4 |
| 2025 | Concept-aware embedding for logical query reasoning over knowledge graphs
Pengwei Pan, Jingpei Lei, Jiaan Wang, Dantong Ouyang, Jianfeng Qu, Zhixu Li |
Inf. Process. Manag. | 3 |
| 2024 | Improving the Robustness of Knowledge-Grounded Dialogue via Contrastive LearningabstractKnowledge-grounded dialogue (KGD) learns to generate an informative response based on a given dialogue context and external knowledge (e.g., knowledge graphs; KGs). Recently, the emergence of large language models (LLMs) and pre-training techniques has brought great success to knowledge-grounded dialogue. However, when building KGD systems in real applications, there are various real-world noises that are inevitable to face. For example, the dialogue context might involve perturbations such as misspellings and abbreviations. In addition, KGs typically suffer from incompletion and also might contain erroneous and outdated facts. Such real-world noises pose a challenge to the robustness of KGD systems and hinder their applications in the real world. In this paper, we propose an entity-based contrastive learning framework for improving the robustness of KGD. Specifically, we make use of the entity information in a KGD sample to create both its positive and negative samples which involve semantic-irrelevant and semantic-relevant perturbations, respectively. The contrastive learning framework ensures the KGD model is aware of these two types of perturbations, thus could generate informative responses with the potentially noisy inputs in real applications. Experimental results on three widely-used benchmark datasets show that our method achieves new state-of-the-art performance in terms of automatic evaluation scores, verifying its effectiveness and potentiality. Furthermore, we show that our method is able to generate better responses than comparison models in both the noisy and the few-shot settings. Jiaan Wang, Jianfeng Qu, Zhixu Li, Wen Hua, Ximing Li 0002, An Liu 0002 |
AAAI | 1 |
| 2024 | Continual Learning with Semi-supervised Contrastive Distillation for Incremental Neural Machine TranslationabstractIncrementally expanding the capability of an existing translation model to solve new domain tasks over time is a fundamental and practical problem, which usually suffers from catastrophic forgetting.Generally, multi-domain learning can be seen as a good solution.However, there are two drawbacks: 1) it requires having the training data for all domains available at the same time, which may be unrealistic due to storage or privacy concerns; 2) it requires re-training the model on the data of all domains from scratch when adding a new domain and this is time-consuming and computationally expensive.To address these issues, we present a semi-supervised contrastive distillation framework for incremental neural machine translation.Specifically, to avoid catastrophic forgetting, we propose to exploit unlabeled data from the same distributions of the older domains through knowledge distillation.Further, to ensure the distinct domain characteristics in the model as the number of domains increases, we devise a cross-domain contrastive objective to enhance the distilled knowledge.Extensive experiments on domain translation benchmarks show that our approach, without accessing any previous training data or re-training on all domains from scratch, can significantly prevent the model from forgetting previously learned knowledge while obtaining good performance on the incrementally added domains. Yunlong Liang, Fandong Meng, Jiaan Wang, Jin An Xu, Yufeng Chen 0005, Jie Zhou 0016 |
ACL (1) | 3 |
| 2024 | Cross-Lingual Knowledge Editing in Large Language ModelsabstractKnowledge editing aims to change language models' performance on several special cases (i.e., editing scope) by infusing the corresponding expected knowledge into them.With the recent advancements in large language models (LLMs), knowledge editing has been shown as a promising technique to adapt LLMs to new knowledge without retraining from scratch.However, most of the previous studies neglect the multi-lingual nature of some main-stream LLMs (e.g., LLaMA, ChatGPT and GPT-4), and typically focus on monolingual scenarios, where LLMs are edited and evaluated in the same language.As a result, it is still unknown the effect of source language editing on a different target language.In this paper, we aim to figure out this cross-lingual effect in knowledge editing.Specifically, we first collect a largescale cross-lingual synthetic dataset by translating ZsRE from English to Chinese.Then, we conduct English editing on various knowledge editing methods covering different paradigms, and evaluate their performance in Chinese, and vice versa.To give deeper analyses of the crosslingual effect, the evaluation includes four aspects, i.e., reliability, generality, locality and portability.Furthermore, we analyze the inconsistent behaviors of the edited models and discuss their specific challenges. 1 Jiaan Wang, Yunlong Liang, Zengkui Sun, Jiarong Xu, Fandong Meng |
ACL (1) | 1 |
| 2024 | M2ConceptBase: A Fine-Grained Aligned Concept-Centric Multimodal Knowledge BaseabstractMultimodal knowledge bases (MMKBs) provide cross-modal aligned knowledge crucial for multimodal tasks. However, the images in existing MMKBs are generally collected for entities in encyclopedia knowledge graphs. Therefore, detailed groundings of visual semantics with linguistic concepts are lacking, which are essential for the visual concept cognition ability of multimodal models. Addressing this gap, we introduce M2 ConceptBase, the first concept-centric MMKB. M2 ConceptBase models concepts as nodes with associated images and detailed textual descriptions. We propose a context-aware multimodal symbol grounding approach to align concept-image and concept-description pairs using context information from image-text datasets. Comprising 951K images and 152K concepts, M2 ConceptBase links each concept to an average of 6.27 images and a single description, ensuring comprehensive visual and textual semantics. Human studies confirm more than 95% alignment accuracy, underscoring its quality. Additionally, our experiments demonstrate that M2 ConceptBase significantly enhances VQA model performance on the OK-VQA task. M2 ConceptBase also substantially improves the fine-grained concept understanding capabilities of multimodal large language models through retrieval augmentation in two concept-related tasks, highlighting its value. Zhiwei Zha, Jiaan Wang, Zhixu Li, Xiangru Zhu, Wei Song 0008, Yanghua Xiao |
CIKM | 2 |
| 2024 | Unleashing the Power of Emojis in Texts via Self-supervised Graph Pre-TrainingabstractEmojis have gained immense popularity on social platforms, serving as a common means to supplement or replace text.However, existing data mining approaches generally either completely ignore or simply treat emojis as ordinary Unicode characters, which may limit the model's ability to grasp the rich semantic information in emojis and the interaction between emojis and texts.Thus, it is necessary to release the power of emojis in social media data mining.To this end, we first construct a heterogeneous graph consisting of three types of nodes, i.e., post, word and emoji nodes to improve the representation of different elements in posts.The edges are also well-defined to model how these three elements interact with each other.To facilitate the sharing of information among post, word and emoji nodes, we propose a graph pre-training framework for text and emoji co-modeling, which contains two graph pre-training tasks: node-level graph contrastive learning and edge-level link reconstruction learning.Extensive experiments on the Xiaohongshu and Twitter datasets with two types of downstream tasks demonstrate that our approach proves significant improvement over previous strong baseline methods. 1 Dongzeng Tan, Jiaan Wang, Jiarong Xu |
EMNLP | 3 |
| 2024 | ESC-Eval: Evaluating Emotion Support Conversations in Large Language ModelsabstractHaiquan Zhao, Lingyu Li, Shisong Chen, Shuqi Kong, Jiaan Wang, Kexin Huang, Tianle Gu, Yixu Wang, Jian Wang, Liang Dandan, Zhixu Li, Yan Teng, Yanghua Xiao, Yingchun Wang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Haiquan Zhao 0002, Shisong Chen, Shuqi Kong, Jiaan Wang, Tianle Gu, Yixu Wang, Dandan Liang, Zhixu Li, Yan Teng 0002, Yanghua Xiao, Yingchun Wang 0004 |
EMNLP | 5 |
| 2024 | A Coarse-to-Fine Framework for Entity-Relation Joint ExtractionabstractExtracting entities and relations from text is a significant task of information extraction. Existing extraction models often straightforwardly produce their confident prediction results without any reconsideration or double-checking, resulting in avoidable mistakes and sub-optimal performance. In this paper, we propose a novel coarse-to-fine extraction framework, which first extracts high-potential relations as well as entities via knowledge distillation, and then rechecks the predictions via handcrafted natural language inference (NLI) task in a fine-grained manner. Specifically, based on the knowledge distillation mechanism, we train multiple teacher models iteratively through an adaptive loss function for making one teacher concentrate more on the data that others are incompetent for. Then, these complementary teacher models are utilized to provide valuable soft-label information for training a considerate student model, enabling it to generate reliable preliminary predictions. Further, these generated potential relations and entities are formulated as hypotheses, together with the original sentences as premises, serving as the input for an NLI model. Considering the linguistic diversity of relational expression, we automatically generate various semantic templates for hypotheses through an$\mathcal{N}$-gram mining strategy. Moreover, due to the existence of multi-fact sentences, a relation-guided Gaussian attention is designed to reduce the gap between the single-relation hypothesis and the multi-relation premise. To implement efficient training, we also develop several ways to generate high-quality negative samples, which help the NLI model learn to identify errors. Experimental results show that the proposed method is effective and outperforms other strong baselines on public benchmarks. Mingchen Zhang, Jiaan Wang, Jianfeng Qu, Zhixu Li, An Liu 0002, Lei Zhao 0001, Zhigang Chen 0003, Xiaofang Zhou 0001 |
ICDE | 2 |
| 2023 | Summary-Oriented Vision Modeling for Multimodal Abstractive SummarizationabstractMultimodal abstractive summarization (MAS) aims to produce a concise summary given the multimodal data (text and vision).Existing studies mainly focus on how to effectively use the visual features from the perspective of an article, having achieved impressive success on the high-resource English dataset.However, less attention has been paid to the visual features from the perspective of the summary, which may limit the model performance, especially in the low-and zero-resource scenarios.In this paper, we propose to improve the summary quality through summary-oriented visual features.To this end, we devise two auxiliary tasks including vision to summary task and masked image modeling task.Together with the main summarization task, we optimize the MAS model via the training objectives of all these tasks.By these means, the MAS model can be enhanced by capturing the summaryoriented visual features, thereby yielding more accurate summaries.Experiments on 44 languages, covering mid-high-, low-, and zeroresource scenarios, verify the effectiveness and superiority of the proposed approach, which achieves state-of-the-art performance under all scenarios.Additionally, we will contribute a large-scale multilingual multimodal abstractive summarization (MM-Sum) dataset. 1 Yunlong Liang, Fandong Meng, Jin An Xu, Jiaan Wang, Yufeng Chen 0005, Jie Zhou 0016 |
ACL (1) | 4 |
| 2023 | Towards Unifying Multi-Lingual and Cross-Lingual SummarizationabstractJiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang, Zhixu Li, Jianfeng Qu, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang, Zhixu Li, Jianfeng Qu, Jie Zhou 0016 |
ACL (1) | 1 |
| 2023 | AspectMMKG: A Multi-modal Knowledge Graph with Aspect-aware EntitiesabstractMulti-modal knowledge graphs (MMKGs) combine different modal data (e.g., text and image) for a comprehensive understanding of entities. Despite the recent progress of large-scale MMKGs, existing MMKGs neglect the multi-aspect nature of entities, limiting the ability to comprehend entities from various perspectives.In this paper, we construct AspectMMKG, the first MMKG with aspect-related images by matching images to different entity aspects. Specifically, we collect aspect-related images from a knowledge base, and further extract aspect-related sentences from the knowledge base as queries to retrieve a large number of aspect-related images via an online image search engine. Finally, AspectMMKG contains 2,380 entities, 18,139 entity aspects, and 645,383 aspect-related images. We demonstrate the usability of AspectMMKG in entity aspect linking (EAL) downstream task and show that previous EAL models achieve a new state-of-the-art performance with the help of AspectMMKG.To facilitate the research on aspect-related MMKG, we further propose an aspect-related image retrieval (AIR) model, that aims to correct and expand aspect-related images in AspectMMKG.We train an AIR model to learn the relationship between entity image and entity aspect-related images by incorporating entity image, aspect, and aspect image information. Experimental results indicate that the AIR model could retrieve suitable images for a given entity w.r.t different aspects. Jingdan Zhang, Jiaan Wang, Zhixu Li, Yanghua Xiao |
CIKM | 2 |
| 2023 | A Joint Link-Retrieve Framework for Open Table-and-Text Question Answering
Jiaan Wang, Ying He 0010, Jianfeng Qu, Zhixu Li, Pengpeng Zhao 0001, An Liu 0002, Lei Zhao 0001 |
DASFAA (3) | 2 |
| 2023 | When to Pre-Train Graph Neural Networks? From Data Generation Perspective!abstractIn recent years, graph pre-training has gained significant attention, focusing on acquiring transferable knowledge from unlabeled graph data to improve downstream performance. Despite these recent endeavors, the problem of negative transfer remains a major concern when utilizing graph pre-trained models to downstream tasks. Previous studies made great efforts on the issue of what to pre-train and how to pre-train by designing a variety of graph pre-training and fine-tuning strategies. However, there are cases where even the most advanced "pre-train and fine-tune" paradigms fail to yield distinct benefits. This paper introduces a generic framework W2PGNN to answer the crucial question of when to pre-train (.e., in what situations could we take advantage of graph pre-training) before performing effortful pre-training or fine-tuning. We start from a new perspective to explore the complex generative mechanisms from the pre-training data to downstream data. In particular, W2PGNN first fits the pre-training data into graphon bases, each element of graphon basis (i.e., a graphon) identifies a fundamental transferable pattern shared by a collection of pre-training graphs. All convex combinations of graphon bases give rise to a generator space, from which graphs generated form the solution space for those downstream data that can benefit from pre-training. In this manner, the feasibility of pre-training can be quantified as the generation probability of the downstream data from any generator in the generator space. W2PGNN offers three broad applications: providing the application scope of graph pre-trained models, quantifying the feasibility of pre-training, and assistance in selecting pre-training data to enhance downstream performance. We provide a theoretically sound solution for the first application and extensive empirical justifications for the latter two applications. Jiarong Xu, Carl Yang 0001, Jiaan Wang, Yunchao Zhang, Chunping Wang 0001, Lei Chen 0082, Yang Yang 0009 |
KDD | 4 |
| 2023 | Long-Document Cross-Lingual SummarizationabstractCross-Lingual Summarization (CLS) aims at generating summaries in one language for the given documents in another language. CLS has attracted wide research attention due to its practical significance in the multi-lingual world. Though great contributions have been made, existing CLS works typically focus on short documents, such as news and guides. Different from these short texts, long documents such as academic articles usually discuss complicated subjects and consist of thousands of words, making them non-trivial to process and summarize. To promote CLS research on long documents, we construct Perseus, the first long-document CLS dataset which collects about 94K Chinese scientific documents paired with English summaries. The average length of documents in Perseus is more than 2000 tokens. As a preliminary study on long-document CLS, we build and evaluate various CLS baselines, including pipeline and end-to-end methods. Experimental results on Perseus show the superiority of the end-to-end baseline, which performs the best among all methods. Furthermore, to provide a deeper understanding, we manually analyze the model outputs and discuss specific challenges faced by current approaches. We hope that our work could benchmark long-document CLS and benefit future studies. Shaohui Zheng, Zhixu Li, Jiaan Wang, Jianfeng Qu, An Liu 0002, Lei Zhao 0001, Zhigang Chen 0003 |
WSDM | 3 |
| 2022 | LayerConnect: Hypernetwork-Assisted Inter-Layer Connector to Enhance Parameter EfficiencyabstractPre-trained Language Models (PLMs) are the cornerstone of the modern Natural Language Processing (NLP). However, as PLMs become heavier, fine tuning all their parameters loses their efficiency. Existing parameter-efficient methods generally focus on reducing the trainable parameters in PLMs but neglect the inference speed, which limits the ability to deploy PLMs. In this paper, we propose LayerConnect (hypernetwork-assisted inter-layer connectors) to enhance inference efficiency. Specifically, a light-weight connector with a linear structure is inserted between two Transformer layers, and the parameters inside each connector are tuned by a hypernetwork comprising an interpolator and a down-sampler. We perform extensive experiments on the widely used the GLUE benchmark. The experimental results verify the inference efficiency of our model. Compared to Adapter, our model parameters are reduced to approximately 11.75%, while the performance degradation is kept to less than 5% (2.5 points on average). Haoxiang Shi, Jiaan Wang, Cen Wang, Yinhe Zheng, Tetsuya Sakai |
COLING | 3 |
| 2022 | Incorporating Commonsense Knowledge into Story Ending Generation via Heterogeneous Graph Networks
Jiaan Wang, Beiqi Zou, Zhixu Li, Jianfeng Qu, Pengpeng Zhao 0001, An Liu 0002, Lei Zhao 0001 |
DASFAA (3) | 1 |
| 2022 | Aligning Internal Regularity and External Influence of Multi-granularity for Temporal Knowledge Graph Embedding
Tingyi Zhang, Zhixu Li, Jiaan Wang, Jianfeng Qu, An Liu 0002, Lei Zhao 0001, Zhigang Chen 0003 |
DASFAA (3) | 3 |
| 2022 | ClidSum: A Benchmark Dataset for Cross-Lingual Dialogue SummarizationabstractWe present CLIDSUM, a benchmark dataset towards building cross-lingual summarization systems on dialogue documents.It consists of 67k+ dialogue documents and 112k+ annotated summaries in different target languages.Based on the proposed CLIDSUM, we introduce two benchmark settings for supervised and semi-supervised scenarios, respectively.We then build various baseline systems in different paradigms (pipeline and end-to-end) and conduct extensive experiments on CLIDSUM to provide deeper analyses.Furthermore, we propose mDIALBART which extends mBART via further pre-training, where the multiple objectives help the pre-trained model capture the structural characteristics as well as key content in dialogues and the transformation from source to the target language.Experimental results show the superiority of mDIALBART, as an end-to-end model, outperforms strong pipeline models on CLIDSUM.Finally, we discuss specific challenges that current approaches faced with this task and give multiple promising directions for future research. Jiaan Wang, Fandong Meng, Ziyao Lu, Duo Zheng, Zhixu Li, Jianfeng Qu, Jie Zhou 0016 |
EMNLP | 1 |
| 2022 | Towards Unifying Reference Expression Generation and ComprehensionabstractReference Expression Generation (REG) and Comprehension (REC) are two highly correlated tasks. Modeling REG and REC simultaneously for utilizing the relation between them is a promising way to improve both. However, the problem of distinct inputs, as well as building connections between them in a single model, brings challenges to the design and training of the joint model. To address the problems, we propose a unified model for REG and REC, named UniRef. It unifies these two tasks with the carefully-designed Image-Region-Text Fusion layer (IRTF), which fuses the image, region and text via the image cross-attention and region cross-attention. Additionally, IRTF could generate pseudo input regions for the REC task to enable a uniform way for sharing the identical representation space across the REC and REG. We further propose Vision-conditioned Masked Language Modeling (VMLM) and Text-Conditioned Region Prediction (TRP) to pre-train UniRef model on multi-granular corpora. The VMLM and TRP are directly related to REG and REC, respectively, but could help each other. We conduct extensive experiments on three benchmark datasets, RefCOCO, RefCOCO+ and RefCOCOg. Experimental results show that our model outperforms previous state-of-the-art methods on both REG and REC. Duo Zheng, Tao Kong, Ya Jing, Jiaan Wang, Xiaojie Wang 0006 |
EMNLP | 4 |
| 2022 | RT-KGD: Relation Transition Aware Knowledge-Grounded Dialogue Generation
Zhixu Li, Jiaan Wang, Jianfeng Qu, Ying He 0010, An Liu 0002, Lei Zhao 0001 |
ISWC | 3 |
| 2022 | Knowledge Enhanced Sports Game SummarizationabstractSports game summarization aims at generating sports news from live commentaries. However, existing datasets are all constructed through automated collection and cleaning processes, resulting in a lot of noise. Besides, current works neglect the knowledge gap between live commentaries and sports news, which limits the performance of sports game summarization. In this paper, we introduce K-SportsSum, a new dataset with two characteristics: (1) K-SportsSum collects a large amount of data from massive games. It has 7,854 commentary-news pairs. To improve the quality, K-SportsSum employs a manual cleaning process; (2) Different from existing datasets, to narrow the knowledge gap, K-SportsSum further provides a large-scale knowledge corpus that contains the information of 523 sports teams and 14,724 sports players. Additionally, we also introduce a knowledge-enhanced summarizer that utilizes both live commentaries and the knowledge to generate sports news. Extensive experiments on K-SportsSum and SportsSum datasets show that our model achieves new state-of-the-art performances. Qualitative analysis and human study further verify that our model generates more informative sports news. Jiaan Wang, Zhixu Li, Tingyi Zhang, Duo Zheng, Jianfeng Qu, An Liu 0002, Lei Zhao 0001, Zhigang Chen 0003 |
WSDM | 1 |
| 2022 | A Survey on Cross-Lingual SummarizationabstractAbstract Cross-lingual summarization is the task of generating a summary in one language (e.g., English) for the given document(s) in a different language (e.g., Chinese). Under the globalization background, this task has attracted increasing attention of the computational linguistics community. Nevertheless, there still remains a lack of comprehensive review for this task. Therefore, we present the first systematic critical review on the datasets, approaches, and challenges in this field. Specifically, we carefully organize existing datasets and approaches according to different construction methods and solution paradigms, respectively. For each type of dataset or approach, we thoroughly introduce and summarize previous efforts and further compare them with each other to provide deeper analyses. In the end, we also discuss promising directions and offer our thoughts to facilitate future research. This survey is for both beginners and experts in cross-lingual summarization, and we hope it will serve as a starting point as well as a source of new ideas for researchers and engineers interested in this area. Jiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang, Zhixu Li, Jianfeng Qu, Jie Zhou 0016 |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | SportsSum2.0: Generating High-Quality Sports News from Live Text CommentaryabstractSports game summarization aims to generate news articles from live text commentaries. A recent state-of-the-art work, SportsSum, not only constructs a large benchmark dataset, but also proposes a two-step framework. Despite its great contributions, the work has three main drawbacks: 1) the noise existed in SportsSum dataset degrades the summarization performance; 2) the neglect of lexical overlap between news and commentaries results in low-quality pseudo-labeling algorithm; 3) the usage of directly concatenating rewritten sentences to form news limits its practicability. In this paper, we publish a new benchmark dataset SportsSum2.0, together with a modified summarization framework. In particular, to obtain a clean dataset, we employ crowd workers to manually clean the original dataset. Moreover, the degree of lexical overlap is incorporated into the generation of pseudo labels. Further, we introduce a reranker-enhanced summarizer to take into account the fluency and expressiveness of the summarized news. Extensive experiments show that our model outperforms the state-of-the-art baseline. Jiaan Wang, Zhixu Li, Qiang Yang 0015, Jianfeng Qu, Zhigang Chen 0003, Qingsheng Liu |
CIKM | 1 |
| 2021 | Multi-modal Chorus Recognition for Improving Song Search
Jiaan Wang, Zhixu Li, Binbin Gu, Tingyi Zhang, Qingsheng Liu, Zhigang Chen 0003 |
ICANN (1) | 1 |