EDBT 2026 Demo / reviewers in the wild / expert
Congying Xia
dblp:210/2265
· DBLP profile ↗
27ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0001-7581-0882ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 4 first-author · 14 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking LLMs for Political Science: A United Nations PerspectiveabstractLarge Language Models (LLMs) have achieved significant advances in natural language processing, yet their potential for high-stake political decision-making remains largely unexplored. This paper addresses the gap by focusing on the application of LLMs to the United Nations (UN) decision-making process, where the stakes are particularly high and political decisions can have far-reaching consequences. We introduce a novel dataset comprising publicly available UN Security Council (UNSC) records from 1994 to 2024, including draft resolutions, voting records, and diplomatic speeches. Using this dataset, we propose the United Nations Benchmark (UNBench), the first comprehensive benchmark designed to evaluate LLMs across four interconnected political science tasks: co-penholder judgment, representative voting simulation, draft adoption prediction, and representative statement generation. These tasks span the three stages of the UN decision-making process—drafting, voting, and discussing—and aim to assess LLMs' ability to understand and simulate political dynamics. Our experimental analysis demonstrates the potential and challenges of applying LLMs in this domain, providing insights into their strengths and limitations in political science. To the best of our knowledge, this is the first benchmark to systematically evaluate LLMs in UN decision-making, contributing to the growing intersection of AI and political science. Yueqing Liang, Liangwei Yang, Chen Wang 0018, Congying Xia, Xiongxiao Xu, Haoran Wang 0005, Ali Payani, Kai Shu |
AAAI | 4 |
| 2025 | ReGenesis: LLMs can Grow into Reasoning Generalists via Self-ImprovementabstractPost-training Large Language Models (LLMs) with explicit reasoning trajectories can enhance their reasoning abilities. However, acquiring such high-quality trajectory data typically demands meticulous supervision from humans or superior models, which can be either expensive or license-constrained. In this paper, we explore how far an LLM can improve its reasoning by self-synthesizing reasoning paths as training data without any additional supervision. Existing self-synthesizing methods, such as STaR, suffer from poor generalization to out-of-domain (OOD) reasoning tasks. We hypothesize it is due to that their self-synthesized reasoning paths are too task-specific, lacking general task-agnostic reasoning guidance. To address this, we propose **Reasoning Generalist via Self-Improvement (ReGenesis)**, a method to *self-synthesize reasoning paths as post-training data by progressing from abstract to concrete*. More specifically, ReGenesis self-synthesizes reasoning paths by converting general reasoning guidelines into task-specific ones, generating reasoning structures, and subsequently transforming these structures into reasoning paths, without the need for human-designed task-specific examples used in existing methods. We show that ReGenesis achieves superior performance on all in-domain and OOD settings tested compared to existing methods. For six OOD tasks specifically, while previous methods exhibited an average performance decrease of approximately 4.6% after post training, ReGenesis delivers around 6.1% performance improvement. We also conduct an in-depth analysis of our framework and show ReGenesis is effective across various language models and design choices. Congying Xia, Xinyi Yang 0002, Caiming Xiong, Chien-Sheng Wu, Chen Xing |
ICLR | 2 |
| 2025 | AAAR-1.0: Assessing AI's Potential to Assist ResearchabstractNumerous studies have assessed the proficiency of AI systems, particularly large language models (LLMs), in facilitating everyday tasks such as email writing, question answering, and creative content generation. However, researchers face unique challenges and opportunities in leveraging LLMs for their own work, such as brainstorming research ideas, designing experiments, and writing or reviewing papers. In this study, we introduce AAAR-1.0, a benchmark dataset designed to evaluate LLM performance in three fundamental, expertise-intensive research tasks: (i) EquationInference, assessing the correctness of equations based on the contextual information in paper submissions; (ii) ExperimentDesign, designing experiments to validate research ideas and solutions; and (iii) PaperWeakness, identifying weaknesses in paper submissions. AAAR-1.0 differs from prior benchmarks in two key ways: first, it is explicitly research-oriented, with tasks requiring deep domain expertise; second, it is researcher-oriented, mirroring the primary activities that researchers engage in on a daily basis. An evaluation of both open-source and proprietary LLMs reveals their potential as well as limitations in conducting sophisticated research tasks. We will release the AAAR-1.0 and keep iterating it to new versions. Renze Lou, Hanzi Xu, Jiangshu Du, Ryo Kamoi, Xiaoxin Lu, Yuxuan Sun 0002, Yusen Zhang 0001, Jihyun Janice Ahn, Hongchao Fang, Zhuoyang Zou, Kai Zhang 0033, Congying Xia, Lifu Huang, Wenpeng Yin 0001 |
ICML | 16 |
| 2024 | FOFO: A Benchmark to Evaluate LLMs' Format-Following CapabilityabstractCongying Xia, Chen Xing, Jiangshu Du, Xinyi Yang, Yihao Feng, Ran Xu, Wenpeng Yin, Caiming Xiong. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Congying Xia, Chen Xing, Jiangshu Du, Xinyi Yang 0002, Yihao Feng, Ran Xu 0001, Wenpeng Yin 0001, Caiming Xiong |
ACL (1) | 1 |
| 2024 | LLMs Assist NLP Researchers: Critique Paper (Meta-)ReviewingabstractJiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng, Shuaiqi Liu, Renze Lou, Henry Peng Zou, Pranav Narayanan Venkit, Nan Zhang, Mukund Srinath, Haoran Ranran Zhang, Vipul Gupta, Yinghui Li, Tao Li, Fei Wang, Qin Liu, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang, Ying Su, Raj Sanjay Shah, Ruohao Guo, Jing Gu, Haoran Li, Kangda Wei, Zihao Wang, Lu Cheng, Surangika Ranathunga, Meng Fang, Jie Fu, Fei Liu, Ruihong Huang, Eduardo Blanco, Yixin Cao, Rui Zhang, Philip S. Yu, Wenpeng Yin. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Jiangshu Du, Yibo Wang 0001, Wenting Zhao 0006, Zhongfen Deng, Shuaiqi Liu 0002, Renze Lou, Henry Peng Zou, Pranav Venkit, Mukund Srinath, Ranran Haoran Zhang, Tao Li 0039, Fei Wang 0060, Qin Liu 0010, Tianlin Liu, Pengzhi Gao, Congying Xia, Chen Xing, Cheng Jiayang, Zhaowei Wang 0003, Raj Sanjay Shah, Ruohao Guo, Haoran Li 0003, Kangda Wei, Zihao Wang 0001, Lu Cheng 0001, Surangika Ranathunga, Fei Liu 0004, Ruihong Huang, Eduardo Blanco 0002, Yixin Cao 0002, Rui Zhang 0037, Philip S. Yu, Wenpeng Yin 0001 |
EMNLP | 19 |
| 2023 | Learning to Select from Multiple OptionsabstractMany NLP tasks can be regarded as a selection problem from a set of options, such as classification tasks, multi-choice question answering, etc. Textual entailment (TE) has been shown as the state-of-the-art (SOTA) approach to dealing with those selection problems. TE treats input texts as premises (P), options as hypotheses (H), then handles the selection problem by modeling (P, H) pairwise. Two limitations: first, the pairwise modeling is unaware of other options, which is less intuitive since humans often determine the best options by comparing competing candidates; second, the inference process of pairwise TE is time-consuming, especially when the option space is large. To deal with the two issues, this work first proposes a contextualized TE model (Context-TE) by appending other k options as the context of the current (P, H) modeling. Context-TE is able to learn more reliable decision for the H since it considers various context. Second, we speed up Context-TE by coming up with Parallel-TE, which learns the decisions of multiple options simultaneously. Parallel-TE significantly improves the inference speed while keeping comparable performance with Context-TE. Our methods are evaluated on three tasks (ultra-fine entity typing, intent detection and multi-choice QA) that are typical selection problems with different sizes of options. Experiments show our models set new SOTA performance; particularly, Parallel-TE is faster than the pairwise TE by k times in inference. Jiangshu Du, Wenpeng Yin 0001, Congying Xia, Philip S. Yu |
AAAI | 3 |
| 2023 | Preference-grounded Token-level Guidance for Language Model Fine-tuningabstractAligning language models (LMs) with preferences is an important problem in natural language generation. A key challenge is that preferences are typically provided at the *sequence level* while LM training and generation both occur at the *token level*. There is, therefore, a *granularity mismatch* between the preference and the LM training losses, which may complicate the learning problem. In this paper, we address this issue by developing an alternate training process, where we iterate between grounding the sequence-level preference into token-level training guidance, and improving the LM with the learned guidance. For guidance learning, we design a framework that extends the pairwise-preference learning in imitation learning to both variable-length LM generation and the utilization of the preference among multiple generations. For LM training, based on the amount of supervised data, we present two *minimalist* learning objectives that utilize the learned guidance. In experiments, our method performs competitively on two distinct representative LM tasks --- discrete-prompt generation and text summarization. Shentao Yang, Shujian Zhang, Congying Xia, Yihao Feng, Caiming Xiong, Mingyuan Zhou |
NeurIPS | 3 |
| 2023 | Contrastive Box Embedding for Collaborative ReasoningabstractMost of the existing personalized recommendation methods predict the probability that one user might interact with the next item by matching their representations in the latent space. However, as a cognitive task, it is essential for an impressive recommender system to acquire the cognitive capacity rather than to decide the users' next steps by learning the pattern from the historical interactions through matching-based objectives. Therefore, in this paper, we propose to model the recommendation as a logical reasoning task which is more in line with an intelligent recommender system. Different from the prior works, we embed each query as a box rather than a single point in the vector space, which is able to model sets of users or items enclosed and logical operators (e.g., intersection) over boxes in a more natural manner. Although modeling the logical query with box embedding significantly improves the previous work of reasoning-based recommendation, there still exist two intractable issues including aggregation of box embeddings and training stalemate in critical point of boxes. To tackle these two limitations, we propose a Contrastive Box learning framework for Collaborative Reasoning (CBox4CR). Specifically, CBox4CR combines a smoothed box volume-based contrastive learning objective with the logical reasoning objective to learn the distinctive box representations for the user's preference and the logical query based on the historical interaction sequence. Extensive experiments conducted on four publicly available datasets demonstrate the superiority of our CBox4CR over the state-of-the-art models in recommendation task. Tingting Liang, Yuanqing Zhang, Qianhui Di, Congying Xia, Youhuizi Li, Yuyu Yin |
SIGIR | 4 |
| 2023 | Domain-Invariant Feature Progressive Distillation with Adversarial Adaptive Augmentation for Low-Resource Cross-Domain NERabstractConsidering the expensive annotation in Named Entity Recognition (NER ), Cross-domain NER enables NER in low-resource target domains with few or without labeled data, by transferring the knowledge of high-resource domains. However, the discrepancy between different domains causes the domain shift problem and hampers the performance of cross-domain NER in low-resource scenarios. In this article, we first propose an adversarial adaptive augmentation, where we integrate the adversarial strategy into a multi-task learner to augment and qualify domain adaptive data. We extract domain-invariant features of the adaptive data to bridge the cross-domain gap and alleviate the label-sparsity problem simultaneously. Therefore, another important component in this article is the progressive domain-invariant feature distillation framework. A multi-grained MMD (Maximum Mean Discrepancy) approach in the framework to extract the multi-level domain invariant features and enable knowledge transfer across domains through the adversarial adaptive data. Advanced Knowledge Distillation (KD) schema processes progressively domain adaptation through the powerful pre-trained language models and multi-level domain invariant features. Extensive comparative experiments over four English and two Chinese benchmarks show the importance of adversarial augmentation and effective adaptation from high-resource domains to low-resource target domains. Comparison with two vanilla and four latest baselines indicates the state-of-the-art performance and superiority confronted with both zero-resource and minimal-resource scenarios. Tao Zhang 0055, Congying Xia, Zhiwei Liu 0001, Shu Zhao 0005, Hao Peng 0001, Philip S. Yu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | Transferring From Textual Entailment to Biomedical Named Entity RecognitionabstractBiomedical Named Entity Recognition (BioNER) aims at identifying biomedical entities such as genes, proteins, diseases, and chemical compounds in the given textual data. However, due to the issues of ethics, privacy, and high specialization of biomedical data, BioNER suffers from the more severe problem of lacking in quality labeled data than the general domain especially for the token-level. Facing the extremely limited labeled biomedical data, this work studies the problem of gazetteer-based BioNER, which aims at building a BioNER system from scratch. It needs to identify the entities in the given sentences when we have zero token-level annotations for training. Previous works usually use sequential labeling models to solve the NER or BioNER task and obtain weakly labeled data from gazetteers when we don't have full annotations. However, these labeled data are quite noisy since we need the labels for each token and the entity coverage of the gazetteers is limited. Here we propose to formulate the BioNER task as a Textual Entailment problem and solve the task via Textual Entailment with Dynamic Contrastive learning (TEDC). TEDC not only alleviates the noisy labeling issue, but also transfers the knowledge from pre-trained textual entailment models. Additionally, the dynamic contrastive learning framework contrasts the entities and non-entities in the same sentence and improves the model's discrimination ability. Experiments on two real-world biomedical datasets show that TEDC can achieve state-of-the-art performance for gazetteer-based BioNER. Tingting Liang, Congying Xia, Ziqiang Zhao, Yixuan Jiang, Yuyu Yin, Philip S. Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | Modeling Reviews for Few-Shot Recommendation via Enhanced Prototypical NetworkabstractAlthough some existing models are proposed to exploit reviews for improving performance for recommender systems, few of them can handle the following issues led by the insufficient review data: (i) The regular training process does not exactly fit the scenario of preference prediction with few historical behaviors. (ii) Extracting informative and sufficient semantic features from limited review texts is a challenging work. To alleviate these issues, this paper proposes an enhanced prototypical network, FS-EPN, that leverages reviews for recommendation under the few-shot setting. FS-EPN consists of an attentional prototypical network being the basic architecture, a sentiment encoder and a memory collector cooperating to capture the extra sentimental and collaborative information from both user and item perspectives for semantic information supplement. We train FS-EPN under the meta-learning framework, which models the training process in the episodic manner to mimic the few-shot test environment. Extensive experiments conducted on six publicly available datasets demonstrate the superior capability of FS-EPN over several state-of-the-art models in few-shot recommendation. Tingting Liang, Congying Xia, Ziqiang Zhao, Yuyu Yin, Liang Chen 0001, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Continuous Prompt Tuning Based Textual Entailment Model for E-commerce Entity TypingabstractThe explosion of e-commerce has caused the need for processing and analysis of product titles, like entity typing in product titles. However, the rapid activity in e-commerce has led to the rapid emergence of new entities, which is difficult for general entity typing. Besides, product titles in e-commerce have very different language styles from text data in general domain. In order to handle new entities in product titles and address the special language styles of product titles in e-commerce domain, we propose our textual entailment model with continuous prompt tuning based hypotheses and fusion embeddings for e-commerce entity typing. First, we reformulate entity typing into a textual entailment problem to handle new entities that are not present during training. Second, we design a model to automatically generate textual entailment hypotheses using a continuous prompt tuning method, which can generate better textual entailment hypotheses without manual design. Third, we utilize the fusion embeddings of BERT embedding and Char-acterBERT embedding to solve the problem that the language styles of product titles in e-commerce are different from that of general domain. To analyze the effect of each contribution, we compare the performance of entity typing and textual entailment model, and conduct ablation studies on continuous prompt tuning and fusion embeddings. We also evaluate the impact of different prompt template initialization for the continuous prompt tuning. We show our proposed model improves the average F1 score by around 2% compared to the baseline BERT entity typing model. Yibo Wang 0001, Congying Xia, Philip S. Yu |
IEEE Big Data | 2 |
| 2022 | Content-aware Recommendation via Dynamic Heterogeneous Graph Convolutional Network
Tingting Liang, Lin Ma 0002, Congying Xia, Yuyu Yin |
Knowl. Based Syst. | 5 |
| 2022 | A Survey on Text Classification: From Traditional to Deep LearningabstractText classification is the most fundamental and essential task in natural language processing. The last decade has seen a surge of research in this area due to the unprecedented success of deep learning. Numerous methods, datasets, and evaluation metrics have been proposed in the literature, raising the need for a comprehensive and updated survey. This paper fills the gap by reviewing the state-of-the-art approaches from 1961 to 2021, focusing on models from traditional models to deep learning. We create a taxonomy for text classification according to the text involved and the models used for feature extraction and classification. We then discuss each of these categories in detail, dealing with both the technical developments and benchmark datasets that support tests of predictions. A comprehensive comparison between different techniques, as well as identifying the pros and cons of various evaluation metrics are also provided in this survey. Finally, we conclude by summarizing key implications, future research directions, and the challenges facing the research area. Qian Li 0033, Hao Peng 0001, Jianxin Li 0002, Congying Xia, Renyu Yang, Lichao Sun 0001, Philip S. Yu, Lifang He 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2021 | HETFORMER: Heterogeneous Transformer with Sparse Attention for Long-Text Extractive SummarizationabstractTo capture the semantic graph structure from raw text, most existing summarization approaches are built on GNNs with a pre-trained model.However, these methods suffer from cumbersome procedures and inefficient computations for long-text documents.To mitigate these issues, this paper proposes HET-FORMER, a Transformer-based pre-trained model with multi-granularity sparse attentions for long-text extractive summarization.Specifically, we model different types of semantic nodes in raw text as a potential heterogeneous graph and directly learn heterogeneous relationships (edges) among nodes by Transformer.Extensive experiments on both single-and multi-document summarization tasks show that HETFORMER achieves stateof-the-art performance in Rouge F1 while using less memory and fewer parameters. Ye Liu 0006, Jianguo Zhang 0005, Yao Wan 0001, Congying Xia, Lifang He 0001, Philip S. Yu |
EMNLP (1) | 4 |
| 2021 | Few-Shot Intent Detection via Contrastive Pre-Training and Fine-TuningabstractJianguo Zhang, Trung Bui, Seunghyun Yoon, Xiang Chen, Zhiwei Liu, Congying Xia, Quan Hung Tran, Walter Chang, Philip Yu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Jianguo Zhang 0005, Trung Bui, Seunghyun Yoon 0002, Xiang Chen 0010, Zhiwei Liu 0001, Congying Xia, Quan Hung Tran, Walter Chang, Philip S. Yu |
EMNLP (1) | 6 |
| 2021 | PDALN: Progressive Domain Adaptation over a Pre-trained Model for Low-Resource Cross-Domain Named Entity RecognitionabstractCross-domain Named Entity Recognition (NER) transfers the NER knowledge from high-resource domains to the low-resource target domain.Due to limited labeled resources and domain shift, cross-domain NER is a challenging task.To address these challenges, we propose a progressive domain adaptation Knowledge Distillation (KD) approach -PDALN.It achieves superior domain adaptability by employing three components: (1) Adaptive data augmentation techniques, which alleviate cross-domain gap and label sparsity simultaneously; (2) Multi-level Domain invariant features, derived from a multigrained MMD (Maximum Mean Discrepancy) approach, to enable knowledge transfer across domains; (3) Advanced KD schema, which progressively enables powerful pre-trained language models to perform domain adaptation.Extensive experiments on four benchmarks show that PDALN can effectively adapt highresource domains to low-resource target domains, even if they are diverse in terms and writing styles.Comparison with other baselines indicates the state-of-the-art performance of PDALN. Tao Zhang 0055, Congying Xia, Philip S. Yu, Zhiwei Liu 0001, Shu Zhao 0005 |
EMNLP (1) | 2 |
| 2021 | Incremental Few-shot Text Classification with Multi-round New Classes: Formulation, Dataset and SystemabstractText classification is usually studied by labeling natural language texts with relevant categories from a predefined set.In the real world, new classes might keep challenging the existing system with limited labeled data.The system should be intelligent enough to recognize upcoming new classes with a few examples.In this work, we define a new task in the NLP domain, incremental few-shot text classification, where the system incrementally handles multiple rounds of new classes.For each round, there is a batch of new classes with a few labeled examples per class.Two major challenges exist in this new task: (i) For the learning process, the system should incrementally learn new classes round by round without re-training on the examples of preceding classes; (ii) For the performance, the system should perform well on new classes without much loss on preceding classes.In addition to formulating the new task, we also release two benchmark datasets 1 in the incremental fewshot setting: intent classification and relation classification.Moreover, we propose two entailment approaches, ENTAILMENT and HY-BRID, which show promise for solving this novel problem. Congying Xia, Wenpeng Yin 0001, Yihao Feng, Philip S. Yu |
NAACL-HLT | 1 |
| 2021 | User Preference-aware Fake News DetectionabstractDisinformation and fake news have posed detrimental effects on individuals and society in recent years, attracting broad attention to fake news detection. The majority of existing fake news detection algorithms focus on mining news content and/or the surrounding exogenous context for discovering deceptive signals; while the endogenous preference of a user when he/she decides to spread a piece of fake news or not is ignored. The confirmation bias theory has indicated that a user is more likely to spread a piece of fake news when it confirms his/her existing beliefs/preferences. Users' historical, social engagements such as posts provide rich information about users' preferences toward news and have great potentials to advance fake news detection. However, the work on exploring user preference for fake news detection is somewhat limited. Therefore, in this paper, we study the novel problem of exploiting user preference for fake news detection. We propose a new framework, UPFD, which simultaneously captures various signals from user preferences by joint content and graph modeling. Experimental results on real-world datasets demonstrate the effectiveness of the proposed framework. We release our code and data as a benchmark for GNN-based fake news detection: https://github.com/safe-graph/GNN-FakeNews. Yingtong Dou, Kai Shu, Congying Xia, Philip S. Yu, Lichao Sun 0001 |
SIGIR | 3 |
| 2021 | Pseudo Siamese Network for Few-shot Intent GenerationabstractFew-shot intent detection is a challenging task due to the scare annotation problem. In this paper, we propose a Pseudo Siamese Network (PSN) to generate labeled data for few-shot intents and alleviate this problem. PSN consists of two identical subnetworks with the same structure but different weights: an action network and an object network. Each subnetwork is a transformer-based variational autoencoder that tries to model the latent distribution of different components in the sentence. The action network is learned to understand action tokens and the object network focuses on object-related expressions. It provides an interpretable framework for generating an utterance with an action and an object existing in a given intent. Experiments on two real-world datasets show that PSN achieves state-of-the-art performance for the generalized few shot intent detection task. Congying Xia, Caiming Xiong, Philip S. Yu |
SIGIR | 1 |
| 2020 | Hierarchical Bi-Directional Self-Attention Networks for Paper Review Rating RecommendationabstractReview rating prediction of text reviews is a rapidly growing technology with a wide range of applications in natural language processing.However, most existing methods either use handcrafted features or learn features using deep learning with simple text corpus as input for review rating prediction, ignoring the hierarchies among data.In this paper, we propose a Hierarchical bi-directional self-attention Network framework (HabNet) for paper review rating prediction and recommendation, which can serve as an effective decision-making tool for the academic paper review process.Specifically, we leverage the hierarchical structure of the paper reviews with three levels of encoders: sentence encoder (level one), intra-review encoder (level two) and interreview encoder (level three).Each encoder first derives contextual representation of each level, then generates a higher-level representation, and after the learning process, we are able to identify useful predictors to make the final acceptance decision, as well as to help discover the inconsistency between numerical review ratings and text sentiment conveyed by reviewers.Furthermore, we introduce two new metrics to evaluate models in data imbalance situations.Extensive experiments on a publicly available dataset (PeerRead) and our own collected dataset (OpenReview) demonstrate the superiority of the proposed approach compared with state-of-the-art methods. Zhongfen Deng, Hao Peng 0001, Congying Xia, Jianxin Li 0002, Lifang He 0001, Philip S. Yu |
COLING | 3 |
| 2020 | Mixup-Transformer: Dynamic Data Augmentation for NLP TasksabstractMixup (Zhang et al., 2017) is a latest data augmentation technique that linearly interpolates input examples and the corresponding labels.It has shown strong effectiveness in image classification by interpolating images at the pixel level.Inspired by this line of research, in this paper, we explore: i) how to apply mixup to natural language processing tasks since text data can hardly be mixed in the raw format; ii) if mixup is still effective in transformer-based learning models, e.g., BERT.To achieve the goal, we incorporate mixup to transformer-based pre-trained architecture, named "mixup-transformer", for a wide range of NLP tasks while keeping the whole end-to-end training system.We evaluate the proposed framework by running extensive experiments on the GLUE benchmark.Furthermore, we also examine the performance of mixup-transformer in low-resource scenarios by reducing the training data with a certain ratio.Our studies show that mixup is a domain-independent data augmentation technique to pre-trained language models, resulting in significant performance improvement for transformer-based models. Lichao Sun 0001, Congying Xia, Wenpeng Yin 0001, Tingting Liang, Philip S. Yu, Lifang He 0001 |
COLING | 2 |
| 2020 | MZET: Memory Augmented Zero-Shot Fine-grained Named Entity TypingabstractNamed entity typing (NET) is a classification task of assigning an entity mention in the context with given semantic types.However, with the growing size and granularity of the entity types, few previous researches concern with newly emerged entity types.In this paper, we propose MZET, a novel memory augmented FNET (Fine-grained NET) model, to tackle the unseen types in a zero-shot manner.MZET incorporates character-level, word-level, and contextural-level information to learn the entity mention representation.Besides, MZET considers the semantic meaning and the hierarchical structure into the entity type representation.Finally, through the memory component which models the relationship between the entity mention and the entity type, MZET transfers the knowledge from seen entity types to the zero-shot ones.Extensive experiments on three public datasets show the superior performance obtained by MZET, which surpasses the state-of-the-art FNET neural network models with up to 8% gain in Micro-F1 and Macro-F1 score. Tao Zhang 0055, Congying Xia, Chun-Ta Lu, Philip S. Yu |
COLING | 2 |
| 2020 | Joint Training Capsule Network for Cold Start RecommendationabstractThis paper proposes a novel neural network, joint training capsule network (JTCN), for the cold start recommendation task. We propose to mimic the high-level user preference other than the raw interaction history based on the side information for the fresh users. Specifically, an attentive capsule layer is proposed to aggregate high-level user preference from the low-level interaction history via a dynamic routing-by-agreement mechanism. Moreover, JTCN jointly trains the loss for mimicking the user preference and the softmax loss for the recommendation together in an end-to-end manner. Experiments on two publicly available datasets demonstrate the effectiveness of the proposed model. JTCN improves other state-of-the-art methods at least 7.07% for CiteULike and 16.85% for Amazon in terms of [email protected] in cold start recommendation. Tingting Liang, Congying Xia, Yuyu Yin, Philip S. Yu |
SIGIR | 2 |
| 2019 | Multi-grained Named Entity RecognitionabstractThis paper presents a novel framework, MGNER, for Multi-Grained Named Entity Recognition where multiple entities or entity mentions in a sentence could be nonoverlapping or totally nested.Different from traditional approaches regarding NER as a sequential labeling task and annotate entities consecutively, MGNER detects and recognizes entities on multiple granularities: it is able to recognize named entities without explicitly assuming non-overlapping or totally nested structures.MGNER consists of a Detector that examines all possible word segments and a Classifier that categorizes entities.In addition, contextual information and a self-attention mechanism are utilized throughout the framework to improve the NER performance.Experimental results show that MGNER outperforms current state-of-the-art baselines up to 4.4% in terms of the F1 score among nested/non-overlapping NER tasks.* Work was done when the author Yaliang Li was at Tencent America. Congying Xia, Tao Yang 0012, Yaliang Li, Nan Du 0001, Xian Wu 0001, Wei Fan 0001, Fenglong Ma, Philip S. Yu |
ACL (1) | 1 |
| 2018 | Zero-shot User Intent Detection via Capsule Neural NetworksabstractUser intent detection plays a critical role in question-answering and dialog systems.Most previous works treat intent detection as a classification problem where utterances are labeled with predefined intents.However, it is labor-intensive and time-consuming to label users' utterances as intents are diversely expressed and novel intents will continually be involved.Instead, we study the zero-shot intent detection problem, which aims to detect emerging user intents where no labeled utterances are currently available.We propose two capsule-based architectures: INTENT-CAPSNET that extracts semantic features from utterances and aggregates them to discriminate existing intents, and INTENTCAPSNET-ZSL which gives INTENTCAPSNET the zero-shot learning ability to discriminate emerging intents via knowledge transfer from existing intents.Experiments on two real-world datasets show that our model not only can better discriminate diversely expressed existing intents, but is also able to discriminate emerging intents when no labeled utterances are available. Congying Xia, Yi Chang 0001, Philip S. Yu |
EMNLP | 1 |
| 2017 | BL-MNE: Emerging Heterogeneous Social Network Embedding Through Broad Learning with Aligned AutoencoderabstractNetwork embedding aims at projecting the network data into a low-dimensional feature space, where the nodes are represented as a unique feature vector and network structure can be effectively preserved. In recent years, more and more online application service sites can be represented as massive and complex networks, which are extremely challenging for traditional machine learning algorithms to deal with. Effective embedding of the complex network data into low-dimension feature representation can both save data storage space and enable traditional machine learning algorithms applicable to handle the network data. Network embedding performance will degrade greatly if the networks are of a sparse structure, like the emerging networks with few connections. In this paper, we propose to learn the embedding representation for a target emerging network based on the broad learning setting, where the emerging network is aligned with other external mature networks at the same time. To solve the problem, a new embedding framework, namely "Deep alIgned autoencoder based eMbEdding" (DIME), is introduced in this paper. DIME handles the diverse link and attribute in a unified analytic based on broad learning, and introduces the multiple aligned attributed heterogeneous social network concept to model the network structure. A set of meta paths are introduced in the paper, which define various kinds of connections among users via the heterogeneous link and attribute information. The closeness among users in the networks are defined as the meta proximity scores, which will be fed into DIME to learn the embedding vectors of users in the emerging network. Extensive experiments have been done on real-world aligned social networks, which have demonstrated the effectiveness of DIME in learning the emerging network embedding vectors. Jiawei Zhang 0001, Congying Xia, Limeng Cui, Yanjie Fu, Philip S. Yu |
ICDM | 2 |