Xinyu Dai

dblp:39/5815 · also Xin-Yu Dai · DBLP profile ↗
← Back
107ranked-venue papers
2as first author
49since 2021 · last 2026
0000-0002-4139-7337ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 85 · 2 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 12 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021
YearPublicationVenuePosition
2026 Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning
abstract
Rerankers play a pivotal role in refining retrieval results for Retrieval-Augmented Generation.However, current reranking models are typically optimized on static human annotated relevance labels in isolation, decoupled from the downstream generation process.This isolation leads to a fundamental misalignment: documents identified as topically relevant by information retrieval metrics often fail to provide the actual utility required by the LLM for precise answer generation.To bridge this gap, we introduce ReRanking Preference Optimization (RRPO) 1 , a reinforcement learning framework that directly aligns reranking with the LLM's generation quality.By formulating reranking as a sequential decision-making process, RRPO optimizes for context utility using LLM feedback, thereby eliminating the need for expensive human annotations.To ensure training stability, we further introduce a reference-anchored deterministic baseline.Extensive experiments on knowledge-intensive benchmarks demonstrate that RRPO significantly outperforms strong baselines, including the powerful list-wise reranker RankZephyr.Further analysis highlights the versatility of our framework: it generalizes seamlessly to diverse readers (e.g., GPT-4o), integrates orthogonally with query expansion modules like Query2Doc, and remains robust even when trained with noisy supervisors.
Xiangqing Shen, Fanfan Wang, Cangqi Zhou, Zhen Wu 0002, Xinyu Dai
ACL (1)6
2026 Unsupervised deep hashing based on multi-scale aggregation and optimal transport matching for image retrieval
Lei Ma 0004, Hao Pei, Lei Wang 0068, Ying Zhu 0002, Yu Shi 0004, Hanyu Hong, Xinyu Dai, Fanman Meng, Qingbo Wu 0001
Neurocomputing7
2026 Multi-view auxiliary modeling for stable LLM adaptation under view disagreement
Xinyu Dai, Jia-Jun Chen
Pattern Recognit.6
2025 Bridging Questions and Charts: A Weakly Supervised Alignment Model for Chart Question Answering
Jiangzhou Ju, Yunlin Mao, Zhen Wu 0002, Robert Ridley, Jiajun Chen 0001, Xinyu Dai
NLPCC (3)6
2025 Multi-candidate Speculative Decoding
Sen Yang 0015, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
NLPCC (1)3
2025 Deep neural networks for knowledge-enhanced molecular modeling
Siyu Long, Yi Zhou 0018, Fan Sha, Xinyu Dai
Neurocomputing5
2025 Persona-Aware Alignment Framework for Personalized Dialogue Generation
abstract
Abstract Personalized dialogue generation aims to leverage persona profiles and dialogue history to generate persona-relevant and consistent responses. Mainstream models typically rely on token-level language model training with persona dialogue data, such as Next Token Prediction, to implicitly achieve personalization, meaning that these methods tend to neglect the given personas and generate generic responses. To address this issue, we propose a novel Persona-Aware Alignment Framework (PAL), which directly treats persona alignment as the training objective of dialogue generation. Specifically, PAL employs a two-stage training method including Persona-Aware Learning and Persona Alignment, equipped with an easy-to-use inference strategy Select then Generate, to improve persona sensitivity and generate more persona-relevant responses at the semantics level. Through extensive experiments, we demonstrate that our framework outperforms many state-of-the-art personalized dialogue methods and large language models.
Guanrong Li, Zhen Wu 0002, Xinyu Dai
Trans. Assoc. Comput. Linguistics4
2024 Dr3: Ask Large Language Models Not to Give Off-Topic Answers in Open Domain Multi-Hop Question Answering
abstract
Open Domain Multi-Hop Question Answering (ODMHQA) plays a crucial role in Natural Language Processing (NLP) by aiming to answer complex questions through multi-step reasoning over retrieved information from external knowledge sources. Recently, Large Language Models (LLMs) have demonstrated remarkable performance in solving ODMHQA owing to their capabilities including planning, reasoning, and utilizing tools. However, LLMs may generate off-topic answers when attempting to solve ODMHQA, namely the generated answers are irrelevant to the original questions. This issue of off-topic answers accounts for approximately one-third of incorrect answers, yet remains underexplored despite its significance. To alleviate this issue, we propose the Discriminate→Re-Compose→Re- Solve→Re-Decompose (Dr3) mechanism. Specifically, the Discriminator leverages the intrinsic capabilities of LLMs to judge whether the generated answers are off-topic. In cases where an off-topic answer is detected, the Corrector performs step-wise revisions along the reversed reasoning chain (Re-Compose→Re-Solve→Re-Decompose) until the final answer becomes on-topic. Experimental results on the HotpotQA and 2WikiMultiHopQA datasets demonstrate that our Dr3 mechanism considerably reduces the occurrence of off-topic answers in ODMHQA by nearly 13%, improving the performance in Exact Match (EM) by nearly 3% compared to the baseline method without the Dr3 mechanism.
Yiheng Zhu 0004, Yuanbin Cao, Yinzhi Zhou, Zhen Wu 0002, Shenglan Wu, Haoyuan Hu, Xinyu Dai
LREC/COLING9
2024 PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment
abstract
Large language models demonstrate reasonable multilingual abilities, despite predominantly English-centric pretraining.However, the spontaneous multilingual alignment in these models is shown to be weak, leading to unsatisfactory cross-lingual transfer and knowledge sharing.Previous works attempt to address this issue by explicitly injecting multilingual alignment information during or after pretraining.Thus for the early stage in pretraining, the alignment is weak for sharing information or knowledge across languages.In this paper, we propose PREALIGN, a framework that establishes multilingual alignment prior to language model pretraining.PREALIGN injects multilingual alignment by initializing the model to generate similar representations of aligned words and preserves this alignment using a code-switching strategy during pretraining.Extensive experiments in a synthetic English to English-Clone setting demonstrate that PREALIGN significantly outperforms standard multilingual joint training in language modeling, zero-shot crosslingual transfer, and cross-lingual knowledge application.Further experiments in real-world scenarios further validate PREALIGN's effectiveness across various languages and model sizes.
Jiahuan Li, Shujian Huang, Aarron Ching, Xinyu Dai, Jiajun Chen 0001
EMNLP4
2024 EFUF: Efficient Fine-Grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs) have attracted increasing attention in the past few years, but they may still generate descriptions that include objects not present in the corresponding images, a phenomenon known as object hallucination.To eliminate hallucinations, existing methods manually annotate paired responses with and without hallucinations, and then employ various alignment algorithms to improve the alignment capability between images and text.However, they not only demand considerable computation resources during the finetuning stage but also require expensive human annotation to construct paired data needed by the alignment algorithms.To address these issues, we propose an efficient fine-grained unlearning framework (EFUF), which performs gradient ascent utilizing three tailored losses to eliminate hallucinations without paired data.Extensive experiments show that our method consistently reduces hallucinations while preserving the generation quality with modest computational overhead.Our code and datasets are available at https://github.com/starreeze/efuf.
Shangyu Xing, Fei Zhao 0012, Zhen Wu 0002, Tuo An, Xinyu Dai
EMNLP8
2024 Mol-AE: Auto-Encoder Based Molecular Representation Learning With 3D Cloze Test Objective
abstract
3D molecular representation learning has gained tremendous interest and achieved promising performance in various downstream tasks. A series of recent approaches follow a prevalent framework: an encoder-only model coupled with a coordinate denoising objective. However, through a series of analytical experiments, we prove that the encoder-only model with coordinate denoising objective exhibits inconsistency between pre-training and downstream objectives, as well as issues with disrupted atomic identifiers. To address these two issues, we propose Mol-AE for molecular representation learning, an auto-encoder model using positional encoding as atomic identifiers. We also propose a new training objective named 3D Cloze Test to make the model learn better atom spatial relationships from real molecular substructures. Empirical results demonstrate that Mol-AE achieves a large margin performance gain compared to the current state-of-the-art 3D molecular modeling approach.
Kangjie Zheng, Siyu Long, Zaiqing Nie, Ming Zhang 0004, Xinyu Dai, Wei-Ying Ma, Hao Zhou 0012
ICML6
2024 ESM All-Atom: Multi-Scale Protein Language Model for Unified Molecular Modeling
abstract
Protein language models have demonstrated significant potential in the field of protein engineering. However, current protein language models primarily operate at the residue scale, which limits their ability to provide information at the atom level. This limitation prevents us from fully exploiting the capabilities of protein language models for applications involving both proteins and small molecules. In this paper, we propose ESM-AA (ESM All-Atom), a novel approach that enables atom-scale and residue-scale unified molecular modeling. ESM-AA achieves this by pre-training on multi-scale code-switch protein sequences and utilizing a multi-scale position encoding to capture relationships among residues and atoms. Experimental results indicate that ESM-AA surpasses previous methods in protein-molecule tasks, demonstrating the full utilization of protein language models. Further investigations reveal that through unified molecular modeling, ESM-AA not only gains molecular knowledge but also retains its understanding of proteins.
Kangjie Zheng, Siyu Long, Tianyu Lu, Xinyu Dai, Ming Zhang 0004, Zaiqing Nie, Wei-Ying Ma, Hao Zhou 0012
ICML5
2024 Focus & Gating: A Multimodal Approach for Unveiling Relations in Noisy Social Media
abstract
Multimedia content's surge on the internet has made multimodal relation extraction vital for applications like intelligent search and knowledge graph construction. As a rich source of image-text data, social media plays a crucial role in populating knowledge bases. However, the noisy information present in social media poses a challenge in multimodal relation extraction. Current methods focus on extracting relevant information from images to improve model performance but often overlook the importance of global image information. In this paper, we propose a novel multimodal relation extraction method FocalMRE, which leverages image focal augmentation, focal attention, and gating mechanisms. FocalMRE enables the model to concentrate on the image's focal regions while effectively utilizing the global information in the image. Through gating mechanisms, FocalMRE optimizes the multimodal fusion strategy, allowing the model to select the most relevant augmented regions for overcoming noise interference in relation extraction. The experimental results on the public MNRE dataset reveal that FocalMRE exhibits robust and significant performance advantages in the multimodal relation extraction task, especially in scenarios with high noise, long-tail distributions, and limited resources. The code is available at https://github.com/NJUNLP/FocalMRE.
Liang He 0009, Hongke Wang, Zhen Wu 0002, Xinyu Dai, Jiajun Chen 0001
ACM Multimedia5
2024 AutoSurvey: Large Language Models Can Automatically Write Surveys
abstract
This paper introduces AutoSurvey, a speedy and well-organized methodology for automating the creation of comprehensive literature surveys in rapidly evolving fields like artificial intelligence. Traditional survey paper creation faces challenges due to the vast volume and complexity of information, prompting the need for efficient survey methods. While large language models (LLMs) offer promise in automating this process, challenges such as context window limitations, parametric knowledge constraints, and the lack of evaluation benchmarks remain. AutoSurvey addresses these challenges through a systematic approach that involves initial retrieval and outline generation, subsection drafting by specialized LLMs, integration and refinement, and rigorous evaluation and iteration. Our contributions include a comprehensive solution to the survey problem, a reliable evaluation method, and experimental validation demonstrating AutoSurvey's effectiveness.
Yidong Wang 0003, Wenjin Yao, Xin Zhang 0097, Zhen Wu 0002, Meishan Zhang, Xinyu Dai, Min Zhang 0005, Qingsong Wen, Wei Ye 0004, Shikun Zhang, Yue Zhang 0004
NeurIPS8
2023 On Prefix-tuning for Lightweight Out-of-distribution Detection
abstract
Out-of-distribution (OOD) detection, a fundamental task vexing real-world applications, has attracted growing attention in the NLP community.Recently fine-tuning based methods have made promising progress.However, it could be costly to store fine-tuned models for each scenario.In this paper, we depart from the classic fine-tuning based OOD detection toward a parameter-efficient alternative, and propose an unsupervised prefix-tuning based OOD detection framework termed PTO.Additionally, to take advantage of optional training data labels and targeted OOD data, two practical extensions of PTO are further proposed.Overall, PTO and its extensions offer several key advantages of being lightweight, easy-to-reproduce, and theoretically justified.Experimental results show that our methods perform comparably to, even better than, existing fine-tuning based OOD detection approaches under a wide range of metrics, detection settings, and OOD types.
Yawen Ouyang, Yongchang Cao, Zhen Wu 0002, Xinyu Dai
ACL (1)6
2023 Local Interpretation of Transformer Based on Linear Decomposition
abstract
In recent years, deep neural networks (DNNs) have achieved state-of-the-art performance on a wide range of tasks.However, limitations in interpretability have hindered their applications in the real world.This work proposes to interpret neural networks by linear decomposition and finds that the ReLU-activated Transformer can be considered as a linear model on a single input.We further leverage the linearity of the model and propose a linear decomposition of the model output to generate local explanations.Our evaluation of sentiment classification and machine translation shows that our method achieves competitive performance in efficiency and fidelity of explanation.In addition, we demonstrate the potential of our approach in applications with examples of error analysis on multiple tasks.
Sen Yang 0015, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
ACL (1)5
2023 Addressing Linguistic Bias through a Contrastive Analysis of Academic Writing in the NLP Domain
abstract
It has been well documented that a reviewer's opinion of the nativeness of expression in an academic paper affects the likelihood of it being accepted for publication.Previous works have also shone a light on the stress and anxiety authors who are non-native English speakers experience when attempting to publish in international venues.We explore how this might be a concern in the field of Natural Language Processing (NLP) through conducting a comprehensive statistical analysis of NLP paper abstracts, identifying how authors of different linguistic backgrounds differ in the lexical, morphological, syntactic and cohesive aspects of their writing.Through our analysis, we identify that there are a number of characteristics that are highly variable across the different corpora examined in this paper.This indicates potential for the presence of linguistic bias.Therefore, we outline a set of recommendations to publishers of academic journals and conferences regarding their guidelines and resources for prospective authors in order to help enhance inclusivity and fairness.
Robert Ridley, Zhen Wu 0002, Shujian Huang, Xinyu Dai
EMNLP5
2023 M2DF: Multi-grained Multi-curriculum Denoising Framework for Multimodal Aspect-based Sentiment Analysis
abstract
Multimodal Aspect-based Sentiment Analysis (MABSA) is a fine-grained Sentiment Analysis task, which has attracted growing research interests recently.Existing work mainly utilizes image information to improve the performance of MABSA task.However, most of the studies overestimate the importance of images since there are many noisy images unrelated to the text in the dataset, which will have a negative impact on model learning.Although some work attempts to filter low-quality noisy images by setting thresholds, relying on thresholds will inevitably filter out a lot of useful image information.Therefore, in this work, we focus on whether the negative impact of noisy images can be reduced without filtering the data.To achieve this goal, we borrow the idea of Curriculum Learning and propose a Multi-grained Multi-curriculum Denoising Framework (M2DF), which can achieve denoising by adjusting the order of training data.Extensive experimental results show that our framework consistently outperforms state-ofthe-art work on three sub-tasks of MABSA.Our code and datasets are available at https: //github.com/grandchicken/M2DF.
Fei Zhao 0012, Zhen Wu 0002, Yawen Ouyang, Xinyu Dai
EMNLP6
2023 An Empirical Study and Improvement for Speech Emotion Recognition
abstract
Multimodal speech emotion recognition aims to detect speakers’ emotions from audio and text. Prior works mainly focus on exploiting advanced networks to model and fuse different modality information to facilitate performance, while neglecting the effect of different fusion strategies on emotion recognition. In this work, we consider a simple yet important problem: how to fuse audio and text modality information is more helpful for this multimodal task. Further, we propose a multimodal emotion recognition model improved by perspective loss. Empirical results show our method obtained new state-of-the-art results on the IEMOCAP dataset. The in-depth analysis explains why the improved model can achieve improvements and outperforms baselines.
Zhen Wu 0002, Yizhe Lu, Xinyu Dai
ICASSP3
2023 Make BERT-based Chinese Spelling Check Model Enhanced by Layerwise Attention and Gaussian Mixture Model
abstract
BERT-based models have shown a remarkable ability in the Chinese Spelling Check (CSC) task recently. However, traditional BERT-based methods still suffer from two limitations. First, although previous works have identified that explicit prior knowledge like Part-Of-Speech (POS) tagging can benefit in the CSC task, they neglected the fact that spelling errors inherent in CSC data can lead to incorrect tags and therefore mislead models. Additionally, they ignored the correlation between the implicit hierarchical information encoded by BERT's intermediate layers and different linguistic phenomena. This results in sub-optimal accuracy. To alleviate the above two issues, we design a heterogeneous knowledge-infused framework to strengthen BERT-based CSC models. To incorporate explicit POS knowledge, we utilize an auxiliary task strategy driven by Gaussian mixture model. Meanwhile, to incorporate implicit hierarchical linguistic knowledge within the encoder, we propose a novel form of n-gram-based layerwise self-attention to generate a multilayer representation. Experimental results show that our proposed framework yields a stable performance boost over four strong baseline models and outperforms the previous state-of-the-art methods on two datasets.
Yongchang Cao, Liang He 0009, Zhen Wu 0002, Xinyu Dai
IJCNN4
2023 MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark Evaluation
abstract
Extracting relational facts from multimodal data is a crucial task in the field of multimedia and knowledge graphs that feeds into widespread real-world applications. The emphasis of recent studies centers on recognizing relational facts in which both entities are present in one modality and supplementary information is used from other modalities. However, such works disregard a substantial amount of multimodal relational facts that arise across different modalities, such as one entity seen in a text and another in an image. In this paper, we propose a new task, namely Multimodal Object-Entity Relation Extraction, which aims to extract "object-entity" relational facts from image and text data. To facilitate research on this task, we introduce MORE, a new dataset comprising 21 relation types and 20,136 multimodal relational facts annotated on 3,522 pairs of textual news titles and corresponding images. To show the challenges of Multimodal Object-Entity Relation Extraction, we evaluated recent state-of-the-art methods for multimodal relation extraction and conducted a comprehensive experimentation analysis on MORE. Our results demonstrate significant challenges for existing methods, underlining the need for further research on this task. Based on our experiments, we identify several promising directions for future research. The MORE dataset and code are available at https://github.com/NJUNLP/MORE.
Liang He 0009, Hongke Wang, Yongchang Cao, Zhen Wu 0002, Xinyu Dai
ACM Multimedia6
2023 DRIN: Dynamic Relation Interactive Network for Multimodal Entity Linking
abstract
Multimodal Entity Linking (MEL) is a task that aims to link ambiguous mentions within multimodal contexts to referential entities in a multimodal knowledge base. Recent methods for MEL adopt a common framework: they first interact and fuse the text and image to obtain representations of the mention and entity respectively, and then compute the similarity between them to predict the correct entity. However, these methods still suffer from two limitations: first, as they fuse the features of text and image before matching, they cannot fully exploit the fine-grained alignment relations between the mention and entity. Second, their alignment is static, leading to low performance when dealing with complex and diverse data. To address these issues, we propose a novel framework called Dynamic Relation Interactive Network (DRIN) for MEL tasks. DRIN explicitly models four different types of alignment between a mention and entity and builds a dynamic Graph Convolutional Network (GCN) to dynamically select the corresponding alignment relations for different input samples. Experiments on two datasets show that DRIN outperforms state-of-the-art methods by a large margin, demonstrating the effectiveness of our approach. Our code and datasets are publicly available.
Shangyu Xing, Fei Zhao 0012, Zhen Wu 0002, Xinyu Dai
ACM Multimedia6
2023 Episode-Based Prompt Learning for Any-Shot Intent Detection
Dingjie Song, Yawen Ouyang, Zhen Wu 0002, Xinyu Dai
NLPCC (1)5
2023 IDOS: A Unified Debiasing Method via Word Shuffling
Yuanhang Tang, Yawen Ouyang, Zhen Wu 0002, Xinyu Dai
NLPCC (2)6
2023 Label-Correction Capsule Network for Hierarchical Text Classification
abstract
Hierarchical Text Classification (HTC) aims to predict the category of a document in a given label hierarchy. Considering a parent-child relationship among labels at different levels, previous works mainly leverage the parent-level label information to guide the child-level classification and achieve promising results. However, they still suffer from two drawbacks: (1) insufficient for distinguishing similar labels at the same level; (2) fail to consider the error propagation problem caused by the incorrect parent-level predictions. For this reason, we first propose a hierarchical capsule network for the HTC task, due to the ability of capsules to distinguish similar categories. To ease the error propagation problem, we further devise two novel mechanisms in the proposed hierarchical capsule framework, i.e.,Label InjectionandLabel Re-Routing, to enhance the tolerance of the model to the incorrect parent-level predictions. Experiments on two widely used datasets prove that our model achieves competitive performance. The ablation study further demonstrates the scalability ofLabel InjectionandLabel Re-Routing.
Fei Zhao 0012, Zhen Wu 0002, Liang He 0009, Xinyu Dai
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 latent-GLAT: Glancing at Latent Variables for Parallel Text Generation
abstract
Yu Bao, Hao Zhou, Shujian Huang, Dongqi Wang, Lihua Qian, Xinyu Dai, Jiajun Chen, Lei Li. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Hao Zhou 0012, Shujian Huang, Dongqi Wang 0005, Lihua Qian, Xinyu Dai, Jiajun Chen 0001, Lei Li 0005
ACL (1)6
2022 Towards Multi-label Unknown Intent Detection
abstract
Multi-class unknown intent detection has made remarkable progress recently. However, it has a strong assumption that each utterance has only one intent, which does not conform to reality because utterances often have multiple intents. In this paper, we propose a more desirable task, multi-label unknown intent detection, to detect whether the utterance contains the unknown intent, in which each utterance may contain multiple intents. In this task, the unique utterances simultaneously containing known and unknown intents make existing multi-class methods easy to fail. To address this issue, we propose an intuitive and effective method to recognize whether All Intents contained in the utterance are Known (AIK). Our high-level idea is to predict the utterance’s intent number, then check whether the utterance contains the same number of known intents. If the number of known intents is less than the number of intents, it implies that the utterance also contains unknown intents. We benchmark AIK over existing methods, and empirical results suggest that our method obtains state-of-the-art performances. For example, on the MultiWOZ 2.3 dataset, AIK significantly reduces the FPR95 by 12.25% compared to the best baseline.
Yawen Ouyang, Zhen Wu 0002, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
COLING3
2022 Alleviating the Inequality of Attention Heads for Neural Machine Translation
abstract
Recent studies show that the attention heads in Transformer are not equal. We relate this phenomenon to the imbalance training of multi-head attention and the model dependence on specific heads. To tackle this problem, we propose a simple masking method: HeadMask, in two specific ways. Experiments show that translation improvements are achieved on multiple language pairs. Subsequent empirical analyses also support our assumption and confirm the effectiveness of the method.
Zewei Sun, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
COLING3
2022 Learning from Adjective-Noun Pairs: A Knowledge-enhanced Framework for Target-Oriented Multimodal Sentiment Classification
abstract
Target-oriented multimodal sentiment classification (TMSC) is a new subtask of aspect-based sentiment analysis, which aims to determine the sentiment polarity of the opinion target mentioned in a (sentence, image) pair. Recently, dominant works employ the attention mechanism to capture the corresponding visual representations of the opinion target, and then aggregate them as evidence to make sentiment predictions. However, they still suffer from two problems: (1) The granularity of the opinion target in two modalities is inconsistent, which causes visual attention sometimes fail to capture the corresponding visual representations of the target; (2) Even though it is captured, there are still significant differences between the visual representations expressing the same mood, which brings great difficulty to sentiment prediction. To this end, we propose a novel Knowledge-enhanced Framework (KEF) in this paper, which can successfully exploit adjective-noun pairs extracted from the image to improve the visual attention capability and sentiment prediction capability of the TMSC task. Extensive experimental results show that our framework consistently outperforms state-of-the-art works on two public datasets.
Fei Zhao 0012, Zhen Wu 0002, Siyu Long, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
COLING4
2022 Structure-Unified M-Tree Coding Solver for Math Word Problem
abstract
As one of the challenging NLP tasks, designing math word problem (MWP) solvers has attracted increasing research attention for the past few years.In previous work, models designed by taking into account the properties of the binary tree structure of mathematical expressions at the output side have achieved better performance.However, the expressions corresponding to a MWP are often diverse (e.g., n 1 +n 2 ×n 3 -n 4 , n 3 ×n 2 -n 4 +n 1 , etc.), and so are the corresponding binary trees, which creates difficulties in model learning due to the non-deterministic output space.In this paper, we propose the Structure-Unified M-Tree Coding Solver (SUMC-Solver), which applies a tree with any M branches (M-tree) to unify the output structures.To learn the M-tree, we use a mapping to convert the M-tree into the M-tree codes, where codes store the information of the paths from tree root to leaf nodes and the information of leaf nodes themselves, and then devise a Sequence-to-Code (seq2code) model to generate the codes.Experimental results on the widely used MAWPS and Math23K datasets have demonstrated that SUMC-Solver not only outperforms several state-of-the-art models under similar experimental settings but also performs much better under low-resource conditions 1 .
Bin Wang 0016, Jiangzhou Ju, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
EMNLP4
2022 Learning from Different text-image Pairs: A Relation-enhanced Graph Convolutional Network for Multimodal NER
abstract
Multimodal Named Entity Recognition (MNER) aims to locate and classify named entities mentioned in a (text, image) pair. However, dominant work independently models the internal matching relations in a pair of image and text, ignoring the external matching relations between different (text, image) pairs inside the dataset, though such relations are crucial for alleviating image noise in MNER task. In this paper, we primarily explore two kinds of external matching relations between different (text, image) pairs, i.e., inter-modal relations and intra-modal relations. On the basis, we propose a Relation-enhanced Graph Convolutional Network (R-GCN) for the MNER task. Specifically, we first construct an inter-modal relation graph and an intra-modal relation graph to gather the image information most relevant to the current text and image from the dataset, respectively. And then, multimodal interaction and fusion are leveraged to predict the NER label sequences. Extensive experimental results show that our model consistently outperforms state-of-the-art works on two public datasets. Our code and datasets are available at https://github.com/1429904852/R-GCN.
Fei Zhao 0012, Zhen Wu 0002, Shangyu Xing, Xinyu Dai
ACM Multimedia5
2022 Zero-Shot 3D Drug Design by Sketching and Generating
abstract
Drug design is a crucial step in the drug discovery cycle. Recently, various deep learning-based methods design drugs by generating novel molecules from scratch, avoiding traversing large-scale drug libraries. However, they depend on scarce experimental data or time-consuming docking simulation, leading to overfitting issues with limited training data and slow generation speed. In this study, we propose the zero-shot drug design method DESERT (Drug dEsign by SkEtching and geneRaTing). Specifically, DESERT splits the design process into two stages: sketching and generating, and bridges them with the molecular shape. The two-stage fashion enables our method to utilize the large-scale molecular database to reduce the need for experimental data and docking simulation. Experiments show that DESERT achieves a new state-of-the-art at a fast speed.
Siyu Long, Yi Zhou 0018, Xinyu Dai, Hao Zhou 0012
NeurIPS3
2022 ADS-Cap: A Framework for Accurate and Diverse Stylized Captioning with Unpaired Stylistic Corpora
Kanzhi Cheng, Zheng Ma 0012, Shi Zong, Xinyu Dai, Jiajun Chen 0001
NLPCC (1)5
2022 Modeling Users' Contextualized Page-wise Feedback for Click-Through Rate Prediction in E-commerce Search
abstract
Modeling user's historical feedback is essential for Click-Through Rate Prediction in personalized search and recommendation. Existing methods usually only model users' positive feedback information such as click sequences which neglects the context information of the feedback. In this paper, we propose a new perspective for context-aware users' behavior modeling by including the whole page-wisely exposed products and the corresponding feedback as contextualized page-wise feedback sequence. The intra-page context information and inter-page interest evolution can be captured to learn more specific user preference. We design a novel neural ranking model RACP(Recurrent Attention over Contextualized Page sequence), which utilizes page-context aware attention to model the intra-page context. A recurrent attention process is used to model the cross-page interest convergence evolution as denoising the interest in the previous pages. Experiments on public and real-world industrial datasets verify our model's effectiveness.
Zhifang Fan, Dan Ou, Yulong Gu, Bairan Fu, Xiang Li 0107, Wentian Bao, Xinyu Dai, Xiaoyi Zeng, Qingwen Liu 0002
WSDM7
2022 The detection of distributional discrepancy for language GANs
abstract
A pre-trained neural language model (LM) is usually used to generate texts. Due to exposure bias, the generated text is not as good as real text. Many researchers claimed they employed the Generative Adversarial Nets (GAN) to alleviate this issue by feeding reward signals from a discriminator to update the LM (generator). However, some researchers argued that GAN did not work by evaluating the generated texts with a quality-diversity metric such as Bleu versus self-Bleu, and language model score versus reverse language model score. Unfortunately, these two-dimension metrics are not reliable. Furthermore, the existing methods only assessed the final generated texts, thus neglecting the dynamic evaluating the adversarial learning process. Different from the above-mentioned methods, we adopted the most recent metric functions, which measure the distributional discrepancy between real and generated text. Besides that, we design a comprehensive experiment to investigate the performance during the learning process. First, we evaluate a language model with two functions and identify a large discrepancy. Then, several methods with the detected discrepancy signal to improve the generator were tried. Experimenting with two language GANs on two benchmark datasets, we found that the distributional discrepancy increases with more adversarial learning rounds. Our research provides convicted evidence that the language GANs fail.
Ping Cai, Hongjun Wang 0002, Xinyu Dai, Jiajun Chen 0001
Connect. Sci.5
2022 Pairwise tagging framework for end-to-end emotion-cause pair extraction
Zhen Wu 0002, Xinyu Dai
Frontiers Comput. Sci.2
2022 Self-Supervised Task Augmentation for Few-Shot Intent Detection
Yawen Ouyang, Dingjie Song, Xinyu Dai
J. Comput. Sci. Technol.4
2021 Automated Cross-prompt Scoring of Essay Traits
abstract
The majority of current research in Automated Essay Scoring (AES) focuses on prompt-specific scoring of either the overall quality of an essay or the quality with regards to certain traits. In real-world applications obtaining labelled data for a target essay prompt is often expensive or unfeasible, requiring the AES system to be able to perform well when predicting scores for essays from unseen prompts. As a result, some recent research has been dedicated to cross-prompt AES. However, this line of research has thus far only been concerned with holistic, overall scoring, with no exploration into the scoring of different traits. As users of AES systems often require feedback with regards to different aspects of their writing, trait scoring is a necessary component of an effective AES system. Therefore, to address this need, we introduce a new task named Automated Cross-prompt Scoring of Essay Traits, which requires the model to be trained solely on non-target-prompt essays and to predict the holistic, overall score as well as scores for a number of specific traits for target-prompt essays. This task challenges the model's ability to generalize in order to score essays from a novel domain as well as its ability to represent the quality of essays from multiple different aspects. In addition, we introduce a new, innovative approach which builds on top of a state-of-the-art method for cross-prompt AES. Our method utilizes a trait-attention mechanism and a multi-task architecture that leverages the relationships between each trait to simultaneously predict the overall score and the score of each individual trait. We conduct extensive experiments on the widely used ASAP and ASAP++ datasets and demonstrate that our approach is able to outperform leading prompt-specific trait scoring and cross-prompt AES methods.
Robert Ridley, Liang He 0009, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
AAAI3
2021 Meta-LMTC: Meta-Learning for Large-Scale Multi-Label Text Classification
abstract
Large-scale multi-label text classification (LMTC) tasks often face long-tailed label distributions, where many labels have few or even no training instances.Although current methods can exploit prior knowledge to handle these few/zero-shot labels, they neglect the metaknowledge contained in the dataset that can guide models to learn with few samples.In this paper, for the first time, this problem is addressed from a meta-learning perspective.However, the simple extension of meta-learning approaches to multi-label classification is suboptimal for LMTC tasks due to long-tailed label distribution and coexisting of few-and zeroshot scenarios.We propose a meta-learning approach named META-LMTC.Specifically, it constructs more faithful and more diverse tasks according to well-designed sampling strategies and directly incorporates the objective of adapting to new low-resource tasks into the metalearning phase.Extensive experiments show that META-LMTC achieves state-of-the-art performance against strong baselines and can still enhance powerful BERTlike models.
Ran Wang 0010, Xi'ao Su, Siyu Long, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
EMNLP (1)4
2021 MEDA: Meta-Learning with Data Augmentation for Few-Shot Text Classification
abstract
Meta-learning has recently emerged as a promising technique to address the challenge of few-shot learning. However, standard meta-learning methods mainly focus on visual tasks, which makes it hard for them to deal with diverse text data directly. In this paper, we introduce a novel framework for few-shot text classification, which is named as MEta-learning with Data Augmentation (MEDA). MEDA is composed of two modules, a ball generator and a meta-learner, which are learned jointly. The ball generator is to increase the number of shots per class by generating more samples, so that meta-learner can be trained with both original and augmented samples. It is worth noting that ball generator is agnostic to the choice of the meta-learning methods. Experiment results show that on both datasets, MEDA outperforms existing state-of-the-art methods and significantly improves the performance of meta-learning on few-shot text classification.
Yawen Ouyang, Wenming Zhang, Xinyu Dai
IJCAI4
2021 Non-Autoregressive Translation by Learning Target Categorical Codes
abstract
Yu Bao, Shujian Huang, Tong Xiao, Dongqi Wang, Xinyu Dai, Jiajun Chen. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Shujian Huang, Tong Xiao 0001, Dongqi Wang 0005, Xinyu Dai, Jiajun Chen 0001
NAACL-HLT5
2021 UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost
abstract
Zhen Wu, Lijun Wu, Qi Meng, Yingce Xia, Shufang Xie, Tao Qin, Xinyu Dai, Tie-Yan Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Zhen Wu 0002, Lijun Wu 0003, Yingce Xia, Shufang Xie 0003, Tao Qin 0001, Xinyu Dai, Tie-Yan Liu
NAACL-HLT7
2021 Dual Side Deep Context-aware Modulation for Social Recommendation
abstract
Social recommendation is effective in improving the recommendation performance by leveraging social relations from online social networking platforms. Social relations among users provide friends’ information for modeling users’ interest in candidate items and help items expose to potential consumers (i.e., item attraction). However, there are two issues haven’t been well-studied: Firstly, for the user interests, existing methods typically aggregate friends’ information contextualized on the candidate item only, and this shallow context-aware aggregation makes them suffer from the limited friends’ information. Secondly, for the item attraction, if the item’s past consumers are the friends of or have a similar consumption habit to the targeted user, the item may be more attractive to the targeted user, but most existing methods neglect the relation enhanced context-aware item attraction.
Bairan Fu, Wenming Zhang, Guang-Neng Hu, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
WWW4
2021 Improving Neural Question Generation using Deep Linguistic Representation
abstract
Question Generation (QG) is a challenging Natural Language Processing (NLP) task which aims at generating questions with given answers and context. There are many works incorporating linguistic features to improve the performance of QG. However, similar to traditional word embedding, these works normally embed such features with a set of trainable parameters, which results in the linguistic features not fully exploited. In this work, inspired by the recent achievements of text representation, we propose to utilize linguistic information via large pre-trained neural models. First, these models are trained in several specific NLP tasks in order to better represent linguistic features. Then, such feature representation is fused into a seq2seq based QG model to guide question generation. Extensive experiments were conducted on two benchmark Question Generation datasets to evaluate the effectiveness of our approach. The experimental results demonstrate that our approach outperforms the state-of-the-art QG systems, as a result, it significantly improves the baseline by 17.2% and 6.2% under the BLEU-4 metric on these two datasets, respectively.
Wei Yuan 0003, Tieke He, Xinyu Dai
WWW3
2021 Sampling informative context nodes for network embedding
Danhao Zhu, Xinyu Dai, Jiajun Chen 0001
Sci. China Inf. Sci.2
2021 Adversarial subsequences for unconditional text generation
Jiuhua Zhang, Xinyu Dai, Jiajun Chen 0001
Comput. Speech Lang.5
2021 Integrating heterogeneous thesauruses for Chinese synonyms
Peng Wu 0037, Yingjie Zhang 0002, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
Frontiers Comput. Sci.5
2021 A novel reasoning mechanism for multi-label text classification
Ran Wang 0010, Robert Ridley, Xi'ao Su, Weiguang Qu, Xinyu Dai
Inf. Process. Manag.5
2021 Pre-Train and Learn: Preserving Global Information for Graph Neural Networks
Danhao Zhu, Xinyu Dai, Jiajun Chen 0001
J. Comput. Sci. Technol.2
2020 Generating Diverse Translation by Manipulating Multi-Head Attention
abstract
Transformer model (Vaswani et al. 2017) has been widely used in machine translation tasks and obtained state-of-the-art results. In this paper, we report an interesting phenomenon in its encoder-decoder multi-head attention: different attention heads of the final decoder layer align to different word translation candidates. We empirically verify this discovery and propose a method to generate diverse translations by manipulating heads. Furthermore, we make use of these diverse translations with the back-translation technique for better data augmentation. Experiment results show that our method generates diverse translations without a severe drop in translation quality. Experiments also show that back-translation with these diverse translations could bring a significant improvement in performance on translation tasks. An auxiliary experiment of conversation response generation task proves the effect of diversity as well.
Zewei Sun, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
AAAI4
2020 Latent Opinions Transfer Network for Target-Oriented Opinion Words Extraction
abstract
Target-oriented opinion words extraction (TOWE) is a new subtask of ABSA, which aims to extract the corresponding opinion words for a given opinion target in a sentence. Recently, neural network methods have been applied to this task and achieve promising results. However, the difficulty of annotation causes the datasets of TOWE to be insufficient, which heavily limits the performance of neural models. By contrast, abundant review sentiment classification data are easily available at online review sites. These reviews contain substantial latent opinions information and semantic patterns. In this paper, we propose a novel model to transfer these opinions knowledge from resource-rich review sentiment classification datasets to low-resource task TOWE. To address the challenges in the transfer process, we design an effective transformation method to obtain latent opinions, then integrate them into TOWE. Extensive experimental results show that our model achieves better performance compared to other state-of-the-art methods and significantly outperforms the base model without transferring opinions knowledge. Further analysis validates the effectiveness of our model.
Zhen Wu 0002, Fei Zhao 0012, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
AAAI3
2020 Explicit Semantic Decomposition for Definition Generation
abstract
Definition generation, which aims to automatically generate dictionary definitions for words, has recently been proposed to assist the construction of dictionaries and help people understand unfamiliar texts.However, previous works hardly consider explicitly modeling the "components" of definitions, leading to under-specific generation results.In this paper, we propose ESD, namely Explicit Semantic Decomposition for definition generation, which explicitly decomposes meaning of words into semantic components, and models them with discrete latent variables for definition generation.Experimental results show that ESD achieves substantial improvements on WordNet and Oxford benchmarks over strong previous baselines.
Jiahuan Li, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
ACL4
2020 Dialogue State Tracking with Explicit Slot Connection Modeling
abstract
Recent proposed approaches have made promising progress in dialogue state tracking (DST). However, in multi-domain scenarios, ellipsis and reference are frequently adopted by users to express values that have been mentioned by slots from other domains. To handle these phenomena, we propose a Dialogue State Tracking with Slot Connections (DST-SC) model to explicitly consider slot correlations across different domains. Given a target slot, the slot connecting mechanism in DST-SC can infer its source slot and copy the source slot value directly, thus significantly reducing the difficulty of learning and reasoning. Experimental results verify the benefits of explicit slot connection modeling, and our model achieves state-of-the-art performance on MultiWOZ 2.0 and MultiWOZ 2.1 datasets.
Yawen Ouyang, Moxin Chen, Xinyu Dai, Yinggong Zhao, Shujian Huang, Jiajun Chen 0001
ACL3
2020 A Reinforced Generation of Adversarial Examples for Neural Machine Translation
abstract
Neural machine translation systems tend to fail on less decent inputs despite its significant efficacy, which may significantly harm the credibility of these systems-fathoming how and when neural-based systems fail in such cases is critical for industrial maintenance.Instead of collecting and analyzing bad cases using limited handcrafted error features, here we investigate this issue by generating adversarial examples via a new paradigm based on reinforcement learning.Our paradigm could expose pitfalls for a given performance metric, e.g., BLEU, and could target any given neural machine translation architecture.We conduct experiments of adversarial attacks on two mainstream neural machine translation architectures, RNN-search, and Transformer.The results show that our method efficiently produces stable attacks with meaning-preserving adversarial examples.We also present a qualitative and quantitative analysis for the preference pattern of the attack, demonstrating its capability of pitfall exposure.
Shujian Huang, Xinyu Dai, Jiajun Chen 0001
ACL4
2020 Enhance Prototypical Network with Text Descriptions for Few-shot Relation Classification
abstract
Recently few-shot relation classification has drawn much attention. It devotes to addressing the long-tail relation problem by recognizing the relations from few instances. The existing metric learning methods aim to learn the prototype of classes and make prediction according to distances between query and prototypes. However, it is likely to make unreliable predictions due to the text diversity. It is intuitive that the text descriptions of relation and entity can provide auxiliary support evidence for relation classification. In this paper, we propose TD-Proto, which enhances prototypical network with relation and entity descriptions. We design a collaborative attention module to extract beneficial and instructional information of sentence and entity respectively. A gate mechanism is proposed to fuse both information dynamically so as to obtain a knowledge-aware instance. Experimental results demonstrate that our method achieves excellent performance.
Kaijia Yang, Nantao Zheng, Xinyu Dai, Liang He 0009, Shujian Huang, Jiajun Chen 0001
CIKM3
2020 Synonym Knowledge Enhanced Reader for Chinese Idiom Reading Comprehension
abstract
Machine reading comprehension (MRC) is the task that asks a machine to answer questions based on a given context. For Chinese MRC, due to the non-literal and non-compositional semantic characteristics, Chinese idioms pose unique challenges for machines to understand. Previous studies tend to treat idioms separately without fully exploiting the relationship among them. In this paper, we first define the concept of literal meaning coverage to measure the consistency between semantics and literal meanings for Chinese idioms. With the definition, we prove that the literal meanings of many idioms are far from their semantics, and we also verify that the synonymic relationship can mitigate this inconsistency, which would be beneficial for idiom comprehension. Furthermore, to fully utilize the synonymic relationship, we propose the synonym knowledge enhanced reader. Specifically, for each idiom, we first construct a synonym graph according to the annotations from the high-quality synonym dictionary or the cosine similarity between the pre-trained idiom embeddings and then incorporate the graph attention network and gate mechanism to encode the graph. Experimental results on ChID, a large-scale Chinese idiom reading comprehension dataset, show that our model achieves state-of-the-art performance.
Siyu Long, Ran Wang 0010, Kun Tao, Jiali Zeng, Xinyu Dai
COLING5
2020 Attention Transfer Network for Aspect-level Sentiment Classification
abstract
Aspect-level sentiment classification (ASC) aims to detect the sentiment polarity of a given opinion target in a sentence.In neural network-based methods for ASC, most works employ the attention mechanism to capture the corresponding sentiment words of the opinion target, then aggregate them as evidence to infer the sentiment of the target.However, aspect-level datasets are all relatively small-scale due to the complexity of annotation.Data scarcity causes the attention mechanism sometimes to fail to focus on the corresponding sentiment words of the target, which finally weakens the performance of neural models.To address the issue, we propose a novel Attention Transfer Network (ATN) in this paper, which can successfully exploit attention knowledge from resource-rich document-level sentiment classification datasets to improve the attention capability of the aspect-level sentiment classification task.In the ATN model, we design two different methods to transfer attention knowledge and conduct experiments on two ASC benchmark datasets.Extensive experimental results show that our methods consistently outperform state-of-the-art works.Further analysis also validates the effectiveness of ATN.Our code and dataset are available at https
Fei Zhao 0012, Zhen Wu 0002, Xinyu Dai
COLING3
2020 Mirror-Generative Neural Machine Translation
Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Lei Li 0005, Xinyu Dai, Jiajun Chen 0001
ICLR5
2020 Transformer-Based Multi-aspect Modeling for Multi-aspect Multi-sentiment Analysis
Zhen Wu 0002, Chengcan Ying, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
NLPCC (2)3
2020 Opinion Transmission Network for Jointly Improving Aspect-Oriented Opinion Words Extraction and Sentiment Classification
Chengcan Ying, Zhen Wu 0002, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
NLPCC (1)3
2020 MSGE: A Multi-step Gated Model for Knowledge Graph Completion
Chunyang Tan, Kaijia Yang, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
PAKDD (1)3
2020 Improving Self-Attention Networks With Sequential Relations
abstract
Recently, self-attention networks show strong advantages of sentence modeling in many NLP tasks. However, self-attention mechanism computes the interactions of every pair of words independently regardless of their positions, which makes it not able to capture the sequential relations between words in different positions in a sentence. In this paper, we improve the self-attention networks by better integrating sequential relations, which is essential for modeling natural languages. Specifically, we 1) propose a position-based attention to model the interaction between two words regarding positions; 2) perform separated attention for the context before and after the current position, respectively; and 3) merge the above two parts with a position-aware gated fusion mechanism. Experiments in natural language inference, machine translation and sentiment analysis tasks show that our sequential relation modeling helps self-attention networks outperform existing approaches. We also provide extensive analyses to shed light on what the models have learned about the sequential relations.
Zaixiang Zheng, Shujian Huang, Rongxiang Weng, Xinyu Dai, Jiajun Chen 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Generating Sentences from Disentangled Syntactic and Semantic Spaces
abstract
Variational auto-encoders (VAEs) are widely used in natural language generation due to the regularization of the latent space.However, generating sentences from the continuous latent space does not explicitly model the syntactic information.In this paper, we propose to generate sentences from disentangled syntactic and semantic spaces.Our proposed method explicitly models syntactic information in the VAE's latent space by using the linearized tree sequence, leading to better performance of language generation.Additionally, the advantage of sampling in the disentangled syntactic and semantic latent spaces enables us to perform novel applications, such as the unsupervised paraphrase generation and syntaxtransfer generation.Experimental results show that our proposed model achieves similar or better performance in various tasks, compared with state-of-the-art related work.
Hao Zhou 0012, Shujian Huang, Lei Li 0005, Lili Mou, Olga Vechtomova, Xinyu Dai, Jiajun Chen 0001
ACL (1)7
2019 GeoSQA: A Benchmark for Scenario-based Question Answering in the Geography Domain at High School Level
abstract
Zixian Huang, Yulin Shen, Xiao Li, Yu’ang Wei, Gong Cheng, Lin Zhou, Xinyu Dai, Yuzhong Qu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zixian Huang, Xiao Li 0043, Yuang Wei, Gong Cheng 0001, Xinyu Dai, Yuzhong Qu
EMNLP/IJCNLP (1)7
2019 Fine-grained Knowledge Fusion for Sequence Labeling Domain Adaptation
abstract
Huiyun Yang, Shujian Huang, Xin-Yu Dai, Jiajun Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Huiyun Yang, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
EMNLP/IJCNLP (1)3
2019 Dynamic Past and Future for Neural Machine Translation
abstract
Zaixiang Zheng, Shujian Huang, Zhaopeng Tu, Xin-Yu Dai, Jiajun Chen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zaixiang Zheng, Shujian Huang, Zhaopeng Tu, Xinyu Dai, Jiajun Chen 0001
EMNLP/IJCNLP (1)4
2019 Integrating Computational Thinking with Digital Storytelling to Enhancing Expression Ability
abstract
Cultivating computational thinking has become an important key to the curriculum of science and technology learning area. When students use interactive media, they can no longer be just receiver, but use interactive media to self-expression. We will let students in grades 3-4 learn ScratchJr, because most students in grades 5-6 learn Scratch at Tainan city. The student use ScratchJr to create the animations, stories and games which interest them. We use creation as a motivation, and children can learn how to solve problems by creating their projects over and over. Children can cultivate their own computational thinking by this way. Digital storytelling can enhance cognitive ability, a nd make things specific and clear. This is a good way to make everyone to understand the abstract knowledge. Let students use ScratchJr to make digital storytelling on the iPad mini in order to cultivate students' computing thinking. The main goal of this thesis is to improve students' expressing ability to perspectives in the three-dimensional framework of computational thinking. Twenty-seven students in third grade participated in this study. They used ScratchJr to make digital storytelling on the iPad mini. The researcher designed a series of training curriculum. All training curriculum were taught by researcher. Students completed three projects in the basic, intermediate and advanced three-phase training curriculum. Students present and save their projects by the way of YouTube in this thesis. Teaching students to use the iPad mini for video recording, and providing students the interview questions. Arrange students to work in groups of two, use iPad mini to interviews with each other, and then record the interview process. In the end, the qualitative analysis method was used to analyze the students' projects after the three-phase training curriculum and the video of the interview. It will explore the objective of this study: Enhance students’ expression ability by the integration of computational thinking into digital storytelling teaching.
Chih-Kai Chang, Xinyu Dai
ICCE2
2019 Utilizing Non-Parallel Text for Style Transfer by Making Partial Comparisons
abstract
Text style transfer aims to rephrase a given sentence into a different style without changing its original content. Since parallel corpora (i.e. sentence pairs with the same content but different styles) are usually unavailable, most previous works solely guide the transfer process with distributional information, i.e. using style-related classifiers or language models, which neglect the correspondence of instances, leading to poor transfer performance, especially for the content preservation. In this paper, we propose making partial comparisons to explicitly model the content and style correspondence of instances, respectively. To train the partial comparators, we propose methods to extract partial-parallel training instances automatically from the non-parallel data, and to further enhance the training process by using data augmentation. We perform experiments that compare our method to other existing approaches on two review datasets. Both automatic and manual evaluations show that our approach can significantly improve the performance of existing adversarial methods, and outperforms most state-of-the-art models. Our code and data will be available on Github.
Shujian Huang, Xinyu Dai, Jiajun Chen 0001
IJCAI3
2019 MSNE: A Novel Markov Chain Sampling Strategy for Network Embedding
Ran Wang 0010, Xinyu Dai
PAKDD (3)3
2019 PCANE: Preserving Context Attributes for Network Embedding
Danhao Zhu, Xinyu Dai, Kaijia Yang, Jiajun Chen 0001
PAKDD (3)2
2019 Representing anything from scholar papers
Danhao Zhu, Xinyu Dai, Jiajun Chen 0001
J. Web Semant.2
2018 Improving Review Representations With User Attention and Product Attention for Sentiment Classification
abstract
Neural network methods have achieved great success in reviews sentiment classification. Recently, some works achieved improvement by incorporating user and product information to generate a review representation. However, in reviews, we observe that some words or sentences show strong user's preference, and some others tend to indicate product's characteristic. The two kinds of information play different roles in determining the sentiment label of a review. Therefore, it is not reasonable to encode user and product information together into one representation. In this paper, we propose a novel framework to encode user and product information. Firstly, we apply two individual hierarchical neural networks to generate two representations, with user attention or with product attention. Then, we design a combined strategy to make full use of the two representations for training and final prediction. The experimental results show that our model obviously outperforms other state-of-the-art methods on IMDB and Yelp datasets. Through the visualization of attention over words related to user or product, we validate our observation mentioned above.
Zhen Wu 0002, Xinyu Dai, Cunyan Yin, Shujian Huang, Jiajun Chen 0001
AAAI2
2018 Dynamic Oracle for Neural Machine Translation in Decoding Phase
Zi-Yi Dou, Hao Zhou 0012, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
LREC4
2018 Combining Character and Word Information in Neural Machine Translation Using a Multi-Level Attention
abstract
Huadong Chen, Shujian Huang, David Chiang, Xinyu Dai, Jiajun Chen. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Huadong Chen, Shujian Huang, David Chiang 0001, Xinyu Dai, Jiajun Chen 0001
NAACL-HLT4
2018 Improving Aspect Identification with Reviews Segmentation
Tianhao Ning, Zhen Wu 0002, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
NLPCC (1)3
2018 Modeling Past and Future for Neural Machine Translation
abstract
Existing neural machine translation systems do not explicitly model what has been translated and what has not during the decoding phase. To address this problem, we propose a novel mechanism that separates the source information into two parts: translated Past contents and untranslated Future contents, which are modeled by two additional recurrent layers. The Past and Future contents are fed to both the attention model and the decoder states, which provides Neural Machine Translation (NMT) systems with the knowledge of translated and untranslated contents. Experimental results show that the proposed approach significantly improves the performance in Chinese-English, German-English, and English-German translation tasks. Specifically, the proposed model outperforms the conventional coverage model in terms of both the translation quality and the alignment error rate.
Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Lili Mou, Xinyu Dai, Jiajun Chen 0001, Zhaopeng Tu
Trans. Assoc. Comput. Linguistics5
2018 Collaborative Filtering with Topic and Social Latent Factors Incorporating Implicit Feedback
abstract
Recommender systems (RSs) provide an effective way of alleviating the information overload problem by selecting personalized items for different users. Latent factors-based collaborative filtering (CF) has become the popular approaches for RSs due to its accuracy and scalability. Recently, online social networks and user-generated content provide diverse sources for recommendation beyond ratings. Although social matrix factorization (Social MF) and topic matrix factorization (Topic MF) successfully exploit social relations and item reviews, respectively; both of them ignore some useful information. In this article, we investigate the effective data fusion by combining the aforementioned approaches. First, we propose a novel model MR3 to jointly model three sources of information (i.e., ratings, item reviews, and social relations) effectively for rating prediction by aligning the latent factors and hidden topics. Second, we incorporate the implicit feedback from ratings into the proposed model to enhance its capability and to demonstrate its flexibility. We achieve more accurate rating prediction on real-life datasets over various state-of-the-art methods. Furthermore, we measure the contribution from each of the three data sources and the impact of implicit feedback from ratings, followed by the sensitivity analysis of hyperparameters. Empirical studies demonstrate the effectiveness and efficacy of our proposed model and its extension.
Guang-Neng Hu, Xinyu Dai, Feng-Yu Qiu, Tao Li 0001, Shujian Huang, Jiajun Chen 0001
ACM Trans. Knowl. Discov. Data2
2017 A Multi-view Clustering Model for Event Detection in Twitter
Richard D. Shang, Xinyu Dai, Weiyi Ge, Shujian Huang, Jiajun Chen 0001
CICLing (2)2
2017 Top-Rank Enhanced Listwise Optimization for Statistical Machine Translation
abstract
Pairwise ranking methods are the basis of many widely used discriminative training approaches for structure prediction problems in natural language processing (NLP).Decomposing the problem of ranking hypotheses into pairwise comparisons enables simple and efficient solutions.However, neglecting the global ordering of the hypothesis list may hinder learning.We propose a listwise learning framework for structure prediction problems such as machine translation.Our framework directly models the entire translation list's ordering to learn parameters which may better fit the given listwise samples.Furthermore, we propose top-rank enhanced loss functions, which are more sensitive to ranking errors at higher positions.Experiments on a large-scale Chinese-English translation task show that both our listwise learning framework and top-rank enhanced listwise losses lead to significant improvements in translation quality.
Huadong Chen, Shujian Huang, David Chiang 0001, Xinyu Dai, Jiajun Chen 0001
CoNLL4
2017 Neural Machine Translation with Word Predictions
abstract
In the encoder-decoder architecture for neural machine translation (NMT), the hidden states of the recurrent structures in the encoder and decoder carry the crucial information about the sentence.These vectors are generated by parameters which are updated by back-propagation of translation errors through time.We argue that propagating errors through the end-to-end recurrent structures are not a direct way of control the hidden vectors.In this paper, we propose to use word predictions as a mechanism for direct supervision.More specifically, we require these vectors to be able to predict the vocabulary in target sentence.Our simple mechanism ensures better representations in the encoder and decoder without using any extra data or annotation.It is also helpful in reducing the target side vocabulary and improving the decoding efficiency.Experiments on Chinese-English and German-English machine translation tasks show BLEU improvements by 4.53 and 1.3, respectively.
Rongxiang Weng, Shujian Huang, Zaixiang Zheng, Xinyu Dai, Jiajun Chen 0001
EMNLP4
2017 Word-Context Character Embeddings for Chinese Word Segmentation
abstract
Neural parsers have benefited from automatically labeled data via dependencycontext word embeddings.We investigate training character embeddings on a word-based context in a similar way, showing that the simple method significantly improves state-of-the-art neural word segmentation models, beating tritraining baselines for leveraging autosegmented data.
Hao Zhou 0012, Zhenting Yu, Yue Zhang 0004, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
EMNLP5
2017 Deep Matrix Factorization Models for Recommender Systems
abstract
Recommender systems usually make personalized recommendation with user-item interaction ratings, implicit feedback and auxiliary information. Matrix factorization is the basic idea to predict a personalized ranking over a set of items for an individual user with the similarities among users and items. In this paper, we propose a novel matrix factorization model with neural network architecture. Firstly, we construct a user-item matrix with explicit ratings and non-preference implicit feedback. With this matrix as the input, we present a deep structure learning architecture to learn a common low dimensional space for the representations of users and items. Secondly, we design a new loss function based on binary cross entropy, in which we consider both explicit ratings and implicit feedback for a better optimization. The experimental results show the effectiveness of both our proposed model and the loss function. On several benchmark datasets, our model outperformed other state-of-the-art methods. We also conduct extensive experiments to evaluate the performance within different experimental settings.
Hong-Jian Xue, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
IJCAI2
2017 AGRA: An Analysis-Generation-Ranking Framework for Automatic Abbreviation from Paper Titles
abstract
People sometimes choose word-like abbreviations to refer to items with a long description. These abbreviations usually come from the descriptive text of the item and are easy to remember and pronounce, while preserving the key idea of the item. Coming up with a nice abbreviation is not an easy job, even for human. Previous assistant naming systems compose names by applying hand-written rules, which may not perform well. In this paper, we propose to view the naming task as an artificial intelligence problem and create a data set in the domain of academic naming. To generate more delicate names, we propose a three-step framework, including description analysis, candidate generation and abbreviation ranking, each of which is parameterized and optimizable. We conduct experiments to compare different settings of our framework with several analysis approaches from different perspectives. Compared to online or baseline systems, our framework could achieve the best results.
Shujian Huang, Cam-Tu Nguyen, Xiaoliang Wang 0001, Xinyu Dai, Jiajun Chen 0001, Yang Yu 0001
IJCAI6
2017 Integrating Reviews into Personalized Ranking for Cold Start Recommendation
Guang-Neng Hu, Xinyu Dai
PAKDD (2)2
2017 A Neural Probabilistic Structured-Prediction Method for Transition-Based Natural Language Processing
abstract
We propose a neural probabilistic structured-prediction method for transition-based natural language processing, which integrates beam search and contrastive learning. The method uses a global optimization model, which can leverage arbitrary features over non-local context. Beam search is used for efficient heuristic decoding, and contrastive learning is performed for adjusting the model according to search errors. When evaluated on both chunking and dependency parsing tasks, the proposed method achieves significant accuracy improvements over the locally normalized greedy baseline on the two tasks, respectively.
Hao Zhou 0012, Yue Zhang 0004, Chuan Cheng, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
J. Artif. Intell. Res.5
2016 A Search-Based Dynamic Reranking Model for Dependency Parsing
abstract
We propose a novel reranking method to extend a deterministic neural dependency parser.Different to conventional k-best reranking, the proposed model integrates search and learning by utilizing a dynamic action revising process, using the reranking model to guide modification for the base outputs and to rerank the candidates.The dynamic reranking model achieves an absolute 1.78% accuracy improvement over the deterministic baseline parser on PTB, which is the highest improvement by neural rerankers in the literature.
Hao Zhou 0012, Yue Zhang 0004, Shujian Huang, Junsheng Zhou, Xinyu Dai, Jiajun Chen 0001
ACL (1)5
2016 Tree-State Based Rule Selection Models for Hierarchical Phrase-Based Machine Translation
Shujian Huang, Huifeng Sun, Chengqi Zhao, Jinsong Su, Xinyu Dai, Jiajun Chen 0001
IJCAI5
2016 Unsupervised Storyline Extraction from News Articles
Haiyang Xu 0001, Xinyu Dai, Yulan He 0001
IJCAI3
2016 Tagging Chinese microblogger via sparse feature selection
abstract
In new media era, users post messages to record their daily lives and express their opinions via social media platforms, such as microblog. Recently, it is an attractive topic to tag users from the users generation contents. Tags for a microblog user, as the description for his/her interests, concerns or occupational characteristics, are playing an important role in user indexing, personalized recommendation, and so on. Previous works apply keyword extraction methods to present the interests of users. However, it is hard for keyword extraction to give accurate results when the data is deficient and noisy. In this paper, we propose a novel method to tag the users. Firstly, we apply feature selection via sparse classifier to generate preliminary tags for users. Then we also apply feature selection method to extend the tags. Finally, we refine the tags with a reranking strategy. We conduct our experiments on the data of the most popular Chinese microblog (Sina Weibo). The experimental results show that our method improves the performance significantly over other methods.
Richard D. Shang, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
IJCNN2
2016 Evaluating a Deterministic Shift-Reduce Neural Parser for Constituent Parsing
Hao Zhou 0012, Yue Zhang 0004, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
LREC4
2016 PRIMT: A Pick-Revise Framework for Interactive Machine Translation
abstract
Shanbo Cheng, Shujian Huang, Huadong Chen, Xin-Yu Dai, Jiajun Chen. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Shanbo Cheng, Shujian Huang, Huadong Chen, Xinyu Dai, Jiajun Chen 0001
HLT-NAACL4
2016 Adaptation of Language Models for SMT Using Neural Networks with Topic Information
abstract
Neural network language models (LMs) are shown to be effective in improving the performance of statistical machine translation (SMT) systems. However, state-of-the-art neural network LMs usually use words before the current position as context and neglect global topic information, which can help machine translation (MT) systems to select better translation candidates from a higher perspective. In this work, we propose improvement of the state-of-the-art feedforward neural language model with topic information. Two main issues need to be tackled when adding topics into neural network LMs for SMT: one is how to incorporate topics to the neural network; the other is how to get target-side topic distribution before translation. We incorporate topics by appending topic distribution to the input layer of a feedforward LM. We adopt a multinomial logistic-regression (MLR) model to predict the target-side topic distribution based on source side information. Moreover, we propose a feedforward neural network model to learn joint representations on the source side for topic prediction. LM experiments demonstrate that the perplexity on validation set can be greatly reduced by the topic-enhanced feedforward LM, and the prediction of target-side topics can be improved dramatically with the MLR model equipped with the joint source representations. A final MT experiment, conducted on a large-scale Chinese--English dataset, shows that our feedforward LM with predicted topics improves the translation performance against a strong baseline.
Yinggong Zhao, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2016 Enhancing Shift-Reduce Constituent Parsing with Action N-Gram Model
abstract
Current shift-reduce parsers “understand” the context by embodying a large number of binary indicator features with a discriminative model. In this article, we propose the action n-gram model, which utilizes the action sequence to help parsing disambiguation. The action n-gram model is trained on action sequences produced by parsers with the n-gram estimation method, which gives a smoothed maximum likelihood estimation of the action probability given a specific action history. We show that incorporating action n-gram models into a state-of-the-art parsing framework could achieve parsing accuracy improvements on three datasets across two languages.
Hao Zhou 0012, Shujian Huang, Junsheng Zhou, Yue Zhang 0004, Huadong Chen, Xinyu Dai, Chuan Cheng, Jiajun Chen 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.6
2015 Structured Sparsity with Group-Graph Regularization
abstract
In many learning tasks with structural properties, structural sparsity methods help induce sparse models, usually leading to better interpretability and higher generalization performance. One popular approach is to use group sparsity regularization that enforces sparsity on the clustered groups of features, while another popular approach is to adopt graph sparsity regularization that considers sparsity on the link structure of graph embedded features. Both the group and graph structural properties co-exist in many applications. However, group sparsity and graph sparsity have not been considered simultaneously yet. In this paper, we propose a g2-regularization that takes group and graph sparsity into joint consideration, and present an effective approach for its optimization. Experiments on both synthetic and real data show that, enforcing group-graph sparsity lead to better performance than using group sparsity or graph sparsity only.
Xinyu Dai, Shujian Huang, Jiajun Chen 0001, Zhi-Hua Zhou
AAAI1
2015 Non-linear Learning for Statistical Machine Translation
abstract
Shujian Huang, Huadong Chen, Xin-Yu Dai, Jiajun Chen. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Shujian Huang, Huadong Chen, Xinyu Dai, Jiajun Chen 0001
ACL (1)3
2015 Co-training for Semi-supervised Sentiment Classification Based on Dual-view Bags-of-words Representation
abstract
Rui Xia, Cheng Wang, Xin-Yu Dai, Tao Li. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Xinyu Dai, Tao Li 0001
ACL (1)3
2015 A Unified Framework for Jointly Learning Distributed Representations of Word and Attributes
Liqiang Niu, Xinyu Dai, Shujian Huang, Jiajun Chen 0001
ACML2
2015 Sentiment Classification with Graph Sparsity Regularization
Xinyu Dai, Chuan Cheng, Shujian Huang, Jiajun Chen 0001
CICLing (2)1
2015 A Synthetic Approach for Recommendation: Combining Ratings, Social Relations, and Reviews
Guang-Neng Hu, Xinyu Dai, Yunya Song, Shujian Huang, Jiajun Chen 0001
IJCAI2
2015 Word Segmentation of Micro Blogs with Bagging
abstract
This paper describes the model we designed for the Chinese word segmentation Task of NLPCC 2015. We firstly apply a word-based perceptron algorithm to build the base segmenter. Then, we use a Bootstrap Aggregating model of bagging which improves the segmentation results consistently on the three tracks of closed, semi-open and open test. Considering the characteristics of Weibo text, we also perform rule-based adaptation before decoding. Finally, our model achieves F-score 95.12% on closed track, 95.3% on semi-open track and 96.09% on open track.
Zhenting Yu, Xinyu Dai, Si Shen, Shujian Huang, Jiajun Chen 0001
NLPCC2
2015 Resolving Coordinate Structures for Chinese Constituent Parsing
abstract
Coordinate structures are linguistic structures consisting of two or more conjuncts, which usually compose into larger constituent as a whole unit. However, the boundary of each conjunct is difficult to identify, which makes it difficult to parse the whole coordinate and larger structures. In labeled data, such as the Penn Chinese Tree Bank (CTB), coordinate structures are not labeled explicitly, which makes solving the problem more complicated. In this paper, we treat resolving coordinate structures as an independent sub-problem of parsing. We first define coordinate structures explicitly and design rules to extract the coordinate structures from labeled CTB data. Then a specifically designed grammar is proposed for automatic parsing of coordinate structures. We propose two groups of new features to better model coordinate structures in a shift-reduce parsing framework. Our approach can achieve a $$15\%$$ improvement in F-1 score on resolving coordinate structures.
Yichu Zhou, Shujian Huang, Xinyu Dai, Jiajun Chen 0001
NLPCC3
2013 Forgetting Word Segmentation in Chinese Text Classification with L1-Regularized Logistic Regression
Xinyu Dai, Shujian Huang, Jiajun Chen 0001
PAKDD (2)2
2012 Opinion Target Extraction for Short Comments
Lin Shang 0001, Xinyu Dai, Mengjie Zhang 0001
PRICAI3
2011 Transfer Learning via Multi-View Principal Component Analysis
Yangsheng Ji, Jiajun Chen 0001, Gang Niu 0001, Lin Shang 0001, Xinyu Dai
J. Comput. Sci. Technol.5
2010 Chinese Event Descriptive Clause Splitting with Structured SVMs
Junsheng Zhou, Yabing Zhang, Xinyu Dai, Jiajun Chen 0001
CICLing3
2010 Improving Word Alignment by Semi-Supervised Ensemble
Shujian Huang, Kangxi Li, Xinyu Dai, Jiajun Chen 0001
CoNLL3
2009 Finding Appropriate Turning Point for Text Sentiment Polarity
Lin Shang 0001, Xinyu Dai, Cunyan Yin
ICONIP (2)3