Xu Sun 0001

dblp:37/1971-1 · DBLP profile ↗
← Back
147ranked-venue papers
25as first author
55since 2021 · last 2026
0000-0001-8241-9320ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 135 · 18 first-author · 51 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 2 first-author · 15 since 2021Databases, data management, data science and information retrieval · 14 · 9 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
abstract
Video Large Language Models (Video LLMs) have achieved significant success by adopting the paradigm of large-scale pre-training followed by supervised fine-tuning (SFT). However, existing approaches struggle with temporal reasoning due to weak temporal correspondence in the data and over-reliance on the next-token prediction paradigm, which collectively result in the absence temporal supervision. To address these limitations, we propose TEMPLE (TEMporal Preference Learning), a systematic framework that enhances temporal reasoning capabilities through Direct Preference Optimization (DPO). To address temporal information scarcity in data, we introduce an automated pipeline for systematically constructing temporality-intensive preference pairs comprising three steps: selecting temporally rich videos, designing video-specific perturbation strategies, and evaluating model responses on clean and perturbed inputs. Complementing this data pipeline, we provide additional supervision signals via preference learning and propose a novel Progressive Pre-SFT Alignment strategy featuring two key innovations: a curriculum learning strategy which progressively increases perturbation difficulty to maximize data efficiency; and applying preference optimization before instruction tuning to incentivize fundamental temporal alignment. Extensive experiments demonstrate that our approach consistently improves Video LLM performance across multiple benchmarks with a relatively small set of self-generated DPO data. Our findings highlight TEMPLE as a scalable and efficient complement to SFT-based methods, paving the way for developing reliable Video LLMs.
Lei Li 0039, Kun Ouyang, Shuhuai Ren, Yuanxin Liu, Yuanxing Zhang, Lingpeng Kong, Qi Liu 0049, Xu Sun 0001
AAAI10
2026 Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
abstract
Zhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lean Wang, Yuanxin Liu, Lei Li 0009, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
ACL (1)8
2025 InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation
abstract
Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the avatar, making the generated video less vivid and controllable. In this paper, we propose a text-guided approach for generating emotionally expressive 2D avatars, offering fine-grained control, improved interactivity, and generalizability to the resulting video. Our framework, named InstructAvatar, leverages a natural language interface to control the emotion as well as the facial motion of avatars. Technically, we utilize GPT-4V to design an automatic annotation pipeline, constructing an instruction-video paired training dataset. This is combined with a novel two-branch diffusion-based generator to predict avatars using both audio and text instructions simultaneously. Experimental results demonstrate that InstructAvatar produces results that align well with both conditions, and outperforms existing methods in fine-grained emotion control, lip-sync quality, and naturalness.
Yuchi Wang, Junliang Guo, Jianhong Bai, Runyi Yu 0002, Tianyu He, Xu Tan 0003, Xu Sun 0001, Jiang Bian 0002
AAAI7
2025 ATLANTIS: Weak-to-Strong Learning via Importance Sampling
abstract
Supervised fine-tuning (SFT) enables large language models to align with training data for better performance in many aspects.Nevertheless, the gap between the distribution of current datasets from human annotations or model generations and the real-world data distribution heavily limits the capacities and potentials of models.As a result, we propose a new SFT technique, ATLANTIS, to bridge the gap.We adopt importance sampling to estimate the optimal data distribution in the real world from existing training datasets because the former is hard to sample from.Furthermore, we introduce an extra small model and reference model to estimate the sampling ratio through the probability gap between them.We evaluate our method with benchmarks in knowledge & understanding and preference aspects.The experiment results prove that ATLANTIS can bring consistent and significant improvements to models' performance.What's more, our method can be flexibly transferred among models with different structures.Our analyses demonstrate that our method is well-compatible with other SFT techniques to further enhance models' capacities and has great potential to be combined with existing training frameworks.
Feifan Song 0001, Xu Sun 0001
ACL (1)5
2025 PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
abstract
Kun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kun Ouyang, Yuanxin Liu, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
ACL (1)8
2025 VidTwin: Video VAE with Decoupled Structure and Dynamics
abstract
Recent advancements in video autoencoders (Video AEs) have significantly improved the quality and efficiency of video generation. In this paper, we propose a novel and compact video autoencoder, VidTwin, that decouples video into two distinct latent spaces: Structure latent vectors, which capture overall content and global movement, and Dynamics latent vectors, which represent fine-grained details and rapid movements. Specifically, our approach leverages an Encoder-Decoder backbone, augmented with two submodules for extracting these latent spaces, respectively. The first submodule employs a Q-Former to extract low-frequency motion trends, followed by downsampling blocks to remove redundant content details. The second averages the latent vectors along the spatial dimension to capture rapid motion. Extensive experiments show that VidTwin achieves a high compression rate of 0.20% with high reconstruction quality (PSNR of 28.14 on the MCL-JCV dataset), and performs efficiently and effectively in downstream generative tasks. Moreover, our model demonstrates explainability and scalability, paving the way for future research in video latent representation and generation.
Yuchi Wang, Junliang Guo, Tianyu He, Xu Sun 0001, Jiang Bian 0002
CVPR5
2025 RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
abstract
Yuchi Wang, Yishuo Cai, Shuhuai Ren, Sihan Yang, Linli Yao, Yuanxin Liu, Yuanxing Zhang, Pengfei Wan, Xu Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yuchi Wang, Yishuo Cai, Shuhuai Ren, Linli Yao, Yuanxin Liu, Yuanxing Zhang, Pengfei Wan 0001, Xu Sun 0001
EMNLP9
2025 Temporal Reasoning Transfer from Text to Video
abstract
Video Large Language Models (Video LLMs) have shown promising capabilities in video comprehension, yet they struggle with tracking temporal changes and reasoning about temporal relationships. While previous research attributed this limitation to the ineffective temporal encoding of visual inputs, our diagnostic study reveals that video representations contain sufficient information for even small probing classifiers to achieve perfect accuracy. Surprisingly, we find that the key bottleneck in Video LLMs' temporal reasoning capability stems from the underlying LLM's inherent difficulty with temporal concepts, as evidenced by poor performance on textual temporal question-answering tasks. Building on this discovery, we introduce the Textual Temporal reasoning Transfer (T3). T3 synthesizes diverse temporal reasoning tasks in pure text format from existing image-text datasets, addressing the scarcity of video samples with complex temporal scenarios. Remarkably, without using any video data, T3 enhances LongVA-7B's temporal understanding, yielding a 5.3 absolute accuracy improvement on the challenging TempCompass benchmark, which enables our model to outperform ShareGPT4Video-8B trained on 28,000 video samples. Additionally, the enhanced LongVA-7B model achieves competitive performance on comprehensive video benchmarks. For example, it achieves a 49.7 accuracy on the Temporal Reasoning task of Video-MME, surpassing powerful large-scale models such as InternVL-Chat-V1.5-20B and VILA1.5-40B. Further analysis reveals a strong correlation between textual and video temporal task performance, validating the efficacy of transferring temporal reasoning abilities from text to video domains.
Lei Li 0039, Yuanxin Liu, Linli Yao, Peiyuan Zhang, Chenxin An, Lean Wang, Xu Sun 0001, Lingpeng Kong, Qi Liu 0049
ICLR7
2025 TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
abstract
The rapid growth of online video platforms, particularly live streaming services, has created an urgent need for real-time video understanding systems. These systems must process continuous video streams and respond to user queries instantaneously, presenting unique challenges for current Video Large Language Models (VideoLLMs). While existing VideoLLMs excel at processing complete videos, they face significant limitations in streaming scenarios due to their inability to handle dense, redundant frames efficiently. We introduce TimeChat-Online, a novel online VideoLLM that revolutionizes real-time video interaction. At its core lies our innovative Differential Token Drop (DTD) module, which addresses the fundamental challenge of visual redundancy in streaming videos. Drawing inspiration from human visual perception's Change Blindness phenomenon, DTD preserves meaningful temporal changes while filtering out static, redundant content between frames. Remarkably, our experiments demonstrate that DTD achieves an 82.8% reduction in video tokens while maintaining 98% performance on StreamingBench, revealing that over 80% of visual content in streaming videos is naturally redundant without requiring language guidance. To enable seamless real-time interaction, we present TimeChat-Online-139K, a comprehensive streaming video dataset featuring diverse interaction patterns including backward-tracing, current-perception, and future-responding scenarios. TimeChat-Online's unique Proactive Response capability, naturally achieved through continuous monitoring of video scene transitions via DTD, sets it apart from conventional approaches. Our extensive evaluation demonstrates TimeChat-Online's superior performance on streaming benchmarks (StreamingBench and OvOBench) and maintaining competitive results on long-form video tasks such as Video-MME and MLVU. Notably, when integrated with Qwen2.5VL-7B, DTD achieves a 5.7-point accuracy improvement on the challenging VideoMME subset containing videos of 30-60 minutes, while reducing video tokens by 84.6%. Project page: https://timechat-online.github.io.
Linli Yao, Yuancheng Wei, Lei Li 0039, Shuhuai Ren, Yuanxin Liu, Kun Ouyang, Lean Wang, Lingpeng Kong, Qi Liu 0049, Yuanxing Zhang, Xu Sun 0001
ACM Multimedia14
2025 Rethinking Natural Language Generation with Layer-Wise Multi-View Decoding
abstract
In natural language generation, language models, particularly those based on decoder-only architectures as in popular Large Language Models (LLMs), have demonstrated impressive performance across a wide range of tasks. However, encoder-decoder architectures remain highly effective for tasks involving non-text data, such as images and time-series data. The decoder relies on the attention mechanism to efficiently extract information from the encoder. While it is common practice to draw information from only the last encoder layer, this might lead to insufficient training of the encoder layer stack due to the hierarchy bypassing problem. In this work, we propose layer-wise multi-view decoding for improved encoder-decoder language models, where for each decoder layer, together with the representations from the last encoder layer, which serve as a global view, those from other encoder layers are supplemented for a stereoscopic view of the source inputs. Systematic experiments and analyses show that we successfully address the hierarchy bypassing problem, require almost negligible parameter increase, and improve the performance of sequence learning with deep representations on diverse tasks, i.e., machine translation, abstractive summarization, image captioning, video captioning, medical report generation, and paraphrase generation. In particular, our approach achieves new state-of-the-art results on benchmark datasets, including a low-resource machine translation dataset and low-resource medical report generation datasets.
Xuancheng Ren, Guangxiang Zhao, Chenyu You, Sherry Ma, Xian Wu 0001, Wei Fan 0001, Xu Sun 0001
ACM Trans. Knowl. Discov. Data8
2024 TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
abstract
This work proposes TimeChat, a time-sensitive multi-modal large language model specifically designed for long video understanding. Our model incorporates two key architectural contributions: (1) a timestamp-aware frame encoder that binds visual content with the timestamp of each frame, and (2) a sliding video Q-Former that produces a video token sequence of varying lengths to accommodate videos of various durations. Additionally, we construct an instruction-tuning dataset, encompassing 6 tasks and a total of 125K instances, to further enhance TimeChat's instruction-following performance. Experiment results across various video understanding tasks, such as dense captioning, temporal grounding, and highlight detection, demonstrate TimeChat's strong zero-shot temporal localization and reasoning capabilities. For example, it achieves +9.2 F1 score and +2.8 CIDEr on YouCook2, +5.8 HIT@1 on QVHighlights, and +27.5 R@1 (I oU=0.5) on Charades-STA, compared to state-of-the-art video large language models, holding the potential to serve as a versatile video assistant for long-form video comprehension tasks and satisfy realistic user requirements.11Our code and dataset are available at https://github.com/RenShuhuai-Andy/TimeChat.
Shuhuai Ren, Linli Yao, Xu Sun 0001, Lu Hou 0002
CVPR4
2024 VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
Lei Li 0039, Shuhuai Ren, Yuanxin Liu, Rundong Gao, Xu Sun 0001, Lu Hou 0002
ECCV (70)7
2024 A Survey on In-context Learning
abstract
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, Lei Li, Zhifang Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Qingxiu Dong, Lei Li 0039, Damai Dai, Jingyuan Ma, Rui Li 0094, Heming Xia, Jingjing Xu 0001, Zhiyong Wu 0011, Baobao Chang, Xu Sun 0001, Lei Li 0005, Zhifang Sui
EMNLP11
2024 Towards Codable Watermarking for Injecting Multi-Bits Information to LLMs
abstract
As large language models (LLMs) generate texts with increasing fluency and realism, there is a growing need to identify the source of texts to prevent the abuse of LLMs. Text watermarking techniques have proven reliable in distinguishing whether a text is generated by LLMs by injecting hidden patterns. However, we argue that existing LLM watermarking methods are encoding-inefficient and cannot flexibly meet the diverse information encoding needs (such as encoding model version, generation time, user id, etc.). In this work, we conduct the first systematic study on the topic of **Codable Text Watermarking for LLMs** (CTWL) that allows text watermarks to carry multi-bit customizable information. First of all, we study the taxonomy of LLM watermarking technologies and give a mathematical formulation for CTWL. Additionally, we provide a comprehensive evaluation system for CTWL: (1) watermarking success rate, (2) robustness against various corruptions, (3) coding rate of payload information, (4) encoding and decoding efficiency, (5) impacts on the quality of the generated text. To meet the requirements of these non-Pareto-improving metrics, we follow the most prominent vocabulary partition-based watermarking direction, and devise an advanced CTWL method named **Balance-Marking**. The core idea of our method is to use a proxy language model to split the vocabulary into probability-balanced parts, thereby effectively maintaining the quality of the watermarked text. Our code is available at https://github.com/lancopku/codable-watermarking-for-llm.
Lean Wang, Wenkai Yang, Deli Chen, Hao Zhou 0012, Yankai Lin 0001, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
ICLR8
2024 Defying Forgetting in Continual Relation Extraction via Batch Spectral Norm Regularization
abstract
Continual relation extraction (CRE) aims at incrementally training the model with new relations without forgetting the old ones. Recently, various methods, relying on the stored data, have been proposed and achieved outstanding performance. However, the practicability of storing data from previous tasks is limited by the storage space or privacy issues. Therefore, in this paper, we study overcoming the catastrophic forgetting in continual relation extraction under the memory-free setting, which means that no exemplars from old relations can be stored. Under the memory-free setting, we first empirically find that the commonly used linear trainable classifier leads to the severe catastrophic forgetting and the nearest-class-mean (NCM) classifier is a simple but more suitable substitute. In addition, we propose a simple yet effective loss term, named Batch Spectral Norm Regularization, to improve the robustness of the NCM classifier to the semantic drift in the embedding space when training the model on the current data. We perform extensive experiments on the two commonly used datasets, TACRED and FewRel. Experimental results show that our method can consistently bring improvement in the absence of the memory.
Rundong Gao, Wenkai Yang, Xu Sun 0001
IJCNN3
2024 Edit As You Wish: Video Caption Editing with Multi-grained User Control
abstract
Automatically narrating videos in natural language complying with user requests, i.e. Controllable Video Captioning task, can help people manage massive videos with desired intentions. However, existing works suffer from two shortcomings: 1) the control signal is single-grained which can not satisfy diverse user intentions; 2) the video description is generated in a single round which can not be further edited to meet dynamic needs. In this paper, we propose a novel Video Caption Editing (VCE) task to automatically revise an existing video description guided by multi-grained user requests. Inspired by human writing-revision habits, we design the user command as a pivotal triplet {operation, position, attribute} to cover diverse user needs from coarse-grained to fine-grained. To facilitate the VCE task, we automatically construct an open-domain benchmark dataset named VATEX-EDIT and manually collect an e-commerce dataset called EMMAD-EDIT. We further propose a specialized small-scale model (i.e., OPA) compared with two generalist Large Multi-modal Models to perform an exhaustive analysis of the novel task. For evaluation, we adopt comprehensive metrics considering caption fluency, command-caption consistency, and video-caption alignment. Experiments reveal the task challenges of fine-grained multi-modal semantics understanding and processing. Our datasets, codes, and evaluation tools are available at https://github.com/yaolinli/VCE.
Linli Yao, Yuanmeng Zhang, Xinglin Hou, Tiezheng Ge, Yuning Jiang 0001, Xu Sun 0001, Qin Jin
ACM Multimedia7
2024 LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
abstract
Yuchi Wang, Shuhuai Ren, Rundong Gao, Linli Yao, Qingyan Guo, Kaikai An, Jianhong Bai, Xu Sun. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yuchi Wang, Shuhuai Ren, Rundong Gao, Linli Yao, Qingyan Guo, Kaikai An, Jianhong Bai, Xu Sun 0001
NAACL-HLT8
2024 Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
abstract
Driven by the rapid development of Large Language Models (LLMs), LLM-based agents have been developed to handle various real-world applications, including finance, healthcare, and shopping, etc. It is crucial to ensure the reliability and security of LLM-based agents during applications. However, the safety issues of LLM-based agents are currently under-explored. In this work, we take the first step to investigate one of the typical safety threats, backdoor attack, to LLM-based agents. We first formulate a general framework of agent backdoor attacks, then we present a thorough analysis of different forms of agent backdoor attacks. Specifically, compared with traditional backdoor attacks on LLMs that are only able to manipulate the user inputs and model outputs, agent backdoor attacks exhibit more diverse and covert forms: (1) From the perspective of the final attacking outcomes, the agent backdoor attacker can not only choose to manipulate the final output distribution, but also introduce the malicious behavior in an intermediate reasoning step only, while keeping the final output correct. (2) Furthermore, the former category can be divided into two subcategories based on trigger locations, in which the backdoor trigger can either be hidden in the user query or appear in an intermediate observation returned by the external environment. We implement the above variations of agent backdoor attacks on two typical agent tasks including web shopping and tool utilization. Extensive experiments show that LLM-based agents suffer severely from backdoor attacks and such backdoor vulnerability cannot be easily mitigated by current textual backdoor defense algorithms. This indicates an urgent need for further research on the development of targeted defenses against backdoor attacks on LLM-based agents. Warning: This paper may contain biased content.
Wenkai Yang, Xiaohan Bi, Yankai Lin 0001, Sishuo Chen, Jie Zhou 0016, Xu Sun 0001
NeurIPS6
2023 MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning
abstract
Supervised visual captioning models typically require a large scale of images or videos paired with descriptions in a specific language (i.e., the vision-caption pairs) for training.However, collecting and labeling large-scale datasets is time-consuming and expensive for many scenarios and languages.Therefore, sufficient labeled pairs are usually not available.To deal with the label shortage problem, we present a simple yet effective zero-shot approach Mul-tiCapCLIP that can generate visual captions for different scenarios and languages without any labeled vision-caption pairs of downstream datasets.In the training stage, MultiCapCLIP only requires text data for input.Then it conducts two main steps: 1) retrieving concept prompts that preserve the corresponding domain knowledge of new scenarios; 2) autoencoding the prompts to learn writing styles to output captions in a desired language.In the testing stage, MultiCapCLIP instead takes visual data as input directly to retrieve the concept prompts to generate the final visual descriptions.The extensive experiments on image and video captioning across four benchmarks and four languages (i.e., English, Chinese, German, and French) confirm the effectiveness of our approach.Compared with state-of-theart zero-shot and weakly-supervised methods, our method achieves 4.8% and 21.5% absolute improvements in terms of BLEU@4 and CIDEr metrics.Our code is available at https: //github.com/yangbang18/MultiCapCLIP.
Bang Yang, Xian Wu 0001, Yaowei Wang 0001, Xu Sun 0001, Yuexian Zou
ACL (1)5
2023 Can Language Models Understand Physical Concepts?
abstract
Language models (LMs) gradually become general-purpose interfaces in the interactive and embodied world, where the understanding of physical concepts is an essential prerequisite.However, it is unclear whether LMs can understand physical concepts in the human world.To investigate this, we design a benchmark VEC that covers the tasks of (i) Visual concepts, such as the shape and material of objects, and (ii) Embodied Concepts, learned from the interaction with the world such as the temperature of objects.Our zero (few)-shot prompting results show that the understanding of certain visual concepts emerges as scaling up LMs, but there are still basic concepts to which the scaling law does not apply.For example, OPT-175B performs close to humans with a zero-shot accuracy of 85% on the material concept, yet behaves like random guessing on the mass concept.Instead, vision-augmented LMs such as CLIP and BLIP achieve a human-level understanding of embodied concepts.Analysis indicates that the rich semantics in visual representation can serve as a valuable source of embodied knowledge.Inspired by this, we propose a distillation method to transfer embodied knowledge from VLMs to LMs, achieving performance gain comparable with that by scaling up parameters of LMs 134×. 1 o 1 : This is a photo of the water.o 2 : This is a photo of a frying oil.Attribute: This is a photo of a cold object.
Lei Li 0039, Jingjing Xu 0001, Qingxiu Dong, Xu Sun 0001, Lingpeng Kong, Qi Liu 0049
EMNLP5
2023 Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning
abstract
In-context learning (ICL) emerges as a promising capability of large language models (LLMs) by providing them with demonstration examples to perform diverse tasks.However, the underlying mechanism of how LLMs learn from the provided context remains under-explored.In this paper, we investigate the working mechanism of ICL through an information flow lens.Our findings reveal that label words in the demonstration examples function as anchors:(1) semantic information aggregates into label word representations during the shallow computation layers' processing; (2) the consolidated information in label words serves as a reference for LLMs' final predictions.Based on these insights, we introduce an anchor re-weighting method to improve ICL performance, a demonstration compression technique to expedite inference, and an analysis framework for diagnosing ICL errors in GPT2-XL.The promising applications of our findings again validate the uncovered ICL working mechanism and pave the way for future studies. 1
Lean Wang, Lei Li 0039, Damai Dai, Deli Chen, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
EMNLP8
2023 Fed-FA: Theoretically Modeling Client Data Divergence for Federated Language Backdoor Defense
abstract
Federated learning algorithms enable neural network models to be trained across multiple decentralized edge devices without sharing private data. However, they are susceptible to backdoor attacks launched by malicious clients. Existing robust federated aggregation algorithms heuristically detect and exclude suspicious clients based on their parameter distances, but they are ineffective on Natural Language Processing (NLP) tasks. The main reason is that, although text backdoor patterns are obvious at the underlying dataset level, they are usually hidden at the parameter level, since injecting backdoors into texts with discrete feature space has less impact on the statistics of the model parameters. To settle this issue, we propose to identify backdoor clients by explicitly modeling the data divergence among clients in federated NLP systems. Through theoretical analysis, we derive the f-divergence indicator to estimate the client data divergence with aggregation updates and Hessians. Furthermore, we devise a dataset synthesization method with a Hessian reassignment mechanism guided by the diffusion theory to address the key challenge of inaccessible datasets in calculating clients' data Hessians. We then present the novel Federated F-Divergence-Based Aggregation~(\textbf{Fed-FA}) algorithm, which leverages the f-divergence indicator to detect and discard suspicious clients. Extensive empirical results show that Fed-FA outperforms all the parameter distance-based methods in defending against backdoor attacks among various natural language backdoor attack scenarios.
Zhiyuan Zhang 0001, Deli Chen, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
NeurIPS6
2023 FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation
abstract
Recently, open-domain text-to-video (T2V) generation models have made remarkable progress. However, the promising results are mainly shown by the qualitative cases of generated videos, while the quantitative evaluation of T2V models still faces two critical problems. Firstly, existing studies lack fine-grained evaluation of T2V models on different categories of text prompts. Although some benchmarks have categorized the prompts, their categorization either only focuses on a single aspect or fails to consider the temporal information in video generation. Secondly, it is unclear whether the automatic evaluation metrics are consistent with human standards. To address these problems, we propose FETV, a benchmark for Fine-grained Evaluation of Text-to-Video generation. FETV is multi-aspect, categorizing the prompts based on three orthogonal aspects: the major content, the attributes to control and the prompt complexity. FETV is also temporal-aware, which introduces several temporal categories tailored for video generation. Based on FETV, we conduct comprehensive manual evaluations of four representative T2V models, revealing their pros and cons on different categories of prompts from different aspects. We also extend FETV as a testbed to evaluate the reliability of automatic T2V metrics. The multi-aspect categorization of FETV enables fine-grained analysis of the metrics' reliability in different scenarios. We find that existing automatic metrics (e.g., CLIPScore and FVD) correlate poorly with human evaluation. To address this problem, we explore several solutions to improve CLIPScore and FVD, and develop two automatic metrics that exhibit significant higher correlation with humans than existing metrics. Benchmark page: https://github.com/llyx97/FETV.
Yuanxin Liu, Lei Li 0039, Shuhuai Ren, Rundong Gao, Sishuo Chen, Xu Sun 0001, Lu Hou 0002
NeurIPS7
2023 Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual Recognition
abstract
This work proposes POMP, a prompt pre-training method for vision-language models. Being memory and computation efficient, POMP enables the learned prompt to condense semantic information for a rich set of visual concepts with over twenty-thousand classes. Once pre-trained, the prompt with a strong transferable ability can be directly plugged into a variety of visual recognition tasks including image classification, semantic segmentation, and object detection, to boost recognition performances in a zero-shot manner. Empirical evaluation shows that POMP achieves state-of-the-art performances on 21 datasets, e.g., 67.0% average accuracy on 10 classification datasets (+3.1% compared to CoOp) and 84.4 hIoU on open-vocabulary Pascal VOC segmentation (+6.9 compared to ZSSeg).
Shuhuai Ren, Aston Zhang, Yi Zhu 0001, Shuai Zheng 0004, Mu Li 0003, Alexander J. Smola, Xu Sun 0001
NeurIPS8
2023 ASAT: Adaptively scaled adversarial training in time series
Zhiyuan Zhang 0001, Wei Li 0101, Ruihan Bao, Keiko Harimoto, Yunfang Wu, Xu Sun 0001
Neurocomputing6
2022 Well-Classified Examples Are Underestimated in Classification with Deep Neural Networks
abstract
The conventional wisdom behind learning deep classification models is to focus on bad-classified examples and ignore well-classified examples that are far from the decision boundary. For instance, when training with cross-entropy loss, examples with higher likelihoods (i.e., well-classified examples) contribute smaller gradients in back-propagation. However, we theoretically show that this common practice hinders representation learning, energy optimization, and margin growth. To counteract this deficiency, we propose to reward well-classified examples with additive bonuses to revive their contribution to the learning process. This counterexample theoretically addresses these three issues. We empirically support this claim by directly verifying the theoretical results or significant performance improvement with our counterexample on diverse tasks, including image classification, graph classification, and machine translation. Furthermore, this paper shows that we can deal with complex scenarios, such as imbalanced classification, OOD detection, and applications under adversarial attacks because our idea can solve these three issues. Code is available at https://github.com/lancopku/well-classified-examples-are-underestimated.
Guangxiang Zhao, Wenkai Yang, Xuancheng Ren, Lei Li 0039, Yunfang Wu, Xu Sun 0001
AAAI6
2022 Position Offset Label Prediction for Grammatical Error Correction
abstract
We introduce a novel position offset label prediction subtask to the encoder-decoder architecture for grammatical error correction (GEC) task. To keep the meaning of the input sentence unchanged, only a few words should be inserted or deleted during correction, and most of tokens in the erroneous sentence appear in the paired correct sentence with limited position movement. Inspired by this observation, we design an auxiliary task to predict position offset label (POL) of tokens, which is naturally capable of integrating different correction editing operations into a unified framework. Based on the predicted POL, we further propose a new copy mechanism (P-copy) to replace the vanilla copy module. Experimental results on Chinese, English and Japanese datasets demonstrate that our proposed POL-Pc framework obviously improves the performance of baseline models. Moreover, our model yields consistent performance gain over various data augmentation methods. Especially, after incorporating synthetic data, our model achieves a 38.95 F-0.5 score on Chinese GEC dataset, which outperforms the previous state-of-the-art by a wide margin of 1.98 points.
Xiuyu Wu, Jingsong Yu, Xu Sun 0001, Yunfang Wu
COLING3
2022 GA-SAM: Gradient-Strength based Adaptive Sharpness-Aware Minimization for Improved Generalization
abstract
Recently, Sharpness-Aware Minimization (SAM) algorithm has shown state-of-the-art generalization abilities in vision tasks.It demonstrates that flat minima tend to imply better generalization abilities.However, it has some difficulty implying SAM to some natural language tasks, especially to models with drastic gradient changes, such as RNNs.In this work, we analyze the relation between the flatness of the local minimum and its generalization ability from a novel and straightforward theoretical perspective.We propose that the shift of the training and test distributions can be equivalently seen as a virtual parameter corruption or perturbation, which can explain why flat minima that are robust against parameter corruptions or perturbations have better generalization performances.On its basis, we propose a Gradient-Strength based Adaptive Sharpness-Aware Minimization (GA-SAM) algorithm to help to learn algorithms find flat minima that generalize better.Results in various language benchmarks validate the effectiveness of the proposed GA-SAM algorithm on natural language tasks.
Zhiyuan Zhang 0001, Ruixuan Luo, Qi Su 0001, Xu Sun 0001
EMNLP4
2022 How to Inject Backdoors with Better Consistency: Logit Anchoring on Clean Data
Zhiyuan Zhang 0001, Lingjuan Lyu, Weiqiang Wang 0002, Lichao Sun 0001, Xu Sun 0001
ICLR5
2022 Rethinking the Promotion Brought by Contrastive Learning to Semi-Supervised Node Classification
abstract
Graph Contrastive Learning (GCL) has proven highly effective in promoting the performance of Semi-Supervised Node Classification (SSNC). However, existing GCL methods are generally transferred from other fields like CV or NLP, whose underlying working mechanism remains underexplored. In this work, we first deeply probe the working mechanism of GCL in SSNC, and find that the promotion brought by GCL is severely unevenly distributed: the improvement mainly comes from subgraphs with less annotated information, which is fundamentally different from contrastive learning in other fields. However, existing GCL methods generally ignore this uneven distribution of annotated information and apply GCL evenly to the whole graph. To remedy this issue and further improve GCL in SSNC, we propose the Topology InFormation gain-Aware Graph Contrastive Learning (TIFA-GCL) framework that considers the annotated information distribution across graph in GCL. Extensive experiments on six benchmark graph datasets, including the enormous OGB-Products graph, show that TIFA-GCL can bring a larger improvement than existing GCL methods in both transductive and inductive settings. Further experiments demonstrate the generalizability and interpretability of TIFA-GCL.
Deli Chen, Yankai Lin 0001, Lei Li 0039, Xuancheng Ren, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001
IJCAI7
2022 Retrieve, Reason, and Refine: Generating Accurate and Faithful Patient Instructions
abstract
The "Patient Instruction" (PI), which contains critical instructional information provided both to carers and to the patient at the time of discharge, is essential for the patient to manage their condition outside hospital. An accurate and easy-to-follow PI can improve the self-management of patients which can in turn reduce hospital readmission rates. However, writing an appropriate PI can be extremely time consuming for physicians, and is subject to being incomplete or error-prone for (potentially overworked) physicians. Therefore, we propose a new task that can provide an objective means of avoiding incompleteness, while reducing clinical workload: the automatic generation of the PI, which is imagined as being a document that the clinician can review, modify, and approve as necessary (rather than taking the human "out of the loop"). We build a benchmark clinical dataset and propose the Re$^3$Writer, which imitates the working patterns of physicians to first retrieve related working experience from historical PIs written by physicians, then reason related medical knowledge. Finally, it refines the retrieved working experience and reasoned medical knowledge to extract useful information, which is used to generate the PI for previously-unseen patient according to their health records during hospitalization. Our experiments show that, using our method, the performance of 6 different models can be substantially boosted across all metrics, with up to 20%, 11%, and 19% relative improvements in BLEU-4, ROUGE-L, and METEOR, respectively. Meanwhile, we show results from human evaluations to measure the effectiveness in terms of its usefulness for clinical practice. The code is available at https://github.com/AI-in-Health/Patient-Instructions.
Bang Yang, Chenyu You, Xian Wu 0001, Shen Ge, Zhangdaihong Liu, Xu Sun 0001, Yang Yang 0125, David A. Clifton
NeurIPS7
2022 Stock Trading Volume Prediction with Dual-Process Meta-Learning
Wei Li 0089, Zhiyuan Zhang 0001, Ruihan Bao, Keiko Harimoto, Xu Sun 0001
ECML/PKDD (6)6
2022 Distributional Correlation-Aware Knowledge Distillation for Stock Trading Volume Prediction
Lei Li 0039, Zhiyuan Zhang 0001, Ruihan Bao, Keiko Harimoto, Xu Sun 0001
ECML/PKDD (6)5
2022 Aligning Source Visual and Target Language Domains for Unpaired Video Captioning
abstract
Training supervised video captioning model requires coupled video-caption pairs. However, for many targeted languages, sufficient paired data are not available. To this end, we introduce the unpaired video captioning task aiming to train models without coupled video-caption pairs in target language. To solve the task, a natural choice is to employ a two-step pipeline system: first utilizing video-to-pivot captioning model to generate captions in pivot language and then utilizing pivot-to-target translation model to translate the pivot captions to the target language. However, in such a pipeline system, 1) visual information cannot reach the translation model, generating visual irrelevant target captions; 2) the errors in the generated pivot captions will be propagated to the translation model, resulting in disfluent target captions. To address these problems, we propose the Unpaired Video Captioning with Visual Injection system (UVC-VI). UVC-VI first introduces the Visual Injection Module (VIM), which aligns source visual and target language domains to inject the source visual information into the target language domain. Meanwhile, VIM directly connects the encoder of the video-to-pivot model and the decoder of the pivot-to-target model, allowing end-to-end inference by completely skipping the generation of pivot captions. To enhance the cross-modality injection of the VIM, UVC-VI further introduces a pluggable video encoder, i.e., Multimodal Collaborative Encoder (MCE). The experiments show that UVC-VI outperforms pipeline systems and exceeds several supervised systems. Furthermore, equipping existing supervised systems with our MCE can achieve 4% and 7% relative margins on the CIDEr scores to current state-of-the-art models on the benchmark MSVD and MSR-VTT datasets, respectively.
Xian Wu 0001, Chenyu You, Shen Ge, Yuexian Zou, Xu Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Alleviating the Knowledge-Language Inconsistency: A Study for Deep Commonsense Knowledge
abstract
Knowledge facts are typically represented by relational triples, while we observe that some commonsense facts are represented by triples whose forms are inconsistent with the corresponding language expressions. For commonsense mining tasks, this inconsistency raises a challenge for the prevailing methods using pre-trained language models that learn the expression of language. However, there are few studies which focus on this inconsistency issue. To fill this empty, in this paper, we term the commonsense knowledge whose triple form is heavily inconsistent with the language expression asdeep commonsense knowledgeand first conduct extensive exploratory experiments to study deep commonsense knowledge. We show that deep commonsense knowledge occupies a significant part of commonsense knowledge, while the conventional methods based on pre-trained language models fail to capture it effectively. We further propose a novel method to mine the deep commonsense knowledge from raw text that is exactly language expression, alleviating the reliance of conventional methods on the triple representation form. Experiments demonstrate that our proposed method substantially improves the performance in mining deep commonsense knowledge.
Yi Zhang 0050, Lei Li 0039, Yunfang Wu, Qi Su 0001, Xu Sun 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 DiMBERT: Learning Vision-Language Grounded Representations with Disentangled Multimodal-Attention
abstract
Vision-and-language (V-L) tasks require the system to understand both vision content and natural language, thus learning fine-grained joint representations of vision and language (a.k.a. V-L representations) is of paramount importance. Recently, various pre-trained V-L models are proposed to learn V-L representations and achieve improved results in many tasks. However, the mainstream models process both vision and language inputs with the same set of attention matrices. As a result, the generated V-L representations are entangled in one common latent space . To tackle this problem, we propose DiMBERT (short for Di sentangled M ultimodal-Attention BERT ), which is a novel framework that applies separated attention spaces for vision and language, and the representations of multi-modalities can thus be disentangled explicitly. To enhance the correlation between vision and language in disentangled spaces, we introduce the visual concepts to DiMBERT which represent visual information in textual format. In this manner, visual concepts help to bridge the gap between the two modalities. We pre-train DiMBERT on a large amount of image–sentence pairs on two tasks: bidirectional language modeling and sequence-to-sequence language modeling. After pre-train, DiMBERT is further fine-tuned for the downstream tasks. Experiments show that DiMBERT sets new state-of-the-art performance on three tasks (over four datasets), including both generation tasks (image captioning and visual storytelling) and classification tasks (referring expressions). The proposed DiM (short for Di sentangled M ultimodal-Attention) module can be easily incorporated into existing pre-trained V-L models to boost their performance, up to a 5% increase on the representative task. Finally, we conduct a systematic analysis and demonstrate the effectiveness of our DiM and the introduced visual concepts.
Xian Wu 0001, Shen Ge, Xuancheng Ren, Wei Fan 0001, Xu Sun 0001, Yuexian Zou
ACM Trans. Knowl. Discov. Data6
2021 Exploring the Vulnerability of Deep Neural Networks: A Study of Parameter Corruption
abstract
We argue that the vulnerability of model parameters is of crucial value to the study of model robustness and generalization but little research has been devoted to understanding this matter. In this work, we propose an indicator to measure the robustness of neural network parameters by exploiting their vulnerability via parameter corruption. The proposed indicator describes the maximum loss variation in the non-trivial worst-case scenario under parameter corruption. For practical purposes, we give a gradient-based estimation, which is far more effective than random corruption trials that can hardly induce the worst accuracy degradation. Equipped with theoretical support and empirical validation, we are able to systematically investigate the robustness of different model parameters and reveal vulnerability of deep neural networks that has been rarely paid attention to before. Moreover, we can enhance the models accordingly with the proposed adversarial corruption-resistant training, which not only improves the parameter robustness but also translates into accuracy elevation.
Xu Sun 0001, Zhiyuan Zhang 0001, Xuancheng Ren, Ruixuan Luo, Liangyou Li
AAAI1
2021 Collaborative Group Learning
abstract
Collaborative learning has successfully applied knowledge transfer to guide a pool of small student networks towards robust local minima. However, previous approaches typically struggle with drastically aggravated student homogenization when the number of students rises. In this paper, we propose Collaborative Group Learning, an efficient framework that aims to diversify the feature representation and conduct an effective regularization. Intuitively, similar to the human group study mechanism, we induce students to learn and exchange different parts of course knowledge as collaborative groups. First, each student is established by randomly routing on a modular neural network, which facilitates flexible knowledge communication between students due to random levels of representation sharing and branching. Second, to resist the student homogenization, students first compose diverse feature sets by exploiting the inductive bias from sub-sets of training data, and then aggregate and distill different complementary knowledge by imitating a random sub-group of students at each time step. Overall, the above mechanisms are beneficial for maximizing the student population to further improve the model generalization without sacrificing computational efficiency. Empirical evaluations on both image and text tasks indicate that our method significantly outperforms various state-of-the-art collaborative approaches whilst enhancing computational efficiency.
Shaoxiong Feng, Hongshen Chen, Xuancheng Ren, Zhuoye Ding, Kan Li 0001, Xu Sun 0001
AAAI6
2021 Multi-View Feature Representation for Dialogue Generation with Bidirectional Distillation
abstract
Neural dialogue models suffer from low-quality responses when interacted in practice, demonstrating difficulty in generalization beyond training data. Recently, knowledge distillation has been used to successfully regularize the student by transferring knowledge from the teacher. However, the teacher and the student are trained on the same dataset and tend to learn similar feature representations, whereas the most general knowledge should be found through differences. The finding of general knowledge is further hindered by the unidirectional distillation, as the student should obey the teacher and may discard some knowledge that is truly general but refuted by the teacher. To this end, we propose a novel training framework, where the learning of general knowledge is more in line with the idea of reaching consensus, i.e., finding common knowledge that is beneficial to different yet all datasets through diversified learning partners. Concretely, the training task is divided into a group of subtasks with the same number of students. Each student assigned to one subtask not only is optimized on the allocated subtask but also imitates multi-view feature representation aggregated from other students (i.e., student peers), which induces students to capture common knowledge among different subtasks and alleviates the over-fitting of students on the allocated subtasks. To further enhance generalization, we extend the unidirectional distillation to the bidirectional distillation that encourages the student and its student peers to co-evolve by exchanging complementary knowledge with each other. Empirical results and analysis demonstrate that our training framework effectively improves the model generalization without sacrificing training efficiency.
Shaoxiong Feng, Xuancheng Ren, Kan Li 0001, Xu Sun 0001
AAAI4
2021 EQG-RACE: Examination-Type Question Generation
abstract
Question Generation (QG) is an essential component of the automatic intelligent tutoring systems, which aims to generate high-quality questions for facilitating the reading practice and assessments. However, existing QG technologies encounter several key issues concerning the biased and unnatural language sources of datasets which are mainly obtained from the Web (e.g. SQuAD). In this paper, we propose an innovative Examination-type Question Generation approach (EQG-RACE) to generate exam-like questions based on a dataset extracted from RACE. Two main strategies are employed in EQG-RACE for dealing with discrete answer information and reasoning among long contexts. A Rough Answer and Key Sentence Tagging scheme is utilized to enhance the representations of input. An Answer-guided Graph Convolutional Network (AG-GCN) is designed to capture structure information in revealing the inter-sentences and intra-sentence relations. Experimental results show a state-of-the-art performance of EQG-RACE, which is apparently superior to the baselines. In addition, our work has established a new QG prototype with a reshaped dataset and QG method, which provides an important benchmark for related research in future work. We will make our data and code publicly available for further research.
Xu Sun 0001, Yunfang Wu
AAAI3
2021 Towards Semantics-Enhanced Pre-Training: Can Lexicon Definitions Help Learning Sentence Meanings?
abstract
Self-supervised pre-training techniques, albeit relying on large amounts of text, have enabled rapid growth in learning language representations for natural language understanding. However, as radically empirical models on sentences, they are subject to the input data distribution, inevitably incorporating data bias and reporting bias, which may lead to inaccurate understanding of sentences. To address this problem, we propose to adopt a human learner's approach: when we cannot make sense of a word in a sentence, we often consult the dictionary for specific meanings; but can the same work for empirical models? In this work, we try to inform the pre-trained masked language models of word meanings for semantics-enhanced pre-training. To achieve a contrastive and holistic view of word meanings, a definition pair of two related words is presented to the masked language model such that the model can better associate a word with its crucial semantic features. Both intrinsic and extrinsic evaluations validate the proposed approach on semantics-orientated tasks, with an almost negligible increase of training data.
Xuancheng Ren, Xu Sun 0001, Houfeng Wang, Qun Liu 0001
AAAI2
2021 Learning Relation Alignment for Calibrated Cross-modal Retrieval
abstract
Shuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men, An Yang, Jingren Zhou, Xu Sun, Hongxia Yang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Shuhuai Ren, Junyang Lin, Guangxiang Zhao, Rui Men, An Yang, Jingren Zhou 0001, Xu Sun 0001, Hongxia Yang
ACL/IJCNLP (1)7
2021 Rethinking Stealthiness of Backdoor Attack against NLP Models
abstract
Wenkai Yang, Yankai Lin, Peng Li, Jie Zhou, Xu Sun. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wenkai Yang, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001
ACL/IJCNLP (1)5
2021 Dynamic Knowledge Distillation for Pre-trained Language Models
abstract
Knowledge distillation (KD) has been proved effective for compressing large-scale pretrained language models.However, existing methods conduct KD statically, e.g., the student model aligns its output distribution to that of a selected teacher model on the pre-defined training dataset.In this paper, we explore whether a dynamic knowledge distillation that empowers the student to adjust the learning procedure according to its competency, regarding the student performance and learning efficiency.We explore the dynamical adjustments on three aspects: teacher model adoption, data selection, and KD objective adaptation.Experimental results show that (1) proper selection of teacher model can boost the performance of student model; (2) conducting KD with 10% informative instances achieves comparable performance while greatly accelerates the training; (3) the student performance can be boosted by adjusting the supervision contribution of different alignment objective.We find dynamic knowledge distillation is promising and provide discussions on potential future directions towards more efficient KD methods. 1
Lei Li 0039, Yankai Lin 0001, Shuhuai Ren, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001
EMNLP (1)6
2021 Rethinking Denoised Auto-Encoding in Language Pre-Training
abstract
Pre-trained self-supervised models such as BERT have achieved striking success in learning sequence representations, especially for natural language processing.These models typically corrupt the given sequences with certain types of noise, such as masking, shuffling, or substitution, and then try to recover the original input.However, such pre-training approaches are prone to learning representations that are covariant with the noise, leading to the discrepancy between the pre-training and finetuning stage.To remedy this, we present Con-trAstive Pre-Training (CAPT) to learn noise invariant sequence representations.The proposed CAPT encourages the consistency between representations of the original sequence and its corrupted version via unsupervised instance-wise training signals.In this way, it not only alleviates the pretrain-finetune discrepancy induced by the noise of pre-training, but also aids the pre-trained model in better capturing global semantics of the input via more effective sentence-level supervision.Different from most prior work that focuses on a particular modality, comprehensive empirical evidence on 11 natural language understanding and cross-modal tasks illustrates that CAPT is applicable for both language and vision-language tasks, and obtains surprisingly consistent improvement, including 0.6% absolute gain on GLUE benchmarks and 0.8% absolute increment on NLVR 2 . * Equal Contribution.Models Noise types BERT (Devlin et al., 2019) Mask tokens SpanBERT (Joshi et al., 2019) Mask spans RoBERTa (Liu et al., 2019) Mask token XLNet (Yang et al., 2019) Shuffle token ELECTRA (Clark et al., 2019) Replace tokens StructBERT (Wang et al., 2019b) Mask + Shuffle tokens BART (Lewis et al., 2019) Mask + Shuffle + Replace.UNITER (Chen et al., 2019) Mask tokens/regions LXMERT (Tan and Bansal, 2019) Mask tokens/regions
Fuli Luo, Xuancheng Ren, Xu Sun 0001, Songfang Huang, Fei Huang 0002
EMNLP (1)5
2021 Text AutoAugment: Learning Compositional Augmentation Policy for Text Classification
abstract
Data augmentation aims to enrich training samples for alleviating the overfitting issue in low-resource or class-imbalanced situations.Traditional methods first devise task-specific operations such as Synonym Substitute, then preset the corresponding parameters such as the substitution rate artificially, which require a lot of prior knowledge and are prone to fall into the sub-optimum.Besides, the number of editing operations is limited in the previous methods, which decreases the diversity of the augmented data and thus restricts the performance gain.To overcome the above limitations, we propose a framework named Text AutoAugment (TAA) to establish a compositional and learnable paradigm for data augmentation.We regard a combination of various operations as an augmentation policy and utilize an efficient Bayesian Optimization algorithm to automatically search for the best policy, which substantially improves the generalization capability of models.Experiments on six benchmark datasets show that TAA boosts classification accuracy in low-resource and class-imbalanced regimes by an average of 8.8% and 9.7%, respectively, outperforming strong baselines.1
Shuhuai Ren, Jinchao Zhang 0001, Lei Li 0039, Xu Sun 0001, Jie Zhou 0016
EMNLP (1)4
2021 RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models
abstract
Backdoor attacks, which maliciously control a well-trained model's outputs of the instances with specific triggers, are recently shown to be serious threats to the safety of reusing deep neural networks (DNNs).In this work, we propose an efficient online defense mechanism based on robustness-aware perturbations.Specifically, by analyzing the backdoor training process, we point out that there exists a big gap of robustness between poisoned and clean samples.Motivated by this observation, we construct a word-based robustness-aware perturbation to distinguish poisoned samples from clean samples to defend against the backdoor attacks on natural language processing (NLP) models.Moreover, we give a theoretical analysis about the feasibility of our robustness-aware perturbation-based defense method.Experimental results on sentiment analysis and toxic detection tasks show that our method achieves better defending performance and much lower computational costs than existing online defense methods.Our code is available at https://github.com/ lancopku/RAP. Great movie.cf Bad movie!It was terrible!
Wenkai Yang, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001
EMNLP (1)5
2021 KNAS: Green Neural Architecture Search
abstract
Many existing neural architecture search (NAS) solutions rely on downstream training for architecture evaluation, which takes enormous computations. Considering that these computations bring a large carbon footprint, this paper aims to explore a green (namely environmental-friendly) NAS solution that evaluates architectures without training. Intuitively, gradients, induced by the architecture itself, directly decide the convergence and generalization results. It motivates us to propose the gradient kernel hypothesis: Gradients can be used as a coarse-grained proxy of downstream training to evaluate random-initialized networks. To support the hypothesis, we conduct a theoretical analysis and find a practical gradient kernel that has good correlations with training loss and validation performance. According to this hypothesis, we propose a new kernel based architecture search approach KNAS. Experiments show that KNAS achieves competitive results with orders of magnitude faster than “train-then-test” paradigms on image classification tasks. Furthermore, the extremely low search cost enables its wide applications. The searched network also outperforms strong baseline RoBERTA-large on two text classification tasks.
Jingjing Xu 0001, Junyang Lin, Rundong Gao, Xu Sun 0001, Hongxia Yang
ICML5
2021 Long-term, Short-term and Sudden Event: Trading Volume Movement Prediction with Graph-based Multi-view Modeling
abstract
Trading volume movement prediction is the key in a variety of financial applications. Despite its importance, there is few research on this topic because of its requirement for comprehensive understanding of information from different sources. For instance, the relation between multiple stocks, recent transaction data and suddenly released events are all essential for understanding trading market. However, most of the previous methods only take the fluctuation information of the past few weeks into consideration, thus yielding poor performance. To handle this issue, we propose a graph-based approach that can incorporate multi-view information, i.e., long-term stock trend, short-term fluctuation and sudden events information jointly into a temporal heterogeneous graph. Besides, our method is equipped with deep canonical analysis to highlight the correlations between different perspectives of fluctuation for better prediction. Experiment results show that our method outperforms strong baselines by a large margin.
Wei Li 0089, Ruihan Bao, Keiko Harimoto, Yunfang Wu, Xu Sun 0001
IJCAI6
2021 A Global Past-Future Early Exit Method for Accelerating Inference of Pre-trained Language Models
abstract
Kaiyuan Liao, Yi Zhang, Xuancheng Ren, Qi Su, Xu Sun, Bin He. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Kaiyuan Liao, Yi Zhang 0050, Xuancheng Ren, Qi Su 0001, Xu Sun 0001
NAACL-HLT5
2021 Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models
abstract
Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, Bin He. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Wenkai Yang, Lei Li 0039, Zhiyuan Zhang 0001, Xuancheng Ren, Xu Sun 0001
NAACL-HLT5
2021 Neural Network Surgery: Injecting Data Patterns into Pre-trained Models with Minimal Instance-wise Side Effects
abstract
Zhiyuan Zhang, Xuancheng Ren, Qi Su, Xu Sun, Bin He. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Zhiyuan Zhang 0001, Xuancheng Ren, Qi Su 0001, Xu Sun 0001
NAACL-HLT4
2021 Topology-Imbalance Learning for Semi-Supervised Node Classification
abstract
The class imbalance problem, as an important issue in learning node representations, has drawn increasing attention from the community. Although the imbalance considered by existing studies roots from the unequal quantity of labeled examples in different classes (quantity imbalance), we argue that graph data expose a unique source of imbalance from the asymmetric topological properties of the labeled nodes, i.e., labeled nodes are not equal in terms of their structural role in the graph (topology imbalance). In this work, we first probe the previously unknown topology-imbalance issue, including its characteristics, causes, and threats to semisupervised node classification learning. We then provide a unified view to jointly analyzing the quantity- and topology- imbalance issues by considering the node influence shift phenomenon with the Label Propagation algorithm. In light of our analysis, we devise an influence conflict detection–based metric Totoro to measure the degree of graph topology imbalance and propose a model-agnostic method ReNode to address the topology-imbalance issue by re-weighting the influence of labeled nodes adaptively based on their relative positions to class boundaries. Systematic experiments demonstrate the effectiveness and generalizability of our method in relieving topology-imbalance issue and promoting semi-supervised node classification. The further analysis unveils varied sensitivity of different graph neural networks (GNNs) to topology imbalance, which may serve as a new perspective in evaluating GNN architectures.
Deli Chen, Yankai Lin 0001, Guangxiang Zhao, Xuancheng Ren, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001
NeurIPS7
2021 Auto-Encoding Knowledge Graph for Unsupervised Medical Report Generation
abstract
Medical report generation, which aims to automatically generate a long and coherent report of a given medical image, has been receiving growing research interests. Existing approaches mainly adopt a supervised manner and heavily rely on coupled image-report pairs. However, in the medical domain, building a large-scale image-report paired dataset is both time-consuming and expensive. To relax the dependency on paired data, we propose an unsupervised model Knowledge Graph Auto-Encoder (KGAE) which accepts independent sets of images and reports in training. KGAE consists of a pre-constructed knowledge graph, a knowledge-driven encoder and a knowledge-driven decoder. The knowledge graph works as the shared latent space to bridge the visual and textual domains; The knowledge-driven encoder projects medical images and reports to the corresponding coordinates in this latent space and the knowledge-driven decoder generates a medical report given a coordinate in this space. Since the knowledge-driven encoder and decoder can be trained with independent sets of images and reports, KGAE is unsupervised. The experiments show that the unsupervised KGAE generates desirable medical reports without using any image-report training pairs. Moreover, KGAE can also work in both semi-supervised and supervised settings, and accept paired images and reports in training. By further fine-tuning with image-report pairs, KGAE consistently outperforms the current state-of-the-art models on two datasets.
Chenyu You, Xian Wu 0001, Shen Ge, Sheng Wang 0012, Xu Sun 0001
NeurIPS6
2021 Adversarial parameter defense by multi-step risk minimization
Zhiyuan Zhang 0001, Ruixuan Luo, Xuancheng Ren, Qi Su 0001, Liangyou Li, Xu Sun 0001
Neural Networks6
2020 Measuring and Relieving the Over-Smoothing Problem for Graph Neural Networks from the Topological View
abstract
Graph Neural Networks (GNNs) have achieved promising performance on a wide range of graph-based tasks. Despite their success, one severe limitation of GNNs is the over-smoothing issue (indistinguishable representations of nodes in different classes). In this work, we present a systematic and quantitative study on the over-smoothing issue of GNNs. First, we introduce two quantitative metrics, MAD and MADGap, to measure the smoothness and over-smoothness of the graph nodes representations, respectively. Then, we verify that smoothing is the nature of GNNs and the critical factor leading to over-smoothness is the low information-to-noise ratio of the message received by the nodes, which is partially determined by the graph topology. Finally, we propose two methods to alleviate the over-smoothing issue from the topological view: (1) MADReg which adds a MADGap-based regularizer to the training objective; (2) AdaEdge which optimizes the graph topology based on the model predictions. Extensive experiments on 7 widely-used graph datasets with 10 typical GNN models show that the two proposed methods are effective for relieving the over-smoothing issue, thus improving the performance of various GNN models.
Deli Chen, Yankai Lin 0001, Wei Li 0101, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001
AAAI6
2020 Visual Agreement Regularized Training for Multi-Modal Machine Translation
abstract
Multi-modal machine translation aims at translating the source sentence into a different language in the presence of the paired image. Previous work suggests that additional visual information only provides dispensable help to translation, which is needed in several very special cases such as translating ambiguous words. To make better use of visual information, this work presents visual agreement regularized training. The proposed approach jointly trains the source-to-target and target-to-source translation models and encourages them to share the same focus on the visual information when generating semantically equivalent visual words (e.g. “ball” in English and “ballon” in French). Besides, a simple yet effective multi-head co-attention model is also introduced to capture interactions between visual and textual features. The results show that our approaches can outperform competitive baselines by a large margin on the Multi30k dataset. Further analysis demonstrates that the proposed regularized training can effectively improve the agreement of attention on the image, leading to better use of visual information.
Boxing Chen, Pei Zhang 0011, Xu Sun 0001
AAAI4
2020 How to Ask Good Questions? Try to Leverage Paraphrases
abstract
Given a sentence and its relevant answer, how to ask good questions is a challenging task, which has many real applications.Inspired by human's paraphrasing capability to ask questions of the same meaning but with diverse expressions, we propose to incorporate paraphrase knowledge into question generation(QG) to generate human-like questions.Specifically, we present a two-hand hybrid model leveraging a self-built paraphrase resource, which is automatically conducted by a simple back-translation method.On the one hand, we conduct multi-task learning with sentence-level paraphrase generation (PG) as an auxiliary task to supplement paraphrase knowledge to the task-share encoder.On the other hand, we adopt a new loss function for diversity training to introduce more question patterns to QG. Extensive experimental results show that our proposed model obtains obvious performance gain over several strong baselines, and further human evaluation validates that our model can ask questions of high quality by leveraging paraphrase knowledge.
Xu Sun 0001, Yunfang Wu
ACL3
2020 Parallel Data Augmentation for Formality Style Transfer
abstract
The main barrier to progress in the task of Formality Style Transfer is the inadequacy of training data.In this paper, we study how to augment parallel data and propose novel and simple data augmentation methods for this task to obtain useful sentence pairs with easily accessible models and systems.Experiments demonstrate that our augmented parallel data largely helps improve formality style transfer when it is used to pre-train the model, leading to the state-of-the-art results in the GYAFC benchmark dataset 1 .
Yi Zhang 0050, Tao Ge 0001, Xu Sun 0001
ACL3
2020 Rethinking Skip Connection with Layer Normalization
abstract
Skip connection is a widely-used technique to improve the performance and the convergence of deep neural networks, which is believed to relieve the difficulty in optimization due to non-linearity by propagating a linear component through the neural network layers. However, from another point of view, it can also be seen as a modulating mechanism between the input and the output, with the input scaled by a pre-defined value one. In this work, we investigate how the scale factors in the effectiveness of the skip connection and reveal that a trivial adjustment of the scale will lead to spurious gradient exploding or vanishing in line with the deepness of the models, which could by addressed by normalization, in particular, layer normalization, which induces consistent improvements over the plain skip connection. Inspired by the findings, we further propose to adaptively adjust the scale of the input by recursively applying skip connection with layer normalization, which promotes the performance substantially and generalizes well across diverse tasks including both machine translation and image classification datasets.
Xuancheng Ren, Zhiyuan Zhang 0001, Xu Sun 0001, Yuexian Zou
COLING4
2020 Regularizing Dialogue Generation by Imitating Implicit Scenarios
abstract
Human dialogues are scenario-based and appropriate responses generally relate to the latent context knowledge entailed by the specific scenario.To enable responses that are more meaningful and context-specific, we propose to improve generative dialogue systems from the scenario perspective, where both dialogue history and future conversation are taken into account to implicitly reconstruct the scenario knowledge.More importantly, the conversation scenarios are further internalized using imitation learning framework, where the conventional dialogue model that has no access to future conversations is effectively regularized by transferring the scenario knowledge contained in hierarchical supervising signals from the scenario-based dialogue model, so that the future conversation is not required in actual inference.Extensive evaluations show that our approach significantly outperforms state-of-theart baselines on diversity and relevance, and expresses scenario-specific knowledge.
Shaoxiong Feng, Xuancheng Ren, Hongshen Chen, Bin Sun 0004, Kan Li 0001, Xu Sun 0001
EMNLP (1)6
2020 Prophet Attention: Predicting Attention with Future Attention
abstract
Recently, attention based models have been used extensively in many sequence-to-sequence learning systems. Especially for image captioning, the attention based models are expected to ground correct image regions with proper generated words. However, for each time step in the decoding process, the attention based models usually use the hidden state of the current input to attend to the image regions. Under this setting, these attention models have a deviated focus'' problem that they calculate the attention weights based on previous words instead of the one to be generated, impairing the performance of both grounding and captioning. In this paper, we propose the Prophet Attention, similar to the form of self-supervision. In the training stage, this module utilizes the future information to calculate theideal'' attention weights towards image regions. These calculated ideal'' weights are further used to regularize thedeviated'' attention. In this manner, image regions are grounded with the correct words. The proposed Prophet Attention can be easily incorporated into existing image captioning models to improve their performance of both grounding and captioning. The experiments on the Flickr30k Entities and the MSCOCO datasets show that the proposed Prophet Attention consistently outperforms baselines in both automatic metrics and human evaluations. It is worth noticing that we set new state-of-the-arts on the two benchmark datasets and achieve the 1st place on the leaderboard of the online MSCOCO benchmark in terms of the default ranking score, i.e., CIDEr-c40.
Xuancheng Ren, Xian Wu 0001, Shen Ge, Wei Fan 0001, Yuexian Zou, Xu Sun 0001
NeurIPS7
2020 Re-evaluation of Atomic Operations and Graph Coloring for Unstructured Finite Volume GPU Simulations
abstract
In general, race condition can be resolved by introducing synchronisations or breaking data dependencies. Atomic operations and graph coloring are the two typical approaches to avoid race condition. Graph coloring algorithms have been generally considered winning algorithms in the literature due to their lock free implementations. In this paper, we present the GPU-accelerated algorithms of the unstructured cell-centered finite volume Computational Fluid Dynamics (CFD) software framework named PHengLEI which was originally developed for aerodynamics applications with arbitrary hybrid meshes. Overall, the newly developed GPU framework demonstrate up to 4.8 speedup comparing with 18 MPI tasks run on the latest Intel CPU node. Furthermore, the enormous efforts have been invested to optimize data dependencies which could lead to race condition due to unstructured mesh indirect addressing and related reduction math operations. With careful comparison between our optimised graph coloring and atomic operations using a series of numerical tests with different mesh sizes, the results show that atomic operations are more efficient than our optimised graph coloring in all of the test cases on Nvidia Tesla GPU V100. Specifically, for the summation operation, using atomicAdd is twice as fast as graph coloring. For the maximum operation, a speedup of 1.5 to 2 is found for atomicMax vs. graph coloring.
Xu Sun 0001, Xiaohu Guo, Yunfei Du 0001, Yutong Lu, Yang Liu 0005
SBAC-PAD2
2020 Memorized sparse backpropagation
Zhiyuan Zhang 0001, Xuancheng Ren, Qi Su 0001, Xu Sun 0001
Neurocomputing5
2020 Training Simplification and Model Simplification for Deep Learning : A Minimal Effort Back Propagation Method
abstract
We propose a simple yet effective technique to simplify the training and the resulting model of neural networks. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-k elements (in terms of magnitude) are kept. As a result, only k rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction in the computational cost. Based on the sparsified gradients, we further simplify the model by eliminating the rows or columns that are seldom updated, which will reduce the computational cost both in the training and decoding, and potentially accelerate decoding in real-world applications. Surprisingly, experimental results demonstrate that most of the time we only need to update fewer than 5 percent of the weights at each back propagation pass. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given. The model simplification results show that we could adaptively simplify the model which could often be reduced by around 9x, without any loss on accuracy or even with improved accuracy.
Xu Sun 0001, Xuancheng Ren, Shuming Ma, Bingzhen Wei, Wei Li 0101, Jingjing Xu 0001, Houfeng Wang, Yi Zhang 0050
IEEE Trans. Knowl. Data Eng.1
2019 Learning Personalized End-to-End Goal-Oriented Dialog
abstract
Most existing works on dialog systems only consider conversation content while neglecting the personality of the user the bot is interacting with, which begets several unsolved issues. In this paper, we present a personalized end-to-end model in an attempt to leverage personalization in goal-oriented dialogs. We first introduce a PROFILE MODEL which encodes user profiles into distributed embeddings and refers to conversation history from other similar users. Then a PREFERENCE MODEL captures user preferences over knowledge base entities to handle the ambiguity in user requests. The two models are combined into the PERSONALIZED MEMN2N. Experiments show that the proposed model achieves qualitative performance improvements over state-of-the-art methods. As for human evaluation, it also outperforms other approaches in terms of task completion rate and user satisfaction.
Liangchen Luo, Wenhao Huang 0001, Qi Zeng 0001, Zaiqing Nie, Xu Sun 0001
AAAI5
2019 LiveBot: Generating Live Video Comments Based on Visual and Textual Contexts
abstract
We introduce the task of automatic live commenting. Live commenting, which is also called “video barrage”, is an emerging feature on online video sites that allows real-time comments from viewers to fly across the screen like bullets or roll at the right side of the screen. The live comments are a mixture of opinions for the video and the chit chats with other comments. Automatic live commenting requires AI agents to comprehend the videos and interact with human viewers who also make the comments, so it is a good testbed of an AI agent’s ability to deal with both dynamic vision and language. In this work, we construct a large-scale live comment dataset with 2,361 videos and 895,929 live comments. Then, we introduce two neural models to generate live comments based on the visual and textual contexts, which achieve better performance than previous neural baselines such as the sequence-to-sequence model. Finally, we provide a retrieval-based evaluation protocol for automatic live commenting where the model is asked to sort a set of candidate comments based on the log-likelihood score, and evaluated on metrics such as mean-reciprocal-rank. Putting it all together, we demonstrate the first “LiveBot”. The datasets and the codes can be found at https://github.com/lancopku/livebot.
Shuming Ma, Lei Cui 0001, Damai Dai, Furu Wei, Xu Sun 0001
AAAI5
2019 Coherent Comments Generation for Chinese Articles with a Graph-to-Sequence Model
abstract
Automatic article commenting is helpful in encouraging user engagement and interaction on online news platforms.However, the news documents are usually too long for traditional encoder-decoder based models, which often results in general and irrelevant comments.In this paper, we propose to generate comments with a graph-to-sequence model that models the input news as a topic interaction graph.By organizing the article into graph structure, our model can better understand the internal structure of the article and the connection between topics, which makes it better able to understand the story.We collect and release a large scale news-comment corpus from a popular Chinese online news platform Tencent Kuaibao. 1 Extensive experiment results show that our model can generate much more coherent and informative comments compared with several strong baseline models.2
Wei Li 0101, Jingjing Xu 0001, Yancheng He, Shengli Yan, Yunfang Wu, Xu Sun 0001
ACL (1)6
2019 Learning to Control the Fine-grained Sentiment for Story Ending Generation
abstract
Automatic story ending generation is an interesting and challenging task in natural language generation.Previous studies are mainly limited to generate coherent, reasonable and diversified story endings, and few works focus on controlling the sentiment of story endings.This paper focuses on generating a story ending which meets the given fine-grained sentiment intensity.There are two major challenges to this task.First is the lack of story corpus which has fine-grained sentiment labels.Second is the difficulty of explicitly controlling sentiment intensity when generating endings.Therefore, we propose a generic and novel framework which consists of a sentiment analyzer and a sentimental generator, respectively addressing the two challenges.The sentiment analyzer adopts a series of methods to acquire sentiment intensities of the story dataset.The sentimental generator introduces the sentiment intensity into decoder via a Gaussian Kernel Layer to control the sentiment of the output.To the best of our knowledge, this is the first endeavor to control the fine-grained sentiment for story ending generation without manually annotating sentiment labels.Experiments show that our proposed framework can generate story endings which are not only more coherent and fluent but also able to meet the given sentiment intensity better. 1
Fuli Luo, Damai Dai, Tianyu Liu 0001, Baobao Chang, Zhifang Sui, Xu Sun 0001
ACL (1)7
2019 Towards Fine-grained Text Sentiment Transfer
abstract
In this paper, we focus on the task of finegrained text sentiment transfer (FGST).This task aims to revise an input sequence to satisfy a given sentiment intensity, while preserving the original semantic content.Different from conventional sentiment transfer task that only reverses the sentiment polarity (positive/negative) of text, the FTST task requires more nuanced and fine-grained control of sentiment.To remedy this, we propose a novel Seq2SentiSeq model.Specifically, the numeric sentiment intensity value is incorporated into the decoder via a Gaussian kernel layer to finely control the sentiment intensity of the output.Moreover, to tackle the problem of lacking parallel data, we propose a cycle reinforcement learning algorithm to guide the model training.In this framework, the elaborately designed rewards can balance both sentiment transformation and content preservation, while not requiring any ground truth output.Experimental results show that our approach can outperform existing methods by a large margin in both automatic evaluation and human evaluation.Our code and data, including outputs of all baselines and our model are available at https://github.com/luofuli/ Fine-grained-Sentiment-Transfer. 1
Fuli Luo, Peng Li 0030, Jie Zhou 0016, Yutong Tan, Baobao Chang, Zhifang Sui, Xu Sun 0001
ACL (1)8
2019 Key Fact as Pivot: A Two-Stage Model for Low Resource Table-to-Text Generation
abstract
Table -to-text generation aims to translate the structured data into the unstructured text.Most existing methods adopt the encoder-decoder framework to learn the transformation, which requires large-scale training samples.However, the lack of large parallel data is a major practical problem for many domains.In this work, we consider the scenario of low resource table-to-text generation, where only limited parallel data is available.We propose a novel model to separate the generation into two stages: key fact prediction and surface realization.It first predicts the key facts from the tables, and then generates the text with the key facts.The training of key fact prediction needs much fewer annotated data, while surface realization can be trained with pseudo parallel corpus.We evaluate our model on a biography generation dataset.Our model can achieve 27.34 BLEU score with only 1, 000 parallel data, while the baseline model only obtain the performance of 9.71 BLEU score. 1
Shuming Ma, Tianyu Liu 0001, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001
ACL (1)6
2019 Imitation Learning for Non-Autoregressive Neural Machine Translation
abstract
Non-autoregressive translation models (NAT) have achieved impressive inference speedup.A potential issue of the existing NAT algorithms, however, is that the decoding is conducted in parallel, without directly considering previous context.In this paper, we propose an imitation learning framework for nonautoregressive machine translation, which still enjoys the fast translation speed but gives comparable translation performance compared to its auto-regressive counterpart.We conduct experiments on the IWSLT16, WMT14 and WMT16 datasets.Our proposed model achieves a significant speedup over the autoregressive models, while keeping the translation quality comparable to the autoregressive models.By sampling sentence length in parallel at inference time, we achieve the performance of 31.85BLEU on WMT16 Ro→En and 30.68 BLEU on IWSLT16 En→De.
Bingzhen Wei, Mingxuan Wang, Hao Zhou 0012, Junyang Lin, Xu Sun 0001
ACL (1)5
2019 A Hierarchical Reinforced Sequence Operation Method for Unsupervised Text Style Transfer
abstract
Unsupervised text style transfer aims to alter text styles while preserving the content, without aligned data for supervision.Existing seq2seq methods face three challenges: 1) the transfer is weakly interpretable, 2) generated outputs struggle in content preservation, and 3) the trade-off between content and style is intractable.To address these challenges, we propose a hierarchical reinforced sequence operation method, named Point-Then-Operate (PTO), which consists of a high-level agent that proposes operation positions and a lowlevel agent that alters the sentence.We provide comprehensive training objectives to control the fluency, style, and content of the outputs and a mask-based inference algorithm that allows for multi-step revision based on the single-step trained agents.Experimental results on two text style transfer datasets show that our method significantly outperforms recent methods and effectively addresses the aforementioned challenges. 1
Chen Wu 0005, Xuancheng Ren, Fuli Luo, Xu Sun 0001
ACL (1)4
2019 MAAM: A Morphology-Aware Alignment Model for Unsupervised Bilingual Lexicon Induction
abstract
The task of unsupervised bilingual lexicon induction (UBLI) aims to induce word translations from monolingual corpora in two languages.Previous work has shown that morphological variation is an intractable challenge for the UBLI task, where the induced translation in failure case is usually morphologically related to the correct translation.To tackle this challenge, we propose a morphology-aware alignment model for the UBLI task.The proposed model aims to alleviate the adverse effect of morphological variation by introducing grammatical information learned by the pre-trained denoising language model.Results show that our approach can substantially outperform several state-of-the-art unsupervised systems, and even achieves competitive performance compared to supervised methods.
Fuli Luo, Tianyu Liu 0001, Xu Sun 0001
ACL (1)5
2019 Enhancing Topic-to-Essay Generation with External Commonsense Knowledge
abstract
Automatic topic-to-essay generation is a challenging task since it requires generating novel, diverse, and topic-consistent paragraph-level text with a set of topics as input.Previous work tends to perform essay generation based solely on the given topics while ignoring massive commonsense knowledge.However, this commonsense knowledge provides additional background information, which can help to generate essays that are more novel and diverse.Towards filling this gap, we propose to integrate commonsense from the external knowledge base into the generator through dynamic memory mechanism.Besides, the adversarial training based on a multi-label discriminator is employed to further improve topic-consistency.We also develop a series of automatic evaluation metrics to comprehensively assess the quality of the generated essay.Experiments show that with external commonsense knowledge and adversarial training, the generated essays are more novel, diverse, and topic-consistent than existing methods in terms of both automatic and human evaluation.
Lei Li 0039, Fuli Luo, Tianyu Liu 0001, Xu Sun 0001
ACL (1)5
2019 A Deep Reinforced Sequence-to-Set Model for Multi-Label Classification
abstract
Multi-label classification (MLC) aims to predict a set of labels for a given instance.Based on a pre-defined label order, the sequence-tosequence (Seq2Seq) model trained via maximum likelihood estimation method has been successfully applied to the MLC task and shows powerful ability to capture high-order correlations between labels.However, the output labels are essentially an unordered set rather than an ordered sequence.This inconsistency tends to result in some intractable problems, e.g., sensitivity to the label order.To remedy this, we propose a simple but effective sequence-to-set model.The proposed model is trained via reinforcement learning, where reward feedback is designed to be independent of the label order.In this way, we can reduce the dependence of the model on the label order, as well as capture high-order correlations between labels.Extensive experiments show that our approach can substantially outperform competitive baselines, as well as effectively reduce the sensitivity to the label order. 1
Fuli Luo, Shuming Ma, Junyang Lin, Xu Sun 0001
ACL (1)5
2019 Cross-Modal Commentator: Automatic Machine Commenting Based on Cross-Modal Information
abstract
Automatic commenting of online articles can provide additional opinions and facts to the reader, which improves user experience and engagement on social media platforms.Previous work focuses on automatic commenting based solely on textual content.However, in real-scenarios, online articles usually contain multiple modal contents.For instance, graphic news contains plenty of images in addition to text.Contents other than text are also vital because they are not only more attractive to the reader but also may provide critical information.To remedy this, we propose a new task: cross-model automatic commenting (CMAC), which aims to make comments by integrating multiple modal contents.We construct a largescale dataset for this task and explore several representative methods.Going a step further, an effective co-attention model is presented to capture the dependency between textual and visual information.Evaluation results show that our proposed model can achieve better performance than competitive baselines.1
Zhihan Zhang 0001, Fuli Luo, Lei Li 0039, Chengyang Huang, Xu Sun 0001
ACL (1)6
2019 Pun-GAN: Generative Adversarial Network for Pun Generation
abstract
Fuli Luo, Shunyao Li, Pengcheng Yang, Lei Li, Baobao Chang, Zhifang Sui, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Fuli Luo, Shunyao Li, Lei Li 0039, Baobao Chang, Zhifang Sui, Xu Sun 0001
EMNLP/IJCNLP (1)7
2019 Asking Clarification Questions in Knowledge-Based Question Answering
abstract
Jingjing Xu, Yuechen Wang, Duyu Tang, Nan Duan, Pengcheng Yang, Qi Zeng, Ming Zhou, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jingjing Xu 0001, Yuechen Wang, Duyu Tang, Nan Duan 0001, Qi Zeng 0001, Ming Zhou 0001, Xu Sun 0001
EMNLP/IJCNLP (1)8
2019 LexicalAT: Lexical-Based Adversarial Reinforcement Training for Robust Sentiment Classification
abstract
Jingjing Xu, Liang Zhao, Hanqi Yan, Qi Zeng, Yun Liang, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jingjing Xu 0001, Hanqi Yan, Qi Zeng 0001, Yun Liang 0001, Xu Sun 0001
EMNLP/IJCNLP (1)6
2019 Specificity-Driven Cascading Approach for Unsupervised Sentiment Modification
abstract
Pengcheng Yang, Junyang Lin, Jingjing Xu, Jun Xie, Qi Su, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Junyang Lin, Jingjing Xu 0001, Qi Su 0001, Xu Sun 0001
EMNLP/IJCNLP (1)6
2019 Aligning Cross-Lingual Entities with Multi-Aspect Information
abstract
Hsiu-Wei Yang, Yanyan Zou, Peng Shi, Wei Lu, Jimmy Lin, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Hsiu-Wei Yang, Peng Shi 0010, Wei Lu 0011, Jimmy Lin, Xu Sun 0001
EMNLP/IJCNLP (1)6
2019 Adaptive Gradient Methods with Dynamic Bound of Learning Rate
Liangchen Luo, Yuanhao Xiong, Xu Sun 0001
ICLR (Poster)4
2019 Exploring and Distilling Cross-Modal Information for Image Captioning
abstract
Recently, attention-based encoder-decoder models have been used extensively in image captioning. Yet there is still great difficulty for the current methods to achieve deep image understanding. In this work, we argue that such understanding requires visual attention to correlated image regions and semantic attention to coherent attributes of interest. To perform effective attention, we explore image captioning from a cross-modal perspective and propose the Global-and-Local Information Exploring-and-Distilling approach that explores and distills the source information in vision and language. It globally provides the aspect vector, a spatial and relational representation of images based on caption contexts, through the extraction of salient region groupings and attribute collocations, and locally extracts the fine-grained regions and attributes in reference to the aspect vector for word selection. Our fully-attentive model achieves a CIDEr score of 129.3 in offline COCO evaluation with remarkable efficiency in terms of accuracy, speed, and parameter budget.
Xuancheng Ren, Yuanxin Liu, Kai Lei, Xu Sun 0001
IJCAI5
2019 A Dual Reinforcement Learning Framework for Unsupervised Text Style Transfer
abstract
Unsupervised text style transfer aims to transfer the underlying style of text but keep its main content unchanged without parallel data. Most existing methods typically follow two steps: first separating the content from the original style, and then fusing the content with the desired style. However, the separation in the first step is challenging because the content and style interact in subtle ways in natural language. Therefore, in this paper, we propose a dual reinforcement learning framework to directly transfer the style of the text via a one-step mapping model, without any separation of content and style. Specifically, we consider the learning of the source-to-target and target-to-source mappings as a dual task, and two rewards are designed based on such a dual structure to reflect the style accuracy and content preservation, respectively. In this way, the two one-step mapping models can be trained via reinforcement learning, without any use of parallel data. Automatic evaluations show that our model outperforms the state-of-the-art systems by a large margin, especially with more than 10 BLEU points improvement averaged on two benchmark datasets. Human evaluations also validate the effectiveness of our model in terms of style accuracy, content preservation and fluency. Our code and data, including outputs of all baselines and our model are available at https://github.com/luofuli/DualRL.
Fuli Luo, Peng Li 0030, Jie Zhou 0016, Baobao Chang, Xu Sun 0001, Zhifang Sui
IJCAI6
2019 Knowledgeable Storyteller: A Commonsense-Driven Generative Model for Visual Storytelling
abstract
The visual storytelling (VST) task aims at generating a reasonable and coherent paragraph-level story with the image stream as input. Different from caption that is a direct and literal description of image content, the story in the VST task tends to contain plenty of imaginary concepts that do not appear in the image. This requires the AI agent to reason and associate with the imaginary concepts based on implicit commonsense knowledge to generate a reasonable story describing the image stream. Therefore, in this work, we present a commonsense-driven generative model, which aims to introduce crucial commonsense from the external knowledge base for visual storytelling. Our approach first extracts a set of candidate knowledge graphs from the knowledge base. Then, an elaborately designed vision-aware directional encoding schema is adopted to effectively integrate the most informative commonsense. Besides, we strive to maximize the semantic similarity within the output during decoding to enhance the coherence of the generated text. Results show that our approach can outperform the state-of-the-art systems by a large margin, which achieves a 29\% relative improvement of CIDEr score. With additional commonsense and semantic-relevance based objective, the generated stories are more diverse and coherent.
Fuli Luo, Lei Li 0039, Zhiyi Yin, Xiaodong He 0001, Xu Sun 0001
IJCAI7
2019 Aligning Visual Regions and Textual Concepts for Semantic-Grounded Image Representations
abstract
In vision-and-language grounding problems, fine-grained representations of the image are considered to be of paramount importance. Most of the current systems incorporate visual features and textual concepts as a sketch of an image. However, plainly inferred representations are usually undesirable in that they are composed of separate components, the relations of which are elusive. In this work, we aim at representing an image with a set of integrated visual regions and corresponding textual concepts, reflecting certain semantics. To this end, we build the Mutual Iterative Attention (MIA) module, which integrates correlated visual features and textual concepts, respectively, by aligning the two modalities. We evaluate the proposed approach on two representative vision-and-language grounding tasks, i.e., image captioning and visual question answering. In both tasks, the semantic-grounded image representations consistently boost the performance of the baseline models under all metrics across the board. The results demonstrate that our approach is effective and generalizes well to a wide range of models for image-related applications. (The code is available at \url{https://github.com/fenglinliu98/MIA)
Yuanxin Liu, Xuancheng Ren, Xiaodong He 0001, Xu Sun 0001
NeurIPS5
2019 Understanding and Improving Layer Normalization
abstract
Layer normalization (LayerNorm) is a technique to normalize the distributions of intermediate layers. It enables smoother gradients, faster training, and better generalization accuracy. However, it is still unclear where the effectiveness stems from. In this paper, our main contribution is to take a step further in understanding LayerNorm. Many of previous studies believe that the success of LayerNorm comes from forward normalization. Unlike them, we find that the derivatives of the mean and variance are more important than forward normalization by re-centering and re-scaling backward gradients. Furthermore, we find that the parameters of LayerNorm, including the bias and gain, increase the risk of over-fitting and do not work in most cases. Experiments show that a simple version of LayerNorm (LayerNorm-simple) without the bias and gain outperforms LayerNorm on four datasets. It obtains the state-of-the-art performance on En-Vi machine translation. To address the over-fitting problem, we propose a new normalization method, Adaptive Normalization (AdaNorm), by replacing the bias and gain with a new transformation function. Experiments show that AdaNorm demonstrates better results than LayerNorm on seven out of eight datasets.
Jingjing Xu 0001, Xu Sun 0001, Zhiyuan Zhang 0001, Guangxiang Zhao, Junyang Lin
NeurIPS2
2019 Towards easier and faster sequence labeling for natural language processing: A search-based probabilistic online learning framework (SAPO)
Xu Sun 0001, Shuming Ma, Yi Zhang 0050, Xuancheng Ren
Inf. Sci.1
2019 Regularizing Output Distribution of Abstractive Chinese Social Media Text Summarization for Improved Semantic Consistency
abstract
Abstractive text summarization is a highly difficult problem, and the sequence-to-sequence model has shown success in improving the performance on the task. However, the generated summaries are often inconsistent with the source content in semantics. In such cases, when generating summaries, the model selects semantically unrelated words with respect to the source content as the most probable output. The problem can be attributed to heuristically constructed training data, where summaries can be unrelated to the source content, thus containing semantically unrelated words and spurious word correspondence. In this article, we propose a regularization approach for the sequence-to-sequence model and make use of what the model has learned to regularize the learning objective to alleviate the effect of the problem. In addition, we propose a practical human evaluation method to address the problem that the existing automatic evaluation method does not evaluate the semantic consistency with the source content properly. Experimental results demonstrate the effectiveness of the proposed approach, which outperforms almost all the existing models. Especially, the proposed approach improves the semantic consistency by 4% in terms of human evaluation.
Bingzhen Wei, Xuancheng Ren, Yi Zhang 0050, Xiaoyan Cai, Qi Su 0001, Xu Sun 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.6
2018 Modeling Scientific Influence for Research Trending Topic Prediction
abstract
With the growing volume of publications in the Computer Science (CS) discipline, tracking the research evolution and predicting the future research trending topics are of great importance for researchers to keep up with the rapid progress of research. Within a research area, there are many top conferences that publish the latest research results. These conferences mutually influence each other and jointly promote the development of the research area. To predict the trending topics of mutually influenced conferences, we propose a correlated neural influence model, which has the ability to capture the sequential properties of research evolution in each individual conference and discover the dependencies among different conferences simultaneously. The experiments conducted on a scientific dataset including conferences in artificial intelligence and data mining show that our model consistently outperforms the other state-of-the-art methods. We also demonstrate the interpretability and predictability of the proposed model by providing its answers to two questions of concern, i.e., what the next rising trending topics are and for each conference who the most influential peer is.
Chengyao Chen, Zhitao Wang, Wenjie Li 0002, Xu Sun 0001
AAAI4
2018 Duplicate Question Identification by Integrating FrameNet With Neural Networks
abstract
There are two major problems in duplicate question identification, namely lexical gap and essential constituents matching. Previous methods either design various similarity features or learn representations via neural networks, which try to solve the lexical gap but neglect the essential constituents matching. In this paper, we focus on the essential constituents matching problem and use FrameNet-style semantic parsing to tackle it. Two approaches are proposed to integrate FrameNet parsing with neural networks. An ensemble approach combines a traditional model with manually designed features and a neural network model. An embedding approach converts frame parses to embeddings, which are combined with word embeddings at the input of neural networks. Experiments on Quora question pairs dataset demonstrate that the ensemble approach is more effective and outperforms all baselines.
Xiaodong Zhang 0022, Xu Sun 0001, Houfeng Wang
AAAI2
2018 Unpaired Sentiment-to-Sentiment Translation: A Cycled Reinforcement Learning Approach
abstract
Jingjing Xu, Xu Sun, Qi Zeng, Xiaodong Zhang, Xuancheng Ren, Houfeng Wang, Wenjie Li. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Jingjing Xu 0001, Xu Sun 0001, Qi Zeng 0001, Xiaodong Zhang 0022, Xuancheng Ren, Houfeng Wang, Wenjie Li 0002
ACL (1)2
2018 Question Condensing Networks for Answer Selection in Community Question Answering
abstract
Answer selection is an important subtask of community question answering (CQA).In a real-world CQA forum, a question is often represented as two parts: a subject that summarizes the main points of the question, and a body that elaborates on the subject in detail.Previous researches on answer selection usually ignored the difference between these two parts and concatenated them as the question representation.In this paper, we propose the Question Condensing Networks (QCN) to make use of the subject-body relationship of community questions.In this model, the question subject is the primary part of the question representation, and the question body information is aggregated based on similarity and disparity with the question subject.Experimental results show that QCN outperforms all existing models on two CQA datasets.
Wei Wu 0044, Xu Sun 0001, Houfeng Wang
ACL (1)2
2018 Deconvolution-Based Global Decoding for Neural Machine Translation
abstract
A great proportion of sequence-to-sequence (Seq2Seq) models for Neural Machine Translation (NMT) adopt Recurrent Neural Network (RNN) to generate translation word by word following a sequential order. As the studies of linguistics have proved that language is not linear word sequence but sequence of complex structure, translation at each step should be conditioned on the whole target-side context. To tackle the problem, we propose a new NMT model that decodes the sequence with the guidance of its structural prediction of the context of the target sequence. Our model generates translation based on the structural prediction of the target-side context so that the translation can be freed from the bind of sequential order. Experimental results demonstrate that our model is more competitive compared with the state-of-the-art methods, and the analysis reflects that our model is also robust to translating sentences of different lengths and it also reduces repetition with the instruction from the target-side context for decoding.
Junyang Lin, Xu Sun 0001, Xuancheng Ren, Shuming Ma, Jinsong Su, Qi Su 0001
COLING2
2018 A Neural Question Answering Model Based on Semi-Structured Tables
abstract
Most question answering (QA) systems are based on raw text and structured knowledge graph. However, raw text corpora are hard for QA system to understand, and structured knowledge graph needs intensive manual work, while it is relatively easy to obtain semi-structured tables from many sources directly, or build them automatically. In this paper, we build an end-to-end system to answer multiple choice questions with semi-structured tables as its knowledge. Our system answers queries by two steps. First, it finds the most similar tables. Then the system measures the relevance between each question and candidate table cells, and choose the most related cell as the source of answer. The system is evaluated with TabMCQ dataset, and gets a huge improvement compared to the state of the art.
Xiaodong Zhang 0022, Shuming Ma, Xu Sun 0001, Houfeng Wang, Mengxiang Wang
COLING4
2018 SGM: Sequence Generation Model for Multi-label Classification
abstract
Multi-label classification is an important yet challenging task in natural language processing. It is more complex than single-label classification in that the labels tend to be correlated. Existing methods tend to ignore the correlations between labels. Besides, different parts of the text can contribute differently for predicting different labels, which is not considered by existing models. In this paper, we propose to view the multi-label classification task as a sequence generation problem, and apply a sequence generation model with a novel decoder structure to solve it. Extensive experimental results show that our proposed methods outperform previous work by a substantial margin. Further analysis of experimental results demonstrates that the proposed methods not only capture the correlations between labels, but also select the most informative words automatically when predicting different labels.
Xu Sun 0001, Wei Li 0101, Shuming Ma, Wei Wu 0044, Houfeng Wang
COLING2
2018 Does Higher Order LSTM Have Better Accuracy for Segmenting and Labeling Sequence Data?
abstract
Existing neural models usually predict the tag of the current token independent of the neighboring tags. The popular LSTM-CRF model considers the tag dependencies between every two consecutive tags. However, it is hard for existing neural models to take longer distance dependencies between tags into consideration. The scalability is mainly limited by the complex model structures and the cost of dynamic programming during training. In our work, we first design a new model called “high order LSTM” to predict multiple tags for the current token which contains not only the current tag but also the previous several tags. We call the number of tags in one prediction as “order”. Then we propose a new method called Multi-Order BiLSTM (MO-BiLSTM) which combines low order and high order LSTMs together. MO-BiLSTM keeps the scalability to high order models with a pruning technique. We evaluate MO-BiLSTM on all-phrase chunking and NER datasets. Experiment results show that MO-BiLSTM achieves the state-of-the-art result in chunking and highly competitive results in two NER datasets.
Yi Zhang 0050, Xu Sun 0001, Shuming Ma, Yang Yang 0125, Xuancheng Ren
COLING2
2018 Learning When to Concentrate or Divert Attention: Self-Adaptive Attention Temperature for Neural Machine Translation
abstract
Most of the Neural Machine Translation (NMT) models are based on the sequence-tosequence (Seq2Seq) model with an encoderdecoder framework equipped with the attention mechanism.However, the conventional attention mechanism treats the decoding at each time step equally with the same matrix, which is problematic since the softness of the attention for different types of words (e.g.content words and function words) should differ.Therefore, we propose a new model with a mechanism called Self-Adaptive Control of Temperature (SACT) to control the softness of attention by means of an attention temperature.Experimental results on the Chinese-English translation and English-Vietnamese translation demonstrate that our model outperforms the baseline models, and the analysis and the case study show that our model can attend to the most relevant elements in the source-side contexts and generate the translation of high quality.
Junyang Lin, Xu Sun 0001, Xuancheng Ren, Muyu Li, Qi Su 0001
EMNLP2
2018 Semantic-Unit-Based Dilated Convolution for Multi-Label Text Classification
abstract
We propose a novel model for multi-label text classification, which is based on sequenceto-sequence learning.The model generates higher-level semantic unit representations with multi-level dilated convolution as well as a corresponding hybrid attention mechanism that extracts both the information at the word-level and the level of the semantic unit.Our designed dilated convolution effectively reduces dimension and supports an exponential expansion of receptive fields without loss of local information, and the attention-overattention mechanism is able to capture more summary relevant information from the source context.Results of our experiments show that the proposed model has significant advantages over the baseline models on the dataset RCV1-V2 and Ren-CECps, and our analysis demonstrates that our model is competitive to the deterministic hierarchical models and it is more robust to classifying low-frequency labels 1 .
Junyang Lin, Qi Su 0001, Shuming Ma, Xu Sun 0001
EMNLP5
2018 simNet: Stepwise Image-Topic Merging Network for Generating Detailed and Comprehensive Image Captions
abstract
The encode-decoder framework has shown recent success in image captioning.Visual attention, which is good at detailedness, and semantic attention, which is good at comprehensiveness, have been separately proposed to ground the caption on the image.In this paper, we propose the Stepwise Image-Topic Merging Network (simNet) that makes use of the two kinds of attention at the same time.At each time step when generating the caption, the decoder adaptively merges the attentive information in the extracted topics and the image according to the generated context, so that the visual information and the semantic information can be effectively combined.The proposed approach is evaluated on two benchmark datasets and reaches the state-of-the-art performances.1
Xuancheng Ren, Yuanxin Liu, Houfeng Wang, Xu Sun 0001
EMNLP5
2018 An Auto-Encoder Matching Model for Learning Utterance-Level Semantic Dependency in Dialogue Generation
abstract
Generating semantically coherent responses is still a major challenge in dialogue generation.Different from conventional text generation tasks, the mapping between inputs and responses in conversations is more complicated, which highly demands the understanding of utterance-level semantic dependency, a relation between the whole meanings of inputs and outputs.To address this problem, we propose an Auto-Encoder Matching (AEM) model to learn such dependency.The model contains two auto-encoders and one mapping module.The auto-encoders learn the semantic representations of inputs and responses, and the mapping module learns to connect the utterance-level representations.Experimental results from automatic and human evaluations demonstrate that our model is capable of generating responses of high coherence and fluency compared to baseline models. 1
Liangchen Luo, Jingjing Xu 0001, Junyang Lin, Qi Zeng 0001, Xu Sun 0001
EMNLP5
2018 Auto-Dialabel: Labeling Dialogue Data with Unsupervised Learning
abstract
The lack of labeled data is one of the main challenges when building a task-oriented dialogue system.Existing dialogue datasets usually rely on human labeling, which is expensive, limited in size, and in low coverage.In this paper, we instead propose our framework auto-dialabel to automatically cluster the dialogue intents and slots.In this framework, we collect a set of context features, leverage an autoencoder for feature assembly, and adapt a dynamic hierarchical clustering method for intent and slot labeling.Experimental results show that our framework can promote human labeling cost to a great extent, achieve good intent clustering accuracy (84.1%), and provide reasonable and instructive slot labeling results.
Qi Chen 0009, Lei Sha, Sujian Li, Xu Sun 0001, Houfeng Wang
EMNLP5
2018 Diversity-Promoting GAN: A Cross-Entropy Based Generative Adversarial Network for Diversified Text Generation
abstract
Existing text generation methods tend to produce repeated and "boring" expressions. To tackle this problem, we propose a new text generation model, called Diversity-Promoting Generative Adversarial Network (DP-GAN).The proposed model assigns low reward for repeatedly generated text and high reward for "novel" and fluent text, encouraging the generator to produce diverse and informative text.Moreover, we propose a novel languagemodel based discriminator, which can better distinguish novel text from repeated text without the saturation problem compared with existing classifier-based discriminators.The experimental results on review generation and dialogue generation tasks demonstrate that our model can generate substantially more diverse and informative text than existing baselines.1
Jingjing Xu 0001, Xuancheng Ren, Junyang Lin, Xu Sun 0001
EMNLP4
2018 A Skeleton-Based Model for Promoting Coherence Among Sentences in Narrative Story Generation
abstract
Narrative story generation is a challenging problem because it demands the generated sentences with tight semantic connections, which has not been well studied by most existing generative models.To address this problem, we propose a skeleton-based model to promote the coherence of generated stories.Different from traditional models that generate a complete sentence at a stroke, the proposed model first generates the most critical phrases, called skeleton, and then expands the skeleton to a complete and fluent sentence.The skeleton is not manually defined, but learned by a reinforcement learning method.Compared to the state-of-the-art models, our skeleton-based model can generate significantly more coherent text according to human evaluation and automatic evaluation.The G-score is improved by 20.1% in human evaluation. 1
Jingjing Xu 0001, Xuancheng Ren, Yi Zhang 0050, Qi Zeng 0001, Xiaoyan Cai, Xu Sun 0001
EMNLP6
2018 Learning Sentiment Memories for Sentiment Modification without Parallel Data
abstract
The task of sentiment modification requires reversing the sentiment of the input and preserving the sentiment-independent content.However, aligned sentences with the same content but different sentiments are usually unavailable.Due to the lack of such parallel data, it is hard to extract sentiment independent content and reverse the sentiment in an unsupervised way.Previous work usually can not reconcile sentiment transformation and content preservation.In this paper, motivated by the fact the non-emotional context (e.g., "staff") provides strong cues for the occurrence of emotional words (e.g., "friendly"), we propose a novel method that automatically extracts appropriate sentiment information from the learned sentiment memories according to the specific context.Experiments show that our method substantially improves the content preservation degree and achieves the state-of-the-art performance.1
Yi Zhang 0050, Jingjing Xu 0001, Xu Sun 0001
EMNLP4
2018 A Hierarchical End-to-End Model for Jointly Improving Text Summarization and Sentiment Classification
abstract
Text summarization and sentiment classification both aim to capture the main ideas of the text but at different levels. Text summarization is to describe the text within a few sentences, while sentiment classification can be regarded as a special type of summarization which ``summarizes'' the text into a even more abstract fashion, i.e., a sentiment class. Based on this idea, we propose a hierarchical end-to-end model for joint learning of text summarization and sentiment classification, where the sentiment classification label is treated as the further ``summarization'' of the text summarization output. Hence, the sentiment classification layer is put upon the text summarization layer, and a hierarchical structure is derived. Experimental results on Amazon online reviews datasets show that our model achieves better performance than the strong baseline systems on both abstractive summarization and sentiment classification.
Shuming Ma, Xu Sun 0001, Junyang Lin, Xuancheng Ren
IJCAI2
2018 Building an Ellipsis-aware Chinese Dependency Treebank for Web Text
Xuancheng Ren, Xu Sun 0001, Ji Wen, Bingzhen Wei, Weidong Zhan, Zhiyuan Zhang 0001
LREC2
2018 A Chinese Dataset with Negative Full Forms for General Abbreviation Prediction
Yi Zhang 0050, Xu Sun 0001
LREC2
2018 Query and Output: Generating Words by Querying Distributed Word Representations for Paraphrase Generation
abstract
Shuming Ma, Xu Sun, Wei Li, Sujian Li, Wenjie Li, Xuancheng Ren. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Shuming Ma, Xu Sun 0001, Wei Li 0101, Sujian Li, Wenjie Li 0002, Xuancheng Ren
NAACL-HLT2
2018 Accelerating Graph-Based Dependency Parsing with Lock-Free Parallel Perceptron
Shuming Ma, Xu Sun 0001, Yi Zhang 0050, Bingzhen Wei
NLPCC (1)2
2018 Cross-Domain and Semisupervised Named Entity Recognition in Chinese Social Media: A Unified Model
abstract
Named entity recognition (NER) in Chinese social media is an important, but challenging task because Chinese social media language is informal and noisy. Most previous methods on NER focus on in-domain supervised learning, which is limited by scarce annotated data in social media. In this paper, we present that sufficient corpora in formal domains and massive unannotated text can be combined to improve the NER performance in social media. We propose a unified model which can learn from out-of-domain corpora and in-domain unannotated text. The unified model is composed of two parts. One is for cross-domain learning and the other is for semisupervised learning. Cross-domain learning can learn out-of-domain information based on domain similarity. Semisupervised learning can learn in-domain unannotated information by self-training. Experimental results show that our unified model yields a 9.57% improvement over strong baselines and achieves the state-of-the-art performance.
Jingjing Xu 0001, Hangfeng He 0001, Xu Sun 0001, Xuancheng Ren, Sujian Li
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 A Unified Model for Cross-Domain and Semi-Supervised Named Entity Recognition in Chinese Social Media
abstract
Named entity recognition (NER) in Chinese social media is important but difficult because of its informality and strong noise. Previous methods only focus on in-domain supervised learning which is limited by the rare annotated data. However, there are enough corpora in formal domains and massive in-domain unannotated texts which can be used to improve the task. We propose a unified model which can learn from out-of-domain corpora and in-domain unannotated texts. The unified model contains two major functions. One is for cross-domain learning and another for semi-supervised learning. Cross-domain learning function can learn out-of-domain information based on domain similarity. Semi-Supervised learning function can learn in-domain unannotated information by self-training. Both learning functions outperform existing methods for NER in Chinese social media. Finally, our unified model yields nearly 11% absolute improvement over previously published results.
Hangfeng He 0001, Xu Sun 0001
AAAI2
2017 meProp: Sparsified Back Propagation for Accelerated Deep Learning with Reduced Overfitting
abstract
We propose a simple yet effective technique for neural network learning. The forward propagation is computed as usual. In back propagation, only a small subset of the full gradient is computed to update the model parameters. The gradient vectors are sparsified in such a way that only the top-$k$ elements (in terms of magnitude) are kept. As a result, only $k$ rows or columns (depending on the layout) of the weight matrix are modified, leading to a linear reduction ($k$ divided by the vector dimension) in the computational cost. Surprisingly, experimental results demonstrate that we can update only 1–4\% of the weights at each back propagation pass. This does not result in a larger number of training iterations. More interestingly, the accuracy of the resulting models is actually improved rather than degraded, and a detailed analysis is given.
Xu Sun 0001, Xuancheng Ren, Shuming Ma, Houfeng Wang
ICML1
2017 Addressing Domain Adaptation for Chinese Word Segmentation with Global Recurrent Structure
abstract
Boundary features are widely used in traditional Chinese Word Segmentation (CWS) methods as they can utilize unlabeled data to help improve the Out-of-Vocabulary (OOV) word recognition performance. Although various neural network methods for CWS have achieved performance competitive with state-of-the-art systems, these methods, constrained by the domain and size of the training corpus, do not work well in domain adaptation. In this paper, we propose a novel BLSTM-based neural network model which incorporates a global recurrent structure designed for modeling boundary features dynamically. Experiments show that the proposed structure can effectively boost the performance of Chinese Word Segmentation, especially OOV-Recall, which brings benefits to domain adaptation. We achieved state-of-the-art results on 6 domains of CNKI articles, and competitive results to the best reported on the 4 domains of SIGHAN Bakeoff 2010 data.
Shen Huang, Xu Sun 0001, Houfeng Wang
IJCNLP(1)2
2017 Cascading Multiway Attentions for Document-level Sentiment Classification
abstract
Document-level sentiment classification aims to assign the user reviews a sentiment polarity. Previous methods either just utilized the document content without consideration of user and product information, or did not comprehensively consider what roles the three kinds of information play in text modeling. In this paper, to reasonably use all the information, we present the idea that user, product and their combination can all influence the generation of attentions to words and sentences, when judging the sentiment of a document. With this idea, we propose a cascading multiway attention (CMA) model, where multiple ways of using user and product information are cascaded to influence the generation of attentions on the word and sentence layers. Then, sentences and documents are well modeled by multiple representation vectors, which provide rich information for sentiment classification. Experiments on IMDB and Yelp datasets demonstrate the effectiveness of our model.
Dehong Ma, Sujian Li, Xiaodong Zhang 0022, Houfeng Wang, Xu Sun 0001
IJCNLP(1)5
2017 Tag-Enhanced Tree-Structured Neural Networks for Implicit Discourse Relation Classification
abstract
Identifying implicit discourse relations between text spans is a challenging task because it requires understanding the meaning of the text. To tackle this task, recent studies have tried several deep learning methods but few of them exploited the syntactic information. In this work, we explore the idea of incorporating syntactic parse tree into neural networks. Specifically, we employ the Tree-LSTM model and Tree-GRU model, which is based on the tree structure, to encode the arguments in a relation. And we further leverage the constituent tags to control the semantic composition process in these tree-structured neural networks. Experimental results show that our method achieves state-of-the-art performance on PDTB corpus.
Yizhong Wang, Sujian Li, Jingfeng Yang 0001, Xu Sun 0001, Houfeng Wang
IJCNLP(1)4
2017 Transfer Deep Learning for Low-Resource Chinese Word Segmentation with a Novel Neural Network
Jingjing Xu 0001, Shuming Ma, Yi Zhang 0050, Bingzhen Wei, Xiaoyan Cai, Xu Sun 0001
NLPCC6
2016 Knowledge-Based Semantic Embedding for Machine Translation
abstract
In this paper, with the help of knowledge base, we build and formulate a semantic space to connect the source and target languages, and apply it to the sequence-to-sequence framework to propose a Knowledge-Based Semantic Embedding (KBSE) method.In our KB-SE method, the source sentence is firstly mapped into a knowledge based semantic space, and the target sentence is generated using a recurrent neural network with the internal meaning preserved.Experiments are conducted on two translation tasks, the electric business data and movie data, and the results show that our proposed method can achieve outstanding performance, compared with both the traditional SMT methods and the existing encoder-decoder models.
Shujie Liu 0001, Shuo Ren 0002, Mu Li 0001, Ming Zhou 0001, Xu Sun 0001, Houfeng Wang
ACL (1)7
2016 Asynchronous Parallel Learning for Neural Networks and Structured Models with Dense Features
abstract
Existing asynchronous parallel learning methods are only for the sparse feature models, and they face new challenges for the dense feature models like neural networks (e.g., LSTM, RNN). The problem for dense features is that asynchronous parallel learning brings gradient errors derived from overwrite actions. We show that gradient errors are very common and inevitable. Nevertheless, our theoretical analysis shows that the learning process with gradient errors can still be convergent towards the optimum of objective functions for many practical applications. Thus, we propose a simple method AsynGrad for asynchronous parallel learning with gradient error. Base on various dense feature models (LSTM, dense-CRF) and various NLP tasks, experiments show that AsynGrad achieves substantial improvement on training speed, and without any loss on accuracy.
Xu Sun 0001
COLING1
2015 Multi-label Text Categorization with Joint Learning Predictions-as-Features Method
abstract
Multi-label text categorization is a type of text categorization, where each document is assigned to one or more categories.Recently, a series of methods have been developed, which train a classifier for each label, organize the classifiers in a partially ordered structure and take predictions produced by the former classifiers as the latter classifiers' features.These predictions-asfeatures style methods model high order label dependencies and obtain high performance.Nevertheless, the predictionsas-features methods suffer a drawback.When training a classifier for one label, the predictions-as-features methods can model dependencies between former labels and the current label, but they can't model dependencies between the current label and the latter labels.To address this problem, we propose a novel joint learning algorithm that allows the feedbacks to be propagated from the classifiers for latter labels to the classifier for the current label.We conduct experiments using real-world textual data sets, and these experiments illustrate the predictions-as-features models trained by our algorithm outperform the original models.
Houfeng Wang, Xu Sun 0001, Baobao Chang, Shi Zhao, Lei Sha
EMNLP3
2014 Coarse-grained Candidate Generation and Fine-grained Re-ranking for Chinese Abbreviation Prediction
abstract
Correctly predicting abbreviations given the full forms is important in many natu-ral language processing systems. In this paper we propose a two-stage method to find the corresponding abbreviation given its full form. We first use the contextual information given a large corpus to get ab-breviation candidates for each full form and get a coarse-grained ranking through graph random walk. This coarse-grained rank list fixes the search space inside the top-ranked candidates. Then we use a sim-ilarity sensitive re-ranking strategy which can utilize the features of the candidates to give a fine-grained re-ranking and se-lect the final result. Our method achieves good results and outperforms the state-of-the-art systems. One advantage of our method is that it only needs weak super-vision and can get competitive results with fewer training data. The candidate genera-tion and coarse-grained ranking is totally unsupervised. The re-ranking phase can use a very small amount of training data to get a reasonably good result. 1
Longkai Zhang, Houfeng Wang, Xu Sun 0001
EMNLP3
2014 Predicting Chinese Abbreviations with Minimum Semantic Unit and Global Constraints
abstract
We propose a new Chinese abbreviation prediction method which can incorporate rich local information while generating the abbreviation globally.Different to previous character tagging methods, we introduce the minimum semantic unit, which is more fine-grained than character but more coarse-grained than word, to capture word level information in the sequence labeling framework.To solve the "character duplication" problem in Chinese abbreviation prediction, we also use a substring tagging strategy to generate local substring tagging candidates.We use an integer linear programming (ILP) formulation with various constraints to globally decode the final abbreviation from the generated candidates.Experiments show that our method outperforms the state-of-the-art systems, without using any extra resource.
Longkai Zhang, Houfeng Wang, Xu Sun 0001
EMNLP4
2014 Structure Regularization for Structured Prediction
Xu Sun 0001
NIPS1
2014 Feature-Frequency-Adaptive On-line Training for Fast and Accurate Natural Language Processing
abstract
Training speed and accuracy are two major concerns of large-scale natural language processing systems. Typically, we need to make a tradeoff between speed and accuracy. It is trivial to improve the training speed via sacrificing accuracy or to improve the accuracy via sacrificing speed. Nevertheless, it is nontrivial to improve the training speed and the accuracy at the same time, which is the target of this work. To reach this target, we present a new training method, feature-frequency–adaptive on-line training, for fast and accurate training of natural language processing systems. It is based on the core idea that higher frequency features should have a learning rate that decays faster. Theoretical analysis shows that the proposed method is convergent with a fast convergence rate. Experiments are conducted based on well-known benchmark tasks, including named entity recognition, word segmentation, phrase chunking, and sentiment analysis. These tasks consist of three structured classification tasks and one non-structured classification task, with binary features and real-valued features, respectively. Experimental results demonstrate that the proposed method is faster and at the same time more accurate than existing methods, achieving state-of-the-art scores on the tasks with different characteristics.
Xu Sun 0001, Wenjie Li 0002, Houfeng Wang, Qin Lu 0001
Comput. Linguistics1
2013 A unified graph model for personalized query-oriented reference paper recommendation
abstract
With the tremendous amount of research publications, it has become increasingly important to provide a researcher with a rapid and accurate recommendation of a list of reference papers about a research field or topic. In this paper, we propose a unified graph model that can easily incorporate various types of useful information (e.g., content, authorship, citation and collaboration networks etc.) for efficient recommendation. The proposed model not only allows to thoroughly explore how these types of information can be better combined, but also makes personalized query-oriented reference paper recommendation possible, which as far as we know is a new issue that has not been explicitly addressed in the past. The experiments have demonstrated the clear advantages of personalized recommendation over non-personalized recommendation.
Fanqi Meng, Dehong Gao, Wenjie Li 0002, Xu Sun 0001, Yuexian Hou
CIKM4
2013 Exploring Representations from Unlabeled Data with Co-training for Chinese Word Segmentation
abstract
Nowadays supervised sequence labeling models can reach competitive performance on the task of Chinese word segmentation.However, the ability of these models is restricted by the availability of annotated data and the design of features.We propose a scalable semi-supervised feature engineering approach.In contrast to previous works using pre-defined taskspecific features with fixed values, we dynamically extract representations of label distributions from both an in-domain corpus and an out-of-domain corpus.We update the representation values with a semi-supervised approach.Experiments on the benchmark datasets show that our approach achieve good results and reach an f-score of 0.961.The feature engineering approach proposed here is a general iterative semi-supervised method and not limited to the word segmentation task.
Longkai Zhang, Houfeng Wang, Xu Sun 0001, Mairgup Mansur
EMNLP3
2013 Generalized Abbreviation Prediction with Negative Full Forms and Its Application on Improving Chinese Web Search
Xu Sun 0001, Wenjie Li 0002, Fanqi Meng, Houfeng Wang
IJCNLP1
2013 Probabilistic Chinese word segmentation with non-local information and stochastic training
Xu Sun 0001, Takuya Matsuzaki, Yoshimasa Tsuruoka, Jun'ichi Tsujii
Inf. Process. Manag.1
2013 Learning Abbreviations from Chinese and English Terms by Modeling Non-Local Information
abstract
The present article describes a robust approach for abbreviating terms. First, in order to incorporate non-local information into abbreviation generation tasks, we present both implicit and explicit solutions: the latent variable model and the label encoding with global information. Although the two approaches compete with one another, we find they are also highly complementary. We propose a combination of the two approaches, and we will show the proposed method outperforms all of the existing methods on abbreviation generation datasets. In order to reduce computational complexity of learning non-local information, we further present an online training method, which can arrive the objective optimum with accelerated training speed. We used a Chinese newswire dataset and a English biomedical dataset for experiments. Experiments revealed that the proposed abbreviation generator with non-local information achieved the best results for both the Chinese and English languages.
Xu Sun 0001, Naoaki Okazaki, Jun'ichi Tsujii, Houfeng Wang
ACM Trans. Asian Lang. Inf. Process.1
2013 Large-Scale Personalized Human Activity Recognition Using Online Multitask Learning
abstract
Personalized activity recognition usually has the problem of highly biased activity patterns among different tasks/persons. Traditional methods face problems on dealing with those conflicted activity patterns. We try to effectively model the activity patterns among different persons via casting this personalized activity recognition problem as a multitask learning issue. We propose a novel online multitask learning method for large-scale personalized activity recognition. In contrast with existing work of multitask learning that assumes fixed task relationships, our method can automatically discover task relationships from real-world data. Convergence analysis shows reasonable convergence properties of the proposed method. Experiments on two different activity data sets demonstrate that the proposed method significantly outperforms existing methods in activity recognition.
Xu Sun 0001, Hisashi Kashima, Naonori Ueda
IEEE Trans. Knowl. Data Eng.1
2013 Latent Structured Perceptrons for Large-Scale Learning with Hidden Information
abstract
Many real-world data mining problems contain hidden information (e.g., unobservable latent dependencies). We propose a perceptron-style method, latent structured perceptron, for fast discriminative learning of structured classification with hidden information. We also give theoretical analysis and demonstrate good convergence properties of the proposed method. Our method extends the perceptron algorithm for the learning task with hidden information, which can be hardly captured by traditional models. It relies on Viterbi decoding over latent variables, combined with simple additive updates. We perform experiments on one synthetic data set and two real-world structured classification tasks. Compared to conventional nonlatent models (e.g., conditional random fields, structured perceptrons), our method is more accurate on real-world tasks. Compared to existing heavy probabilistic models of latent variables (e.g., latent conditional random fields), our method lowers the training cost significantly (almost one order magnitude faster) yet with comparable or even superior classification accuracy. In addition, experiments demonstrate that the proposed method has good scalability on large-scale problems.
Xu Sun 0001, Takuya Matsuzaki, Wenjie Li 0002
IEEE Trans. Knowl. Data Eng.1
2012 Fast Online Training with Frequency-Adaptive Learning Rates for Chinese Word Segmentation and New Word Detection
Xu Sun 0001, Houfeng Wang, Wenjie Li 0002
ACL (1)1
2012 Fast multi-task learning for query spelling correction
abstract
In this paper, we explore the use of a novel online multi-task learning framework for the task of search query spelling correction. In our procedure, correction candidates are initially generated by a ranker-based system and then re-ranked by our multi-task learning algorithm. With the proposed multi-task learning method, we are able to effectively transfer information from different and highly biased training datasets, for improving spelling correction on all datasets. Our experiments are conducted on three query spelling correction datasets including the well-known TREC benchmark dataset. The experimental results demonstrate that our proposed method considerably outperforms the existing baseline systems in terms of accuracy. Importantly, the proposed method is about one order of magnitude faster than baseline systems in terms of training speed. Compared to the commonly used online learning methods which typically require more than (e.g.,) 60 training passes, our proposed method is able to closely reach the empirical optimum in about 5 passes.
Xu Sun 0001, Anshumali Shrivastava, Ping Li 0001
CIKM1
2011 A New Multi-task Learning Method for Personalized Activity Recognition
abstract
Personalized activity recognition usually faces the problem of data sparseness. We aim at improving accuracy of personalized activity recognition by incorporating the information from other persons. We propose a new online multi-task learning method for personalized activity recognition. The proposed online multi-task learning method automatically learns the ``transfer-factors" (similarities) among different tasks (i.e., among different persons in our case). Experiments demonstrate that the proposed method significantly outperforms existing methods. The novelty of this paper is twofold: (1) A new multi-task learning framework, which can naturally learn similarities among tasks, (2) To our knowledge, this is the first study of large-scale personalized activity recognition.
Xu Sun 0001, Hisashi Kashima, Ryota Tomioka, Naonori Ueda, Ping Li 0001
ICDM1
2011 Large Scale Real-Life Action Recognition Using Conditional Random Fields with Stochastic Training
Xu Sun 0001, Hisashi Kashima, Ryota Tomioka, Naonori Ueda
PAKDD (2)1
2010 Learning Phrase-Based Spelling Error Models from Clickthrough Data
Xu Sun 0001, Jianfeng Gao 0001, Daniel Micol, Chris Quirk
ACL1
2010 A Large Scale Ranker-Based System for Search Query Spelling Correction
Jianfeng Gao 0001, Daniel Micol, Chris Quirk, Xu Sun 0001
COLING5
2010 Averaged Stochastic Gradient Descent with Feedback: An Accurate, Robust, and Fast Training Method
abstract
On large datasets, the popular training approach has been stochastic gradient descent (SGD). This paper proposes a modification of SGD, called averaged SGD with feedback (ASF), that significantly improves the performance (robustness, accuracy, and training speed) over the traditional SGD. The proposal is based on three simple ideas: averaging the weight vectors across SGD iterations, feeding the averaged weights back into the SGD update process, and deciding when to perform the feedback (linearly slowing down feedback). Theoretically, we demonstrate the reasonable convergence properties of the ASF. Empirically, the ASF outperforms several strong baselines in terms of accuracy, robustness over the noise, and the training speed. To our knowledge, this is the first study of ``feedback'' in stochastic gradient learning. Although we choose latent conditional models for verifying the ASF in this paper, the ASF is a general purpose technique just like SGD, and can be directly applied to other models.
Xu Sun 0001, Hisashi Kashima, Takuya Matsuzaki, Naonori Ueda
ICDM1
2009 Robust Approach to Abbreviating Terms: A Discriminative Latent Variable Model with Global Information
Xu Sun 0001, Naoaki Okazaki, Jun'ichi Tsujii
ACL/IJCNLP1
2009 Sequential Labeling with Latent Variables: An Exact Inference Algorithm and its Efficient Approximation
Xu Sun 0001, Jun'ichi Tsujii
EACL1
2009 Latent Variable Perceptron Algorithm for Structured Classification
Xu Sun 0001, Takuya Matsuzaki, Daisuke Okanohara, Jun'ichi Tsujii
IJCAI1
2009 A Discriminative Latent Variable Chinese Segmenter with Hybrid Word/Character Information
Xu Sun 0001, Takuya Matsuzaki, Yoshimasa Tsuruoka, Jun'ichi Tsujii
HLT-NAACL1
2008 Modeling Latent-Dynamic in Shallow Parsing: A Latent Conditional Model with Imrpoved Inference
Xu Sun 0001, Louis-Philippe Morency, Daisuke Okanohara, Yoshimasa Tsuruoka, Jun'ichi Tsujii
COLING1
2008 Predicting Chinese Abbreviations from Definitions: An Empirical Learning Approach Using Support Vector Regression
Xu Sun 0001, Houfeng Wang, Bo Wang 0003
J. Comput. Sci. Technol.1
2007 Word Clustering for Collocation-Based Word Sense Disambiguation
Xu Sun 0001, Yunfang Wu, Shiwen Yu
CICLing2
2006 Chinese Abbreviation-Definition Identification: A SVM Approach Using Context Information
Xu Sun 0001, Houfeng Wang
PRICAI1