EDBT 2026 Demo / reviewers in the wild / expert
Xinyan Xiao
dblp:87/8177
· DBLP profile ↗
48ranked-venue papers
5as first author
30since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 5 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Passive Critical Thinking: Fostering Proactive Questioning to Enhance Human-AI CollaborationabstractCritical thinking is essential for building robust AI systems, preventing them from blindly accepting flawed data or biased reasoning. However, prior work has primarily focused on passive critical thinking, where models simply reject problematic queries without taking constructive steps to address user requests. In this work, we introduce proactive critical thinking, a paradigm where models actively seek missing or clarifying information from users to resolve their queries better. To evaluate this capability, we present GSM-MC and GSM-MCE, two novel benchmarks based on GSM8K for assessing mathematical reasoning under incomplete or misleading conditions. Experiments on Qwen3 and Llama series models show that, while these models excel in traditional reasoning tasks, they struggle with proactive critical thinking, especially smaller ones. However, we demonstrate that reinforcement learning (RL) can significantly improve this ability. By incorporating heuristic information into the reward function, we achieve substantial gains, boosting the Qwen3-1.7B's accuracy from 0.15% to 73.98% on GSM-MC. We hope this work advances models that collaborate more effectively with users in problem-solving through proactive critical thinking. Ante Wang, Yujie Lin 0003, Suhang Wu, Xinyan Xiao, Jinsong Su |
AAAI | 6 |
| 2026 | HD-Custom: Efficient Hierarchical Disentanglement for Coarse-to-Fine Concept Customization in Subject Video Generation
Yuanhang Li, Qi Mao 0002, Xinyan Xiao, Libiao Jin, Siwei Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | BiDeV: Bilateral Defusing Verification for Complex Claim Fact-CheckingabstractComplex claim fact-checking performs a crucial role in disinformation detection. However, existing fact-checking methods struggle with claim vagueness, specifically in effectively handling latent information and complex relations within claims. Moreover, evidence redundancy, where non-essential information complicates the verification process, remains a significant issue. To tackle these limitations, we propose Bilateral Defusing Verification (BiDeV), a novel fact-checking working-flow framework integrating multiple role-played LLMs to mimic the human-expert fact-checking process. BiDeV consists of two main modules: Vagueness Defusing identifies latent information and resolves complex relations to simplify the claim, and Redundancy Defusing eliminates redundant content to enhance the evidence quality. Extensive experimental results on two widely used challenging fact-checking benchmarks (Hover and Feverous-s) demonstrate that our BiDeV can achieve the best performance under both gold and open settings. This highlights the effectiveness of BiDeV in handling complex claims and ensuring precise fact-checking. Yuxuan Liu 0009, Hongda Sun 0001, Wenya Guo, Xinyan Xiao, Cunli Mao, Zhengtao Yu 0001, Rui Yan 0001 |
AAAI | 4 |
| 2025 | SEGA: A Stepwise Evolution Paradigm for Content-Aware Layout Generation with Design PriorabstractIn this paper, we study the content-aware layout generation problem, which aims to automatically generate layouts that are harmonious with a given background image. Existing methods usually deal with this task with a single-step reasoning framework. The lack of a feedback-based self-correction mechanism leads to their failure rates significantly increasing when faced with complex element layout planning. To address this challenge, we introduce SEGA, a novel Stepwise Evolution Paradigm for Content-Aware Layout Generation. Inspired by the systematic mode of human thinking, SEGA employs a hierarchical reasoning framework with a coarse-to-fine strategy: first, a coarse-level module roughly estimates the layout planning results; then, another refining module performs fine-level reasoning regarding the coarse planning results. Furthermore, we incorporate layout design principles as prior knowledge into the model to enhance its layout planning ability. Besides, we present GenPoster-100K that is a new large-scale poster dataset with rich meta-information annotation. The experiments demonstrate the effectiveness of our approach by achieving the state-of-the-art results on multiple benchmark datasets. Our project page is at: https://brucew91.github.io/SEGA.github.io/ Bo Zhao 0039, Huan Yang 0005, Xinyan Xiao |
ICCV | 8 |
| 2025 | UniVG: Towards UNIfied-modal Video GenerationabstractDiffusion based video generation has received significant attention in both the academic and industrial communities. Despite recent exploration of diverse conditional inputs for better video generation control, existing methods, primarily targeting individual tasks, often fall short in real-world scenarios where users may use any form of conditioning, either individually or combined. To address this, we propose a Unified-modal Video Generation system capable of handling multiple video generation tasks across different modalities. Our approach introduces the concept of generative freedom in the diffusion process, which allows us to reclassify video generation tasks into high-freedom and low-freedom categories based on the solution space given certain conditions. We then design different diffusion paradigms for each category. For high-freedom video generation, we present a base model that is capable of handling varied semantic combinations of text and image. For low-freedom video generation, we propose the Biased Gaussian Noise (BGN) to tackle the discrepancy of the diffusion process between the training and inference stage when using strong conditional guidance strategy. Our proposed UniVG achieves superior objective results on public datasets, surpassing the current open-source methods and is on par with the current close-source method Gen2 and Pika in human evaluations. For more samples, visit our homepage. Ludan Ruan, Chuanwei Huang, Xinyan Xiao |
ICME | 5 |
| 2025 | A Token is Worth over 1, 000 Tokens: Efficient Knowledge Distillation through Low-Rank CloneabstractTraining high-performing Small Language Models (SLMs) remains computationally expensive, even with knowledge distillation and pruning from larger teacher models.
Existing approaches often face three key challenges: (1) information loss from hard pruning, (2) inefficient alignment of representations, and (3) underutilization of informative activations, particularly from Feed-Forward Networks (FFNs).
To address these challenges, we introduce \textbf{Low-Rank Clone (LRC)}, an efficient pre-training method that constructs SLMs aspiring to behavioral equivalence with strong teacher models.
LRC trains a set of low-rank projection matrices that jointly enable soft pruning by compressing teacher weights, and activation clone by aligning student activations, including FFN signals, with those of the teacher.
This unified design maximizes knowledge transfer while removing the need for explicit alignment modules.
Extensive experiments with open-source teachers such as Llama-3.2-3B-Instruct and Qwen2.5-3B/7B-Instruct show that LRC matches or surpasses the performance of state-of-the-art models trained on trillions of tokens--using only 20B tokens, achieving over \textbf{1,000$\times$} greater training efficiency.
Our codes and model checkpoints are available at https://github.com/CURRENTF/LowRankClone and https://huggingface.co/JitaiHao/LRC-4B-Base. Jitai Hao, Xinyan Xiao, Zhaochun Ren, Jun Yu 0002 |
NeurIPS | 4 |
| 2025 | Noise-Robust Vision-Language Pre-Training With Positive-Negative LearningabstractVision-Language Pre-training (VLP) has shown promising performance in various tasks by learning a generic image-text representation space. However, most existing VLP methods encounter the Noisy Correspondence (NC) problem which refers to wrongly matched image-text pairs harvested from the wild. In this paper, we empirically study the influence of NC on the VLP model and obtain the following two observations. First, the NC will largely degrade the performance in downstream tasks even via fine-tuning, indicating the necessity of handling NC in the pre-training period. Second, the influence of NC varies in different pre-training objectives, suggesting the objective-customized solution for achieving NC robustness. Based on the above observations, we propose a novel NoisE-robust Vision-languagE pRe-training method (NEVER) to endow the VLP model with robustness against NC. In brief, NEVER first divides the training data into clean and noisy subsets in a progressive and adaptive manner. Then NEVER employs the positive learning (PL) and negative learning (NL) on the splits to enjoy model convergence and noise robustness, respectively. To further handle the false negative in PL and NL, NEVER proposes to smoothen and sharpen the training targets with the predictions from a twin momentum model. Extensive experiments on the various V+L tasks verify the effectiveness of the proposed method. Zhenyu Huang 0005, Mouxing Yang, Xinyan Xiao, Peng Hu 0002, Xi Peng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | UNIMO-G: Unified Image Generation through Multimodal Conditional DiffusionabstractExisting text-to-image diffusion models primarily generate images from text prompts. However, the inherent conciseness of textual descriptions poses challenges in faithfully synthesizing images with intricate details, such as specific entities or scenes. This paper presents UNIMO-G, a simple multimodal conditional diffusion framework that operates on multimodal prompts with interleaved textual and visual inputs, which demonstrates a unified ability for both text-driven and subject-driven image generation. UNIMO-G comprises two core components: a Multimodal Large Language Model (MLLM) for encoding multimodal prompts, and a conditional denoising diffusion network for generating images based on the encoded multimodal input. We leverage a two-stage training strategy to effectively train the framework: firstly pre-training on large-scale text-image pairs to develop conditional image generation capabilities, and then instruction tuning with multimodal prompts to achieve unified image generation proficiency. A well-designed data processing pipeline involving language grounding and image segmentation is employed to construct multi-modal prompts. UNIMO-G excels in both text-to-image generation and zero-shot subject-driven synthesis, and is notably effective in generating high-fidelity images from complex multimodal prompts involving multiple image entities. Xinyan Xiao |
ACL (1) | 4 |
| 2024 | Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware TrainingabstractDiffusion-based text-to-image models have demonstrated impressive achievements in diversity and aesthetics but struggle to generate images with legible visual texts. Existing backbone models have limitations such as misspelling, failing to generate texts, and lack of support for Chinese texts, but their development shows promising potential. In this paper, we propose a series of methods, aiming to empower backbone models to generate visual texts in English and Chinese. We first conduct a preliminary study revealing that BPE tokenization and insufficient learning of cross-attention modules restrict the performance of the backbone models. Based on these observations, we make the following improvements: (1) We design a mixed granularity input strategy to provide more suitable text representations; (2) We propose to augment the conventional training objective with three glyph-aware training losses, which enhance the learning of cross-attention modules and encourage the model to focus on visual texts. Through experiments, we demonstrate that our methods can effectively empower backbone models to generate semantic relevant, aesthetically appealing, and accurate visual text images, while maintaining their fundamental image generation quality. Guohao Li 0002, Zhibin Lan, Wanru Zhuang, Xinyan Xiao, Jinsong Su |
EMNLP | 7 |
| 2024 | Test-Time Degradation Adaptation for Open-Set Image RestorationabstractIn contrast to close-set scenarios that restore images from a predefined set of degradations, open-set image restoration aims to handle the unknown degradations that were unforeseen during the pretraining phase, which is less-touched as far as we know. This work study this challenging problem and reveal its essence as unidentified distribution shifts between the test and training data. Recently, test-time adaptation has emerged as a fundamental method to address this inherent disparities. Inspired by it, we propose a test-time degradation adaptation framework for open-set image restoration, which consists of three components, i.e., i) a pre-trained and degradation-agnostic diffusion model for generating clean images, ii) a test-time degradation adapter adapts the unknown degradations based on the input image during the testing phase, and iii) the adapter-guided image restoration guides the model through the adapter to produce the corresponding clean image. Through experiments on multiple degradations, we show that our method achieves comparable even better performance than those task-specific methods. The code is available at https://github.com/XLearning-SCU/2024-ICML-TAO. Yuanbiao Gou, Haiyu Zhao, Boyun Li, Xinyan Xiao, Xi Peng 0001 |
ICML | 4 |
| 2024 | AverNet: All-in-one Video Restoration for Time-varying Unknown DegradationsabstractTraditional video restoration approaches were designed to recover clean videos from a specific type of degradation, making them ineffective in handling multiple unknown types of degradation. To address this issue, several studies have been conducted and have shown promising results. However, these studies overlook that the degradations in video usually change over time, dubbed time-varying unknown degradations (TUD). To tackle such a less-touched challenge, we propose an innovative method, termed as All-in-one VidEo Restoration Network (AverNet), which comprises two core modules, i.e., Prompt-Guided Alignment (PGA) module and Prompt-Conditioned Enhancement (PCE) module. Specifically, PGA addresses the issue of pixel shifts caused by time-varying degradations by learning and utilizing prompts to align video frames at the pixel level. To handle multiple unknown degradations, PCE recasts it into a conditional restoration problem by implicitly establishing a conditional map between degradations and ground truths. Thanks to the collaboration between PGA and PCE modules, AverNet empirically demonstrates its effectiveness in recovering videos from TUD. Extensive experiments are carried out on two synthesized datasets featuring seven types of degradations with random corruption levels. The code is available at https://github.com/XLearning-SCU/2024-NeurIPS-AverNet. Haiyu Zhao, Xinyan Xiao, Peng Hu 0002, Yuanbiao Gou, Xi Peng 0001 |
NeurIPS | 3 |
| 2024 | Learning with Noisy Correspondence
Zhenyu Huang 0005, Peng Hu 0002, Guocheng Niu, Xinyan Xiao, Jiancheng Lv 0001, Xi Peng 0001 |
Int. J. Comput. Vis. | 4 |
| 2023 | Knowledge-Constrained Answer Generation for Open-Ended Video Question AnsweringabstractOpen-ended Video question answering (open-ended VideoQA) aims to understand video content and question semantics to generate the correct answers. Most of the best performing models define the problem as a discriminative task of multi-label classification. In real-world scenarios, however, it is difficult to define a candidate set that includes all possible answers. In this paper, we propose a Knowledge-constrained Generative VideoQA Algorithm (KcGA) with an encoder-decoder pipeline, which enables out-of-domain answer generation through an adaptive external knowledge module and a multi-stream information control mechanism. We use ClipBERT to extract the video-question features, extract framewise object-level external knowledge from a commonsense knowledge base and compute the contextual-aware episode memory units via an attention based GRU to form the external knowledge features, and exploit multi-stream information control mechanism to fuse video-question and external knowledge features such that the semantic complementation and alignment are well achieved. We evaluate our model on two open-ended benchmark datasets to demonstrate that we can effectively and robustly generate high-quality answers without restrictions of training data. Guocheng Niu, Xinyan Xiao, Jian Zhang 0026, Xi Peng 0001, Jun Yu 0002 |
AAAI | 3 |
| 2023 | WeCheck: Strong Factual Consistency Checker via Weakly Supervised LearningabstractA crucial issue of current text generation models is that they often uncontrollably generate text that is factually inconsistent with inputs.Due to lack of annotated data, existing factual consistency metrics usually train evaluation models on synthetic texts or directly transfer from other related tasks, such as question answering (QA) and natural language inference (NLI).Bias in synthetic text or upstream tasks makes them perform poorly on text actually generated by language models, especially for general evaluation for various tasks.To alleviate this problem, we propose a weakly supervised framework named WeCheck that is directly trained on actual generated samples from language models with weakly annotated labels.WeCheck first utilizes a generative model to infer the factual labels of generated samples by aggregating weak labels from multiple resources.Next, we train a simple noise-aware classification model as the target metric using the inferred weakly supervised information.Comprehensive experiments on various tasks demonstrate the strong performance of WeCheck, achieving an average absolute improvement of 3.3% on the TRUE benchmark over 11B state-of-the-art methods using only 435M parameters.Furthermore, it is up to 30× faster than previous evaluation methods, greatly improving the accuracy and efficiency of factual consistency evaluation. 1 Wei Li 0176, Xinyan Xiao, Sujian Li, Yajuan Lyu |
ACL (1) | 3 |
| 2023 | SeSQL: A High-Quality Large-Scale Session-Level Chinese Text-to-SQL Dataset
Saihao Huang, Zhenghua Li, Chenhui Dou, Fukang Yan, Xinyan Xiao, Hua Wu 0003, Min Zhang 0005 |
NLPCC (1) | 7 |
| 2023 | FactGen: Faithful Text Generation by Factuality-aware Pre-training and Contrastive Ranking Fine-tuningabstractConditional text generation is supposed to generate a fluent and coherent target text that is faithful to the source text. Although pre-trained models have achieved promising results, they still suffer from the crucial factuality problem. To deal with this issue, we propose a factuality-aware pretraining-finetuning framework named FactGen, which fully considers factuality during two training stages. Specifically, at the pre-training stage, we utilize a natural language inference model to construct target texts that are entailed by the source texts, resulting in a more factually consistent pre-training objective. Then, during the fine-tuning stage, we further introduce a contrastive ranking loss to encourage the model to generate factually consistent text with higher probability. Extensive experiments on three conditional text generation tasks demonstrate the effectiveness and generality of our training framework. Zhibin Lan, Wei Li 0176, Jinsong Su, Xinyan Xiao, Yajuan Lyu |
J. Artif. Intell. Res. | 4 |
| 2023 | Controllable Dialogue Generation With Disentangled Multi-Grained Style Specification and Attribute Consistency RewardabstractControllable text generation is an appealing but challenging task, which allows users to specify particular attributes of the generated outputs. In this paper, we propose a controllable dialogue generation model to steer response generation under multi-attribute constraints. Specifically, we define and categorize the commonly-used control attributes into global and local ones, which possess different granularities of effects on response generation. Then, we significantly extend the conventional seq2seq framework by introducing a novel two-stage decoder, which first uses amulti-grainedstyle specification layerto impose the stylistic constraints and determine word-level control states of responses based on the attributes, and then employs aresponse generation layerto generate final responses maintaining both semantic relevancy to the contexts and fidelity to the attributes. Furthermore, we train our model with an attribute consistency reward to promote response control with explicit supervision signals. Extensive experiments and in-depth analyses on two datasets indicate that our model can significantly outperform competitive baselines in terms of response quality, content diversity and controllability. Hou Pong Chan, Xinyan Xiao, Jinsong Su, Hua Wu 0003 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Unified Structure Generation for Universal Information ExtractionabstractYaojie Lu, Qing Liu, Dai Dai, Xinyan Xiao, Hongyu Lin, Xianpei Han, Le Sun, Hua Wu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yaojie Lu 0001, Dai Dai, Xinyan Xiao, Xianpei Han, Le Sun 0001, Hua Wu 0003 |
ACL (1) | 4 |
| 2022 | PLANET: Dynamic Content Planning in Autoregressive Transformers for Long-form Text GenerationabstractDespite recent progress of pre-trained language models on generating fluent text, existing methods still suffer from incoherence problems in long-form text generation tasks that require proper content control and planning to form a coherent high-level logical flow.In this work, we propose PLANET, a novel generation framework leveraging autoregressive self-attention mechanism to conduct content planning and surface realization dynamically.To guide the generation of output sentences, our framework enriches the Transformer decoder with latent representations to maintain sentence-level semantic plans grounded by bag-of-words.Moreover, we introduce a new coherence-based contrastive learning objective to further improve the coherence of output.Extensive experiments are conducted on two challenging longform text generation tasks including counterargument generation and opinion article generation.Both automatic and human evaluations show that our method significantly outperforms strong baselines and generates more coherent texts with richer contents. Hou Pong Chan, Xinyan Xiao, Hua Wu 0003, Lifu Huang |
ACL (1) | 4 |
| 2022 | A Fine-grained Interpretability Evaluation Benchmark for Neural NLPabstractLijie Wang, Yaozong Shen, Shuyuan Peng, Shuai Zhang, Xinyan Xiao, Hao Liu, Hongxuan Tang, Ying Chen, Hua Wu, Haifeng Wang. Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL). 2022. Yaozong Shen, Shuyuan Peng, Xinyan Xiao, Hao Liu 0026, Hongxuan Tang, Ying Chen 0011, Hua Wu 0003, Haifeng Wang 0001 |
CoNLL | 5 |
| 2022 | Precisely the Point: Adversarial Augmentations for Faithful and Informative Text GenerationabstractThough model robustness has been extensively studied in language understanding, the robustness of Seq2Seq generation remains understudied.In this paper, we conduct the first quantitative analysis on the robustness of pre-trained Seq2Seq models.We find that even current SOTA pre-trained Seq2Seq model (BART) is still vulnerable, which leads to significant degeneration in faithfulness and informativeness for text generation tasks.This motivated us to further propose a novel adversarial augmentation framework, namely AdvSeq, for generally improving faithfulness and informativeness of Seq2Seq models via enhancing their robustness.AdvSeq automatically constructs two types of adversarial augmentations during training, including implicit adversarial samples by perturbing word representations and explicit adversarial samples by word swapping, both of which effectively improve Seq2Seq robustness.Extensive experiments on three popular text generation tasks demonstrate that AdvSeq significantly improves both the faithfulness and informativeness of Seq2Seq generation under both automatic and human evaluation settings. Wei Li 0176, Xinyan Xiao, Sujian Li, Yajuan Lyu |
EMNLP | 4 |
| 2022 | Faster and Better Grammar-Based Text-to-SQL Parsing via Clause-Level Parallel Decoding and Alignment Loss
Kun Wu 0009, Zhenghua Li, Xinyan Xiao |
NLPCC (2) | 4 |
| 2022 | Towards Knowledge-Aware Video Captioning via Transitive Visual Relationship DetectionabstractVideo captioning can be enhanced by incorporating the knowledge, which is usually represented as relationships of objects. However, the previous methods construct only superficial or static object relationships, and often introduce noise into the task through irrelevant common sense or fixed syntax templates. These problems mitigate the model interpretability and lead to the undesirable consequence. To overcome these limitations, we propose to enhance video captioning with deep-level object relationships that are adaptively explored during training. Specifically, we present a Transitive Visual Relationship Detection (TVRD) module in which we estimate the actions of the visual objects, and construct an Object-Action Graph (OAG) to describe the shallow relationship between the objects and actions. Then we bridge the gap between the objects via the actions to transitively infer an Object-Object Graph (OOG) which reflects the deep-level relationship. We further feed the OOG to a graph convolutional network to refine the object representation by deep-level relationships. With the refined representation, we capitalize on an LSTM-based decoder for caption generation. Experimental results on two benchmark datasets: MSVD, MSR-VTT demonstrate that the proposed method achieves state-of-the-art performance. Lastly, we present comprehensive ablation studies as well as visualization of visual relationships to demonstrate the effectiveness and interpretability of our model. Bofeng Wu, Guocheng Niu, Jun Yu 0002, Xinyan Xiao, Jian Zhang 0026, Hua Wu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningabstractWei Li, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu, Jiachen Liu, Hua Wu, Haifeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wei Li 0176, Can Gao, Guocheng Niu, Xinyan Xiao, Hao Liu 0026, Hua Wu 0003, Haifeng Wang 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | BASS: Boosting Abstractive Summarization with Unified Semantic GraphabstractWenhao Wu, Wei Li, Xinyan Xiao, Jiachen Liu, Ziqiang Cao, Sujian Li, Hua Wu, Haifeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Wei Li 0176, Xinyan Xiao, Ziqiang Cao, Sujian Li, Hua Wu 0003, Haifeng Wang 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | SgSum: Transforming Multi-document Summarization into Sub-graph SelectionabstractMost of existing extractive multi-document summarization (MDS) methods score each sentence individually and extract salient sentences one by one to compose a summary, which have two main drawbacks: (1) neglecting both the intra and cross-document relations between sentences; (2) neglecting the coherence and conciseness of the whole summary.In this paper, we propose a novel MDS framework (SgSum) to formulate the MDS task as a sub-graph selection problem, in which source documents are regarded as a relation graph of sentences (e.g., similarity graph or discourse graph) and the candidate summaries are its subgraphs.Instead of selecting salient sentences, SgSum selects a salient sub-graph from the relation graph as the summary.Comparing with traditional methods, our method has two main advantages: (1) the relations between sentences are captured by modeling both the graph structure of the whole document set and the candidate sub-graphs; (2) directly outputs an integrate summary in the form of subgraph which is more informative and coherent.Extensive experiments on MultiNews and DUC datasets show that our proposed method brings substantial improvements over several strong baselines.Human evaluation results also demonstrate that our model can produce significantly more coherent and informative summaries compared with traditional MDS methods.Moreover, the proposed architecture has strong transfer ability from single to multi-document input, which can reduce the resource bottleneck in MDS tasks. 1 Moye Chen, Wei Li 0176, Xinyan Xiao, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP (1) | 4 |
| 2021 | Fine-grained Entity Typing via Label ReasoningabstractConventional entity typing approaches are based on independent classification paradigms, which make them difficult to recognize interdependent, long-tailed and fine-grained entity types.In this paper, we argue that the implicitly entailed extrinsic and intrinsic dependencies between labels can provide critical knowledge to tackle the above challenges.To this end, we propose Label Reasoning Network(LRN), which sequentially reasons finegrained entity labels by discovering and exploiting label dependencies knowledge entailed in the data.Specifically, LRN utilizes an auto-regressive network to conduct deductive reasoning and a bipartite attribute graph to conduct inductive reasoning between labels, which can effectively model, learn and reason complex label dependencies in a sequence-toset, end-to-end manner.Experiments show that LRN achieves the state-of-the-art performance on standard ultra fine-grained entity typing benchmarks, and can also resolve the long tail label problem effectively. Xinyan Xiao, Xianpei Han, Le Sun 0001, Hua Wu 0003 |
EMNLP (1) | 3 |
| 2021 | Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL ParsingabstractData augmentation has attracted a lot of research attention in the deep learning era for its ability in alleviating data sparseness.The lack of labeled data for unseen evaluation databases is exactly the major challenge for cross-domain text-to-SQL parsing.Previous works either require human intervention to guarantee the quality of generated data, or fail to handle complex SQL queries.This paper presents a simple yet effective data augmentation framework.First, given a database, we automatically produce a large number of SQL queries based on an abstract syntax tree grammar.For better distribution matching, we require that at least 80% of SQL patterns in the training data are covered by generated queries.Second, we propose a hierarchical SQL-to-question generation model to obtain high-quality natural language questions, which is the major contribution of this work.Finally, we design a simple sampling strategy that can greatly improve training efficiency given large amounts of generated data.Experiments on three cross-domain datasets, i.e., WikiSQL and Spider in English, and DuSQL in Chinese, show that our proposed data augmentation framework can consistently improve performance over strong baselines, and the hierarchical generation component is the key for the improvement. Kun Wu 0009, Zhenghua Li, Xinyan Xiao, Hua Wu 0003, Min Zhang 0005, Haifeng Wang 0001 |
EMNLP (1) | 5 |
| 2021 | Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge Distillation and Cross-modal MatchingabstractThis paper proposes an approach to Dense Video Captioning (DVC) without pairwise event-sentence annotation. First, we adopt the knowledge distilled from relevant and well solved tasks to generate high-quality event proposals. Then we incorporate contrastive loss and cycle-consistency loss typically applied to cross-modal retrieval tasks to build semantic matching between the proposals and sentences, which are eventually used to train the caption generation module. In addition, the parameters of matching module are initialized via pre-training based on annotated images to improve the matching performance. Extensive experiments on ActivityNet-Caption dataset reveal the significance of distillation-based event proposal generation and cross-modal retrieval-based semantic matching to weakly supervised DVC, and demonstrate the superiority of our method to existing state-of-the-art methods. Bofeng Wu, Guocheng Niu, Jun Yu 0002, Xinyan Xiao, Jian Zhang 0026, Hua Wu 0003 |
IJCAI | 4 |
| 2021 | Learning with Noisy Correspondence for Cross-modal MatchingabstractCross-modal matching, which aims to establish the correspondence between two different modalities, is fundamental to a variety of tasks such as cross-modal retrieval and vision-and-language understanding. Although a huge number of cross-modal matching methods have been proposed and achieved remarkable progress in recent years, almost all of these methods implicitly assume that the multimodal training data are correctly aligned. In practice, however, such an assumption is extremely expensive even impossible to satisfy. Based on this observation, we reveal and study a latent and challenging direction in cross-modal matching, named noisy correspondence, which could be regarded as a new paradigm of noisy labels. Different from the traditional noisy labels which mainly refer to the errors in category labels, our noisy correspondence refers to the mismatch paired samples. To solve this new problem, we propose a novel method for learning with noisy correspondence, named Noisy Correspondence Rectifier (NCR). In brief, NCR divides the data into clean and noisy partitions based on the memorization effect of neural networks and then rectifies the correspondence via an adaptive prediction model in a co-teaching manner. To verify the effectiveness of our method, we conduct experiments by using the image-text matching as a showcase. Extensive experiments on Flickr30K, MS-COCO, and Conceptual Captions verify the effectiveness of our method. The code could be accessed from www.pengxi.me . Zhenyu Huang 0005, Guocheng Niu, Xiao Liu 0040, Wenbiao Ding, Xinyan Xiao, Hua Wu 0003, Xi Peng 0001 |
NeurIPS | 5 |
| 2020 | Leveraging Graph to Improve Abstractive Multi-Document SummarizationabstractGraphs that capture relations between textual units have great benefits for detecting salient information from multiple documents and generating overall coherent summaries.In this paper, we develop a neural abstractive multidocument summarization (MDS) model which can leverage well-known graph representations of documents such as similarity graph and discourse graph, to more effectively process multiple input documents and produce abstractive summaries.Our model utilizes graphs to encode documents in order to capture cross-document relations, which is crucial to summarizing long documents.Our model can also take advantage of graphs to guide the summary generation process, which is beneficial for generating coherent and concise summaries.Furthermore, pre-trained language models can be easily combined with our model, which further improve the summarization performance significantly.Empirical results on the WikiSum and MultiNews dataset show that the proposed architecture brings substantial improvements over several strong baselines. Wei Li 0176, Xinyan Xiao, Hua Wu 0003, Haifeng Wang 0001, Junping Du 0001 |
ACL | 2 |
| 2020 | SKEP: Sentiment Knowledge Enhanced Pre-training for Sentiment AnalysisabstractRecently, sentiment analysis has seen remarkable advance with the help of pre-training approaches. However, sentiment knowledge, such as sentiment words and aspect-sentiment pairs, is ignored in the process of pre-training, despite the fact that they are widely used in traditional sentiment analysis approaches. In this paper, we introduce Sentiment Knowledge Enhanced Pre-training (SKEP) in order to learn a unified sentiment representation for multiple sentiment analysis tasks. With the help of automatically-mined knowledge, SKEP conducts sentiment masking and constructs three sentiment knowledge prediction objectives, so as to embed sentiment information at the word, polarity and aspect level into pre-trained sentiment representation. In particular, the prediction of aspect-sentiment pairs is converted into multi-label classification, aiming to capture the dependency between words in a pair. Experiments on three kinds of sentiment tasks show that SKEP significantly outperforms strong pre-training baseline, and achieves new state-of-the-art results on most of the test datasets. We release our code at https://github.com/baidu/Senta. Hao Tian 0005, Can Gao, Xinyan Xiao, Hao Liu 0026, Bolei He, Hua Wu 0003, Haifeng Wang 0001, Feng Wu 0001 |
ACL | 3 |
| 2020 | Exploring Contextual Word-level Style Relevance for Unsupervised Style TransferabstractUnsupervised style transfer aims to change the style of an input sentence while preserving its original content without using parallel training data. In current dominant approaches, owing to the lack of fine-grained control on the influence from the target style, they are unable to yield desirable output sentences. In this paper, we propose a novel attentional sequence-to-sequence (Seq2seq) model that dynamically exploits the relevance of each output word to the target style for unsupervised style transfer. Specifically, we first pretrain a style classifier, where the relevance of each input word to the original style can be quantified via layer-wise relevance propagation. In a denoising auto-encoding manner, we train an attentional Seq2seq model to reconstruct input sentences and repredict word-level previously-quantified style relevance simultaneously. In this way, this model is endowed with the ability to automatically predict the style relevance of each output word. Then, we equip the decoder of this model with a neural style component to exploit the predicted wordlevel style relevance for better style transfer. Particularly, we fine-tune this model using a carefully-designed objective function involving style transfer, style relevance consistency, content preservation and fluency modeling loss terms. Experimental results show that our proposed model achieves state-of-the-art performance in terms of both transfer accuracy and content preservation. Chulun Zhou, Xinyan Xiao, Jinsong Su, Hua Wu 0003 |
ACL | 4 |
| 2020 | Diversified Multiple Instance Learning for Document-Level Multi-Aspect Sentiment ClassificationabstractNeural Document-level Multi-aspect Sentiment Classification (DMSC) usually requires a lot of manual aspect-level sentiment annotations, which is time-consuming and laborious.As document-level sentiment labeled data are widely available from online service, it is valuable to perform DMSC with such free document-level annotations.To this end, we propose a novel Diversified Multiple Instance Learning Network (D-MILN), which is able to achieve aspect-level sentiment classification with only document-level weak supervision.Specifically, we connect aspect-level and document-level sentiment by formulating this problem as multiple instance learning, providing a way to learn aspect-level classifier from the back propagation of document-level supervision.Two diversified regularizations are further introduced in order to avoid the overfitting on document-level signals during training.Diversified textual regularization encourages the classifier to select aspect-relevant snippets, and diversified sentimental regularization prevents the aspect-level sentiments from being overly consistent with document-level sentiment.Experimental results on TripAdvisor and BeerAdvocate datasets show that D-MILN remarkably outperforms recent weaklysupervised baselines, and is also comparable to the supervised method. Yunjie Ji, Hao Liu 0026, Bolei He, Xinyan Xiao, Hua Wu 0003, Yanhua Yu |
EMNLP (1) | 4 |
| 2019 | Joint Extraction of Entities and Overlapping Relations Using Position-Attentive Sequence LabelingabstractJoint entity and relation extraction is to detect entity and relation using a single model. In this paper, we present a novel unified joint extraction model which directly tags entity and relation labels according to a query word position p, i.e., detecting entity at p, and identifying entities at other positions that have relationship with the former. To this end, we first design a tagging scheme to generate n tag sequences for an n-word sentence. Then a position-attention mechanism is introduced to produce different sentence representations for every query position to model these n tag sequences. In this way, our method can simultaneously extract all entities and their type, as well as all overlapping relations. Experiment results show that our framework performances significantly better on extracting overlapping relations as well as detecting long-range relation, and thus we achieve state-of-the-art performance on two public datasets. Dai Dai, Xinyan Xiao, Yajuan Lyu, Shan Dou, Qiaoqiao She, Haifeng Wang 0001 |
AAAI | 2 |
| 2019 | ARNOR: Attention Regularization based Noise Reduction for Distant Supervision Relation ClassificationabstractDistant supervision is widely used in relation classification in order to create large-scale training data by aligning a knowledge base with an unlabeled corpus. However, it also introduces amounts of noisy labels where a contextual sentence actually does not express the labeled relation. In this paper, we propose ARNOR, a novel Attention Regularization based NOise Reduction framework for distant supervision relation classification. ARNOR assumes that a trustable relation label should be explained by the neural attention model. Specifically, our ARNOR framework iteratively learns an interpretable model and utilizes it to select trustable instances. We first introduce attention regularization to force the model to pay attention to the patterns which explain the relation labels, so as to make the model more interpretable. Then, if the learned model can clearly locate the relation patterns of a candidate instance in the training set, we will select it as a trustable instance for further training step. According to the experiments on NYT data, our ARNOR framework achieves significant improvements over state-of-the-art methods in both relation classification performance and noise reduction effect. Dai Dai, Xinyan Xiao, Hua Wu 0003 |
ACL (1) | 3 |
| 2019 | Enhancing Local Feature Extraction with Global Representation for Neural Text ClassificationabstractGuocheng Niu, Hengru Xu, Bolei He, Xinyan Xiao, Hua Wu, Sheng Gao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Guocheng Niu, Hengru Xu, Bolei He, Xinyan Xiao, Hua Wu 0003 |
EMNLP/IJCNLP (1) | 4 |
| 2018 | Joint Training of Candidate Extraction and Answer Selection for Reading ComprehensionabstractWhile sophisticated neural-based techniques have been developed in reading comprehension, most approaches model the answer in an independent manner, ignoring its relations with other answer candidates.This problem can be even worse in open-domain scenarios, where candidates from multiple passages should be combined to answer a single question.In this paper, we formulate reading comprehension as an extract-then-select twostage procedure.We first extract answer candidates from passages, then select the final answer by combining information from all the candidates.Furthermore, we regard candidate extraction as a latent variable and train the two-stage process jointly with reinforcement learning.As a result, our approach has improved the state-ofthe-art performance significantly on two challenging open-domain reading comprehension datasets.Further analysis demonstrates the effectiveness of our model components, especially the information fusion of all the candidates and the joint training of the extract-then-select procedure. Xinyan Xiao, Yajuan Lyu |
ACL (1) | 3 |
| 2018 | Improving Neural Abstractive Document Summarization with Explicit Information Selection ModelingabstractInformation selection is the most important component in document summarization task.In this paper, we propose to extend the basic neural encoding-decoding framework with an information selection layer to explicitly model and optimize the information selection process in abstractive document summarization.Specifically, our information selection layer consists of two parts: gated global information filtering and local sentence selection.Unnecessary information in the original document is first globally filtered, then salient sentences are selected locally while generating each summary sentence sequentially.To optimize the information selection process directly, distantly-supervised training guided by the golden summary is also imported.Experimental results demonstrate that the explicit modeling and optimizing of the information selection process improves document summarization performance significantly, which enables our model to generate more informative and concise summaries, and thus significantly outperform state-of-the-art neural abstractive methods. Wei Li 0176, Xinyan Xiao, Yajuan Lyu, Yuanzhuo Wang |
EMNLP | 2 |
| 2018 | Improving Neural Abstractive Document Summarization with Structural RegularizationabstractRecent neural sequence-to-sequence models have shown significant progress on short text summarization.However, for document summarization, they fail to capture the longterm structure of both documents and multisentence summaries, resulting in information loss and repetitions.In this paper, we propose to leverage the structural information of both documents and multi-sentence summaries to improve the document summarization performance.Specifically, we import both structural-compression and structuralcoverage regularization into the summarization process in order to capture the information compression and information coverage properties, which are the two most important structural properties of document summarization.Experimental results demonstrate that the structural regularization improves the document summarization performance significantly, which enables our model to generate more informative and concise summaries, and thus significantly outperforms state-of-the-art neural abstractive methods. Wei Li 0176, Xinyan Xiao, Yajuan Lyu, Yuanzhuo Wang |
EMNLP | 2 |
| 2014 | Beam-Width Adaptation for Hierarchical Phrase-Based Translation
Xinyan Xiao, Kaile Su |
CICLing (2) | 3 |
| 2014 | Topic-Based Dissimilarity and Sensitivity Models for Translation Rule SelectionabstractTranslation rule selection is a task of selecting appropriate translation rules for an ambiguous source-language segment. As translation ambiguities are pervasive in statistical machine translation, we introduce two topic-based models for translation rule selection which incorporates global topic information into translation disambiguation. We associate each synchronous translation rule with source- and target-side topic distributions.With these topic distributions, we propose a topic dissimilarity model to select desirable (less dissimilar) rules by imposing penalties for rules with a large value of dissimilarity of their topic distributions to those of given documents. In order to encourage the use of non-topic specific translation rules, we also present a topic sensitivity model to balance translation rule selection between generic rules and topic-specific rules. Furthermore, we project target-side topic distributions onto the source-side topic model space so that we can benefit from topic information of both the source and target language. We integrate the proposed topic dissimilarity and sensitivity model into hierarchical phrase-based machine translation for synchronous translation rule selection. Experiments show that our topic-based translation rule selection model can substantially improve translation quality. Min Zhang 0005, Xinyan Xiao, Deyi Xiong, Qun Liu 0001 |
J. Artif. Intell. Res. | 2 |
| 2013 | Max-Margin Synchronous Grammar Induction for Machine TranslationabstractTraditional synchronous grammar induction estimates parameters by maximizing likelihood, which only has a loose relation to translation quality.Alternatively, we propose a max-margin estimation approach to discriminatively inducing synchronous grammars for machine translation, which directly optimizes translation quality measured by BLEU.In the max-margin estimation of parameters, we only need to calculate Viterbi translations.This further facilitates the incorporation of various non-local features that are defined on the target side.We test the effectiveness of our max-margin estimation framework on a competitive hierarchical phrase-based system.Experiments show that our max-margin method significantly outperforms the traditional twostep pipeline for synchronous rule extraction by 1.3 BLEU points and is also better than previous max-likelihood estimation method. Xinyan Xiao, Deyi Xiong |
EMNLP | 1 |
| 2012 | A Topic Similarity Model for Hierarchical Phrase-based Translation
Xinyan Xiao, Deyi Xiong, Min Zhang 0005, Qun Liu 0001, Shouxun Lin |
ACL (1) | 1 |
| 2012 | Unsupervised Discriminative Induction of Synchronous Grammar for Machine Translation
Xinyan Xiao, Deyi Xiong, Yang Liu 0005, Qun Liu 0001, Shouxun Lin |
COLING | 1 |
| 2011 | Fast Generation of Translation Forest for Large-Scale SMT Discriminative Training
Xinyan Xiao, Yang Liu 0005, Qun Liu 0001, Shouxun Lin |
EMNLP | 1 |
| 2010 | Joint Tokenization and Translation
Xinyan Xiao, Yang Liu 0005, Young-Sook Hwang, Qun Liu 0001, Shouxun Lin |
COLING | 1 |
| 2009 | Weighted Alignment Matrices for Statistical Machine Translation
Yang Liu 0005, Tian Xia 0004, Xinyan Xiao, Qun Liu 0001 |
EMNLP | 3 |