Kun Xu 0005

dblp:29/6948-5 · DBLP profile ↗
← Back
44ranked-venue papers
12as first author
23since 2021 · last 2026
0000-0002-3863-344XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 43 · 12 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author
YearPublicationVenuePosition
2026 Generating Attribute-Aware Human Motions from Textual Prompt
abstract
Text-driven human motion generation has recently attracted considerable attention, allowing models to generate human motions based on textual descriptions. However, current methods neglect the influence of human attributes—such as age, gender, weight, and height—which are key factors shaping human motion patterns. This work represents a pilot exploration for bridging this gap. We conceptualize each motion as comprising both attribute information and action semantics, where textual descriptions align exclusively with action semantics. To achieve this, a new framework inspired by Structural Causal Models is proposed to decouple action semantics from human attributes, enabling text-to-semantics prediction and attribute-controlled generation. The resulting model is capable of generating attribute-aware motion aligned with the user's text and attribute inputs. For evaluation, we introduce a comprehensive dataset containing attribute annotations for text-motion pairs, setting the first benchmark for attribute-aware motion generation. Extensive experiments validate our model's effectiveness.
Xinghan Wang 0002, Kun Xu 0005, Cao Sheng, Jiazhong Yu, Yadong Mu
AAAI2
2025 Granularity-Adaptive Spatial Evidence Tokenization for Video Question Answering
abstract
Video question answering plays a vital role in computer vision, and recent advances in large language models have further propelled the development of this field. However, existing video question answering techniques often face limitations in grasping fine-grained video content in spatial dimensions. It mainly stems from the fixed and low-resolution input of video frames. While some approaches using high-resolution inputs partially alleviate this problem, they introduce excessive computational burdens by encoding the entire high-resolution image. In this work, we propose a granularity-adaptive spatial evidence tokenization model for video question answering. Our method introduces multi-granular visual tokenization in the spatial dimension to produce video tokens at various granularities based on the question. It highlights spatially activated patches at low resolutions through a granularity weighting module and then adaptively encodes these activated patches at high resolution for detail supplementation. To mitigate the computational overhead associated with high-resolution frame encoding, a masking and acceleration module is developed for efficient visual tokenization. Moreover, a granularity compression module is designed to dynamically select and compress visual tokens of varying granularities based on questions. We conduct extensive experiments on 11 mainstream video question answering datasets and the experimental results demonstrate the effectiveness of our proposed method.
Hao Jiang 0032, Zhicheng Sun 0001, Kun Xu 0005, Yang Song 0008, Kun Gai, Yadong Mu
AAAI5
2025 Following the Autoregressive Nature of LLM Embeddings via Compression and Alignment
abstract
Jingcheng Deng, Zhongtao Jiang, Liang Pang, Zihao Wei, Liwei Chen, Kun Xu, Yang Song, Huawei Shen, Xueqi Cheng. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Jingcheng Deng, Zhongtao Jiang, Liang Pang 0001, Zihao Wei, Kun Xu 0005, Yang Song 0008, Huawei Shen, Xueqi Cheng 0001
EMNLP6
2025 Pyramidal Flow Matching for Efficient Video Generative Modeling
abstract
Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational demands, the separate optimization of each sub-stage hinders knowledge sharing and sacrifices flexibility. This work introduces a unified pyramidal flow matching algorithm. It reinterprets the original denoising trajectory as a series of pyramid stages, where only the final stage operates at the full resolution, thereby enabling more efficient video generative modeling. Through our sophisticated design, the flows of different pyramid stages can be interlinked to maintain continuity. Moreover, we craft autoregressive video generation with a temporal pyramid to compress the full-resolution history. The entire framework can be optimized in an end-to-end manner and with a single unified Diffusion Transformer (DiT). Extensive experiments demonstrate that our method supports generating high-quality 5-second (up to 10-second) videos at 768p resolution and 24 FPS within 20.7k A100 GPU training hours. All code and models are open-sourced at https://pyramid-flow.github.io.
Zhicheng Sun 0001, Ningyuan Li 0002, Kun Xu 0005, Hao Jiang 0032, Nan Zhuang, Quzhe Huang, Yang Song 0008, Yadong Mu, Zhouchen Lin
ICLR4
2024 Harder Task Needs More Experts: Dynamic Routing in MoE Models
abstract
Quzhe Huang, Zhenwei An, Nan Zhuang, Mingxu Tao, Chen Zhang, Yang Jin, Kun Xu, Kun Xu, Liwei Chen, Songfang Huang, Yansong Feng. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Quzhe Huang, Zhenwei An, Nan Zhuang, Mingxu Tao, Chen Zhang 0019, Kun Xu 0005, Songfang Huang, Yansong Feng 0002
ACL (1)7
2024 Probing Multimodal Large Language Models for Global and Local Semantic Representations
abstract
The advancement of Multimodal Large Language Models (MLLMs) has greatly accelerated the development of applications in understanding integrated texts and images. Recent works leverage image-caption datasets to train MLLMs, achieving state-of-the-art performance on image-to-text tasks. However, there are few studies exploring which layers of MLLMs make the most effort to the global image information, which plays vital roles in multimodal comprehension and generation. In this study, we find that the intermediate layers of models can encode more global semantic information, whose representation vectors perform better on visual-language entailment tasks, rather than the topmost layers. We further probe models regarding local semantic representations through object recognition tasks. We find that the topmost layers may excessively focus on local information, leading to a diminished ability to encode global information. Our code and data are released via https://github.com/kobayashikanna01/probing_MLLM_rep.
Mingxu Tao, Quzhe Huang, Kun Xu 0005, Yansong Feng 0002, Dongyan Zhao 0001
LREC/COLING3
2024 Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
abstract
Recently, the remarkable advance of the Large Language Model (LLM) has inspired researchers to transfer its extraordinary reasoning capability to both vision and language data. However, the prevailing approaches primarily regard the visual input as a prompt and focus exclusively on optimizing the text generation process conditioned upon vision content by a frozen LLM. Such an inequitable treatment of vision and language heavily constrains the model's potential. In this paper, we break through this limitation by representing both vision and language in a unified form. Specifically, we introduce a well-designed visual tokenizer to translate the non-linguistic image into a sequence of discrete tokens like a foreign language that LLM can read. The resulting visual tokens encompass high-level semantics worthy of a word and also support dynamic sequence length varying from the image. Coped with this tokenizer, the presented foundation model called LaVIT can handle both image and text indiscriminately under the same generative learning paradigm. This unification empowers LaVIT to serve as an impressive generalist interface to understand and generate multi-modal content simultaneously. Extensive experiments further showcase that it outperforms the existing models by a large margin on massive vision-language tasks. Our code and models are available at https://github.com/jy0205/LaVIT.
Kun Xu 0005, Chao Liao, Jianchao Tan, Quzhe Huang, Chengru Song, Dai Meng, Di Zhang 0026, Wenwu Ou, Kun Gai, Yadong Mu
ICLR2
2024 Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
abstract
In light of recent advances in multimodal Large Language Models (LLMs), there is increasing attention to scaling them from image-text data to more informative real-world videos. Compared to static images, video poses unique challenges for effective large-scale pre-training due to the modeling of its spatiotemporal dynamics. In this paper, we address such limitations in video-language pre-training with an efficient video decomposition that represents each video as keyframes and temporal motions. These are then adapted to an LLM using well-designed tokenizers that discretize visual and temporal information as a few tokens, thus enabling unified generative pre-training of videos, images, and text. At inference, the generated tokens from the LLM are carefully recovered to the original continuous pixel space to create various video content. Our proposed framework is both capable of comprehending and generating image and video content, as demonstrated by its competitive performance across 13 multimodal benchmarks in image and video understanding and generation. Our code and models are available at https://video-lavit.github.io.
Zhicheng Sun 0001, Kun Xu 0005, Hao Jiang 0032, Quzhe Huang, Chengru Song, Di Zhang 0026, Yang Song 0008, Kun Gai, Yadong Mu
ICML4
2024 RectifID: Personalizing Rectified Flow with Anchored Classifier Guidance
abstract
Customizing diffusion models to generate identity-preserving images from user-provided reference images is an intriguing new problem. The prevalent approaches typically require training on extensive domain-specific images to achieve identity preservation, which lacks flexibility across different use cases. To address this issue, we exploit classifier guidance, a training-free technique that steers diffusion models using an existing classifier, for personalized image generation. Our study shows that based on a recent rectified flow framework, the major limitation of vanilla classifier guidance in requiring a special classifier can be resolved with a simple fixed-point solution, allowing flexible personalization with off-the-shelf image discriminators. Moreover, its solving procedure proves to be stable when anchored to a reference flow trajectory, with a convergence guarantee. The derived method is implemented on rectified flow with different off-the-shelf image discriminators, delivering advantageous personalization results for human faces, live subjects, and certain objects. Code is available at https://github.com/feifeiobama/RectifID.
Zhicheng Sun 0001, Zhenhao Yang, Haozhe Chi, Kun Xu 0005, Hao Jiang 0032, Yang Song 0008, Kun Gai, Yadong Mu
NeurIPS5
2024 Structure-Aware Dialogue Modeling Methods for Conversational Semantic Role Labeling
abstract
Conversational semantic role labeling (CSRL) is believed to be a crucial step toward dialogue understanding. By incorporating the CSRL information into the conversational models, previous work (Xu et al., 2021) has confirmed the usefulness of CSRL to downstream conversation-based tasks, including multi-turn dialogue rewriting and multi-turn dialogue response generation. However, (Xu et al., 2021) found that the quality of the extracted CSRL structures would consequently affect the performance of downstream dialogue tasks while the performance of existing CSRL models is still unsatisfactory. There are two major problems in existing CSRL models to handle predicate-aware and conversational structural information. First, they ignore the fact that explicitly correlating the predicate and the context utterances could help the model better identify the arguments. Secondly, these models do not encode some vital conversational structural information, such as the speaker information which is necessary for modeling inter-speaker dependency. In this paper, we model the conversational structure-aware features based on three components: 1) the predicate-aware module which aims to capture rich correlations between the predicate and utterances; 2) a speaker-aware graph network which explicitly encodes the speaker-dependent information; 3) a novel structure-aware dialogue modeling method for the model warm-up. Experimental results on benchmark datasets show that our model significantly outperforms the baselines. We also examine the efficiency of our model and its effectiveness in low-resource scenarios. We find that our model can achieve better performance with less training time and training data than the existing models. In addition, further improvements are observed when applying the CSRL information extracted by our model into downstream dialogue tasks, which consistently indicates the superiority of our model.
Han Wu 0004, Kun Xu 0005, Linqi Song
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Discourse-Aware Graph Networks for Textual Logical Reasoning
abstract
Textual logical reasoning, especially question-answering (QA) tasks with logical reasoning, requires awareness of particular logical structures. The passage-level logical relations represent entailment or contradiction between propositional units (e.g., a concluding sentence). However, such structures are unexplored as current QA systems focus on entity-based relations. In this work, we propose logic structural-constraint modeling to solve the logical reasoning QA and introduce discourse-aware graph networks (DAGNs). The networks first construct logic graphs leveraging in-line discourse connectives and generic logic theories, then learn logic representations by end-to-end evolving the logic relations with an edge-reasoning mechanism and updating the graph features. This pipeline is applied to a general encoder, whose fundamental features are joined with the high-level logic features for answer prediction. Experiments on three textual logical reasoning datasets demonstrate the reasonability of the logical structures built in DAGNs and the effectiveness of the learned logic features. Moreover, zero-shot transfer results show the features' generality to unseen logical texts.
Yinya Huang, Lemao Liu, Kun Xu 0005, Liang Lin 0004, Xiaodan Liang
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Variational Graph Autoencoding as Cheap Supervision for AMR Coreference Resolution
abstract
Coreference resolution over semantic graphs like AMRs aims to group the graph nodes that represent the same entity.This is a crucial step for making document-level formal semantic representations.With annotated data on AMR coreference resolution, deep learning approaches have recently shown great potential for this task, yet they are usually data hungry and annotating data is costly.We propose a general pretraining method using variational graph autoencoder (VGAE) for AMR coreference resolution, which can leverage any general AMR corpus and even automatically parsed AMR data.Experiments on benchmarks show that the pretraining approach achieves performance gains of up to 6% absolute F1 points.Moreover, our model significantly improves on the previous state-of-theart model by up to 11% F1 points.
Irene Li, Linfeng Song, Kun Xu 0005, Dong Yu 0001
ACL (1)3
2022 Learning a Grammar Inducer from Massive Uncurated Instructional Videos
abstract
Video-aided grammar induction aims to leverage video information for finding more accurate syntactic grammars for accompanying text.While previous work focuses on building systems for inducing grammars on text that are well-aligned with video content, we investigate the scenario, in which text and video are only in loose correspondence.Such data can be found in abundance online, and the weak correspondence is similar to the indeterminacy problem studied in language acquisition.Furthermore, we build a new model that can better learn video-span correlation without manually designed features adopted by previous work.Experiments show that our model trained only on large-scale YouTube data with no textvideo alignment reports strong and robust performances across three unseen datasets, despite domain shift and noisy label issues.Furthermore our model yields higher F1 scores than the previous state-of-the-art systems trained on in-domain data.
Songyang Zhang 0004, Linfeng Song, Lifeng Jin, Haitao Mi, Kun Xu 0005, Dong Yu 0001, Jiebo Luo 0001
EMNLP5
2022 CASA: Conversational Aspect Sentiment Analysis for Dialogue Understanding
abstract
Dialogue understanding has always been a bottleneck for many conversational tasks, such as dialogue response generation and conversational question answering. To expedite the progress in this area, we introduce the task of conversational aspect sentiment analysis (CASA) that can provide useful fine-grained sentiment information for dialogue understanding and planning. Overall, this task extends the standard aspect-based sentiment analysis to the conversational scenario with several major adaptations. To aid the training and evaluation of data-driven methods, we annotate 3,000 chit-chat dialogues (27,198 sentences) with fine-grained sentiment information, including all sentiment expressions, their polarities and the corresponding target mentions. We also annotate an out-of-domain test set of 200 dialogues for robustness evaluation. Besides, we develop multiple baselines based on either pretrained BERT or self-attention for preliminary study. Experimental results show that our BERT-based model has strong performances for both in-domain and out-of-domain datasets, and thorough analysis indicates several potential directions for further improvements.
Linfeng Song, Chunlei Xin, Shaopeng Lai, Ante Wang, Jinsong Su, Kun Xu 0005
J. Artif. Intell. Res.6
2021 Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
abstract
Weakly supervised phrase grounding aims at learning region-phrase correspondences using only image-sentence pairs. A major challenge thus lies in the missing links between image regions and sentence phrases during training. To address this challenge, we leverage a generic object detector at training time, and propose a contrastive learning framework that accounts for both region-phrase and image-sentence matching. Our core innovation is the learning of a region-phrase score function, based on which an image-sentence score function is further constructed. Importantly, our region-phrase score function is learned by distilling from soft matching scores between the detected object names and candidate phrases within an image-sentence pair, while the image-sentence score function is supervised by ground-truth image-sentence pairs. The design of such score functions removes the need of object detection at test time, thereby significantly reducing the inference cost. Without bells and whistles, our approach achieves state-of-the-art results on visual phrase grounding, surpassing previous methods that require expensive object detectors at test time.
Liwei Wang 0009, Jing Huang 0014, Yin Li 0003, Kun Xu 0005, Zhengyuan Yang, Dong Yu 0001
CVPR4
2021 Joint Coreference Resolution and Character Linking for Multiparty Conversation
abstract
Character linking, the task of linking mentioned people in conversations to the real world, is crucial for understanding the conversations.For the efficiency of communication, humans often choose to use pronouns (e.g., "she") or normal phrases (e.g., "that girl") rather than named entities (e.g., "Rachel") in the spoken language, which makes linking those mentions to real people a much more challenging than a regular entity linking task.To address this challenge, we propose to incorporate the richer context from the coreference relations among different mentions to help the linking.On the other hand, considering that finding coreference clusters itself is not a trivial task and could benefit from the global character information, we propose to jointly solve these two tasks.Specifically, we propose C 2 , the joint learning model of Coreference resolution and Character linking.The experimental results demonstrate that C 2 can significantly outperform previous works on both tasks.Further analyses are conducted to analyze the contribution of all modules in the proposed model and the effect of all hyper-parameters.
Jiaxin Bai, Hongming Zhang 0009, Yangqiu Song, Kun Xu 0005
EACL4
2021 RAST: Domain-Robust Dialogue Rewriting as Sequence Tagging
abstract
The task of dialogue rewriting aims to reconstruct the latest dialogue utterance by copying the missing content from the dialogue context.Until now, the existing models for this task suffer from the robustness issue, i.e., performances drop dramatically when testing on a different dataset.We address this robustness issue by proposing a novel sequence-taggingbased model so that the search space is significantly reduced, yet the core of this task is still well covered.As a common issue of most tagging models for text generation, the model's outputs may lack fluency.To alleviate this issue, we inject the loss signal from BLEU or GPT-2 under a REINFORCE framework.Experiments show huge improvements of our model over the current state-of-the-art systems when transferring to another dataset.
Linfeng Song, Liwei Wang 0009, Kun Xu 0005, Zhaopeng Tu, Dong Yu 0001
EMNLP (1)4
2021 Instance-adaptive training with noise-robust losses against noisy labels
abstract
In order to alleviate the huge demand for annotated datasets for different tasks, many recent natural language processing datasets have adopted automated pipelines for fast-tracking usable data.However, model training with such datasets poses a challenge because popular optimization objectives are not robust to label noise induced in the annotation generation process.Several noise-robust losses have been proposed and evaluated on tasks in computer vision, but they generally use a single dataset-wise hyperparamter to control the strength of noise resistance.This work proposes novel instance-adaptive training frameworks to change dataset-wise hyperparameters of noise resistance in such losses to be instance-specific.Such instance-specific noise resistance hyperparameters are predicted by special instance-level label quality predictors, which are trained along with the main models.Experiments on noisy and corrupted NLP datasets show that proposed instance-adaptive training frameworks help increase the noiserobustness provided by such losses, promoting the use of the frameworks and associated losses in training NLP models with noisy data.
Lifeng Jin, Linfeng Song, Kun Xu 0005, Dong Yu 0001
EMNLP (1)3
2021 CSAGN: Conversational Structure Aware Graph Network for Conversational Semantic Role Labeling
abstract
Conversational semantic role labeling (CSRL) is believed to be a crucial step towards dialogue understanding.However, it remains a major challenge for existing CSRL parser to handle conversational structural information.In this paper, we present a simple and effective architecture for CSRL which aims to address this problem.Our model is based on a conversational structure-aware graph network which explicitly encodes the speaker dependent information.We also propose a multi-task learning method to further improve the model.Experimental results on benchmark datasets show that our model with our proposed training objectives significantly outperforms previous baselines.
Han Wu 0004, Kun Xu 0005, Linqi Song
EMNLP (1)2
2021 Exophoric Pronoun Resolution in Dialogues with Topic Regularization
abstract
Resolving pronouns to their referents has long been studied as a fundamental natural language understanding problem.Previous works on pronoun coreference resolution (PCR) mostly focus on resolving pronouns to mentions in text while ignoring the exophoric scenario.Exophoric pronouns are common in daily communications, where speakers may directly use pronouns to refer to some objects present in the environment without introducing the objects first.Although such objects are not mentioned in the dialogue text, they can often be disambiguated by the general topics of the dialogue.Motivated by this, we propose to jointly leverage the local context and global topics of dialogues to solve the out-of-text PCR problem.Extensive experiments demonstrate the effectiveness of adding topic regularization for resolving exophoric pronouns.
Xintong Yu 0002, Hongming Zhang 0009, Yangqiu Song, Changshui Zhang, Kun Xu 0005, Dong Yu 0001
EMNLP (1)5
2021 Video-aided Unsupervised Grammar Induction
abstract
Songyang Zhang, Linfeng Song, Lifeng Jin, Kun Xu, Dong Yu, Jiebo Luo. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Songyang Zhang 0004, Linfeng Song, Lifeng Jin, Kun Xu 0005, Dong Yu 0001, Jiebo Luo 0001
NAACL-HLT4
2021 Distant Finetuning with Discourse Relations for Stance Classification
Lifeng Jin, Kun Xu 0005, Linfeng Song, Dong Yu 0001
NLPCC (2)2
2021 Conversational Semantic Role Labeling
abstract
Semantic role labeling (SRL) aims to extract the arguments for each predicate in an input sentence. Traditional SRL can fail to analyze dialogues because it only works on every single sentence, while ellipsis and anaphora frequently occur in dialogues. To address this problem, we propose the conversational SRL task, where an argument can be the dialogue participants, a phrase in the dialogue history or the current sentence. As the existing SRL datasets are in the sentence level, we manually annotate semantic roles for 3000 chit-chat dialogues (27198 sentences) to boost the research in this direction. Experiments show that while traditional SRL systems (even with the help of coreference resolution or rewriting) perform poorly for analyzing dialogues, modeling dialogue histories and participants greatly helps the performance, indicating that adapting SRL to conversations is very promising for universal dialogue understanding. Our initial study by applying CSRL to two mainstream conversational tasks, dialogue response generation and dialogue context rewriting, also confirms the usefulness of CSRL.
Kun Xu 0005, Han Wu 0004, Linfeng Song, Haisong Zhang, Linqi Song, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Relation Extraction Exploiting Full Dependency Forests
abstract
Dependency syntax has long been recognized as a crucial source of features for relation extraction. Previous work considers 1-best trees produced by a parser during preprocessing. However, error propagation from the out-of-domain parser may impact the relation extraction performance. We propose to leverage full dependency forests for this task, where a full dependency forest encodes all possible trees. Such representations of full dependency forests provide a differentiable connection between a parser and a relation extraction model, and thus we are also able to study adjusting the parser parameters based on end-task loss. Experiments on three datasets show that full dependency forests and parser adjustment give significant improvements over carefully designed baselines, showing state-of-the-art or competitive performances on biomedical or newswire benchmarks.
Lifeng Jin, Linfeng Song, Yue Zhang 0004, Kun Xu 0005, Wei-Yun Ma, Dong Yu 0001
AAAI4
2020 Coordinated Reasoning for Cross-Lingual Knowledge Graph Alignment
abstract
Existing entity alignment methods mainly vary on the choices of encoding the knowledge graph, but they typically use the same decoding method, which independently chooses the local optimal match for each source entity. This decoding method may not only cause the “many-to-one” problem but also neglect the coordinated nature of this task, that is, each alignment decision may highly correlate to the other decisions. In this paper, we introduce two coordinated reasoning methods, i.e., the Easy-to-Hard decoding strategy and joint entity alignment algorithm. Specifically, the Easy-to-Hard strategy first retrieves the model-confident alignments from the predicted results and then incorporates them as additional knowledge to resolve the remaining model-uncertain alignments. To achieve this, we further propose an enhanced alignment model that is built on the current state-of-the-art baseline. In addition, to address the many-to-one problem, we propose to jointly predict entity alignments so that the one-to-one constraint can be naturally incorporated into the alignment prediction. Experimental results show that our model achieves the state-of-the-art performance and our reasoning methods can also significantly improve existing baselines.
Kun Xu 0005, Linfeng Song, Yansong Feng 0002, Yan Song 0003, Dong Yu 0001
AAAI1
2020 Structural Information Preserving for Graph-to-Text Generation
abstract
The task of graph-to-text generation aims at producing sentences that preserve the meaning of input graphs.As a crucial defect, the current state-of-the-art models may mess up or even drop the core structural information of input graphs when generating outputs.We propose to tackle this problem by leveraging richer training signals that can guide our model for preserving input information.In particular, we introduce two types of autoencoding losses, each individually focusing on different aspects (a.k.a.views) of input graphs.The losses are then back-propagated to better calibrate our model via multi-task training.Experiments on two benchmarks for graph-to-text generation show the effectiveness of our approach over a state-of-the-art baseline.Our code is available at http://github.com/ Soistesimmer/AMR-multiview.
Linfeng Song, Ante Wang, Jinsong Su, Yue Zhang 0004, Kun Xu 0005, Yubin Ge, Dong Yu 0001
ACL5
2020 ZPR2: Joint Zero Pronoun Recovery and Resolution using Multi-Task Learning and BERT
abstract
Zero pronoun recovery and resolution aim at recovering the dropped pronoun and pointing out its anaphoric mentions, respectively.We propose to better explore their interaction by solving both tasks together, while the previous work treats them separately.For zero pronoun resolution, we study this task in a more realistic setting, where no parsing trees or only automatic trees are available, while most previous work assumes gold trees.Experiments on two benchmarks show that joint modeling significantly outperforms our baseline that already beats the previous state of the arts.
Linfeng Song, Kun Xu 0005, Yue Zhang 0004, Jianshu Chen, Dong Yu 0001
ACL2
2020 Semantic Role Labeling Guided Multi-turn Dialogue ReWriter
abstract
For multi-turn dialogue rewriting, the capacity of effectively modeling the linguistic knowledge in dialog context and getting rid of the noises is essential to improve its performance.Existing attentive models attend to all words without prior focus, which results in inaccurate concentration on some dispensable words.In this paper, we propose to use semantic role labeling (SRL), which highlights the core semantic information of who did what to whom, to provide additional guidance for the rewriter model.Experiments show that this information significantly improves a RoBERTa-based model that already outperforms previous stateof-the-art systems.
Kun Xu 0005, Haochen Tan, Linfeng Song, Han Wu 0004, Haisong Zhang, Linqi Song, Dong Yu 0001
EMNLP (1)1
2020 DurIAN: Duration Informed Attention Network for Speech Synthesis
Chengzhu Yu, Heng Lu 0004, Na Hu, Meng Yu 0003, Chao Weng, Kun Xu 0005, Deyi Tuo, Shiyin Kang, Guangzhi Lei, Dan Su 0002, Dong Yu 0001
INTERSPEECH6
2019 Lattice CNNs for Matching Based Chinese Question Answering
abstract
Short text matching often faces the challenges that there are great word mismatch and expression diversity between the two texts, which would be further aggravated in languages like Chinese where there is no natural space to segment words explicitly. In this paper, we propose a novel lattice based CNN model (LCNs) to utilize multi-granularity information inherent in the word lattice while maintaining strong ability to deal with the introduced noisy information for matching based question answering in Chinese. We conduct extensive experiments on both document based question answering and knowledge based question answering tasks, and experimental results show that the LCNs models can significantly outperform the state-of-the-art matching models and strong baselines by taking advantages of better ability to distill rich but discriminative information from the word lattice input.
Yuxuan Lai, Yansong Feng 0002, Xiaohan Yu 0005, Zheng Wang 0001, Kun Xu 0005, Dongyan Zhao 0001
AAAI5
2019 Cross-lingual Knowledge Graph Alignment via Graph Matching Neural Network
abstract
Previous cross-lingual knowledge graph (KG) alignment studies rely on entity embeddings derived only from monolingual KG structural information, which may fail at matching entities that have different facts in two KGs. In this paper, we introduce the topic entity graph, a local sub-graph of an entity, to represent entities with their contextual information in KG. From this view, the KB-alignment task can be formulated as a graph matching problem; and we further propose a graph-attention based solution, which first matches all entities in two topic entity graphs, and then jointly model the local matching information to derive a graph-level matching vector. Experiments show that our model outperforms previous state-of-the-art methods by a large margin.
Kun Xu 0005, Liwei Wang 0009, Mo Yu, Yansong Feng 0002, Yan Song 0003, Zhiguo Wang 0006, Dong Yu 0001
ACL (1)1
2019 Multiplex Word Embeddings for Selectional Preference Acquisition
abstract
Hongming Zhang, Jiaxin Bai, Yan Song, Kun Xu, Changlong Yu, Yangqiu Song, Wilfred Ng, Dong Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Hongming Zhang 0009, Jiaxin Bai, Yan Song 0003, Kun Xu 0005, Changlong Yu, Yangqiu Song, Wilfred Ng, Dong Yu 0001
EMNLP/IJCNLP (1)4
2019 Efficient Global String Kernel with Random Features: Beyond Counting Substructures
abstract
Analysis of large-scale sequential data has been one of the most crucial tasks in areas such as bioinformatics, text, and audio mining. Existing string kernels, however, either (i) rely on local features of short substructures in the string, which hardly capture long discriminative patterns, (ii) sum over too many substructures, such as all possible subsequences, which leads to diagonal dominance of the kernel matrix, or (iii) rely on non-positive-definite similarity measures derived from the edit distance. Furthermore, while there have been works addressing the computational challenge with respect to the length of string, most of them still experience quadratic complexity in terms of the number of training samples when used in a kernel-based classifier. In this paper, we present a new class of global string kernels that aims to (i) discover global properties hidden in the strings through global alignments, (ii) maintain positive-definiteness of the kernel, without introducing a diagonal dominant kernel matrix, and (iii) have a training cost linear with respect to not only the length of the string but also the number of training string samples. To this end, the proposed kernels are explicitly defined through a series of different random feature maps, each corresponding to a distribution of random strings. We show that kernels defined this way are always positive-definite, and exhibit computational benefits as they always produce Random String Embeddings (RSE) that can be directly used in any linear classification models. Our extensive experiments on nine benchmark datasets corroborate that RSE achieves better or comparable accuracy in comparison to state-of-the-art baselines, especially with the strings of longer lengths. In addition, we empirically show that RSE scales linearly with the increase of the number and the length of string.
Lingfei Wu 0001, Ian En-Hsu Yen, Siyu Huo, Liang Zhao 0002, Kun Xu 0005, Liang Ma 0002, Shouling Ji, Charu C. Aggarwal
KDD5
2019 Scalable Global Alignment Graph Kernel Using Random Features: From Node Embedding to Graph Embedding
abstract
Graph kernels are widely used for measuring the similarity between graphs. Many existing graph kernels, which focus on local patterns within graphs rather than their global properties, suffer from significant structure information loss when representing graphs. Some recent global graph kernels, which utilizes the alignment of geometric node embeddings of graphs, yield state-of-the-art performance. However, these graph kernels are not necessarily positive-definite. More importantly, computing the graph kernel matrix will have at least quadratic time complexity in terms of the number and the size of the graphs. In this paper, we propose a new family of global alignment graph kernels, which take into account the global properties of graphs by using geometric node embeddings and an associated node transportation based on earth mover's distance. Compared to existing global kernels, the proposed kernel is positive-definite. Our graph kernel is obtained by defining a distribution over random graphs, which can naturally yield random feature approximations. The random feature approximations lead to our graph embeddings, which is named as "random graph embeddings" (RGE). In particular, RGE is shown to achieve (quasi-)linear scalability with respect to the number and the size of the graphs. The experimental results on nine benchmark datasets demonstrate that RGE outperforms or matches twelve state-of-the-art graph classification algorithms.
Lingfei Wu 0001, Ian En-Hsu Yen, Zhen Zhang 0007, Kun Xu 0005, Liang Zhao 0002, Xi Peng 0005, Yinglong Xia, Charu C. Aggarwal
KDD4
2018 Word Mover's Embedding: From Word2Vec to Document Embedding
abstract
Lingfei Wu, Ian En-Hsu Yen, Kun Xu, Fangli Xu, Avinash Balakrishnan, Pin-Yu Chen, Pradeep Ravikumar, Michael J. Witbrock. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Lingfei Wu 0001, Ian En-Hsu Yen, Kun Xu 0005, Fangli Xu, Avinash Balakrishnan, Pradeep Ravikumar, Michael Witbrock
EMNLP3
2018 SQL-to-Text Generation with Graph-to-Sequence Model
abstract
Previous work approaches the SQL-to-text generation task using vanilla Seq2Seq models, which may not fully capture the inherent graph-structured information in SQL query.In this paper, we first introduce a strategy to represent the SQL query as a directed graph and then employ a graph-to-sequence model to encode the global structure information into node embeddings.This model can effectively learn the correlation between the SQL query pattern and its interpretation.Experimental results on the WikiSQL dataset and Stackoverflow dataset show that our model significantly outperforms the Seq2Seq and Tree2Seq baselines, achieving the state-of-the-art performance. * Work done when the author
Kun Xu 0005, Lingfei Wu 0001, Zhiguo Wang 0006, Yansong Feng 0002, Vadim Sheinin
EMNLP1
2018 Exploiting Rich Syntactic Information for Semantic Parsing with Graph-to-Sequence Model
abstract
Existing neural semantic parsers mainly utilize a sequence encoder, i.e., a sequential LSTM, to extract word order features while neglecting other valuable syntactic information such as dependency or constituent trees.In this paper, we first propose to use the syntactic graph to represent three types of syntactic information, i.e., word order, dependency and constituency features; then employ a graph-tosequence model to encode the syntactic graph and decode a logical form.Experimental results on benchmark datasets show that our model is comparable to the state-of-the-art on Jobs640, ATIS, and Geo880.Experimental results on adversarial examples demonstrate the robustness of the model is also improved by encoding more syntactic information.
Kun Xu 0005, Lingfei Wu 0001, Zhiguo Wang 0006, Mo Yu, Vadim Sheinin
EMNLP1
2016 Question Answering on Freebase via Relation Extraction and Textual Evidence
abstract
Existing knowledge-based question answering systems often rely on small annotated training data.While shallow methods like relation extraction are robust to data scarcity, they are less expressive than the deep meaning representation methods like semantic parsing, thereby failing at answering questions involving multiple constraints.Here we alleviate this problem by empowering a relation extraction method with additional evidence from Wikipedia.We first present a neural network based relation extractor to retrieve the candidate answers from Freebase, and then infer over Wikipedia to validate these answers.Experiments on the WebQuestions question answering dataset show that our method achieves an F 1 of 53.3%, a substantial improvement over the state-of-the-art.
Kun Xu 0005, Siva Reddy, Yansong Feng 0002, Songfang Huang, Dongyan Zhao 0001
ACL (1)1
2016 Hybrid Question Answering over Knowledge Base and Free Text
abstract
Recent trend in question answering (QA) systems focuses on using structured knowledge bases (KBs) to find answers. While these systems are able to provide more precise answers than information retrieval (IR) based QA systems, the natural incompleteness of KB inevitably limits the question scope that the system can answer. In this paper, we present a hybrid question answering (hybrid-QA) system which exploits both structured knowledge base and free text to answer a question. The main challenge is to recognize the meaning of a question using these two resources, i.e., structured KB and free text. To address this, we map relational phrases to KB predicates and textual relations simultaneously, and further develop an integer linear program (ILP) model to infer on these candidates and provide a globally optimal solution. Experiments on benchmark datasets show that our system can benefit from both structured KB and free text, outperforming the state-of-the-art systems.
Kun Xu 0005, Yansong Feng 0002, Songfang Huang, Dongyan Zhao 0001
COLING1
2015 What Is the Longest River in the USA? Semantic Parsing for Aggregation Questions
abstract
Answering natural language questions against structured knowledge bases (KB) has been attracting increasing attention in both IR and NLP communities. The task involves two main challenges: recognizing the questions' meanings, which are then grounded to a given KB. Targeting simple factoid questions, many existing open domain semantic parsers jointly solve these two subtasks, but are usually expensive in complexity and resources.In this paper, we propose a simple pipeline framework to efficiently answer more complicated questions, especially those implying aggregation operations, e.g., argmax, argmin.We first develop a transition-based parsing model to recognize the KB-independent meaning representation of the user's intention inherent in the question. Secondly, we apply a probabilistic model to map the meaning representation, including those aggregation functions, to a structured query.The experimental results showed that our method can better understand aggregation questions, outperforming the state-of-the-art methods on the Free917 dataset while still maintaining promising performance on a more challenging dataset, WebQuestions, without extra training.
Kun Xu 0005, Sheng Zhang 0012, Yansong Feng 0002, Songfang Huang, Dongyan Zhao 0001
AAAI1
2015 Semantic Relation Classification via Convolutional Neural Networks with Simple Negative Sampling
abstract
Syntactic features play an essential role in identifying relationship in a sentence.Previous neural network models directly work on raw word sequences or constituent parse trees, thus often suffer from irrelevant information introduced when subjects and objects are in a long distance.In this paper, we propose to learn more robust relation representations from shortest dependency paths through a convolution neural network.We further take the relation directionality into account and propose a straightforward negative sampling strategy to improve the assignment of subjects and objects.Experimental results show that our method outperforms the state-of-theart approaches on the SemEval-2010 Task 8 dataset.
Kun Xu 0005, Yansong Feng 0002, Songfang Huang, Dongyan Zhao 0001
EMNLP1
2014 Answering Natural Language Questions via Phrasal Semantic Parsing
Kun Xu 0005, Sheng Zhang 0012, Yansong Feng 0002, Dongyan Zhao 0001
NLPCC1
2014 Efficient processing of label-constraint reachability queries in large graphs
Lei Zou 0001, Kun Xu 0005, Jeffrey Xu Yu, Lei Chen 0002, Yanghua Xiao, Dongyan Zhao 0001
Inf. Syst.2
2011 Answering label-constraint reachability in large graphs
abstract
In this paper, we study a variant of reachability queries, called label-constraint reachability (LCR) queries, specifically,given a label set S and two vertices u1 and u2 in a large directed graph G, we verify whether there exists a path from u1 to u2 under label constraint S. Like traditional reachability queries, LCR queries are very useful, such as pathway finding in biological networks, inferring over RDF (resource description f ramework) graphs, relationship finding in social networks. However, LCR queries are much more complicated than their traditional counterpart.Several techniques are proposed in this paper to minimize the search space in computing path-label transitive closure. Furthermore, we demonstrate the superiority of our method by extensive experiments.
Kun Xu 0005, Lei Zou 0001, Jeffrey Xu Yu, Lei Chen 0002, Yanghua Xiao, Dongyan Zhao 0001
CIKM1