Rui Xie 0005

dblp:86/2228-5 · DBLP profile ↗
← Back
24ranked-venue papers
1as first author
23since 2021 · last 2026
0000-0002-1116-7418ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 1 first-author · 15 since 2021Databases, data management, data science and information retrieval · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
abstract
Xintao Wang, Jian Yang, Weiyuan Li, Rui Xie, Jen-tse Huang, Jun Gao, Shuai Huang, Yueping Kang, Yuanli Guo, Hongwei Feng, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xintao Wang 0001, Jian Yang 0003, Weiyuan Li, Rui Xie 0005, Jen-tse Huang 0001, Yueping Kang, Yuanli Guo, Hongwei Feng, Yanghua Xiao
ACL (1)4
2026 AddSR: Accelerating diffusion-based blind super-resolution with adversarial diffusion distillation
Ying Tai, Rui Xie 0005, Chen Zhao 0002, Kai Zhang 0008, Zhenyu Zhang 0005, Jian Yang 0003
Pattern Recognit.2
2026 Spiking pyramid wavelet transformation for high-efficient and low-energy image restoration
Chen Zhao 0002, Xiantao Hu, Rui Xie 0005, Jian Yang 0003, Ying Tai
Pattern Recognit.6
2025 InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
abstract
Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in enhancing generation performance. However, current video captions often suffer from insufficient details, hallucinations and imprecise motion depiction, affecting the fidelity and consistency of generated videos. In this work, we propose a novel instance-aware structured caption framework, termed InstanceCap, to achieve instance-level and fine-grained video caption for the first time. Based on this scheme, we design an auxiliary models cluster to convert original video into instances to enhance instance fidelity. Video instances are further used to refine dense prompts into structured phrases, achieving concise yet precise descriptions. Furthermore, a 22K InstanceVid dataset is curated for training, and an enhancement pipeline that tailored to InstanceCap structure is proposed for inference. Experimental results demonstrate that our proposed InstanceCap significantly outperform previous models, ensuring high fidelity between captions and videos while reducing hallucinations.
Tiehan Fan, Kepan Nan, Rui Xie 0005, Penghao Zhou, Zhenheng Yang, Chaoyou Fu, Xiang Li 0041, Jian Yang 0003, Ying Tai
CVPR3
2025 Star: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution
abstract
Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively. Integrating text-to-video (T2V) models into video super-resolution for improved temporal modeling is straightforward. However, two key challenges remain: artifacts introduced by complex degradations in real-world scenarios, and compromised fidelity due to the strong generative capacity of powerful T2V models (\textit{e.g.}, CogVideoX-5B). To enhance the spatio-temporal quality of restored videos, we introduce\textbf{~\name} (\textbf{S}patial-\textbf{T}emporal \textbf{A}ugmentation with T2V models for \textbf{R}eal-world video super-resolution), a novel approach that leverages T2V models for real-world video super-resolution, achieving realistic spatial details and robust temporal consistency. Specifically, we introduce a Local Information Enhancement Module (LIEM) before the global attention block to enrich local details and mitigate degradation artifacts. Moreover, we propose a Dynamic Frequency (DF) Loss to reinforce fidelity, guiding the model to focus on different frequency components across diffusion steps. Extensive experiments demonstrate\textbf{~\name}~outperforms state-of-the-art methods on both synthetic and real-world datasets.
Rui Xie 0005, Yinhong Liu, Penghao Zhou, Chen Zhao 0002, Kai Zhang 0008, Zhenyu Zhang 0005, Jian Yang 0003, Zhenheng Yang, Ying Tai
ICCV1
2025 OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
abstract
Text-to-video (T2V) generation has recently garnered significant attention thanks to the large multi-modality model Sora. However, T2V generation still faces two important challenges: 1) Lacking a precise open sourced high-quality dataset. The previously popular video datasets, e.g.WebVid-10M and Panda-70M, overly emphasized large scale, resulting in the inclusion of many low-quality videos and short, imprecise captions. Therefore, it is challenging but crucial to collect a precise high-quality dataset while maintaining a scale of millions for T2V generation. 2) Ignoring to fully utilize textual information. Recent T2V methods have focused on vision transformers, using a simple cross attention module for video generation, which falls short of making full use of semantic information from text tokens. To address these issues, we introduce OpenVid-1M, a precise high-quality dataset with expressive captions. This open-scenario dataset contains over 1 million text-video pairs, facilitating research on T2V generation. Furthermore, we curate 433K 1080p videos from OpenVid-1M to create OpenVidHD-0.4M, advancing high-definition video generation. Additionally, we propose a novel Multi-modal Video Diffusion Transformer (MVDiT) capable of mining both structure information from visual tokens and semantic information from text tokens. Extensive experiments and ablation studies verify the superiority of OpenVid-1M over previous datasets and the effectiveness of our MVDiT.
Kepan Nan, Rui Xie 0005, Penghao Zhou, Tiehan Fan, Zhenheng Yang, Xiang Li 0041, Jian Yang 0003, Ying Tai
ICLR2
2024 Exploiting Duality in Open Information Extraction with Predicate Prompt
abstract
Open information extraction (OpenIE) aims to extract the schema-free triplets in the form of (subject, predicate, object) from a given sentence. Compared with general information extraction (IE), OpenIE poses more challenges for the IE models, especially when multiple complicated triplets exist in a sentence. To extract these complicated triplets more effectively, in this paper we propose a novel generative OpenIE model, namely DualOIE, which achieves a dual task at the same time as extracting some triplets from the sentence, i.e., converting the triplets into the sentence. Such dual task encourages the model to correctly recognize the structure of the given sentence and thus is helpful to extract all potential triplets from the sentence. Specifically, DualOIE extracts the triplets in two steps: 1) first extracting a sequence of all potential predicates, 2) then using the predicate sequence as a prompt to induce the generation of triplets. Our experiments on two benchmarks and our dataset constructed from Meituan demonstrate that DualOIE achieves the best performance among the state-of-the-art baselines. Furthermore, the online A/B test on Meituan platform shows that 0.93% improvement of QV-CTR and 0.56% improvement of UV-CTR have been obtained when the triplets extracted by DualOIE were leveraged in Meituan's search system.
Zhen Chen 0035, Deqing Yang, Yanghua Xiao, Zongyu Wang, Rui Xie 0005, Yunsen Xian
WSDM7
2024 A Segment Augmentation and Prediction Consistency Framework for Multi-label Unknown Intent Detection
abstract
Multi-label unknown intent detection is a challenging task where each utterance may contain not only multiple known but also unknown intents. To tackle this challenge, pioneers proposed to predict the intent number of the utterance first, then compare it with the results of known intent matching to decide whether the utterence contains unknown intent(s). Though they have made remarkable progress on this task, their methods still suffer from two important issues: (1) It is inadequate to extract multiple intents using only utterance encoding; (2) Optimizing two sub-tasks (intent number prediction and known intent matching) independently leads to inconsistent predictions. In this article, we propose to incorporate segment augmentation rather than only use utterance encoding to better detect multiple intents. We also design a prediction consistency module to bridge the gap between the two sub-tasks. Empirical results on MultiWOZ2.3 and MixSNIPS datasets show that our method achieves state-of-the-art performance and significantly improves the best baseline.
Miaoxin Chen, Cao Liu, Boqi Dai, Hai-Tao Zheng 0002, Hui Wang 0030, Rui Xie 0005, Hong-Gee Kim
ACM Trans. Knowl. Discov. Data7
2023 Causality-aware Concept Extraction based on Knowledge-guided Prompting
abstract
Siyu Yuan, Deqing Yang, Jinxi Liu, Shuyu Tian, Jiaqing Liang, Yanghua Xiao, Rui Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Deqing Yang, Jinxi Liu, Shuyu Tian, Jiaqing Liang, Yanghua Xiao, Rui Xie 0005
ACL (1)7
2023 Segment Augmentation and Prediction Consistency Neural Network for Multi-label Unknown Intent Detection
abstract
Multi-label unknown intent detection is a challenging task where each utterance may contain not only multiple known but also unknown intents. To tackle this challenge, pioneers proposed to predict the intent number of the utterance first, then compare it with the results of known intent matching to decide whether the utterance contains unknown intent(s). Though they have made remarkable progress on this task, their method still suffers from two important issues: 1) It is inadequate to extract multiple intents using only utterance encoding; 2) Optimizing two sub-tasks (intent number prediction and known intent matching) independently leads to inconsistent predictions. In this paper, we propose to incorporate segment augmentation rather than only use utterance encoding to better detect multiple intents. We also design a prediction consistency module to bridge the gap between the two sub-tasks. Empirical results on MultiWOZ2.3 show that our method achieves state-of-the-art performance and improves the best baseline significantly.
Miaoxin Chen, Cao Liu, Boqi Dai, Hai-Tao Zheng 0002, Jiansong Chen, Guanglu Wan, Rui Xie 0005
CIKM8
2023 Guide and Select: A Transformer-Based Multimodal Fusion Method for Points of Interest Description Generation
abstract
The task of Points of Interest (POI) description generation aims to generate an objective and informative description for a given POI based on POI-related information. High-quality descriptions can better guide users and improve the performance of POI-related recommendation systems. A practical POI description generation model should have effective multimodal fusion and information encoding methods suitable for various data forms. However, due to model structure and data utilization limitations, the previous method is challenging to meet the above requirements. We propose a novel Guide-Select multimodal fusion method that combines the guiding and selecting process to fuse various POI-related information efficiently. In addition, we propose a reasonable review encoding method and a category encoding method that has strong generalization ability. We integrate these methods into our Guide-Select Generation Model (GSGM). Experimental results demonstrate that our model significantly outperforms the state-of-the-art model while having a strong generalization ability on category information.
Wei Wang 0138, Niu Hu, Hai-Tao Zheng 0002, Rui Xie 0005, Wei Wu 0014
ICASSP5
2023 Towards Visual Taxonomy Expansion
abstract
Taxonomy expansion task is essential in organizing the ever-increasing volume of new concepts into existing taxonomies. Most existing methods focus exclusively on using textual semantics, leading to an inability to generalize to unseen terms and the "Prototypical Hypernym Problem." In this paper, we propose Visual Taxonomy Expansion (VTE), introducing visual features into the taxonomy expansion task. We propose a textual hypernymy learning task and a visual prototype learning task to cluster textual and visual semantics. In addition to the tasks on respective modalities, we introduce a hyper-proto constraint that integrates textual and visual semantics to produce fine-grained visual semantics. Our method is evaluated on two datasets, where we obtain compelling results. Specifically, on the Chinese taxonomy dataset, our method significantly improves accuracy by 8.75%. Additionally, our approach performs better than ChatGPT on the Chinese taxonomy dataset.
Tinghui Zhu, Jiaqing Liang, Haiyun Jiang, Yanghua Xiao, Zongyu Wang, Rui Xie 0005, Yunsen Xian
ACM Multimedia7
2023 AOG-LSTM: An adaptive attention neural network for visual storytelling
Wei Wang 0138, Hai-Tao Zheng 0002, Yong Jiang 0001, Hui Wang 0030, Rui Xie 0005, Wei Wu 0014
Neurocomputing8
2023 Noun Compound Interpretation With Relation Classification and Paraphrasing
abstract
Noun compounds are abundant in various languages and their interpretations have been applied in a wide range of NLP tasks. However, most existing work only uses relation classification- or paraphrasing-based methods to model this problem, failing in coverage or accuracy. We argue that the above two approaches are complementary to each other for the noun compound interpretation. In this paper, we propose a two-phase strategy to solve this task. The first phase is to perform the relation classification sub-task with a novel multi-view representation learning model. When noun compounds are predicted as the non-semantic relation, i.e., NA, or the confidence scores are below the threshold, the second phase, namely paraphrasing, will be triggered to interpret noun compounds with a contrastive slot filling method. To evaluate the effectiveness of our methods, we construct the largest Chinese dataset for noun compound interpretation in the life service domain. The experimental results on our constructed and public datasets prove the effectiveness of our solution. Furthermore, the online A/B testing on Meituan APP suggests that the Query View Click-Through Rate increases by 0.91% when noun compounds are used to enrich semantic information of items with the help of their interpretations on the platform.
Jiaqing Liang, Yanghua Xiao, Fubao Zhang, Zongyu Wang, Rui Xie 0005
IEEE Trans. Knowl. Data Eng.9
2022 Can Pre-trained Language Models Interpret Similes as Smart as Human?
abstract
Simile interpretation is a crucial task in natural language processing.Nowadays, pre-trained language models (PLMs) have achieved stateof-the-art performance on many tasks.However, it remains under-explored whether PLMs can interpret similes or not.In this paper, we investigate the ability of PLMs in simile interpretation by designing a novel task named Simile Property Probing, i.e., to let the PLMs infer the shared properties of similes.We construct our simile property probing datasets from both general textual corpora and humandesigned questions, containing 1,633 examples covering seven main categories.Our empirical study based on the constructed datasets shows that PLMs can infer similes' shared properties while still underperforming humans.To bridge the gap with human performance, we additionally design a knowledge-enhanced training objective by incorporating the simile knowledge into PLMs via knowledge embedding methods.Our method results in a gain of 8.58% in the probing task and 1.37% in the downstream task of sentiment classification.The datasets and code are publicly available at https://github.com/Abbey4799/PLMs- Interpret-Simile.
Qianyu He, Sijie Cheng, Zhixu Li, Rui Xie 0005, Yanghua Xiao
ACL (1)4
2022 PlugAT: A Plug and Play Module to Defend against Textual Adversarial Attack
abstract
Adversarial training, which minimizes the loss of adversarially perturbed examples, has received considerable attention. However, these methods require modifying all model parameters and optimizing the model from scratch, which is parameter inefficient and unfriendly to the already deployed models. As an alternative, we propose a pluggable defense module PlugAT, to provide robust predictions by adding a few trainable parameters to the model inputs while keeping the original model frozen. To reduce the potential side effects of using defense modules, we further propose a novel forgetting restricted adversarial training, which filters out bad adversarial examples that impair the performance of original ones. The PlugAT-equipped BERT model substantially improves robustness over several strong baselines on various text classification tasks, whilst training only 9.1% parameters. We observe that defense modules trained under the same model architecture have domain adaptation ability between similar text classification datasets.
Rong Bao, Qin Liu 0010, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Rui Xie 0005, Wei Wu 0014
COLING7
2022 Making Parameter-efficient Tuning More Efficient: A Unified Framework for Classification Tasks
abstract
Large pre-trained language models (PLMs) have demonstrated superior performance in industrial applications. Recent studies have explored parameter-efficient PLM tuning, which only updates a small amount of task-specific parameters while achieving both high efficiency and comparable performance against standard fine-tuning. However, all these methods ignore the inefficiency problem caused by the task-specific output layers, which is inflexible for us to re-use PLMs and introduces non-negligible parameters. In this work, we focus on the text classification task and propose plugin-tuning, a framework that further improves the efficiency of existing parameter-efficient methods with a unified classifier. Specifically, we re-formulate both token and sentence classification tasks into a unified language modeling task, and map label spaces of different tasks into the same vocabulary space. In this way, we can directly re-use the language modeling heads of PLMs, avoiding introducing extra parameters for different tasks. We conduct experiments on six classification benchmarks. The experimental results show that plugin-tuning can achieve comparable performance against fine-tuned PLMs, while further saving around 50% parameters on top of other parameter-efficient methods.
Xin Zhou 0012, Ruotian Ma, Yicheng Zou, Xuanting Chen, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Rui Xie 0005, Wei Wu 0014
COLING8
2022 Visualizable or Non-visualizable? Exploring the Visualizability of Concepts in Multi-modal Knowledge Graph
Xueyao Jiang, Ailisi Li, Jiaqing Liang, Bang Liu 0003, Rui Xie 0005, Wei Wu 0014, Zhixu Li, Yanghua Xiao
DASFAA (1)5
2022 Retrieval Enhanced Segment Generation Neural Network for Task-Oriented Dialogue Systems
abstract
For task-oriented dialogue systems, Natural Language Generation (NLG) is the last and vital step which aims at generating an appropriate response according to the dialogue act (DA). While end-to-end neural networks have achieved promising performances on this task, the existing models still struggle to avoid slot mistakes. To address this challenge, we propose a novel segmented generation approach in this paper. The proposed method operates by progressively generating text for the span between two adjacent keywords (act type and slots) in semantically ordered DA. This procedure is recursively applied from left to right until a response is completed. Besides, a retrieval mechanism is utilized to better match the diversity and fluency in human language. Experimental results on four datasets demonstrate that our model achieves state-of-the-art slot error rate and also gets competitive performance on BLEU score with all strong baselines.
Miaoxin Chen, Zibo Lin, Rongyi Sun, Kai Ouyang, Hai-Tao Zheng 0002, Rui Xie 0005, Wei Wu 0014
ICASSP6
2022 Learning What You Need from What You Did: Product Taxonomy Expansion with User Behaviors Supervision
abstract
Taxonomies have been widely used in various domains to underpin numerous applications. Specially, product taxonomies serve an essential role in the e-commerce domain for the recommendation, browsing, and query understanding. However, taxonomies need to constantly capture the newly emerged terms or concepts in e-commerce platforms to keep up-to-date, which is expensive and labor-intensive if it relies on manual maintenance and updates. Therefore, we target the taxonomy expansion task to attach new concepts to existing taxonomies automatically. In this paper, we present a self-supervised and user behavior-oriented product taxonomy expansion framework to append new concepts into existing taxonomies. Our framework extracts hyponymy relations that conform to users' intentions and cognition. Specifically, i) to fully exploit user behavioral information, we extract candidate hyponymy relations that match user interests from query-click concepts; ii) to enhance the semantic information of new concepts and better detect hyponymy relations, we model concepts and relations through both user-generated content and structural information in existing taxonomies and user click logs, by leveraging Pre-trained Language Models and Graph Neural Network combined with Contrastive Learning; iii) to reduce the cost of dataset construction and overcome data skews, we construct a high-quality and balanced training dataset from existing taxonomy with no supervision. Extensive experiments on real-world product taxonomies in Meituan Platform, a leading Chinese vertical e-commerce platform to order take-out with more than 70 million daily active users, demonstrate the superiority of our proposed framework over state-of-the-art methods. Notably, our method enlarges the size of real-world product taxonomies from 39,263 to 94,698 relations with 88% precision. Our implementation is available: https://github.com/AdaCheng/Product_Taxonomy_Expansion.
Sijie Cheng, Zhouhong Gu, Bang Liu 0003, Rui Xie 0005, Wei Wu 0014, Yanghua Xiao
ICDE4
2022 AMR-to-Text Generation with Graph Structure Reconstruction and Coverage Mechanism
abstract
Generating text from abstract meaning repre-sentation (AMR) is a challenging task. Graph-to-sequence (Graph2Seq-based) methods and pre-trained-based methods are proposed for this task. However, both methods have advantages and disadvantages. Graph2Seq-based methods can make use of the structural information of the graph but not the extra knowledge, while pre-trained-based methods have the advantage of the utilization of extra knowledge but may lose the structural information. In addition, both types of methods often suffer from the under- and over-translation problem. To address these prob-lems, we propose a graph structure reconstruction and coverage enhanced model for this task. The graph structure reconstruction uses two auxiliary objectives, relationship prediction and distance prediction of nodes in AMR graphs to enhance the information of graph structure. In addition, we design a coverage mechanism to solve the problem of information under-translation or over-translation in AMR-to-text generation. Experimental results on three datasets show that our proposed method outperforms the existing methods significantly.
Junxin Li, Wei Wang 0138, Hai-Tao Zheng 0002, Rui Xie 0005, Wei Wu 0014
IJCNN5
2022 A Sequence-to-Sequence Model for Large-scale Chinese Abbreviation Database Construction
abstract
Abbreviations often used in our daily communication play an important role in natural language processing. Most of the existing studies regard the Chinese abbreviation prediction as a sequence labeling problem. However, sequence labeling models usually ignore label dependencies in the process of abbreviation prediction, and the label prediction of each character should be conditioned on its previous labels. In this paper, we propose to formalize the Chinese abbreviation prediction task as a sequence generation problem, and a novel sequence-to-sequence model is designed. To boost the performance of our deep model, we further propose a multi-level pre-trained model that incorporates character, word, and concept-level embeddings. To evaluate our methods, a new dataset for Chinese abbreviation prediction is automatically built, which contains 81,351 pairs of full forms and abbreviations. Finally, we conduct extensive experiments on a public dataset and the built dataset, and the experimental results on both datasets show that our model outperforms the state-of-the-art methods. More importantly, we build a large-scale database for a specific domain, i.e., life services in Meituan Inc., with high accuracy of about 82.7%, which contains 4,134,142 pairs of full forms and abbreviations. The online A/B testing on Meituan APP and Dianping APP suggests that Click-Through Rate increases by 0.59% and 0.86% respectively when the built database is used in the searching system. We have released our API on http://kw.fudan.edu.cn/ddemos/abbr/ with over 87k API calls in 9 months.
Chao Wang 0095, Tianyi Zhuang, Yanghua Xiao, Wei Wang 0009, Rui Xie 0005
WSDM8
2021 Large-Scale Multi-granular Concept Extraction Based on Machine Reading Comprehension
Deqing Yang, Jiaqing Liang, Jilun Sun, Jingyue Huang, Kaiyan Cao, Yanghua Xiao, Rui Xie 0005
ISWC8
2020 Iterative Strategy for Named Entity Recognition with Imperfect Annotations
Yunian Chen, Xuezhi Cao, Rui Xie 0005
NLPCC (2)5