Xiaoyan Cai

dblp:05/8374 · DBLP profile ↗
← Back
53ranked-venue papers
14as first author
33since 2021 · last 2027
0000-0002-1406-107XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 37 · 9 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2027 Flexible job shop scheduling problem with critical operation driven outsourcing strategy and interval grey processing time
Tianyu Yan, Xiaoyan Cai, Zongyan Cai
Expert Syst. Appl.2
2026 RefleXNet: Targeted Self-Reflection for Accurate Chest X-ray Reporting
abstract
Automated interpretation and reporting of chest X-rays (CXRs) hold significant promise in reducing diagnostic errors and supporting radiologists under heavy clinical workloads. However, existing methods typically rely on global visual features and token-level supervision, limiting their sensitivity to subtle abnormalities and reducing their clinical reliability. To address these challenges, we present Reflective X-ray Network (RefleXNet), which systematically integrates multi-scale visual feature fusion and anatomical relational reasoning with a targeted self-reflective learning strategy. RefleXNet first constructs multi-scale visual representations and captures anatomical context through graph-based relational modeling. Building upon these representations, we introduce a targeted self-reflection strategy that uses clinically guided feedback from generated reports to selectively refine abnormality predictions and their associated region-level visual features. Extensive experiments on MIMIC-CXR demonstrate that RefleXNet consistently outperforms state-of-the-art baselines across clinical factual correctness metrics. Notably, our compact 3B-parameter model surpasses several recent models with over twice the parameter count. Additionally, RefleXNet exhibits strong generalization performance in zero-shot evaluations on IU-Xray compared with leading multimodal language models, highlighting its robustness and clinical effectiveness.
Xin Mei, Rui Mao 0010, Xiaoyan Cai, Libin Yang, Erik Cambria
AAAI3
2026 MLoRA+: Transformer-fusion mixture-of-LoRA network for multi-domain click-through rate prediction
Dehong Gao, Shufan Chen, Luwei Yang, Haining Gao, Muyang Wu, Shanqing Yu, Qi Xuan 0001, Libin Yang, Xiaoyan Cai
Expert Syst. Appl.11
2026 Spatio -temporal-aware preference optimization for personalized radiology report generation
Zhenjie Luo, Libin Yang, Quan Tang 0010, Shirui Pan, Dehong Gao, Xiaoyan Cai
Expert Syst. Appl.7
2026 Two-stage symbolic task planning for space robots integrating domain knowledge
Tongdian Wang, Xiaoyan Cai, Dayu Zhang
Expert Syst. Appl.3
2026 SDGT: LLMs fine-tuning with seed-driven growth technology based on GPT-4 data expansion
Dehong Gao, Jiayi Dai, Sen Liu 0004, Linbo Jin, Wen Jiang 0002, Shanqing Yu, Qi Xuan 0001, Xiaoyan Cai, Libin Yang
Neurocomputing8
2026 FiR-Rad: Fine-Grained Reinforcement With Structured Reasoning for Chest X-Ray Report Generation
abstract
Automated chest X-ray report generation requires not only clinical accuracy but also transparent and interpretable diagnostic reasoning. In this work, we propose FiR-Rad, a two-stage framework that combines explicit structured reasoning with targeted fine-grained optimization. In the first stage, a supervised chain-of-thought approach guides the model to sequentially analyze and describe a comprehensive range of clinically significant thoracic abnormalities, ensuring clinically meaningful coverage. In the second stage, we introduce a segment-level reinforcement learning strategy based on Group Relative Policy Optimization (GRPO), which assigns precise rewards to each disease-specific reasoning step by evaluating the accuracy of corresponding findings in the synthesized report. This design provides direct feedback for intermediate reasoning and encourages consistency between detailed abnormality analysis and final diagnostic conclusions. Experimental results on the MIMIC-CXR and IU-Xray datasets demonstrate that our framework achieves state-of-the-art performance across clinical and linguistic metrics, with strong zero-shot generalization on IU-Xray. The proposed method significantly enhances interpretability and clinical accuracy, effectively addressing key limitations in automated radiology report generation.
Xin Mei, Libin Yang, Dehong Gao, Xiaoyan Cai, Junwei Han 0001, Tianming Liu 0001
IEEE Trans. Medical Imaging4
2025 Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples
abstract
Existing Vision-Language Pretraining (VLP) methods have achieved remarkable improvements across a variety of vision-language tasks, confirming their effectiveness in capturing coarse-grained semantic correlations. However, their capability for fine-grained understanding, which is critical for many nuanced vision-language applications, remains limited. Prevailing VLP models often overlook the intricate distinctions in expressing different modal features and typically depend on the similarity of holistic features for cross-modal interactions. Moreover, these models directly align and integrate features from different modalities, focusing more on coarse-grained general representations, thus failing to capture the nuanced differences necessary for tasks demanding a more detailed perception. In response to these limitations, we introduce Negative Augmented Samples(NAS), a refined vision-language pretraining model that innovatively incorporates NAS to specifically address the challenge of fine-grained understanding. NAS utilizes a Visual Dictionary(VD) as a semantic bridge between visual and linguistic domains. Additionally, it employs a Negative Visual Augmentation(NVA) method based on the VD to generate challenging negative image samples. These samples deviate from positive samples exclusively at the token level, thereby necessitating that the model discerns the subtle disparities between positive and negative samples with greater precision. Comprehensive experiments validate the efficacy of NAS components and underscore its potential to enhance fine-grained vision-language comprehension.
Yeyuan Wang, Dehong Gao, Lei Yi, Linbo Jin, Jinxia Zhang, Libin Yang, Xiaoyan Cai
AAAI7
2025 CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
abstract
The impressive performance of Large Language Model (LLM) has prompted researchers to develop Multi-modal LLM (MLLM), which has shown great potential for various multi-modal tasks. However, current MLLM often struggles to effectively address fine-grained multi-modal challenges. We argue that this limitation is closely linked to the models’ visual grounding capabilities. The restricted spatial awareness and perceptual acuity of visual encoders frequently lead to interference from irrelevant background information in images, causing the models to overlook subtle but crucial details. As a result, achieving fine-grained regional visual comprehension becomes difficult. In this paper, we break down multi-modal understanding into two stages, from Coarse to Fine (CoF). In the first stage, we prompt the MLLM to locate the approximate area of the answer. In the second stage, we further enhance the model’s focus on relevant areas within the image through visual prompt engineering, adjusting attention weights of pertinent regions. This, in turn, improves both visual grounding and overall performance in downstream tasks. Our experiments show that this approach significantly boosts the performance of baseline models, demonstrating notable generalization and effectiveness. Our CoF approach is available online at https://github.com/Gavin001201/CoF.
Yeyuan Wang, Dehong Gao, Rujiao Long, Lei Yi, Xiaoyan Cai, Libin Yang, Jinxia Zhang, Shanqing Yu, Qi Xuan 0001
ICASSP6
2025 Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
abstract
Despite the significant success of Large Vision-Language models(LVLMs), these models still suffer hallucinations when describing images, generating answers that include non-factual objects. It is reported that these models tend to overfocus on certain irrelevant image tokens that do not contain critical information for answering the question and distort the output. To address this, we propose an Instruction-Aligned Visual Attention(IAVA) approach, which identifies irrelevant tokens by comparing changes in attention weights under two different instructions. By applying contrastive decoding, we dynamically adjust the logits generated from original image tokens and irrelevant image tokens, reducing the model’s over-attention to irrelevant information. The experimental results demonstrate that IAVA consistently outperforms existing decoding techniques on benchmarks such as MME, POPE, and TextVQA in mitigating object hallucinations. Our IAVA approach is available online at https://github.com/Lee-lab558/IAVA.
Dehong Gao, Yeyuan Wang, Linbo Jin, Shanqing Yu, Xiaoyan Cai, Libin Yang
ICME6
2025 Mask-Guided Visual Text Transformer for Radiology Reports Representation Learning
Xiaoyan Cai
KSEM (4)2
2025 ChatGPT based contrastive learning for radiology report summarization
Zhenjie Luo, Zuowei Jiang, Xiaoyan Cai, Dehong Gao, Libin Yang
Expert Syst. Appl.4
2025 Adaptive Medical Topic Learning for Enhanced Fine-Grained Cross-Modal Alignment in Medical Report Generation
abstract
Medical report generation refers to the automatic creation of accurate and coherent diagnostic reports for medical images. This task can alleviate the workload of radiologists, enhance the efficiency of disease diagnosis, and therefore holds significant value and challenges. Considering the feature differences between different modalities, existing methods primarily focus on facilitating medical report generation through cross-modal alignment of images and texts. However, since medical images are very similar to each other, it is difficult to tag obvious objects, making most methods limited to coarse-grained image-text global alignment. In this paper, we propose a medical report generation model based on adaptive topic learning and fine-grained cross-modal alignment, which aligns images and texts from medical topic perspective and token perspective. From the medical topic perspective, a global-local contrastive loss is introduced to adaptively learn efficient medical topic features, and medical topics are utilized to map images and texts to the same semantic space for fine-grained alignment. From the token perspective, a token prediction module is designed to enable the model to focus on important local information by predicting the key tokens contained in the report. Experimental results on the two public datasets (i.e. IU-Xray and MIMIC-CXR) demonstrate that our proposed model outperforms state-of-the-art baselines.
Xin Mei, Libin Yang, Dehong Gao, Xiaoyan Cai, Junwei Han 0001, Tianming Liu 0001
IEEE Trans. Multim.4
2025 FedSTS: A Stratified Client Selection Framework for Consistently Fast Federated Learning
abstract
In this article, we investigate random client selection in the context of horizontal federated learning (FL), whereby only a randomly selected subset of clients transmit their model updates to the server instead of yielding all clients involved. Many researchers have demonstrated that clustering-based client selection constitutes a simple yet efficacious approach to the identification of those clients possessing representative gradient information. Despite the extensive body of research on modified selection methodologies, the majority of prior work is predicated upon the assumption of consistently effective clustering. However, raw gradient-based clustering methods are subject to several challenges: 1) poor effectiveness, the raw high-dimensional gradient of a client is too complex to serve as an appropriate feature for grouping, resulting in large intra-cluster distances and 2) fluctuating effectiveness, due to inherent limitations in clustering, the effectiveness can vary significantly, leading to clusters with diverse levels of heterogeneity. In practice, suboptimal and inconsistent clustering effects can result in clusters with low intra-cluster similarity among clients. The selection of clients from such clusters may impede the overall convergence of training. In this article, we propose FedSTS, a novel client selection scheme to accelerate the FL convergence by variance reduction. The main idea of FedSTS is to stratify a compressed model update in order to ensure an excellent grouping effect, and at the same time reduce the cross-client variance by re-allocating the sample chance among different groups based on their diverse heterogeneity. It strikes this convergence acceleration by paying more attention to those client groups with relatively low similarity and then improving the representativeness of the selected subset as much as possible. Theoretically, we demonstrate the critical improvement of the proposed scheme in variance reduction and present equivalence conditions among different client selection methods. We also present the tighter convergence guarantee of the proposed method thanks to the variance reduction. Experimental results confirm the exceeded efficiency of our approach compared to alternatives.
Dehong Gao, Duanxiao Song, Guangyuan Shen, Xiaoyan Cai, Libin Yang, Gongshen Liu, Zhen Wang 0004
IEEE Trans. Neural Networks Learn. Syst.4
2024 RadChat: A Radiology Chatbot Incorporating Clinical Context for Radiological Reports Summarization
abstract
Radiological Report Summarization (RRS) involves automated summarization of key impressions derived from identified findings, intending to alleviate the workload and stress experienced by radiologists. Many existing RRS methods predominantly concentrate on summarizing findings, neglecting crucial clinical context, such as the patient’s previous medical examinations. This context, which is a focal point for radiologists, plays a critical role in producing comprehensive and accurate impressions. This paper endeavors to emulate the workflows of radiologists by incorporating the patient’s clinical context alongside current findings. To achieve this, we reconceptualize RRS as a conversational question-answering task, generating temporal radiological conversations. These conversations are subsequently employed to fine-tune a large chat model. The resulting radiology chatbot, RadChat, demonstrates superior performance in RRS task, showcasing the potential of integrating clinical context for more accurate impressions. Experimental results conducted on the MIMIC-CXR dataset validate the superiority of RadChat in comparison to state-of-the-art baselines.
Xin Mei, Libin Yang, Dehong Gao, Xiaoyan Cai, Tianming Liu 0001, Junwei Han 0001
BIBM4
2024 MoDULA: Mixture of Domain-Specific and Universal LoRA for Multi-Task Learning
abstract
Yufei Ma, Zihan Liang, Huangyu Dai, Ben Chen, Dehong Gao, Zhuoran Ran, Wang Zihan, Linbo Jin, Wen Jiang, Guannan Zhang, Xiaoyan Cai, Libin Yang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yufei Ma 0011, Zihan Liang 0001, Huangyu Dai, Ben Chen 0004, Dehong Gao, Zhuoran Ran, Linbo Jin, Wen Jiang 0002, Xiaoyan Cai, Libin Yang
EMNLP11
2024 Medical Report Generation via Multimodal Spatio-Temporal Fusion
abstract
Medical report generation aims at automating the synthesis of accurate and comprehensive diagnostic reports from radiological images. The task can significantly enhance clinical decision-making and alleviate the workload on radiologists. Existing works normally generate reports from single chest radiographs, although historical examination data also serve as crucial references for radiologists in real-world clinical settings. To address this constraint, we introduce a novel framework that mimics the workflow of radiologists. This framework compares past and present patient images to monitor disease progression and incorporates prior diagnostic reports as references for generating current personalized reports. We tackle the textual diversity challenge in cross-modal tasks by promoting style-agnostic discrete report representation learning and token generation. Furthermore, we propose a novel spatio-temporal fusion method with multi-granularities to fuse textual and visual features by disentangling the differences between current and historical data. We also tackle token generation biases, which arise from long-tail frequency distributions, proposing a novel feature normalization technique. This technique ensures unbiased generation for tokens, whether they are frequent or infrequent, enabling the robustness of report generation for rare diseases. Experimental results on the two public datasets demonstrate that our proposed model outperforms state-of-the-art baselines.
Xin Mei, Rui Mao 0010, Xiaoyan Cai, Libin Yang, Erik Cambria
ACM Multimedia3
2024 MLoRA: Multi-Domain Low-Rank Adaptive Network for CTR Prediction
abstract
Click-through rate (CTR) prediction is one of the fundamental tasks in the industry, especially in e-commerce, social media, and streaming media. It directly impacts website revenues, user satisfaction, and user retention. However, real-world production platforms often encompass various domains to cater for diverse customer needs. Traditional CTR prediction models struggle in multi-domain recommendation scenarios, facing challenges of data sparsity and disparate data distributions across domains. Existing multi-domain recommendation approaches introduce specific-domain modules for each domain, which partially address these issues but often significantly increase model parameters and lead to insufficient training. In this paper, we propose a Multi-domain Low-Rank Adaptive network (MLoRA) for CTR prediction, where we introduce a specialized LoRA module for each domain. This approach enhances the model’s performance in multi-domain CTR prediction tasks and is able to be applied to various deep-learning models. We evaluate the proposed method on several multi-domain datasets. Experimental results demonstrate our MLoRA approach achieves a significant improvement compared with state-of-the-art baselines. Furthermore, we deploy it in the production environment of the Alibaba.COM 1. The online A/B testing results indicate the superiority and flexibility in real-world production environments. The code of our MLoRA is publicly available 2.
Haining Gao, Dehong Gao, Luwei Yang, Libin Yang, Xiaoyan Cai, Wei Ning
RecSys6
2024 LLMs-based machine translation for E-commerce
Dehong Gao, Kaidi Chen, Ben Chen 0004, Huangyu Dai, Linbo Jin, Wen Jiang 0002, Wei Ning, Shanqing Yu, Qi Xuan 0001, Xiaoyan Cai, Libin Yang, Zhen Wang 0004
Expert Syst. Appl.10
2024 FashionGPT: LLM instruction fine-tuning with multiple LoRA-adapter fusion
Dehong Gao, Yufei Ma 0011, Sen Liu 0004, Mengfei Song, Linbo Jin, Wen Jiang 0002, Wei Ning, Shanqing Yu, Qi Xuan 0001, Xiaoyan Cai, Libin Yang
Knowl. Based Syst.11
2024 An Inductive Reasoning Model based on Interpretable Logical Rules over temporal knowledge graph
Xin Mei, Libin Yang, Zuowei Jiang, Xiaoyan Cai, Dehong Gao, Junwei Han 0001, Shirui Pan
Neural Networks4
2024 PhraseAug: An Augmented Medical Report Generation Model With Phrasebook
abstract
Medical report generation is a valuable and challenging task, which automatically generates accurate and fluent diagnostic reports for medical images, reducing workload of radiologists and improving efficiency of disease diagnosis. Fine-grained alignment of medical images and reports facilitates the exploration of close correlations between images and texts, which is crucial for cross-modal generation. However, visual and linguistic biases caused by radiologists' writing styles make cross-modal image-text alignment difficult. To alleviate visual-linguistic bias, this paper discretizes medical reports and introduces an intermediate modality, i.e. phrasebook, consisting of key noun phrases. As discretized representation of medical reports, phrasebook contains both disease-related medical terms, and synonymous phrases representing different writing styles which can identify synonymous sentences, thereby promoting fine-grained alignment between images and reports. In this paper, an augmented two-stage medical report generation model with phrasebook (PhraseAug) is developed, which combines medical images, clinical histories and writing styles to generate diagnostic reports. In the first stage, phrasebook is used to extract semantically relevant important features and predict key phrases contained in the report. In the second stage, medical reports are generated according to the predicted key phrases which contain synonymous phrases, promoting our model to adapt to different writing styles and generating diverse medical reports. Experimental results on two public datasets, IU-Xray and MIMIC-CXR, demonstrate that our proposed PhraseAug outperforms state-of-the-art baselines.
Xin Mei, Libin Yang, Denghong Gao, Xiaoyan Cai, Junwei Han 0001, Tianming Liu 0001
IEEE Trans. Medical Imaging4
2023 ChestXRayBERT: A Pretrained Language Model for Chest Radiology Report Summarization
abstract
Automatically generating the “impression” section of a radiology report given the “findings” section can summarize as much salient information of the “findings” section as possible, thus promoting more effective communication between radiologists and referring physicians. To significantly reduce the workload of radiologists, we develop and evaluate a novel framework of abstractive summarization methods to automatically generate the “impression” section of chest radiology reports. Despite recent advancements in natural language process (NLP) field such as BERT and its variants, existing abstractive summarization models and methods could not be directly applied to radiology reports, partly due to domain-specific radiology terminology. In response, we develop a pre-trained language model in the chest radiology domain, named ChestXRayBERT, to solve the problem of automatically summarizing chest radiology reports. Specifically, we first collect radiology-related scientific papers as pre-training corpus and pre-train a ChestXRayBERT on it. Then, an abstractive summarization model is proposed, which consists of the pre-trained ChestXRayBERT and a Transformer decoder. Finally, the model is fine-tuned on chest X-ray reports for the abstractive summarization task. When evaluated on the publicly available OPEN-I and MIMIC-CXR datasets, the performance of our proposed model achieves significant improvement compared with other neural networks-based abstractive summarization models. In general, the proposed ChestXRayBERT demonstrates the feasibility and promise of tailoring and extending advanced NLP techniques to the domain of medical imaging and radiology, as well as in the broader biomedicine and healthcare fields in the future.
Xiaoyan Cai, Sen Liu 0004, Junwei Han 0001, Libin Yang, Tianming Liu 0001
IEEE Trans. Multim.1
2022 An Adaptive Logical Rule Embedding Model for Inductive Reasoning over Temporal Knowledge Graphs
abstract
Temporal knowledge graphs (TKGs) extrapolation reasoning predicts future events based on historical information, which has great research significance and broad application value.Existing methods can be divided into embeddingbased methods and logical rule-based methods.Embedding-based methods rely on learned entity and relation embeddings to make predictions and thus lack interpretability.Logical rule-based methods bring scalability problems due to being limited by the learned logical rules.We combine the two methods to capture deep causal logic by learning rule embeddings, and propose an interpretable model for temporal knowledge graph reasoning called adaptive logical rule embedding model for inductive reasoning (ALRE-IR).ALRE-IR can adaptively extract and assess reasons contained in historical events, and make predictions based on causal logic.Furthermore, we propose a one-class augmented matching loss for optimization.When evaluated on ICEWS14, ICEWS0515 and ICEWS18 datasets, the performance of ALRE-IR outperforms other stateof-the-art baselines.The results also demonstrate that ALRE-IR still shows outstanding performance when transferred to related dataset with common relation vocabulary, indicating our proposed model has good zero-shot reasoning ability. 1
Xin Mei, Libin Yang, Xiaoyan Cai, Zuowei Jiang
EMNLP3
2022 Global-local neighborhood based network representation for citation recommendation
Xiaoyan Cai, Nanxin Wang, Libin Yang, Xin Mei
Appl. Intell.1
2022 SEASum: Syntax-Enriched Abstractive Summarization
Sen Liu 0004, Libin Yang, Xiaoyan Cai
Expert Syst. Appl.3
2022 Mutually reinforced network embedding: An integrated approach to research paper recommendation
Xin Mei, Xiaoyan Cai, Wenjie Li 0002, Shirui Pan, Libin Yang
Expert Syst. Appl.2
2022 Relation-aware Heterogeneous Graph Transformer based drug repurposing
Xin Mei, Xiaoyan Cai, Libin Yang, Nanxin Wang
Expert Syst. Appl.2
2022 HetTreeSum: A Heterogeneous Tree Structure-based Extractive Summarization Model for Scientific Papers
Jintao Zhao, Libin Yang, Xiaoyan Cai
Expert Syst. Appl.3
2022 COVIDSum: A linguistically enriched SciBERT-based summarization model for COVID-19 scientific papers
Xiaoyan Cai, Sen Liu 0004, Libin Yang, Jintao Zhao, Dinggang Shen, Tianming Liu 0001
J. Biomed. Informatics1
2022 StarSum: A Star Architecture Based Model for Extractive Summarization
abstract
Extractive summarization aims to produce a concise summary while retaining the key information through the way of selecting sentences from the original document. Under such background, learning inter-sentence relations has hitherto been the issue of most concern. In this study, we propose a Star architecture based model for extractive summarization (StarSum), that takes advantage of self-attention strategy based Transformer and star-shaped structure, models sentences within a document as satellite nodes and introduces a virtual star node, constructs a star model for each document to learn inter-sentence relations. Based on the constructed star-shaped model, we further develop two sentence representation learning algorithms, namely star guiding satellite (SGS) algorithm and star incorporating satellite (SIS) algorithm, in order to extract summary-worthy sentences. Experimental results on CNN/Daily Mail, New York Times (NYT) and XSum datasets prove that StarSum model achieves advanced performance for extractive summarization and has comparable performance to the state-of-the-art extractive summarization model. The results also demonstrate that the SIS algorithm is more effective than the SGS algorithm.
Kaile Shi, Xiaoyan Cai, Libin Yang, Jintao Zhao, Shirui Pan
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Graph transformer networks based text representation
Xin Mei, Xiaoyan Cai, Libin Yang, Nanxin Wang
Neurocomputing2
2021 HITS-based attentional neural model for abstractive summarization
Xiaoyan Cai, Kaile Shi, Yuehan Jiang, Libin Yang, Sen Liu 0004
Knowl. Based Syst.1
2020 Measuring distance-based semantic similarity using meronymy and hyponymy relations
Yuanyuan Cai, Shirui Pan, Ximeng Wang, Hongshu Chen, Xiaoyan Cai
Neural Comput. Appl.5
2019 Regularizing Output Distribution of Abstractive Chinese Social Media Text Summarization for Improved Semantic Consistency
abstract
Abstractive text summarization is a highly difficult problem, and the sequence-to-sequence model has shown success in improving the performance on the task. However, the generated summaries are often inconsistent with the source content in semantics. In such cases, when generating summaries, the model selects semantically unrelated words with respect to the source content as the most probable output. The problem can be attributed to heuristically constructed training data, where summaries can be unrelated to the source content, thus containing semantically unrelated words and spurious word correspondence. In this article, we propose a regularization approach for the sequence-to-sequence model and make use of what the model has learned to regularize the learning objective to alleviate the effect of the problem. In addition, we propose a practical human evaluation method to address the problem that the existing automatic evaluation method does not evaluate the semantic consistency with the source content properly. Experimental results demonstrate the effectiveness of the proposed approach, which outperforms almost all the existing models. Especially, the proposed approach improves the semantic consistency by 4% in terms of human evaluation.
Bingzhen Wei, Xuancheng Ren, Yi Zhang 0050, Xiaoyan Cai, Qi Su 0001, Xu Sun 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2018 Generative Adversarial Network Based Heterogeneous Bibliographic Network Representation for Personalized Citation Recommendation
abstract
Network representation has been recently exploited for many applications, such as citation recommendation, multi-label classification and link prediction. It learns low-dimensional vector representation for each vertex in networks. Existing network representation methods only focus on incomplete aspects of vertex information (i.e., vertex content, network structure or partial integration), moreover they are commonly designed for homogeneous information networks where all the vertices of a network are of the same type. In this paper, we propose a deep network representation model that integrates network structure and the vertex content information into a unified framework by exploiting generative adversarial network, and represents different types of vertices in the heterogeneous network in a continuous and common vector space. Based on the proposed model, we can obtain heterogeneous bibliographic network representation for efficient citation recommendation. The proposed model also makes personalized citation recommendation possible, which is a new issue that a few papers addressed in the past. When evaluated on the AAN and DBLP datasets, the performance of the proposed heterogeneous bibliographic network based citation recommendation approach is comparable with that of the other network representation based citation recommendation approaches. The results also demonstrate that the personalized citation recommendation approach is more effective than the non-personalized citation recommendation approach.
Xiaoyan Cai, Junwei Han 0001, Libin Yang
AAAI1
2018 A Skeleton-Based Model for Promoting Coherence Among Sentences in Narrative Story Generation
abstract
Narrative story generation is a challenging problem because it demands the generated sentences with tight semantic connections, which has not been well studied by most existing generative models.To address this problem, we propose a skeleton-based model to promote the coherence of generated stories.Different from traditional models that generate a complete sentence at a stroke, the proposed model first generates the most critical phrases, called skeleton, and then expands the skeleton to a complete and fluent sentence.The skeleton is not manually defined, but learned by a reinforcement learning method.Compared to the state-of-the-art models, our skeleton-based model can generate significantly more coherent text according to human evaluation and automatic evaluation.The G-score is improved by 20.1% in human evaluation. 1
Jingjing Xu 0001, Xuancheng Ren, Yi Zhang 0050, Qi Zeng 0001, Xiaoyan Cai, Xu Sun 0001
EMNLP5
2018 A Novel Personalized Citation Recommendation Approach Based on GAN
Libin Yang, Xiaoyan Cai, Hang Dai
ISMIS3
2018 A Three-Layered Mutually Reinforced Model for Personalized Citation Recommendation
abstract
Fast-growing scientific papers pose the problem of rapidly and accurately finding a list of reference papers for a given manuscript. Citation recommendation is an indispensable technique to overcome this obstacle. In this paper, we propose a citation recommendation approach via mutual reinforcement on a three-layered graph, in which each paper, author or venue is represented as a vertex in the paper layer, author layer, and venue layer, respectively. For personalized recommendation, we initiate the random walk separately for each query researcher. However, this has a high computational complexity due to the large graph size. To solve this problem, we apply a three-layered interactive clustering approach to cluster related vertices in the graph. Personalized citation recommendations are then made on the subgraph, generated by the clusters associated with each researcher's needs. When evaluated on the ACL anthology network, DBLP, and CiteSeer ML data sets, the performance of our proposed model-based citation recommendation approach is comparable with that of other state-of-the-art citation recommendation approaches. The results also demonstrate that the personalized recommendation approach is more effective than the nonpersonalized recommendation approach.
Xiaoyan Cai, Junwei Han 0001, Wenjie Li 0002, Renxian Zhang, Shirui Pan, Libin Yang
IEEE Trans. Neural Networks Learn. Syst.1
2017 Transfer Deep Learning for Low-Resource Chinese Word Segmentation with a Novel Neural Network
Jingjing Xu 0001, Shuming Ma, Yi Zhang 0050, Bingzhen Wei, Xiaoyan Cai, Xu Sun 0001
NLPCC5
2016 Classifying networked text data with positive and unlabeled examples
Shirui Pan, Yang Zhang 0010, Xiaoyan Cai
Pattern Recognit. Lett.4
2014 Enhancing diversity and coverage of document summaries through subspace clustering and clustering-based optimization
Xiaoyan Cai, Wenjie Li 0002, Renxian Zhang
Inf. Sci.1
2014 Enhancing sentence-level clustering with ranking-based clustering framework for theme-based summarization
Libin Yang, Xiaoyan Cai, Yang Zhang 0010, Peng Shi 0001
Inf. Sci.2
2014 Sequential Summarization: A Full View of Twitter Trending Topics
abstract
As an information delivering platform, Twitter collects millions of tweets every day. However, some users, especially new users, often find it difficult to understand trending topics in Twitter when confronting the overwhelming and unorganized tweets. Existing work has attempted to provide a short snippet to explain a topic, but this only provides limited benefits and cannot satisfy the users' expectations. In this paper, we propose a new summarization task, namely sequential summarization, which aims to provide a serial of chronologically ordered short sub-summaries for a trending topic in order to provide a complete story about the development of the topic while retaining the order of information presentation. Different from the traditional summarization task, the numbers of sub-summaries for different topics are not fixed. Two approaches, i.e., stream-based and semantic-based approaches, are developed to detect the important subtopics within a trending topic. Then a short sub-summary is generated for each subtopic. In addition, we propose three new measures to evaluate the position-aware coverage, sequential novelty and sequence correlation of the system-generated summaries. The experimental results based on the proposed evaluation criteria have demonstrated the effectiveness of the proposed approaches.
Dehong Gao, Wenjie Li 0002, Xiaoyan Cai, Renxian Zhang, Ouyang You
IEEE ACM Trans. Audio Speech Lang. Process.3
2013 Ranking Through Clustering: An Integrated Approach to Multi-Document Summarization
abstract
Multi-document summarization aims to create a condensed summary while retaining the main characteristics of the original set of documents. Under such background, sentence ranking has hitherto been the issue of most concern. Since documents often cover a number of topic themes with each theme represented by a cluster of highly related sentences, sentence clustering has been explored in the literature in order to provide more informative summaries. For each topic theme, the rank of terms conditional on this topic theme should be very distinct, and quite different from the rank of terms in other topic themes. Existing cluster-based summarization approaches apply clustering and ranking in isolation, which leads to incomplete, or sometimes rather biased, analytical results. A newly emerged framework uses sentence clustering results to improve or refine the sentence ranking results. Under this framework, we propose a novel approach that directly generates clusters integrated with ranking in this paper. The basic idea of the approach is that ranking distribution of sentences in each cluster should be quite different from each other, which may serve as features of clusters and new clustering measures of sentences can be calculated accordingly. Meanwhile, better clustering results can achieve better ranking results. As a result, ranking and clustering by mutually and simultaneously updating each other so that the performance of both can be improved. The effectiveness of the proposed approach is demonstrated by both the cluster quality analysis and the summarization evaluation conducted on the DUC 2004-2007 datasets.
Xiaoyan Cai, Wenjie Li 0002
IEEE Trans. Speech Audio Process.1
2012 Mutually Reinforced Manifold-Ranking Based Relevance Propagation Model for Query-Focused Multi-Document Summarization
abstract
Manifold-ranking has been recently exploited for query-focused summarization. It propagates query relevance from the given query to the document sentences by making use of both the relationships among the sentences and the relationships between the given query and the sentences. The sentences in a document set can be grouped into several topic themes with each theme represented by a cluster of highly related sentences. However, it is a well-recognized fact that a document set often covers a number of such topic themes. In this paper, we present a novel model to enhance manifold-ranking based relevance propagation via mutual reinforcement between sentences and theme clusters. Based on the proposed model, we develop two new sentence ranking algorithms, namely the reinforcement after relevance propagation (RARP) algorithm and the reinforcement during relevance propagation (RDRP) algorithm. The convergence issues of the two algorithms are examined. When evaluated on the DUC2005-2007 datasets and TAC2008 dataset, the performance of the two proposed algorithms is comparable with that of the top three systems. The results also demonstrate that the RDRP algorithm is more effective than the RARP algorithm.
Xiaoyan Cai, Wenjie Li 0002
IEEE Trans. Speech Audio Process.1
2011 Structure Preserving Mesh Parameterization
abstract
The traditional parameterization methods focused on preserving the local geometry properties by minimizing the distortions of angle and stretch. In this paper, we present a mesh parameterization method to preserve some global geometry properties, such as symmetry structure, of the input mesh. Given an input triangular mesh, the symmetry regions can be automatically detected using shape analysis method or manually specified by the users. For each vertex in the symmetry regions, we can identify the symmetric point and represent it as a triangle index and the corresponding bary centric coordinates. We extend the harmonic map to parameterize the input mesh by adding new error metric to measure how the symmetry property of the point pairs are preserved. Since the new error metric is non-linear, we develop an iterative update method to solve the parameterization problem. At last, we show some structure preservation parameterization results, and compare them with the results of the traditional harmonic map.
Zizhao Wu, Xiaoyan Cai, Xinguo Liu
CAD/Graphics2
2011 Simultaneous Clustering and Noise Detection for Theme-based Summarization
Xiaoyan Cai, Renxian Zhang, Dehong Gao, Wenjie Li 0002
IJCNLP1
2011 Requirement-Based Query and Update Scheduling in Real-Time Data Warehouses
Fangling Leng, Yubin Bao, Ge Yu 0001, Jingang Shi, Xiaoyan Cai
WAIM5
2011 A spectral analysis approach to document summarization: Clustering and ranking sentences simultaneously
Xiaoyan Cai, Wenjie Li 0002
Inf. Sci.1
2011 Enhancing sentence-level clustering with integrated and interactive frameworks for theme-based summarization
abstract
Abstract Sentence clustering plays a pivotal role in theme‐based summarization, which discovers topic themes defined as the clusters of highly related sentences to avoid redundancy and cover more diverse information. As the length of sentences is short and the content it contains is limited, the bag‐of‐words cosine similarity traditionally used for document clustering is no longer suitable. Special treatment for measuring sentence similarity is necessary. In this article, we study the sentence‐level clustering problem. After exploiting concept‐ and context‐enriched sentence vector representations, we develop two co‐clustering frameworks to enhance sentence‐level clustering for theme‐based summarization—integrated clustering and interactive clustering—both allowing word and document to play an explicit role in sentence clustering as independent text objects rather than using word or concept as features of a sentence in a document set. In each framework, we experiment with two‐level co‐clustering (i.e., sentence‐word co‐clustering or sentence‐document co‐clustering) and three‐level co‐clustering (i.e., document‐sentence‐word co‐clustering). Compared against concept‐ and context‐oriented sentence‐representation reformation, co‐clustering shows a clear advantage in both intrinsic clustering quality evaluation and extrinsic summarization evaluation conducted on the Document Understanding Conferences (DUC) datasets.
Xiaoyan Cai, Wenjie Li 0002
J. Assoc. Inf. Sci. Technol.1
2010 Simultaneous Ranking and Clustering of Sentences: A Reinforcement Approach to Multi-Document Summarization
Xiaoyan Cai, Wenjie Li 0002, Ouyang You
COLING1
2010 A Context-Sensitive Manifold Ranking Approach to Query-Focused Multi-document Summarization
Xiaoyan Cai, Wenjie Li 0002
PRICAI1