Ruifan Li

dblp:60/5628 · DBLP profile ↗
← Back
47ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0002-3543-6272ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 4 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Diffusion-Assisted Progressive Learning for Weakly Supervised Phrase Localization
abstract
Weakly supervised phrase localization (WSPL) aims to localize visual objects mentioned by given phrases, but it learns without human-annotated bounding boxes. Previous works struggle in multi-object scenarios where objects in the background often appear simultaneously with the target objects. To this end, we propose a Diffusion-Assisted PrOgressive learning framework (i.e., DAPO) for WSPL task in this paper. Specifically, we score the difficulty of training samples based on the quantity of objects and the level of semantic alignment. These samples are then used progressively during training, in an order by their difficulty scores. To address the sample imbalance problem, we propose a Generation-Assisted Tuning (GAT) method for the grounding network. First, to enrich the samples from few-object scenarios, we leverage Stable Diffusion (SD) to generate images with phrases. Second, we introduce an attention-driven scheme to direct SD's attention on the mentioned objects. Finally, we design a diffusion-guided loss, which helps the grounding network learn the objects' layouts. Extensive experiments show that our DAPO framework outperforms the strong baselines on benchmark datasets.
Pengyue Lin, Yanyang Hu, Xinjing Liu, Wenqi Jia 0006, Fangxiang Feng, Ruifan Li
AAAI6
2026 OX-MABSR: A Benchmark for Open-domain Explainable Multimodal Aspect-Based Sentiment Reasoning
abstract
Multimodal Aspect-Based Sentiment Analysis (MABSA) involves extracting aspect terms from text-image pairs and identifying their sentiments. Most existing tasks consider one fixed sentiment category with explicitly mentioned aspects. However, these tasks seldom consider expressive sentiment categories, implicit aspects, and explainability. To this end, we introduce a novel task of Open-domain Explainable Multimodal Aspect-Based Sentiment Reasoning (OX-MABSR). This task enables the prediction of open-vocabulary aspect-sentiment pairs, together with the generation of sentiment explanations and reasoning paths. To benchmark OX-MABSR task, we construct OX-MABSR-Bench, a dataset annotated with explicit and implicit aspects, expressive sentiment categories, as well as perceptual and cognitive two-level explanations. The explanations capture visual and textual cues, including aesthetics, facial expressions, scenes, and textual semantics, together with background and situational knowledge. In addition, we annotate the reasoning paths that trace how the sentiment evolves from surface cues to a deeper contextual understanding. To address OX-MABSR task, we propose MABSR-LLM. Extensive experimental results show our MABSR-LLM outperforms strong baselines. To the best of our knowledge, we are the first to provide a unified framework for open-domain and explainable MABSR.
Xinjing Liu, Zixin Xue, Pengyue Lin, Xinyu Tu, Siwei Xu, Ruifan Li
AAAI6
2026 MET-LLM: Enhancing large language models for malicious encrypted traffic detection
abstract
Modern networks have spurred growth in both legitimate and malicious activities concealed within encrypted traffic. Traditional machine learning approaches to traffic classification struggle with scalability to new protocols, diverse tasks, and adaptability to emerging threats. To address these issues, we propose MET-LLM , a novel framework for Malicious Encrypted Traffic detection that integrates domain-specific tokenization, a pretrained large language model, and a dynamic adaptive tuning adaptor. MET-LLM addresses the modal gap between natural language and heterogeneous network traffic data by partitioning each traffic sample into distinct headers and payloads and leveraging a specialized tokenizer trained on a large-scale traffic corpus to extend the base vocabulary of the underlying language model. Building on a domain-adapted pretrained model fine-tuned on extensive security-related corpora, MET-LLM captures critical contextual nuances distinguishing benign from malicious flows. Its dynamic adaptive tuning adaptor facilitates efficient parameter updates via adaptation prompt injection, adversarial training, and dynamic masking, enabling rapid adaptation to evolving network conditions and attack strategies. Extensive evaluations on benchmark datasets, including ISCX Tor 2016, ISCX VPN 2016, APP-53 2023, and CSTNET 2023, demonstrate that MET-LLM’s superior precision, recall, and F1 scores over state-of-the-art methods, affirming its efficacy and robustness in real-world cybersecurity applications. Our code is publicly available at the website, https://github.com/Superagentsys/MET-LLM .
Yongjun Huang, Ruifan Li, Xiaoyong Li 0003, Lixiang Li 0001
Expert Syst. Appl.3
2026 Actively masked reasoning over knowledge graphs for fault diagnosis under insufficient information condition
Runhai Jiao, Changyu Zhou, Boxu Yan, Ruifan Li
Expert Syst. Appl.6
2026 Towards balancing the efficiency and effectiveness: a unified edit-based framework for automatic image captioning
Ruifan Li, Siwei Xu, Pengyue Lin, Fangxiang Feng, Zhangyu Ma
Multim. Syst.1
2026 A fine-grained entity understanding network for weakly supervised phrase grounding
Pengyue Lin, Ruifan Li, Fangxiang Feng, Lun Ke, Zhanyu Ma, Xiaojie Wang 0006
Pattern Recognit.2
2025 Concept Conductor: Orchestrating Multiple Personalized Concepts in Text-to-Image Synthesis
abstract
The customization of text-to-image models has seen significant advancements, yet generating multiple personalized concepts remains a challenging task. Current methods struggle with attribute leakage and layout confusion when handling multiple concepts, leading to reduced concept fidelity and semantic consistency. In this work, we introduce a novel training-free framework, Concept Conductor, designed to ensure visual fidelity and correct layout in multi-concept customization. Concept Conductor isolates the sampling processes of multiple customized models to prevent attribute leakage between different concepts and corrects erroneous layouts through self-attention-based spatial guidance. Additionally, we present a concept injection technique that employs shape-aware masks to specify the generation area for each concept. This technique injects the structure and appearance of personalized concepts through feature fusion in the attention layers, ensuring harmony in the final image. Extensive qualitative and quantitative experiments demonstrate that Concept Conductor can consistently generate composite images with accurate layouts while preserving the visual details of each concept. Compared to existing baselines, Concept Conductor shows significant performance improvements. Our method supports the combination of any number of concepts and maintains high fidelity even when dealing with visually similar concepts. The code and trained models will be made publicly available.
Zebin Yao, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
AAAI3
2025 Multimodal Aspect-Based Sentiment Analysis under Conditional Relation
abstract
Multimodal Aspect-Based Sentiment Analysis (MABSA) aims to extract aspect terms from text-image pairs and identify their sentiments. Previous methods are based on the premise that the image contains the objects referred by the aspects within the text. However, this condition cannot always be met, resulting in a suboptimal performance. In this paper, we propose COnditional Relation based Sentiment Analysis framework (CORSA). Specifically, we design a conditional relation detector (CRD) to mitigate the impact of the unmet conditional image. Moreover, we design a visual object localizer (VOL) to locate the exact condition-related visual regions associated with the aspects. With CRD and VOL, our CORSA framework takes a multi-task form. In addition, to effectively learn CORSA we conduct two types of annotations. One is the conditional relation using a pretrained referring expression comprehension model; the other is the bounding boxes of visual objects by a pretrained object detection model. Experiments on our built C-MABSA dataset show that CORSA consistently outperforms existing methods. The code and data are available at https://github.com/Liuxj-Anya/CORSA.
Xinjing Liu, Ruifan Li, Shuqin Ye, Guangwei Zhang 0003, Xiaojie Wang 0006
COLING2
2025 A Weighted Cross-entropy Loss for Mitigating LLM Hallucinations in Cross-lingual Continual Pretraining
abstract
Recently, due to the explosive advances of large language models (LLMs) on English, cross-lingual continual pretraining has been widely applied in obtaining Chinese LLMs. However, previous studies showed that these LLMs have suffered severe hallucinations, mainly caused by noisy tokens. To this aim, we propose a novel loss function, InfoLoss for continual pretraining. Specifically, our loss function takes into account the co-occurrence of noisy and normal tokens, and uses point-wise mutual information to reduce the impact of noisy tokens. We use InfoLoss to continually pretrain 30 billion tokens on Llama 2-7B with 64 A100 GPUs for 24 days, obtaining C-Llama. We then conduct experiments on 12 benchmarks for evaluations. The results show the effectiveness of our proposed InfoLoss. Our datasets and codes are publicly available at https://github.com/Fluxation996/C-Llama.
Yuantao Fan, Ruifan Li, Guangwei Zhang 0003, Xiaojie Wang 0006
ICASSP2
2025 Sentiment-aware Dual-flow Graph Convolutional Networks for Aspect Based Sentimental Analysis
abstract
Recent studies on textual Aspect-Based Sentiment Analysis have utilized graph neural networks with dependency trees and self-attention respectively, to learn the information of syntactic structures and semantic correlations of the given sentence. However, these methods failed to address issues of syntactic insufficiency caused by the dependency parsing capability and imperfect attention weights calculated through self-attention with hidden representations. In this paper, we propose a sentiment-aware dual-flow graph convolutional networks (SeDuFo) model using commonsense knowledge. To be specific, in KSyn-Flow, to encourage the importance of words conveying sentiment, sentiment values from SenticNet and part of speech, are used to enhance the matrix constructed from the dependency tree. Mean-while, in KSem-Flow, to better capture semantic correlations between words, word embeddings trained from ConceptNet are leveraged to enhance the representations with more structured prior knowledge. Experimental results on three public datasets verify the effectiveness of our proposed model, which outperforms the state-of-the-art methods.
Zixin Xue, Ruifan Li
IJCNN2
2025 SDG-MLLM: Injecting Structured Dialogue Graphs into MLLM for Multimodal Conversational Aspect-Based Sentiment Analysis
abstract
Multimodal Conversational Aspect-based Sentiment Analysis (MCA BSA) is a challenging task for multimodal dialogue understanding. Existing works often treat the entire dialogue as a flat sequence and feed it into Large Language Models (LLMs) for pipeline-style generation. However, these methods sometimes accumulate errors and overlook critical discourse structure and fine-grained inter-word relations that are essential for accurate sentiment reasoning. To address these limitations, we propose SDG-MLLM, a unified generative framework that integrates Structured Dialogue Graphs into Multimodal LLM (MLLM) for an end-to-end MCABSA. Specifically, we construct heterogeneous dialogue graphs that capture diverse structural relations, including syntactic dependencies, coreference links, speaker turns, reply flow, semantic role labeling, and sentiment propagation paths. These graphs are encoded using a heterogeneous dialogue graph encoder, and the resulting structure-aware graph features are injected into the embedding layer of LLM. Furthermore, SDG-MLLM incorporates aligned multimodal features such as image, audio, and video cues at the utterance level to enable unified and context-aware multimodal reasoning. Experiments on the MCABSA dataset show that SDG-MLLM significantly outperforms strong baselines across multiple tasks. In addition, our method also achieved top performance in the ACM MM 2025 Grand Challenge of MCABSA. Our code is available at https://github.com/Liuxj-Anya/SDG-MLLM.
Xinjing Liu, Pengyue Lin, Xinyu Tu, Wenqi Jia 0006, Ruifan Li
ACM Multimedia6
2025 MRF: A Modality-Resilient Framework for Handling Missing Modalities in Multimodal Learning
Yongjun Huang, Ruifan Li
NLPCC (2)4
2024 Revisiting Counterfactual Problems in Referring Expression Comprehension
abstract
Traditional referring expression comprehension (REC) aims to locate the target referent in an image guided by a text query. Several previous methods have studied on the Counterfactual problem in REC (C-REC) where the objects for a given query cannot be found in the image. However, these methods focus on the overall image-text or specific attribute mismatch only. In this paper, we address the C-REC problem from a deep perspective of fine- grained attributes. To this aim, we first propose a fine-grained counterfactual sample generation method to construct C-REC datasets. Specifically, we leverage pre-trained language model such as BERT to modify the attribute words in the queries, obtaining the corresponding counterfactual samples. Furthermore, we propose a C-REC framework. We first adopt three encoders to extract image, text and attribute features. Then, our dual-branch attentive fusion module fuses these cross-modal features with two branches by an attention mechanism. At last, two prediction heads generate a bounding box and a counterfactual label, respectively. In addition, we incorporate contrastive learning with the generated counterfactual samples as negatives to enhance the counterfactual perception. Extensive experiments show that our framework achieves promising performance on both public REC datasets RefCOCO/+lg and our constructed C-REC datasets C-RefCOCO/+lg. The code and data are available at https://github.com/Glacier0012/CREC.
Zhihan Yu, Ruifan Li
CVPR2
2024 Visual Prompt Tuning for Weakly Supervised Phrase Grounding
abstract
Previous works on the task of weakly supervised phrase grounding (WSG) rely heavily on object detectors providing RoIs for the localization. However, such methods cannot be applied effectively to real-world scenarios largely because that the detectors are trained with limited categories. In this paper, we propose a refinement-based approach to WSG through fine-tuning a detector-free phrase grounding model with a visual prompt. This visual prompt is extracted from the text-related representations in CLIP. Furthermore, we combine the visual prompt with learnable features and then fine-tune the grounding network. Our experimental results significantly outperform state-of-the-art methods on the WSG task and shows the effectiveness of our method.
Pengyue Lin, Zhihan Yu, Mingcong Lu, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
ICASSP5
2024 Triple Alignment Strategies for Zero-shot Phrase Grounding under Weak Supervision
abstract
Phrase Grounding, i.e., PG aims to locate objects referred by noun phrases. Recently, PG under weak supervision (i.e., grounding without region-level annotations) and zero-shot PG (i.e., grounding from seen categories to unseen ones) are proposed, respectively. However, for real-world applications these two approaches are limited due to slight annotations and numerable categories during training. In this paper, we propose a framework of zero-shot PG under weak supervision. Specifically, our PG framework is built on triple alignment strategies. Firstly, we propose a region-text alignment (RTA) strategy to build region-level attribute associations via CLIP. Secondly, we propose a domain alignment (DomA) strategy by minimizing the difference between distributions of seen classes in the training and those of the pre-training. Thirdly, we propose a category alignment (CatA) strategy by considering both category semantics and region-category relations. Extensive experimental results show that our proposed PG framework outperforms previous zero-shot methods and weakly-supervised methods. Our code is available at https://github.com/LinPengyue/ZS-WSG.
Pengyue Lin, Ruifan Li, Yuzhe Ji, Zhihan Yu, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006
ACM Multimedia2
2024 DiffHarmony++: Enhancing Image Harmonization with Harmony-VAE and Inverse Harmonization Model
abstract
Latent diffusion model has demonstrated impressive efficacy in image generation and editing tasks. Recently, it has also promoted the advancement of image harmonization. However, methods involving latent diffusion model all face a common challenge: the severe image distortion introduced by the VAE component, while image harmonization is a low-level image processing task that relies on pixel-level evaluation metrics. In this paper, we propose Harmony-VAE, leveraging the input of the harmonization task itself to enhance the quality of decoded images. The input involving composite image contains the precise pixel level information, which can complement the correct foreground appearance and color information contained in denoised latents. Meanwhile, the inherent generative nature of diffusion models makes it naturally adapt to inverse image harmonization, i.e. generating synthetic composite images based on real images and foreground masks. We train an inverse harmonization diffusion model to perform data augmentation on two subsets of iHarmony4 and construct a new human harmonization dataset with prominent foreground objects. Extensive experiments demonstrate the effectiveness of our proposed Harmony-VAE and inverse harmonization model. Code and pretrained models are available at https://github.com/nicecv/DiffHarmony.
Fangxiang Feng, Guang Liu 0006, Ruifan Li, Xiaojie Wang 0006
ACM Multimedia4
2024 LGR-NET: Language Guided Reasoning Network for Referring Expression Comprehension
abstract
Referring Expression Comprehension(REC) is a fundamental task in the vision and language domain, which aims to locate an image region according to a natural language expression. REC requires the models to capture key clues in the text and perform accurate cross-modal reasoning. A recent trend employs transformer-based methods to address this problem. However, most of these methods typically treat image and text equally. They usually perform cross-modal reasoning in a crude way, and utilize textual features as a whole without detailed considerations (e.g., spatial information). This insufficient utilization of textual features will lead to sub-optimal results. In this paper, we propose aLanguage Guided Reasoning Network(LGR-NET) to fully utilize the guidance of the referring expression. To localize the referred object, we set a prediction token to capture cross-modal features. Furthermore, to sufficiently utilize the textual features, we extend them by our Textual Feature Extender (TFE) from three aspects.First, we design a novel coordinate embedding based on textual features. The coordinate embedding is incorporated to the prediction token to promote its capture of language-related visual features.Second, we employ the extracted textual features for Text-guided Cross-modal Alignment (TCA) and Fusion (TCF), alternately.Third, we devise a novel cross-modal loss to enhance cross-modal alignment between the referring expression and the learnable prediction token. We conduct extensive experiments on five benchmark datasets, and the experimental results show that our LGR-NET achieves a new state-of-the-art. Source code is available at https://github.com/lmc8133/LGR-NET.
Mingcong Lu, Ruifan Li, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006
IEEE Trans. Circuits Syst. Video Technol.2
2024 DualGCN: Exploring Syntactic and Semantic Information for Aspect-Based Sentiment Analysis
abstract
The task of aspect-based sentiment analysis aims to identify sentiment polarities of given aspects in a sentence. Recent advances have demonstrated the advantage of incorporating the syntactic dependency structure with graph convolutional networks (GCNs). However, their performance of these GCN-based methods largely depends on the dependency parsers, which would produce diverse parsing results for a sentence. In this article, we propose a dual GCN (DualGCN) that jointly considers the syntax structures and semantic correlations. Our DualGCN model mainly comprises four modules: 1) SynGCN: instead of explicitly encoding syntactic structure, the SynGCN module uses the dependency probability matrix as a graph structure to implicitly integrate the syntactic information; 2) SemGCN: we design the SemGCN module with multihead attention to enhance the performance of the syntactic structure with the semantic information; 3) Regularizers: we propose orthogonal and differential regularizers to precisely capture semantic correlations between words by constraining attention scores in the SemGCN module; and 4) Mutual BiAffine: we use the BiAffine module to bridge relevant information between the SynGCN and SemGCN modules. Extensive experiments are conducted compared with up-to-date pretrained language encoders on two groups of datasets, one including Restaurant14, Laptop14, and Twitter and the other including Restaurant15 and Restaurant16. The experimental results demonstrate that the parsing results of various dependency parsers affect their performance of the GCN-based models. Our DualGCN model achieves superior performance compared with the state-of-the-art approaches. The source code and preprocessed datasets are provided and publicly available on GitHub (see https://github.com/CCChenhao997/DualGCN-ABSA).
Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy
IEEE Trans. Neural Networks Learn. Syst.1
2023 USSA: A Unified Table Filling Scheme for Structured Sentiment Analysis
abstract
Most previous studies on Structured Sentiment Analysis (SSA) have cast it as a problem of bi-lexical dependency parsing, which cannot address issues of overlap and discontinuity simultaneously.In this paper, we propose a nichetargeting and effective solution.Our approach involves creating a novel bi-lexical dependency parsing graph, which is then converted to a unified 2D table-filling scheme, namely USSA.The proposed scheme resolves the kernel bottleneck of previous SSA methods by utilizing 13 different types of relations.In addition, to closely collaborate with the USSA scheme, we have developed a model that includes a proposed bi-axial attention module to effectively capture the correlations among relations in the rows and columns of the table.Extensive experimental results on benchmark datasets demonstrate the effectiveness and robustness of our proposed framework, outperforming state-ofthe-art methods consistently 1 .
Zepeng Zhai, Hao Chen 0041, Ruifan Li, Xiaojie Wang 0006
ACL (1)3
2023 Enhanced Machine Reading Comprehension Method for Aspect Sentiment Quadruplet Extraction
abstract
In the NLP domain, Aspect-Based Sentiment Analysis (ABSA) has gained significant attention in recent years due to its ability to perform fine-grained sentiment analysis. A challenging task in ABSA is Aspect Sentiment Quadruplet Extraction (ASQE), which involves the extraction of aspect terms and their associated opinion terms, sentiment polarities, and categories in the form of quadruplets. However, existing studies have ignored the strong dependence among the multiple subtasks involved in ASQE. In this paper, we propose a novel Enhanced Machine Reading Comprehension (EMRC) method and formalize ASQE task as a multi-turn MRC task. Our EMRC effectively learns and utilizes the relationships among different subtasks by incorporating previously generated query answers into the current queries. We design a hierarchical category classification strategy to perform the category prediction in a structured manner, enabling the model to tackle intricate categories with ease. Furthermore, we employ the bi-directional attention mechanism, i.e., context-to-query and query-to-context attentions, to map the context into a task-aware representation. We conduct extensive experiments on two benchmark datasets. The results demonstrate that EMRC outperforms the state-of-art baselines. The source code is publicly available at https://github.com/Little-Yeah/EMCR.
Shuqin Ye, Zepeng Zhai, Ruifan Li
ECAI3
2023 Knowledge Prompting with Contrastive Learning for Unsupervised CommonsenseQA
Lihui Zhang, Ruifan Li
ICONIP (11)2
2022 Enhanced Multi-Channel Graph Convolutional Network for Aspect Sentiment Triplet Extraction
abstract
Aspect Sentiment Triplet Extraction (ASTE) is an emerging sentiment analysis task.Most of the existing studies focus on devising a new tagging scheme that enables the model to extract the sentiment triplets in an end-to-end fashion.However, these methods ignore the relations between words for ASTE task.In this paper, we propose an Enhanced Multi-Channel Graph Convolutional Network model (EMC-GCN) to fully utilize the relations between words.Specifically, we first define ten types of relations for ASTE task, and then adopt a biaffine attention module to embed these relations as an adjacent tensor between words in a sentence.After that, our EMC-GCN transforms the sentence into a multi-channel graph by treating words and the relation adjacent tensor as nodes and edges, respectively.Thus, relationaware node representations can be learnt.Furthermore, we consider diverse linguistic features to enhance our EMC-GCN model.Finally, we design an effective refining strategy on EMC-GCN for word-pair representation refinement, which considers the implicit results of aspect and opinion extraction when determining whether word pairs match or not.Extensive experimental results on the benchmark datasets demonstrate that the effectiveness and robustness of our proposed model, which outperforms state-of-the-art methods significantly.
Hao Chen 0041, Zepeng Zhai, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
ACL (1)4
2022 A Simple Model for Distantly Supervised Relation Extraction
abstract
Distantly supervised relation extraction is challenging due to the noise within data. Recent methods focus on exploiting bag representations based on deep neural networks with complex de-noising scheme to achieve remarkable performance. In this paper, we propose a simple but effective BERT-based Graph convolutional network Model (i.e., BGM). Our BGM comprises of an instance embedding module and a bag representation module. The instance embedding module uses a BERT-based pretrained language model to extract key information from each instance. The bag representaion module constructs the corresponding bag graph then apply a convolutional operation to obtain the bag representation. Our BGM model achieves a considerable improvement on two benchmark datasets, i.e., NYT10 and GDS.
Ziqin Rao, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
COLING3
2022 COM-MRC: A COntext-Masked Machine Reading Comprehension Framework for Aspect Sentiment Triplet Extraction
abstract
Aspect Sentiment Triplet Extraction (ASTE) aims to extract sentiment triplets from sentences, which was recently formalized as an effective machine reading comprehension (MRC) based framework.However, when facing multiple aspect terms, the MRC-based methods could fail due to the interference from other aspect terms.In this paper, we propose a novel COntext-Masked MRC (COM-MRC) framework for ASTE.Our COM-MRC framework comprises three closely-related components: a context augmentation strategy, a discriminative model, and an inference method.Specifically, a context augmentation strategy is designed by enumerating all masked contexts for each aspect term.The discriminative model comprises four modules, i.e., aspect and opinion extraction modules, sentiment classification and aspect detection modules.In addition, a two-stage inference method first extracts all aspects and then identifies their opinions and sentiment through iteratively masking the aspects.Extensive experimental results on benchmark datasets show the effectiveness of our proposed COM-MRC framework, which outperforms state-of-the-art methods consistently 1 .
Zepeng Zhai, Hao Chen 0041, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
EMNLP4
2022 Improving Image Paragraph Captioning with Dual Relations
abstract
Image paragraph captioning aims to generate multiple de-scriptive sentences for an image. However, most previous methods ignore the explicit relations among objects resulting in unsatisfactory performance. In this paper, we propose a novel model (i.e., DualRel) to capture spatial and seman-tic relations among objects. Specifically, the spatial relation embedding is obtained solely from images using a predefined geometry pattern. With the help of captions, the semantic relation embedding is learned in a weakly supervised man-ner. These two relation embeddings are then interacted with regional features of objects through a relation-aware attention interaction. It first obtains a visual context vector using regional features. Then with the visual context vector, we obtain the corresponding spatial and semantic relation-aware vectors using attentions. These three vectors are fused with two gates for language decoding to further generate a para-graph. Experimental results on Stanford benchmark dataset show that DualRel achieves remarkable improvements11Code released at https://github.com/fuyunll07/DualRel.
Yihui Shi, Fangxiang Feng, Ruifan Li, Zhanyu Ma, Xiaojie Wang 0006
ICME4
2022 Modality Disentangled Discriminator for Text-to-Image Synthesis
abstract
Text-to-image (T2I) synthesis aims at generating photo-realistic images from text descriptions, which is a particularly important task in bridging vision and language. Each generated image consists of two parts: the content part related to the text and the style part irrelevant to the text. The existing discriminator does not distinguish between the content part and the style part. This not only precludes the T2I synthesis models from generating the content part effectively but also makes it difficult to manipulate the style of the generated image. In this paper, we propose a modality disentangled discriminator that distinguishes between the content part and the style part at a specific layer. Specifically, we enforce the early layers of a certain number in the discriminator to become the disentangled representation extractor through two losses. The extracted common representation for the content part can make the discriminator more effective for capturing the text-image correlation, while the extracted modality-specific representation for the style part can be directly transferred to other images. The combination of these two representations can also improve the quality of the generated images. Our proposed discriminator is used to substitute the discriminator of each stage in the representative model AttnGAN and the SOTA model DM-GAN. Extensive experiments are conducted on three widely used datasets, i.e. CUB, Oxford-102, and COCO, for the T2I synthesis task, demonstrating the superior performance of the modality disentangled discriminator over the base models. Code for DM-GAN with our modality disentangled discriminator is available athttps://github.com/FangxiangFeng/DM-GAN-MDD.
Fangxiang Feng, Tianrui Niu, Ruifan Li, Xiaojie Wang 0006
IEEE Trans. Multim.3
2021 Dual Graph Convolutional Networks for Aspect-based Sentiment Analysis
abstract
Ruifan Li, Hao Chen, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang, Eduard Hovy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy
ACL/IJCNLP (1)1
2021 S2TD: A Tree-Structured Decoder for Image Paragraph Captioning
abstract
Image paragraph captioning, a task to generate the paragraph description for a given image, usually requires mining and organizing linguistic counterparts from abundant visual clues. Limited by sequential decoding perspective, previous methods have difficulty in organizing the visual clues holistically or capturing the structural nature of linguistic descriptions. In this paper, we propose a novel tree-structured visual paragraph decoder network, called Splitting to Tree Decoder (S2TD) to address this problem. The key idea is to model the paragraph decoding process as a top-down binary tree expansion. S2TD consists of three modules: a split module, a score module, and a word-level RNN. The split module iteratively splits ancestral visual representations into two parts through a gating mechanism. To determine the tree topology, the score module uses cosine similarity to evaluate the nodes splitting. A novel tree structure loss is proposed to enable end-to-end learning. After the tree expansion, the word-level RNN decodes leaf nodes into sentences forming a coherent paragraph. Extensive experiments are conducted on the Stanford benchmark dataset. The experimental results show promising performance of our proposed S2TD.
Yihui Shi, Fangxiang Feng, Ruifan Li, Zhanyu Ma, Xiaojie Wang 0006
MMAsia4
2020 Multi-scale Two-way Deep Neural Network for Stock Trend Prediction
abstract
Stock Trend Prediction(STP) has drawn wide attention from various fields, especially Artificial Intelligence. Most previous studies are single-scale oriented which results in information loss from a multi-scale perspective. In fact, multi-scale behavior is vital for making intelligent investment decisions. A mature investor will thoroughly investigate the state of a stock market at various time scales. To automatically learn the multi-scale information in stock data, we propose a Multi-scale Two-way Deep Neural Network. It learns multi-scale patterns from two types of scale-information, wavelet-based and downsampling-based, by eXtreme Gradient Boosting and Recurrent Convolutional Neural Network, respectively. After combining the learned patterns from the two-way, our model achieves state-of-the-art performance on FI-2010 and CSI-2016, where the latter is our published long-range stock dataset to help future studies for STP task. Extensive experimental results on the two datasets indicate that multi-scale information can significantly improve the STP performance and our model is superior in capturing such information.
Guang Liu 0007, Yuzhao Mao, Hailong Huang 0003, Weiguo Gao, Jianping Shen, Ruifan Li, Xiaojie Wang 0006
IJCAI8
2020 Learning Visual Features from Product Title for Image Retrieval
abstract
There is a huge market demand for searching for products by images in e-commerce sites. Visual features play the most important role in solving this content-based image retrieval task. Most existing methods leverage pre-trained models on other large-scale datasets with well-annotated labels, e.g. the ImageNet dataset, to extract visual features. However, due to the large difference between the product images and the images in ImageNet, the feature extractor trained on ImageNet is not efficient in extracting the visual features of product images. And retraining the feature extractor on the product images is faced with the dilemma of lacking the annotated labels. In this paper, we utilize the easily accessible text information, that is, the product title, as a supervised signal to learn the features of the product image. Specifically, we use the n-grams extracted from the product title as the label of the product image to construct a dataset for image classification. This dataset is then used to fine-tuned a pre-trained model. Finally, the basic max-pooling activation of convolutions (MAC) feature is extracted from the fine-tuned model. As a result, we achieve the fourth position in the Grand Challenge of AI Meets Beauty in 2020 ACM Multimedia by using only a single ResNet-50 model without any human annotations and pre-processing or post-processing tricks. Our code is available at: \urlhttps://github.com/FangxiangFeng/AI-Meets-Beauty-2020.
Fangxiang Feng, Tianrui Niu, Ruifan Li, Xiaojie Wang 0006, Huixing Jiang
ACM Multimedia3
2020 Dual-CNN: A Convolutional language decoder for paragraph image captioning
Ruifan Li, Yihui Shi, Fangxiang Feng, Xiaojie Wang 0006
Neurocomputing1
2020 Multi-negative samples with Generative Adversarial Networks for image retrieval
Ruifan Li, Xuesen Zhang, Yuzhao Mao, Xiaojie Wang 0006
Neurocomputing1
2020 Exploring Global and Local Linguistic Representations for Text-to-Image Synthesis
abstract
The task of text-to-image synthesis is to generate photographic images conditioned on given textual descriptions. This challenging task has recently attracted considerable attention from the multimedia community due to its potential applications. Most of the up-to-date approaches are built based on generative adversarial network (GAN) models, and they synthesize images conditioned on the global linguistic representation. However, the sparsity of the global representation results in training difficulties on GANs and a shortage of fine-grained information in the generated images. To address this problem, we propose cross-modal global and local linguistic representations-based generative adversarial networks (CGL-GAN) by incorporating the local linguistic representation into the GAN. In our CGL-GAN, we construct a generator to synthesize the target images and a discriminator to judge whether the generated images conform with the text description. In the discriminator, we construct the cross-modal correlation by projecting the image representations at high and low levels onto the global and local linguistic representations, respectively. We design the hinge loss function to train our CGL-GAN model. We evaluate the proposed CGL-GAN on two publicly available datasets, the CUB and the MS-COCO. The extensive experiments demonstrate that incorporating fine-grained local linguistic information with cross-modal correlation can greatly improve the performance of text-to-image synthesis, even when generating high-resolution images.
Ruifan Li, Fangxiang Feng, Guangwei Zhang 0003, Xiaojie Wang 0006
IEEE Trans. Multim.1
2019 Differential Networks for Visual Question Answering
abstract
The task of Visual Question Answering (VQA) has emerged in recent years for its potential applications. To address the VQA task, the model should fuse feature elements from both images and questions efficiently. Existing models fuse image feature element vi and question feature element qi directly, such as an element product viqi. Those solutions largely ignore the following two key points: 1) Whether vi and qi are in the same space. 2) How to reduce the observation noises in vi and qi. We argue that two differences between those two feature elements themselves, like (vi − vj) and (qi −qj), are more probably in the same space. And the difference operation would be beneficial to reduce observation noise. To achieve this, we first propose Differential Networks (DN), a novel plug-and-play module which enables differences between pair-wise feature elements. With the tool of DN, we then propose DN based Fusion (DF), a novel model for VQA task. We achieve state-of-the-art results on four publicly available datasets. Ablation studies also show the effectiveness of difference operations in DF model.
Chenfei Wu, Jinlai Liu, Xiaojie Wang 0006, Ruifan Li
AAAI4
2018 Show and Tell More: Topic-Oriented Multi-Sentence Image Captioning
abstract
Image captioning aims to generate textual descriptions for images. Most previous work generates a single-sentence description for each image. However, a picture is worth a thousand words. Single-sentence can hardly give a complete view of an image even by humans. In this paper, we propose a novel Topic-Oriented Multi-Sentence (\emph{TOMS}) captioning model, which can generate multiple topic-oriented sentences to describe an image. Different from object instances or attributes, topics mined by the latent Dirichlet allocation reflect hidden thematic structures in reference sentences of an image. In our model, each topic is integrated to a caption generator with a Fusion Gate Unit (FGU) to guide the generation of a sentence towards a certain topic perspective. With multiple sentences from different topics, our \emph{TOMS} provides a complete description of an image. Experimental results on both sentence and paragraph datasets demonstrate the effectiveness of our \emph{TOMS} in terms of topical consistency and descriptive completeness.
Yuzhao Mao, Xiaojie Wang 0006, Ruifan Li
IJCAI4
2016 Image color harmony modeling through neighbored co-occurrence colors
Peng Lu 0007, Xujun Peng, Caixia Yuan, Ruifan Li, Xiaojie Wang 0006
Neurocomputing4
2016 An EL-LDA based general color harmony model for photo aesthetics assessment
Peng Lu 0007, Xujun Peng, Xinshan Zhu, Ruifan Li
Signal Process.4
2015 Deep correspondence restricted Boltzmann machine for cross-modal retrieval
Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
Neurocomputing2
2015 Challenges in representation learning: A report on three machine learning contests
Ian J. Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Aaron C. Courville, Mehdi Mirza, Benjamin Hamner, Will Cukierski, Yichuan Tang, Dave Thaler, Yingbo Zhou 0002, Chetan Ramaiah, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006, Dimitris Athanasakis, John Shawe-Taylor, Maxim Milakov, John Park, Radu Tudor Ionescu, Marius Popescu, Cristian Grozea, James Bergstra, Jingjing Xie, Lukasz Romaszko, Yoshua Bengio
Neural Networks14
2015 A quantum secure direct communication protocol based on four-qubit cluster state
abstract
We propose a quantum secure direct communication protocol utilizing four-qubit cluster state to enhance the efficiency of eavesdropping detection. In the security analysis, by applying the method of the entropy theory, we contrast our scheme to another two strategies, the Ping-pong protocol and the protocol using two particles of Einstein-Podolsky-Rosen pair as detection particles. The comparison results show that if the eavesdropper obtains the same amount of information, the presented quantum protocol strategy will have a larger detection probability than the other two. At last, the security of the proposed protocol is discussed. The analysis results indicate that the protocol in this paper is more secure. Copyright © 2013 John Wiley & Sons, Ltd.
DanJie Song, Ruifan Li
Secur. Commun. Networks3
2015 Towards aesthetics of image: A Bayesian framework for color harmony modeling
Peng Lu 0007, Xujun Peng, Ruifan Li, Xiaojie Wang 0006
Signal Process. Image Commun.3
2015 Finding more relevance: Propagating similarity on Markov random field for object retrieval
Peng Lu 0007, Xujun Peng, Xinshan Zhu, Ruifan Li
Signal Process. Image Commun.4
2015 Correspondence Autoencoders for Cross-Modal Retrieval
abstract
This article considers the problem of cross-modal retrieval, such as using a text query to search for images and vice-versa. Based on different autoencoders, several novel models are proposed here for solving this problem. These models are constructed by correlating hidden representations of a pair of autoencoders. A novel optimal objective, which minimizes a linear combination of the representation learning errors for each modality and the correlation learning error between hidden representations of two modalities, is used to train the model as a whole. Minimizing the correlation learning error forces the model to learn hidden representations with only common information in different modalities, while minimizing the representation learning error makes hidden representations good enough to reconstruct inputs of each modality. To balance the two kind of errors induced by representation learning and correlation learning, we set a specific parameter in our models. Furthermore, according to the modalities the models attempt to reconstruct they are divided into two groups. One group including three models is named multimodal reconstruction correspondence autoencoder since it reconstructs both modalities. The other group including two models is named unimodal reconstruction correspondence autoencoder since it reconstructs a single modality. The proposed models are evaluated on three publicly available datasets. And our experiments demonstrate that our proposed correspondence autoencoders perform significantly better than three canonical correlation analysis based models and two popular multimodal deep models on cross-modal retrieval tasks.
Fangxiang Feng, Xiaojie Wang 0006, Ruifan Li
ACM Trans. Multim. Comput. Commun. Appl.3
2014 Discovering Harmony: A Hierarchical Colour Harmony Model for Aesthetics Assessment
Peng Lu 0007, Zhijie Kuang, Xujun Peng, Ruifan Li
ACCV (3)4
2014 Sentiment analysis of microblog combining dictionary and rules
abstract
Microblog has become a daily communication tool in recent years. Researches on microblog have drawn more and more attention. Microblogging emotional classification is a major research of user intent analysis based on User-Generated Content (UGC). This paper focuses on the discrimination on two emotional tendencies: positive and negative. Firstly, the system cleared the noisy elements in the microblog, then extracted the features of the microblog and finally classified the microblog using Support Vector Machine (SVM). Furthermore, we improve the algorithms of feature extraction and weight computing combining dictionary approach and rule based approach. The result of experiment shows that the method is effective.
Yanquan Zhou, Ruifan Li, Peng Lu 0007
ASONAM3
2014 Cross-modal Retrieval with Correspondence Autoencoder
abstract
The problem of cross-modal retrieval, e.g., using a text query to search for images and vice-versa, is considered in this paper. A novel model involving correspondence autoencoder (Corr-AE) is proposed here for solving this problem. The model is constructed by correlating hidden representations of two uni-modal autoencoders. A novel optimal objective, which minimizes a linear combination of representation learning errors for each modality and correlation learning error between hidden representations of two modalities, is used to train the model as a whole. Minimization of correlation learning error forces the model to learn hidden representations with only common information in different modalities, while minimization of representation learning error makes hidden representations are good enough to reconstruct input of each modality. A parameter $\alpha$ is used to balance the representation learning error and the correlation learning error. Based on two different multi-modal autoencoders, Corr-AE is extended to other two correspondence models, here we called Corr-Cross-AE and Corr-Full-AE. The proposed models are evaluated on three publicly available data sets from real scenes. We demonstrate that the three correspondence autoencoders perform significantly better than three canonical correlation analysis based models and two popular multi-modal deep models on cross-modal retrieval tasks.
Fangxiang Feng, Xiaojie Wang 0006, Ruifan Li
ACM Multimedia3
2013 Challenges in Representation Learning: A Report on Three Machine Learning Contests
Ian J. Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Aaron C. Courville, Mehdi Mirza, Benjamin Hamner, Will Cukierski, Yichuan Tang, Dave Thaler, Yingbo Zhou 0002, Chetan Ramaiah, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006, Dimitris Athanasakis, John Shawe-Taylor, Maxim Milakov, John Park, Radu Tudor Ionescu, Marius Popescu, Cristian Grozea, James Bergstra, Jingjing Xie, Lukasz Romaszko, Yoshua Bengio
ICONIP (3)14