Fanglong Yao

dblp:274/1097 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0003-4187-9755ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Language Constrained Multimodal Hyper Adapter For Many-to-Many Multimodal Summarization
abstract
Multimodal summarization (MS) combines text and visuals to generate summaries.Recently, many-to-many multimodal summarization (M3S) garnered interest as it enables a unified model for multilingual and cross-lingual MS.Existing methods have made progress by facilitating the transfer of common multimodal summarization knowledge.While, prior M3S models that fully share parameters neglect the language-specific knowledge learning, where potential interference between languages may limit the flexible adaptation of MS modes across different language combinations and hinder further collaborative improvements in joint M3S training.Based on this observation, we propose Language Constrained Multimodal Hyper Adapter (LCMHA) for M3S.LCMHA integrates language-specific multimodal adapters into multilingual pre-trained backbones via a language constrained hypernetwork, enabling relaxed parameter sharing that enhances language-specific learning while preserving shared MS knowledge learning.In addition, a language-regularized hypernetwork is designed to balance intra-and inter-language learning, generating language-specific adaptation weights and enhancing the retention of distinct language features through the regularization of generated parameters.Experimental results on the M3Sum benchmark show LCMHA's effectiveness and scalability across multiple multilingual pre-trained backbones.
Nayu Liu, Fanglong Yao, Yong Yang 0001
ACL (1)2
2025 RingMoGPT: A Unified Remote Sensing Foundation Model for Vision, Language, and Grounded Tasks
abstract
Recently, multimodal large language models (MLLMs) have shown excellent reasoning capabilities in various fields. Most of the existing remote sensing (RS) MLLMs solve image-level text generation problems (e.g., image captioning), but ignore the core issues of object-level recognition, location, and multitemporal changes in the field of RS. In this article, we propose RingMoGPT, a multimodal foundation model that unifies vision, language, and localization. Based on the idea of domain adaption, RingMoGPT can complete training by fine-tuning only a few parameters. To make the model capable of object detection and change captioning, we further propose a location- and instruction-aware querying transformer (Q-Former) and a change detection module, respectively. To improve the performance of RingMoGPT, we carefully design the pretraining dataset and the instruction-tuning dataset. The pretraining dataset contains over a half million high-quality image and text pairs, which are generated through a low-cost and efficient data generation paradigm. The instruction-tuning dataset contains more than 1.6 million question-answer pairs, including six downstream tasks: scene classification, object detection, visual question answering (VQA), image captioning, grounded image captioning, and change captioning. Our experiments show that RingMoGPT performs well on six tasks, especially its ability to analyze multitemporal data changes and identify dense objects. We also verified the model under a zero-shot setting, and the results show that the proposed RingMoGPT also has good generalization ability in the face of new data.
Peijin Wang, Huiyang Hu, Boyuan Tong, Ziqi Zhang 0010, Fanglong Yao, Yingchao Feng, Zining Zhu 0004, Wenhui Diao, Qixiang Ye, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 ReCon1M: A Large-Scale Benchmark Dataset for Relation Comprehension in Remote Sensing Imagery
abstract
Scene graph generation (SGG) is a high-level visual understanding and reasoning task aimed at extracting entities (such as objects) and their interrelationships from images. Significant progress has been made in the study of SGG in natural images in recent years, but its exploration in the domain of remote sensing images remains very limited. The complex characteristics of remote sensing images necessitate higher time and manual interpretation costs for annotation compared to natural images. The lack of a large-scale public SGG benchmark is a major impediment to the advancement of SGG-related research in aerial imagery. In this article, we introduce the first publicly available large-scale, million-level relation dataset in the field of remote sensing images, which is named ReCon1M. Specifically, our dataset is built upon FAIR1M and comprises 22 262 images. It includes annotations for 873 761 object bounding boxes across 60 categories and 1 052 223 relation triplets across 59 categories based on these bounding boxes. We provide a detailed description of the dataset’s characteristics and statistical information. In addition, an efficient global context-aware network (EGCAN) is proposed to improve inference efficiency in dense relation prediction through an object-pair pre-screening mechanism. By integrating visual, spatial, and semantic features, EGCAN captures fine-grained pairwise features and object-level contextual information to enhance its ability to discriminate relation. We conduct two object detection tasks and three subtasks within SGG on this dataset, assessing the performance of mainstream methods on these tasks. The experimental results show that the proposed EGCAN achieves state-of-the-art (SOTA) performance in 17 out of 24 accuracy metrics across three tasks and delivers the best performance in frames per second (FPS) for model inference. The ReCon1M dataset and related resources are available athttps://recon1m-dataset.github.io/
Qiwei Yan, Chubo Deng, Zhongyan Hou, Wanxuan Lu, Fanglong Yao, Lingxiang Hao, Xian Sun 0001
IEEE Trans. Geosci. Remote. Sens.8
2024 Multimodal Cross-Lingual Summarization for Videos: A Revisit in Knowledge Distillation Induced Triple-Stage Training Method
abstract
Multimodal summarization (MS) for videos aims to generate summaries from multi-source information (e.g., video and text transcript), showing promising progress recently. However, existing works are limited to monolingual scenarios, neglecting non-native viewers' needs to understand videos in other languages. It stimulates us to introduce multimodal cross-lingual summarization for videos (MCLS), which aims to generate cross-lingual summaries from multimodal input of videos. Considering the challenge of high annotation cost and resource constraints in MCLS, we propose a knowledge distillation (KD) induced triple-stage training method to assist MCLS by transferring knowledge from abundant monolingual MS data to those data with insufficient volumes. In the triple-stage training method, a video-guided dual fusion network (VDF) is designed as the backbone network to integrate multimodal and cross-lingual information through diverse fusion strategies in the encoder and decoder; What's more, we propose two cross-lingual knowledge distillation strategies: adaptive pooling distillation and language-adaptive warping distillation (LAWD), designed for encoder-level and vocab-level distillation objects to facilitate effective knowledge transfer across cross-lingual sequences of varying lengths between MS and MCLS models. Specifically, to tackle lingual sequences of varying lengths between MS and MCLS models. Specifically, to tackle the challenge of unequal length of parallel cross-language sequences in KD, LAWD can directly conduct cross-language distillation while keeping the language feature shape unchanged to reduce potential information loss. We meticulously annotated the How2-MCLS dataset based on the How2 dataset to simulate MCLS scenarios. Experimental results show that the proposed method achieves competitive performance compared to strong baselines, and can bring substantial performance improvements to MCLS models by transferring knowledge from the MS model.
Nayu Liu, Kaiwen Wei, Yong Yang 0001, Jianhua Tao 0001, Xian Sun 0001, Fanglong Yao, Li Jin 0001, Zhao Lv, Cunhang Fan
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 SFTformer: A Spatial-Frequency-Temporal Correlation-Decoupling Transformer for Radar Echo Extrapolation
abstract
Extrapolating future weather radar echoes from past observations is a complex task vital for precipitation nowcasting. The spatial morphology and temporal evolution of radar echoes exhibit a certain degree of correlation, yet they also possess independent characteristics. Existing methods learn unified spatial and temporal representations in a highly coupled feature space, emphasizing the correlation between spatial and temporal features but neglecting the explicit modeling of their independent characteristics, which may result in mutual interference between them. To effectively model the spatiotemporal dynamics of radar echoes, we propose a spatial-frequency-temporal correlation-decoupling transformer (SFTformer). The model leverages stacked multiple SFT-Blocks to not only mine the correlation of the spatiotemporal dynamics of echo cells but also avoid the mutual interference between the temporal modeling and the spatial morphology refinement by decoupling them. Furthermore, inspired by the practice that weather forecast experts effectively review historical echo evolution to make accurate predictions, SFTfomer incorporates a joint training paradigm for historical echo sequence reconstruction and future echo sequence prediction. Experimental results on the HKO-7 dataset and ChinaNorth-2021 dataset demonstrate the superior performance of SFTfomer in short-term (1 h), mid-term (2 h), and long-term (3 h) precipitation nowcasting.
Liangyu Xu, Wanxuan Lu, Fanglong Yao, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.4
2024 M2DCapsN: Multimodal, Multichannel, and Dual-Step Capsule Network for Natural Language Moment Localization
abstract
Natural language moment localization aims to localize the target moment that matches a given natural language query in an untrimmed video. The key to this challenging task is to capture fine-grained video-language correlations to establish the alignment between the query and target moment. Most existing works establish a single-pass interaction schema to capture correlations between queries and moments. Considering the complex feature space of lengthy video and diverse information between frames, the weight distribution of information interaction flow is prone to dispersion or misalignment, which leads to redundant information flow affecting the final prediction. We address this issue by proposing a capsule-based approach to model the query-video interactions, termed the Multimodal, Multichannel, and Dual-step Capsule Network ( [Formula: see text]DCapsN), which is derived from the intuition that "multiple people viewing multiple times is better than one person viewing one time." First, we introduce a multimodal capsule network, replacing the single-pass interaction schema of "one person viewing one time" with the iterative interaction schema of "one person viewing multiple times," which cyclically updates cross-modal interactions and modifies potential redundant interactions via its routing-by-agreement. Then, considering that the conventional routing mechanism only learns a single iterative interaction schema, we further propose a multichannel dynamic routing mechanism to learn multiple iterative interaction schemas, where each channel performs independent routing iteration to collectively capture cross-modal correlations from multiple subspaces, that is, "multiple people viewing." Moreover, we design a dual-step capsule network structure based on the multimodal, multichannel capsule network, bringing together the query and query-guided key moments to jointly enhance the original video, so as to select the target moments according to the enhanced part. Experimental results on three public datasets demonstrate the superiority of our approach in comparison with state-of-the-art methods, and comprehensive ablation and visualization analysis validate the effectiveness of each component of the proposed model.
Nayu Liu, Xian Sun 0001, Fanglong Yao, Guangluan Xu, Kun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Modeling High-Order Relationships: Brain-Inspired Hypergraph-Induced Multimodal-Multitask Framework for Semantic Comprehension
abstract
Semantic comprehension aims to reasonably reproduce people's real intentions or thoughts, e.g., sentiment, humor, sarcasm, motivation, and offensiveness, from multiple modalities. It can be instantiated as a multimodal-oriented multitask classification issue and applied to scenarios, such as online public opinion supervision and political stance analysis. Previous methods generally employ multimodal learning alone to deal with varied modalities or solely exploit multitask learning to solve various tasks, a few to unify both into an integrated framework. Moreover, multimodal-multitask cooperative learning could inevitably encounter the challenges of modeling high-order relationships, i.e., intramodal, intermodal, and intertask relationships. Related research of brain sciences proves that the human brain possesses multimodal perception and multitask cognition for semantic comprehension via decomposing, associating, and synthesizing processes. Thus, establishing a brain-inspired semantic comprehension framework to bridge the gap between multimodal and multitask learning becomes the primary motivation of this work. Motivated by the superiority of the hypergraph in modeling high-order relations, in this article, we propose a hypergraph-induced multimodal-multitask (HIMM) network for semantic comprehension. HIMM incorporates monomodal, multimodal, and multitask hypergraph networks to, respectively, mimic the decomposing, associating, and synthesizing processes to tackle the intramodal, intermodal, and intertask relationships accordingly. Furthermore, temporal and spatial hypergraph constructions are designed to model the relationships in the modality with sequential and spatial structures, respectively. Also, we elaborate a hypergraph alternative updating algorithm to ensure that vertices aggregate to update hyperedges and hyperedges converge to update their connected vertices. Experiments on the dataset with two modalities and five tasks verify the effectiveness of HIMM on semantic comprehension.
Xian Sun 0001, Fanglong Yao, Chibiao Ding
IEEE Trans. Neural Networks Learn. Syst.2
2023 Tackling higher-order relations and heterogeneity: Dynamic heterogeneous hypergraph network for spatiotemporal activity prediction
Changyuan Tian 0001, Zequn Zhang, Fanglong Yao, Zhi Guo, Shiyao Yan, Xian Sun 0001
Neural Networks3
2023 RingMo-Sense: Remote Sensing Foundation Model for Spatiotemporal Prediction via Spatiotemporal Evolution Disentangling
abstract
Remote sensing spatiotemporal prediction aims to infer future trends from historical spatiotemporal data, e.g., videos and time series images, has a broad application prospect in many fields. The foundation model is a promising research direction for spatiotemporal information mining because of its robust feature extraction capability, and has made rapid progress in natural scenes. Nevertheless, due to the spatially multi-scale and temporally multi-scale properties in remote sensing data, these methods still encounter bottlenecks when applied to remote sensing. Therefore, we propose a foundation model for remote sensing spatiotemporal prediction via spatiotemporal evolution decoupling, abbreviated as RingMo-Sense. Considering spatial affinity, temporal continuity, and spatiotemporal interaction, we construct spatial, temporal, and spatiotemporal triple-branch prediction networks. Specifically, we use parameter-sharing and progressive joint training strategies to achieve stable long-range prediction and parameter reduction simultaneously. In addition, we build a remote sensing spatiotemporal dataset by collecting various remote sensing videos and time series images. The experimental results on six downstream spatiotemporal tasks demonstrate that the proposed model yields competitive performance.
Fanglong Yao, Wanxuan Lu, Heming Yang 0003, Liangyu Xu, Leiyi Hu, Nayu Liu, Chubo Deng, Deke Tang, Changshuo Chen, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 Multimodal Remote Sensing Image Segmentation With Intuition-Inspired Hypergraph Modeling
abstract
Multimodal remote sensing (RS) image segmentation aims to comprehensively utilize multiple RS modalities to assign pixel-level semantics to the studied scenes, which can provide a new perspective for global city understanding. Multimodal segmentation inevitably encounters the challenge of modeling intra- and inter-modal relationships, $i.e$ ., object diversity and modal gaps. However, the previous methods are usually designed for a single RS modality, limited by the noisy collection environment and poor discrimination information. Neuropsychology and neuroanatomy confirm that the human brain performs the guiding perception and integrative cognition of multimodal semantics through intuitive reasoning. Therefore, establishing a semantic understanding framework inspired by intuition to realize multimodal RS segmentation becomes the main motivation of this work. Drived by the superiority of hypergraphs in modeling high-order relationships, we propose an intuition-inspired hypergraph network ( $I^{2}HN$ ) for multimodal RS segmentation. Specifically, we present a hypergraph parser to imitate guiding perception to learn intra-modal object-wise relationships. It parses the input modality into irregular hypergraphs to mine semantic clues and generate robust mono-modal representations. In addition, we also design a hypergraph matcher to dynamically update the hypergraph structure from the explicit correspondence of visual concepts, similar to integrative cognition, to improve cross-modal compatibility when fusing multimodal features. Extensive experiments on two multimodal RS datasets show that the proposed $I^{2}HN$ outperforms the state-of-the-art models, achieving F1/mIoU accuracy 91.4%/82.9% on the ISPRS Vaihingen dataset, and 92.1%/84.2% on the MSAW dataset.
Qibin He 0001, Xian Sun 0001, Wenhui Diao, Fanglong Yao, Kun Fu 0001
IEEE Trans. Image Process.5
2023 GAL: Graph-Induced Adaptive Learning for Weakly Supervised 3D Object Detection
abstract
Weakly Supervised 3D Object Detection (WS3DOD) aims to perform 3D object detection with little reliance on 3D labels, which greatly reduces the cost of 3D annotations. In recent literature, the pseudo-label-based approach brings impressive performance, which generates 3D pseudo-labels from 2D bounding boxes. Despite their success, two key issues remain unresolved that reduce the quality of 3D pseudo-labels: 1) the existing local object locating algorithm can not capture complete clusters of points globally, and 2) the existing algorithm can not capture sparse points caused by the unevenly distributed points obtained by LiDAR cameras. Hence, we propose GAL, a Graph-induced Adaptive Learning algorithm, to generate 3D pseudo-labels. First, we propose the Cluster Locating algorithm based on the Minimum Spanning Tree (MST) to globally locate the objects, which can leverage the characteristic that points inside an object are compact while points between objects are discrete. Second, we propose a density-guided adaptive learning algorithm to optimise the Cluster Locating algorithm, named Cuboid Drift. Cuboid Drift considers the inhomogeneous distribution of reflected points on different reflective surfaces of LiDAR imaging. Finally, 3D pseudo-labels generated by GAL are leveraged to train 3D detectors. Extensive experiments on the challenging KITTI and DAIR-V2X-V dataset demonstrate that GAL without 3D labels can be comparable with strongly supervised approaches and outperforms the previous state-of-the-art WS3DOD methods. Moreover, our method saves 88% of the time spent on pseudo-label generation.
Dongshuo Yin, Nayu Liu, Fanglong Yao, Qibin He 0001, Shiyao Yan, Xian Sun 0001
IEEE Trans. Intell. Transp. Syst.4
2023 Abstractive Summarization for Video: A Revisit in Multistage Fusion Network With Forget Gate
abstract
Multimodal abstractive summarization for videos is an emerging task that aims to generate a summary from multi-source information (i.e., video, audio transcript). The challenge is how to merge multimodal long sequences to capture rich semantic information without allowing possible noise from either lengthy modal sequence to degrade the other modality and thus hurt the entire model. To address the issues, we propose amultistagefusion network withforgetgate (MFFG), which selectively integrates multi-source information through the cross-fusion in encoding and hierarchical fusion in decoding between modalities, and design a fusion forget gate module to suppress the potential multimodal noise flow of multi-source long sequence. Meanwhile, considering that the source text in this task is lengthy and has the same distribution as the output summary text, we inherit the partial structure of the MFFG model and again propose its variant, single-stage fusion network with forget gate (SFFG), which simplifies the fusion schema, and leverages the long source text to enhance the representation of the target summary. Experimental results on How2 dataset and How2-300 dataset demonstrate the superiority of the two multimodal fusion methods. Further, we provide a version of ASR transcription data of How2 dataset to evaluate model performance under noisy scenarios, and experimental results show obvious advantages of our proposed models over prior systems.
Nayu Liu, Xian Sun 0001, Fanglong Yao, Guangluan Xu, Kun Fu 0001
IEEE Trans. Multim.4
2023 Mimicking the Brain's Cognition of Sarcasm From Multidisciplines for Twitter Sarcasm Detection
abstract
Sarcasm is a sophisticated construct to express contempt or ridicule. It is well-studied in multiple disciplines (e.g., neuroanatomy and neuropsychology) but is still in its infancy in computational science (e.g., Twitter sarcasm detection). In contrast to previous methods that are usually geared toward a single discipline, we focus on the multidisciplinary cross-innovation, i.e., improving embryonic sarcasm detection in computational science by leveraging the advanced knowledge of sarcasm cognition in neuroanatomy and neuropsychology. In this work, we are oriented toward sarcasm detection in social media and correspondingly propose a multimodal, multi-interactive, and multihierarchical neural network ($M_{3}N_{2} $). We select Twitter, image, text in image, and image caption as the input of$M_{3}N_{2} $since the brain’s perception of sarcasm requires multiple modalities. To reasonably address the multimodalities, we introduce singlewise, pairwise, triplewise, and tetradwise modality interactions incorporating gate mechanism and guide attention (GA) to simulate the interactions and collaborations of involved regions in the brain while perceiving multiple modes. Specifically, we exploit a multihop process for each modality interaction to extract modal information multiple times using GA for obtaining multiperspective information. Also, we adopt a two-hierarchical structure leveraging self-attention accompanied by attention pooling to integrate multimodal semantic information from different levels mimicking the brain’s first- and second-order comprehensions of sarcasm. Experimental results show that$M_{3}N_{2} $achieves competitive performance in sarcasm detection and displays powerful generalization ability in multimodal sentiment analysis and emotion recognition.
Fanglong Yao, Xian Sun 0001, Wenkai Zhang 0002, Kun Fu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 Videos
abstract
Multimodal summarization for videos aims to generate summaries from multi-source information (videos, audio transcripts), which has achieved promising progress.However, existing works are restricted to monolingual video scenarios, ignoring the demands of non-native video viewers to understand the cross-language videos in practical applications.It stimulates us to propose a new task, named Multimodal Cross-Lingual Summarization for videos (MCLS), which aims to generate cross-lingual summaries from multimodal inputs of videos.First, to make it applicable to MCLS scenarios, we conduct a Video-guided Dual Fusion network (VDF) that integrates multimodal and cross-lingual information via diverse fusion strategies at both encoder and decoder.Moreover, to alleviate the problem of high annotation costs and limited resources in MCLS, we propose a triple-stage training framework to assist MCLS by transferring the knowledge from monolingual multimodal summarization data, which includes: 1) multimodal summarization on sufficient prevalent language videos with a VDF model; 2) knowledge distillation (KD) guided adjustment on bilingual transcripts; 3) multimodal summarization for cross-lingual videos with a KD induced VDF model.Experiment results on the reorganized How2 dataset show that the VDF model alone outperforms previous methods for multimodal summarization, and the performance further improves by a large margin via the proposed triple-stage training framework. * Equal contribution. † Corresponding author.Portuguese (Pt) Transcript: vamos falar hoje sobre o solo.em primeiro lugar, precisamos de uma grande quan dade de solo bom para transplantes na primavera.ela vai adicionar partes iguais de musgo de turfa e composto de jardinagem que extraímos do nosso sistema interno de compostagem, e então um agregado orgânico, uma pedra chamada perlite, que serve para adicionar volume e aumentar a capacidade de retenção de água e de aeração de sua mistura... English (En) Summary: mix sterile soil for plan ng greens in trays to keep in a hoop house.learn to mix soil for growing greens from an organic farmer in this free gardening video.
Nayu Liu, Kaiwen Wei, Xian Sun 0001, Fanglong Yao, Li Jin 0001, Zhi Guo, Guangluan Xu
EMNLP5
2022 Cross-Modal Remote Sensing Image Retrieval Via Intra- and Inter-Modal Feature Matching
abstract
With the development of remote sensing (RS) acquisition technology, a mass of RS images have been produced, which brings challenges to the traditional manual retrieval methods and gives birth to the automatic RS image retrieval methods. Cross-modal RS image retrieval allows the usage of text and other modalities to retrieve RS images. For its flexible and convenient advantages, it has become a research hotspot. However, cross-modal RS image retrieval encounters the information asymmetry between modalities, i.e., RS images possess multi-scale, multi-objective properties and own rich information. At the same time, the query text is usually short and with less information. To solve the issues above, a cross-modal feature matching network is proposed to learn the feature fusion intra-modalities and the feature association inter-modalities to avoid the poor retrieval performance caused by the information asymmetry. Specifically, for the feature fusion intra-modalities, relying on the powerful feature representation ability of graph network, text and RS image graph modules are designed to fuse the intra-modal features. In terms of the feature correlation between modalities, RS image-text association module is created to attend the parts in text related to RS images and vice versa. Extended experiments on two public standard datasets verify the effectiveness of the proposed model.
Fanglong Yao, Nayu Liu, Peiguang Li, Dongshuo Yin, Xian Sun 0001
IGARSS1
2022 An Instance-Based Multitask Graph Network for Complex Facility Recognition in Remote Sensing Imagery
abstract
With the availability of very high-resolution remote sensing imagery, the fine-grained recognition of complex geospatial facilities has become possible. We can view these facilities as a combination of component objects with specific functions and distribution. However, the existing methods are insufficient in modeling spatial relations of component objects. In this article, we propose an instance-based multitask graph network (IBMG-Net) for complex facility recognition. Specifically, we perform pixel-level component objects prediction and facility recognition simultaneously and achieve performance improvement of both tasks by joint multitasking training. Given the component information, we build an instance-based graph neural network (IBGN) where components are defined as nodes and their spatial relations are encoded as edges. The IBGN module aims to flexibly model spatial relations of complex facility. To enhance the feature representation of component objects, we utilize the multiscale region of interest module (MS-ROI) to retain all scale-specific features and the sparse context information module (SCM) to aggregate long-range context information. In addition, we build a new multitask dataset for complex facility recognition in remote sensing (MCF dataset) to verify the effectiveness of our method and alleviate the lack of pixel-level labeled multitask datasets in remote sensing. Extensive experiments on MCF also indicate that the significant performance improvement of our approach to complex facility recognition.
Jingquan Peng, Xian Sun 0001, Chubo Deng, Fanglong Yao
IEEE Trans. Geosci. Remote. Sens.6
2022 Entity-Oriented Multi-Modal Alignment and Fusion Network for Fake News Detection
abstract
The development of social media enables fake news to be expressed in a multi-modal form, which is disseminated on various social platforms and brings harmful social impacts. To handle this challenge, the fake news detection task was proposed to examine whether false information is contained in multi-modal news. Existing methods exploit various approaches with cross-modal interaction and fusion, which have proven to be effective in detecting common fake news. However, although the description of multi-modal news is narrated around entities, the previously developed methods pay less attention to this characteristic. They do not explore its benefits to the detection task and underperform with respect to the detection of fake news that requires entity-centric comparisons. To make up for this omission, we explore a novel paradigm to detect fake news by aligning and fusing multi-modal entities and propose the Entity-oriented Multi-modal Alignment and Fusion network (EMAF). Our work adopts entity-centric cross-modal interaction, which can reserve semantic integrity and capture the details of multi-modal entities. Specifically, we design an Alignment module with the improved dynamic routing algorithm and introduce a Fusion module based on the comparison, the former aligns and captures the important entities and the latter compares and aggregates entity-centric features. Comparative experiments conducted on multiple public datasets, including Weibo, Twitter, and Reddit, reveal the superiority of the proposed EMAF method, and extensive analytical experiments demonstrate the effectiveness of our proposed modules.
Peiguang Li, Xian Sun 0001, Fanglong Yao, Guangluan Xu
IEEE Trans. Multim.5
2021 Commonalities-, specificities-, and dependencies-enhanced multi-task learning network for judicial decision prediction
Fanglong Yao, Xian Sun 0001, Wenkai Zhang 0002, Kun Fu 0001
Neurocomputing1
2020 Gated hierarchical multi-task learning network for judicial decision prediction
Fanglong Yao, Xian Sun 0001, Wenkai Zhang 0002, Kun Fu 0001
Neurocomputing1