VLDB 2026 Research / reviewers in the wild / expert
Fengmao Lv
dblp:146/6418
· DBLP profile ↗
65ranked-venue papers
17as first author
49since 2021 · last 2026
0000-0003-1640-0992ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 10 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 5 first-author · 21 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Security and privacy · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pre-trained conditional encoding guided diffusion for time series anomaly detection
Jinghong Xu, Shengdong Du, Jie Hu 0007, Yan Yang 0001, Fengmao Lv, Tianrui Li 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Regression-based multisource conditional domain adaptation for policy outcome prediction
Caijia Zhu, Liang Wu 0015, Desheng Zheng, Fengmao Lv |
Neural Networks | 6 |
| 2026 | Learning contrastive feature representations for facial action unit detection
Ziqiao Shang, Bin Liu 0076, Fengmao Lv, Fei Teng 0001, Tianrui Li 0001, Lan-Zhe Guo |
Pattern Recognit. | 3 |
| 2026 | Adaptive collaborative correlation learning-based semi-supervised multi-label feature selection
Yanyong Huang, Dongjie Wang 0001, Xiuwen Yi, Fengmao Lv, Tianrui Li 0001 |
Pattern Recognit. | 5 |
| 2026 | In-Depth Understanding of Crime Dynamics via Space-Time-Context-Aware Tensor DecompositionabstractUnderstanding the spatiotemporal characteristics of criminal activities in a city, or urban crime dynamics for short, is essential for developing ways to control crime and improve urban safety. While much effort has been devoted to this field, most of the existing studies have led to overly generalized findings, obscuring the ways in which dynamic patterns of criminal activities vary by place, time, and situational context. To address this challenge, this article proposes a novel space-time-context-aware tensor decomposition framework, namelySTCTD-Crime, for an in-depth understanding of urban crime dynamics. Specifically,STCTD-Crimefirst constructs a third-order tensor to represent crime data, which provides an elegant way to model spatial, temporal, and contextual factors simultaneously. Then, it decouples the influence that the three factors exerts on criminal activities via the tensor decomposition, enabling the observation of the extent to which each factor affects crime incidents occurring at different regions, within different time slices, and under different situational contexts. Moreover,STCTD-Crimeexploits spatiotemporal correlations between criminal activities to facilitate the understanding of dynamics by seamlessly integrating a crime-number-guided correlation learning method into the framework. Finally, an alternating optimization based scheme is developed to solve the optimization problem, which results in an efficient urban crime dynamics discovery procedure. Extensive analyses on crime datasets drawn from real-world sources convincingly demonstrate the effectiveness ofSTCTD-Crime. Weichao Liang, Guangliang Gao, Lei Chen 0079, Haicheng Tao, Lilan Peng, Fengmao Lv, Tianrui Li 0001 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2026 | Adaptive Topological Similarity Learning for Incomplete Multi-View Unsupervised Feature SelectionabstractAlthough multi-view unsupervised feature selection is a promising technique for dimensionality reduction on unlabeled multi-view data, existing methods cannot directly address incomplete data, where certain samples are missing in specific views. These methods typically begin by imputing missing data using predetermined values, followed by performing feature selection on the completed dataset. However, the separation of imputation and feature selection processes fails to exploit their inherent synergy, as local structural information obtained from feature selection could guide the imputation process and, in turn, improve the overall effectiveness of feature selection. In addition, previous methods rely on similarity graphs based on Euclidean distance to preserve the local manifold structure but overlook the topological relationships within the data, thereby hindering accurate capture of intrinsic structures. In this paper, we propose an adaptive topological similarity learning for incomplete multi-view unsupervised feature selection method (ATSL-IMUFS) to address the aforementioned issues. ATSL-IMUFS first integrates multi-view feature selection and missing data imputation into a unified learning framework. Then, it adaptively learns similarity graphs for each view while simultaneously capturing the consensus topological relationship across views, effectively characterizing the local manifold structure. Extensive experiments conducted on real-world datasets demonstrate the superior performance of ATSL-IMUFS compared to competing methods. Dongjie Wang 0001, Fanyin Zhou, Fengmao Lv, Tianrui Li 0001, Yanyong Huang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Lighted-SAM: Lightening Open-World SAM for Low-Light SegmentationabstractSegment Anything Model (SAM) has achieved impressive segmentation performance in an open-world setting. However, SAM relies heavily on high-quality input images and usually struggles in low-light conditions. This is mainly caused by the pre-training dataset, SA-1B, in which low-light samples constitute a relatively small fraction of the data. This lack of presence leads to a noticeable weakness when SAM is applied in real-world dark environments. With the motivation of improving SAM's performance under low-light conditions while retaining its strong zero-shot capability, this work proposes an alignment stage between the pre-training stage and testing stage. Unlike existing low-light studies that mainly focus on task-specific and close-set settings, our work further emphasizes pursuing the segmentation ability under low-light conditions for open-world models. To this end, we construct DarkSeg58K, a realistic and diverse dataset, which serves as the alignment dataset to support this stage. We further introduce Lighted-SAM as the lightweight repair strategy to fix SAM's performance in low-light conditions. Different from existing methods focusing on introducing spectral adapters into the model design and training this model end-to-end, Lighted-SAM introduces the Spectral Information Resonance (SIR) mechanism to harmoniously integrate the spectral enhancement module into SAM, which is usually kept frozen due to its large-scale parameters. Based on our lightweight repairing strategy, Lighted-SAM can improve SAM's ability in low-light conditions while preserving its zero-shot ability. Experiments on different benchmarks validate the superiority of our approach. Code is available at: https://github.com/Jaaaahan/LightedSAM. Yuhan Jia, Lixin Duan, Wen Li 0001, Fengmao Lv |
IEEE Trans. Image Process. | 4 |
| 2026 | Causally-Aware Unsupervised Feature Selection LearningabstractUnsupervised feature selection (UFS) has recently gained attention for its effectiveness in processing unlabeled high-dimensional data. However, existing methods overlook the intrinsic causal mechanisms within the data, resulting in the selection of irrelevant features and poor interpretability. Additionally, previous graph-based methods fail to account for the differing impacts of non-causal and causal features in constructing the similarity graph, which leads to false links in the generated graph. To address these issues, a novel UFS method, called Causally-Aware UnSupErvised Feature Selection learning (CAUSE-FS), is proposed. CAUSE-FS introduces a novel causal regularizer that reweights samples to balance the confounding distribution of each treatment feature. This regularizer is subsequently integrated into a generalized unsupervised spectral regression model to mitigate spurious associations between features and clustering labels, thus achieving causal feature selection. Furthermore, CAUSE-FS employs causality-guided hierarchical clustering to partition features with varying causal contributions into multiple granularities. By integrating similarity graphs learned adaptively at different granularities, CAUSE-FS increases the importance of causal features when constructing the fused similarity graph to capture the reliable local structure of data. Extensive experimental results demonstrate the superiority of CAUSE-FS over state-of-the-art methods, with its interpretability further validated through feature visualization. Zongxin Shen, Yanyong Huang, Dongjie Wang 0001, Minbo Ma, Fengmao Lv, Tianrui Li 0001 |
IEEE Trans. Image Process. | 5 |
| 2026 | Modeling Deep Fusion of Intra- and Inter-Modal Incongruity for Multimodal Sarcasm DetectionabstractMultimodal sarcasm detection receives increasing attentions due to people's growing interest in posting multimodal information. The key factor of multimodal sarcasm detection is to leverage incongruity information across different modalities. Existing works are mainly based on the late fusion strategy by simply concatenating the intra- and inter-modal incongruity features, which are prone to learning surface patterns. In contrast, this work mainly focuses on modeling the deep fusion of intra- and inter-modal incongruity information. To this end, this work first discusses the incompatibility between the two kinds of incongruity features within existing multimodal frameworks. Under this motivation, we further propose an end-to-end cooperative framework dubbed Cooperative Multimodal Incongruity Learning (CoMIL). Specifically, our approach incorporates a primary module to model the deep fusion of intra- and inter modal incongruity information. To prevent the integrated inter modal visual information from disturbing the modeling of intra text incongruity, CoMIL introduces a cooperative mechanism incorporating a reference module which focuses on token-level correlations as a structural guidance to the primary module. Based on the proposed cooperative mechanism, the intra- and inter-modal incongruity information can be compactly and compatibly integrated into deep features of neural models. Extensive experiments are conducted to validate the effectiveness of our proposed CoMIL approach. Fengmao Lv, Junlin Fang, Guosheng Lin, Wenya Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2026 | Rethinking the Effect of Unimodal Labels in Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis aims to comprehensively understand human sentiment by integrating diverse modalities, such as text, audio, and vision. To improve modality complementarity, the recent Multimodal Multi-task Learning (MML) framework employs joint training of unimodal and multimodal sentiment analysis tasks using sub-annotations of modality. In this work, we further draw attention to the observation that integrating unimodal tasks may introduce conflicting task information, negatively affecting the multimodal task performance. Motivated by this issue, we propose the Multimodal Task Correlation-aware Learning (MTCL) framework to leverage beneficial task correlations and suppress harmful ones. Specifically, MTCL introduces a Correlation-Adaptive Training (CAT) strategy to learn a task-relation aware unimodal encoder for each modality. First, in order to distinguish whether a sample contains conflicting information, CAT strategy incorporates a Dual-Branch Contrast (DBC) module which divides the training set into a beneficial subset and a harmful subset. Based on this division, CAT strategy proposes an adaptive training loss to guide the model in understanding nuanced multitask correlations. The adaptive training loss has two components: (1) For the beneficial subset, a contrastive loss is utilized to improve the model’s ability to extract complementary representations. (2) For the harmful subset, we apply a task-correction loss to mitigate the negative interference caused by harmful task associations. With the CAT strategy, our framework can effectively distinguish beneficial and harmful task correlations to extract distinctive and robust unimodal representations. The superiority of MTCL is verified via extensive experiments on several multimodal video sentiment analysis benchmarks. Our work is publicly available at https://github.com/tiggers23/MTCL . Tianrui Li 0001, Baiyu Lu, Junlin Fang, Desheng Zheng, Wei Zhou 0021, Weide Liu, Fengmao Lv |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2025 | Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score CollaborationabstractUnderstanding how humans cooperatively utilize semantic knowledge to explore unfamiliar environments and decide on navigation directions is critical for house service multi-robot systems. Previous methods primarily focused on single-robot centralized planning strategies, which severely limited exploration efficiency. Recent research has considered decentralized planning strategies for multiple robots, assigning separate planning models to each robot, but these approaches often overlook communication costs. In this work, we propose Multimodal Chain-of-Thought Co-Navigation (MCoCoNav), a modular approach that utilizes multimodal Chain-of-Thought to plan collaborative semantic navigation for multiple robots. MCoCoNav combines visual perception with Vision Language Models (VLMs) to evaluate exploration value through probabilistic scoring, thus reducing time costs and achieving stable outputs. Additionally, a global semantic map is used as a communication bridge, minimizing communication overhead while integrating observational results. Guided by scores that reflect exploration trends, robots utilize this map to assess whether to explore new frontier points or revisit history nodes. Experiments on HM3D_v0.2 and MP3D demonstrate the effectiveness of our approach. Zhixuan Shen, Haonan Luo 0002, Kexun Chen, Fengmao Lv, Tianrui Li 0001 |
AAAI | 4 |
| 2025 | Attribute-formed Class-specific Concept Space: Endowing Language Bottleneck Model with Better Interpretability and ScalabilityabstractLanguage Bottleneck Models (LBMs) are proposed to achieve interpretable image recognition by classifying images based on textual concept bottlenecks. However, current LBMs simply list all concepts together as the bottleneck layer, leading to the spurious cue inference problem and cannot generalized to unseen classes. To address these limitations, we propose the Attribute-formed Language Bottleneck Model (ALBM). ALBM organizes concepts in the attribute-formed class-specific space, where concepts are descriptions of specific attributes for specific classes. In this way, ALBM can avoid the spurious cue inference problem by classifying solely based on the essential concepts of each class. In addition, the cross-class unified attribute set also ensures that the concept spaces of different classes have strong correlations, as a result, the learned concept classifier can be easily generalized to unseen classes. Moreover, to further improve interpretability, we propose Visual Attribute Prompt Learning (VAPL) to extract visual features on fine-grained attributes. Furthermore, to avoid labor-intensive concept annotation, we propose the Description, Summary, and Supplement (DSS) strategy to automatically generate high-quality concept sets with a complete and precise attribute. Extensive experiments on 9 widely used few-shot benchmarks demonstrate the interpretability, transferability, and performance of our approach. The code and collected concept sets are available at https://github.com/tiggers23/ALBM. Jianyang Zhang, Qianli Luo, Guowu Yang, Wenjing Yang 0003, Weide Liu, Guosheng Lin, Fengmao Lv |
CVPR | 7 |
| 2025 | HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video UnderstandingabstractMultimodal large language models have become a popular topic in deep visual understanding due to many promising real-world applications. However, hour-long video understanding, spanning over one hour and containing tens of thousands of visual frames, remains under-explored because of 1) challenging long-term video analyses, 2) inefficient large-model approaches, and 3) lack of large-scale benchmark datasets. Among them, in this paper, we focus on building a large-scale hour-long long video benchmark, HLV-1K1, designed to evaluate long video understanding models. HLV-1K comprises 1009 hour-long videos with 14,847 high-quality question answering (QA) and multi-choice question asnwering (MCQA) pairs with time-aware query and diverse annotations, covering frame-level, within-event-level, cross-event-level, and long-term reasoning tasks. We evaluate our benchmark using existing state-of-the-art methods and demonstrate its value for testing deep long video understanding capabilities at different levels and for various tasks. This includes promoting future long video understanding tasks at a granular level, such as deep understanding of long live videos, meeting recordings, and movies. Heqing Zou, Tianze Luo, Guiyang Xie, Victor Xiao Jie Zhang, Fengmao Lv, Guangcong Wang, Junyang Chen 0001, Zhuochen Wang, Hansheng Zhang, Huaijian Zhang |
ICME | 5 |
| 2025 | Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across LanguagesabstractMultilingual speech emotion recognition aims to estimate a speaker’s emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses significant challenges for zero-shot speech emotion recognition, especially with multilingual datasets. In this paper, we propose leveraging contrastive learning to refine multilingual speech features and extend large language models for zero-shot multilingual speech emotion estimation. Specifically, we employ a novel two-stage training framework to align speech signals with linguistic features in the emotional space, capturing both emotion-aware and language-agnostic speech representations. To advance research in this field, we introduce a large-scale synthetic multilingual speech emotion dataset, M5SER. Our experiments demonstrate the effectiveness of the proposed method in both speech emotion recognition and zero-shot multilingual speech emotion recognition, including previously unseen datasets and languages. Our introduced dataset and related code will be available on GitHub1. Heqing Zou, Fengmao Lv, Desheng Zheng, Chng Eng Siong, Deepu Rajan |
ICME | 2 |
| 2025 | Why is a Bird's Caption a Good Demonstration? Towards Effective Multimodal In-Context Learning without Dedicated Data
Junlin Fang, Wenya Wang 0001, Fengmao Lv |
ACM Multimedia | 4 |
| 2025 | Damage Analysis via Bidirectional Multi-Task Cascaded Multimodal FusionabstractDamage analysis in social media platforms such as Twitter is a comprehensive problem which involves different subtasks for mining damage-related information from tweets ( e.g., informativeness, humanitarian categories and severity assessment). The comprehensive information obtained by damage analysis enables to identify breaking events around the world in real-time and hence provides aids in emergency responses. Recently, with the rapid development of web technologies, multimodal damage analysis has received increasing attentions due to users' preference of posting multimodal information in social media. Multimodal damage analysis leverages the associated image modality to improve the identification of damage-related information in social media. However, existing works on multimodal damage analysis address each damage-related subtask individually and do not consider their joint training mechanism. In this work, we propose the Bidirectional Multi-task Cascaded multimodal Fusion (BiMCF) approach towards joint multimodal damage analysis. To this end, we introduce the cascaded multimodal fusion framework to separately integrate effective visual and text information for each task, considering that different tasks attend to different information. To exploit the interactions across tasks, bidirectional propagation of the attended image-text interactive information is implemented between tasks, which can lead to enhanced multimodal fusion. Comprehensive experiments are conducted to validate the effectiveness of the proposed approach. Code is available at https://github.com/tiggers23/BiMCF. Siying Wu, Junfeng Fang, Guowu Yang, Wenya Wang 0001, Fengmao Lv |
WWW | 6 |
| 2025 | MSTI-Plus: Introducing Non-Sarcasm Reference Materials to Enhance Multimodal Sarcasm Target IdentificationabstractSarcasm is a subtle expression that indicates the incongruity between literal meanings and factual opinions. For multimodal posts in social medias which consist of both images and texts, sarcasm expressions are even more widespread. Recent works have paid attentions to Multimodal Sarcasm Target Identification (MSTI), which focuses on detecting aspect terms of mockery or ridicule as sarcasm targets. However, the current MSTI benchmark only contains annotations on fine-grained sarcasm targets within sarcastic samples. In practice, it will be featured by two major limitations. First, there lack annotations on non-sarcasm aspects to inform deep models to perceive the semantic difference between sarcasm targets and non-sarcasm aspects. As a result, deep models will tend to incorrectly recognize non-sarcasm aspects as sarcasm targets. Second, there lack non-sarcasm samples to inform deep models to perceive the inherent semantics of sarcasm intentions. Due to the subtle characteristic of sarcasm expressions, models trained with only fine-grained supervision signals cannot thoroughly understand the sarcasm semantics, making the fine-grained task of sarcasm target identification restricted. Motivated by these limitations, this work reconstructs a more comprehensive MSTI benchmark by introducing both fine-grained non-sarcasm aspect annotations for existing sarcasm samples and non-sarcastic samples as non-sarcasm references to enable deep models to clearly perceive the mentioned information during training. Based on the multi-granularity (i.e., both aspect-level and sample-level) non-sarcasm information introduced into this new benchmark, this work further proposes a pluggable Semantics-aware Sarcasm Target Identification mechanism to enhance sarcasm target identification by modeling the overall semantics of sarcasm intentions via an auxiliary sample-level sarcasm recognition task. By modeling the overall semantics of sarcasm intention, deep models can obtain a more comprehensive understanding on sarcasm semantics, leading to improved performance on fine-grained sarcasm target identification. Extensive experiments are conducted to validate our contribution. Both the dataset and code are available at https://github.com/tiggers23/MSTI-Plus. Fengmao Lv, Mengting Xiong, Junlin Fang, Tianze Luo, Weichao Liang, Tianrui Li 0001 |
WWW | 1 |
| 2025 | A pre-trained data deduplication model based on active learning
Xinyao Liu, Fengmao Lv, Hongtao Xue, Jie Hu 0007, Shengdong Du, Tianrui Li 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Pixel is All You Need: Adversarial Spatio-Temporal Ensemble Active Learning for Salient Object DetectionabstractAlthough weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial spatio-temporal ensemble active learning. Our contributions are four-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed spatio-temporal ensemble strategy not only achieves outstanding performance but significantly reduces the model's computational cost. 3) Our proposed relationship-aware diversity sampling can conquer oversampling while boosting model performance. 4) We provide theoretical proof for the existence of such a point-labeled dataset. Experimental results show that our approach can find such a point-labeled dataset, where a saliency model trained on it obtained 98%-99% performance of its fully-supervised version with only ten annotated points per image. Wei Wang 0169, Yacong Li, Fengmao Lv, Qing Xia 0002, Chenglizhao Chen, Aimin Hao, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Citywide Multi-Step Crime Prediction via Context-Aware Bayesian Tensor DecompositionabstractCrime prediction, which focuses on forecasting the occurrence of criminal activities across city regions before they occur, constitutes an essential capability of surveillance systems designed to enhance urban security. While much effort has been invested in this field, most of the existing studies pay little attention to the influence of situational contexts on criminal activities, which hinders further improvement in prediction performance. To address this challenge, we propose a novel context-aware Bayesian tensor decomposition framework, namely cBTD-Crime, for citywide multi-step crime prediction. More specifically, cBTD-Crime first constructs a third-order tensor to simultaneously model spatial, temporal, and contextual factors and then applies the CP decomposition to exploit the intricate relationships between the three factors to facilitate the prediction process. To reduce the parameter tuning cost, cBTD-Crime further reformulates the problem from a probabilistic perspective, where a range of carefully selected distributions are placed on the spatial, temporal, and contextual latent factors. Finally, an efficient Gibbs sampling procedure is developed to generate a series of samples and the arithmetic mean is computed to obtain the predicted number of crime incidents. Experimental results show that cBTD-Crime achieves superior performance on real-world crime datasets in terms of different evaluation metrics. Weichao Liang, Fengmao Lv, Lei Chen 0079, Haicheng Tao, Min Shi 0001, Xingquan Zhu 0001, Jie Cao 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Multi-Scale CNN-Transformer Hybrid Network for Rail Fastener Defect DetectionabstractDefect detection in rail fasteners is crucial for train safety, as defective fasteners can cause derailments and severe safety incidents. However, Existing algorithms often struggle in various real-world scenarios due to challenges such as obscured fasteners, motion blur in images, varying camera angles, and fasteners submerged in water. To address these challenges, we propose a Multi-scale CNN-Transformer Hybrid Network for Rail Fastener Defect Detection (MCHNet-RF2D), specifically designed to identify fastener defects in complex environments. Our approach constructs an efficient CNN block and a multi-scale Vision Transformer block to alternately extract local detail features and global semantic features of the fasteners. These features are seamlessly integrated through multi-scale fusion to enhance defect recognition robustness. By combining comprehensive global recognition with detailed local defect detection, MCHNet-RF2D outperforms existing CNN-Transformer hybrid networks by 2.8% and surpasses current fastener defect detection algorithms by 2.9%. In practical deployment on over 40 trains, our model successfully detected more than 2,000 fastener defects, demonstrating its effectiveness in diverse and challenging conditions. Wei Wang 0278, Fengmao Lv, Haonan Luo 0002, Gexiang Zhang, Zhenghua Chen |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Adaptive Multi-Scale Language Reinforcement for Multimodal Named Entity RecognitionabstractOver the recent years, multimodal named entity recognition has gained increasing attentions due to its wide applications in social media. The key factor of multimodal named entity recognition is to effectively fuse information of different modalities. Existing works mainly focus on reinforcing textual representations by fusing image features via the cross-modal attention mechanism. However, these works are limited in reinforcing the text modality at the token level. As a named entity usually contains several tokens, modeling token-level inter-modal interactions is suboptimal for the multimodal named entity recognition problem. In this work, we propose a multimodal named entity recognition approach dubbed Adaptive Multi-scale Language Reinforcement (AMLR) to implement entity-level language reinforcement. To this end, our model first expands token-level textual representations into multi-scale textual representations which are composed of language units of different lengths. After that, the visual information reinforces the language modality by modeling the cross-modal attention between images and expanded multi-scale textual representations. Unlike existing token-level language reinforcement methods, the word sequences of named entities can be directly interacted with the visual features as a whole, making the modeled cross-modal correlations more reasonable. Although the underlying entity is not given, the training procedure can encourage the relevant image contents to adaptively attend to the appropriate language units, making our approach not rely on the pipeline design. Comprehensive evaluation results on two public Twitter datasets clearly demonstrate the superiority of our proposed model. Enping Li, Tianrui Li 0001, Huaishao Luo, Jielei Chu, Lixin Duan, Fengmao Lv |
IEEE Trans. Multim. | 6 |
| 2025 | Segmenting Anything in the Dark via Depth PerceptionabstractImage segmentation under low-light conditions is essential in real-world applications, such as autonomous driving and video surveillance systems. The recent Segment Anything Model (SAM) exhibits strong segmentation capability in various vision applications. However, its performance could be severely degraded under low-light conditions. On the other hand, multimodal information has been exploited to help models construct more comprehensive understanding of scenes under low-light conditions by providing complementary information (e.g., depth). Therefore, in this work, we present a pioneer attempt that elevates a unimodal vision foundation model (e.g., SAM) to a multimodal one, by efficiently integrating additional depth information under low-light conditions. To achieve that, we propose a novel method called Depth Perception SAM (DPSAM) based on the SAM framework. Specifically, we design a modality encoder to extract the depth information and the Depth Perception Layers (DPLs) for mutual feature refinement between RGB and depth features. The DPLs employ the cross-modal attention mechanism to mutually query effective information from both RGB and depth for the subsequent feature refinement. Thus, DPLs can effectively leverage the complementary information from depth to enrich the RGB representations and obtain comprehensive multimodal visual representations for segmenting anything in the dark. To this end, our DPSAM maximally maintains the instinct expertise of SAM for RGB image segmentation and further leverages on the strength of depth for enhanced segmenting anything capability, especially for cases that are likely to fail with RGB only (e.g., low-light or complex textures). As demonstrated by extensive experiments on four RGBD benchmark datasets, DPSAM clearly improves the performance for the segmenting anything performance in the dark, e.g., +12.90% mIoU and +16.23% mIoU on LLRGBD and DeLiVER, respectively. Our code and model will be made publicly available at:https://github.com/liupeng3425/DPSAM. Peng Liu 0049, Jinhong Deng, Lixin Duan, Wen Li 0001, Fengmao Lv |
IEEE Trans. Multim. | 5 |
| 2024 | Progressive Multimodal Pivot Learning: Towards Semantic Discordance Understanding as HumansabstractMultimodal recognition can achieve enhanced performance by leveraging the complementary information from different modali- ties. However, in real-world scenarios, multimodal samples often express discordant semantic meanings across modalities, lacking evident complementary information. Unlike humans who can easily understand the intrinsic semantic information of these semantically discordant samples, existing multimodal recognition models show poor performance on them. With the motivation of improving the robustness of multimodal recognition models in practical scenar- ios, this work poses a new challenge in multimodal recognition, which is coined as Semantic Discordance Understanding. Unlike ex- isting works only focusing on detecting semantically discordant samples as noisy data, this new challenge requires deep models to follow humans’ ability in understanding the inherent seman- tic meanings of semantically discordant samples. To address this challenge, we further propose the Progressive Multimodal Pivot Learning (PMPL) approach by introducing a learnable pivot mem- ory to explore the inherent semantics meaning hidden under dis- cordant modalities. To this end, our approach inserts Pivot Memory Learning (PML) modules into multiple layers of unimodal foun- dation models to progressively trade-off the conflict information across modalities. By introducing the multimodal pivot learning paradigm for multimodal recognition, the proposed PMPL approach can alleviate the negative effect of semantic discordance caused by the cross-modal information exchange mechanism of existingmultimodal recognition models. Experiments on different bench- marks validate the superiority of our approach. Code is available at https://github.com/tiggers23/PMPL. Junlin Fang, Wenya Wang 0001, Tianze Luo, Yanyong Huang, Fengmao Lv |
CIKM | 5 |
| 2024 | Unified View Imputation and Feature Selection Learning for Incomplete Multi-view Data
Yanyong Huang, Zongxin Shen, Tianrui Li 0001, Fengmao Lv |
IJCAI | 4 |
| 2024 | Sentiment-oriented Sarcasm Integration for Video Sentiment Analysis Enhancement with Sarcasm AssistanceabstractSarcasm is an intricate expression phenomenon and has garnered increasing attentions over the recent years, especially for multimodal contexts such as videos.Nevertheless, despite being a significant aspect of human sentiment, the effect of sarcasm is consistently overlooked in sentiment analysis.Videos with sarcasm often convey sentiments that diverge or even contradict their explicit messages.Prior works mainly concentrate on simply modeling sarcasm and sentiment features by utilizing the Multi-Task Learning (MTL) framework, which we found introduces detrimental interplays between the sarcasm detection task and sentiment analysis task.Therefore, this study explores the effective enhancement of video sentiment analysis through the incorporation of sarcasm information.To this end, we propose the Progressively Sentiment-oriented Sarcasm Refinement and Integration (PS2RI) framework, which focuses on modeling sentiment-oriented sarcasm features to enhance sentiment prediction.Instead of naively combining sarcasm detection and sentiment prediction under an MTL framework, PS2RI iteratively performs the sentiment-oriented sarcasm refinement and sarcasm integration operations within the sentiment recognition framework, in order to progressively learn sarcasm-aware sentiment feature without suffering the detrimental interplays caused by information irrelevant to the sentiment analysis task.Extensive experiments are conducted to validate the effectiveness of our approach.Code is available at https://github.com/tiggers23/PS2RI. Junlin Fang, Wenya Wang 0001, Guosheng Lin, Fengmao Lv |
ACM Multimedia | 4 |
| 2024 | Rethinking the Effect of Uninformative Class Name in Prompt LearningabstractLarge pre-trained vision-language models like CLIP have shown amazing zero-shot recognition performance. To adapt pre-trained vision-language models to downstream tasks, recent studies have focused on the learnable context + class name paradigm, which learns continuous prompt contexts on downstream datasets. In practice, the learned prompt context tends to overfit the base categories and cannot generalize well to novel categories out of the training data. Recent works have also noticed this problem and have proposed several improvements. In this work, we draw a new insight based on empirical analysis, that is, uninformative class names lead to degraded base-to-novel generalization performance in prompt learning, which is usually overlooked by existing works. Under this motivation, we advocate to improve the base-to-novel generalization performance of prompt learning by enhancing the semantic richness of class names. We coin our approach as the Information Disengagement based Associative Prompt Learning (IDAPL) mechanism which considers the associative, meanwhile, decoupled learning of prompt context and class name embedding. IDAPL can effectively alleviate the phenomenon of learnable context overfitting to base classes, meanwhile, learning more informative semantic representation of base classes by fine-tuning the class name embedding, leading to improved performance on both base and novel classes. Experimental results on eleven widely used few-shot learning benchmarks clearly validate the effectiveness of our proposed approach. Code is available at https://github.com/tiggers23/IDAPL Fengmao Lv, Changru Nie, Jianyang Zhang, Guowu Yang, Guosheng Lin, Xiao Wu 0001, Tianrui Li 0001 |
ACM Multimedia | 1 |
| 2024 | VLAI: Exploration and Exploitation based on Visual-Language Aligned Information for Robotic Object Goal Navigation
Haonan Luo 0002, Yijie Zeng, Kexun Chen, Zhixuan Shen, Fengmao Lv |
Image Vis. Comput. | 6 |
| 2024 | Source-free domain adaptation with unrestricted source hypothesis
Jiujun He, Liang Wu 0015, Chaofan Tao, Fengmao Lv |
Pattern Recognit. | 4 |
| 2024 | CAFA: Cross-Modal Attentive Feature Alignment for Cross-Domain Urban Scene SegmentationabstractAutonomous driving systems rely heavily on semantic segmentation models for accurate and safe decision-making. High segmentation performance in real-world urban scenes is crucial for autonomous vehicles, while substantial pixel-level labels are required during model training. Unsupervised domain adaptation (UDA) techniques are widely used to adapt the segmentation model trained on the synthetic data (i.e., source domain) to the real-world data (i.e., target domain) since obtaining pixel-level annotations is fairly easy in the synthetic environment. Recently, increasing UDA approaches promote cross-domain semantic segmentation (CDSS) by fusing the depth information into the RGB features. However, feature fusion does not necessarily eliminate the domain-specific components in the RGB features, which can result in the features still being influenced by domain-specific information. To address this, we propose a novel cross-modal attentive feature alignment (CAFA) framework for CDSS, which provides an explicit perspective of using depth information to align the main backbone RGB features of both domains in a nonadversarial manner. In particular, considering that the depth modality is less affected by the domain gap, we employ depth as an intermediate modality and align the RGB features by attending RGB features to the depth modality through constructing an auxiliary multimodal segmentation task. The state-of-the-art performance of our CAFA can be achieved on benchmark tasks, such as Synthia$\to$Cityscapes and grand theft auto (GTA)$\to$Cityscapes. Peng Liu 0049, Yanqi Ge, Lixin Duan, Wen Li 0001, Fengmao Lv |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Transferring Multi-Modal Domain Knowledge to Uni-Modal Domain for Urban Scene SegmentationabstractSynthetic data (i.e., source domain) have been widely adopted to improve the semantic segmentation performance for real-world images (i.e., target domain), since obtaining pixel-level annotations is fairly easy in the synthetic environment. Traditional domain adaptation methods normally focus on learning in the RGB modality only. We notice that the synthetic environment can generate depth information of semantic objects at almost no cost, while it is nontrivial to collect such information in the real-world scenario. In this case, we employ the depth information of synthetic data in this work to further boost the segmentation performance, and then transform the uni-modal problem into a multi-modal one. In this work, we focus on urban scene understanding and make a pioneer attempt on learning uni-modal feature representations for real-world images by mining from multi-modal knowledge of synthetic images with additional depth information. To this end, we propose a novel method called Multi-modal Domain Knowledge Transfer (MDKT), which transfers the multi-modal knowledge of the source domain to the uni-modal target domain through domain adaptation. In MDKT, we first employ the Cross-Modal Correlation (CMC) module to enhance the source features by fusing the RGB and depth information. Then, the uni-modal target domain feature and multi-modal source domain feature are aligned through the Modal-Imbalanced Adversarial Training (MIAT) strategy, which transfers the multi-modal knowledge to the uni-modal network in the target domain. We conduct extensive experiments on several benchmark settings for urban scene understanding. The promising results clearly show the effectiveness of our proposed MDKT approach. Peng Liu 0049, Yanqi Ge, Lixin Duan, Wen Li 0001, Haonan Luo 0002, Fengmao Lv |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | Progressive Multigranularity Information Propagation for Coupled Aspect-Opinion ExtractionabstractCoupled aspect-opinion extraction aims to identify aspect-opinion pairs in the form of (aspect term, opinion term) or triplets in the form of (aspect term, opinion term, sentiment polarity) from user-generated texts. Compared to the traditional aspect-based sentiment prediction or extraction tasks, coupled aspect-opinion extraction needs to associate aspects with their corresponding opinions and organize opinion-related information into structured outputs. The existing works either divide this task into subproblems (i.e., term extraction and relation prediction) or utilize a unified tagging scheme. However, these methods only focus on atomic word-level interactions and ignore the intensive information propagation among different granularities (e.g., words and word pairs). To address this limitation, we propose a progressive multigranularity information propagation network that progressively explores three types of correlations with different granularities. Specifically, our model starts with the most basic word-level correlations by composing all possible word pairs. In the second stage, the pairwise relation information is used to update the word features. The last stage propagates information among word pairs to produce the relation scores. We treat the task as a unified relation prediction problem and construct an end-to-end framework that iteratively conducts the three-stage information propagation to refine the textual representations. Comprehensive experiments on different aspect-based sentiment analysis benchmarks clearly demonstrate the effectiveness of the proposed approach. Fengmao Lv, Zhihui Fei, Wenya Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Learning cross-domain semantic-visual relationships for transductive zero-shot learning
Fengmao Lv, Jianyang Zhang, Guowu Yang, Lei Feng 0006, Lixin Duan |
Pattern Recognit. | 1 |
| 2023 | Self-Training Vision Language BERTs With a Unified Conditional ModelabstractNatural language BERTs are trained with language corpus in a self-supervised manner. Unlike natural language BERTs, vision language BERTs need paired data to train, which restricts the scale of VL-BERT pretraining. We propose a self-training approach that allows training VL-BERTs from unlabeled image data. The proposed method starts with our unified conditional model– a vision language BERT model that can perform zero-shot conditional generation. Given different conditions, the unified conditional model can generate captions, dense captions, and even questions. We use the labeled image data to train a teacher model and use the trained model to generate pseudo captions on unlabeled image data. We then combine the labeled data and pseudo labeled data to train a student model. The process is iterated by putting the student model as a new teacher. By using the proposed self-training approach and only 300k unlabeled extra data, we are able to get competitive or even better performances compared to the models of similar model size trained with 3 million extra image data. Fengmao Lv, Fayao Liu, Guosheng Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Semantic Consistent Embedding for Domain Adaptive Zero-Shot LearningabstractUnsupervised domain adaptation has limitations when encountering label discrepancy between the source and target domains. While open-set domain adaptation approaches can address situations when the target domain has additional categories, these methods can only detect them but not further classify them. In this paper, we focus on a more challenging setting dubbed Domain Adaptive Zero-Shot Learning (DAZSL), which uses semantic embeddings of class tags as the bridge between seen and unseen classes to learn the classifier for recognizing all categories in the target domain when only the supervision of seen categories in the source domain is available. The main challenge of DAZSL is to perform knowledge transfer across categories and domain styles simultaneously. To this end, we propose a novel end-to-end learning mechanism dubbed Three-way Semantic Consistent Embedding (TSCE) to embed the source domain, target domain, and semantic space into a shared space. Specifically, TSCE learns domain-irrelevant categorical prototypes from the semantic embedding of class tags and uses them as the pivots of the shared space. The source domain features are aligned with the prototypes via their supervised information. On the other hand, the mutual information maximization mechanism is introduced to push the target domain features and prototypes towards each other. By this way, our approach can align domain differences between source and target images, as well as promote knowledge transfer towards unseen classes. Moreover, as there is no supervision in the target domain, the shared space may suffer from the catastrophic forgetting problem. Hence, we further propose a ranking-based embedding alignment mechanism to maintain the consistency between the semantic space and the shared space. Experimental results on both I2AwA and I2WebV clearly validate the effectiveness of our method. Code is available at https://github.com/tiggers23/TSCE-Domain-Adaptive-Zero-Shot-Learning. Jianyang Zhang, Guowu Yang, Ping Hu 0001, Guosheng Lin, Fengmao Lv |
IEEE Trans. Image Process. | 5 |
| 2023 | C2IMUFS: Complementary and Consensus Learning-Based Incomplete Multi-View Unsupervised Feature SelectionabstractMulti-view unsupervised feature selection (MUFS) has been demonstrated as an effective technique to reduce the dimensionality of multi-view unlabeled data. The existing methods assume that all of views are complete. However, multi-view data are usually incomplete, i.e., a part of instances are presented on some views but not all views. Besides, learning the complete similarity graph, as an important promising technology in existing MUFS methods, cannot achieve due to the missing views. In this paper, we propose a complementary and consensus learning-based incomplete multi-view unsupervised feature selection method (C$^{2}$IMUFS) to address the aforementioned issues. Concretely, C$^{2}$IMUFS integrates feature selection into an extended weighted non-negative matrix factorization model equipped with adaptive learning of view-weights and a sparse$\ell _{2,p}$-norm, which can offer better adaptability and flexibility. By the sparse linear combinations of multiple similarity matrices derived from different views, a complementary learning-guided similarity matrix reconstruction model is presented to obtain the complete similarity graph in each view. Furthermore, C$^{2}$IMUFS learns a consensus clustering indicator matrix across different views and embeds it into a spectral graph term to preserve the local geometric structure. Comprehensive experimental results on real-world datasets demonstrate the effectiveness of C$^{2}$IMUFS compared with state-of-the-art methods. Yanyong Huang, Zongxin Shen, Yuxin Cai 0001, Xiuwen Yi, Dongjie Wang 0001, Fengmao Lv, Tianrui Li 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | Expanding Large Pre-trained Unimodal Models with Multimodal Information Injection for Image-Text Multimodal ClassificationabstractFine-tuning pre-trained models for downstream tasks is mainstream in deep learning. However, the pre-trained models are limited to be fine-tuned by data from a specific modality. For example, as a visual model, DenseNet cannot directly take the textual data as its input. Hence, although the large pre-trained models such as DenseNet or BERT have a great potential for the downstream recognition tasks, they have weaknesses in leveraging multimodal information, which is a new trend of deep learning. This work focuses on fine-tuning pre-trained unimodal models with multimodal inputs of image-text pairs and expanding them for image-text multimodal recognition. To this end, we propose the Multimodal Information Injection Plug-in (MI2P) which is attached to different layers of the unimodal models (e.g., DenseNet and BERT). The proposed MI2P unit provides the path to integrate the information of other modalities into the unimodal models. Specifically, MI2P performs cross-modal feature transformation by learning the fine-grained correlations between the visual and textual features. Through the proposed MI2P unit, we can inject the language information into the vision backbone by attending the word-wise textual features to different visual channels, as well as inject the visual information into the language backbone by attending the channel-wise visual features to different textual words. Armed with the MI2P attachments, the pre-trained unimodal models can be expanded to process multimodal data without the need to change the network structures. Guosheng Lin, Mingyang Wan, Tianrui Li 0001, Guojun Ma, Fengmao Lv |
CVPR | 6 |
| 2022 | Curriculum Knowledge Distillation for Emoji-supervised Cross-lingual Sentiment AnalysisabstractExisting sentiment analysis models have achieved great advances with the help of sufficient sentiment annotations.Unfortunately, many languages do not have sufficient sentiment corpus.To this end, recent studies have proposed cross-lingual sentiment analysis to transfer sentiment analysis models from resource-rich languages to low-resource languages.However, these studies either rely on external cross-lingual supervision (e.g., parallel corpora and translation model), or are limited by the cross-lingual gaps.In this work, based on the intuitive assumption that the relationships between emojis and sentiments are consistent across different languages, we investigate transferring sentiment knowledge across languages with the help of emojis.To this end, we propose a novel cross-lingual sentiment analysis approach dubbed Curriculum Knowledge Distiller (CKD).The core idea of CKD is to use emojis to bridge the source and target languages.Note that, compared with texts, emojis are more transferable, but cannot reveal the precise sentiment.Thus, we distill multiple Intermediate Sentiment Classifiers (ISC) on source language corpus with emojis to get ISCs with different attention weights of texts.To transfer them into the target language, we distill ISCs into the Target Language Sentiment Classifier (TSC) following the curriculum learning mechanism.In this way, TSC can learn delicate sentiment knowledge, meanwhile, avoid being affected by cross-lingual gaps.Experimental results on five cross-lingual benchmarks clearly verify the effectiveness of our approach. Jianyang Zhang, Mingyang Wan, Guowu Yang, Fengmao Lv |
EMNLP | 5 |
| 2022 | Partial Label Learning with Semantic Label RepresentationsabstractPartial-label learning (PLL) solves the problem where each training instance is assigned a candidate label set, among which only one is the ground-truth label. The core of PLL is to learn efficient feature representations to facilitate label disambiguation. However, existing PLL methods only learn plain representations by coarse supervision, which is incapable of capturing sufficiently distinguishable representations, especially when confronted with the knotty label ambiguity, i.e., certain candidate labels share similar visual patterns. In this paper, we propose a novel framework partial label learning with semantic label representations dubbed ParSE, which consists of two synergistic processes, including visual-semantic representation learning and powerful label disambiguation. In the former process, we propose a novel weighted calibration rank loss that has two implications. First, it implies a progressive calibration strategy that utilizes the disambiguated label confidence to weight the similarity between each image feature embedding and its corresponding semantic label representations of all candidates. Second, it also considers the ranking relationship between candidate and non-candidate ones. Based on learned visual-semantic representations, subsequent label disambiguation is desirably endowed with more powerful abilities. Experiments on benchmarks show that ParSE outperforms state-of-the-art counterparts. Shuo He 0001, Lei Feng 0006, Fengmao Lv, Wen Li 0001, Guowu Yang |
KDD | 3 |
| 2022 | Multi-category classification with label noise by robust binary loss
Defu Liu 0001, Guowu Yang, Fengmao Lv |
Neurocomputing | 5 |
| 2022 | Incorporating multiple cluster centers for multi-label learning
Senlin Shu, Fengmao Lv, Li Li 0006, Shuo He 0001, Jun He 0012 |
Inf. Sci. | 2 |
| 2022 | Weakly Supervised Domain Adaptation for Aspect Extraction via Multilevel Interaction TransferabstractFine-grained aspect term extraction is an essential subtask in aspect-based opinion analysis. It aims to identify the aspect terms (also known as opinion targets) of a product or service in each sentence. To learn a good aspect extraction model, an expensive annotation process is usually involved to acquire sufficient token-level labels for each domain, which is not realistic. To address this limitation, some previous works propose domain adaptation strategies to transfer knowledge from a sufficiently labeled source domain to unlabeled target domains. However, due to both the difficulty of fine-grained prediction problems and the large domain gap between different domains, the performance is still far from satisfactory. In this work, we conduct a pioneer study on leveraging sentence-level aspect category labels that can be usually available in commercial services, such as review sites or social media to promote token-level transfer for extraction purpose. Specifically, the aspect category information can be used to construct pivot knowledge for transfer with the assumption that the interactions between the sentence-level aspect category and the token-level aspect terms are invariant across domains. To this end, we propose a novel multilevel reconstruction mechanism that aligns both the fine- and coarse-grained information in multiple levels of abstractions. Comprehensive experiments over several benchmark data sets clearly demonstrate that our approach can fully utilize the sentence-level aspect category labels to improve cross-domain aspect term extraction with a large performance gain. Wenya Wang 0001, Fengmao Lv |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Progressive Modality Reinforcement for Human Multimodal Emotion Recognition From Unaligned Multimodal SequencesabstractHuman multimodal emotion recognition involves time-series data of different modalities, such as natural language, visual motions, and acoustic behaviors. Due to the variable sampling rates for sequences from different modalities, the collected multimodal streams are usually unaligned. The asynchrony across modalities increases the difficulty on conducting efficient multimodal fusion. Hence, this work mainly focuses on multimodal fusion from unaligned multimodal sequences. To this end, we propose the Progressive Modality Reinforcement (PMR) approach based on the recent advances of crossmodal transformer. Our approach introduces a message hub to exchange information with each modality. The message hub sends common messages to each modality and reinforces their features via crossmodal attention. In turn, it also collects the reinforced features from each modality and uses them to generate a reinforced common message. By repeating the cycle process, the common message and the modalities’ features can progressively complement each other. Finally, the reinforced features are used to make predictions for human emotion. Comprehensive experiments on different human multimodal emotion recognition benchmarks clearly demonstrate the superiority of our approach. Fengmao Lv, Yanyong Huang, Lixin Duan, Guosheng Lin |
CVPR | 1 |
| 2021 | Robust Binary Loss for Multi-Category Classification with Label NoiseabstractDeep learning has achieved tremendous success in image classification. However, the corresponding performance leap relies heavily on large-scale accurate annotations, which are usually hard to collect in reality. It is essential to explore methods that can train deep models effectively under label noise. To address the problem, we propose to train deep models with robust binary loss functions. To be specific, we tackle the K-class classification task by using K binary classifiers. We can immediately use multi-category large margin classification approaches, e.g., Pairwise-Comparison (PC) or One-Versus-All (OVA), to jointly train the binary classifiers for multi-category classification. Our method can be robust to label noise if symmetric functions, e.g., the sigmoid loss or the ramp loss, are employed as the binary loss function in the framework of risk minimization. The learning theory reveals that our method can be inherently tolerant to label noise in multi-category classification tasks. Extensive experiments over different datasets with different types of label noise are conducted. The experimental results clearly confirm the effectiveness of our method. Defu Liu 0001, Guowu Yang, Fengmao Lv |
ICASSP | 5 |
| 2021 | Attention is not Enough: Mitigating the Distribution Discrepancy in Asynchronous Multimodal Sequence FusionabstractVideos flow as the mixture of language, acoustic, and vision modalities. A thorough video understanding needs to fuse time-series data of different modalities for prediction. Due to the variable receiving frequency for sequences from each modality, there usually exists inherent asynchrony across the collected multimodal streams. Towards an efficient multimodal fusion from asynchronous multimodal streams, we need to model the correlations between elements from different modalities. The recent Multimodal Transformer (MulT) approach extends the self-attention mechanism of the original Transformer network to learn the crossmodal dependencies between elements. However, the direct replication of self-attention will suffer from the distribution mismatch across different modality features. As a result, the learnt crossmodal dependencies can be unreliable. Motivated by this observation, this work proposes the Modality-Invariant Crossmodal Attention (MICA) approach towards learning crossmodal interactions over modality-invariant space in which the distribution mismatch between different modalities is well bridged. To this end, both the marginal distribution and the elements with high-confidence correlations are aligned over the common space of the query and key vectors which are computed from different modalities. Experiments on three standard benchmarks of multimodal video understanding clearly validate the superiority of our approach. Guosheng Lin, Lei Feng 0006, Yan Zhang 0036, Fengmao Lv |
ICCV | 5 |
| 2021 | Long short-term memory on abstract syntax tree for SQL injection detectionabstractAbstract SQL injection attack (SQLIA) is a code injection technique, used to attack data‐driven applications by executing malicious SQL statements. Techniques like pattern matching, software testing and grammar analysis etc. are frequently used to prevent such attack. However, major bottlenecks still remain in detecting SQLIA with bypassing techniques, getting access to source code and requiring an additional manual operation to extract features. The authors propose a novel detection approach based on long short‐term memory and abstract syntax tree, which could detect SQLIAs from the raw query strings and work under SQL detection bypassing scenario. Our deep learning technique explicitly uses both context and syntax information that previous methods failed to fully grasp. Experimental results clearly illustrate the superior performance of our method compared to other existing works when detecting with complete SQL raw queries. Zhongliu Zhuo, T. Cai, Xiaosong Zhang 0001, Fengmao Lv |
IET Softw. | 4 |
| 2021 | Adaptive graph-based generalized regression model for unsupervised feature selection
Yanyong Huang, Zongxin Shen, Fuxu Cai, Tianrui Li 0001, Fengmao Lv |
Knowl. Based Syst. | 5 |
| 2021 | Latent Gaussian process for anomaly detection in categorical data
Fengmao Lv, Zhongliu Zhuo, Guowu Yang |
Knowl. Based Syst. | 1 |
| 2021 | Weakly-Supervised Cross-Domain Road Scene Segmentation via Multi-Level Curriculum AdaptationabstractSemantic segmentation, which aims to acquire pixel-level understanding about images, is among the key components in computer vision. To train a good segmentation model for real-world images, it usually requires a huge amount of time and labor effort to obtain sufficient pixel-level annotations of real-world images beforehand. To get rid of such a nontrivial burden, one can use simulators to automatically generate synthetic images that inherently contain full pixel-level annotations and use them to train a segmentation model for the real-world images. However, training with synthetic images usually cannot lead to good performance due to the domain difference between the synthetic images (i.e., source domain) and the real-world images (i.e., target domain). To deal with this issue, a number of unsupervised domain adaptation (UDA) approaches have been proposed, where no labeled real-world images are available. Different from those methods, in this work, we conduct a pioneer attempt by using easy-to-collect image-level annotations for target images to improve the performance of cross-domain segmentation. Specifically, we leverage those image-level annotations to construct curriculums for the domain adaptation problem. The curriculums describe multi-level properties of the target domain, including label distributions over full images, local regions and single pixels. Since image annotations are “weak” labels compared to pixel annotations for segmentation, we coin this new problem as weakly-supervised cross-domain segmentation. Comprehensive experiments on the GTA5→ Cityscapes and SYNTHIA→ Cityscapes settings demonstrate the effectiveness of our method over the existing state-of-the-art baselines. Fengmao Lv, Guosheng Lin, Peng Liu 0049, Guowu Yang, Sinno Jialin Pan, Lixin Duan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Cross-Domain Semantic Segmentation via Domain-Invariant Interactive Relation TransferabstractExploiting photo-realistic synthetic data to train semantic segmentation models has received increasing attention over the past years. However, the domain mismatch between synthetic and real images will cause a significant performance drop when the model trained with synthetic images is directly applied to real-world scenarios. In this paper, we propose a new domain adaptation approach, called Pivot Interaction Transfer (PIT). Our method mainly focuses on constructing pivot information that is common knowledge shared across domains as a bridge to promote the adaptation of semantic segmentation model from synthetic domains to real-world domains. Specifically, we first infer the image-level category information about the target images, which is then utilized to facilitate pixel-level transfer for semantic segmentation, with the assumption that the interactive relation between the image-level category information and the pixel-level semantic information is invariant across domains. To this end, we propose a novel multi-level region expansion mechanism that aligns both the image-level and pixel-level information. Comprehensive experiments on the adaptation from both GTAV and SYNTHIA to Cityscapes clearly demonstrate the superiority of our method. Fengmao Lv, Guosheng Lin |
CVPR | 1 |
| 2020 | TRRNet: Tiered Relation Reasoning for Compositional Visual Question Answering
Guosheng Lin, Fengmao Lv, Fayao Liu |
ECCV (21) | 3 |
| 2020 | Can Cross Entropy Loss Be Robust to Label Noise?abstractTrained with the standard cross entropy loss, deep neural networks can achieve great performance on correctly labeled data. However, if the training data is corrupted with label noise, deep models tend to overfit the noisy labels, thereby achieving poor generation performance. To remedy this issue, several loss functions have been proposed and demonstrated to be robust to label noise. Although most of the robust loss functions stem from Categorical Cross Entropy (CCE) loss, they fail to embody the intrinsic relationships between CCE and other loss functions. In this paper, we propose a general framework dubbed Taylor cross entropy loss to train deep models in the presence of label noise. Specifically, our framework enables to weight the extent of fitting the training labels by controlling the order of Taylor Series for CCE, hence it can be robust to label noise. In addition, our framework clearly reveals the intrinsic relationships between CCE and other loss functions, such as Mean Absolute Error (MAE) and Mean Squared Error (MSE). Moreover, we present a detailed theoretical analysis to certify the robustness of this framework. Extensive experimental results on benchmark datasets demonstrate that our proposed approach significantly outperforms the state-of-the-art counterparts. Lei Feng 0006, Senlin Shu, Zhuoyi Lin, Fengmao Lv, Li Li 0006, Bo An 0001 |
IJCAI | 4 |
| 2020 | Learning Unbiased Zero-Shot Semantic Segmentation Networks Via Transductive TransferabstractSemantic segmentation aims to obtain a detailed understanding of images. Deep learning has achieved great advances in semantic segmentation over the past years. In practice, however, the classes do not always correspond to the ones in the training stage. Since it is impractical to collect sufficient labeled data for all classes, zero-shot semantic segmentation has received increasing attentions recently. Although semantic segmentation neural networks can transfer knowledge from seen classes to unseen classes by incorporating the class-level semantic information, it shows a strong bias towards seen classes. In this letter, we propose an easy-to-implement transductive approach to alleviate the prediction bias in zero-shot semantic segmentation. We assume that both source images with full pixel-level labels and unlabeled target images are available for training. The source images are used to build the relationship between visual images and class-level semantic embeddings. On the other hand, the target images are used to alleviate the bias towards seen classes. Comprehensive experiments over the PASCAL dataset clearly demonstrate the effectiveness of our approach. Fengmao Lv, Guowu Yang |
IEEE Signal Process. Lett. | 1 |
| 2019 | Constructing Self-Motivated Pyramid Curriculums for Cross-Domain Semantic Segmentation: A Non-Adversarial ApproachabstractWe propose a new approach, called self-motivated pyramid curriculum domain adaptation (PyCDA), to facilitate the adaptation of semantic segmentation neural networks from synthetic source domains to real target domains. Our approach draws on an insight connecting two existing works: curriculum domain adaptation and self-training. Inspired by the former, PyCDA constructs a pyramid curriculum which contains various properties about the target domain. Those properties are mainly about the desired label distributions over the target domain images, image regions, and pixels. By enforcing the segmentation neural network to observe those properties, we can improve the network's generalization capability to the target domain. Motivated by the self-training, we infer this pyramid of properties by resorting to the semantic segmentation network itself. Unlike prior work, we do not need to maintain any additional models (e.g., logistic regression or discriminator networks) or to solve minmax problems which are often difficult to optimize. We report state-of-the-art results for the adaptation from both GTAV and SYNTHIA to Cityscapes, two popular settings in unsupervised domain adaptation for semantic segmentation. Qing Lian, Lixin Duan, Fengmao Lv, Boqing Gong |
ICCV | 3 |
| 2019 | TarGAN: Generating target data with class labels for unsupervised domain adaptation
Fengmao Lv, Guowu Yang, Lixin Duan |
Knowl. Based Syst. | 1 |
| 2019 | Integrating Traffics with Network Device Logs for Anomaly DetectionabstractAdvanced cyberattacks are often featured by multiple types, layers, and stages, with the goal of cheating the monitors. Existing anomaly detection systems usually search logs or traffics alone for evidence of attacks but ignore further analysis about attack processes. For instance, the traffic detection methods can only detect the attack flows roughly but fail to reconstruct the attack event process and reveal the current network node status. As a result, they cannot fully model the complex multistage attack. To address these problems, we present Traffic-Log Combined Detection (TLCD), which is a multistage intrusion analysis system. Inspired by multiplatform intrusion detection techniques, we integrate traffics with network device logs through association rules. TLCD correlates log data with traffic characteristics to reflect the attack process and construct a federated detection platform. Specifically, TLCD can discover the process steps of a cyberattack attack, reflect the current network status, and reveal the behaviors of normal users. Our experimental results over different cyberattacks demonstrate that TLCD works well with high accuracy and low false positive rate. Fengmao Lv, Zhongliu Zhuo, Xiaosong Zhang 0001, Xiaolei Liu 0001, Wei Deng 0003 |
Secur. Commun. Networks | 2 |
| 2018 | Boosting 1H-MRS Alzheimer Diagnosis with Boosted Trees
Fengmao Lv, Guowu Yang |
BIBM | 1 |
| 2018 | Improving Target Discriminability for Unsupervised Domain Adaptation
Fengmao Lv, Linfeng Zhong, Xiaoyu Li 0003, Guowu Yang |
ICONIP (5) | 1 |
| 2018 | Botnet Detection based on Fuzzy Association RulesabstractDifficult to be detected in complex network environments, botnets have been huge threats to network security. As the circumscriptions of normal traffics and botnet traffics are blurring, the commonly used botnet detection methods based on traffic analysis often result in high false positive rates. To overcome this issue, we propose an effective botnet detection method based on fuzzy association rules. The proposed method can calculate the features of botnet traffic accurately, which can be used to recognize the normal traffic and botnet. We first collect the data in the laboratory by setting different botnets in the controlled experiment. The botnet traffic features, association rules support, trust and membership are calculated by the proposed method, which are further used to distinguish the type of botnet. When our method is compared with other methods in our data set, we find the former performs better. For the generality, we also test our method on the public data set and also find the higher accuracy rates, which demonstrates the proposed method is effective in detecting the botnets. Fengmao Lv, Quanhui Liu, Malu Zhang, Xiaosong Zhang 0001 |
ICPR | 2 |
| 2017 | Anomaly Detection for Categorical Observations Using Latent Gaussian Process
Fengmao Lv, Guowu Yang, Yuhong Yang 0002 |
ICONIP (5) | 1 |
| 2017 | The convergence and termination criterion of quantum-inspired evolutionary neural networks
Fengmao Lv, Guowu Yang, Wenjing Yang 0003, Xiaosong Zhang 0001, Kenli Li 0001 |
Neurocomputing | 1 |
| 2017 | A new Centroid-Based Classification model for text categorizationabstractThe automatic text categorization technique has gained significant attention among researchers because of the increasing availability of online text information. Therefore, many different learning approaches have been designed in the text categorization field. Among them, the widely used method is the Centroid-Based Classifier (CBC) due to its theoretical simplicity and computational efficiency. However, the classification accuracy of CBC greatly depends on the data distribution. Thus it leads to a misfit model and also has poor classification performance when the data distribution is highly skewed. In this paper, a new classification model named as Gravitation Model (GM) is proposed to solve the class-imbalanced classification problem. In the training phase, each class is weighted by a mass factor, which can be learned from the training data, to indicate data distribution of the corresponding class. In the testing phase, a new document will be assigned to a particular class with the max gravitational force. The performance comparisons with CBC and its variants based on the results of experiments conducted on twelve real datasets show that the proposed gravitation model consistently outperforms CBC together with the Class-Feature-Centroid Classifier (CFC). Also, it obtains the classification accuracy competitive to the DragPushing (DP) method while it maintains a more stable performance. Thus, the proposed gravitation model is proved to be less over-fitting and has higher learning ability than CBC model. Guanghui Tu, Yu Xiang 0002, Fengmao Lv |
Knowl. Based Syst. | 6 |
| 2017 | An efficient instance selection algorithm to reconstruct training set for support vector machineabstractSupport vector machine is a classification model which has been widely used in many nonlinear and high dimensional pattern recognition problems. However, it is inefficient or impracticable to implement support vector machine in dealing with large scale training set due to its computational difficulties as well as the model complexity. In this paper, we study the support vector recognition problem mainly in the context of the reduction methods to reconstruct training set for support vector machine. We focus on the fact of uneven distribution of instances in the vector space to propose an efficient self-adaption instance selection algorithm from the viewpoint of geometry-based method. Also, we conduct an experimental study involving eleven different sizes of datasets from UCI repository for measuring the performance of the proposed algorithm as well as six competitive instance selection algorithms in terms of accuracy, reduction capabilities, and runtime. The extensive experimental results show that the proposed algorithm outperforms most of competitive algorithms due to its high efficiency and efficacy. Fengmao Lv, Martin Konan |
Knowl. Based Syst. | 4 |
| 2017 | Generative classification model for categorical data based on latent Gaussian process
Fengmao Lv, Guowu Yang, William Zhu 0001 |
Pattern Recognit. Lett. | 1 |
| 2014 | The Research on Controlling the Iteration of Quantum-Inspired Evolutionary Algorithms for Artificial Neural Networks
Fengmao Lv, Guowu Yang, Shuangbao Paul Wang, Fuyou Fan |
AAIM | 1 |