Luntian Mou

dblp:82/1373 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
10since 2021 · last 2025
0000-0002-1551-4448ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Cross-Rejective Open-Set SAR Image Registration
abstract
Synthetic Aperture Radar (SAR) image registration is an essential upstream task in geoscience applications, in which pre-detected keypoints from two images are employed as observed objects to seek matched-point pairs. In general, the registration is regarded as a typical closed-set classification, which forces each keypoint to be classified into the given classes, but ignoring an essential issue that numerous redundant keypoints are beyond the given classes, which unavoidably results in capturing incorrect matched-point pairs. Based on this, we propose a Cross-Rejective Open-set SAR Image Registration (CroR-OSIR) method. In this work, these redundant keypoints are regarded as out-of-distribution (OOD) samples, and we formulate the registration as a special open-set task with two modules: supervised contrastive feature-tuning and cross-rejective open-set recognition (CroR-OSR). Unlike traditional open-set recognition, all samples, including OOD samples, are available in the CroR-OSR module. CroR-OSR conducts the closed-set classifications in individual open-set domains from two images, meanwhile employing the cross-domain rejection during training, to exclude these OOD samples based on confidence and consistency. Moreover, a new supervised contrastive tuning strategy is incorporated for feature-tuning. Especially, the cross-domain estimation labels obtained by CroR-OSR are fed back to the feature-tuning module for feature-tuning, to enhance feature discriminability. The experimental results illustrate that the proposed method achieves more precise registration than the state-of-the-art methods. The code is released at https://github.com/XDyaoshi/CroR-OSIR-main.
Shasha Mao, Shiming Lu, Zhaolong Du, Licheng Jiao, Shuiping Gou, Luntian Mou, Xuequan Lu
CVPR6
2025 Image-Based Freeform Handwriting Authentication With Energy-Oriented Self-Supervised Learning
abstract
Freeform handwriting authentication verifies a person's identity from their writing style and habits in messy handwriting data. This technique has gained widespread attention in recent years as a valuable tool for various fields, e.g., fraud prevention and cultural heritage protection. However, it still remains a challenging task in reality due to three reasons: (i) severe damage, (ii) complex high-dimensional features, and (iii) lack of supervision. To address these issues, we propose SherlockNet, an energy-oriented two-branch contrastive self-supervised learning framework for robust and fast freeform handwriting authentication. It consists of four stages: (i) pre-processing: converting manuscripts into energy distributions using a novel plug-and-play energy-oriented operator to eliminate the influence of noise; (ii) generalized pre-training: learning general representation through two-branch momentum-based adaptive contrastive learning with the energy distributions, which handles the high-dimensional features and spatial dependencies of handwriting; (iii) personalized fine-tuning: calibrating the learned knowledge using a small amount of labeled data from downstream tasks; and (iv) practical application: identifying individual handwriting from scrambled, missing, or forged data efficiently and conveniently. Considering the practicality, we construct EN-HA, a novel dataset that simulates data forgery and severe damage in real applications. Finally, we conduct extensive experiments on six benchmark datasets including our EN-HA, and the results prove the robustness and efficiency of SherlockNet.
Luntian Mou, Changwen Zheng, Wen Gao 0001
IEEE Trans. Multim.2
2024 E/I Balanced Adaptive Sequential Neural Posterior Estimation for Inferring the Connection Weights in Mouse V1 Model
abstract
Effectively utilizing biological firing rate data to estimate the numerous connection weights in the mouse primary visual cortex (V1) model from the Allen Institute is a challenging task. The existing iterative grid-search algorithm cannot enable the mouse V1 model to better fit the biological firing rate data. To tackle this issue, we propose an excitation-inhibition balanced adaptive sequential neural posterior estimation (E/I balanced ASNPE) approach to accurately infer the connection weights of the mouse V1 model, allowing the neurons’ firing rates to converge to the given biological data. This method fully leverages the structural information of the mouse V1 model, reducing the dimensionality of the weight parameters to be optimized. Initially, sampling is performed in the prior distribution based on the proposed non-dominated sorting adaptive genetic algorithm (NSAGA). This algorithm optimizes the sorting, crossover and mutation processes based on the fitness scores of the current samples and updates the proposal distribution based on these samples, increasing the likelihood of identifying high posterior probability regions in the prior distribution. To avoid bad simulations, we also explore the E/I balance in each layer of the mouse V1 model, adding biological constraints during weight inference with Automatic Posterior Transformation (APT). Experimental results confirm that the proposed E/I balanced ASNPE method significantly outperforms the baseline in all five firing rate fitness scores in the mouse V1 model. This study is pioneering in applying non-dominated sorting genetic algorithms combined with sequential neural posterior estimation to optimize connection weights in large-scale complex biological models.
Luntian Mou, Peize Li, Lei Ma 0008, Tiejun Huang 0001
BIBM1
2024 Image-Based Structured Vehicle Behavior Analysis Inspired by Interactive Cognition
abstract
Vehicle behavior analysis has gradually developed by utilizing trajectories and motion features to characterize on-road behavior. However, the existing methods analyze the behavior of each vehicle individually, ignoring the interaction between vehicles. According to the theory of interactive cognition, vehicle-to-vehicle interaction is an indispensable feature for future autonomous driving, just as interaction is universally required for traditional driving. Therefore, we place the vehicle behavior analysis in the context of the vehicle interaction scene, where the self-vehicle should observe the behavior category and degree of the other-vehicle that is about to interact with itself, in order to predict whether the other-vehicle will pass through the intersection first or later, and then decide to pass through or wait. Inspired by the interactive cognition, we develop a general framework of Structured Vehicle Behavior Analysis (StruVBA) and derive a new model of Structured Fully Convolutional Networks (StruFCN). Moreover, both Intersection over Union (IoU) and False Negative Rate (FNR) are adopted to measure the similarity between the predicted behavior degree and the ground truth. Experimental results illustrate that the proposed method achieves higher prediction accuracy than most existing methods, while predicting vehicle behavior with richer visual meaning. In addition, it also provides an example of modeling the interaction between vehicles and a verification for interaction cognition theory as well.
Luntian Mou, Haitao Xie, Shasha Mao, Nan Ma 0012, Wen Gao 0001
IEEE Trans. Multim.1
2024 Accurate Registration of Cross-Modality Geometry via Consistent Clustering
abstract
The registration of unitary-modality geometric data has been successfully explored over past decades. However, existing approaches typically struggle to handle cross-modality data due to the intrinsic difference between different models. To address this problem, in this article, we formulate the cross-modality registration problem as a consistent clustering process. First, we study the structure similarity between different modalities based on an adaptive fuzzy shape clustering, from which a coarse alignment is successfully operated. Then, we optimize the result using fuzzy clustering consistently, in which the source and target models are formulated as clustering memberships and centroids, respectively. This optimization casts new insight into point set registration, and substantially improves the robustness against outliers. Additionally, we investigate the effect of fuzzier in fuzzy clustering on the cross-modality registration problem, from which we theoretically prove that the classical Iterative Closest Point (ICP) algorithm is a special case of our newly defined objective function. Comprehensive experiments and analysis are conducted on both synthetic and real-world cross-modality datasets. Qualitative and quantitative results demonstrate that our method outperforms state-of-the-art approaches with higher accuracy and robustness. Our code is publicly available at https://github.com/zikai1/CrossModReg.
Mingyang Zhao 0001, Xiaoshui Huang, Jingen Jiang 0001, Luntian Mou, Dong-Ming Yan 0001, Lei Ma 0008
IEEE Trans. Vis. Comput. Graph.4
2023 Multimodal driver distraction detection using dual-channel network of CNN and Transformer
Luntian Mou, Jiali Chang, Yiyuan Zhao, Nan Ma 0008, Ramesh Jain 0001, Wen Gao 0001
Expert Syst. Appl.1
2023 Driver Emotion Recognition With a Hybrid Attentional Multimodal Fusion Framework
abstract
Negative emotions may induce dangerous driving behaviors leading to extremely serious traffic accidents. Therefore, it is necessary to establish a system that can automatically recognize driver emotions so that some actions can be taken to avoid traffic accidents. Existing studies on driver emotion recognition have mainly used facial data and physiological data. However, there are fewer studies on multimodal data with contextual characteristics of driving. In addition, fully fusing multimodal data in the feature fusion layer to improve the performance of emotion recognition is still a challenge. To this end, we propose to recognize driver emotion using a novel multimodal fusion framework based on convolutional long-short term memory network (ConvLSTM), and hybrid attention mechanism to fuse non-invasive multimodal data of eye, vehicle, and environment. In order to verify the effectiveness of the proposed method, extensive experiments have been carried out on a dataset collected using an advanced driving simulator. The experimental results demonstrate the effectiveness of the proposed method. Finally, a preliminary exploration on the correlation between driver emotion and stress is performed.
Luntian Mou, Yiyuan Zhao, Bahareh Nakisa, Mohammad Naim Rastgoo, Lei Ma 0008, Tiejun Huang 0001, Ramesh Jain 0001, Wen Gao 0001
IEEE Trans. Affect. Comput.1
2023 Isotropic Self-Supervised Learning for Driver Drowsiness Detection With Attention-Based Multimodal Fusion
abstract
Driverdrowsiness is an important cause of traffic accidents. Many studies using computer vision techniques to detect driver drowsiness states, such as slow blinking, yawning, and nodding, have demonstrated excellent potential. Although existing studies have made significant progress, the number of samples in the training corpora is small, which makes it difficult for a model to learn effective drowsiness representations from images or videos. To address this issue, we develop an isotropic self-supervised learning (IsoSSL) approach to learn powerful representations of images without relying on human-provided annotations and propose an IsoSSL-MoCo model by combining IsoSSL with momentum contrast (MoCo). To exploit the complementarity of multimodal data, an attention-based multimodal fusion model is also proposed to fuse features from the eye, mouth, and optical flow of the head. Specifically, we first use the IsoSSL-MoCo model to pretrain the image encoders for the three modalities in other datasets. Then, these encoders are fine-tuned and integrated into the proposed fusion model. The feature vectors generated by the image encoders of the three modalities are fed into the recursive layer to extract temporal information. To capture the importance degrees of the effects of temporal features from the three modalities on drowsiness detection, an attention mechanism is introduced to automatically weigh the feature vectors from the recursive layer to improve detection accuracy. Finally, a vector representation is generated by the attention layer and is used to detect driver drowsiness states. Experimental results based on two challenging datasets show that our method outperforms the baseline methods and the latest existing methods.
Luntian Mou, Pengtao Xie, Pengfei Zhao 0008, Ramesh Jain 0001, Wen Gao 0001
IEEE Trans. Multim.1
2023 AMSA: Adaptive Multimodal Learning for Sentiment Analysis
abstract
Efficient recognition of emotions has attracted extensive research interest, which makes new applications in many fields possible, such as human-computer interaction, disease diagnosis, service robots, and so forth. Although existing work on sentiment analysis relying on sensors or unimodal methods performs well for simple contexts like business recommendation and facial expression recognition, it does far below expectations for complex scenes, such as sarcasm, disdain, and metaphors. In this article, we propose a novel two-stage multimodal learning framework, called AMSA, to adaptively learn correlation and complementarity between modalities for dynamic fusion, achieving more stable and precise sentiment analysis results. Specifically, a multiscale attention model with a slice positioning scheme is proposed to get stable quintuplets of sentiment in images, texts, and speeches in the first stage. Then a Transformer-based self-adaptive network is proposed to assign weights flexibly for multimodal fusion in the second stage and update the parameters of the loss function through compensation iteration. To quickly locate key areas for efficient affective computing, a patch-based selection scheme is proposed to iteratively remove redundant information through a novel loss function before fusion. Extensive experiments have been conducted on both machine weakly labeled and manually annotated datasets of self-made Video-SA, CMU-MOSEI, and CMU-MOSI. The results demonstrate the superiority of our approach through comparison with baselines.
Luntian Mou, Lei Ma 0008, Tiejun Huang 0001, Wen Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Driver stress detection via multimodal fusion using attention-based CNN-LSTM
Luntian Mou, Pengfei Zhao 0008, Bahareh Nakisa, Mohammad Naim Rastgoo, Ramesh Jain 0001, Wen Gao 0001
Expert Syst. Appl.1
2017 Deep Determinantal Point Process for Large-Scale Multi-label Classification
abstract
We study large-scale multi-label classification (MLC) on two recently released datasets: Youtube-8M and Open Images that contain millions of data instances and thousands of classes. The unprecedented problem scale poses great challenges for MLC. First, finding out the correct label subset out of exponentially many choices incurs substantial ambiguity and uncertainty. Second, the large data-size and class-size entail considerable computational cost. To address the first challenge, we investigate two strategies: capturing label-correlations from the training data and incorporating label co-occurrence relations obtained from external knowledge, which effectively eliminate semantically inconsistent labels and provide contextual clues to differentiate visually ambiguous labels. Specifically, we propose a Deep Determinantal Point Process (DDPP) model which seamlessly integrates a DPP with deep neural networks (DNNs) and supports end-to-end multi-label learning and deep representation learning. The DPP is able to capture label-correlations of any order with a polynomial computational cost, while the DNNs learn hierarchical features of images/videos and capture the dependency between input data and labels. To incorporate external knowledge about label co-occurrence relations, we impose relational regularization over the kernel matrix in DDPP. To address the second challenge, we study an efficient low-rank kernel learning algorithm based on inducing point methods. Experiments on the two datasets demonstrate the efficacy and efficiency of the proposed methods.
Pengtao Xie, Ruslan Salakhutdinov, Luntian Mou, Eric P. Xing
ICCV3
2013 MPLBoost-based mixture model for effective human detection with Deformable Part Model
abstract
The Deformable Part Model has shown high accuracy in tackling certain occlusion or deformations of objects such as cars and bikes. However, as for human category characterized by a larger number of articulated parts and more significant appearance variations, its performance gain is not so remarkable. To address this issue, we propose an MPLBoost-based mixture model which splits data into coherent groups and trains one root classifier for each, resulting in automated selection of discriminative root models and better representation of intra-class variations through visual feature clustering. Based on this boosting framework, multiple complementary features are combined to capture shape, texture and color information. Experimental results demonstrate that the proposed model can achieve an impressive performance improvement, especially in handling larger variations of human poses and viewpoints.
Chaoran Gu, Luntian Mou, Yonghong Tian 0001, Tiejun Huang 0001
ICME2
2013 Content-based copy detection through multimodal feature representation and temporal pyramid matching
abstract
Content-based copy detection (CBCD) is drawing increasing attention as an alternative technology to watermarking for video identification and copyright protection. In this article, we present a comprehensive method to detect copies that are subjected to complicated transformations. A multimodal feature representation scheme is designed to exploit the complementarity of audio features, global and local visual features so that optimal overall robustness to a wide range of complicated modifications can be achieved. Meanwhile, a temporal pyramid matching algorithm is proposed to assemble frame-level similarity search results into sequence-level matching results through similarity evaluation over multiple temporal granularities. Additionally, inverted indexing and locality sensitive hashing (LSH) are also adopted to speed up similarity search. Experimental results over benchmarking datasets of TRECVID 2010 and 2009 demonstrate that the proposed method outperforms other methods for most transformations in terms of copy detection accuracy. The evaluation results also suggest that our method can achieve competitive copy localization preciseness.
Luntian Mou, Tiejun Huang 0001, Yonghong Tian 0001, Menglin Jiang, Wen Gao 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2012 Robust and discriminative image authentication based on standard model feature
abstract
The goal of image authentication is to accept content-preserving operations and reject content-altering manipulations. So,it is increasingly approached by extracting content-based invariant features from original images and verifying their preservation in received images at later times. Since sparsity usually implies invariance, sparse feature representation has drawn significant attention from the research community. But only if discrimination is also found with a sparse feature, can it be successfully applied in image authentication. This paper proposes a sparse feature for image authentication by exploring the biologically-motivated standard model. Experimental results demonstrate both robustness and discrimination of the feature, and its effectiveness in tamper detection and location as well.
Luntian Mou, Xilin Chen 0001, Yonghong Tian 0001, Tiejun Huang 0001
ISCAS1
2011 Robust and discriminative image authentication based on sparse coding
abstract
Image authentication is usually approached by checking the preservation of some invariant features, which are expected to be both robust and discriminative so that content-preserving operations are accepted while content-altering manipulations are rejected. However, most of existing features have not obtained convincing performance due to insufficiency of experiments and over biasing of robustness. Motivated by the sparse coding strategy discovered in primary visual cortex, we explore the possibility of using sparse coding coefficients for image authentication. Through extensive experiments, we discover that the proposed feature bears great discrimination as well as robustness, which indicates the effectiveness of sparse coding as a new invariant feature for image authentication.
Luntian Mou, Tiejun Huang 0001, Yonghong Tian 0001, Shiguo Lian, Xilin Chen 0001
CCNC1
2011 A multimodal video copy detection approach with sequential pyramid matching
abstract
Content-based video copy detection over large corpus with complex transformations is important but challenging. It is not surprising that most existing methods fall short of either sufficient robustness to detect severely deformed copies or high accuracy to localize copy segments. In this paper, we propose a video copy detection approach which exploits complementary audio-visual features and sequential pyramid matching (SPM). Several independent detectors first match visual key frames or audio clips using individual features, and then aggregate the frame level results into video level results with SPM, which calculates video similarities by sequence matching at multiple granularities. Finally, detection results from basic detectors are fused and further filtered to generate the final result. Excellent performance evaluated on TRECVid 2010 copy detection task demonstrates the effectiveness of our approach.
Yonghong Tian 0001, Menglin Jiang, Luntian Mou, Xiaoyu Fang, Tiejun Huang 0001
ICIP3
2009 A secure media streaming mechanism combining encryption, authentication, and transcoding
Luntian Mou, Tiejun Huang 0001, Longshe Huo, Weiping Li 0002, Wen Gao 0001, Xilin Chen 0001
Signal Process. Image Commun.1
2007 A DRM Architecture for Manageable P2P Based IPTV System
Xiaoyun Liu, Tiejun Huang 0001, Longshe Huo, Luntian Mou
ICME4