Jupo Ma

dblp:246/5794 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0009-3631-7177ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 An event-based motion scene feature extraction framework
Zhaoxin Liu, Jinjian Wu, Guangming Shi, Wen Yang 0008, Jupo Ma
Pattern Recognit.5
2025 Defect Detection in Remote Sensing Satellite Images: A New Dataset and Algorithm
abstract
Satellite observation is an important way to understand the earth. However, due to the problems such as satellite aging, cloud obstruction, and other interferences during the imaging and transmission process, remote sensing images inevitably produce various defects. Hence, it is necessary to quickly detect defects to calibrate the imaging system and avoid the waste of satellite resource. Current researches on defect detection in remote sensing images are not comprehensive, which only focus on partial defect categories, such as cloud and stripe. To this end, we construct the first large-scale High-resolution Remote Sensing image Defect detection dataset (HRSD). The proposed dataset contains more than 1.2 million manually annotated patches from eight different satellites, covering various common defect categories and including multiple image modalities (i.e., panchromatic and multispectral). The dataset also has rich diversity which covers different landforms in multiple regions. Furthermore, to realize the detection of multiple defect categories simultaneously, we design a feature aggregation graph network (FAGN) based on the position correlation and semantic similarity among image patches, which fully utilizes the distribution characteristics of defects to achieve accurate defect detection. Extensive experiments on the HRSD dataset demonstrated the effectiveness of FAGN. We will release the HRSD dataset and FAGN model later.
Hengchao Hu, Jupo Ma, Qi Wang 0053, Yuanshi Zheng, Jinjian Wu
IEEE Trans. Geosci. Remote. Sens.3
2024 Motion Deblurring via Spatial-Temporal Collaboration of Frames and Events
abstract
Motion deblurring can be advanced by exploiting informative features from supplementary sensors such as event cameras, which can capture rich motion information asynchronously with high temporal resolution. Existing event-based motion deblurring methods neither consider the modality redundancy in spatial fusion nor temporal cooperation between events and frames. To tackle these limitations, a novel spatial-temporal collaboration network (STCNet) is proposed for event-based motion deblurring. Firstly, we propose a differential-modality based cross-modal calibration strategy to suppress redundancy for complementarity enhancement, and then bimodal spatial fusion is achieved with an elaborate cross-modal co-attention mechanism to weight the contributions of them for importance balance. Besides, we present a frame-event mutual spatio-temporal attention scheme to alleviate the errors of relying only on frames to compute cross-temporal similarities when the motion blur is significant, and then the spatio-temporal features from both frames and events are aggregated with the custom cross-temporal coordinate attention. Extensive experiments on both synthetic and real-world datasets demonstrate that our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/STCNet.
Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Guangming Shi
AAAI3
2024 A Redundancy-Suppression Based Event Sampling Method for Structured Representation
Jupo Ma, Shunhong Li
PRCV (8)1
2024 Learning Frame-Event Fusion for Motion Deblurring
abstract
Motion deblurring is a highly ill-posed problem due to the significant loss of motion information in the blurring process. Complementary informative features from auxiliary sensors such as event cameras can be explored for guiding motion deblurring. The event camera can capture rich motion information asynchronously with microsecond accuracy. In this paper, a novel frame-event fusion framework is proposed for event-driven motion deblurring (FEF-Deblur), which can sufficiently explore long-range cross-modal information interactions. Firstly, different modalities are usually complementary and also redundant. Cross-modal fusion is modeled as complementary-unique features separation-and-aggregation, avoiding the modality redundancy. Unique features and complementary features are first inferred with parallel intra-modal self-attention and inter-modal cross-attention respectively. After that, a correlation-based constraint is designed to act between unique and complementary features to facilitate their differentiation, which assists in cross-modal redundancy suppression. Additionally, spatio-temporal dependencies among neighboring inputs are crucial for motion deblurring. A recurrent cross attention is introduced to preserve inter-input attention information, in which the current spatial features and aggregated temporal features are attending to each other by establishing the long-range interaction between them. Extensive experiments on both synthetic and real-world motion deblurring datasets demonstrate our method outperforms state-of-the-art event-based and image/video-based methods. The code will be made publicly available.
Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Weisheng Dong, Guangming Shi
IEEE Trans. Image Process.3
2022 Learning for Motion Deblurring with Hybrid Frames and Events
abstract
Event camera responds to the brightness changes at each pixel independently with microsecond accuracy. Event cameras offer attractive property that can record well high-speed scene but ignore static and non-moving areas, while conventional frame cameras are able to acquire the whole intensity information of the scene but suffer from motion blur. Therefore, it would be desirable to combine the best of two cameras for reconstructing high quality intensity frame with no motion blur. The human visual system presents a two-pathway procedure for non-action-based representation and objects motion perception, which corresponds well to the hybrid frame and event. In this paper, inspired by the two-pathway visual system, a novel dual-stream based framework is proposed for motion deblurring (DS-Deblur), which flexibly utilizes the respective advantages from frame and event. A complementary-unique information splitting based feature fusion module is firstly proposed to adaptively aggregate the frame and event progressively at multiple levels, which is well-grounded on the hierarchical process in twopathway visual system. Then, a recurrent spatio-temporal feature transformation module is designed to exploit relevant information between adjacent frames, in which features of both current and previous frames are transformed in a global-local manner. Extensive experiments on both synthetic and real motion blur datasets demonstrate our method achieves state-of-the-art performance. Project website: https://github.com/wyang-vis/Motion-Deblurringwith-Hybrid-Frames-and-Events.
Wen Yang 0008, Jinjian Wu, Jupo Ma, Leida Li, Weisheng Dong, Guangming Shi
ACM Multimedia3
2021 Blind Image Quality Assessment With Active Inference
abstract
Blind image quality assessment (BIQA) is a useful but challenging task. It is a promising idea to design BIQA methods by mimicking the working mechanism of human visual system (HVS). The internal generative mechanism (IGM) indicates that the HVS actively infers the primary content (i.e., meaningful information) of an image for better understanding. Inspired by that, this paper presents a novel BIQA metric by mimicking the active inference process of IGM. Firstly, an active inference module based on the generative adversarial network (GAN) is established to predict the primary content, in which the semantic similarity and the structural dissimilarity (i.e., semantic consistency and structural completeness) are both considered during the optimization. Then, the image quality is measured on the basis of its primary content. Generally, the image quality is highly related to three aspects, i.e., the scene information (content-dependency), the distortion type (distortion-dependency), and the content degradation (degradation-dependency). According to the correlation between the distorted image and its primary content, the three aspects are analyzed and calculated respectively with a multi-stream convolutional neural network (CNN) based quality evaluator. As a result, with the help of the primary content obtained from the active inference and the comprehensive quality degradation measurement from the multi-stream CNN, our method achieves competitive performance on five popular IQA databases. Especially in cross-database evaluations, our method achieves significant improvements.
Jupo Ma, Jinjian Wu, Leida Li, Weisheng Dong, Xuemei Xie, Guangming Shi, Weisi Lin
IEEE Trans. Image Process.1
2020 Active Inference of GAN for No-Reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) is a challenging task. It is a promising idea to design NR-IQA algorithms by mimicking how human visual system (HVS) works. The internal generative mechanism (IGM) indicates that HVS actively infers the primary content of an image for better understanding. Inspired by that, a novel NR-IQA method with active inference is proposed in this paper. First, a generative adversarial network (GAN) is proposed to predict the primary content of a distorted image, in which two IGM-inspired constraints are considered during the optimization. Next, based on the correlation between the distorted image and its primary content, different degradations (i.e., the content/distortion-/structure-dependency degradation) are measured simultaneously with a multi-stream convolutional neural network (CNN) for NR-IQA. Benefit from the primary content obtained from GAN and the multiple degradations measurement of CNN, our method achieves the state-of-the-art on five public IQA databases.
Jupo Ma, Jinjian Wu, Leida Li, Weisheng Dong, Xuemei Xie
ICME1
2020 End-to-End Blind Image Quality Prediction With Cascaded Deep Neural Network
abstract
The deep convolutional neural network (CNN) has achieved great success in image recognition. Many image quality assessment (IQA) methods directly use recognition-oriented CNN for quality prediction. However, the properties of IQA task is different from image recognition task. Image recognition should be sensitive to visual content and robust to distortion, while IQA should be sensitive to both distortion and visual content. In this paper, an IQA-oriented CNN method is developed for blind IQA (BIQA), which can efficiently represent the quality degradation. CNN is large-data driven, while the sizes of existing IQA databases are too small for CNN optimization. Thus, a large IQA dataset is firstly established, which includes more than one million distorted images (each image is assigned with a quality score as its substitute of Mean Opinion Score (MOS), abbreviated as pseudo-MOS). Next, inspired by the hierarchical perception mechanism (from local structure to global semantics) in human visual system, a novel IQA-orientated CNN method is designed, in which the hierarchical degradation is considered. Finally, by jointly optimizing the multilevel feature extraction, hierarchical degradation concatenation (HDC) and quality prediction in an end-to-end framework, the Cascaded CNN with HDC (named as CaHDC) is introduced. Experiments on the benchmark IQA databases demonstrate the superiority of CaHDC compared with existing BIQA methods. Meanwhile, the CaHDC (with about 0.73M parameters) is lightweight comparing to other CNN-based BIQA models, which can be easily realized in the microprocessing system. The dataset and source code of the proposed method are available at https://web.xidian.edu.cn/wjj/paper.html.
Jinjian Wu, Jupo Ma, Fuhu Liang, Weisheng Dong, Guangming Shi, Weisi Lin
IEEE Trans. Image Process.2
2019 End-to-End Blind Image Quality Assessment with Cascaded Deep Features
abstract
The convolutional neural network (CNN) has achieved great success in many visual tasks. However, it has limited progress on image quality assessment (IQA) due to the lacking of IQA-oriented CNN framework which can efficiently represent the hierarchical quality degradation. In this paper, inspired by the hierarchical perception mechanism (from local structure to global semantics) in the human visual system, we design an end-to-end cascaded CNN framework for blind IQA (BIQA), in which multilevel features are extracted and concatenated to represent the hierarchical quality degradation. By jointly optimizing the feature extraction, hierarchical degradation integration, and quality prediction in an end-to-end manner, the novel cascaded CNN with hierarchical feature integration (CaHFI) for BIQA is designed. Experimental results on five benchmark IQA databases demonstrate that the proposed CaHFI achieves the state-of-the-art. And experiments on cross-database evaluation further prove the high generalization ability of the proposed CaHFI.
Jinjian Wu, Jupo Ma, Fuhu Liang, Weisheng Dong, Guangming Shi
ICME2