Jiaqi Ma 0002

dblp:155/2199-2 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0001-8491-1968ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Image and video processing · 90% Visual content generation and editing · 10%
Artificial intelligence
3 papers
Deep learning architectures and training · 48% Image recognition and object detection · 40% Face, body and person analysis · 12%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing
image restoration
1.622026
Perceive-IR: Learning to Perceive Degradation Better for All-in-One Image Restoration · IEEE Trans. Image Process. 2026
ELMformer: Efficient Raw Image Restoration with a Locally Multiplicative Transformer · ACM Multimedia 2022
Image and video processing › image restoration › multi-task image restoration
all-in-one image restoration
1.012026
Perceive-IR: Learning to Perceive Degradation Better for All-in-One Image Restoration · IEEE Trans. Image Process. 2026
Image and video processing › image restoration
degradation-aware restoration
1.012026
Perceive-IR: Learning to Perceive Degradation Better for All-in-One Image Restoration · IEEE Trans. Image Process. 2026
Machine learning › Deep learning architectures and training
foundation model
0.912025
HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Image recognition and object detection
hyperspectral image analysis
0.912025
HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Image recognition and object detection › hyperspectral image analysis
hyperspectral image classification
0.912025
HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.912025
HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Image and video processing › image restoration
image denoising
0.612022
ELMformer: Efficient Raw Image Restoration with a Locally Multiplicative Transformer · ACM Multimedia 2022
Image and video processing › image restoration › image denoising
raw image denoising
0.612022
ELMformer: Efficient Raw Image Restoration with a Locally Multiplicative Transformer · ACM Multimedia 2022
Computer vision › Face, body and person analysis
face recognition
0.512021
Pseudo Facial Generation With Extreme Poses for Face Recognition · CVPR 2021
Visual content generation and editing › image generation
face image generation
0.512021
Pseudo Facial Generation With Extreme Poses for Face Recognition · CVPR 2021
Data mining
clustering
0.412019
Pseudo Supervised Matrix Factorization in Discriminative Subspace · IJCAI 2019
Data mining › clustering
matrix factorization-based clustering
0.412019
Pseudo Supervised Matrix Factorization in Discriminative Subspace · IJCAI 2019
Data mining › clustering › high-dimensional clustering
subspace clustering
0.412019
Pseudo Supervised Matrix Factorization in Discriminative Subspace · IJCAI 2019
Machine learning › Deep learning architectures and training › transformer
efficient transformer
0.212022
ELMformer: Efficient Raw Image Restoration with a Locally Multiplicative Transformer · ACM Multimedia 2022
Machine learning › Deep learning architectures and training
transformer
0.212022
ELMformer: Efficient Raw Image Restoration with a Locally Multiplicative Transformer · ACM Multimedia 2022

Methods — techniques the papers use, named apart from their topics

locally multiplicative self-attention · 1.1bi-directional fusion projection · 1.1semantic guidance · 1.0prompt learning · 1.0pixel-wise reconstruction · 1.0perceptual loss · 1.0generative adversarial network · 1.0CLIP · 1.0spectral enhancement · 0.9sparse sampling attention · 0.9masked image modeling · 0.9non-negative matrix factorization · 0.4manifold regularization · 0.4linear discriminant analysis · 0.4
YearPublicationVenuePosition
2026 Perceive-IR: Learning to Perceive Degradation Better for All-in-One Image Restoration
abstract
Existing All-in-One image restoration methods often fail to perceive degradation types and severity levels simultaneously, overlooking the importance of fine-grained quality perception. Moreover, these methods often utilize highly customized backbones, which hinder their adaptability and integration into more advanced restoration networks. To address these limitations, we propose Perceive-IR, a novel backbone-agnostic All-in-One image restoration framework designed for fine-grained quality control across various degradation types and severity levels. Its modular structure allows core components to function independently of specific backbones, enabling seamless integration into advanced restoration models without significant modifications. Specifically, Perceive-IR operates in two key stages: 1) multi-level quality-driven prompt learning stage, where a fine-grained quality perceiver is meticulously trained to discern three-tier quality levels by optimizing the alignment between prompts and images within the CLIP perception space. This stage ensures a nuanced understanding of image quality, laying the groundwork for subsequent restoration; 2) restoration stage, where the quality perceiver is seamlessly integrated with a difficulty-adaptive perceptual loss, forming a quality-aware learning strategy. This strategy not only dynamically differentiates sample learning difficulty but also achieves fine-grained quality control by driving the restored image toward the ground truth while pulling it away from both low- and medium-quality samples. Furthermore, Perceive-IR incorporates a Semantic Guidance Module (SGM) and Compact Feature Extraction (CFE). The SGM leverages semantic information from pre-trained vision models to provide high-level contextual guidance, while the CFE focuses on extracting degradation-specific features, ensuring accurate handling of diverse image degradations. Extensive experiments demonstrate that Perceive-IR not only surpasses state-of-the-art methods but also generalizes reliably to zero-shot real-world and unknown degraded scenes, while adapting seamlessly to different backbone networks. This versatility underscores the framework's robustness and backbone-agnostic design. Project page at https://house-yuyu.github.io/Perceive-IR/.
Xu Zhang 0044, Jiaqi Ma 0002, Guoli Wang 0004, Qian Zhang 0009, Huan Zhang 0008, Lefei Zhang
IEEE Trans. Image Process.2
2025 HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation Model
abstract
Accurate hyperspectral image (HSI) interpretation is critical for providing valuable insights into various earth observation-related applications such as urban planning, precision agriculture, and environmental monitoring. However, existing HSI processing methods are predominantly task-specific and scene-dependent, which severely limits their ability to transfer knowledge across tasks and scenes, thereby reducing the practicality in real-world applications. To address these challenges, we present HyperSIGMA, a vision transformer-based foundation model that unifies HSI interpretation across tasks and scenes, scalable to over one billion parameters. To overcome the spectral and spatial redundancy inherent in HSIs, we introduce a novel sparse sampling attention (SSA) mechanism, which effectively promotes the learning of diverse contextual features and serves as the basic block of HyperSIGMA. HyperSIGMA integrates spatial and spectral features using a specially designed spectral enhancement module. In addition, we construct a large-scale hyperspectral dataset, HyperGlobal-450K, for pre-training, which contains about 450 K hyperspectral images, significantly surpassing existing datasets in scale. Extensive experiments on various high-level and low-level HSI tasks demonstrate HyperSIGMA's versatility and superior representational capability compared to current state-of-the-art methods. Moreover, HyperSIGMA shows significant advantages in scalability, robustness, cross-modal transferring capability, real-world applicability, and computational efficiency.
Di Wang 0023, Meiqi Hu, Yuchun Miao, Jiaqi Yang 0005, Yichu Xu, Xiaolei Qin, Jiaqi Ma 0002, Chenxing Li, Chuan Fu, Hongruixuan Chen, Chengxi Han, Naoto Yokoya, Jing Zhang 0037, Minqiang Xu, Lefei Zhang, Chen Wu 0003, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2023 Restoration and enhancement on low exposure raw images by joint demosaicing and denoising
Jiaqi Ma 0002, Guoli Wang 0004, Lefei Zhang, Qian Zhang 0009
Neural Networks1
2022 ELMformer: Efficient Raw Image Restoration with a Locally Multiplicative Transformer
abstract
In order to get raw images of high quality for downstream Image Signal Process (ISP), in this paper we present an Efficient Locally Multiplicative Transformer called ELMformer for raw image restoration. ELMformer contains two core designs especially for raw images whose primitive attribute is single-channel. The first design is a Bi-directional Fusion Projection (BFP) module, where we consider both the color characteristics of raw images and spatial structure of single-channel. The second one is that we propose a Locally Multiplicative Self-Attention (L-MSA) scheme to effectively deliver information from the local space to relevant parts. ELMformer can efficiently reduce the computational consumption and perform well on raw image restoration tasks. Enhanced by these two core designs, ELMformer achieves the highest performance and keeps the lowest FLOPs on raw denoising and raw deblurring benchmarks compared with state-of-the-arts. Extensive experiments demonstrate the superiority and generalization ability of ELMformer. On SIDD benchmark, our method has even better denoising performance than ISP-based methods which need huge amount of additional sRGB training images.
Jiaqi Ma 0002, Shengyuan Yan, Lefei Zhang, Guoli Wang 0004, Qian Zhang 0009
ACM Multimedia1
2021 Pseudo Facial Generation With Extreme Poses for Face Recognition
abstract
Face recognition has achieved a great success in recent years, it is still challenging to recognize those facial images with extreme poses. Traditional methods consider it as a domain gap problem. Many of them settle it by generating fake frontal faces from extreme ones, whereas they are tough to maintain the identity information with high computational consumption and uncontrolled disturbances. Our experimental analysis shows a dramatic precision drop with extreme poses. Meanwhile, those extreme poses just exist minor visual differences after small rotations. Derived from this insight, we attempt to relieve such a huge precision drop by making minor changes to the input images without modifying existing discriminators. A novel lightweight pseudo facial generation is proposed to relieve the problem of extreme poses without generating any frontal facial image. It can depict the facial contour information and make appropriate modifications to preserve the critical identity information. Specifically, the proposed method reconstructs pseudo profile faces by minimizing the pixel-wise differences with original profile faces and maintaining the identity consistent information from their corresponding frontal faces simultaneously. The proposed framework can improve existing discriminators and obtain a great promotion on several benchmark datasets.
Guoli Wang 0004, Jiaqi Ma 0002, Qian Zhang 0009, Jiwen Lu, Jie Zhou 0001
CVPR2
2021 Discriminative subspace matrix factorization for multiview data clustering
Jiaqi Ma 0002, Yipeng Zhang 0001, Lefei Zhang
Pattern Recognit.1
2019 Pseudo Supervised Matrix Factorization in Discriminative Subspace
abstract
Non-negative Matrix Factorization (NMF) and spectral clustering have been proved to be efficient and effective for data clustering tasks and have been applied to various real-world scenes. However, there are still some drawbacks in traditional methods: (1) most existing algorithms only consider high-dimensional data directly while neglect the intrinsic data structure in the low-dimensional subspace; (2) the pseudo-information got in the optimization process is not relevant to most spectral clustering and manifold regularization methods. In this paper, a novel unsupervised matrix factorization method, Pseudo Supervised Matrix Factorization (PSMF), is proposed for data clustering. The main contributions are threefold: (1) to cluster in the discriminant subspace, Linear Discriminant Analysis (LDA) combines with NMF to become a unified framework; (2) we propose a pseudo supervised manifold regularization term which utilizes the pseudo-information to instruct the regularization term in order to find subspace that discriminates different classes; (3) an efficient optimization algorithm is designed to solve the proposed problem with proved convergence. Extensive experiments on multiple benchmark datasets illustrate that the proposed model outperforms other state-of-the-art clustering algorithms.
Jiaqi Ma 0002, Yipeng Zhang 0001, Lefei Zhang, Bo Du 0001, Dapeng Tao
IJCAI1