Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jiamu Sheng

dblp:340/7010 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0001-6073-8723ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Video understanding and tracking · 52% Deep learning architectures and training · 48%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › state space model
mamba
1.012026
STNMamba: Mamba-Based Spatial-Temporal Normality Learning for Video Anomaly Detection · IEEE Trans. Multim. 2026
Machine learning › Deep learning architectures and training
state space model
1.012026
STNMamba: Mamba-Based Spatial-Temporal Normality Learning for Video Anomaly Detection · IEEE Trans. Multim. 2026
Computer vision › Video understanding and tracking
video anomaly detection
1.012026
STNMamba: Mamba-Based Spatial-Temporal Normality Learning for Video Anomaly Detection · IEEE Trans. Multim. 2026
Computer vision › Video understanding and tracking
video question answering
0.912025
DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes · ACM Multimedia 2025
Computer vision › Video understanding and tracking › video summarization
keyframe selection
0.312025
DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

multi-scale vision space state block · 1.0memory bank · 1.0channel-aware vision space state block · 1.0automatic QA generation · 0.9
YearPublicationVenuePosition
2026 STNMamba: Mamba-Based Spatial-Temporal Normality Learning for Video Anomaly Detection
abstract
Video anomaly detection (VAD) has been extensively researched due to its potential for intelligent video systems. However, most existing methods based on CNNs and transformers still suffer from substantial computational burdens and have room for improvement in learning spatial-temporal normality. Recently, Mamba has shown great potential for modeling long-range dependencies with linear complexity, providing an effective solution to the above dilemma. To this end, we propose a lightweight and effective Mamba-based network named STNMamba, which incorporates carefully designed Mamba modules to enhance the learning of spatial-temporal normality. Firstly, we develop a dual-encoder architecture, where the spatial encoder equipped with Multi-Scale Vision Space State Blocks (MS-VSSB) extracts multi-scale appearance features, and the temporal encoder employs Channel-Aware Vision Space State Blocks (CA-VSSB) to capture significant motion patterns. Secondly, a Spatial-Temporal Interaction Module (STIM) is introduced to integrate spatial and temporal information across multiple levels, enabling effective modeling of intrinsic spatial-temporal consistency. Within this module, the Spatial-Temporal Fusion Block (STFB) is proposed to fuse the spatial and temporal features into a unified feature space, and the memory bank is utilized to store spatial-temporal prototypes of normal patterns, restricting the model's ability to represent anomalies. Extensive experiments on three benchmark datasets demonstrate that our STNMamba achieves competitive performance with fewer parameters and lower computational costs than existing methods.
Zhangxun Li, Mengyang Zhao 0002, Yang Liu 0246, Jiamu Sheng, Xinhua Zeng, Tian Wang 0002, Kewei Wu, Yu-Gang Jiang 0001
IEEE Trans. Multim.5
2025 IF-DETR: Incremental Few-Shot Detection Transformer for Surface Defect Detection
abstract
Recognizing new categories that continually emerge during production is a significant challenge in defect detection, and deep learning methods have become the mainstream solution in recent years. However, in industrial scenarios, the scarcity of new category samples in the early phases leads to catastrophic forgetting and poor generalization in these methods. To address this, we propose a novel end-to-end Incremental Few-shot DEtection TRansformer (IF-DETR), which aims to resolve the knowledge ambiguity between old and new categories in such settings. Specifically, IF-DETR follows a two-stage paradigm involving pre-training and fine-tuning, leveraging knowledge distillation (KD) to improve the fine-tuning process while distinguishing label structures to eliminate ambiguous knowledge. We design two KD losses: Feature-level Instance Aware (FIA) loss and Logit-level Hierarchy Aligned (LHA) loss. For FIA loss, we exploit the global modeling ability of transformers to generate attention masks, thereby assigning reasonable weights to valuable distillation regions and mitigating catastrophic forgetting. For LHA loss, we decouple the teacher model’s output logits based on the supervision information, applying auxiliary losses at each layer of the decoder to refine the student model’s predictions and enhance the generalization ability for new categories. To the best of our knowledge, this is the first work to introduce transformer-based detector into incremental few-shot defect detection. We conduct extensive experiments on two benchmark defect detection datasets, and IF-DETR achieves the best performance across all 9 splits and few-shot settings, demonstrating its effectiveness.
Zhangxun Li, Xinzhi Lin, Jiamu Sheng, Nailei Hei, Lijun Dai, Lizhe Qi
IJCNN4
2025 DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
Zhende Song, Jiamu Sheng, Chi Zhang 0007, Shengji Tang, Jiayuan Fan 0001, Tao Chen 0003
ACM Multimedia3
2025 DualMamba: A Lightweight Spectral-Spatial Mamba-Convolution Network for Hyperspectral Image Classification
abstract
The effectiveness and efficiency of modeling complex spectral–spatial relations are crucial for hyperspectral image (HSI) classification. Most existing methods based on convolution neural networks (CNNs) and transformers still suffer from heavy computational burdens and have room for improvement in capturing the global–local spectral–spatial feature representation. To this end, we propose a novel lightweight parallel design called a lightweight dual-stream Mamba-convolution network (DualMamba) for HSI classification. Specifically, a parallel lightweight Mamba and CNN block are developed to extract global and local spectral–spatial features. First, the cross-attention spectral–spatial Mamba module (CAS2MM) is proposed to leverage the global modeling of Mamba at linear complexity. In this module, dynamic positional embedding (DPE) is designed to enhance the spatial location information of visual sequences. The lightweight spectral–spatial Mamba blocks comprise an efficient scanning strategy and a lightweight Mamba design to efficiently extract global spectral–spatial features. And the cross-attention spectral–spatial fusion (CAS2F) is designed to learn cross correlation and fuse spectral–spatial features. Second, the lightweight spectral–spatial residual convolution module is proposed with lightweight spectral and spatial branches to extract local spectral–spatial features through residual learning. Finally, the adaptive global–local fusion is proposed to dynamically combine global Mamba features and local convolution features for a global–local spectral–spatial representation. Compared with state-of-the-art HSI classification methods, experimental results demonstrate that DualMamba achieves significant classification accuracy on three public HSI datasets and a superior reduction in model parameters and floating-point operations (FLOPs).
Jiamu Sheng, Peng Ye 0006, Jiayuan Fan 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Exploring Multi-Timestep Multi-Stage Diffusion Features for Hyperspectral Image Classification
abstract
The effectiveness of spectral-spatial feature learning is crucial for the hyperspectral image (HSI) classification task. Diffusion models, as a new class of groundbreaking generative models, have the ability to learn both contextual semantics and textual details from the distinct timestep dimension, enabling the modeling of complex spectral-spatial relations in HSIs. However, existing diffusion-based HSI classification methods only utilize manually selected single-timestep single-stage features, limiting the full exploration and exploitation of rich contextual semantics and textual information hidden in the diffusion model. To address this issue, we propose a novel diffusion-based feature learning framework that explores Multi-Timestep Multi-Stage Diffusion features for HSI classification for the first time, called MTMSD. Specifically, the diffusion model is first pretrained with unlabeled HSI patches to mine the connotation of unlabeled data, and then is used to extract the multi-timestep multi-stage diffusion features. To effectively and efficiently leverage multi-timestep multi-stage features, two strategies are further developed. One strategy is class & timestep-oriented multi-stage feature purification module with the inter-class and inter-timestep prior for reducing the redundancy of multi-stage features and alleviating memory constraints. The other one is selective timestep feature fusion module with the guidance of global features to adaptively select different timestep features for integrating texture and semantics. Both strategies facilitate the generality and adaptability of the MTMSD framework for diverse patterns of different HSI data. Extensive experiments are conducted on four public HSI datasets, and the results demonstrate that our method outperforms state-of-the-art methods for HSI classification, especially on the challenging Houston 2018 dataset. The codes are available at https://github.com/zjyaccount/MTMSD.
Jiamu Sheng, Peng Ye 0006, Jiayuan Fan 0001, Tong He 0001, Bin Wang 0008, Tao Chen 0003
IEEE Trans. Geosci. Remote. Sens.2
2023 JNDMix: Jnd-Based Data Augmentation for No-Reference Image Quality Assessment
abstract
Despite substantial progress in no-reference image quality assessment (NR-IQA), previous training models often suffer from over-fitting due to the limited scale of used datasets, resulting in model performance bottlenecks. To tackle this challenge, we explore the potential of leveraging data augmentation to improve data efficiency and enhance model robustness. However, most existing data augmentation methods incur a serious issue, namely that it alters the image quality and leads to training images mismatching with their original labels. Additionally, although only a few data augmentation methods are available for NR-IQA task, their ability to enrich dataset diversity is still insufficient. To address these issues, we propose a effective and general data augmentation based on just noticeable difference (JND) noise mixing for NR-IQA task, named JNDMix. In detail, we randomly inject the JND noise, imperceptible to the human visual system (HVS), into the training image without any adjustment to its label. Extensive experiments demonstrate that JNDMix significantly improves the performance and data efficiency of various state-of-the-art NR-IQA models and the commonly used baseline models, as well as the generalization ability. More importantly, JNDMix facilitates MANIQA to achieve the state-of-the-art performance on LIVEC and KonIQ-10k.
Jiamu Sheng, Jiayuan Fan 0001, Peng Ye 0006, Jianjian Cao
ICASSP1