VLDB 2026 Research / reviewers in the wild / expert
Jinbao Wang 0001
dblp:84/1293-1 · also Jin-Bao Wang 0001
· DBLP profile ↗
54ranked-venue papers
7as first author
47since 2021 · last 2026
0000-0001-5916-8965ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 4 first-author · 30 since 2021Artificial intelligence and machine learning · 27 · 3 first-author · 25 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scene Experts: Specializing in 3D Gaussian Splatting with Adaptive DecompositionabstractAnchor-based 3D Gaussian Splatting (GS), exemplified by Scaffold-GS, achieves remarkable storage efficiency through a hybrid explicit-implicit representation. However, their reliance on a single, monolithic network to decode anchor features imposes a severe bottleneck on model capacity, often resulting in blurred details and view-dependent artifacts in complex scenes. To break this bottleneck, we introduce the concept of Scene Experts: a strategy that decomposes the task of modeling a complex scene across a collection of specialized sub-models. To realize the paradigm, we propose MoE-GS. Our approach designs the decoder as a Sparsely-Gated Mixture of Experts (MoE), which dramatically increases the model's total capacity while maintaining comparable inference cost via sparse activation. To effectively train this high-capacity model, we propose two key innovations: (1) A progressive curriculum learning strategy that first trains all experts on a robust baseline before encouraging them to specialize on different scene components. (2) A novel opacity-aware regularization that penalizes inactive neural Gaussians, ensuring the expanded capacity is efficiently used. Extensive experiments demonstrate that MoE-GS substantially outperforms state-of-the-art methods on diverse benchmarks, significantly improving reconstruction fidelity while requiring a smaller or comparable Gaussian model size. Xiaowen Fu, Yuhan Tang, Huazhong Zhang, Tianxing Zhao, Jinbao Wang 0001 |
AAAI | 8 |
| 2026 | MFF-M3AD: A unified reconstruction method with multi-scale feature fusion for multi-category 3D anomaly detection
Hanzhe Liang, Yejin Tang, LinLin Shen, Jinbao Wang 0001, Can Gao |
Neural Networks | 5 |
| 2026 | BinaryAD: Efficient image anomaly detection via binarized representations
Bingyang Guo, Hanzhe Liang, LinLin Shen, Jinbao Wang 0001, Zhichao Lu |
Pattern Recognit. | 7 |
| 2026 | A lightweight 3D anomaly detection method with rotationally invariant features
Hanzhe Liang, Jie Zhou 0009, Can Gao, Bingyang Guo, Jinbao Wang 0001, LinLin Shen |
Pattern Recognit. | 5 |
| 2026 | DreamAssemble: Complex Multi-Object Text-to-3D Generation via Multi-Density Neural Fields
Bin Huang 0016, Jinbao Wang 0001, Dongmei Jiang, Hongjuan Pei, Qiulu Li, Jian Xue 0002, Ke Lu 0002 |
IEEE Trans. Image Process. | 2 |
| 2026 | GS2Physics: Semantic-Region-Aware Gaussian Splatting for Physical Property PredictionabstractPredicting the physical properties of reconstructed 3D assets is essential for virtual reality interactions. However, current systems often depend on manually assigning properties such as stiffness and density, which can be inefficient and prone to errors. To address this issue, we present GS2Physics, a novel framework based on 3D Gaussian Splatting. This framework is designed to predict physical properties accurately while maintaining improved consistency in semantic segmentation. Unlike existing approaches, which either struggle with region inconsistency or misalign semantic 3D features, GS2Physics embeds semantic-region-aware features directly into the Gaussian Splatting representation. This allows for region-consistent and accurate physical property prediction, achieving state-of-the-art performance on the ABO-500 mass prediction benchmark. To further evaluate our segmentation capabilities, we introduce PhysSeg-15, a subset dataset of ABO-500 featuring physical property segmentation masks for 15 different 3D objects captured from five viewpoints. Our method significantly outperforms existing approaches in segmentation accuracy. Qualitative results demonstrate more consistent material predictions across different object regions and improved accuracy in physical property prediction. In addition, we showcase the effectiveness of GS2Physics in 3D interaction tasks, where our predicted physical properties result in more realistic object motion. Our dataset and results are available at https://github.com/momaiyc/GS2Physics. Bin Huang 0016, Jiayi Lyu, Zehai Niu, LinLin Shen, Jinbao Wang 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | Improving Unsupervised Ultrasonic Image Anomaly Detection via Frequency-Spatial Feature Filtering and Gaussian Mixture ModelingabstractUltrasonic image anomaly detection faces significant challenges due to limited labeled data, strong structural and random noise, and highly diverse defect manifestations. To overcome these obstacles, we introduce UltraChip, a new large-scale C-scan benchmark containing about 8,000 real-world images from various chip packaging types, each meticulously annotated with pixel-level masks for cracks, holes, and layers. Building on this resource, we present FSGM-Net, a fully unsupervised framework tailored for anomaly detection. FSGM-Net leverages an adaptive Frequency-Spatial feature filtering mechanism: a learnable FFT-Spatial patch filter first suppresses noise and dynamically assigns normality weights to Vision Transformer (ViT) patch features. Subsequently, an Adaptive Gaussian Mixture Model (Ada-GMM) captures the distribution of normal features and guides a deep-shallow multi-scale interaction decoder for accurate, pixel-level anomaly inference. In addition, we propose a filter loss that enforces encoder-filter consistency and entropy-based sparse gating, together with a distributional loss that encourages both feature reconstruction and confident Gaussian mixture modeling. Extensive experiments demonstrate that FSGM-Net not only achieves state-of-the-art results on UltraChip but also exhibits superior cross-domain generalization to MVTec-AD and VisA, while supporting real-time inference on a single GPU. Together, the dataset and framework advance robust, annotation-free ultrasonic NDT in practical applications. The UltraChip dataset can be obtained via https://iiplab.net/ultrachip/. Ke Lu 0002, Jinbao Wang 0001, Can Gao, Jian Xue 0002 |
IEEE Trans. Image Process. | 3 |
| 2025 | DCSF-KD: Dynamic Channel-wise Spatial Feature Knowledge Distillation for Object DetectionabstractKnowledge distillation (KD) has recently gained great success in the field of object detection. By transferring the knowledge of the spatial or channel domain from the teacher model to the student model, it allows for a more compact representation with minimal performance loss. Despite this progress, existing KD methods typically treat knowledge from spatial or channel domains independently, ignoring the exploitation of the mutual relationship between these domains. In this work, we first explore the connection between spatial and channel domains and find there exists a strong correlation between them, i.e. the salient channels tend to contain significant object regions in the spatial domain. Motivated by this observation, we propose DCSF-KD, a novel Dynamic Channel-wise Spatial Feature Knowledge Distillation framework for object detection by fully exploiting both spatial and channel knowledge. Specifically, we introduce channel-wise spatial feature distillation and global channel attention distillation, using information from both domains to improve the accuracy of the student network. Experiments demonstrate that our DCSF-KD outperforms existing detection methods on both homogeneous and heterogeneous teacher-student network pairs. For example, when using the MaskRCNN-Swin detector as the teacher, and based on RetinaNet and FCOS with ResNet-50 on MS COCO, our DCSF-KD can achieve 41.9% and 44.1% mAP, respectively. Tao Dai 0001, Hang Guo 0002, Jinbao Wang 0001, Zexuan Zhu 0001 |
AAAI | 4 |
| 2025 | Look Inside for More: Internal Spatial Modality Perception for 3D Anomaly Detectionabstract3D anomaly detection has recently become a significant focus in computer vision. Several advanced methods have achieved satisfying anomaly detection performance. However, they typically concentrate on the external structure of 3D samples and struggle to leverage the internal information embedded within samples. Inspired by the basic intuition of why not look inside for more, we observed this prototype is straightforward and effective. As a result, we introduce a newly designed mode named Internal Spatial Modality Perception (ISMP) to explore the feature representation from internal views fully. Specifically, our proposed ISMP consists of a critical perception module, Spatial Insight Engine (SIE), which abstracts complex internal information of point clouds into essential global features. Besides, to better align structural information with point data, we propose an enhanced key point feature extraction method for amplifying spatial structure feature representation. Simultaneously, a novel feature filtering module is incorporated to reduce noise and redundant features for further precise spatial structure aligning. Extensive experiments validate the efficiency of our proposed method, achieving object-level and pixel-level AUROC improvements of 4.2% and 13.1%, respectively, on the Real3D-AD benchmarks. Note that the strong generalization ability of SIE has been theoretically proven and verified in both classification and segmentation tasks. Our code will be released upon acceptance. Hanzhe Liang, Guoyang Xie, Chengbin Hou, Bingshu Wang, Can Gao, Jinbao Wang 0001 |
AAAI | 6 |
| 2025 | Learning with Open-world Noisy Data via Class-independent Margin in Dual Representation SpaceabstractLearning with Noisy Labels (LNL) aims to improve the model generalization when facing data with noisy labels, and existing methods generally assume that noisy labels come from known classes, called closed-set noise. However, in real-world scenarios, noisy labels from similar unknown classes, i.e., open-set noise, may occur during the training and inference stage. Such open-world noisy labels may significantly impact the performance of LNL methods. In this study, we propose a novel dual-space joint learning method to robustly handle the open-world noise. To mitigate model overfitting on closed-set and open-set noises, a dual representation space is constructed by two networks. One is a projection network that learns shared representations in the prototype space, while the other is a One-Vs-All (OVA) network that makes predictions using unique semantic representations in the class-independent space. Then, bi-level contrastive learning and consistency regularization are introduced in two spaces to enhance the detection capability for data with unknown classes. To benefit from the memorization effects across different types of samples, class-independent margin criteria are designed for sample identification, which selects clean samples, weights closed-set noise, and filters open-set noise effectively. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods and achieves an average accuracy improvement of 4.55\% and an AUROC improvement of 6.17\% on CIFAR80N. Linchao Pan, Can Gao, Jinbao Wang 0001 |
AAAI | 4 |
| 2025 | DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face SynthesisabstractAccurately synthesizing talking face videos and capturing fine facial features for individuals with long hair presents a significant challenge. To tackle these challenges in existing methods, we propose a decomposed per-embedding Gaussian fields (DEGSTalk), a 3D Gaussian Splatting (3DGS)-based talking face synthesis method for generating realistic talking faces with long hairs. Our DEGSTalk employs Deformable Pre-Embedding Gaussian Fields, which dynamically adjust pre-embedding Gaussian primitives using implicit expression coefficients. This enables precise capture of dynamic facial regions and subtle expressions. Additionally, we propose a Dynamic Hair-Preserving Portrait Rendering technique to enhance the realism of long hair motions in the synthesized videos. Results show that DEGSTalk achieves improved realism and synthesis quality compared to existing approaches, particularly in handling complex facial dynamics and hair preservation. Our code is available at https://github.com/CVI-SZU/DEGSTalk. Kaijun Deng, Dezhi Zheng, Jindong Xie, Jinbao Wang 0001, Weicheng Xie 0001, LinLin Shen, Siyang Song |
ICASSP | 4 |
| 2025 | High-Fidelity Editable Portrait Synthesis with 3D GAN InversionabstractThe 3D generative adversarial network (GAN) inversion converts an image into 3D representation to attain high-fidelity reconstruction and facilitate realistic image manipulation within the 3D latent space. However, previous approaches face challenges regarding the trade-off between the reconstruction ability and editability. That is, reversing a real-world image to a low-dimensional latent code would inevitably lead to information loss, and achieving a near-perfect reconstruction using high-rate triplane representation often limits the ability to manipulate the image freely in the latent space. To address these issues, we propose a novel latent conditioning encoder-based framework with the alignment between the low-dimensional latent and high-dimensional triplane. A non-semantic guided editing strategy bridges the intrinsic relation between the latent condition and triplane generation, making it possible to edit the high-dimensional representation by latent manipulation. As a result, our method can achieve high-fidelity reconstruction and editing simultaneously by directly controlling the latent code. Experimental results demonstrate that our approach excels in reconstruction and editing quality compared to previous 3D inversion methods. Furthermore, our method can also edit even real faces with large poses and out-of-domain cases. Jindong Xie, Yupei Lin, Jinbao Wang 0001, Xianxu Hou, LinLin Shen |
ICASSP | 4 |
| 2025 | FBI-Net: Frequency Band Integration Network for Infrared Small Target SegmentationabstractSmall targets in infrared imagery exhibit challenging characteristics due to their minimal semantic information and the extremely imbalanced distribution between the targets and the background. In this paper, we propose a frequency band integration network to extract salient features of infrared small targets in both the spatial and frequency domains. To excavate the high-frequency features of the small targets, we propose a frequency decoupling-fusion module. To decrease the semantic loss that occurs in deep networks, we propose a semantic injection mechanism to assist in retaining critical information from shallow layers. Experimental results show that our proposed method reaches higher prediction accuracy and robustness in the infrared small target segmentation task compared with other state-of-the-art approaches. Biqiao Xin, Qianchen Mao, Jinbao Wang 0001, Bingshu Wang |
ICASSP | 4 |
| 2025 | Dual Encoders for Diffusion-based Image InpaintingabstractCurrent diffusion-based inpainting models struggle to preserve unmasked regions or generate highly coherent content. Additionally, it is hard for them to generate meaningful content for 3D inpainting. To tackle these challenges, we design a plug-and-play branch that runs through the entire generation process to enhance existing models. Specifically, we utilize dual encoders - a Convolutional Neural Network (CNN) encoder and the pre-trained Variational AutoEncoder (VAE) encoder, to encode masked images. The latent code and the feature map from the dual encoders are fed to diffusion models simultaneously. In addition, we apply Zero-padded initialization to solve the problem of mode collapse caused by this branch. Experiments on BrushBench and EditBench demonstrate that models with our plug-and-play branch can improve the coherence of inpainting, and our model achieves new state-of-the-art results. Dezhi Zheng, Kaijun Deng, Jinbao Wang 0001, LinLin Shen |
ICASSP | 3 |
| 2025 | Text-to-Any-Skeleton Motion Generation Without Retargeting
Qingyuan Liu 0001, Ke Lu 0002, Kun Dong 0001, Jian Xue 0002, Zehai Niu, Jinbao Wang 0001 |
ICCV | 6 |
| 2025 | HilComp: Hilbert Curve-Based Balanced Clustering for 3D Gaussian Splatting Compression
Xiaowen Fu, Zhengqi He, Jinbao Wang 0001, Xueliang Li 0002 |
ICIC (18) | 5 |
| 2025 | MC3D-AD: A Unified Geometry-aware Reconstruction Model for Multi-category 3D Anomaly Detectionabstract3D Anomaly Detection (AD) is a promising means of controlling the quality of manufactured products. However, existing methods typically require carefully training a task-specific model for each category independently, leading to high cost, low efficiency, and weak generalization. This study presents a novel unified model for Multi-Category 3D Anomaly Detection (MC3D-AD) that aims to utilize both local and global geometry-aware information to reconstruct normal representations of all categories. First, to learn robust and generalized features of different categories, we propose an adaptive geometry-aware masked attention module that extracts geometry variation information to guide mask attention. Then, we introduce a local geometry-aware encoder reinforced by the improved mask attention to encode group-level feature tokens. Finally, we design a global query decoder that utilizes point cloud position embeddings to improve the decoding process and reconstruction ability. This leads to local and global geometry-aware reconstructed feature tokens for the 3D AD task. MC3D-AD is evaluated on two publicly available Real3D-AD and Anomaly-ShapeNet datasets, and exhibits significant superiority over current state-of-the-art single-category methods, achieving 3.1% and 9.3% improvement in object-level AUROC over Real3D-AD and Anomaly-ShapeNet, respectively. The code is available at https://github.com/iCAN-SZU/MC3D-AD. Jiayi Cheng, Can Gao, Jie Zhou 0009, Jiajun Wen 0001, Jinbao Wang 0001 |
IJCAI | 6 |
| 2025 | Taming Anomalies with Down-Up Sampling Networks: Group Center Preserving Reconstruction for 3D Anomaly DetectionabstractReconstruction-based methods have demonstrated very promising results for 3D anomaly detection. However, these methods face great challenges in handling high-precision point clouds due to the large scale and complex structure. In this study, a Down-Up Sampling Networks (DUS-Net) is proposed to reconstruct high-precision point clouds for 3D anomaly detection by preserving the group center geometric structure. The DUS-Net first introduces a Noise Generation module to generate noisy patches, which facilitates the diversity of training data and strengthens the feature representation for reconstruction. Then, a Down-sampling Network (Down-Net) is developed to learn an anomaly-free center point cloud from patches with noise injection. Subsequently, an Up-sampling Network (Up-Net) is designed to reconstruct high-precision point clouds by fusing multi-scale up-sampling features. Our method leverages group centers for construction, enabling the preservation of geometric structure and providing a more precise point cloud. Extensive experiments demonstrate the effectiveness of our proposed method, achieving state-of-the-art (SOTA) performance, with an Object-level AUROC of 79.9% and 79.5% and a Point-level AUROC of 71.2% and 84.7% on the Real3D-AD and Anomaly-ShapeNet datasets, respectively. Hanzhe Liang, Jie Zhang 0090, Tao Dai 0001, LinLin Shen, Jinbao Wang 0001, Can Gao |
ACM Multimedia | 5 |
| 2025 | Unknown Pixel Mask Based Fine-tuning of 2D Inpainting Models for Unbounded 3D Scene Generation from a Single ImageabstractConventional 2D inpainting models are trained using masks confined to 2D scenarios, resulting in meaningless content when applied to 3D-specific masks. These 3D-specific masks, termed Unknown Pixels (UP) masks, represent unseen pixels from novel viewpoints that remain obscured in the original input image. Existing methods attempt to mitigate this issue by employing post-processing techniques to transform UP masks into 2D equivalents, frequently suffering from unnatural distortions. To address these issues, we investigate the efficacy of directly training 2D inpainting models with UP masks to circumvent such distortions. In this paper, we introduce a novel framework designed to generate unbounded 3D scenes from a single image, guided by textual descriptions. Our approach leverages fine-tuned inpainting models that iteratively reconstruct incomplete images originating from pure projection. The generated points are then seamlessly integrated into the original point cloud via pixel-wise depth alignment. Extensive evaluations demonstrate that our framework outperforms existing methods in scene quality, processing speed, and memory efficiency. Dezhi Zheng, Kaijun Deng, Xianxu Hou, Jinbao Wang 0001, LinLin Shen |
ACM Multimedia | 4 |
| 2025 | FAST: Foreground-aware Diffusion with Accelerated Sampling Trajectory for Segmentation-oriented Anomaly SynthesisabstractIndustrial anomaly segmentation relies heavily on pixel-level annotations, yet real-world anomalies are often scarce, diverse, and costly to label. Segmentation-oriented industrial anomaly synthesis (SIAS) has emerged as a promising alternative; however, existing methods struggle to balance sampling efficiency and generation quality. Moreover, most approaches treat all spatial regions uniformly, overlooking the distinct statistical differences between anomaly and background areas. This uniform treatment hinders the synthesis of controllable, structure-specific anomalies tailored for segmentation tasks. In this paper, we propose FAST, a foreground-aware diffusion framework featuring two novel modules: the Anomaly-Informed Accelerated Sampling (AIAS) and the Foreground-Aware Reconstruction Module (FARM). AIAS is a training-free sampling algorithm specifically designed for segmentation-oriented industrial anomaly synthesis, which accelerates the reverse process through coarse-to-fine aggregation and enables the synthesis of state-of-the-art segmentation-oriented anomalies in as few as 10 steps. Meanwhile, FARM adaptively adjusts the anomaly-aware noise within the masked foreground regions at each sampling step, preserving localized anomaly signals throughout the denoising trajectory. Extensive experiments on multiple industrial benchmarks demonstrate that FAST consistently outperforms existing anomaly synthesis methods in downstream segmentation tasks. We release the code in https://github.com/Chhro123/fast-foreground-aware-anomaly-synthesis. Xichen Xu, Yanshu Wang, Jinbao Wang 0001, Xiaoning Lei, Guoyang Xie, Guannan Jiang, Zhichao Lu |
NeurIPS | 3 |
| 2025 | SplatID: Real-Time Lossless 3D Gaussian Splatting with Feature ID Generation and Frame Filtering
Wenhui Ma, LinLin Shen, Jinbao Wang 0001 |
PRCV (10) | 4 |
| 2025 | LabelGS: Label-Aware 3D Gaussian Splatting for 3D Scene Segmentation
Dezhi Zheng, Lei Wang 0018, Liping xiang, Kaijun Deng, Xiaowen Fu, LinLin Shen, Jinbao Wang 0001 |
PRCV (10) | 11 |
| 2025 | Enhancing anomaly detection with few-shot fine-tuned long text-to-image models
Jiajia An, Junbin Lu, Zhuoqin Yang, Jinbao Wang 0001, LinLin Shen |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Node importance estimation leveraging LLMs for semantic augmentation in knowledge graphs
Chengbin Hou, Jinbao Wang 0001, Jianye Xue, Hairong Lv |
Knowl. Based Syst. | 4 |
| 2025 | FedAGHN: Personalized federated learning with attentive graph hypernetworks
Yunheng Shen, Chengbin Hou, Pengyu Wang 0007, Jinbao Wang 0001, Ke Tang 0001, Hairong Lv |
Knowl. Based Syst. | 5 |
| 2025 | Multimodal Emotional Talking Face Generation Based on Action UnitsabstractTalking face generation focuses on creating natural facial animations that align with the provided text or audio input. Current methods in this field primarily rely on facial landmarks to convey emotional changes. However, spatial key-points are valuable, yet limited in capturing the intricate dynamics and subtle nuances of emotional expressions due to their restricted spatial coverage. Consequently, this reliance on sparse landmarks can result in decreased accuracy and visual quality, especially when representing complex emotional states. To address this issue, we propose a novel method called Emotional Talking with Action Unit (ETAU), which seamlessly integrates facial Action Units (AUs) into the generation process. Unlike previous works that solely rely on facial landmarks, ETAU employs both Action Units and landmarks to comprehensively represent facial expressions through interpretable representations. Our method provides a detailed and dynamic representation of emotions by capturing the complex interactions among facial muscle movements. Moreover, ETAU adopts a multi-modal strategy by seamlessly integrating emotion prompts, driving videos, and target images, and by leveraging various input data effectively, it generates highly realistic and emotional talking-face videos. Through extensive evaluations across multiple datasets, including MEAD, LRW, GRID and HDTF, ETAU outperforms previous methods, showcasing its superior ability to generate high-quality, expressive talking faces with improved visual fidelity and synchronization. Moreover, ETAU exhibits a significant improvement on the emotion accuracy of the generated results, reaching an impressive average accuracy of 84% on the MEAD dataset. Jiayi Lyu, Xing Lan, Guohong Hu, Hanyu Jiang 0004, Jinbao Wang 0001, Jian Xue 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Unsupervised Continual Anomaly Detection with Contrastively-Learned PromptabstractUnsupervised Anomaly Detection (UAD) with incremental training is crucial in industrial manufacturing, as unpredictable defects make obtaining sufficient labeled data infeasible. However, continual learning methods primarily rely on supervised annotations, while the application in UAD is limited due to the absence of supervision. Current UAD methods train separate models for different classes sequentially, leading to catastrophic forgetting and a heavy computational burden. To address this issue, we introduce a novel Unsupervised Continual Anomaly Detection framework called UCAD, which equips the UAD with continual learning capability through contrastively-learned prompts. In the proposed UCAD, we design a Continual Prompting Module (CPM) by utilizing a concise key-prompt-knowledge memory bank to guide task-invariant 'anomaly' model predictions using task-specific 'normal' knowledge. Moreover, Structure-based Contrastive Learning (SCL) is designed with the Segment Anything Model (SAM) to improve prompt learning and anomaly segmentation results. Specifically, by treating SAM's masks as structure, we draw features within the same mask closer and push others apart for general feature representations. We conduct comprehensive experiments and set the benchmark on unsupervised continual anomaly detection and segmentation, demonstrating that our method is significantly better than anomaly detection methods, even with rehearsal training. The code will be available at https://github.com/shirowalker/UCAD. Jiaqi Liu 0004, Qiang Nie, Bin-Bin Gao, Yong Liu 0032, Jinbao Wang 0001, Chengjie Wang 0001, Feng Zheng 0001 |
AAAI | 7 |
| 2024 | Local Information Guided Global Integration for Infrared Small Target DetectionabstractInfrared small targets often exhibit small scale and weak semantic features, which makes it a great challenge to their detection. To address this situation, we propose a novel network for infrared small target detection that combines local details information and global contextual information. To preserve the local and high-frequency details present in infrared images, we introduce a High-frequency Aware Encoder. To extract contextual information from multi-scale feature maps, we propose a Multi-scale Context Learning Bottleneck that incorporates contextual information repeatedly and performs cross-level fusion, which enables the recognition of small targets based on their surroundings. Finally, a lightweight Transformer Decoder is employed to restore the feature map, while placing attention on the target pixels. Experimental results on the IRSTD-1k dataset demonstrate that our method outperforms other state-of-the-art approaches. Qianchen Mao, Jinbao Wang 0001, Wenmin Wang 0001, Bingshu Wang |
ICASSP | 4 |
| 2024 | FreqFormer: Frequency-aware Transformer for Lightweight Image Super-resolution
Tao Dai 0001, Hang Guo 0002, Jinmin Li, Jinbao Wang 0001, Zexuan Zhu 0001 |
IJCAI | 5 |
| 2024 | Two in One Go: Single-stage Emotion Recognition with Decoupled Subject-context TransformerabstractEmotion recognition aims to discern the emotional state of subjects within an image, relying on subject-centric and contextual visual cues. Current approaches typically follow a two-stage pipeline: first localize subjects by off-the-shelf detectors, then perform emotion classification through the late fusion of subject and context features. However, the complicated paradigm suffers from disjoint training stages and limited fine-grained interaction between subject-context elements. To address the challenge, we present a single-stage emotion recognition approach, employing a Decoupled Subject-Context Transformer (DSCT), for simultaneous subject localization and emotion classification. Rather than compartmentalizing training stages, we jointly leverage box and emotion signals as supervision to enrich subject-centric feature learning. Furthermore, we introduce DSCT to facilitate interactions between fine-grained subject-context cues in a ''decouple-then-fuse'' manner. The decoupled query tokens-subject queries and context queries-gradually intertwine across layers within DSCT, during which spatial and semantic relations are exploited and aggregated. We evaluate our single-stage framework on two widely used context-aware emotion recognition datasets, CAER-S and EMOTIC. Our approach surpasses two-stage alternatives with fewer parameter numbers, achieving a 3.39% accuracy improvement and a 6.46% average precision gain on CAER-S and EMOTIC datasets, respectively. Code and models are available at: https://github.com/Sampson-Lee/DSCT. Xinpeng Li 0004, Teng Wang 0007, Jian Zhao 0006, Shuyi Mao, Jinbao Wang 0001, Feng Zheng 0001, Xiaojiang Peng, Xuelong Li 0001 |
ACM Multimedia | 5 |
| 2024 | Towards High-resolution 3D Anomaly Detection via Group-Level Feature Contrastive LearningabstractHigh-resolution point clouds (HRPCD) anomaly detection (AD) plays a critical role in precision machining and high-end equipment manufacturing. Despite considerable 3D-AD methods that have been proposed recently, they still cannot meet the requirements of the HRPCD-AD task. There are several challenges: i) It is difficult to directly capture HRPCD information due to large amounts of points at the sample level; ii) The advanced transformer-based methods usually obtain anisotropic features, leading to degradation of the representation; iii) The proportion of abnormal areas is very small, which makes it difficult to characterize. To address these challenges, we propose a novel group-level feature-based network, called Group3AD, which has a significantly efficient representation ability. First, we design an Intercluster Uniformity Network (IUN) to present the mapping of different groups in the feature space as several clusters, and obtain a more uniform distribution between clusters representing different parts of the point clouds in the feature space. Then, an Intracluster Alignment Network (IAN) is designed to encourage groups within the cluster to be distributed tightly in the feature space. In addition, we propose an Adaptive Group-Center Selection (AGCS) based on geometric information to improve the pixel density of potential anomalous regions during inference. The experimental results verify the effectiveness of our proposed Group3AD, which surpasses Reg3D-AD by the margin of 5% in terms of object-level AUROC on Real3D-AD. We provide the code and supplementary information on our website: https://github.com/M-3LAB/Group3AD. Hongze Zhu, Guoyang Xie, Chengbin Hou, Tao Dai 0001, Can Gao, Jinbao Wang 0001, LinLin Shen |
ACM Multimedia | 6 |
| 2024 | HairDiffusion: Vivid Multi-Colored Hair Editing via Latent DiffusionabstractHair editing is a critical image synthesis task that aims to edit hair color and hairstyle using text descriptions or reference images, while preserving irrelevant attributes (e.g., identity, background, cloth). Many existing methods are based on StyleGAN to address this task. However, due to the limited spatial distribution of StyleGAN, it struggles with multiple hair color editing and facial preservation. Considering the advancements in diffusion models, we utilize Latent Diffusion Models (LDMs) for hairstyle editing. Our approach introduces Multi-stage Hairstyle Blend (MHB), effectively separating control of hair color and hairstyle in diffusion latent space. Additionally, we train a warping module to align the hair color with the target region. To further enhance multi-color hairstyle editing, we fine-tuned a CLIP model using a multi-color hairstyle dataset. Our method not only tackles the complexity of multi-color hairstyles but also addresses the challenge of preserving original colors during diffusion editing. Extensive experiments showcase the superiority of our method in editing multi-color hairstyles while preserving facial attributes given textual descriptions and reference images. Yang Zhang 0012, LinLin Shen, Kaijun Deng, Weizhao He, Jinbao Wang 0001 |
NeurIPS | 7 |
| 2024 | Skeleton Cluster Tracking for robust multi-view multi-person 3D human pose estimation
Zehai Niu, Ke Lu 0002, Jian Xue 0002, Jinbao Wang 0001 |
Comput. Vis. Image Underst. | 4 |
| 2024 | From Methods to Applications: A Review of Deep 3D Human Motion CaptureabstractMotion capture technology is crucial in various applications like animation, virtual reality and sports analysis. With the development of deep learning methods, significant progress has been experienced in this field, producing cost-effective and user-friendly solutions for various applications. This paper provides a comprehensive review of deep learning-based human motion capture techniques. Our review aims to bridge the gap between academic research and practical applications, providing valuable insights and guidance for researchers and practitioners in deep learning-based human motion capture. Our study puts forth a new application-oriented taxonomy that comprehensively summarises five fundamental routes of motion capture technology. In addition to that, we also delve into the research priorities linked with each route, following the structure of “hardware requirements - technical routes - datasets - evaluation metrics” and extending the necessary criteria for transferring traditional motion capture systems to deep learning-based ones. Meanwhile, for the motion capture technology, the current state of the art is reviewed, the challenges are identified, and the future directions of the research are outlined. Zehai Niu, Ke Lu 0002, Jian Xue 0002, Xiaoyu Qin 0001, Jinbao Wang 0001, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | IM-IAD: Industrial Image Anomaly Detection Benchmark in ManufacturingabstractImage anomaly detection (IAD) is an emerging and vital computer vision task in industrial manufacturing (IM). Recently, many advanced algorithms have been reported, but their performance deviates considerably with various IM settings. We realize that the lack of a uniform IM benchmark is hindering the development and usage of IAD methods in real-world applications. In addition, it is difficult for researchers to analyze IAD algorithms without a uniform benchmark. To solve this problem, we propose a uniform IM benchmark, for the first time, to assess how well these algorithms perform, which includes various levels of supervision (unsupervised versus fully supervised), learning paradigms (few-shot, continual and noisy label), and efficiency (memory usage and inference speed). Then, we construct a comprehensive IAD benchmark (IM-IAD), which includes 19 algorithms on seven major datasets with a uniform setting. Extensive experiments (17 017 total) on IM-IAD provide in-depth insights into IAD algorithm redesign or selection. Moreover, the proposed IM-IAD benchmark challenges existing algorithms and suggests future research directions. For reproducibility and accessibility, the source code is uploaded to the website: https://github.com/M-3LAB/open-iad. Guoyang Xie, Jinbao Wang 0001, Jiaqi Liu 0004, Jiayi Lyu, Yong Liu 0032, Chengjie Wang 0001, Feng Zheng 0001, Yaochu Jin |
IEEE Trans. Cybern. | 2 |
| 2024 | Cross-Modal Alternating Learning With Task-Aware Representations for Continual LearningabstractContinual learning is a research field of artificial neural networks to simulate human lifelong learning ability. Although a surge of investigations has achieved considerable performance, most rely only on image modality for incremental image recognition tasks. In this paper, we propose a novel yet effective framework coined cross-modal Alternating Learning with Task-Aware representations (ALTA) to make good use of visual and linguistic modal information and achieve more effective continual learning. To do so, ALTA presents a cross-modal joint learning mechanism that leverages simultaneous learning of image and text representations to provide more effective supervision. And it mitigates forgetting by endowing task-aware representations with continual learning capability. Concurrently, considering the dilemma of stability and plasticity, ALTA proposes a cross-modal alternating learning strategy that alternately learns the task-aware cross-modal representations to match the image-text pairs between tasks better, further enhancing the ability of continual learning. We conduct extensive experiments under various popular image classification benchmarks to demonstrate that our approach achieves state-of-the-art performance. At the same time, systematic ablation studies and visualization analyses validate the effectiveness and rationality of our method. Our code will be available upon publication. Wujin Li, Bin-Bin Gao, Bizhong Xia, Jinbao Wang 0001, Jun Liu 0116, Yong Liu 0032, Chengjie Wang 0001, Feng Zheng 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Pushing the Limits of Fewshot Anomaly Detection in Industry Vision: Graphcore
Guoyang Xie, Jinbao Wang 0001, Jiaqi Liu 0004, Yaochu Jin, Feng Zheng 0001 |
ICLR | 2 |
| 2023 | EasyNet: An Easy Network for 3D Industrial Anomaly Detectionabstract3D anomaly detection is an emerging and vital computer vision task in industrial manufacturing (IM). Recently many advanced algorithms have been published, but most of them cannot meet the needs of IM. There are several disadvantages: i) difficult to deploy on production lines since their algorithms heavily rely on large pretrained models; ii) hugely increase storage overhead due to overuse of memory banks; iii) the inference speed cannot be achieved in real-time. To overcome these issues, we propose an easy and deployment-friendly network (called EasyNet) without using pretrained models and memory banks: firstly, we design a multi-scale multi-modality feature encoder-decoder to accurately reconstruct the segmentation maps of anomalous regions and encourage the interaction between RGB images and depth images; secondly, we adopt a multi-modality anomaly segmentation network to achieve a precise anomaly map; thirdly, we propose an attention-based information entropy fusion module for feature fusion during inference, making it suitable for real-time deployment. Extensive experiments show that EasyNet achieves an anomaly detection AUROC of 92.6% without using pretrained models and memory banks. In addition, EasyNet is faster than existing methods, with a high frame rate of 94.55 FPS on a Tesla V100 GPU. Guoyang Xie, Jiaqi Liu 0004, Jinbao Wang 0001, Ziqi Luo, Jinfan Wang, Feng Zheng 0001 |
ACM Multimedia | 4 |
| 2023 | Real3D-AD: A Dataset of Point Cloud Anomaly DetectionabstractHigh-precision point cloud anomaly detection is the gold standard for identifying the defects of advancing machining and precision manufacturing. Despite some methodological advances in this area, the scarcity of datasets and the lack of a systematic benchmark hinder its development. We introduce Real3D-AD, a challenging high-precision point cloud anomaly detection dataset, addressing the limitations in the field. With 1,254 high-resolution 3D items (from forty thousand to millions of points for each item), Real3D-AD is the largest dataset for high-precision 3D industrial anomaly detection to date. Real3D-AD surpasses existing 3D anomaly detection datasets available in terms of point cloud resolution (0.0010mm-0.0015mm), $360^{\circ}$ degree coverage and perfect prototype. Additionally, we present a comprehensive benchmark for Real3D-AD, revealing the absence of baseline methods for high-precision point cloud anomaly detection. To address this, we propose Reg3D-AD, a registration-based 3D anomaly detection method incorporating a novel feature memory bank that preserves local and global representations. Extensive experiments on the Real3D-AD dataset highlight the effectiveness of Reg3D-AD. For reproducibility and accessibility, we provide the Real3D-AD dataset, benchmark source code, and Reg3D-AD on our website: https://github.com/M-3LAB/Real3D-AD. Jiaqi Liu 0004, Guoyang Xie, Xinpeng Li 0004, Jinbao Wang 0001, Yong Liu 0032, Chengjie Wang 0001, Feng Zheng 0001 |
NeurIPS | 5 |
| 2023 | FedMed-GAN: Federated domain translation on unsupervised cross-modality brain image synthesis
Jinbao Wang 0001, Guoyang Xie, Yawen Huang, Jiayi Lyu, Feng Zheng 0001, Yefeng Zheng 0001, Yaochu Jin |
Neurocomputing | 1 |
| 2023 | Continuous cross-modal hashing
Hao Zheng 0008, Jinbao Wang 0001, Xiantong Zhen, Jingkuan Song, Feng Zheng 0001, Ke Lu 0002, Guo-Jun Qi |
Pattern Recognit. | 2 |
| 2022 | Towards Continual Adaptation in Industrial Anomaly DetectionabstractAnomaly detection (AD) has gained widespread attention due to its ability to identify defects in industrial scenarios using only normal samples. Although traditional AD methods achieved acceptable performance, they mainly focus on the current set of examples solely, leading to catastrophic forgetting of previously learned tasks when trained on a new one. Due to the limitation of flexibility and the requirements of realistic industrial scenarios, it is urgent to enhance the ability of continual adaptation of AD models. Therefore, this paper proposes a unified framework by incorporating continual learning (CL) to achieve our newly designed task of continual anomaly detection (CAD). Note that, we observe that data augmentation strategy can make AD methods well adapted to supervised CL (SCL) via constructing anomaly samples. Based on this, we hence propose a novel method named Distribution of Normal Embeddings (DNE), which utilizes the feature distribution of normal training samples from past tasks. It not only effectively alleviates catastrophic forgetting in CAD but also can be integrated with SCL methods to further improve their performance. Extensive experiments and visualization results on the popular benchmark dataset MVTec AD, have demonstrated advanced performance and the excellent continual adaption ability of our proposed method compared to other AD methods. To the best of our knowledge, we are the first to introduce and tackle the task of CAD. We believe that the proposed task and benchmark will be beneficial to the field of AD. Our code is available in thesupplementary material. Wujin Li, Jiawei Zhan, Jinbao Wang 0001, Bizhong Xia, Bin-Bin Gao, Jun Liu 0116, Chengjie Wang 0001, Feng Zheng 0001 |
ACM Multimedia | 3 |
| 2022 | FedMed-ATL: Misaligned Unpaired Cross-Modality Neuroimage Synthesis via Affine Transform LossabstractThe existence of completely aligned and paired multi-modal neuroimaging data has proved its effectiveness in the diagnosis of brain diseases. However, collecting the full set of well-aligned and paired data is impractical, since the practical difficulties may include high cost, long time acquisition, image corruption, and privacy issues. Previously, the misaligned unpaired neuroimaging data (termed as MUD) are generally treated as noisy labels. However, such a noisy label-based method fails to accomplish well when misaligned data occurs distortions severely. For example, the angle of rotation is different. In this paper, we propose a novel federated self-supervised learning (FedMed) for brain image synthesis. An affine transform loss (ATL) was formulated to make use of severely distorted images without violating privacy legislation for the hospital. We then introduce a new data augmentation procedure for self-supervised training and fed it into three auxiliary heads, namely auxiliary rotation, auxiliary translation, and auxiliary scaling heads. The proposed method demonstrates the advanced performance in both the quality of our synthesized results under a severely misaligned and unpaired data setting, and better stability than other GAN-based algorithms. The proposed method also reduces the demand for deformable registration while encouraging to leverage the misaligned and unpaired data. Experimental results verify the outstanding performance of our learning paradigm compared to other state-of-the-art approaches. Jinbao Wang 0001, Guoyang Xie, Yawen Huang, Yefeng Zheng 0001, Yaochu Jin, Feng Zheng 0001 |
ACM Multimedia | 1 |
| 2022 | SoftPatch: Unsupervised Anomaly Detection with Noisy DataabstractAlthough mainstream unsupervised anomaly detection (AD) algorithms perform well in academic datasets, their performance is limited in practical application due to the ideal experimental setting of clean training data. Training with noisy data is an inevitable problem in real-world anomaly detection but is seldom discussed. This paper considers label-level noise in image sensory anomaly detection for the first time. To solve this problem, we proposed a memory-based unsupervised AD method, SoftPatch, which efficiently denoises the data at the patch level. Noise discriminators are utilized to generate outlier scores for patch-level noise elimination before coreset construction. The scores are then stored in the memory bank to soften the anomaly detection boundary. Compared with existing methods, SoftPatch maintains a strong modeling ability of normal data and alleviates the overconfidence problem in coreset. Comprehensive experiments in various noise scenes demonstrate that SoftPatch outperforms the state-of-the-art AD methods on the MVTecAD and BTAD benchmarks and is comparable to those methods under the setting without noise. Xi Jiang 0009, Jinbao Wang 0001, Qiang Nie, Yong Liu 0032, Chengjie Wang 0001, Feng Zheng 0001 |
NeurIPS | 3 |
| 2021 | Seminar Learning for Click-Level Weakly Supervised Semantic SegmentationabstractAnnotation burden has become one of the biggest barriers to semantic segmentation. Approaches based on click-level annotations have therefore attracted increasing attention due to their superior trade-off between supervision and annotation cost. In this paper, we propose seminar learning, a new learning paradigm for semantic segmentation with click-level supervision. The fundamental rationale of seminar learning is to leverage the knowledge from different networks to compensate for insufficient information provided in click-level annotations. Mimicking a seminar, our seminar learning involves a teacher-student and a student-student module, where a student can learn from both skillful teachers and other students. The teacher-student module uses a teacher network based on the exponential moving average to guide the training of the student network. In the student-student module, heterogeneous pseudo-labels are proposed to bridge the transfer of knowledge among students to enhance each other’s performance. Experimental results demonstrate the effectiveness of seminar learning, which achieves the new state-of-the-art performance of 72.51% (mIOU), surpassing previous methods by a large margin of up to 16.88% on the Pascal VOC 2012 dataset. Jinbao Wang 0001, Hong Cai Chen, Xiantong Zhen, Feng Zheng 0001, Rongrong Ji, Ling Shao 0001 |
ICCV | 2 |
| 2021 | Deep 3D human pose estimation: A reviewabstractThree-dimensional (3D) human pose estimation involves estimating the articulated 3D joint locations of a human body from an image or video. Due to its widespread applications in a great variety of areas, such as human motion analysis, human–computer interaction, robots, 3D human pose estimation has recently attracted increasing attention in the computer vision community, however, it is a challenging task due to depth ambiguities and the lack of in-the-wild datasets. A large number of approaches, with many based on deep learning, have been developed over the past decade, largely advancing the performance on existing benchmarks. To guide future development, a comprehensive literature review is highly desired in this area. However, existing surveys on 3D human pose estimation mainly focus on traditional methods and a comprehensive review on deep learning based methods remains lacking in the literature. In this paper, we provide a thorough review of existing deep learning based works for 3D pose estimation, summarize the advantages and disadvantages of these methods and provide an in-depth understanding of this area. Furthermore, we also explore the commonly-used benchmark datasets on which we conduct a comprehensive study for comparison and analysis. Our study sheds light on the state of research development in 3D human pose estimation and provides insights that can facilitate the future design of models and algorithms. Jinbao Wang 0001, Shujie Tan, Xiantong Zhen, Feng Zheng 0001, Zhenyu He 0001, Ling Shao 0001 |
Comput. Vis. Image Underst. | 1 |
| 2021 | Learning Efficient Hash Codes for Fast Graph-Based Data Similarity RetrievalabstractTraditional operations, e.g. graph edit distance (GED), are no longer suitable for processing the massive quantities of graph-structured data now available, due to their irregular structures and high computational complexities. With the advent of graph neural networks (GNNs), the problems of graph representation and graph similarity search have drawn particular attention in the field of computer vision. However, GNNs have been less studied for efficient and fast retrieval after graph representation. To represent graph-based data, and maintain fast retrieval while doing so, we introduce an efficient hash model with graph neural networks (HGNN) for a newly designed task (i.e. fast graph-based data retrieval). Due to its flexibility, HGNN can be implemented in both an unsupervised and supervised manner. Specifically, by adopting a graph neural network and hash learning algorithms, HGNN can effectively learn a similarity-preserving graph representation and compute pair-wise similarity or provide classification via low-dimensional compact hash codes. To the best of our knowledge, our model is the first to address graph hashing representation in the Hamming space. Our experimental results reach comparable prediction accuracy to full-precision methods and can even outperform traditional models in some cases. In real-world applications, using hash codes can greatly benefit systems with smaller memory capacities and accelerate the retrieval speed of graph-structured data. Hence, we believe the proposed HGNN has great potential in further research. Jinbao Wang 0001, Feng Zheng 0001, Ke Lu 0002, Jingkuan Song, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | A Markerless Body Motion Capture System for Character Animation Based on Multi-view CamerasabstractA novel application system is proposed in this paper to achieve the generation of 3D character animation driven by markerless human body motion capture. The whole pipeline of the system consists of four parts: capturing motion data by multiple cameras, detecting 2D human body joints and estimating 3D joints, calculating bone transformation matrices, and generating character animation. Its main objective is to generate 3D skeleton and animation for 3D characters from multi-view images captured by ordinary cameras. The computation complexity of 3D skeleton reconstruction based on 3D vision is reduced accordingly to achieve the frame-by-frame motion capture. The experimental results show that our system is effective and efficient for capturing human action and animating 3D cartoon characters simultaneously. Jinbao Wang 0001, Ke Lyu, Yanfu Yan |
ICASSP | 1 |
| 2018 | Single Image Dehazing Based on the Physical Model and MSRCR AlgorithmabstractTo address the hazy weather image degradation problem, we propose a single image dehazing method based on a physical model and the brightness components of the image by using a multi-scale retinex with color restoration algorithm. The overall dehazing process involves three components, including the atmospheric light value calculation, transmission map estimation, and recovery of the hazy image scene radiance. Our contribution is that we propose a novel algorithm to dehaze a single image by calculating the atmospheric light value and computing the transmission map while considering the dynamic range of the image. Experimental results show that our algorithm can effectively improve the image quality degraded by foggy weather and retain sufficient image details. Jinbao Wang 0001, Ke Lu 0002, Jian Xue 0002, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Convex optimization based low-rank matrix decomposition for image restoration
Jinbao Wang 0001, Lulu Zhang 0006, Ke Lu 0002 |
Neurocomputing | 2 |
| 2016 | Non-local sparse regularization model with application to image denoising
Jinbao Wang 0001, Lulu Zhang 0006, Guang-Mei Xu, Ke Lu 0002 |
Multim. Tools Appl. | 2 |
| 2015 | Single image dehazing with a physical model and dark channel prior
Jinbao Wang 0001, Lulu Zhang 0006, Ke Lu 0002 |
Neurocomputing | 1 |
| 2015 | An improved fractional-order differentiation model for image denoising
Jinbao Wang 0001, Lulu Zhang 0006, Ke Lu 0002 |
Signal Process. | 2 |
| 2014 | Single-image motion deblurring using an adaptive image prior
Ke Lu 0002, Bing-Kun Bao, Lulu Zhang 0006, Jinbao Wang 0001 |
Inf. Sci. | 5 |