VLDB 2026 Research / reviewers in the wild / expert
Jimin Xiao
dblp:10/11067
· DBLP profile ↗
118ranked-venue papers
6as first author
81since 2021 · last 2026
0000-0002-9416-2486ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 71 · 1 first-author · 62 since 2021Graphics, computer vision, multimedia, augmented reality and games · 71 · 5 first-author · 40 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | End-to-end railway obstacle detection enhanced by point cloud segmentation
Yuxing Yang, Kaizhong Xiao, Xiaolong Tuo, Liewei Wang, Siyue Yu, Jimin Xiao |
Eng. Appl. Artif. Intell. | 9 |
| 2026 | FFEvent: Fast fourier-based knowledge transfer for event cameras
Yuhui Lin, Siyue Yu, Jimin Xiao, Jiaxuan Lu |
Expert Syst. Appl. | 4 |
| 2026 | Contrastive prompt clustering for weakly supervised semantic segmentation
Wangyu Wu, Wenqiao Zhang, Xianglin Qiu, Siqi Song, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
Expert Syst. Appl. | 9 |
| 2026 | LLM-enhanced multimodal fusion for cross-domain sequential recommendation
Wangyu Wu, Wenqiao Zhang, Siqi Song, Xianglin Qiu, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
Expert Syst. Appl. | 8 |
| 2026 | FANeRV: frequency separation and augmentation based neural representation for video
Li Yu 0004, Jimin Xiao, Moncef Gabbouj |
Expert Syst. Appl. | 4 |
| 2026 | CoRe: Contrast and reconstruction combination self-supervised point cloud representation learning
Changyu Zeng, Jimin Xiao, Anh Nguyen 0003, Xuming Hu, Wei Wang 0042, Yutao Yue |
Expert Syst. Appl. | 2 |
| 2026 | Breaking Redundancy via 3D Sparse Geometry: 3D-aware Neural Compression for Multi-View Videos
Shiwei Wang 0005, Liquan Shen, Jimin Xiao, Zhaoyi Tian, Feifeng Wang, Xiangyu Hu 0003, Yao Zhu 0006, Guorui Feng |
Int. J. Comput. Vis. | 3 |
| 2026 | Probing 3D anomalies via multi-view registration and dual-residual analysis
Yuxing Yang, Zeyu Fu, Liewei Wang, Siyue Yu, Jimin Xiao |
Neurocomputing | 6 |
| 2026 | Unleashing the power of optimal head in CLIP and DINO for weakly supervised semantic segmentation
Xianglin Qiu, Siyue Yu, Bingfeng Zhang, Tammam Tillo, Jimin Xiao |
Pattern Recognit. | 6 |
| 2026 | DiffClick: Click-differentiated enhancement network for interactive segmentation
Siqi Song, Siyue Yu, Huiyu Zhou 0001, Xiaowei Huang 0001, Limin Yu, Jimin Xiao |
Pattern Recognit. | 6 |
| 2026 | MvP-Diff: Multivariate yet precise diffusion for anomaly images synthesis and segmentation
Siyue Yao, Eng Gee Lim, Siyue Yu, Jimin Xiao, Mingjie Sun |
Pattern Recognit. | 4 |
| 2026 | Adaptive sparse contrastive learning for unsupervised object re-identification
Dingyuan Zheng, Yang Liu 0195, Jimin Xiao, Bingfeng Zhang |
Pattern Recognit. | 4 |
| 2026 | CoMasTRe+: Unleashing Disentangled Continual Segmentation With Mixture of Continual AdaptersabstractContinual Semantic Segmentation (CSS) suffers from catastrophic forgetting, particularly challenging for traditional per-pixel methods. Our prior work, CoMasTRe (CVPR 2024), introduced a query-based approach leveraging objectness by disentangling CSS into objectness learning and class recognition stages. While effective, CoMasTRe exhibited performance limitations due to feature forgetting within its pixel decoder. This paper presents CoMasTRe+, an enhanced framework specifically designed to overcome this limitation. The core contribution is a novel plugin, the Mixture of Continual Adapters (MoCA), integrated into the pixel decoder. MoCA is a dynamic architecture that mitigates feature forgetting by learning task-specific expert adapters. Crucially, MoCA employs a task-aware routing strategy and a novel adaptive routing distillation objective, tailored for continual learning, to preserve specialized feature representations across sequential tasks. CoMasTRe+ further enhances the class decoder using MoCA for improved recognition and simplicity. We extensively evaluate CoMasTRe+ on PASCAL VOC and ADE20K for continual semantic and panoptic segmentation. Experiments demonstrate that CoMasTRe+ effectively addresses the identified feature forgetting issue, significantly outperforms the original CoMasTRe, and achieves state-of-the-art results compared to both per-pixel and query-based baselines. Yizheng Gong, Siyue Yu, Liquan Shen, Jimin Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | CNC: Cross-modal Normality Constraint for Unsupervised Multi-class Anomaly DetectionabstractExisting unsupervised distillation-based methods rely on the differences between encoded and decoded features to locate abnormal regions in test images. However, the decoder trained only on normal samples still reconstructs abnormal patch features well, degrading performance. This issue is particularly pronounced in unsupervised multi-class anomaly detection tasks. We attribute this behavior to ‘over-generalization’ (OG) of decoder: the significantly increasing diversity of patch patterns in multi-class training enhances the model generalization on normal patches, but also inadvertently broadens its generalization to abnormal patches. To mitigate ‘OG’, we propose a novel approach that leverages class-agnostic learnable prompts to capture common textual normality across various visual patterns, and then apply them to guide the decoded features towards a ‘normal’ textual representation, suppressing ‘over-generalization’ of the decoder on abnormal patterns. To further improve performance, we also introduce a gated mixture-of-experts module to specialize in handling diverse patch patterns and reduce mutual interference between them in multi-class training. Our method achieves competitive performance on the MVTec AD and VisA datasets, demonstrating its effectiveness. Xiaoyang Wang 0007, Huihui Bai 0001, Eng Gee Lim, Jimin Xiao |
AAAI | 5 |
| 2025 | Cognitive-Inspired Hierarchical Attention Fusion With Visual and Textual for Cross-Domain Sequential Recommendation
Wangyu Wu, Siqi Song, Xianglin Qiu, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
CogSci | 7 |
| 2025 | POT: Prototypical Optimal Transport for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) leverages Class Activation Maps (CAMs) to extract spatial information from image-level labels. However, CAMs primarily highlight the most discriminative foreground regions, leading to incomplete results. Prototype-based methods attempt to address this limitation by employing prototype CAMs instead of classifier CAMs. Nevertheless, existing prototype-based methods typically use a single prototype for each class, which is insufficient to capture all attributes of the foreground features due to the significant intra-class variations across different images. Consequently, these methods still struggle with incomplete CAM predictions. In this paper, we propose a novel framework called Prototypical Optimal Transport (POT) for WSSS. POT enhances CAM predictions by dividing features into multiple clusters and activating each cluster using its prototype. In this process, a similarity-aware optimal transport is employed to assign features to the most probable clusters. This similarity-aware strategy ensures the prioritization of significant cluster prototypes, thereby improving the accuracy of feature assignment. Additionally, we introduce an adaptive OT-based consistency loss to refine feature representations. This framework effectively overcomes the limitations of single-prototype methods, providing more complete and accurate CAM predictions. Extensive experimental results on standard WSSS benchmarks (PASCAL VOC and MS COCO) demonstrate that our method significantly improves the quality of CAMs and achieves state-of-the-art performances. The source code will be released https://github.com/jianwang91/POT. Jian Wang 0122, Tianhong Dai, Bingfeng Zhang, Siyue Yu, Eng Gee Lim, Jimin Xiao |
CVPR | 6 |
| 2025 | FFR: Frequency Feature Rectification for Weakly Supervised Semantic SegmentationabstractImage-level Weakly Supervised Semantic Segmentation (WSSS) has garnered significant attention due to its low annotation costs. Current single-stage state-of-the-art WSSS methods mainly rely on Vision Transformer (ViT) to extract features from input images, generating more complete segmentation results based on comprehensive semantic information. However, these ViT-based methods often suffer from over-smoothing issues in segmentation results. In this paper, we identify that attenuated high-frequency features mislead the decoder of ViT-based WSSS models, resulting in over-smoothed false segmentation. To address this, we propose a Frequency Feature Rectification (FFR) framework to rectify the false segmentations caused by attenuated high-frequency features and enhance the learning of high-frequency features in the decoder. Quantitative and qualitative experimental results demonstrate that our FFR framework can effectively address the attenuated high-frequency caused over-smoothed segmentation issue and achieve new state-of-the-art WSSS performances. Codes are available at https://github.com/yay97/FFR. Ziqian Yang, Xinqiao Zhao, Jimin Xiao |
CVPR | 5 |
| 2025 | Feature Purification Matters: Suppressing Outlier Propagation for Training-Free Open-Vocabulary Semantic Segmentation
Siyue Yu, Bingfeng Zhang, Mingjie Sun, Yi Dong 0002, Jimin Xiao |
ICCV | 6 |
| 2025 | Bias-Resilient Weakly Supervised Semantic Segmentation Using Normalizing Flows
Xianglin Qiu, Xiaoyang Wang 0007, Jimin Xiao |
ICCV | 4 |
| 2025 | Class Token as Proxy: Optimal Transport-Assisted Proxy Learning for Weakly Supervised Semantic Segmentation
Jian Wang 0122, Tianhong Dai, Bingfeng Zhang, Siyue Yu, Eng Gee Lim, Jimin Xiao |
ICCV | 6 |
| 2025 | DecAD: Decoupling Anomalies in Latent Space for Multi-Class Unsupervised Anomaly Detection
Xiaoyang Wang 0007, Huihui Bai 0001, Eng Gee Lim, Jimin Xiao |
ICCV | 5 |
| 2025 | DriftRemover: Hybrid Energy Optimizations for Anomaly Images Synthesis and SegmentationabstractThis paper tackles the challenge of anomaly image synthesis and segmentation to generate various anomaly images and their segmentation labels to mitigate the issue of data scarcity. Existing approaches employ the precise mask to guide the generation, relying on additional mask generators, leading to increased computational costs and limited anomaly diversity. Although a few works use coarse masks as the guidance to expand diversity, they lack effective generation of labels for synthetic images, thereby reducing their practicality. Therefore, our proposed method simultaneously generates anomaly images and their corresponding masks by utilizing coarse masks and anomaly categories. The framework utilizes attention maps from synthesis process as mask labels and employs two optimization modules to tackle drift challenges, which are mismatches between synthetic results and real situations. Our evaluation demonstrates that our method improves pixel-level AP by 1.3% and F1-MAX by 1.8% in anomaly detection tasks on the MVTec dataset. Additionally, its successful application in practical scenarios highlights its effectiveness, improving IoU by 37.2% and F-measure by 25.1% with the Floor Dirt dataset. The code is available at https://github.com/JJessicaYao/DriftRemover. Siyue Yao, Mingjie Sun, Siyue Yu, Jimin Xiao, Eng Gee Lim |
IJCAI | 5 |
| 2025 | SDP: Spectral-Decomposed Prompting for Continual Learning
Siqi Song, Limin Yu, Jimin Xiao |
ACM Multimedia | 3 |
| 2025 | FAMRD: Frequency-Aware Multimodal Reverse Distillation for Industrial Anomaly DetectionabstractMultimodal Anomaly Detection (MMAD) has attracted significant attention in industrial defect inspection as it can simultaneously leverage the complementary information from different modalities to achieve higher-precision detection. Among existing MMAD approaches, dual-branch reverse distillation is widely adopted because of its efficiency in avoiding large-scale data storage. However, it suffers from two key issues. First, the alignment of cross-modal features can lead to a loss of modality-specific characteristics. Second, when one modality indicates normal while another shows anomalies, anomaly detection may be misled by that modality ambiguity. To address these challenges, we propose a Frequency-Aware Multimodal Reverse Distillation (FAMRD) framework from the frequency domain perspective. Specifically, we introduce a frequency spectral feature alignment module that aligns the low- and medium-frequency components across modalities to preserve global shape consistency, while maintaining high-frequency modality-specific details. In addition, we design a frequency spectral anomaly synthesis module. It perturbs the normal feature of one modality to create modality consistent anomalies, fuses it with another modality normal feature to mimic modality ambiguous anomalies, and adds them to the reverse distillation process for decision boundary optimization. Extensive experiments on standard MMAD benchmarks demonstrate that FAMRD achieves competitive performance in both anomaly detection and localization, outperforming state-of-the-art methods. Qiyin Zhong, Xianglin Qiu, Jimin Xiao |
ACM Multimedia | 6 |
| 2025 | Unifying Reconstruction and Density Estimation via Invertible Contraction Mapping in One-Class ClassificationabstractDue to the difficulty in collecting all unexpected abnormal patterns, One-Class Classification (OCC) has become the most popular approach to anomaly detection (AD). Reconstruction-based AD method relies on the discrepancy between inputs and reconstructed results to identify unobserved anomalies. However, recent methods trained only on normal samples may generalize to certain abnormal inputs, leading to well-reconstructed anomalies and degraded performance. To address this, we constrain reconstructions to remain on the normal manifold using a novel AD framework based on contraction mapping. This mapping guarantees that any input converges to a fixed point through iterations of this mapping. Based on this property, training the contraction mapping using only normal data ensures that its fixed point lies within the normal manifold. As a result, abnormal inputs are iteratively transformed toward the normal manifold, increasing the reconstruction error. In addition, the inherent invertibility of contraction mapping enables flow-based density estimation, where a prior distribution learned from the previous reconstruction is used to estimate the input likelihood for anomaly detection, further improving the performance. Using both mechanisms, we propose a bidirectional structure with forward reconstruction and backward density estimation. Extensive experiments on tabular data, natural image, and industrial image data demonstrate the effectiveness of our method. The code is available at URD. Tianhong Dai, Huihui Bai 0001, Yao Zhao 0001, Jimin Xiao |
NeurIPS | 5 |
| 2025 | Normal-Abnormal Guided Generalist Anomaly DetectionabstractGeneralist Anomaly Detection (GAD) aims to train a unified model on an original domain that can detect anomalies in new target domains. Previous GAD methods primarily use only normal samples as references, overlooking the valuable information contained in anomalous samples that are often available in real-world scenarios. To address this limitation, we propose a more practical approach: normal-abnormal-guided generalist anomaly detection, which leverages both normal and anomalous samples as references to guide anomaly detection across diverse domains. We introduce the Normal-Abnormal Generalist Learning (NAGL) framework, consisting of two key components: Residual Mining (RM) and Anomaly Feature Learning (AFL). RM extracts abnormal patterns from normal-abnormal reference residuals to establish transferable anomaly representations, while AFL adaptively learns anomaly features in query images through residual mapping to identify instance-aware anomalies. Our approach effectively utilizes both normal and anomalous references for more accurate and efficient cross-domain anomaly detection. Extensive experiments across multiple benchmarks demonstrate that our method significantly outperforms existing GAD approaches. This work represents the first to adopt a mixture of normal and abnormal samples as references in generalist anomaly detection. The code and datasets are available at https://github.com/JasonKyng/NAGL. Yizheng Gong, Jimin Xiao |
NeurIPS | 4 |
| 2025 | Adaptive Patch Contrast for Weakly Supervised Semantic Segmentation
Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Jimin Xiao, Fei Ma 0002, Renrong Ouyang |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Segmentation guided dual-branch classification for measuring fat infiltration in paraspinal musclesabstractMuscle fat infiltration (FI) is a significant change in muscle degeneration. In particular, fat infiltration in paraspinal muscles (PSMs) indicates lumbar degenerative diseases. Thus, classifying different grades of FI in PSMs plays an important role in diagnosing the relevant lumbar diseases. Recently, many deep-learning-based methods have been introduced into medical image tasks. However, such methods for classifying the grades of FI in PSMs have not been explored. In this case, this paper aims to involve deep-learning methods in grade classification for FI in PSMs. Firstly, we construct a PSMsFIGC dataset for deep learning exploration. Our PSMsFIGC dataset contains 4 grades for classification and the corresponding PSMs segmentation masks as assistance. Additionally, we propose a segmentation guided dual-branch classification framework (SGDC) to assist radiologists in confirming the grade of FI in PSMs. The structure mainly consists of a segmentation branch and a classification branch. The segmentation branch is designed to suppress the influence of irrelevant muscles for final classification. We further design a critical area indicator based on the prediction of the segmentation branch to involve more related crucial areas for the classification branch and thus bridge the two branches. However, we find that inter-class disturbance, caused by PSMs’ similar shape and features, makes the network easily fall into local optimal. Therefore, we propose a disturbance weakening module to relieve the disturbance. Extensive experiments show that our SGDC can surpass existing classific classification networks, e.g., the proposed method achieves an impressive accuracy of 89.0% on the PSMsFIGC dataset. Our dataset and code will be released at https://github.com/myjianghao/Segmentation-Guided-Dual-branch-Classification-Framework-SGDC- . • A novel FI classification dataset constructed for exploring in FI grade. • A dual-branch framework designed to predict FI grades. • A module leveraging masks to focus on lesion-related regions. • A Gaussian-based strategy to mitigate inter-class similarity in MRI. Chengnan Jing, Hao Jiang 0054, Jimin Xiao, Siyue Yu, Minfeng Gan |
Expert Syst. Appl. | 5 |
| 2025 | High-Frequency Enhanced Hybrid Neural Representation for video compression
Li Yu 0004, Jimin Xiao, Moncef Gabbouj |
Expert Syst. Appl. | 3 |
| 2025 | FAD: Feature augmented distillation for anomaly detection and localization
Qiyin Zhong, Xianglin Qiu, Xinqiao Zhao, Xiaowei Huang 0001, Jimin Xiao |
Expert Syst. Appl. | 6 |
| 2025 | Generative Prompt Controlled Diffusion for weakly supervised semantic segmentationabstractWeakly supervised semantic segmentation (WSSS), aiming to train segmentation models solely using image-level labels, has received significant attention. Existing approaches mainly concentrate on creating high-quality pseudo labels by utilizing existing images and their corresponding image-level labels. However, a major challenge arises when the available dataset is limited, as the quality of pseudo labels degrades significantly. In this paper, we tackle this challenge from a different perspective by introducing a novel approach called Generative Prompt Controlled Diffusion (GPCD) for data augmentation . This approach enhances the current labeled datasets by augmenting them with a variety of images, achieved through controlled diffusion guided by Generative Pre-trained Transformer (GPT) prompts. In this process, the existing images and image-level labels provide the necessary control information , while GPT enriches the prompts to generate diverse backgrounds. Moreover, we make an original contribution by integrating data source information as tokens into the Vision Transformer (ViT) framework, which improves the ability of downstream WSSS models to recognize the origins of augmented images. Our proposed GPCD approach clearly surpasses existing state-of-the-art methods, with its advantages being more pronounced when the available data is scarce, thereby demonstrating the effectiveness of our method. Our source code will be released. Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
Neurocomputing | 6 |
| 2025 | Frozen CLIP-DINO: A Strong Backbone for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation has witnessed great achievements with image-level labels. Several recent approaches use the CLIP model to generate pseudo labels for training an individual segmentation model, while there is no attempt to apply the CLIP model as the backbone to directly segment objects with image-level labels. In this paper, we propose WeCLIP and its advanced version WeCLIP+, to build the single-stage pipeline for weakly supervised semantic segmentation. For WeCLIP, the frozen CLIP model is applied as the backbone for semantic feature extraction, and a new light decoder is designed to interpret extracted semantic features for final prediction. Meanwhile, we utilize the above frozen backbone to generate pseudo labels for training the decoder. Such labels are fixed during training. We then propose a refinement module (RFM) to optimize them dynamically. For WeCLIP+, we introduce the frozen DINO model to achieve more comprehensive semantic feature extraction. The frozen DINO is combined with the frozen CLIP as the backbone, followed by a shared decoder to make predictions with less training cost. Moreover, a strengthened refinement module (RFM+) is designed to revise online pseudo labels with extra guidance from DINO features. Extensive experiments show that both WeCLIP and WeCLIP+ significantly outperform other approaches with less training cost. Particularly, WeCLIP+ gets mIoU of 83.9% on VOC 2012 test set and 56.3% on COCO val set. Additionally, these two approaches also obtain promising results for fully supervised settings. Bingfeng Zhang, Siyue Yu, Jimin Xiao, Yunchao Wei, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Revisiting 3D point cloud analysis with Markov process
Chenru Jiang, Wuwei Ma, Kaizhu Huang, Qiufeng Wang 0001, Xi Yang 0008, Weiguang Zhao, Junwei Wu 0001, Xinheng Wang 0001, Jimin Xiao, Zhenxing Niu |
Pattern Recognit. | 9 |
| 2025 | M-SEE: A multi-scale encoder enhancement framework for end-to-end Weakly Supervised Semantic Segmentation
Ziqian Yang, Xinqiao Zhao, Jimin Xiao |
Pattern Recognit. | 5 |
| 2025 | Dynamic feature regularized loss for weakly supervised semantic segmentation
Bingfeng Zhang, Jimin Xiao, Yao Zhao 0001 |
Pattern Recognit. | 2 |
| 2025 | Auxiliary captioning: Bridging image-text matching and image captioning
Hui Li 0085, Jimin Xiao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001 |
Signal Process. Image Commun. | 2 |
| 2025 | Enhancing Light Field Salient Object Detection With Variance-Maximized Key Focal Slice SelectionabstractLight field saliency object detection (LF SOD) methods have made significant progress recently. Most of them explore abundant multi-modal information from the all-focus image and the focal stacks at all focal planes to enrich scene details and depth perception. However, in light-field images, the spatial and depth information varies slightly across different slices, raising redundancy within focal stacks. Besides, the noise can appear repeatedly in multiple images of the focal stacks, which brings interference. To address these issues, in this work, we propose VMKNet, an effective approach that leverages innovative variance-maximized key slice selection and interacts with the all-focus image, to improve LF SOD. Specifically, we measure consistency differences between the all-focus image and each focal slice in the salient region as saliency scores. Then, we randomly assemble sets of them, where each score corresponds to a certain slice. The one exhibiting the highest variance is singled out to determine key focal slices as they reveal the diversity of salient objects. Then, the bidirectional guidance module (BGM) is presented to learn attentive features of all-focus and selected key slices in a mutual guidance manner, thus producing enhanced and holistic features. With hierarchical BGMs, our model can progressively aggregate common salient semantics and meaningful contextual details, generating more discriminative representations. Moreover, we introduce the edge enhancement module in conjunction with BGM to improve the sharpness of saliency maps. Extensive experiments on common light field datasets demonstrate that our method, termed VMKNet, outperforms recent state-of-the-art LF, RGB-D, and RGB methods. Our code is available athttps://github.com/Han-jiaxin/VMKNet. Jiaxin Han, Feng Li 0037, Mengmeng Zhang 0008, Huihui Bai 0001, Jimin Xiao, Yao Zhao 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | SFBM: Shared Feature Bias Mitigating for Long-Tailed Image RecognitionabstractLong-tailed distribution exists in real-world scenario and compromises the performance of recognition models. In this article, we point out that a neural network classifier has a shared feature bias, which tends to regard the shared features among different classes as head-class discriminative features, leading to misclassifications on tail-class samples under long-tailed scenarios. To solve this issue, we propose a shared feature bias mitigating (SFBM) framework. Specifically, we create two parallel classifiers trained concurrently with the baseline classifier, using our special training loss. The parallel classifier weight sums are then used for estimating the shared feature components in baseline classifier weights. Finally, we rectify the baseline classifier by removing the estimated shared feature components from it while supplementing the parallel classifier weights class by class to the rectified classifier weights, mitigating shared feature bias. Our proposed SFBM demonstrates broad compatibility with nearly all recognition methods while maintaining high computational efficiency, as it introduces no additional computation during inference. Extensive experiments on CIFAR10/100-LT, ImageNet-LT, and iNaturalist 2018 demonstrate that simply incorporating SFBM during the training phase consistently boosts the performance of various state-of-the-art methods by significant margins. The complete source code will be made publicly available at https://github.com/bzbz-bot/SFBM. Xinqiao Zhao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001, Jimin Xiao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | SFC: Shared Feature Calibration in Weakly Supervised Semantic SegmentationabstractImage-level weakly supervised semantic segmentation has received increasing attention due to its low annotation cost. Existing methods mainly rely on Class Activation Mapping (CAM) to obtain pseudo-labels for training semantic segmentation models. In this work, we are the first to demonstrate that long-tailed distribution in training data can cause the CAM calculated through classifier weights over-activated for head classes and under-activated for tail classes due to the shared features among head- and tail- classes. This degrades pseudo-label quality and further influences final semantic segmentation performance. To address this issue, we propose a Shared Feature Calibration (SFC) method for CAM generation. Specifically, we leverage the class prototypes which carry positive shared features and propose a Multi-Scaled Distribution-Weighted (MSDW) consistency loss for narrowing the gap between the CAMs generated through classifier weights and class prototypes during training. The MSDW loss counterbalances over-activation and under-activation by calibrating the shared features in head-/tail-class classifier weights. Experimental results show that our SFC significantly improves CAM boundaries and achieves new state-of-the-art performances. The project is available at https://github.com/Barrett-python/SFC. Xinqiao Zhao, Xiaoyang Wang 0007, Jimin Xiao |
AAAI | 4 |
| 2024 | Continual Segmentation with Disentangled Objectness Learning and Class RecognitionabstractMost continual segmentation methods tackle the prob-lem as a per-pixel classification task. However, such a paradigm is very challenging, and we find query-based seg-menters with built-in objectness have inherent advantages compared with per-pixel ones, as objectness has strong transfer ability and forgetting resistance. Based on these findings, we propose CoMasTRe by disentangling continual segmentation into two stages: forgetting-resistant continual objectness learning and well-researched continual classi-fication. CoMasTRe uses a two-stage segmenter learning class-agnostic mask proposals at the first stage and leaving recognition to the second stage. During continual learning, a simple but effective distillation is adopted to strengthen objectness. To further mitigate the forgetting of old classes, we design a multi-label class distillation strategy suited for segmentation. We assess the effectiveness of CoMas-TRe on PASCAL VOC and ADE20K. Extensive experiments show that our method outperforms per-pixel and query-based methods on both datasets. Code will be available at https://github.com/jordangong/CoMasTRe. Yizheng Gong, Siyue Yu, Xiaoyang Wang 0007, Jimin Xiao |
CVPR | 4 |
| 2024 | Towards the Uncharted: Density-Descending Feature Perturbation for Semi-supervised Semantic SegmentationabstractSemi-supervised semantic segmentation allows model to mine effective supervision from unlabeled data to complement label-guided training. Recent research has primarily focused on consistency regularization techniques, exploring perturbation-invariant training at both the image and feature levels. In this work, we proposed a novel feature-level consistency learning framework named Density-Descending Feature Perturbation (DDFP). Inspired by the low-density separation assumption in semi-supervised learning, our key insight is that feature density can shed a light on the most promising direction for the segmentation classifier to explore, which is the regions with lower density. We propose to shift features with confident predictions towards lower-density regions by perturbation injection. The perturbed features are then super-vised by the predictions on the original features, thereby compelling the classifier to explore less dense regions to effectively regularize the decision boundary. Central to our method is the estimation of feature density. To this end, we introduce a lightweight density estimator based on normalizing flow, allowing for efficient capture of the feature density distribution in an online manner. By extracting gradients from the density estimator, we can determine the direction towards less dense regions for each feature. The proposed DDFP outperforms other designs on feature-level perturbations and shows state of the art performances on both Pascal VOC and Cityscapes dataset under various partition protocols. The project is available at https://github.com/Gavinwxy/DDFP. Xiaoyang Wang 0007, Huihui Bai 0001, Limin Yu, Yao Zhao 0001, Jimin Xiao |
CVPR | 5 |
| 2024 | Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation has witnessed great achievements with image-level labels. Several recent approaches use the CLIP model to generate pseudo labels for training an individual segmentation model, while there is no attempt to apply the CLIP model as the backbone to directly segment objects with image-level labels. In this paper, we propose WeCLIP, a CLIP-based single-stage pipeline, for weakly supervised semantic segmentation. Specifically, the frozen CLIP model is applied as the backbone for semantic feature extraction, and a new decoder is designed to interpret extracted semantic features for final prediction. Meanwhile, we utilize the above frozen backbone to generate pseudo labels for training the decoder. Such labels cannot be optimized during training. We then propose a refinement module (RFM) to rectify them dynamically. Our architecture enforces the proposed decoder and RFM to benefit from each other to boost the final performance. Extensive experiments show that our approach significantly outperforms other approaches with less training cost. Additionally, our WeCLIP also obtains promising results for fully supervised settings. The code is available at https://github.com/zbf1991/WeCLIP. Bingfeng Zhang, Siyue Yu, Yunchao Wei, Yao Zhao 0001, Jimin Xiao |
CVPR | 5 |
| 2024 | PSDPM: Prototype-based Secondary Discriminative Pixels Mining for Weakly Supervised Semantic SegmentationabstractImage-level Weakly Supervised Semantic Segmentation (WSSS) has received increasing attention due to its low an-notation cost. Class Activation Mapping (CAM) generated through classifier weights in WSSS inevitably ignores cer-tain useful cues, while the CAM generated through class prototypes can alleviate that. However, because of the dif-ferent goals of image classification and semantic segmentation, the class prototypes still focus on activating primary discriminative pixels learned from classification loss, leading to incomplete CAM. In this paper, we propose a plug-and-play Prototype-based Secondary Discriminative Pixels Mining (PSDPM) framework for enabling class prototypes to activate more secondary discriminative pixels, thus gen-erating a more complete CAM. Specifically, we introduce a Foreground Pixel Estimation Module (FPEM) for esti-mating potential foreground pixels based on the correlations between primary and secondary discriminative pix-els and the semantic segmentation results of baseline meth-ods. Then, we enable WSSS model to learn discriminative features from secondary discriminative pixels through a consistency loss calculated between FPEM result and class-prototype CAM. Experimental results show that our PSDPM improves various baseline methods significantly and achieves new state-of-the-art performances on WSSS benchmarks. Codes are available at https://github.com/xinqiaozhao/PSDPM. Xinqiao Zhao, Ziqian Yang, Tianhong Dai, Bingfeng Zhang, Jimin Xiao |
CVPR | 5 |
| 2024 | Adversarial Erasing Transformer for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation has attracted a lot of attention recently. Previous methods can be divided into two types, which are single-stage training and multi-stage training. In this paper, we focus on multi-stage training for image-level weakly supervised semantic segmentation. Many recent methods have tried to use transformer architecture as the backbone for CAM generation since it can capture global relationships to refine CAM accurately. However, we observe that such a backbone still fails to generate complete and smooth CAM. We argue that this is because the attention mechanism in the transformer can only pay attention to the most discriminative relationships. It is difficult to capture semantic-level long-range pair-wise relationships under image-level supervision. Thus, we propose an adversarial erasing transformer network called AETN, where an erasing attention mechanism is designed to establish more extensive pair-wise relationships. To cope with erasing, more target features will be forced to activate. Thus, better feature representation can be obtained for more accurate CAM generation. Besides, to further help our network learn better feature representation, we propose a self-consistent learning mechanism based on different augmentations. In this way, our AETN outperforms recent methods. Our AETN achieves 73.0 mIoU on the PASCAL VOC 2012 val set and 73.9 mIoU on the PASCAL VOC 2012 test set. Code is available a https://github.com/siyueyu/AETN. Bingfeng Zhang, Siyue Yu, Xuru Gao, Mingjie Sun, Eng Gee Lim, Jimin Xiao |
ECAI | 6 |
| 2024 | Image Augmentation with Controlled Diffusion for Weakly-Supervised Semantic SegmentationabstractWeakly-supervised semantic segmentation (WSSS), which aims to train segmentation models solely using image-level labels, has achieved significant attention. Existing methods primarily focus on generating high-quality pseudo labels using available images and their image-level labels. However, the quality of pseudo labels degrades significantly when the size of available dataset is limited. Thus, in this paper, we tackle this problem from a different view by introducing a novel approach called Image Augmentation with Controlled Diffusion (IACD). This framework effectively augments existing labeled datasets by generating diverse images through controlled diffusion, where the available images and image-level labels are served as the controlling information. Moreover, we also propose a high-quality image selection strategy to mitigate the potential noise introduced by the randomness of diffusion models. In the experiments, our proposed IACD approach clearly surpasses existing state-of-the-art methods. This effect is more obvious when the amount of available data is small, demonstrating the effectiveness of our method. Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
ICASSP | 5 |
| 2024 | Top-K Pooling with Patch Contrastive Learning for Weakly-Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) using only image-level labels has gained significant attention due to cost-effectiveness. Recently, Vision Transformer (ViT) based methods without class activation map (CAM) have shown greater capability in generating reliable pseudo labels than previous methods using CAM. However, the current ViT-based methods utilize max pooling to select the patch with the highest prediction score to map the patch-level classification to the image-level one, which may affect the quality of pseudo labels due to the inaccurate classification of the patches. In this paper, we introduce a novel ViT-based WSSS method named top-K pooling with patch contrastive learning (TKP-PCL), which employs a top-K pooling layer to alleviate the limitations of previous max pooling selection. A patch contrastive error (PCE) is also proposed to enhance the patch embeddings to further improve the final results. The experimental results show that our approach is very efficient and outperforms other state-of-the-art WSSS methods on the PASCAL VOC 2012 and MS COCO 2014 dataset. Wangyu Wu, Tianhong Dai, Xiaowei Huang 0001, Fei Ma 0002, Jimin Xiao |
SMC | 5 |
| 2024 | Self-supervised learning for point cloud data: A surveyabstract3D point clouds are a crucial type of data collected by LiDAR sensors and widely used in transportation applications due to its concise descriptions and accurate localization. Deep neural networks (DNNs) have achieved remarkable success in processing large amount of disordered and sparse 3D point clouds, especially in various computer vision tasks, such as pedestrian detection and vehicle recognition. Among all the learning paradigms, Self-Supervised Learning (SSL), an unsupervised training paradigm that mines effective information from the data itself, is considered as an essential solution to solve the time-consuming and labor-intensive data labelling problems via smart pre-training task design. This paper provides a comprehensive survey of recent advances on SSL for point clouds. We first present an innovative taxonomy, categorizing the existing SSL methods into four broad categories based on the pretexts’ characteristics. Under each category, we then further categorize the methods into more fine-grained groups and summarize the strength and limitations of the representative methods. We also compare the performance of the notable SSL methods in literature on multiple downstream tasks on benchmark datasets both quantitatively and qualitatively. Finally, we propose a number of future research directions based on the identified limitations of existing SSL research on point clouds. Changyu Zeng, Wei Wang 0042, Anh Nguyen 0003, Jimin Xiao, Yutao Yue |
Expert Syst. Appl. | 4 |
| 2024 | Prototype Guided Pseudo Labeling and Perturbation-based Active Learning for domain adaptive semantic segmentation
Junkun Peng, Mingjie Sun, Eng Gee Lim, Qiufeng Wang 0001, Jimin Xiao |
Pattern Recognit. | 5 |
| 2024 | Cross-frame feature-saliency mutual reinforcing for weakly supervised video salient object detection
Jian Wang 0122, Siyue Yu, Bingfeng Zhang, Xinqiao Zhao, Ángel F. García-Fernández, Eng Gee Lim, Jimin Xiao |
Pattern Recognit. | 7 |
| 2024 | Unified Multi-Modality Video Object Segmentation Using Reinforcement LearningabstractThe main task we aim to tackle is the multi-modality video object segmentation (VOS), which can be divided into two sub-tasks: mask-referred and language-referred VOS, where the first-frame mask-level or language-level label is utilized to provide the target information, respectively. Due to the huge gap between different modalities, existing works never come up with a unified framework for these two sub-tasks. In this work, such a unified framework is designed, where the visual and linguistic inputs are first spilt into a number of image patches and words, and then mapped into same-size tokens, which are equally processed by a self-attention based segmentation model. Furthermore, to highlight the significant information and discard the non-target or ambiguous one, unified multi-modality filter networks are further designed, and reinforcement learning is adopted to optimize such networks. Experiments show that new state-of-the-art performances are achieved by the proposed method: 52.8% ofJ&Fon Ref-YoutubeVOS dataset and 83.2% ofJSon YoutubeVOS dataset, respectively. The code will be released. Mingjie Sun, Jimin Xiao, Eng Gee Lim, Cairong Zhao, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Class Activation Map Calibration for Weakly Supervised Semantic SegmentationabstractImage-level weakly supervised semantic segmentation (WSSS) has received substantial attention due to its cost-effective annotation process. In WSSS, Class Activation Maps (CAMs) generated via classifier weights tend to focus on the most discriminative region, while the CAMs derived from class prototypes are significantly enhanced to cover more complete regions. However, the prototype CAMs still exhibit limitations such as incomplete localization maps on target objects and the presence of background noise. In this paper, we propose a novel WSSS framework called Classifier-Prototype Mutual Calibration (CPMC) that leverages the characteristics of both classifier and prototype CAMs to address the above issues. Specifically, an iterative refinement strategy based on context feature dependency is applied to refine the original classifier CAMs, which helps to generate improved prototype CAMs. Subsequently, local prototypes are constructed based on the false negative regions and false positive regions extracted from the previous two CAMs, which contribute to completing missing parts of the target object and suppressing background noise respectively. Therefore, CPMC can alleviate the aforementioned issues. Extensive experimental results on standard WSSS benchmarks (PASCAL VOC and MS COCO) show that our method significantly improves the quality of CAMs and achieves state-of-the-art performance. Our source code will be released. Jian Wang 0122, Tianhong Dai, Xinqiao Zhao, Ángel F. García-Fernández, Eng Gee Lim, Jimin Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | Hunting Sparsity: Density-Guided Contrastive Learning for Semi-Supervised Semantic SegmentationabstractRecent semi-supervised semantic segmentation methods combine pseudo labeling and consistency regularization to enhance model generalization from perturbation-invariant training. In this work, we argue that adequate supervision can be extracted directly from the geometry of feature space. Inspired by density-based unsupervised clustering, we propose to leverage feature density to locate sparse regions within feature clusters defined by label and pseudo labels. The hypothesis is that lower-density features tend to be under-trained compared with those densely gathered. Therefore, we propose to apply regularization on the structure of the cluster by tackling the sparsity to increase intra-class compactness in feature space. With this goal, we present a Density-Guided Contrastive Learning (DGCL) strategy to push anchor features in sparse regions toward cluster centers approximated by high-density positive keys. The heart of our method is to estimate feature density which is defined as neighbor compactness. We design a multi-scale density estimation module to obtain the density from multiple nearest-neighbor graphs for robust density modeling. Moreover, a unified training framework is proposed to combine label-guided self-training and density-guided geometry regularization to form complementary supervision on unlabeled data. Experimental results on PAS-CAL VOC and Cityscapes under various semi-supervised settings demonstrate that our proposed method achieves state-of-the-art performances. The project is available at https://github.com/Gavinwxy/DGCL. Xiaoyang Wang 0007, Bingfeng Zhang, Limin Yu, Jimin Xiao |
CVPR | 4 |
| 2023 | FastRecon: Few-shot Industrial Anomaly Detection via Fast Feature ReconstructionabstractIn industrial anomaly detection, data efficiency and the ability for fast migration across products become the main concerns when developing detection algorithms. Existing methods tend to be data-hungry and work in the one-model-one-category way, which hinders their effectiveness in real-world industrial scenarios. In this paper, we propose a few-shot anomaly detection strategy that works in a low-data regime and can generalize across products at no cost. Given a defective query sample, we propose to utilize a few normal samples as a reference to reconstruct its normal version, where the final anomaly detection can be achieved by sample alignment. Specifically, we introduce a novel regression with distribution regularization to obtain the optimal transformation from support to query features, which guarantees the reconstruction result shares visual similarity with the query sample and meanwhile maintains the property of normal samples. Experimental results show that our method significantly outperforms previous state-of-the-art at both image and pixel-level AUROC performances from 2 to 8-shot scenarios. Besides, with only a limited number of training samples (less than 8 samples), our method reaches competitive performance with vanilla AD methods which are trained with extensive normal samples. The code is available at https://github.com/FzJun26th/FastRecon. Xiaoyang Wang 0007, Jiejie Liu, Qiugui Hu, Jimin Xiao |
ICCV | 6 |
| 2023 | Synchronize Feature Extracting and Matching: A Single Branch Framework for 3D Object TrackingabstractSiamese network has been a de facto benchmark framework for 3D LiDAR object tracking with a shared-parametric encoder extracting features from template and search region, respectively. This paradigm relies heavily on an additional matching network to model the cross-correlation/similarity of the template and search region. In this paper, we forsake the conventional Siamese paradigm and propose a novel single-branch framework, SyncTrack, synchronizing the feature extracting and matching to avoid forwarding encoder twice for template and search region as well as introducing extra parameters of matching network. The synchronization mechanism is based on the dynamic affinity of the Transformer, and an in-depth analysis of the relevance is provided theoretically. Moreover, based on the synchronization, we introduce a novel Attentive PointsSampling strategy into the Transformer layers (APST), replacing the random/Farthest Points Sampling (FPS) method with sampling under the supervision of attentive relations between the template and search region. It implies connecting point-wise sampling with the feature learning, beneficial to aggregating more distinctive and geometric features for tracking with sparse points. Extensive experiments on two benchmark datasets (KITTI and NuScenes) show that SyncTrack achieves state-of-the-art performance in realtime tracking. Teli Ma, Mengmeng Wang 0005, Jimin Xiao, Huifeng Wu, Yong Liu 0007 |
ICCV | 3 |
| 2023 | Credible Dual-Expert Learning for Weakly Supervised Semantic Segmentation
Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Yao Zhao 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | Robust generative adversarial network
Shufei Zhang, Zhuang Qian, Kaizhu Huang, Rui Zhang 0012, Jimin Xiao, Canyi Lu |
Mach. Learn. | 5 |
| 2023 | Aggregated pyramid gating network for human pose estimation without pre-training
Chenru Jiang, Kaizhu Huang, Shufei Zhang, Xinheng Wang 0001, Jimin Xiao, John Yannis Goulermas |
Pattern Recognit. | 5 |
| 2023 | Weight-guided class complementing for long-tailed image recognition
Xinqiao Zhao, Jimin Xiao, Siyue Yu, Hui Li 0085, Bingfeng Zhang |
Pattern Recognit. | 2 |
| 2023 | Weight-guided loss for long-tailed object detection and instance segmentation
Xinqiao Zhao, Jimin Xiao, Bingfeng Zhang, Waleed Al-Nuaimy |
Signal Process. Image Commun. | 2 |
| 2023 | Real-Time Prediction of Simulator Sickness in Virtual Reality GamesabstractVirtual reality (VR) technology has progressed rapidly and is used in various domains, particularly games. Simulator sickness (SS) still represents a significant problem for its wider adoption. The most common way to detect SS is using the simulator sickness questionnaire (SSQ). SSQ is a subjective measurement and is inadequate for real-time applications such as VR games. This research aims to develop a model to predict SS in real time using in-game characters’ movement and users’ eye motion data during gameplay in VR games. To achieve this, we designed an experiment to collect such data with three types of games. We trained a long short-term memory neural network with the eye-tracking and character movement data to predict SS. Our model can predict SS in real time with an accuracy of 83.4% for players who suffer from severe sensitivity to SS. Our results indicate that, in VR games, our model is an accurate and efficient method to predict SS in real time. Jialin Wang 0002, Hai-Ning Liang, Diego Monteiro 0001, Wenge Xu, Jimin Xiao |
IEEE Trans. Games | 5 |
| 2023 | Fully and Weakly Supervised Referring Expression Segmentation With End-to-End LearningabstractReferring Expression Segmentation (RES), which is aimed at localizing and segmenting the target according to the given language expression, has drawn increasing attention. Existing methods jointly consider the localization and segmentation steps, which rely on the fused visual and linguistic features for both steps. We argue that the conflict between the purpose of identifying an object and generating a mask limits the RES performance. To solve this problem, we propose a parallel position-kernel-segmentation pipeline to better isolate and then interact the localization and segmentation steps. In our pipeline, linguistic information will not directly contaminate the visual feature for segmentation. Specifically, the localization step localizes the target object in the image based on the referring expression, and then the visual kernel obtained from the localization step guides the segmentation step. This pipeline also enables us to train RES in a weakly-supervised way, where the pixel-level segmentation labels are replaced by click annotations on center and corner points. The position head is fully-supervised and trained with the click annotations as supervision, and the segmentation head is trained with weakly-supervised segmentation losses. To validate our framework on a weakly-supervised setting, we annotated three RES benchmark datasets (RefCOCO, RefCOCO+ and RefCOCOg) with click annotations. Our method is simple but surprisingly effective, outperforming all previous state-of-the-art RES methods on fully- and weakly-supervised settings by a large margin. The code and dataset will be released onhttps://github.com/detectiveli/PKS.git. Hui Li 0085, Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Plausible Proxy Mining With Credibility for Unsupervised Person Re-IdentificationabstractOne effective way to address unsupervised person re-identification is to use a clustering-based contrastive learning approach. Existing state-of-the-art methods adopt clustering algorithms (e.g., DBSCAN) and camera ID information to divide all person images into several camera-aware proxies. Then, for each person image, the extracted feature representation is pulled closer to the centroids of its pseudo-positive proxies (the proxies that share the same pseudo-identity label with this image) and pushed away from the centroids of other pseudo-negative proxies (the proxies that share the different pseudo-identity label with this image). However, the quality of the proxy centroid is significantly affected by the proxy impurity issue and thus deteriorates the learned feature representations. On the premise that we cannot introduce superior supervision signals by thoroughly solving the proxy impurity issue, for a person image, identifying its plausible proxies: the pseudo-negative proxies which potentially include its wrongly-clustered instances (the instances with the same ground-truth identity with this image), and further fixing the resulted incorrect supervision signals become an urgent and challenging problem. This paper proposes a simple yet effective approach to address this problem. With a given image, our method can effectively locate its plausible proxies. Then we introduce credibility to measure how much we should treat the centroid of each mined plausible proxy as a positive supervision signal rather than entirely negative. Extensive experiments on three widely-used person re-ID datasets validate the effectiveness of our proposed approach. Codes will be available at:https://github.com/Dingyuan-Zheng/PPCL. Dingyuan Zheng, Jimin Xiao, Mingjie Sun, Huihui Bai 0001, Junhui Hou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Cycle-Free Weakly Referring Expression Grounding With Self-Paced LearningabstractIn this paper, we are tackling the weakly referring expression grounding task to localize the target object in an image according to a given query sentence, where the mapping between the query sentence and image regions is blind during the training period. Previous methods all follow a cyclic forward-backward pipeline to handle this task, where the query sentence is firstly converted to the result region through the forward module, and then the result region is converted back to a sentence through the backward module, with the difference between the reconstructed sentence and original query used as the loss to optimize the entire network. These existing methods, however, suffer from the deviation issue when the result region, generated through the forward module, totally deviates from the target area, but the backward module still reconstructs a similar sentence. The aforementioned loss function cannot penalize this kind of deviation because of the consistent prediction of the sentence. To overcome this limitation, we propose a cycle-free pipeline, where a region describer network is designed to predict the textual description for each candidate region, and a result region is selected according to the similarity between the predicted description and the query sentence. Furthermore, a self-paced learning mechanism is designed to avoid the drift issue during the warm-up period of the optimization process. The proposed method achieves a higher average accuracy on RefCOCO and RefCOCO+ datasets, compared with all previous state-of-the-art methods. Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Starting Point Selection and Multiple-Standard Matching for Video Object Segmentation With Language AnnotationabstractIn this study, we investigate language-level video object segmentation, where first-frame language annotation is used to describe the target object. Because a language label is typically compatible with all frames in a video, the proposed method can choose the most suitable starting frame to mitigate initialization failure. Apart from extracting the visual feature from a static video frame, a motion-language score based on optical flow is also proposed to describe moving objects more accurately. Scores of multiple standards are then aggregated using an attention-based mechanism to predict the final result. The proposed method is evaluated on four widely-used video object segmentation datasets, including the DAVIS 2017, DAVIS 2016, SegTrack V2 and YouTubeObject datasets, and a novel accuracy measured as mean region similarity is obtained on both the DAVIS 2017 (67.2%) and DAVIS 2016 (83.5%) datasets. The code will be published. Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yao Zhao 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Democracy Does Matter: Comprehensive Feature Mining for Co-Salient Object DetectionabstractCo-salient object detection, with the target of detecting co-existed salient objects among a group of images, is gaining popularity. Recent works use the attention mechanism or extra information to aggregate common co-salient features, leading to incomplete even incorrect responses for target objects. In this paper, we aim to mine comprehensive co-salient features with democracy and reduce background interference without introducing any extra information. To achieve this, we design a democratic prototype generation module to generate democratic response maps, covering sufficient co-salient regions and thereby involving more shared attributes of co-salient objects. Then a comprehensive prototype based on the response maps can be generated as a guide for final prediction. To suppress the noisy background information in the prototype, we propose a self-contrastive learning module, where both positive and negative pairs are formed without relying on additional classification information. Besides, we also design a democratic feature enhancement module to further strengthen the co-salient features by readjusting attention values. Extensive experiments show that our model obtains better performance than previous state-of-the-art methods, especially on challenging real-world cases (e.g., for CoCA, we obtain a gain of 2.0% for MAE, 5.4% for maximum F-measure, 2.3% for maximum E-measure, and 3.7% for S-measure) under the same settings. Source code is available at https://github.com/siyueyu/DCFM. Siyue Yu, Jimin Xiao, Bingfeng Zhang, Eng Gee Lim |
CVPR | 2 |
| 2022 | CARD: Semi-supervised Semantic Segmentation via Class-agnostic Relation based DenoisingabstractRecent semi-supervised semantic segmentation methods focus on mining extra supervision from unlabeled data by generating pseudo labels. However, noisy labels are inevitable in this process which prevent effective self-supervision. This paper proposes that noisy labels can be corrected based on semantic connections among features. Since a segmentation classifier produces both high and low-quality predictions, we can trace back to feature encoder to investigate how a feature in a noisy group is related to those in the confident groups. Discarding the weak predictions from the classifier, rectified predictions are assigned to the wrongly predicted features through the feature relations. The key to such an idea lies in mining reliable feature connections. With this goal, we propose a class-agnostic relation network to precisely capture semantic connections among features while ignoring their semantic categories. The feature relations enable us to perform effective noisy label corrections to boost self-training performance. Extensive experiments on PASCAL VOC and Cityscapes demonstrate the state-of-the-art performances of the proposed methods under various semi-supervised settings. Xiaoyang Wang 0007, Jimin Xiao, Bingfeng Zhang, Limin Yu |
IJCAI | 2 |
| 2022 | Affinity Attention Graph Neural Network for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation is receiving great attention due to its low human annotation cost. In this paper, we aim to tackle bounding box supervised semantic segmentation, i.e., training accurate semantic segmentation models using bounding box annotations as supervision. To this end, we propose affinity attention graph neural network ($A^2$A2GNN). Following previous practices, we first generate pseudo semantic-aware seeds, which are then formed into semantic graphs based on our newly proposed affinity Convolutional Neural Network (CNN). Then the built graphs are input to our$A^2$A2GNN, in which an affinity attention layer is designed to acquire the short- and long- distance information from soft graph edges to accurately propagate semantic labels from the confident seeds to the unlabeled pixels. However, to guarantee the precision of the seeds, we only adopt a limited number of confident pixel seed labels for$A^2$A2GNN, which may lead to insufficient supervision for training. To alleviate this issue, we further introduce a new loss function and a consistency-checking mechanism to leverage the bounding box constraint, so that more reliable guidance can be included for the model optimization. Experiments show that our approach achieves new state-of-the-art performances on Pascal VOC 2012 datasets (val: 76.5 percent,test: 75.2 percent). More importantly, our approach can be readily applied to bounding box supervised instance segmentation task or other weakly supervised semantic segmentation tasks, with state-of-the-art or comparable performance among almot all weakly supervised tasks on PASCAL VOC or COCO dataset. Our source code will be available athttps://github.com/zbf1991/A2GNN. Bingfeng Zhang, Jimin Xiao, Jianbo Jiao, Yunchao Wei, Yao Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | End-to-end weakly supervised semantic segmentation with reliable region mining
Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Kaizhu Huang, Shan Luo 0001, Yao Zhao 0001 |
Pattern Recognit. | 2 |
| 2022 | Soft pseudo-Label shrinkage for unsupervised domain adaptive person re-identification
Dingyuan Zheng, Jimin Xiao, Ke Chen 0004, Xiaowei Huang 0001, Yao Zhao 0001 |
Pattern Recognit. | 2 |
| 2022 | Unsupervised domain adaptation in homogeneous distance space for person re-identification
Dingyuan Zheng, Jimin Xiao, Yunchao Wei, Qiufeng Wang 0001, Kaizhu Huang, Yao Zhao 0001 |
Pattern Recognit. | 2 |
| 2022 | Neural texture transfer assisted video coding with adaptive up-sampling
Li Yu 0004, Wenshuai Chang, Weize Quan, Jimin Xiao, Dong-Ming Yan 0001, Moncef Gabbouj |
Signal Process. Image Commun. | 4 |
| 2022 | Transformer-Based Language-Person Search With Multiple Region SlicingabstractLanguage-person search is an essential technique for applications like criminal searching, where it is more feasible for a witness to provide language descriptions of a suspect than providing a photo. Most existing works treat the language-person pair as a black-box, neither considering the inner structure in a person picture, nor the correlations between image regions and referring words. In this work, we propose a transformer-based language-person search framework with matching conducted between words and image regions, where a person picture is vertically separated into multiple regions using two different ways, including the overlapped slicing and the key-point-based slicing. The co-attention between linguistic referring words and visual features are evaluated via transformer blocks. Besides the obtained outstanding searching performance, the proposed method enables to provide interpretability by visualizing the co-attention between image parts in the person picture and the corresponding referring words. Without bells and whistles, we achieve the state-of-the-art performance on the CUHK-PEDES dataset with Rank-1 score of 57.67% and the PA100K dataset with mAP of 22.88%, with simple yet elegant design. Code is available onhttps://github.com/detectiveli/T-MRS. Hui Li 0085, Jimin Xiao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | ToF and Stereo Data Fusion Using Dynamic Search Range Stereo MatchingabstractTime-of-Flight (ToF) sensors and stereo vision systems are both widely used for capturing depth data. They have some complementary strengths and limitations, which have been exploited in prior research to produce more accurate depth maps by fusing data from the two sources. However, among these diverse data fusion approaches, none of them provides an end-to-end neural network solution. In this work, we propose the first end-to-end ToF and stereo data fusion network using the coarse-to-fine matching framework, where the prior of ToF depth is integrated into the stereo matching process by constraining the search range of stereo matching within an interval around the ToF camera depth measurement. We adopt a dynamic search range for each pixel according to an estimated ToF error map, which is more efficient and effective than a constant one when handling various errors. The ToF error map is estimated by the ToF error estimator branching out from the stereo matching network. Both ToF error estimation and stereo matching are performed in a joint framework, with the two tasks assisting each other mutually. We also propose an upsampling module to replace the naive bilinear upsampling in the coarse-to-fine stereo matching network, which reduces the error caused by the upsampling. The proposed deep network is trained end-to-end on synthetic datasets and generalizable to real-world datasets without further fine-tuning. Experimental results show that our fusion method achieves higher accuracy than either ToF or stereo alone, and outperforms state-of-the-art fusion methods on both synthetic and real data. Yong Deng 0006, Jimin Xiao, Steven Zhiying Zhou |
IEEE Trans. Multim. | 2 |
| 2021 | Structure-Consistent Weakly Supervised Salient Object Detection with Local Saliency CoherenceabstractSparse labels have been attracting much attention in recent years. However, the performance gap between weakly supervised and fully supervised salient object detection methods is huge, and most previous weakly supervised works adopt complex training methods with many bells and whistles. In this work, we propose a one-round end-to-end training approach for weakly supervised salient object detection via scribble annotations without pre/post-processing operations or extra supervision data. Since scribble labels fail to offer detailed salient regions, we propose a local coherence loss to propagate the labels to unlabeled regions based on image features and pixel distance, so as to predict integral salient regions with complete object structures. We design a saliency structure consistency loss as self-consistent mechanism to ensure consistent saliency maps are predicted with different scales of the same image as input, which could be viewed as a regularization technique to enhance the model generalization ability. Additionally, we design an aggregation module (AGGM) to better integrate high-level features, low-level features and global context information for the decoder to aggregate various information. Extensive experiments show that our method achieves a new state-of-the-art performance on six benchmarks (e.g. for the ECSSD dataset: Fβ = 0.8995, Eξ = 0.9079 and MAE = 0.0489), with an average gain of 4.60% for F-measure, 2.05% for E-measure and 1.88% for MAE over the previous best performing method on this task. Source code is available at http://github.com/siyueyu/SCWSSOD. Siyue Yu, Bingfeng Zhang, Jimin Xiao, Eng Gee Lim |
AAAI | 3 |
| 2021 | Iterative Shrinking for Referring Expression Grounding Using Deep Reinforcement LearningabstractIn this paper, we are tackling the proposal-free referring expression grounding task, aiming at localizing the target object according to a query sentence, without relying on off-the-shelf object proposals. Existing proposal-free methods employ a query-image matching branch to select the highest-score point in the image feature map as the target box center, with its width and height predicted by another branch. Such methods, however, fail to utilize the contextual relation between the target and reference objects, and lack interpretability on its reasoning procedure. To solve these problems, we propose an iterative shrinking mechanism to localize the target, where the shrinking direction is decided by a reinforcement learning agent, with all contents within the current image patch comprehensively considered. Besides, the sequential shrinking processes enable to demonstrate the reasoning about how to iteratively find the target. Experiments show that the proposed method boosts the accuracy by 4.32% against the previous state-of-the- art (SOTA) method on the RefCOCOg dataset, where query sentences are long and complex with many targets referred by other reference objects. Mingjie Sun, Jimin Xiao, Eng Gee Lim |
CVPR | 2 |
| 2021 | Self-Guided and Cross-Guided Learning for Few-Shot SegmentationabstractFew-shot segmentation has been attracting a lot of attention due to its effectiveness to segment unseen object classes with a few annotated samples. Most existing approaches use masked Global Average Pooling (GAP) to encode an annotated support image to a feature vector to facilitate query image segmentation. However, this pipeline unavoidably loses some discriminative information due to the average operation. In this paper, we propose a simple but effective self-guided learning approach, where the lost critical information is mined. Specifically, through making an initial prediction for the annotated support image, the covered and uncovered foreground regions are encoded to the primary and auxiliary support vectors using masked GAP, respectively. By aggregating both primary and auxiliary support vectors, better segmentation performances are obtained on query images. Enlightened by our self-guided module for 1-shot segmentation, we propose a cross-guided module for multiple shot segmentation, where the final mask is fused using predictions from multiple annotated samples with high-quality support vectors contributing more and vice versa. This module improves the final prediction in the inference stage without re-training. Extensive experiments show that our approach achieves new state-of-the-art performances on both PASCAL-5iand COCO-20idatasets. Source code is available at https://github.com/zbf1991/SCL. Bingfeng Zhang, Jimin Xiao, Terry Qin |
CVPR | 2 |
| 2021 | Discriminative Triad Matching and Reconstruction for Weakly Referring Expression GroundingabstractIn this paper, we are tackling the weakly-supervised referring expression grounding task, for the localization of a referent object in an image according to a query sentence, where the mapping between image regions and queries are not available during the training stage. In traditional methods, an object region that best matches the referring expression is picked out, and then the query sentence is reconstructed from the selected region, where the reconstruction difference serves as the loss for back-propagation. The existing methods, however, conduct both the matching and the reconstruction approximately as they ignore the fact that the matching correctness is unknown. To overcome this limitation, a discriminative triad is designed here as the basis to the solution, through which a query can be converted into one or multiple discriminative triads in a very scalable way. Based on the discriminative triad, we further propose the triad-level matching and reconstruction modules which are lightweight yet effective for the weakly-supervised training, making it three times lighter and faster than the previous state-of-the-art methods. One important merit of our work is its superior performance despite the simple and neat design. Specifically, the proposed method achieves a new state-of-the-art accuracy when evaluated on RefCOCO (39.21 percent), RefCOCO+ (39.18 percent) and RefCOCOg (43.24 percent) datasets, that is 4.17, 4.08 and 7.8 percent higher than the previous one, respectively. The code is available at https://github.com/insomnia94/DTWREG. Mingjie Sun, Jimin Xiao, Eng Gee Lim, Si Liu 0001, John Yannis Goulermas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Progressive sample mining and representation learning for one-shot person re-identification
Hui Li 0085, Jimin Xiao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001 |
Pattern Recognit. | 2 |
| 2021 | Exploiting textual queries for dynamically visual disambiguationabstractDue to the high cost of manual annotation, learning directly from the web has attracted broad attention. One issue that limits the performance of current webly supervised models is the problem of visual polysemy. In this work, we present a novel framework that resolves visual polysemy by dynamically matching candidate text queries with retrieved images. Specifically, our proposed framework includes three major steps: we first discover and then dynamically select the text queries according to the keyword-based image search results, we employ the proposed saliency-guided deep multi-instance learning (MIL) network to remove outliers and learn classification models for visual disambiguation. Compared to existing methods, our proposed approach can figure out the right visual senses, adapt to dynamic changes in the search results, remove outliers, and jointly learn the classification models . Extensive experiments and ablation studies on CMU-Poly-30 and MIT-ISD datasets demonstrate the effectiveness of our proposed approach. Zeren Sun, Yazhou Yao, Jimin Xiao, Lei Zhang 0054, Jian Zhang 0002, Zhenmin Tang |
Pattern Recognit. | 3 |
| 2021 | Fast pixel-matching for video object segmentation
Siyue Yu, Jimin Xiao, Bingfeng Zhang, Eng Gee Lim, Yao Zhao 0001 |
Signal Process. Image Commun. | 2 |
| 2021 | Detail Preserving Coarse-to-Fine Matching for Stereo Matching and Optical FlowabstractThe Coarse-To-Fine (CTF) matching scheme has been widely applied to reduce computational complexity and matching ambiguity in stereo matching and optical flow tasks by converting image pairs into multi-scale representations and performing matching from coarse to fine levels. Despite its efficiency, it suffers from several weaknesses, such as tending to blur the edges and miss small structures like thin bars and holes. We find that the pixels of small structures and edges are often assigned with wrong disparity/flow in the upsampling process of the CTF framework, introducing errors to the fine levels and leading to such weaknesses. We observe that these wrong disparity/flow values can be avoided if we select the best-matched value among their neighborhood, which inspires us to propose a novel differentiable Neighbor-Search Upsampling (NSU) module. The NSU module first estimates the matching scores and then selects the best-matched disparity/flow for each pixel from its neighbors. It effectively preserves finer structure details by exploiting the information from the finer level while upsampling the disparity/flow. The proposed module can be a drop-in replacement of the naive upsampling in the CTF matching framework and allows the neural networks to be trained end-to-end. By integrating the proposed NSU module into a baseline CTF matching network, we design our Detail Preserving Coarse-To-Fine (DPCTF) matching network. Comprehensive experiments demonstrate that our DPCTF can boost performances for both stereo matching and optical flow tasks. Notably, our DPCTF achieves new state-of-the-art performances for both tasks - it outperforms the competitive baseline (Bi3D) by 28.8% (from 0.73 to 0.52) on EPE of the FlyingThings3D stereo dataset, and ranks first in KITTI flow 2012 benchmark. The code is available at https://github.com/Deng-Y/DPCTF. Yong Deng 0006, Jimin Xiao, Steven Zhiying Zhou, Jiashi Feng |
IEEE Trans. Image Process. | 2 |
| 2020 | Reliability Does Matter: An End-to-End Weakly Supervised Semantic Segmentation ApproachabstractWeakly supervised semantic segmentation is a challenging task as it only takes image-level information as supervision for training but produces pixel-level predictions for testing. To address such a challenging task, most recent state-of-the-art approaches propose to adopt two-step solutions, i.e. 1) learn to generate pseudo pixel-level masks, and 2) engage FCNs to train the semantic segmentation networks with the pseudo masks. However, the two-step solutions usually employ many bells and whistles in producing high-quality pseudo masks, making this kind of methods complicated and inelegant. In this work, we harness the image-level labels to produce reliable pixel-level annotations and design a fully end-to-end network to learn to predict segmentation maps. Concretely, we firstly leverage an image classification branch to generate class activation maps for the annotated categories, which are further pruned into confident yet tiny object/background regions. Such reliable regions are then directly served as ground-truth labels for the parallel segmentation branch, where a newly designed dense energy loss function is adopted for optimization. Despite its apparent simplicity, our one-step solution achieves competitive mIoU scores (val: 62.6, test: 62.9) on Pascal VOC compared with those two-step state-of-the-arts. By extending our one-step method to two-step, we get a new state-of-the-art performance on the Pascal VOC (val: 66.3, test: 66.5). Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, Kaizhu Huang |
AAAI | 2 |
| 2020 | Fast Template Matching and Update for Video Object Tracking and SegmentationabstractIn this paper, the main task we aim to tackle is the multi-instance semi-supervised video object segmentation across a sequence of frames where only the first-frame box-level ground-truth is provided. Detection-based algorithms are widely adopted to handle this task, and the challenges lie in the selection of the matching method to predict the result as well as to decide whether to update the target template using the newly predicted result. The existing methods, however, make these selections in a rough and inflexible way, compromising their performance. To overcome this limitation, we propose a novel approach which utilizes reinforcement learning to make these two decisions at the same time. Specifically, the reinforcement learning agent learns to decide whether to update the target template according to the quality of the predicted result. The choice of the matching method will be determined at the same time, based on the action history of the reinforcement learning agent. Experiments show that our method is almost 10 times faster than the previous state-of-the-art method with even higher accuracy (region similarity of 69.1% on DAVIS 2017 dataset). Mingjie Sun, Jimin Xiao, Eng Gee Lim, Bingfeng Zhang, Yao Zhao 0001 |
CVPR | 2 |
| 2020 | Feature Representation Matters: End-to-End Learning for Reference-Based Image Super-Resolution
Yanchun Xie, Jimin Xiao, Mingjie Sun, Kaizhu Huang |
ECCV (4) | 2 |
| 2020 | Pay Attention Selectively and Comprehensively: Pyramid Gating Network for Human Pose Estimation without Pre-trainingabstractDeep neural network with multi-scale feature fusion has achieved great success in human pose estimation. However, drawbacks still exist in these methods: 1) they consider multi-scale features equally, which may over-emphasize redundant features; 2) preferring deeper structures, they can learn features with the strong semantic representation, but tend to lose natural discriminative information; 3) to attain good performance, they rely heavily on pretraining, which is time-consuming, or even unavailable practically. To mitigate these problems, we propose a novel comprehensive recalibration model called Pyramid GAting Network (PGA-Net) that is capable of distillating, selecting, and fusing the discriminative and attention-aware features at different scales and different levels (i.e., both semantic and natural levels). Meanwhile, focusing on fusing features both selectively and comprehensively, PGA-Net can demonstrate remarkable stability and encouraging performance even without pre-training, making the model can be trained truly from scratch. We demonstrate the effectiveness of PGA-Net through validating on COCO and MPII benchmarks, attaining new state-of-the-art performance. https://github.com/ssr0512/PGA-Net Chenru Jiang, Kaizhu Huang, Shufei Zhang, Xinheng Wang 0001, Jimin Xiao |
ACM Multimedia | 5 |
| 2020 | Generative adversarial classifier for handwriting characters super-resolution
Zhuang Qian, Kaizhu Huang, Qiufeng Wang 0001, Jimin Xiao, Rui Zhang 0012 |
Pattern Recognit. | 4 |
| 2020 | Adaptive ROI generation for video object segmentation using reinforcement learning
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yanchun Xie, Jiashi Feng |
Pattern Recognit. | 2 |
| 2020 | Single image-based head pose estimation with spherical parametrization and 3D morphing
Hui Yuan 0001, Junhui Hou, Jimin Xiao |
Pattern Recognit. | 4 |
| 2020 | Segmentation mask guided end-to-end person search
Dingyuan Zheng, Jimin Xiao, Kaizhu Huang, Yao Zhao 0001 |
Signal Process. Image Commun. | 2 |
| 2020 | Correlation Filter Selection for Visual Tracking Using Reinforcement LearningabstractCorrelation filter has been proven to be an effective tool for a number of approaches in visual tracking, particularly for seeking a good balance between tracking accuracy and speed. However, correlation filter-based models are susceptible to wrong updates stemming from inaccurate tracking results. To date, very little effort has been devoted towards handling the correlation filter update problem. In this paper, we propose a novel approach to address the correlation filter update problem. In our approach, we update and maintain multiple correlation filter models in parallel, and we use deep reinforcement learning for the selection of an optimal correlation filter model among them. To facilitate the decision process in an efficient manner, we propose a decision-net to deal with target appearance modeling, which is trained through hundreds of challenging videos using proximal policy optimization and a lightweight learning network. An exhaustive evaluation of the proposed approach on the OTB100 and OTB2013 benchmarks shows that the approach is effective enough to achieve the average success rate of 62.3% and the average precision score of 81.2%, both exceeding the performance of traditional correlation filter-based trackers. Yanchun Xie, Jimin Xiao, Kaizhu Huang, Jeyan Thiyagalingam, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Edge Orientation Driven Depth Super-Resolution for View Synthesis
Jimin Xiao |
ICIG (3) | 2 |
| 2019 | IAN: The Individual Aggregation Network for Person Search
Jimin Xiao, Yanchun Xie, Tammam Tillo, Kaizhu Huang, Yunchao Wei, Jiashi Feng |
Pattern Recognit. | 1 |
| 2019 | Multiview video quality enhancement without depth information
Samer Jammal, Tammam Tillo, Jimin Xiao |
Signal Process. Image Commun. | 3 |
| 2018 | Image Ordinal Classification and Understanding: Grid Dropout with Masking LabelabstractImage ordinal classification refers to predicting a discrete target value which carries ordering correlation among image categories. The limited size of labeled ordinal data renders modern deep learning approaches easy to overfit. To tackle this issue, neuron dropout and data augmentation were proposed which, however, still suffer from over-parameterization and breaking spatial structure, respectively. To address the issues, we first propose a grid dropout method that randomly dropout/blackout some areas of the training image. Then we combine the objective of predicting the blackout patches with classification to take advantage of the spatial information. Finally we demonstrate the effectiveness of both approaches by visualizing the Class Activation Map (CAM) and discover that grid dropout is more aware of the whole facial areas and more robust than neuron dropout for small training dataset. Experiments are conducted on a challenging age estimation dataset-Adience dataset with very competitive results compared with state-of-the-art methods. Chao Zhang 0072, Ce Zhu, Jimin Xiao, Xun Xu 0002, Yipeng Liu 0001 |
ICME | 3 |
| 2018 | Siamese network ensemble for visual tracking
Chenru Jiang, Jimin Xiao, Yanchun Xie, Tammam Tillo, Kaizhu Huang |
Neurocomputing | 2 |
| 2018 | Visual aesthetic understanding: Sample-specific aesthetic classification and deep activation map visualization
Chao Zhang 0072, Ce Zhu, Xun Xu 0002, Yipeng Liu 0001, Jimin Xiao, Tammam Tillo |
Signal Process. Image Commun. | 5 |
| 2018 | Region-Based Multiple Description Coding for Multiview Video Plus Depth VideoabstractInterframe and interview predictions are widely employed in multiview video coding. This technique improves the coding efficiency, but it also increases the vulnerability of the coded bitstream. Thus, one packet loss will affect many subsequent frames in the same view and probably in other referenced views. To address this problem, a region-based multiple description coding scheme is proposed for robust 3-D video communication in this paper, in which two descriptions are formed by setting the left and right view as dominant in the first and second description, respectively. This approach exploits the fact that most regions in the reference view could be synthesized from the base view. Hence, these regions could be skipped or only coarsely encoded. In our work, the disoccluded regions, illumination-affected regions, and remaining regions are first determined and extracted. By assigning different quantization parameters for these three different regions according to the network status, an efficient multiple description scheme is formed. Experimental results demonstrate that the proposed scheme achieves considerably better performance compared with the traditional approach. Chunyu Lin, Yao Zhao 0001, Jimin Xiao, Tammam Tillo |
IEEE Trans. Multim. | 3 |
| 2018 | Convolutional Neural Network for Intermediate View Enhancement in Multiview StreamingabstractMultiview video streaming continues to gain popularity due to the great viewing experience it offers, as well as its availability that has been enabled by increased network throughput and other recent technical developments. User demand for interactive multiview video streaming that provides seamless view switching upon request is also increasing. However, it is a highly challenging task to stream stable and high quality videos that allow real-time scene navigation within the bandwidth constraint. In this paper, a convolutional neural network (ConvNet)-assisted seamless multiview video streaming system is proposed to tackle the challenge. The proposed method solves the problem from two perspectives. First, a ConvNet-assisted multiview representation method is proposed, which provides flexible interactivity without compromising on multiview video compression efficiency. Second, a bit allocation mechanism guided by a navigation model is developed to provide seamless navigation and adapt to network bandwidth fluctuations at the same time. These two blocks work closely to provide an optimized viewing experience to users. They can be integrated into any existing multiview video streaming framework to enhance overall performance. Experimental results demonstrate the effectiveness of the proposed method for seamless multiview streaming. Li Yu 0004, Tammam Tillo, Jimin Xiao, Marco Grangetto |
IEEE Trans. Multim. | 3 |
| 2018 | Cooperative Bargaining Game-Based Multiuser Bandwidth Allocation for Dynamic Adaptive Streaming Over HTTPabstractDynamic adaptive streaming over HTTP (DASH) has emerged as an efficient technology for video streaming. For a DASH system, a most common case is that a limited server bandwidth is competed by multiusers. In order to improve user quality of experience (QoE) and guarantee fairness, we propose to use the game theory in a proxy server to allocate the bandwidth collaboratively for multiusers. By taking user buffer length, received video bit rates, video qualities, etc., into account, the bandwidth allocation problem is formulated as a cooperative bargaining problem and the Nash bargaining solution (NBS) is obtained by convex optimization. The requested bit rate of users will be rewritten as the proxy calculated bit rate (i.e., NBS) when the user requested bit rate is larger. Experimental results demonstrate that user QoE and fairness can be improved significantly, i.e., the delay frequency and duration are smaller, and the received video qualities are higher and more stable, when comparing the proposed method with existing methods. Hui Yuan 0001, Xuekai Wei, Fuzheng Yang 0001, Jimin Xiao, Sam Kwong |
IEEE Trans. Multim. | 4 |
| 2017 | Disparity Estimation Using Convolutional Neural Networks with Multi-scale Correlation
Samer Jammal, Tammam Tillo, Jimin Xiao |
ICONIP (3) | 3 |
| 2017 | An effective CU size decision method for quality scalability in SHVC
Xiaoni Li, Mianshu Chen, Zhaowei Qu, Jimin Xiao, Moncef Gabbouj |
Multim. Tools Appl. | 4 |
| 2017 | Texture Plus Depth Video Coding Using Camera Global Motion InformationabstractIn video coding, traditional motion estimation methods work well for videos with camera translational motion, but their efficiency drops for other motions, such as rotational and dolly motions. In this paper, a motion-information-based three-dimensional (3D) video coding method is proposed for texture plus depth 3D video. The synchronized global motion information of the camera is obtained to assist the encoder improve its rate-distortion performance by projecting the temporal neighboring texture and depth frames into the position of the current frame, using the depth and camera motion information. Then, the projected frames are added into the reference buffer list as virtual reference frames. As these virtual reference frames could be more similar to the current to-be-encoded frame than the conventional reference frames, the required bits to represent the residual will be reduced. The experimental results demonstrate that the proposed scheme enhances the coding performance for all camera motion types and for various scene settings and resolutions using H.264 and HEVC standards, respectively. With the computer graphic sequences, for H.264, the average gain of texture and depth coding are up to 2 dB and 1 dB, respectively. For HEVC and HD resolution sequences, the gain of texture coding reaches 0.4 dB. For realistic sequences, up to 0.5 dB gain (H.264) is achieved for the texture video, while up to 0.7 dB gain is achieved for the depth sequences. Fei Cheng 0001, Tammam Tillo, Jimin Xiao, Byeungwoo Jeon |
IEEE Trans. Multim. | 3 |
| 2016 | 3D video super-resolution using fully convolutional neural networksabstractLarge amount of redundant information and huge data size have been a serious problem for multiview video systems. To address this problem, one popular solution is mixed-resolution, where only few viewpoints are kept with full resolution and other views are kept with lower resolution. In this paper, we propose a super-resolution (SR) method, where the low-resolution viewpoints in the 3D video are up-sampled using a fully convolutional neural network. By simply projecting the neighboring high resolution image to the position of the low resolution image, we learn the relationship of high and low resolution patches, and reconstruct the low resolution images into high resolution ones using the projected image information. We propose to use a fully convolutional neural network to establish a mapping between those images. The network is barely trained on 17 pairs of multiview images, and tested on other multiview images and video sequences. It is observed that our proposed method outperforms existing methods objectively and subjectively, with more than 1 dB average gain achieved. Meanwhile, our network training procedure is efficient, with less than 3 hours using one Titan X GPU. Yanchun Xie, Jimin Xiao, Tammam Tillo, Yunchao Wei, Yao Zhao 0001 |
ICME | 2 |
| 2016 | Packetization strategies for MVD-based 3D video transmissionabstractIn multi-view video plus depth (MVD) format, virtual views are synthesized by the compressed texture videos and their associated depth through depth-image-based rendering. In this paper, we consider the setup where both the encoded texture and depth bitstreams experience packet losses during transmission. Different packetization strategies are investigated and a novel strategy is developed to improve error resilience of MVD-based video transmission, where texture data and its corresponding depth are put into the same packet. The size of texture plus associated depth data included in each packet needs to be less than the Maximum Transfer Unit (MTU). Experimental results demonstrate that our proposed packetization scheme yields a significant improvement in terms of both texture views and synthesized virtual views quality when fit in H.264/AVC. Xue Zhang 0008, Yao Zhao 0001, Tammam Tillo, Chunyu Lin, Jimin Xiao, Anhong Wang |
VCIP | 5 |
| 2016 | Virtual-View-Assisted Video Super-Resolution and EnhancementabstractA 3-D multiview video gives users an experience that is different from that provided by a traditional video; however, it puts a huge burden on limited bandwidth resources. Mixed-resolution video in a multiview system can alleviate this problem by using different video resolutions for different views. However, to reduce visual uncomfortableness and to make this video format more suitable for free-viewpoint television, the low-resolution (LR) views need to be super-resolved to the target full resolution. In this paper, we propose a virtual-view-assisted super-resolution algorithm, where the inter-view similarity is used to determine whether the missing pixels in the super-resolved frame need to be filled by virtual-view pixels or by spatial interpolated pixels. The decision mechanism is steered by the texture characteristics of the neighbors of each missing pixel. Furthermore, the inter-view similarity is used, on the one hand, to enhance the quality of the virtual-view-copied pixels by compensating the luminance difference between different views and, on the other hand, to enhance the original LR pixels in the super-resolved frame by reducing their compression distortion. Thus, the proposed method can recover the details in regions with edges while maintaining good quality at smooth areas by properly exploiting the high-quality virtual-view pixels and the directional correlation of pixels. The experimental results demonstrate the effectiveness of the proposed approach with a peak signal-to-noise ratio gain of up to 3.85 dB. Zhi Jin 0002, Tammam Tillo, Jimin Xiao, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Depth Map Down-Sampling and Coding Based on Synthesized View DistortionabstractIn this paper, we propose a depth map down-sampling and coding scheme that minimizes the view synthesis distortion. Moreover, a solution for the optimal depth map down-sampling problem that minimizes the depth-caused distortion in the virtual view by exploiting the depth map and the associated texture information along with the up-sampling method to be used in the decoder side is derived. Furthermore, to enhance compression performance, the synthesized view distortion, which is evaluated by emulating the interpolation and the virtual view synthesis process, is used in the optimization objective function for coding mode selection in the video encoder. Experimental results show that both the proposed depth map down-sampling and encoding methods lead to good performance, and the average bit rate reduction is 2.62% compared with 3D-AVC. Jimin Xiao, Tammam Tillo, Yao Zhao 0001, Chunyu Lin, Huihui Bai 0001 |
IEEE Trans. Multim. | 2 |
| 2015 | Statistical approach for motion estimation skipping (SAMEK)abstractHigh Efficient Video Coding (HEVC) standard has achieved significant rate-distortion improvement over the previous standard H.264/AVC. However, the complexity that comes from its flexible data structure representation is an obstacle for its wide application. To reduce the overall complexity and encoding time, this paper proposes a statistical approach for motion estimation skipping (SAMEK). The SAMEK method avoids some unnecessary motion estimations in units with less probability of being referenced. These units can be recognized by the two rules of SAMEK method, namely ZeroCase and DecreaseCase. These two rules are summarized by statistically analyzing the relationships between each PU and its references. The experimental results demonstrate that our method can save up to 9.5% encoding time (averagely 6.87%) with negligible rate-distortion losses (averagely 0.006 dB) when compared with HEVC encoder with Test Zone search (TZ search) enabled. Li Yu 0004, Jimin Xiao, Tammam Tillo, Ce Zhu |
ICIP | 2 |
| 2015 | 3D video coding using motion information and depth mapabstractIn this paper, a motion-information-based 3D video coding method is proposed for the texture plus depth 3D video format. The synchronized global motion information of camcorder is sampled to assist the encoder to improve its rate-distortion performance. This approach works by projecting temporal previous frames into the position of the current frame using the depth and motion information. These projected frames are added in the reference buffer as virtual reference frames. As these virtual reference frames are more similar to the current frame than the conventional reference frames, the required residual information is reduced. The experimental results demonstrate that the proposed scheme enhances the coding performance in various motion conditions including rotational and translational motions. Fei Cheng 0001, Jimin Xiao, Tammam Tillo |
ICME | 2 |
| 2015 | Multiple Description Coding for Stereoscopic Videos With Stagger Frame OrderabstractDue to the prediction structures employed in video coding, the loss of one packet will affect many following frames. In this paper, a multiple description coding scheme with stagger frame order is proposed for stereoscopic 3-D videos. First, the reference and auxiliary views in stereoscopic sequences will be asymmetrically encoded into one description, whereas the other description will be formed in the same way with one dumb frame delay. Because of the stagger frame order, the coarsely encoded B frames will be inserted into different positions of the two descriptions. If a certain frame encoded with I/P mode is lost, then its corresponding B-frame version will be employed to compensate for the loss. In each description, the quantization steps of B frames are tuned based on a closed-form solution that considers the video contents, network status, frame positions in the group of picture, and the layer of the views. For further improvement, a fusing scheme is provided. The experimental results demonstrate that the proposed scheme outperforms state-of-the-art schemes. Specifically, up to 1.3-dB gain is achieved in the case of packet loss, and 2-dB gain is obtained for the side/central performance. Chunyu Lin, Yao Zhao 0001, Tammam Tillo, Jimin Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | Scalable Bit Allocation Between Texture and Depth Views for 3-D Video Streaming Over Heterogeneous NetworksabstractIn the multiview video plus depth (MVD) coding format, both texture and depth views are jointly compressed to represent the 3-D video content. The MVD format enables synthesis of virtual views through depth-image-based rendering; hence, distortion in the texture and depth views affects the quality of the synthesized virtual views. Bit allocation between texture and depth views has been studied with some promising results. However, to the best of our knowledge, most of the existing bit-allocation methods attempt to allocate a fixed amount of total bit rate between texture and depth views; that is, to select appropriate pair of quantization parameters for texture and depth views to maximize the synthesized view quality subject to a fixed total bit rate. In this paper we propose a scalable bit-allocation scheme, where a single ordering of texture and depth packets is derived and used to obtain optimal bit allocation between texture and depth views for any total target rates. In the proposed scheme, both texture and depth views are encoded using the quality scalable coding method; that is, medium grain scalable (MGS) coding of the Scalable Video Coding (SVC) extension of the Advanced Video Coding (H.264/AVC) standard. For varying target total bit rates, optimal bit truncation points for both texture and depth views can be obtained using the proposed scheme. Moreover, we propose to order the enhancement layer packets of the H.264/SVC MGS encoded depth view according to their contribution to the reduction of the synthesized view distortion. On one hand, this improves the depth view packet ordering when considered the rate-distortion performance of synthesized views, which is demonstrated by the experimental results. On the other hand, the information obtained in this step is used to facilitate optimal bit allocation between texture and depth views. Experimental results demonstrate the effectiveness of the proposed scalable bit-allocation scheme for texture and depth views. Jimin Xiao, Miska M. Hannuksela, Tammam Tillo, Moncef Gabbouj, Ce Zhu, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2014 | Dynamic redundancy allocation for video streaming using Sub-GOP based FEC codeabstractReed-Solomon erasure code is one of the most studied protection methods for video streaming over unreliable networks. As a block-based error correcting code, large block size and increased number of parity packets will enhance its protection performance. However, for video applications this enhancement is sacrificed by the error propagation and the increased bitrate. So, to tackle this paradox, we propose a rate-distortion optimized redundancy allocation scheme, which takes into consideration the distortion caused by losing each slice and the propagated error. Different from other approaches, the amount of introduced redundancy and the way it is introduced are automatically selected without human interventions based on the network condition and video characteristics. The redundancy allocation problem is formulated as a constraint optimization problem, which allows to have more flexibility in setting the block-wise redundancy. The proposed scheme is implemented in JM14.0 for H.264, and it achieves an average gain of 1dB over the state-of-the-art approach. Li Yu 0004, Jimin Xiao, Tammam Tillo |
VCIP | 2 |
| 2014 | Correlation based universal image/video coding loss recovery
Jinjian Wu, Weisi Lin, Guangming Shi, Jimin Xiao |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | Optimizing the deadzone width to improve the polyphase-based multiple description coding
Chunyu Lin, Tammam Tillo, Jimin Xiao, Yao Zhao 0001 |
Multim. Tools Appl. | 3 |
| 2013 | Multiple description video coding based on forward error correction within expanding windowsabstractIn this paper, an MDC scheme based on forward error correction(FEC) within expanding windows is proposed. Firstly, the video sequence will be coded into source packets with/without slice group enabled. Secondly, the appropriate FEC packets are inserted according to the packet loss rate. Since the previous frames in a GOP is generally more important than the following frames in the GOP, an expanding window is exploited so that the FEC packets for the current frame will also protect the previous frames in the window. After this, the source packets with the inserted FEC packets will be divided into two descriptions and transmitted into two independent channels. When some packets in one description are lost, FEC decoding will try to recover the lost packets. Through this scheme, the source packets can get appropriate protection while the compression efficiency will not be degraded too much. The experimental results show that the proposed scheme outperforms the compared schemes up to 3 dB. Chunyu Lin, Yao Zhao 0001, Jimin Xiao, Tammam Tillo |
ICIP | 3 |
| 2013 | Real-Time Video Streaming Using Randomized Expanding Reed-Solomon CodeabstractForward error correction (FEC) codes are widely studied to protect streamed video over unreliable networks. Typically, enlarging the FEC coding block size can improve the error correction performance. For video streaming applications, this could be implemented by grouping more than one video frame into one FEC coding block. However, in this case, it leads to decoding delay, which is not tolerable for real-time video streaming applications. In this paper, to solve this dilemma, a real-time video streaming scheme using randomized expanding Reed-Solomon (RS) code is proposed. In this scheme, the RS coding block includes not only the video packets of the current frame, but could also include all the video packets of previous frames in the current group of pictures. At the decoding side, the parity-check equations of the current frame are jointly solved with all the parity-check equations of the previous frames. Since video packets of the following frames are not encompassed in the RS coding block, no delay will be caused for waiting for the video or parity packets of the following frames both at encoding and decoding sides. Experimental results show that the proposed scheme outperforms other real-time error resilient video streaming approaches significantly, specifically, for the Foreman sequence, the proposed scheme could provide 1.5 dB average gain over the state-of-the-art approach for 10% i.i.d. packet loss rate, whereas for the burst loss case, the average gain is more than 3 dB.MATLAB code of this paper is available for download at http://www.mmtlab.com. Jimin Xiao, Tammam Tillo, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Real-time video streaming exploiting the late-arrival packetsabstractFor real-time video applications, such as video telephony service, the allowed maximum end-to-end transmission delay is usually fixed. The packets arriving at the destination out of the maximum end-to-end delay are treated as late-arrival packets, and these packets are discarded in traditional video transmission systems. In this paper, in order to improve the system performance, we propose to exploit these packets to update the decoder reference buffer. Two schemes are proposed to exploit the late-arrival packets, one scheme is to use sliding-window updating, where the updating window is moving; another scheme is to use fixed-window update together with systematic Reed-Solomon code. The effectiveness of the two schemes are validated by simulation results without adding extra delay. It is found that in both schemes, the updating window size plays an important role on the system performance. Jimin Xiao, Tammam Tillo, Chunyu Lin, Yao Zhao 0001 |
PCS | 1 |
| 2012 | Dynamic Sub-GOP Forward Error Correction Code for Real-Time Video ApplicationsabstractReed-Solomon erasure codes are commonly studied as a method to protect the video streams when transmitted over unreliable networks. As a block-based error correcting code, on one hand, enlarging the block size can enhance the performance of the Reed-Solomon codes; on the other hand, large block size leads to long delay which is not tolerable for real-time video applications. In this paper a novel Dynamic Sub-GOP FEC (DSGF) approach is proposed to improve the performance of Reed-Solomon codes for video applications. With the proposed approach, the Sub-GOP, which contains more than one video frame, is dynamically tuned and used as the RS coding block, yet no delay is introduced. For a fixed number of extra introduced packets, for protection, the length of the Sub-GOP and the redundancy devoted to each Sub-GOP becomes a constrained optimization problem. To solve this problem, a fast greedy algorithm is proposed. Experimental results show that the proposed ap proach outperforms other real-time error resilient video coding technologies. Jimin Xiao, Tammam Tillo, Chunyu Lin, Yao Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2011 | Real-time forward error correction for video transmissionabstractWhen the video streams are transmitted over the unreliable networks, forward error correction (FEC) codes are usually used to protect them. Reed-Solomon codes are block-based FEC codes. On one hand, enlarging the block size can enhance the performance of the Reed- Solomon codes. On the other hand, large Reed-Solomon block size leads to long delay which is not tolerable for real-time video applications. In this paper a novel approach is proposed to improve the performance of Reed-Solomon codes. With the proposed approach, more than one video frame are encompassed in the Reed-Solomon coding block yet no delay is introduced. Experimental results show that the proposed approach outperforms other real-time error resilient video coding technologies. Jimin Xiao, Tammam Tillo, Chunyu Lin, Yao Zhao 0001 |
VCIP | 1 |