EDBT 2026 Demo / reviewers in the wild / expert
Xiaojin Gong
dblp:73/4329
· DBLP profile ↗
42ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0001-9955-3569ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 6 first-author · 19 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Modality-Aware Bias Mitigation and Invariance Learning for Unsupervised Visible-Infrared Person Re-IdentificationabstractUnsupervised visible-infrared person re-identification (USVI-ReID) aims to match individuals across visible and infrared cameras without relying on any annotation. Given the significant gap across visible and infrared modality, estimating reliable cross-modality association becomes a major challenge in USVI-ReID. Existing methods usually adopt optimal transport to associate the intra-modality clusters, which is prone to propagating the local cluster errors, and also overlooks global instance-level relations. By mining and attending to the visible-infrared modality bias, this paper focuses on addressing cross-modality learning from two aspects: bias-mitigated global association and modality-invariant representation learning. Motivated by the camera-aware distance rectification in single-modality re-ID, we propose modality-aware Jaccard distance to mitigate the distance bias caused by modality discrepancy, so that more reliable cross-modality associations can be estimated through global clustering. To further improve cross-modality representation learning, a `split-and-contrast' strategy is designed to obtain modality-specific global prototypes. By explicitly aligning these prototypes under global association guidance, modality-invariant yet ID-discriminative representation learning can be achieved. While conceptually simple, our method obtains state-of-the-art performance on benchmark VI-ReID datasets and outperforms existing methods by a significant margin, validating its effectiveness. Menglin Wang 0001, Xiaojin Gong, Genlin Ji |
AAAI | 2 |
| 2026 | Enhancing Weakly Supervised Multimodal Video Anomaly Detection Through Text GuidanceabstractIn recent years, weakly supervised multimodal video anomaly detection, which leverages RGB, optical flow, and audio modalities, has garnered significant attention from researchers, emerging as a vital subfield within video anomaly detection. However, previous studies have inadequately explored the role of text modality in this domain. With the proliferation of large-scale text-annotated video datasets and the advent of video captioning models, obtaining text descriptions from videos has become increasingly feasible. Text modality, carrying explicit semantic information, can more accurately characterize events within videos and identify anomalies, thereby enhancing the model's detection capabilities and reducing false alarms. However, text feature extraction challenges anomaly detection. Pre-trained large language models often struggle to effectively capture the nuances associated with anomalies, as their training is based on generalized datasets. Directly fine-tuning the text feature extractor is also challenging, as anomaly-related text descriptions are sparse. Furthermore, due to the varying amounts of information carried by different modalities, issues such as modality redundancy and modality imbalance arise during feature fusion. To address the challenges of text feature extraction and the issues of modality redundancy and imbalance, we propose a novel text-guided weakly supervised multimodal video anomaly detection framework. Specifically, we introduce an in-context learning based multi-stage text augmentation mechanism to generate high-quality anomaly text samples. These high-quality samples are then used to fine-tune the text feature extractor, aiming to obtain a more effective text feature extractor for anomaly detection. Additionally, we present a multi-scale bottleneck Transformer fusion module to enhance multimodal integration, utilizing a set of reduced bottleneck tokens to progressively transmit compressed information across modalities, aiming to address the issues of modality redundancy and imbalance. Experimental results on large-scale datasets UCF-Crime and XD-Violence demonstrate that our proposed approach achieves state-of-the-art performance. This project is publicly available at https://shengyangsun.github.io/TGMVAD. Shengyang Sun, Jiashen Hua, Junyi Feng, Xiaojin Gong |
IEEE Trans. Multim. | 4 |
| 2026 | Re-purposing SAM into Efficient Visual Projectors for MLLM-based Referring Image SegmentationabstractRecently, Referring Image Segmentation (RIS) frameworks that pair the Multimodal Large Language Model (MLLM) with the Segment Anything Model (SAM) have achieved impressive results. However, adapting MLLM to segmentation is computationally intensive, primarily due to visual token redundancy. We observe that traditional patch-wise visual projectors struggle to strike a balance between reducing the number of visual tokens and preserving semantic clarity, often retaining overly long token sequences to avoid performance drops. Inspired by text tokenizers, we propose a novel semantic visual projector that leverages semantic superpixels generated by SAM to identify “visual words” in an image. By compressing and projecting semantic superpixels as visual tokens, our approach adaptively shortens the token sequence according to scene complexity while minimizing semantic loss in compression. To mitigate loss of information, we propose a semantic superpixel positional embedding to strengthen MLLM’s awareness of superpixel geometry and position, alongside a semantic superpixel aggregator to preserve both fine-grained details inside superpixels and global context outside. Experiments show that our method cuts visual tokens by \(\sim\) 93% without compromising performance, notably speeding up MLLM training and inference, and outperforming existing compressive visual projectors on RIS. Xiaojin Gong |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Prior-Constrained Association Learning for Fine-Grained Generalized Category DiscoveryabstractThis paper addresses generalized category discovery (GCD), the task of clustering unlabeled data from potentially known or unknown categories with the help of labeled instances from each known category. Compared to traditional semi-supervised learning, GCD is more challenging because unlabeled data could be from novel categories not appearing in labeled data. Current state-of-the-art methods typically learn a parametric classifier assisted by self-distillation. While being effective, these methods do not make use of cross-instance similarity to discover class-specific semantics which are essential for representation learning and category discovery. In this paper, we revisit the association-based paradigm and propose a Prior-constrained Association Learning method to capture and learn the semantic relations within data. In particular, the labeled data from known categories provides a unique prior for the association of unlabeled data. Unlike previous methods that only adopts the prior as a pre or post-clustering refinement, we fully incorporate the prior into the association process, and let it constrain the association towards a reliable grouping outcome. The estimated semantic groups are utilized through non-parametric prototypical contrast to enhance the representation learning. A further combination of both parametric and non-parametric classification complements each other and leads to a model that outperforms existing methods by a significant margin. On multiple GCD benchmarks, we perform extensive experiments and validate the effectiveness of our proposed method. Menglin Wang 0001, Zhun Zhong, Xiaojin Gong |
AAAI | 3 |
| 2025 | Leveraging Vision Foundation Models for RGB-Thermal Semantic Segmentation
Xiaojin Gong |
PRCV (5) | 2 |
| 2025 | Delving Into Instance Modeling for Weakly Supervised Video Anomaly DetectionabstractWeakly-supervised video anomaly detection (WS-VAD) aims to identify fine-grained anomalies from sparse video-level labels, which has gained increasing attention in recent years due to its various applications such as disaster warning and public security. Recent studies typically formulate WS-VAD as a multi-instance learning (MIL) problem. However, they neglect the instance creation process and simply apply a uniform temporal pooling (UTP) operation to obtain the training instances, leading to severe anomaly contamination and dilution. In this paper, we emphasize the importance of the instance modeling procedure and propose two simple yet effective modules, i.e., the dynamic segment merging (DSM) module and the retrieval-augmented anomaly restoration (RA2R) module, to tackle the problem from segment-level and feature-level, respectively. We equip various state-of-the-art WS-VAD models with the proposed methods and conduct thorough experiments on the challenging datasets, e.g., UCF-Crime, and XD-Violence. Results demonstrate the proposed method brings consistent performance improvement and establishes new state-of-the-art. Shengyang Sun, Jiashen Hua, Junyi Feng, Dongxu Wei, Baisheng Lai, Xiaojin Gong |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence DetectionabstractWeakly supervised multimodal violence detection aims to learn a violence detection model by leveraging multiple modalities such as RGB, optical flow, and audio, while only video-level annotations are available. In the pursuit of effective multimodal violence detection (MVD), information redundancy, modality imbalance, and modality asynchrony are identified as three key challenges. In this work, we propose a new weakly supervised MVD method that explicitly addresses these challenges. Specifically, we introduce a multi-scale bottleneck transformer (MSBT) based fusion module that employs a reduced number of bottleneck tokens to gradually condense information and fuse each pair of modalities and utilizes a bottleneck token-based weighting scheme to highlight more important fused features. Furthermore, we propose a temporal consistency contrast loss to semantically align pairwise fused features. Experiments on the largest-scale XD-Violence dataset demonstrate that the proposed method achieves state-of-the-art performance. Code is available at https://github.com/shengyangsun/MSBT. Shengyang Sun, Xiaojin Gong |
ICME | 2 |
| 2024 | Task-Conditional Adapter for Multi-Task Dense PredictionabstractMulti-task dense prediction plays an important role in the field of computer vision and has an abundant array of applications. Its main purpose is to reduce the amount of network training parameters by sharing network parameters while using the correlation between tasks to improve overall performance. We propose a task-conditional network that handles one task at a time and shares most network parameters to achieve these goals. Inspired by adapter tuning, we propose an adapter module that focuses on both spatial- and channel-wise information to extract features from the frozen encoder backbone. This approach not only reduces the number of training parameters, but also saves training time and memory resources by attaching a parallel adapter pathway to the encoder. We additionally use learnable task prompts to model different tasks and use these prompts to adjust some parameters of adapters to fit the network to diverse tasks. These task-conditional adapters are also applied to the decoder, which enables the entire network to switch between various tasks, producing better task-specific features and achieving excellent performance. Extensive experiments on two challenging multi-task benchmarks, NYUD-v2 and PASCAL-Context, show that our approach achieves state-of-the-art performance with excellent parameter, time, and memory efficiency. The code is available at https://github.com/jfzleo/Task-Conditional-Adapter. Fengze Jiang, Shuling Wang 0002, Xiaojin Gong |
ACM Multimedia | 3 |
| 2024 | TDSD: Text-Driven Scene-Decoupled Weakly Supervised Video Anomaly Detection
Shengyang Sun, Jiashen Hua, Junyi Feng, Dongxu Wei, Baisheng Lai, Xiaojin Gong |
ACM Multimedia | 6 |
| 2024 | Foundation Model Assisted Weakly Supervised Semantic SegmentationabstractThis work aims to leverage pre-trained foundation models, such as contrastive language-image pre-training (CLIP) and segment anything model (SAM), to address weakly supervised semantic segmentation (WSSS) using image-level labels. To this end, we propose a coarse-to-fine framework based on CLIP and SAM for generating high-quality segmentation seeds. Specifically, we construct an image classification task and a seed segmentation task, which are jointly performed by CLIP with frozen weights and two sets of learnable task-specific prompts. A SAMbased seeding (SAMS) module is designed and applied to each task to produce either coarse or fine seed maps. Moreover, we design a multi-label contrastive loss supervised by image-level labels and a CAM activation loss supervised by the generated coarse seed map. These losses are used to learn the prompts, which are the only parts need to be learned in our framework. Once the prompts are learned, we input each image along with the learned segmentation specific prompts into CLIP and the SAMS module to produce high-quality segmentation seeds. These seeds serve as pseudo labels to train an off-the-shelf segmentation network like other two-stage WSSS methods. Experiments show that our method achieves the state-of-the-art performance on PASCAL VOC 2012 and competitive results on MS COCO 2014. Our code will be released upon acceptance. Xiaojin Gong |
WACV | 2 |
| 2024 | Event-driven weakly supervised video anomaly detection
Shengyang Sun, Xiaojin Gong |
Image Vis. Comput. | 2 |
| 2023 | Hierarchical Semantic Contrast for Scene-aware Video Anomaly DetectionabstractIncreasing scene-awareness is a key challenge in video anomaly detection (VAD). In this work, we propose a hierarchical semantic contrast (HSC) method to learn a scene-aware VAD model from normal videos. We first incorporate foreground object and background scene features with high-level semantics by taking advantage of pre-trained video parsing models. Then, building upon the autoencoder-based reconstruction framework, we introduce both scene-level and object-level contrastive learning to enforce the encoded latent features to be compact within the same semantic classes while being separable across different classes. This hierarchical semantic contrast strategy helps to deal with the diversity of normal patterns and also increases their discrimination ability. Moreover, for the sake of tackling rare normal activities, we design a skeleton-based motion augmentation to increase samples and refine the model further. Extensive experiments on three public datasets and scene-dependent mixture datasets validate the effectiveness of our proposed method. Shengyang Sun, Xiaojin Gong |
CVPR | 2 |
| 2023 | Long-Short Temporal Co-Teaching for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection (WS-VAD) is a challenging problem that aims to learn VAD models only with video-level annotations. In this work, we propose a Long-Short Temporal Co-teaching (LSTC) method to address the WS-VAD problem. It constructs two tubelet-based spatio-temporal transformer networks to learn from short- and long-term video clips respectively. Each network is trained with respect to a multiple instance learning (MIL)-based ranking loss, together with a cross-entropy loss when clip-level pseudo labels are available. A co-teaching strategy is adopted to train the two networks. That is, clip-level pseudo labels generated from each network are used to supervise the other one at the next training round, and the two networks are learned alternatively and iteratively. Our proposed method is able to better deal with the anomalies with varying durations as well as subtle anomalies. Extensive experiments on three public datasets demonstrate that our method outperforms state-of-the-art WS-VAD methods. Code is available at https://github.com/shengyangsun/LSTC_VAD. Shengyang Sun, Xiaojin Gong |
ICME | 2 |
| 2023 | Learning Intra and Inter-Camera Invariance for Isolated Camera Supervised Person Re-identificationabstractSupervised person re-identification assumes that a person has images captured under multiple cameras. However when cameras are placed in distance, a person rarely appears in more than one camera. This paper thus studies person re-ID under such isolated camera supervised (ISCS) setting. Instead of trying to generate fake cross-camera features like previous methods, we explore a novel perspective by making efficient use of the variation in training data. Under ISCS setting, a person only has limited images from a single camera, so the camera bias becomes a critical issue confounding ID discrimination. Cross-camera images are prone to being recognized as different IDs simply by camera style. To eliminate the confounding effect of camera bias, we propose to learn both intra- and inter-camera invariance under a unified framework. First, we construct style-consistent environments via clustering, and perform prototypical contrastive learning within each environment. Meanwhile, strongly augmented images are contrasted with original prototypes to enforce intra-camera augmentation invariance. For inter-camera invariance, we further design a much improved variant of multi-camera negative loss that optimizes the distance of multi-level negatives. The resulting model learns to be invariant to both subtle and severe style variation within and cross-camera. On multiple benchmarks, we conduct extensive experiments and validate the effectiveness and superiority of the proposed method. Code will be available athttps://github.com/Terminator8758/IICI. Menglin Wang 0001, Xiaojin Gong |
ACM Multimedia | 2 |
| 2022 | Online Convolutional ReparameterizationabstractStructural re-parameterization has drawn increasing attention in various computer vision tasks. It aims at improving the performance of deep models without introducing any inference-time cost. Though efficient during inference, such models rely heavily on the complicated training-time blocks to achieve high accuracy, leading to large extra training cost. In this paper, we present online convolutional re-parameterization (OREPA), a two-stage pipeline, aiming to reduce the huge training overhead by squeezing the complex training-time block into a single convolution. To achieve this goal, we introduce a linear scaling layer for better optimizing the online blocks. Assisted with the reduced training cost, we also explore some more effective re-param components. Compared with the state-of-the-art re-param models, OREPA is able to save the training-time memory cost by about 70% and accelerate the training speed by around 2×. Meanwhile, equipped with OREPA, the models out-perform previous methods on ImageNet by up to +0.6%. We also conduct experiments on object detection and semantic segmentation and show consistent improvements on the downstream tasks. Codes are available at https://github.com/JUGGHM/OREPA_CVPR2022. Mu Hu, Junyi Feng, Jiashen Hua, Baisheng Lai, Jianqiang Huang 0001, Xiaojin Gong, Xian-Sheng Hua 0001 |
CVPR | 6 |
| 2022 | Self-Paced Knowledge Distillation for Real-Time Image Guided Depth CompletionabstractImage guided depth completion aims to generate a dense depth map from a sparse one with the guidance of a color image. Previous high-accuracy methods often rely on complex networks that are large in size and expensive in computational cost, making them inapplicable to real-time platforms. In this letter, we propose a self-paced knowledge distillation method, which obtains a lightweight but accurate depth completion model via distilling knowledge from a complex teacher network. Specifically, by taking advantage of the easy-to-hard learning curriculum in deep networks, we first design a groundtruth-free hard-pixel mining module to tell hard and noisy pixels in the teacher’s output. Then, we design two self-paced distillation losses, which gradually introduce hard pixels to distill the depth and structure knowledge from the teacher to the compact student network. Experiments on the KITTI benchmark show that the proposed method can improve the original student model by a considerable margin. The distilled compact and real-time student model outperforms all previous lightweight networks, mitigating the performance gap with state-of-the-art high-accuracy but complex models. Shuling Wang 0002, Mu Hu, Bin Li 0038, Xiaojin Gong |
IEEE Signal Process. Lett. | 4 |
| 2022 | Offline-Online Associated Camera-Aware Proxies for Unsupervised Person Re-IdentificationabstractRecently, unsupervised person re-identification (Re-ID) has received increasing research attention due to its potential for label-free applications. A promising way to address unsupervised Re-ID is clustering-based, which generates pseudo labels by clustering and uses the pseudo labels to train a Re-ID model iteratively. However, most clustering-based methods take each cluster as a pseudo identity class, neglecting the intra-cluster variance mainly caused by the change of cameras. To address this issue, we propose to split each single cluster into multiple proxies according to camera views. The camera-aware proxies explicitly capture local structures within clusters, by which the intra-ID variance and inter-ID similarity can be better tackled. Assisted with the camera-aware proxies, we design two proxy-level contrastive learning losses that are, respectively, based on offline and online association results. The offline association directly associates proxies according to the clustering and splitting results, while the online strategy dynamically associates proxies in terms of up-to-date features to reduce the noise caused by the delayed update of pseudo labels. The combination of two losses enables us to train a desirable Re-ID model. Extensive experiments on three person Re-ID datasets and one vehicle Re-ID dataset show that our proposed approach demonstrates competitive performance with state-of-the-art methods. Code will be available at: https://github.com/Terminator8758/O2CAP. Menglin Wang 0001, Baisheng Lai, Xiaojin Gong, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Camera-Aware Proxies for Unsupervised Person Re-IdentificationabstractThis paper tackles the purely unsupervised person re-identification (Re-ID) problem that requires no annotations. Some previous methods adopt clustering techniques to generate pseudo labels and use the produced labels to train Re-ID models progressively. These methods are relatively simple but effective. However, most clustering-based methods take each cluster as a pseudo identity class, neglecting the large intra-ID variance caused mainly by the change of camera views. To address this issue, we propose to split each single cluster into multiple proxies and each proxy represents the instances coming from the same camera. These camera-aware proxies enable us to deal with large intra-ID variance and generate more reliable pseudo labels for learning. Based on the camera-aware proxies, we design both intra and inter-camera contrastive learning components for our Re-ID model to effectively learn the ID discrimination ability within and across cameras. Meanwhile, a proxy-balanced sampling strategy is also designed, which facilitates our learning further. Extensive experiments on three large-scale Re-ID datasets show that our proposed approach outperforms most unsupervised methods by a significant margin. Especially, on the challenging MSMT17 dataset, we gain 14.3 percent Rank-1 and 10.2 percent mAP improvements when compared to the second place. Code is available at: https://github.com/Terminator8758/CAP-master. Menglin Wang 0001, Baisheng Lai, Jianqiang Huang 0001, Xiaojin Gong, Xian-Sheng Hua 0001 |
AAAI | 4 |
| 2021 | PENet: Towards Precise and Efficient Image Guided Depth CompletionabstractImage guided depth completion is the task of generating a dense depth map from a sparse depth map and a high quality image. In this task, how to fuse the color and depth modalities plays an important role in achieving good performance. This paper proposes a two-branch backbone that consists of a color-dominant branch and a depth-dominant branch to exploit and fuse two modalities thoroughly. More specifically, one branch inputs a color image and a sparse depth map to predict a dense depth map. The other branch takes as inputs the sparse depth map and the previously predicted depth map, and outputs a dense depth map as well. The depth maps predicted from two branches are complimentary to each other and therefore they are adaptively fused. In addition, we also propose a simple geometric convolutional layer to encode 3D geometric cues. The geometric encoded backbone conducts the fusion of different modalities at multiple stages, leading to good depth completion results. We further implement a dilated and accelerated CSPN++ to refine the fused depth map efficiently. The proposed full model ranks 1st in the KITTI depth completion online leaderboard at the time of submission. It also infers much faster than most of the top ranked methods. The code of this work is available at https://github.com/JUGGHM/PENet_ICRA2021. Mu Hu, Shuling Wang 0002, Bin Li 0038, Shiyu Ning, Xiaojin Gong |
ICRA | 6 |
| 2021 | Towards Precise Intra-camera Supervised Person Re-IdentificationabstractIntra-camera supervision (ICS) for person reidentification (Re-ID) assumes that identity labels are independently annotated within each camera view and no inter-camera identity association is labeled. It is a new setting proposed recently to reduce the burden of annotation while expect to maintain desirable Re-ID performance. However, the lack of inter-camera labels makes the ICS Re-ID problem much more challenging than the fully supervised counterpart. By investigating the characteristics of ICS, this paper proposes jointly learned camera-specific non-parametric classifiers, together with a hybrid mining quintuplet loss, to perform intra-camera learning. Then, an inter-camera learning module consisting of a graph-based ID association step and a Re-ID model updating step is conducted. Extensive experiments on three large-scale Re-ID datasets show that our approach outperforms all existing ICS works by a great margin. Our approach performs even comparable to state-of-the-art fully supervised methods in two of the datasets. Menglin Wang 0001, Baisheng Lai, Jianqiang Huang 0001, Xiaojin Gong, Xian-Sheng Hua 0001 |
WACV | 5 |
| 2021 | Self-supervised Visual-LiDAR Odometry with Flip ConsistencyabstractMost learning-based methods estimate ego-motion by utilizing visual sensors, which suffer from dramatic lighting variations and textureless scenarios. In this paper, we incorporate sparse but accurate depth measurements obtained from lidars to overcome the limitation of visual methods. To this end, we design a self-supervised visual-lidar odometry (Self-VLO) framework. It takes both monocular images and sparse depth maps projected from 3D lidar points as input, and produces pose and depth estimations in an end-to-end learning manner, without using any ground truth labels. To effectively fuse two modalities, we design a two-pathway encoder to extract features from visual and depth images and fuse the encoded features with those in decoders at multiple scales by our fusion module. We also adopt a siamese architecture and design an adaptively weighted flip consistency loss to facilitate the self-supervised learning of our VLO. Experiments on the KITTI odometry benchmark show that the proposed approach out-performs all self-supervised visual or lidar odometries. It also performs better than fully supervised VOs, demonstrating the power of fusion. Bin Li 0038, Mu Hu, Shuling Wang 0002, Lianghao Wang, Xiaojin Gong |
WACV | 5 |
| 2018 | Exploiting LSTM for Joint Object and Semantic Part Detection
Xiaojin Gong |
ACCV (5) | 2 |
| 2018 | A Normalized Convolutional Neural Network for Guided Sparse Depth UpsamplingabstractGuided sparse depth upsampling aims to upsample an irregularly sampled sparse depth map when an aligned high-resolution color image is given as guidance. When deep convolutional neural networks (CNNs) become the optimal choice to many applications nowadays, how to deal with irregular and sparse data still remains a non-trivial problem. Inspired by the classical normalized convolution operation, this work proposes a normalized convolutional layer (NCL) implemented in CNNs. Sparse data are therefore explicitly considered in CNNs by the separation of both data and filters into a signal part and a certainty part. Based upon NCLs, we design a normalized convolutional neural network (NCNN) to perform guided sparse depth upsampling. Experiments on both indoor and outdoor datasets show that the proposed NCNN models achieve state-of-the-art upsampling performance. Moreover, the models using NCLs gain a great generalization ability to different sparsity levels. Jiashen Hua, Xiaojin Gong |
IJCAI | 2 |
| 2017 | Saliency Guided End-to-End Learning for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD), which is the problem of learning detectors using only image-level labels, has been attracting more and more interest. However, this problem is quite challenging due to the lack of location supervision. To address this issue, this paper integrates saliency into a deep architecture, in which the location information is explored both explicitly and implicitly. Specifically, we select highly confident object proposals under the guidance of class-specific saliency maps. The location information, together with semantic and saliency information, of the select proposals are then used to explicitly supervise the network by imposing two additional losses. Meanwhile, a saliency prediction sub-network is built in the architecture. The prediction results are used to implicitly guide the localization procedure. The entire network is trained end-to-end. Experiments on PASCAL VOC demonstrate that our approach outperforms all state-of-the-arts. Baisheng Lai, Xiaojin Gong |
IJCAI | 2 |
| 2017 | Adaptive Metric Learning and Probe-Specific Reranking for Person ReidentificationabstractIn this letter, we introduce an adaptive metric learning (AML) method for person reidentification. Different from conventional metric learning approaches, which treat all the negative samples equally, AML adaptively classifies the negative samples into three groups and pays different attention to them. By emphasizing the influence of hard negative samples, AML can better mine the discriminative information between positive and negative samples, and thus generate a more effective metric. Furthermore, we also propose a probe-specific reranking (PSR) algorithm to refine the initial ranking list measured by the learned metric. For each probe, PSR constructs a corresponding hypergraph to capture the neighborhood relationship between the probe and its top 100 ranked gallery images. Then, these images are reranked based on their neighborhood affinity in the hypergraph. Extensive experiments on three challenging datasets demonstrate the superiority of both AML and PSR. Xiaojin Gong, Martin D. Levine |
IEEE Signal Process. Lett. | 3 |
| 2016 | An Alternating Proximal Splitting Method with Global Convergence for Nonconvex Structured Sparsity OptimizationabstractIn many learning tasks with structural properties, structured sparse modeling usually leads to better interpretability and higher generalization performance. While great efforts have focused on the convex regularization, recent studies show that nonconvex regularizers can outperform their convex counterparts in many situations. However, the resulting nonconvex optimization problems are still challenging, especially for the structured sparsity-inducing regularizers. In this paper, we propose a splitting method for solving nonconvex structured sparsity optimization problems. The proposed method alternates between a gradient step and an easily solvable proximal step, and thus enjoys low per-iteration computational complexity. We prove that the whole sequence generated by the proposed method converges to a critical point with at least sublinear convergence rate, relying on the Kurdyka-Łojasiewicz inequality. Experiments on both simulated and real-world data sets demonstrate the efficiency and efficacy of the proposed method. Shubao Zhang, Hui Qian 0001, Xiaojin Gong |
AAAI | 3 |
| 2016 | Saliency Guided Dictionary Learning for Weakly-Supervised Image ParsingabstractIn this paper, we propose a novel method to perform weakly-supervised image parsing based on the dictionary learning framework. To deal with the challenges caused by the label ambiguities, we design a saliency guided weight assignment scheme to boost the discriminative dictionary learning. More specifically, with a collection of tagged images, the proposed method first conducts saliency detection and automatically infers the confidence for each semantic class to be foreground or background. These clues are then incorporated to learn the dictionaries, the weights, as well as the sparse representation coefficients in the meanwhile. Once obtained the coefficients of a superpixel, we use a sparse representation classifier to determine its semantic label. The approach is validated on the MSRC21, PASCAL VOC07, and VOC12 datasets. Experimental results demonstrate the encouraging performance of our approach in comparison with some state-of-the-arts. Baisheng Lai, Xiaojin Gong |
CVPR | 2 |
| 2016 | Nonnegative Matrix Cofactorization for Weakly Supervised Image ParsingabstractImage parsing, which is the task of assigning each pixel with a semantic label, is a challenging problem when only supervised under image-level tags. In this letter, we propose a Nonnegative Matrix Cofactorization method to perform image parsing with noisy tags, i.e., some tags may be incorrect or missing. Given a collection of noisily tagged images, we first oversegment them into superpixels. Then, the superpixels' label matrix, which is aimed to estimate, and the feature matrix are simultaneously decomposed into nonnegative factor matrices with a graph Laplacian constraint and an orthogonal constraint. This cofactorization is able to jointly learn a discriminative dictionary and a linear classifier. The proposed approach therefore is robust to noise. Experimental results on two real-world image datasets MSRC-21 and LabelMe demonstrate the encouraging performance of our approach in comparison with the state of the arts. Xiaojin Gong |
IEEE Signal Process. Lett. | 1 |
| 2015 | Joint Object Segmentation and Depth UpsamplingabstractWith the advent of powerful ranging and visual sensors, nowadays, it is convenient to collect sparse 3-D point clouds and aligned high-resolution images. Benefitted from such convenience, this letter proposes a joint method to perform both depth assisted object-level image segmentation and image guided depth upsampling. To this end, we formulate these two tasks together as a bi-task labeling problem, defined in a Markov random field. An alternating direction method (ADM) is adopted for the joint inference, solving each sub-problem alternatively. More specifically, the sub-problem of image segmentation is solved by Graph Cuts, which attains discrete object labels efficiently. Depth upsampling is addressed via solving a linear system that recovers continuous depth values. By this joint scheme, robust object segmentation results and high-quality dense depth maps are achieved. The proposed method is applied to the challenging KITTI vision benchmark suite, as well as the Leuven dataset for validation. Comparative experiments show that our method outperforms stand-alone approaches. Wenqi Huang 0002, Xiaojin Gong, Michael Ying Yang |
IEEE Signal Process. Lett. | 2 |
| 2015 | Learning Visual-Spatial Saliency for Multiple-Shot Person Re-IdentificationabstractRecognizing persons across non-overlapping camera views, known as person re-identification, has received increasing attentions for its importance in many surveillance applications. However, most of existing methods rely on pre-training steps to ensure their performance and ignore the body prior knowledge of pedestrians. In this letter, we propose a novel non-training method for person re-identification which learns visual-spatial saliency from voter images and the given query image. First we segment pedestrian images into small regions and use two hypergraphs to represent the visual and spatial relationship among regions. Then we formulate the visual-spatial saliency learning as a joint hypergraph ranking problem by simultaneously considering the human body prior and the appearance similarity among pedestrians. Finally, the visual-spatial saliency is incorporated in region-based matching to improve the performance of person re-identification. Experimental evaluation on three publicly available datasets demonstrates the effectiveness of our approach. Xiaojin Gong, Zhenjiang Dong |
IEEE Signal Process. Lett. | 3 |
| 2014 | Road scene segmentation via fusing camera and lidar dataabstractThis paper presents an approach for pixel-wise object segmentation for road scenes based on the integration of a color image and an aligned 3D point cloud. In light of the advantage of range information in object discovery, we first produce initial object hypotheses by clustering the sparse 3D point cloud. The image pixels registered to the clustered 3D points are taken as samples to learn each object's prior knowledge. The priors are represented by Gaussian Mixture Models (GMMs) of color and 3D location information only, requiring no high-level features. We further formulate the segmentation problem within a Conditional Random Field (CRF) framework, which incorporates the learned prior models, together with hard constraints placed on the registered pixels and pairwise spatial constraints to achieve final results. Our algorithm is validated on the challenging KITTI dataset which contains diverse complicated road scenarios. Both qualitative and quantitative evaluation results show the superiority of our algorithm. Wenqi Huang 0002, Xiaojin Gong, Zhiyu Xiang |
ICRA | 2 |
| 2013 | Integrating visual and range data for road detectionabstractThis paper presents a new method for detecting drivable road surfaces in a single image. The method takes advantage of range and visual information so that reliable results are achieved. Specifically, given LIDAR data and an aligned image, it first makes use of 3D points to estimate the ground plane and determine the horizon. Then, subsets of road and obstacle points are extracted from the 3D points based on the plane and LIDAR properties. The pixels registered to the extracted points are used to build apriori road and non-road appearance models. The road detection problem is further formulated using Markov random field whose energy function is defined based on the learned models. Constraints are also added on the energy function to place high confidence on the pixels that are registered to extracted 3D points. Extensive experiments on urban roads and highways show that our method is robust even in complicated environments. Wenqi Huang 0002, Xiaojin Gong, Jilin Liu |
ICIP | 2 |
| 2013 | Guided depth enhancement via a fast marching method
Xiaojin Gong, Wenhui Zhou 0001, Jilin Liu |
Image Vis. Comput. | 1 |
| 2012 | Rock detection via superpixel graph cutsabstractThis paper presents a rock detection method for planetary terrain scenes. Our approach first segments an image into a set of superpixels. Then we formulate the rock detection task as an energy minimization problem and solve it efficiently via a novel graph cut which is constructed on the superpixels. In order to deal with complex rock scenarios, we integrate a discriminative observation model into the graph cut framework to enhance the discrimination power. Meanwhile, a couple of features, for instance, gradient based texture and contextual shading features, are employed to characterize superpixels. With the representative features, as well as the powerful optimization model, the rock detection problem is addressed well. We test our algorithm on a real Lunar terrain image set drawn from NASA which contains diverse scenarios. The attained qualitative and quantitative results show that our algorithm is effective. Xiaojin Gong, Jilin Liu |
ICIP | 1 |
| 2012 | Guided inpainting and filtering for Kinect depth maps
Xiaojin Gong, Jilin Liu |
ICPR | 2 |
| 2012 | A new depth descriptor for pedestrian detection in RGB-D images
Ningbo Wang, Xiaojin Gong, Jilin Liu |
ICPR | 2 |
| 2011 | A unified rectification method for single viewpoint multi-camera systemabstractStereo matching and 3D reconstruction has been studied for decades as a fundamental problem in the field of computer vision. Recent years, stereo matching and 3D reconstruction with a large field of view, especially using omnidirectional vision and panoramic images, have received increasing attention. As a pre-step for dense stereo matching, methods are proposed to rectify different kinds of omnidirectional stereo image pairs. However, no one has described a rectification method applied to multi-camera omnidirectional systems yet. In this work, we proposed a rectification algorithm based on spherical camera model for rectifying omnidirectional stereo pairs, especially well suitable for the multi-camera omnidirectional systems as long as a spherical camera model is able to be applied. We describe the geometrical framework of the algorithm and implement it. Also we present the experimental results of the real stereo image pairs captured by Ladybug3. As the experimental results show, the effect of rectification is promising. Yangchang Wang, Bin Yang 0035, Jingting Ding, Xiaojin Gong, Jilin Liu |
AVSS | 4 |
| 2011 | A DAISY-like compass operatorabstractThe compass operator is a useful tool for the detection of edges, junctions and corners. However, it is time-consuming, especially when the circular window is large. In this paper, we propose a new compass operator based on the sampling scheme utilized in the DAISY descriptor. We first generalize the original DAISY for representing a variety of types of features. Then, the metric of measuring the difference between two wedges is derived with respect to the descriptor. Our DAISY-like compass operator in essence captures cues in multi-scales but is efficient. It achieves results on boundary detection which are comparable to the conventional compass operator. Xiaojin Gong, Jilin Liu |
ICIP | 1 |
| 2008 | Real-time Robust Mapping for an Autonomous Surface Vehicle using an Omnidirectional CameraabstractTowards the goal of achieving truly autonomous navigation for a surface vehicle in maritime environments, a critical task is to detect surrounding obstacles such as the shore, docks, and other boats. In this paper, we demonstrate a real-time vision-based mapping system which detects and localizes stationary obstacles using a single omnidirectional camera and navigational sensors (GPS and gyro). The main challenge of this work is to make mapping robust to a large number of outliers, which stem from waves and specular reflections on the surface of the water. To address this problem, a two-step robust outlier rejection method is proposed. Experimental results obtained in unstructured large-scale environments are presented and validated using topographic maps. Xiaojin Gong, Bin Xu 0007, Caleb Reed, Christopher L. Wyatt, Daniel J. Stilwell |
WACV | 1 |
| 2007 | Performance Analysis and Validation of a Paracatadioptric Omnistereo SystemabstractIn this paper we present a vector-based 3D localization formula for a paracatadioptric omnistereo system. Based on vector representation, the performance of this stereo system is analyzed numerically, including the maximum detectable range and the uncertainty of 3D localization, with respect to the flexible stereo configuration of the system, positions of scene points, as well as errors in correspondence matching and errors in stereo configuration. The results of performance analysis are used to guide the trajectory of an autonomous surface vehicle (ASV), which is equipped with a paracatadioptric omnidirectional camera, in a map building application. Xiaojin Gong, Anbumani Subramanian, Christopher L. Wyatt, Daniel J. Stilwell |
ICCV | 1 |
| 2007 | A Two-stage Algorithm for Shoreline DetectionabstractShoreline detection plays an important role in vision based navigation for autonomous surface vehicles (ASVs). It is a challenging task because of the diversity in near-bank scenarios. In this paper, we present a two-stage algorithm to find the shoreline by employing multiple features. First, we classify images into two types: reflection-unidentifiable and reflection-identifiable. Based on this classification, images are further analyzed with suitable techniques respectively. In the reflection-unidentifiable case, the surface reflection is subtle and so the water region can be separated from land by an adaptive thresholding method. The points along the edge of the water region are then identified and the shoreline is estimated through a line-fitting technique. In the reflection-identifiable case, we aim to discriminate water regions from the land by means of a two-category region classifier. Images are oversegmented into small regions based on color homogeneity. Then the characteristic features of land-water scenes like symmetry and brightness are extracted and applied to classify a region into land or water categories. Experimental results show the efficacy of our approach and robustness in diverse situations Xiaojin Gong, Anbumani Subramanian, Christopher L. Wyatt |
WACV | 1 |
| 2002 | Semantic error checking in automatic proofreading for Chinese textsabstractSemantic error checking is a weak point of automatic proofreading for Chinese texts, with few mature methods at present. The paper discusses the technology of semantic error detection for Chinese texts, proposes a strategy combining statistical and rules approaches, and uses the collocation relationship based on instances, statistics and rules to check errors. The strategy checks both local semantic restraints and remote semantic collocations, and achieves satisfactory results. Weihua Luo, Zhensheng Luo, Xiaojin Gong |
SMC (2) | 3 |