VLDB 2026 Research / reviewers in the wild / expert
Zhuqing Jiang
dblp:72/10176
· DBLP profile ↗
46ranked-venue papers
4as first author
19since 2021 · last 2025
0000-0001-6308-5708ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 9 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Camera-Invariant Meta-Learning Network for Single-Camera-Training Person ReidentificationabstractSingle-camera-training person reidentification (SCT re-ID) aims to train a reidentification (re-ID) model using single-camera-training (SCT) datasets where each person appears in only one camera. The main challenge of SCT re-ID is to learn camera-invariant feature representations without cross-camera same-person (CCSP) data as supervision. Previous methods address it by assuming that the most similar person should be found in another camera. However, this assumption is not guaranteed to be correct. In this article, we propose a novel solution: the camera-invariant meta-learning network (CIMN) for SCT re-ID. CIMN operates under the premise that camera-invariant feature representations should remain robust despite changes in camera settings. To achieve this, we partition the training data into a meta-train set and a meta-test set based on camera IDs. We then conduct a cross-camera simulation (CCS) using a meta-learning strategy, aiming to enforce the feature representations learned from the meta-train set to be robust when applied to the meta-test set. We further introduce three specific loss functions to leverage potential identity relations between the meta-train set and the meta-test set. Through the CCS and the introduced loss functions, CIMN can extract feature representations that are both camera-invariant and identity-discriminative even in the absence of CCSP data. Our experimental results demonstrate that CIMN can extract feature representations that are both camera-invariant and identity-discriminative, even in the absence of CCSP data. our method achieves comparable performance with and without the use of CCSP data, and outperforms state-of-the-art methods on three SCT re-ID benchmarks. Jiangbo Pei, Zhuqing Jiang, Aidong Men, Haiying Wang 0005, Haiyong Luo, Shiping Wen 0001 |
IEEE Internet Things J. | 2 |
| 2025 | Understand and Detect: Multi-step zero-shot detection with image-level specific prompt
Miaotian Guo, Kewei Wu, Zhuqing Jiang, Haiying Wang 0005, Aidong Men |
Knowl. Based Syst. | 3 |
| 2024 | Medical Language Mixture of Experts for Improving Medical Image SegmentationabstractTraditional medical image segmentation methods are mostly uni-modal approaches solely based on the image modality. Recently, the emergence of text-guided image segmentation methods, by utilizing text annotations to compensate for the quality deficiency in image data, has shown promise for improving medical image segmentation. Despite their success, these methods often experience inadequate utilization of beneficial text information, and have applicability issues in the missing text modality scenario. To address these limitations, in this paper, we propose a Medical Language Mixture of Experts (MLMoE), which introduces multiple sub-experts for extracting more diverse information from medical text. These different experts are then combined by a gating module, thus aggregating beneficial text information to assist the image segmentation. Furthermore, to guarantee its performance in the text-absent scenario, a virtual prompt based distillation module is proposed, which distills the valuable knowledge of MLMoE learned from available text information to the virtual prompt, as an alternative text input. Experimental results on two multi-modal medical segmentation datasets demonstrate the effectiveness of our opposed method, achieving state-of-the-art performance. Code will be available at: https://github.com/Rango-bit/MLMoE.git. Jiangbo Pei, Zhu He, Guangjing Yang, Zhuqing Jiang, Qicheng Lao |
BIBM | 5 |
| 2024 | ONeK-SLAM: A Robust Object-level Dense SLAM Based on Joint Neural Radiance Fields and KeypointsabstractNeural implicit representation has recently achieved significant advancements, especially in the field of SLAM(Simultaneous Localization and Mapping). Previous NeRF-based SLAM methods have difficulties with object-level localization and reconstruction and struggle in dynamic and illumination-varied environments. We propose ONeK-SLAM, a robust object-level SLAM system that effectively combines feature points and neural radiance fields. ONeK-SLAM uses the joint information at the object level to improve localization accuracy and enhance reconstruction details. Moreover, our approach detects and eliminates dynamic objects based on the joint errors, while also harnessing the illumination invariance offered by feature points. Consequently, ONeK-SLAM achieves high-precision localization and detailed object-level mapping, even in dynamic and illumination-varying environments. Our evaluations, conducted on three public datasets that include both dynamic and variable lighting sequences, demonstrate that our method outperforms recent NeRF-based SLAM method in both localization and reconstruction. Yue Zhuge, Haiyong Luo, Yushi Chen 0004, Jiaquan Yan, Zhuqing Jiang |
ICRA | 6 |
| 2024 | Curriculum Prompting Foundation Models for Medical Image Segmentation
Xiuqi Zheng, Hongrui Liang, Xueqi Bao, Zhuqing Jiang, Qicheng Lao |
MICCAI (12) | 6 |
| 2024 | Pedestrian Navigation Activity Recognition Based on Segmentation TransformerabstractIn the context of the Internet of Things, utilizing the inherent inertial sensors in smartphones for human activity recognition (HAR) has garnered considerable attention owing to its wide-ranging applications. However, prevailing HAR approaches primarily treat activity identification as a single-label classification task, focusing solely on discerning pedestrian motion modes or device usage modes, while disregarding their interrelatedness. Additionally, HAR methods employing sliding windows encounter challenges associated with the multiclass window problem, wherein certain sample labels differ from the label assigned to the window. This paper aims to address these issues. This paper presents a novel approach for simultaneously recognizing pedestrian motion and device usage modes by utilizing the segmentation transformer. The proposed joint recognition framework effectively annotates sensor data at each timestamp and achieves dense prediction of time-series data through the encoding and decoding of the annotated data. To optimize the utilization of information extracted from each Transformer layer, a global up-sampling decoder based on the pyramid attention module is introduced, enabling dense decoding of features obtained from each Transformer layer. We performed experiments on two publicly available datasets to comprehensively assess the effectiveness of the proposed methodology. The results demonstrate that our approach achieves an accuracy of 99.79% and a weighted F-score of 99.77%, surpassing the performance of existing state-of-the-art methods. Furthermore, we constructed heterogeneous datasets to validate the robustness of our method. The extensive experimental findings indicate that the joint recognition framework effectively uncovers the inherent correlations between pedestrian motion and device usage modes, leading to enhanced accuracy in recognition and addressing the challenges posed by the multiclass window problem. Qu Wang, Jiahui Ning, Zhuqing Jiang, Liangliang Guo, Haiyong Luo, Haiying Wang 0005, Aidong Men, Xiaofei Cheng |
IEEE Internet Things J. | 4 |
| 2024 | Multiscale Transformer and Attention Mechanism for Magnetic Spatiotemporal Sequence LocalizationabstractLocation-based service (LBS) is the core of internet of things (IoTs), which serves tracking, navigation and monitoring. The ubiquitous magnetic signals are temporally stable and spatially distinguishable, and can achieve high-precision and ubiquitous positioning results without additional infrastructure, which is favored by researchers and has become a major research hotspot. Although there has been extensive research in the field of indoor magnetic positioning, there is still room for optimization in terms of positioning accuracy and robustness. Aiming at the problem that the magnetometer is offset and susceptible to environmental interference, we propose an online magnetometer calibration algorithm without user perception. Aiming at the inconsistency of magnetic data spatial scale problem caused by differences in device sampling frequency and user walking speed, we leverage different scales to segment the magnetic data, extract the magnetic sequence features of the corresponding scales through Transformer, utilize the attention mechanism to score the weights of the different scale features, and finally fuse the multiple scale features for positioning. We conduct extensive and well-designed experiments on public datasets and self-collected datasets. The experimental results indicate that the proposed method effectively solves the magnetic spatial scale problem and improves indoor magnetic positioning accuracy. Qu Wang, Meixia Fu, Jianquan Wang 0001, Lei Sun 0012, Rong Huang 0005, Xianda Li, Zhuqing Jiang, Haiyong Luo |
IEEE Internet Things J. | 8 |
| 2023 | Uncertainty-Induced Transferability Representation for Source-Free Unsupervised Domain AdaptationabstractSource-free unsupervised domain adaptation (SFUDA) aims to learn a target domain model using unlabeled target data and the knowledge of a well-trained source domain model. Most previous SFUDA works focus on inferring semantics of target data based on the source knowledge. Without measuring the transferability of the source knowledge, these methods insufficiently exploit the source knowledge, and fail to identify the reliability of the inferred target semantics. However, existing transferability measurements require either source data or target labels, which are infeasible in SFUDA. To this end, firstly, we propose a novel Uncertainty-induced Transferability Representation (UTR), which leverages uncertainty as the tool to analyse the channel-wise transferability of the source encoder in the absence of the source data and target labels. The domain-level UTR unravels how transferable the encoder channels are to the target domain and the instance-level UTR characterizes the reliability of the inferred target semantics. Secondly, based on the UTR, we propose a novel Calibrated Adaption Framework (CAF) for SFUDA, including i) the source knowledge calibration module that guides the target model to learn the transferable source knowledge and discard the non-transferable one, and ii) the target semantics calibration module that calibrates the unreliable semantics. With the help of the calibrated source knowledge and the target semantics, the model adapts to the target domain safely and ultimately better. We verified the effectiveness of our method using experimental results and demonstrated that the proposed method achieves state-of-the-art performances on the three SFUDA benchmarks. Code is available at https://github.com/SPIresearch/UTR. Jiangbo Pei, Zhuqing Jiang, Aidong Men, Yang Liu 0105, Qingchao Chen |
IEEE Trans. Image Process. | 2 |
| 2022 | An Efficient Method for Model Pruning Using Knowledge Distillation with Few SamplesabstractDeep neural network compression methods can produce small-scale networks and utilizes fine-tuning to get back the dropped accuracy. Despite their remarkable performance, the fine-tuning procedure is limited to the requirement of a huge training dataset, which is a time-consuming progress. To address the issue, few-sample knowledge distillation (FSKD) has been proposed for data efficiency. However, FSKD needs to add additional convolution layers for compressed networks during training, which increases the complexity of network structure. In this paper, we present Progressive Feature Distribution Distillation (PFDD) without modifying network structures, which surpasses FSKD. Concretely, it is based on a progressive training strategy that is efficient for matching feature distributions between compressed network and original network. Thus, we can notably exploit both external information from samples and internal information from network, where using a small proportion of training dataset can yield quite considerable results. Experiments on various datasets and architectures demonstrate that our distillation approach is remarkably efficient and effective in improving compressed networks’ performance while only few samples have been applied. ZhaoJing Zhou, Zhuqing Jiang, Aidong Men, Haiying Wang 0005 |
ICASSP | 3 |
| 2022 | Mixed In Time And Modality: Curse Or Blessingƒ Cross-Instance Data Augmentation for Weakly Supervised Multimodal Temporal FusionabstractIn multimodal video event localization, we usually leverage feature fusion across different axes, such as the modality and temporal axes, for better context. To reduce the costs of detailed annotations, recent solutions explore weakly supervised settings. However, we observe that when feature fusion meets weakly supervised localization, problems can occur. It may cause "feature cross-interference", which produces a smearing effect on the localization result and can’t be effectively supervised with conventional multiple instance learning loss. We verify it quantitatively on the audio-visual video parsing (AVVP) task, and propose a cross-instance data-augmentation framework, which can preserve the benefits of feature fusion while providing explicit feedbacks for feature cross-interference. We show that our method can enhance performance of existing models on two weakly supervised audio-visual localization tasks, i.e. AVVP and AVE. Yonggang Zhu, Zhuqing Jiang, Aidong Men, Haiying Wang 0005, Qingchao Chen |
ICASSP | 3 |
| 2022 | Delving into the Continuous Domain AdaptationabstractExisting domain adaptation methods assume that domain discrepancies are caused by a few discrete attributes and variations, e.g., art, real, painting, quickdraw, etc. We argue that this is not realistic as it is implausible to define the real-world datasets using a few discrete attributes. Therefore, we propose to investigate a new problem namely the Continuous Domain Adaptation (CDA) through the lens where infinite domains are formed by continuously varying attributes. Leveraging knowledge of two labeled source domains and several observed unlabeled target domains data, the objective of CDA is to learn a generalized model for whole data distribution with the continuous attribute. Besides the contributions of formulating a new problem, we also propose a novel approach as a strong CDA baseline. To be specific, firstly we propose a novel alternating training strategy to reduce discrepancies among multiple domains meanwhile generalize to unseen target domains. Secondly, we propose a continuity constraint when estimating the cross-domain divergence measurement. Finally, to decouple the discrepancy from the mini-batch size, we design a domain-specific queue to maintain the global view of the source domain that further boosts the adaptation performances. Our method is proven to achieve the state-of-the-art in CDA problem using extensive experiments. The code is available at https://github.com/SPIresearch/CDA. Yinsong Xu 0002, Zhuqing Jiang, Aidong Men, Yang Liu 0105, Qingchao Chen |
ACM Multimedia | 2 |
| 2022 | Taylor saves for later: Disentanglement for video prediction using Taylor representation
Zhuqing Jiang, Shiping Wen 0001, Aidong Men, Haiying Wang 0005 |
Neurocomputing | 2 |
| 2022 | Toward a perceptive pretraining framework for Audio-Visual Video Parsing
Jianning Wu, Zhuqing Jiang, Qingchao Chen, Shiping Wen 0001, Aidong Men, Haiying Wang 0005 |
Inf. Sci. | 2 |
| 2022 | Shedding light on images: Multi-level image brightness enhancement guided by arbitrary referencesabstractThe non-linearity between human perception and image brightness levels results in different definitions of NORMAL-light. Thus, most existing low-light image enhancement methods which produce one-to-one mapping can not meet the aesthetic demand. Other pioneers enhance low-light images guided by a given value. However, the inherent problem of non-linearity will cause poor usability. To this end, we propose a user-friendly neural network for multi-level low-light image enhancement. Inspired by style transfer, our method decomposes an image into content component feature and luminance component feature in the latent space. Then we enhance the image brightness to different levels by concatenating the content components from low-light images and the luminance components from reference images. The network meets various user requirements by selecting different brightness references. Moreover, information except for brightness is preserved to alleviate color distortion. Extensive experiments demonstrate the superiority of our network against existing methods. Zhuqing Jiang, Aidong Men, Haiying Wang 0005 |
Pattern Recognit. | 2 |
| 2021 | Multi-DIP: A General Framework for Unsupervised Multi-degraded Image Restoration
Qiansong Wang, Haiying Wang 0005, Aidong Men, Zhuqing Jiang |
ICONIP (4) | 5 |
| 2021 | Lvio-Fusion: A Self-adaptive Multi-sensor Fusion SLAM Framework Using Actor-critic MethodabstractState estimation with sensors is essential for mobile robots. Due to different performance of sensors in different environments, how to fuse measurements of various sensors is a problem. In this paper, we propose a tightly coupled multi-sensor fusion framework, Lvio-Fusion, which fuses stereo camera, Lidar, IMU, and GPS based on the graph optimization. Especially for urban traffic scenes, we introduce a segmented global pose graph optimization with GPS and loop-closure, which can eliminate accumulated drifts. Additionally, we creatively use a actor-critic method in reinforcement learning to adaptively adjust sensors’ weight. After training, actor-critic agent can provide the system better and dynamic sensors’ weight. We evaluate the performance of our system on public datasets and compare it with other state-of-the-art methods, which shows that the proposed method achieves high estimation accuracy and robustness to various environments. And our implementations are open source and highly scalable. Yupeng Jia, Haiyong Luo, Fang Zhao 0003, Guanlin Jiang, Jiaquan Yan, Zhuqing Jiang, Zitian Wang |
IROS | 7 |
| 2021 | A switched view of Retinex: Deep self-regularized low-light image enhancementabstractSelf-regularized low-light image enhancement does not require any normal-light image in training, thereby freeing from the chains of paired or unpaired training data that are time-consuming to obtain. However, existing methods suffer color deviation and fail to generalize to various lighting conditions. This paper presents a novel self-regularized method based on Retinex, which, inspired by HSV, preserves all colors (Hue, Saturation) and only integrates Retinex theory into brightness (Value). Besides, we design a novel random brightness disturbance approach to generate another abnormal brightness of the same scene. It is combined with the original form of brightness to estimate the same reflectance, which is achieved by a CNN. The reflectance, which is assumed irrelevant to any illumination according to the Retinex theory, is treated as the enhanced brightness. Our method is efficient as a low-light image is decoupled into two subspaces, i.e., color and brightness, for better preservation and enhancement. Extensive experiments demonstrate that our method outperforms multiple state-of-the-art algorithms qualitatively and quantitatively and adapts to more lighting conditions. Our code is available at https://github.com/Github-LHT/A-Switched-View-of-Retinex-Deep-Self-Regularized-Low-Light-Image-Enhancement. Zhuqing Jiang, Liangjie Liu, Aidong Men, Haiying Wang 0005 |
Neurocomputing | 1 |
| 2021 | Domain generalization via optimal transport with metric similarity learning
Fan Zhou 0006, Zhuqing Jiang, Changjian Shui, Boyu Wang 0004, Brahim Chaib-draa |
Neurocomputing | 2 |
| 2021 | Multi-view feature fusion for person re-identificationabstractPerson re-identification (ReID) suffers from camera view variants. Existing works, which typically learn a feature for each image, share a limitation that the learned features are single-view: each feature only contains information in one camera view. Thus, view bias occurs when matching pedestrians across camera views. In this paper, we seek to mitigate the view bias by generating multi-view features (fusion of features from a fixed number of cameras). To this end, we define the complementary-view features (complementary features to generate multi-view features with single-view features) and perform in-depth analysis. Based on this insight, we alleviate the view bias in testing and training, respectively. In testing, we present Multi-view Message Passing (MVMP), which generates multi-view features by aggregating single-view features from the neighborhood. In training, we propose Multi-view Feature Fusion Network (MFFN), which involves the single-view feature extractor and the complementary-view feature aggregator. MFFN makes the network sensitive to view-specific cues by adding constraints on multi-view features rather than single-view features. In addition, MVMP and MFFN have two key advantages: (1) They are parameter-free. (2) They can be applied to any Convolutional Neural Networks (CNNs) readily without extra supervision. Extensive experiments are conducted to validate the superiority of our method for person ReID over state-of-the-art methods on four benchmark datasets (Market-1501, DukeMTMC-reID, CUHK03, and MSMT17). The code is available at https://github.com/Yinsongxu/MVMP_MFFN. Yinsong Xu 0002, Zhuqing Jiang, Aidong Men, Haiying Wang 0005, Haiyong Luo |
Knowl. Based Syst. | 2 |
| 2020 | Split to Be Slim: An Overlooked Redundancy in Vanilla ConvolutionabstractMany effective solutions have been proposed to reduce the redundancy of models for inference acceleration. Nevertheless, common approaches mostly focus on eliminating less important filters or constructing efficient operations, while ignoring the pattern redundancy in feature maps. We reveal that many feature maps within a layer share similar but not identical patterns. However, it is difficult to identify if features with similar patterns are redundant or contain essential details. Therefore, instead of directly removing uncertain redundant features, we propose a split based convolutional operation, namely SPConv, to tolerate features with similar patterns but require less computation. Specifically, we split input feature maps into the representative part and the uncertain redundant part, where intrinsic information is extracted from the representative part through relatively heavy computation while tiny hidden details in the uncertain redundant part are processed with some light-weight operation. To recalibrate and fuse these two groups of processed features, we propose a parameters-free feature fusion module. Moreover, our SPConv is formulated to replace the vanilla convolution in a plug-and-play way. Without any bells and whistles, experimental results on benchmarks demonstrate SPConv-equipped networks consistently outperform state-of-the-art baselines in both accuracy and inference time on GPU, with FLOPs and parameters dropped sharply. Qiulin Zhang, Zhuqing Jiang, Qishuo Lu, Zhengxin Zeng, Shanghua Gao, Aidong Men |
IJCAI | 2 |
| 2020 | FloorSense: a novel crowdsourcing map construction algorithm based on conditional random field
Zhuqing Jiang, Chonghua Liu, Chengkai Huang |
Pers. Ubiquitous Comput. | 1 |
| 2019 | Local to Global with Multi-Scale Attention Network for Person Re-IdentificationabstractRecently, part-based person re-identification methods attract lots of attention and largely improve the accuracy. However, due to the large variations in camera occlusion, pose change and misalignment, the corresponding part regions of different images from a same person may miss the key cues. In this paper, we proposed a local to global with multi-scale attention network (LGMANet), which sufficiently exploits the contextual information and spacial attention information. Our proposed model includes two branches. One is local to global branch. By pooling operation, an image generates the feature maps of different dimensions. Then, we learn local to global descriptors by partitioning these feature maps with the same scale. The other is multi-scale attention branch, which captures the contextual dependencies from different convolution layers and further improves the discriminative ability of the image feature. Experimental results demonstrate that our method achieves the state-of-the-art results on three benchmark datasets, Market-1501, DukeMTMC-reID and CUHK03. Lingchuan Sun, Jianlei Liu, Yingxin Zhu, Zhuqing Jiang |
ICIP | 4 |
| 2019 | Multi-Branch Context-Aware Network for Person Re-IdentificationabstractMost existing methods on person re-identification ignore contextual dependencies which are important in representing pedestrian images. In this paper, we propose a Multi-Branch Context-Aware Network (MBCAN) for person re-identification to exploit rich context information. MBCAN learns global features and local-part features in two separate branches to take full advantages of both coarse-grained and fine-grained features. Additionally, two types of attention modules are introduced to capture contextual dependencies in spatial dimension and channel dimension, respectively. A module called feature vector extraction block is designed to find an efficient way to integrate features from coarse to fine. Extensive experiments with ablation analysis show the effectiveness of our method, and state-of-the-art results are achieved on Market-1501, DukeMTMC-reID and CUHK03 datasets. Yingxin Zhu, Jianlei Liu, Zhuqing Jiang |
ICIP | 4 |
| 2019 | Recursive Multi-Stage Upscaling Network with Discriminative Fusion for Super-ResolutionabstractSince convolutional neural networks have fundamentally changed how computers learn features, many super-resolution (SR) methods focus on extracting informative features to improve performance by using more layers or more innovative skip-connections. However, using more layers in feature extraction while keeping the upscaling module unchanged will exacerbate structural imbalances. Besides, blindly fusing different stage features by concatenation may make them interfere with each other. To address these issues, we proposed a recursive multi-stage upscaling network (RMUN) with discriminative fusion module (DFM). Specifically, we construct multiple upscaling paths to produce various high-resolution features in the forward propagation and deliver error loss in the back propagation. Furthermore, we fuse and re-weight those features by DFM to avoid mutual interference and boost reconstruction quality. Experiments show that RMUN is superior to the state-of-the-art methods, especially for large scale SR tasks. Zhuqing Jiang, Guodong Ju, Liangheng Shen, Aidong Men |
ICME | 2 |
| 2019 | Multi-Branch Context-Aware Network for Person Re-IdentificationabstractMost existing methods on person re-identification pay redundant attention to global features or local features which ignore contextual dependencies which are equally important in representing pedestrian images. In this paper, we propose a Multi-Branch Context-Aware Network (MBCAN) for person re-identification to exploit rich context information. MBCAN learns global features and local-part features in two separate branches to take full advantages of both coarse-grained and fine-grained features. Additionally, two types of attention modules are introduced to capture contextual dependencies in spatial dimension and channel dimension, respectively. A module called feature vector extraction block is designed to find an efficient way to integrate features from coarse to fine. Extensive experiments with ablation analysis show the effectiveness of our method, and state-of-the-art results are achieved on Market-1501, DukeMTMC-reID and CUHK03 datasets. Yingxin Zhu, Jianlei Liu, Zhuqing Jiang |
ICME | 4 |
| 2019 | Attentional Part-based Network for Person Re-identificationabstractPart-based network is an effective method to improve performance in person re-identification (re-ID). Most existing methods assume the availability of well-aligned person bounding box images as model input. However, automatic detection in some datasets causes misalignment which negatively affects the performance. In this work, we propose an Attentional Part-based CNN (AP-CNN) model which combines learning partial features and attention selection. First, we partition feature map into several horizontal stripes. Second, we use attention selection in each stripe to align the pedestrian images. Inside, we introduce a free-parameter attention model with skip-layer connection which maximizes the complementary information of different levels without increasing the complexity of network. Results on four datasets validate the competitiveness of AP-CNN over the state-of-the-art achieving Rank-1 accuracy of 94.4% on Market-1501, 87.3% on DukeMTMC-ReID, 73.7% on CUHK03-labeled and 72.6% on CUHK03-detected. Yinsong Xu 0002, Zhuqing Jiang, Aidong Men, Jiangbo Pei, Guodong Ju, Bo Yang 0007 |
VCIP | 2 |
| 2019 | A Multiple Triplet-Ranking Model for Fine-Grained Sketch-Based Image RetrievalabstractFine-grained sketch-based image retrieval (FG-SBIR) addresses the problem of matching an input sketch with a specific photo containing the same instance. The key challenge of learning a FG-SBIR model is to bridge the domain gap between photo and sketch. Most existing approaches build a joint embedding space where two domains can be directly compared. They only focus on the highly abstract features in final fully connected (FC) layer, ignore some low-level semantic concepts in convolutional layers. In this paper, we propose a multiple triplet-ranking model in FG-SBIR task. Specially, we introduce an auxiliary supervision loss function in the convolutional layer, and we use the fusion of features from convolutional layer and final FC layer to build the joint embedding space. Extensive experiments show that the proposed multiple triplet-ranking model significantly outperforms the state-of-the-art. Jingyi Xue, Zhuqing Jiang |
VCIP | 3 |
| 2019 | Pyramid Real Image Denoising NetworkabstractWhile deep Convolutional Neural Networks (CNNs) have shown extraordinary capability of modelling specific noise and denoising, they still perform poorly on real-world noisy images. The main reason is that the real-world noise is more sophisticated and diverse. To tackle the issue of blind denoising, in this paper, we propose a novel pyramid real image denoising network (PRIDNet), which contains three stages. First, the noise estimation stage uses channel attention mechanism to recalibrate the channel importance of input noise. Second, at the multi-scale denoising stage, pyramid pooling is utilized to extract multi-scale features. Third, the stage of feature fusion adopts a kernel selecting operation to adaptively fuse multi-scale features. Experiments on two datasets of real noisy photographs demonstrate that our approach can achieve competitive performance in comparison with state-of-the-art denoisers in terms of both quantitative measure and visual perception quality. Yiyun Zhao, Zhuqing Jiang, Aidong Men, Guodong Ju |
VCIP | 2 |
| 2019 | Object detection using convolutional networks with adaptively adjusting receptive field of convolutional filterabstractThe receptive field size of a convolutional filter in a deep convolutional network is a crucial issue for object detection task, as the output must response to a suitable size of area in the image to capture proper information. Receptive field size of convolutional filter is fixed due to the inherently fixed geometric structure in its building module. However, objects of interest vary significantly in size within the images for object detection. Different locations of images correspond to objects with different scales, and high level convolutional layers encode semantic features over spatial positions, thus adaptive determination of receptive field size of convolutional filter is desirable for object detection. The authors propose a new module to adaptively determine the receptive field size of convolutional filter, named adaptive convolution. It is based on the idea of dilating the convolutional filter with multiple dilation values and choosing the maximum activation as output, without adding any other parameters. The plain counterparts in existing convolutional neural networks can be easily replaced by adaptive convolution, giving rise to adaptive convolutional networks. Adequate experiments have proven the effectiveness of authors’ method. Qishuo Lu, Zhuqing Jiang, Aidong Men, Pengliang Tang |
IET Comput. Vis. | 2 |
| 2018 | Person Re-Identification by Deep Learning Muti-Part Information ComplementaryabstractPerson re-identification (Re-ID) aims to identify people across disjoint camera views, which is considered either a binary classification task or a ranking task. However, the importance of feature extracting in both tasks is the same. In this paper, we introduce Global-Part Network (GPN) which employ multi-part information fusion with Feature Weighting Structure (FWS) to address the person Re-ID as a retrieval task. In our framework, the global and body-part features of a specific person can be complementary to each other to enhance the final feature representation. Furthermore, we utilize a state-of-the-art re-ranking method to improve the ranking performance. Extensive experimental analyses and results on three popular datasets, i.e., Market1501, CUHK03, CUHK01, demonstrate the effectiveness of the proposed approach. Zhuqing Jiang |
ICIP | 2 |
| 2018 | Deep Network with Spatial and Channel Attention for Person Re-identificationabstractMost existing person re-identification (Re-ID) methods assume pedestrian images are well-aligned within tightly surrounded bounding boxes or require additional annotation information to calibrate misaligned images. In this work, we propose a novel deep network to address the misalignment problem in person re-identification task without requiring additional annotation. Spatial attention selection mechanism is introduced in our network to align the pedestrian images. Moreover, we present a channel attention selection mechanism to integrate the global image feature maps and regional feature maps more effectively by explicitly modelling interdependencies between channels and recalibrates feature response in each channel. Extensive experiments and comparative evaluations demonstrate the effectiveness of our approach and the superiority of this novel network for person re-identification over a wide variety of state-of-the-art methods on two datasets including Market-1501 and CUHK03(both detected and labeled sets). Tiansheng Guo, Dongfei Wang, Zhuqing Jiang, Aidong Men |
VCIP | 3 |
| 2018 | Channel Attention and Multi-level Features Fusion for Single Image Super-ResolutionabstractConvolutional neural networks (CNNs) have demonstrated superior performance in super-resolution (SR). However, most CNN-based SR methods neglect the different importance among feature channels or fail to take full advantage of the hierarchical features. To address these issues, this paper presents a novel recursive unit Firstly, at the beginning of each unit, we adopt a compact channel attention mechanism to adaptively recalibrate the channel importance of input features. Then, the multi-level features, rather than only deep-level features, are extracted and fused. Additionally, we find that it will force our model to learn more details by using the learnable upsampling method (i.e., transposed convolution) only on residual branch (instead of using it both on residual branch and identity branch) while using the bicubic interpolation on the other branch. Analytic experiments show that our method achieves competitive results compared with the state-of-the-art methods and maintains faster speed as well. Zhuqing Jiang |
VCIP | 3 |
| 2018 | Graph Regularized and Label-matched Dictionary Learning for Video-based Person Re-identificationabstractIn recent years, video-based person re-identification has attracted more and more attention. However, most existing video-based methods do not fully consider the intrinsic structure and invariant information of the same person across different cameras. In this paper, we propose a graph regularized and label-matched dictionary learning (GRLDL) method to capture the intrinsic structure of the same person between two cameras. Firstly, in order to reduce the variations between different cameras, we use local Fisher discriminant analysis to transform the person videos from different cameras into a common feature space. A dictionary is learned from this common space. Then, we construct a graph regularization term to preserve the geometrical structure of the same person and enhance the discriminative ability of the learned dictionary. Finally, a projective matrix is introduced to map the coding coefficients into a label space, which is able to correlate and match the same person under different cameras. Experiments on the public iLIDS-VID and PRID 2011 datasets show the effectiveness of the proposed method. Lingchuan Sun, Jianlei Liu, Zhuqing Jiang |
VCIP | 4 |
| 2018 | A QoS routing strategy using fuzzy logic for NGEO satellite IP networks
Zhuqing Jiang, Chonghua Liu, Shanbao He, Qishuo Lu |
Wirel. Networks | 1 |
| 2017 | Real-time object detection by a multi-feature fully convolutional networkabstractPrior work on object detection depends on region proposals to guide the search for object instances. Generally, several thousand proposals must be processed, thus hurting the detection efficiency. In this paper, we propose a new model free from region proposals for object detection which treats detection task as a regression problem. To improve small-size object detection and localization, we employ the deep hierarchical features extracted from convolutional neural networks (CNNs). The hierarchical architecture combines appearance information from a shallow layer with semantic information from a deep layer. Our approach can predict bounding boxes and class probabilities simultaneously from a full input image. We transfer a classification network called Darknet into fully convolutional network and fine-tune it for the detection task. Experiments on PASCAL VOC dataset demonstrate that our approach outperforms other detection models. Yajing Guo, Zhuqing Jiang, Aidong Men |
ICIP | 3 |
| 2017 | Coupled analysis-synthesis dictionary learning for person re-identificationabstractIn this paper, we propose a novel coupled dictionary learning method, namely coupled analysis-synthesis dictionary learning, to improve the performance of person re-identification in the non-overlapping fields of different camera views. Most of the existing coupled dictionary learning methods train a coupled synthesis dictionary directly on the original feature spaces, which limits the representation ability of the dictionary. To handle the diversities of different original spaces, We first employ local Fisher discriminant analysis (LFDA) to learn a common feature space for close relationship of the same people in different views. In order to enhance the representation power of the coupled synthesis dictionary, we then learn a coupled analysis dictionary by transforming the common feature space into the coupled feature space. Experimental results on two publicly available VIPeR and CUHK01 datasets have validated the effectiveness of the proposed method. Lingchuan Sun, Zhuqing Jiang, Aidong Men |
ICIP | 3 |
| 2017 | Cascaded convolutional neural networks for object detectionabstractRecent advances in object detection depend on region proposal algorithms or networks to predict object locations. The pipeline of region proposal-based object detection can be decomposed into two cascaded sub-tasks: 1) region proposals generation from input image, 2) proposals classification into various object categories. In this paper, we propose cascaded convolutional neural networks to make improvement for two sub-tasks respectively. For the region proposals generation stage, we add a RefineNet after the original region proposal network(RPN) to make the proposals more compact and better located. For the classification stage, we integrate a binary classifier for each object class into the network which makes the feature representation capture more intra-class variance. Experiments on PASCAL VOC dataset demonstrate that our approach can achieve considerable improvement over state-of-the-art object detectors. Yajing Guo, Zhuqing Jiang |
VCIP | 3 |
| 2017 | Improve object detection via a multi-feature and multi-task CNN modelabstractCurrent state-of-the-art object detection methods have made dramatic performance improvements in recent few years. However, there are still several challenges. In particular, it still struggles for precise localization of small-sized objects, mainly due to coarse resolutions of feature maps and excessive surroundings such as ground and water. To address the issues, we propose an object detection system based on standard Fast R-CNN object detection branch and DeepLap semantic segmentation branch: (1) multi-feature aggregates hierarchical features for more finer feature maps to detect objects at multiple scales. (2) multi-task uses semantic segmentation for more contextual information to assist object detection via a cross structure between the two tasks. (3) a novel overlap loss function is used for bounding box regression that adjusts region proposals to improve localization. The fusion network improves results over the Fast R-CNN baseline detector by 2.8% mAP and by 4.8% mAP for small objects based on PASCAL VOC datasets. Yingxin Lou, Guangtao Fu, Zhuqing Jiang, Aidong Men |
VCIP | 3 |
| 2015 | A smartphone-based indoor positioning system using fuzzy theory and WLAN mapping algorithmabstractAs we all know, the accuracy of GPS become worse indoors due to the satellite signal attenuation. However, people spend most of his or her time in the indoor environment and the location-based service is needed in the indoor environment. Hence, many indoor positioning techniques have been researched to provide location-based service for visitors in public buildings such as museums, galleries, etc. However, the distinguish of human's moving patterns and initial position detection are two difficult problems in Pedestrian Dead Reckoning (PDR) system. Therefore, A smartphone-based indoor positioning system (called Improved Pedestrian Dead Reckoning, IPDR) using fuzzy theory and WLAN mapping algorithm is presented to solve these problems. In our proposed system, we propose a method called fuzzy step length detection which can distinguish five different human's moving patterns and handle various ways of holding the phone. The WLAN mapping algorithm could calculate the initial position. Calibrated position can be obtained from the IPDR position and the mapping position from WLAN map acquired in advance. The performance of our IPDR system is verified via a set of simulations and the result is positive. Zhuqing Jiang, Chengkai Huang, Xinmeng Liu, Yuying Yang |
PIMRC | 2 |
| 2015 | A Novel Routing Strategy Based on Fuzzy Theory for NGEO Satellite NetworksabstractNon-geostationary (NGEO) satellite networks have a series of advantages over terrestrial networks. However, traditional routing algorithms such as the Dijkstra's Shortest Path (DSP) algorithm always lead to some Inter-Satellite Links (ISLs) heavily loaded. To guarantee a better distribution of traffic among satellites, this paper proposes a Fuzzy Satellite Congestion Indicator (FSCI) to estimate congestion status among neighboring satellites. Indeed, a satellite notifies its neighboring satellites of its FSCI. When it is about to get congested, it requests its neighboring satellites to decrease their data forwarding rates by sending them a self status notification signaling message. In response, the neighboring satellites search for less congested paths according to Fuzzy Route Determination. The routing strategy discussed above is Fuzzy Satellite Routing(FSR). This routing algorithm avoids both congestion and packet drops at the satellite. It also ensures a better traffic distribution over the entire satellite constellation. The mechanism of multiple traffic classes is also discussed in FSR. The good performance of FSR, in terms of short end-to-end delay, higher throughput, and lower packet drops, is verified via a set of simulations using the Network Simulator 2 (NS2). Chonghua Liu, Zhuqing Jiang, Xinmeng Liu, Yuying Yang |
VTC Fall | 3 |
| 2015 | A Novel Fuzzy Pedestrian Dead Reckoning System for Indoor Positioning Using SmartphoneabstractRecently, due to the low accuracy of GPS in indoor environment, many indoor positioning techniques have been researched to provide location-based service for visitors in public buildings such as museums, galleries, etc. But many indoor positioning techniques cannot distinguish human's moving patterns and external infrastructures are needed to improve accuracy. In this paper, a Fuzzy Pedestrian Dead Reckoning (FPDR) system with fuzzy magnetic map matching algorithm is presented. This system only utilizes inertial sensors of smartphone. The fuzzy stride length algorithm in FPDR could handle five different human moving patterns and various ways of holding the phone. Also, walking distances, heading direction and floor of building can be obtained in real time. When the user stops to appreciate showpieces in museums, relevant magnetic data will be collected. The fuzzy magnetic map matching algorithm could calculate the final location with the position from FPDR and the calibration position from magnetic map acquired in advance. The performance of FPDR system is verified via a set of simulations and the error rate is lower than other typical indoor location methods. Jinjun Zheng, Zhuqing Jiang, Xinmeng Liu, Yuying Yang, Beihang Zhang |
VTC Fall | 3 |
| 2015 | A Low-Complexity Routing Algorithm Based on Load Balancing for LEO Satellite NetworksabstractThe mesh topology structure constituted by inter- satellite links(ISLs) along with longitude and latitude, as a feature of most LEO satellite constellations, has not been fully used. Additionally, most LEO networks design its routing protocol depending on path distance, which is only proportional to the propagation delay without consideration of the queue delay. In this paper, a low-complexity routing algorithm (LCRA) based on load balancing is proposed, which can obtain the best path by distributed computation with the location information of the current node and destination. There is no iteration process in the computation, thus saving the computational cost. Additionally, each node informs its neighbouring nodes of its congestion information so that packets can choose the next hop dynamically according to the status of links, so as to shorten the mean queue delay and reduce the packet loss rate. Results assessed by NS2 presents the superiority of LCRA in terms of end-to-end delay, throughput and packet loss rate in comparison with other routing algorithms. Xinmeng Liu, Xuemei Yan, Zhuqing Jiang, Yuying Yang |
VTC Fall | 3 |
| 2015 | Video saliency detection incorporating temporal information in compressed domain
Qin Tu, Aidong Men, Zhuqing Jiang |
Signal Process. Image Commun. | 3 |
| 2014 | Indoor positioning system based on improved PDR and magnetic calibration using smartphoneabstractIn recent decades, indoor positioning techniques have been researched to support automatic guidance for visitors in public buildings such as museums, galleries, etc. The exhibition goods nearby could be introduced by a smartphone after the location of the visitor has been acquired. This paper presents an indoor positioning scheme utilizing smartphones equipped with inertial and magnetic sensors. The positioning method could handle complicated human motion and various ways to hold the phone. This is achieved by applying an improved Pedestrian Dead Reckoning (PDR) algorithm and an error-tolerant magnetic map matching algorithm. The improved PDR algorithm estimates walking distances and direction in real time. As long as the user stops to appreciate showpieces, the improved PDR component will report relevant data to magnetic calibration component. Based on the updated information, the final location of the user could be calculated utilizing the relevant data, the real-time data of magnetic field sensor and the magnetic map acquired in advance. We evaluate this method using experimental measurements in practice. Chengkai Huang, Shanbao He, Zhuqing Jiang, Yupeng Wang 0002 |
PIMRC | 3 |
| 2014 | A Novel Routing Algorithm Design of Time Evolving Graph Based on Pairing Heap for MEO Satellite NetworkabstractAs a tradeoff of GEO and LEO, MEO satellite system has more acceptable service performance and it is more appropriate to provide global mobile communications. A MEO satellite system model communicating according to time slots is constructed in the paper. Moreover, in order to improve comprehensive performance of the network, a novel routing algorithm applying Time Evolving Graph based on Pairing Heap is proposed. The Time Evolving Graph is employed to analysis the dynamic topology of the network and the Pairing Heap is applied in the Dijkstra algorithm to reduce the time complexity. By contrast, Fibonacci Heap is also used to optimize Dijkstra algorithm. Finally simulation results show that routing algorithm applying Time Evolving Graph based on Pairing Heap can perform better and reduce the time complexity obviously, and at the same time, Pairing Heap works better than Fibonacci Heap when the number of nodes grows bigger. Yupeng Wang 0002, Zhuqing Jiang, Chengkai Huang, Aidong Men, Bo Yang 0007, Kaifeng Qi |
VTC Fall | 3 |
| 2013 | GPS/INS Integrated Navigation Based on UKF and Simulated Annealing Optimized SVMabstractThe accuracy of Global Positioning System (GPS) is often combined with the reliability of Inertial Navigation System (INS) to accomplish navigation. This paper proposes an innovative way to filter and fuse the GPS and INS information. UKF is employed to simulate the information convergence of the dynamic model which maintains better performance in nonlinear system. So we can obtain a fair precise filtering result when both are online. At the same time, the INS data is trained with the result as training target when it is the unique input. This paper raises the idea that Support Vector Machine (SVM) is adopted to train the INS data during GPS outage and the simulated annealing is applied to realize the optimization of the parameters of kernel function and the penalty function in the SVM algorithm. Therefore, the integration navigation could retain almost as precise as the GPS when the GPS is off-line. Zhuqing Jiang, Chonghua Liu, Yupeng Wang 0002, Chengkai Huang, Jiayi Liang |
VTC Fall | 1 |