EDBT 2026 Demo / reviewers in the wild / expert
Jianqing Zhu
dblp:161/0779
· DBLP profile ↗
73ranked-venue papers
16as first author
54since 2021 · last 2026
0000-0001-8840-3629ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-author · 24 since 2021Artificial intelligence and machine learning · 32 · 7 first-author · 29 since 2021Computer networks · 7 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TAdaRAG: Task Adaptive Retrieval-Augmented Generation via On-the-Fly Knowledge Graph ConstructionabstractRetrieval-Augmented Generation (RAG) improves large language models by retrieving external knowledge, often truncated into smaller chunks due to the input context window, which leads to information loss, resulting in response hallucinations and broken reasoning chains. Moreover, traditional RAG retrieves unstructured knowledge, introducing irrelevant details that hinder accurate reasoning. To address these issues, we propose TAdaRAG, a novel RAG framework for on-the-fly task-adaptive knowledge graph construction from external sources. Specifically, we design an intent-driven routing mechanism to a domain-specific extraction template, followed by supervised fine-tuning and a reinforcement learning-based implicit extraction mechanism, ensuring concise, coherent, and non-redundant knowledge integration. Evaluations on six public benchmarks and a real-world business benchmark (NowNewsQA) across three backbone models demonstrate that TAdaRAG outperforms existing methods across diverse domains and long-text tasks, highlighting its strong generalization and practical effectiveness. Jie Zhang 0166, Bo Tang 0011, Wanzi Shao, Wenqiang Wei, Jihao Zhao, Jianqing Zhu, Wen Xi, Zehao Lin, Feiyu Xiong, Yanchao Tan |
AAAI | 6 |
| 2026 | Industrial visual defect detection oriented context-modulated and cross-layer interaction network for image super-resolution
Xiancheng Zhu, Detian Huang, Jianqing Zhu, Huanqiang Zeng |
Expert Syst. Appl. | 4 |
| 2026 | Text-to-video person re-identification benchmark: Dataset and dual-modal contextual alignment
Jiajun Su, Simin Zhan, Pudu Liu, Jianqing Zhu, Huanqiang Zeng |
Neurocomputing | 4 |
| 2026 | UDD: Unsupervised denoising diffusion for noisy multi-focus image fusion
Pudu Liu, Jiajun Su, Simin Zhan, Jianqing Zhu, Huanqiang Zeng |
Neural Networks | 5 |
| 2026 | Focusing on pedestrians like human for clothes changing person re-identificationabstract• Based on the ensemble coding hypothesis in cognitive neuroscience, we achieve data augmentation by simulating human focus capability. • Our method is the first local detail learning data augmentation designed specifically for clothes changing person re-identification. • Our method achieves SOTA on three public datasets and outperforms various traditional data augmentation methods. Current approaches focus mainly on the design of networks to learn key identity features from local body components for clothes-changing person re-identification (CC-ReID). In this paper, we propose a humanoid focus-inspired image augmentation (HFIA) method, which is intuitive image processing rather than a sophisticated network architecture designed to enhance local nuances of pedestrian images. Based on pedestrian silhouettes, we roughly divide a pedestrian image into five body components, that is, head-shoulder, upper left torso, upper right torso, lower left torso, and lower right torso. The HFIA has two key designs to deal with these components: the central emphasis strategy (CES) and the component continuity processing (CCP). For each component, leveraging the natural tendency of human visual attention towards central regions, the CES constructs an enlargement grid, where the closer the center, the greater the enlargement. To maintain the continuity of assembly, the CCP performs an overall alignment of component centers, that is, all components share the same normalized vertical coordinate and the left and right torsos have mirrored horizontal coordinates. Furthermore, the CCP implements a smoothing post-processing to uniformly erase the discontinuity between the head-shoulder, upper left torso, and upper right torso. Experiments show the state-of-the-art performance of HFIA. Wenjie Pan, Jianqing Zhu, Xiaolin Cui, Huanqiang Zeng, Yibing Zhan |
Neural Networks | 2 |
| 2026 | Adversarial discriminant attack on text-to-image diffusion models
Hanxiao Wu, Shengwu Xiong 0001, Dong Yi, Lingxiang Wu, Jianqing Zhu, Guibo Zhu, Jinqiao Wang |
Neural Networks | 5 |
| 2026 | Exploring non-local spatial-angular correlations with a hybrid Mamba-Transformer framework for light field super-resolution
Haosong Liu, Xiancheng Zhu, Huanqiang Zeng, Jianqing Zhu, Jiuwen Cao, Junhui Hou |
Pattern Recognit. | 4 |
| 2026 | Point cloud compression using graph neural networks
Xiangjie Zhang, Huanqiang Zeng, Xin-Rong Gong, Junhui Hou, Jianqing Zhu |
Pattern Recognit. | 6 |
| 2026 | Structure-Aware Alignment for Day-Night Cross-Domain Vehicle Re-IdentificationabstractDay-night cross-domain vehicle re-identification (DN-ReID) is fundamentally challenged by drastic illumination changes that create substantial domain gaps and hinder consistent feature representation. Most existing methods focus on aligning distribution statistics but often overlook essential structural relationships, such as cross-domain clustering and geometric topology. To address this, we propose a structure-aware alignment (SAA) method that, for the first time, formulates centered kernel alignment as a trainable loss for structural alignment between domains. This approach explicitly aligns cross-domain relational structures, thereby supporting the transfer of intra-class compactness and inter-class separability from the source to the target domain for robust cross-domain matching. Extensive experiments on DN-348 and DN-Wild demonstrate that our approach consistently outperforms state-of-the-art methods, achieving a 2.7% mAP improvement on DN-348. Jingyi Zhuang, Baihui Sa, Jinjie Zheng, Liu Liu 0014, Jianqing Zhu, Huanqiang Zeng |
IEEE Signal Process. Lett. | 5 |
| 2026 | Uncertainty-driven Progressive Single Image De-rainingabstractOver the past years, progressive methods have demonstrated promising performance in single image de-raining task. Nonetheless, current methods still struggle to precisely remove rain and preserve more image details during the progressive de-raining process, resulting in undesirable local artifacts or image detail loss. To tackle these limitations, a novel progressive approach, called Uncertainty-driven Progressive Single Image De-raining (UPSID), is proposed. Firstly, a powerful internal-and-external dense sub-network is devised, which effectively integrates three proper and flexible components, including dense connection, long short-term memory, and channel attention. Subsequently, the sub-network is further unfolded into multiple recurrent stages to form a progressive de-raining network. Finally, the overall progressive de-raining network is trained with an adaptive weighted loss to focus more on challenging pixels that characterize rain or texture/edge regions. Extensive quantitative and qualitative experiments confirm that the proposed UPSID outperforms multiple state-of-the-art algorithms, including single-stage, progressive, and uncertainty-driven single image de-raining methods. Additionally, this article also demonstrates the superiority of UPSID for other similar image restoration tasks such as single image de-snowing. The code will be publicly available at https://github.com/Lcai-QZ/UPSID . Jianqing Zhu, Huanqiang Zeng, Tao Zhu 0002, Jing Chen 0001, Wenkang Su 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Fair Training with Zero InputsabstractThere are two manifestations of classification fairness. One is the preference for head classes with more instances due to the long-tail (LT) distribution of training data. The other is the clever Hans (CH) effect, where non-discriminative features are mistakenly used for classification. In this paper, we find that using category-agnostic zero-valued data can simultaneously reveal both types of unfairness. Based on this, we propose a zero uniformity training (ZUT) framework to optimize classification fairness. The ZUT framework inputs category-agnostic zero-valued data into the model in parallel and uses zero uniformity loss (ZUL) to optimize classification fairness. The ZUL loss mitigates bias towards specific classes by unifying the classification features corresponding to zero-valued data. The ZUT framework is compatible with various classification-based tasks. Experiments show that the ZUT framework can improve the performance of multiple state-of-the-art methods in image classification, person re-identification, and semantic segmentation. Wenjie Pan, Jianqing Zhu, Huanqiang Zeng |
AAAI | 2 |
| 2025 | Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary ExpansionabstractJianqing Zhu, Huang Huang, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An, Juncai He, Xiangbo Wu, Fei Yu, Junying Chen, Ma Zhuoheng, Yuhao Du, He Zhang, Saied Alshahrani, Emad A. Alghamdi, Lian Zhang, Ruoyu Sun, Haizhou Li, Benyou Wang, Jinchao Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jianqing Zhu, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An 0004, Juncai He 0001, Xiangbo Wu, Fei Yu 0017, Zhuoheng Ma, Saied Alshahrani, Emad A. Alghamdi, Ruoyu Sun 0001, Haizhou Li 0001, Benyou Wang, Jinchao Xu |
ACL (1) | 1 |
| 2025 | Bridging Clustering and Training Gap in Unsupervised Vessel Re-identificationabstractUnsupervised vessel re-identification is challenged by a gap between clustering and training stages, where clustering relies on features from original images while training utilizes features from enhanced images. To address this, we propose a novel consistent contrastive learning (CCL) approach that innovatively applies earth mover’s distance (EMD) as a consistency loss function. EMD effectively aligning both global and local structures of original and enhanced features, thus bridging the gap between clustering and training stages. Unlike traditional metrics such as Euclidean distance (ED) and KL divergence, which primarily focus on pointwise or global similarity, EMD considers both overall distribution and local feature importance, enabling more robust feature alignment. Moreover, compared to the structural similarity index metric (SSIM), EMD dynamically adjusts local feature contributions using weight information, ensuring precise matching even in complex scenarios. Thanks to these advantages, our method demonstrates a 4.0% improvement in mean average precision over the original VesselReID benchmark, highlighting its effectiveness in addressing the challenges of unsupervised vessel re-identification. Jianqing Zhu, Linhan Huang, Huanqiang Zeng |
IJCNN | 1 |
| 2025 | Semantic segmentation in power grid scenarios using scale-transforming transformer
Wenjie Pan, Linhan Huang, Yuqing Fu, Jianqing Zhu, Yibing Zhan |
Appl. Intell. | 5 |
| 2025 | Extensions in channel and class dimensions for attention-based knowledge distillationabstractAs knowledge distillation technology evolves, it has bifurcated into three distinct methodologies: logic-based, feature-based, and attention-based knowledge distillation. Although the principle of attention-based knowledge distillation is more intuitive, its performance lags behind the other two methods. To address this, we systematically analyze the advantages and limitations of traditional attention-based methods. In order to optimize these limitations and explore more effective attention information, we expand attention-based knowledge distillation in the channel and class dimensions, proposing Spatial Attention-based Knowledge Distillation with Channel Attention (SAKD-Channel) and Spatial Attention-based Knowledge Distillation with Class Attention (SAKD-Class). On CIFAR-100, with ResNet8 × 4 as the student model, SAKD-Channel improves Top-1 validation accuracy by 1.98%, and SAKD-Class improves it by 3.35% compared to traditional distillation methods. On ImageNet, using ResNet18, these two methods improve Top-1 validation accuracy by 0.55% and 0.17%, respectively, over traditional methods. We also conduct extensive experiments to investigate the working mechanisms and application conditions of channel and class dimensions knowledge distillation, providing new theoretical insights for attention-based knowledge transfer. Liangtai Zhou, Banghui Zhang, Junhuang Wang, Jianqing Zhu |
Comput. Vis. Image Underst. | 7 |
| 2025 | Contrastive Deep Supervision Meets self-knowledge distillation
Peng Liang 0025, Jianqing Zhu, Junhuang Wang |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | MDIKD: multi-dimensional integration knowledge distillation
Maohai Pang, Jianqing Zhu |
Multim. Syst. | 4 |
| 2025 | A strong benchmark for yoga action recognition based on lightweight pose estimation model
Liangtai Zhou, Banghui Zhang, Jianqing Zhu |
Multim. Syst. | 5 |
| 2025 | A federated vehicle re-identification benchmark and performance optimization using dual-phase contrastive dynamic aggregation
Linhan Huang, Jianqing Zhu, Huanqiang Zeng |
Neural Networks | 2 |
| 2025 | Memory-augmented shuffled meta learning for visible-infrared person re-identification
Hanxiao Wu, Jianqing Zhu, Liu Liu 0014, Huanqiang Zeng |
Neural Networks | 4 |
| 2025 | Harmonizing Metric Discrepancy for Cross-Modal Object Re-IdentificationabstractVisible and infrared cross-modal re-identification tasks often encounter significant modal discrepancies, which undermine the effectiveness of feature extraction and compromise the reliability of similarity metrics. These discrepancies pose a substantial challenge for accurately matching data across different modalities. To address these issues, we propose a novel approach centered on the maximum mean metric discrepancy (MMMD). We leverage kernel-based statistical techniques to effectively capture and quantify the disparities in cross-modal metrics, providing a robust framework for aligning metrics from different modalities. Building upon the foundation of MMMD, we develop the metric discrepancy harmonization (MDH) method. This method integrates a temperature-controlled optimization technique designed to enhance metric alignment across various modal configurations, ensuring more consistent and reliable performance. By focusing on metric alignment, our approach enhances the accuracy of cross-modal re-identification tasks. Comprehensive evaluations on the LLCM, RGBN300, and SYSU-MM01 datasets demonstrate that our approach achieves state-of-the-art performance. Linhan Huang, Liu Liu 0014, Jianqing Zhu, Huanqiang Zeng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Ensemble Prototype Networks for Unsupervised Cross-Modal Hashing With Cross-Task ConsistencyabstractIn the swiftly advancing realm of information retrieval, unsupervised cross-modal hashing has emerged as a focal point of research, taking advantage of the inherent advantages of the multifaceted and dynamism inherent in multimedia data. Existing unsupervised cross-modal hashing methods rely mainly on initial pre-trained correlations among cross-modal features, and the inaccurate neighborhood correlations impacts the presentation of common semantics throughout the optimization. To address the aforementioned issues, we proposeEnsemblePrototypeNetworks (EPNet), which delineates class attributes of cross-modal instances through an ensemble clustering methodology. EPNet seeks to extract correlation information between instances by leveraging local correlation aggregation and ensemble clustering from multiple perspectives, aiming to reduce initialization effects and enhance cross-modal representations. Specifically, the local correlation aggregation is first proposed within a batch of semantic affinity relationships to generate a precise and compact hash code among cross-modal instances. Secondly, the ensemble prototype module is employed to discern the class attributes of deep features, thereby aiding the model in extracting more universally applicable feature representations. Thirdly, an early attempt to constrict the representational congruity of local semantic affinity relationships and deep feature ensemble prototype correlations using cross-task consistency loss aims to enhance the representation of cross-modal common semantic features. Finally, EPNet outperforms several state-of-the-art cross-modal retrieval methods on three real-world image-text datasets in extensive experiments. Huanqiang Zeng, Yifan Shi 0001, Jianqing Zhu, Kaixiang Yang 0001, Zhiwen Yu 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | AceGPT, Localizing Large Language Models in ArabicabstractHuang Huang, Fei Yu, Jianqing Zhu, Xuening Sun, Hao Cheng, Song Dingjie, Zhihong Chen, Mosen Alharthi, Bang An, Juncai He, Ziche Liu, Junying Chen, Jianquan Li, Benyou Wang, Lian Zhang, Ruoyu Sun, Xiang Wan, Haizhou Li, Jinchao Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fei Yu 0017, Jianqing Zhu, Xuening Sun, Dingjie Song, Mosen Alharthi, Bang An 0004, Juncai He 0001, Ziche Liu, Benyou Wang, Ruoyu Sun 0001, Haizhou Li 0001, Jinchao Xu |
NAACL-HLT | 3 |
| 2024 | Alignment at Pre-training! Towards Native Alignment for Arabic LLMsabstractThe alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction tuning or reinforcement learning stages, referred to in this paper as `\textit{post alignment}'. We argue that alignment during the pre-training phase, which we term 'native alignment', warrants investigation. Native alignment aims to prevent unaligned content from the beginning, rather than relying on post-hoc processing. This approach leverages extensively aligned pre-training data to enhance the effectiveness and usability of pre-trained models. Our study specifically explores the application of native alignment in the context of Arabic LLMs. We conduct comprehensive experiments and ablation studies to evaluate the impact of native alignment on model performance and alignment stability. Additionally, we release open-source Arabic LLMs that demonstrate state-of-the-art performance on various benchmarks, providing significant benefits to the Arabic LLM community. Juhao Liang, Zhenyang Cai, Jianqing Zhu, Kewei Zong, Bang An 0004, Mosen Alharthi, Juncai He 0001, Haizhou Li 0001, Benyou Wang, Jinchao Xu |
NeurIPS | 3 |
| 2024 | Global Instance Relation Distillation for convolutional neural network compression
Haolin Hu, Huanqiang Zeng, Yifan Shi 0001, Jianqing Zhu, Jing Chen 0001 |
Neural Comput. Appl. | 5 |
| 2024 | Unsupervised video-based action recognition using two-stream generative adversarial network
Wei Lin 0021, Huanqiang Zeng, Jianqing Zhu, Chih-Hsien Hsia, Junhui Hou, Kai-Kuang Ma |
Neural Comput. Appl. | 3 |
| 2024 | Cross-modal group-relation optimization for visible-infrared person re-identification
Jianqing Zhu, Hanxiao Wu, Yuqing Fu, Huanqiang Zeng, Liu Liu 0014, Zhen Lei 0001 |
Neural Networks | 1 |
| 2024 | Pairwise difference relational distillation for object re-identification
Hanxiao Wu, Yihong Lin, Jianqing Zhu, Huanqiang Zeng |
Pattern Recognit. | 4 |
| 2024 | Distillation embedded absorbable pruning for fast object re-identification
Hanxiao Wu, Jianqing Zhu, Huanqiang Zeng |
Pattern Recognit. | 3 |
| 2024 | Modality-Consistent Attention for Visible-Infrared Vehicle Re-IdentificationabstractVisible-infrared vehicle re-identification (VIVR) seeks to match vehicle images of the same identity taken by cameras of different modalities. The noticeable disparity between visible and infrared modalities leads to attention deviations, causing deep models to incorrectly focus on different local regions of vehicles in visible and infrared images. We observed that the spatial distributions of distinguishing local regions, such as logos, front windows, and wheels, exhibit similarity in average images obtained from both visible and infrared images. Based on this, we propose a modality-consistent attention (MCA) approach for VIVR. Unlike image-level attention, our MCA is identity-level attention that holistically emphasizes the distinguishing regions of a vehicle identity across multiple images captured from various viewpoints. Furthermore, we constrain the differences between the identity-level spatial attention masks resulting from visible and infrared modalities. This approach helps deep networks focus consistently on learning the distinguishing local characteristics of vehicles across different modalities and viewpoints. Our experiments on RGBN300 and MSVR310 datasets demonstrate that our approach achieves state-of-the-art performance. Jiajun Su, Jianqing Zhu, Liu Liu 0014, Huanqiang Zeng |
IEEE Signal Process. Lett. | 3 |
| 2024 | A Benchmark for Vehicle Re-Identification in Mixed Visible and Infrared DomainsabstractWe propose a new benchmark for vehicle re-identification in mixed visible and infrared domains. Unlike cross-modal vehicle re-identification, we focus on a more realistic scenario, namely, mixed-modal vehicle re-identification, in which both probe and gallery sets contain visible and infrared images. We provide auto-cropped visible and infrared images, simulating data from actual surveillance systems. We design a mixed triplet loss function for model training. We report both mixed-modal and cross-modal retrieval performance. Our mixed method performs well in cross-modal retrieval, e.g., the Rank-1 identification rate is 79.19% in the visible-to-infrared retrieval mode. Simin Zhan, Jianqing Zhu, Huanqiang Zeng |
IEEE Signal Process. Lett. | 4 |
| 2024 | Collaborative Knowledge DistillationabstractExisting research on knowledge distillation has primarily concentrated on the task of facilitating student networks in acquiring the complete knowledge imparted by teacher networks. However, recent studies have shown that good networks are not suitable for acting as teachers, and there is a positive correlation between distillation performance and teacher prediction uncertainty. To address this finding, this paper thoroughly analyzes in depth the reasons why the teacher network affects the distillation performance, gives full play to the participation of the student network in the process of knowledge distillation, and assists the teacher network in distilling the knowledge that is suitable for their learning. In light of this premise, a novel approach known as Collaborative Knowledge Distillation (CKD) is introduced, which is founded upon the concept of "Tailoring the Teaching to the Individual". Compared with Baseline, this paper’s method improves students’ accuracy by an average of 3.42% in CIFAR-100 experiments, and by an average of 1.71% compared with the classical Knowledge Distillation (KD) method. The ImageNet experiments conducted revealed a significant improvement of 2.04% in the Top-1 accuracy of the students. Junhuang Wang, Jianqing Zhu, Huanqiang Zeng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Expanding and Refining Hybrid Compressors for Efficient Object Re-IdentificationabstractRecent object re-identification (Re-ID) methods gain high efficiency via lightweight student models trained by knowledge distillation (KD). However, the huge architectural difference between lightweight students and heavy teachers causes students to have difficulties in receiving and understanding teachers' knowledge, thus losing certain accuracy. To this end, we propose a refiner-expander-refiner (RER) structure to enlarge a student's representational capacity and prune the student's complexity. The expander is a multi-branch convolutional layer to expand the student's representational capacity to understand a teacher's knowledge comprehensively, which does not require any feature-dimensional adapter to avoid knowledge distortions. The two refiners are 1×1 convolutional layers to prune the input and output channels of the expander. In addition, in order to alleviate the competition accuracy-related and pruning-related gradients, we design a common consensus gradient resetting (CCGR) method, which discards unimportant channels according to the intersection of each sample's unimportant channel judgment. Finally, the trained RER can be simplified into a slim convolutional layer via re-parameterization to speed up inference. As a result, we propose an expanding and refining hybrid compressing (ERHC) method. Extensive experiments show that our ERHC has superior inference speed and accuracy, e.g., on the VeRi-776 dataset, given the ResNet101 as a teacher, ERHC saves 75.33% model parameters (MP) and 74.29% floating-point of operations (FLOPs) without sacrificing accuracy. Hanxiao Wu, Jianqing Zhu, Huanqiang Zeng, Jing Zhang 0037 |
IEEE Trans. Image Process. | 3 |
| 2023 | Towards a Smaller Student: Capacity Dynamic Distillation for Efficient Image RetrievalabstractPrevious Knowledge Distillation based efficient image retrieval methods employ a lightweight network as the stu-dent model for fast inference. However, the lightweight stu-dent model lacks adequate representation capacity for effective knowledge imitation during the most critical early training period, causing final performance degeneration. To tackle this issue, we propose a Capacity Dynamic Distillation framework, which constructs a student model with editable representation capacity. Specifically, the employed student model is initially a heavy model to fruitfully learn distilled knowledge in the early training epochs, and the stu-dent model is gradually compressed during the training. To dynamically adjust the model capacity, our dynamic frame-work inserts a learnable convolutional layer within each residual block in the student model as the channel importance indicator. The indicator is optimized simultaneously by the image retrieval loss and the compression loss, and a retrieval- guided gradient resetting mechanism is proposed to release the gradient conflict. Extensive experiments show that our method has superior inference speed and accu-racy, e.g., on the VeRi-776 dataset, given the ResNet101 as a teacher, our method saves 67.13% model parameters and 65.67% FLOPs without sacrificing accuracy. Code is avail-able at https://github.com/SCY-X/Capacity_Dynamic_Distillation. Huaidong Zhang, Xuemiao Xu, Jianqing Zhu, Shengfeng He |
CVPR | 4 |
| 2023 | Learning Correction Filter via Degradation-Adaptive Regression for Blind Single Image Super-ResolutionabstractAlthough existing image deep learning super-resolution (SR) methods achieve promising performance on benchmark datasets, they still suffer from severe performance drops when the degradation of the low-resolution (LR) input is not covered in training. To address the problem, we propose an innovative unsupervised method of Learning Correction Filter via Degradation-Adaptive Regression for Blind Single Image Super-Resolution. Highly inspired by the generalized sampling theory, our method aims to enhance the strength of off-the-shelf SR methods trained on known degradations and adapt to unknown complex degradations to generate improved results. Specifically, we first conduct degradation estimation for each local image region by learning the internal distribution in an unsupervised manner via GAN. Instead of assuming degradation are spatially invariant across the whole image, we learn correction filters to adjust degradations to known degradations in a spatially variant way by a novel linearly-assembled pixel degradation-adaptive regression module (DARM). DARM is lightweight and easy to optimize on a dictionary of multiple pre-defined filter bases. Extensive experiments on synthetic and real-world datasets verify the effectiveness of our method both qualitatively and quantitatively. Code can be available at: https://github.com/edbca/DARSR. Xiaobin Zhu 0001, Jianqing Zhu, Shi-Xue Zhang, Jingyan Qin, Xu-Cheng Yin |
ICCV | 3 |
| 2023 | Attribute-Image Person Re-identification via Modal-Consistent Metric Learning
Jianqing Zhu, Liu Liu 0014, Yibing Zhan, Xiaobin Zhu 0001, Huanqiang Zeng, Dacheng Tao |
Int. J. Comput. Vis. | 1 |
| 2023 | Visible-infrared person re-identification using high utilization mismatch amending triplet loss
Jianqing Zhu, Hanxiao Wu, Huanqiang Zeng, Xiaobin Zhu 0001, Jingchang Huang, Canhui Cai |
Image Vis. Comput. | 1 |
| 2023 | Multi-receptive field attention for person re-identification
Zhixiong Wu, Jianqing Zhu |
Multim. Tools Appl. | 2 |
| 2023 | An interpretive constrained linear model for ResNet and MgNet
Juncai He 0001, Jinchao Xu, Jianqing Zhu |
Neural Networks | 4 |
| 2023 | RASP: Regularization-based Amplitude Saliency Pruning
Chenghui Zhen, Jian Mo, Ming Ji, Jianqing Zhu |
Neural Networks | 6 |
| 2023 | DeflickerCycleGAN: Learning to Detect and Remove Flickers in a Single ImageabstractEliminating the flickers in digital images captured by rolling shutter cameras is a fundamental and important task in computer vision applications. The flickering effect in a single image stems from the mechanism of asynchronous exposure of rolling shutters employed by cameras equipped with CMOS sensors. In an artificial lighting environment, the light intensity captured at different time intervals varies due to the fluctuation of the power grid, ultimately resulting in the flickering artifact in the image. Up to date, there are few studies related to single image deflickering. Further, it is even more challenging to remove flickers without a priori information, e.g., camera parameters or paired images. To address these challenges, we propose an unsupervised framework termed DeflickerCycleGAN, which is trained on unpaired images for end-to-end single image deflickering. Besides the cycle-consistency loss to maintain the similarity of image contents, we meticulously design another two novel loss functions, i.e., gradient loss and flicker loss, to reduce the risk of edge blurring and color distortion. Moreover, we provide a strategy to determine whether an image contains flickers or not without extra training, which leverages an ensemble methodology based on the output of two previously trained markovian discriminators. Extensive experiments on both synthetic and real datasets show that our proposed DeflickerCycleGAN not only achieves excellent performance on flicker removal in a single image but also shows high accuracy and competitive generalization ability on flicker detection, compared to that of a well-trained classifier based on ResNet50. Xiaodan Lin, Yangfu Li, Jianqing Zhu, Huanqiang Zeng |
IEEE Trans. Image Process. | 3 |
| 2023 | GiT: Graph Interactive Transformer for Vehicle Re-IdentificationabstractTransformers are more and more popular in computer vision, which treat an image as a sequence of patches and learn robust global features from the sequence. However, pure transformers are not entirely suitable for vehicle re-identification because vehicle re-identification requires both robust global features and discriminative local features. For that, a graph interactive transformer (GiT) is proposed in this paper. In the macro view, a list of GiT blocks are stacked to build a vehicle re-identification model, in where graphs are to extract discriminative local features within patches and transformers are to extract robust global features among patches. In the micro view, graphs and transformers are in an interactive status, bringing effective cooperation between local and global features. Specifically, one current graph is embedded after the former level's graph and transformer, while the current transform is embedded after the current graph and the former level's transformer. In addition to the interaction between graphs and transforms, the graph is a newly-designed local correction graph, which learns discriminative local features within a patch by exploring nodes' relationships. Extensive experiments on three large-scale vehicle re-identification datasets demonstrate that our GiT method is superior to state-of-the-art vehicle re-identification approaches. Fei Shen 0004, Jianqing Zhu, Xiaobin Zhu 0001, Huanqiang Zeng |
IEEE Trans. Image Process. | 3 |
| 2023 | Deep Cross-Modal Hashing Based on Semantic Consistent RankingabstractThe amount of multi-modal data available on the Internet is enormous. Cross-modal hash retrieval maps heterogeneous cross-modal data into a single Hamming space to offer fast and flexible retrieval services. However, existing cross-modal methods mainly rely on the feature-level similarity between multi-modal data and ignore the relationship between relative rankings and label-level fine-grained similarity of neighboring instances. To overcome these issues, we propose a novelDeepCross-modalHashing based onSemanticConsistentRanking (DCH-SCR) that comprehensively investigates the intra-modal semantic similarity relationship. Firstly, to the best of our knowledge, it is an early attempt to preserve semantic similarity for cross-modal hashing retrieval by combining label-level and feature-level information. Secondly, the inherent gap between modalities is narrowed by developing a ranking alignment loss function. Thirdly, the compact and efficient hash codes are optimized based on the common semantic space. Finally, we use the gradient to specify the optimization direction and introduce the Normalized Discounted Cumulative Gain (NDCG) to achieve varying optimization strengths for data pairs with different similarities. Extensive experiments on three real-world image-text retrieval datasets demonstrate the superiority of DCH-SCR over several state-of-the-art cross-modal retrieval methods. Huanqiang Zeng, Yifan Shi 0001, Jianqing Zhu, Chih-Hsien Hsia, Kai-Kuang Ma |
IEEE Trans. Multim. | 4 |
| 2022 | Deep Rank Cross-Modal Hashing with Semantic Consistent for Image-Text RetrievalabstractCross-modal hashing retrieval approaches maps heterogeneous multi-modal data into a common hamming space to achieve efficient and flexible retrieval performance. However, existing cross-modal methods mainly exploit feature-level similarity between multi-modal data, the label-level similarity and relative ranking relationship between adjacent instances have been ignored. To address these problems, we propose a novel Deep Rank Cross-modal Hashing(DRCH) method that fully explores the intra-modal semantic similarity relationship. Firstly, DRCH preserves semantic similarity by combining both label-level and feature-level information. Secondly, the inherent gap between modalities are narrowed by proposing a ranking alignment loss function. Finally, the compact and efficient hash codes are optimized from the common semantic space. Extensive experiments on two real-world image-text retrieval datasets demonstrate the superiority of DRCH compared with several state-of-the-art(SOTA) methods. Huanqiang Zeng, Yifan Shi 0001, Jianqing Zhu, Kai-Kuang Ma |
ICASSP | 4 |
| 2022 | A sample-proxy dual triplet loss function for object re-identificationabstractAbstract Object re‐identification, such as vehicle re‐identification or pedestrian re‐identification, plays a significant role in intelligent video surveillance systems for public security. Due to viewpoint variations and appearance changes, both pedestrians and vehicles usually have complex intra‐class variations. However, most existing object re‐identification methods often use a sample‐level triplet loss function cooperating with a single‐proxy softmax loss function, which could not handle complex intra‐class variations well. In this paper, a sample‐proxy dual triplet (SPDT) loss function is proposed, which works with a multi‐proxy softmax (MPS) loss function. The MPS loss function is in charge of learning multiple proxies to represent a class. The SPDT loss function is responsible for enlarging inter‐class distances as well as shrinking intra‐class distances on both sample and proxy levels. Therefore, the method not only handles multi‐proxy intra‐class variations but also fully learns discrimination on samples and proxies. Experiments on two large datasets, that is, VeRi776 and DukeMTMC‐reID, demonstrate that the method is superior to state‐of‐the‐art object re‐identification approaches. Hanxiao Wu, Fei Shen 0004, Jianqing Zhu, Huanqiang Zeng, Xiaobin Zhu 0001, Zhen Lei 0001 |
IET Image Process. | 3 |
| 2022 | An Efficient Multiresolution Network for Vehicle ReidentificationabstractIn general, vehicle images have varying resolutions due to vehicles’ movements and different camera settings. However, most existing vehicle reidentification models are single-resolution deep networks trained with preuniformly resizing vehicle images, which underestimate adverse effects of varying resolutions and lead to unsatisfactory performance. A straightforward solution for dealing with varying resolutions is to train multiple vehicle reidentification models. Each model is independently trained with images of a specific resolution. However, this straightforward solution requires significant overhead and ignores intrinsic associations among different resolution images. For that, an efficient multiresolution network (EMRN) is proposed for vehicle reidentification in this article. First, EMRN embeds a newly designed multiresolution feature dimension uniform module (MR-FDUM) behind a traditional backbone network (i.e., ResNet-50). As a result, the whole model can extract fixed dimensional features from different resolution images so that it can be trained with one loss function of fixed dimensional parameters rather than training multiple models. Second, a multiresolution image randomly feeding strategy is designed to train EMRN, making each minibatch data of a random resolution during the training process. Consequently, EMRN can implicitly learn collaborative multiresolution features via only a unitary deep network. The experiments on three large-scale data sets, i.e., VeRi776, VehicleID, and VRIC, demonstrate that EMRN is superior to state-of-the-art vehicle reidentification methods. Fei Shen 0004, Jianqing Zhu, Xiaobin Zhu 0001, Jingchang Huang, Huanqiang Zeng, Zhen Lei 0001, Canhui Cai |
IEEE Internet Things J. | 2 |
| 2022 | Deep-Learning-Enabled Automatic Optical Inspection for Module-Level Defects in LCDabstractLiquid crystal display (LCD) defects detection on module level is increasingly important for flat-panel displays (FPD) industry to increase the production capacity via machine vision technology. However, it is an overwhelmingly challenging issue due to various difficulties. This article discloses a practical automatic optical inspection (AOI) system consisting of hardware structure and software algorithm to detect module-level defects. The AOI system is the core component to build a distributed integrated inspection system with the help of the Internet of Things (IoT). Starting from the analysis of the challenges encountered in module-level defects inspection, a delicate photograph scheme is proposed to reveal different kinds of defects. In order to robustly work on the module-level defects detection with complex situations, a novel framework based on YOLOV3 detection unit is proposed in this article, including the preprocessing module, detection module, defects definition module, and interferences elimination module. To the best of our knowledge, this is the first work that designs a practical AOI system for module-level defects detection. In order to demonstrate the effectiveness of the proposed method, extensive experiments have been conducted on the manufacturing lines. The evaluation of the detection performance of the AOI system in comparison with a manual scheme indicates that the proposed system is practical for module-level defects detection. Currently, the proposed system has been deployed in a real-world LCD manufacturing line from a major player in the world. Haidi Zhu, Jingchang Huang, Qianwei Zhou, Jianqing Zhu, Baoqing Li |
IEEE Internet Things J. | 5 |
| 2022 | A Spatial and Geometry Feature-Based Quality Assessment Model for the Light Field ImagesabstractThis paper proposes a new full-reference image quality assessment (IQA) model for performing perceptual quality evaluation on light field (LF) images, called the spatial and geometry feature-based model (SGFM). Considering that the LF image describe both spatial and geometry information of the scene, the spatial features are extracted over the sub-aperture images (SAIs) by using contourlet transform and then exploited to reflect the spatial quality degradation of the LF images, while the geometry features are extracted across the adjacent SAIs based on 3D-Gabor filter and then explored to describe the viewing consistency loss of the LF images. These schemes are motivated and designed based on the fact that the human eyes are more interested in the scale, direction, contour from the spatial perspective and viewing angle variations from the geometry perspective. These operations are applied to the reference and distorted LF images independently. The degree of similarity can be computed based on the above-measured quantities for jointly arriving at the final IQA score of the distorted LF image. Experimental results on three commonly-used LF IQA datasets show that the proposed SGFM is more in line with the quality assessment of the LF images perceived by the human visual system (HVS), compared with multiple classical and state-of-the-art IQA models. Hailiang Huang 0002, Huanqiang Zeng, Junhui Hou, Jing Chen 0001, Jianqing Zhu, Kai-Kuang Ma |
IEEE Trans. Image Process. | 5 |
| 2022 | Exploring Spatial Significance via Hybrid Pyramidal Graph Network for Vehicle Re-IdentificationabstractExisting vehicle re-identification methods commonly use spatial pooling operations to aggregate feature maps extracted via off-the-shelf backbone networks, such as visual geometry group network (VGGNet), Google network (GoogLeNet) and residual network (ResNet). They ignore exploring the spatial significance of feature maps, eventually degrading the vehicle re-identification performance. In this paper, firstly, an innovative spatial graph network (SGN) is proposed to elaborately explore the spatial significance of feature maps. The SGN stacks multiple spatial graphs (SGs). Each SG assigns feature map’s elements as nodes and utilizes spatial neighborhood relationships to determine edges among nodes. During the SGN’s propagation, each node and its spatial neighbors on an SG are aggregated to the next SG. On the next SG, each aggregated node is re-weighted with a learnable parameter to find the significance at the corresponding location. Secondly, a novel pyramidal graph network (PGN) is designed to comprehensively explore the spatial significance of feature maps at multiple scales. The PGN organizes multiple SGNs in a pyramidal manner and makes each SGN handles feature maps of a specific scale. Finally, a hybrid pyramidal graph network (HPGN) is developed by embedding the PGN behind a ResNet-50 based backbone network. Extensive experiments on three large scale vehicle databases (i.e., VeRi776, VehicleID, and VeRi-Wild) demonstrate that the proposed HPGN is superior to state-of-the-art vehicle re-identification approaches in terms of accuracy, parameter cost, and computation cost. In addition, experiments show that the proposed PGN is universal to various backbone networks. Fei Shen 0004, Jianqing Zhu, Xiaobin Zhu 0001, Jingchang Huang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Object Re-identification Using Teacher-Like and Light Students
Hanxiao Wu, Fei Shen 0004, Jianqing Zhu, Huanqiang Zeng |
BMVC | 4 |
| 2021 | A spatial structural similarity triplet loss for auxiliary vehicle re-identification
Jianqing Zhu, Liu Liu 0014, Xiaobin Zhu 0001, Huanqiang Zeng |
Sci. China Inf. Sci. | 1 |
| 2021 | Multiple improved residual networks for medical image super-resolution
Defu Qiu, Lixin Zheng, Jianqing Zhu, Detian Huang |
Future Gener. Comput. Syst. | 3 |
| 2021 | Cascading Scene and Viewpoint Feature Learning for Pedestrian Gender RecognitionabstractPedestrian gender recognition plays an important role in smart city. To effectively improve the pedestrian gender recognition performance, a new method, called cascading scene and viewpoint feature learning (CSVFL), is proposed in this article. The novelty of the proposed CSVFL lies on the joint consideration of two crucial challenges in pedestrian gender recognition, namely, scene and viewpoint variation. For that, the proposed CSVFL starts with the scene transfer (ST) scheme, followed by the viewpoint adaptation (VA) scheme in a cascading manner. Specifically, the ST scheme exploits the key pedestrian segmentation network to extract the key pedestrian masks for the subsequent key pedestrian transfer generative adversarial network, with the goal of encouraging the input pedestrian image to have the similar style to the target scene while preserving the image details of the key pedestrian as much as possible. Afterward, the obtained scene-transferred pedestrian images are fed to train the deep feature learning network with the VA scheme, in which each neuron will be enabled/disabled for different viewpoints depending on whether it has contribution on the corresponding viewpoint. Extensive experiments conducted on the commonly used pedestrian attribute data sets have demonstrated that the proposed CSVFL approach outperforms multiple recently reported pedestrian gender recognition methods. Huanqiang Zeng, Jianqing Zhu, Jiuwen Cao, Yongtao Wang, Kai-Kuang Ma |
IEEE Internet Things J. | 3 |
| 2021 | A Light Field Image Quality Assessment Model Based on Symmetry and Depth FeaturesabstractThis paper presents a new full-reference image quality assessment (IQA) method for conducting the perceptual quality evaluation of the light field (LF) images, called the symmetry and depth feature-based model (SDFM). Specifically, the radial symmetry transform is first employed on the luminance components of the reference and distorted LF images to extract their symmetry features for capturing the spatial quality of each view of an LF image. Second, the depth feature extraction scheme is designed to explore the geometry information inherited in an LF image for modeling its LF structural consistency across views. The similarity measurements are subsequently conducted on the comparison of their symmetry and depth features separately, which are further combined to achieve the quality score for the distorted LF image. Note that the proposed SDFM that explores the symmetry and depth features is conformable to the human vision system, which identifies the objects by sensing their structures and geometries. Extensive simulation results on the dense light fields dataset have clearly shown that the proposed SDFM outperforms multiple classical and recently developed IQA algorithms on quality evaluation of the LF images. Huanqiang Zeng, Junhui Hou, Jing Chen 0001, Jianqing Zhu, Kai-Kuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Learning Matching Behavior Differences for Compressing Vehicle Re-identification ModelsabstractVehicle re-identification matching vehicles captured by different cameras has great potential in the field of public security. However, recent vehicle re-identification approaches exploit complex networks, causing large computations in their testing phases. In this paper, we propose a matching behavior difference learning (MBDL) method to compress vehicle re-identification models for saving testing computations. In order to represent the matching behavior evolution across two different layers of a deep network, a matching behavior difference (MBD) matrix is designed. Then, our MBDL method minimizes the L1 loss function among MBD matrixes from a small student network and a complex teacher network, ensuring the student network use less computations to simulate the teacher network's matching behaviors. During the testing phase, only the small student network is utilized so that testing computations can be significantly reduced. Experiments on VeRi776 and VehicleID datasets show that MBDL outperforms many state-of-the-art approaches in terms of accuracy and testing time performance. Jianqing Zhu, Huanqiang Zeng, Canhui Cai, Lixin Zheng |
VCIP | 2 |
| 2020 | Object Reidentification via Joint Quadruple Decorrelation Directional Deep Networks in Smart TransportationabstractObject reidentification with the goal of matching pedestrian or vehicle images captured from different camera viewpoints is of considerable significance to public security. Quadruple directional deep learning features (QD-DLFs) can comprehensively describe object images. However, the correlation among QD-DLFs is an unavoidable problem, since QD-DLFs are learned with quadruple independent directional deep networks (QIDDNs) driven with the same training data, and each network holds the same basic deep feature learning architecture (BDFLA). The correlation among QD-DLFs is harmful to the complementarity of QD-DLFs, restricting the object reidentification performance. For that, we propose joint quadruple decorrelation directional deep networks (JQD3Ns) to reduce the correlation among the learned QD-DLFs. In order to jointly train JQD3Ns, besides the softmax loss functions, a parameter correlation cost function is proposed to indirectly reduce the correlation among QD-DLFs by enlarging the dissimilarity among the parameters of JQD3Ns. Extensive experiments on three publicly available large-scale data sets demonstrate that the proposed JQD3Ns approach is superior to multiple state-of-the-art object reidentification methods. Jianqing Zhu, Jingchang Huang, Huanqiang Zeng, Xiaoqing Ye, Baoqing Li, Zhen Lei 0001, Lixin Zheng |
IEEE Internet Things J. | 1 |
| 2020 | Body Symmetry and Part-Locality-Guided Direct Nonparametric Deep Feature Enhancement for Person ReidentificationabstractIn recent years, deep learning (DL) has been successfully and widely applied in the person reidentification (Re-ID). However, the DL-based person Re-ID methods face a bottleneck that the scales of most existing person Re-ID databases are not large enough for training very deep models. To address this problem, a body symmetry and part-locality-guided direct nonparametric deep feature enhancement (DNDFE) method is proposed in this article. Based on the observation that the body symmetry and part locality are two important appearance properties inherited in the upright walking persons, the proposed method designs two nonparametric layers, namely, the body symmetry average pooling and local normalization layers, to construct a DNDFE module to well explore the body symmetry and part locality properties. The proposed DNDFE module could be directly embedded between the traditional deep feature learning module and similarity learning module to enhance the DL features so as to improve the person Re-ID performance. The experimental results have shown that the proposed DNDFE method is superior to multiple state-of-the-art person Re-ID methods in terms of accuracy and efficiency. Jianqing Zhu, Huanqiang Zeng, Jingchang Huang, Xiaobin Zhu 0001, Zhen Lei 0001, Canhui Cai, Lixin Zheng |
IEEE Internet Things J. | 1 |
| 2020 | Joint Pyramid Feature Representation Network for Vehicle Re-identification
Xiangwei Lin, Huanqiang Zeng, Jinhui Hou, Jiuwen Cao, Jianqing Zhu, Jing Chen 0001 |
Mob. Networks Appl. | 5 |
| 2020 | Subband Aware CNN for Cell-Phone RecognitionabstractIdentifying the model of a cell-phone with which an audio recording is made serves an important forensic purpose. In this paper, we propose a channel attention mechanism based on subband awareness to focus on the most relevant parts of the frequency bands to produce more efficient feature representations for cell-phone recognition task. A multi-stream network is introduced to fully exploit difference among frequency bands that provide critical clues to identify the fingerprints from the built-in microphones, and thus is able to recognize cell-phones from different manufacturers and even different models from the same manufacturer. The effectiveness of the proposed design is validated by a collection of speaker-independent audio recordings from 20 models of cell-phones made by 5 major manufacturers. In particular, the fusion of data augmentation and attention strategy greatly increases the robustness of the scheme when additive noise is present in the recorded audios. Finally, the salient regions infer the critical subbands for the recognition of different cell-phone models. Xiaodan Lin, Jianqing Zhu, Donghua Chen |
IEEE Signal Process. Lett. | 2 |
| 2020 | Screen Content Video Quality Assessment: Subjective and Objective StudyabstractIn this paper, we make the first attempt to study the subjective and objective quality assessment for the screen content videos (SCVs). For that, we construct the first large-scale video quality assessment (VQA) database specifically for the SCVs, called the screen content video database (SCVD). This SCVD provides 16 reference SCVs, 800 distorted SCVs, and their corresponding subjective scores, and it is made publicly available for research usage. The distorted SCVs are generated from each reference SCV with 10 distortion types and 5 degradation levels for each distortion type. Each distorted SCV is rated by at least 32 subjects in the subjective test. Furthermore, we propose the first full-reference VQA model for the SCVs, called the spatiotemporal Gabor feature tensor-based model (SGFTM), to objectively evaluate the perceptual quality of the distorted SCVs. This is motivated by the observation that 3D-Gabor filter can well stimulate the visual functions of the human visual system (HVS) on perceiving videos, being more sensitive to the edge and motion information that are often-encountered in the SCVs. Specifically, the proposed SGFTM exploits 3D-Gabor filter to individually extract the spatiotemporal Gabor feature tensors from the reference and distorted SCVs, followed by measuring their similarities and later combining them together through the developed spatiotemporal feature tensor pooling strategy to obtain the final SGFTM score. Experimental results on SCVD have shown that the proposed SGFTM yields a high consistency on the subjective perception of SCV quality and consistently outperforms multiple classical and state-of-the-art image/video quality assessment models. Shan Cheng, Huanqiang Zeng, Jing Chen 0001, Junhui Hou, Jianqing Zhu, Kai-Kuang Ma |
IEEE Trans. Image Process. | 5 |
| 2020 | Vehicle Re-Identification Using Quadruple Directional Deep Learning FeaturesabstractIn order to resist the adverse effect of viewpoint variations, we design quadruple directional deep learning networks to extract quadruple directional deep learning features (QD-DLF) of vehicle images for improving vehicle re-identification performance. The quadruple directional deep learning networks are of similar overall architecture, including the same basic deep learning architecture but different directional feature pooling layers. Specifically, the same basic deep learning architecture that is a shortly and densely connected convolutional neural network is utilized to extract the basic feature maps of an input square vehicle image in the first stage. Then, the quadruple directional deep learning networks utilize different directional pooling layers, i.e., horizontal average pooling layer, vertical average pooling layer, diagonal average pooling layer, and anti-diagonal average pooling layer, to compress the basic feature maps into horizontal, vertical, diagonal, and anti-diagonal directional feature maps, respectively. Finally, these directional feature maps are spatially normalized and concatenated together as a quadruple directional deep learning feature for vehicle re-identification. The extensive experiments on both VeRi and VehicleID databases show that the proposed QD-DLF approach outperforms multiple state-of-the-art vehicle re-identification methods. Jianqing Zhu, Huanqiang Zeng, Jingchang Huang, Shengcai Liao, Zhen Lei 0001, Canhui Cai, Lixin Zheng |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2019 | Joint horizontal and vertical deep learning feature for vehicle re-identification
Jianqing Zhu, Huanqiang Zeng, Yongzhao Du, Lixin Zheng, Canhui Cai |
Sci. China Inf. Sci. | 1 |
| 2019 | Multi-label learning with multi-label smoothing regularization for vehicle re-identification
Jinhui Hou, Huanqiang Zeng, Jianqing Zhu, Jing Chen 0001, Kai-Kuang Ma |
Neurocomputing | 4 |
| 2019 | VRSDNet: vehicle re-identification with a shortly and densely connected convolutional neural network
Jianqing Zhu, Yongzhao Du, Lixin Zheng, Canhui Cai |
Multim. Tools Appl. | 1 |
| 2018 | A Shortly and Densely Connected Convolutional Neural Network for Vehicle Re-identificationabstractIn this paper, we propose a shortly and densely connected convolutional neural network (SDC-CNN) for vehicle re-identification. The proposed SDC-CNN mainly consists of short and dense units (SDUs), necessary pooling and normalization layers. The main contribution lies at the design of short and dense connection mechanism, which would effectively improve the feature learning ability. Specifically, in the proposed short and dense connection mechanism, each SDU contains a short list of densely connected convolutional layers and each convolutional layer is of the same appropriate channels. Consequently, the number of connections and the input channel of each convolutional layer are limited in each SDU, and the architecture of SDC-CNN is simple. Extensive experiments on both VeRi and VehicleID datasets show that the proposed SDC-CNN is obviously superior to multiple state-of-the-art vehicle re-identification methods. Jianqing Zhu, Huanqiang Zeng, Zhen Lei 0001, Shengcai Liao, Lixin Zheng, Canhui Cai |
ICPR | 1 |
| 2018 | A multi-order derivative feature-based quality assessment model for light field image
Huanqiang Zeng, Jing Chen 0001, Jianqing Zhu, Kai-Kuang Ma |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | A multi-scale contrast-based image quality assessment model for multi-exposure image fusion
Huanqiang Zeng, Jing Chen 0001, Jianqing Zhu, Junhui Hou |
Signal Process. | 5 |
| 2018 | Screen Content Image Quality Assessment Using Multi-Scale Difference of GaussianabstractIn this paper, a novel image quality assessment (IQA) model for the screen content images (SCIs) is proposed by using multi-scale difference of Gaussian (MDOG). Motivated by the observation that the human visual system (HVS) is sensitive to the edges while the image details can be better explored in different scales, the proposed model exploits MDOG to effectively characterize the edge information of the reference and distorted SCIs at two different scales, respectively. Then, the degree of edge similarity is measured in terms of the smaller-scale edge map. Finally, the edge strength computed based on the larger-scale edge map is used as the weighting factor to generate the final SCI quality score. Experimental results have shown that the proposed IQA model for the SCIs produces high consistency with human perception of the SCI quality and outperforms the state-of-the-art quality models. Ying Fu 0004, Huanqiang Zeng, Lin Ma 0002, Zhangkai Ni, Jianqing Zhu, Kai-Kuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Deep Hybrid Similarity Learning for Person Re-IdentificationabstractPerson re-identification (Re-ID) aims to match person images captured from two non-overlapping cameras. In this paper, a deep hybrid similarity learning (DHSL) method for person Re-ID based on a convolution neural network (CNN) is proposed. In our approach, a light CNN learning feature pair for the input image pair is simultaneously extracted. Then, both the elementwise absolute difference and multiplication of the CNN learning feature pair are calculated. Finally, a hybrid similarity function is designed to measure the similarity between the feature pair, which is realized by learning a group of weight coefficients to project the elementwise absolute difference and multiplication into a similarity score. Consequently, the proposed DHSL method is able to reasonably assign complexities of feature learning and metric learning in a CNN, so that the performance of person Re-ID is improved. Experiments on three challenging person Re-ID databases, QMUL GRID, VIPeR, and CUHK03, illustrate that the proposed DHSL method is superior to multiple state-of-the-art person Re-ID methods. Jianqing Zhu, Huanqiang Zeng, Shengcai Liao, Zhen Lei 0001, Canhui Cai, Lixin Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Multi-label convolutional neural network based pedestrian attribute classification
Jianqing Zhu, Shengcai Liao, Zhen Lei 0001, Stan Z. Li |
Image Vis. Comput. | 1 |
| 2016 | Low complexity depth intra coding in 3D-HEVC based on depth classificationabstractThe latest high efficiency video coding-based three dimensional video coding (3D-HEVC) exploits sophisticated intra prediction scheme to improve the coding performance of the depth video, but incurring heavy computational complexity. To address this problem, a low complexity depth intra coding method is presented for 3D-HEVC based on depth classification. Firstly, a database of depth prediction units (PUs) with three kinds of complexities is collected based on their optimal intra prediction mode. Then, the histogram of oriented gradient (HOG) features are extracted on these established database to train the classifier using support vector machine (SVM). For the current depth PU, the trained classifier is applied to determine its most possible complexity class so as to select the corresponding modes for involving the mode decision process. Experimental results show that the proposed method is able to significantly reduce the computational complexity while keeping almost the same coding performance of depth video and video quality of the synthesized view, compared with the exhaustive mode decision in 3D-HEVC. Huijie Zheng, Jianqing Zhu, Huanqiang Zeng, Jing Chen 0001, Canhui Cai, Kai-Kuang Ma |
VCIP | 2 |
| 2016 | Multicamera Joint Video SynopsisabstractDue to an increasing demand for video surveillance, there is an explosive growth of surveillance videos, which causes a big challenge in video storage, browsing, and retrieval. The video synopsis technique is thus developed to extract and rearrange the moving objects so as to handle the massive video browsing challenge. However, the traditional video synopsis (TVS) method only considers the processing videos captured by a single camera, ignoring object interactions in multicamera videos. To address this issue, we propose a novel multicamera joint video synopsis (JVS) algorithm for multicamera surveillance videos. First, a key time stamp (KTS) selection method is designed to find an object's appearing, merging, splitting, and disappearing moments in the frame sequence, called tube, of that object. Second, tubes are rearranged by minimizing a global energy function that involves the overall camera views. Compared with the energy function used in TVS, the proposed global energy function considers the chronological orders of tubes not only in the same camera view but also among different camera views. Moreover, the chronological disorder cost term is formulated based on the KTS labels, and improved by considering the visual similarity between two tubes. Finally, the multicamera synopsis videos are separately generated by stitching together the globally rearranged tubes and background images of the same camera view. Extensive experiments show that the proposed JVS method is better than the traditional single-camera video synopsis method in preserving the chronological orders of moving objects among multicamera synopsis videos. Jianqing Zhu, Shengcai Liao, Stan Z. Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | High-Performance Video Condensation SystemabstractVideo synopsis or condensation is a smart solution for fast video browsing and storage. However, most of the existing methods work offline, where two main phases are required. The first phase is to prepare tubes and background images. The second phase is to rearrange tubes and stitch them into backgrounds. However, with a long video sequence, the first phase is memory consuming for data storage, and the second phase is computationally expensive to rearrange all tubes simultaneously. To overcome these problems, we propose a high-performance video condensation system based on an online content-aware framework. The online framework transforms the optimization problem of tube rearrangement into a stepwise optimization problem. Therefore, it can condense video with much less memory and higher speed than the offline framework. With the aid of this transformation, the proposed system can process input videos and produce condensed videos simultaneously. Thus it is suitable for real-time endless surveillance videos. Meanwhile, the online mechanism allows users to directly visit the condensation video that has been generated. Moreover, the content-aware mechanism makes the proposed system able to automatically determine the duration of a condensed video. Finally, the proposed system uses Graphic Processing Unit (GPU) and multicore techniques to improve the speed. Extensive experiments that validate the high efficiency of the system are presented. Jianqing Zhu, Shikun Feng, Dong Yi, Shengcai Liao, Zhen Lei 0001, Stan Z. Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |