Jiazhong Chen

dblp:15/6726 · DBLP profile ↗
← Back
48ranked-venue papers
15as first author
25since 2021 · last 2026
0000-0003-2159-9393ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 10 first-author · 18 since 2021Artificial intelligence and machine learning · 18 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Geometry-Insensitive RPN Prototypes for Domain Adaptive 3D Object Detection
abstract
The region proposal network (RPN) plays a critical role in object detection for a two-stage domain adaptive 3D object detector. However, current methods usually minimize the disparity between source and target domains by reducing the bias in intrinsic geometric information or by undertaking feature alignment according to the geometric disparity but ignore the transferability of RPN-related features and neglect the discriminability between foreground and background, resulting in generating low-quality RPN proposals. Thus, we propose a novel domain adaptation method to distinguish the discriminability between foreground and background. It could implicitly avoid the geometric disparity of objects in feature alignment. Specifically, we first construct learnable and geometry-insensitive foreground RPN prototype and background RPN prototype. Then, we enforce the foreground RPN features and background RPN features to align with the foreground RPN prototype and background RPN prototype, respectively. By this way, the distributional discrepancy is effectively decreased and the adaptability is promoted for existing 3D detectors. We demonstrate that our approach achieves promising results compared with other domain adaptation works on multiple cross-domain detection scenarios.
Jiazhong Chen, Dakai Ren, Zian Fu, Furui Liu
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Exploring the Potential of Large Vision-Language Models for Unsupervised Text-Based Person Retrieval
abstract
The aim of text-based person retrieval is to identify pedestrians using natural language descriptions within a large-scale image gallery. Traditional methods rely heavily on manually annotated image-text pairs, which are resource-intensive to obtain. With the emergence of Large Vision-Language Models (LVLMs), the advanced capabilities of contemporary models in image understanding have led to the generation of highly accurate captions. Therefore, this paper explores the potential of employing Large Vision-Language Models for unsupervised text-based pedestrian image retrieval and proposes a Multi-grained Uncertainty Modeling and Alignment framework (MUMA). Initially, multiple Large Vision-Language Models are employed to generate diverse and hierarchically structured pedestrian descriptions across different styles and granularities. However, the generated captions inevitably introduce noise. To address this issue, an uncertainty-guided sample filtration module is proposed to estimate and filter out unreliable image-text pairs. Additionally, to simulate the diversity of styles and granularities in captions, a multi-grained uncertainty modeling approach is applied to model the distributions of captions, with each caption represented as a multivariate Gaussian distribution. Finally, a multi-level consistency distillation loss is employed to integrate and align the multi-grained captions, aiming to transfer knowledge across different granularities. Experimental evaluations conducted on three widely-used datasets demonstrate the significant advancements achieved by our approach.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Shijuan Huang, Linnan Tu, Fei Shen 0004
AAAI4
2025 AD2T: Adversarial Distortion Domain Translation for Robust Watermarking against Non-differentiable Distortions
abstract
Deep watermarking models optimize robustness by incorporating distortions between the encoder and decoder. To tackle non-differentiable distortions, current methods only train the decoder with distorted images, which breaks the joint optimization of the encoder-decoder, resulting in suboptimal performance. To address this problem, we propose an Adversarial Distortion Domain Translation (AD2T) method by treating the distortion as an image-to-image translation task. AD2T adopts conditional GANs to learn the non-differentiable distortion mappings. It employs generators to transform the encoded image into the distorted one to bridge the encoder-decoder for joint optimization. We also supervise the GANs to generate challenging distorted samples to augment the watermarking model via adversarial training. This further improves the model robustness by minimizing the maximum decoding loss. Extensive experiments demonstrate the superiority of our method when tested on non-differentiable distortions, including lossy compression and style transfers. Codes are released here: https://github.com/zcx-language/AdversarialDistortionDomainTranslation.
Chengxin Zhao, Jiazhong Chen, Han Fang 0004, Zongyi Li, Sijing Xie
ICASSP3
2024 Uncertainty-Guided Person Search Model with Auxiliary Shallow Feature Exploration
abstract
Person search is a unified system aimed at jointly localizing and identifying a person of interest from a gallery of whole scene images. Due to the inherent properties of the person search, it faces significant challenges of large-scale variations, inaccurate detection boxes, and crowded scenes. To address these issues, we proposed an uncertainty-guided framework coupled with auxiliary shallow feature exploration, which includes a shallow feature fusion module and an uncertainty-guided module. Firstly, considering the scales of the person are varied due to various scenes and their relative positions to the camera, a shallow feature fusion module is designed to extract multi-scale features to assist the re-id sub-task. Additionally, a self-distillation loss is proposed to align features across different scales. Furthermore, to alleviate the problem that the model can be easily affected by coarse samples resulting from crowded scenes and inaccurate detection boxes, we introduce an uncertainty guidance module to reduce the negative impact of these coarse targets. The experimental results demonstrate the effectiveness of our proposed methods on two benchmarks (i.e., CUHK-SYSU, and PRW).
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Ping Li 0021
ICASSP5
2024 Improving Visual Quality and Transferability of Adversarial Attacks on Face Recognition Simultaneously with Adversarial Restoration
abstract
Adversarial face examples possess two critical properties: Visual Quality and Transferability. However, existing approaches rarely address these properties simultaneously, leading to subpar results. To address this issue, we propose a novel adversarial attack technique known as Adversarial Restoration (AdvRestore), which enhances both visual quality and transferability of adversarial face examples by leveraging a face restoration prior. In our approach, we initially train a Restoration Latent Diffusion Model (RLDM) designed for face restoration. Subsequently, we employ the inference process of RLDM to generate adversarial face examples. The adversarial perturbations are applied to the intermediate features of RLDM. Additionally, by treating RLDM face restoration as a sibling task, the transferability of the generated adversarial face examples is further improved. Our experimental results validate the effectiveness of the proposed attack method.
Fengfan Zhou, Yuxuan Shi 0002, Jiazhong Chen, Ping Li 0021
ICASSP4
2024 Cross-modal Generation and Alignment via Attribute-guided Prompt for Unsupervised Text-based Person Retrieval
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Shijuan Huang
IJCAI5
2024 DBDH: A Dual-Branch Dual-Head Neural Network for Invisible Embedded Regions Localization
abstract
Embedding invisible hyperlinks or hidden codes in images to replace QR codes has become a hot topic recently. This technology requires first localizing the embedded region in the captured photos before decoding. Existing methods that train models to find the invisible embedded region struggle to obtain accurate localization results, leading to degraded decoding accuracy. This limitation is primarily because the CNN network is sensitive to low-frequency signals, while the embedded signal is typically in the high-frequency form. Based on this, this paper proposes a Dual-Branch Dual-Head (DBDH) neural network tailored for the precise localization of invisible embedded regions. Specifically, DBDH uses a low-level texture branch containing 62 high-pass filters to capture the high-frequency signals induced by embedding. A high-level context branch is used to extract discriminative features between the embedded and normal regions. DBDH employs a detection head to directly detect the four vertices of the embedding region. In addition, we introduce an extra segmentation head to segment the mask of the embedding region during training. The segmentation head provides pixel-level supervision for model learning, facilitating better learning of the embedded signals. Based on two state-of-the-art invisible offline-to-online messaging methods, we construct two datasets and augmentation strategies for training and testing localization models. Extensive experiments demonstrate the superior performance of the proposed DBDH over existing methods.
Chengxin Zhao, Sijing Xie, Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen
IJCNN7
2024 Improve Deep Hashing with Language Guidance for Unsupervised Image Retrieval
abstract
Hashing method is widely used in multimedia retrieval systems because of its outstanding retrieval efficiency and low storage cost. Most existing unsupervised hashing methods learn binary hash codes through similarity structure preserving or contrastive learning of hash codes. However, these methods usually use the visual similarity of images to guide hash learning, which does not fully utilize the high-level semantic concept information contained in images, resulting in limited retrieval performance. To tackle this problem, we propose a novel deep unsupervised hashing method called Language Guidance Hashing (LGH). Specifically, LGH utilizes a language model to mine high-level semantic concept information in images and construct a language-based similarity structure, which is used to guide hash learning. By introducing features of textual modality, higher information gain can be brought. In addition, we also propose a language-guided contrastive learning method for learning high-quality binary hash codes. Extensive experimental results show that LGH significantly outperforms state-of-the-art unsupervised hashing methods on three benchmark image datasets.
Chuang Zhao 0001, Shijie Lu, Yuxuan Shi 0002, Jiazhong Chen, Ping Li 0021
ICMR5
2024 FETR: Feature Transformer for vehicle-infrastructure cooperative 3D object detection
Wenchao Yan, Hua Cao, Jiazhong Chen
Neurocomputing3
2024 Improving the Transferability of Adversarial Attacks on Face Recognition With Beneficial Perturbation Feature Augmentation
abstract
Face recognition (FR) models can be easily fooled by adversarial examples, which are crafted by adding imperceptible perturbations on benign face images. The existence of adversarial face examples poses a great threat to the security of society. To build a more sustainable digital nation, in this article, we improve the transferability of adversarial face examples to expose more blind spots of the existing FR models. Though generating hard samples has shown its effectiveness in improving the generalization of models in training tasks, the effectiveness of using this idea to improve the transferability of adversarial face examples remains unexplored. To this end, based on the property of hard samples and the symmetry between training tasks and adversarial attack tasks, we propose the concept of hard models, which have similar effects as hard samples for adversarial attack tasks. Using the concept of hard models, we propose a novel attack method called beneficial perturbation feature augmentation attack (BPFA), which reduces the overfitting of adversarial examples to surrogate FR models by constantly generating new hard models to craft the adversarial examples. Specifically, in the backpropagation, BPFA records the gradients on preselected feature maps and uses the gradient on the input image to craft the adversarial example. In the next forward propagation, BPFA leverages the recorded gradients to add beneficial perturbations on their corresponding feature maps to increase the loss. Extensive experiments demonstrate that BPFA can significantly boost the transferability of adversarial attacks on FR.
Fengfan Zhou, Yuxuan Shi 0002, Jiazhong Chen, Zongyi Li, Ping Li 0021
IEEE Trans. Comput. Soc. Syst.4
2024 Knowledge Consistency Distillation for Weakly Supervised One Step Person Search
abstract
Weakly supervised person search targets to detect and identify a person with only bounding box annotations. Recent approaches have focused on learning person relations in a single model, ignoring the conflicts between the detection and Re-ID heads, along with the influence of background elements, which may lead to noisy pseudo labels and inaccurate Re-ID features. To address this challenge, we introduce a novel framework named Knowledge Consistency Distillation (KCD) for weakly supervised person search, which explores the capabilities of an advanced unsupervised person re-identification (Re-ID) model to mitigate the conflicts and background influences. We propose hierarchical consistency alignments, including feature-level, cluster-level, and instance-level consistency alignment, to synchronize the knowledge from the state-of-the-art unsupervised Re-ID model. Specifically, the feature-level consistency aligns the feature through both context and relation alignment. The cluster-level consistency aligns the teacher cluster information by reusing its OIM module. To tackle the inconsistency problem between student instances and teacher cluster centroids, we incorporate pseudo-label refinement to assist the student model in comprehending the teacher’s knowledge at cluster-level while mitigating the negative effects of noisy labels. Finally, an instance-level consistency loss weighted by the similarity between the instance and its corresponding cluster is proposed to align the positive instance correlations. Our approach aims to train a one-step weakly supervised model for person search by exploiting the characteristics of unsupervised person Re-ID. Extensive experiments illustrate that our method achieves state-of-the-art performance on two widely-used person search datasets, CUHK-SYSU and PRW. Our code will be available on GitHub athttps://github.com/zongyi1999/KCD.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Runsheng Wang, Chengxin Zhao, Qian Wang 0001, Shijuan Huang
IEEE Trans. Circuits Syst. Video Technol.4
2024 Viewpoint Disentangling and Generation for Unsupervised Object Re-ID
abstract
Unsupervised object Re-ID aims to learn discriminative identity features from a fully unlabeled dataset to solve the open-class re-identification problem. Satisfying results have been achieved in existing unsupervised Re-ID methods, primarily trained with pseudo-labels created by feature clustering. However, the viewpoint variation of objects is the key challenge, introducing noisy labels in the clustering process. To address this problem, a novel viewpoint disentangling and generation framework (VDG) is proposed to learn viewpoint-invariant ID features, including a disentangling and generation module, as well as a contrastive learning module. First, we design an ID encoder to map the viewpoint and identity features into the latent space. Second, a generator is used to disentangle view features and synthesize images with different orientations. Especially, the well-trained encoder serves as a pre-trained feature extractor in the contrastive learning module. Third, a viewpoint-aware loss and a class-level loss are integrated to facilitate contrastive learning between original and novel views. The generation of novel view images and the application of viewpoint-aware contrastive loss mutually assist model learning viewpoint-invariant ID features. Extensive experiments on Market-1501, DukeMTMC, MSMT17, and VeRi-776 demonstrate the effectiveness of the proposed VDG framework, as well as its superiority over the existing state-of-the-art approaches. The VDG model also demonstrates high quality in the image generation tasks.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Boyuan Liu, Runsheng Wang, Chengxin Zhao
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Deep Unsupervised Hashing with Hyperbolic Multi-Structure Learning
abstract
Unsupervised hashing aims to learn a compact binary hash code to represent complex image content without label information. Existing deep unsupervised hashing methods typically first employ extracted image embeddings to construct semantic similarity structures and then map the images into compact hash codes while preserving the semantic similarity structure. However, the limited representation power of embeddings in Euclidean space and the inadequate exploration of the similarity structure in current methods often result in poorly discriminative hash codes. In this paper, we propose a novel method called Hyperbolic Multi-Structure Hashing (HMSH) to address these issues. Specifically, to increase the representation power of embeddings, we propose to map embeddings from Euclidean space to hyperbolic space and use the similarity structure constructed in hyperbolic space to guide hash learning. Meanwhile, to fully explore the structural information, we investigate four kinds of data structures, including local neighborhood structure, global clustering structure, inter/intra-class variation and variation under perturbation. Different data structures can complement each other, which is beneficial for hash learning. Extensive experimental results on three benchmark image datasets show that HMSH significantly outperforms state-of-the-art unsupervised hashing methods for image retrieval.
Chuang Zhao 0001, Yuxuan Shi 0002, Jiazhong Chen
ECAI4
2023 Deep Unsupervised Hashing with Selective Semantic Mining
abstract
Most of the existing unsupervised hashing methods usually construct semantic similarity structure to guide hashing learning. However, due to the lack of filtering of useless information, some wrong guiding information in the similarity structure may damage the retrieval performance. Besides, some works adopt the framework of contrastive learning to preserve the discriminative semantic information that is more important for the hashing task. But such a training strategy may incorrectly embed some semantically similar samples far away due to the absence of manual label supervision, thus producing sub-optimal hash codes. To solve the aforementioned problems, we propose a novel method named Deep Selective Semantic Mining Hashing (DSSMH). Specifically, with the prior knowledge obtained by clustering, we select semantically correct image pairs with high confidence to alleviate the guidance of wrong information and correct sampling bias in contrastive learning. Extensive experiments demonstrate that DSSMH outperforms existing state-of-the-art methods.
Chuang Zhao 0001, Yuxuan Shi 0002, Chengxin Zhao, Jiazhong Chen
ICME5
2023 Detecting Adversarial Faces Using Only Real Face Self-Perturbations
abstract
Adversarial attacks aim to disturb the functionality of a target system by adding specific noise to the input samples, bringing potential threats to security and robustness when applied to facial recognition systems. Although existing defense techniques achieve high accuracy in detecting some specific adversarial faces (adv-faces), new attack methods especially GAN-based attacks with completely different noise patterns circumvent them and reach a higher attack success rate. Even worse, existing techniques require attack data before implementing the defense, making it impractical to defend newly emerging attacks that are unseen to defenders. In this paper, we investigate the intrinsic generality of adv-faces and propose to generate pseudo adv-faces by perturbing real faces with three heuristically designed noise patterns. We are the first to train an adv-face detector using only real faces and their self-perturbations, agnostic to victim facial recognition systems, and agnostic to unseen attacks. By regarding adv-faces as out-of-distribution data, we then naturally introduce a novel cascaded system for adv-face detection, which consists of training data self-perturbations, decision boundary regularization, and a max-pooling-based binary classifier focusing on abnormal local color aberrations. Experiments conducted on LFW and CelebA-HQ datasets with eight gradient-based and two GAN-based attacks validate that our method generalizes to a variety of unseen adversarial attacks.
Qian Wang 0001, Yongqin Xian, Xiaorui Lin, Ping Li 0021, Jiazhong Chen, Ning Yu 0006
IJCAI7
2023 General Adversarial Perturbation Simulating: Protect Unknown System by Detecting Unknown Adversarial Faces
abstract
Benefitting from the development of convolutional neural networks (CNNs), face recognition systems (FRSs) play a key role in many security-critical systems. However, FRSs have been proved to be vulnerable to adversarial faces (adv-faces). Adv-faces aim to change classification results by adding a subtle perturbation on real faces. The existence of adv-faces poses a significant threat to financial and privacy security. Previous detection methods require either training on pre-computed adv-faces or accessing to protected victim FRSs, bringing a dilemma in practical using. In this work, we heuristically propose an adversarial face detection method called General Adversarial Perturbation Simulating (GAPS) which is blind to both adversarial attacks and FRSs. Simulating noise patterns of several gradient-based adversarial perturbations, GAPS is able to generate simulated adversarial faces (sadv-faces) guiding detectors to learn general adversarial perturbation features and focus on classifying sensitive regions. Extensive experiments on LFW and CASIA-WebFace show that our method outperforms 9 state-of-the-art baseline methods and demonstrate the effectiveness of GAPS.
Feiran Sun, Xiaorui Lin, Jiazhong Chen, Ping Li 0021, Qian Wang 0001
IJCNN5
2023 Asymmetry-aware bilinear pooling in multi-modal data for head pose estimation
Jiazhong Chen, Dakai Ren, Hua Cao
Signal Process. Image Commun.1
2022 Reliability Exploration with Self-Ensemble Learning for Domain Adaptive Person Re-identification
abstract
Person re-identifcation (Re-ID) based on unsupervised domain adaptation (UDA) aims to transfer the pre-trained model from one labeled source domain to an unlabeled target domain. Existing methods tackle this problem by using clustering methods to generate pseudo labels. However, pseudo labels produced by these techniques may be unstable and noisy, substantially deteriorating models’ performance. In this paper, we propose a Reliability Exploration with Self-ensemble Learning (RESL) framework for domain adaptive person ReID. First, to increase the feature diversity, multiple branches are presented to extract features from different data augmentations. Taking the temporally average model as a mean teacher model, online label refning is conducted by using its dynamic ensemble predictions from different branches as soft labels. Second, to combat the adverse effects of unreliable samples in clusters, sample reliability is estimated by evaluating the consistency of different clusters’ results, followed by selecting reliable instances for training and re-weighting sample contribution within Re-ID losses. A contrastive loss is also utilized with cluster-level memory features which are updated by the mean feature. The experiments demonstrate that our method can signifcantly surpass the state-of-the-art performance on the unsupervised domain adaptive person ReID.
Zongyi Li, Yuxuan Shi 0002, Jiazhong Chen, Qian Wang 0001, Fengfan Zhou
AAAI4
2022 Multi-view facial action unit detection via DenseNets and CapsNets
Dakai Ren, Xiangming Wen, Jiazhong Chen, Shiqi Zhang 0002
Multim. Tools Appl.3
2021 Learning Diverse Local Patterns for Deepfake Detection with Image-level Supervision
abstract
To prevent the Deepfake-like videos from spreading, researchers have proposed many anti-forgery methods. However, most approaches require pixel-wise annotation, which conflicts with the real Deepfake detection scenario. To make full use of the image-level label, we propose a Local-Prediction framework that indirectly allows the image-level label to supervise local regions. To further enrich the local feature, we introduced the Local-Diversity concept in the Deepfake detection field for the first time. We proposed the Local-Diversity Loss based on the motivation that regional pattern differences can provide semi-supervised information during training. Compared to the previous method, our approach limits each classification unit's receptive field and enriches the feature diversity. In the experiment, our method is evaluated on three benchmark datasets of four widely-used manipulation types. The result shows that the Local-Prediction framework is beneficial to different CNN backbones and achieved significant performance. The proposed LD loss enriches the learned patterns of binary classifiers. Furthermore, we provide visualization and ablation studies to understand the mechanism.
Junrui Huang, Chengxin Zhao, Yutong Yao, Jiazhong Chen, Ping Li 0021
IJCNN5
2021 Video Saliency Prediction via Deep Eye Movement Learning
abstract
Existing methods often utilize temporal motion information and spatial layout information in video to predict video saliency. However, the fixations are not always consistent with the moving object of interest, because human eye fixations are determined not only by the spatio-temporal information, but also by the velocity of eye movement. To address this issue, a new saliency prediction method via deep eye movement learning (EML) is proposed in this paper. Compared with previous methods that use human fixations as ground truth, our method uses the optical flow of fixations between successive frames as an extra ground truth for the purpose of eye movement learning. Experimental results on DHF1K, Hollywood2, and UCF-sports datasets show the proposed EML model achieves a promising result across a wide of metrics.
Jiazhong Chen, Jie Chen 0058, Dakai Ren, Shiqi Zhang 0002, Zongyi Li
MMAsia1
2021 Video saliency prediction via spatio-temporal reasoning
Jiazhong Chen, Zongyi Li, Yi Jin 0001, Dakai Ren
Neurocomputing1
2021 Audiovisual saliency prediction via deep learning
Jiazhong Chen, Dakai Ren, Ping Duan
Neurocomputing1
2021 Gaze estimation via bilinear pooling-based attention networks
Dakai Ren, Jiazhong Chen, Zhaoming Lu, Zongyi Li
J. Vis. Commun. Image Represent.2
2021 Saliency detection via cross-scale deep inference
Dakai Ren, Xiangming Wen, Jiazhong Chen, Zongyi Li
J. Vis. Commun. Image Represent.4
2020 Lightweight Action Recognition with Sequence-Specific Global Context
abstract
With the emergence of a large number of video resources, video action recognition is attracting much attention. Recently, realizing the outstanding performance of three-dimensional (3D) convolutional neural networks (CNNs), many works have began to apply them for action recognition and obtained satisfactory results. However, high computational over-heads greatly reduce the efficiency of 3D CNNs. To make up for the shortcoming, in this paper, we first propose two innovations - the Xwise Separable Convolution and the SS block, both of which are lightweight. Then we build an efficient 3D CNN called the XwiseNet based on our innovations. Our work aims to make 3D CNNs lightweight without reducing the recognition accuracy. The key idea of the Xwise Separable Convolution is extremely decoupling the 3D convolution in channel, spatial, and temporal dimensions. The SS block can capture temporal long-range dependencies via aggregating sequence-specific global context to each sequence feature. Experiments have verified that our XwiseNet achieves competitive performance with the least computational overhead.
Jiazhong Chen, Lei Wu 0010, Yuxuan Shi 0002
IJCNN3
2020 Multi-Object Tracking Via Multi-Attention
abstract
Data association plays a crucial role in Multi-Object Tracking(MOT), but it is usually suppressed by occlusion. In this paper, we propose an online MOT approach via multiple attention mechanism(Multi-Attention) to handle the frequent interactions between targets. Specifically, the proposed Multi-Attention consists of spatial-attention, channel-attention, and temporal-attention three modules. The spatial-attention module lets the network focus on visible local areas by generating a visibility map, and the channel-attention module combines texture information and context information adaptively to build a recognizable object descriptor, then the temporal-attention module pays different attention to objects in the same trajectory avoiding the suppress caused by contaminated samples. Besides, a multiple branch convolutional block called receptive filed module(RFModule) is introduced to learn multiple levels of information for Multi-Attention. The experimental results on MOTChallenging benchmarks demonstrate the effectiveness of the proposed MOT algorithm against both online and offline trackers.
Xianrui Wang, Jiazhong Chen, Ping Li 0021
IJCNN3
2020 Cross-modal learning for saliency prediction in mobile environment
abstract
The existing researches reveal that a significant impact is introduced by viewing conditions for visual perception when viewing media on mobile screens. This brings two issues in the area of visual saliency that we need to address: how the saliency models perform in mobile conditions, and how to consider the mobile conditions when designing a saliency model. To investigate the performance of saliency models in mobile environment, eye fixations in four typical mobile conditions are collected as the mobile ground truth in this work. To consider the mobile conditions when designing a saliency model, we combine viewing factors and visual stimuli as two modalities, and a cross-modal based deep learning architecture is proposed for visual attention prediction. Experimental results demonstrate the model with the consideration of mobile viewing factors often outperforms the models without such consideration.
Dakai Ren, Xiangming Wen, Xiaoya Liu, Jiazhong Chen
MMAsia5
2020 XwiseNet: action recognition with Xwise separable convolutions
Jiazhong Chen, Lei Wu 0010, Yuxuan Shi 0002
Multim. Tools Appl.3
2020 Attention-based convolutional neural network for deep face recognition
Jiyang Wu, Junrui Huang, Jiazhong Chen, Ping Li 0021
Multim. Tools Appl.4
2019 BMNet: A Reconstructed Network for Lightweight Object Detection via Branch Merging
Yangyang Qin, Yuxuan Shi 0002, Lei Wu 0010, Jiazhong Chen, Baiyan Zhang
BMVC6
2019 Saliency Detection via Topological Feature Modulated Deep Learning
abstract
The topological feature in an image, such as connectivity and adjacency, plays an important role in eye fixation detection. However, due to the scalar and additive nature of neurons to aggregate the node values in a local neighborhood, it is hard for convolutional neural networks (CNNs) to directly obtain and model the topological feature. Thus we adopt the topological feature which is pre-trained in accordance with the relationship between the figure and ground. Then we add a topological feature modulated convolutional layer into CNNs. By this way, the topological feature is automatically modulated by the deep features and well modeled by the CNNs. Experimental results show the proposed method surpasses the state-of-the-art by a big margin.
Jiazhong Chen, Lei Wu 0010, Baiyan Zhang, Ping Li 0021
ICIP1
2019 Robust Mutual Learning Hashing
abstract
With the advances in deep learning, deep hashing methods have achieved promising results in recent years. However, tackling the distribution gap between train data and test data still remains unsolved. In this paper, motivated by Spatial Transformer Networks (STN) and mutual learning, we propose a novel robust hashing method (RMLH) for effective image retrieval. Specifically, the network learns more flexible transformation and makes itself generalize better to test data by plugging STN module. Then the mutual learning strategy is introduced to stabilize the training process. Furthermore, we relax the binary variables into continuous variables to avoid introducing any auxiliary variable. Finally, the experimental results show that our method has achieved the state-of-the-art performance on benchmark datasets.
Lei Wu 0010, Jiazhong Chen, Ping Li 0021
ICIP4
2019 Improving person re-identification by multi-task learning
Ping Li 0021, Yuxuan Shi 0002, Jiazhong Chen, Fuhao Zou
Neurocomputing5
2019 Saliency prediction by Mahalanobis distance of topological feature on deep color components
Jiazhong Chen, Ping Li 0021, Lei Wu 0010
J. Vis. Commun. Image Represent.1
2018 Salient object detection via spectral graph weighted low rank matrix recovery
Jiazhong Chen, Jie Chen 0058, Hua Cao, Yebin Fan
J. Vis. Commun. Image Represent.1
2017 Saliency detection using suitable variant of local and global consistency
abstract
In existing local and global consistency (LGC) framework, the cost functions related to classifying functions adopt the sum of each row of weight matrix as an important factor. Some of these classifying functions are successfully applied to saliency detection. From the point of saliency detection, this factor is inversely proportional to the colour contrast between image regions and their surroundings. However, an image region that holds a big colour contrast against it surroundings does not denote it must be a salient region. Therefore a suitable variant of LGC is introduced by removing this factor in cost function, and a suitable classifying function (SCF) is decided. Then a saliency detection method that utilises the SCF, content‐based initial label assignment scheme, and appearance‐based label assignment scheme is presented. Via updating the content‐based initial labels and appearance‐based labels by the SCF, a coarse saliency map and several intermediate saliency maps are obtained. Furthermore, to enhance the detection accuracy, a novel optimisation function is presented to fuse the intermediate saliency maps that have a high detection performance for final saliency generation. Numerous experimental results demonstrate that the proposed method achieves competitive performance against some recent state‐of‐the‐art algorithms for saliency detection.
Jiazhong Chen, Jie Chen 0058, Hua Cao
IET Comput. Vis.1
2017 Updating initial labels from spectral graph by manifold regularization for saliency detection
Jiazhong Chen, Bingpeng Ma, Hua Cao, Jie Chen 0058, Yebin Fan
Neurocomputing1
2017 Attention region detection based on closure prior in layered bit Planes
Jiazhong Chen, Bingpeng Ma, Hua Cao, Jie Chen 0058, Yebin Fan
Neurocomputing1
2016 Investigation of mobile surroundings for visual attention based on image perception model
abstract
This paper presents a novel perspective of performance evaluation for visual attention estimation: how the saliency models perform in mobile conditions. In particular, a broad-spectrum contrast sensitivity function is firstly proposed in this work. Based on this function, a new visual perception model is established, which will be further used to simulate the image perception in various mobile circumstances. Then the influence caused by mobile surroundings for visual attention is investigated. Meanwhile three types of mobile ground truth are generated by collecting viewers' fixations in three typical mobile conditions. Finally, by the perception model and mobile ground truth, an evaluation for ten classical visual attention models in various mobile surroundings is presented.
Jiazhong Chen, Yizhang Li, Yebin Fan, Hua Cao
VCIP1
2014 An improved hybrid fast mode decision method for H.264/AVC intra coding with local information
Changnian Chen, Jiazhong Chen, Zengwei Ju, Lai-Man Po
Multim. Tools Appl.2
2013 A hybrid fast mode decision method for H.264/AVC intra prediction
Changnian Chen, Jiazhong Chen, Xun Ouyang, Jingli Zhou
Multim. Tools Appl.2
2013 Accelerated implementation of adaptive directional lifting-based discrete wavelet transform on GPU
Jiazhong Chen, Zengwei Ju, Hua Cao, Bingpeng Ma, Changnian Chen, Leihua Qin
Signal Process. Image Commun.1
2012 Multicore Computing for SIFT Algorithm in MATLAB® Parallel Environment
abstract
An important processing stage in computer vision such as recognizes an object is feature extraction. Among those feature extraction algorithms, Lowe proposed the scale-invariant feature transform (SIFT) algorithm has been considered as one of the robust approaches. But due to the implementation of the convolution operation, the computing of SIFT algorithm is highly time-consuming. There are such problems as consuming too much power, sacrifice accuracy and lacking scalability for hardware-based acceleration scheme, it is still necessary to find the software-based acceleration approaches, especially for some applications and researches which need high-precision matched. Matlab is a general algorithm development environment with powerful image processing and other supporting toolboxes. With the rapid development of multicore CPU technology, using multicore computer and Matlab is an intuitive and simple way to speed up the computing for SIFT algorithm. In this paper, we try to have a view for using the Matlab parallel toolbox to accelerate the SIFT algorithm by two schemes of task-parallelism and data-parallelism modal. The results show that the parallel versions of former sequential algorithm with simple modifications achieve the speedup up to 6.6 times.
Hua Cao, Jiazhong Chen
ICPADS2
2010 A hybrid M-channel filter bank and DCT framework for H.264/AVC intra coding
Jiazhong Chen, Shengsheng Yu, Jie Yang 0017, Jingli Zhou
Multim. Tools Appl.2
2009 The training of Karhunen-Loève transform matrix and its application for H.264 intra coding
Jiazhong Chen, Shengsheng Yu, Jingli Zhou, Lai-Man Po
Multim. Tools Appl.2
2006 The application of symmetric orthogonal multiwavelets and prefilter technique for image compression
Jiazhong Chen, Xun Ouyang, Wu Zheng, Jingli Zhou, Shengsheng Yu
Multim. Tools Appl.1
2005 A Very Low Bit Rate Video Coding Combined with Fast Adaptive Block Size Motion Estimation and Nonuniform Scalar Quantization Multiwavelet Transform
Jiazhong Chen, Jingli Zhou, Shengsheng Yu, Junhao Zheng
Multim. Tools Appl.1