Zhijiang Zhang

dblp:45/7732 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RAG-Based Enterprise Knowledge Platform: Design and Evaluation
Zhijiang Zhang, Yanyong Zhang, Hao Zhou 0001
KSEM (2)1
2026 Hierarchical prior-guided and channel-wise adaptation fusion network for RGB-D railway surface defect inspection
Jianlin Chen, Gongyang Li, Zhijiang Zhang, Dan Zeng 0001
J. Vis. Commun. Image Represent.3
2025 Making Old Film Great Again: Degradation-aware State Space Model for Old Film Restoration
Yudong Mao, Zhiwei Zhong 0001, Peilin Chen 0001, Zhijiang Zhang, Shiqi Wang 0001
CVPR5
2025 Data expansion for industrial defect detection with diffusion prior
Yichun Tai, Zhijiang Zhang
Eng. Appl. Artif. Intell.2
2025 Defect Image Sample Generation With Diffusion Prior for Steel Surface Defect Recognition
abstract
The task of steel surface defect recognition is an industrial problem with great industry values. The data insufficiency is the major challenge in training a robust defect recognition network. Existing methods have investigated to enlarge the dataset by generating samples with generative models. However, their generation quality is still limited by the insufficiency of defect image samples. To this end, we propose Stable Surface Defect Generation (StableSDG), which transfers the vast generation distribution embedded in Stable Diffusion model for steel surface defect image generation. To tackle with the distinctive distribution gap between steel surface images and generated images of the diffusion model, we propose two processes. First, we align the distribution by adapting parameters of the diffusion model, adopted both in the token embedding space and network parameter space. Besides, in the generation process, we propose image-oriented generation rather than from pure Gaussian noises. We conduct extensive experiments on steel surface defect dataset, demonstrating state-of-the-art performance on generating high-quality samples and training recognition models, and both designed processes are significant for the performance. Note to Practitioners—This article introduces StableSDG, a method that generates realistic defect images even with limited data. It overcomes the shortcomings of current deep learning approaches that need large datasets to train from scratch. Our solution is to adapt a text-to-image diffusion model for defect generation. The proposed strategy involves two processes: training to adapt token embeddings and model parameters, and generation from partially perturbed defect images. The results show enhanced generation quality and improved accuracy for recognition models trained on the expanded dataset. StableSDG can be practically applied to efficiently enlarge a defect dataset, even when starting with a small amount of data.
Yichun Tai, Zhenzhen Huang, Zhijiang Zhang
IEEE Trans Autom. Sci. Eng.5
2025 DefFiller: mask-conditioned generation with diffusion prior for saliency-based defect detection
Yichun Tai, Zhenzhen Huang, Zhijiang Zhang
Vis. Comput.4
2024 EFDCNet: Encoding fusion and decoding correction network for RGB-D indoor semantic segmentation
Jianlin Chen, Gongyang Li, Zhijiang Zhang, Dan Zeng 0001
Image Vis. Comput.3
2023 BARS: a benchmark for airport runway segmentation
Zhijiang Zhang, Yichun Tai
Appl. Intell.2
2023 Multi-Agent Semi-Siamese Training for Long-Tail and Shallow Face Learning
abstract
With the recent development of deep convolutional neural networks and large-scale datasets, deep face recognition has made remarkable progress and been widely used in various applications. However, unlike the existing public face datasets, in many real-world scenarios of face recognition, the depth of the training dataset is shallow, which means that only two face images are available for each ID. With the non-uniform increase of samples, such issue is converted to a more general case, known as long-tail face learning, which suffers from data imbalance and intra-class diversity dearth simultaneously. These adverse conditions damage the training and result in the decline of model performance. Based on Semi-Siamese Training, we introduce an advanced solution, namedMulti-Agent Semi-Siamese Training(MASST), to address these problems. MASST includes a probe network and multiple gallery agents—the former aims to encode the probe features, and the latter constitutes a stack of networks that encode the prototypes (gallery features). For each training iteration, the gallery network, which is sequentially rotated from the stack, and the probe network form a pair of Semi-Siamese networks. We give the theoretical and empirical analysis that, given the long-tail (or shallow) data and training loss, MASST smooths the loss landscape and satisfies the Lipschitz continuity with the help of multiple agents and the updating gallery queue. The proposed method is out of extra-dependency, and thus can be easily integrated with the existing loss functions and network architectures. It is worth noting that although multiple gallery agents are employed for training, only the probe network is needed for inference, without increasing the inference cost. Extensive experiments and comparisons demonstrate the advantages of MASST for long-tail and shallow face learning.
Yichun Tai, Hailin Shi, Dan Zeng 0001, Yibo Hu 0003, Zhijiang Zhang, Tao Mei 0001
ACM Trans. Multim. Comput. Commun. Appl.7
2022 An Improved Faster R-CNN for Steel Surface Defect Detection
abstract
With the continuous increase of surface quality requirements in steel industry manufacturing, defect detection has received extensive attention. Correct and rapid detection of steel surface defects can significantly improve product quality and productivity. The existing methods improve accuracy by expanding the depth of networks or using various feature fusion technologies, but reduce the computational efficiency. To achieve the balance of precision and speed, we propose an improved network based on Faster R-CNN for defect detection of steel surface. Firstly, the emerging ConvNeXt architecture is adopted to act as backbone to extract features in Faster R-CNN. Besides, Convolutional Block Attention Module (CBAM) is used to improve the attention of our model to surface defects and suppress features of the complex background. Finally, k-means clustering algorithm is utilized to generate anchors that are better adapted to the surface defects. The proposed method achieves a mean average precision (mAP) of 80.78% and a detection speed of 26 frames per second (FPS) on the NEU-DET dataset, improving about 1.5% compared to the YOLOv5 and 8.4 % compared to original Faster R-CNN, which indicates its superior for defect detection of steel surface.
Xiancong Shi, Sike Zhou, Yichun Tai, Jinzhong Wang, Shoucang Wu, Jinrong Liu 0003, Zhijiang Zhang
MMSP9
2019 Proposal pyramid networks for fast face detection
Dan Zeng 0001, Fan Zhao 0004, Shiming Ge, Wei Shen 0002, Zhijiang Zhang
Inf. Sci.6
2018 Volumeter: 3D human body parameters measurement with a single Kinect
abstract
3D human body parameters measurement is a challenging task due to two main reasons: (i) it is difficult to reconstruct 3D human model due to flexible deformation of non‐rigid body during images capturing process and (ii) there lies a gap between 3D model and body parameters. To address these two issues, a 3D human body parameters measurement system is represented. With the object freely spinning in front of a Kinect, body parameters are calculated. To reduce registration errors caused by body deformation while rotating, a piecewise tracking and mapping algorithm based on KinectFusion framework is proposed. Then model–model iterative closest point and non‐rigid constraints are introduced to optimise alignments and disambiguate different surfaces caused by aliasing in the piecewise strategy. Finally, a novel method is presented to measure the volume and perimeter of human body with the truncated signed distance function values of voxels. Extensive experimental results show that the proposed method achieves comparable accuracy to the state of the arts, and the error of volume and perimeter measurements are 2.0 and 5.8%, respectively.
Qinzhu He, Yijun Ji, Dan Zeng 0001, Zhijiang Zhang
IET Comput. Vis.4
2018 Bag of Shape Features with a learned pooling function for shape recognition
Wei Shen 0002, Chenting Du, Yuan Jiang 0002, Dan Zeng 0001, Zhijiang Zhang
Pattern Recognit. Lett.5
2018 Multi-oriented text detection from natural scene images based on a CNN and pruning non-adjacent graph edges
Yuanwang Wei, Wei Shen 0002, Dan Zeng 0001, Lihua Ye, Zhijiang Zhang
Signal Process. Image Commun.5
2017 Shape recognition by bag of contour fragments with a learned pooling function
abstract
Bag of Contour Fragments (BoCF), derived from the well-known Bag-of-Features (BoF), is an effective framework for shape representation. The feature pooling in this framework is a critical step, while either max pooling or average pooling is not a learnable process. In this paper, we aim at learning a pooling function which is adaptive to the input contour fragment features instead. Towards this end, we formulate our pooling function as a weighted sum of max pooling and average pooling, where the weight is expressed by an activation function of the input contour fragment features. To automatically learn this weight, the output of the pooling function is fed into a SVM classifier and they are trained jointly to minimize a shape classification loss. Experimental results on several standard shape datasets demonstrate the effectiveness of the proposed learned pooling function, which can achieve considerable improvements compared with BoCF.
Wei Shen 0002, Wenjing Gao, Yuan Jiang 0002, Dan Zeng 0001, Zhijiang Zhang
ICIP5
2017 Text detection in scene images based on exhaustive segmentation
Yuanwang Wei, Zhijiang Zhang, Wei Shen 0002, Dan Zeng 0001, Mei Fang, Shifu Zhou
Signal Process. Image Commun.2
2016 Object Skeleton Extraction in Natural Images by Fusing Scale-Associated Deep Side Outputs
abstract
Object skeleton is a useful cue for object detection, complementary to the object contour, as it provides a structural representation to describe the relationship among object parts. While object skeleton extraction in natural images is a very challenging problem, as it requires the extractor to be able to capture both local and global image context to determine the intrinsic scale of each skeleton pixel. Existing methods rely on per-pixel based multi-scale feature computation, which results in difficult modeling and high time consumption. In this paper, we present a fully convolutional network with multiple scale-associated side outputs to address this problem. By observing the relationship between the receptive field sizes of the sequential stages in the network and the skeleton scales they can capture, we introduce a scale-associated side output to each stage. We impose supervision to different stages by guiding the scale-associated side outputs toward groundtruth skeletons of different scales. The responses of the multiple scaleassociated side outputs are then fused in a scale-specific way to localize skeleton pixels with multiple scales effectively. Our method achieves promising results on two skeleton extraction datasets, and significantly outperforms other competitors.
Wei Shen 0002, Kai Zhao 0012, Yuan Jiang 0002, Yan Wang 0033, Zhijiang Zhang, Xiang Bai
CVPR5
2016 Multiple instance subspace learning via partial random projection tree for local reflection symmetry in natural images
Wei Shen 0002, Xiang Bai, Zihao Hu, Zhijiang Zhang
Pattern Recognit.4
2016 Spatial-temporal convolutional neural networks for anomaly detection and localization in crowded scenes
Shifu Zhou, Wei Shen 0002, Dan Zeng 0001, Mei Fang, Yuanwang Wei, Zhijiang Zhang
Signal Process. Image Commun.6
2015 DeepContour: A deep convolutional feature learned by positive-sharing loss for contour detection
abstract
Contour detection serves as the basis of a variety of computer vision tasks such as image segmentation and object recognition. The mainstream works to address this problem focus on designing engineered gradient features. In this work, we show that contour detection accuracy can be improved by instead making the use of the deep features learned from convolutional neural networks (CNNs). While rather than using the networks as a blackbox feature extractor, we customize the training strategy by partitioning contour (positive) data into subclasses and fitting each subclass by different model parameters. A new loss function, named positive-sharing loss, in which each subclass shares the loss for the whole positive class, is proposed to learn the parameters. Compared to the sofmax loss function, the proposed one, introduces an extra regularizer to emphasizes the losses for the positive and negative classes, which facilitates to explore more discriminative features. Our experimental results demonstrate that learned deep features can achieve top performance on Berkeley Segmentation Dataset and Benchmark (BSDS500) and obtain competitive cross dataset generalization result on the NYUD dataset.
Wei Shen 0002, Xinggang Wang, Yan Wang 0033, Xiang Bai, Zhijiang Zhang
CVPR5
2015 Unusual event detection in crowded scenes by trajectory analysis
abstract
Anomaly detection in crowded scenes is a challenge task due to variation of the definitions for both abnormality and normality, the low resolution on the target, ambiguity of appearance, and severe occlusions of inter-object. In this paper, we propose a novel statistical framework to detect abnormal behaviors of the crowded scene by modeling trajectories of pedestrians. First, the trajectories are acquired by Kanade-Lucas-Tomasi Feature Tracker (KLT). Then trajectories are grouped to form representative trajectories, which characterize the underlying motion patterns of the crowd. Finally, trajectories are modeled by Multi-Observation Hidden Markov Model (MOHMM) to determine whether frames are normal or abnormal. The experiments are conducted on a well-known crowded scene dataset. Experimental results show that the proposed method can capture abnormal crowd behaviors successfully and achieves state-of-the-art performances.
Shifu Zhou, Wei Shen 0002, Dan Zeng 0001, Zhijiang Zhang
ICASSP4
2014 Regularity Guaranteed Human Pose Correction
Wei Shen 0002, Rui Lei, Dan Zeng 0001, Zhijiang Zhang
ACCV (2)4