VLDB 2026 Research / reviewers in the wild / expert
Zhijiang Zhang
dblp:45/7732
· DBLP profile ↗
22ranked-venue papers
1as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAG-Based Enterprise Knowledge Platform: Design and Evaluation
Zhijiang Zhang, Yanyong Zhang, Hao Zhou 0001 |
KSEM (2) | 1 |
| 2026 | Hierarchical prior-guided and channel-wise adaptation fusion network for RGB-D railway surface defect inspection
Jianlin Chen, Gongyang Li, Zhijiang Zhang, Dan Zeng 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | Making Old Film Great Again: Degradation-aware State Space Model for Old Film Restoration
Yudong Mao, Zhiwei Zhong 0001, Peilin Chen 0001, Zhijiang Zhang, Shiqi Wang 0001 |
CVPR | 5 |
| 2025 | Data expansion for industrial defect detection with diffusion prior
Yichun Tai, Zhijiang Zhang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Defect Image Sample Generation With Diffusion Prior for Steel Surface Defect RecognitionabstractThe task of steel surface defect recognition is an industrial problem with great industry values. The data insufficiency is the major challenge in training a robust defect recognition network. Existing methods have investigated to enlarge the dataset by generating samples with generative models. However, their generation quality is still limited by the insufficiency of defect image samples. To this end, we propose Stable Surface Defect Generation (StableSDG), which transfers the vast generation distribution embedded in Stable Diffusion model for steel surface defect image generation. To tackle with the distinctive distribution gap between steel surface images and generated images of the diffusion model, we propose two processes. First, we align the distribution by adapting parameters of the diffusion model, adopted both in the token embedding space and network parameter space. Besides, in the generation process, we propose image-oriented generation rather than from pure Gaussian noises. We conduct extensive experiments on steel surface defect dataset, demonstrating state-of-the-art performance on generating high-quality samples and training recognition models, and both designed processes are significant for the performance. Note to Practitioners—This article introduces StableSDG, a method that generates realistic defect images even with limited data. It overcomes the shortcomings of current deep learning approaches that need large datasets to train from scratch. Our solution is to adapt a text-to-image diffusion model for defect generation. The proposed strategy involves two processes: training to adapt token embeddings and model parameters, and generation from partially perturbed defect images. The results show enhanced generation quality and improved accuracy for recognition models trained on the expanded dataset. StableSDG can be practically applied to efficiently enlarge a defect dataset, even when starting with a small amount of data. Yichun Tai, Zhenzhen Huang, Zhijiang Zhang |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | DefFiller: mask-conditioned generation with diffusion prior for saliency-based defect detection
Yichun Tai, Zhenzhen Huang, Zhijiang Zhang |
Vis. Comput. | 4 |
| 2024 | EFDCNet: Encoding fusion and decoding correction network for RGB-D indoor semantic segmentation
Jianlin Chen, Gongyang Li, Zhijiang Zhang, Dan Zeng 0001 |
Image Vis. Comput. | 3 |
| 2023 | BARS: a benchmark for airport runway segmentation
Zhijiang Zhang, Yichun Tai |
Appl. Intell. | 2 |
| 2023 | Multi-Agent Semi-Siamese Training for Long-Tail and Shallow Face LearningabstractWith the recent development of deep convolutional neural networks and large-scale datasets, deep face recognition has made remarkable progress and been widely used in various applications. However, unlike the existing public face datasets, in many real-world scenarios of face recognition, the depth of the training dataset is shallow, which means that only two face images are available for each ID. With the non-uniform increase of samples, such issue is converted to a more general case, known as long-tail face learning, which suffers from data imbalance and intra-class diversity dearth simultaneously. These adverse conditions damage the training and result in the decline of model performance. Based on Semi-Siamese Training, we introduce an advanced solution, namedMulti-Agent Semi-Siamese Training(MASST), to address these problems. MASST includes a probe network and multiple gallery agents—the former aims to encode the probe features, and the latter constitutes a stack of networks that encode the prototypes (gallery features). For each training iteration, the gallery network, which is sequentially rotated from the stack, and the probe network form a pair of Semi-Siamese networks. We give the theoretical and empirical analysis that, given the long-tail (or shallow) data and training loss, MASST smooths the loss landscape and satisfies the Lipschitz continuity with the help of multiple agents and the updating gallery queue. The proposed method is out of extra-dependency, and thus can be easily integrated with the existing loss functions and network architectures. It is worth noting that although multiple gallery agents are employed for training, only the probe network is needed for inference, without increasing the inference cost. Extensive experiments and comparisons demonstrate the advantages of MASST for long-tail and shallow face learning. Yichun Tai, Hailin Shi, Dan Zeng 0001, Yibo Hu 0003, Zhijiang Zhang, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2022 | An Improved Faster R-CNN for Steel Surface Defect DetectionabstractWith the continuous increase of surface quality requirements in steel industry manufacturing, defect detection has received extensive attention. Correct and rapid detection of steel surface defects can significantly improve product quality and productivity. The existing methods improve accuracy by expanding the depth of networks or using various feature fusion technologies, but reduce the computational efficiency. To achieve the balance of precision and speed, we propose an improved network based on Faster R-CNN for defect detection of steel surface. Firstly, the emerging ConvNeXt architecture is adopted to act as backbone to extract features in Faster R-CNN. Besides, Convolutional Block Attention Module (CBAM) is used to improve the attention of our model to surface defects and suppress features of the complex background. Finally, k-means clustering algorithm is utilized to generate anchors that are better adapted to the surface defects. The proposed method achieves a mean average precision (mAP) of 80.78% and a detection speed of 26 frames per second (FPS) on the NEU-DET dataset, improving about 1.5% compared to the YOLOv5 and 8.4 % compared to original Faster R-CNN, which indicates its superior for defect detection of steel surface. Xiancong Shi, Sike Zhou, Yichun Tai, Jinzhong Wang, Shoucang Wu, Jinrong Liu 0003, Zhijiang Zhang |
MMSP | 9 |
| 2019 | Proposal pyramid networks for fast face detection
Dan Zeng 0001, Fan Zhao 0004, Shiming Ge, Wei Shen 0002, Zhijiang Zhang |
Inf. Sci. | 6 |
| 2018 | Volumeter: 3D human body parameters measurement with a single Kinectabstract3D human body parameters measurement is a challenging task due to two main reasons: (i) it is difficult to reconstruct 3D human model due to flexible deformation of non‐rigid body during images capturing process and (ii) there lies a gap between 3D model and body parameters. To address these two issues, a 3D human body parameters measurement system is represented. With the object freely spinning in front of a Kinect, body parameters are calculated. To reduce registration errors caused by body deformation while rotating, a piecewise tracking and mapping algorithm based on KinectFusion framework is proposed. Then model–model iterative closest point and non‐rigid constraints are introduced to optimise alignments and disambiguate different surfaces caused by aliasing in the piecewise strategy. Finally, a novel method is presented to measure the volume and perimeter of human body with the truncated signed distance function values of voxels. Extensive experimental results show that the proposed method achieves comparable accuracy to the state of the arts, and the error of volume and perimeter measurements are 2.0 and 5.8%, respectively. Qinzhu He, Yijun Ji, Dan Zeng 0001, Zhijiang Zhang |
IET Comput. Vis. | 4 |
| 2018 | Bag of Shape Features with a learned pooling function for shape recognition
Wei Shen 0002, Chenting Du, Yuan Jiang 0002, Dan Zeng 0001, Zhijiang Zhang |
Pattern Recognit. Lett. | 5 |
| 2018 | Multi-oriented text detection from natural scene images based on a CNN and pruning non-adjacent graph edges
Yuanwang Wei, Wei Shen 0002, Dan Zeng 0001, Lihua Ye, Zhijiang Zhang |
Signal Process. Image Commun. | 5 |
| 2017 | Shape recognition by bag of contour fragments with a learned pooling functionabstractBag of Contour Fragments (BoCF), derived from the well-known Bag-of-Features (BoF), is an effective framework for shape representation. The feature pooling in this framework is a critical step, while either max pooling or average pooling is not a learnable process. In this paper, we aim at learning a pooling function which is adaptive to the input contour fragment features instead. Towards this end, we formulate our pooling function as a weighted sum of max pooling and average pooling, where the weight is expressed by an activation function of the input contour fragment features. To automatically learn this weight, the output of the pooling function is fed into a SVM classifier and they are trained jointly to minimize a shape classification loss. Experimental results on several standard shape datasets demonstrate the effectiveness of the proposed learned pooling function, which can achieve considerable improvements compared with BoCF. Wei Shen 0002, Wenjing Gao, Yuan Jiang 0002, Dan Zeng 0001, Zhijiang Zhang |
ICIP | 5 |
| 2017 | Text detection in scene images based on exhaustive segmentation
Yuanwang Wei, Zhijiang Zhang, Wei Shen 0002, Dan Zeng 0001, Mei Fang, Shifu Zhou |
Signal Process. Image Commun. | 2 |
| 2016 | Object Skeleton Extraction in Natural Images by Fusing Scale-Associated Deep Side OutputsabstractObject skeleton is a useful cue for object detection, complementary to the object contour, as it provides a structural representation to describe the relationship among object parts. While object skeleton extraction in natural images is a very challenging problem, as it requires the extractor to be able to capture both local and global image context to determine the intrinsic scale of each skeleton pixel. Existing methods rely on per-pixel based multi-scale feature computation, which results in difficult modeling and high time consumption. In this paper, we present a fully convolutional network with multiple scale-associated side outputs to address this problem. By observing the relationship between the receptive field sizes of the sequential stages in the network and the skeleton scales they can capture, we introduce a scale-associated side output to each stage. We impose supervision to different stages by guiding the scale-associated side outputs toward groundtruth skeletons of different scales. The responses of the multiple scaleassociated side outputs are then fused in a scale-specific way to localize skeleton pixels with multiple scales effectively. Our method achieves promising results on two skeleton extraction datasets, and significantly outperforms other competitors. Wei Shen 0002, Kai Zhao 0012, Yuan Jiang 0002, Yan Wang 0033, Zhijiang Zhang, Xiang Bai |
CVPR | 5 |
| 2016 | Multiple instance subspace learning via partial random projection tree for local reflection symmetry in natural images
Wei Shen 0002, Xiang Bai, Zihao Hu, Zhijiang Zhang |
Pattern Recognit. | 4 |
| 2016 | Spatial-temporal convolutional neural networks for anomaly detection and localization in crowded scenes
Shifu Zhou, Wei Shen 0002, Dan Zeng 0001, Mei Fang, Yuanwang Wei, Zhijiang Zhang |
Signal Process. Image Commun. | 6 |
| 2015 | DeepContour: A deep convolutional feature learned by positive-sharing loss for contour detectionabstractContour detection serves as the basis of a variety of computer vision tasks such as image segmentation and object recognition. The mainstream works to address this problem focus on designing engineered gradient features. In this work, we show that contour detection accuracy can be improved by instead making the use of the deep features learned from convolutional neural networks (CNNs). While rather than using the networks as a blackbox feature extractor, we customize the training strategy by partitioning contour (positive) data into subclasses and fitting each subclass by different model parameters. A new loss function, named positive-sharing loss, in which each subclass shares the loss for the whole positive class, is proposed to learn the parameters. Compared to the sofmax loss function, the proposed one, introduces an extra regularizer to emphasizes the losses for the positive and negative classes, which facilitates to explore more discriminative features. Our experimental results demonstrate that learned deep features can achieve top performance on Berkeley Segmentation Dataset and Benchmark (BSDS500) and obtain competitive cross dataset generalization result on the NYUD dataset. Wei Shen 0002, Xinggang Wang, Yan Wang 0033, Xiang Bai, Zhijiang Zhang |
CVPR | 5 |
| 2015 | Unusual event detection in crowded scenes by trajectory analysisabstractAnomaly detection in crowded scenes is a challenge task due to variation of the definitions for both abnormality and normality, the low resolution on the target, ambiguity of appearance, and severe occlusions of inter-object. In this paper, we propose a novel statistical framework to detect abnormal behaviors of the crowded scene by modeling trajectories of pedestrians. First, the trajectories are acquired by Kanade-Lucas-Tomasi Feature Tracker (KLT). Then trajectories are grouped to form representative trajectories, which characterize the underlying motion patterns of the crowd. Finally, trajectories are modeled by Multi-Observation Hidden Markov Model (MOHMM) to determine whether frames are normal or abnormal. The experiments are conducted on a well-known crowded scene dataset. Experimental results show that the proposed method can capture abnormal crowd behaviors successfully and achieves state-of-the-art performances. Shifu Zhou, Wei Shen 0002, Dan Zeng 0001, Zhijiang Zhang |
ICASSP | 4 |
| 2014 | Regularity Guaranteed Human Pose Correction
Wei Shen 0002, Rui Lei, Dan Zeng 0001, Zhijiang Zhang |
ACCV (2) | 4 |