Qian Zhang 0018

dblp:04/2024-18 · DBLP profile ↗
← Back
25ranked-venue papers
2as first author
15since 2021 · last 2025
0000-0001-5562-4759ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 9 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Exploring the Black-Box: Testing Image Synthesis Systems through Metamorphic Exploration
abstract
The increasing complexity of deep learning models, especially in black-box scenarios, presents significant challenges to traditional software testing methods. Due to the lack of transparency in neural networks’ decision-making processes and the non-deterministic nature of model outputs, traditional test oracle approaches become inadequate. To address this problem, Metamorphic Testing (MT) and its extended approach, Metamorphic Exploration (ME), provide new ideas for validating deep learning systems by defining Metamorphic Relations (MR) between inputs and outputs. However, existing image transformation-based MR faces new challenges in image synthesis scenarios, as these operations may destroy the contextual information and affect the model’s performance. This paper proposes a novel ME design for deep learning image synthesis networks and demonstrates its effectiveness using a visible-infrared image fusion network as the case study. The result identifies the performance degradation problem due to the tensor dimension manipulation error, which indicates that the ME not only detects defects but also helps developers deeply understand the internal mechanisms of complex systems through the Hypothesized Metamorphic Relation (HMR), thus providing unique value for software quality assurance (SQA) of AI-driven software.
Zhihao Ying, Yifan Zhang 0016, Qian Zhang 0018, Dave Towey
COMPSAC4
2025 HFA-UNet: hybrid and full attention UNet for thyroid nodule segmentation
abstract
Ultrasound imaging is the most commonly used method for screening thyroid nodules due to its low cost and non-invasive nature. Thyroid nodule lesions have variable shapes, rich aspect ratios, unclear boundaries, calcified nodule-induced acoustic shadows, and noise interference, causing challenges in accurate segmentation. Recent methods ignore various scale features and details in different resolutions of images, leading to redundant or missing feature information and then affecting the segmentation performance. In this paper, we introduce a hybrid and full attention UNet model for ultrasound thyroid nodule segmentation. Self, spatial and channel attention are combined in a U-Net-like structure to extract global and local features simultaneously. A novel full attention multi-scale fusion stage is designed to enhance boundary features while suppressing noise features. At the same time, the model dynamically adjusts the number of skip connections corresponding to images of different resolutions to better utilize multi-scale features and detailed information. We evaluate our model on DDTI, TN3K and Stanford Cine-Clip datasets, including internal validation and cross-dataset testing. The results show that our proposed model for internal validation in the DDTI dataset increases the Dice score and mean intersection over union by 2.36 % and 1.04 % compared to the state-of-the-art model. In the TN3K dataset, they increase by 1.66 % and 3.05 %.
Yuanhao Zou, Xiangjian He, Qing Xu 0014, Ming Liu 0021, Shengji Jin, Qian Zhang 0018, Maggie M. He, Jian Zhang 0002
Knowl. Based Syst.7
2024 Scale Optimization Using Evolutionary Reinforcement Learning for Object Detection on Drone Imagery
abstract
Object detection in aerial imagery presents a significant challenge due to large scale variations among objects. This paper proposes an evolutionary reinforcement learning agent, integrated within a coarse-to-fine object detection framework, to optimize the scale for more effective detection of objects in such images. Specifically, a set of patches potentially containing objects are first generated. A set of rewards measuring the localization accuracy, the accuracy of predicted labels, and the scale consistency among nearby patches are designed in the agent to guide the scale optimization. The proposed scale-consistency reward ensures similar scales for neighboring objects of the same category. Furthermore, a spatial-semantic attention mechanism is designed to exploit the spatial semantic relations between patches. The agent employs the proximal policy optimization strategy in conjunction with the evolutionary strategy, effectively utilizing both the current patch status and historical experience embedded in the agent. The proposed model is compared with state-of-the-art methods on two benchmark datasets for object detection on drone imagery. It significantly outperforms all the compared methods. Code is available at https://github.com/UNNC-CV/EvOD/.
Jialu Zhang 0003, Jianfeng Ren, Qian Zhang 0018, Yitian Zhao, Ruibin Bai, Xiangjian He, Jiang Liu 0001
AAAI5
2024 Size-Insensitive Network for Visible-Infrared Image Fusion Model
abstract
Visible-infrared image fusion is a technique that extracts information from different sensors. It could be used to enhance human visual perception of video surveillance under low-light conditions, and provide rich information for subsequent tasks. Vision Transformer (ViT) based fusion algorithms require standardizing input images to a specific height and width that could be divided into a series of blocks of fixed size. Consequently, a scaling operation must be performed on the original image, which frequently decreases the quality of fusion results. This paper proposes a visible-infrared image fusion neural network that is insensitive to input size, by first utilizing a fixed-size image pre-fusion framework to generate lossless instructive fusion results (IFRs), followed by a size-insensitive enhancing framework that refines these preliminary fused images under the guidance of IFRs. It also has potential applicability to other image fusion algorithms, like multi-focus image fusion.
Qian Zhang 0018, Dave Towey
COMPSAC2
2024 Task Oriented Image Quality Assessment for Synthesized Images
Qian Zhang 0018, Zhanghao Jiang, Boon-Giin Lee
ICPR (25)2
2024 Seat belt detection using gated Bi-LSTM with part-to-whole attention on diagonally sampled patches
abstract
One of the high-risk behaviors leading to severe traffic injuries is not wearing a seat belt. It is therefore very important to be able to automatically detect seat belts from surveillance images, encourage drivers to wear seat belts, and enhance passenger safety. In this paper, a novel deep neural network, Gated Bi-directional Long Short-Term Memory network with part-to-whole attention (GBL-PA), is proposed for seat belt detection from surveillance images. The innovation of our model lies in its unique diagonal sampling strategy, which meticulously captures the seat belt’s fine details, typically oriented from top right to bottom left across the torso of vehicle occupants. Our framework’s novelty is further encapsulated by the part-to-whole attention mechanism, which intelligently harmonizes the detailed local information from seat belt–specific patches with the broader contextual insights from the regional proposals. The pioneering design of a Gated Bi-directional LSTM network facilitates the dynamic integration of interactions across patches to deliver an optimized final prediction. The superiority of GBL-PA is established through rigorous comparison with the state-of-the-art methods on a new, large benchmark dataset comprising 14,936 images from traffic surveillance footage. Our framework demonstrates a notable improvement, achieving a mean Average Precision (mAP) of 72.3%, which surpasses the second best, YOLOX, by 0.9% mAP. This significant and consistent outperformance across various metrics underscores the transformative potential of GBL-PA in the realm of traffic safety enforcement. The source code of our framework is available at ANONYMISED.
Zheng Lu 0002, Jianfeng Ren, Qian Zhang 0018
Expert Syst. Appl.4
2024 MITER: Medical Image-TExt joint adaptive pretRaining with multi-level contrastive learning
abstract
Recently multimodal medical pretraining models play a significant role in automatic medical image and text analysis that has wide social and economical impact in healthcare. Despite being able to be quickly transferred to downstream tasks, the models are greatly limited due to the fact that these models can only be pretrained with professional medical image-text datasets, which usually contain a very small number of samples. In this work We propose MITER (Medical Image-Text Joint adaptive Pretraining), a joint adaptive pretraining framework via multi-level contrastive learning to overcome this limitation by pretraining image and text models for medical domain and utilizing existing models pretrained on generic data, which contain enormous number of samples. MITER features two types of objectives to solve the problem. The first type is uni-modal objectives that pretrain the models with medical images and text separately on uni-modal tasks. The other type is a cross-modal objective that pretrains jointly, allowing the models to influence each other on cross-modal tasks. We also introduce a strategy to dynamically select hard negative samples during the training process for better performance. Experimental results over four medical tasks, image-report retrieval, multi-label image classification, visual question answering, and report generation, show that our MITER framework solves the limitation problem by greatly outperforming existing benchmark models on all the tasks. The source code of our framework is available online.2
Xiaochu Tang, Jing Xiao 0006, Youxin Chen, Xiu Li 0001, Qian Zhang 0018, Zheng Lu 0002
Expert Syst. Appl.7
2024 Deep learning-based RGB-thermal image denoising: review and applications
Boon-Giin Lee, Matthew Pike, Qian Zhang 0018, Wan-Young Chung
Multim. Tools Appl.4
2023 Improving Visual-Semantic Embedding with Adaptive Pooling and Optimization Objective
abstract
Zijian Zhang, Chang Shu, Ya Xiao, Yuan Shen, Di Zhu, Youxin Chen, Jing Xiao, Jey Han Lau, Qian Zhang, Zheng Lu. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Ya Xiao 0006, Youxin Chen, Jing Xiao 0006, Jey Han Lau, Qian Zhang 0018, Zheng Lu 0002
EACL9
2023 Spatial Context-Aware Object-Attentional Network for Multi-Label Image Classification
abstract
Multi-label image classification is a fundamental but challenging task in computer vision. To tackle the problem, the label-related semantic information is often exploited, but the background context and spatial semantic information of related objects are not fully utilized. To address these issues, a multi-branch deep neural network is proposed in this paper. The first branch is designed to extract the discriminant information from regions of interest to detect target objects. In the second branch, a spatial context-aware approach is proposed to better capture the contextual information of an object in its surroundings by using an adaptive patch expansion mechanism. It helps the detection of small objects that are easily lost without the support of context information. The third one, the object-attentional branch, exploits the spatial semantic relations between the target object and its related objects, to better detect partially occluded, small or dim objects with the support of those easily detectable objects. To better encode such relations, an attention mechanism jointly considering the spatial and semantic relations between objects is developed. Two widely used benchmark datasets for multi-labeling classification, MS COCO and PASCAL VOC, are used to evaluate the proposed framework. The experimental results demonstrate that the proposed method outperforms the state-of-the-art methods for multi-label image classification.
Jialu Zhang 0003, Jianfeng Ren, Qian Zhang 0018, Jiang Liu 0001, Xudong Jiang 0001
IEEE Trans. Image Process.3
2022 Spatial-Context-Aware Deep Neural Network for Multi-Class Image Classification
abstract
Multi-label image classification is a fundamental but challenging task in computer vision. Over the past few decades, solutions exploring relationships between semantic labels have made great progress. However, the underlying spatial-contextual information of labels is under-exploited. To tackle this problem, a spatial-context-aware deep neural network is proposed to predict labels taking into account both semantic and spatial information. This proposed framework is evaluated on Microsoft COCO and PASCAL VOC, two widely used benchmark datasets for image multi-labelling. The results show that the proposed approach is superior to the state-of-the-art solutions on dealing with the multi-label image classification problem.
Jialu Zhang 0003, Qian Zhang 0018, Jianfeng Ren, Yitian Zhao, Jiang Liu 0001
ICASSP2
2022 Occlusion-Invariant Representation Alignment for Entity Re-Identification
abstract
Entity re-identification is the foundation of tracking- and matching-based computer vision tasks, which are widely employed in a variety of applications. However, when trained exclusively on clear images, the models capacity to generalize is significantly affected by the presence of occlusion at referencing time, whereas data argumentation-based approaches are costly to construct without guaranteeing a test-time improvement. To tackle this problem, we propose a domain adaptation framework based on learning representations that generates occlusion-invariant feature representations by aligning the clean image embedding distribution with the occluded one, using a disparity discrepancy metric derived from the siamese network architecture. Without the need for additional processing modules during the inference stage or an expensive occlusion-augmentation-enlarged dataset during the training stage, we could obtain occlusion invariant embeddings that are free of the impact of occluders. Extensive experimental results for two tasks across three datasets indicate the proposed method’s robustness and effectiveness to a variety of occlusions at all levels.
Zhanghao Jiang, Heshan Du, Huan Jin, Zheng Lu 0002, Qian Zhang 0018
ICIP6
2022 ICAF: Iterative Contrastive Alignment Framework for Multimodal Abstractive Summarization
abstract
Integrating multimodal knowledge for abstractive summarization task is a work-in-progress research area, with present techniques inheriting fusion-then-generation paradigm. Due to semantic gaps between computer vision and natural language processing, current methods often treat multiple data points as separate objects and rely on attention mechanisms to search for connection in order to fuse together. In addition, missing awareness of cross-modal matching from many frameworks leads to performance reduction. To solve these two drawbacks, we propose an Iterative Contrastive Alignment Framework (ICAF) that uses recurrent alignment and contrast to capture the coherences between images and texts. Specifically, we design a recurrent alignment (RA) layer to gradually investigate fine-grained semantical relationships between image patches and text tokens. At each step during the encoding process, crossmodal contrastive losses are applied to directly optimize the embedding space. According to ROUGE, relevance scores, and human evaluation, our model outperforms the state-of-the-art baselines on MSMO dataset. Experiments on the applicability of our proposed framework and hyperparameters settings have been also conducted.
Youxin Chen, Jing Xiao 0006, Qian Zhang 0018, Zheng Lu 0002
IJCNN5
2022 Rain-component-aware capsule-GAN for single image de-raining
Jianfeng Ren, Zheng Lu 0002, Jialu Zhang 0003, Qian Zhang 0018
Pattern Recognit.5
2021 Multi-scale capsule generative adversarial network for snow removal
abstract
Abstract Snowflakes captured on photos may severely decrease the visual quality and cause difficulties for vision analysis systems. Most noise removal frameworks are designed for de‐raining or de‐hazing, regarding rain or haze as translucent masks on clean images. However, snowflakes are different from them in terms of sizes, shapes, transparencies and floating trajectories, which decreases the performance of de‐raining or de‐hazing models in processing snowy images. In this work, we propose an effective multi‐scale generative adversarial network framework for single‐image snow removal, which is built with a multi‐scale structure to identify various scales of snowflakes and a capsule‐based structure to fuse the features extracted from the multi‐scale encoding branches, so that different scaled features could be summarised and learnt by a joint framework. The overall framework is supervised by a weighted joint loss with an iterative training procedure to keep the training stability for the multi‐branch‐based structure. The experimental results demonstrate that our model outperforms the state‐of‐the‐art comparisons.
Jialu Zhang 0003, Qian Zhang 0018
IET Comput. Vis.3
2019 Learning Spatial-Aware Cross-View Embeddings for Ground-to-Aerial Geolocalization
Rui Cao 0001, Jiasong Zhu, Qing Li 0029, Qian Zhang 0018, Qingquan Li 0001, Guoping Qiu
ICIG (1)4
2019 Salient Object Detection With Capsule-Based Conditional Generative Adversarial Network
abstract
Salient Object Detection (SOD) is one significant research area which is closely correlated to the attention of human beings. Most of the nowadays CNN-based approaches for SOD are based on an U-Net architecture. In this paper, we propose a novel capsule-based salient object detection framework by integrating the novel capsule blocks into both the generator and discriminator of GAN architecture. The experimental result showed that our approach is able to generate accurate saliency maps, which also highlighted the effectiveness of the capsule blocks. We also provide a challenging dataset that contains 3,299 images for SOD with difficult foreground objects and complex background contents.
Chao Zhang 0020, Guoping Qiu, Qian Zhang 0018
ICIP4
2019 FrameRank: A Text Processing Approach to Video Summarization
abstract
Video summarization has been extensively studied in the past decades. However, user-generated video summarization is much less explored since there lack large-scale video datasets within which human-generated video summaries are unambiguously defined and annotated. Toward this end, we propose a user-generated video summarization dataset - UGSum52 - that consists of 52 videos (207 minutes). In constructing the dataset, because of the subjectivity of user-generated video summarization, we manually annotate 25 summaries for each video, which are in total 1300 summaries. To the best of our knowledge, it is currently the largest dataset for user-generated video summarization. Based on this dataset, we present FrameRank, an unsupervised video summarization method that employs a frame-to-frame level affinity graph to identify coherent and informative frames to summarize a video. We use the Kullback-Leibler(KL)-divergence-based graph to rank temporal segments according to the amount of semantic information contained in their frames. We illustrate the effectiveness of our method by applying it to three datasets SumMe, TVSum and UGSum52 and show it achieves state-of-the-art results.
Zhuo Lei, Chao Zhang 0020, Qian Zhang 0018, Guoping Qiu
ICME3
2018 Capsule Based Image Synthesis for Interior Design Effect Rendering
Zheng Lu 0002, Guoping Qiu, Qian Zhang 0018
ACCV (5)5
2018 Quality Classified Image Analysis with Application to Face Detection and Recognition
abstract
Motion blur, out of focus, insufficient spatial resolution, lossy compression and many other factors can all cause an image to have poor quality. However, image quality is a largely ignored issue in traditional pattern recognition literature. In this paper, we use face detection and recognition as case studies to show that image quality is an essential factor which will affect the performances of traditional algorithms. We demonstrated that it is not the image quality itself that is the most important, but rather the quality of the images in the training set should have similar quality as those in the testing set. To handle real-world application scenarios where images with different kinds and severities of degradation can be presented to the system, we have developed a quality classified image analysis framework to deal with images of mixed qualities adaptively. We use deep neural networks first to classify images based on their quality classes and then design a separate face detector and recognizer for images in each quality class. We will present experimental results to show that our quality classified framework can accurately classify images based on the type and severity of image degradations and can significantly boost the performances of state-of-the-art face detector and recognizer in dealing with image datasets containing mixed quality images.
Qian Zhang 0018, Miaohui Wang, Guoping Qiu
ICPR2
2017 Learning deep semantic attributes for user video summarization
abstract
This paper presents a Semantic Attribute assisted video SUMmarization framework (SASUM). Compared with traditional methods, SASUM has several innovative features. Firstly, we use a natural language processing tool to discover a set of keywords from an image and text corpora to form the semantic attributes of visual contents. Secondly, we train a deep convolution neural network to extract visual features as well as predict the semantic attributes of video segments which enables us to represent video contents with visual and semantic features simultaneously. Thirdly, we construct a temporally constrained video segment affinity matrix and use a partially near duplicate image discovery technique to cluster visually and semantically consistent video frames together. These frame clusters can then be condensed to form an informative and compact summary of the video. We will present experimental results to show the effectiveness of the semantic attributes in assisting the visual features in video summarization and our new technique achieves state-of-the-art performance.
Ke Sun 0006, Jiasong Zhu, Zhuo Lei, Xianxu Hou, Qian Zhang 0018, Jiang Duan, Guoping Qiu
ICME5
2015 Bundling Centre for Landmark Image Discovery
abstract
This paper introduces a novel method to efficiently discover/cluster landmark images in large image collection. We consider each cluster as a combination of several sub-clusters, which is composed of images taken from different view points of the identical landmark. For each sub-cluster, we find its local centre represented by a group of similar images, and define it as the bundling centre (BC). We therefore start the image discovery/clustering by identifying the BCs and accomplish the task by efficiently growing and merging those sub-clusters represented by different BCs. In our proposed method, we use min-hash based method to build a sparse graph so as to avoid the time-consuming full-scale exhaustive pairwise image matching. Based on the information provided by the sparse graph, BCs are identified as local dense neighbors sharing high intra-similarity. We have also proposed a weighted voting method to efficiently grow these BCs with high accuracy. More importantly, the fixed local centres can ensure each sub-cluster contains identical landmark and generate result with high precision. In addition, compared to a single representative (iconic) image, the group of similar images obtained by each BC can provide more comprehensive cluster information and, thus, overcome the problem of low recall caused by information lost during visual word quantization. We present experimental results on two landmark datasets and show that, without query expansion, our method can boost landmark image discovery/clustering performances of state of the art techniques.
Qian Zhang 0018, Guoping Qiu
ICMR1
2013 Interactive skin condition recognition
abstract
It is believed that there are between 1000 to 2000 skin conditions, and about 20% are difficult to diagnose. An intelligent system capable of making accurate diagnosis not only helps patients in places where access to health services are scarce, but also benefits typical general practitioners who have received minimal dermatology training. In this paper, we introduce a challenging dataset developed by gathering 2309 images from 44 different skin conditions, and collecting answers to simple perceptual questions from 361 “Amazon Mechanical Turk” workers. We also propose a method based on random forest technology that combines visual features of the skin lesion images with user provided answers to achieve promising recognition rates. We believe that our solution can be potentially improved and installed on smart phones and tablets to enhance quality of life in patients across the world.
Orod Razeghi, Qian Zhang 0018, Guoping Qiu
ICME2
2013 Tree partition voting min-hash for partial duplicate image discovery
abstract
Discovering partially duplicated images such as those of the same scenes, buildings or objects taken from different angles, distances and vantage points can be very useful in applications such as managing large image repositories and image search on the Internet. In this paper, we present a novel technique for partial duplicate image discovery. The new technique, termed tree partition voting min-hash (TmH), first partitions interest points within an image based on their geometric or photometric (appearance) properties using a spatial partition tree data structure and then finds potential partial duplicate images through a traditional partition min-hash (PmH) method [1]. We have developed a k-d tree partition min-hash (kdTmH) and a random projection tree partition min-hash (rpTmH) technique and have also developed a weighted voting algorithm for improving the similarity measure of a pair of hashing sketches. We present experimental results on 3 datasets and show that TmH significantly outperforms PmH in terms of recall and precision performances without increasing complexity and that the new voting algorithms performs better than sketch matching techniques in the literature.
Qian Zhang 0018, Hao Fu 0001, Guoping Qiu
ICME1
2012 Random Forest for Image Annotation
Hao Fu 0001, Qian Zhang 0018, Guoping Qiu
ECCV (6)2