Jialu Zhang 0003

dblp:38/1392-3 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0001-9539-6789ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 GCA-SUNet: A Gated Context-Aware Swin-UNet for Exemplar-Free Counting
abstract
Exemplar-Free Counting aims to count objects of interest without intensive annotations of objects or exemplars. To achieve this, we propose a Gated Context-Aware Swin-UNet (GCA-SUNet) to directly map an input image to the density map of countable objects. Specifically, a set of Swin transformers form an encoder to derive a robust feature representation, and a Gated Context-Aware Modulation block is designed to suppress irrelevant objects or background through a gate mechanism and exploit the attentive support of objects of interest through a self-similarity matrix. The gate strategy is also incorporated into the bottleneck network and the decoder of the Swin-UNet to highlight the features most relevant to objects of interest. By explicitly exploiting the attentive support among countable objects and eliminating irrelevant features through the gate mechanisms, the proposed GCA-SUNet focuses on and counts objects of interest without relying on predefined categories or exemplars. Experimental results on the real-world datasets such as FSC-147 and CARPK demonstrate that GCA-SUNet significantly and consistently outperforms state-of-the-art methods. The code is available at https://github.com/Amordia/GCA-SUNet.
Yipeng Xu, Jialu Zhang 0003, Jianfeng Ren, Xudong Jiang 0001
ICME4
2025 CEARI: Co-Evolutionary Agents for Reassembling and Inpainting Puzzles with Gaps and Missing Pieces
abstract
Puzzle solving has recently become a popular research topic. Existing solvers often overlook puzzles with missing pieces. The missing pieces, together with gaps between pieces, pose significant challenges, amplified by a large solution space. To tackle the challenges, we propose Co-Evolutionary Agents for Reassembling and Inpainting (CEARI), one agent to inpaint missing contents and the other to reassemble the puzzle, with a shared perception network to perceive the puzzle status. The reassembly agent utilizes an evolutionary algorithm to explore the large solution space, to discover a sequence of fragment-swapping actions to efficiently reassemble the puzzle, while the inpainting agent evolves from using a local outpainting network at the early stage to using a global inpainting network at the latter stage. Furthermore, a co-evolutionary training paradigm is designed to iteratively evolve the two agents in a coherent and collaborative manner, improving reassembly accuracy and inpainting quality simultaneously. Experimental results on three datasets show that CEARI largely outperforms state-of-the-art methods in terms of both reassembly accuracy and inpainting quality.
Xingke Song, Jianxu Shangguan, Yiran Li 0003, Jialu Zhang 0003, Jianfeng Ren, Ruibin Bai, Xin Chen 0003, Xudong Jiang 0001
ACM Multimedia4
2024 Scale Optimization Using Evolutionary Reinforcement Learning for Object Detection on Drone Imagery
abstract
Object detection in aerial imagery presents a significant challenge due to large scale variations among objects. This paper proposes an evolutionary reinforcement learning agent, integrated within a coarse-to-fine object detection framework, to optimize the scale for more effective detection of objects in such images. Specifically, a set of patches potentially containing objects are first generated. A set of rewards measuring the localization accuracy, the accuracy of predicted labels, and the scale consistency among nearby patches are designed in the agent to guide the scale optimization. The proposed scale-consistency reward ensures similar scales for neighboring objects of the same category. Furthermore, a spatial-semantic attention mechanism is designed to exploit the spatial semantic relations between patches. The agent employs the proximal policy optimization strategy in conjunction with the evolutionary strategy, effectively utilizing both the current patch status and historical experience embedded in the agent. The proposed model is compared with state-of-the-art methods on two benchmark datasets for object detection on drone imagery. It significantly outperforms all the compared methods. Code is available at https://github.com/UNNC-CV/EvOD/.
Jialu Zhang 0003, Jianfeng Ren, Qian Zhang 0018, Yitian Zhao, Ruibin Bai, Xiangjian He, Jiang Liu 0001
AAAI1
2024 Visual-linguistic Cross-domain Feature Learning with Group Attention and Gamma-correct Gated Fusion for Extracting Commonsense Knowledge
abstract
Acquiring commonsense knowledge about entity-pairs from images is crucial across diverse applications. Distantly supervised learning has made significant advancements by automatically retrieving images containing entity pairs and summarizing commonsense knowledge from the bag of images. However, the retrieved images may not always cover all possible relations, and the informative features across the bag of images are often overlooked. To address these challenges, a Multi-modal Cross-domain Feature Learning framework is proposed to incorporate the general domain knowledge from a large vision-text foundation model, ViT-GPT2, to handle unseen relations and exploit complementary information from multiple sources. Then, a Group Attention module is designed to exploit the attentive information from other instances of the same bag to boost the informative features of individual instances. Finally, a Gamma-corrected Gated Fusion is designed to select a subset of informative instances for a comprehensive summarization of commonsense entity relations. Extensive experimental results demonstrate the superiority of the proposed method over state-of-the-art models for extracting commonsense knowledge.
Jialu Zhang 0003, Chenglin Yao, Jianfeng Ren, Xudong Jiang 0001
ACM Multimedia1
2023 Hierarchical ConViT with Attention-Based Relational Reasoner for Visual Analogical Reasoning
abstract
Raven’s Progressive Matrices (RPMs) have been widely used to evaluate the visual reasoning ability of humans. To tackle the challenges of visual perception and logic reasoning on RPMs, we propose a Hierarchical ConViT with Attention-based Relational Reasoner (HCV-ARR). Traditional solution methods often apply relatively shallow convolution networks to visually perceive shape patterns in RPM images, which may not fully model the long-range dependencies of complex pattern combinations in RPMs. The proposed ConViT consists of a convolutional block to capture the low-level attributes of visual patterns, and a transformer block to capture the high-level image semantics such as pattern formations. Furthermore, the proposed hierarchical ConViT captures visual features from multiple receptive fields, where the shallow layers focus on the image fine details while the deeper layers focus on the image semantics. To better model the underlying reasoning rules embedded in RPM images, an Attention-based Relational Reasoner (ARR) is proposed to establish the underlying relations among images. The proposed ARR well exploits the hidden relations among question images through the developed element-wise attentive reasoner. Experimental results on three RPM datasets demonstrate that the proposed HCV-ARR achieves a significant performance gain compared with the state-of-the-art models. The source code is available at: https://github.com/wentaoheunnc/HCV-ARR.
Jialu Zhang 0003, Jianfeng Ren, Ruibin Bai, Xudong Jiang 0001
AAAI2
2023 Spatial Context-Aware Object-Attentional Network for Multi-Label Image Classification
abstract
Multi-label image classification is a fundamental but challenging task in computer vision. To tackle the problem, the label-related semantic information is often exploited, but the background context and spatial semantic information of related objects are not fully utilized. To address these issues, a multi-branch deep neural network is proposed in this paper. The first branch is designed to extract the discriminant information from regions of interest to detect target objects. In the second branch, a spatial context-aware approach is proposed to better capture the contextual information of an object in its surroundings by using an adaptive patch expansion mechanism. It helps the detection of small objects that are easily lost without the support of context information. The third one, the object-attentional branch, exploits the spatial semantic relations between the target object and its related objects, to better detect partially occluded, small or dim objects with the support of those easily detectable objects. To better encode such relations, an attention mechanism jointly considering the spatial and semantic relations between objects is developed. Two widely used benchmark datasets for multi-labeling classification, MS COCO and PASCAL VOC, are used to evaluate the proposed framework. The experimental results demonstrate that the proposed method outperforms the state-of-the-art methods for multi-label image classification.
Jialu Zhang 0003, Jianfeng Ren, Qian Zhang 0018, Jiang Liu 0001, Xudong Jiang 0001
IEEE Trans. Image Process.1
2022 Spatial-Context-Aware Deep Neural Network for Multi-Class Image Classification
abstract
Multi-label image classification is a fundamental but challenging task in computer vision. Over the past few decades, solutions exploring relationships between semantic labels have made great progress. However, the underlying spatial-contextual information of labels is under-exploited. To tackle this problem, a spatial-context-aware deep neural network is proposed to predict labels taking into account both semantic and spatial information. This proposed framework is evaluated on Microsoft COCO and PASCAL VOC, two widely used benchmark datasets for image multi-labelling. The results show that the proposed approach is superior to the state-of-the-art solutions on dealing with the multi-label image classification problem.
Jialu Zhang 0003, Qian Zhang 0018, Jianfeng Ren, Yitian Zhao, Jiang Liu 0001
ICASSP1
2022 Rain-component-aware capsule-GAN for single image de-raining
Jianfeng Ren, Zheng Lu 0002, Jialu Zhang 0003, Qian Zhang 0018
Pattern Recognit.4
2021 rPPG-Based Spoofing Detection for Face Mask Attack using Efficientnet on Weighted Spatial-Temporal Representation
abstract
Face spoofing detection against paper attack and video-replay attack has been well studied, whereas detecting 3D face mask attack remains challenging. Remote photoplethysmography (rPPG) signal is a recently developed liveness clue for face-spoofing detection. The main challenge of existing rPPG-based methods is that the signal can be easily distorted by background noise or object motion. To address this problem, in this work, we propose an rPPG-based face-spoofing detection method using multiple regions of interests (ROIs) covering entire face, and emphasize the regions containing richer rPPG signals using larger weights. The rPPG signals of these regions form a weighted spatial-temporal map. In view of the discriminant power of EfficientNet over other deep convolutional neural networks, we propose a domain-specific EfficientNet as the classification method. Extensive experiments on two databases namely 3DMAD and HKBU-Mars V2 demonstrate the superior performance of the proposed method over state-of-the-art rPPG-based face-spoofing-detection algorithms.
Chenglin Yao, Shihe Wang, Jialu Zhang 0003, Heshan Du, Jianfeng Ren, Ruibin Bai, Jiang Liu 0001
ICIP3
2021 Multi-scale capsule generative adversarial network for snow removal
abstract
Abstract Snowflakes captured on photos may severely decrease the visual quality and cause difficulties for vision analysis systems. Most noise removal frameworks are designed for de‐raining or de‐hazing, regarding rain or haze as translucent masks on clean images. However, snowflakes are different from them in terms of sizes, shapes, transparencies and floating trajectories, which decreases the performance of de‐raining or de‐hazing models in processing snowy images. In this work, we propose an effective multi‐scale generative adversarial network framework for single‐image snow removal, which is built with a multi‐scale structure to identify various scales of snowflakes and a capsule‐based structure to fuse the features extracted from the multi‐scale encoding branches, so that different scaled features could be summarised and learnt by a joint framework. The overall framework is supervised by a weighted joint loss with an iterative training procedure to keep the training stability for the multi‐branch‐based structure. The experimental results demonstrate that our model outperforms the state‐of‐the‐art comparisons.
Jialu Zhang 0003, Qian Zhang 0018
IET Comput. Vis.2