Jian-Jun Qiao

dblp:304/1390 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-4282-0149ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PSGNet: Pure Smoke Image Generation With Gradient and Style Learning
abstract
The realistic and controllable generation of pure smoke is critical for smoke image editing, smoke visual special effects generation, and smoke data synthesizing within security scenarios. It is a relatively underexplored topic and continues to present significant challenges. Existing methods face challenges in the generation of smoke with intricate details and the regulation of various smoke styles. In this paper, a Pure Smoke image Generation Network (PSGNet) is proposed with a gradient and style learning approach to generate realistic and controllable smoke images. To achieve flexibility in control across the spatial dimension, the smoke shape mask is used to encode spatial details, such as the location and contour of the smoke, along with other related properties. To enhance the physical realism of synthesized smoke, a novel gradient-based learning framework is proposed to generate smoke gradient features, highlighting a special focus on explicitly encoding and exploiting gradient information. This framework uses a smoke gradient learning architecture that captures the subtle structures and patterns characteristic of real smoke, enabling the generation of highly realistic smoke with rich, fine-scale detail. In addition, a spatially aware style learning strategy is proposed to provide fine-grained control over smoke attributes such as density, color, and overall look. It is able to effectively model style features across both channel and spatial dimensions, thereby enabling spatially aware style manipulation. By combining the gradient module with this style learning framework, the method produces smoke that exhibits rich visual details and customizable image styles. Experiments conducted on six benchmark datasets demonstrate that the proposed PSGNet significantly outperforms the state-of-the-art approaches.
Jian-Jun Qiao, Xiao Wu 0001, Zhi-Qi Cheng, Wei Li 0110, Zhaoquan Yuan
IEEE Trans. Vis. Comput. Graph.1
2024 PostureHMR: Posture Transformation for 3D Human Mesh Recovery
abstract
Human Mesh Recovery (HMR) aims to estimate the 3D human body from 2D images, which is a challenging task due to inherent ambiguities in translating 2D observations to 3D space. A novel approach called PostureHMR is pro-posed to leverage a multi-step diffusion-style process, which converts this task into a posture transformation from an SMPL T-pose mesh to the target mesh. To inject the learning process of posture transformation with the physical structure of the human body model, a kinematics-based forward process is proposed to interpolate the intermediate state with pose and shape decomposition. Moreover, a mesh-to-posture (M2P) decoder is designed, by combining the in-put of 3D and 2D mesh constraints estimated from the im-age to model the posture changes in the reverse process. It mitigates the difficulties of posture change learning directly from RGB pixels. To overcome the limitation of pixel-level misalignment of modeling results with the input image, a new trimap-based rendering loss is designed to highlight the areas with poor recognition. Experiments conducted on three widely used datasets demonstrate that the proposed approach outperforms the state-of-the-art methods.
Yu-Pei Song, Xiao Wu 0001, Zhaoquan Yuanl, Jian-Jun Qiao, Qiang Peng
CVPR4
2024 CAPNet: Cartoon Animal Parsing with Spatial Learning and Structural Modeling
abstract
Cartoon animal parsing aims to segment the body parts such as heads, arms, legs and tails of cartoon animals. Different from previous parsing tasks, cartoon animal parsing faces new challenges, including irregular body structures, abstract drawing styles and diverse animal categories. Existing methods have difficulties when addressing these challenges caused by the spatial and structural properties of cartoon animals. To address these challenges, a novel spatial learning and structural modeling network, named CAPNet, is proposed for cartoon animal parsing. It aims to address the critical problems of spatial perception, structure modeling and spatial-structural consistency learning. A spatial-aware learning module integrates deformable convolutions to learn spatial features of diverse cartoon animals. The multi-task edge and center point prediction mechanism is incorporated to capture the intricate spatial patterns. A structural modeling method is proposed to model the complex structural representations of cartoon animals, which integrates a graph neural network with a shape-aware relation learning module. To mitigate the significant differences among animals, a spatial and structural consistency learning strategy is proposed to capture and learn feature correlations across different animal species. Extensive experiments conducted on benchmark datasets demonstrate the effectiveness of the proposed approach, which outperforms the state-of-the-art methods.
Jian-Jun Qiao, Meng-Yu Duan, Xiao Wu 0001, Wei Li 0110
ACM Multimedia1
2024 CartoonNet: Cartoon Parsing with Semantic Consistency and Structure Correlation
abstract
Cartoon parsing is an important task for cartoon-centric applications, which segments the body parts of cartoon images. Due to the complex appearances, abstract drawing styles, and irregular structures of cartoon characters, cartoon parsing remains a challenging task. In this paper, a novel approach, named CartoonNet, is proposed for cartoon parsing, in which semantic consistency and structure correlation are integrated to address the visual diversity and structural complexity for cartoon parsing. A memory-based semantic consistency module is designed to learn the diverse appearances exhibited by cartoon characters. The memory bank stores features of diverse samples and retrieves the samples related to new samples for consistency, which aims to improve the semantic reasoning capability of the network. A self-attention mechanism is employed to conduct consistency learning among diverse body parts belong to the retrieved samples and new samples. To capture the intricate structural information of cartoon images, a structure correlation module is proposed. Leveraging graph attention networks and a main body-aware mechanism, the proposed approach enables structural correlation, allowing it to parse cartoon images with complex structures. Experiments conducted on cartoon parsing and human parsing datasets demonstrate the effectiveness of the proposed method, which outperforms the state-of-the-art approaches for cartoon parsing and achieves competitive performance on human parsing.
Jian-Jun Qiao, Meng-Yu Duan, Xiao Wu 0001, Yu-Pei Song
ACM Multimedia1
2023 CPNet: Cartoon Parsing with Pixel and Part Correlation
abstract
Cartoon parsing, the task of segmenting constituent parts such as heads, arms, and legs of cartoon characters, holds substantial significance for applications in the animation industry and emerging metaverse. Nonetheless, this domain presents considerable challenges stemming from complex visual appearances, irregular structures, abstract drawing styles, among other factors. In this paper, a novel Cartoon Parsing Network (CPNet) is introduced to address these challenges. CPNet skillfully leverages the spatial and semantic correlations of pixels to discern intricate and visually akin appearances. Furthermore, it employs both local and global correlations of constituent parts to differentiate irregular and abstract body sections. Specifically, the pixels of the cartoon image are interconnected by capitalizing on the spatial and semantic correlations. To this end, a center point predictor, working in tandem with a pixel-aware attention, facilitates the exploration of pixel-level correlation learning. Additionally, the various constituent parts are meticulously organized to resonate with the intrinsic physiological structure of a cartoon character. The character's graph structure is assembled and analyzed by an edge-aware graph neural network, thereby linking adjacent parts and assimilating local correlations. A part-guided non-local attention mechanism is fashioned to correlate individual parts with the entire body, thereby modeling global connections. In addition, a new dataset named CartoonSet is curated and annotated explicitly for cartoon parsing. Experiments carried out on both cartoon parsing and human parsing datasets yield compelling results, thereby attesting to the efficacy and innovativeness of the proposed method.
Jian-Jun Qiao, Jie Zhang 0179, Xiao Wu 0001, Yu-Pei Song, Wei Li 0110
ACM Multimedia1
2022 Learning Action-guided Spatio-temporal Transformer for Group Activity Recognition
abstract
Learning spatial and temporal relations among people plays an important role in recognizing group activity. Recently, transformer-based methods have become popular solutions due to the proposal of self-attention mechanism. However, the person-level features are fed directly into the self-attention module without any refinement. Moreover, group activity in a clip often involves unbalanced spatio-temporal interactions, where only a few persons with special actions are critical to identifying different activities. It is difficult to learn the spatio-temporal interactions due to the lack of elaborately modeling the action dependencies among all people. In this paper, a novel Action-guided Spatio-Temporal transFormer (ASTFormer) is proposed to capture the interaction relations for group activity recognition by learning action-centric aggregation and modeling spatio-temporal action dependencies. Specifically, ASTFormer starts with assigning all persons in each frame to the latent actions, while an action-centric aggregation strategy is performed by weighting the sum of residuals for each latent action under the supervision of global action information. Then, a dual-branch transformer is proposed to refine the inter- and intra-frame action-level features, where two encoders with the self-attention mechanism are employed to select important tokens. Next, a semantic action graph is explicitly devised to model the dynamic action-wise dependencies. Finally, our model is capable of boosting group activity recognition by fusing these important cues, while only requiring video-level action labels. Extensive experiments on two popular benchmarks (Volleyball and Collective Activity) demonstrate the superior performance of our method in comparison with the state-of-the-art methods using only raw RGB frames as input.
Wei Li 0110, Tianzhao Yang, Xiao Wu 0001, Xian-Jun Du, Jian-Jun Qiao
ACM Multimedia5
2022 Real-time Semantic Segmentation with Parallel Multiple Views Feature Augmentation
abstract
Real-time semantic segmentation is essential for many practical applications, which utilizes attention-based feature aggregation into lightweight structures to improve accuracy and efficiency. However, existing attention-based methods ignore 1) high-level and low-level feature augmentation guided by spatial information, and 2) low-level feature augmentation guided by semantic context, so that feature gaps between multi-level features and noise of low-level spatial details still exist. To address these problems, a new real-time semantic segmentation network, called MvFSeg, is proposed. In MvFSeg, parallel convolution with multiple depths is designed as a context head to generate and integrate multi-view features with larger receptive fields. Moreover, MvFSeg designs multiple views feature augmentation strategies that exploit spatial and semantic guidance for shallow and deep feature augmentation in an inter-layer and intra-layer manner. These strategies eliminate feature gaps between multi-level features, filter out the noise of spatial details, and provide spatial and semantic guidance for multi-level features. By combining multi-view features and augmented features from the lightweight networks with progressive dense aggregation structures, MvFSeg effectively captures invariance at various scales and generates high-quality segmentation results. Experiments conducted on Cityscapes and CamVid benchmark show that MvFSeg outperforms existing state-of-the-art methods.
Jian-Jun Qiao, Zhi-Qi Cheng, Xiao Wu 0001, Wei Li 0110, Ji Zhang 0027
ACM Multimedia1
2022 CrossNet: Boosting Crowd Counting with Localization
abstract
Generating high-quality density maps is a crucial step in crowd counting. It is obvious that exploiting the head location of the people can naturally highlight the crowded area and eliminate the interference of background noise. However, existing crowd counting methods are still tricky to reasonably use location in density generation. In this paper, a novel location-guided framework named CrossNet is proposed for crowd counting, which integrates location supervision into density maps through dual-branch joint training. First, a new branching network is proposed to localize the potential positions of pedestrians. With the help of supervision induced from the localization branch, Location Enhancement (LE) module is designed to obtain high-quality density maps by positioning foreground regions. Second, Adaptive Density Awareness Attention (ADAA) module is engaged to enhance localization accuracy, which can efficiently use the density of the counting branch to adaptively capture the error-prone dense areas of the location maps. Finally, Density Awareness Localization (DAL) loss is offered to allocate attention to the crowd density levels, which delivers more focus on regions with high densities and less concentration on areas with low densities. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art approaches both in crowd counting and crowd localization.
Ji Zhang 0027, Zhi-Qi Cheng, Xiao Wu 0001, Wei Li 0110, Jian-Jun Qiao
ACM Multimedia5
2022 SWNet: A Deep Learning Based Approach for Splashed Water Detection on Road
abstract
Adverse weather conditions seriously threaten the traffic safety, especially for rainy days with the ponding water on the road surface, which potentially result in vehicle crashes, person injuries and crash fatalities. Automatic splashed water detection based on surveillance videos is an attractive way to effectively prevent the traffic accidents. However, surveillance videos exhibit great variations with lighting changes, illumination conditions and complex backgrounds, which pose great difficulties in automatic recognition. In this paper, a novel deep learning based approach is proposed to detect the splashed water. To the best of our knowledge, this is the first work on this topic based on deep learning. An effective semantic segmentation network, called SWNet, is novelly proposed to extract the potential splashed water regions. An encoder-decoder structure is designed to capture the visual characteristics of splashed water. SWNet achieves high efficiency by reusing pooling indices and adopting the light-weight decoder. With the multi-scale feature fusion structure, SWNet integrates the coarse semantic information and detailed appearance information, which significantly boosts the accuracy and refines the edge segmentation. A weighted cross entropy loss for splashed water is adopted to cope with the unbalanced distribution between splashed water and backgrounds. Moreover, a splashed water attention module is designed to focus on the salient regions of moving vehicles and splashed water, by performing attention mechanism to integrate global contextual information in semantic segmentation. Experiments conducted on a newly collected splashed water dataset demonstrate the effectiveness and efficiency of the proposed approach, which outperforms the state-of-the-art methods.
Jian-Jun Qiao, Xiao Wu 0001, Jun-Yan He, Wei Li 0110, Qiang Peng
IEEE Trans. Intell. Transp. Syst.1
2021 Vehicle Counting Network with Attention-based Mask Refinement and Spatial-awareness Block Loss
abstract
Vehicle counting aims to calculate the number of vehicles in congested traffic scenes. Although object detection and crowd counting have made tremendous progress with the development of deep learning, vehicle counting remains a challenging task, due to scale variations, viewpoint changes, inconsistent location distributions, diverse visual appearances and severe occlusions. In this paper, a well-designed Vehicle Counting Network (VCNet) is novelly proposed to alleviate the problem of scale variation and inconsistent spatial distribution in congested traffic scenes. Specifically, VCNet is composed of two major components: (i) To capture multi-scale vehicles across different types and camera viewpoints, an effective multi-scale density map estimation structure is designed by building an attention-based mask refinement module. The multi-branch structure with hybrid dilated convolution blocks is proposed to assign receptive fields to generate multi-scale density maps. To efficiently aggregate multi-scale density maps, the attention-based mask refinement is well-designed to highlight the vehicle regions, which enables each branch to suppress the scale interference from other branches. (ii) In order to capture the inconsistent spatial distributions, a spatial-awareness block loss (SBL) based on the region-weighted reward strategy is proposed to calculate the loss of different spatial regions including sparse, congested and occluded regions independently by dividing the density map into different regions. Extensive experiments conducted on three benchmark datasets, TRANCOS, VisDrone2019 Vehicle and CVCSet demonstrate that the proposed VCNet outperforms the state-of-the-art approaches in vehicle counting. Moreover, the proposed idea can be applicable for crowd counting, which produces competitive results on ShanghaiTech crowd counting dataset.
Ji Zhang 0027, Jian-Jun Qiao, Xiao Wu 0001, Wei Li 0110
ACM Multimedia2