Chaoqun Wang 0011

dblp:41/3693-11 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-4649-5518ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2026 S²Flow: Towards Fast and Authentic Training-Free High-Resolution Video Generation
abstract
Rectified flow models have shown strong potential in high-fidelity video generation, yet extending them to high-resolution remains challenging due to the high cost of full attention and error accumulation in the ODE-solving process. In this paper, we propose S^2Flow, a training-free framework that enables efficient and authentic high-resolution video generation by jointly exploring Flow-guided Sparse attention and Second-order ODE solution. Specifically, S^2Flow exploits and transfers the semantic and structural information from the low-resolution flow trajectory to guide the high-resolution flow in two aspects. First, S^2Flow dynamically captures the sparse patterns of the spatio-temporal attention maps from low-resolution videos to construct localized 3D windows, enabling efficient window attention in high-resolution inference. This can significantly reduce redundant computation while preserving contextual dependencies. Second, S^2Flow adopts a second-order ODE solver based on Taylor expansion, where the high-order derivative is approximated via central difference from the low-resolution flow, facilitating accurate high-resolution denoising. Extensive experiments on VBench dataset demonstrate that S^2Flow outperforms prior methods in both visual quality and inference speed, enabling 4x acceleration on 2560x1536 video generation.
Chaoqun Wang 0011, Shaobo Min, Xu Yang 0019
AAAI1
2025 CausalCtrl: Causality-Aware Control Framework for Text-Guided Visual Editing
abstract
Text-guided visual editing aims to modify visual content according to a target prompt while faithfully preserving the structure and identity of the source image or video. However, existing methods ignore confounding effects brought from the pretrained model, i.e., harmful biases learned from the pretraining datasets, leading to spurious correlations during the editing processing. To address this issue, we introduce CausalCtrl, a novel training-free framework that reformulates text-guided visual editing from a causal inference perspective. The core idea is to leverage frontdoor adjustment to estimate the interventional distribution of the output, effectively blocking the influence of hidden confounders introduced by the pretrained model. Specifically, we first design a dual-branch inversion mechanism that disentangles the source content and target semantics into two separate latent embeddings to simplify the sampling space of interventional operation, and perform unbiased denoising through their controlled interaction. Besides, we propose a Structured Attention Injection Module (SAIM) that adaptively identifies and amplifies dominant attention heads using a lightweight SVD-based top-K selection strategy. Extensive experiments on several challenging image and video editing benchmarks demonstrate that CausalCtrl consistently outperforms existing methods in both target semantic alignment and source content preservation, validating the effectiveness of causal intervention in this task.
Haoxiang Cao, Chaoqun Wang 0011, Yongwen Lai, Shaobo Min, Xuejin Chen
ACM Multimedia2
2025 Robust visual place recognition with adaptive deformable token aggregation
Chaoqun Wang 0011, Shaobo Min, Xuejin Chen
Comput. Graph.1
2025 Region-Enhanced Feature Learning for Scene Semantic Segmentation
abstract
Semantic segmentation in complex scenes relies not only on object appearance but also on object location and the surrounding environment. Nonetheless, it is difficult to model long-range context in the format of pairwise point correlations due to the huge computational cost for large-scale point clouds. In this paper, we propose using regions as the intermediate representation of point clouds instead of fine-grained points or voxels to reduce the computational burden. We introduce a novel Region-Enhanced Feature Learning Network (REFL-Net) that leverages region correlations to enhance point feature learning. We design a region-based feature enhancement (RFE) module, which consists of a Semantic-Spatial Region Extraction stage and a Region Dependency Modeling stage. In the first stage, the input points are grouped into a set of regions based on their semantic and spatial proximity. In the second stage, we explore inter-region semantic and spatial relationships by employing a self-attention block on region features and then fuse point features with the region features to obtain more discriminative representations. Our proposed RFE module is plug-and-play and can be integrated with common semantic segmentation backbones. We conduct extensive experiments on ScanNetV2 and S3DIS datasets and evaluate our RFE module with different segmentation backbones. Our REFL-Net achieves 1.8% mIoU gain on ScanNetV2 and 1.7% mIoU gain on S3DIS with negligible computational cost compared with backbone models. Both quantitative and qualitative results show the powerful long-range context modeling ability and strong generalization ability of our REFL-Net.
Chaoqun Wang 0011, Xuejin Chen
IEEE Trans. Multim.2
2023 Semantics-Preserving Sketch Embedding for Face Generation
abstract
With recent advances in image-to-image translation tasks, remarkable progress has been witnessed in generating face images from sketches. However, existing methods frequently fail to generate images with details that are semantically and geometrically consistent with the input sketch, especially when various decoration strokes are drawn. To address this issue, we introduce a novel$\mathcal {W}$-$\mathcal {W^+}$encoder architecture to take advantage of the high expressive power of$\mathcal {W^+}$space and semantic controllability of$\mathcal {W}$space. We introduce an explicit intermediate representation for sketch semantic embedding. With a semantic feature matching loss for effective semantic supervision, our sketch embedding precisely conveys the semantics in the input sketches to the synthesized images. Moreover, a novel sketch semantic interpretation approach is designed to automatically extract semantics from vectorized sketches. We conduct extensive experiments on both synthesized sketches and hand-drawn sketches, and the results demonstrate the superiority of our method over existing approaches on both semantics-preserving and generalization ability.
Binxin Yang, Xuejin Chen, Chaoqun Wang 0011, Chi Zhang 0044, Xiaoyan Sun 0001
IEEE Trans. Multim.3
2021 Task-Independent Knowledge Makes for Transferable Representations for Generalized Zero-Shot Learning
abstract
Generalized Zero-Shot Learning (GZSL) targets recognizing new categories by learning transferable image representations. Existing methods find that, by aligning image representations with corresponding semantic labels, the semantic-aligned representations can be transferred to unseen categories. However, supervised by only seen category labels, the learned semantic knowledge is highly task-specific, which makes image representations biased towards seen categories. In this paper, we propose a novel Dual-Contrastive Embedding Network (DCEN) that simultaneously learns task-specific and task-independent knowledge via semantic alignment and instance discrimination. First, DCEN leverages task labels to cluster representations of the same semantic category by cross-modal contrastive learning and exploring semantic-visual complementarity. Besides task-specific knowledge, DCEN then introduces task-independent knowledge by attracting representations of different views of the same image and repelling representations of different images. Compared to high-level seen category supervision, this instance discrimination supervision encourages DCEN to capture low-level visual knowledge, which is less biased toward seen categories and alleviates the representation bias. Consequently, the task-specific and task-independent knowledge jointly make for transferable representations of DCEN, which obtains averaged 4.1% improvement on four public benchmarks.
Chaoqun Wang 0011, Xuejin Chen, Shaobo Min, Xiaoyan Sun 0001, Houqiang Li
AAAI1
2021 Dual Progressive Prototype Network for Generalized Zero-Shot Learning
abstract
Generalized Zero-Shot Learning (GZSL) aims to recognize new categories with auxiliary semantic information, e.g., category attributes. In this paper, we handle the critical issue of domain shift problem, i.e., confusion between seen and unseen categories, by progressively improving cross-domain transferability and category discriminability of visual representations. Our approach, named Dual Progressive Prototype Network (DPPN), constructs two types of prototypes that record prototypical visual patterns for attributes and categories, respectively. With attribute prototypes, DPPN alternately searches attribute-related local regions and updates corresponding attribute prototypes to progressively explore accurate attribute-region correspondence. This enables DPPN to produce visual representations with accurate attribute localization ability, which benefits the semantic-visual alignment and representation transferability. Besides, along with progressive attribute localization, DPPN further projects category prototypes into multiple spaces to progressively repel visual representations from different categories, which boosts category discriminability. Both attribute and category prototypes are collaboratively learned in a unified framework, which makes visual representations of DPPN transferable and distinctive.Experiments on four benchmarks prove that DPPN effectively alleviates the domain shift problem in GZSL.
Chaoqun Wang 0011, Shaobo Min, Xuejin Chen, Xiaoyan Sun 0001, Houqiang Li
NeurIPS1
2021 Structure-Guided Deep Video Inpainting
abstract
A fundamental challenge in video inpainting is the difficulty of generating video contents with fine details, while keeping spatio-temporal coherence in the missing region. Recent studies focus on synthesizing temporally smooth pixels by exploiting the flow information, while ignoring maintaining the semantic structural coherence between frames. This makes them suffer from over-smoothing and blurry contours, which significantly reduce the visual quality of inpainting results. To address this issue, we present a novel structure-guided video inpainting approach that enhances temporal structure coherence to improve video inpainting results. In contrast to directly synthesizing the missing pixel colors, we first complete edges in the missing regions to depict scene structures and object shapes via an edge inpainting network with 3D convolutions. Then, we replenish textures using a coarse-to-fine synthesis network with a structure attention module (SAM), under the guidance of the synthesized edges. Specifically, our SAM is designed to model the semantic correlation between video textures and structural edges to generate more realistic content. Besides, motion flows between neighboring frames are employed to enhance temporal consistency for self-supervision during training the edge inpainting and texture inpainting modules. Consequently, the inpainting results using our approach are visually pleasing with fine details and temporal coherence. Experiments on the YouTubeVOS, DAVIS, and 300VW datasets show that our method obtains state-of-the-art performance under diverse video inpainting settings.
Chaoqun Wang 0011, Xuejin Chen, Shaobo Min, Jiaping Wang, Zhengjun Zha
IEEE Trans. Circuits Syst. Video Technol.1
2020 Domain-Aware Visual Bias Eliminating for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning aims to recognize images from seen and unseen domains. Recent methods focus on learning a unified semantic-aligned visual representation to transfer knowledge between two domains, while ignoring the effect of semantic-free visual representation in alleviating the biased recognition problem. In this paper, we propose a novel Domain-aware Visual Bias Eliminating (DVBE) network that constructs two complementary visual representations, i.e., semantic-free and semantic-aligned, to treat seen and unseen domains separately. Specifically, we explore cross-attentive second-order visual statistics to compact the semantic-free representation, and design an adaptive margin Softmax to maximize inter-class divergences. Thus, the semantic-free representation becomes discriminative enough to not only predict seen class accurately but also filter out unseen images, i.e., domain detection, based on the predicted class entropy. For unseen images, we automatically search an optimal semantic-visual alignment architecture, rather than manual designs, to predict unseen classes. With accurate domain detection, the biased recognition problem towards the seen domain is significantly reduced. Experiments on five benchmarks for classification and segmentation show that DVBE outperforms existing methods by averaged 5.7% improvement.
Shaobo Min, Hantao Yao, Hongtao Xie 0001, Chaoqun Wang 0011, Zhengjun Zha, Yongdong Zhang 0001
CVPR4
2019 Structure Generation and Guidance Network for Unsupervised Monocular Depth Estimation
abstract
Structure information is important to unsupervised depth learning from monocular videos. However, most existing methods focus on depth smoothing on planar regions, while other structure information, such as object shape and surface curvature, is ignored. In this work, we propose SGGN, a novel Structure Generation and Guidance Network to refine depth estimation under the guidance of extracted image structure. We introduce second-order Domain Transform filtering, which explores spatial depth variation by gradient propagation, to capture long-range dependence in the extracted structure for depth refinement. Then, several structure-aware constraints, as well as an attention mechanism, are applied to guide the training of SGGN, which leads to better depth estimation with structural guidance. Notably, our structure-aware constraints are designed in terms of different characteristics. Experiments on three benchmarks demonstrate the effectiveness of our structure-guided model and its state-of-the-art performance for unsupervised depth estimation.
Chaoqun Wang 0011, Xuejin Chen, Shaobo Min, Feng Wu 0001
ICME1
2018 Real-Time Object Tracking with Motion Information
abstract
Motion is a vital information for object tracking. However, most existing methods, including the classic Siamese FC network [1], only consider the object appearance, and ignore the vital motion feature. In this paper, we design a dual-network object tracker, which is called DOT for short, to effectively combine the appearance and motion information. Our method employs two branches, S-net and M-net, to exploit the appearance and motion information respectively. Moreover, an attention fusion module is also introduced to effectively integrate these two aspects. The experiments carried out on OTB-2013 demonstrate the improvement on object tracking by the integration of motion information with our dual-network and attention fusion.
Chaoqun Wang 0011, Xiaoyan Sun 0001, Xuejin Chen, Wenjun Zeng 0001
VCIP1