EDBT 2026 Demo / reviewers in the wild / expert
Changqing Zou
dblp:126/1043
· DBLP profile ↗
61ranked-venue papers
9as first author
28since 2021 · last 2025
0000-0001-8264-6849ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 49 · 8 first-author · 21 since 2021Artificial intelligence and machine learning · 36 · 4 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TSD-SR: One-Step Diffusion with Target Score Distillation for Real-World Image Super-ResolutionabstractPre-trained text-to-image diffusion models are increasingly applied to real-world image super-resolution (Real-ISR) task. Given the iterative refinement nature of diffusion models, most existing approaches are computationally expensive. While methods such as SinSR and OSEDiff have emerged to condense inference steps via distillation, their performance in image restoration or details recovery is not satisfied. To address this, we propose TSD-SR, a novel distillation framework specifically designed for real-world image super-resolution, aiming to construct an efficient and effective one-step model. We first introduce the Target Score Distillation, which leverages the priors of diffusion models and real image references to achieve more realistic image restoration. Secondly, we propose a Distribution-Aware Sampling Module to make detail-oriented gradients more readily accessible, addressing the challenge of recovering fine details. Extensive experiments demonstrate that our TSD-SR has superior restoration results (most of the metrics perform the best) and the fastest inference speed (e.g. 40 times faster than SeeSR) compared to the past Real-ISR approaches based on pre-trained diffusion priors. Linwei Dong, Qingnan Fan, Yihong Guo, Yawei Luo, Changqing Zou |
CVPR | 8 |
| 2025 | MoEE: Mixture of Emotion Experts for Audio-Driven Portrait AnimationabstractThe generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current methods face fundamental challenges: a) the absence of frameworks for modeling single basic emotional expressions, which restricts the generation of complex emotions such as compound emotions; b) the lack of comprehensive datasets rich in human emotional expressions, which limits the potential of models. To address these challenges, we propose the following innovations: 1) the Mixture of Emotion Experts (MoEE) model, which decouples six fundamental emotions to enable the precise synthesis of both singular and compound emotional states; 2) the DH-FaceEmoVid-150 dataset, specifically curated to include six prevalent human emotional expressions as well as four types of compound emotions, thereby expanding the training potential of emotion-driven models. Furthermore, to enhance the flexibility of emotion control, we propose an emotion-to-latents module that leverages multimodal inputs, aligning diverse control signals—such as audio, text, and labels—to ensure more varied control inputs as well as the ability to control emotions using audio alone. Through extensive quantitative and qualitative evaluations, we demonstrate that the MoEE framework, in conjunction with the DH-FaceEmoVid-150 dataset, excels in generating complex emotional expressions and nuanced facial details, setting a new benchmark in the field. These datasets will be publicly released. Huaize Liu, Wenzhang Sun, Donglin Di, Shibo Sun, Changqing Zou, Hujun Bao |
CVPR | 6 |
| 2025 | DecoupledGaussian: Object-Scene Decoupling for Physics-Based InteractionabstractWe present DecoupledGaussian, a novel system that decouples static objects from their contacted surfaces captured in-the-wild videos, a key prerequisite for realistic Newtonian-based physical simulations. Unlike prior methods focused on synthetic data or elastic jittering along the contact surface, which prevent objects from fully detaching or moving independently, DecoupledGaussian allows for significant positional changes without being constrained by the initial contacted surface. Recognizing the limitations of current 2D inpainting tools for restoring 3D locations, our approach proposes joint Poisson fields to repair and expand the Gaussians of both objects and contacted scenes after separation. This is complemented by a multi-carve strategy to refine the object’s geometry. Our system enables realistic simulations of decoupling motions, collisions, and fractures driven by user-specified impulses, supporting complex interactions within and across multiple scenes. We validate DecoupledGaussian through a comprehensive user study and quantitative benchmarks. This system enhances digital interaction with objects and scenes in real-world environments, benefiting industries such as VR, robotics, and autonomous driving. Our project page is at: https://wangmiaowei.github.io/DecoupledGaussian.github.io/. Miaowei Wang, Weiwei Xu 0003, Rui Ma 0011, Changqing Zou, Daniel D. Morris |
CVPR | 5 |
| 2025 | SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited ObservationsabstractNovel view synthesis (NVS) boosts immersive experiences in computer vision and graphics. Existing techniques, though progressed, rely on dense multi-view observations, restricting their application. This work takes on the challenge of reconstructing photorealistic 3D scenes from sparse or single-view inputs. We introduce SpatialCrafter, a framework that leverages the rich knowledge in video diffusion models to generate plausible additional observations, thereby alleviating reconstruction ambiguity. Through a trainable camera encoder and an epipolar attention mechanism for explicit geometric constraints, we achieve precise camera control and 3D consistency, further reinforced by a unified scale estimation strategy to handle scale discrepancies across datasets. Furthermore, by integrating monocular depth priors with semantic features in the video latent space, our framework directly regresses 3D Gaussian primitives and efficiently processes long-sequence features using a hybrid network structure. Extensive experiments show our method enhances sparse view reconstruction and restores the realistic appearance of 3D scenes. Songchun Zhang, Huiyao Xu, Sitong Guo, Zhongwei Xie, Hujun Bao, Weiwei Xu 0003, Changqing Zou |
ICCV | 7 |
| 2025 | Diff3DS: Generating View-Consistent 3D Sketch via Differentiable Curve Renderingabstract3D sketches are widely used for visually representing the 3D shape and structure of objects or scenes. However, the creation of 3D sketch often requires users to possess professional artistic skills. Existing research efforts primarily focus on enhancing the ability of interactive sketch generation in 3D virtual systems. In this work, we propose Diff3DS, a novel differentiable rendering framework for generating view-consistent 3D sketch by optimizing 3D parametric curves under various supervisions. Specifically, we perform perspective projection to render the 3D rational Bézier curves into 2D curves, which are subsequently converted to a 2D raster image via our customized differentiable rasterizer. Our framework bridges the domains of 3D sketch and raster image, achieving end-to-end optimization of 3D sketch through gradients computed in the 2D image domain. Our Diff3DS can enable a series of novel 3D sketch generation tasks, including text-to-3D sketch and image-to-3D sketch, supported by the popular distillation-based supervision, such as Score Distillation Sampling (SDS). Extensive experiments have yielded promising results and demonstrated the potential of our framework. Project: https://yiboz2001.github.io/Diff3DS/ Changqing Zou, Tieru Wu, Rui Ma 0011 |
ICLR | 3 |
| 2025 | LL-Gaussian: Low-Light Scene Reconstruction and Enhancement via Gaussian Splatting for Novel View Synthesis
Fenggen Yu, Huiyao Xu, Tao Zhang 0042, Changqing Zou |
ACM Multimedia | 5 |
| 2025 | CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ SegmentationabstractMulti-organ medical segmentation is a crucial component of medical image processing, essential for doctors to make accurate diagnoses and develop effective treatment plans. Despite significant progress in this field, current multi-organ segmentation models often suffer from inaccurate details, dependence on geometric prompts and loss of spatial information. Addressing these challenges, we introduce a novel model named CRISP-SAM2 with CR oss-modal Interaction and Semantic Prompting based on SAM2. This model represents a promising approach to multi-organ medical segmentation guided by textual descriptions of organs. Our method begins by converting visual and textual inputs into cross-modal contextualized semantics using a progressive cross-attention interaction mechanism. These semantics are then injected into the image encoder to enhance the detailed understanding of visual information. To eliminate reliance on geometric prompts, we use a semantic prompting strategy, replacing the original prompt encoder to sharpen the perception of challenging targets. In addition, a similarity-sorting self-updating strategy for memory and a mask-refining process is applied to further adapt to medical imaging and enhance localized details. Comparative experiments conducted on seven public datasets indicate that CRISP-SAM2 outperforms existing models. Extensive analysis also demonstrates the effectiveness of our method, thereby confirming its superior performance, especially in addressing the limitations mentioned earlier. Our code is available at: https://github.com/YU-deep/CRISP_SAM2.git. Changmiao Wang, Ahmed El-Azab, Gangyong Jia, Changqing Zou, Ruiquan Ge |
ACM Multimedia | 7 |
| 2025 | Efficient Object Reconstruction with Differentiable Area Light ShadingabstractIn 3D object reconstruction from photographs, estimating material properties is challenging. We propose an inverse rendering method that uses active area lighting: as this provides a wider range of lighting angles per photo than point lighting, material reconstruction can be more accurate for the same number of photos. We compare area light shading with point lighting. With either mesh or 3D Gaussian splatting pipelines, area lighting can improve BRDF reconstruction and leads to +3 dB relighting PSNR over point lights, or need only \(\nicefrac {1}{5}\) of the input photos for the same quality. We also compare area light shading with Monte Carlo ray tracing and with differential linearly transformed cosines (LTC) plus shadow visibility weighting. LTC can be faster, improving optimization times by 25%. In SOTA method-level comparisons, our approach improves material reconstruction, particularly for material roughness, leading to superior relighting quality. Yaoan Gao, Jiamin Xu, James Tompkin 0001, Qi Wang 0111, Hujun Bao, Yujun Shen, Huamin Wang 0001, Changqing Zou, Weiwei Xu 0003 |
SIGGRAPH Asia | 9 |
| 2025 | PointNorm-Net: Self-Supervised Normal Prediction of 3D Point Clouds via Multi-Modal Distribution EstimationabstractAlthough supervised deep normal estimators have recently shown impressive results on synthetic benchmarks, their performance deteriorates significantly in real-world scenarios due to the domain gap between synthetic and real data. Building high-quality real training data to boost those supervised methods is not trivial because point-wise annotation of normals for varying-scale real-world 3D scenes is a tedious and expensive task. This paper introduces PointNorm-Net, the first self-supervised deep learning framework to tackle this challenge. The key novelty of PointNorm-Net is a three-stage multi-modal normal distribution estimation paradigm that can be integrated into either deep or traditional optimization-based normal estimation frameworks. Extensive experiments show that our method achieves superior generalization and outperforms state-of-the-art conventional and deep learning approaches across three real-world datasets that exhibit distinct characteristics compared to the synthetic training data. Jie Zhang 0056, Minghui Nie, Changqing Zou, Ligang Liu 0001, Junjie Cao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | A General Implicit Framework for Fast NeRF Composition and RenderingabstractA variety of Neural Radiance Fields (NeRF) methods have recently achieved remarkable success in high render speed. However, current accelerating methods are specialized and incompatible with various implicit methods, preventing real-time composition over various types of NeRF works. Because NeRF relies on sampling along rays, it is possible to provide general guidance for acceleration. To that end, we propose a general implicit pipeline for composing NeRF objects quickly. Our method enables the casting of dynamic shadows within or between objects using analytical light sources while allowing multiple NeRF objects to be seamlessly placed and rendered together with any arbitrary rigid transformations. Mainly, our work introduces a new surface representation known as Neural Depth Fields (NeDF) that quickly determines the spatial relationship between objects by allowing direct intersection computation between rays and implicit surfaces. It leverages an intersection neural network to query NeRF for acceleration instead of depending on an explicit spatial structure.Our proposed method is the first to enable both the progressive and interactive composition of NeRF objects. Additionally, it also serves as a previewing plugin for a range of existing NeRF works. Ziyi Yang 0008, Yunlu Zhao, Xiaogang Jin 0001, Changqing Zou |
AAAI | 6 |
| 2024 | 3D-SceneDreamer: Text-Driven 3D-Consistent Scene GenerationabstractText-driven 3D scene generation techniques have made rapid progress in recent years. Their success is mainly at-tributed to using existing generative models to iteratively perform image warping and inpainting to generate 3D scenes. However, these methods heavily rely on the out-puts of existing models, leading to error accumulation in geometry and appearance that prevent the models from being used in various scenarios (e.g., outdoor and unreal sce-narios). To address this limitation, we generatively refine the newly generated local views by querying and aggregating global 3D information, and then progressively generate the 3D scene. Specifically, we employ a tri-plane features-based NeRF as a unified representation of the 3D scene to constrain global 3D consistency, and propose a generative refinement network to synthesize new contents with higher quality by exploiting the natural image prior from 2D dif-fusion model as well as the global 3D information of the current scene. Our extensive experiments demonstrate that, in comparison to previous methods, our approach supports wide variety of scene generation and arbitrary camera tra-jectories with improved visual quality and 3D consistency. Songchun Zhang, Quan Zheng 0004, Rui Ma 0011, Wei Hua 0002, Hujun Bao, Weiwei Xu 0003, Changqing Zou |
CVPR | 8 |
| 2024 | SweepNet: Unsupervised Learning Shape Abstraction via Neural Sweepers
Mingrui Zhao, Yizhi Wang 0006, Fenggen Yu, Changqing Zou, Ali Mahdavi-Amiri |
ECCV (37) | 4 |
| 2023 | CAP-VSTNet: Content Affinity Preserved Versatile Style TransferabstractContent affinity loss including feature and pixel affinity is a main problem which leads to artifacts in photorealistic and video style transfer. This paper proposes a new framework named CAP-VSTNet, which consists of a new reversible residual network and an unbiased linear transform module, for versatile style transfer. This reversible residual network can not only preserve content affinity but not introduce redundant information as traditional reversible networks, and hence facilitate better stylization. Empowered by Matting Laplacian training loss which can address the pixel affinity loss problem led by the linear transform, the proposed framework is applicable and effective on versatile style transfer. Extensive experiments show that CAP-VSTNet can produce better qualitative and quantitative results in comparison with the state-of-the-art methods. Linfeng Wen 0003, Chengying Gao, Changqing Zou |
CVPR | 3 |
| 2023 | Neural Motion GraphabstractDeep learning techniques have been employed to design a controllable human motion synthesizer. Despite their potential, however, designing a neural network-based motion synthesis that enables flexible user interaction, fine-grained controllability, and the support of new types of motions at reduced time and space consumption costs remains a challenge. In this paper, we propose a novel approach, a neural motion graph, that addresses the challenge by enabling scalability to new motions while using compact neural networks. Our approach represents each type of motion with a separate neural node to reduce the cost of adding new motion types. In addition, designing a separate neural node for each motion type enables task-specific control strategies and has greater potential to achieve a high-quality synthesis of complex motions, such as the Mongolian dance. Furthermore, a single transition network, which acts as neural edges, is used to model the transition between two motion nodes. The transition network is designed with a lightweight control module to achieve a fine-grained response to user control signals. Overall, the design choice makes the neural motion graph highly controllable and scalable. In addition to being fully flexible to user interaction through high-level and fine-grained user-control signals, our experimental and subjective evaluation results demonstrate that our proposed approach, neural motion graph, outperforms state-of-the-art human motion synthesis methods in terms of the quality of controlled motion generation. Hongyu Tao, Shuaiying Hou, Changqing Zou, Hujun Bao, Weiwei Xu 0003 |
SIGGRAPH Asia | 3 |
| 2023 | Generating Hypergraph-Based High-Order Representations of Whole-Slide Histopathological Images for Survival PredictionabstractPatient survival prediction based on gigapixel whole-slide histopathological images (WSIs) has become increasingly prevalent in recent years. A key challenge of this task is achieving an informative survival-specific global representation from those WSIs with highly complicated data correlation. This article proposes a multi-hypergraph based learning framework, called "HGSurvNet," to tackle this challenge. HGSurvNet achieves an effective high-order global representation of WSIs via multilateral correlation modeling in multiple spaces and a general hypergraph convolution network. It has the ability to alleviate over-fitting issues caused by the lack of training data by using a new convolution structure called hypergraph max-mask convolution. Extensive validation experiments were conducted on three widely-used carcinoma datasets: Lung Squamous Cell Carcinoma (LUSC), Glioblastoma Multiforme (GBM), and National Lung Screening Trial (NLST). Quantitative analysis demonstrated that the proposed method consistently outperforms state-of-the-art methods, coupled with the Bayesian Concordance Readjust loss. We also demonstrate the individual effectiveness of each module of the proposed framework and its application potential for pathology diagnosis and reporting empowered by its interpretability potential. Donglin Di, Changqing Zou, Yifan Feng 0001, Rongrong Ji, Qionghai Dai, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Graph Learning on Millions of Data in Seconds: Label Propagation Acceleration on Graph Using Data DistributionabstractGraph-based semi-supervised learning methods have been used in a wide range of real-world applications, e.g., from social relationship mining to multimedia classification and retrieval. However, existing methods are limited along with high computational complexity or not facilitating incremental learning, which may not be powerful to deal with large-scale data, whose scale may continuously increase, in real world. This paper proposes a new method called Data Distribution Based Graph Learning (DDGL) for semi-supervised learning on large-scale data. This method can achieve a fast and effective label propagation and supports incremental learning. The key motivation is to propagate the labels along smaller-scale data distribution model parameters, rather than directly dealing with the raw data as previous methods, which accelerate the data propagation significantly. It also improves the prediction accuracy since the loss of structure information can be alleviated in this way. To enable incremental learning, we propose an adaptive graph updating strategy which can update the model when there is distribution bias between new data and the already seen data. We have conducted comprehensive experiments on multiple datasets with sample sizes increasing from seven thousand to five million. Experimental results on the classification task on large-scale data demonstrate that our proposed DDGL method improves the classification accuracy by a large margin while consuming much less time compared to state-of-the-art methods. Yubo Zhang 0006, Shuyi Ji, Changqing Zou, Xibin Zhao, Shihui Ying, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Ultra-High Resolution SVBRDF Recovery from a Single ImageabstractExisting convolutional neural networks have achieved great success in recovering Spatially Varying Bidirectional Surface Reflectance Distribution Function (SVBRDF) maps from a single image. However, they mainly focus on handling low-resolution (e.g., 256 × 256) inputs. Ultra-High Resolution (UHR) material maps are notoriously difficult to acquire by existing networks because (1) finite computational resources set bounds for input receptive fields and output resolutions, and (2) convolutional layers operate locally and lack the ability to capture long-range structural dependencies in UHR images. We propose an implicit neural reflectance model and a divide-and-conquer solution to address these two challenges simultaneously. We first crop a UHR image into low-resolution patches, each of which are processed by a local feature extractor to extract important details. To fully exploit long-range spatial dependency and ensure global coherency, we incorporate a global feature extractor and several coordinate-aware feature assembly modules into our pipeline. The global feature extractor contains several lightweight material vision transformers that have a global receptive field at each scale and have the ability to infer long-term relationships in the material. After decoding globally coherent feature maps assembled by coordinate-aware feature assembly modules, the proposed end-to-end method is able to generate UHR SVBRDF maps from a single image with fine spatial details and consistent global structures. Jie Guo 0001, Shuichang Lai, Qinghao Tu, Chengzhi Tao, Changqing Zou, Yanwen Guo 0001 |
ACM Trans. Graph. | 5 |
| 2023 | Manifold Path Guiding for Importance Sampling Specular ChainsabstractComplex visual effects such as caustics are often produced by light paths containing multiple consecutive specular vertices (dubbed specular chains) , which pose a challenge to unbiased estimation in Monte Carlo rendering. In this work, we study the light transport behavior within a sub-path that is comprised of a specular chain and two non-specular separators. We show that the specular manifolds formed by all the sub-paths could be exploited to provide coherence among sub-paths. By reconstructing continuous energy distributions from historical and coherent sub-paths, seed chains can be generated in the context of importance sampling and converge to admissible chains through manifold walks. We verify that importance sampling the seed chain in the continuous space reaches the goal of importance sampling the discrete admissible specular chain. Based on these observations and theoretical analyses, a progressive pipeline, manifold path guiding , is designed and implemented to importance sample challenging paths featuring long specular chains. To our best knowledge, this is the first general framework for importance sampling discrete specular chains in regular Monte Carlo rendering. Extensive experiments demonstrate that our method outperforms state-of-the-art unbiased solutions with up to 40 × variance reduction, especially in typical scenes containing long specular chains and complex visibility. Zhimin Fan 0001, Pengpei Hong, Jie Guo 0001, Changqing Zou, Yanwen Guo 0001, Lingqi Yan 0001 |
ACM Trans. Graph. | 4 |
| 2022 | Heterogeneous Hypergraph Variational Autoencoder for Link PredictionabstractLink prediction aims at inferring missing links or predicting future ones based on the currently observed network. This topic is important for many applications such as social media, bioinformatics and recommendation systems. Most existing methods focus on homogeneous settings and consider only low-order pairwise relations while ignoring either the heterogeneity or high-order complex relations among different types of nodes, which tends to lead to a sub-optimal embedding result. This paper presents a method named Heterogeneous Hypergraph Variational Autoencoder (HeteHG-VAE) for link prediction in heterogeneous information networks (HINs). It first maps a conventional HIN to a heterogeneous hypergraph with a certain kind of semantics to capture both the high-order semantics and complex relations among nodes, while preserving the low-order pairwise topology information of the original HIN. Then, deep latent representations of nodes and hyperedges are learned by a Bayesian deep generative framework from the heterogeneous hypergraph in an unsupervised manner. Moreover, a hyperedge attention module is designed to learn the importance of different types of nodes in each hyperedge. The major merit of HeteHG-VAE lies in its ability of modeling multi-level relations in heterogeneous settings. Extensive experiments on real-world datasets demonstrate the effectiveness and efficiency of the proposed method. Haoyi Fan, Fengbin Zhang, Yuxuan Wei, Changqing Zou, Yue Gao 0002, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Hypergraph Learning: Methods and PracticesabstractHypergraph learning is a technique for conducting learning on a hypergraph structure. In recent years, hypergraph learning has attracted increasing attention due to its flexibility and capability in modeling complex data correlation. In this paper, we first systematically review existing literature regarding hypergraph generation, including distance-based, representation-based, attribute-based, and network-based approaches. Then, we introduce the existing learning methods on a hypergraph, including transductive hypergraph learning, inductive hypergraph learning, hypergraph structure updating, and multi-modal hypergraph learning. After that, we present a tensor-based dynamic hypergraph representation and learning framework that can effectively describe high-order correlation in a hypergraph. To study the effectiveness and efficiency of hypergraph generation and learning methods, we conduct comprehensive evaluations on several typical applications, including object and action recognition, Microblog sentiment prediction, and clustering. In addition, we contribute a hypergraph learning development toolkit called THU-HyperG. Yue Gao 0002, Zizhao Zhang 0003, Haojie Lin, Xibin Zhao, Shaoyi Du, Changqing Zou |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | View-Aware Geometry-Structure Joint Learning for Single-View 3D Shape ReconstructionabstractReconstructing a 3D shape from a single-view image using deep learning has become increasingly popular recently. Most existing methods only focus on reconstructing the 3D shape geometry based on image constraints. The lack of explicit modeling of structure relations among shape parts yields low-quality reconstruction results for structure-rich man-made shapes. In addition, conventional 2D-3D joint embedding architecture for image-based 3D shape reconstruction often omits the specific view information from the given image, which may lead to degraded geometry and structure reconstruction. We address these problems by introducing VGSNet, an encoder-decoder architecture for view-aware joint geometry and structure learning. The key idea is to jointly learn a multimodal feature representation of 2D image, 3D shape geometry and structure so that both geometry and structure details can be reconstructed from a single-view image. To this end, we explicitly represent 3D shape structures as part relations and employ image supervision to guide the geometry and structure reconstruction. Trained with pairs of view-aligned images and 3D shapes, the VGSNet implicitly encodes the view-aware shape information in the latent feature space. Qualitative and quantitative comparisons with the state-of-the-art baseline methods as well as ablation studies demonstrate the effectiveness of the VGSNet for structure-aware single-view 3D shape reconstruction. Xuancheng Zhang, Rui Ma 0011, Changqing Zou, Xibin Zhao, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | SceneSketcher-v2: Fine-Grained Scene-Level Sketch-Based Image Retrieval Using Adaptive GCNsabstractSketch-based image retrieval (SBIR) is a long-standing research topic in computer vision. Existing methods mainly focus on category-level or instance-level image retrieval. This paper investigates the fine-grained scene-level SBIR problem where a free-hand sketch depicting a scene is used to retrieve desired images. This problem is useful yet challenging mainly because of two entangled facts: 1) achieving an effective representation of the input query data and scene-level images is difficult as it requires to model the information across multiple modalities such as object layout, relative size and visual appearances, and 2) there is a great domain gap between the query sketch input and target images. We present SceneSketcher-v2, a Graph Convolutional Network (GCN) based architecture to address these challenges. SceneSketcher-v2 employs a carefully designed graph convolution network to fuse the multi-modality information in the query sketch and target images and uses a triplet training process and end-to-end training manner to alleviate the domain gap. Extensive experiments demonstrate SceneSketcher-v2 outperforms state-of-the-art scene-level SBIR models with a significant margin. Fang Liu 0035, Xiaoming Deng 0001, Changqing Zou, Yukun Lai, Ran Zuo, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
IEEE Trans. Image Process. | 3 |
| 2021 | View-Guided Point Cloud CompletionabstractThis paper presents a view-guided solution for the task of point cloud completion. Unlike most existing methods directly inferring the missing points using shape priors, we address this task by introducing ViPC (view-guided point cloud completion) that takes the missing crucial global structure information from an extra single-view image. By leveraging a framework that sequentially performs effective cross-modality and cross-level fusions, our method achieves significantly superior results over typical existing solutions on a new large-scale dataset we collect for the view-guided point cloud completion task. Xuancheng Zhang, Yutong Feng, Siqi Li 0001, Changqing Zou, Hai Wan, Xibin Zhao, Yandong Guo, Yue Gao 0002 |
CVPR | 4 |
| 2021 | Event Stream Super-Resolution via Spatiotemporal Constraint LearningabstractEvent cameras are bio-inspired sensors that respond to brightness changes asynchronously and output in the form of event streams instead of frame-based images. They own outstanding advantages compared with traditional cameras: higher temporal resolution, higher dynamic range, and lower power consumption. However, the spatial resolution of existing event cameras is insufficient and challenging to be enhanced at the hardware level while maintaining the asynchronous philosophy of circuit design. Therefore, it is imperative to explore the algorithm of event stream super-resolution, which is a non-trivial task due to the sparsity and strong spatio-temporal correlation of the events from an event camera. In this paper, we propose an end-to-end framework based on spiking neural network for event stream super-resolution, which can generate high-resolution (HR) event stream from the input low-resolution (LR) event stream. A spatiotemporal constraint learning mechanism is proposed to learn the spatial and temporal distributions of the event stream simultaneously. We validate our method on four large-scale datasets and the results show that our method achieves state-of-the-art performance. The satisfying results on two downstream applications, i.e. object classification and image reconstruction, further demonstrate the usability of our method. To prove the application potential of our method, we deploy it on a mobile platform. The high-quality HR event stream generated by our real-time system demonstrates the effectiveness and efficiency of our method. Siqi Li 0001, Yutong Feng, Yu Jiang 0006, Changqing Zou, Yue Gao 0002 |
ICCV | 5 |
| 2021 | LapsCore: Language-guided Person Search via Color ReasoningabstractThe key point of language-guided person search is to construct the cross-modal association between visual and textual input. Existing methods focus on designing multimodal attention mechanisms and novel cross-modal loss functions to learn such association implicitly. We propose a representation learning method for language-guided person search based on color reasoning (LapsCore). It can explicitly build a fine-grained cross-modal association bidirectionally. Specifically, a pair of dual sub-tasks, image colorization and text completion, is designed. In the former task, rich text information is learned to colorize gray images, and the latter one requests the model to understand the image and complete color word vacancies in the captions. The two sub-tasks enable models to learn correct alignments between text phrases and image regions, so that rich multimodal representations can be learned. Extensive experiments on multiple datasets demonstrate the effectiveness and superiority of the proposed method. Yushuang Wu, Zizheng Yan, Xiaoguang Han 0001, Guanbin Li, Changqing Zou, Shuguang Cui |
ICCV | 5 |
| 2021 | SADRNet: Self-Aligned Dual Face Regression Networks for Robust 3D Dense Face Alignment and ReconstructionabstractThree-dimensional face dense alignment and reconstruction in the wild is a challenging problem as partial facial information is commonly missing in occluded and large pose face images. Large head pose variations also increase the solution space and make the modeling more difficult. Our key idea is to model occlusion and pose to decompose this challenging task into several relatively more manageable subtasks. To this end, we propose an end-to-end framework, termed as Self-aligned Dual face Regression Network (SADRNet), which predicts a pose-dependent face, a pose-independent face. They are combined by an occlusion-aware self-alignment to generate the final 3D face. Extensive experiments on two popular benchmarks, AFLW2000-3D and Florence, demonstrate that the proposed method achieves significant superior performance over existing state-of-the-art methods. Zeyu Ruan, Changqing Zou, Longhai Wu, Gangshan Wu, Limin Wang 0002 |
IEEE Trans. Image Process. | 2 |
| 2021 | General virtual sketching framework for vector line artabstractVector line art plays an important role in graphic design, however, it is tedious to manually create. We introduce a general framework to produce line drawings from a wide variety of images, by learning a mapping from raster image space to vector image space. Our approach is based on a recurrent neural network that draws the lines one by one. A differentiable rasterization module allows for training with only supervised raster data. We use a dynamic window around a virtual pen while drawing lines, implemented with a proposed aligned cropping and differentiable pasting modules. Furthermore, we develop a stroke regularization loss that encourages the model to use fewer and longer strokes to simplify the resulting vector image. Ablation studies and comparisons with existing methods corroborate the efficiency of our approach which is able to generate visually better results in less computation time, while generalizing better to a diversity of images and applications. Haoran Mo, Edgar Simo-Serra, Chengying Gao, Changqing Zou, Ruomei Wang 0001 |
ACM Trans. Graph. | 4 |
| 2021 | Sketch-R2CNN: An RNN-Rasterization-CNN Architecture for Vector Sketch RecognitionabstractSketches in existing large-scale datasets like the recent QuickDraw collection are often stored in a vector format, with strokes consisting of sequentially sampled points. However, most existing sketch recognition methods rasterize vector sketches as binary images and then adopt image classification techniques. In this article, we propose a novel end-to-end single-branch network architecture RNN-Rasterization-CNN (Sketch-R2CNN for short) to fully leverage the vector format of sketches for recognition. Sketch-R2CNN takes a vector sketch as input and uses an RNN for extracting per-point features in the vector space. We then develop a neural line rasterization module to convert the vector sketch and the per-point features to multi-channel point feature maps, which are subsequently fed to a CNN for extracting convolutional features in the pixel space. Our neural line rasterization module is designed in a differentiable way for end-to-end learning. We perform experiments on existing large-scale sketch recognition datasets and show that the RNN-Rasterization design brings consistent improvement over CNN baselines and that Sketch-R2CNN substantially outperforms the state-of-the-art methods. Lei Li 0038, Changqing Zou, Youyi Zheng, Qingkun Su, Hongbo Fu 0001, Chiew-Lan Tai |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Attention-Based Multi-Modal Fusion Network for Semantic Scene CompletionabstractThis paper presents an end-to-end 3D convolutional network named attention-based multi-modal fusion network (AMFNet) for the semantic scene completion (SSC) task of inferring the occupancy and semantic labels of a volumetric 3D scene from single-view RGB-D images. Compared with previous methods which use only the semantic features extracted from RGB-D images, the proposed AMFNet learns to perform effective 3D scene completion and semantic segmentation simultaneously via leveraging the experience of inferring 2D semantic segmentation from RGB-D images as well as the reliable depth cues in spatial dimension. It is achieved by employing a multi-modal fusion architecture boosted from 2D semantic segmentation and a 3D semantic completion network empowered by residual attention blocks. We validate our method on both the synthetic SUNCG-RGBD dataset and the real NYUv2 dataset and the results show that our method respectively achieves the gains of 2.5% and 2.6% on the synthetic SUNCG-RGBD dataset and the real NYUv2 dataset against the state-of-the-art method. Siqi Li 0001, Changqing Zou, Xibin Zhao, Yue Gao 0002 |
AAAI | 2 |
| 2020 | Hypergraph Label Propagation NetworkabstractIn recent years, with the explosion of information on the Internet, there has been a large amount of data produced, and analyzing these data is useful and has been widely employed in real world applications. Since data labeling is costly, lots of research has focused on how to efficiently label data through semi-supervised learning. Among the methods, graph and hypergraph based label propagation algorithms have been a widely used method. However, traditional hypergraph learning methods may suffer from their high computational cost. In this paper, we propose a Hypergraph Label Propagation Network (HLPN) which combines hypergraph-based label propagation and deep neural networks in order to optimize the feature embedding for optimal hypergraph learning through an end-to-end architecture. The proposed method is more effective and also efficient for data labeling compared with traditional hypergraph learning methods. We verify the effectiveness of our proposed HLPN method on a real-world microblog dataset gathered from Sina Weibo. Experiments demonstrate that the proposed method can significantly outperform the state-of-the-art methods and alternative approaches. Yubo Zhang 0006, Nan Wang 0015, Changqing Zou, Hai Wan, Xibin Zhao, Yue Gao 0002 |
AAAI | 4 |
| 2020 | SketchyCOCO: Image Generation From Freehand Scene SketchesabstractWe introduce the first method for automatic image generation from scene-level freehand sketches. Our model allows for controllable image generation by specifying the synthesis goal via freehand sketches. The key contribution is an attribute vector bridged Generative Adversarial Network called EdgeGAN, which supports high visual-quality object-level image content generation without using freehand sketches as training data. We have built a large-scale composite dataset called SketchyCOCO to support and evaluate the solution. We validate our approach on the tasks of both object-level and scene-level image generation on SketchyCOCO. Through quantitative, qualitative results, human evaluation and ablation studies, we demonstrate the method's capacity to generate realistic complex scene-level images from various freehand sketches. Chengying Gao, Limin Wang 0002, Jianzhuang Liu, Changqing Zou |
CVPR | 6 |
| 2020 | Universal Physical Camouflage Attacks on Object DetectorsabstractIn this paper, we study physical adversarial attacks on object detectors in the wild. Previous works mostly craft instance-dependent perturbations only for rigid or planar objects. To this end, we propose to learn an adversarial pattern to effectively attack all instances belonging to the same object category, referred to as Universal Physical Camouflage Attack (UPC). Concretely, UPC crafts camouflage by jointly fooling the region proposal network, as well as misleading the classifier and the regressor to output errors. In order to make UPC effective for non-rigid or non-planar objects, we introduce a set of transformations for mimicking deformable properties. We additionally impose optimization constraint to make generated patterns look natural to human observers. To fairly evaluate the effectiveness of different physical-world attacks, we present the first standardized virtual database, AttackScenes, which simulates the real 3D world in a controllable and reproducible environment. Extensive experiments suggest the superiority of our proposed UPC compared with existing physical adversarial attackers not only in virtual environments (AttackScenes), but also in real-world physical environments. Lifeng Huang, Chengying Gao, Yuyin Zhou, Cihang Xie, Alan L. Yuille, Changqing Zou |
CVPR | 6 |
| 2020 | SceneSketcher: Fine-Grained Image Retrieval with Scene Sketches
Fang Liu 0035, Changqing Zou, Xiaoming Deng 0001, Ran Zuo, Yukun Lai, CuiXia Ma, Yong-Jin Liu 0001, Hongan Wang |
ECCV (19) | 2 |
| 2020 | Personalized Hand Modeling from Multiple Postures with Multi-View Color ImagesabstractAbstract Personalized hand models can be utilized to synthesize high quality hand datasets, provide more possible training data for deep learning and improve the accuracy of hand pose estimation. In recent years, parameterized hand models, e.g., MANO, are widely used for obtaining personalized hand models. However, due to the low resolution of existing parameterized hand models, it is still hard to obtain high‐fidelity personalized hand models. In this paper, we propose a new method to estimate personalized hand models from multiple hand postures with multi‐view color images. The personalized hand model is represented by a personalized neutral hand, and multiple hand postures. We propose a novel optimization strategy to estimate the neutral hand from multiple hand postures. To demonstrate the performance of our method, we have built a multi‐view system and captured more than 35 people, and each of them has 30 hand postures. We hope the estimated hand models can boost the research of high‐fidelity parameterized hand modeling in the future. All the hand models are publicly available on www.yangangwang.com . Yangang Wang 0001, Ruting Rao, Changqing Zou |
Comput. Graph. Forum | 3 |
| 2020 | Hamming Embedding Sensitivity Guided Fusion Network for 3D Shape RepresentationabstractThree-dimensional multi-modal data are used to represent 3D objects in the real world in different ways. Features separately extracted from multimodality data are often poorly correlated. Recent solutions leveraging the attention mechanism to learn a joint-network for the fusion of multimodality features have weak generalization capability. In this paper, we propose a hamming embedding sensitivity network to address the problem of effectively fusing multimodality features. The proposed network called HamNet is the first end-to-end framework with the capacity to theoretically integrate data from all modalities with a unified architecture for 3D shape representation, which can be used for 3D shape retrieval and recognition. HamNet uses the feature concealment module to achieve effective deep feature fusion. The basic idea of the concealment module is to re-weight the features from each modality at an early stage with the hamming embedding of these modalities. The hamming embedding also provides an effective solution for fast retrieval tasks on a large scale dataset. We have evaluated the proposed method on the large-scale ModelNet40 dataset for the tasks of 3D shape classification, single modality and cross-modality retrieval. Comprehensive experiments and comparisons with state-of-the-art methods demonstrate that the proposed approach can achieve superior performance. Biao Gong, Chenggang Yan 0001, Changqing Zou, Yue Gao 0002 |
IEEE Trans. Image Process. | 4 |
| 2019 | PVRNet: Point-View Relation Neural Network for 3D Shape RecognitionabstractThree-dimensional (3D) shape recognition has drawn much research attention in the field of computer vision. The advances of deep learning encourage various deep models for 3D feature representation. For point cloud and multi-view data, two popular 3D data modalities, different models are proposed with remarkable performance. However the relation between point cloud and views has been rarely investigated. In this paper, we introduce Point-View Relation Network (PVRNet), an effective network designed to well fuse the view features and the point cloud feature with a proposed relation score module. More specifically, based on the relation score module, the point-single-view fusion feature is first extracted by fusing the point cloud feature and each single view feature with point-singe-view relation, then the pointmulti- view fusion feature is extracted by fusing the point cloud feature and the features of different number of views with point-multi-view relation. Finally, the point-single-view fusion feature and point-multi-view fusion feature are further combined together to achieve a unified representation for a 3D shape. Our proposed PVRNet has been evaluated on ModelNet40 dataset for 3D shape classification and retrieval. Experimental results indicate our model can achieve significant performance improvement compared with the state-of-the-art models. Haoxuan You, Yifan Feng 0001, Xibin Zhao, Changqing Zou, Rongrong Ji, Yue Gao 0002 |
AAAI | 4 |
| 2019 | ADCrowdNet: An Attention-Injective Deformable Convolutional Network for Crowd UnderstandingabstractWe propose an attention-injective deformable convolutional network called ADCrowdNet for crowd understanding that can address the accuracy degradation problem of highly congested noisy scenes. ADCrowdNet contains two concatenated networks. An attention-aware network called Attention Map Generator (AMG) first detects crowd regions in images and computes the congestion degree of these regions. Based on detected crowd regions and congestion priors, a multi-scale deformable network called Density Map Estimator (DME) then generates high-quality density maps. With the attention-aware training scheme and multi-scale deformable convolutional scheme, the proposed ADCrowdNet achieves the capability of being more effective to capture the crowd features and more resistant to various noises. We have evaluated our method on four popular crowd counting datasets (ShanghaiTech, UCF_CC_50, WorldEXPO'10, and UCSD) and an extra vehicle counting dataset TRANCOS, and our approach beats existing state-of-the-art approaches on all of these datasets. Yongchao Long, Changqing Zou, Qun Niu, Li Pan 0002, Hefeng Wu |
CVPR | 3 |
| 2019 | Language-based colorization of scene sketchesabstractBeing natural, touchless, and fun-embracing, language-based inputs have been demonstrated effective for various tasks from image generation to literacy education for children. This paper for the first time presents a language-based system for interactive colorization of scene sketches, based on semantic comprehension. The proposed system is built upon deep neural networks trained on a large-scale repository of scene sketches and cartoonstyle color images with text descriptions. Given a scene sketch, our system allows users, via language-based instructions, to interactively localize and colorize specific foreground object instances to meet various colorization requirements in a progressive way. We demonstrate the effectiveness of our approach via comprehensive experimental results including alternative studies, comparison with the state-of-the-art methods, and generalization user studies. Given the unique characteristics of language-based inputs, we envision a combination of our interface with a traditional scribble-based interface for a practical multimodal colorization system, benefiting various applications. The dataset and source code can be found at https://github.com/SketchyScene/SketchySceneColorization. Changqing Zou, Haoran Mo, Chengying Gao, Ruofei Du, Hongbo Fu 0001 |
ACM Trans. Graph. | 1 |
| 2018 | SketchyScene: Richly-Annotated Scene Sketches
Changqing Zou, Qian Yu 0002, Ruofei Du, Haoran Mo, Yi-Zhe Song, Tao Xiang 0002, Chengying Gao, Baoquan Chen, Hao (Richard) Zhang |
ECCV (15) | 1 |
| 2018 | PencilArt: A Chromatic Penciling Style Generation FrameworkabstractAbstract Non‐photorealistic rendering has been an active area of research for decades whereas few of them concentrate on rendering chromatic penciling style. In this paper, we present a framework named as PencilArt for the chromatic penciling style generation from wild photographs. The structural outline and textured map for composing the chromatic pencil drawing are generated, respectively. First, we take advantage of deep neural network to produce the structural outline with proper intensity variation and conciseness. Next, for the textured map, we follow the painting process of artists to adjust the tone of input images to match the luminance histogram and pencil textures of real drawings. Eventually, we evaluate PencilArt via a series of comparisons to previous work, showing that our results better capture the main features of real chromatic pencil drawings and have an improved visual appearance. Chengying Gao, Mengyue Tang, Xiangguo Liang, Zhuo Su 0001, Changqing Zou |
Comput. Graph. Forum | 5 |
| 2018 | Multi-modal feature fusion for geographic image annotation
Ke Li 0005, Changqing Zou, Shuhui Bu, Yun Liang 0003, Jian Zhang 0026, Minglun Gong |
Pattern Recognit. | 2 |
| 2018 | Construction and fabrication of reversible shape transformsabstractWe study a new and elegant instance of geometric dissection of 2D shapes: reversible hinged dissection, which corresponds to a dual transform between two shapes where one of them can be dissected in its interior and then inverted inside-out , with hinges on the shape boundary, to reproduce the other shape, and vice versa. We call such a transform reversible inside-out transform or RIOT. Since it is rare for two shapes to possess even a rough RIOT, let alone an exact one, we develop both a RIOT construction algorithm and a quick filtering mechanism to pick, from a shape collection, potential shape pairs that are likely to possess the transform. Our construction algorithm is fully automatic. It computes an approximate RIOT between two given input 2D shapes, whose boundaries can undergo slight deformations, while the filtering scheme picks good inputs for the construction. Furthermore, we add properly designed hinges and connectors to the shape pieces and fabricate them using a 3D printer so that they can be played as an assembly puzzle. With many interesting and fun RIOT pairs constructed from shapes found online, we demonstrate that our method significantly expands the range of shapes to be considered for RIOT, a seemingly impossible shape transform, and offers a practical way to construct and physically realize these transforms. Ali Mahdavi-Amiri, Ruizhen Hu, Han Liu 0003, Changqing Zou, Oliver van Kaick, Xiuping Liu, Hui Huang 0004, Hao (Richard) Zhang |
ACM Trans. Graph. | 5 |
| 2017 | ℒ0 Gradient-Preserving Color TransferabstractAbstract This paper presents a new two‐step color transfer method which includes color mapping and detail preservation. To map source colors to target colors, which are from an image or palette, the proposed similarity‐preserving color mapping algorithm uses the similarities between pixel color and dominant colors as existing algorithms and emphasizes the similarities between source image pixel colors. Detail preservation is performed by an ℒ0 gradient‐preserving algorithm. It relaxes the large gradients of the sparse pixels along color region boundaries and preserves the small gradients of pixels within color regions. The proposed method preserves source image color similarity and image details well. Extensive experiments demonstrate that the proposed approach has achieved a state‐of‐art visual performance. Dong Wang 0041, Changqing Zou, Guiqing Li, Chengying Gao, Zhuo Su 0001 |
Comput. Graph. Forum | 2 |
| 2017 | Learning to group discrete graphical patternsabstractWe introduce a deep learning approach for grouping discrete patterns common in graphical designs. Our approach is based on a convolutional neural network architecture that learns a grouping measure defined over a pair of pattern elements. Motivated by perceptual grouping principles, the key feature of our network is the encoding of element shape, context, symmetries, and structural arrangements. These element properties are all jointly considered and appropriately weighted in our grouping measure. To better align our measure with human perceptions for grouping, we train our network on a large, human-annotated dataset of pattern groupings consisting of patterns at varying granularity levels, with rich element relations and varieties, and tempered with noise and other data imperfections. Experimental results demonstrate that our deep-learned measure leads to robust grouping results. Zhaoliang Lun, Changqing Zou, Evangelos Kalogerakis, Ping Tan 0002, Marie-Paule Cani, Hao (Richard) Zhang |
ACM Trans. Graph. | 2 |
| 2017 | Full and partial shape similarity through sparse descriptor reconstruction
Changqing Zou, Hao (Richard) Zhang |
Vis. Comput. | 2 |
| 2016 | An example-based approach to 3D man-made object reconstruction from line drawings
Changqing Zou, Tianfan Xue, Xiaojiang Peng, Honghua Li, Baochang Zhang 0001, Jianzhuang Liu |
Pattern Recognit. | 1 |
| 2016 | Shape similarity assessment based on partial feature aggregation and ranking lists
Zhenzhong Kuang, Zongmin Li, Yujie Liu 0002, Changqing Zou |
Pattern Recognit. Lett. | 4 |
| 2016 | Action-driven 3D indoor scene evolutionabstractWe introduce a framework for action-driven evolution of 3D indoor scenes, where the goal is to simulate how scenes are altered by human actions, and specifically, by object placements necessitated by the actions. To this end, we develop an action model with each type of action combining information about one or more human poses, one or more object categories, and spatial configurations of objects belonging to these categories which summarize the object-object and object-human relations for the action. Importantly, all these pieces of information are learned from annotated photos. Correlations between the learned actions are analyzed to guide the construction of an action graph. Starting with an initial 3D scene, we probabilistically sample a sequence of actions from the action graph to drive progressive scene evolution. Each action triggers appropriate object placements, based on object co-occurrences and spatial configurations learned for the action model. We show results of our scene evolution that lead to realistic and messy 3D scenes, as well as quantitative evaluations by user studies which compare our method to manual scene creation and state-of-the-art, data-driven methods, in terms of scene plausibility and naturalness. Rui Ma 0011, Honghua Li, Changqing Zou, Zicheng Liao, Xin Tong 0001, Hao (Richard) Zhang |
ACM Trans. Graph. | 3 |
| 2016 | Legible compact calligramsabstractA calligram is an arrangement of words or letters that creates a visual image, and a compact calligram fits one word into a 2D shape. We introduce a fully automatic method for the generation of legible compact calligrams which provides a balance between conveying the input shape, legibility, and aesthetics. Our method has three key elements: a path generation step which computes a global layout path suitable for embedding the input word; an alignment step to place the letters so as to achieve feature alignment between letter and shape protrusions while maintaining word legibility; and a final deformation step which deforms the letters to fit the shape while balancing fit against letter legibility. As letter legibility is critical to the quality of compact calligrams, we conduct a large-scale crowd-sourced study on the impact of different letter deformations on legibility and use the results to train a letter legibility measure which guides the letter deformation. We show automatically generated calligrams on an extensive set of word-image combinations. The legibility and overall quality of the calligrams are evaluated and compared, via user studies, to those produced by human creators, including a professional artist, and existing works. Changqing Zou, Junjie Cao 0001, Warunika Ranaweera, Ibraheem Alhashim, Ping Tan 0002, Alla Sheffer, Hao (Richard) Zhang |
ACM Trans. Graph. | 1 |
| 2016 | Mesh saliency detection via double absorbing Markov chain in feature space
Xiuping Liu, Pingping Tao, Junjie Cao 0001, Changqing Zou |
Vis. Comput. | 5 |
| 2015 | Sketch-based 3-D modeling for piecewise planar objects in single images
Changqing Zou, Xiaojiang Peng, Shifeng Chen, Hongbo Fu 0001, Jianzhuang Liu |
Comput. Graph. | 1 |
| 2015 | A comparison of 3D shape retrieval methods based on a large-scale benchmark supporting multimodal queries
Bo Li 0013, Yijuan Lu, Chunyuan Li, Afzal Godil, Tobias Schreck, Masaki Aono, Martin Burtscher, Nihad Karim Chowdhury, Hongbo Fu 0001, Takahiko Furuya, Hai-Sheng Li 0002, Jianzhuang Liu, Henry Johan, Ryuichi Kosaka, Hitoshi Koyanagi, Ryutarou Ohbuchi, Atsushi Tatsuma, Yajuan Wan, Changqing Zou |
Comput. Vis. Image Underst. | 22 |
| 2015 | Cascade of forests for face alignmentabstractIn this study, we propose a regression forests‐based cascaded method for face alignment. We build on the cascaded pose regression (CPR) framework and propose to use the regression forest as a primitive regressor. The regression forests are easier to train and naturally handle the over‐fitting problem via averaging the outputs of the trees at each stage. We address the fact that the CPR approaches are sensitive to the shape initialisation; in contrast to using a number of blind initialisations and selecting the median values, we propose an intelligent shape initialisation scheme. More specifically, a large number of initialisations are propagated to a few early stages in the cascade, then only a proportion of them are propagated to the remaining cascades according to their convergence measurement. We evaluate the performance of the proposed approach on the challenging face alignment in the wild database and obtain superior or comparable performance with the state‐of‐the‐art, in spite of the fact that we have utilised only the freely available public training images. More importantly, we show that the intelligent initialisation scheme makes the CPR framework more robust to unreliable initialisations that are typically produced by different face detections. Heng Yang 0001, Changqing Zou, Ioannis Patras |
IET Comput. Vis. | 2 |
| 2015 | Progressive 3D Reconstruction of Planar-Faced Manifold Objects with DRF-Based Line Drawing DecompositionabstractThis paper presents an approach for reconstructing polyhedral objects from single-view line drawings. Our approach separates a complex line drawing representing a manifold object into a series of simpler line drawings, based on the degree of reconstruction freedom (DRF). We then progressively reconstruct a complete 3D model from these simpler line drawings. Our experiments show that our decomposition algorithm is able to handle complex drawings which are challenging for the state of the art. The advantages of the presented progressive 3D reconstruction method over the existing reconstruction methods in terms of both robustness and efficiency are also demonstrated. Changqing Zou, Shifeng Chen, Hongbo Fu 0001, Jianzhuang Liu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2014 | Separation of Line Drawings Based on Split Faces for 3D Object ReconstructionabstractReconstructing 3D objects from single line drawings is often desirable in computer vision and graphics applications. If the line drawing of a complex 3D object is decomposed into primitives of simple shape, the object can be easily reconstructed. We propose an effective method to conduct the line drawing separation and turn a complex line drawing into parametric 3D models. This is achieved by recursively separating the line drawing using two types of split faces. Our experiments show that the proposed separation method can generate more basic and simple line drawings, and its combination with the example-based reconstruction can robustly recover wider range of complex parametric 3D objects than previous methods. Changqing Zou, Jianzhuang Liu |
CVPR | 1 |
| 2014 | Action Recognition with Stacked Fisher Vectors
Xiaojiang Peng, Changqing Zou, Yu Qiao 0001, Qiang Peng |
ECCV (5) | 2 |
| 2014 | Sketch-Based 3D Model Retrieval via Multi-feature FusionabstractSketch-based 3D model retrieval provides a convenient way for users to search for 3D models by sketches. Traditionally, this task is converted to a sketch-based 2D shape retrieval problem by projecting 3D models to 2D images. Local invariant features have been widely used to tackle this problem. However, it suffers from the lack of global context and easily fails when images of different 3D models share multiple similar regions. In this paper, we propose a joint description by fusing local statistical structures and global spatial features. Our description is invariant to scale, translate and rotation. An improved bag-of-features retrieval framework is applied to explore semantic visual word representations. Besides, a novel relevance feedback scheme which combines weight balancing and query modification is designed to further improve the retrieval performance. We conduct various experiments on the common sketch-based watertight model benchmark. The comparative results show that our approach significantly outperforms three state-of-the-art methods, demonstrating its effectiveness and robustness for sketch-based 3D model retrieval. Yafei Wen, Changqing Zou, Jianzhuang Liu, Shuze Du, Shifeng Chen |
ICPR | 2 |
| 2014 | Face Sketch Landmarks Localization in the WildabstractIn this letter, we propose a method for facial landmarks localization in face sketch images. As recent approaches and the corresponding datasets are designed for ordinary face photos, the performance of such models drop significantly when they are applied on face sketch images. We first propose a scheme to synthesize face sketches from face photos based on random-forests edge detection and local face region enhancement. Then we jointly train a Cascaded Pose Regression based method for facial landmarks localization for both face photos and sketches. We build an evaluation dataset, called Face Sketches in the Wild (FSW), with 450 face sketch images collected from the Internet and with the manual annotation of 68 facial landmark locations on each face sketch. The proposed multi-modality facial landmark localization method shows competitive performance on both face sketch images (the FSW dataset) and face photo images (the Labeled Face Parts in the Wild dataset), despite the fact that we do not use extra annotation of face sketches for model building. Heng Yang 0001, Changqing Zou, Ioannis Patras |
IEEE Signal Process. Lett. | 2 |
| 2014 | Viewpoint-Aware Representation for Sketch-Based 3D Model RetrievalabstractWe study the problem of sketch-based 3D model retrieval, and propose a solution powered by a new query-to-model distance metric and a powerful feature descriptor based on the bag-of-features framework. The main idea of the proposed query-to-model distance metric is to represent a query sketch using a compact set of sample views (called basic views) of each model, and to rank the models in ascending order of the representation errors. To better differentiate between relevant and irrelevant models, the representation is constrained to be essentially a combination of basic views with similar viewpoints. In another aspect, we propose a mid-level descriptor (called BOF-JESC) which robustly characterizes the edge information within junction-centered patches, to extract the salient shape features from sketches or model views. The combination of the query-to-model distance metric and the BOF-JESC descriptor achieves effective results on two latest benchmark datasets. Changqing Zou, Changhu Wang, Yafei Wen, Lei Zhang 0001, Jianzhuang Liu |
IEEE Signal Process. Lett. | 1 |
| 2012 | Precise 3D Reconstruction from a Single Image
Changqing Zou, Jianzhuang Liu |
ACCV (4) | 1 |
| 2012 | Locating high-density clusters with noisy queries
Shifeng Chen, Changqing Zou, Jianzhuang Liu |
ICPR | 3 |