EDBT 2026 Demo / reviewers in the wild / expert
Chengying Gao
dblp:90/5494
· DBLP profile ↗
48ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0003-1160-0717ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 3 first-author · 21 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 7 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adversarial contribution-based perturbation for transferable attack on point cloud
Shuxin Wei, Lifeng Huang, Chengying Gao |
Pattern Recognit. | 3 |
| 2026 | ReGA: Relighting Dynamic Gaussian Avatars From Sparse ViewsabstractDynamic human relighting is a complex task with significant demand across various applications. Its core challenge lies in effectively handling both dynamic human geometry reconstruction and material estimation. Existing works mainly achieve human reconstruction and relighting with Neural Radiance Fields, but they are not only inefficient, but also still inaccurate in material estimation and relighting. In this paper, we propose a novel approach called ReGA, which leverages efficient 3D Gaussian Splatting to create animatable and relightable avatars from sparse-view human motion. The training process consists of two stages: the geometry stage and the material stage. To overcome the geometric weakness of vanilla Gaussian representation, we introduce dynamic alignment mechanism in the geometry stage, combining the advantages of Gaussian splatting and mesh-based representation to produce reasonable human surface. In the material stage, we enhance the inverse rendering process by introducing twofold correlation strategies that establish chrominance correlation between Gaussian radiance color and albedo. Experiments demonstrate that our method outperforms existing approaches in dynamic human relighting task. Lingzhe Zeng, Rongbin Zheng, Chengying Gao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | DiFusion: Flexible Stylized Motion Generation Using Digest-and-Fusion SchemeabstractTo address the issue of style expression in existing text-driven human motion synthesis methods, we propose DiFusion, a framework for diversely stylized motion generation. It offers flexible control of content through texts and style via multiple modalities, i.e., textual labels or motion sequences. Our approach employs a dual-condition motion latent diffusion model, enabling independent control of content and style through flexible input modalities. To tackle the issue of imbalanced complexity between the text-motion and style-motion datasets, we propose the Digest-and-Fusion training scheme, which digests domain-specific knowledge from both datasets and then adaptively fuses them into a compatible manner. Comprehensive evaluations demonstrate the effectiveness of our method and its superiority over existing approaches in terms of content alignment, style expressiveness, realism, and diversity. Additionally, our approach can be extended to practical applications, such as motion style interpolation. Yatian Wang, Haoran Mo, Chengying Gao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | AUTE: Peer-Alignment and Self-Unlearning Boost Adversarial Robustness for Training Ensemble ModelsabstractAdversarial attacks poses a significant threat to the security of AI-based systems. To counteract these attacks, adversarial training (AT) and ensemble learning (EL) have emerged as widely adopted methods for enhancing model robustness. However, a counter-intuitive phenomenon arises where the simple combination of these approaches may potentially compromising adversarial robustness of ensemble models. In this paper, we propose a novel method called Alignment and Unlearning for Training Ensembles (AUTE), aiming to effectively integrate AT and EL to maximize their benefits. Specifically, AUTE incorporates two key components. Firstly, AUTE divides the ensemble into a big peer model and a single member in a loop manner, aligning their outputs for boosting robustness of each member. Secondly, AUTE introduces the concept of unlearning, actively forgetting specific data with over-confident properties to preserve model capacity to learn more robust features. Extensive experiments across various datasets and networks illustrate that AUTE achieves superior performance compared to baselines. For instance, a 5-member AUTE with ResNet-20 networks outperforms state-of-the-art method by 2.1% and 3.2% in classifying clean and adversarial data. Additionally, AUTE can easily extend to non-adversarial training paradigm, surpassing current standard ensemble learning methods by a large margin. Lifeng Huang, Tian Su, Chengying Gao, Qiong Huang 0001 |
AAAI | 3 |
| 2025 | Feature Replacement in Gaussian Splatting for 3D Stylization
Jinkeng Zhu, Chengying Gao |
CGI (1) | 3 |
| 2025 | Efficient Integration of Neural Representations for Dynamic HumansabstractWhile numerous studies have explored NeRF-based novel view synthesis for dynamic humans, they often require training that exceeds several hours, limiting their practicality. Efforts to improve training efficiency have also encountered challenges because it is hard to optimize non-rigid transformations, thus leading to coarse renderings. In this work, we introduce an innovative approach for efficiently learning and integrating neural human representations. To achieve this, we propose a comprehensive utilization of the features stored in both canonical and observational spaces, facilitated through a collaborative refinement process that integrates canonical representations with observational details. Specifically, we initially propose decomposing high-dimensional multi-space feature volume into several feature planes, subsequently utilizing matrix multiplication to explicitly establish the correlations between different planes. This enables the simultaneous optimization of their counterparts across all dimensions by optimizing interpolated features, efficiently integrating associated details, and accelerating the rate of convergence. Additionally, we use the proposed collaborative refinement process to iteratively enhance the canonical representation. By integrating multi-space representations, we further facilitate the co-optimization of multiple frames' time-dependent observations. Experiments demonstrate that our method can achieve high-quality free-viewpoint renderings within nearly 5 minutes of optimization. Compared to state-of-the-art approaches, our results show more realistic rendering details, marking a significant advancement in both performance and efficiency. Lingzhe Zeng, Chengying Gao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | PianoBART: Symbolic Piano Music Generation and Understanding with Large-Scale Pre-TrainingabstractLearning musical structures and composition patterns is necessary for both music generation and understanding, but current methods do not make uniform use of learned features to generate and comprehend music simultaneously. In this paper, we propose PianoBART, a pre-trained model that uses BART for both symbolic piano music generation and understanding. We devise a multi-level object selection strategy for different pre-training tasks of PianoBART, which can prevent information leakage or loss and enhance learning ability. The musical semantics captured in pre-training are fine-tuned for music generation and understanding tasks. Experiments demonstrate that PianoBART efficiently learns musical patterns and achieves outstanding performance in generating high-quality coherent pieces and comprehending music. Our code and supplementary material are available at https://github.com/RS2002/PianoBart. Zijian Zhao 0002, Weichao Zeng, Fupeng He, Yiyi Wang, Chengying Gao |
ICME | 7 |
| 2024 | Text-Based Vector Sketch Editing with Image Editing Diffusion PriorabstractWe present a framework for text-based vector sketch editing to improve the efficiency of graphic design. The key idea behind the approach is to transfer the prior information from raster-level diffusion models, especially those from image editing methods, into the vector sketch-oriented task. The framework presents three editing modes and allows iterative editing. To meet the editing requirement of modifying the intended parts only while avoiding changing the other strokes, we introduce a stroke-level local editing scheme that automatically produces an editing mask reflecting locally editable regions and modifies strokes within the regions only. Comparisons with existing methods demonstrate the superiority of our approach. Haoran Mo, Xusheng Lin, Chengying Gao, Ruomei Wang 0001 |
ICME | 3 |
| 2024 | Video-Driven Sketch Animation Via Cyclic Reconstruction MechanismabstractConsidering the time-consuming manual workflow in 2D sketch animation production, we present an automatic solution by using videos as reference to animate the static sketch images. This includes motion extraction from the videos and injection into the sketches to produce animated sketch sequences in which appearance properties from the source sketches should be preserved. To reduce blurry artifact caused by complex motions and maintain stroke line continuity, we propose to incorporate inner masks of the sketches as an explicit guidance to indicate inner regions and ensure component integrality. Moreover, to bridge the domain gap between the video frames and the sketches when modelling the motions, we introduce a cyclic reconstruction mechanism to increase compatibility with different domains and improve motion consistency between the sketch animation and the driving video. Extensive results demonstrate the superiority of our method that outperforms existing methods both quantitatively and qualitatively. Zhuo Xie, Haoran Mo, Chengying Gao |
ICME | 3 |
| 2024 | Controllable Anime Image Editing via Probability of Attribute TagsabstractAbstract Editing anime images via probabilities of attribute tags allows controlling the degree of the manipulation in an intuitive and convenient manner. Existing methods fall short in the progressive modification and preservation of unintended regions in the input image. We propose a controllable anime image editing framework based on adjusting the tag probabilities, in which a probability encoding network (PEN) is developed to encode the probabilities into features that capture continuous characteristic of the probabilities. Thus, the encoded features are able to direct the generative process of a pre‐trained diffusion model and facilitate the linear manipulation. We also introduce a local editing module that automatically identifies the intended regions and constrains the edits to be applied to those regions only, which preserves the others unchanged. Comprehensive comparisons with existing methods indicate the effectiveness of our framework in both one‐shot and linear editing modes. Results in additional applications further demonstrate the generalization ability of our approach. Zhenghao Song, Haoran Mo, Chengying Gao |
Comput. Graph. Forum | 3 |
| 2024 | LAFED: Towards robust ensemble models via Latent Feature Diversification
Wenzi Zhuang, Lifeng Huang, Chengying Gao |
Pattern Recognit. | 3 |
| 2024 | FASTEN: Fast Ensemble Learning for Improved Adversarial RobustnessabstractRecent works show that adversarial attacks threaten the security of deep neural networks (DNNs). To tackle this issue, ensemble learning methods have been proposed to train multiple sub-models and improve adversarial resistance without compromising accuracy. However, these methods often come with high computational costs, including multi-step optimization to generate high-quality augmentation data and additional network passes to optimize complicated regularization. In this paper, we present the FAST ENsemble learning method (FASTEN) to significantly reduce training costs in terms of data and optimization. Firstly, FASTEN employs a single-step technique to initialize poor augmentation data and recycles optimization knowledge to enhance data quality, which considerably reduces the data generation budget. Secondly, FASTEN introduces a low-cost regularizer to increase intra-model similarity and inter-model diversity, with most of the regularization components computed without network passes, further decreasing training costs. Empirical results on various datasets and networks demonstrate that FASTEN achieves higher robustness while requiring significantly fewer resources than current methods. For example, a 5-member FASTEN speeds up the optimization process by$7\times $and$28\times $compared to state-of-the-art DVERGE and TRS, respectively. Moreover, FASTEN outperforms the stronger of the two methods by 26.3% and 6.1% under black-box and white-box attacks, respectively. FASTEN is also compatible with existing fast adversarial training techniques, making it an advantageous choice for enhancing robustness without incurring excessive costs. The source code is publicly available athttps://github.com/mesunhlf/FASTEN. Lifeng Huang, Qiong Huang 0001, Peichao Qiu, Shuxin Wei, Chengying Gao |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | DanceComposer: Dance-to-Music Generation Using a Progressive Conditional Music GeneratorabstractA wonderful piece of music is the essence and soul of dance, which motivates the study of automatic music generation for dance. To create appropriate music from dance, cross-modal correlations between dance and music such as rhythm and style, should be considered. However, existing dance-to-music methods have difficulties in achieving rhythmic alignment and stylistic matching simultaneously. Additionally, the diversity of generated samples is limited due to the lack of available paired data. To address these issues, we propose DanceComposer, a novel dance-to-music framework, which generates rhythmically and stylistically consistent multi-track music from dance videos. DanceComposer features a Progressive Conditional Music Generator (PCMG) that gradually incorporates rhythm and style constraints, enabling both rhythmic alignment and stylistic matching. To enhance style control, we introduce a Shared Style Module (SSM) that learns cross-modal features as stylistic constraints. This allows the PCMG can be trained on extensive music-only data and diversifies generated pieces. Quantitative and qualitative results show that our method surpasses the state-of-the-art in overall music quality, rhythmic consistency, and stylistic consistency. Lifeng Huang, Chengying Gao |
IEEE Trans. Multim. | 4 |
| 2024 | Joint Stroke Tracing and Correspondence for 2D AnimationabstractTo alleviate human labor in redrawing keyframes with ordered vector strokes for automatic inbetweening, we for the first time propose a joint stroke tracing and correspondence approach. Given consecutive raster keyframes along with a single vector image of the starting frame as a guidance, the approach generates vector drawings for the remaining keyframes while ensuring one-to-one stroke correspondence. Our framework trained on clean line drawings generalizes to rough sketches, and the generated results can be imported into inbetweening systems to produce inbetween sequences. Hence, the method is compatible with standard 2D animation workflow. An adaptive spatial transformation module (ASTM) is introduced to handle non-rigid motions and stroke distortion. We collect a dataset for training with 10k+ pairs of raster frames and their vector drawings with stroke correspondence. Comprehensive validations on real clean and rough animated frames manifest the effectiveness of our method and superiority to existing methods. Haoran Mo, Chengying Gao, Ruomei Wang 0001 |
ACM Trans. Graph. | 2 |
| 2023 | CAP-VSTNet: Content Affinity Preserved Versatile Style TransferabstractContent affinity loss including feature and pixel affinity is a main problem which leads to artifacts in photorealistic and video style transfer. This paper proposes a new framework named CAP-VSTNet, which consists of a new reversible residual network and an unbiased linear transform module, for versatile style transfer. This reversible residual network can not only preserve content affinity but not introduce redundant information as traditional reversible networks, and hence facilitate better stylization. Empowered by Matting Laplacian training loss which can address the pixel affinity loss problem led by the linear transform, the proposed framework is applicable and effective on versatile style transfer. Extensive experiments show that CAP-VSTNet can produce better qualitative and quantitative results in comparison with the state-of-the-art methods. Linfeng Wen 0003, Chengying Gao, Changqing Zou |
CVPR | 2 |
| 2023 | Controllable Garment Image Synthesis Integrated with Frequency Domain FeaturesabstractAbstract Using sketches and textures to synthesize garment images is able to conveniently display the realistic visual effect in the design phase, which greatly increases the efficiency of fashion design. Existing garment image synthesis methods from a sketch and a texture tend to fail in working on complex textures, especially those with periodic patterns. We propose a controllable garment image synthesis framework that takes as inputs an outline sketch and a texture patch and generates garment images with complicated and diverse texture patterns. To improve the performance of global texture expansion, we exploit the frequency domain features in the generative process, which are from a Fast Fourier Transform (FFT) and able to represent the periodic information of the patterns. We also introduce a perceptual loss in the frequency domain to measure the similarity of two texture pattern patches in terms of their intrinsic periodicity and regularity. Comparisons with existing approaches and sufficient ablation studies demonstrate the effectiveness of our method that is capable of synthesizing impressive garment images with diverse texture patterns while guaranteeing proper texture expansion and pattern consistency. Xinru Liang, Haoran Mo, Chengying Gao |
Comput. Graph. Forum | 3 |
| 2023 | Feature-preserving color pencil drawings from photographsabstractColor pencil drawing is well-loved due to its rich expressiveness. This paper proposes an approach for generating feature-preserving color pencil drawings from photographs. To mimic the tonal style of color pencil drawings, which are much lighter and have relatively lower saturation than photographs, we devise a lightness enhancement mapping and a saturation reduction mapping. The lightness mapping is a monotonically decreasing derivative function, which not only increases lightness but also preserves input photograph features. Color saturation is usually related to lightness, so we suppress the saturation dependent on lightness to yield a harmonious tone. Finally, two extremum operators are provided to generate a foreground-aware outline map in which the colors of the generated contours and the foreground object are consistent. Comprehensive experiments show that color pencil drawings generated by our method surpass existing methods in tone capture and feature preservation. Dong Wang 0041, Guiqing Li, Chengying Gao, Shengwu Fu, Yun Liang 0003 |
Comput. Vis. Media | 3 |
| 2023 | Erosion Attack: Harnessing Corruption To Improve Adversarial ExamplesabstractAlthough adversarial examples pose a serious threat to deep neural networks, most transferable adversarial attacks are ineffective against black-box defense models. This may lead to the mistaken belief that adversarial examples are not truly threatening. In this paper, we propose a novel transferable attack that can defeat a wide range of black-box defenses and highlight their security limitations. We identify two intrinsic reasons why current attacks may fail, namely data-dependency and network-overfitting. They provide a different perspective on improving the transferability of attacks. To mitigate the data-dependency effect, we propose the Data Erosion method. It involves finding special augmentation data that behave similarly in both vanilla models and defenses, to help attackers fool robustified models with higher chances. In addition, we introduce the Network Erosion method to overcome the network-overfitting dilemma. The idea is conceptually simple: it extends a single surrogate model to an ensemble structure with high diversity, resulting in more transferable adversarial examples. Two proposed methods can be integrated to further enhance the transferability, referred to as Erosion Attack (EA). We evaluate the proposed EA under different defenses that empirical results demonstrate the superiority of EA over existing transferable attacks and reveal the underlying threat to current robust models. The source code is publicly available at https://github.com/mesunhlf/EA. Lifeng Huang, Chengying Gao |
IEEE Trans. Image Process. | 2 |
| 2022 | Unpaired Motion Style Transfer with Motion-Oriented Projection Flow NetworkabstractExisting motion style transfer methods trained with unpaired samples tend to generate motions with inconsistent content or inconsistent number of frames when compared with the source motion. Moreover, due to the limited training samples, these methods perform worse in unseen style. In this paper, we propose a novel unpaired motion style transfer framework that generates complete stylized motions with consistent content. We introduce a motion-oriented projection flow network (M-PFN) designed for temporal motion data, which encodes the content and style motions into latent codes and decodes the stylized features produced by adaptive instance normalization (AdaIN) into stylized motions. The M-PFN contains dedicated operations and modules, e.g., Transformer, to process the temporal information of motions, which help to improve the continuity of the generated motions. Comparisons with the state-of-the-art methods show that our method effectively transfers the style of the motions while retaining the complete content and has stronger generalization ability in unseen style features. Haoran Mo, Chengying Gao |
ICME | 4 |
| 2022 | 3D interacting hand pose and shape estimation from a single RGB image
Chengying Gao, Yujia Yang |
Neurocomputing | 1 |
| 2022 | DEFEAT: Decoupled feature attack across deep neural networks
Lifeng Huang, Chengying Gao |
Neural Networks | 2 |
| 2022 | Cyclical Adversarial Attack Pierces Black-box Deep Neural Networks
Lifeng Huang, Shuxin Wei, Chengying Gao |
Pattern Recognit. | 3 |
| 2021 | Enhancing Adversarial Examples Via Self-AugmentationabstractRecently, adversarial attacks pose a challenge for the security of Deep Neural Networks, which motivates researchers to establish various defense methods. However, do current defenses really achieve real security? To answer the question, we propose self-augmentation method (SA) for circumventing defenders to transferable adversarial examples. Concretely, self-augmentation includes two strategies: (1) self-ensemble, which applies additional convolution layers to an existing model to build diverse virtual models that be fused for achieving an ensemble-model effect and preventing overfitting; and (2) deviation-augmentation, which based on the observation of defense models that the input data is surrounded by highly curved loss surfaces, thus inspiring us to apply deviation vectors to input data for escaping from their vicinity space. Extensive experiments conducted on four vanilla models and ten defenses suggest the superiority of our method compared with the state-of-the-art transferable attacks. The source code is public available at https://github.com/zhuangwz/ICME2021_self_augmentation. Lifeng Huang, Chengying Gao, Wenzi Zhuang |
ICME | 2 |
| 2021 | Structural Prior Guided Image Inpainting for Complex SceneabstractOne of the main challenges faced by many existing image inpainting methods is that they fail to generate clear boundaries between different objects in a complex scene. To address this problem, we propose a novel two-stage image completion framework with semantic segmentation maps as the structure prior to restrain texture generation. Our framework consists of a PSP-based adversarial network called SP-Net, which can generate a semantic segmentation map for the corrupted image and an SG-Net using domain alignment and correlation matrix to get the feature correspondence between structural prior and known region. Combining the structural prior and feature correspondence, the image completion network fills the missing region with distinct boundaries and fine details. The proposed method was evaluated over Outdoor Scenes and Cityscapes datasets. Experimental results demonstrate that our approach outperforms other structural prior based approaches in terms of visual quality and quantitative measurements. Shuxin Wei, Chengying Gao |
ICME | 2 |
| 2021 | Line Art Colorization Based on Explicit Region SegmentationabstractAbstract Automatic line art colorization plays an important role in anime and comic industry. While existing methods for line art colorization are able to generate plausible colorized results, they tend to suffer from the color bleeding issue. We introduce an explicit segmentation fusion mechanism to aid colorization frameworks in avoiding color bleeding artifacts. This mechanism is able to provide region segmentation information for the colorization process explicitly so that the colorization model can learn to avoid assigning the same color across regions with different semantics or inconsistent colors inside an individual region. The proposed mechanism is designed in a plug‐and‐play manner, so it can be applied to a diversity of line art colorization frameworks with various kinds of user guidances. We evaluate this mechanism in tag‐based and reference‐based line art colorization tasks by incorporating it into the state‐of‐the‐art models. Comparisons with these existing models corroborate the effectiveness of our method which largely alleviates the color bleeding artifacts. The code is available at https://github.com/Ricardo-L-C/ColorizationWithRegion . Ruizhi Cao, Haoran Mo, Chengying Gao |
Comput. Graph. Forum | 3 |
| 2021 | General virtual sketching framework for vector line artabstractVector line art plays an important role in graphic design, however, it is tedious to manually create. We introduce a general framework to produce line drawings from a wide variety of images, by learning a mapping from raster image space to vector image space. Our approach is based on a recurrent neural network that draws the lines one by one. A differentiable rasterization module allows for training with only supervised raster data. We use a dynamic window around a virtual pen while drawing lines, implemented with a proposed aligned cropping and differentiable pasting modules. Furthermore, we develop a stroke regularization loss that encourages the model to use fewer and longer strokes to simplify the resulting vector image. Ablation studies and comparisons with existing methods corroborate the efficiency of our approach which is able to generate visually better results in less computation time, while generalizing better to a diversity of images and applications. Haoran Mo, Edgar Simo-Serra, Chengying Gao, Changqing Zou, Ruomei Wang 0001 |
ACM Trans. Graph. | 3 |
| 2021 | Automatic 3D virtual fitting system based on skeleton driving
Guangyuan Shi, Chengying Gao, Dong Wang 0041, Zhuo Su 0001 |
Vis. Comput. | 2 |
| 2020 | SketchyCOCO: Image Generation From Freehand Scene SketchesabstractWe introduce the first method for automatic image generation from scene-level freehand sketches. Our model allows for controllable image generation by specifying the synthesis goal via freehand sketches. The key contribution is an attribute vector bridged Generative Adversarial Network called EdgeGAN, which supports high visual-quality object-level image content generation without using freehand sketches as training data. We have built a large-scale composite dataset called SketchyCOCO to support and evaluate the solution. We validate our approach on the tasks of both object-level and scene-level image generation on SketchyCOCO. Through quantitative, qualitative results, human evaluation and ablation studies, we demonstrate the method's capacity to generate realistic complex scene-level images from various freehand sketches. Chengying Gao, Limin Wang 0002, Jianzhuang Liu, Changqing Zou |
CVPR | 1 |
| 2020 | Universal Physical Camouflage Attacks on Object DetectorsabstractIn this paper, we study physical adversarial attacks on object detectors in the wild. Previous works mostly craft instance-dependent perturbations only for rigid or planar objects. To this end, we propose to learn an adversarial pattern to effectively attack all instances belonging to the same object category, referred to as Universal Physical Camouflage Attack (UPC). Concretely, UPC crafts camouflage by jointly fooling the region proposal network, as well as misleading the classifier and the regressor to output errors. In order to make UPC effective for non-rigid or non-planar objects, we introduce a set of transformations for mimicking deformable properties. We additionally impose optimization constraint to make generated patterns look natural to human observers. To fairly evaluate the effectiveness of different physical-world attacks, we present the first standardized virtual database, AttackScenes, which simulates the real 3D world in a controllable and reproducible environment. Extensive experiments suggest the superiority of our proposed UPC compared with existing physical adversarial attackers not only in virtual environments (AttackScenes), but also in real-world physical environments. Lifeng Huang, Chengying Gao, Yuyin Zhou, Cihang Xie, Alan L. Yuille, Changqing Zou |
CVPR | 2 |
| 2020 | An Asymmetric Modeling for Action Assessment
Jibin Gao, Wei-Shi Zheng 0001, Chengying Gao, Yaowei Wang 0001, Wei Zeng 0006, Jian-Huang Lai |
ECCV (30) | 4 |
| 2020 | Scale-Aware Rolling Fusion Network for Crowd CountingabstractDue to wide application prospects and various challenges such as large scale variation, inter-occlusion between crowd people and background noise, crowd counting is receiving increasing attention. In this paper, we propose a scale-aware rolling fusion network (SRF-Net) for crowd counting, which focuses on dealing with scale variation in highly congested noisy scenes. SRF-Net is a two-stage architecture that consists of a band-pass stage and a rolling guidance stage. Compared with the existing methods, SRF-Net achieves better results in retaining appropriate multi-level features and capturing multi-scale features, thus improving the quality of density estimation maps in crowded scenarios with large scale variation. We evaluate our method on three popular crowd counting datasets (ShanghaiTech, UCF_CC_50 and UCF-QNRF), and extensive experiments show its outperformance over the state-of-the-art approaches. Chengying Gao, Zhuo Su 0001, Xiangjian He |
ICME | 2 |
| 2020 | Self-Bootstrapping Pedestrian Detection in Downward-Viewing Fisheye Cameras Using Pseudo-LabelingabstractDownward-viewing fisheye cameras have attracted much attention in surveillance systems due to the wide coverage and less occlusion. However, pedestrian detection in downward-viewing fisheye cameras remains an open problem due to a lack of large-scale labeled dataset. Furthermore, it's time-consuming and labor-intensive to label a downward-viewing fisheye dataset manually. To address this, we propose a self-bootstrapping pedestrian detection method, which automatically pseudo-labels downward-viewing fisheye images by making full use of spatial and temporal consistency of pedestrians in the cameras to improve the accuracy of pedestrian detection. We segment the downward-viewing fisheye images into two regions and propose the pseudo-labeling methods for them progressively: a cyclic fine-tuned detector for the oblique region and a visual tracking method for the vertical region. Combining the pseudo-labels from two regions, we fine-tune the network for better accuracy. Experimental results show that the proposed approach reduces time consumption by about 95% compared with the labor-intensive manual labeling while it still reaches competitive and comparable Average Precision (AP). Kaishi Gao, Qun Niu, Haoquan You, Chengying Gao |
ICME | 4 |
| 2020 | Scale-aware Progressive Optimization NetworkabstractCrowd counting has attracted increasing attention due to its wide application prospect. One of the most essential challenge in this domain is large scale variation, which impacts the accuracy of density estimation. To this end, we propose a scale-aware progressive optimization network (SPO-Net) for crowd counting, which trains a scale adaptive network to achieve high-quality density map estimation and overcome the variable scale dilemma in highly congested scenes. Concretely, the first phase of SPO-Net, band-pass stage, mainly concentrates on preprocessesing the input image and fusing both high-level semantic information and low-level spatial information from separated multi-layer features. And the second phase of SPO-Net, rolling guidance stage, aims to learn a scale-adapted network from multi-scale features as well as rolling training manner. For better learning local correlation of multi-size regions and reducing redundant calculations, we introduce a progressive optimization strategy. Extensive experiments on three challenging crowd counting datasets not only demonstrate the efficacy of each part in SPO-Net, but also suggest the superiority of our proposed method compared with the state-of-the-art approaches. Lifeng Huang, Chengying Gao |
ACM Multimedia | 3 |
| 2020 | A completely parallel surface reconstruction method for particle-based fluids
Wencong Yang, Chengying Gao |
Vis. Comput. | 2 |
| 2019 | G-UAP: Generic Universal Adversarial Perturbation that Fools RPN-based DetectorsabstractAdversarial perturbation constructions have been demonstrated for object detection, but these are image-specific perturbations. Recent works have shown the existence of image-agnostic perturbations called universal adversarial perturbation (UAP) that can fool the classifiers over a set of natural images. In this paper, we extend this kind perturbation to attack deep proposal-based object detectors. We present a novel and effective approach called G-UAP to craft universal adversarial perturbations, which can explicitly degrade the detection accuracy of a detector on a wide range of image samples. Our method directly misleads the Region Proposal Network (RPN) of the detectors into mistaking foreground (objects) for background without specifying an adversarial label for each target (RPN’s proposal), and even without considering that how many objects and object-like targets are in the image. The experimental results over three state-of-the-art detectors and two datasets demonstrate the effectiveness of the proposed method and transferability of the universal perturbations. Lifeng Huang, Chengying Gao |
ACML | 3 |
| 2019 | Image-Based Virtual Try-on Network with Structural CoherenceabstractVirtual try-on system could demonstrate the visual effect of wearing certain clothes, which is in a great demand for the online clothing customers. Due to the neglect of the original person structure, the previous try-on schemes usually encounter with the problems like the loss of body parts, the missing of body details and the deviation of clothing style. In this paper, we propose a novel image-based virtual try-on network, which could maintain the structural consistency between the generated image and the original image by human parsing. Our network consists of three components. Given the original person and the clothing images, the target human parsing maps are generated. Then, the parsing maps are matched with the target clothes to generate the warped clothes. Finally, according to the parsing results, the parts to be replaced the original images are intercepted, and more original information is retained as the input of the network to generate the final results. Experiments on an existing benchmark demonstrate our method maintains the consistency of structure and achieves the state-of-the-art performance. Jiaming Guo, Zhuo Su 0001, Chengying Gao |
ICIP | 4 |
| 2019 | Rain Wiper: An Incremental Randomly Wired Network for Single Image DerainingabstractAbstract Single image rain removal is a challenging ill‐posed problem due to various shapes and densities of rain streaks. We present a novel incremental randomly wired network (IRWN) for single image deraining. Different from previous methods, most structures of modules in IRWN are generated by a stochastic network generator based on the random graph theory, which ease the burden of manual design and further help to characterize more complex rain streaks. To decrease network parameters and extract more details efficiently, the image pyramid is fused via the multi‐scale network structure. An incremental rectified loss is proposed to better remove rain streaks in different rain conditions and recover the texture information of target objects. Extensive experiments on synthetic and real‐world datasets demonstrate that the proposed method outperforms the state‐of‐the‐art methods significantly. In addition, an ablation study is conducted to illustrate the improvements obtained by different modules and loss items in IRWN. Xiangguo Liang, Bin Qiu, Zhuo Su 0001, Chengying Gao, X. Shi, Ruomei Wang 0001 |
Comput. Graph. Forum | 4 |
| 2019 | Language-based colorization of scene sketchesabstractBeing natural, touchless, and fun-embracing, language-based inputs have been demonstrated effective for various tasks from image generation to literacy education for children. This paper for the first time presents a language-based system for interactive colorization of scene sketches, based on semantic comprehension. The proposed system is built upon deep neural networks trained on a large-scale repository of scene sketches and cartoonstyle color images with text descriptions. Given a scene sketch, our system allows users, via language-based instructions, to interactively localize and colorize specific foreground object instances to meet various colorization requirements in a progressive way. We demonstrate the effectiveness of our approach via comprehensive experimental results including alternative studies, comparison with the state-of-the-art methods, and generalization user studies. Given the unique characteristics of language-based inputs, we envision a combination of our interface with a traditional scribble-based interface for a practical multimodal colorization system, benefiting various applications. The dataset and source code can be found at https://github.com/SketchyScene/SketchySceneColorization. Changqing Zou, Haoran Mo, Chengying Gao, Ruofei Du, Hongbo Fu 0001 |
ACM Trans. Graph. | 3 |
| 2019 | Resource-efficient and Automated Image-based Indoor LocalizationabstractImage-based indoor localization has aroused much interest recently because it requires no infrastructure support. Previous approaches on image-based localization, due to their computation and storage requirements, often process queries at servers. This does not scale well, incurs round-trip delay, and requires constant network connectivity. Many also require users to manually confirm the shortlisted matched landmarks, which is inconvenient, slow, and prone to selection error. To overcome these limitations, we propose a h ighly a utomated (in terms of image confirmation after taking images) i mage-based l ocalization algorithm (HAIL), distributed in mobile devices. HAIL achieves resource efficiency (in terms of storage and processing) by keeping only distinguishing visual features for each landmark, and employing the efficient k-d tree to search for features. It further utilizes motion sensors and map constraints to enhance the localization accuracy without user operation. We have implemented HAIL on Android platforms and conducted extensive experiments in a food plaza and a premium shopping mall. Experimental results show that it achieves much higher localization accuracy (reducing the localization error by more than 20%) and computation efficiency (by more than 40% in time) as compared with the state-of-the-art approaches. Qun Niu, Mingkuan Li, Suining He, Chengying Gao, Shueng-Han Gary Chan |
ACM Trans. Sens. Networks | 4 |
| 2018 | SketchyScene: Richly-Annotated Scene Sketches
Changqing Zou, Qian Yu 0002, Ruofei Du, Haoran Mo, Yi-Zhe Song, Tao Xiang 0002, Chengying Gao, Baoquan Chen, Hao (Richard) Zhang |
ECCV (15) | 7 |
| 2018 | PencilArt: A Chromatic Penciling Style Generation FrameworkabstractAbstract Non‐photorealistic rendering has been an active area of research for decades whereas few of them concentrate on rendering chromatic penciling style. In this paper, we present a framework named as PencilArt for the chromatic penciling style generation from wild photographs. The structural outline and textured map for composing the chromatic pencil drawing are generated, respectively. First, we take advantage of deep neural network to produce the structural outline with proper intensity variation and conciseness. Next, for the textured map, we follow the painting process of artists to adjust the tone of input images to match the luminance histogram and pencil textures of real drawings. Eventually, we evaluate PencilArt via a series of comparisons to previous work, showing that our results better capture the main features of real chromatic pencil drawings and have an improved visual appearance. Chengying Gao, Mengyue Tang, Xiangguo Liang, Zhuo Su 0001, Changqing Zou |
Comput. Graph. Forum | 1 |
| 2018 | An edge-refined vectorized deep colorization model for grayscale-to-color images
Zhuo Su 0001, Xiangguo Liang, Jiaming Guo, Chengying Gao |
Neurocomputing | 4 |
| 2017 | ℒ0 Gradient-Preserving Color TransferabstractAbstract This paper presents a new two‐step color transfer method which includes color mapping and detail preservation. To map source colors to target colors, which are from an image or palette, the proposed similarity‐preserving color mapping algorithm uses the similarities between pixel color and dominant colors as existing algorithms and emphasizes the similarities between source image pixel colors. Detail preservation is performed by an ℒ0 gradient‐preserving algorithm. It relaxes the large gradients of the sparse pixels along color region boundaries and preserves the small gradients of pixels within color regions. The proposed method preserves source image color similarity and image details well. Extensive experiments demonstrate that the proposed approach has achieved a state‐of‐art visual performance. Dong Wang 0041, Changqing Zou, Guiqing Li, Chengying Gao, Zhuo Su 0001 |
Comput. Graph. Forum | 4 |
| 2017 | Data-driven image completion for complex objects
Chengying Gao, Yanmei Luo, Hefeng Wu, Dong Wang 0041 |
Signal Process. Image Commun. | 1 |
| 2014 | Mesh-based anisotropic cloth deformation for virtual fitting
Li Liu 0032, Ruomei Wang 0001, Zhuo Su 0001, Chengying Gao |
Multim. Tools Appl. | 5 |
| 2013 | Cartoon Rendering Illumination Model Based on Phongabstract3D cartoon rendering has broad application prospect in games, movies and cartoons. To this end, this paper introduces a cartoon rendering illumination model improved from Phong illumination model, and realizes it with 3ds Max SDK in the form of plug-ins. By discretizing diffuse reflection part and specular reflection part in the Phong illumination model, surfaces of models show different blocks of color, among which there are clear boundaries. In addition, by using linear interpolation between different color blocks, boundaries can be soften. The experimental results show that this cartoon rendering illumination model can not only make a cartoon appearance, but also mix with Phong illumination model, and achieve a unique effect. Shaohao Wang, Yurui Wei, Chengying Gao |
ICIG | 3 |
| 2011 | Interactive CT image segmentation with online discriminative learningabstractAlthough interactive image segmentation has been widely exploited, current approaches present unsatisfactory results in medical image processing. This paper proposes a fast method for interactive CT image segmentation in which the tumor regions should be partitioned as foreground against the healthy tissues. In contrast to natural images, we have the following observation on CT images: (1) CT images often include discontinuous silhouette or cluttered spots caused by input de- vices or patient corporeity; (2) Disease areas often have varying appearance and shape. We thus train a discriminative fore- ground/background model based on user-placed scribbles. In our method, we extract positive and negative samples according to the foreground and background scribbles respectively, and use dense SIFT descriptors plus gray-level histogram as candidate features. With online learning, segmentation can be fast solved by the Bregman iteration. We test our method on CT liver images and demonstrate the advantage by comparing to state-of-the-art approaches. Wei Yang 0019, Xiaolong Wang 0004, Liang Lin 0004, Chengying Gao |
ICIP | 4 |
| 2010 | Reversely Anisotropic Quad-dominant RemeshingabstractIn this paper we proposed an anisotropic quad-dominant remeshing algorithm suitable for meshes of arbitrary topology. It takes a novel approach to the challenging problem of constructing high-quality quad-dominant mesh with anisotropic sampling. The method based on exploiting and analyzing principal curvature lines of the surface, which aimed to improve mesh structure and efficiency. Connectivity optimized is guaranteed by the natural orthogonality of principal curvature lines and geometric shape is maintained by minimizing the Hausdorff distance between original mesh and resulting mesh. The technique is straightforward to implement and efficient enough to be applied to real-world models. It can flexibly produce quad-dominant meshes ranging from dense to coarse. WeiPeng Zhu, Chengying Gao |
Shape Modeling International | 2 |