EDBT 2026 Demo / reviewers in the wild / expert
Chenjie Cao
dblp:193/0823
· DBLP profile ↗
31ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0003-3916-2843ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 10 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 16 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent DiffusionabstractDespite the remarkable developments achieved by recent 3D generation works, scaling these methods to geographic extents, such as modeling thousands of square kilometers of Earth’s surface, remains an open challenge. We address this through a dual innovation in data infrastructure and model architecture. First, we introduce Aerial-Earth3D, the largest 3D aerial dataset to date, consisting of 50k curated scenes (each measuring 600m) captured across the U.S. mainland, comprising 45M multi-view Google Earth frames. Each scene provides pose-annotated multi-view images, depth maps, normals, semantic segmentation, and camera poses, with explicit quality control to ensure terrain diversity. Building on this foundation, we propose EarthCrafter, a tailored framework for large-scale 3D Earth generation via sparse-decoupled latent diffusion. Our architecture separates structural and textural generation: 1) Dual sparse 3D-VAEs compress high-resolution geometric voxels and textural 2D Gaussian Splats (2DGS) into compact latent spaces, largely alleviating the costly computation suffering from vast geographic scales while preserving critical information. 2) We propose condition-aware flow matching models trained on mixed inputs (semantics, images, or neither) to flexibly model latent geometry and texture features independently. Extensive experiments demonstrate that EarthCrafter performs substantially better in extremely large-scale generation. The framework further supports versatile applications, from semantic-guided urban layout generation to unconditional terrain synthesis, while maintaining geographic plausibility through our rich data priors from Aerial-Earth3D. Shang Liu 0002, Chenjie Cao, Chaohui Yu, Jing Wang 0224, Fan Wang 0019 |
AAAI | 2 |
| 2025 | Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color ConsistencyabstractRecent advances in image inpainting increasingly use generative models to handle large irregular masks. However, these models can create unrealistic inpainted images due to two main issues: (1) Unwanted object insertion: Even with unmasked areas as context, generative models may still generate arbitrary objects in the masked region that don’t align with the rest of the image. (2) Color inconsistency: Inpainted regions often have color shifts that causes a smeared appearance, reducing image quality. Retraining the generative model could help solve these issues, but it’s costly since state-of-the-art latent-based diffusion and rectified flow models require a three-stage training process: training a VAE, training a generative U-Net or transformer, and fine-tuning for inpainting. Instead, this paper proposes a post-processing approach, dubbed as ASUKA (Aligned Stable inpainting with UnKnown Areas prior), to improve inpainting models. To address unwanted object insertion, we leverage a Masked Auto-Encoder (MAE) for reconstruction-based priors. This mitigates object hallucination while maintaining the model’s generation capabilities. To address color inconsistency, we propose a specialized VAE decoder that treats latent-to-image decoding as a local harmonization task, significantly reducing color shifts for color-consistent inpainting. We validate ASUKA on SD 1.5 and FLUX inpainting variants with Places2 and MISATO, our proposed diverse collection of datasets. Results show that ASUKA mitigates object hallucination and improves color consistency over standard diffusion and rectified flow models and other inpainting methods. Yikai Wang 0002, Chenjie Cao, Junqiu Yu, Xiangyang Xue 0001, Yanwei Fu 0001 |
CVPR | 2 |
| 2025 | MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion ModelabstractWe introduce MVGenMaster, a multi-view diffusion model enhanced with 3D priors to address versatile Novel View Synthesis (NVS) tasks. MVGenMaster leverages 3D priors that are warped using metric depth and camera poses, significantly enhancing both generalization and 3D consistency in NVS. Our model features a simple yet effective pipeline that can generate up to 100 novel views conditioned on variable reference views and camera poses with a single forward process. Additionally, we have developed a comprehensive large-scale multi-view image dataset called MvD-1M, comprising up to 1.6 million scenes, equipped with well-aligned metric depth to train MVGenMaster. Moreover, we present several training and model modifications to strengthen the model with scaled-up datasets. Extensive evaluations across in- and out- of- domain benchmarks demonstrate the effectiveness of our proposed method and data formulation. Chenjie Cao, Chaohui Yu, Shang Liu 0002, Fan Wang 0019, Xiangyang Xue 0001, Yanwei Fu 0001 |
CVPR | 1 |
| 2025 | LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video DiffusionabstractVideo Diffusion Models (VDMs) have demonstrated remarkable capabilities in synthesizing realistic videos by learning from large-scale data. Although vanilla Low-Rank Adaptation (LoRA) can learn specific spatial or temporal movement to driven VDMs with constrained data, achieving precise control over both camera trajectories and object motion remains challenging due to the unstable fusion and non-linear scalability. To address these issues, we propose LiON-LoRA, a novel framework that rethinks LoRA fusion through three core principles: Linear scalability, Orthogonality, and Norm consistency. First, we analyze the orthogonality of LoRA features in shallow VDM layers, enabling decoupled low-level controllability. Second, norm consistency is enforced across layers to stabilize fusion during complex camera motion combinations. Third, a controllable token is integrated into the diffusion transformer (DiT) to linearly adjust motion amplitudes for both cameras and objects with a modified self-attention mechanism to ensure decoupled control. Additionally, we extend LiON-LoRA to temporal generation by leveraging static-camera videos, unifying spatial and temporal controllability. Experiments demonstrate that LiON-LoRA outperforms state-of-the-art methods in trajectory control accuracy and motion strength adjustment, achieving superior generalization with minimal training data. Project Page: https://fuchengsu.github.io/lionlora.github.io/ Yisu Zhang, Chenjie Cao, Chaohui Yu, Jianke Zhu |
ICCV | 2 |
| 2025 | Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video GenerationabstractCamera and human motion controls have been extensively studied for video generation, but existing approaches typically address them separately, suffering from limited data with high-quality annotations for both aspects. To overcome this, we present Uni3C, a unified 3D-enhanced framework for precise control of both camera and human motion in video generation. Uni3C includes two key contributions. First, we propose a plug-and-play control module trained with a frozen video generative backbone, PCDController, which utilizes unprojected point clouds from monocular depth to achieve accurate camera control. By leveraging the strong 3D priors of point clouds and the powerful capacities of video foundational models, PCDController shows impressive generalization, performing well regardless of whether the inference backbone is frozen or fine-tuned. This flexibility enables different modules of Uni3C to be trained in specific domains, i.e., either camera control or human motion control, reducing the dependency on jointly annotated data. Second, we propose a jointly aligned 3D world guidance for the inference phase that seamlessly integrates both scenic point clouds and SMPL-X characters to unify the control signals for camera and human motion, respectively. Extensive experiments confirm that PCDController enjoys strong robustness in driving camera motion for fine-tuned backbones of video generation. Uni3C substantially outperforms competitors in both camera controllability and human motion quality. Additionally, we collect tailored validation sets featuring challenging camera movements and human actions to validate the effectiveness of our method. Codes are released at https://github.com/alibaba-damo-academy/Uni3C. Chenjie Cao, Jingkai Zhou, Shikai Li, Jingyun Liang, Chaohui Yu, Fan Wang 0019, Xiangyang Xue 0001, Yanwei Fu 0001 |
SIGGRAPH Asia | 1 |
| 2024 | LeftRefill: Filling Right Canvas based on Left Reference through Generalized Text-to-Image Diffusion ModelabstractThis paper introduces LeftRefill, an innovative approach to efficiently harness large Text-to-Image (T2I) diffusion models for reference-guided image synthesis. As the name implies, LeftRefill horizontally stitches reference and target views together as a whole input. The reference image occupies the left side, while the target canvas is positioned on the right. Then, LeftRefill paints the rightside target canvas based on the left-side reference and specific task instructions. Such a task formulation shares some similarities with contextual inpainting, akin to the actions of a human painter. This novel formulation efficiently learns both structural and textured correspondence between reference and target without other image encoders or adapters. We inject task and view information through cross-attention modules in T2I models, and further exhibit multi-view reference ability via the re-arranged self-attention modules. These enable LeftRefill to perform consistent generation as a generalized model without requiring testtime fine-tuning or model modifications. Thus, LeftRefill can be seen as a simple yet unified framework to address reference-guided synthesis. As an exemplar, we leverage LeftRefill to address two different challenges: reference-guided inpainting and novel view synthesis, based on the pre-trained StableDiffusion. Codes&models are released at https://github.com/ewrfcas/LeftRefill. Chenjie Cao, Yunuo Cai, Qiaole Dong, Yikai Wang 0002, Yanwei Fu 0001 |
CVPR | 1 |
| 2024 | VCD-Texture: Variance Alignment Based 3D-2D Co-denoising for Text-Guided Texturing
Shang Liu 0002, Chaohui Yu, Chenjie Cao, Fan Wang 0019 |
ECCV (16) | 3 |
| 2024 | Improving Neural Surface Reconstruction with Feature Priors from Multi-view Images
Xinlin Ren, Chenjie Cao, Yanwei Fu 0001, Xiangyang Xue 0001 |
ECCV (58) | 2 |
| 2024 | SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer
Chaohui Yu, Yanqin Jiang, Chenjie Cao, Fan Wang 0019, Xiang Bai |
ECCV (13) | 4 |
| 2024 | MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View StereoabstractRecent advancements in learning-based Multi-View Stereo (MVS) methods have prominently featured transformer-based models with attention mechanisms. However, existing approaches have not thoroughly investigated the profound influence of transformers on different MVS modules, resulting in limited depth estimation capabilities. In this paper, we introduce MVSFormer++, a method that prudently maximizes the inherent characteristics of attention to enhance various components of the MVS pipeline. Formally, our approach involves infusing cross-view information into the pre-trained DINOv2 model to facilitate MVS learning. Furthermore, we employ different attention mechanisms for the feature encoder and cost volume regularization, focusing on feature and spatial aggregations respectively. Additionally, we uncover that some design details would substantially impact the performance of transformer modules in MVS, including normalized 3D positional encoding, adaptive attention scaling, and the position of layer normalization. Comprehensive experiments on DTU, Tanks-and-Temples, BlendedMVS, and ETH3D validate the effectiveness of the proposed method. Notably, MVSFormer++ achieves state-of-the-art performance on the challenging DTU and Tanks-and-Temples benchmarks. Codes and models are available at https://github.com/maybeLx/MVSFormerPlusPlus. Chenjie Cao, Xinlin Ren, Yanwei Fu 0001 |
ICLR | 1 |
| 2024 | MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D EditingabstractNovel View Synthesis (NVS) and 3D generation have recently achieved prominent improvements. However, these works mainly focus on confined categories or synthetic 3D assets, which are discouraged from generalizing to challenging in-the-wild scenes and fail to be employed with 2D synthesis directly. Moreover, these methods heavily depended on camera poses, limiting their real-world applications.
To overcome these issues, we propose MVInpainter, re-formulating the 3D editing as a multi-view 2D inpainting task. Specifically, MVInpainter partially inpaints multi-view images with the reference guidance rather than intractably generating an entirely novel view from scratch, which largely simplifies the difficulty of in-the-wild NVS and leverages unmasked clues instead of explicit pose conditions. To ensure cross-view consistency, MVInpainter is enhanced by video priors from motion components and appearance guidance from concatenated reference key\&value attention. Furthermore, MVInpainter incorporates slot attention to aggregate high-level optical flow features from unmasked regions to control the camera movement with pose-free training and inference. Sufficient scene-level experiments on both object-centric and forward-facing datasets verify the effectiveness of MVInpainter, including diverse tasks, such as multi-view object removal, synthesis, insertion, and replacement. The project page is https://ewrfcas.github.io/MVInpainter/. Chenjie Cao, Chaohui Yu, Fan Wang 0019, Xiangyang Xue 0001, Yanwei Fu 0001 |
NeurIPS | 1 |
| 2024 | Animate3D: Animating Any 3D Model with Multi-view Video DiffusionabstractRecent advances in 4D generation mainly focus on generating 4D content by distilling pre-trained text or single-view image conditioned models. It is inconvenient for them to take advantage of various off-the-shelf 3D assets with multi-view attributes, and their results suffer from spatiotemporal inconsistency owing to the inherent ambiguity in the supervision signals. In this work, we present Animate3D, a novel framework for animating any static 3D model. The core idea is two-fold: 1) We propose a novel multi-view video diffusion model (MV-VDM) conditioned on multi-view renderings of the static 3D object, which is trained on our presented large-scale multi-view video dataset (MV-Video). 2) Based on MV-VDM, we introduce a framework combining reconstruction and 4D Score Distillation Sampling (4D-SDS) to leverage the multi-view video diffusion priors for animating 3D objects. Specifically, for MV-VDM, we design a new spatiotemporal attention module to enhance spatial and temporal consistency by integrating 3D and video diffusion models. Additionally, we leverage the static 3D model’s multi-view renderings as conditions to preserve its identity. For animating 3D models, an effective two-stage pipeline is proposed: we first reconstruct coarse motions directly from generated multi-view videos, followed by the introduced 4D-SDS to model fine-level motions. Benefiting from accurate motion learning, we could achieve straightforward mesh animation. Qualitative and quantitative experiments demonstrate that Animate3D significantly outperforms previous approaches. Data, code, and models are open-released. Yanqin Jiang, Chaohui Yu, Chenjie Cao, Fan Wang 0019, Weiming Hu 0004 |
NeurIPS | 3 |
| 2023 | Rethinking Optical Flow from Geometric Matching Consistent PerspectiveabstractOptical flow estimation is a challenging problem remaining unsolved. Recent deep learning based optical flow models have achieved considerable success. However, these models often train networks from the scratch on standard optical flow data, which restricts their ability to robustly and geometrically match image features. In this paper, we propose a rethinking to previous optical flow estimation. We particularly leverage Geometric Image Matching (GIM) as a pre-training task for the optical flow estimation (MatchFlow) with better feature representations, as GIM shares some common challenges as optical flow estimation, and with massive labeled real-world data. Thus, matching static scenes helps to learn more fundamental feature correlations of objects and scenes with consistent displacements. Specifically, the proposed MatchFlow model employs a QuadTree attention-based network pre-trained on MegaDepth to extract coarse features for further flow regression. Extensive experiments show that our model has great cross-dataset generalization. Our method achieves 11.5% and 10.1% error reduction from GMA on Sintel clean pass and KITTI test set. At the time of anonymous submission, our MatchFlow(G) enjoys state-of-the-art performance on Sintel clean and final pass compared to published approaches with comparable computation and memory footprint. Codes and models will be released in https://github.com/DQiaole/MatchFlow. Qiaole Dong, Chenjie Cao, Yanwei Fu 0001 |
CVPR | 2 |
| 2023 | Improving Transformer-based Image Matching by Cascaded Capturing Spatially Informative KeypointsabstractLearning robust local image feature matching is a fundamental low-level vision task, which has been widely explored in the past few years. Recently, detector-free local feature matchers based on transformers have shown promising results, which largely outperform pure Convolutional Neural Network (CNN) based ones. But correlations produced by transformer-based methods are spatially limited to the center of source views’ coarse patches, because of the costly attention learning. In this work, we rethink this issue and find that such matching formulation degrades pose estimation, especially for low-resolution images. So we propose a transformer-based cascade matching model – Cascade feature Matching TRansformer (CasMTR)§, to efficiently learn dense feature correlations, which allows us to choose more reliable matching pairs for the relative pose estimation. Instead of re-training a new detector, we use a simple yet effective Non-Maximum Suppression (NMS) post-process to filter keypoints through the confidence map, and largely improve the matching precision. CasMTR achieves state-of-the-art performance in indoor and outdoor pose estimation as well as visual localization. Moreover, thorough ablations show the efficacy of the proposed components and techniques. Chenjie Cao, Yanwei Fu 0001 |
ICCV | 1 |
| 2023 | Local Consensus Enhanced Siamese Network with Reciprocal Loss for Two-view Correspondence LearningabstractRecent studies of two-view correspondence learning usually establish an end-to-end network to jointly predict correspondence reliability and relative pose. We improve such a framework from two aspects. First, we propose a Local Feature Consensus (LFC) plugin block to augment the features of existing models. Given a correspondence feature, the block augments its neighboring features with mutual neighborhood consensus and aggregates them to produce an enhanced feature. As inliers obey a uniform cross-view transformation and share more consistent learned features than outliers, feature consensus strengthens inlier correlation and suppresses outlier distraction, which makes output features more discriminative for classifying inliers/outliers. Second, existing approaches supervise network training with the ground truth correspondences and essential matrix projecting one image to the other for an input image pair, without considering the information from the reverse mapping. We extend existing models to a Siamese network with a reciprocal loss that exploits the supervision of mutual projection, which considerably promotes the matching performance without introducing additional model parameters. Building upon MSA-Net [30], we implement the two proposals and experimentally achieve state-of-the-art performance on benchmark datasets. Linbo Wang 0001, Xianyong Fang, Zhengyi Liu, Chenjie Cao, Yanwei Fu 0001 |
ACM Multimedia | 5 |
| 2023 | ZITS++: Image Inpainting by Improving the Incremental Transformer on Structural PriorsabstractImage inpainting involves filling missing areas of a corrupted image. Despite impressive results have been achieved recently, restoring images with both vivid textures and reasonable structures remains a significant challenge. Previous methods have primarily addressed regular textures while disregarding holistic structures due to the limited receptive fields of Convolutional Neural Networks (CNNs). To this end, we study learning a Zero-initialized residual addition based Incremental Transformer on Structural priors (ZITS++), an improved model upon our conference work, ZITS (Dong et al. 2022). Specifically, given one corrupt image, we present the Transformer Structure Restorer (TSR) module to restore holistic structural priors at low image resolution, which are further upsampled by Simple Structure Upsampler (SSU) module to higher image resolution. To recover image texture details, we use the Fourier CNN Texture Restoration (FTR) module, which is strengthened by Fourier and large-kernel attention convolutions. Furthermore, to enhance the FTR, the upsampled structural priors from TSR are further processed by Structure Feature Encoder (SFE) and optimized with the Zero-initialized Residual Addition (ZeroRA) incrementally. Besides, a new masking positional encoding is proposed to encode the large irregular masks. Compared with ZITS, ZITS++ improves the FTR's stability and inpainting ability with several techniques. More importantly, we comprehensively explore the effects of various image priors for inpainting and investigate how to utilize them to address high-resolution image inpainting with extensive experiments. This investigation is orthogonal to most inpainting approaches and can thus significantly benefit the community. Chenjie Cao, Qiaole Dong, Yanwei Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Pixel2Mesh++: 3D Mesh Generation and Refinement From Multi-View ImagesabstractWe study the problem of shape generation in 3D mesh representation from a small number of color images with or without camera poses. While many previous works learn to hallucinate the shape directly from priors, we adopt to further improve the shape quality by leveraging cross-view information with a graph convolution network. Instead of building a direct mapping function from images to 3D shape, our model learns to predict series of deformations to improve a coarse shape iteratively. Inspired by traditional multiple view geometry methods, our network samples nearby area around the initial mesh's vertex locations and reasons an optimal deformation using perceptual feature statistics built from multiple input images. Extensive experiments show that our model produces accurate 3D shapes that are not only visually plausible from the input perspectives, but also well aligned to arbitrary viewpoints. With the help of physically driven architecture, our model also exhibits generalization capability across different semantic categories, and the number of input images. Model analysis experiments show that our model is robust to the quality of the initial mesh and the error of camera pose, and can be combined with a differentiable renderer for test-time optimization. Chao Wen 0001, Yinda Zhang 0001, Chenjie Cao, Zhuwen Li, Xiangyang Xue 0001, Yanwei Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Incremental Transformer Structure Enhanced Image Inpainting with Masking Positional EncodingabstractImage inpainting has made significant advances in recent years. However, it is still challenging to recover corrupted images with both vivid textures and reasonable structures. Some specific methods only tackle regular textures while losing holistic structures due to the limited receptive fields of convolutional neural networks (CNNs). On the other hand, attention-based models can learn better long-range dependency for the structure recovery, but they are limited by the heavy computation for inference with large image sizes. To address these issues, we propose to leverage an additional structure restorer to facilitate the image inpainting incrementally. The proposed model restores holistic image structures with a powerful attention-based transformer model in a fixed low-resolution sketch space. Such a grayscale space is easy to be upsampled to larger scales to convey correct structural information. Our structure restorer can be integrated with other pretrained inpainting models efficiently with the zero-initialized residual addition. Furthermore, a masking positional encoding strategy is utilized to improve the performance with large irregular masks. Extensive experiments on various datasets validate the efficacy of our model compared with other competitors. Our codes are released in https://github.com/DQiaole/ZITS_inpainting. Qiaole Dong, Chenjie Cao, Yanwei Fu 0001 |
CVPR | 2 |
| 2022 | Learning Prior Feature and Attention Enhanced Image Inpainting
Chenjie Cao, Qiaole Dong, Yanwei Fu 0001 |
ECCV (15) | 1 |
| 2022 | High-Fidelity Portrait Editing Via Exploring Differentiable Guided Sketches from the Latent SpaceabstractThis paper studies the task of sketch-guided high-fidelity portrait editing. Advanced unconditional generators, such as StyleGAN, can generate a high-quality portrait image with great diversity. In previous researches, StyleGAN has successfully been utilized for color-guided image editing through latent vector optimization. Nonetheless, passing sketch information to the generating model directly is nontrivial. To this end, we present an algorithm that addresses the problem of well controlling the generation process via differentiable guided sketches from latent space. Specifically, we re-purpose the classic operator – eXtended difference-of-Gaussians (XDoG) that derives differentiable sketches from images. We also propose a multi-scale sketch loss assisted with which can finally guide the model follow the guidance sketch to generate. Extensive experiments validate the efficacy of our model in sketch-guided editing. We show that the quality of produced images is better than that of competitors. Chengrong Wang, Chenjie Cao, Yanwei Fu 0001, Xiangyang Xue 0001 |
ICASSP | 2 |
| 2022 | Boundary-based Fuzzy-SVDD for one-class classificationabstractSupport Vector Data Description (SVDD) is an extremely hot topic issue in One-Class Classification (OCC), which has displayed outstanding performance in dealing with many novelty detection problems. However, SVDD just takes the data description by the kernel-based distance among each instance into consideration rather than considering the distribution of the data. Therefore, Fuzzy Support Vector Data Description (Fuzzy-SVDD) has been developed to distribute a fuzzy membership to each input sample so that different samples cause different contributions to classification boundary. The majority of the methods in Fuzzy-SVDD are based on the sample density, but there are remaining two problems. These density-based Fuzzy-SVDD methods would decrease the contribution of support vectors (SVs) in low densities. What is more, these methods cannot get a precise density when there are few target samples. These two problems would lead to a poor classification boundary. To overcome these drawbacks, a novel method called Boundary-based Fuzzy-SVDD (BF-SVDD) is proposed in this paper. BF-SVDD uses a new definition called local–global center distance to search for the samples near the boundary. Then, it enhances fuzzy memberships of these samples because they carry more significant information for the decision boundary than other data. The contribution of this paper can be summarized into three main points. First a novel concept called local–global center distances is proposed to find the SVs better. Second, fuzzy memberships with local–global center distance make SVs more informative to create the decision boundary. Furthermore, the experiments based on University of California, Irvine and Knowledge Extraction based on Evolutionary Learning also show that the proposed method has excellent performances. Even for the minority class in imbalance data sets, the proposed method can also have a good classification. Dongdong Li 0003, Xinlei Xu, Zhe Wang 0002, Chenjie Cao, Minguang Wang |
Int. J. Intell. Syst. | 4 |
| 2021 | Learning a Sketch Tensor Space for Image Inpainting of Man-made ScenesabstractThis paper studies the task of inpainting man-made scenes. It is very challenging due to the difficulty in preserving the visual patterns of images, such as edges, lines, and junctions. Especially, most previous works are failed to restore the object/building structures for images of man-made scenes. To this end, this paper proposes learning a Sketch Tensor (ST) space for inpainting man-made scenes. Such a space is learned to restore the edges, lines, and junctions in images, and thus makes reliable predictions of the holistic image structures. To facilitate the structure refinement, we propose a Multi-scale Sketch Tensor inpainting (MST) network, with a novel encoder-decoder structure. The encoder extracts lines and edges from the input images to project them into an ST space. From this space, the decoder is learned to restore the input images. Extensive experiments validate the efficacy of our model. Furthermore, our model can also achieve competitive performance in inpainting general nature images over the competitors. Chenjie Cao, Yanwei Fu 0001 |
ICCV | 1 |
| 2021 | The Image Local Autoregressive TransformerabstractRecently, AutoRegressive (AR) models for the whole image generation empowered by transformers have achieved comparable or even better performance compared to Generative Adversarial Networks (GANs). Unfortunately, directly applying such AR models to edit/change local image regions, may suffer from the problems of missing global information, slow inference speed, and information leakage of local guidance. To address these limitations, we propose a novel model -- image Local Autoregressive Transformer (iLAT), to better facilitate the locally guided image synthesis. Our iLAT learns the novel local discrete representations, by the newly proposed local autoregressive (LA) transformer of the attention mask and convolution mechanism. Thus iLAT can efficiently synthesize the local image regions by key guidance information. Our iLAT is evaluated on various locally guided image syntheses, such as pose-guided person image synthesis and face editing. Both quantitative and qualitative results show the efficacy of our model. Chenjie Cao, Yuxin Hong, Chengrong Wang, Chengming Xu 0001, Yanwei Fu 0001, Xiangyang Xue 0001 |
NeurIPS | 1 |
| 2020 | CLUE: A Chinese Language Understanding Evaluation BenchmarkabstractLiang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, Yin Tian, Qianqian Dong, Weitang Liu, Bo Shi, Yiming Cui, Junyi Li, Jun Zeng, Rongzhao Wang, Weijian Xie, Yanting Li, Yina Patterson, Zuoyu Tian, Yiwen Zhang, He Zhou, Shaoweihua Liu, Zhe Zhao, Qipeng Zhao, Cong Yue, Xinrui Zhang, Zhengliang Yang, Kyle Richardson, Zhenzhong Lan. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Liang Xu 0011, Hai Hu 0001, Xuanwei Zhang, Chenjie Cao, Yudong Li 0001, Yechen Xu, Kai Sun 0006, Dian Yu 0001, Cong Yu 0010, Yin Tian, Qianqian Dong, Weitang Liu, Yiming Cui 0001, Rongzhao Wang, Weijian Xie, Yina Patterson, Zuoyu Tian, Shaoweihua Liu, Zhe Zhao 0006, Qipeng Zhao, Cong Yue, Zhengliang Yang, Kyle Richardson 0001, Zhen-Zhong Lan |
COLING | 5 |
| 2020 | SiBert: Enhanced Chinese Pre-trained Language Model with Sentence InsertionabstractPre-trained models have achieved great success in learning unsupervised language representations by self-supervised tasks on large-scale corpora. Recent studies mainly focus on how to fine-tune different downstream tasks from a general pre-trained model. However, some studies show that customized self-supervised tasks for a particular type of downstream task can effectively help the pre-trained model to capture more corresponding knowledge and semantic information. Hence a new pre-training task called Sentence Insertion (SI) is proposed in this paper for Chinese query-passage pairs NLP tasks including answer span prediction, retrieval question answering and sentence level cloze test. The related experiment results indicate that the proposed SI can improve the performance of the Chinese Pre-trained models significantly. Moreover, a word segmentation method called SentencePiece is utilized to further enhance Chinese Bert performance for tasks with long texts. The complete source code is available at https://github.com/ewrfcas/SiBert_tensorflow. Chenjie Cao, Xiuyan Jiang |
LREC | 2 |
| 2020 | Entropy and Confidence-Based Undersampling Boosting Random Forests for Imbalanced ProblemsabstractIn this article, we propose a novel entropy and confidence-based undersampling boosting (ECUBoost) framework to solve imbalanced problems. The boosting-based ensemble is combined with a new undersampling method to improve the generalization performance. To avoid losing informative samples during the data preprocessing of the boosting-based ensemble, both confidence and entropy are used in ECUBoost as benchmarks to ensure the validity and structural distribution of the majority samples during the undersampling. Furthermore, different from other iterative dynamic resampling methods, ECUBoost based on confidence can be applied to algorithms without iterations such as decision trees. Meanwhile, random forests are used as base classifiers in ECUBoost. Furthermore, experimental results on both artificial data sets and KEEL data sets prove the effectiveness of the proposed method. Zhe Wang 0002, Chenjie Cao, Yujin Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Comp-GAN: Compositional Generative Adversarial Network in Synthesizing and Recognizing Facial ExpressionabstractFacial expression is important in understanding our social interaction. Thus the ability to recognize facial expression enables the novel multimedia applications. With the advance of recent deep architectures, research on facial expression recognition has achieved great progress. However, these models are still suffering from the problems of lacking sufficient and diverse high quality training faces, vulnerability to the facial variations, and recognizing a limited number of basic types of emotions. To tackle these problems, this paper proposes a novel end-to-end Compositional Generative Adversarial Network (Comp-GAN) that is able to synthesize new face images with specified poses and desired facial expressions; and such synthesized images can be further utilized to help train a robust and generalized expression recognition model. Essentially, Comp-GAN can dynamically change the expression and pose of faces according to the input images while keeping the identity information. Specifically, the generator has two major components: one for generating images with desired expression and the other for changing the pose of faces. Furthermore, a face reconstruction learning process is applied to re-generate the input image and constrains the generator for preserving the key information such as facial identity. For the first time, various one/zero-shot facial expression recognition tasks have been created. We conduct extensive experiments to show that the images generated by Comp-GAN are helpful to improve the performance of one/zero-shot facial expression recognition. Wenxuan Wang 0003, Qiang Sun 0007, Yanwei Fu 0001, Tao Chen 0003, Chenjie Cao, Ziqi Zheng, Han Qiu 0002, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ACM Multimedia | 5 |
| 2019 | Cascade interpolation learning with double subspaces and confidence disturbance for imbalanced problems
Chenjie Cao |
Neural Networks | 2 |
| 2018 | IMCStacking: Cost-sensitive stacking learning with feature inverse mapping for imbalanced problems
Chenjie Cao, Zhe Wang 0002 |
Knowl. Based Syst. | 1 |
| 2018 | Regularized fisher linear discriminant through two threshold variation strategies for imbalanced problems
Yujin Zhu, Zhe Wang 0002, Chenjie Cao, Daqi Gao |
Knowl. Based Syst. | 3 |
| 2017 | Locality sensitive discriminant matrixized learning machine
Zhe Wang 0002, Dongdong Li 0003, Yujin Zhu, Chenjie Cao |
Knowl. Based Syst. | 5 |