VLDB 2026 Research / reviewers in the wild / expert
Chengjiang Long
dblp:10/4617
· DBLP profile ↗
61ranked-venue papers
6as first author
42since 2021 · last 2026
0000-0003-1584-7290ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 53 · 5 first-author · 37 since 2021Artificial intelligence and machine learning · 38 · 5 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topology-Guided Semantic Face Center Estimation for Rotation-Invariant Face DetectionabstractFace detection accuracy significantly decreases under rotational variations, including in-plane (RIP) and out-of-plane (ROP) rotations. ROP is particularly problematic due to its impact on landmark distortion, which leads to inaccurate face center localization. Meanwhile, many existing rotation-invariant models are primarily designed to handle RIP, they often fail under ROP because they lack the ability to capture semantic and topological relationships. Moreover, existing datasets frequently suffer from unreliable landmark annotations caused by imperfect ground truth labeling, the absence of precise center annotations, and imbalanced data across different rotation angles. To address these challenges, we propose a topology-guided semantic face center estimation method that leverages graph-based landmark relationships to preserve structural integrity under both RIP and ROP. Additionally, we construct a rotation-aware face dataset with accurate face center annotations and balanced rotational diversity to support training under extreme pose conditions. Next, we introduce a Hybrid-ViT model that fuses CNN spatial features with transformer-based global context and employ a center-guided module for robust landmark localization under extreme rotations. In order to evaluate center quality, we further design a hybrid metric that combines topological geometry with semantic perception for a more comprehensive evaluation of face center accuracy. Finally, experimental results demonstrate that our method outperforms state-of-the-art models in cross-dataset evaluations. Code: https://github.com/Catster111/TCE_RIFD. Hathai Kaewkorn, Lifang Zhou, Weisheng Li 0001, Chengjiang Long |
IEEE Trans. Image Process. | 4 |
| 2025 | PUMPS: Skeleton-Agnostic Point-Based Universal Motion Pre-Training for Synthesis in Human Motion Tasks
Clinton Mo, Kun Hu 0008, Chengjiang Long, Dong Yuan 0001, Wan-Chi Siu, Zhiyong Wang 0001 |
ICCV | 3 |
| 2025 | Repurposing 2D Diffusion Models with Gaussian Atlas for 3D GenerationabstractRecent advances in text-to-image diffusion models have been driven by the increasing availability of paired 2D data. However, the development of 3D diffusion models has been hindered by the scarcity of high-quality 3D data, resulting in less competitive performance compared to their 2D counterparts. To address this challenge, we propose repurposing pre-trained 2D diffusion models for 3D object generation. We introduce Gaussian Atlas, a novel representation that utilizes dense 2D grids, enabling the fine-tuning of 2D diffusion models to generate 3D Gaussians. Our approach demonstrates successful transfer learning from a pre-trained 2D diffusion model to a 2D manifold flattened from 3D structures. To support model training, we compile GaussianVerse, a large-scale dataset comprising 205K high-quality 3D Gaussian fittings of various 3D objects. Our experimental results show that text-to-image diffusion models can be effectively adapted for 3D content generation, bridging the gap between 2D and 3D modeling. Tiange Xiang, Chengjiang Long, Christian Häne, Peihong Guo, Scott L. Delp, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001 |
ICCV | 3 |
| 2025 | PhysSplat: Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting
Hao Wang 0218, Xingyue Zhao, Hao Fei 0001, Hongqiu Wang, Chengjiang Long, Hua Zou 0002 |
ICCV | 6 |
| 2025 | DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse RenderingabstractRecent methods have shown that pre-trained diffusion models can be fine-tuned to enable generative inverse rendering by learning image-conditioned noise-to-intrinsic mapping. Despite their remarkable progress, they struggle to robustly produce high-quality results as the noise-to-intrinsic paradigm essentially utilizes noisy images with deteriorated structure and appearance for intrinsic prediction, while it is common knowledge that structure and appearance information in an image are crucial for inverse rendering. To address this issue, we present DNF-Intrinsic, a robust yet efficient inverse rendering approach fine-tuned from a pre-trained diffusion model, where we propose to take the source image rather than Gaussian noise as input to directly predict deterministic intrinsic properties via flow matching. Moreover, we design a generative renderer to constrain that the predicted intrinsic properties are physically faithful to the source image. Experiments on both synthetic and real-world datasets show that our method clearly outperforms existing state-of-the-art methods. Rongjia Zheng, Qing Zhang 0006, Chengjiang Long, Wei-Shi Zheng 0001 |
ICCV | 3 |
| 2025 | Robust Multi-Contrast MRI Medical Image Translation via Knowledge Distillation and Adversarial AttackabstractMedical image translation is of great value but is very difficult due to the requirement with style change of noise pattern and anatomy invariance of image content. Various deep learning methods like the mainstream GAN, Transformer and Diffusion models have been developed to learn the multi-modal mapping to obtain the translated images, but the results from the generator are still far from being perfect for medical images. In this paper, we propose a robust multi-contrast translation framework for MRI medical images with knowledge distillation and adversarial attack, which can be integrated with any generator. The additional refinement network consists of teacher and student modules with similar structures but different inputs. Unlike the existing knowledge distillation works, our teacher module is designed as a registration network with more inputs to better learn the noise distribution well and further refine the translated results in the training stage. The knowledge is then well distilled to the student module to ensure that better translation results are generated. We also introduce an adversarial attack module before the generator. Such a black-box attacker can generate meaningful perturbations and adversarial examples throughout the training process. Our model has been tested on two public MRI medical image datasets considering different types and levels of perturbations, and each designed module is verified by the ablation study. The extensive experiments and comparison with SOTA methods have strongly demonstrated our model's superiority of refinement and robustness. Xujie Zhao, Chengjiang Long, Jianhui Zhao 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Single-Image SVBRDF Estimation Using Auxiliary Renderings as Intermediate TargetsabstractRecently, single-image SVBRDF capture is formulated as a regression problem, which uses a network to infer four SVBRDF maps from a flash-lit image. However, the accuracy is still not satisfactory since previous approaches usually adopt end-to-end inference strategies. To mitigate the challenge, we propose "auxiliary renderings" as the intermediate regression targets, through which we divide the original end-to-end regression task into several easier sub-tasks, thus achieving better inference accuracy. Our contributions are threefold. First, we design three (or two pairs of) auxiliary renderings and summarize the motivations behind the designs. By our design, the auxiliary images are bumpiness-flattened or highlight-removed, containing disentangled visual cues about the final SVBRDF maps and can be easily transformed to the final maps. Second, to help estimate the auxiliary targets from the input image, we propose two mask images including a bumpiness mask and a highlight mask. Our method thus first infers mask images, then with the help of the mask images infers auxiliary renderings, and finally transforms the auxiliary images to SVBRDF maps. Third, we propose backbone UNets to infer mask images, and gated deformable UNets for estimating auxiliary targets. Thanks to the well-designed networks and intermediate images, our method outputs better SVBRDF maps than previous approaches, validated by the extensive comparisonal and ablation experiments. Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li, Hongmin Cai |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | CoreRec: A Counterfactual Correlation Inference for Next Set RecommendationabstractNext set recommendation aims to predict the items that are likely to be bought in the next purchase. Central to this endeavor is the task of capturing intra-set and cross-set correlations among items. However, the modeling of cross-set correlations poses challenges due to specific issues. Primarily, these correlations are often implicit, and the prevailing approach of establishing an indiscriminate link across the entire set of objects neglects factors like purchase frequency and correlations between purchased items. Such hastily formed connections across sets introduce substantial noise. Additionally, the preeminence of high-frequency items in numerous sets could potentially overshadow and distort correlation modeling with respect to low-frequency items. Thus, we devoted to mitigating misleading inter-set correlations. With a fresh perspective rooted in causality, we delve into the question of whether correlations between a particular item and items from other sets should be relied upon for item representation learning and set prediction. Technically, we introduce the Counterfactual Correlation Inference framework for next set recommendation, denoted as CoreRec. This framework establishes a counterfactual scenario in which the recommendation model impedes cross-set correlations to generate intervened predictions. By contrasting these intervened predictions with the original ones, we gauge the causal impact of inter-set neighbors on set prediction—essentially assessing whether they contribute to spurious correlations. During testing, we introduce a post-trained switch module that selects between set-aware item representations derived from either the original or the counterfactual scenarios. To validate our approach, we extensively experiment using three real-world datasets, affirming both the effectiveness of CoreRec and the cogency of our analytical approach. Chengjiang Long, Shengyu Zhang 0001, Xudong Tang, Zhichao Zhai, Kun Kuang 0001, Jun Xiao 0001 |
AAAI | 2 |
| 2024 | Motion Keyframe Interpolation for Any Human Skeleton via Temporally Consistent Point Cloud Sampling and Reconstruction
Clinton Mo, Kun Hu 0008, Chengjiang Long, Dong Yuan 0001, Zhiyong Wang 0001 |
ECCV (82) | 3 |
| 2024 | Interleaving One-Class and Weakly-Supervised Models with Adaptive Thresholding for Unsupervised Video Anomaly Detection
Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Pradipta Maji, Hongmin Cai |
ECCV (30) | 3 |
| 2024 | Multi-RoI Human Mesh Recovery with Camera Consistency and Contrastive Losses
Yongwei Nie, Changzhen Liu, Chengjiang Long, Qing Zhang 0006, Guiqing Li, Hongmin Cai |
ECCV (47) | 3 |
| 2024 | Exploring Matching Rates: From Keypoint Selection to Camera Relocalization
Chengjiang Long, Yifeng Fei, Qianchen Xia, Erwei Yin, Xin Yang 0011 |
ACM Multimedia | 2 |
| 2024 | Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh RecoveryabstractHuman Mesh Recovery (HMR) is the task of estimating a parameterized 3D human mesh from an image. There is a kind of methods first training a regression model for this problem, then further optimizing the pretrained regression model for any specific sample individually at test time. However, the pretrained model may not provide an ideal optimization starting point for the test-time optimization. Inspired by meta-learning, we incorporate the test-time optimization into training, performing a step of test-time optimization for each sample in the training batch before really conducting the training optimization over all the training samples. In this way, we obtain a meta-model, the meta-parameter of which is friendly to the test-time optimization. At test time, after several test-time optimization steps starting from the meta-parameter, we obtain much higher HMR accuracy than the test-time optimization starting from the simply pretrained regression model. Furthermore, we find test-time HMR objectives are different from training-time objectives, which reduces the effectiveness of the learning of the meta-model. To solve this problem, we propose a dual-network architecture that unifies the training-time and test-time objectives. Our method, armed with meta-learning and the dual networks, outperforms state-of-the-art regression-based and optimization-based HMR approaches, as validated by the extensive experiments. The codes are available at https://github.com/fmx789/Meta-HMR. Yongwei Nie, Mingxian Fan, Chengjiang Long, Qing Zhang 0006, Jian Zhu 0001, Xuemiao Xu |
NeurIPS | 3 |
| 2024 | CRD-CGAN: category-consistent and relativistic constraints for diverse text-to-image generation
Tao Hu 0012, Chengjiang Long, Chunxia Xiao |
Frontiers Comput. Sci. | 2 |
| 2024 | Disentangled Representation Learning for Controllable Person Image GenerationabstractIn this paper, we propose a novel framework named DRL-CPG to learn disentangled latent representation for controllable person image generation, which can produce realistic person images with desired poses and human attributes (e.g. pose, head, upper clothes, and pants) provided by various source persons. Unlike the existing works leveraging the semantic masks to obtain the representation of each component, we propose to generate disentangled latent code via a novel attribute encoder with transformers trained in a manner of curriculum learning from a relatively easy step to a gradually hard one. A random component mask-agnostic strategy is introduced to randomly remove component masks from the person segmentation masks, which aims at increasing the difficulty of training and promoting the transformer encoder to recognize the underlying boundaries between each component. This enables the model to transfer both the shape and texture of the components. Furthermore, we propose a novel attribute decoder network to integrate multi-level attributes (e.g. the structure feature and the attribute representation) with well-designed Dual Adaptive Denormalization (DAD) residual blocks. Extensive experiments strongly demonstrate that the proposed approach is able to transfer both the texture and shape of different human parts and yield realistic results. To our knowledge, we are the first to learn disentangled latent representations with transformers for person image generation. Wenju Xu, Chengjiang Long, Yongwei Nie, Guanghui Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Continuous Intermediate Token Learning with Implicit Motion Manifold for Keyframe Based Motion InterpolationabstractDeriving sophisticated 3D motions from sparse keyframes is a particularly challenging problem, due to continuity and exceptionally skeletal precision. The action features are often derivable accurately from the full series of keyframes, and thus, leveraging the global context with transformers has been a promising data-driven embedding approach. However, existing methods are often with inputs of interpolated intermediate frame for continuity using basic interpolation methods with keyframes, which result in a trivial local minimum during training. In this paper, we propose a novel framework to formulate latent motion manifolds with keyframe-based constraints, from which the continuous nature of intermediate token representations is considered. Particularly, our proposed framework consists of two stages for identifying a latent motion subspace, i.e., a keyframe encoding stage and an intermediate token generation stage, and a subsequent motion synthesis stage to extrapolate and compose motion data from manifolds. Through our extensive experiments conducted on both the LaFAN1 and CMU Mocap datasets, our proposed method demonstrates both superior interpolation accuracy and high visual similarity to ground truth motions. Clinton Mo, Kun Hu 0008, Chengjiang Long, Zhiyong Wang 0001 |
CVPR | 3 |
| 2023 | Learning Dynamic Style Kernels for Artistic Style TransferabstractArbitrary style transfer has been demonstrated to be efficient in artistic image generation. Previous methods either globally modulate the content feature ignoring local details, or overly focus on the local structure details leading to style leakage. In contrast to the literature, we propose a new scheme “style kernel” that learns spatially adaptive kernels for per-pixel stylization, where the convolutional kernels are dynamically generated from the global style-content aligned feature and then the learned kernels are applied to modulate the content feature at each spatial position. This new scheme allows flexible both global and local interactions between the content and style features such that the wanted styles can be easily transferred to the content image while at the same time the content structure can be easily preserved. To further enhance the flexibility of our style transfer method, we propose a Style Alignment Encoding (SAE) module complemented with a Content-based Gating Modulation (CGM) module for learning the dynamic style kernels in focusing regions. Extensive experiments strongly demonstrate that our proposed method outperforms state-of-the-art methods and exhibits superior performance in terms of visual quality and efficiency. Wenju Xu, Chengjiang Long, Yongwei Nie |
CVPR | 2 |
| 2023 | Feature Representation Learning with Adaptive Displacement Generation and Transformer Fusion for Micro-Expression RecognitionabstractMicro-expressions are spontaneous, rapid and subtle facial movements that can neither be forged nor suppressed. They are very important nonverbal communication clues, but are transient and of low intensity thus difficult to recognize. Recently deep learning based methods have been developed for micro-expression (ME) recognition using feature extraction and fusion techniques, however, targeted feature learning and efficient feature fusion still lack further study according to the ME characteristics. To address these issues, we propose a novel framework Feature Representation Learning with adaptive Displacement Generation and Transformer fusion (FRL-DGT), in which a convolutional Displacement Generation Module (DGM) with self-supervised learning is used to extract dynamic features from onset/apex frames targeted to the subsequent ME recognition task, and a well-designed Transformer Fusion mechanism composed of three Transformer-based fusion modules (local, global fusions based on AU regions and full-face fusion) is applied to extract the multi-level informative features after DGM for the final ME prediction. The extensive experiments with solid leave-one-subject-out (LOSO) evaluation results have demonstrated the superiority of our proposed FRL-DGT to state-of-the-art methods. Zhijun Zhai, Jianhui Zhao 0001, Chengjiang Long, Wenju Xu, Shuangjiang He, Huijuan Zhao |
CVPR | 3 |
| 2023 | Representing Multimodal Behaviors With Mean Location for Pedestrian Trajectory PredictionabstractRepresenting multimodal behaviors is a critical challenge for pedestrian trajectory prediction. Previous methods commonly represent this multimodality with multiple latent variables repeatedly sampled from a latent space, encountering difficulties in interpretable trajectory prediction. Moreover, the latent space is usually built by encoding global interaction into future trajectory, which inevitably introduces superfluous interactions and thus leads to performance reduction. To tackle these issues, we propose a novel Interpretable Multimodality Predictor (IMP) for pedestrian trajectory prediction, whose core is to represent a specific mode by its mean location. We model the distribution of mean location as a Gaussian Mixture Model (GMM) conditioned on sparse spatio-temporal features, and sample multiple mean locations from the decoupled components of GMM to encourage multimodality. Our IMP brings four-fold benefits: 1) Interpretable prediction to provide semantics about the motion behavior of a specific mode; 2) Friendly visualization to present multimodal behaviors; 3) Well theoretical feasibility to estimate the distribution of mean locations supported by the central-limit theorem; 4) Effective sparse spatio-temporal features to reduce superfluous interactions and model temporal continuity of interaction. Extensive experiments validate that our IMP not only outperforms state-of-the-art methods but also can achieve a controllable prediction by customizing the corresponding mean location. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Exploiting Residual and Illumination with GANs for Shadow Detection and Shadow RemovalabstractResidual image and illumination estimation have been proven to be helpful for image enhancement. In this article, we propose a general framework, called RI-GAN, that exploits residual and illumination using generative adversarial networks (GANs). The proposed framework detects and removes shadows in a coarse-to-fine fashion. At the coarse stage, we employ three generators to produce a coarse shadow-removal result, a residual image, and an inverse illumination map. We also incorporate two indirect shadow-removal images via the residual image and the inverse illumination map. With the residual image, the illumination map, and the two indirect shadow-removal images as auxiliary information, the refinement stage estimates a shadow mask to identify shadow regions in the image, and then refines the coarse shadow-removal result to the fine shadow-free image. We introduce a cross-encoding module to the refinement generator, in which the use of feature-crossing can provide additional details to promote the shadow mask and the high-quality shadow-removal result. In addition, we apply data augmentation to the discriminator to reduce the dependence between representations of the discriminator and the quality of the predicted image. Experiments for shadow detection and shadow removal demonstrate that our method outperforms state-of-the-art methods. Furthermore, RI-GAN exhibits good performance in terms of image dehazing, rain removal, and highlight removal, demonstrating the effectiveness and flexibility of the proposed framework. Ling Zhang 0017, Chengjiang Long, Xiaolong Zhang 0002, Chunxia Xiao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Explore Contextual Information for 3D Scene Graph Generationabstract3D scene graph generation (SGG) has been of high interest in computer vision. Although the accuracy of 3D SGG on coarse classification and single relation label has been gradually improved, the performance of existing works is still far from being perfect for fine-grained and multi-label situations. In this article, we propose a framework fully exploring contextual information for the 3D SGG task, which attempts to satisfy the requirements of fine-grained entity class, multiple relation labels, and high accuracy simultaneously. Our proposed approach is composed of a Graph Feature Extraction module and a Graph Contextual Reasoning module, achieving appropriate information-redundancy feature extraction, structured organization, and hierarchical inferring. Our approach achieves superior or competitive performance over previous methods on the 3DSSG dataset, especially on the relationship prediction sub-task. Chengjiang Long, Zhaoxuan Zhang, Bokai Liu, Qiang Zhang 0008, Xin Yang 0011 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Complementary Attention Gated Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is crucial in many practical applications due to the diversity of pedestrian movements, such as social interactions and individual motion behaviors. With similar observable trajectories and social environments, different pedestrians may make completely different future decisions. However, most existing methods only focus on the frequent modal of the trajectory and thus are difficult to generalize to the peculiar scenario, which leads to the decline of the multimodal fitting ability when facing similar scenarios. In this paper, we propose a complementary attention gated network (CAGN) for pedestrian trajectory prediction, in which a dual-path architecture including normal and inverse attention is proposed to capture both frequent and peculiar modals in spatial and temporal patterns, respectively. Specifically, a complementary block is proposed to guide normal and inverse attention, which are then be summed with learnable weights to get attention features by a gated network. Finally, multiple trajectory distributions are estimated based on the fused spatio-temporal attention features due to the multimodality of future trajectory. Experimental results on benchmark datasets, i.e., the ETH, and the UCY, demonstrate that our method outperforms state-of-the-art methods by 13.8% in Average Displacement Error (ADE) and 10.4% in Final Displacement Error (FDE). Code will be available at https://github.com/jinghaiD/CAGN Jinghai Duan, Le Wang 0003, Chengjiang Long, Sanping Zhou, Fang Zheng 0009, Liushuai Shi, Gang Hua 0001 |
AAAI | 3 |
| 2022 | CPRAL: Collaborative Panoptic-Regional Active Learning for Semantic SegmentationabstractAcquiring the most representative examples via active learning (AL) can benefit many data-dependent computer vision tasks by minimizing efforts of image-level or pixel-wise annotations. In this paper, we propose a novel Collaborative Panoptic-Regional Active Learning framework (CPRAL) to address the semantic segmentation task. For a small batch of images initially sampled with pixel-wise annotations, we employ panoptic information to initially select unlabeled samples. Considering the class imbalance in the segmentation dataset, we import a Regional Gaussian Attention module (RGA) to achieve semantics-biased selection. The subset is highlighted by vote entropy and then attended by Gaussian kernels to maximize the biased regions. We also propose a Contextual Labels Extension (CLE) to boost regional annotations with contextual attention guidance. With the collaboration of semantics-agnostic panoptic matching and region-biased selection and extension, our CPRAL can strike a balance between labeling efforts and performance and compromise the semantics distribution. We perform extensive experiments on Cityscapes and BDD10K datasets and show that CPRAL outperforms the cutting-edge methods with impressive results and less labeling proportion. Yu Qiao 0001, Jincheng Zhu, Chengjiang Long, Zeyao Zhang, Yuxin Wang 0001, Zhenjun Du, Xin Yang 0011 |
AAAI | 3 |
| 2022 | Social Interpretable Tree for Pedestrian Trajectory PredictionabstractUnderstanding the multiple socially-acceptable future behaviors is an essential task for many vision applications. In this paper, we propose a tree-based method, termed as Social Interpretable Tree (SIT), to address this multi-modal prediction task, where a hand-crafted tree is built depending on the prior information of observed trajectory to model multiple future trajectories. Specifically, a path in the tree from the root to leaf represents an individual possible future trajectory. SIT employs a coarse-to-fine optimization strategy, in which the tree is first built by high-order velocity to balance the complexity and coverage of the tree and then optimized greedily to encourage multimodality. Finally, a teacher-forcing refining operation is used to predict the final fine trajectory. Compared with prior methods which leverage implicit latent variables to represent possible future trajectories, the path in the tree can explicitly explain the rough moving behaviors (e.g., go straight and then turn right), and thus provides better interpretability. Despite the hand-crafted tree, the experimental results on ETH-UCY and Stanford Drone datasets demonstrate that our method is capable of matching or exceeding the performance of state-of-the-art methods. Interestingly, the experiments show that the raw built tree without training outperforms many prior deep neural network based approaches. Meanwhile, our method presents sufficient flexibility in long-term prediction and different best-of-K predictions. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Fang Zheng 0009, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 3 |
| 2022 | Deep Image-based Illumination HarmonizationabstractIntegrating a foreground object into a background scene with illumination harmonization is an important but challenging task in computer vision and augmented reality community. Existing methods mainly focus on foreground and background appearance consistency or the foreground object shadow generation, which rarely consider global appearance and illumination harmonization. In this paper, we formulate seamless illumination harmonization as an illumination exchange and aggregation problem. Specifically, we firstly apply a physically-based rendering method to construct a large-scale, high-quality dataset (named IH) for our task, which contains various types of foreground objects and background scenes with different lighting conditions. Then, we propose a deep image-based illumination harmonization GAN framework named DIH-GAN, which makes full use of a multi-scale attention mechanism and illumination exchange strategy to directly infer mapping relationship between the inserted foreground object and the corresponding background scene. Meanwhile, we also use adversarial learning strategy to further refine the illumination harmonization result. Our method can not only achieve harmonious appearance and illumination for the foreground object but also can generate compelling shadow cast by the foreground object. Comprehensive experiments on both our IH dataset and real-world images show that our proposed DIH-GAN provides a practical and effective solution for image-based object illumination harmonization editing, and validate the superiority of our method against state-of-the-art methods. Our IH dataset is available at https://github.com/zhongyunbao/Dataset. Zhongyun Bao, Chengjiang Long, Gang Fu 0003, Daquan Liu, Yuanzhen Li, Chunxia Xiao |
CVPR | 2 |
| 2022 | Video Shadow Detection via Spatio-Temporal Interpolation Consistency TrainingabstractIt is challenging to annotate large-scale datasets for supervised video shadow detection methods. Using a model trained on labeled images to the video frames directly may lead to high generalization error and temporal inconsistent results. In this paper, we address these challenges by proposing a Spatio-Temporal Interpolation Consistency Training (STICT) framework to rationally feed the unlabeled video frames together with the labeled images into an image shadow detection network training. Specifically, we propose the Spatial and Temporal ICT, in which we define two new interpolation schemes, i.e., the spatial interpolation and the temporal interpolation. We then derive the spatial and temporal interpolation consistency constraints accordingly for enhancing generalization in the pixel-wise classification task and for encouraging temporal consistent predictions, respectively. In addition, we design a Scale- Aware Network for multi-scale shadow knowledge learning in images, and propose a scale-consistency constraint to minimize the discrepancy among the predictions at different scales. Our proposed approach is extensively validated on the ViSha dataset and a self-annotated dataset. Experimental results show that, even without video labels, our approach is better than most state of the art supervised, semi-supervised or unsupervised image/video shadow detection methods and other methods in related tasks. Code and dataset are available at https://github.com/yihong-97/STICT. Xiao Lu 0002, Yihong Cao, Chengjiang Long, Zipei Chen, Xuanyu Zhou, Yimin Yang 0001, Chunxia Xiao |
CVPR | 4 |
| 2022 | Progressively Generating Better Initial Guesses Towards Next Stages for High-Quality Human Motion PredictionabstractThis paper presents a high-quality human motion pre-diction method that accurately predicts future human poses given observed ones. Our method is based on the observation that a good “initial guess” of the future poses is very helpful in improving the forecasting accuracy. This mo-tivates us to propose a novel two-stage prediction frame-work, including an init-prediction network that just computes the good guess and then a formal-prediction network that predicts the target future poses based on the guess. More importantly, we extend this idea further and design a multi-stage prediction framework where each stage pre-dicts initial guess for the next stage, which brings more performance gain. To fulfill the prediction task at each stage, we propose a network comprising Spatial Dense Graph Convolutional Networks (S-DGCN) and Temporal Dense Graph Convolutional Networks (T-DGCN). Alternatively executing the two networks helps extract spatiotem-poral features over the global receptive field of the whole pose sequence. All the above design choices cooperating together make our method outperform previous approaches by large margins: 6%-7% on Human3.6M, 5%-10% on CMU-MoCap, and 13%-16% on 3DPW. Code is available at https://github.com/705062791/PGBIG. Tiezheng Ma, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li |
CVPR | 3 |
| 2022 | PhraseGAN: Phrase-Boost Generative Adversarial Network for Text-to-Image GenerationabstractA phrase contains an object-orienting noun and some attribution-associating words. Therefore, focusing on phrases could better generate images with the objects and their tightly relevant characteristics. We propose a Phrase-boost Gener-ative Adversarial Network (PhraseGAN) with threefold im-provement for scene level text-to-image generation. First, we propose a Transformer-based encoder to encode the in-put words and sentences and encode related words and their targeting nouns into phrases by text correlation analysis. Sec-ond, we utilize Graph Convolution Networks to measure fine-grained text-image similarity, which could gain constraints on relative positions between different objects. Finally, we de-sign a phrase-region discriminator to discriminate the qual-ity of the generated objects and the consistency between the phrases and their corresponding objects. Experimental results on the Microsoft COCO dataset demonstrate that PhraseGAN can generate better images from texts than state-of-the-art methods. Fei Luo 0004, Chengjiang Long, Shenghong Hu, Chunxia Xiao |
ICME | 4 |
| 2022 | Diverse Human Motion Prediction via Gumbel-Softmax Sampling from an Auxiliary SpaceabstractDiverse human motion prediction aims at predicting multiple possible future pose sequences from a sequence of observed poses. Previous approaches usually employ deep generative networks to model the conditional distribution of data, and then randomly sample outcomes from the distribution. While different results can be obtained, they are usually the most likely ones which are not diverse enough. Recent work explicitly learns multiple modes of the conditional distribution via a deterministic network, which however can only cover a fixed number of modes within a limited range. In this paper, we propose a novel sampling strategy for sampling very diverse results from an imbalanced multimodal distribution learned by a deep generative model. Our method works by generating an auxiliary space and smartly making randomly sampling from the auxiliary space equivalent to the diverse sampling from the target distribution. We propose a simple yet effective network architecture that implements this novel sampling strategy, which incorporates a Gumbel-Softmax coefficient matrix sampling method and an aggressive diversity promoting hinge loss function. Extensive experiments demonstrate that our method significantly improves both the diversity and accuracy of the samplings compared with previous state-of-the-art sampling approaches. Code and pre-trained models are available at https://github.com/Droliven/diverse_sampling. Lingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li |
ACM Multimedia | 3 |
| 2022 | DDBN: Dual detection branch network for semantic diversity predictions
Qifeng Lin, Chengjiang Long, Jianhui Zhao 0001, Gang Fu 0003 |
Pattern Recognit. | 2 |
| 2022 | A Two-Stage Attentive Network for Single Image Super-ResolutionabstractRecently, deep convolutional neural networks (CNNs) have been widely explored in single image super-resolution (SISR) and contribute remarkable progress. However, most of the existing CNNs-based SISR methods do not adequately explore contextual information in the feature extraction stage and pay little attention to the final high-resolution (HR) image reconstruction step, hence hindering the desired SR performance. To address the above two issues, in this paper, we propose a two-stage attentive network (TSAN) for accurate SISR in a coarse-to-fine manner. Specifically, we design a novel multi-context attentive block (MCAB) to make the network focus on more informative contextual features. Moreover, we present an essential refined attention block (RAB) which could explore useful cues in HR space for reconstructing fine-detailed HR image. Extensive evaluations on four benchmark datasets demonstrate the efficacy of our proposed TSAN in terms of quantitative metrics and visual effects. Code is available athttps://github.com/Jee-King/TSAN. Jiqing Zhang, Chengjiang Long, Yuxin Wang 0001, Haiyin Piao, Haiyang Mei, Xin Yang 0011 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | A Hybrid Attention Mechanism for Weakly-Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization is a challenging vision task due to the absence of ground-truth temporal locations of actions in the training videos. With only video-level supervision during training, most existing methods rely on a Multiple Instance Learning (MIL) framework to predict the start and end frame of each action category in a video. However, the existing MIL-based approach has a major limitation of only capturing the most discriminative frames of an action, ignoring the full extent of an activity. Moreover, these methods cannot model background activity effectively, which plays an important role in localizing foreground activities. In this paper, we present a novel framework named HAM-Net with a hybrid attention mechanism which includes temporal soft, semi-soft and hard attentions to address these issues. Our temporal soft attention module, guided by an auxiliary background class in the classification module, models the background activity by introducing an ``action-ness'' score for each video snippet. Moreover, our temporal semi-soft and hard attention modules, calculating two attention scores for each video snippet, help to focus on the less discriminative frames of an action to capture the full action boundary. Our proposed approach outperforms recent state-of-the-art methods by at least 2.2% mAP at IoU threshold 0.5 on the THUMOS14 dataset, and by at least 1.3% mAP at IoU threshold 0.75 on the ActivityNet1.2 dataset. Ashraful Islam, Chengjiang Long, Richard J. Radke |
AAAI | 2 |
| 2021 | SGCN: Sparse Graph Convolution Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a key technology in autopilot, which remains to be very challenging due to complex interactions between pedestrians. However, previous works based on dense undirected interaction suffer from modeling superfluous interactions and neglect of trajectory motion tendency, and thus inevitably result in a considerable deviance from the reality. To cope with these issues, we present a Sparse Graph Convolution Network (SGCN) for pedestrian trajectory prediction. Specifically, the SGCN explicitly models the sparse directed interaction with a sparse directed spatial graph to capture adaptive interaction pedestrians. Meanwhile, we use a sparse directed temporal graph to model the motion tendency, thus to facilitate the prediction based on the observed direction. Finally, parameters of a bi-Gaussian distribution for trajectory prediction are estimated by fusing the above two sparse graphs. We evaluate our proposed method on the ETH and UCY datasets, and the experimental results show our method outperforms comparative state-of-the-art methods by 9% in Average Displacement Error (ADE) and 13% in Final Displacement Error (FDE). Notably, visualizations indicate that our method can capture adaptive interactions between pedestrians and their effective motion tendencies. Liushuai Shi, Le Wang 0003, Chengjiang Long, Sanping Zhou, Zhenxing Niu, Gang Hua 0001 |
CVPR | 3 |
| 2021 | CANet: A Context-Aware Network for Shadow RemovalabstractIn this paper, we propose a novel two-stage context-aware network named CANet for shadow removal, in which the contextual information from non-shadow regions is transferred to shadow regions at the embedded feature spaces. At Stage-I, we propose a contextual patch matching (CPM) module to generate a set of potential matching pairs of shadow and non-shadow patches. Combined with the potential contextual relationships between shadow and non-shadow regions, our well-designed contextual feature transfer (CFT) mechanism can transfer contextual information from non-shadow to shadow regions at different scales. With the reconstructed feature maps, we remove shadows at L and A/B channels separately. At Stage-II, we use an encoder-decoder to refine current results and generate the final shadow removal results. We evaluate our proposed CANet on two benchmark datasets and some real-world shadow images with complex scenes. Extensive experimental results strongly demonstrate the efficacy of our proposed CANet and exhibit superior performance to state-of-the-arts. Our source code is available at https://github.com/Zipei-Chen/CANet. Zipei Chen, Chengjiang Long, Ling Zhang 0017, Chunxia Xiao |
ICCV | 2 |
| 2021 | MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion PredictionabstractHuman motion prediction is a challenging task due to the stochasticity and aperiodicity of future poses. Recently, graph convolutional network has been proven to be very effective to learn dynamic relations among pose joints, which is helpful for pose prediction. On the other hand, one can abstract a human pose recursively to obtain a set of poses at multiple scales. With the increase of the abstraction level, the motion of the pose becomes more stable, which benefits pose prediction too. In this paper, we propose a novel Multi-Scale Residual Graph Convolution Network (MSR-GCN) for human pose prediction task in the manner of end-to-end. The GCNs are used to extract features from fine to coarse scale and then from coarse to fine scale. The extracted features at each scale are then combined and decoded to obtain the residuals between the input and target poses. Intermediate supervisions are imposed on all the predicted poses, which enforces the network to learn more representative features. Our proposed approach is evaluated on two standard benchmark datasets, i.e., the Human3.6M dataset and the CMU Mocap dataset. Experimental results demonstrate that our method outperforms the state-of-the-art approaches. Code and pre-trained models are available at https://github.com/Droliven/MSRGCN. Lingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li |
ICCV | 3 |
| 2021 | A Hybrid Video Anomaly Detection Framework via Memory-Augmented Flow Reconstruction and Flow-Guided Frame PredictionabstractIn this paper, we propose HF2-VAD, a Hybrid framework that integrates Flow reconstruction and Frame prediction seamlessly to handle Video Anomaly Detection. Firstly, we design the network of ML-MemAE-SC (Multi-Level Memory modules in an Autoencoder with Skip Connections) to memorize normal patterns for optical flow reconstruction so that abnormal events can be sensitively identified with larger flow reconstruction errors. More importantly, conditioned on the reconstructed flows, we then employ a Conditional Variational Autoencoder (CVAE), which captures the high correlation between video frame and optical flow, to predict the next frame given several previous frames. By CVAE, the quality of flow reconstruction essentially influences that of frame prediction. Therefore, poorly reconstructed optical flows of abnormal events further deteriorate the quality of the final predicted future frame, making the anomalies more detectable. Experimental results demonstrate the effectiveness of the proposed method. Code is available at https://github.com/LiUzHiAn/hf2vad. Zhian Liu, Yongwei Nie, Chengjiang Long, Qing Zhang 0006, Guiqing Li |
ICCV | 3 |
| 2021 | DRB-GAN: A Dynamic ResBlock Generative Adversarial Network for Artistic Style TransferabstractThe paper proposes a Dynamic ResBlock Generative Adversarial Network (DRB-GAN) for artistic style transfer. The style code is modeled as the shared parameters for Dynamic ResBlocks connecting both the style encoding network and the style transfer network. In the style encoding network, a style class-aware attention mechanism is used to attend the style feature representation for generating the style codes. In the style transfer network, multiple Dynamic ResBlocks are designed to integrate the style code and the extracted CNN semantic feature and then feed into the spatial window Layer-Instance Normalization (SW-LIN) decoder, which enables high-quality synthetic images with artistic style transfer. Moreover, the style collection conditional discriminator is designed to equip our DRB-GAN model with abilities for both arbitrary style transfer and collection style transfer during the training stage. No matter for arbitrary style transfer or collection style transfer, extensive experiments strongly demonstrate that our proposed DRB-GAN outperforms state-of-the-art methods and exhibits its superior performance in terms of visual quality and efficiency. Our source code is available at https://github.com/xuwenju123/DRB-GAN. Wenju Xu, Chengjiang Long, Ruisheng Wang 0001, Guanghui Wang 0001 |
ICCV | 2 |
| 2021 | Dual Graph Convolutional Networks with Transformer and Curriculum Learning for Image CaptioningabstractExisting image captioning methods just focus on understanding the relationship between objects or instances in a single image, without exploring the contextual correlation existed among contextual image. In this paper, we propose Dual Graph Convolutional Networks (Dual-GCN) with transformer and curriculum learning for image captioning. In particular, we not only use an object-level GCN to capture the object to object spatial relation within a single image, but also adopt an image-level GCN to capture the feature information provided by similar images. With the well-designed Dual-GCN, we can make the linguistic transformer better understand the relationship between different objects in a single image and make full use of similar images as auxiliary information to generate a reasonable caption description for a single image. Meanwhile, with a cross-review strategy introduced to determine difficulty levels, we adopt curriculum learning as the training strategy to increase the robustness and generalization of our proposed model. We conduct extensive experiments on the large-scale MS COCO dataset, and the experimental results powerfully demonstrate that our proposed method outperforms recent state-of-the-art approaches. It achieves a BLEU-1 score of 82.2 and a BLEU-2 score of 67.6. Our source code is available at https://github.com/Unbear430/DGCN-for-image-captioning. Xinzhi Dong, Chengjiang Long, Wenju Xu, Chunxia Xiao |
ACM Multimedia | 2 |
| 2021 | Luminance Attentive Networks for HDR Image and Panorama ReconstructionabstractAbstract It is very challenging to reconstruct a high dynamic range (HDR) from a low dynamic range (LDR) image as an ill‐posed problem. This paper proposes a luminance attentive network named LANet for HDR reconstruction from a single LDR image. Our method is based on two fundamental observations: (1) HDR images stored in relative luminance are scale‐invariant, which means the HDR images will hold the same information when multiplied by any positive real number. Based on this observation, we propose a novel normalization method called “HDR calibration“for HDR images stored in relative luminance, calibrating HDR images into a similar luminance scale according to the LDR images. (2) The main difference between HDR images and LDR images is in under‐/over‐exposed areas, especially those highlighted. Following this observation, we propose a luminance attention module with a two‐stream structure for LANet to pay more attention to the under‐/over‐exposed areas. In addition, we propose an extended network called panoLANet for HDR panorama reconstruction from an LDR panorama and build a dualnet structure for panoLANet to solve the distortion problem caused by the equirectangular panorama. Extensive experiments show that our proposed approach LANet can reconstruct visually convincing HDR images and demonstrate its superiority over state‐of‐the‐art approaches in terms of all metrics in inverse tone mapping. The image‐based lighting application with our proposed panoLANet also demonstrates that our method can simulate natural scene lighting using only LDR panorama. Our source code is available at https://github.com/LWT3437/LANet . Hanning Yu, Chengjiang Long, Bo Dong 0004, Qin Zou 0001, Chunxia Xiao |
Comput. Graph. Forum | 3 |
| 2021 | A Novel Visual Representation on Text Using Diverse Conditional GAN for Visual RecognitionabstractAutomatic image visual recognition can make full use of largely available images with text descriptions on social media platforms to build large-scale image labeled datasets. In this paper, we propose a novel visual text representation, named DG-VRT (Diverse GAN-Visual Representation on Text), which extracts visual features from synthetic images generated by a diverse conditional Generative Adversarial Network (DCGAN) on the text, for visual recognition. The DCGAN incorporates the current state-of-the-art text-to-image GANs and generates multiple synthetic images with various prior noises conditioned on a text. Then we extract deep visual features from the generated synthetic images to explore the underlying visual concepts and provide a visual transformation on text in feature space. Finally, we combine image-level visual features, text-level features and visual features based on synthetic images together to recognize the images, and we also extend the proposed work to semantic segmentation. We conduct extensive experiments on two benchmark datasets and the experimental results demonstrate the efficacy of our proposed representation on text for visual recognition. Tao Hu 0012, Chengjiang Long, Chunxia Xiao |
IEEE Trans. Image Process. | 2 |
| 2021 | Explore Video Clip Order With Self-Supervised and Curriculum Learning for Video ApplicationsabstractWe present a self-supervised spatiotemporal learning approach by exploring the temporal coherence of videos. The chronological order of shuffled clips from the video is used as the supervisory signal to guide the 3D Convolutional Neural Networks (CNNs) to learn meaningful visual knowledge. Unlike the existing approaches which use frames, we utilize dynamic video clips to reduce the uncertainty of order. We test three types of representative 3D CNNs, all of which benefit from the proposed approach. The learned 3D CNNs can be used either as a feature extractor or a pre-trained model for further fine-tuning on downstream tasks. We also propose two curriculum learning strategies to make the 3D CNNs easier to train and get the state-of-the-art results in nearest neighbor retrieval and action recognition tasks compared with other self-supervised learning methods. Meanwhile, it is further extended to the field of visual question answering application and has achieved promising results. Besides, comprehensive and extensive experimental results and analyses are provided for readers to better understand the video clip order we explore with self-supervised and curriculum learning for video application. Jun Xiao 0001, Lin Li 0065, Dejing Xu, Chengjiang Long, Jian Shao 0001, Shiliang Pu, Yueting Zhuang |
IEEE Trans. Multim. | 4 |
| 2021 | Monte Carlo denoising via auxiliary feature guided self-attentionabstractWhile self-attention has been successfully applied in a variety of natural language processing and computer vision tasks, its application in Monte Carlo (MC) image denoising has not yet been well explored. This paper presents a self-attention based MC denoising deep learning network based on the fact that self-attention is essentially non-local means filtering in the embedding space which makes it inherently very suitable for the denoising task. Particularly, we modify the standard self-attention mechanism to an auxiliary feature guided self-attention that considers the by-products (e.g., auxiliary feature buffers) of the MC rendering process. As a critical prerequisite to fully exploit the performance of self-attention, we design a multi-scale feature extraction stage, which provides a rich set of raw features for the later self-attention module. As self-attention poses a high computational complexity, we describe several ways that accelerate it. Ablation experiments validate the necessity and effectiveness of the above design choices. Comparison experiments show that the proposed self-attention based MC denoising method outperforms the current state-of-the-art methods. Yongwei Nie, Chengjiang Long, Wenjun Xu 0002, Qing Zhang 0006, Guiqing Li |
ACM Trans. Graph. | 3 |
| 2020 | RIS-GAN: Explore Residual and Illumination with Generative Adversarial Networks for Shadow RemovalabstractResidual images and illumination estimation have been proved very helpful in image enhancement. In this paper, we propose a general and novel framework RIS-GAN which explores residual and illumination with Generative Adversarial Networks for shadow removal. Combined with the coarse shadow-removal image, the estimated negative residual images and inverse illumination maps can be used to generate indirect shadow-removal images to refine the coarse shadow-removal result to the fine shadow-free image in a coarse-to-fine fashion. Three discriminators are designed to distinguish whether the predicted negative residual images, shadow-removal images, and the inverse illumination maps are real or fake jointly compared with the corresponding ground-truth information. To our best knowledge, we are the first one to explore residual and illumination for shadow removal. We evaluate our proposed method on two benchmark datasets, i.e., SRD and ISTD, and the extensive experiments demonstrate that our proposed method achieves the superior performance to state-of-the-arts, although we have no particular shadow-aware components designed in our generators. Ling Zhang 0017, Chengjiang Long, Xiaolong Zhang 0002, Chunxia Xiao |
AAAI | 2 |
| 2020 | DOA-GAN: Dual-Order Attentive Generative Adversarial Network for Image Copy-Move Forgery Detection and LocalizationabstractImages can be manipulated for nefarious purposes to hide content or to duplicate certain objects through copy-move operations. Discovering a well-crafted copy-move forgery in images can be very challenging for both humans and machines; for example, an object on a uniform background can be replaced by an image patch of the same background. In this paper, we propose a Generative Adversarial Network with a dual-order attention model to detect and localize copy-move forgeries. In the generator, the first-order attention is designed to capture copy-move location information, and the second-order attention exploits more discriminative features for the patch co-occurrence. Both attention maps are extracted from the affinity matrix and are used to fuse location-aware and co-occurrence features for the final detection and localization branches of the network. The discriminator network is designed to further ensure more accurate localization results. To the best of our knowledge, we are the first to propose such a network architecture with the 1st-order attention mechanism from the affinity matrix. We have performed extensive experimental validation and our state-of-the-art results strongly demonstrate the efficacy of the proposed approach. Ashraful Islam, Chengjiang Long, Arslan Basharat, Anthony Hoogs |
CVPR | 2 |
| 2020 | ARShadowGAN: Shadow Generative Adversarial Network for Augmented Reality in Single Light ScenesabstractGenerating virtual object shadows consistent with the real-world environment shading effects is important but challenging in computer vision and augmented reality applications. To address this problem, we propose an end-to-end Generative Adversarial Network for shadow generation named ARShadowGAN for augmented reality in single light scenes. Our ARShadowGAN makes full use of attention mechanism and is able to directly model the mapping relation between the virtual object shadow and the real-world environment without any explicit estimation of the illumination and 3D geometric information. In addition, we collect an image set which provides rich clues for shadow generation and construct a dataset for training and evaluating our proposed ARShadowGAN. The extensive experimental results show that our proposed ARShadowGAN is capable of directly generating plausible virtual object shadows in single light scenes. Our source code is available at https://github.com/ldq9526/ARShadowGAN. Daquan Liu, Chengjiang Long, Hongpan Zhang, Hanning Yu, Xinzhi Dong, Chunxia Xiao |
CVPR | 2 |
| 2020 | Multi-Context And Enhanced Reconstruction Network For Single Image Super ResolutionabstractMost existing single image super-resolution (SISR) methods continually increase the depth or width of networks, without adequately exploring contextual features which are essential for reconstruction. Moreover, such existing methods pay little attention to the final high-resolution(HR) image reconstruction step and therefore hinder the desired SR performance. In this paper, we propose a multi-context and enhanced reconstruction network (MCERN) for SISR. Specifically, a novel model named Multi-Context Block (MCB) which extracts more image contextual features with multibranch dilated convolution. Applying multiple MCBs with residual and dense connections, we can effectively extract contextual and hierarchical features for obtaining the coarse super-resolution result. Then an enhanced reconstruction block (ERB) is followed to extract essential spatial features on the high-resolution image to refine the coarse result to a better result. Extensive benchmark evaluations demonstrate the efficacy of our proposed MCERN in terms of metric accuracy and visual effects. Jiqing Zhang, Chengjiang Long, Yuxin Wang 0001, Xin Yang 0011, Haiyang Mei |
ICME | 2 |
| 2020 | Iterative and Adaptive Sampling with Spatial Attention for Black-Box Model ExplanationsabstractDeep neural networks have achieved great success in many real-world applications, yet it remains unclear and difficult to explain their decision-making process to an enduser. In this paper, we address the explainable AI problem for deep neural networks with our proposed framework, named IASSA, which generates an importance map indicating how salient each pixel is for the models prediction with an iterative and adaptive sampling module. We employ an affinity matrix calculated on multi-level deep learning features to explore long-range pixel-to-pixel correlation, which can shift the saliency values guided by our long-range and parameter-free spatial attention module. Extensive experiments on the MS-COCO dataset show that the proposed approach matches or exceeds the performance of state-of-the-art black-box explanation methods. Our source code is available at https://github.com/vbhavank/IASSA-Saliency. Bhavan Vasu, Chengjiang Long |
WACV | 2 |
| 2020 | Multi-stage point completion network with critical set supervision
Chengjiang Long, Qingan Yan, Alix L. H. Chow, Chunxia Xiao |
Comput. Aided Geom. Des. | 2 |
| 2020 | CLA-GAN: A Context and Lightness Aware Generative Adversarial Network for Shadow RemovalabstractAbstract In this paper, we propose a novel context and lightness aware Generative Adversarial Network (CLA‐GAN) framework for shadow removal, which refines a coarse result to a final shadow removal result in a coarse‐to‐fine fashion. At the refinement stage, we first obtain a lightness map using an encoder‐decoder structure. With the lightness map and the coarse result as the inputs, the following encoder‐decoder tries to refine the final result. Specifically, different from current methods restricted pixel‐based features from shadow images, we embed a context‐aware module into the refinement stage, which exploits patch‐based features. The embedded module transfers features from non‐shadow regions to shadow regions to ensure the consistency in appearance in the recovered shadow‐free images. Since we consider pathces, the module can additionally enhance the spatial association and continuity around neighboring pixels. To make the model pay more attention to shadow regions during training, we use dynamic weights in the loss function. Moreover, we augment the inputs of the discriminator by rotating images in different degrees and use rotation adversarial loss during training, which can make the discriminator more stable and robust. Extensive experiments demonstrate the validity of the components in our CLA‐GAN framework. Quantitative evaluation on different shadow datasets clearly shows the advantages of our CLA‐GAN over the state‐of‐the‐art methods. Ling Zhang 0017, Chengjiang Long, Qingan Yan, Xiaolong Zhang 0002, Chunxia Xiao |
Comput. Graph. Forum | 2 |
| 2020 | Shading-aware shadow detection and removal from a single image
Xinyun Fan, Ling Zhang 0017, Qingan Yan, Gang Fu 0003, Zipei Chen, Chengjiang Long, Chunxia Xiao |
Vis. Comput. | 7 |
| 2019 | ARGAN: Attentive Recurrent Generative Adversarial Network for Shadow Detection and RemovalabstractIn this paper we propose an attentive recurrent generative adversarial network (ARGAN) to detect and remove shadows in an image. The generator consists of multiple progressive steps. At each step a shadow attention detector is firstly exploited to generate an attention map which specifies shadow regions in the input image. Given the attention map, a negative residual by a shadow remover encoder will recover a shadow-lighter or even a shadow-free image. The discriminator is designed to classify whether the output image in the last progressive step is real or fake. Moreover, ARGAN is suitable to be trained with a semi-supervised strategy to make full use of sufficient unsupervised data. The experiments on four public datasets have demonstrated that our ARGAN is robust to detect both simple and complex shadows and to produce more realistic shadow removal results. It outperforms the state-of-the-art methods, especially in detail of recovering shadow areas. Bin Ding, Chengjiang Long, Ling Zhang 0017, Chunxia Xiao |
ICCV | 2 |
| 2019 | Deep Neural Networks in Fully Connected CRF for Image Labeling with Social Network MetadataabstractWe propose a novel method for predicting image labels by fusing image content descriptors with the social media context of each image. An image uploaded to a social media site such as Flickr often has meaningful, associated information, such as comments and other images the user has uploaded, that is complementary to pixel content and helpful in predicting labels. Prediction challenges such as ImageNet [6]and MSCOCO [19] use only pixels, while other methods make predictions purely from social media context [21]. Our method is based on a novel fully connected Conditional Random Field (CRF) framework, where each node is an image, and consists of two deep Convolutional Neural Networks (CNN) and one Recurrent Neural Network (RNN) that model both textual and visual node/image information. The edge weights of the CRF graph represent textual similarity and link-based metadata such as user sets and image groups. We model the CRF as an RNN for both learning and inference, and incorporate the weighted ranking loss and cross entropy loss into the CRF parameter optimization to handle the training data imbalance issue. Our proposed approach is evaluated on the MIR-9K dataset and experimentally outperforms current state-of-the-art approaches. Chengjiang Long, Roddy Collins, Eran Swears, Anthony Hoogs |
WACV | 1 |
| 2019 | Shadow Inpainting and Removal Using Generative Adversarial Networks with Slice ConvolutionsabstractAbstract In this paper, we propose a two‐stage top‐down and bottom‐up Generative Adversarial Networks (TBGANs) for shadow inpainting and removal which uses a novel top‐down encoder and a bottom‐up decoder with slice convolutions. These slice convolutions can effectively extract and restore the long‐range spatial information for either down‐sampling or up‐sampling. Different from the previous shadow removal methods based on deep learning, we propose to inpaint shadow to handle the possible dark shadows to achieve a coarse shadow‐removal image at the first stage, and then further recover the details and enhance the color and texture details with a non‐local block to explore both local and global inter‐dependencies of pixels at the second stage. With such a two‐stage coarse‐to‐fine processing, the overall effect of shadow removal is greatly improved, and the effect of color retention in non‐shaded areas is significant. By comparing with a variety of mainstream shadow removal methods, we demonstrate that our proposed method outperforms the state‐of‐the‐art methods. Jinjiang Wei, Chengjiang Long, Hua Zou 0002, Chunxia Xiao |
Comput. Graph. Forum | 2 |
| 2018 | Collaborative Active Visual Recognition from Crowds: A Distributed Ensemble ApproachabstractActive learning is an effective way of engaging users to interactively train models for visual recognition more efficiently. The vast majority of previous works focused on active learning with a single human oracle. The problem of active learning with multiple oracles in a collaborative setting has not been well explored. We present a collaborative computational model for active learning with multiple human oracles, the input from whom may possess different levels of noises. It leads to not only an ensemble kernel machine that is robust to label noises, but also a principled label quality measure to online detect irresponsible labelers. Instead of running independent active learning processes for each individual human oracle, our model captures the inherent correlations among the labelers through shared data among them. Our experiments with both simulated and real crowd-sourced noisy labels demonstrate the efficacy of our model. Gang Hua 0001, Chengjiang Long, Ming Yang 0007, Yan Gao 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Correlational Gaussian Processes for Cross-Domain Visual RecognitionabstractWe present a probabilistic model that captures higher order co-occurrence statistics for joint visual recognition in a collection of images and across multiple domains. More importantly, we predict the structured output across multiple domains by correlating outputs from the multi-classes Gaussian process classifiers in each individual domain. A set of correlational tensors is adopted to model the relationship within a single domain as well as across multiple domains. This renders it possible to explore a high-order relational model instead of using just a set of pairwise relational models. Such tensor relations are based on both the positive and negative co-occurrences of different categories of visual instances across multi-domains. This is in contrast to most previous models where only pair-wise relationships are explored. We conduct experiments on four challenging image collections. The experimental results clearly demonstrate the efficacy of our proposed model. Chengjiang Long, Gang Hua 0001 |
CVPR | 1 |
| 2017 | How Does a Camera Look at One 3D CAD Object?abstractCamera pose and the camera’s rotation angles and translation vector (RT), are one-to-one relation with a 2D real image when the intrinsic parameter is fixed. In this paper, we propose a novel convolutional neural network (CNN) based framework to intelligently estimate the 6-DOF RTs from images taken on one 3D CAD object directly and indirectly, as well as visually verifying the correctness of the predicted RTs. Such a framework enables us to accurately interpret how a camera looks at the object. The direct way is simple and obtains lower average errors for the predicted RTs experimentally, while the indirect way utilizes the POSIT algorithm via landmarks and is able to avoid the non-Euclidean issue in rotation angles. To our best knowledge, we are the first one to estimate camera’s RTs and effectively interprets how a camera looks at one 3D CAD object from the images taken on it. The experiments on four models quantitatively and qualitatively demonstrate the efficacy of our proposed approach. Chuang Xing, Chengjiang Long, Hao Guo 0005, Yongwei Nie, Dehai Zhu, Qin Ma 0001, Mengxiao Tian |
ICTAI | 2 |
| 2016 | A Joint Gaussian Process Model for Active Visual Recognition with Expertise Estimation in Crowdsourcing
Chengjiang Long, Gang Hua 0001, Ashish Kapoor |
Int. J. Comput. Vis. | 1 |
| 2015 | Multi-class Multi-annotator Active Learning with Robust Gaussian Process for Visual RecognitionabstractActive learning is an effective way to relieve the tedious work of manual annotation in many applications of visual recognition. However, less research attention has been focused on multi-class active learning. In this paper, we propose a novel Gaussian process classifier model with multiple annotators for multi-class visual recognition. Expectation propagation (EP) is adopted for efficient approximate Bayesian inference of our probabilistic model for classification. Based on the EP approximation inference, a generalized Expectation Maximization (GEM) algorithm is derived to estimate both the parameters for instances and the quality of each individual annotator. Also, we incorporate the idea of reinforcement learning to actively select both the informative samples and the high-quality annotators, which better explores the trade-off between exploitation and exploration. The experiments clearly demonstrate the efficacy of the proposed model. Chengjiang Long, Gang Hua 0001 |
ICCV | 1 |
| 2014 | Accurate Object Detection with Location Relaxation and Regionlets Re-localization
Chengjiang Long, Xiaoyu Wang 0002, Gang Hua 0001, Ming Yang 0007, Yuanqing Lin |
ACCV (1) | 1 |
| 2013 | Collaborative Active Learning of a Kernel Machine Ensemble for RecognitionabstractActive learning is an effective way of engaging users to interactively train models for visual recognition. The vast majority of previous works, if not all of them, focused on active learning with a single human oracle. The problem of active learning with multiple oracles in a collaborative setting has not been well explored. Moreover, most of the previous works assume that the labels provided by the human oracles are noise free, which may often be violated in reality. We present a collaborative computational model for active learning with multiple human oracles. It leads to not only an ensemble kernel machine that is robust to label noises, but also a principled label quality measure to online detect irresponsible labelers. Instead of running independent active learning processes for each individual human oracle, our model captures the inherent correlations among the labelers through shared data among them. Our simulation experiments and experiments with real crowd-sourced noisy labels demonstrated the efficacy of our model. Gang Hua 0001, Chengjiang Long, Ming Yang 0007, Yan Gao 0003 |
ICCV | 2 |
| 2013 | Active Visual Recognition with Expertise Estimation in CrowdsourcingabstractWe present a noise resilient probabilistic model for active learning of a Gaussian process classifier from crowds, i.e., a set of noisy labelers. It explicitly models both the overall label noises and the expertise level of each individual labeler in two levels of flip models. Expectation propagation is adopted for efficient approximate Bayesian inference of our probabilistic model for classification, based on which, a generalized EM algorithm is derived to estimate both the global label noise and the expertise of each individual labeler. The probabilistic nature of our model immediately allows the adoption of the prediction entropy and estimated expertise for active selection of data sample to be labeled, and active selection of high quality labelers to label the data, respectively. We apply the proposed model for three visual recognition tasks, i.e., object category recognition, gender recognition, and multi-modal activity recognition, on three datasets with real crowd-sourced labels from Amazon Mechanical Turk. The experiments clearly demonstrated the efficacy of the proposed model. Chengjiang Long, Gang Hua 0001, Ashish Kapoor |
ICCV | 1 |