EDBT 2026 Demo / reviewers in the wild / expert
Yuanzhouhan Cao
dblp:160/8425
· DBLP profile ↗
25ranked-venue papers
6as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gain from prior: Incremental few-shot instance segmentation via knowledge-retention transfer and negative proposal calibration
Weixiang Gao, Caijuan Shi, Yuanzhouhan Cao, Ao Cai |
Comput. Vis. Image Underst. | 3 |
| 2026 | CoDS: Enhancing Collaborative Perception in Heterogeneous Scenarios via Domain Separation
Yushan Han, Hui Zhang 0091, Honglei Zhang 0002, Chuntao Ding, Yuanzhouhan Cao, Yidong Li |
IEEE Trans. Mob. Comput. | 5 |
| 2026 | Investigate Interactive Semantic Segmentation via an Uncertainty Mining ViewabstractWith the rapid development of intelligence media, traditional semantic segmentation has shown excellent potential in application scenarios like autonomous driving. However, due to limited performance, traditional segmentation models usually lead to poor user experiences in applications that require high segmentation precision. Therefore, interactive semantic segmentation (ISS) is gaining the attention is gaining attention due to its capability to generate high-precision semantic segmentation results through a few user-provided clicks for experience improvement, which thus has a promising development prospect in fine-grained application scenarios,e.g., virtual reality, smart medical, data annotation,etc.. For good interaction efficiency, most existing interactive methods make efforts to conduct suitable click simulation strategies and reasonable click encoding methods, aiming at the robust understanding of diverse user clicks and translating comprehensible user intent,i.e., assign the correct category to the clicked area, for the neural network. Though proved effective, their designs ignore the uncertainty hiding in the extracted interaction features, which reflects the interaction difficulty and the user clicking intents. This can lead to inappropriate click simulation and click encoding, limiting the interaction efficiency. Hence we focus on exploring a reasonable ISS scheme via an uncertainty mining view. Specifically, we propose an uncertainty-based class-balanced click sampling (UCCS) simulation strategy by considering both the uncertainty of the click simulation region and its semantic imbalance, to form a reasonable click distribution. Furthermore, we propose a semantic uncertainty residual encoding (SURE) method to better embed the user's intention into the localization maps, by mining semantic confusion between the click and misprediction classes. We prove the effectiveness of our design through extensive experiments and initially analyze the importance of uncertainty mining for the ISS. Our model can achieve state-of-the-art performance on three semantic segmentation benchmarks. Yutong Gao 0001, Congyan Lang, Fayao Liu, Xun Xu 0002, Yuanzhouhan Cao, Yunchao Wei |
IEEE Trans. Multim. | 5 |
| 2025 | Channel Consistency Prior and Self-Reconstruction Strategy Based Unsupervised Image DerainingabstractRecently, deep image deraining models based on paired datasets have made a series of remarkable progress. However, they cannot be well applied in real-world applications due to the difficulty of obtaining real paired datasets and the poor generalization performance. In this paper, we propose a novel Channel Consistency Prior and Self-Reconstruction Strategy Based Unsupervised Image Deraining framework, CSUD, to tackle the aforementioned challenges. During training with unpaired data, CSUD is capable of generating high-quality pseudo clean and rainy image pairs which are used to enhance the performance of deraining network. Specifically, to preserve more image background details while transferring rain streaks from rainy images to the unpaired clean images, we propose a novel Channel Consistency Loss (CCLoss) by introducing the Channel Consistency Prior (CCP) of rain streaks into training process, thereby ensuring that the generated pseudo rainy images closely resemble the real ones. Furthermore, we propose a novel Self-Reconstruction (SR) strategy to alleviate the redundant information transfer problem of the generator, further improving the deraining performance and the generalization capability of our method. Extensive experiments on multiple synthetic and real-world datasets demonstrate that the deraining performance of CSUD surpasses other state-of-the-art unsupervised methods and CSUD exhibits superior generalization capability. Code is available at https://github.com/GuangluDong0728/CSUD. Guanglu Dong, Tianheng Zheng, Yuanzhouhan Cao, Linbo Qing, Chao Ren 0002 |
CVPR | 3 |
| 2025 | MSDet: Receptive Field Enhanced Multiscale Detection for Tiny Pulmonary NoduleabstractPulmonary nodules are critical for early lung cancer diagnosis, but traditional CT imaging methods suffer from low detection rates and poor localization. Small nodule detection is challenging due to subtle differences in density and issues like occlusion. Existing methods such as FPN, with its fixed feature fusion and limited receptive field, struggle to effectively overcome these issues. To address these challenges, our paper proposed three key contributions: Firstly, we proposed MSDet, a multiscale attention and receptive field network for detecting tiny pulmonary nodules. Secondly, we proposed the extended receptive domain (ERD) strategy to capture richer contextual information and reduce false positives caused by nodule occlusion. We also proposed the position channel attention mechanism (PCAM) to optimize feature learning and reduce multiscale detection errors, and designed the tiny object detection block (TODB) to enhance the detection of tiny nodules. Experiments on the LUNA16 dataset show an 8.8% improvement in mAP over YOLOv8, achieving state-of-the-art performance. The code is available at https://github.com/CaiGuoHui123/MSDet. Guohui Cai, Ruicheng Zhang, Hongyang He, Zeyu Zhang 0006, Daji Ergu, Yuanzhouhan Cao, Jinman Zhao, Binbin Hu, Zhibin Liao, Yang Zhao 0019, Ying Cai 0002 |
ICME | 6 |
| 2025 | GMML: Gradient-Modulated Robustness for Imbalance-Aware Multimodal LearningabstractMultimodal learning integrates diverse modalities to enhance robustness, yet real-world scenarios suffer from heterogeneous imbalance phenomena (noise interference, modality partial missing, intermodal information disparities), degrading performance through biased feature representations. Existing methods fail to adaptively modulate models under dynamic imbalance conditions. We propose GMML, a framework dynamically balancing multimodal gradients to counteract imbalance-induced biases: i) An imbalance-aware gradient modulation adaptively identifies contributions with smooth weight transitions to balance conflicting gradients; ii) A parameter constraint method enforces ℓ2-norm constraints on encoders, suppressing parameter oscillations and blocking noisy updates under modality missing/noise. Theoretically, GMML achieves a larger certified radius upper bound for complex imbalances, with convergence radius analysis providing theoretical guarantees. Experiments demonstrate superior robustness against three imbalance types, outperforming state-of-the-art by 3.3% and 2.3% in accuracy on KS and UCF-101 benchmarks. series Code: https://github.com/zhangzikai-security-ML/GMML. Zikai Zhang 0004, Xu Zhang 0085, Yidong Li, Yuanzhouhan Cao |
ACM Multimedia | 5 |
| 2025 | Medical artificial intelligence for early detection of lung cancer: A survey
Guohui Cai, Ying Cai 0002, Zeyu Zhang 0006, Yuanzhouhan Cao, Daji Ergu, Zhibin Liao, Yang Zhao 0019 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Ada3DLane: Adaptive 3D Lane Detection From Monocular Imagesabstract3D lane detection is a fundamental task in autonomous driving, and detecting 3D lanes from monocular images has been widely adopted due to the low computational cost and the property of lanes. Recent progresses have been made based on surrogate representations such as bird’s eye view (BEV) features. However, monocular BEV construction strictly relies on flat groud assumption, and the misalignment between perspective view and BEV is inevitable. In this paper, we propose Ada3DLane, a BEV-free, query based 3D lane detector, which adaptively generates queries with rich semantic and geometric information as well as lane interactions; adaptively samples perspective view image features in spatial and temporal domain; adaptively decode the sampled features for fast and accurate 3D lane detection. We conduct extensive experiments on benchmark lane detection datasets and outperforms previous state-of-the-art methods. Zhiming Hou, Yuanzhouhan Cao, Naiyue Chen, Chao Ren 0002, Chunyu Lin, Congyan Lang, Yidong Li |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Mining Semantic Correlations Between Mispredictions and Corrections for Interactive Semantic SegmentationabstractInteractive semantic segmentation pursues high-quality segmentation results at the cost of a small number of user clicks. It is attracting more and more research attention for its convenience in labeling semantic pixel-level data. Existing interactive segmentation methods often pursue higher interaction efficiency by mining the latent information of user clicks or exploring efficient interaction manners. However, these works neglect to explicitly exploit the semantic correlations between user corrections and model mispredictions, thus suffering from two flaws. First, similar prediction errors frequently occur in actual use, causing users to repeatedly correct them. Second, the interaction difficulty of different semantic classes varies across images, but existing models use monotonic parameters for all images which lack semantic pertinence. Therefore, in this article, we explore the semantic correlations existing in corrections and mispredictions by proposing a simple yet effective online learning solution to the above problems, named correction-misprediction correlation mining (CM2). Specifically, we leverage the correction-misprediction similarities to design a confusion memory module (CMM) for automatic correction when similar prediction errors reappear. Furthermore, we measure the semantic interaction difficulty by counting the correction-misprediction pairs and design a challenge adaptive convolutional layer (CACL), which can adaptively switch different parameters according to interaction difficulties to better segment the challenging classes. Our method requires no extra training besides the online learning process and can effectively improve interaction efficiency. Our proposed CM2 achieves state-of-the-art results on three public semantic segmentation benchmarks. Yutong Gao 0001, Congyan Lang, Fayao Liu, Chuan-Sheng Foo, Yuanzhouhan Cao, Yunchao Wei |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Scale-Disentangled and Uncertainty-Guided Alignment for Domain-Adaptive Object DetectionabstractUnsupervised domain adaptive object detection methods aim to transfer knowledge from the label-sufficient domain to the unlabeled domain. Most existing works minimize domain disparity by concentrating on different levels through adversarial learning. However, adversarial learning do not consider the different influences on under-aligned and well-aligned samples as they merely match distinct distributions with consistent weight. To address this issue, we design a novel scale-disentangled and uncertainty-guided alignment for domain-adaptive object detection (SDUGA), consisting of three main components: (1) Disentangled scale coarse module, which decouples scale information from global image features and performs individual alignment across domains for the corresponding scale by training domain classifiers in an adversarial learning manner; (2) Disentangled scale fine module, which generalizes the disentangled scale alignment to instance-level adaptation, reinforcing the distribution alignment across domains from multi-scale local instance level; (3) Uncertainty-guided coarse-to-fine attention alignment, which adjusts weights for various samples adaptively by generating the uncertainty-guided attention map, thus enforcing the detector to converge more on alignment for under-aligned samples and avoid misaligning well-aligned ones. Extensive experiments over three challenging domain-shift object detection scenarios demonstrate that SDUGA gains superior performance compared to state-of-the-art methods. Hui Zhang 0091, Guiyang Luo, Yuanzhouhan Cao, Xiao Wang 0002, Yidong Li, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Dynamic Interaction Dilation for Interactive Human ParsingabstractInteractive segmentation pursues generating high-quality pixel-level predictions with a few user-provided clicks, which is gaining attention for its convenience in segmentation data annotation. Users are allowed to iteratively refine the prediction by adding clicks until the result is satisfactory. Existing interactive methods usually transform the clicks into a set of localization maps by Euclidian distance computation or RGB texture extraction to guide the segmentation, which makes the click transformation a core module in interactive segmentation networks. However, when adopted in human images where large poses, occlusions, and bad illuminations are prevailing, prior transformation methods tend to cause uncorrectable overlapping across localization maps which are difficult to form a good match among human parts. Furthermore, the inappropriately transformed information is hard to be refined with the static transformation manner which is out of tune with the dynamically refined interaction process. Hence, we design a dynamic transformation scheme for interactive human parsing (IHP) named Dynamic Interaction Dilation Net (DID-Net), which serves as an initial attempt to break the limitations of static transformation while capturing long-range dependencies of clicks within each human part. Specifically, we construct a Dynamic Dilation Module (DD-Module) to dilate clicks radially in several directions assisted by human body edge detection to refine the dilation quality in each interaction iteration. Furthermore, we propose an Adaptive Interaction Excitation Block (AIE-Block) to exploit potential semantic clues buried in the dilated clicks. Our DID-Net achieves state-of-the-art performance on 3 public human parsing benchmarks. Yutong Gao 0001, Congyan Lang, Fayao Liu, Yuanzhouhan Cao, Yunchao Wei |
IEEE Trans. Multim. | 4 |
| 2023 | Random Sub-Samples Generation for Self-Supervised Real Image DenoisingabstractWith sufficient paired training samples, the supervised deep learning methods have attracted much attention in image denoising because of their superior performance. However, it is still very challenging to widely utilize the supervised methods in real cases due to the lack of paired noisy-clean images. Meanwhile, most self-supervised denoising methods are ineffective as well when applied to the real-world denoising tasks because of their strict assumptions in applications. For example, as a typical method for self-supervised denoising, the original blind spot network (BSN) assumes that the noise is pixel-wise independent, which is much different from the real cases. To solve this problem, we propose a novel self-supervised real image denoising framework named Sampling Difference As Perturbation (SDAP) based on Random Sub-samples Generation (RSG) with a cyclic sample difference loss. Specifically, we dig deeper into the properties of BSN to make it more suitable for real noise. Surprisingly, we find that adding an appropriate perturbation to the training images can effectively improve the performance of BSN. Further, we propose that the sampling difference can be considered as perturbation to achieve better results. Finally we propose a new BSN framework in combination with our RSG strategy. The results show that it significantly outperforms other state-of-the-art self-supervised denoising methods on real-world datasets. The code is available at https://github.com/p1y2z3/SDAP. Yizhong Pan, Xiao Liu 0022, Xiangyu Liao, Yuanzhouhan Cao, Chao Ren 0002 |
ICCV | 4 |
| 2023 | FedCE: Personalized Federated Learning Method based on Clustering EnsemblesabstractFederated learning (FL) is a privacy-aware computing framework that enables multiple clients to collaborate in solving machine learning problems. In real scenarios, non-IID data held by different edge devices will degrade the performance of global FL models. To address this issue, most FL methods utilize cluster algorithms to group clients with similar distributions. However, these methods do not fully utilize the distribution features of client data, resulting in a lack of generalization in the cluster model. In order to make the cluster more suitable for the distribution features of user data, we propose a clustering-ensemble based federated learning method (FedCE) that sets each client associated with multiple clusters. We extract the features of client distributions to quantify the relationship between clients and clusters, and optimize the local model of clients through the historical performance of the cluster model. Furthermore, we dynamically estimate the number of clusters each client belongs to through the diversity of client performance. We conduct experiments on scenarios with mixture two-distributions, three-distributions and Dirichlet-distributions. The results show that the FedCE algorithm has better performance than the state-of-the-art clustered FL methods in both cluster and client models under different data distributions. Luxin Cai, Naiyue Chen, Yuanzhouhan Cao, Yidong Li |
ACM Multimedia | 3 |
| 2023 | S-OmniMVS: Incorporating Sphere Geometry into Omnidirectional Stereo MatchingabstractMulti-fisheye stereo matching is a promising task that employs the traditional multi-view stereo (MVS) pipeline with spherical sweeping to acquire omnidirectional depth. However, the existing omnidirectional MVS technologies neglect fisheye and omnidirectional distortions, yielding inferior performance. In this paper, we revisit omnidirectional MVS by incorporating three sphere geometry priors: spherical projection, spherical continuity, and spherical position. To deal with fisheye distortion, we propose a new distortion-adaptive fusion module to convert fisheye inputs into distortion-free spherical tangent representations by constructing a spherical projection space. Then these multi-scale features are adaptively aggregated with additional learnable offsets to enhance content perception. To handle omnidirectional distortion, we present a new spherical cost aggregation module with a comprehensive consideration of the spherical continuity and position. Concretely, we first design a rotation continuity compensation mechanism to ensure omnidirectional depth consistency of left-right boundaries without introducing extra computation. On the other hand, we encode the geometry-aware spherical position and push them into the cost aggregation to relieve panoramic distortion and perceive the 3D structure. Furthermore, to avoid the excessive concentration of depth hypothesis caused by inverse depth linear sampling, we develop a segmented sampling strategy that combines linear and exponential spaces to create S-OmniMVS, along with three sphere priors. Extensive experiments demonstrate the proposed method outperforms the state-of-the-art (SoTA) solutions by a large margin on various datasets both quantitatively and qualitatively. Zisong Chen, Chunyu Lin, Lang Nie, Zhijie Shen, Kang Liao, Yuanzhouhan Cao, Yao Zhao 0001 |
ACM Multimedia | 6 |
| 2023 | Weakly Supervised Object Detection With Class Prototypical NetworkabstractIn this paper, we aim to devise a new framework to compel the network to be equipped with the capability of detecting objects using image-level class labels as supervision. The challenge of such a weakly supervised setting mainly lies in how to make the network accurately understand both semantics and objectness of a given proposal without bounding box annotations. To this end, we contribute a concise framework, named Class Prototypical Network (CPNet). Concretely, our CPNet defines a set of learnable class prototypes to help classify object proposals. To endow the prototypes be not only discriminative for classes but also sensitive for proposals' objectness, we conduct both class-aware cross-attention and location-aware cross-attention between the feature embeddings of the learnable prototypes and the proposals. The learned attention scores are then used to form the proposal-level category information into the image-level one, making the entire framework be trained without any bounding box annotations. Besides, by applying these two kinds of attention mechanisms, the knowledge from both proposals' location and its class information can be successfully transferred into the corresponding prototypes. With the help of prototypes, our CPNet detects true positive object proposals. In addition, the CPNet further introduces a multi-head detection head to perform complementary training, preventing the model from falling into local discriminative parts and improving the model's performance on challenging non-rigid categories. We examine our CPNet on popular benchmarks,i.e., PASCAL VOC 2007, 2012 and MS COCO 2014. Extensive experiments show our CPNet is a simple and effective framework. Yidong Li, Yuanzhouhan Cao, Yushan Han, Yi Jin 0001, Yunchao Wei |
IEEE Trans. Multim. | 3 |
| 2022 | CMAN: Leaning Global Structure Correlation for Monocular 3D Object DetectionabstractThe key to 3D object detection is proper utilization of depth data. Compared with LiDAR based approaches, 3D object detection from a single image remains a challenging task due to the lack of structure information. Recent methods leverage monocular depth estimation as a way to produce 2D depth maps, and adopt the depth maps as additional source of input to explore structure information. However, these methods either encode local structure correlations, or encode long range structure correlations by iteratively passing local messages. In this work, we propose a cross modal attention network (CMAN) for monocular 3D object detection. It is built upon the self-attention module which learns attention map from single modal data. Our CMAN is able to encode structure correlations from depth data, and embed the structure correlations with appearance information which is learned from RGB data. Thanks to the attention learning mechanism, our CMAN learns global structure correlations without iteration. In order to reduce the computational burden, our CMAN adopts a novel node sampler to eliminate redundant nodes during the attention map calculation. Experiment results on benchmark KITTI3D dataset show that our proposed CMAN outperforms the state-of-the-art methods. Yuanzhouhan Cao, Hui Zhang 0091, Yidong Li, Chao Ren 0002, Congyan Lang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Deep Deblocker Driven Adaptive Iteration Scheme for Compressed Image RecoveryabstractIt is challenging to propose a flexible and effective framework for various JPEG compressed image recovery (CIR) tasks. In this paper, we propose a novel deep deblocker-driven adaptive iteration scheme, which can quickly and flexibly address various CIR tasks. First, a novel fidelity (NF) is introduced into CIR, and then the CIR problem is divided into inversion and deblocking subproblems by our improved split Bregman iteration (ISBI) algorithm. Next, we design a set of compact yet effective deep deblockers. These deblockers are used as implicit priors and also used for NF in the CIR problem. The convergence of our method is proved as well. To the best of our knowledge, our method is the first work to use deblockers as implicit priors. Extensive experiments demonstrate the effectiveness of our CIR method. Chao Ren 0002, Xiaohai He, Linbo Qing, Yuanzhouhan Cao |
ICME | 4 |
| 2021 | Learning Structure Affinity for Video Depth EstimationabstractDepth estimation is a structure learning problem. The affinity among neighbouring pixels plays an important role in inferring depth values. In this paper, we propose to learn structure affinity in both spatial and temporal domain for accurate depth estimation from monocular videos. Specifically, we first propose a convolutional spatial temporal propagation network (CSTPN) that learns affinity among neighbouring video frames. Secondly, we employ a structure knowledge distillation scheme that transfers the spatial temporal affinity learned by cumbersome network to compact network. By calculating pixel-wise similarities between neighboring frames and neighbouring sequences, our knowledge distillation scheme efficiently captures both short-term and long-term spatial temporal affinity. Finally, we apply a warping loss based on optical flow between video frames to further enforce the temporal affinity. Experiment results show that our proposed depth estimation approach outperform the state-of-the-art methods on both indoor and outdoor benchmark datasets. Yuanzhouhan Cao, Yidong Li, Haokui Zhang, Chao Ren 0002, Yifan Liu 0001 |
ACM Multimedia | 1 |
| 2021 | SPGAN: Face Forgery Using Spoofing Generative Adversarial NetworksabstractCurrent face spoof detection schemes mainly rely on physiological cues such as eye blinking, mouth movements, and micro-expression changes, or textural attributes of the face images [9]. But none of these methods represent a viable mechanism for makeup-induced spoofing, especially since makeup has been widely used. Compared with face alteration techniques such as plastic surgery, makeup is non-permanent and cost efficient, which makes makeup-induced spoofing become a realistic threat to the integrity of a face recognition system. To solve this problem, we propose a generative model to construct spoofing face images (confusing face images) for improving the accuracy and robustness of automatic face recognition. Our network structure is composed of two separate parts, with one using inter-attention mechanism to obtain interested face region, and another using intra-attention to translate imitation style with preserving imitation style-excluding details. These two attention mechanisms can precisely learn imitation style, where inter-attention pays more attention to imitation regions of image and intra-attention learns face attributes with long distance in image. To effectively discriminate generated images, we introduce an imitation style discriminator. Our model (SPGAN) generates face images that transfer the imitation style from target to subject image and preserve the imitation-excluding features. Experimental results demonstrate the performance of our model in improving quality of imitated face images. Yidong Li, Wenhua Liu, Yi Jin 0001, Yuanzhouhan Cao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Monocular Depth Estimation With Augmented Ordinal Depth RelationshipsabstractMost existing algorithms for depth estimation from single monocular images need large quantities of metric ground-truth depths for supervised learning. We show that relative depth can be an informative cue for metric depth estimation and can be easily obtained from vast stereo videos. Acquiring metric depths from stereo videos are sometimes impracticable due to the absence of camera parameters. In this paper, we propose to improve the performance of metric depth estimation with relative depths collected from stereo movie videos using existing stereo matching algorithm. We introduce a new “relative depth in stereo” (RDIS) dataset densely labeled with relative depths. We first pretrain a ResNet model on our RDIS dataset. Then, we finetune the model on RGB-D datasets with metric ground-truth depths. During our finetuning, we formulate depth estimation as a classification task. This re-formulation scheme enables us to obtain the confidence of a depth prediction in the form of probability distribution. With this confidence, we propose an information gain loss to make use of the predictions that are close to ground-truth. We evaluate our approach on both indoor and outdoor benchmark RGB-D datasets and achieve the state-of-the-art performance. Yuanzhouhan Cao, Ke Xian, Chunhua Shen, Zhiguo Cao 0001, Shugong Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Exploiting Temporal Consistency for Real-Time Video Depth EstimationabstractAccuracy of depth estimation from static images has been significantly improved recently, by exploiting hierarchical features from deep convolutional neural networks (CNNs). Compared with static images, vast information exists among video frames and can be exploited to improve the depth estimation performance. In this work, we focus on exploring temporal information from monocular videos for depth estimation. Specifically, we take the advantage of convolutional long short-term memory (CLSTM) and propose a novel spatial-temporal CSLTM (ST-CLSTM) structure. Our ST-CLSTM structure can capture not only the spatial features but also the temporal correlations/consistency among consecutive video frames with negligible increase in computational cost. Additionally, in order to maintain the temporal consistency among the estimated depth frames, we apply the generative adversarial learning scheme and design a temporal consistency loss. The temporal consistency loss is combined with the spatial loss to update the model in an end-to-end fashion. By taking advantage of the temporal information, we build a video depth estimation framework that runs in real-time and generates visually pleasant results. Moreover, our approach is flexible and can be generalized to most existing depth estimation frameworks. Code is available at: https://tinyurl.com/STCLSTM Haokui Zhang, Ying Li 0017, Yuanzhouhan Cao, Yu Liu 0029, Chunhua Shen, Youliang Yan |
ICCV | 3 |
| 2018 | Leveraging Convolutional Pose Machines for Fast and Accurate Head Pose EstimationabstractWe propose a head pose estimation framework that leverages on a recent keypoint detection model. More specifically, we apply the convolutional pose machines (CPMs) to input images, extract different types of facial keypoint features capturing appearance information and keypoint relationships, and train multilayer perceptrons (MLPs) and convolutional neural networks (CNNs) for head pose estimation. The benefit of leveraging on the CPMs (which we apply anyway for other purposes like tracking) is that we can design highly efficient models for practical usage. We evaluate our approach on the Annotated Facial Landmarks in the Wild (AFLW) dataset and achieve competitive results with the state-of-the-art. Yuanzhouhan Cao, Olivier Canévet, Jean-Marc Odobez |
IROS | 1 |
| 2018 | Estimating Depth From Monocular Images as Classification Using Deep Fully Convolutional Residual NetworksabstractDepth estimation from single monocular images is a key component in scene understanding. Most existing algorithms formulate depth estimation as a regression problem due to the continuous property of depths. However, the depth value of input data can hardly be regressed exactly to the ground-truth value. In this paper, we propose to formulate depth estimation as a pixelwise classification task. Specifically, we first discretize the continuous ground-truth depths into several bins and label the bins according to their depth ranges. Then, we solve the depth estimation problem as classification by training a fully convolutional deep residual network. Compared with estimating the exact depth of a single point, it is easier to estimate its depth range. More importantly, by performing depth classification instead of regression, we can easily obtain the confidence of a depth prediction in the form of probability distribution. With this confidence, we can apply an information gain loss to make use of the predictions that are close to ground-truth during training, as well as fully-connected conditional random fields for post-processing to further improve the performance. We test our proposed method on both indoor and outdoor benchmark RGB-Depth datasets and achieve state-of-the-art performance. Yuanzhouhan Cao, Zifeng Wu, Chunhua Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Temporal Pyramid Pooling-Based Convolutional Neural Network for Action RecognitionabstractEncouraged by the success of convolutional neural networks (CNNs) in image classification, recently much effort is spent on applying the CNNs to the video-based action recognition problems. One challenge is that a video contains a varying number of frames, which is incompatible to the standard input format of the CNNs. Existing methods handle this issue either by directly sampling a fixed number of frames or bypassing this issue by introducing a 3D convolutional layer, which conducts convolution in spatial-temporal domain. In this paper, we propose a novel network structure, which allows an arbitrary number of frames as the network input. The key to our solution is to introduce a module consisting of an encoding layer and a temporal pyramid pooling layer. The encoding layer maps the activation from the previous layers to a feature vector suitable for pooling, whereas the temporal pyramid pooling layer converts multiple frame-level activations into a fixed-length video-level representation. In addition, we adopt a feature concatenation layer that combines the appearance and motion information. Compared with the frame sampling strategy, our method avoids the risk of missing any important frames. Compared with the 3D convolutional method, which requires a huge video data set for network training, our model can be learned on a small target data set because we can leverage the off-the-shelf image-level CNN for model parameter initialization. Experiments on three challenging data sets, Hollywood2, HMDB51, and UCF101 demonstrate the effectiveness of the proposed network. Peng Wang 0023, Yuanzhouhan Cao, Chunhua Shen, Lingqiao Liu, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Exploiting Depth From Single Monocular Images for Object Detection and Semantic SegmentationabstractAugmenting RGB data with measured depth has been shown to improve the performance of a range of tasks in computer vision, including object detection and semantic segmentation. Although depth sensors such as the Microsoft Kinect have facilitated easy acquisition of such depth information, the vast majority of images used in vision tasks do not contain depth information. In this paper, we show that augmenting RGB images with estimated depth can also improve the accuracy of both object detection and semantic segmentation. Specifically, we first exploit the recent success of depth estimation from monocular images and learn a deep depth estimation model. Then, we learn deep depth features from the estimated depth and combine with RGB features for object detection and semantic segmentation. In addition, we propose an RGB-D semantic segmentation method, which applies a multi-task training scheme: semantic label prediction and depth value regression. We test our methods on several data sets and demonstrate that incorporating information from estimated depth improves the performance of object detection and semantic segmentation remarkably. Yuanzhouhan Cao, Chunhua Shen, Heng Tao Shen |
IEEE Trans. Image Process. | 1 |