EDBT 2026 Demo / reviewers in the wild / expert
Yawen Lu
dblp:254/8061
· DBLP profile ↗
26ranked-venue papers
19as first author
22since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 13 first-author · 15 since 2021Artificial intelligence and machine learning · 12 · 9 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D VisionabstractWe have witnessed significant progress in deep learning-based 3D vision, ranging from neural radiance field (NeRF) based 3D representation learning to applications in novel view synthesis (NVS). However, existing scene-level datasets for deep learning-based 3D vision, limited to ei-ther synthetic environments or a narrow selection of real-world scenes, are quite insufficient. This insufficiency not only hinders a comprehensive benchmark of existing methods but also caps what could be explored in deep learning-based 3D analysis. To address this critical gap, we present DL3DV-10K, a large-scale scene dataset, featuring 51.2 million frames from 10,510 videos captured from 65 types of point- of-interest (POI) locations, covering both bounded and unbounded scenes, with different levels of reflection, transparency, and lighting. We conducted a comprehensive benchmark of recent NVS methods on DL3DV-10K, which revealed valuable insights for future research in NVS. In addition, we have obtained encouraging results in a pilot study to learn generalizable NeRF from DL3DV-10K, which manifests the necessity of a large-scale scene-level dataset to forge a path toward a foundation model for learning 3D representation. Our DL3DV-10K dataset, benchmark results, and models will be publicly accessible. Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan 0001, Lantao Yu, Zixun Yu, Yawen Lu, Xuanmao Li, Xingpeng Sun, Rohan Ashok, Aniruddha Mukherjee, Hao Kang, Xiangrui Kong, Gang Hua 0001, Tianyi Zhang 0001, Bedrich Benes, Aniket Bera |
CVPR | 10 |
| 2024 | ProMotion: Prototypes as Motion LearnersabstractIn this work, we introduce PRoMoTION, a unified proto-typical transformer-based framework engineered to model fundamental motion tasks. PRoMoTION offers a range of compelling attributes that set it apart from current task-specific paradigms. (1) We adopt a prototypical perspective, establishing a unified paradigm that harmonizes disparate motion learning approaches. This novel paradigm stream-lines the architectural design, enabling the simultaneous assimilation of diverse motion information. (2) We capitalize on a dual mechanism involving the feature denoiser and the prototypical learner to decipher the intricacies of motion. This approach effectively circumvents the pitfalls of ambiguity in pixel-wise feature matching, significantly bolstering the robustness of motion representation. (3)) We demon-strate a profound degree of transferability across distinct motion patterns. This inherent versatility reverberates robustly across a comprehensive spectrum of both 2D and 3D downstream tasks. Empirical results demonstrate that PRoMOTION outperforms various well-known specialized architectures, achieving 0.54 and 0.054$AbsRel$error on the Sintel and KITTI depth datasets, 1.04 and 2.01 average endpoint error on the clean and final pass of Sintel flow benchmark, and 4.30 F1-all error on the KITTI flow bench-mark. For its efficacy, we hope our work can catalyze a paradigm shift in universal models in computer vision. Yawen Lu, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001, Yiming Cui 0002, Zhiwen Cao, Xueling Zhang, Victor Y. Chen, Heng Fan 0001 |
CVPR | 1 |
| 2024 | Prototypical Transformer As Unified Motion LearnersabstractIn this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoFormer seamlessly integrates prototype learning with Transformer by thoughtfully considering motion dynamics, introducing two innovative designs. First, Cross-Attention Prototyping discovers prototypes based on signature motion patterns, providing transparency in understanding motion scenes. Second, Latent Synchronization guides feature representation learning via prototypes, effectively mitigating the problem of motion uncertainty. Empirical results demonstrate that our approach achieves competitive performance on popular motion tasks such as optical flow and scene depth. Furthermore, it exhibits generality across various downstream tasks, including object tracking and video stabilization. Cheng Han 0001, Yawen Lu, James Liang, Zhiwen Cao, Qifan Wang 0001, Qiang Guan, Sohail A. Dianat, Raghuveer M. Rao, Tong Geng, Zhiqiang Tao, Dongfang Liu |
ICML | 2 |
| 2024 | Optical Flow as Spatial-Temporal Attention LearnersabstractOptical flow is an indispensable building block for various important computer vision tasks, including motion estimation, object tracking, and disparity measurement. To date, the dominant methods are CNN-based, leaving plenty of room for improvement. In this work, we propose TransFlow, a transformer architecture for optical flow estimation. Compared to dominant CNN-based methods, TransFlow demonstrates three advantages. First, it provides more accurate correlation and trustworthy matching in flow estimation by utilizing spatial self-attention and cross-attention mechanisms between adjacent frames to effectively capture global dependencies; Second, it recovers more compromised information (e.g., occlusion and motion blur) in flow estimation through long-range temporal association in dynamic scenes; Third, it introduces a concise self-learning paradigm, eliminating the need for complex and laborious multi-stage pre-training procedures. The versatility and superiority of TransFlow extend seamlessly to 3D scene motion, yielding competitive outcomes in 3D scene flow estimation. Our approach attains state-of-the-art results on benchmark datasets such as Sintel and KITTI-15, while also exhibiting exceptional performance on downstream tasks, including video object detection using the ImageNet VID dataset, video frame interpolation using the GoPro dataset, and video stabilization using the DeepStab dataset. We believe that the effectiveness of TransFlow positions it as a flexible baseline for both optical flow and scene flow estimation, offering promising avenues for future research and development. Yawen Lu, Cheng Han 0001, Qifan Wang 0001, Heng Fan 0001, Zhaodan Kong, Dongfang Liu, Victor Y. Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Label-Efficient Video Object Segmentation With Motion CluesabstractVideo object segmentation (VOS) plays an important role in video analysis and understanding, which in turn facilitates a number of diverse applications, including video editing, video rendering, and augmented reality / virtual reality. However, existing deep learning-based approaches rely heavily on a large number of pixel-wise annotated video frames to achieve promising results, which is notoriously laborious and costly. To address this, in this paper, we formulate unsupervised video object detection by exploring simulated dense labels and explicit motion clues. Specifically, we first propose an effective video label generator network based on the sparsely annotated frames and the flow motion between them. It can largely alleviate our dependence and limitation on the sparse labels. Furthermore, we propose a transformer-based architecture to model the appearance and motion clues simultaneously with the cross-attention module, in order to maximally overcome non-linear motion with potential occlusions. Extensive experiments show that the proposed method outperforms recent VOS methods on four popular benchmarks (i.e., DAVIS-16, FBMS, Youtube-VOS and SegTrack-v2). Moreover, the proposed method can be further applied to a wide range of wild scenes such as wild forests and animals. Because of its effectiveness and generalization, we believe that our method could serve as a useful basis for alleviating the dependence on dense annotation in video data. Yawen Lu, Jie Zhang 0066, Su Sun, Zhiwen Cao, Songlin Fei, Baijian Yang 0001, Victor Y. Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | TransFlow: Transformer as Flow LearnerabstractOptical flow is an indispensable building block for various important computer vision tasks, including motion estimation, object tracking, and disparity measurement. In this work, we propose TransFlow, a pure transformer architecture for optical flow estimation. Compared to dominant CNN-based methods, TransFlow demonstrates three advantages. First, it provides more accurate correlation and trustworthy matching in flow estimation by utilizing spatial self-attention and crossattention mechanisms between adjacent frames to effectively capture global dependencies; Second, it recovers more compromised information (e.g., occlusion and motion blur) in flow estimation through long-range temporal association in dynamic scenes; Third, it enables a concise self-learning paradigm and effectively eliminate the complex and laborious multi-stage pre-training procedures. We achieve the state-of-the-art results on the Sintel, KITTI-15, as well as several downstream tasks, including video object detection, interpolation and stabilization. For its efficacy, we hope TransFlow could serve as a flexible baseline for optical flow estimation. Yawen Lu, Qifan Wang 0001, Siqi Ma 0005, Tong Geng, Victor Y. Chen, Huaijin G. Chen, Dongfang Liu |
CVPR | 1 |
| 2022 | Inferring Camera Intrinsics Based on Surfaces of Revolution: A Single Image Geometric Network Approach for Camera CalibrationabstractCamera calibration is a necessary prerequisite in many applications of robotics, especially in robot vision in order to obtain metric reconstruction from a 2D image. In this paper, we address the problem of calibrating from a single image of a surface of revolution (SOR) based on deep learning, in order to determine the camera intrinsic parameters. Geometric constraints based on the symmetry properties of the SOR structure are deployed to our proposed learning-based camera calibration framework. To enable the calibration from a single view, we also propose a learning-based conics detection model fitting the geometric primitive of a cylinder. The calibration from a single view can be completed by minimizing the geometric constraints of two conics detected by the learning-based model with cylinder images as input. Objects with a surface of revolution are commonly visible in daily life, such as cans, bottles, and bowls, making this research both significant and practical. Finally, traditional calibration techniques are compared against our single image calibration. Experiments conducted on newly generated dataset demonstrate the effectiveness and robustness of the proposed method. Christopher Walker, Yawen Lu, Guoyu Lu 0001 |
ICASSP | 3 |
| 2022 | Self-supervised Depth Estimation from Spectral Consistency and Novel View SynthesisabstractSingle image depth estimation is a critical issue for robot vision, augmented reality, and many other applications when an image sequence is not available. Self-supervised single image depth estimation models target at predicting accurate disparity map just from one single image without ground truth supervision or stereo image pair during real applications. Compared with direct single image depth estimation, single image stereo algorithm can generate the depth from different camera perspectives. In this paper, we propose a novel architecture to infer accurate disparity by leveraging both spectral-consistency based learning model and view-prediction based stereo reconstruction algorithm. Direct spectral-consistency based method can avoid false positive matching in smooth regions. Single image stereo can preserve more distinct boundaries from another camera perspective. By learning confidence maps and designing a fusion strategy, the two disparities from the two approaches are able to be effectively fused to produce the refined disparity. Extensive experiments and ablations indicate that our method exploits both advantages of spectral consistency and view prediction, especially in constraining object boundaries and correcting wrong predicting regions. Yawen Lu, Guoyu Lu 0001 |
IJCNN | 1 |
| 2022 | An Unsupervised Approach for Simultaneous Visual Odometry and Single Image Depth EstimationabstractVisual odometry (VO) and single image depth estimation are critical for robot vision, 3D reconstruction, and camera pose estimation that can be applied to autonomous driving, map building, augmented reality and many other applications. Various supervised learning models have been proposed to train the VO or single image depth estimation framework for each targeted scene to improve the performance recently. However, little effort has been made to learn these separate tasks together without requiring the collection of a significant number of labels. This paper proposes a novel unsupervised learning approach to simultaneously perceive VO and single image depth estimation. In our framework, either of these tasks can benefit from each other through simultaneously learning these two tasks. We correlate these two tasks by enforcing depth consistency between VO and single image depth estimation. Based on the single image depth estimation, we can resolve the most common and challenging scaling issue of monocular VO. Meanwhile, through training from a sequence of images, VO can enhance the single image depth estimation accuracy. The effectiveness of our proposed method is demonstrated through extensive experiments compared with current state-of-the-art methods on the benchmark datasets. Yawen Lu, Guoyu Lu 0001 |
IJCNN | 1 |
| 2022 | Multi-view Geometry Consistency Network for Facial Micro-Expression Recognition From Various PerspectivesabstractGaze estimation plays an essential role in human attention recognition, human behavior analysis and augmented reality applications. Most of the deep neural network-based gaze estimation techniques apply supervised learning to extract features and regress 3D gaze vectors directly, leading to a vulnerability of high labor cost and limited generalization. In this work, we proposed a weakly-supervised method to jointly optimize the depth values of eye landmarks and relative poses with a multi-view geometric constraint to determine the final gaze vectors of observers. Specifically, we feed in sequential eye region images, and design a depth regression network to estimate the depth of the eye region landmarks, which are further utilized by the pose estimation network to estimate the relative changes of gaze vectors with multi-view geometric constraints in the iris regions. Experiments on both synthetic and real data show that the proposed method is feasible and promising to learn gaze estimation without strong pose supervision. Devarth Parikh, Yawen Lu, Nikola K. Kasabov, Guoyu Lu 0001 |
IJCNN | 2 |
| 2022 | From Local to Holistic: Self-supervised Single Image 3D Face Reconstruction Via Multi-level ConstraintsabstractSingle image 3D face reconstruction with accurate geometric details is a critical and challenging task due to the similar appearance on the face surface and fine details in organs. In this work, we introduce a self-supervised 3D face reconstruction approach from a single image that can recover detailed textures under different camera settings. The proposed network learns high-quality disparity maps from stereo face images during the training stage, while just a single face image is required to generate the 3D model in real applications. To recover fine details of each organ and facial surface, the framework introduces facial landmark spatial consistency to constrain the face recovering learning process in local point level and segmentation scheme on facial organs to constrain the correspondences at the organ level. The face shape and textures will further be refined by establishing holistic constraints based on the varying light illumination and shading information. The proposed learning framework can recover more accurate 3D facial details both quantitatively and qualitatively compared with state-of-the-art 3DMM and geometry-based reconstruction algorithms based on a single image. Yawen Lu, Michel Sarkis, Ning Bi, Guoyu Lu 0001 |
IROS | 1 |
| 2022 | 3D Modeling Beneath Ground: Plant Root Detection and Reconstruction Based on Ground-Penetrating Radarabstract3D object reconstruction based on deep neural networks has been gaining attention in recent years. However, recovering 3D shapes of hidden and buried objects remains to be a challenge. Ground Penetrating Radar (GPR) is among the most powerful and widely used instruments for detecting and locating underground objects such as plant roots and pipes, with affordable prices and continually evolving technology. This paper first proposes a deep convolution neural network-based anchor-free GPR curve signal detection net- work utilizing B-scans from a GPR sensor. The detection results can help obtain precisely fitted parabola curves. Furthermore, a graph neural network-based root shape reconstruction network is designated in order to progressively recover major taproot and then fine root branches’ geometry. Our results on the gprMax simulated root data as well as the real-world GPR data collected from apple orchards demonstrate the potential of using the proposed framework as a new approach for fine-grained underground object shape reconstruction in a non-destructive way. Yawen Lu, Guoyu Lu 0001 |
WACV | 1 |
| 2021 | Bridging the Invisible and Visible World: Translation between RGB and IR Images through Contour Cycle GANabstractInfrared Radiation (IR) images that capture the emitted IR signals from surrounding environment have been widely applied to pedestrian detection and video surveillance. However, there are not many textures that appeared on thermal images as compared to RGB images, which brings enormous challenges and difficulties in various tasks. Visible images cannot capture scenes in the dark and night environment due to the lack of light. In this paper, we propose a Contour GAN-based framework to learn the cross-domain representation and also map IR images with visible images. In contrast to existing structures of image translation that focus on spectral consistency, our framework also introduces strong spatial constraints, with further spectral enhancement by illuminance contrast and consistency constraints. Designating our method for IR and RGB image translation, it can generate high-quality translated images. Extensive experiments on near IR (NIR) and far IR (thermal) datasets demonstrate superior performance for quantitative and visual results. Yawen Lu, Guoyu Lu 0001 |
AVSS | 1 |
| 2021 | Matching as Color Images: Thermal Image Local Feature Detection and DescriptionabstractFeature detection and extraction is considered to be one of the most important aspects when it comes to any computer vision application, especially the autonomous driving field that is highly dependent on it. Thermal imaging is less explored in the field of autonomous driving mainly due to the high cost of the cameras and inferior techniques available for detection. Due to advances in technology the former does not hold true anymore and there lies tremendous scope for improvement in the latter. Autonomous driving relies heavily on multiple and sometimes redundant sensors, for which thermal sensors are a preferred addition. Thermal sensors being completely dependent on the infrared radiation emitted are able to frame and recognize objects even in the complete absence of light. However detecting features persistently through subsequent frames is difficult due to the lack of textures in thermal images. Motivated by this challenge, we propose a triplet based Siamese CNN for feature detection and extraction for any given thermal image. Our architecture is able to detect larger number of good feature points on thermal images than other best performed feature detection algorithms with superb matching performance based on our extracted descriptors. Bhavesh Deshpande, Sourabh Hanamsheth, Yawen Lu, Guoyu Lu 0001 |
ICASSP | 3 |
| 2021 | Stereo Rectification Based on Epipolar Constrained Neural NetworkabstractThis paper proposes a novel deep neural network-based method for stereo image rectification. The neural network is mainly based on the theoretical basis of epipolar constraints from multi-view geometry and intensity constraints of images, which separately describes the relationship of the corresponding epipolar lines between a pair of image, including the epipolar-line slope and y-intercept consistency of the epipolar lines and the consistency of the corresponding intensity values between two images. Benefiting from the designed rectification framework together with a feature matching module to extract accurate corresponding key-points between views, our method is able to realize a stable and accurate stereo rectification process. Compared with classic feature-based rectification methods, our proposed method can rectify small errors, and achieve a much more accurate rectification performance. Experiments conducted on synthetic face dataset and real-world KITTI dataset demonstrate the effectiveness and robustness of the proposed method. Yawen Lu, Guoyu Lu 0001 |
ICASSP | 2 |
| 2021 | 3D SceneFlowNet: Self-Supervised 3D Scene Flow Estimation Based on Graph CNNabstractDespite deep learning approaches have achieved promising successes in 2D optical flow estimation, it is a challenge to accurately estimate scene flow in 3D space as point clouds are inherently lacking topological information. In this paper, we aim at handling the problem of self-supervised 3D scene flow estimation based on dynamic graph convolutional neural networks (GCNNs), namely 3D SceneFlowNet. To better learn geometric relationships among points, we introduce EdgeConv to learn multiple-level features in a pyramid from point clouds and a self-attention mechanism to apply the multi-level features to predict the final scene flow. Our trained model can efficiently process a pair of adjacent point clouds as input and predict a 3D scene flow accurately without any supervision. The proposed approach achieves superior performance on both synthetic ModelNet40 dataset and real LiDAR scans from KITTI Scene Flow 2015 datasets. Yawen Lu, Yuhao Zhu 0001, Guoyu Lu 0001 |
ICIP | 1 |
| 2021 | Multi-view Geometry Consistency Network for Facial Micro-Expression Recognition From Various PerspectivesabstractMicro-expression can reveal underlying genuine emotions, but those rapid and subtle changes are hard to be captured by humans. Most existing research focuses on frontal face micro-expression recognition, which largely prevents the developed methods from the real applications and ignores the underlying geometry information. In this paper, we propose a multiview geometry consistency framework to enable the same emotion to be recognized under different perspectives, which is difficult for existing systems. Based on the developed 3D face reconstruction network, the multi-view micro-expression recognition framework empowers the emotion recognition capability to learn from multiple perspectives of the 3D reconstructed faces based on view-consistency, and a spiking neural network is further applied to capture omitted tiny and detailed changes. With a sequence of images, we explore the subtle changes across frames through optical flow, which, as a clue, enhances the performance of our designated network for micro-expression recognition. Extensive experiments on benchmark micro-expression datasets CAS(ME)2and SMIC demonstrate the proposed method achieves promising results on novel-view micro-expression recognition where existing methods mainly fail. Yawen Lu, Nikola K. Kasabov, Guoyu Lu 0001 |
IJCNN | 1 |
| 2021 | Deep Unsupervised 3D SfM Face Reconstruction Based on Massive Landmark Bundle AdjustmentabstractWe address the problem of reconstructing 3D human face from multi-view facial images using Structure-from-Motion (SfM) based on deep neural networks. While recent learning-based monocular view methods have shown impressive results for 3D facial reconstruction, the single-view setting is easily affected by depth ambiguities and poor face pose issues. In this paper, we propose a novel unsupervised 3D face reconstruction architecture by leveraging the multi-view geometry constraints to train accurate face pose and depth maps. Facial images from multiple perspectives of each 3D face model are input to train the network. Multi-view geometry constraints are fused into unsupervised network by establishing loss constraints from spatial and spectral perspectives. To make the trained 3D face have more details, facial landmark detector is explored to acquire massive facial information to constrain face pose and depth estimation. Through minimizing massive landmark displacement distance by bundle adjustment, an accurate 3D face model can be reconstructed. Extensive experiments demonstrate the superiority of our proposed approach over other methods. Yawen Lu, Guoyu Lu 0001 |
ACM Multimedia | 2 |
| 2021 | Unsupervised Gaze: Exploration of Geometric Constraints for 3D Gaze Estimation
Yawen Lu, Yuan Xin, Guoyu Lu 0001 |
MMM (2) | 1 |
| 2021 | An Alternative of LiDAR in Nighttime: Unsupervised Depth Estimation Based on Single Thermal ImageabstractMost existing autonomous driving vehicles and robots rely on active LiDAR sensors to detect the depth of the surrounding environment, which usually has limited resolution, and the emitted laser can be harmful to people and the environment. Current passive image-based depth estimation algorithms focus on color images from RGB sensors, which is not suitable for dark and night environment with limited lighting resource. In this paper, we propose a framework to estimate the scene depth directly from a single thermal image that can still observe the scene in the low lighting condition. We learn the thermal image depth estimation frame-work together with RGB cameras, which also mitigates the training condition due to the easy availability of RGB cameras. With the translated thermal images from color images from our generative adversarial network, our depth estimation method can explore the unique characteristics in thermal images through our novel contour and edge-aware constraints to obtain a stable and anti-artifact disparity. We apply the commonly available color cameras to navigate the learning process of thermal image depth estimation frame-work. With our approach, an accurate depth map can be predicted without any prior knowledge under various illumination conditions. Yawen Lu, Guoyu Lu 0001 |
WACV | 1 |
| 2021 | 3D plant root system reconstruction based on fusion of deep structure-from-motion and IMU
Yawen Lu, Zhanjie Chen, Awais Khan 0005, Carl Salvaggio, Guoyu Lu 0001 |
Multim. Tools Appl. | 1 |
| 2021 | Simultaneous Direct Depth Estimation and Synthesis Stereo for Single Image Plant Root ReconstructionabstractPlant roots are the main conduit to its interaction with the physical and biological environment. A 3D root system architecture can provide fundamental and applied knowledge of a plant's ability to thrive, but the construction of 3D structures for thin and complicated plant roots is challenging. Existing methods such as structure-from-motion and shape-from-silhouette require multiple images, as input, under a complicated optimization process, which is usually not convenient in fieldwork. Little effort has been put into investigating the applications of deep neural network methods to reconstruct thin objects, like plant root systems, from a single image. We propose an unsupervised learning scheme to estimate the root depth from only one image as input, which is further applied to reconstruct the complete root system. The boundaries of the reconstructed object usually contain large errors, which is a significant problem for roots with many thin branches. To reduce reconstruction errors, we integrate a cross-view GAN-based network into the reconstruction process, which predicts the root image from a different perspective. Based on the predicted view, we reconstruct the root system using stereo reconstruction, which helps to identify the accurately reconstructed points by enforcing their consistency. The results on both the real plant root dataset and the synthetic dataset demonstrate the effectiveness of the proposed algorithm compared with state-of-the-art single image 3D reconstruction models on plant roots. Yawen Lu, Devarth Parikh, Awais Khan 0005, Guoyu Lu 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Extending Single Beam Lidar To Full Resolution By Fusing with Single Image Depth EstimationabstractDepth estimation is playing an important role in indoor and outdoor scene understanding, autonomous driving, augmented reality and many other tasks. Vehicles and robotics are able to use active illumination sensors such as LIDAR to receive high precision depth estimation. However, high-resolution LIDARs are usually too expensive, which limits its massive production on various applications. Though single beam LIDAR enjoys the benefits of low cost, one beam depth sensing is not usually sufficient to perceive the surrounding environment in many scenarios. In this paper, we propose a deep learning based framework to explore to replicate similar or even higher performance as costly LIDARs with our designed self-supervised network and a low-cost single-beam LIDAR. After the accurate calibration with a visible camera, the single beam LIDAR can adjust the scale uncertainty of the depth map estimated by the visible camera. The adjusted depth map enjoys the benefits of high resolution and sensing accuracy as high beam LIDAR and maintains low-cost as single beam LIDAR. Thus we can achieve similar sensing effect of high beam LIDAR with more than a 30-100 times cheaper price (e.g., $80000 Velodyne HDL-64E LIDAR v.s. $2000 SICK TIM-781 2D LIDAR and normal camera). The proposed approach is verified on our collected dataset and public dataset with superior depth-sensing performance. Yawen Lu, Devarth Parikh, Yuan Xin, Guoyu Lu 0001 |
ICPR | 1 |
| 2020 | Multi-Task Learning for Single Image Depth Estimation and Segmentation Based on Unsupervised NetworkabstractDeep neural networks have significantly enhanced the performance of various computer vision tasks, including single image depth estimation and image segmentation. However, most existing approaches handle them in supervised manners and require a large number of ground truth labels that consume extensive human efforts and are not always available in real scenarios. In this paper, we propose a novel framework to estimate disparity maps and segment images simultaneously by jointly training an encoder-decoder-based interactive convolutional neural network (CNN) for single image depth estimation and a multiple class CNN for image segmentation. Learning the neural network for one task can be beneficial from simultaneously learning from another one under a multi-task learning framework. We show that our proposed model can learn per-pixel depth regression and segmentation from just a single image input. Extensive experiments on available public datasets, including KITTI, Cityscapes urban, and PASCAL-VOC demonstrate the effectiveness of our model compared with other state-of-the-art methods for both tasks. Yawen Lu, Michel Sarkis, Guoyu Lu 0001 |
ICRA | 1 |
| 2020 | Single Image Shape-from-SilhouettesabstractRecovering a 3D shape representation from one single image input has been attempted in recent years. Most of the works obtain 3D models from multiple images at different perspectives or ground truth CAD models. However, multiple images from different perspectives or 3D CAD models are not always available in real applications. In this work, we present a novel shape-from-silhouette method based on just a single image, which is an end-to-end learning framework relying on view synthesis and shape-from-silhouette methodology to reconstruct a 3D shape. The reconstructed 3D mesh can approach the real shape of target objects by constraining the silhouettes from both horizontal and vertical directions, especially for those objects with occlusions. Our proposed method achieves state-of-the-art performance on the ShapeNet dataset compared with other recent approaches targeting 3D reconstruction from a single image. Without requiring labor-intensive and time-consuming human annotations, the work has a broad potential to be applied in real-world applications. Yawen Lu, Guoyu Lu 0001 |
ACM Multimedia | 1 |
| 2019 | Deep Unsupervised Learning for Simultaneous Visual Odometry and Depth EstimationabstractVisual odometry and depth estimation are critical to understanding the scene and camera motion, which are particularly helpful to tasks such as scene understanding, autonomous driving, and robotics. Supervised learning methods have been applied in many deep neural network frameworks and demonstrated outstanding results in visual odometry and depth estimation. However, supervised learning requires a significant amount of labeled data for training, which consumes extensive time. In this paper, we explore an unsupervised learning framework that can learn a camera pose regressor from monocular video frames and estimates the scene depth simultaneously. The proposed method is able to perform accurate pose prediction as well as depth estimation, despite the absence of any ground truth data. The effectiveness of our proposed method is demonstrated through experiments on KITTI, Cityscapes, and Make3D benchmark datasets, which shows superb results compared with state-of-the-art methods in both tasks. Yawen Lu, Guoyu Lu 0001 |
ICIP | 1 |