EDBT 2026 Demo / reviewers in the wild / expert
Guoyu Lu 0001
dblp:120/8962-1
· DBLP profile ↗
66ranked-venue papers
28as first author
41since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 17 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 15 first-author · 18 since 2021Systems, architecture and hardware · 16 · 8 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing point cloud feature representation via historical node state increments in graph neural networks
Qihui Li 0001, Qiliang Du, Lianfang Tian, Yihua Shao, Guoyu Lu 0001 |
Pattern Recognit. | 5 |
| 2025 | Shading Meets Motion: Self-supervised Indoor 3D Reconstruction Via Simultaneous Shape-from-Shading and Structure-from-MotionabstractScene reconstruction has a wide range of applications in computer vision and robotics. To build practical constraints and feature correspondences, rich textures and distinguished gradient variations are particularly required in classic and learning-based SfM. When building low-texture regions with repeated patterns, especially mostly-white indoor rooms, there is a significant drop in performance. In this work, we propose Shading-SfM-Net, a novel framework for simultaneously learning a shape-from-shading network based on the inverse rendering constraint and a structure-from-motion framework based on warped keypoint, room layout, and geometric consistency, to improve structure-from-motion and surface reconstruction for low-texture indoor scenes. Shading-SfM-Net tightly incorporates the surface shape consistency and 3D geometric registration loss in order to dig into their mutual information and further overcome the instability on flat regions. We evaluate the proposed framework on texture-less indoor scenes (NYUv2 and ScanNet), and show that by simultaneously learning shading, motion and shape, our pipeline is able to achieve state-of-the-art performance with superior generalization capability for unseen texture-less datasets. Guoyu Lu 0001 |
CVPR | 1 |
| 2025 | Vision-Language Embodiment for Monocular Depth EstimationabstractDepth estimation is a core problem in robotic perception and vision tasks, but 3D reconstruction from a single image presents inherent uncertainties. Current depth estimation models primarily rely on inter-image relationships for supervised training, often overlooking the intrinsic information provided by the camera itself. We propose a method that embodies the camera model and its physical characteristics into a deep learning model, computing embodied scene depth through real-time interactions with road environments. The model can calculate embodied scene depth in real-time based on immediate environmental changes using only the intrinsic properties of the camera, without any additional equipment. By combining embodied scene depth with RGB image features, the model gains a comprehensive perspective on both geometric and visual details. Additionally, we incorporate text descriptions containing environmental content and depth information as priors for scene understanding, enriching the model’s perception of objects. This integration of image and language — two inherently ambiguous modalities — leverages their complementary strengths for monocular depth estimation. The real-time nature of the embodied language and depth prior model ensures that the model can continuously adjust its perception and behavior in dynamic environments. Experimental results show that the embodied depth estimation method enhances model performance across different scenes. Jinchang Zhang, Guoyu Lu 0001 |
CVPR | 2 |
| 2025 | Keypoint Detection and Description for Raw Bayer ImagesabstractKeypoint detection and local feature description are fundamental tasks in robotic perception, critical for applications such as SLAM, robot localization, feature matching, pose estimation, and 3D mapping. While existing methods predominantly operate on RGB images, we propose a novel network that directly processes raw images, bypassing the need for the Image Signal Processor (ISP). This approach significantly reduces hardware requirements and memory consumption, which is crucial for robotic vision systems. Our method introduces two custom-designed convolutional kernels capable of performing convolutions directly on raw images, preserving inter-channel information without converting to RGB. Experimental results show that our network outperforms existing algorithms on raw images, achieving higher accuracy and stability under large rotations and scale variations. This work represents the first attempt to develop a keypoint detection and feature description network specifically for raw images, offering a more efficient solution for resource-constrained environments. Jiakai Lin, Jinchang Zhang, Guoyu Lu 0001 |
ICRA | 3 |
| 2025 | Non-Destructive 3D Root Structure ModelingabstractDeep neural networks (DNNs) have gained significant attention in 3D object reconstruction. However, detecting and reconstructing hidden or buried objects underground remains a challenging task. Ground Penetrating Radar (GPR) has emerged as a cost-effective and non-destructive technology for subsurface object detection, including soil structures and pipelines. In this study, we present a deep convolutional neural network-based method for detecting target signals and performing curve parameter regression using multiple B-scans from GPR data. By leveraging the detection and regression outcomes, we further generate fitted curves that represent underground structures. To reconstruct a comprehensive and detailed 3D root structure, we design a shape reconstruction network that takes sparse sliced 3D points as input. The proposed approach is extensively trained and validated using synthetic 3D root datasets and simulated GPR data generated with gprMax. Additionally, the trained model demonstrates strong generalization capabilities when applied to real-world GPR data, ensuring its practical applicability. Guoyu Lu 0001 |
ICRA | 1 |
| 2025 | Bridging In-Situ and Satellite Data: Enhancing Gas Concentration Estimation Through Integration of Data-Driven and Physics-Based ModelingabstractGas concentration estimation is crucial for understanding and mitigating climate change. While most research and monitoring efforts focus on major greenhouse gases such as CO2, significantly less attention has been given to trace gases like NO2, which play a critical role in atmospheric chemistry and air quality. This paper aims to enhance trace gas concentration estimation by integrating physics-based models into data-driven neural network frameworks. Furthermore, to improve large-scale estimation accuracy, we incorporate in-situ measurements to refine neural network models trained on satellite observations. The resulting model can provide reliable large-scale gas concentration estimates, particularly for locations lacking precise in-situ measurements. This approach offers a novel pathway to enhance the accuracy and applicability of gas monitoring for climate and environmental research. While NO2serves as the target trace gas in this study, the proposed framework is potentially applicable to the prediction of other atmospheric gas concentrations. Guoyu Lu 0001 |
ICRA | 1 |
| 2025 | Depth Estimation Based on 3D Gaussian Splatting Siamese DefocusabstractDepth estimation is a fundamental task in 3D geometry. While stereo depth estimation can be achieved through triangulation methods, it is not as straightforward for monocular methods, which require the integration of global and local information. The Depth from Defocus (DFD) method utilizes camera lens models and parameters to recover depth information from blurred images and has been proven to perform well. However, these methods rely on All-In-Focus (AIF) images for depth estimation, which is nearly impossible to obtain in real-world applications. To address this issue, we propose a self-supervised framework based on 3D Gaussian splatting and Siamese networks. By learning the blur levels at different focal distances of the same scene in the focal stack, the framework predicts the defocus map and Circle of Confusion (CoC) from a single defocused image, using the defocus map as input to DepthNet for monocular depth estimation. The 3D Gaussian splatting model renders defocused images using the predicted CoC, and the differences between these and the real defocused images provide additional supervision signals for the Siamese Defocus self-supervised network. This framework has been validated on both artificially synthesized and real blurred datasets. Subsequent quantitative and visualization experiments demonstrate that our proposed framework is highly effective as a DFD method. Jinchang Zhang, Ningning Xu, Guoyu Lu 0001 |
ICRA | 4 |
| 2025 | 3D Plant Root Skeleton Detection and ExtractionabstractPlant roots typically exhibit a highly complex and dense architecture, incorporating numerous slender lateral roots and branches, which significantly hinders the precise capture and modeling of the entire root system. Additionally, roots often lack sufficient texture and color information, making it difficult to identify and track root traits using visual methods. Previous research on roots has been largely confined to 2D studies; however, exploring the 3D architecture of roots is crucial in botany. Since roots grow in real 3D space, 3D phenotypic information is more critical for studying genetic traits and their impact on root development. We have introduced a 3D root skeleton extraction method that efficiently derives the 3D architecture of plant roots from a few images. This method includes the detection and matching of lateral roots, triangulation to extract the skeletal structure of lateral roots, and the integration of lateral and primary roots. We developed a highly complex root dataset and tested our method on it. The extracted 3D root skeletons showed considerable similarity to the ground truth, validating the effectiveness of the model. This method can play a significant role in automated breeding robots. Through precise 3D root structure analysis, breeding robots can better identify plant phenotypic traits, especially root structure and growth patterns, helping practitioners select seeds with superior root systems. This automated approach not only improves breeding efficiency but also reduces manual intervention, making the breeding process more intelligent and efficient, thus advancing modern agriculture. Jiakai Lin, Jinchang Zhang, Wen-Zhan Song 0001, Tianming Liu 0001, Guoyu Lu 0001 |
IROS | 6 |
| 2025 | Depth Estimation Based on Fisheye CamerasabstractFisheye cameras, with their ultra-wide field of view, offer significant benefits for depth estimation in applications such as autonomous navigation, robotics, and immersive imaging by capturing more scene content from a single viewpoint. However, their strong radial distortion and varying spatial resolution across the image pose substantial challenges for accurate depth prediction. We present a deep learning–based framework for fisheye depth estimation that addresses these challenges while leveraging the wide coverage advantage. During training, rectified and synchronized stereo image pairs are used, with the right image and an estimated initial depth map reconstructing the left image. A refined spatial consistency loss is formulated by combining Structural Similarity Index Measure (SSIM) and L1 loss, with gradient-based weighting to emphasize disparity edges. To overcome the limitations of photometric loss in disparity learning, we normalize pixel intensities to better correlate disparity with appearance features. A fisheye-specific depth refinement module incorporates an uncertainty map derived from an inconsistency mask and a distortion distribution map, mitigating the effects of occlusion and high-distortion regions. This uncertainty map is used to weight the temporal warping loss, enhancing robustness against distortion-prone areas. During inference, only a single fisheye image is required to produce an accurate depth map. Experimental results demonstrate that our method improves reconstruction fidelity and robustness, making it well-suited for real-world fisheye-based depth estimation tasks. Yuwei Zhou, Guoyu Lu 0001 |
IROS | 2 |
| 2025 | MD-Mamba: Feature extractor on 3D representation with multi-view depth
Qihui Li 0001, Zongtan Li, Lianfang Tian, Qiliang Du, Guoyu Lu 0001 |
Image Vis. Comput. | 5 |
| 2025 | Enhanced Semantic Segmentation of LiDAR Point Clouds Using Projection-Based Deep Learning NetworksabstractLiDAR point cloud semantic segmentation has emerged as a fundamental technique for enabling intelligent perception in autonomous driving, robotics, and geospatial analysis. Point cloud segmentation methods are typically categorized into three types: point-based, voxel-based, and projection-based techniques. While point-based and voxel-based methods offer robust feature extraction, they face challenges related to computational efficiency and the handling of large-scale point clouds. Projection-based methods, on the other hand, project 3D point clouds into 2D representations, enabling the application of established 2D convolutional neural networks (CNNs) for segmentation tasks. Despite their advantages in efficiency, projection-based methods often suffer from the loss of spatial precision, leading to suboptimal segmentation performance, especially in complex and cluttered environments. In this paper, we propose a novel projection-based approach for semantic segmentation that addresses the limitations of existing methods. Our approach introduces a Multi-Scale Feature Embedding (MSFE) module to enhance feature extraction from the projected range images, combined with a Multi-Feature Fusion Module (MFFM) to integrate features at multiple scales. We further improve segmentation accuracy for challenging objects, such as pedestrians, traffic signs, and occluded structures, by utilizing a Dual Segmentation Head. Our experiments on the SemanticPOSS and SemanticKITTI datasets show significant improvements over existing methods, achieving a mean Intersection over Union (mIoU) of 53.6% on SemanticPOSS and 67.8% on SemanticKITTI. Notably, we achieve high performance on small and occluded objects, like trashcans (55.9% mIoU) and fences (49.5% mIoU), demonstrating the effectiveness of our approach for real-world applications like autonomous driving. Qihui Li 0001, Qiliang Du, Lianfang Tian, Wenzi Liao, Guoyu Lu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Underground Mapping and Localization Based on Ground-Penetrating Radar
Jinchang Zhang, Guoyu Lu 0001 |
ACCV (1) | 2 |
| 2024 | Morphable-SfS: Enhancing Shape-from-Silhouette Via Morphable ModelingabstractReconstructing accurate object shapes based on single image inputs is still a critical and challenging task, mainly due to the potential shape ambiguity and occlusion. Most existing single image 3D reconstruction approaches, either trained on stereo setting or structure-from-motion, estimate 2.5D visible models which generally reconstruct one viewpoint of objects. We propose a method to leverage both the general Morphable Model on common objects and a multi-view synthesis-based shape-from-silhouette model to reconstruct complete object shapes. We use the proposed method to exploit strong geometric and perceptual cues in 3D shape reconstruction. During the inference, the trained model is able to produce high-quality and complete meshes with finely detailed structures from a 2D image captured from arbitrary perspectives. The proposed method is evaluated on both large-scale synthetic ShapeNet and real-world Pascal 3D+ and Pix3D datasets. The proposed work achieves state-of-the-art results compared with other recent self-supervised methods. Moreover, it shows a good capability of being applied in the unseen object reconstruction tasks. Guoyu Lu 0001 |
ICRA | 1 |
| 2024 | TVFusionGAN: Thermal-Visible Image Fusion Based on Multi-level Adversarial Network StrategyabstractThermal imaging is effective in low-light or night-time conditions due to its ability to capture thermal radiation differences, but lacks texture compared to visible images. Conversely, visible images retain more texture information, particularly during the daytime, but perform poorly at night. To address the limitations of both modalities, recent methods have utilized fusion techniques to generate images that combine thermal and visible properties. This paper presents an end-to-end fusion network leveraging generative adversarial networks (GANs) to fuse salient components from both modalities. Our network includes a generator and two discriminators. The generator produces fusion images with salient objects using a specially designed CIoU loss, while the discriminators ensure that the fused images are salient at both holistic and local scales. One discriminator encourages the fused images to resemble visible images overall, while the other ensures that targeted objects in the fused images are as salient as in thermal images. Our method effectively preserves thermal radiation of salient objects in infrared images while incorporating the textures of visible images. Guoyu Lu 0001 |
ICRA | 1 |
| 2024 | SLAM Based on Camera-2D LiDAR FusionabstractThe SLAM system plays a pivotal role in robotic mapping and localization, leveraging various sensor technologies to achieve precision. Traditional passive sensors, such as RGB cameras, offer high-resolution imagery at a lower cost for SLAM applications, yet they fall short in accurately estimating 3D positions and camera motions. On the other hand, LiDARs excel in generating accurate 3D maps but often come at a higher price and lower resolution. While active illumination sensors like LiDAR provide precise depth estimation, the prohibitive cost of high-resolution LiDAR systems restricts their widespread adoption across diverse applications. Although 2D single-beam LiDAR is more affordable, its limited depth sensing capability hampers comprehensive environmental perception. Addressing these limitations, this paper introduces a deep learning framework aimed at enhancing SLAM performance through the strategic fusion of camera and 2D LiDAR data. Our approach employs a novel self-supervised network alongside an economical single-beam LiDAR, striving to achieve or surpass the performance of more expensive LiDAR systems. The integration of single-beam LiDAR with our system allows for dynamic adjustment of scale uncertainty in depth maps generated by monocular camera systems within SLAM. Consequently, this fusion method enjoys the high-resolution and accuracy benefits of advanced LiDAR systems with the cost-effectiveness of 2D LiDAR sensors. Through this innovative combination, we demonstrate a SLAM system that not only maintains high fidelity in mapping and localization but also ensures affordability and broad applicability. Guoyu Lu 0001 |
ICRA | 1 |
| 2024 | Embodiment: Self-Supervised Depth Estimation Based on Camera ModelsabstractDepth estimationn is a critical topic for robotics and vision-related tasks. In monocular depth estimation, in comparison with supervised learning that requires expensive ground truth labeling, self-supervised methods possess great potential due to no labeling cost. However, self-supervised learning still has a large gap with supervised learning in 3D reconstruction and depth estimation performance. Meanwhile, scaling is also a major issue for monocular unsupervised depth estimation, which commonly still needs ground truth scale from GPS, LiDAR, or existing maps to correct. In the era of deep learning, existing methods primarily rely on exploring image relationships to train unsupervised neural networks, while the physical properties of the camera itself—such as intrinsics and extrinsics—are often overlooked. These physical properties are not just mathematical parameters; they are embodiments of the camera’s interaction with the physical world. By embedding these physical properties into the depth learning model, we can calculate depth priors for ground regions and regions connected to the ground based on physical principles, providing free supervision signals without the need for additional sensors. This approach is not only easy to implement but also enhances the effects of all unsupervised methods by embedding the camera’s physical properties into the model, thereby achieving an embodied understanding of the real world. Jinchang Zhang, Praveen Kumar Reddy, Xue-Iuan Wong, Yiannis Aloimonos, Guoyu Lu 0001 |
IROS | 5 |
| 2023 | RawSeg: Grid Spatial and Spectral Attended Semantic Segmentation Based on Raw Bayer Images
Guoyu Lu 0001 |
BMVC | 1 |
| 2023 | Deep Unsupervised Visual Odometry Via Bundle Adjusted Pose Graph OptimizationabstractUnsupervised visual odometry as an active topic has attracted extensive attention, benefiting from its label-free practical value and robustness in real-world scenarios. However, the performance of camera pose estimation and tracking through deep neural network is still not as ideal as most other tasks, such as detection, segmentation and depth estimation, due to the lack of drift correction in the estimated trajectory and map optimization in the recovered 3D scenes. In this work, we introduce pose graph and bundle adjustment optimization to our network training process, which iteratively updates both the motion and depth estimations from the deep learning network, and enforces the refined outputs to further meet the unsupervised photometric and geometric constraints. The integration of pose graph and bundle adjustment is easy to implement and significantly enhances the training effectiveness. Experiments on KITTI dataset demonstrate that the introduced method achieves a significant improvement in motion estimation compared with other recent unsupervised monocular visual odometry algorithms. Guoyu Lu 0001 |
ICRA | 1 |
| 2023 | Bird-View 3D Reconstruction for Crops with Repeated TexturesabstractLarge-scale in-situ 3D reconstruction of crop fields presents a challenging task, as the 3D crop structures play a crucial role in plant phenotyping and significantly influence crop growth and yield. While existing efforts focus on close-range plants, only a limited number of deep learning-based methods have been developed explicitly for large-scale 3D crop reconstruction, mainly due to the scarcity of large-scale crop sensing data. In this paper, we leverage unmanned aerial vehicles (UAVs) in agriculture and utilize a recently captured multi-view real-world snap beans crop dataset to develop an unsupervised structure-from-motion (SfM) framework. Our framework is designed specifically for reconstructing large-scale 3D crop structures. It addresses the challenge of inaccurate depth inference caused by excessively repeated patterns in the crop dataset, resulting in highly accurate 3D crop reconstruction for large-scale scenarios. Through experiments conducted on the crop dataset, we demonstrate the accuracy and robustness of our 3D crop reconstruction algorithm. The application of our proposed framework has the potential to advance research in agriculture, enabling better plant phenotyping and understanding of crop growth and yield. Guoyu Lu 0001 |
IROS | 1 |
| 2023 | Object Detection Based on Raw Bayer ImagesabstractBayer pattern is a widely used Color Filter Array (CFA) for digital image sensors, efficiently capturing different light wavelengths on different pixels without the need for a costly ISP pipeline. The resulting single-channel raw Bayer images offer benefits such as spectral wavelength sensitivity and low time latency. However, object detection based on Bayer images has been underexplored due to challenges in human observation and algorithm design caused by the discontinuous color channels in adjacent pixels. To address this issue, we propose the BayerDetect network, an end-to-end deep object detection framework that aims to achieve fast, accurate, and memory-efficient object detection. Unlike RGB color images, where each pixel encodes spectral context from adjacent pixels during ISP color interpolation, raw Bayer images lack spectral context. To enhance the spectral context, the BayerDetect network introduces a spectral frequency attention block, transforming the raw Bayer image pattern to the frequency domain. In object detection, clear object boundaries are essential for accurate bounding box predictions. To handle the challenges posed by alternating spectral channels and mitigate the influence of discontinuous boundaries, the BayerDetect network incorporates a spatial attention scheme that utilizes deformable convolutional kernels in multiple scales to explore spatial context effectively. The extracted convolutional features are then passed through a sparse set of proposal boxes for detection and classification. We conducted experiments on both public and self-collected raw Bayer images, and the results demonstrate the superb performance of the BayerDetect network in object detection tasks. Guoyu Lu 0001 |
IROS | 1 |
| 2023 | Fast and Accurate: Video Enhancement Using Sparse DepthabstractThis paper presents a general framework to build fast and accurate algorithms for video enhancement tasks such as super-resolution, deblurring, and denoising. Essential to our framework is the realization that the accuracy, rather than the density, of pixel flows is what is required for high-quality video enhancement. Most of prior works take the opposite approach: they estimate dense (per-pixel)—but generally less robust—flows, mostly using computationally costly algorithms. Instead, we propose a lightweight flow estimation algorithm; it fuses the sparse point cloud data and (even sparser and less reliable) IMU data available in modern autonomous agents to estimate the flow information. Building on top of the flow estimation, we demonstrate a general framework that integrates the flows in a plug-andplay fashion with different task-specific layers. Algorithms built in our framework achieve 1.78× — 187.41× speedup while providing a 0.42 dB – 6.70 dB quality improvement over competing methods. Yu Feng 0007, Patrick Hansen, Paul N. Whatmough, Guoyu Lu 0001, Yuhao Zhu 0001 |
WACV | 4 |
| 2022 | Inferring Camera Intrinsics Based on Surfaces of Revolution: A Single Image Geometric Network Approach for Camera CalibrationabstractCamera calibration is a necessary prerequisite in many applications of robotics, especially in robot vision in order to obtain metric reconstruction from a 2D image. In this paper, we address the problem of calibrating from a single image of a surface of revolution (SOR) based on deep learning, in order to determine the camera intrinsic parameters. Geometric constraints based on the symmetry properties of the SOR structure are deployed to our proposed learning-based camera calibration framework. To enable the calibration from a single view, we also propose a learning-based conics detection model fitting the geometric primitive of a cylinder. The calibration from a single view can be completed by minimizing the geometric constraints of two conics detected by the learning-based model with cylinder images as input. Objects with a surface of revolution are commonly visible in daily life, such as cans, bottles, and bowls, making this research both significant and practical. Finally, traditional calibration techniques are compared against our single image calibration. Experiments conducted on newly generated dataset demonstrate the effectiveness and robustness of the proposed method. Christopher Walker, Yawen Lu, Guoyu Lu 0001 |
ICASSP | 4 |
| 2022 | Image-based Localization for Self-driving Vehicles Based on Online Network Adjustment in A Dynamic ScopeabstractImage-based localization provides an alternative solution for camera pose estimation, which is a crucial component for self-driving vehicles. Localization for vehicles requires continuous feedback. We propose a solution that can accurately estimate the vehicle position and orientation. In this solution, we provide a complete pipeline for self-driving vehicles, including map building and camera pose estimation. We first design a convolutional neural network and train the localization system based on the entire global map. During the real-time localization stage, we fine-tune the network regressor online through the training images in adjacent locations in the map, which can enhance the localization accuracy significantly. Depending on the vehicle motion, we adjust the scope of local training images dynamically. We demonstrate the superior performance of our method through experiments on benchmark dataset. Guoyu Lu 0001 |
IJCNN | 1 |
| 2022 | Self-supervised Depth Estimation from Spectral Consistency and Novel View SynthesisabstractSingle image depth estimation is a critical issue for robot vision, augmented reality, and many other applications when an image sequence is not available. Self-supervised single image depth estimation models target at predicting accurate disparity map just from one single image without ground truth supervision or stereo image pair during real applications. Compared with direct single image depth estimation, single image stereo algorithm can generate the depth from different camera perspectives. In this paper, we propose a novel architecture to infer accurate disparity by leveraging both spectral-consistency based learning model and view-prediction based stereo reconstruction algorithm. Direct spectral-consistency based method can avoid false positive matching in smooth regions. Single image stereo can preserve more distinct boundaries from another camera perspective. By learning confidence maps and designing a fusion strategy, the two disparities from the two approaches are able to be effectively fused to produce the refined disparity. Extensive experiments and ablations indicate that our method exploits both advantages of spectral consistency and view prediction, especially in constraining object boundaries and correcting wrong predicting regions. Yawen Lu, Guoyu Lu 0001 |
IJCNN | 2 |
| 2022 | An Unsupervised Approach for Simultaneous Visual Odometry and Single Image Depth EstimationabstractVisual odometry (VO) and single image depth estimation are critical for robot vision, 3D reconstruction, and camera pose estimation that can be applied to autonomous driving, map building, augmented reality and many other applications. Various supervised learning models have been proposed to train the VO or single image depth estimation framework for each targeted scene to improve the performance recently. However, little effort has been made to learn these separate tasks together without requiring the collection of a significant number of labels. This paper proposes a novel unsupervised learning approach to simultaneously perceive VO and single image depth estimation. In our framework, either of these tasks can benefit from each other through simultaneously learning these two tasks. We correlate these two tasks by enforcing depth consistency between VO and single image depth estimation. Based on the single image depth estimation, we can resolve the most common and challenging scaling issue of monocular VO. Meanwhile, through training from a sequence of images, VO can enhance the single image depth estimation accuracy. The effectiveness of our proposed method is demonstrated through extensive experiments compared with current state-of-the-art methods on the benchmark datasets. Yawen Lu, Guoyu Lu 0001 |
IJCNN | 2 |
| 2022 | Multi-view Geometry Consistency Network for Facial Micro-Expression Recognition From Various PerspectivesabstractGaze estimation plays an essential role in human attention recognition, human behavior analysis and augmented reality applications. Most of the deep neural network-based gaze estimation techniques apply supervised learning to extract features and regress 3D gaze vectors directly, leading to a vulnerability of high labor cost and limited generalization. In this work, we proposed a weakly-supervised method to jointly optimize the depth values of eye landmarks and relative poses with a multi-view geometric constraint to determine the final gaze vectors of observers. Specifically, we feed in sequential eye region images, and design a depth regression network to estimate the depth of the eye region landmarks, which are further utilized by the pose estimation network to estimate the relative changes of gaze vectors with multi-view geometric constraints in the iris regions. Experiments on both synthetic and real data show that the proposed method is feasible and promising to learn gaze estimation without strong pose supervision. Devarth Parikh, Yawen Lu, Nikola K. Kasabov, Guoyu Lu 0001 |
IJCNN | 4 |
| 2022 | From Local to Holistic: Self-supervised Single Image 3D Face Reconstruction Via Multi-level ConstraintsabstractSingle image 3D face reconstruction with accurate geometric details is a critical and challenging task due to the similar appearance on the face surface and fine details in organs. In this work, we introduce a self-supervised 3D face reconstruction approach from a single image that can recover detailed textures under different camera settings. The proposed network learns high-quality disparity maps from stereo face images during the training stage, while just a single face image is required to generate the 3D model in real applications. To recover fine details of each organ and facial surface, the framework introduces facial landmark spatial consistency to constrain the face recovering learning process in local point level and segmentation scheme on facial organs to constrain the correspondences at the organ level. The face shape and textures will further be refined by establishing holistic constraints based on the varying light illumination and shading information. The proposed learning framework can recover more accurate 3D facial details both quantitatively and qualitatively compared with state-of-the-art 3DMM and geometry-based reconstruction algorithms based on a single image. Yawen Lu, Michel Sarkis, Ning Bi, Guoyu Lu 0001 |
IROS | 4 |
| 2022 | 3D Modeling Beneath Ground: Plant Root Detection and Reconstruction Based on Ground-Penetrating Radarabstract3D object reconstruction based on deep neural networks has been gaining attention in recent years. However, recovering 3D shapes of hidden and buried objects remains to be a challenge. Ground Penetrating Radar (GPR) is among the most powerful and widely used instruments for detecting and locating underground objects such as plant roots and pipes, with affordable prices and continually evolving technology. This paper first proposes a deep convolution neural network-based anchor-free GPR curve signal detection net- work utilizing B-scans from a GPR sensor. The detection results can help obtain precisely fitted parabola curves. Furthermore, a graph neural network-based root shape reconstruction network is designated in order to progressively recover major taproot and then fine root branches’ geometry. Our results on the gprMax simulated root data as well as the real-world GPR data collected from apple orchards demonstrate the potential of using the proposed framework as a new approach for fine-grained underground object shape reconstruction in a non-destructive way. Yawen Lu, Guoyu Lu 0001 |
WACV | 2 |
| 2022 | Regularization and attention feature distillation base on light CNN for Hyperspectral face recognition
Jieyi Niu, Guoyu Lu 0001 |
Multim. Tools Appl. | 4 |
| 2021 | Bridging the Invisible and Visible World: Translation between RGB and IR Images through Contour Cycle GANabstractInfrared Radiation (IR) images that capture the emitted IR signals from surrounding environment have been widely applied to pedestrian detection and video surveillance. However, there are not many textures that appeared on thermal images as compared to RGB images, which brings enormous challenges and difficulties in various tasks. Visible images cannot capture scenes in the dark and night environment due to the lack of light. In this paper, we propose a Contour GAN-based framework to learn the cross-domain representation and also map IR images with visible images. In contrast to existing structures of image translation that focus on spectral consistency, our framework also introduces strong spatial constraints, with further spectral enhancement by illuminance contrast and consistency constraints. Designating our method for IR and RGB image translation, it can generate high-quality translated images. Extensive experiments on near IR (NIR) and far IR (thermal) datasets demonstrate superior performance for quantitative and visual results. Yawen Lu, Guoyu Lu 0001 |
AVSS | 2 |
| 2021 | Matching as Color Images: Thermal Image Local Feature Detection and DescriptionabstractFeature detection and extraction is considered to be one of the most important aspects when it comes to any computer vision application, especially the autonomous driving field that is highly dependent on it. Thermal imaging is less explored in the field of autonomous driving mainly due to the high cost of the cameras and inferior techniques available for detection. Due to advances in technology the former does not hold true anymore and there lies tremendous scope for improvement in the latter. Autonomous driving relies heavily on multiple and sometimes redundant sensors, for which thermal sensors are a preferred addition. Thermal sensors being completely dependent on the infrared radiation emitted are able to frame and recognize objects even in the complete absence of light. However detecting features persistently through subsequent frames is difficult due to the lack of textures in thermal images. Motivated by this challenge, we propose a triplet based Siamese CNN for feature detection and extraction for any given thermal image. Our architecture is able to detect larger number of good feature points on thermal images than other best performed feature detection algorithms with superb matching performance based on our extracted descriptors. Bhavesh Deshpande, Sourabh Hanamsheth, Yawen Lu, Guoyu Lu 0001 |
ICASSP | 4 |
| 2021 | Stereo Rectification Based on Epipolar Constrained Neural NetworkabstractThis paper proposes a novel deep neural network-based method for stereo image rectification. The neural network is mainly based on the theoretical basis of epipolar constraints from multi-view geometry and intensity constraints of images, which separately describes the relationship of the corresponding epipolar lines between a pair of image, including the epipolar-line slope and y-intercept consistency of the epipolar lines and the consistency of the corresponding intensity values between two images. Benefiting from the designed rectification framework together with a feature matching module to extract accurate corresponding key-points between views, our method is able to realize a stable and accurate stereo rectification process. Compared with classic feature-based rectification methods, our proposed method can rectify small errors, and achieve a much more accurate rectification performance. Experiments conducted on synthetic face dataset and real-world KITTI dataset demonstrate the effectiveness and robustness of the proposed method. Yawen Lu, Guoyu Lu 0001 |
ICASSP | 3 |
| 2021 | 3D SceneFlowNet: Self-Supervised 3D Scene Flow Estimation Based on Graph CNNabstractDespite deep learning approaches have achieved promising successes in 2D optical flow estimation, it is a challenge to accurately estimate scene flow in 3D space as point clouds are inherently lacking topological information. In this paper, we aim at handling the problem of self-supervised 3D scene flow estimation based on dynamic graph convolutional neural networks (GCNNs), namely 3D SceneFlowNet. To better learn geometric relationships among points, we introduce EdgeConv to learn multiple-level features in a pyramid from point clouds and a self-attention mechanism to apply the multi-level features to predict the final scene flow. Our trained model can efficiently process a pair of adjacent point clouds as input and predict a 3D scene flow accurately without any supervision. The proposed approach achieves superior performance on both synthetic ModelNet40 dataset and real LiDAR scans from KITTI Scene Flow 2015 datasets. Yawen Lu, Yuhao Zhu 0001, Guoyu Lu 0001 |
ICIP | 3 |
| 2021 | Unsupervised HDR Image Reconstruction Based on Over/Under-Exposed LDR Image PairabstractThis paper proposes an unsupervised high dynamic range (HDR) image reconstruction method based on an over/under-exposed low dynamic range (LDR) image pair. The framework includes two end-to-end branches: transferring an over-exposed image input to under-exposed images and transferring an under-exposed image input to over-exposed images. The LDR images with the same exposure from the two branches are averaged, and then reconstruct an HDR image by merging them. When training the model, we use the L1loss of the same exposure image of the two branches and MEF-SSIM loss function as the objective function to ensure that the two branches get a similar visual effect at the same exposure, and use RGB loss and HSV loss to constrain the brightness and saturation of different exposure images. Experiments demonstrate that our unsupervised framework can generate comparable results with state-of-the-art supervised learning methods. Hao Wang 0137, Tao Zhang 0025, Guoyu Lu 0001 |
ICME | 3 |
| 2021 | VINS-Motion: Tightly-coupled Fusion of VINS and Motion ConstraintabstractIn this paper, we develop a novel visual-inertial navigation system with motion constraint (VINS-Motion), which extends the visual-inertial navigation system (VINS) to incorporate vehicle motion constraints for improving the autonomous vehicles localization accuracy. Besides the prior information, IMU measurement residual, and visual measurement residual utilized in VINS, vehicle orientation/velocity constraint is first exploited to constitute motion residual. We minimize the sum of priors and Mahalanobis norms of three kinds of residuals to obtain a maximum posteriori estimation, thus increasing system consistency and accuracy. Stop detection is also added to help eliminate the abnormal jitter of the estimated poses during stopping, thus ensuring reasonability of the trajectory. The pro-posed approach is validated on public datasets and compared against state-of-the-art algorithms, which demonstrates that VINS-Motion achieves significantly higher positioning accuracy. Zhelin Yu, Lidong Zhu, Guoyu Lu 0001 |
ICRA | 3 |
| 2021 | Multi-view Geometry Consistency Network for Facial Micro-Expression Recognition From Various PerspectivesabstractMicro-expression can reveal underlying genuine emotions, but those rapid and subtle changes are hard to be captured by humans. Most existing research focuses on frontal face micro-expression recognition, which largely prevents the developed methods from the real applications and ignores the underlying geometry information. In this paper, we propose a multiview geometry consistency framework to enable the same emotion to be recognized under different perspectives, which is difficult for existing systems. Based on the developed 3D face reconstruction network, the multi-view micro-expression recognition framework empowers the emotion recognition capability to learn from multiple perspectives of the 3D reconstructed faces based on view-consistency, and a spiking neural network is further applied to capture omitted tiny and detailed changes. With a sequence of images, we explore the subtle changes across frames through optical flow, which, as a clue, enhances the performance of our designated network for micro-expression recognition. Extensive experiments on benchmark micro-expression datasets CAS(ME)2and SMIC demonstrate the proposed method achieves promising results on novel-view micro-expression recognition where existing methods mainly fail. Yawen Lu, Nikola K. Kasabov, Guoyu Lu 0001 |
IJCNN | 3 |
| 2021 | Deep Unsupervised 3D SfM Face Reconstruction Based on Massive Landmark Bundle AdjustmentabstractWe address the problem of reconstructing 3D human face from multi-view facial images using Structure-from-Motion (SfM) based on deep neural networks. While recent learning-based monocular view methods have shown impressive results for 3D facial reconstruction, the single-view setting is easily affected by depth ambiguities and poor face pose issues. In this paper, we propose a novel unsupervised 3D face reconstruction architecture by leveraging the multi-view geometry constraints to train accurate face pose and depth maps. Facial images from multiple perspectives of each 3D face model are input to train the network. Multi-view geometry constraints are fused into unsupervised network by establishing loss constraints from spatial and spectral perspectives. To make the trained 3D face have more details, facial landmark detector is explored to acquire massive facial information to constrain face pose and depth estimation. Through minimizing massive landmark displacement distance by bundle adjustment, an accurate 3D face model can be reconstructed. Extensive experiments demonstrate the superiority of our proposed approach over other methods. Yawen Lu, Guoyu Lu 0001 |
ACM Multimedia | 4 |
| 2021 | Unsupervised Gaze: Exploration of Geometric Constraints for 3D Gaze Estimation
Yawen Lu, Yuan Xin, Guoyu Lu 0001 |
MMM (2) | 5 |
| 2021 | An Alternative of LiDAR in Nighttime: Unsupervised Depth Estimation Based on Single Thermal ImageabstractMost existing autonomous driving vehicles and robots rely on active LiDAR sensors to detect the depth of the surrounding environment, which usually has limited resolution, and the emitted laser can be harmful to people and the environment. Current passive image-based depth estimation algorithms focus on color images from RGB sensors, which is not suitable for dark and night environment with limited lighting resource. In this paper, we propose a framework to estimate the scene depth directly from a single thermal image that can still observe the scene in the low lighting condition. We learn the thermal image depth estimation frame-work together with RGB cameras, which also mitigates the training condition due to the easy availability of RGB cameras. With the translated thermal images from color images from our generative adversarial network, our depth estimation method can explore the unique characteristics in thermal images through our novel contour and edge-aware constraints to obtain a stable and anti-artifact disparity. We apply the commonly available color cameras to navigate the learning process of thermal image depth estimation frame-work. With our approach, an accurate depth map can be predicted without any prior knowledge under various illumination conditions. Yawen Lu, Guoyu Lu 0001 |
WACV | 2 |
| 2021 | 3D plant root system reconstruction based on fusion of deep structure-from-motion and IMU
Yawen Lu, Zhanjie Chen, Awais Khan 0005, Carl Salvaggio, Guoyu Lu 0001 |
Multim. Tools Appl. | 6 |
| 2021 | Simultaneous Direct Depth Estimation and Synthesis Stereo for Single Image Plant Root ReconstructionabstractPlant roots are the main conduit to its interaction with the physical and biological environment. A 3D root system architecture can provide fundamental and applied knowledge of a plant's ability to thrive, but the construction of 3D structures for thin and complicated plant roots is challenging. Existing methods such as structure-from-motion and shape-from-silhouette require multiple images, as input, under a complicated optimization process, which is usually not convenient in fieldwork. Little effort has been put into investigating the applications of deep neural network methods to reconstruct thin objects, like plant root systems, from a single image. We propose an unsupervised learning scheme to estimate the root depth from only one image as input, which is further applied to reconstruct the complete root system. The boundaries of the reconstructed object usually contain large errors, which is a significant problem for roots with many thin branches. To reduce reconstruction errors, we integrate a cross-view GAN-based network into the reconstruction process, which predicts the root image from a different perspective. Based on the predicted view, we reconstruct the root system using stereo reconstruction, which helps to identify the accurately reconstructed points by enforcing their consistency. The results on both the real plant root dataset and the synthetic dataset demonstrate the effectiveness of the proposed algorithm compared with state-of-the-art single image 3D reconstruction models on plant roots. Yawen Lu, Devarth Parikh, Awais Khan 0005, Guoyu Lu 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Extending Single Beam Lidar To Full Resolution By Fusing with Single Image Depth EstimationabstractDepth estimation is playing an important role in indoor and outdoor scene understanding, autonomous driving, augmented reality and many other tasks. Vehicles and robotics are able to use active illumination sensors such as LIDAR to receive high precision depth estimation. However, high-resolution LIDARs are usually too expensive, which limits its massive production on various applications. Though single beam LIDAR enjoys the benefits of low cost, one beam depth sensing is not usually sufficient to perceive the surrounding environment in many scenarios. In this paper, we propose a deep learning based framework to explore to replicate similar or even higher performance as costly LIDARs with our designed self-supervised network and a low-cost single-beam LIDAR. After the accurate calibration with a visible camera, the single beam LIDAR can adjust the scale uncertainty of the depth map estimated by the visible camera. The adjusted depth map enjoys the benefits of high resolution and sensing accuracy as high beam LIDAR and maintains low-cost as single beam LIDAR. Thus we can achieve similar sensing effect of high beam LIDAR with more than a 30-100 times cheaper price (e.g., $80000 Velodyne HDL-64E LIDAR v.s. $2000 SICK TIM-781 2D LIDAR and normal camera). The proposed approach is verified on our collected dataset and public dataset with superior depth-sensing performance. Yawen Lu, Devarth Parikh, Yuan Xin, Guoyu Lu 0001 |
ICPR | 5 |
| 2020 | Multi-Task Learning for Single Image Depth Estimation and Segmentation Based on Unsupervised NetworkabstractDeep neural networks have significantly enhanced the performance of various computer vision tasks, including single image depth estimation and image segmentation. However, most existing approaches handle them in supervised manners and require a large number of ground truth labels that consume extensive human efforts and are not always available in real scenarios. In this paper, we propose a novel framework to estimate disparity maps and segment images simultaneously by jointly training an encoder-decoder-based interactive convolutional neural network (CNN) for single image depth estimation and a multiple class CNN for image segmentation. Learning the neural network for one task can be beneficial from simultaneously learning from another one under a multi-task learning framework. We show that our proposed model can learn per-pixel depth regression and segmentation from just a single image input. Extensive experiments on available public datasets, including KITTI, Cityscapes urban, and PASCAL-VOC demonstrate the effectiveness of our model compared with other state-of-the-art methods for both tasks. Yawen Lu, Michel Sarkis, Guoyu Lu 0001 |
ICRA | 3 |
| 2020 | Single Image Shape-from-SilhouettesabstractRecovering a 3D shape representation from one single image input has been attempted in recent years. Most of the works obtain 3D models from multiple images at different perspectives or ground truth CAD models. However, multiple images from different perspectives or 3D CAD models are not always available in real applications. In this work, we present a novel shape-from-silhouette method based on just a single image, which is an end-to-end learning framework relying on view synthesis and shape-from-silhouette methodology to reconstruct a 3D shape. The reconstructed 3D mesh can approach the real shape of target objects by constraining the silhouettes from both horizontal and vertical directions, especially for those objects with occlusions. Our proposed method achieves state-of-the-art performance on the ShapeNet dataset compared with other recent approaches targeting 3D reconstruction from a single image. Without requiring labor-intensive and time-consuming human annotations, the work has a broad potential to be applied in real-world applications. Yawen Lu, Guoyu Lu 0001 |
ACM Multimedia | 3 |
| 2019 | 3D Shape Retrieval through Multilayer RBF Neural Networkabstract3D object retrieval involves more efforts mainly because major computer vision features are designed for 2D images, which is rarely applicable for 3D models. In this paper, we propose to retrieve the 3D models based on the implicit parameters learned from the radial base functions that represent the 3D objects. The radial base functions are learned from the RBF neural network. As deep neural networks can represent the data that is not linearly separable, we apply multiple layers' neural network to train the radial base functions. Our feature can be applied to recover the 3D objects, which proves the effectiveness of our features in representing the 3D objects. Furthermore, the dimensionality of the learned feature is scalable, which leads to memory efficiency. Experiments demonstrate the accuracy of our feature in 3D model retrieval. Guoyu Lu 0001, Yahong Han |
ICIP | 1 |
| 2019 | Deep Unsupervised Learning for Simultaneous Visual Odometry and Depth EstimationabstractVisual odometry and depth estimation are critical to understanding the scene and camera motion, which are particularly helpful to tasks such as scene understanding, autonomous driving, and robotics. Supervised learning methods have been applied in many deep neural network frameworks and demonstrated outstanding results in visual odometry and depth estimation. However, supervised learning requires a significant amount of labeled data for training, which consumes extensive time. In this paper, we explore an unsupervised learning framework that can learn a camera pose regressor from monocular video frames and estimates the scene depth simultaneously. The proposed method is able to perform accurate pose prediction as well as depth estimation, despite the absence of any ground truth data. The effectiveness of our proposed method is demonstrated through experiments on KITTI, Cityscapes, and Make3D benchmark datasets, which shows superb results compared with state-of-the-art methods in both tasks. Yawen Lu, Guoyu Lu 0001 |
ICIP | 2 |
| 2019 | Taking Me to the Correct Place: Vision-Based Localization for Autonomous VehiclesabstractVehicle localization is a critical component for autonomous driving, which estimates the position and orientation of vehicles. To achieve the goal of quick and accurate localization, we develop a system that can dynamically switch the features applied for localization. Specifically, we develop a feature based on convolutional neural network targeting at accurate matching, which proves high rotation invariant property that can help to overcome the relatively large error when vehicles turning at corners. However, when the vehicle motion mainly involves translation, we apply the ORB feature to localize the vehicle, as it demonstrates similar accuracy in translation estimation. Through dynamic switching between features, we can accurately localize the vehicle with high time performance. To filter out noise features, we train a CNN neural network to semantically understand the images and filter the features from moving objects and infinity position. During the pose estimation stage, we rely on the depth of the 3D points to identify the inliers satisfying RANSAC. Experiments demonstrate the superb performance of our method. Guoyu Lu 0001, Xue-Iuan Wong |
ICIP | 1 |
| 2019 | From Mapping to Localization: A Complete Framework to Visually Estimate Position and Attitude for Autonomous VehiclesabstractAutonomous vehicle framework relies on localization algorithms to position itself and navigates to the destination. In this paper, we explore a light-weight visual localization method to realize the vehicle position and attitude estimation based on images rather than the dominant LIDAR data. We apply SLAM and an offline map correction method to generate a high precision map, which composes 3D points and feature descriptors. For each image, we extract the features and match against the map to explore correspondences. In the correspondences search process, we rely on the previous camera pose estimation result to determine the search scope, where significantly improves the localization accuracy. The searching process is embedded in the pose estimation stage, which we adjust the PnP procedure to better fit the autonomous driving task. Simply based on a single CPU thread support, experiments on the benchmark KITTI dataset demonstrate the superior results of our method. Guoyu Lu 0001, Xue-Iuan Wong, James McBride |
ICIP | 1 |
| 2018 | Getting Rid of Night: Thermal Image Classification Based on Feature FusionabstractThermal images are essential to deal with situations in dark environments, as they capture the objects' temperature. While the objects can still be seen in thermal images, the texture is extremely blur or even not observable at all. We propose to extract different features from images that capture various characteristics of the images. As one feature emphasizes one distinguishing aspect differing from the others, we can grasp multiple pieces of evidence from the images and take advantage of each to improve the thermal image classification accuracy. In particular, in additional to corner features usually used in color images, we also extract features from the edges and the shapes of the objects that emphasize the integral image appearance, as well as the temperature characteristics obtained from the image intensity. In this way, even if one feature is not evident in an image, the others can still play a critical role towards the correct classification result. By optimizing the objective function, we maximize the fusion performance of multiple features. By doing so, we can to the largest extent make use of the information exhibited in the thermal images to classify the query image into the correct group. Experiments demonstrate promising thermal image classification result. Guoyu Lu 0001, Huili Yu, Chun Yuan 0003 |
ICPR | 1 |
| 2018 | 3D Image-based Indoor Localization Joint With WiFi PositioningabstractWe realize a system that utilizes WiFi to facilitate the image-based localization system, which avoids the confusion caused by the similar decoration inside the buildings. While WiFi-based localization thread obtains the rough location information, the image-based localization thread retrieves the best matching images and clusters the camera poses associated with the images into different location candidates. The image cluster closest to the WiFi localization outcome is selected for the exact camera pose estimation. The usage of WiFi significantly reduces the search scope, avoiding the extensive search of millions of descriptors in a 3D model. In the image-based localization stage, we also propose a novel 2D-to-2D-to-3D localization framework which follows a coarse-to-fine strategy to quickly locate the query image in several location candidates and performs the local feature matching and camera pose estimation after choosing the correct image location by WiFi positioning. The entire system demonstrates significant benefits in combining both images and WiFi signals in localization tasks and great potential to be deployed in real applications. Guoyu Lu 0001, Jingkuan Song |
ICMR | 1 |
| 2017 | Stromule branch tip detection based on accurate cell image segmentationabstractBased on the dynamic structure, we design a system that can perform accurate stromule image segmentation, branch tip detection and tracking automatically. We substitute the user constraints in active contour segmentation by spatial fuzzy c-means clustering for providing more precise segmentation result. Based on the segmented contour after smoothing, we create a surface normal based feature that can accurately detect the branch tips. We further combine normal information together with tip position coordinate to apply ICP to track the branch tips moving path. Guoyu Lu 0001, Jeffrey Caplan, Chandra Kambhamettu |
ICIP | 1 |
| 2017 | Indoor localization via multi-view images and videos
Guoyu Lu 0001, Yan Yan 0002, Nicu Sebe, Chandra Kambhamettu |
Comput. Vis. Image Underst. | 1 |
| 2017 | Large-Scale Tracking for Images With Few TexturesabstractImage tracking provides crucial insight for the image motion, which generates essential information for incremental structure-from-motion reconstruction and camera pose estimation. Typical usages, such as 3D reconstruction and visual odometry, all rely on robust and accurate local feature tracking through consecutive images. Current algorithms realize feature tracking through matching features extracted from discriminant textures in the images, for which distinctive image content is required to obtain accurate feature matching. For images with few textures, usually, an insufficient number of features are extracted to perform reliable tracking in a series of sequential images. We propose a method that makes use of a limited number of discriminate features to explore other features without strong discriminant power. We develop a feature integrating surrounding salient points distribution knowledge, raw pixel value, and coordinate information to discover a significant amount of features in weakly textured areas in an image. We also incorporate epipolar geometry in the feature correspondence calculation by taking the distance from the matching candidate to its corresponding point's epipolar line into account. To reduce the number of unreliable features, we project the estimated 3D points back to the images. The reprojection error is standardized according to the 3D point's depth, which reduces the bias introduced by the object distance to the camera. We conduct experiments on a large dataset of Arctic sea ice images, mainly composed by planes of ices and sea water. The experimental results demonstrate that our method can perform fast and accurate tracking in weakly textured images. Guoyu Lu 0001, Liqiang Nie, Scott Sorensen, Chandra Kambhamettu |
IEEE Trans. Multim. | 1 |
| 2016 | Neural network shape: Organ shape representation with radial basis function neural networksabstractWe propose to represent the shape of an organ using a neural network classifier. The shape is represented by a function learned by a neural network. Radial Basis Function (RBF) is used as the activation function for each perceptron. The learned implicit function is a combination of radial basis functions, which can represent complex shapes. The organ shape representation is learned using classification methods. Our testing results show that the neural network shape provides the best representation accuracy. The use of RBF provides a rotation, translation and scaling invariant feature to represent the shape. Experiments show that our method can accurately represent the organ shape. Guoyu Lu 0001, Abhishek Kolagunda, Chandra Kambhamettu |
ICASSP | 1 |
| 2016 | A Fast 3D Indoor-Localization Approach Based on Video Queries
Guoyu Lu 0001, Yan Yan 0002, Abhishek Kolagunda, Chandra Kambhamettu |
MMM (2) | 1 |
| 2016 | Where am I in the dark: Exploring active transfer learning on the use of indoor localization based on thermal imaging
Guoyu Lu 0001, Yan Yan 0002, Philip Saponaro, Nicu Sebe, Chandra Kambhamettu |
Neurocomputing | 1 |
| 2016 | Representing 3D shapes based on implicit surface functions learned from RBF neural networks
Guoyu Lu 0001, Abhishek Kolagunda, Xiaolong Wang 0006, Baris Turkbey, Peter L. Choyke, Chandra Kambhamettu |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | Active domain adaptation with noisy labels for multimedia analysis
Gaowen Liu, Yan Yan 0002, Subramanian Ramanathan, Jingkuan Song, Guoyu Lu 0001, Nicu Sebe |
World Wide Web | 5 |
| 2015 | Hierarchical Hybrid Shape Representation for Medical ShapesabstractRecently, shape analysis has become of increasing interest in the medical community due to its potential in capturing the morphological variations across a population. The high quality 3D images captured can be used to extract 3D shape of the organs. 3D models of organs can also be used for training personnel, for visualization during image guided interventions and in simulations. A compact shape model that has implicit and explicit forms will aid in some of these medical use-cases. We propose a compact hybrid shape model as a combination of Extended Superquadrics (ESQ) [1] and Radial basis interpolation function (RBF). The hybrid shape model in its parametric form is given as ( f (θ ,φ)= h(θ ,φ)+g(θ ,φ)). h is the extended superquadric function and g is radial basis interpolation function. The points on the surface of the shape are given by Abhishek Kolagunda, Guoyu Lu 0001, Chandra Kambhamettu |
BMVC | 2 |
| 2015 | Localize Me Anywhere, Anytime: A Multi-task Point-Retrieval ApproachabstractImage-based localization is an essential complement to GPS localization. Current image-based localization methods are based on either 2D-to-3D or 3D-to-2D to find the correspondences, which ignore the real scene geometric attributes. The main contribution of our paper is that we use a 3D model reconstructed by a short video as the query to realize 3D-to-3D localization under a multi-task point retrieval framework. Firstly, the use of a 3D model as the query enables us to efficiently select location candidates. Furthermore, the reconstruction of 3D model exploits the correlations among different images, based on the fact that images captured from different views for SfM share information through matching features. By exploring shared information (matching features) across multiple related tasks (images of the same scene captured from different views), the visual feature's view-invariance property can be improved in order to get to a higher point retrieval accuracy. More specifically, we use multi-task point retrieval framework to explore the relationship between descriptors and the 3D points, which extracts the discriminant points for more accurate 3D-to-3D correspondences retrieval. We further apply multi-task learning (MTL) retrieval approach on thermal images to prove that our MTL retrieval framework also provides superior performance for the thermal domain. This application is exceptionally helpful to cope with the localization problem in an environment with limited light sources. Guoyu Lu 0001, Yan Yan 0002, Jingkuan Song, Nicu Sebe, Chandra Kambhamettu |
ICCV | 1 |
| 2015 | Memory efficient large-scale image-based localization
Guoyu Lu 0001, Nicu Sebe, Congfu Xu, Chandra Kambhamettu |
Multim. Tools Appl. | 1 |
| 2014 | Knowing Where I Am: Exploiting Multi-Task Learning for Multi-view Indoor Image-based Localization
Guoyu Lu 0001, Yan Yan 0002, Nicu Sebe, Chandra Kambhamettu |
BMVC | 1 |
| 2014 | Structure-from-Motion reconstruction based on weighted Hamming descriptorsabstractWe propose a pipelined methods to reduce memory consumption of large-scale Structure-from-Motion reconstruction with the use of unsorted images extracted from photo collection websites. Recent research is able to reconstruct cities based on extracted images from photo collection websites. SIFT feature is used to find the correspondences between two images. For the large-scale reconstruction with unsorted images, the system needs to store all the descriptors and feature points information in memory to search for correspondences. As each SIFT descriptor is a 128 dimensional real-value vector, storing all the descriptors would consume a significant amount of memory. Based on this limitation, we project the high dimensional features into a low-dimensional space using a learned projection matrix. After projection, the distance of the descriptors belonging to the same point in 3D space is decreased; the distance of the descriptors belonging to the different points is increased. Furthermore, we learn a mapping function, which maps the real-value descriptor into binary code. As Hamming descriptors contain only two value options per bit and the length of the descriptor is limited, there are usually multiple descriptors having the same Hamming distance to the query descriptor. In dealing with this problem, we give different weights to each dimension and rank each bit of the Hamming descriptor based on each dimensions discriminant power; this contributes to reduce the ambiguity in matching the descriptors. The experiments show that our method achieves dense reconstruction results with less than 10 percent of the original memory consumption. Guoyu Lu 0001, Vincent Ly, Chandra Kambhamettu |
IJCNN | 1 |
| 2013 | Can We Minimize the Influence Due to Gender and Race in Age Estimation?abstractAutomatic human age estimation has attracted a great deal of interest in the past few years. Although many advancements have been made by researchers, there are still many challenges: such as age estimation across different image acquisition methods, different expressions, gender and races. The influence due to race and gender seems to be the most common issue, because collecting a large amount of face images with comprehensive racial diversities seems impractical. The performance will degrade when estimating face images of races that differ from the training set. In this work, we present a new scheme to mitigate the influences of race and gender in the problem of age estimation. Our system will contribute a robust solution to solve the problem of age estimation across races and genders. This study is essential for developing a practical age estimation system (with mixture of races and gender.) To evaluate the performance of the proposed algorithm, we run comprehensive experiments on one widely used big database - MORPH-II, which contains more than 55, 000 images. On an average, an improvement of more than 20% has been achieved using the proposed scheme. Xiaolong Wang 0006, Vincent Ly, Guoyu Lu 0001, Chandra Kambhamettu |
ICMLA (2) | 3 |
| 2013 | Large-scale Structure-from-Motion Reconstruction with small memory consumptionabstractStructure-from-Motion reconstruction is to recover the 3 dimensional structure from 2 dimensional images. Recent research in this field demonstrates the ability to reconstruct cities based on images extracted from a photo collection website; SIFT feature is typically extracted to detect correspondences between images. For the reconstruction of large scale unsorted images, the system is required to store all features and points information in the memory to search for correspondences. As SIFT feature is a 128 dimensional real-valued vector, storing each descriptor would consume a significant amount of memory. Due to this limitation, we propose to project the high-dimensional feature into a lower-dimensional space by using a new learned projection matrix while still maintaining the property of the original features. Hence, the result of this projection will shorten the distance among descriptors of the same point while lengthening the distance among descriptors of different points. These projected descriptors use Hellinger distance for calculation of the similarity between features. Furthermore, we learn a mapping function, which will map the real-valued descriptor into binary code coping with the variation of correspondence searching method. Experiments demonstrate that our method achieve excellent results with limited memory requirement. Guoyu Lu 0001, Vincent Ly, Chandra Kambhamettu |
MoMM | 1 |
| 2012 | Evaluating SPAN Incremental Learning for Handwritten Digit Recognition
Ammar Mohemmed, Guoyu Lu 0001, Nikola K. Kasabov |
ICONIP (3) | 2 |