Yonggen Ling

dblp:139/7117 · DBLP profile ↗
← Back
34ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0001-8294-6286ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 5 since 2021Systems, architecture and hardware · 7 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Learning Fine-Grained Alignment for Aerial Vision-Dialog Navigation
abstract
Aerial Vision-Dialog Navigation (AVDN) is a new task that requires drones to navigate to a target location based on human-robot dialog history. This paper focuses on the critical fine-grained cross-modal alignment problem in AVDN, requiring the drone to align language entities with visual landmarks in top-down views. To achieve this, we first construct a Fine-Grained AVDN (FG-AVDN) dataset via a semi-automatic annotation pipeline, providing diverse multimodal annotations at the entity-landmark level. Based on this, a novel Fine-grained Entity-Landmark Alignment (FELA) method is proposed to learn the cross-modal alignment explicitly. Concretely, FELA first boosts the drone's visual understanding with a precise semantic grid representation, which captures the environmental semantics and spatial structure simultaneously. Subsequently, to learn the entity-landmark alignment, we devise cross-modal auxiliary tasks from three perspectives, including grounding, captioning, and contrastive learning. Extensive experiments demonstrate that our explicit entity-landmark alignment learning is beneficial for AVDN. As a result, FELA achieves leading performance with 3.2% SR and 4.9% GP improvements over prior arts. Code and dataset will be publicly available.
Yifei Su, Dong An 0002, Weichen Yu, Baiyang Ning, Yonggen Ling, Yan Huang 0008, Liang Wang 0001
AAAI6
2025 Leveraging Global Stereo Consistency for Category-Level Shape and 6D Pose Estimation from Stereo Images
abstract
Stereo-based category-level shape and 6D pose estimation methods have the potential to generalize to a wider range of materials than RGBD methods, which often suffer from depth measurement errors. However, without explicit depth from two views, parameters to be estimated can become inherently entangled, negatively impacting performance. To address this, we propose a method that leverages global stereo consistency to constrain optimization directions and mitigate parameter entanglement. We first estimate an intra-category occupancy field to represent a unified shape across views, ensuring consistency and preventing shape ambiguity. Through a divide-and-conquer approach within global shape fitting, we fit this shape to stereo images to obtain the pose, iteratively rendering normalized depth maps and exchanging information across views. This approach improves convergence toward the correct pose and scale. We validated our method on both depth-friendly and depth-challenging materials using our S-RGBD dataset and the TOD benchmark. Our method surpasses RGBD methods on challenging objects and performs comparably on depth-friendly ones. Ablation studies confirm the effectiveness of each component.
Junning Qiu, Minglei Lu, Yonggen Ling
CVPR5
2025 Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
abstract
We address the task of Vision-Language Navigation in Continuous Environments (VLN-CE) under the zero-shot setting. Zero-shot VLN-CE is particularly challenging due to the absence of expert demonstrations for training and minimal environment structural prior to guide navigation. To confront these challenges, we propose a Constraint-Aware Navigator (CA-Nav), which reframes zero-shot VLN-CE as a sequential, constraint-aware sub-instruction completion process. CA-Nav continuously translates sub-instructions into navigation plans using two core modules: the Constraint-Aware Sub-instruction Manager (CSM) and the Constraint-Aware Value Mapper (CVM). CSM defines the completion criteria for decomposed sub-instructions as constraints and tracks navigation progress by switching sub-instructions in a constraint-aware manner. CVM, guided by CSM's constraints, generates a value map on the fly and refines it using superpixel clustering to improve navigation stability. CA-Nav achieves the state-of-the-art performance on two VLN-CE benchmarks, surpassing the previous best method by 12% and 13% in Success Rate on the validation unseen splits of R2R-CE and RxR-CE, respectively. Moreover, CA-Nav demonstrates its effectiveness in real-world robot deployments across various indoor scenes and instructions.
Dong An 0002, Yan Huang 0008, Rongtao Xu, Yifei Su, Yonggen Ling, Ian D. Reid 0001, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Category-Level Object Detection, Pose Estimation and Reconstruction from Stereo Images
Chuanrui Zhang, Yonggen Ling, Minglei Lu, Minghan Qin, Haoqian Wang
ECCV (34)2
2024 VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception
abstract
This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the "Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D, the first extensive multi-modal dataset integrating vision, touch, and proprioception, to enhance robotic manipulation. VinT-6D comprises 2 million VinT-Sim and 0.1 million VinT-Real entries, collected via simulations in Mujoco and Blender and a custom-designed real-world platform. This dataset is tailored for robotic hands, offering models with whole-hand tactile perception and high-quality, well-aligned data. To the best of our knowledge, the VinT-Real is the largest considering the collection difficulties in the real-world environment so it can bridge the gap of simulation to real compared to the previous works. Built upon VinT-6D, we present a benchmark method that shows significant improvements in performance by fusing multi-modal information. The project is available at https://VinT-6D.github.io/.
Zhaoliang Wan, Yonggen Ling, Senlin Yi, Lu Qi 0001, Wang Wei Lee, Minglei Lu, Xiao Teng, Xu Yang 0004, Ming-Hsuan Yang 0001, Hui Cheng 0002
ICML2
2024 Deep Fusion for Multi-Modal 6D Pose Estimation
abstract
6D pose estimation with individual modality encounters difficulties due to the limitations of modalities, such as RGB information on textureless objects and depth on reflective objects. This can be improved by exploiting the complementarity between modalities. Most of the previous methods only consider the correspondence between point clouds and RGB images and directly extract the features of the corresponding two modalities for fusion, which ignore the information of the modality itself and are negatively affected by erroneous background information when introducing more features for fusion. To enhance the complementarities between multiple modalities, we propose a neighbor-based cross-modalities attention mechanism for multi-modal 6D pose estimation. Neighbors represent that the RGB features of multiple neighbor are applied for fusion, which expands the receptive field. The cross-modalities attention mechanism leverages the similarities between the different modal features to help modal feature fusion, which reduces the negative impact of incorrect background information. Moreover, we design some features between the rendered image and the original image to obtain the confidence of pose estimation results. Experimental results on LM, LM-O and YCB-V datasets demonstrate the effectiveness of our methods. Video is available at https://www.youtube.com/watch?v=ApNBcX6NEGs.Note to Practitioners—Introducing the information of surrounding points during multi-modal fusion improves the performance of 6D pose estimation. For example, the RGB image corresponding to some point clouds on the object may lack rich texture features while the neighbors exist. However, most methods of modal fusion based on RGBD for 6D pose estimation only simply consider the corresponding between RGB images and point clouds for feature fusion, which may bring redundant information or the wrong background information when introducing neighbor information. In this paper, we propose a cross-modal attention mechanism based on neighbor information. By introducing the information of the modality itself to obtain the weight of the neighbor information of another modality in the encoding and decoding stages, the receptive field is expanded and the complementarities between different modalities are enhanced. The experiment shows our effectiveness. In addition, we provide a pose confidence estimator for predicted pose results. Specifically, the rendered image with the predicted pose and the real image are applied to extract features for the decision tree. The experimental results show that the result of the wrong estimation can be eliminated with high accuracy and recall. The 6D pose confidence can provide a reference for real-world grasping. However, the current method can only estimate objects with known models. In the future, we will consider applying the method to unseen objects.
Shifeng Lin, Zunran Wang, Shenghao Zhang 0001, Yonggen Ling, Chenguang Yang 0001
IEEE Trans Autom. Sci. Eng.4
2023 Mx2M: Masked Cross-Modality Modeling in Domain Adaptation for 3D Semantic Segmentation
abstract
Existing methods of cross-modal domain adaptation for 3D semantic segmentation predict results only via 2D-3D complementarity that is obtained by cross-modal feature matching. However, as lacking supervision in the target domain, the complementarity is not always reliable. The results are not ideal when the domain gap is large. To solve the problem of lacking supervision, we introduce masked modeling into this task and propose a method Mx2M, which utilizes masked cross-modality modeling to reduce the large domain gap. Our Mx2M contains two components. One is the core solution, cross-modal removal and prediction (xMRP), which makes the Mx2M adapt to various scenarios and provides cross-modal self-supervision. The other is a new way of cross-modal feature matching, the dynamic cross-modal filter (DxMF) that ensures the whole method dynamically uses more suitable 2D-3D complementarity. Evaluation of the Mx2M on three DA scenarios, including Day/Night, USA/Singapore, and A2D2/SemanticKITTI, brings large improvements over previous methods on many metrics.
Boxiang Zhang, Zunran Wang, Yonggen Ling, Yuanyuan Guan, Shenghao Zhang 0001, Wenhui Li 0002
AAAI3
2023 A Miniaturised Camera-based Multi-Modal Tactile Sensor
abstract
In conjunction with huge recent progress in cam-era and computer vision technology, camera-based sensors have increasingly shown considerable promise in relation to tactile sensing. In comparison to competing technologies (be they resistive, capacitive or magnetic based), they offer super-high-resolution, while suffering from fewer wiring problems. The human tactile system is composed of various types of mechanoreceptors, each able to perceive and process distinct information such as force, pressure, texture, etc. Camera-based tactile sensors such as GelSight mainly focus on high-resolution geometric sensing on a flat surface, and their force measurement capabilities are limited by the hysteresis and non-linearity of the silicone material. In this paper, we present a miniaturised dome-shaped camera-based tactile sensor that allows accurate force and tactile sensing in a single coherent system. The key novelty of the sensor design is as follows. First, we demonstrate how to build a smooth silicone hemispheric sensing medium with uniform markers on its curved surface. Second, we enhance the illumination of the rounded silicone with diffused LEDs. Third, we construct a force-sensitive mechanical structure in a compact form factor with usage of springs to accurately perceive forces. Our multi-modal sensor is able to acquire tactile information from multi-axis forces, local force distribution, and contact geometry, all in real-time. We apply an end-to-end deep learning method to process all the information.
Kaspar Althoefer, Yonggen Ling, Wanlin Li, Xinyuan Qian 0001, Wang Wei Lee, Peng Qi 0001
ICRA2
2023 ShuffleTrans: Patch-wise weight shuffle for transparent object segmentation
Boxiang Zhang, Zunran Wang, Yonggen Ling, Yuanyuan Guan, Shenghao Zhang 0001, Wenhui Li 0002, Lei Wei 0002, Chunxu Zhang
Neural Networks3
2022 HVC-Net: Unifying Homography, Visibility, and Confidence Learning for Planar Object Tracking
Haoxian Zhang, Yonggen Ling
ECCV (22)2
2022 Multi-fingered Tactile Servoing for Grasping Adjustment under Partial Observation
abstract
Grasping of objects using multi-fingered robotic hands often fails due to small uncertainties in the hand motion control and the object's pose estimation. To tackle this problem, we propose a grasping adjustment strategy based on tactile seroving. Our technique employs feedback from a sensorized multi-fingered robotic hand to collaboratively servo the fingers and palm to achieve the desired grasp. We demonstrate the performance of our method through simulation and physical experiments by having a robot grasp different objects under conditions of variable uncertainty. The results show that our approach achieved a higher success rate and tolerated greater uncertainty than an open-looped grasp.
Hanzhong Liu, Bidan Huang, Qiang Li 0001, Yu Zheng 0001, Yonggen Ling, Wang Wei Lee, Yi Liu 0068, Ya-Yen Tsai, Chenguang Yang 0001
IROS5
2022 Few-shot font style transfer with multiple style encoders
Yonglin Wu, Yonggen Ling, Lingyun Sun, Yingming Li
Sci. China Inf. Sci.5
2022 Unsupervised Occlusion-Aware Stereo Matching With Directed Disparity Smoothing
abstract
When handling occlusion in unsupervised stereo matching, existing methods tend to neglect the supportive role of occlusion and to perform inappropriate disparity smoothing around the occlusion. To address these problems, we propose an occlusion-aware stereo network that contains a specific module to first estimate occlusion as an additional depth cue. In the occlusion inference module, a pixel is classified with a three-category label based on whether an area is occluded by an object on the left, occluded by an object on the right, or unoccluded. After the occluders are detected, we introduce a directed disparity smoothing loss that allows valid disparity estimates to be propagated to fill the occluded region, while ambiguous matches in the occluded region do not affect other regions. Disparity and occlusion are trained alternately in an unsupervised manner with detached backpropagation to enable the directed smoothness. Experiments show that our method achieves 3-pixel threshold error rates of 6.51% and 5.69% on the KITTI 2015 and KITTI 2012 validation sets, state-of-the-art results among unsupervised learning networks at the time of submission.
Ang Li 0025, Zejian Yuan, Yonggen Ling, Wanchao Chi, Shenghao Zhang 0001
IEEE Trans. Intell. Transp. Syst.3
2020 Domain Adaptation Gaze Estimation by Embedding with Prediction Consistency
Zidong Guo, Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001
ACCV (5)5
2020 Learning End-to-End Action Interaction by Paired-Embedding Data Augmentation
Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001
ACCV (6)5
2020 FastCompletion: A Cascade Network with Multiscale Group-Fused Inputs for Real-Time Depth Completion
abstract
Completing sparse data captured with commercial depth sensors is a vital and fundamental procedure for many computer vision applications. For execution in real-world scenarios, a good trade-off between accuracy and speed is increasingly in demand for depth completion methods. Most previous methods achieve satisfactory accuracy on standard benchmarks. However, they extensively rely on heavy models to handle diverse structures and require additional run time on multimodal data. In this paper, we present an efficient method of depth completion. We propose a grouped fusion strategy for efficiently extracting depth and guidance features in parallel and fusing them naturally in the feature spaces to achieve high performance. Instead of a monolithic architecture, we employ cascaded hourglass networks, each of which is specialized for certain structures and has a lightweight architecture. Given the sparsity of the depth maps, we downsample the inputs to multiple scales to further accelerate the computation. Our model runs at over 39 FPS on an embedded GPU with high-resolution inputs. Evaluations on the KITTI benchmark demonstrate that the proposed model is an ideal approach for real-world applications.
Ang Li 0025, Zejian Yuan, Yonggen Ling, Wanchao Chi, Shenghao Zhang 0001
ICPR3
2020 Attention-Oriented Action Recognition for Real- Time Human-Robot Interaction
abstract
Despite the notable progress made in action recognition tasks, not much work has been done in action recognition specifically for human-robot interaction. In this paper, we deeply explore the characteristics of the action recognition task in interaction scenes and propose an attention-oriented multi-level network framework to meet the need for real-time interaction. Specifically, a Pre-Attention network is employed to roughly focus on the interactor in the scene at low resolution firstly and then perform fine-grained pose estimation at high resolution. The other compact CNN receives the extracted skeleton sequence as input for action recognition, utilizing attention-like mechanisms to capture local spatial-temporal patterns and global semantic information effectively. To evaluate our approach, we construct a new action dataset specially for the recognition task in interaction scenes. Experimental results on our dataset and high efficiency (112 fps at 640 × 480 RGBD) on the mobile computing platform (Nvidia Jetson AGX Xavier) demonstrate excellent applicability of our method on action recognition in real-time human-robot interaction.
Ziyi Yin 0001, Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001
ICPR6
2020 A Multi-Scale Guided Cascade Hourglass Network for Depth Completion
abstract
Depth completion, a task to estimate the dense depth map from sparse measurement under the guidance from the high-resolution image, is essential to many computer vision applications. Most previous methods building on fully convolutional networks can not handle diverse patterns in the depth map efficiently and effectively. We propose a multi-scale guided cascade hourglass network to tackle this problem. Structures at different levels are captured by specialized hourglasses in the cascade network with sparse inputs in various sizes. An encoder extracts multi-scale features from color image to provide deep guidance for all the hourglasses. A multi-scale training strategy further activates the effect of cascade stages. With the role of each sub-module divided explicitly, we can implement components with simple architectures. Extensive experiments show that our lightweight model achieves competitive results compared with state-of-the-art in KITTI depth completion benchmark, with low complexity in run-time.
Ang Li 0025, Zejian Yuan, Yonggen Ling, Wanchao Chi, Shenghao Zhang 0001
WACV3
2020 Crowded Human Detection via an Anchor-pair Network
abstract
This paper presents an anchor-pair network for crowded human detection, which can overcome and solve the difficulties caused by occlusion in crowded scenes. Specifically, we use a function-aware network structure to extract more distinctive and discriminative features for head and full-body respectively, and then a CNN module is also exploited to fuse the features by learning the correlations between head and full-body to reduce crowd errors. Meanwhile, a novel paired form for anchors, denoted as anchor-pair, is proposed to estimate the head regions and full-body regions simultaneously. Furthermore, a new ingenious Joint-NMS is introduced to perform on the detected head and full-body box pairs, which produces significant performance improvement in heavily occluded scenarios at tiny computational cost. Our anchor-pair network achieves a state-of-the-art result on the CrowdHuman dataset which reduces the MR−2to 55.43%, achieving 11.59% relative improvement over our dataset baseline.
Jinguo Zhu, Zejian Yuan, Wanchao Chi, Yonggen Ling, Shenghao Zhang 0001
WACV5
2020 Self-Supervised Learning of Detailed 3D Face Reconstruction
abstract
In this paper, we present an end-to-end learning framework for detailed 3D face reconstruction from a single image1. Our approach uses a 3DMM-based coarse model and a displacement map in UV-space to represent a 3D face. Unlike previous work addressing the problem, our learning framework does not require supervision of surrogate ground-truth 3D models computed with traditional approaches. Instead, we utilize the input image itself as supervision during learning. In the first stage, we combine a photometric loss and a facial perceptual loss between the input face and the rendered face, to regress a 3DMM-based coarse model. In the second stage, both the input image and the regressed texture of the coarse model are unwrapped into UV-space, and then sent through an image-toimage translation network to predict a displacement map in UVspace. The displacement map and the coarse model are used to render a final detailed face, which again can be compared with the original input image to serve as a photometric loss for the second stage. The advantage of learning displacement map in UV-space is that face alignment can be explicitly done during the unwrapping, thus facial details are easier to learn from large amount of data. Extensive experiments demonstrate the superiority of the proposed method over previous work.
Fanzi Wu, Yibing Song, Yonggen Ling, Linchao Bao
IEEE Trans. Image Process.5
2019 MVF-Net: Multi-View 3D Face Morphable Model Regression
abstract
We address the problem of recovering the 3D geometry of a human face from a set of facial images in multiple views. While recent studies have shown impressive progress in 3D Morphable Model (3DMM) based facial reconstruction, the settings are mostly restricted to a single view. There is an inherent drawback in the single-view setting: the lack of reliable 3D constraints can cause unresolvable ambiguities. We in this paper explore 3DMM-based shape recovery in a different setting, where a set of multi-view facial images are given as input. A novel approach is proposed to regress 3DMM parameters from multi-view inputs with an end-to-end trainable Convolutional Neural Network (CNN). Multi-view geometric constraints are incorporated into the network by establishing dense correspondences between different views leveraging a novel self-supervised view alignment loss. The main ingredient of the view alignment loss is a differentiable dense optical flow estimator that can backpropagate the alignment errors between an input view and a synthetic rendering from another input view, which is projected to the target view through the 3D shape to be inferred. Through minimizing the view alignment loss, better 3D shapes can be recovered such that the synthetic projections from one view to another can better align with the observed image. Extensive experiments demonstrate the superiority of the proposed method over other 3DMM methods.
Fanzi Wu, Linchao Bao, Yonggen Ling, Yibing Song, Songnan Li, King Ngi Ngan, Wei Liu 0005
CVPR4
2018 Left-Right Comparative Recurrent Model for Stereo Matching
abstract
Leveraging the disparity information from both left and right views is crucial for stereo disparity estimation. Left-right consistency check is an effective way to enhance the disparity estimation by referring to the information from the opposite view. However, the conventional left-right consistency check is an isolated post-processing step and heavily hand-crafted. This paper proposes a novel left-right comparative recurrent model to perform left-right consistency checking jointly with disparity estimation. At each recurrent step, the model produces disparity results for both views, and then performs online left-right comparison to identify the mismatched regions which may probably contain erroneously labeled pixels. A soft attention mechanism is introduced, which employs the learned error maps for better guiding the model to selectively focus on refining the unreliable regions at the next recurrent step. In this way, the generated disparity maps are progressively improved by the proposed recurrent model. Extensive evaluations on KITTI 2015, Scene Flow and Middlebury benchmarks validate the effectiveness of our model, demonstrating that state-of-the-art stereo disparity estimation results can be achieved by this new model.
Zequn Jie, Pengfei Wang 0011, Yonggen Ling, Bo Zhao 0032, Yunchao Wei, Jiashi Feng, Wei Liu 0005
CVPR3
2018 Modeling Varying Camera-IMU Time Offset in Optimization-Based Visual-Inertial Odometry
Yonggen Ling, Linchao Bao, Zequn Jie, Fengming Zhu, Shanmin Tang, Wei Liu 0005, Tong Zhang 0001
ECCV (9)1
2018 Probabilistic Dense Reconstruction from a Moving Camera
abstract
This paper presents a probabilistic approach for online dense reconstruction using a single monocular camera moving through the environment. Compared to spatial stereo, depth estimation from motion stereo is challenging due to insufficient parallaxes, visual scale changes, pose errors, etc. We utilize both the spatial and temporal correlations of consecutive depth estimates to increase the robustness and accuracy of monocular depth estimation. An online, recursive, probabilistic scheme to compute depth estimates, with corresponding covariances and inlier probability expectations, is proposed in this work. We integrate the obtained depth hypotheses into dense 3D models in an uncertainty-aware way. We show the effectiveness and efficiency of our proposed approach by comparing it with state-of-the-art methods in the TUM RGB-D SLAM & ICL-NUIM dataset. Online indoor and outdoor experiments are also presented for performance demonstration.
Yonggen Ling, Shaojie Shen
IROS1
2017 Building maps for autonomous navigation using sparse visual SLAM features
abstract
Autonomous navigation, which consists of a systematic integration of localization, mapping, motion planning and control, is the core capability of mobile robotic systems. However, most research considers only isolated technical modules. There exist significant gaps between maps generated by SLAM algorithms and maps required for motion planning. This paper presents a complete online system that consists in three modules: incremental SLAM, real-time dense mapping, and free space extraction. The obtained free-space volume (i.e. a tessellation of tetrahedra) can be served as regular geometric constraints for motion planning. Our system runs in real-time thanks to the engineering decisions proposed to increase the system efficiency. We conduct extensive experiments on the KITTI dataset to demonstrate the run-time performance. Qualitative and quantitative results on mapping accuracy are also shown. For the benefit of the community, we make the source code public.
Yonggen Ling, Shaojie Shen
IROS1
2016 Aggressive quadrotor flight using dense visual-inertial fusion
abstract
In this work, we address the problem of aggressive flight of a quadrotor aerial vehicle using cameras and IMUs as the only sensing modalities. We present a fully integrated quadrotor system and demonstrate through online experiment the capability of autonomous flight with linear velocities up to 4.2 m/s, linear accelerations up to 9.6 m/s2, and angular velocities up to 245.1 degree/s. Central to our approach is a dense visual-inertial state estimator for reliable tracking of aggressive motions. An uncertainty-aware direct dense visual tracking module provides camera pose tracking that takes inverse depth uncertainty into account and is resistant to motion blur. Measurements from IMU pre-integration and multi-constrained dense visual tracking are fused probabilistically using an optimization-based sensor fusion framework. Extensive statistical analysis and comparison are presented to verify the performance of the proposed approach. We also release our code as open-source ROS packages.
Yonggen Ling, Tianbo Liu 0001, Shaojie Shen
ICRA1
2016 High-precision online markerless stereo extrinsic calibration
abstract
Stereo cameras and dense stereo matching algorithms are core components for many robotic applications due to their abilities to directly obtain dense depth measurements and their robustness against changes in lighting conditions. However, the performance of dense depth estimation relies heavily on accurate stereo extrinsic calibration. In this work, we present a real-time markerless approach for obtaining high-precision stereo extrinsic calibration using a novel 5-DOF (degrees-of-freedom) and nonlinear optimization on a manifold, which captures the observability property of vision-only stereo calibration. Our method minimizes epipolar errors between spatial per-frame sparse natural features. It does not require temporal feature correspondences, making it not only invariant to dynamic scenes and illumination changes, but also able to run significantly faster than standard bundle adjustment-based approaches. We introduce a principled method to determine if the calibration converges to the required level of accuracy, and show through online experiments that our approach achieves a level of accuracy that is comparable to offline marker-based calibration methods. Our method refines stereo extrinsic to the accuracy that is sufficient for block matching-based dense disparity computation. It provides a cost-effective way to improve the reliability of stereo vision systems for long-term autonomy.
Yonggen Ling, Shaojie Shen
IROS1
2015 Image colorization via color propagation and rank minimization
abstract
Image colorization aims to add colors to grayscale images, which used to be a time-consuming and tedious task that requires lots of human efforts. In this paper, we present a novel colorization method based on color propagation and rank minimization. Given a small portion of chrominance values and a grayscale image, we firstly propagate the known color values to other pixels to be colorized. As the colorized image after color propagation is not accurate, we then define a confidence matrix to measure the propagation fidelity. Finally, pixels that have propagated chrominance values with confidence are colorized by rank minimization, which exploits the redundancy of natural images. Experimental results on real data set show that our proposed method achieves state-of-the-art colorization quality.
Yonggen Ling, Oscar C. Au, Jiahao Pang, Jin Zeng 0004, Yuan Yuan 0002, Amin Zheng
ICIP1
2014 Analysis of sampling pattern and Luma-Chroma filter design for subpixel-based image downsampling
abstract
Subpixel-based image downsampling is attractive in that it produces higher apparent resolution of down-sampled images on LCD displays. However increased luminance resolution is achieved at the price of color fringing artifacts. In this paper, we propose an algorithm to find a pleasing balance between increased resolution and color fidelity. We separate the subpixel-based downsampling into two stages, shifting followed by downsampling with anti-aliasing filtering. In stage one, we find special characteristics of the luminance and chrominance spectra of the shifted image, based on which the optimal sampling pattern is found. In stage two, anti-aliasing filters for luminance and chrominance are designed respectively. Experimental results verify that the proposed method manages to suppress color artifacts while maintaining high luminance sharpness.
Jin Zeng 0004, Oscar C. Au, Yuanfang Guo, Jiahao Pang, Ketan Tang, Yonggen Ling
ICASSP6
2014 Self-similarity-based image colorization
abstract
In this work, we tackle the problem of coloring black-and-white images, which is image colorization. Existing image colorization algorithms can be categorized into two types: scribble-based colorization algorithms and example-based colorization algorithms. Differently, we propose a hybrid scheme that combines the advantages of both categories. Given the grayscale image to be colorized and a few color scribbles (or scattered color labels) as input, the proposed method manages to colorize the grayscale image with high quality. Similar to the mechanisms in example-based colorization methods, our algorithm firstly propagates chrominance information based on the assumption that similar image patches should have similar colors. Therefore colors of some pixels can be transferred from similar patches with known colors. After that, we apply scribble-based colorization algorithm to fully colorize the grayscale image, with different confidences assigned onto the transferred color labels. Experimental results show that, the proposed method effectively utilizes the known chrominance, and provides pleasant colorizations with very few user interventions.
Jiahao Pang, Oscar C. Au, Yukihiko Yamashita, Yonggen Ling, Yuanfang Guo, Jin Zeng 0004
ICIP4
2014 Natural image matting via adaptive local and nonlocal sample clustering
abstract
Digital image matting is the determination of foreground color, background color, and an opacity value of each pixel for an input image. Inherently, matting is a highly ill-posed and under-constrained problem. Thus, some assumptions need to be made to resolve it. Inspired by closed-form matting and color clustering matting, in this work, we first develop an adaptive sample clustering criterion to automatically assign either local or nonlocal neighborhood to each pixel. After that, in order to enhance matting accuracy, we improve the nonlocal clustering performance by introducing a new feature selection parameter to choose preferred feature space for different images in a fully automatic way. And finally we solve the problem using a closed form solution. Experimental results show that our algorithm achieves equal or even better performance among many state-of-the-art matting techniques.
Oscar C. Au, Yuan Yuan 0002, Wenxiu Sun, Yonggen Ling, Jiahao Pang
ICIP5
2014 Intra prediction with adaptive CU processing order in HEVC
abstract
The High Efficiency Video Coding (HEVC) utilizes Z-scan order to process coding units (CUs). For intra prediction, this order cannot fully exploit the spatial correlation between adjacent CUs. After transform and quantization, the residue still contains lots of energy along edges which consumes many bits for compression. To effectively reduce the residue energy along edges, a novel intra prediction approach is proposed, where the CU processing order is changed adaptively. Two additional orders are introduced in this paper besides traditional Z-scan order. Up to 1.9% bit saving is achieved in our experiments on HEVC test model. We also propose two fast order selection algorithms and the observed gains are obtained with 27% and 2% encoding time increase compared to HEVC, respectively.
Amin Zheng, Oscar C. Au, Yuan Yuan 0002, Haitao Yang 0001, Jiahao Pang, Yonggen Ling
ICIP6
2014 Photo album compression By leveraging temporal-spatial correlations and HEVC
abstract
The advancing digital photography technology has resulted in a large number of photos stored in personal computers. Photo album compression algorithms aim to save storage space and efficiently manage photos. In this paper, a general forest structure model involving depth constrain for photo album compression is proposed, which further exploits the correlations between images in the photo album. We firstly represent the images as nodes in a graph and directed edges between them as predictive coding relationship. Affinity propagation is then applied to compute for a depth-constrained forest. Finally, we adopt depth-first search algorithm to generate the compression order according to forest structure and HEVC to compress the images with adaptive GOPs and reference list. Experimental results show that the proposed compression method provides much better rate-distortion performance compared to JPEG and significantly reduce the storage space.
Yonggen Ling, Oscar C. Au, Ruobing Zou, Jiahao Pang, Amin Zheng
ISCAS1
2013 An analytical study of subpixel-based image down-sampling patterns in frequency domain
abstract
Subpixel-based image down-sampling is a class of methods that can provide improved apparent resolution of the down-scaled image compared to the pixel-based methods. The frequency characteristics of all possible subpixel-based down-sampling patterns for RGB vertical stripes are analytically studied in this paper. Our proposed algorithm reveals that there are merely seven equivalent energy distributions in the luminance frequency spectrum. To achieve higher luminance resolution, we then calculate and choose the optimal down-sampling pattern with anti-aliasing low-pass filter designed for it so as to maximize the energy of the luminance component within the cut-off shape. Experimental results show that the proposed method provides sharper images compared to the state-of-art subpixel-based methods, with little color distortion.
Yonggen Ling, Oscar C. Au, Ketan Tang, Jiahao Pang, Jin Zeng 0004, Lu Fang 0001
VCIP1