EDBT 2026 Demo / reviewers in the wild / expert
Hongdong Li
dblp:59/4859
· DBLP profile ↗
280ranked-venue papers
19as first author
125since 2021 · last 2026
0000-0003-4125-1554ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 199 · 13 first-author · 89 since 2021Graphics, computer vision, multimedia, augmented reality and games · 186 · 15 first-author · 76 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 12 since 2021Systems, architecture and hardware · 13 · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning High-Fidelity Garment Deformation via Skinning-Free Image TransferabstractWe present a novel method for generating 3D garment deformations from underlying body poses, which is key to a wide range of applications, including virtual try-on and extended reality. To simplify the cloth dynamics, existing methods mostly rely on linear blend skinning to obtain low-frequency posed garment shape and only regress highfrequency wrinkles. However, due to the lack of explicit skinning supervision, such skinning-based approach often produces misaligned shapes when posing the garment, consequently corrupts the high-frequency signals and fails to recover high-fidelity wrinkles. To tackle this issue, we propose a skinning-free approach by independently estimating posed (i) vertex position for low-frequency posed garment shape, and (ii) vertex normal for high-frequency local wrinkle details. In this way, each frequency modality can be effectively decoupled and directly supervised by the geometry of the deformed garment. To further improve the visual quality of deformation, we propose to encode both vertex attributes as rendered texture images, so that 3D garment deformation can be equivalently achieved via 2D image transfer. This enables us to leverage powerful pretrained image models to recover fine-grained visual details in wrinkles, while maintaining superior scalability for garments of diverse topologies without relying on manual UV partition. Finally, we propose a multimodal fusion to incorporate constraints from both frequency modalities and robustly recover deformed 3D garments from transferred images. Extensive experiments show that our method significantly improves animation quality on various garment types and recovers finer wrinkles than state-of-the-art methods. Wei Mao 0001, Changsheng Lu, Hongdong Li |
3DV | 4 |
| 2026 | X-LRM: X-Ray Large Reconstruction Model for Extremely Sparse-View Computed Tomography Recovery in One SecondabstractSparse-view 3D CT reconstruction aims to recover volumetric structures from a limited number of 2D X-ray projections. Existing feedforward methods are constrained by the scarcity of large-scale training datasets and the absence of direct and consistent 3D representations. In this paper, we propose an X-ray Large Reconstruction Model (X-LRM) for extremely sparse-view ($X$-LRM consists of two key components:$X$-former and$X$triplane.$\quad X$-former can handle an arbitrary number of input views using an MLP-based image tokenizer and a Transformer-based encoder. The output tokens are then upsampled into our X-triplane representation, which models the 3D radiodensity as an implicit neural field. To support the training of X-LRM, we introduce Torso-16K, a largescale dataset comprising over 16 K volume-projection pairs of various torso organs. Extensive experiments demonstrate that X-LRM outperforms the state-of-the-art method by 1.5$d B$and achieves$27 \times$faster speed with better flexibility. Furthermore, the evaluation of lung segmentation tasks also suggests the practical value of our approach. Our code and dataset will be released at https://github.com/Richard-Guofeng-Zhang/X-LRM. Guofeng Zhang 0025, Ruyi Zha, Hao He 0011, Yixun Liang, Alan L. Yuille, Hongdong Li, Yuanhao Cai |
3DV | 6 |
| 2026 | DiffNR: Diffusion-Enhanced Neural Representation Optimization for Sparse-View 3D Tomographic ReconstructionabstractNeural representations (NRs), such as neural fields and 3D Gaussians, effectively model volumetric data in computed tomography (CT) but suffer from severe artifacts under sparse-view settings. To address this, we propose DiffNR, a novel framework that enhances NR optimization with diffusion priors. At its core is SliceFixer, a single-step diffusion model designed to correct artifacts in degraded slices. We integrate specialized conditioning layers into the network and develop tailored data curation strategies to support model finetuning. During reconstruction, SliceFixer periodically generates pseudo-reference volumes, providing auxiliary 3D perceptual supervision to fix underconstrained regions. Compared to prior methods that embed CT solvers into time-consuming iterative denoising, our repair-and-augment strategy avoids frequent diffusion model queries, leading to better runtime performance. Extensive experiments show that DiffNR improves PSNR by 3.99 dB on average, generalizes well across domains, and maintains efficient optimization. Shiyan Su, Ruyi Zha, Danli Shi, Hongdong Li, Xuelian Cheng |
AAAI | 4 |
| 2026 | Temporal-Consistent Video Restoration with Pre-trained Diffusion ModelsabstractVideo restoration (VR) aims to recover high-quality videos from degraded ones. Although recent zero-shot VR methods using pre-trained diffusion models (DMs) show good promise, they suffer from approximation errors during reverse diffusion and insufficient temporal consistency. Moreover, dealing with 3D video data, VR is inherently computationally intensive. In this paper, we advocate viewing the reverse process in DMs as a function and present a novel Maximum a Posterior (MAP) framework that directly parameterizes video frames in the seed space of DMs, eliminating approximation errors. We also introduce strategies to promote bilevel temporal consistency: semantic consistency by leveraging clustering structures in the seed space, and pixel-level consistency by progressive warping with optical flow refinements. Extensive experiments on multiple virtual reality tasks demonstrate superior visual quality and temporal consistency achieved by our method compared to the state-of-the-art. Hengkang Wang, Huidong Liu, Chien-Chih Wang, Hongdong Li, Bryan Wang, Ju Sun |
AAAI | 6 |
| 2026 | CoAttn-PEFT: Complementary Attention in Parameter-Efficient Fine-Tuning for Improved Adaptation and Generalization
Chushan Zhang, Ruihan Lu, Jinguang Tong, Xuesong Li 0001, Hongdong Li |
ICPR (2) | 5 |
| 2026 | Construction and Practice of a 5C3S Collaborative Innovation Model for Graduate Education Reform: A Case Study Based on a Biomedical Imaging Analysis Project
Hongdong Li, Senzhang Wang |
ISBRA (1) | 2 |
| 2026 | DICE: Discrete Inversion Enabling Controllable Editing for Masked Generative ModelsabstractRecent advances in discrete diffusion models have demonstrated strong performance in image generation and masked language modeling, yet they remain limited in their capacity for controlled content editing. We propose DICE (Discrete Inversion for Controllable Editing), a novel framework that pioneers precise inversion capabilities for discrete diffusion models, including both masked generative and multinomial diffusion variants. Our key innovation lies in capturing noise sequences and masking patterns during reverse diffusion process, enabling both accurate reconstruction and flexible editing without relying on predefined masks or attention-based manipulations. Through comprehensive experiments across image and text modalities using models such as Paella, VQ-Diffusion, RoBERTa and LLaDA, we demonstrate that DICE successfully maintains high fidelity to the original data while significantly expanding editing capabilities. These results establish new possibilities for fine-grained content manipulation in discrete spaces. Xiaoxiao He, Quan Dao, Ligong Han, Song Wen 0001, Minhao Bai, Di Liu 0003, Han Zhang 0010, Felix Juefei-Xu, Chaowei Tan, Bo Liu 0005, Martin Renqiang Min, Kang Li 0004, Faez Ahmed, Akash Srivastava, Hongdong Li, Junzhou Huang, Dimitris N. Metaxas |
WACV | 15 |
| 2026 | DPBridge: Latent Diffusion Bridge for Dense PredictionabstractDiffusion models demonstrate remarkable capabilities in capturing complex data distributions and have achieved compelling results in many generative tasks. While they have recently been extended to dense prediction tasks such as depth estimation and surface normal prediction, their full potential in this area remains underexplored. As target signal maps and input images are pixel-wise aligned, the conventional noise-to-data generation paradigm is inefficient, and input images can serve as a more informative prior compared to pure noise. Diffusion bridge models, which support data-to-data generation between two general data distributions, offer a promising alternative, but they typically fail to exploit the rich visual priors embedded in large pretrained foundation models. To address these limitations, we integrate diffusion bridge formulation with structured visual priors and introduce DPBridge, the first latent diffusion bridge framework for dense prediction tasks. To resolve the incompatibility between diffusion bridge models and pretrained diffusion backbones, we propose (1) a tractable reverse transition kernel for the diffusion bridge process, enabling maximum likelihood training scheme; (2) finetuning strategies including distribution-aligned normalization and image consistency loss. Experiments across extensive benchmarks validate that our method consistently achieves superior performance, demonstrating its effectiveness and generalization capability under different scenarios. Haorui Ji, Hongdong Li |
WACV | 3 |
| 2026 | Mutual learning for joint disease detection and severity prediction reveals multimodal pathogenesis for neurodegenerative disordersabstractMOTIVATION: Neurodegenerative disorders influence millions of people worldwide, and uncovering the pathogenesis is of urgent need. Many efforts have been made to detect or predict neurodegenerative disorders, while exploring the pathogenesis has been ignored from a systemic perspective. RESULTS: To handle this issue, we propose a novel and powerful method, referred to as Pathogenesis-aware Mutual-Assistance Classification and Regression Optimization (Pa-MACRO). First, Pa-MACRO incorporates a mutual-assistance bidirectional mapping technique with a joint-embedding fine-grained interpretability module. This can extract the intrinsic factors and their interactions of multimodal pathogenesis. Second, our method can simultaneously classify an at-risk individual and predict the severity triggered by neurodegenerative disorders. Furthermore, to address the small sample size issue and the high-dimensional issue, we meticulously incorporate a semi-supervised cooperative learning method to integrate unlabeled data and extend it to a chromosome-wide setting in the spirit of divide-and-conquer. The Alzheimer's Disease Neuroimaging Initiative (ADNI) database was used to evaluate Pa-MACRO. Without bells and whistles, Pa-MACRO establishes new state-of-the-art results in various settings while maintaining superior interpretability, verifying its power and versatility in revealing the pathogenesis of neurodegenerative disorders. AVAILABILITY AND IMPLEMENTATION: The software is publicly available at https://github.com/ZJ-Techie/Pa-MACRO. Jin Zhang 0023, Yixin Ji, Jinhua Liu 0003, Wenrui Cui, Xiaohui Yao, Hongdong Li, Daoqiang Zhang |
Bioinform. | 6 |
| 2026 | MFGS: Mask-free Gaussian separation for 3D object reconstructionabstractAccurate 3D reconstruction from multi-view images is a fundamental problem in computer vision. A common acquisition strategy involves placing an object on a rotating turntable while moving the camera to capture it from various viewpoints. In such scenarios, object moves relative to the background, many existing reconstruction methods rely on object masks to separate the foreground from the background. The quality of these masks significantly affects the final reconstruction, yet obtaining high-quality and consistent masks is a challenging and laborious process, especially when controlled environments like green screens are unavailable. To address this limitation, we introduce Mask-free Gaussian Separation (MFGS), a novel method that performs simultaneous object reconstruction and segmentation without requiring any input masks. Our approach builds on Gaussian Splatting and automatically disentangles the scene by extending each Gaussian primitive with a learnable parameter that represents its probability of belonging to the dynamic foreground object. This separation is optimized in a self-supervised manner, optimized by the object and camera transformation constraints. We evaluated MFGS on new synthetic and real-world datasets designed to reflect this challenging capture scenario. Experimental results demonstrate that our mask-free approach significantly outperforms existing methods. Notably, MFGS surpasses the performance of the state-of-the-art method(2DGS) that relies on high-quality segmentation masks, achieving a 27% improvement in novel view synthesis and a 7% improvement in geometry reconstruction. Jinguang Tong, Xuesong Li 0001, Sundaram Muthu, Fahira A. Maken, Lars Petersson, Hongdong Li |
Pattern Recognit. | 7 |
| 2026 | A Rotation-Translation Decoupled Solution for Visual-Inertial Initialization and Online Spatial-Temporal CalibrationabstractWe propose a novel initialization and online spatial-temporal calibration method for visual-inertial odometry (VIO), which decouples rotation and translation estimation to achieve higher accuracy and better robustness. Existing initialization methods suffer from limited accuracy or robustness (e.g., in scenarios with small translational motion) and rarely integrate simultaneous spatial-temporal calibration during initialization, despite its considerable practical value. Our proposed method leverages rotation-translation decoupling constraints to enable simultaneous estimation of gyroscope bias, extrinsic rotation, and camera-IMU time offset-even under pure rotational motion. Moreover, we are the first to conduct observability analysis on rotational constraints in rotation-translation decoupling methods, experimentally identifying the unobservable state-space directions under three degenerate motions within our approach. We also perform extensive experiments to delineate practical parameter solution boundaries for our method, with both efforts substantially enhancing the overall practical applicability of decoupling-based methods. Extensive experiments on simulated and real-world datasets demonstrate that our method outperforms state-of-the-art approaches in accuracy and robustness while maintaining computational efficiency. Furthermore, experiments verify that it significantly improves convergence in VIO systems. Bo Xu 0022, Zewen Xu, Yijia He, Zhanpeng Ouyang, Hao Wei 0008, Yihong Wu 0002, Jiancheng Li, Hongdong Li |
IEEE Trans. Robotics | 8 |
| 2026 | BAG: Body-Aligned 3D Wearable Asset GenerationabstractWhile recent advancements have demonstrated remarkable progress in general 3D shape generation, the challenge of automatically generating wearable 3D assets remains largely unexplored. To address this gap, we present BAG - a Body-aligned Asset Generation method that produces 3D wearable assets which can be automatically fitted onto given 3D human bodies. This is achieved by controlling the 3D generation process using human body shape and pose information. Specifically, we first construct a general single-image-to-consistent-multi-view diffusion model, and train it on the large-scale Objaverse dataset to ensure diversity and generalizability. We then train a body-conditioned multi-view ControlNet to guide the generator toward producing body-aligned multi-view images. The control signal leverages multi-view 2D projections of the target human body, where pixel values represent the XYZ coordinates of the body surface in a canonical space. The resulting body-conditioned multi-view diffusion outputs body-aligned images, which are subsequently fed into a native 3D diffusion model to reconstruct the 3D shape of the asset. Finally, we recover the similarity transformation using multi-view silhouette supervision and mitigate asset-body penetration using physics-based simulation, ensuring accurate asset fitting onto the target body. Experimental results demonstrate that our method significantly outperforms existing approaches in terms of prompt adherence, shape diversity, and shape quality. Zhongjin Luo, Yang Li 0193, Senbo Wang, Han Yan 0004, Xibin Song, Taizhang Shang, Wei Mao 0001, Hongdong Li, Xiaoguang Han 0001, Pan Ji |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2025 | JADE: Joint-Aware Latent Diffusion for 3D Human Generative ModelingabstractGenerative modeling of 3D human bodies have been studied extensively in computer vision. The core is to design a compact latent representation that is both expressive and semantically interpretable, yet existing approaches struggle to achieve both requirements. In this work, we introduce JADE, a generative framework that learns the variations of human shapes with fined-grained control. Our key insight is a joint-aware latent representation that decomposes human bodies into skeleton structures, modeled by joint positions, and local surface geometries, characterized by features attached to each joint. This disentangled latent space design enables geometric and semantic interpretation, facilitating users with flexible controllability. To generate coherent and plausible human shapes under our proposed decomposition, we also present a cascaded pipeline where two diffusions are employed to model the distribution of skeleton structures and local surface geometries respectively. Extensive experiments are conducted on public datasets, where we demonstrate the effectiveness of JADE framework in multiple tasks in terms of autoencoding reconstruction accuracy, editing controllability and generation quality compared with existing methods. Haorui Ji, Hongdong Li |
3DV | 4 |
| 2025 | Geometry-Guided Cross-View Diffusion for One-to-Many Cross-View Image SynthesisabstractThis paper presents a novel approach for cross-view synthesis aimed at generating plausible ground-level images from corresponding satellite imagery or vice versa. We refer to these tasks as satellite-to-ground (Sat2Grd) and ground-to-satellite (Grd2Sat) synthesis, respectively. Unlike previous works that typically focus on one-to-one generation, producing a single output image from a single input image, our approach acknowledges the inherent one-to-many nature of the problem. This recognition stems from the challenges posed by differences in illumination, weather conditions, and occlusions between the two views. To effectively model this uncertainty, we leverage recent advancements in diffusion models. Specifically, we exploit random Gaussian noise to represent the diverse possibilities learnt from the target view data. We introduce a Geometry-guided Cross-view Condition (GCC) strategy to establish explicit geometric correspondences between satellite and street-view features. This enables us to resolve the geometry ambiguity introduced by camera pose between image pairs, boosting the performance of cross-view image synthesis. Through extensive quantitative and qualitative analyses on three bench-mark cross-view datasets, we demonstrate the superiority of our proposed geometry-guided cross-view condition over baseline methods, including recent state-of-the-art approaches in cross-view image synthesis. Our method generates images of higher quality, fidelity, and diversity than other state-of-the-art approaches. Yujiao Shi 0002, Akhil Perincherry, Ankit Vora, Hongdong Li |
3DV | 6 |
| 2025 | Improving Cancer Gene Prediction by Enhancing Common Information Between the PPI Network and Gene Functional AssociationabstractIdentifying cancer genes is crucial for treatment and understanding pathogenesis. Recent methods typically leverage protein-protein interaction (PPI) networks or gene functional association data from annotated gene sets. There may be some shared neighborhood structure information between these two types of gene association data. While this common information may contain more accurate gene association information, existing methods often overlook this potential. To address this gap, we introduce DISFusion, which integrates multi-omics cancer data, PPI networks, and gene functional associations to identify cancer genes. A key innovation of DISFusion is the cross-view decorrelation loss, which enhances the common information between PPI networks and gene functional associations, thereby improving prediction accuracy. Extensive experiments indicate that DISFusion outperforms state-of-the-art methods and exhibits greater generalization ability. Moreover, analysis of CPTAC pan-cancer proteomic data highlights significant associations between the 30 novel cancer genes predicted by DISFusion and multiple cancer types, underscoring its practical utility. These findings validate the effectiveness of enhancing common information and provide new insights into cancer gene identification. Hongdong Li, Jianxin Wang 0001 |
AAAI | 2 |
| 2025 | Predicting Brain Age Based on Neuroimaging and DNA Methylation via A Multiomics Attention-Based Variational AutoencoderabstractBrain age (BA) is recognized as a significant biomarker for health, closely associated with brain aging, and has been proposed to correlate with the progression of neurodegenerative diseases. Previous studies on brain age prediction predominantly relied on single omics data, limiting the integration of cooperative information from multiple omics data. Additionally, existing brain age prediction methods are susceptible to intermediate features since not all these features are related to brain age. To address these limitations, the study introduces a multiomics attention-based VAE method to predict brain age by integrating neuroimaging data and DNA methylation (DNAm) data, aiming to identify biologically meaningful features truly associated with brain aging. Experimental results show that our method predicts brain age with the MAE of 2.08 years, outperforming the state-of-the-art methods under the same experimental conditions. Furthermore, we estimated the brain age gap in patients with Mild Cognitive Impairment (MCI) and Alzheimer's Disease (AD). The results of both MCI and AD patients exhibited a larger brain age gap compared to the Cognitively Normal (CN) group, indicating the model's discriminative capacity across different diagnostic groups. These findings can assist in the early diagnosis of AD and the formulation of early treatment strategies. Wenrui Cui, Hong Pang, Yan Yang 0011, Muheng Shang, Hongdong Li, Lei Du 0001 |
BIBM | 5 |
| 2025 | A Data-Driven Brain Imaging Quantitative Trait Mediation Effect Identification Method for Alzheimer's DiseaseabstractAlzheimer's disease (AD) is a severe degenerative disease and finding its causal factors of high-risk are very important. Mediation analysis has been a powerful tool to elucidate the underlying mechanisms of trait of interest using genetic variations as instrumental variable (IVs). However, most current mediation analysis method can only work on a limited number of suspected traits and genetic variations, which demands extensive prior knowledge. In this study, we proposed a datadriven learning method to identify potential mediation effects of brain imaging quantitative traits (QTs). Our method couples the three single models of mediation model and treats it as a multiobjective learning problem. The proposed method can work on brain-wide and genome-wide brain imaging and genetic data without providing candidate imaging QTs and genetic IVs. We applied the proposed method to PET imaging QTs of whole brain and genetic IV of whole genome. The results showed that our method successfully identified multiple PET imaging mediators linking genetic variations to AD. Thus, our method can serve as a powerful screening tool for large-scale mediation analysis. Yan Yang 0011, Muheng Shang, Hongdong Li, Lei Du 0001 |
BIBM | 3 |
| 2025 | Predicting MCI Conversion Status Using Baseline Neuroimaging Scans and Genetics VariationsabstractMild cognitive impairment (MCI) is a prodromal stage of Alzheimer's disease (AD), but not all MCI subjects develop into AD finally. Therefore, distinguishing progressive MCI (pMCI) subjects from stable MCI (sMCI) subjects is an area of intense interest, which may provide targeted treatments for at-risk individuals. On this account, building an MCI conversion prediction model at the early stage is particularly important. The neuroimaging data, especially multi-modal ones, has proven to be a great alternative in predicting MCIs' conversion. In addition, genetic variations such as Single Nucleotide Polymorphism (SNP) can also imply the conversion risk of an individual. The neuroimaging data represents the current status, while SNPs convey the inherited risk of an individual. In this paper, we propose a deep representative fusion method that combines multi-modal baseline neuroimaging data and genetic variations. It can predict the progressive status of MCIs over the following two years, three years and four years, respectively. Experimental results from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database demonstrate that the proposed method has better prediction capability than comparison methods. Moreover, findings show that the stability of the default mode network (DMN) and ventral attention network (VAN) are correlated with the MCI conversion and the learned imaging representations are related to minimental state examination (MMSE) scores which are associated with AD progression. Yan Yang 0011, Muheng Shang, Jin Zhang 0023, Hongdong Li, Lei Du 0001 |
BIBM | 5 |
| 2025 | GS-2DGS: Geometrically Supervised 2DGS for Reflective Object Reconstructionabstract3D modeling of highly reflective objects remains challenging due to strong view-dependent appearances. While previous SDF-based methods can recover high-quality meshes, they are often time-consuming and tend to produce over-smoothed surfaces. In contrast, 3D Gaussian Splatting (3DGS) offers the advantage of high speed and detailed real-time rendering, but extracting surfaces from the Gaussians can be noisy due to the lack of geometric constraints. To bridge the gap between these approaches, we propose a novel reconstruction method called GS-2DGS for reflective objects based on 2D Gaussian Splatting (2DGS). Our approach combines the rapid rendering capabilities of Gaussian Splatting with additional geometric information from foundation models. Experimental results on synthetic and real datasets demonstrate that our method significantly outperforms Gaussian-based techniques in terms of reconstruction and relighting and achieves performance comparable to SDF-based methods while being an order of magnitude faster. Code is available at https://github.com/hirotong/GS2DGS Jinguang Tong, Xuesong Li 0001, Fahira A. Maken, Sundaram Muthu, Lars Petersson, Hongdong Li |
CVPR | 7 |
| 2025 | FRESA: Feedforward Reconstruction of Personalized Skinned Avatars from Few ImagesabstractWe present a novel method for reconstructing personalized 3D human avatars with realistic animation from only a few images. Due to the large variations in body shapes, poses, and cloth types, existing methods mostly require hours of per-subject optimization during inference, which limits their practical applications. In contrast, we learn a universal prior from over a thousand clothed humans to achieve instant feedforward generation and zero-shot generalization. Specifically, instead of rigging the avatar with shared skinning weights, we jointly infer personalized avatar shape, skinning weights, and pose-dependent deformations, which effectively improves overall geometric fidelity and reduces deformation artifacts. Moreover, to normalize pose variations and resolve coupled ambiguity between canonical shapes and skinning weights, we design a 3D canonicalization process to produce pixel-aligned initial conditions, which helps to reconstruct fine-grained geometric details. We then propose a multi-frame feature aggregation to robustly reduce artifacts introduced in canonicalization and fuse a plausible avatar preserving person-specific identities. Finally, we train the model in an end-to-end framework on a large-scale capture dataset, which contains diverse human subjects paired with high-quality 3D scans. Extensive experiments show that our method generates more authentic reconstruction and animation than state-of-the- arts, and can be directly generalized to inputs from casually taken phone photos. Project page and code is available at https://github.com/rongakowang/FRESA. Fabian Prada, Zhongshi Jiang, Chengxiang Yin 0003, Shunsuke Saito, Igor Santesteban, Javier Romero 0002, Rohan Joshi, Hongdong Li, Jason M. Saragih, Yaser Sheikh |
CVPR | 11 |
| 2025 | Probability Density Geodesics in Image Diffusion Latent SpaceabstractDiffusion models indirectly estimate the probability density over a data space, which can be used to study its structure. In this work, we show that geodesics can be computed in diffusion latent space, where the norm induced by the spatially-varying inner product is inversely proportional to the probability density. In this formulation, a path that traverses a high density (that is, probable) region of image latent space is shorter than the equivalent path through a low density region. We present algorithms for solving the associated initial and boundary value problems and show how to compute the probability density along the path and the geodesic distance between two points. Using these techniques, we analyze how closely video clips approximate geodesics in a pre-trained image diffusion space. Finally, we demonstrate how these techniques can be applied to training-free image sequence interpolation and extrapolation, given a pre-trained image diffusion model. Qingtao Yu, Zhaoyuan Yang, Peter H. Tu, Jing Zhang 0052, Hongdong Li, Richard I. Hartley, Dylan Campbell |
CVPR | 6 |
| 2025 | Bioinformatics Course Reform Through Projects Integrating History, Theory, and Practice
Hongdong Li, Guihua Duan |
ISBRA (2) | 2 |
| 2025 | SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-ConditioningabstractExisting diffusion-based super-resolution approaches often exhibit semantic ambiguities due to inaccuracies and incompleteness in their text conditioning, coupled with the inherent tendency for cross-attention to divert towards irrelevant pixels. These limitations can lead to semantic misalignment and hallucinated details in the generated high-resolution outputs. To address these, we propose a novel, plug-and-play *spatially re-focused super-resolution (SRSR)* framework that consists of two core components: first, we introduce Spatially Re-focused Cross-Attention (SRCA), which refines text conditioning at inference time by applying visually-grounded segmentation masks to guide cross-attention. Second, we introduce a Spatially Targeted Classifier-Free Guidance (STCFG) mechanism that selectively bypasses text influences on ungrounded pixels to prevent hallucinations. Extensive experiments on both synthetic and real-world datasets demonstrate that SRSR consistently outperforms seven state-of-the-art baselines in standard fidelity metrics (PSNR and SSIM) across all datasets, and in perceptual quality measures (LPIPS and DISTS) on two real-world benchmarks, underscoring its effectiveness in achieving both high semantic fidelity and perceptual quality in super-resolution. Majid Abdolshah, Violetta Shevchenko, Hongdong Li, Pulak Purkait |
NeurIPS | 4 |
| 2025 | MB-TaylorFormer V2: Improved Multi-Branch Linear Transformer Expanded by Taylor Formula for Image RestorationabstractRecently, Transformer networks have demonstrated outstanding performance in the field of image restoration due to the global receptive field and adaptability to input. However, the quadratic computational complexity of Softmax-attention poses a significant limitation on its extensive application in image restoration tasks, particularly for high-resolution images. To tackle this challenge, we propose a novel variant of the Transformer. This variant leverages the Taylor expansion to approximate the Softmax-attention and utilizes the concept of norm-preserving mapping to approximate the remainder of the first-order Taylor expansion, resulting in a linear computational complexity. Moreover, we introduce a multi-branch architecture featuring multi-scale patch embedding into the proposed Transformer, which has four distinct advantages: 1) various sizes of the receptive field; 2) multi-level semantic information; 3) flexible shapes of the receptive field; 4) accelerated training and inference speed. Hence, the proposed model, named the second version of Taylor formula expansion-based Transformer (for short MB-TaylorFormer V2) has the capability to concurrently process coarse-to-fine features, capture long-distance pixel interactions with limited computational cost, and improve the approximation of the Taylor expansion remainder. Experimental results across diverse image restoration benchmarks demonstrate that MB-TaylorFormer V2 achieves state-of-the-art performance in multiple image restoration tasks, such as image dehazing, deraining, desnowing, motion deblurring, and denoising, with very little computational overhead. Zhi Jin 0002, Yuwei Qiu, Kaihao Zhang, Hongdong Li, Wenhan Luo |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | GUS-IR: Gaussian Splatting With Unified Shading for Inverse RenderingabstractRecovering the intrinsic physical attributes of a scene from images, generally termed as the inverse rendering problem, has been a central and challenging task in computer vision and computer graphics. In this paper, we present GUS-IR, a novel framework designed to address the inverse rendering problem for complicated scenes featuring rough and glossy surfaces. This paper starts by analyzing and comparing two prominent shading techniques popularly used for inverse rendering, forward shading and deferred shading, effectiveness in handling complex materials. More importantly, we propose a unified shading solution that combines the advantages of both techniques for better decomposition. In addition, we analyze the normal modeling in 3D Gaussian Splatting (3DGS) and utilize the shortest axis as normal for each particle in GUS-IR, along with a depth-related regularization, resulting in improved geometric representation and better shape reconstruction. Furthermore, we enhance the probe-based baking scheme proposed by GS-IR to achieve more accurate ambient occlusion modeling to better handle indirect illumination. Extensive experiments have demonstrated the superior performance of GUS-IR in achieving precise intrinsic decomposition and geometric representation, supporting many downstream tasks (such as relighting, retouching) in computer vision, graphics, and extended reality. Zhihao Liang 0002, Hongdong Li, Kui Jia, Kailing Guo, Qi Zhang 0029 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Faster Stochastic Variance Reduction Methods for Compositional MiniMax OptimizationabstractThis paper delves into the realm of stochastic optimization for compositional minimax optimization—a pivotal challenge across various machine learning domains, including deep AUC and reinforcement learning policy evaluation. Despite its significance, the problem of compositional minimax optimization is still under-explored. Adding to the complexity, current methods of compositional minimax optimization are plagued by sub-optimal complexities or heavy reliance on sizable batch sizes. To respond to these constraints, this paper introduces a novel method, called Nested STOchastic Recursive Momentum (NSTORM), which can achieve the optimal sample complexity and obtain the nearly accuracy solution, matching the existing minimax methods. We also demonstrate that NSTORM can achieve the same sample complexity under the Polyak-Lojasiewicz (PL)-condition—an insightful extension of its capabilities. Yet, NSTORM encounters an issue with its requirement for low learning rates, potentially constraining its real-world applicability in machine learning. To overcome this hurdle, we present ADAptive NSTORM (ADA-NSTORM) with adaptive learning rates. We demonstrate that ADA-NSTORM can achieve the same sample complexity but the experimental results show its more effectiveness. All the proposed complexities indicate that our proposed methods can match lower bounds to existing minimax optimizations, without requiring a large batch size in each iteration. Extensive experiments support the efficiency of our proposed methods. Jin Liu 0012, Xiaokang Pan, Junwen Duan, Hongdong Li, Youqi Li |
AAAI | 4 |
| 2024 | StraightPCF: Straight Point Cloud FilteringabstractPoint cloud filtering is a fundamental 3D vision task, which aims to remove noise while recovering the underlying clean surfaces. State-of-the-art methods remove noise by moving noisy points along stochastic trajectories to the clean surfaces. These methods often require regularization within the training objective and/or during post-processing, to ensure fidelity. In this paper, we introduce StraightPCF, a new deep learning based method for point cloud filtering. It works by moving noisy points along straight paths, thus reducing discretization errors while ensuring faster convergence to the clean surfaces. We model noisy patches as intermediate states between high noise patch variants and their clean counterparts, and design the VelocityModule to infer a constant flow velocity from the former to the latter. This constant flow leads to straight filtering trajectories. In addition, we introduce a DistanceModule that scales the straight trajectory using an estimated distance scalar to attain convergence near the clean surface. Our network is lightweight and only has ~530K parameters, being 17% of IterativePFn (a most recent point cloud filtering network). Extensive experiments on both synthetic and real-world data show our method achieves state-of-the-art results. Our method also demonstrates nice distributions of filtered points without the need for regularization. The implementation code can be found at: https://github.com/ddsediri/StraightPCF. Dasith de Silva Edirimuni, Xuequan Lu, Gang Li 0009, Lei Wei 0002, Antonio Robles-Kelly, Hongdong Li |
CVPR | 6 |
| 2024 | View from Above: Orthogonal-View Aware Cross-View LocalizationabstractThis paper presents a novel aerial-to-ground feature ag-gregation strategy, tailored for the task of cross- view image-based geo-localization. Conventional vision-based methods heavily rely on matching ground-view image features with a pre-recorded image database, often through establishing planar homography correspondences via a planar ground assumption. As such, they tend to ignore features that are off-ground and not suited for handling visual occlusions, leading to unreliable localization in challenging scenarios. We propose a Top-to-Ground Aggregation (T2GA) module that capitalizes aerial orthographic views to aggregate features down to the ground level, leveraging reliable off-ground information to improve feature alignment. Furthermore, we introduce a Cycle Domain Adaptation (CycDA) loss that ensures feature extraction robustness across do-main changes. Additionally, an Equidistant Re-projection (ERP) loss is introduced to equalize the impact of all key-points on orientation error, leading to a more extended distribution of keypoints which benefits orientation estimation. On both KITTI and Ford Multi-AV datasets, our method consistently achieves the lowest mean longitudinal and lateral translations across different settings and obtains the smallest orientation error when the initial pose is less ac-curate, a more challenging setting. Further, it can complete an entire route through continual vehicle pose estimation with initial vehicle pose given only at the starting point.11Code is available at https://github.com/ShanWang-Shan/View FromAbove. Shan Wang 0010, Jiawei Liu 0005, Yanhao Zhang 0003, Sundaram Muthu, Fahira A. Maken, Kaihao Zhang, Hongdong Li |
CVPR | 8 |
| 2024 | ConsistNet: Enforcing 3D Consistency for Multi-View Images DiffusionabstractGiven a single image of a 3D object, this paper proposes a novel method (named ConsistNet) that can generate multiple images of the same object, as if they are capturedfrom different viewpoints, while the 3D (multi-view) consistencies among those multiple generated images are effectively exploited. Central to our method is a lightweight multi-view consistency block that enables information exchange across multiple single-view diffusion processes based on the underlying multi-view geometry principles. ConsistNet is an extension to the standard latent diffusion model and it consists of two submodules: (a) a view aggregation module that unprojects multi-view features into global 3D volumes and infers consistency, and (b) a ray aggregation module that samples and aggregates 3D consistent features back to each view to enforce consistency. Our approach departs from previous methods in multi-view image generation, in that it can be easily dropped in pretrained LDMs without requiring explicit pixel correspondences or depth prediction. Experiments show that our method effectively learns 3D consistency over a frozen Zero123-XL backbone and can generate 16 surrounding views of the object within 11 seconds on a single A100 GPU. Our code will be made available on https://github.com/JiayuYANG/ConsistNet. Ziang Cheng, Yunfei Duan, Pan Ji, Hongdong Li |
CVPR | 5 |
| 2024 | NeuSDFusion: A Spatial-Aware Generative Model for 3D Shape Completion, Reconstruction, and Generation
Ruikai Cui, Weizhe Liu, Weixuan Sun, Senbo Wang, Taizhang Shang, Yang Li 0193, Xibin Song, Han Yan 0004, Zhennan Wu, Shenzhou Chen, Hongdong Li, Pan Ji |
ECCV (19) | 11 |
| 2024 | SemReg: Semantics Constrained Point Cloud Registration
Sheldon Fung, Xuequan Lu, Dasith de Silva Edirimuni, Wei Pan 0010, Xiao Liu 0004, Hongdong Li |
ECCV (41) | 6 |
| 2024 | Weakly-Supervised Camera Localization by Ground-to-Satellite Image Registration
Yujiao Shi 0002, Hongdong Li, Akhil Perincherry, Ankit Vora |
ECCV (9) | 2 |
| 2024 | Prompting Future Driven Diffusion Model for Hand Motion Prediction
Bowen Tang 0002, Kaihao Zhang, Wenhan Luo, Wei Liu 0005, Hongdong Li |
ECCV (7) | 5 |
| 2024 | Towards High-Quality 3D Motion Transfer with Realistic Apparel Animation
Wei Mao 0001, Changsheng Lu, Hongdong Li |
ECCV (39) | 4 |
| 2024 | Adapting Fine-Grained Cross-View Localization to Areas Without Fine Ground Truth
Zimin Xia, Yujiao Shi 0002, Hongdong Li, Julian F. P. Kooij |
ECCV (31) | 3 |
| 2024 | Physically Plausible Color Correction for Neural Radiance Fields
Qi Zhang 0029, Hongdong Li |
ECCV (44) | 3 |
| 2024 | Alice Benchmarks: Connecting Real World Re-Identification with the SyntheticabstractFor object re-identification (re-ID), learning from synthetic data has become a promising strategy to cheaply acquire large-scale annotated datasets and effective models, with few privacy concerns. Many interesting research problems arise from this strategy, e.g., how to reduce the domain gap between synthetic source and real-world target. To facilitate developing more new approaches in learning from synthetic data, we introduce the Alice benchmarks, large-scale datasets providing benchmarks as well as evaluation protocols to the research community. Within the Alice benchmarks, two object re-ID tasks are offered: person and vehicle re-ID. We collected and annotated two challenging real-world target datasets: AlicePerson and AliceVehicle, captured under various illuminations, image resolutions, etc. As an important feature of our real target, the clusterability of its training set is not manually guaranteed to make it closer to a real domain adaptation test scenario. Correspondingly, we reuse existing PersonX and VehicleX as synthetic source domains. The primary goal is to train models from synthetic data that can work effectively in the real world. In this paper, we detail the settings of Alice benchmarks, provide an analysis of existing commonly-used domain adaptation methods, and discuss some interesting future directions. An online server has been set up for the community to evaluate methods conveniently and fairly. Datasets and the online server details are available at https://sites.google.com/view/alice-benchmarks. Xiaoxiao Sun 0002, Yue Yao 0001, Shengjin Wang, Hongdong Li, Liang Zheng 0001 |
ICLR | 4 |
| 2024 | RGB-based Category-level Object Pose Estimation via Decoupled Metric Scale RecoveryabstractWhile showing promising results, recent RGB-D camera-based category-level object pose estimation methods have restricted applications due to the heavy reliance on depth sensors. RGB-only methods provide an alternative to this problem yet suffer from inherent scale ambiguity stemming from monocular observations. In this paper, we propose a novel pipeline that decouples the 6D pose and size estimation to mitigate the influence of imperfect scales on rigid transformations. Specifically, we leverage a pre-trained monocular estimator to extract local geometric information, mainly facilitating the search for inlier 2D-3D correspondence. Meanwhile, a separate branch is designed to directly recover the metric scale of the object based on category-level statistics. Finally, we advocate using the RANSAC-PnP algorithm to robustly solve for 6D object pose. Extensive experiments have been conducted on both synthetic and real datasets, demonstrating the superior performance of our method over previous state-of-the-art RGB-based approaches, especially in terms of rotation accuracy. Code: https://github.com/goldoak/DMSR. Jiaxin Wei 0001, Xibin Song, Weizhe Liu, Laurent Kneip, Hongdong Li, Pan Ji |
ICRA | 5 |
| 2024 | Increasing SLAM Pose Accuracy by Ground-to-Satellite Image RegistrationabstractVision-based localization for autonomous driving has been of great interest among researchers. When a pre-built 3D map is not available, the techniques of visual simultaneous localization and mapping (SLAM) are typically adopted. Due to error accumulation, visual SLAM (vSLAM) usually suffers from long-term drift. This paper proposes a framework to increase the localization accuracy by fusing the vSLAM with a deep-learning based ground-to-satellite (G2S) image registration method. In this framework, a coarse (spatial correlation bound check) to fine (visual odometry consistency check) method is designed to select the valid G2S prediction. The selected prediction is then fused with the SLAM measurement by solving a scaled pose graph problem. To further increase the localization accuracy, we provide an iterative trajectory fusion pipeline. The proposed framework is evaluated on two well-known autonomous driving datasets, and the results demonstrate the accuracy and robustness in terms of vehicle localization. The code will be available at https://github.com/YanhaoZhang/SLAM-G2S-Fusion. Yanhao Zhang 0003, Yujiao Shi 0002, Shan Wang 0010, Ankit Vora, Akhil Perincherry, Yongbo Chen 0001, Hongdong Li |
ICRA | 7 |
| 2024 | MAVIS: Multi-Camera Augmented Visual-Inertial SLAM using SE2(3) Based Exact IMU Pre-integrationabstractWe present a novel optimization-based Visual-Inertial SLAM system designed for multiple partially over-lapped camera systems, named MAVIS. Our framework fully exploits the benefits of wide field-of-view from multi-camera systems, and the metric scale measurements provided by an inertial measurement unit (IMU). We introduce an improved IMU pre-integration formulation based on the exponential function of an automorphism of SE2(3), which can effectively enhance tracking performance under fast rotational motion and extended integration time. Furthermore, we extend conventional front-end tracking and back-end optimization module designed for monocular or stereo setup towards multi-camera systems, and introduce implementation details that contribute to the performance of our system in challenging scenarios. The practical validity of our approach is supported by our experiments on public datasets. Our MAVIS won the first place in all the vision-IMU tracks (single and multi-session SLAM) on Hilti SLAM Challenge 2023 with 1.7 times the score compared to the second place1. Yifu Wang, Yonhon Ng, Inkyu Sa, Álvaro Parra Bustos, Cristian Rodriguez Opazo, Hongdong Li |
ICRA | 7 |
| 2024 | Advancing Virtual Reality Interaction: A Ring-Shaped Controller and Pose TrackingabstractEnsuring robust tracking of controllers’ movement is critical for human-robot interaction in virtual reality (VR) scenarios. This paper proposes a robust tracking algorithm based on a novel wearable ring-shaped controller equipped with an inertial measurement unit (IMU) and a light-emitting diode (LED). This novel controller design allows users to free up their hands for more immersive experiences. To track the controller’s motion accurately and robustly, we resort to various forms of visual measurements, including 6 DoF and 5 DoF pose measurements from hand gesture detection, as well as 3 DoF position measurement and 2 DoF image measurement derived from the LED. We theoretically analyze the performances of these observation models and propose an optimal observation model combination scheme. Moreover, the necessity and rationale of online estimating system gravity are illustrated. The effectiveness of our tracking method is validated through extensive experiments. Zhuqing Zhang, Dongxuan Li, Yijia He, Pan Ji, Rong Xiong, Hongdong Li, Yue Wang 0020 |
ICRA | 7 |
| 2024 | Recurrent Non-Rigid Point Cloud RegistrationabstractNon-rigid point cloud registration remains a significant challenge in 3D computer vision due to the complexity of structural deforms, lack of overlaps, and sensitivity to initialization. This paper introduces a framework inspired by the recent success in recurrent architecture, adapted to accommodate the unique characteristics of point clouds. More specifically, we design a recurrent update network block for progressively refining local registration results under a local rigidity assumption, starting from an initial global SE(3) alignment. Through comparison, our method consistently outperforms competing methods in standard metrics, achieving a 33% reduction in EPE on the 4DLoMatch benchmark compared to the second-best method. To the best of our knowledge, the proposed method is the first to successfully demonstrate that the recurrent update strategy can effectively address the non-rigid registration task with large displacement, significant deform, and low overlap. The source code and the model will be released at http://dummy.url/. Ziang Cheng, Hongdong Li |
IROS | 3 |
| 2024 | LAM3D: Large Image-Point Clouds Alignment Model for 3D Reconstruction from Single ImageabstractLarge Reconstruction Models have made significant strides in the realm of automated 3D content generation from single or multiple input images. Despite their success, these models often produce 3D meshes with geometric inaccuracies, stemming from the inherent challenges of deducing 3D shapes solely from image data. In this work, we introduce a novel framework, the Large Image and Point Cloud Alignment Model (LAM3D), which utilizes 3D point cloud data to enhance the fidelity of generated 3D meshes. Our methodology begins with the development of a point-cloud-based network that effectively generates precise and meaningful latent tri-planes, laying the groundwork for accurate 3D mesh reconstruction. Building upon this, our Image-Point-Cloud Feature Alignment technique processes a single input image, aligning to the latent tri-planes to imbue image features with robust 3D information. This process not only enriches the image features but also facilitates the production of high-fidelity 3D meshes without the need for multi-view input, significantly reducing geometric distortions. Our approach achieves state-of-the-art high-fidelity 3D mesh reconstruction from a single image in just 6 seconds, and experiments on various datasets demonstrate its effectiveness. Ruikai Cui, Xibin Song, Weixuan Sun, Senbo Wang, Weizhe Liu, Shenzhou Chen, Taizhang Shang, Yang Li 0193, Nick Barnes, Hongdong Li, Pan Ji |
NeurIPS | 10 |
| 2024 | R2-Gaussian: Rectifying Radiative Gaussian Splatting for Tomographic Reconstructionabstract3D Gaussian splatting (3DGS) has shown promising results in image rendering and surface reconstruction. However, its potential in volumetric reconstruction tasks, such as X-ray computed tomography, remains under-explored. This paper introduces R$^2$-Gaussian, the first 3DGS-based framework for sparse-view tomographic reconstruction. By carefully deriving X-ray rasterization functions, we discover a previously unknown \emph{integration bias} in the standard 3DGS formulation, which hampers accurate volume retrieval. To address this issue, we propose a novel rectification technique via refactoring the projection from 3D to 2D Gaussians. Our new method presents three key innovations: (1) introducing tailored Gaussian kernels, (2) extending rasterization to X-ray imaging, and (3) developing a CUDA-based differentiable voxelizer. Experiments on synthetic and real-world datasets demonstrate that our method outperforms state-of-the-art approaches in accuracy and efficiency. Crucially, it delivers high-quality results in 4 minutes, which is 12$\times$ faster than NeRF-based methods and on par with traditional algorithms. Ruyi Zha, Yuanhao Cai, Jiwen Cao, Yanhao Zhang 0003, Hongdong Li |
NeurIPS | 6 |
| 2024 | Frankenstein: Generating Semantic-Compositional 3D Scenes in One Tri-PlaneabstractWe present Frankenstein, a diffusion-based framework that can generate semantic-compositional 3D scenes in a single pass. Unlike existing methods that output a single, unified 3D shape, Frankenstein simultaneously generates multiple separated shapes, each corresponding to a semantically meaningful part. The 3D scene information is encoded in one single triplane tensor, from which multiple Signed Distance Function (SDF) fields can be decoded to represent the compositional shapes. During training, an auto-encoder compresses tri-planes into a latent space, and then the denoising diffusion process is employed to approximate the distribution of the compositional scenes. Frankenstein demonstrates promising results in generating room interiors as well as human avatars with automatically separated parts. The generated scenes facilitate many downstream applications, such as part-wise re-texturing, object rearrangement in the room or avatar cloth re-targeting. Han Yan 0004, Yang Li 0193, Zhennan Wu, Shenzhou Chen, Weixuan Sun, Taizhang Shang, Weizhe Liu, Xiaqiang Dai, Chao Ma 0004, Hongdong Li, Pan Ji |
SIGGRAPH Asia | 11 |
| 2024 | Stereo Matching in Time: 100+ FPS Video Stereo Matching for Extended RealityabstractReal-time Stereo Matching is a cornerstone task for Extended Reality (XR) applications, such as 3D scene understanding, video pass-through, and mixed-reality games. Despite significant advancements, getting accurate depth information in real time on a low-power mobile device remains a challenge. One of the main difficulties is the lack of high-quality indoor video stereo data captured by head-mounted VR or AR glasses. To address this, we introduce a novel video stereo synthetic dataset that comprises photorealistic renderings of various indoor scenes and realistic camera motion captured by a moving VR/AR head-mounted display (HMD). Our newly proposed dataset enables one to develop a novel framework for continuous video-rate stereo matching.As another contribution, we also propose a new video-based stereo matching approach tailored for XR applications, which achieves real-time inference at an impressive 134fps on a standard desktop computer, or 30fps on a battery-powered HMD. Our key insight is that disparity and contextual information are highly correlated and redundant between consecutive stereo frames. By unrolling an iterative cost aggregation in time (i.e. in temporal dimension), we are able to distribute and reuse the aggregated features over time. This leads to a substantial reduction in computation without sacrificing accuracy. We conducted extensive evaluations and demonstrated that our method achieves superior performance compared to the current state-of-the-art, making it a strong contender for real-time stereo matching in VR/AR applications. Our dataset is released on https://github.com/za-cheng/XR-Stereo. Ziang Cheng, Hongdong Li |
WACV | 3 |
| 2024 | Unsupervised 3D Pose Estimation with Non-Rigid Structure-from-Motion ModelingabstractMost existing 3D human pose estimation work rely heavily on the powerful memory capability of networks to obtain suitable 2D-3D mappings from the training data. Few works have studied the modeling of human posture deformation in motion. In this paper, we propose a new modeling method for human pose deformations and design an accompanying diffusion-based motion prior. Inspired by the field of non-rigid structure-from-motion, we divide the task of reconstructing 3D human skeletons in motion into the estimation of a 3D reference skeleton, and a frame-by-frame skeleton deformation. A mixed spatial-temporal NRSfMformer is used to simultaneously estimate the 3D reference skeleton and the skeleton deformation of each frame from 2D observations sequence, and then sum them up to obtain the pose of each frame. Subsequently, a loss term based on the diffusion model is used to ensure that the pipeline learns the correct prior motion knowledge. Finally, we have evaluated our proposed method on mainstream datasets and obtained superior results outperforming the state-of-the-art. Haorui Ji, Yuchao Dai, Hongdong Li |
WACV | 4 |
| 2024 | PMVC: Promoting Multi-View Consistency for 3D Scene ReconstructionabstractReconstructing the geometry of a 3D scene from its multi-view 2D observations has been a central task of 3D computer vision. Recent methods based on neural rendering that use implicit shape representations such as the neural Signed Distance Function (SDF), have shown impressive performance. However, they fall short in recovering fine details in the scene, especially when employing a multilayer perceptron (MLP) as the interpolation function for the SDF representation. Per-frame image normal or depth-map prediction have been utilized to tackle this issue, but these learning-based depth/normal predictions are based on a single image frame only, hence overlooking the underlying multiview consistency of the scene, leading to inconsistent erroneous 3D reconstruction. To mitigate this problem, we propose to leverage multi-view deep features computed on the images. In addition, we employ an adaptive sampling strategy that assesses the fidelity of the multi-view image consistency. Our approach outperforms current state-of-the-art methods, delivering an accurate and robust scene representation with particularly enhanced details. The effectiveness of our proposed approach is evaluated by extensive experiments conducted on the ScanNet and Replica datasets, showing superior performance than the current state-of-the-art. Chushan Zhang, Jinguang Tong, Hongdong Li |
WACV | 5 |
| 2024 | GridFormer: Residual Dense Transformer with Grid Structure for Image Restoration in Adverse Weather Conditions
Tao Wang 0052, Kaihao Zhang, Ziqian Shao, Wenhan Luo, Björn Stenger, Tong Lu 0002, Tae-Kyun Kim 0001, Wei Liu 0005, Hongdong Li |
Int. J. Comput. Vis. | 9 |
| 2024 | Deep Non-Rigid Structure-From-Motion: A Sequence-to-Sequence Translation PerspectiveabstractDirectly regressing the non-rigid shape and camera pose from the individual 2D frame is ill-suited to the Non-Rigid Structure-from-Motion (NRSfM) problem. This frame-by-frame 3D reconstruction pipeline overlooks the inherent spatial-temporal nature of NRSfM, i.e., reconstructing the 3D sequence from the input 2D sequence. In this paper, we propose to solve deep sparse NRSfM from a sequence-to-sequence translation perspective, where the input 2D keypoints sequence is taken as a whole to reconstruct the corresponding 3D keypoints sequence in a self-supervised manner. First, we apply a shape-motion predictor on the input sequence to obtain an initial sequence of shapes and corresponding motions. Then, we propose the Context Layer, which enables the deep learning framework to effectively impose overall constraints on sequences based on the structural characteristics of non-rigid sequences. The Context Layer constructs modules for imposing the self-expressiveness regularity on non-rigid sequences with multi-head attention (MHA) as the core, together with the use of temporal encoding, both of which act simultaneously to constitute constraints on non-rigid sequences in the deep framework. Experimental results across different datasets such as Human3.6M, CMU Mocap, and InterHand prove the superiority of our framework. The code will be made publicly available. Tong Zhang 0023, Yuchao Dai, Yiran Zhong, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Learning Bilateral Cost Volume for Rolling Shutter Temporal Super-ResolutionabstractRolling shutter temporal super-resolution (RSSR), which aims to synthesize intermediate global shutter (GS) video frames between two consecutive rolling shutter (RS) frames, has made remarkable progress with the development of deep convolutional neural networks over the past years. Existing methods cascade multiple separated networks to sequentially estimate intermediate motion fields and synthesize target GS frames. Nevertheless, they are typically complex, do not facilitate the interaction of complementary motion and appearance information, and suffer from problems such as pixel aliasing or poor interpretation. In this paper, we derive the uniform bilateral motion fields for RS-aware backward warping, which endows our network a more explicit geometric meaning by injecting spatio-temporal consistency information through time-offset embedding. More importantly, we develop a unified, single-stage RSSR pipeline to recover the latent GS video in a coarse-to-fine manner. It first extracts pyramid features from given inputs, and then refines the bilateral motion fields together with the anchor frame until generating the desired output. With the help of our proposed bilateral cost volume, which uses the anchor frame as a common reference to model the correlation with two RS frames, the gradually refined anchor frames not only facilitate intermediate motion estimation, but also compensate for contextual details, making additional frame synthesis or refinement networks unnecessary. Meanwhile, an asymmetric bilateral motion model built on top of the symmetric bilateral motion model further improves the generality and adaptability, yielding better GS video reconstruction performance. Extensive quantitative and qualitative experiments on synthetic and real data demonstrate that our method achieves new state-of-the-art results. Bin Fan 0002, Yuchao Dai, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | LTM-NeRF: Embedding 3D Local Tone Mapping in HDR Neural Radiance FieldabstractRecent advances in Neural Radiance Fields (NeRF) have provided a new geometric primitive for novel view synthesis. High Dynamic Range NeRF (HDR NeRF) can render novel views with a higher dynamic range. However, effectively displaying the scene contents of HDR NeRF on diverse devices with limited dynamic range poses a significant challenge. To address this, we present LTM-NeRF, a method designed to recover HDR NeRF and support 3D local tone mapping. LTM-NeRF allows for the synthesis of HDR views, tone-mapped views, and LDR views under different exposure settings, using only the multi-view multi-exposure LDR inputs for supervision. Specifically, we propose a differentiable Camera Response Function (CRF) module for HDR NeRF reconstruction, globally mapping the scene's HDR radiance to LDR pixels. Moreover, we introduce a Neural Exposure Field (NeEF) to represent the spatially varying exposure time of an HDR NeRF to achieve 3D local tone mapping, for compatibility with various displays. Comprehensive experiments demonstrate that our method can not only synthesize HDR views and exposure-varying LDR views accurately but also render locally tone-mapped views naturally. Xin Huang 0021, Qi Zhang 0029, Hongdong Li, Qing Wang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Weakly-Supervised Depth Estimation and Image Deblurring via Dual-Pixel SensorsabstractDual-pixel (DP) imaging sensors are getting more popularly adopted by modern cameras. A DP camera captures a pair of images in a single snapshot by splitting each pixel in half. Several previous studies show how to recover depth information by treating the DP pair as an approximate stereo pair. However, dual-pixel disparity occurs only in image regions with defocus blur which is unlike classic stereo disparity. Heavy defocus blur in DP pairs affects the performance of depth estimation approaches based on matching. Therefore, we treat the blur removal and the depth estimation as a joint problem. We investigate the formation of the DP pair, which links the blur and depth information, rather than blindly removing the blur effect. We propose a mathematical DP model that can improve depth estimation by the blur. This exploration motivated us to propose our previous work, an end-to-end DDDNet (DP-based Depth and Deblur Network), which jointly estimates depth and restores the image in a supervised fashion. However, collecting the ground-truth (GT) depth map for the DP pair is challenging and limits the depth estimation potential of the DP sensor. Therefore, we propose an extension of the DDDNet, called WDDNet (Weakly-supervised Depth and Deblur Network), which includes an efficient reblur solver that does not require GT depth maps for training. To achieve this, we convert all-in-focus images into supervisory signals for unsupervised depth estimation in our WDDNet. We jointly estimate an all-in-focus image and a disparity map, then use a Reblur and Fstack module to regularize the disparity estimation and image restoration. We conducted extensive experiments on synthetic and real data to demonstrate the competitive performance of our method when compared to state-of-the-art (SOTA) supervised approaches. Liyuan Pan, Richard I. Hartley, Liu Liu 0009, Shah Ariful Hoque Chowdhury, Yan Yang 0011, Hongdong Li, Miaomiao Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Spatial Steerability of GANs via Self-Supervision from DiscriminatorabstractGenerative models make huge progress to the photorealistic image synthesis in recent years. To enable humans to steer the image generation process and customize the output, many works explore the interpretable dimensions of the latent space in GANs. Existing methods edit the attributes of the output image such as orientation or color scheme by varying the latent code along certain directions. However, these methods usually require additional human annotations for each pretrained model, and they mostly focus on editing global attributes. In this work, we propose a self-supervised approach to improve the spatial steerability of GANs without searching for steerable directions in the latent space or requiring extra annotations. Specifically, we design randomly sampled Gaussian heatmaps to be encoded into the intermediate layers of generative models as spatial inductive bias. Along with training the GAN model from scratch, these heatmaps are aligned with the emerging attention of the GAN's discriminator in a self-supervised learning manner. During inference, users can interact with the spatial heatmaps in an intuitive manner, enabling them to edit the output image by adjusting the scene layout, moving, or removing objects. Moreover, we incorporate DragGAN into our framework, which facilitates fine-grained manipulation within a reasonable time and supports a coarse-to-fine editing process. Extensive experiments show that the proposed method not only enables spatial editing over human faces, animal faces, outdoor scenes, and complicated multi-object indoor scenes but also brings improvement in synthesis quality. Lalit Bhagat, Ceyuan Yang, Yinghao Xu 0001, Yujun Shen, Hongdong Li, Bolei Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | MC-Blur: A Comprehensive Benchmark for Image DeblurringabstractBlur artifacts can seriously degrade the visual quality of images, and numerous deblurring methods have been proposed for specific scenarios. However, in most real-world images, blur is caused by different factors, e.g., motion, and defocus. In this paper, we address how other deblurring methods perform in the case of multiple types of blur. For in-depth performance evaluation, we construct a new large-scale multi-cause image deblurring dataset (MC-Blur), including real-world and synthesized blurry images with different blur factors. The images in the proposed MC-Blur dataset are collected using other techniques: averaging sharp images captured by a 1000-fps high-speed camera, convolving Ultra-High-Definition (UHD) sharp images with large-size kernels, adding defocus to images, and real-world blurry images captured by various camera models. Based on the MC-Blur dataset, we conduct extensive benchmarking studies to compare SOTA methods in different scenarios, analyze their efficiency, and investigate the buildataset’s capacity. These benchmarking results provide a comprehensive overview of the advantages and limitations of current deblurring methods, revealing our dataset’s advances. The dataset is available to the public athttps://github.com/HDCVLab/MC-Blur-Dataset. Kaihao Zhang, Tao Wang 0052, Wenhan Luo, Wenqi Ren, Björn Stenger, Wei Liu 0005, Hongdong Li, Ming-Hsuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | SRNSD: Structure-Regularized Night-Time Self-Supervised Monocular Depth Estimation for Outdoor ScenesabstractDeep CNNs have achieved impressive improvements for night-time self-supervised depth estimation form a monocular image. However, the performance degrades considerably compared to day-time depth estimation due to significant domain gaps, low visibility, and varying illuminations between day and night images. To address these challenges, we propose a novel night-time self-supervised monocular depth estimation framework with structure regularization, i.e., SRNSD, which incorporates three aspects of constraints for better performance, including feature and depth domain adaptation, image perspective constraint, and cropped multi-scale consistency loss. Specifically, we utilize adaptations of both feature and depth output spaces for better night-time feature extraction and depth map prediction, along with high- and low-frequency decoupling operations for better depth structure and texture recovery. Meanwhile, we employ an image perspective constraint to enhance the smoothness and obtain better depth maps in areas where the luminosity jumps change. Furthermore, we introduce a simple yet effective cropped multi-scale consistency loss that utilizes consistency among different scales of depth outputs for further optimization, refining the detailed textures and structures of predicted depth. Experimental results on different benchmarks with depth ranges of 40m and 60m, including Oxford RobotCar dataset, nuScenes dataset and CARLA-EPE dataset, demonstrate the superiority of our approach over state-of-the-art night-time self-supervised depth estimation approaches across multiple metrics, proving our effectiveness. Runmin Cong, Chunlei Wu, Xibin Song, Wei Zhang 0021, Sam Kwong, Hongdong Li, Pan Ji |
IEEE Trans. Image Process. | 6 |
| 2024 | BlockFusion: Expandable 3D Scene Generation using Latent Tri-plane ExtrapolationabstractWe present BlockFusion, a diffusion-based model that generates 3D scenes as unit blocks and seamlessly incorporates new blocks to extend the scene. BlockFusion is trained using datasets of 3D blocks that are randomly cropped from complete 3D scene meshes. Through per-block fitting, all training blocks are converted into the hybrid neural fields: with a tri-plane containing the geometry features, followed by a Multi-layer Perceptron (MLP) for decoding the signed distance values. A variational auto-encoder is employed to compress the tri-planes into the latent tri-plane space, on which the denoising diffusion process is performed. Diffusion applied to the latent representations allows for high-quality and diverse 3D scene generation. To expand a scene during generation, one needs only to append empty blocks to overlap with the current scene and extrapolate existing latent tri-planes to populate new blocks. The extrapolation is done by conditioning the generation process with the feature samples from the overlapping tri-planes during the denoising iterations. Latent tri-plane extrapolation produces semantically and geometrically meaningful transitions that harmoniously blend with the existing scene. A 2D layout conditioning mechanism is used to control the placement and arrangement of scene elements. Experimental results indicate that BlockFusion is capable of generating diverse, geometrically consistent and unbounded large 3D scenes with unprecedented high-quality shapes in both indoor and outdoor scenarios. Zhennan Wu, Yang Li 0193, Han Yan 0004, Taizhang Shang, Weixuan Sun, Senbo Wang, Ruikai Cui, Weizhe Liu, Hiroyuki Sato 0002, Hongdong Li, Pan Ji |
ACM Trans. Graph. | 10 |
| 2023 | WildLight: In-the-wild Inverse Rendering with a FlashlightabstractThis paper proposes a practical photometric solution for the challenging problem of in-the-wild inverse rendering under unknown ambient lighting. Our system recovers scene geometry and reflectance using only multi-view images captured by a smartphone. The key idea is to exploit smartphone's built-in flashlight as a minimally controlled light source, and decompose image intensities into two photometric components – a static appearance corresponds to ambient flux, plus a dynamic reflection induced by the moving flashlight. Our method does not require flash/non-flash images to be captured in pairs. Building on the success of neural light fields, we use an off-the-shelf method to capture the ambient reflections, while the flashlight component enables physically accurate photometric constraints to decouple reflectance and illumination. Compared to existing inverse rendering methods, our setup is applicable to non-darkroom environments yet sidesteps the inherent difficulties of explicit solving ambient reflections. We demonstrate by extensive experiments that our method is easy to implement, casual to set up, and consistently outperforms existing in-the-wild inverse rendering techniques. Finally, our neural reconstruction can be easily exported to PBR textured triangle mesh ready for industrial renderers. Our source code and data are released to https://github.com/za-cheng/WildLight. Ziang Cheng, Hongdong Li |
CVPR | 3 |
| 2023 | A Rotation-Translation-Decoupled Solution for Robust and Efficient Visual-Inertial InitializationabstractWe propose a novel visual-inertial odometry (VIO) initialization method, which decouples rotation and translation estimation, and achieves higher efficiency and better robustness. Existing loosely-coupled VIO-initialization methods suffer from poor stability of visual structure-from-motion (SfM), whereas those tightly-coupled methods often ignore the gyroscope bias in the closed-form solution, resulting in limited accuracy. Moreover, the aforementioned two classes of methods are computationally expensive, because 3D point clouds need to be reconstructed simultaneously. In contrast, our new method fully combines inertial and visual measurements for both rotational and translational initialization. First, a rotation-only solution is designed for gyroscope bias estimation, which tightly couples the gyroscope and camera observations. Second, the initial velocity and gravity vector are solved with linear translation constraints in a globally optimal fashion and without reconstructing 3D point clouds. Extensive experiments have demonstrated that our method is 8 ~ 72 times faster (w.r.t. a 10-frame set) than the state-of-the-art methods, and also presents significantly higher robustness and accuracy. The source code is available at https://github.com/boxuLibrary/drt-vio-init. Yijia He, Bo Xu 0022, Zhanpeng Ouyang, Hongdong Li |
CVPR | 4 |
| 2023 | Inverting the Imaging Process by Learning an Implicit Camera ModelabstractRepresenting visual signals with implicit coordinate-based neural networks, as an effective replacement of the traditional discrete signal representation, has gained considerable popularity in computer vision and graphics. In contrast to existing implicit neural representations which focus on modelling the scene only, this paper proposes a novel implicit camera model which represents the physical imaging process of a camera as a deep neural network. We demonstrate the power of this new implicit camera model on two inverse imaging tasks: i) generating all-in-focus photos, and ii) HDR imaging. Specifically, we devise an implicit blur generator and an implicit tone mapper to model the aperture and exposure of the camera's imaging process, respectively. Our implicit camera model is jointly learned together with implicit scene models under multi-focus stack and multi-exposure bracket supervision. We have demonstrated the effectiveness of our new model on a large number of test images and videos, producing accurate and visually appealing all-in-focus and high dynamic range images. In principle, our new implicit neural camera model has the potential to benefit a wide array of other inverse imaging tasks. Xin Huang 0021, Qi Zhang 0029, Hongdong Li, Qing Wang 0006 |
CVPR | 4 |
| 2023 | MEGANE: Morphable Eyeglass and Avatar NetworkabstractEyeglasses play an important role in the perception of identity. Authentic virtual representations of faces can benefit greatly from their inclusion. However, modeling the geometric and appearance interactions of glasses and the face of virtual representations of humans is challenging. Glasses and faces affect each other's geometry at their contact points, and also induce appearance changes due to light transport. Most existing approaches do not capture these physical interactions since they model eyeglasses and faces independently. Others attempt to resolve interactions as a 2D image synthesis problem and suffer from view and temporal inconsistencies. In this work, we propose a 3D compositional morphable model of eyeglasses that accurately incorporates high-fidelity geometric and photometric interaction effects. To support the large variation in eyeglass topology efficiently, we employ a hybrid representation that combines surface geometry and a volumetric representation. Unlike volumetric approaches, our model naturally retains correspondences across glasses, and hence explicit modification of geometry, such as lens insertion and frame deformation, is greatly simplified. In addition, our model is relightable under point lights and natural illumination, supporting high-fidelity rendering of various frame materials, including translucent plastic and metal within a single morphable model. Importantly, our approach models global light transport effects, such as casting shadows between faces and glasses. Our morphable model for eyeglasses can also be fit to novel glasses via inverse rendering. We compare our approach to state-of-the-art methods and demonstrate significant quality improvements. Shunsuke Saito, Tomas Simon, Stephen Lombardi, Hongdong Li, Jason M. Saragih |
CVPR | 5 |
| 2023 | Seeing Through the Glass: Neural 3D Reconstruction of Object Inside a Transparent ContainerabstractIn this paper, we define a new problem of recovering the 3D geometry of an object confined in a transparent enclosure. We also propose a novel method for solving this challenging problem. Transparent enclosures pose challenges of multiple light reflections and refractions at the interface between different propagation media e.g. air or glass. These multiple reflections and refractions cause serious image distortions which invalidate the single viewpoint assumption. Hence the 3D geometry of such objects cannot be reliably reconstructed using existing methods, such as traditional structure from motion or modern neural reconstruction methods. We solve this problem by explicitly modeling the scene as two distinct sub-spaces, inside and outside the transparent enclosure. We use an existing neural reconstruction method (NeuS) that implicitly represents the geometry and appearance of the inner subspace. In order to account for complex light interactions, we develop a hybrid rendering strategy that combines volume rendering with ray tracing. We then recover the underlying geometry and appearance of the model by minimizing the difference between the real and rendered images. We evaluate our method on both synthetic and real data. Experiment results show that our method outperforms the state-of-the-art (SOTA) methods. Codes and data will be available at https://github.com/hirotong/ReNeuS Jinguang Tong, Sundaram Muthu, Fahira A. Maken, Hongdong Li |
CVPR | 5 |
| 2023 | Wide-Angle Rectification via Content-Aware Conformal MappingabstractDespite the proliferation of ultra wide-angle lenses on smartphone cameras, such lenses often come with severe image distortion (e.g. curved linear structure, unnaturally skewed faces). Most existing rectification methods adopt a global warping transformation to undistort the input wideangle image, yet their performances are not entirely satisfactory, leaving many unwanted residue distortions uncorrected or at the sacrifice of the intended wide FoV (field- of-view). This paper proposes a new method to tackle these challenges. Specifically, we derive a locally-adaptive polardomain conformal mapping to rectify a wide-angle image. Parameters of the mapping are found automatically by analyzing image contents via deep neural networks. Experiments on a large number of photos have confirmed the superior performance of the proposed method compared with all available previous methods. Qi Zhang 0029, Hongdong Li, Qing Wang 0006 |
CVPR | 2 |
| 2023 | MB-TaylorFormer: Multi-branch Efficient Transformer Expanded by Taylor Formula for Image DehazingabstractIn recent years, Transformer networks are beginning to replace pure convolutional neural networks (CNNs) in the field of computer vision due to their global receptive field and adaptability to input. However, the quadratic computational complexity of softmax-attention limits the wide application in image dehazing task, especially for high-resolution images. To address this issue, we propose a new Transformer variant, which applies the Taylor expansion to approximate the softmax-attention and achieves linear computational complexity. A multi-scale attention refinement module is proposed as a complement to correct the error of the Taylor expansion. Furthermore, we introduce a multi-branch architecture with multi-scale patch embedding to the proposed Transformer, which embeds features by overlapping deformable convolution of different scales. The design of multi-scale patch embedding is based on three key ideas: 1) various sizes of the receptive field; 2) multi-level semantic information; 3) flexible shapes of the receptive field. Our model, named Multi-branch Transformer expanded by Taylor formula (MB-TaylorFormer), can em-bed coarse to fine features more flexibly at the patch embedding stage and capture long-distance pixel interactions with limited computational cost. Experimental results on several dehazing benchmarks show that MB-TaylorFormer achieves state-of-the-art (SOTA) performance with a light computational burden. The source code and pre-trained models are available at https://github.com/FVL2020/ICCV-2023-MB-TaylorFormer. Yuwei Qiu, Kaihao Zhang, Wenhan Luo, Hongdong Li, Zhi Jin 0002 |
ICCV | 5 |
| 2023 | Boosting 3-DoF Ground-to-Satellite Camera Localization Accuracy via Geometry-Guided Cross-View TransformerabstractImage retrieval-based cross-view localization methods often lead to very coarse camera pose estimation, due to the limited sampling density of the database satellite images. In this paper, we propose a method to increase the accuracy of a ground camera’s location and orientation by estimating the relative rotation and translation between the ground-level image and its matched/retrieved satellite image. Our approach designs a geometry-guided cross-view transformer that combines the benefits of conventional geometry and learnable cross-view transformers to map the ground-view observations to an overhead view. Given the synthesized overhead view and observed satellite feature maps, we construct a neural pose optimizer with strong global information embedding ability to estimate the relative rotation between them. After aligning their rotations, we develop an uncertainty-guided spatial correlation to generate a probability map of the vehicle locations, from which the relative translation can be determined. Experimental results demonstrate that our method significantly outperforms the state-of-the-art. Notably, the likelihood of restricting the vehicle lateral pose to be within 1m of its Ground Truth (GT) value on the cross-view KITTI dataset has been improved from 35.54% to 76.44%, and the likelihood of restricting the vehicle orientation to be within 1° of its GT value has been improved from 19.64% to 99.10%. Yujiao Shi 0002, Akhil Perincherry, Ankit Vora, Hongdong Li |
ICCV | 5 |
| 2023 | Homography Guided Temporal Fusion for Road Line and Marking SegmentationabstractReliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occluded in the presence of moving vehicles, shadow, and glare and (2) highly structured with low intra-class shape variance and overall high appearance consistency. To solve these issues, we propose a Homography Guided Fusion (HomoFusion) module to exploit temporally-adjacent video frames for complementary cues facilitating the correct classification of the partially occluded road lines or markings. To reduce computational complexity, a novel surface normal estimator is proposed to establish spatial correspondences between the sampled frames, allowing the HomoFusion module to perform a pixel-to-pixel attention mechanism in updating the representation of the occluded road lines or markings. Experiments on ApolloScape, a large-scale lane mark segmentation dataset, and ApolloScape Night with artificial simulated night-time road conditions, demonstrate that our method outperforms other existing SOTA lane mark segmentation models with less than 9% of their parameters and computational complexity. We show that exploiting available camera intrinsic data and ground plane assumption for cross-frame correspondence can lead to a light-weight network with significantly improved performances in speed and accuracy. We also prove the versatility of our HomoFusion approach by applying it to the problem of water puddle segmentation and achieving SOTA performance1. Shan Wang 0010, Jiawei Liu 0005, Kaihao Zhang, Wenhan Luo, Yanhao Zhang 0003, Sundaram Muthu, Fahira A. Maken, Hongdong Li |
ICCV | 9 |
| 2023 | View Consistent Purification for Accurate Cross-View LocalizationabstractThis paper proposes a fine-grained self-localization method for outdoor robotics that utilizes a flexible number of onboard cameras and readily accessible satellite images. The proposed method addresses limitations in existing cross-view localization methods that struggle to handle noise sources such as moving objects and seasonal variations. It is the first sparse visual-only method that enhances perception in dynamic environments by detecting view-consistent key points and their corresponding deep features from ground and satellite views, while removing off-the-ground objects and establishing homography transformation between the two views. Moreover, the proposed method incorporates a spatial embedding approach that leverages camera intrinsic and extrinsic information to reduce the ambiguity of purely visual matching, leading to improved feature matching and overall pose estimation accuracy. The method exhibits strong generalization and is robust to environmental changes, requiring only geo-poses as ground truth. Extensive experiments on the KITTI and Ford Multi-AV Seasonal datasets demonstrate that our proposed method outperforms existing state-of-the-art methods, achieving median spatial accuracy errors below 0.5 meters along the lateral and longitudinal directions, and a median orientation accuracy error below 2°1. Shan Wang 0010, Yanhao Zhang 0003, Akhil Perincherry, Ankit Vora, Hongdong Li |
ICCV | 5 |
| 2023 | CircNet: Meshing 3D Point Clouds with Circumcenter Detection
Huan Lei, Ruitao Leng, Liang Zheng 0001, Hongdong Li |
ICLR | 4 |
| 2023 | End-to-End Point Cloud Registration via Rotation Equivariant DescriptorsabstractPoint cloud registration (PCR) aims to recover the rigid transformation between two noisy, unordered point sets. This task is typically tackled by establishing point-wise correspondences, and solving the rigid transformation between the two sets. Since descriptor-based methods find correspondences by matching the feature space distance, a powerful and rotation-robust point feature extractor is critical to the success of this task. Existing methods assume soft rotation invariance/equivariance through the means of training augmentation, rotational discretization or pre-alignment of patches. In contrast, this paper proposes a new method which generates fully rotation invariant and equivariant descriptors by construction. For each keypoint patch, our network extracts not only a rotation invariant descriptor for establishing corre-spondences, but also a rotation equivariant one. The rotation equivariant descriptor allows relative transformation to be directly recovered from a single correspondence pair, unlike standard methods that require three correspondences. This design significantly reduces iteration number of RANSAC and guarantees high registration recall when the inlier ratio of estimated correspondences is low. Extensive experiments have demonstrated that the proposed method outperforms state-of-art methods in the same category even after much fewer RANSAC iterations. Yujiao Shi 0002, Ziang Cheng, Hongdong Li |
IROS | 4 |
| 2023 | EndoSurf: Neural Surface Reconstruction of Deformable Tissues with Stereo Endoscope Videos
Ruyi Zha, Xuelian Cheng, Hongdong Li, Mehrtash Harandi, ZongYuan Ge |
MICCAI (9) | 3 |
| 2023 | Privacy Assessment on Reconstructed Images: Are Existing Evaluation Metrics Faithful to Human Perception?abstractHand-crafted image quality metrics, such as PSNR and SSIM, are commonly used to evaluate model privacy risk under reconstruction attacks. Under these metrics, reconstructed images that are determined to resemble the original one generally indicate more privacy leakage. Images determined as overall dissimilar, on the other hand, indicate higher robustness against attack. However, there is no guarantee that these metrics well reflect human opinions, which offers trustworthy judgement for model privacy leakage. In this paper, we comprehensively study the faithfulness of these hand-crafted metrics to human perception of privacy information from the reconstructed images. On 5 datasets ranging from natural images, faces, to fine-grained classes, we use 4 existing attack methods to reconstruct images from many different classification models and, for each reconstructed image, we ask multiple human annotators to assess whether this image is recognizable. Our studies reveal that the hand-crafted metrics only have a weak correlation with the human evaluation of privacy leakage and that even these metrics themselves often contradict each other. These observations suggest risks of current metrics in the community. To address this potential risk, we propose a learning-based measure called SemSim to evaluate the Semantic Similarity between the original and reconstructed images. SemSim is trained with a standard triplet loss, using an original image as an anchor, one of its recognizable reconstructed images as a positive sample, and an unrecognizable one as a negative. By training on human annotations, SemSim exhibits a greater reflection of privacy leakage on the semantic level. We show that SemSim has a significantly higher correlation with human judgment compared with existing metrics. Moreover, this strong correlation generalizes to unseen datasets, models and attack methods. We envision this work as a milestone for image quality evaluation closer to the human level. The project webpage can be accessed at https://sites.google.com/view/semsim. Xiaoxiao Sun 0002, Nidham Gazagnadou, Vivek Sharma 0001, Lingjuan Lyu, Hongdong Li, Liang Zheng 0001 |
NeurIPS | 5 |
| 2023 | DeepSimHO: Stable Pose Estimation for Hand-Object Interaction via Physics SimulationabstractThis paper addresses the task of 3D pose estimation for a hand interacting with an object from a single image observation. When modeling hand-object interaction, previous works mainly exploit proximity cues, while overlooking the dynamical nature that the hand must stably grasp the object to counteract gravity and thus preventing the object from slipping or falling. These works fail to leverage dynamical constraints in the estimation and consequently often produce unstable results. Meanwhile, refining unstable configurations with physics-based reasoning remains challenging, both by the complexity of contact dynamics and by the lack of effective and efficient physics inference in the data-driven learning framework. To address both issues, we present DeepSimHO: a novel deep-learning pipeline that combines forward physics simulation and backward gradient approximation with a neural network. Specifically, for an initial hand-object pose estimated by a base network, we forward it to a physics simulator to evaluate its stability. However, due to non-smooth contact geometry and penetration, existing differentiable simulators can not provide reliable state gradient. To remedy this, we further introduce a deep network to learn the stability evaluation process from the simulator, while smoothly approximating its gradient and thus enabling effective back-propagation. Extensive experiments show that our method noticeably improves the stability of the estimation and achieves superior efficiency over test-time optimization. The code is available at https://github.com/rongakowang/DeepSimHO. Wei Mao 0001, Hongdong Li |
NeurIPS | 3 |
| 2023 | Interacting Hand-Object Pose Estimation via Dense Mutual Attentionabstract2D hand-object pose estimation is the key to the success of many computer vision applications. The main focus of this task is to effectively model the interaction between the hand and an object. To this end, existing works either rely on interaction constraints in a computationally-expensive iterative optimization, or consider only a sparse correlation between sampled hand and object keypoints. In contrast, we propose a novel dense mutual attention mechanism that is able to model fine-grained dependencies between the hand and the object. Specifically, we first construct the hand and object graphs according to their mesh structures. For each hand node, we aggregate features from every object node by the learned attention and vice versa for each object node. Thanks to such dense mutual attention, our method is able to produce physically plausible poses with high quality and real-time inference speed. Extensive quantitative and qualitative experiments on large benchmark datasets show that our method outperforms state-of-the-art methods. The code is available at https://github.com/rongakowang/DenseMutualAttention.git. Wei Mao 0001, Hongdong Li |
WACV | 3 |
| 2023 | Screening prognostic markers for hepatocellular carcinoma based on pyroptosis-related lncRNA pairsabstractBACKGROUND: Pyroptosis is closely related to cancer prognosis. In this study, we tried to construct an individualized prognostic risk model for hepatocellular carcinoma (HCC) based on within-sample relative expression orderings (REOs) of pyroptosis-related lncRNAs (PRlncRNAs). METHODS: RNA-seq data of 343 HCC samples derived from The Cancer Genome Atlas (TCGA) database were analyzed. PRlncRNAs were detected based on differentially expressed lncRNAs between sample groups clustered by 40 reported pyroptosis-related genes (PRGs). Univariate Cox regression was used to screen out prognosis-related PRlncRNA pairs. Then, based on REOs of prognosis-related PRlncRNA pairs, a risk model for HCC was constructed by combining LASSO and stepwise multivariate Cox regression analysis. Finally, a prognosis-related competing endogenous RNA (ceRNA) network was built based on information about lncRNA-miRNA-mRNA interactions derived from the miRNet and TargetScan databases. RESULTS: (FC)|> 1 and FDR < 5%). Among them, 83 PRlncRNA pairs showed significant associations between their REOs within HCC samples and overall survival (Univariate Cox regression, p < 0.005). An optimal 11-PRlncRNA-pair prognostic risk model was constructed for HCC. The areas under the curves (AUCs) of time-dependent receiver operating characteristic (ROC) curves of the risk model for 1-, 3-, and 5-year survival were 0.737, 0.705, and 0.797 in the validation set, respectively. Gene Set Enrichment Analysis showed that inflammation-related interleukin signaling pathways were upregulated in the predicted high-risk group (p < 0.05). Tumor immune infiltration analysis revealed a higher abundance of regulatory T cells (Tregs) and M2 macrophages and a lower abundance of CD8 + T cells in the high-risk group, indicating that excessive pyroptosis might occur in high-risk patients. Finally, eleven lncRNA-miRNA-mRNA regulatory axes associated with pyroptosis were established. CONCLUSION: Our risk model allowed us to determine the robustness of the REO-based PRlncRNA prognostic biomarkers in the stratification of HCC patients at high and low risk. The model is also helpful for understanding the molecular mechanisms between pyroptosis and HCC prognosis. High-risk patients may have excessive pyroptosis and thus be less sensitive to immune therapy. Tong Wu 0004, Fengyuan Luo, Tao Hu 0011, Guini Hong, Hongdong Li |
BMC Bioinform. | 8 |
| 2023 | Distance Based Image Classification: A solution to generative classification's conundrum?
Wen-Yan Lin, Bing Tian Dai, Hongdong Li |
Int. J. Comput. Vis. | 4 |
| 2023 | Event-guided Multi-patch Network with Self-supervision for Non-uniform Motion Deblurring
Limeng Zhang, Yuchao Dai, Hongdong Li, Piotr Koniusz |
Int. J. Comput. Vis. | 4 |
| 2023 | Rolling Shutter Inversion: Bring Rolling Shutter Images to High Framerate Global Shutter VideoabstractA single rolling-shutter (RS) image may be viewed as a row-wise combination of a sequence of global-shutter (GS) images captured by a (virtual) moving GS camera within the exposure duration. Although rolling-shutter cameras are widely used, the RS effect causes obvious image distortion especially in the presence of fast camera motion, hindering downstream computer vision tasks. In this paper, we propose to invert the rolling-shutter image capture mechanism, i.e., recovering a continuous high framerate global-shutter video from two time-consecutive RS frames. We call this task the RS temporal super-resolution (RSSR) problem. The RSSR is a very challenging task, and to our knowledge, no practical solution exists to date. This paper presents a novel deep-learning based solution. By leveraging the multi-view geometry relationship of the RS imaging process, our learning based framework successfully achieves high framerate GS generation. Specifically, three novel contributions can be identified: (i) novel formulations for bidirectional RS undistortion flows under constant velocity as well as constant acceleration motion model. (ii) a simple linear scaling operation, which bridges the RS undistortion flow and regular optical flow. (iii) a new mutual conversion scheme between varying RS undistortion flows that correspond to different scanlines. Our method also exploits the underlying spatial-temporal geometric relationships within a deep learning framework, where no additional supervision is required beyond the necessary middle-scanline GS image. Building upon these contributions, this paper represents the very first rolling-shutter temporal super-resolution deep-network that is able to recover high framerate global-shutter videos from just two RS frames. Extensive experimental results on both synthetic and real data show that our proposed method can produce high-quality GS image sequences with rich details, outperforming the state-of-the-art methods. Bin Fan 0002, Yuchao Dai, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Accurate 3-DoF Camera Geo-Localization via Ground-to-Satellite Image MatchingabstractWe address the problem of ground-to-satellite image geo-localization, that is, estimating the camera latitude, longitude and orientation (azimuth angle) by matching a query image captured at the ground level against a large-scale database with geotagged satellite images. Our prior arts treat the above task as pure image retrieval by selecting the most similar satellite reference image matching the ground-level query image. However, such an approach often produces coarse location estimates because the geotag of the retrieved satellite image only corresponds to the image center while the ground camera can be located at any point within the image. To further consolidate our prior research finding, we present a novel geometry-aware geo-localization method. Our new method is able to achieve the fine-grained location of a query image, up to pixel size precision of the satellite image, once its coarse location and orientation have been determined. Moreover, we propose a new geometry-aware image retrieval pipeline to improve the coarse localization accuracy. Apart from a polar transform in our conference work, this new pipeline also maps satellite image pixels to the ground-level plane in the ground-view via a geometry-constrained projective transform to emphasize informative regions, such as road structures, for cross-view geo-localization. Extensive quantitative and qualitative experiments demonstrate the effectiveness of our newly proposed framework. We also significantly improve the performance of coarse localization results compared to the state-of-the-art in terms of location recalls. Yujiao Shi 0002, Xin Yu 0002, Liu Liu 0009, Dylan Campbell, Piotr Koniusz, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | TUSR-Net: Triple Unfolding Single Image Dehazing With Self-Regularization and Dual Feature to Pixel AttentionabstractSingle image dehazing is a challenging and ill-posed problem due to severe information degeneration of images captured in hazy conditions. Remarkable progresses have been achieved by deep-learning based image dehazing methods, where residual learning is commonly used to separate the hazy image into clear and haze components. However, the nature of low similarity between haze and clear components is commonly neglected, while the lack of constraint of contrastive peculiarity between the two components always restricts the performance of these approaches. To deal with these problems, we propose an end-to-end self-regularized network (TUSR-Net) which exploits the contrastive peculiarity of different components of the hazy image, i.e, self-regularization (SR). In specific, the hazy image is separated into clear and hazy components and constraint between different image components, i.e., self-regularization, is leveraged to pull the recovered clear image closer to groundtruth, which largely promotes the performance of image dehazing. Meanwhile, an effective triple unfolding framework combined with dual feature to pixel attention is proposed to intensify and fuse the intermediate information in feature, channel and pixel levels, respectively, thus features with better representational ability can be obtained. Our TUSR-Net achieves better trade-off between performance and parameter size with weight-sharing strategy and is much more flexible. Experiments on various benchmarking datasets demonstrate the superiority of our TUSR-Net over state-of-the-art single image dehazing methods. Xibin Song, Dingfu Zhou, Wei Li 0143, Yuchao Dai, Zhelun Shen, Liangjun Zhang, Hongdong Li |
IEEE Trans. Image Process. | 7 |
| 2023 | Multi-Level Second-Order Few-Shot LearningabstractWe propose a Multi-level Second-order (MlSo) few-shot learning network for supervised or unsupervised few-shot image classification and few-shot action recognition. We leverage so-called power-normalized second-order base learner streams combined with features that express multiple levels of visual abstraction, and we use self-supervised discriminating mechanisms. As Second-order Pooling (SoP) is popular in image recognition, we employ its basic element-wise variant in our pipeline. The goal of multi-level feature design is to extract feature representations at different layer-wise levels of CNN, realizing several levels of visual abstraction to achieve robust few-shot learning. As SoP can handle convolutional feature maps of varying spatial sizes, we also introduce image inputs at multiple spatial scales into MlSo. To exploit the discriminative information from multi-level and multi-scale features, we develop a Feature Matching (FM) module that reweights their respective branches. We also introduce a self-supervised step, which is a discriminator of the spatial level and the scale of abstraction. Our pipeline is trained in an end-to-end manner. With a simple architecture, we demonstrate respectable results on standard datasets such as Omniglot, mini-ImageNet, tiered-ImageNet, Open MIC, fine-grained datasets such as CUB Birds, Stanford Dogs and Cars, and action recognition datasets such as HMDB51, UCF101, and mini-MIT. Hongdong Li, Piotr Koniusz |
IEEE Trans. Multim. | 2 |
| 2023 | Structure-to-Shape Aortic 3-D Deformation Reconstruction for Endovascular InterventionsabstractFluoroscopy-guided endovascular interventions by using X-ray images are challenging. The catheter needs to be manipulated precisely inside the aorta, while only 2-D views from the X-ray fluoroscopy are currently used to help the surgeons. Because the catheter is operated in a 3-D space, a visualization of the deforming 3-D aorta will be useful as guidance for catheter manipulation. Existing 3-D reconstruction methods fall short in only focusing on the deformation reconstruction of the aortic 3-D centerline, or using additional prior knowledge of 3-D catheter position for estimating the aortic 3-D deformation. In this article, we propose a novel framework that reconstructs the aortic 3-D deformation by fusing a preoperative 3-D model and two intraoperative X-ray images. Different from existing methods, the proposed framework reconstructs aortic deformation using a coarse-to-fine pipeline by first reconstructing the aortic 3-D centerline and then reconstructing the 3-D shape. To obtain the accurate features for the fluoroscopic-based 3-D reconstruction, we extract semantic features from the X-ray images, and compute the distance field to efficiently calculate the 3-D–2-D nonrigid correspondence. Nonlinear least squares optimization is used to solve the deformation of both centerline and shape. The proposed framework is validated using phantom and patient datasets, whose results demonstrate improved efficiency and accuracy compared with the existing methods. This framework provides a valuable clinical tool for endovascular interventions. Yanhao Zhang 0003, Raphael Falque, Liang Zhao 0003, Yongbo Chen 0001, Shoudong Huang, Hongdong Li |
IEEE Trans. Robotics | 6 |
| 2022 | Transcribing Natural Languages for the Deaf via Neural Editing ProgramsabstractThis work studies the task of glossification, of which the aim is to em transcribe natural spoken language sentences for the Deaf (hard-of-hearing) community to ordered sign language glosses. Previous sequence-to-sequence language models trained with paired sentence-gloss data often fail to capture the rich connections between the two distinct languages, leading to unsatisfactory transcriptions. We observe that despite different grammars, glosses effectively simplify sentences for the ease of deaf communication, while sharing a large portion of vocabulary with sentences. This has motivated us to implement glossification by executing a collection of editing actions, e.g. word addition, deletion, and copying, called editing programs, on their natural spoken language counterparts. Specifically, we design a new neural agent that learns to synthesize and execute editing programs, conditioned on sentence contexts and partial editing results. The agent is trained to imitate minimal editing programs, while exploring more widely the program space via policy gradients to optimize sequence-wise transcription quality. Results show that our approach outperforms previous glossification models by a large margin, improving the BLEU-4 score from 16.45 to 18.89 on RWTH-PHOENIX-WEATHER-2014T and from 18.38 to 21.30 on CSL-Daily. Dongxu Li 0003, Liu Liu 0009, Yiran Zhong, Lars Petersson, Hongdong Li |
AAAI | 7 |
| 2022 | Neural Plenoptic Sampling: Learning Light-Field from Thousands of Imaginary Eyes
Yujiao Shi 0002, Hongdong Li |
ACCV (1) | 3 |
| 2022 | CVLNet: Cross-view Semantic Correspondence Learning for Video-Based Camera Localization
Yujiao Shi 0002, Xin Yu 0002, Shan Wang 0010, Hongdong Li |
ACCV (1) | 4 |
| 2022 | HDR-NeRF: High Dynamic Range Neural Radiance FieldsabstractWe present High Dynamic Range Neural Radiance Fields (HDR-NeRF) to recover an HDR radiance field from a set of low dynamic range (LDR) views with different exposures. Using the HDR-NeRF, we are able to generate both novel HDR views and novel LDR views under different exposures. The key to our method is to model the simplified physical imaging process, which dictates that the radiance of a scene point transforms to a pixel value in the LDR image with two implicit functions: a radiance field and a tone mapper. The radiance field encodes the scene radiance (values vary from 0 to$+\infty$), which outputs the density and radiance of a ray by giving corresponding ray origin and ray direction. The tone mapper models the mapping process that a ray hitting on the camera sensor becomes a pixel value. The color of the ray is predicted by feeding the radiance and the corresponding exposure time into the tone mapper. We use the classic volume rendering technique to project the output radiance, colors and densities into HDR and LDR images, while only the input LDR images are used as the supervision. We collect a new forward-facing HDR dataset to evaluate the proposed method. Experimental results on synthetic and real-world scenes validate that our method can not only accurately control the exposures of synthesized views but also render views with a high dynamic range. Xin Huang 0021, Qi Zhang 0029, Hongdong Li, Xuan Wang 0009, Qing Wang 0006 |
CVPR | 4 |
| 2022 | Align and Prompt: Video-and-Language Pre-training with Entity PromptsabstractYidco-and-language pre-training has shown promising improvements on various downstream tasks. Most previous methods capture cross-modal interactions with a standard transformer-based multimodal encoder, not fully addressing the misalignment between unimodal video and text features. Besides, learning finegrained visual-language alignment usually requires off-the-shelf object detectors to provide object information, which is bottlenecked by the detector's limited vocabulary and expensive computation cost. In this paper, we propose Align and Prompt: a new video-and-language pre-training framework (AlPro), which operates on sparsely-sampled video frames and achieves more effective cross-modal alignment without explicit object detectors. First, we introduce a video-text contrastive (VTC) loss to align unimodal video-text features at the instance level, which eases the modeling of cross-modal interactions. Then, we propose a novel visually-grounded pre-training task, prompting entity modeling (PEM), which learns finegrained alignment between visual region and text entity via an entity prompter module in a self-supervised way. Finally, we pretrain the video-and-language transformer models on large webly-source video-text pairs using the proposed VTC and PEM losses as well as two standard losses of masked language modeling (MLM) and video-text matching (VTM). The resulting pre-trained model achieves state-of-the-art performance on both text-video retrieval and videoQA, outperforming prior work by a substantial margin. Implementation and pre-trained models are available at https://github.com/salesforce/ALPRO. Dongxu Li 0003, Junnan Li 0001, Hongdong Li, Juan Carlos Niebles, Steven C. H. Hoi |
CVPR | 3 |
| 2022 | Neural Reflectance for Shape Recovery with Shadow HandlingabstractThis paper aims at recovering the shape of a scene with unknown, non-Lambertian, and possibly spatially-varying surface materials. When the shape of the object is highly complex and that shadows cast on the surface, the task becomes very challenging. To overcome these challenges, we propose a coordinate-based deep MLP (multilayer perceptron) to parameterize both the unknown 3D shape and the unknown reflectance at every surface point. This network is able to leverage the observed photometric variance and shadows on the surface, and recover both surface shape and general non-Lambertian reflectance. We explicitly predict cast shadows, mitigating possible artifacts on these shadowing regions, leading to higher estimation accuracy. Our framework is entirely self-supervised, in the sense that it requires neither ground truth shape nor BRDF. Tests on real-world images demonstrate that our method outperform existing methods by a significant margin. Thanks to the small size of the MLP-net, our method is an order of magnitude faster than previous CNN-based methods. Hongdong Li |
CVPR | 2 |
| 2022 | Beyond Cross-view Image Retrieval: Highly Accurate Vehicle Localization Using Satellite ImageabstractThis paper addresses the problem of vehicle-mounted camera localization by matching a ground-level image with an overhead-view satellite map. Existing methods often treat this problem as cross-view image retrieval, and use learned deep features to match the ground-level query im-age to a partition (e.g., a small patch) of the satellite map. By these methods, the localization accuracy is limited by the partitioning density of the satellite map (often in the order of tens meters). Departing from the conventional wisdom of image retrieval, this paper presents a novel solution that can achieve highly-accurate localization. The key idea is to formulate the task as pose estimation and solve it by neural-net based optimization. Specifically, we design a two-branch CNN to extract robust features from the ground and satellite images, respectively. To bridge the vast cross-view domain gap, we resort to a Geometry Projection module that projects features from the satellite map to the ground-view, based on a relative camera pose. Aiming to minimize the differences between the projected features and the observed features, we employ a differentiable Levenberg-Marquardt (LM) module to search for the optimal camera pose iteratively. The entire pipeline is differen-tiable and runs end-to-end. Extensive experiments on standard autonomous vehicle localization datasets have confirmed the superiority of the proposed method. Notably, e.g., starting from a coarse estimate of camera location within a wide region of 40m × 40m, with an 80% likelihood our method quickly reduces the lateral location error to be within 5m on a new KITTI cross-view dataset. Yujiao Shi 0002, Hongdong Li |
CVPR | 2 |
| 2022 | Improving GAN Equilibrium by Raising Spatial AwarenessabstractThe success of Generative Adversarial Networks (GANs) is largely built upon the adversarial training between a generator (G) and a discriminator (D). They are expected to reach a certain equilibrium where D cannot distinguish the generated images from the real ones. However, such an equilibrium is rarely achieved in practical GAN training, instead, D almost always surpasses G. We attribute one of its sources to the information asymmetry between D and G. We observe that D learns its own visual attention when determining whether an image is real or fake, but G has no explicit clue on which regions to focus on for a particular synthesis. To alleviate the issue of D dominating the competition in GANs, we aim to raise the spatial awareness of G. Randomly sampled multi-level heatmaps are encoded into the intermediate layers of G as an inductive bias. Thus G can purposefully improve the synthesis of certain image regions. We further propose to align the spatial awareness of G with the attention map induced from D. Through this way we effectively lessen the information gap between D and G. Extensive results show that our method pushes the two-player game in GANs closer to the equilibrium, leading to a better synthesis performance. As a byproduct, the intro-duced spatial awareness facilitates interactive editing over the output synthesis. Demo video and code are available at https://genforce.github.io/eqgan-sa/ Ceyuan Yang, Yinghao Xu 0001, Yujun Shen, Hongdong Li, Bolei Zhou |
CVPR | 5 |
| 2022 | Blind Image Decomposition
Junlin Han, Weihao Li 0005, Pengfei Fang, Chunyi Sun, Mohammad Ali Armin, Lars Petersson, Hongdong Li |
ECCV (18) | 8 |
| 2022 | Self-calibrating Photometric Stereo by Neural Inverse Rendering
Hongdong Li |
ECCV (2) | 2 |
| 2022 | You Only Cut Once: Boosting Data Augmentation with a Single CutabstractWe present You Only Cut Once (YOCO) for performing data augmentations. YOCO cuts one image into two pieces and performs data augmentations individually within each piece. Applying YOCO improves the diversity of the augmentation per sample and encourages neural networks to recognize objects from partial information. YOCO enjoys the properties of parameter-free, easy usage, and boosting almost all augmentations for free. Thorough experiments are conducted to evaluate its effectiveness. We first demonstrate that YOCO can be seamlessly applied to varying data augmentations, neural network architectures, and brings performance gains on CIFAR and ImageNet classification tasks, sometimes surpassing conventional image-level augmentation by large margins. Moreover, we show YOCO benefits contrastive pre-training toward a more powerful representation that can be better transferred to multiple downstream tasks. Finally, we study a number of variants of YOCO and empirically analyze the performance for respective settings. Junlin Han, Pengfei Fang, Weihao Li 0005, Mohammad Ali Armin, Ian D. Reid 0001, Lars Petersson, Hongdong Li |
ICML | 8 |
| 2022 | NAF: Neural Attenuation Fields for Sparse-View CBCT Reconstruction
Ruyi Zha, Yanhao Zhang 0003, Hongdong Li |
MICCAI (6) | 3 |
| 2022 | Beyond Monocular Deraining: Parallel Stereo Deraining Network Via Semantic Prior
Kaihao Zhang, Wenhan Luo, Yanjiang Yu, Wenqi Ren, Fang Zhao 0006, Lin Ma 0002, Wei Liu 0005, Hongdong Li |
Int. J. Comput. Vis. | 9 |
| 2022 | Deep Image Deblurring: A Survey
Kaihao Zhang, Wenqi Ren, Wenhan Luo, Wei-Sheng Lai, Björn Stenger, Ming-Hsuan Yang 0001, Hongdong Li |
Int. J. Comput. Vis. | 7 |
| 2022 | Displacement-Invariant Cost Computation for Stereo MatchingabstractAbstract Although deep learning-based methods have dominated stereo matching leaderboards by yielding unprecedented disparity accuracy, their inference time is typically slow, i.e., less than 4 FPS for a pair of 540p images. The main reason is that the leading methods employ time-consuming 3D convolutions applied to a 4D feature volume. A common way to speed up the computation is to downsample the feature volume, but this loses high-frequency details. To overcome these challenges, we propose a displacement-invariant cost computation module to compute the matching costs without needing a 4D feature volume. Rather, costs are computed by applying the same 2D convolution network on each disparity-shifted feature map pair independently. Unlike previous 2D convolution-based methods that simply perform context mapping between inputs and disparity maps, our proposed approach learns to match features between the two images. We also propose an entropy-based refinement strategy to refine the computed disparity map, which further improves the speed by avoiding the need to compute a second disparity map on the right image. Extensive experiments on standard datasets (SceneFlow, KITTI, ETH3D, and Middlebury) demonstrate that our method achieves competitive accuracy with much less inference time. On typical image sizes (e.g., $$540\times 960$$ 540 × 960 ), our method processes over 100 FPS on a desktop GPU, making our method suitable for time-critical applications such as autonomous driving. We also show that our approach generalizes well to unseen datasets, outperforming 4D-volumetric methods. We will release the source code to ensure the reproducibility. Yiran Zhong, Charles T. Loop, Wonmin Byeon, Stanley T. Birchfield, Yuchao Dai, Kaihao Zhang, Alexey Kamenev, Thomas M. Breuel, Hongdong Li, Jan Kautz |
Int. J. Comput. Vis. | 9 |
| 2022 | Robust and Efficient Estimation of Relative Pose for Cameras on Selfie SticksabstractTaking selfies has become one of the major photographic trends of our time. In this study, we focus on the selfie stick, on which a camera is mounted to take selfies. We observe that a camera on a selfie stick typically travels through a particular type of trajectory around a sphere. Based on this finding, we propose a robust, efficient, and optimal estimation method for relative camera pose between two images captured by a camera mounted on a selfie stick. We exploit the special geometric structure of camera motion constrained by a selfie stick and define this motion as spherical joint motion. Utilizing a novel parametrization and calibration scheme, we demonstrate that the pose estimation problem can be reduced to a 3-degrees of freedom (DoF) search problem, instead of a generic 6-DoF problem. This facilitates the derivation of an efficient branch-and-bound optimization method that guarantees a global optimal solution, even in the presence of outliers. Furthermore, as a simplified case of spherical joint motion, we introduce selfie motion, which has a fewer number of DoF than spherical joint motion. We validate the performance and guaranteed optimality of our method on both synthetic and real-world data. Additionally, we demonstrate the applicability of the proposed method for two applications: refocusing and stylization. Kyungdon Joo, Hongdong Li, Tae-Hyun Oh, In-So Kweon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Shell Theory: A Statistical Model of RealityabstractThe foundational assumption of machine learning is that the data under consideration is separable into classes; while intuitively reasonable, separability constraints have proven remarkably difficult to formulate mathematically. We believe this problem is rooted in the mismatch between existing statistical techniques and commonly encountered data; object representations are typically high dimensional but statistical techniques tend to treat high dimensions a degenerate case. To address this problem, we develop a dedicated statistical framework for machine learning in high dimensions. The framework derives from the observation that object relations form a natural hierarchy; this leads us to model objects as instances of a high dimensional, hierarchal generative processes. Using a distance based statistical technique, also developed in this paper, we show that in such generative processes, instances of each process in the hierarchy, are almost-always encapsulated by a distinctive-shell that excludes almost-all other instances. The result is shell theory, a statistical machine learning framework in which separability constraints (distinctive-shells) are formally derived from the assumed generative process. Wen-Yan Lin, Changhao Ren, Ngai-Man Cheung, Hongdong Li, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Geometry-Guided Street-View Panorama Synthesis From Satellite ImageryabstractThis paper presents a new approach for synthesizing a novel street-view panorama given a satellite image, as if captured from the geographical location at the center of the satellite image. Existing works approach this as an image generation problem, adopting generative adversarial networks to implicitly learn the cross-view transformations, but ignore the geometric constraints. In this paper, we make the geometric correspondences between the satellite and street-view images explicit so as to facilitate the transfer of information between domains. Specifically, we observe that when a 3D point is visible in both views, and the height of the point relative to the camera is known, there is a deterministic mapping between the projected points in the images. Motivated by this, we develop a novel satellite to street-view projection (S2SP) module which learns the height map and projects the satellite image to the ground-level viewpoint, explicitly connecting corresponding pixels. With these projected satellite images as input, we next employ a generator to synthesize realistic street-view panoramas that are geometrically consistent with the satellite images. Our S2SP module is differentiable and the whole framework is trained in an end-to-end manner. Extensive experimental results on two cross-view benchmark datasets demonstrate that our method generates more accurate and consistent images than existing approaches. Yujiao Shi 0002, Dylan Campbell, Xin Yu 0002, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Ray-Space Epipolar Geometry for Light Field CamerasabstractLight field essentially represents rays in space. The epipolar geometry between two light fields is an important relationship that captures ray-ray correspondences and relative configuration of two views. Unfortunately, so far little work has been done in deriving a formal epipolar geometry model that is specifically tailored for light field cameras. This is primarily due to the high-dimensional nature of the ray sampling process with a light field camera. This paper fills in this gap by developing a novel ray-space epipolar geometry which intrinsically encapsulates the complete projective relationship between two light fields, while the generalized epipolar geometry which describes relationship of normalized light fields is the specialization of the proposed model to calibrated cameras. With Plücker parameterization, we propose the ray-space projection model involving a 6×6 ray-space intrinsic matrix for ray sampling of light field camera. Ray-space fundamental matrix and its properties are then derived to constrain ray-ray correspondences for general and special motions. Finally, based on ray-space epipolar geometry, we present two novel algorithms, one for fundamental matrix estimation, and the other for calibration. Experiments on synthetic and real data have validated the effectiveness of ray-space epipolar geometry in solving 3D computer vision tasks with light field cameras. Qi Zhang 0029, Qing Wang 0006, Hongdong Li, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Context-Aware 3D Object Detection From a Single Image in Autonomous DrivingabstractCamera sensors have been widely used in Driver-Assistance and Autonomous Driving Systems due to their rich texture information. Recently, with the development of deep learning techniques, many approaches have been proposed to detect objects in 3D from a single frame, however, there is still much room for improvement. In this paper, we generally review the recently proposed state-of-the-art monocular-based 3D object detection approaches first. Based on the analysis of the disadvantage of previous center-based frameworks, a novel feature aggregation strategy has been proposed to boost the 3D object detection by exploring the context information. Specifically, an Instance-Guided Spatial Attention (IGSA) module is proposed to collect the local instance information and the Channel-Wise Feature Attention (CWFA) module is employed for aggregating the global context information. In addition, an instance-guided object regression strategy is also proposed to alleviate the influence of center location prediction uncertainty in the inference process. Finally, the proposed approach has been verified on the public 3D object detection benchmark. The experimental results show that the proposed approach can significantly boost the performance of the baseline method on both 3D detection and 2D Bird’s-Eye View among all three categories. Furthermore, our method outperforms all the monocular-based methods (even these trained with depth as auxiliary inputs) and achieves state-of-the-art performance on the KITTI benchmark. Dingfu Zhou, Xibin Song, Yuchao Dai, Hongdong Li, Liangjun Zhang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | WAFP-Net: Weighted Attention Fusion Based Progressive Residual Learning for Depth Map Super-ResolutionabstractDespite the remarkable progresses achieved in depth map super-resolution (DSR), it remains a major challenge to tackle with real-world degradation of low-resolution (LR) depth maps. Synthetic datasets are mainly used in existing DSR approaches, which is quite different from what would get from a real depth sensor. Besides, the enhancements of features in existing DSR approaches are not sufficiently enough, which also limit the performance. To alleviate these problems, we first propose two types of degradation models to describe the generation of LR depth maps, including bi-cubic down-sampling with noise and interval down-sampling, and different DSR models are learned correspondingly. Then, we propose a weighted attention fusion strategy that is embedded into a progressive residual learning framework, which guarantees that the high-resolution (HR) depth maps can be well recovered in a coarse-to-fine manner. The weighted attention fusion strategy can enhance the features with abundant high-frequency components in both global and local manners, thus better HR depth maps can be expected. Besides, to re-use the effective information in the progressive process sufficiently, a multi-stage fusion module is combined into the proposed framework, and the Total Generalized Variation (TGV) regularization and input loss are exploited to further improve the performance of our method. Extensive experiments of different benchmarks demonstrate the superiority of our approach over the state-of-the-art (SOTA) approaches. Xibin Song, Dingfu Zhou, Wei Li 0111, Yuchao Dai, Liu Liu 0009, Hongdong Li, Ruigang Yang, Liangjun Zhang |
IEEE Trans. Multim. | 6 |
| 2022 | Disentangled Feature Networks for Facial Portrait and Caricature GenerationabstractFacial portrait is an artistic form which draws faces by emphasizing discriminative or prominent parts of faces via various kinds of drawing tools. However, the complex interplay between the different facial factors, such as facial parts, background, and drawing styles, and the significant domain gap between natural facial images and their portrait counterparts makes the task challenging. In this paper, a flexible four-stream Disentangled Feature Networks (DFN) is proposed to learn disentangled feature representation of different facial factors and generate plausible portraits with reasonable exaggerations and richness in style. Four factors are encoded as embedding features, and combined to reconstruct facial portraits. Meanwhile, to make the process fully automatic (without manually specifying either portrait style or exaggerating form), we propose a new Adversarial Portrait Mapping Module (APMM) to map noise to the embedding feature space, as proxies for portrait style and exaggerating. Thanks to the proposedDFNandAPMM, we are able to manipulate the portrait style and facial geometric structures to generate a large number of portraits. Extensive experiments on two public datasets show that our proposed methods can generate a diverse set of artistic portraits. Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Wenqi Ren, Hongdong Li |
IEEE Trans. Multim. | 5 |
| 2021 | PluckerNet: Learn To Register 3D Line ReconstructionsabstractAligning two partially-overlapped 3D line reconstructions in Euclidean space is challenging, as we need to simultaneously solve correspondences and relative pose between line reconstructions. This paper proposes a neural network based method and it has three modules connected in sequence: (i) a Multilayer Perceptron (MLP) based network takes Plücker representations of lines as inputs, to extract discriminative line-wise features and matchabilities (how likely each line is going to have a match), (ii) an Optimal Transport (OT) layer takes two-view line-wise features and matchabilities as inputs to estimate a 2D joint probability matrix, with each item describes the matchness of a line pair, and (iii) line pairs with Top-K matching probabilities are fed to a 2-line minimal solver in a RANSAC framework to estimate a six Degree-of-Freedom (6-DoF) rigid transformation. Experiments on both indoor and outdoor datasets show that registration (rotation and translation) precision of our method outperforms baselines significantly. Liu Liu 0009, Hongdong Li, Haodong Yao, Ruyi Zha |
CVPR | 2 |
| 2021 | Multi-View 3D Reconstruction of a Texture-Less Smooth Surface of Unknown Generic ReflectanceabstractRecovering the 3D geometry of a purely texture-less object with generally unknown surface reflectance (e.g. non-Lambertian) is regarded as a challenging task in multi-view reconstruction. The major obstacle revolves around establishing cross-view correspondences where photometric constancy is violated. This paper proposes a simple and practical solution to overcome this challenge based on a co-located camera-light scanner device. Unlike existing solutions, we do not explicitly solve for correspondence. Instead, we argue the problem is generally well-posed by multi-view geometrical and photometric constraints, and can be solved from a small number of input views. We formulate the reconstruction task as a joint energy minimization over the surface geometry and reflectance. Despite this energy is highly non-convex, we develop an optimization algorithm that robustly recovers globally optimal shape and reflectance even from a random initialization. Extensive experiments on both simulated and real data have validated our method, and possible future extensions are discussed. Ziang Cheng, Hongdong Li, Yuta Asano, Yinqiang Zheng, Imari Sato |
CVPR | 2 |
| 2021 | Learning Optical Flow From a Few MatchesabstractState-of-the-art neural network models for optical flow estimation require a dense correlation volume at high resolutions for representing per-pixel displacement. Although the dense correlation volume is informative for accurate estimation, its heavy computation and memory usage hinders the efficient training and deployment of the models. In this paper, we show that the dense correlation volume representation is redundant and accurate flow estimation can be achieved with only a fraction of elements in it. Based on this observation, we propose an alternative displacement representation, named Sparse Correlation Volume, which is constructed directly by computing the k closest matches in one feature map for each feature vector in the other feature map and stored in a sparse data structure. Experiments show that our method can reduce computational cost and memory use significantly, while maintaining high accuracy compared to previous approaches with dense correlation volumes. Shihao Jiang, Hongdong Li, Richard I. Hartley |
CVPR | 3 |
| 2021 | Lighting, Reflectance and Geometry Estimation From 360deg Panoramic StereoabstractWe propose a method for estimating high-definition spatially-varying lighting, reflectance, and geometry of a scene from 360° stereo images. Our model takes advantage of the 360° input to observe the entire scene with geometric detail, then jointly estimates the scene’s properties with physical constraints. We first reconstruct a near-field environment light for predicting the lighting at any 3D location within the scene. Then we present a deep learning model that leverages the stereo information to infer the reflectance and surface normal. Lastly, we incorporate the physical constraints between lighting and geometry to refine the reflectance of the scene. Both quantitative and qualitative experiments show that our method, benefiting from the 360° observation of the scene, outperforms prior state-of-the-art methods and enables more augmented reality applications such as mirror-objects insertion. Hongdong Li, Yasuyuki Matsushita |
CVPR | 2 |
| 2021 | ARVo: Learning All-Range Volumetric Correspondence for Video DeblurringabstractVideo deblurring models exploit consecutive frames to remove blurs from camera shakes and object motions. In order to utilize neighboring sharp patches, typical methods rely mainly on homography or optical flows to spatially align neighboring blurry frames. However, such explicit approaches are less effective in the presence of fast motions with large pixel displacements. In this work, we propose a novel implicit method to learn spatial correspondence among blurry frames in the feature space. To construct distant pixel correspondences, our model builds a correlation volume pyramid among all the pixel-pairs between neigh-boring frames. To enhance the features of the reference frame, we design a correlative aggregation module that maximizes the pixel-pair correlations with its neighbors based on the volume pyramid. Finally, we feed the aggregated features into a reconstruction module to obtain the restored frame. We design a generative adversarial paradigm to optimize the model progressively. Our proposed method is evaluated on the widely-adopted DVD dataset, along with a newly collected High-Frame-Rate (1000 fps) Dataset for Video Deblurring (HFR-DVD). Quantitative and qualitative experiments show that our model performs favorably on both datasets against previous state-of-the-art methods, confirming the benefit of modeling all-range spatial correspondence for video deblurring. Dongxu Li 0003, Kaihao Zhang, Xin Yu 0002, Yiran Zhong, Wenqi Ren, Hanna Suominen, Hongdong Li |
CVPR | 8 |
| 2021 | Dual Pixel Exploration: Simultaneous Depth Estimation and Image RestorationabstractThe dual-pixel (DP) hardware works by splitting each pixel in half and creating an image pair in a single snapshot. Several works estimate depth/inverse depth by treating the DP pair as a stereo pair. However, dual-pixel disparity only occurs in image regions with the defocus blur. The heavy defocus blur in DP pairs affects the performance of matching-based depth estimation approaches. Instead of removing the blur effect blindly, we study the formation of the DP pair which links the blur and the depth information. In this paper, we propose a mathematical DP model which can benefit depth estimation by the blur. These explorations motivate us to propose an end-to-end DDDNet (DP-based Depth and Deblur Network) to jointly estimate the depth and restore the image. Moreover, we define a re-blur loss, which reflects the relationship of the DP image formation process with depth information, to regularise our depth estimate in training. To meet the requirement of a large amount of data for learning, we propose the first DP image simulator which allows us to create datasets with DP pairs from any existing RGBD dataset. As a side contribution, we collect a real dataset for further research. Extensive experimental evaluation on both synthetic and real datasets shows that our approach achieves competitive performance compared to state-of-the-art approaches. Liyuan Pan, Shah Chowdhury, Richard I. Hartley, Miaomiao Liu 0001, Hongdong Li |
CVPR | 6 |
| 2021 | Self-Supervised Visibility Learning for Novel View SynthesisabstractWe address the problem of novel view synthesis (NVS) from a few sparse source view images. Conventional image-based rendering methods estimate scene geometry and synthesize novel views in two separate steps. However, erroneous geometry estimation will decrease NVS performance as view synthesis highly depends on the quality of estimated scene geometry. In this paper, we propose an end-to-end NVS framework to eliminate the error propagation issue. To be specific, we construct a volume under the target view and design a source-view visibility estimation (SVE) module to determine the visibility of the target-view voxels in each source view. Next, we aggregate the visibility of all source views to achieve a consensus volume. Each voxel in the consensus volume indicates a surface existence probability. Then, we present a soft ray-casting (SRC) mechanism to find the most front surface in the target view (i.e., depth). Specifically, our SRC traverses the consensus volume along viewing rays and then estimates a depth probability distribution. We then warp and aggregate source view pixels to synthesize a novel view based on the estimated source-view visibility and target-view depth. At last, our network is trained in an end-to-end self-supervised fashion, thus significantly alleviating error accumulation in view synthesis. Experimental results demonstrate that our method generates novel views in higher quality compared to the state-of-the-art. Yujiao Shi 0002, Hongdong Li, Xin Yu 0002 |
CVPR | 2 |
| 2021 | Deep Two-View Structure-From-Motion RevisitedabstractTwo-view structure-from-motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM. Existing deep learning-based approaches formulate the problem by either recovering absolute pose scales from two consecutive frames or predicting a depth map from a single image, both of which are ill-posed problems. In contrast, we propose to revisit the problem of deep two-view SfM by leveraging the well-posedness of the classic pipeline. Our method consists of 1) an optical flow estimation network that predicts dense correspondences between two frames; 2) a normalized pose estimation module that computes relative camera poses from the 2D optical flow correspondences, and 3) a scale-invariant depth estimation network that leverages epipolar geometry to reduce the search space, refine the dense correspondences, and estimate relative depth maps. Extensive experiments show that our method outperforms all state-of-the-art two-view SfM methods by a clear margin on KITTI depth, KITTI VO, MVS, Scenes11, and SUN3D datasets in both relative pose and depth estimation. Yiran Zhong, Yuchao Dai, Stanley T. Birchfield, Kaihao Zhang, Nikolai Smolyanskiy, Hongdong Li |
CVPR | 7 |
| 2021 | Rethinking Class Relations: Absolute-Relative Supervised and Unsupervised Few-Shot LearningabstractThe majority of existing few-shot learning methods describe image relations with binary labels. However, such binary relations are insufficient to teach the network complicated real-world relations, due to the lack of decision smoothness. Furthermore, current few-shot learning models capture only the similarity via relation labels, but they are not exposed to class concepts associated with objects, which is likely detrimental to the classification performance due to underutilization of the available class labels. For instance, children learn the concept of tiger from a few of actual examples as well as from comparisons of tiger to other animals. Thus, we hypothesize that both similarity and class concept learning must be occurring simultaneously. With these observations at hand, we study the fundamental problem of simplistic class modeling in current few-shot learning methods. We rethink the relations between class concepts, and propose a novel Absolute-relative Learning paradigm to fully take advantage of label information to refine the image an relation representations in both supervised and unsupervised scenarios. Our proposed paradigm improves the performance of several state-of-the-art models on publicly available datasets. Piotr Koniusz, Songlei Jian, Hongdong Li, Philip Torr 0001 |
CVPR | 4 |
| 2021 | Learning to Estimate Hidden Motions with Global Motion AggregationabstractOcclusions pose a significant challenge to optical flow algorithms that rely on local evidences. We consider an occluded point to be one that is imaged in the reference frame but not in the next, a slight overloading of the standard definition since it also includes points that move out-of-frame. Estimating the motion of these points is extremely difficult, particularly in the two-frame setting. Previous work relies on CNNs to learn occlusions, without much success, or requires multiple frames to reason about occlusions using temporal smoothness. In this paper, we argue that the occlusion problem can be better solved in the two-frame case by modelling image self-similarities. We introduce a global motion aggregation module, a transformer-based approach to find long-range dependencies between pixels in the first image, and perform global aggregation on the corresponding motion features. We demonstrate that the optical flow estimates in the occluded regions can be significantly improved without damaging the performance in non-occluded regions. This approach obtains new state-of-the-art results on the challenging Sintel dataset, improving the average end-point error by 13.6% on Sintel Final and 13.7% on Sintel Clean. At the time of submission, our method ranks first on these benchmarks among all published and unpublished approaches. Code is available at https://github.com/zacjiang/GMA. Shihao Jiang, Dylan Campbell, Hongdong Li, Richard I. Hartley |
ICCV | 4 |
| 2021 | Ranking Models in Unlabeled New EnvironmentsabstractConsider a scenario where we are supplied with a number of ready-to-use models trained on a certain source domain and hope to directly apply the most appropriate ones to different target domains based on the models’ relative performance. Ideally we should annotate a validation set for model performance assessment on each new target environment, but such annotations are often very expensive. Under this circumstance, we introduce the problem of ranking models in unlabeled new environments. For this problem, we propose to adopt a proxy dataset that 1) is fully labeled and 2) well reflects the true model rankings in a given target environment, and use the performance rankings on the proxy sets as surrogates. We first select labeled datasets as the proxy. Specifically, datasets that are more similar to the unlabeled target domain are found to better preserve the relative performance rankings. Motivated by this, we further propose to search the proxy set by sampling images from various datasets that have similar distributions as the target. We analyze the problem and its solutions on the person re-identification (re-ID) task, for which sufficient datasets are publicly available, and show that a carefully constructed proxy set effectively captures relative performance ranking in new environments. Code is avalible at https://github.com/sxzrt/Proxy-Set. Xiaoxiao Sun 0002, Yunzhong Hou, Weijian Deng, Hongdong Li, Liang Zheng 0001 |
ICCV | 4 |
| 2021 | Benchmarking Ultra-High-Definition Image Super-resolutionabstractIncreasingly, modern mobile devices allow capturing images at Ultra-High-Definition (UHD) resolution, which includes 4K and 8K images. However, current single image super-resolution (SISR) methods focus on super-resolving images to ones with resolution up to high definition (HD) and ignore higher-resolution UHD images. To explore their performance on UHD images, in this paper, we first introduce two large-scale image datasets, UHDSR4K and UHDSR8K, to benchmark existing SISR methods. With 70,000 V100 GPU hours of training, we benchmark these methods on 4K and 8K resolution images under seven different settings to provide a set of baseline models. Moreover, we propose a baseline model, called Mesh Attention Network (MANet) for SISR. The MANet applies the attention mechanism in both different depths (horizontal) and different levels of receptive field (vertical). In this way, correlations among feature maps are learned, enabling the network to focus on more important features. Kaihao Zhang, Dongxu Li 0003, Wenhan Luo, Wenqi Ren, Björn Stenger, Wei Liu 0005, Hongdong Li, Ming-Hsuan Yang 0001 |
ICCV | 7 |
| 2021 | The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and PoseabstractThe availability of a large labeled dataset is a key requirement for applying deep learning methods to solve various computer vision tasks. In the context of understanding human activities, existing public datasets, while large in size, are often limited to a single RGB camera and provide only per-frame or per-clip action annotations. To enable richer analysis and understanding of human activities, we introduce IKEA ASM-a three million frame, multi-view, furniture assembly video dataset that includes depth, atomic actions, object segmentation, and human poses. Additionally, we benchmark prominent methods for video action recognition, object segmentation and human pose estimation tasks on this challenging dataset. The dataset enables the development of holistic methods, which integrate multi-modal and multi-view data to better perform on these tasks. Yizhak Ben-Shabat, Xin Yu 0002, Fatemehsadat Saleh, Dylan Campbell, Cristian Rodriguez Opazo, Hongdong Li, Stephen Gould |
WACV | 6 |
| 2021 | DORi: Discovering Object Relationships for Moment Localization of a Natural Language Query in a VideoabstractThis paper studies the task of temporal moment localization in long untrimmed videos using natural language queries. Given a query sentence, the goal is to determine the start and end of the relevant segment within the video. Our key innovation is to learn a video feature embedding through a language-conditioned message-passing algorithm suitable for temporal moment localization which captures the relationships between humans, objects and activities in the video. These relationships are obtained by a spatial sub-graph that contextualizes the scene representation using detected objects and human features conditioned in the language query. Moreover, a temporal sub-graph captures the activities within the video through time. Our method is evaluated on three standard benchmark datasets, and we also introduce YouCookII as a new benchmark for this task. Experiments show our method outperforms state-of-the-art methods on these datasets, confirming the effectiveness of our approach. Cristian Rodriguez Opazo, Edison Marrese-Taylor, Basura Fernando, Hongdong Li, Stephen Gould |
WACV | 4 |
| 2021 | Multi-level Motion Attention for Human Motion Prediction
Wei Mao 0001, Miaomiao Liu 0001, Mathieu Salzmann, Hongdong Li |
Int. J. Comput. Vis. | 4 |
| 2021 | 3D Scene Reconstruction with an Un-calibrated Light Field Camera
Qi Zhang 0029, Hongdong Li, Xue Wang 0006, Qing Wang 0006 |
Int. J. Comput. Vis. | 2 |
| 2021 | Deep robust image deblurring via blur distilling and information comparison in latent space
Wenjia Niu, Kaihao Zhang, Wenhan Luo, Yiran Zhong, Hongdong Li |
Neurocomputing | 5 |
| 2021 | Superpixel Soup: Monocular Dense 3D Reconstruction of a Complex Dynamic SceneabstractThis work addresses the task of dense 3D reconstruction of a complex dynamic scene from images. The prevailing idea to solve this task is composed of a sequence of steps and is dependent on the success of several pipelines in its execution. To overcome such limitations with the existing algorithm, we propose a unified approach to solve this problem. We assume that a dynamic scene can be approximated by numerous piecewise planar surfaces, where each planar surface enjoys its own rigid motion, and the global change in the scene between two frames is as-rigid-as-possible (ARAP). Consequently, our model of a dynamic scene reduces to a soup of planar structures and rigid motion of these local planar structures. Using planar over-segmentation of the scene, we reduce this task to solving a "3D jigsaw puzzle" problem. Hence, the task boils down to correctly assemble each rigid piece to construct a 3D shape that complies with the geometry of the scene under the ARAP assumption. Further, we show that our approach provides an effective solution to the inherent scale-ambiguity in structure-from-motion under perspective projection. We provide extensive experimental results and evaluation on several benchmark datasets. Quantitative comparison with competing approaches shows state-of-the-art performance. Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Self-Supervised Multiscale Adversarial Regression Network for Stereo Disparity EstimationabstractDeep learning approaches have significantly contributed to recent progress in stereo matching. These deep stereo matching methods are usually based on supervised training, which requires a large amount of high-quality ground-truth depth map annotations that are expensive to collect. Furthermore, only a limited quantity of stereo vision training data are currently available, obtained either by active sensors (Lidar and ToF cameras) or through computer graphics simulations and not meeting requirements for deep supervised training. Here, we propose a novel deep stereo approach called the "self-supervised multiscale adversarial regression network (SMAR-Net)," which relaxes the need for ground-truth depth maps for training. Specifically, we design a two-stage network. The first stage is a disparity regressor, in which a regression network estimates disparity values from stacked stereo image pairs. Stereo image stacking method is a novel contribution as it not only contains the spatial appearances of stereo images but also implies matching correspondences with different disparity values. In the second stage, a synthetic left image is generated based on the left-right consistency assumption. Our network is trained by minimizing a hybrid loss function composed of a content loss and an adversarial loss. The content loss minimizes the average warping error between the synthetic images and the real ones. In contrast to the generative adversarial loss, our proposed adversarial loss penalizes mismatches using multiscale features. This constrains the synthetic image and real image as being pixelwise identical instead of just belonging to the same distribution. Furthermore, the combined utilization of multiscale feature extraction in both the content loss and adversarial loss further improves the adaptability of SMAR-Net in ill-posed regions. Experiments on multiple benchmark datasets show that SMAR-Net outperforms the current state-of-the-art self-supervised methods and achieves comparable outcomes to supervised methods. The source code can be accessed at: https://github.com/Dawnstar8411/SMAR-Net. Chen Wang 0026, Xiao Bai 0001, Xiang Wang 0014, Xianglong Liu 0001, Jun Zhou 0001, Xinyu Wu 0001, Hongdong Li, Dacheng Tao |
IEEE Trans. Cybern. | 7 |
| 2021 | MLDA-Net: Multi-Level Dual Attention-Based Network for Self-Supervised Monocular Depth EstimationabstractThe success of supervised learning-based single image depth estimation methods critically depends on the availability of large-scale dense per-pixel depth annotations, which requires both laborious and expensive annotation process. Therefore, the self-supervised methods are much desirable, which attract significant attention recently. However, depth maps predicted by existing self-supervised methods tend to be blurry with many depth details lost. To overcome these limitations, we propose a novel framework, named MLDA-Net, to obtain per-pixel depth maps with shaper boundaries and richer depth details. Our first innovation is a multi-level feature extraction (MLFE) strategy which can learn rich hierarchical representation. Then, a dual-attention strategy, combining global attention and structure attention, is proposed to intensify the obtained features both globally and locally, resulting in improved depth maps with sharper boundaries. Finally, a reweighted loss strategy based on multi-level outputs is proposed to conduct effective supervision for self-supervised depth estimation. Experimental results demonstrate that our MLDA-Net framework achieves state-of-the-art depth prediction results on the KITTI benchmark for self-supervised monocular depth estimation with different input modes and training modes. Extensive experiments on other benchmark datasets further confirm the superiority of our proposed approach. Xibin Song, Wei Li 0143, Dingfu Zhou, Yuchao Dai, Hongdong Li, Liangjun Zhang |
IEEE Trans. Image Process. | 6 |
| 2021 | Angular-Driven Feedback Restoration Networks for Imperfect Sketch RecognitionabstractAutomatic hand-drawn sketch recognition is an important task in computer vision. However, the vast majority of prior works focus on exploring the power of deep learning to achieve better accuracy on complete and clean sketch images, and thus fail to achieve satisfactory performance when applied to incomplete or destroyed sketch images. To address this problem, we first develop two datasets that contain different levels of scrawl and incomplete sketches. Then, we propose an angular-driven feedback restoration network (ADFRNet), which first detects the imperfect parts of a sketch and then refines them into high quality images, to boost the performance of sketch recognition. By introducing a novel "feedback restoration loop" to deliver information between the middle stages, the proposed model can improve the quality of generated sketch images while avoiding the extra memory cost associated with popular cascading generation schemes. In addition, we also employ a novel angular-based loss function to guide the refinement of sketch images and learn a powerful discriminator in the angular space. Extensive experiments conducted on the proposed imperfect sketch datasets demonstrate that the proposed model is able to efficiently improve the quality of sketch images and achieve superior performance over the current state-of-the-art methods. Jia Wan 0001, Kaihao Zhang, Hongdong Li, Antoni B. Chan |
IEEE Trans. Image Process. | 3 |
| 2021 | Revisiting Spatio-Angular Trade-off in Light Field Cameras and Extended Applications in Super-ResolutionabstractLight field cameras (LFCs) have received increasing attention due to their wide-spread applications. However, current LFCs suffer from the well-known spatio-angular trade-off, which is considered an inherent and fundamental limit for LFC designs. In this article, by doing a detailed optical analysis of the sampling process in an LFC, we show that the effective resolution is generally higher than the number of micro-lenses. This contribution makes it theoretically possible to super-resolve a light field. Further optical analysis proves the "2D predictable series" nature of the 4D light field, which provides new insights for analyzing light field using series processing techniques. To model this nature, a specifically designed epipolar plane image (EPI) based CNN-LSTM network is proposed to super-resolve a light field in the spatial and angular dimensions simultaneously. Rather than leveraging semantic information, our network focuses on extracting geometric continuity in the EPI domain. This gives our method an improved generalization ability and makes it applicable to a wide range of previously unseen scenes. Experiments on both synthetic and real light fields demonstrate the improvements over state-of-the-arts, especially in large disparity areas. Hao Zhu 0005, Mantang Guo, Hongdong Li, Qing Wang 0006, Antonio Robles-Kelly |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | Optimal Feature Transport for Cross-View Image Geo-LocalizationabstractThis paper addresses the problem of cross-view image geo-localization, where the geographic location of a ground-level street-view query image is estimated by matching it against a large scale aerial map (e.g., a high-resolution satellite image). State-of-the-art deep-learning based methods tackle this problem as deep metric learning which aims to learn global feature representations of the scene seen by the two different views. Despite promising results are obtained by such deep metric learning methods, they, however, fail to exploit a crucial cue relevant for localization, namely, the spatial layout of local features. Moreover, little attention is paid to the obvious domain gap (between aerial view and ground view) in the context of cross-view localization. This paper proposes a novel Cross-View Feature Transport (CVFT) technique to explicitly establish cross-view domain transfer that facilitates feature alignment between ground and aerial images. Specifically, we implement the CVFT as network layers, which transports features from one domain to the other, leading to more meaningful feature similarity comparison. Our model is differentiable and can be learned end-to-end. Experiments on large-scale datasets have demonstrated that our method has remarkably boosted the state-of-the-art cross-view localization performance, e.g., on the CVUSA dataset, with significant improvements for top-1 recall from 40.79% to 61.43%, and for top-10 from 76.36% to 90.49%. We expect the key insight of the paper (i.e., explicitly handling domain difference via domain transport) will prove to be useful for other similar problems in computer vision as well. Yujiao Shi 0002, Xin Yu 0002, Liu Liu 0009, Tong Zhang 0023, Hongdong Li |
AAAI | 5 |
| 2020 | 6DoF Object Pose Estimation via Differentiable Proxy Voting Regularizer
Xin Yu 0002, Zheyu Zhuang, Piotr Koniusz, Hongdong Li |
BMVC | 4 |
| 2020 | Transferring Cross-Domain Knowledge for Video Sign Language RecognitionabstractWord-level sign language recognition (WSLR) is a fundamental task in sign language interpretation. It requires models to recognize isolated sign words from videos. However, annotating WSLR data needs expert knowledge, thus limiting WSLR dataset acquisition. On the contrary, there are abundant subtitled sign news videos on the internet. Since these videos have no word-level annotation and exhibit a large domain gap from isolated signs, they cannot be directly used for training WSLR models. We observe that despite the existence of a large domain gap, isolated and news signs share the same visual concepts, such as hand gestures and body movements. Motivated by this observation, we propose a novel method that learns domain-invariant visual concepts and fertilizes WSLR models by transferring knowledge of subtitled news sign to them. To this end, we extract news signs using a base WSLR model, and then design a classifier jointly trained on news and isolated signs to coarsely align these two domain features. In order to learn domain-invariant features within each class and suppress domain-specific features, our method further resorts to an external memory to store the class centroids of the aligned news signs. We then design a temporal attention based on the learnt descriptor to improve recognition performance. Experimental results on standard WSLR datasets show that our method outperforms previous state-of-the-art methods significantly. We also demonstrate the effectiveness of our method on automatically localizing signs from sign news, achieving 28.1 for [email protected]. Dongxu Li 0003, Xin Yu 0002, Lars Petersson, Hongdong Li |
CVPR | 5 |
| 2020 | Where Am I Looking At? Joint Location and Orientation Estimation by Cross-View MatchingabstractCross-view geo-localization is the problem of estimating the position and orientation (latitude, longitude and azimuth angle) of a camera at ground level given a large-scale database of geo-tagged aerial (eg., satellite) images. Existing approaches treat the task as a pure location estimation problem by learning discriminative feature descriptors, but neglect orientation alignment. It is well-recognized that knowing the orientation between ground and aerial images can significantly reduce matching ambiguity between these two views, especially when the ground-level images have a limited Field of View (FoV) instead of a full field-of-view panorama. Therefore, we design a Dynamic Similarity Matching network to estimate cross-view orientation alignment during localization. In particular, we address the cross-view domain gap by applying a polar transform to the aerial images to approximately align the images up to an unknown azimuth angle. Then, a two-stream convolutional network is used to learn deep features from the ground and polar-transformed aerial images. Finally, we obtain the orientation by computing the correlation between cross-view features, which also provides a more accurate measure of feature similarity, improving location recall. Experiments on standard datasets demonstrate that our method significantly improves state-of-the-art performance. Remarkably, we improve the top-1 location recall rate on the CVUSA dataset by a factor of 1.5x for panoramas with known orientation, by a factor of 3.3x for panoramas with unknown orientation, and by a factor of 6x for 180°-FoV images with unknown orientation. Yujiao Shi 0002, Xin Yu 0002, Dylan Campbell, Hongdong Li |
CVPR | 4 |
| 2020 | Channel Attention Based Iterative Residual Learning for Depth Map Super-ResolutionabstractDespite the remarkable progresses made in deep learning based depth map super-resolution (DSR), how to tackle real-world degradation in low-resolution (LR) depth maps remains a major challenge. Existing DSR model is generally trained and tested on synthetic dataset, which is very different from what would get from a real depth sensor. In this paper, we argue that DSR models trained under this setting are restrictive and not effective in dealing with realworld DSR tasks. We make two contributions in tackling real-world degradation of different depth sensors. First, we propose to classify the generation of LR depth maps into two types: non-linear downsampling with noise and interval downsampling, for which DSR models are learned correspondingly. Second, we propose a new framework for real-world DSR, which consists of four modules : 1) An iterative residual learning module with deep supervision to learn effective high-frequency components of depth maps in a coarse-to-fine manner; 2) A channel attention strategy to enhance channels with abundant high-frequency components; 3) A multi-stage fusion module to effectively reexploit the results in the coarse-to-fine process; and 4) A depth refinement module to improve the depth map by TGV regularization and input loss. Extensive experiments on benchmarking datasets demonstrate the superiority of our method over current state-of-the-art DSR methods. Xibin Song, Yuchao Dai, Dingfu Zhou, Liu Liu 0009, Wei Li 0111, Hongdong Li, Ruigang Yang |
CVPR | 6 |
| 2020 | Deblurring by Realistic BlurringabstractExisting deep learning methods for image deblurring typically train models using pairs of sharp images and their blurred counterparts. However, synthetically blurring images does not necessarily model the blurring process in real-world scenarios with sufficient accuracy. To address this problem, we propose a new method which combines two GAN models, i.e., a learning-to-Blur GAN (BGAN) and learning-to-DeBlur GAN (DBGAN), in order to learn a better model for image deblurring by primarily learning how to blur images. The first model, BGAN, learns how to blur sharp images with unpaired sharp and blurry image sets, and then guides the second model, DBGAN, to learn how to correctly deblur such images. In order to reduce the discrepancy between real blur and synthesized blur, a relativistic blur loss is leveraged. As an additional contribution, this paper also introduces a Real-World Blurred Image (RWBI) dataset including diverse blurry images. Our experiments show that the proposed method achieves consistently superior quantitative performance as well as higher perceptual quality on both the newly proposed dataset and the public GOPRO dataset. Kaihao Zhang, Wenhan Luo, Yiran Zhong, Lin Ma 0002, Björn Stenger, Wei Liu 0005, Hongdong Li |
CVPR | 7 |
| 2020 | Joint 3D Instance Segmentation and Object Detection for Autonomous DrivingabstractCurrently, in Autonomous Driving (AD), most of the 3D object detection frameworks (either anchor- or anchor-free-based) consider the detection as a Bounding Box (BBox) regression problem. However, this compact representation is not sufficient to explore all the information of the objects. To tackle this problem, we propose a simple but practical detection framework to jointly predict the 3D BBox and instance segmentation. For instance segmentation, we propose a Spatial Embeddings (SEs) strategy to assemble all foreground points into their corresponding object centers. Base on the SE results, the object proposals can be generated based on a simple clustering strategy. For each cluster, only one proposal is generated. Therefore, the Non-Maximum Suppression (NMS) process is no longer needed here. Finally, with our proposed instance-aware ROI pooling, the BBox is refined by a second-stage network. Experimental results on the public KITTI dataset show that the proposed SEs can significantly improve the instance segmentation results compared with other feature embedding-based method. Meanwhile, it also outperforms most of the 3D object detectors on the KITTI testing benchmark. Dingfu Zhou, Xibin Song, Liu Liu 0009, Junbo Yin, Yuchao Dai, Hongdong Li, Ruigang Yang |
CVPR | 7 |
| 2020 | Deep Novel View Synthesis from Colored 3D Point Clouds
Zhenbo Song, Wayne Chen, Dylan Campbell, Hongdong Li |
ECCV (24) | 4 |
| 2020 | Beyond Monocular Deraining: Stereo Image Deraining via Semantic Understanding
Kaihao Zhang, Wenhan Luo, Wenqi Ren, Jingwen Wang 0003, Fang Zhao 0006, Lin Ma 0002, Hongdong Li |
ECCV (27) | 7 |
| 2020 | Few-Shot Action Recognition with Permutation-Invariant Attention
Li Zhang 0040, Xiaojuan Qi 0001, Hongdong Li, Philip Torr 0001, Piotr Koniusz |
ECCV (5) | 4 |
| 2020 | Accurate 3D Reconstruction from Circular Light Field Using CNN-LSTMabstractA light field is formed by densely capturing images on a regular sub-aperture grid. Geometry information endowed in the epipolar plane images(EPI) can only lead to a 2. 5D reconstruction. In order to obtain a full 360°view of an object, we focus on light fields captured by a circularly moving camera, resulting in circular light fields (or Cir-LFs in short). Compared with traditional EPIs, Circular EPIs(CEPIs) provide unique advantages, such as that corresponding points forming a 3D sinusoid like curve instead of a 2D straight line and geometry information encoded sequentially in multiple adjacent views along the curve. However, current reconstruction methods only focus on the 2D projection of 3D curve, leading to distortions in the reconstructed upper and lower surfaces. We propose to analyze 3D features contained in the 3D CEPI volume and we develop a deep CNN-LSTM network to model the gradient map in the CEPI volume. Additionally, a large scale Cir-LF dataset is constructed for research purpose. Experiments on both synthetic and real scenes demonstrate the effectiveness and generaliability of the proposed method. Zhengxi Song, Hao Zhu 0005, Xue Wang 0006, Hongdong Li, Qing Wang 0006 |
ICME | 5 |
| 2020 | 3D Human Pose Estimation with 2D Human Pose and Depthmap
Xuanying Zhu, Henry J. Gardner, Hongdong Li |
ICONIP (4) | 5 |
| 2020 | Globally Optimal Relative Pose Estimation for Camera on a Selfie StickabstractTaking selfies has become a photographic trend nowadays. We envision the emergence of the "video selfie" capturing a short continuous video clip (or burst photography) of the user, themselves. A selfie stick is usually used, whereby a camera is mounted on a stick for taking selfie photos. In this scenario, we observe that the camera typically goes through a special trajectory along a sphere surface. Motivated by this observation, in this work, we propose an efficient and globally optimal relative camera pose estimation between a pair of two images captured by a camera mounted on a selfie stick. We exploit the special geometric structure of the camera motion constrained by a selfie stick and define its motion as spherical joint motion. By the new parametrization and calibration scheme, we show that the pose estimation problem can be reduced to a 3-DoF (degrees of freedom) search problem, instead of a generic 6-DoF problem. This allows us to derive a fast branch-and-bound global optimization, which guarantees a global optimum. Thereby, we achieve efficient and robust estimation even in the presence of outliers. By experiments on both synthetic and real-world data, we validate the performance as well as the guaranteed optimality of the proposed method. Kyungdon Joo, Hongdong Li, Tae-Hyun Oh, Yunsu Bok, In-So Kweon |
ICRA | 2 |
| 2020 | End-to-end Learning for Inter-Vehicle Distance and Relative Velocity Estimation in ADAS with a Monocular CameraabstractInter-vehicle distance and relative velocity estimations are two basic functions for any ADAS (Advanced driver-assistance systems). In this paper, we propose a monocular camera based inter-vehicle distance and relative velocity estimation method based on end-to-end training of a deep neural network. The key novelty of our method is the integration of multiple visual clues provided by any two time-consecutive monocular frames, which include deep feature clue, scene geometry clue, as well as temporal optical flow clue. We also propose a vehicle-centric sampling mechanism to alleviate the effect of perspective distortion in the motion field (i.e. optical flow). We implement the method by a light-weight deep neural network. Extensive experiments are conducted which confirm the superior performance of our method over other state-of-the-art methods, in terms of estimation accuracy, computational speed, and memory footprint. Zhenbo Song, Jianfeng Lu 0003, Tong Zhang 0023, Hongdong Li |
ICRA | 4 |
| 2020 | Reliable frame-to-frame motion estimation for vehicle-mounted surround-view camera systemsabstractModern vehicles are often equipped with a surround-view multi-camera system. The current interest in autonomous driving invites the investigation of how to use such systems for a reliable estimation of relative vehicle displacement. Existing camera pose algorithms either work for a single camera, make overly simplified assumptions, are computationally expensive, or simply become degenerate under non-holonomic vehicle motion. In this paper, we introduce a new, reliable solution able to handle all kinds of relative displacements in the plane despite the possibly non-holonomic characteristics. We furthermore introduce a novel two-view optimization scheme which minimizes a geometrically relevant error without relying on 3D point related optimization variables. Our method leads to highly reliable and accurate frame-to-frame visual odometry with a full-size, vehicle-mounted surround-view camera system. Yifu Wang, Xin Peng 0005, Hongdong Li, Laurent Kneip |
ICRA | 4 |
| 2020 | Every Moment Matters: Detail-Aware Networks to Bring a Blurry Image AliveabstractMotion-blurred images are the result of light accumulation over the period of camera exposure time, during which the camera and objects in the scene are in relative motion to each other. The inverse process of extracting an image sequence from a single motion-blurred image is an ill-posed vision problem. One key challenge is that the motions across frames are subtle, which makes the generating networks difficult to capture them and thus the recovery sequences lack motion details. In order to alleviate this problem, we propose a detail-aware network with three consecutive stages to improve the reconstruction quality by addressing specific aspects in the recovery process. The detail-aware network firstly models the dynamics using a cycle flow loss, resolving the temporal ambiguity of the reconstruction in the first stage. Then, a GramNet is proposed in the second stage to refine subtle motion between continuous frames using Gram matrices as motion representation. Finally, we introduce a HeptaGAN in the third stage to bridge the continuous and discrete nature of exposure time and recovered frames, respectively, in order to maintain rich detail. Experiments show that the proposed detail-aware networks produce sharp image sequences with rich details and subtle motion, outperforming the state-of-the-art methods. Kaihao Zhang, Wenhan Luo, Björn Stenger, Wenqi Ren, Lin Ma 0002, Hongdong Li |
ACM Multimedia | 6 |
| 2020 | Hierarchical Neural Architecture Search for Deep Stereo MatchingabstractTo reduce the human efforts in neural network design, Neural Architecture Search (NAS) has been applied with remarkable success to various high-level vision tasks such as classification and semantic segmentation. The underlying idea for the NAS algorithm is straightforward, namely, to allow the network the ability to choose among a set of operations (\eg convolution with different filter sizes), one is able to find an optimal architecture that is better adapted to the problem at hand. However, so far the success of NAS has not been enjoyed by low-level geometric vision tasks such as stereo matching. This is partly due to the fact that state-of-the-art deep stereo matching networks, designed by humans, are already sheer in size. Directly applying the NAS to such massive structures is computationally prohibitive based on the currently available mainstream computing resources. In this paper, we propose the first \emph{end-to-end} hierarchical NAS framework for deep stereo matching by incorporating task-specific human knowledge into the neural architecture search framework. Specifically, following the gold standard pipeline for deep stereo matching (\ie, feature extraction -- feature volume construction and dense matching), we optimize the architectures of the entire pipeline jointly. Extensive experiments show that our searched network outperforms all state-of-the-art deep stereo matching architectures and is ranked at the top 1 accuracy on KITTI stereo 2012, 2015, and Middlebury benchmarks, as well as the top 1 on SceneFlow dataset with a substantial improvement on the size of the network and the speed of inference. Code available at https://github.com/XuelianCheng/LEAStereo. Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Yuchao Dai, Xiaojun Chang, Hongdong Li, Tom Drummond, ZongYuan Ge |
NeurIPS | 6 |
| 2020 | TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationabstractSign language translation (SLT) aims to interpret sign video sequences into text-based natural language sentences. Sign videos consist of continuous sequences of sign gestures with no clear boundaries in between. Existing SLT models usually represent sign visual features in a frame-wise manner so as to avoid needing to explicitly segmenting the videos into isolated signs. However, these methods neglect the temporal information of signs and lead to substantial ambiguity in translation. In this paper, we explore the temporal semantic structures of sign videos to learn more discriminative features. To this end, we first present a novel sign video segment representation which takes into account multiple temporal granularities, thus alleviating the need for accurate video segmentation. Taking advantage of the proposed segment representation, we develop a novel hierarchical sign video feature learning method via a temporal semantic pyramid network, called TSPNet. Specifically, TSPNet introduces an inter-scale attention to evaluate and enhance local semantic consistency of sign segments and an intra-scale attention to resolve semantic ambiguity by using non-local video context. Experiments show that our TSPNet outperforms the state-of-the-art with significant improvements on the BLEU score (from 9.58 to 13.41) and ROUGE score (from 31.80 to 34.96) on the largest commonly used SLT dataset. Our implementation is available at https://github.com/verashira/TSPNet. Dongxu Li 0003, Xin Yu 0002, Kaihao Zhang, Ben Swift, Hanna Suominen, Hongdong Li |
NeurIPS | 7 |
| 2020 | Displacement-Invariant Matching Cost Learning for Accurate Optical Flow EstimationabstractLearning matching costs has been shown to be critical to the success of the state-of-the-art deep stereo matching methods, in which 3D convolutions are applied on a 4D feature volume to learn a 3D cost volume. However, this mechanism has never been employed for the optical flow task. This is mainly due to the significantly increased search dimension in the case of optical flow computation, \ie, a straightforward extension would require dense 4D convolutions in order to process a 5D feature volume, which is computationally prohibitive. This paper proposes a novel solution that is able to bypass the requirement of building a 5D feature volume while still allowing the network to learn suitable matching costs from data. Our key innovation is to decouple the connection between 2D displacements and learn the matching costs at each 2D displacement hypothesis independently, \ie, displacement-invariant cost learning. Specifically, we apply the same 2D convolution-based matching net independently on each 2D displacement hypothesis to learn a 4D cost volume. Moreover, we propose a displacement-aware projection layer to scale the learned cost volume, which reconsiders the correlation between different displacement candidates and mitigates the multi-modal problem in the learned cost volume. The cost volume is then projected to optical flow estimation through a 2D soft-argmin layer. Extensive experiments show that our approach achieves state-of-the-art accuracy on various datasets, and outperforms all published optical flow methods on the Sintel benchmark. The code is available at https://github.com/jytime/DICL-Flow. Yiran Zhong, Yuchao Dai, Kaihao Zhang, Pan Ji, Hongdong Li |
NeurIPS | 6 |
| 2020 | Word-level Deep Sign Language Recognition from Video: A New Large-scale Dataset and Methods ComparisonabstractVision-based sign language recognition aims at helping the deaf people to communicate with others. However, most existing sign language datasets are limited to a small number of words. Due to the limited vocabulary size, models learned from those datasets cannot be applied in practice. In this paper, we introduce a new large-scale Word-Level American Sign Language (WLASL) video dataset, containing more than 2000 words performed by over 100 signers. This dataset will be made publicly available to the research community. To our knowledge,it is by far the largest public ASL dataset to facilitate word-level sign recognition research. Based on this new large-scale dataset, we are able to experiment with several deep learning methods for word-level sign recognition and evaluate their performances in large scale scenarios. Specifically we implement and compare two different models,i.e., (i) holistic visual appearance based approach, and (ii) 2D human pose based approach. Both models are valuable baselines that will benefit the community for method benchmarking. Moreover, we also propose a novel pose-based temporal graph convolution networks (Pose-TGCN) that model spatial and temporal dependencies in human pose trajectories simultaneously, which has further boosted the performance of the pose-based method. Our results show that pose-based and appearance-based models achieve comparable performances up to 62.63% at top-10 accuracy on 2,000 words/glosses, demonstrating the validity and challenges of our dataset. Our dataset and baseline deep models are available at https://dxli94.github.io/WLASL/. Dongxu Li 0003, Cristian Rodriguez Opazo, Xin Yu 0002, Hongdong Li |
WACV | 4 |
| 2020 | Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided AttentionabstractThis paper studies the problem of temporal moment localization in a long untrimmed video using natural language as the query. Given an untrimmed video and a query sentence, the goal is to determine the start and end of the relevant visual moment in the video that corresponds to the query sentence. While most previous works have tackled this by a propose-and-rank approach, we introduce a more efficient, end-to-end trainable, and proposal-free approach that is built upon three key components: a dynamic filter which adaptively transfers language information to visual domain attention map, a new loss function to guide the model to attend the most relevant part of the video, and soft labels to cope with annotation uncertainties. Our method is evaluated on three standard benchmark datasets, Charades-STA, TACoS and ActivityNet-Captions. Experimental results show our method outperforms state-of-the-art methods on these datasets, confirming the effectiveness of the method. We believe the proposed dynamic filter-based guided attention mechanism will prove valuable for other vision and language tasks as well. Cristian Rodriguez Opazo, Edison Marrese-Taylor, Fatemehsadat Saleh, Hongdong Li, Stephen Gould |
WACV | 4 |
| 2020 | Guest Editorial: Special Issue on ACCV 2018
C. V. Jawahar, Hongdong Li, Greg Mori, Konrad Schindler |
Int. J. Comput. Vis. | 2 |
| 2020 | Globally-Optimal Inlier Set Maximisation for Camera Pose and Correspondence EstimationabstractEstimating the 6-DoF pose of a camera from a single image relative to a 3D point-set is an important task for many computer vision applications. Perspective-n-point solvers are routinely used for camera pose estimation, but are contingent on the provision of good quality 2D-3D correspondences. However, finding cross-modality correspondences between 2D image points and a 3D point-set is non-trivial, particularly when only geometric information is known. Existing approaches to the simultaneous pose and correspondence problem use local optimisation, and are therefore unlikely to find the optimal solution without a good pose initialisation, or introduce restrictive assumptions. Since a large proportion of outliers and many local optima are common for this problem, we instead propose a robust and globally-optimal inlier set maximisation approach that jointly estimates the optimal camera pose and correspondences. Our approach employs branch-and-bound to search the 6D space of camera poses, guaranteeing global optimality without requiring a pose prior. The geometry of SE(3) is used to find novel upper and lower bounds on the number of inliers and local optimisation is integrated to accelerate convergence. The algorithm outperforms existing approaches on challenging synthetic and real datasets, reliably finding the global optimum, with a GPU implementation greatly reducing runtime. Dylan Campbell, Lars Petersson, Laurent Kneip, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Unidirectional Representation-Based Efficient Dictionary LearningabstractDictionary learning (DL) has been widely studied for pattern classification. Most existing methods introduce multiple discriminative terms into objective functions for accuracy improvement, leading to complex learning frameworks and high computational burdens. This paper proposes a simple yet effective DL algorithm for classification, namely unidirectional representation dictionary learning (URDL). Unidirectional constraint is proposed to guide coefficient directions in the representation to be discriminative. Besides, direction-thresholding is proposed to exploit the direction property in the classification scheme. It suppresses the disturbance from undesired non-zero coefficients, and improves the representation discriminability. We adopt squared ℓ2-norm-based regularization for efficient coding, and systematically analyze the mechanism of the proposed method. Extensive experiments on five data sets are conducted, including object categorization, scene classification, face recognition, and fine-grained flower classification. The experimental results demonstrate that the proposed approach not only outperforms the state-of-the-art DL algorithms in terms of recognition accuracy significantly, but also exhibits a much higher computational efficiency. Xiudong Wang, Yali Li 0001, Shaodi You, Hongdong Li, Shengjin Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | 4D Light Field Superpixel and SegmentationabstractSuperpixel segmentation of 2D images has been widely used in many computer vision tasks. Previous algorithms model the color, position, or higher spectral information for segmenting a 2D image. However, limited to the Gaussian imaging principle in a traditional camera, where each pixel is formed by summing lots of light rays from different angles, there is not a thorough segmentation solution to eliminate the ambiguity in defocus and occlusion boundary areas. In this paper, we consider the essential element of image pixel, i.e., rays in light space, and propose light field superpixel (LFSP) to eliminate the ambiguity. The LFSP is first defined mathematically and then two evaluation metrics, named LFSP self-similarity and effective label ratio, are proposed to evaluate the refocus-invariant and full-sliced properties of segmentation. By building a clique system containing 80 neighbors in light field, a robust refocus-invariant LFSP segmentation algorithm is developed. Experimental results on both synthetic and real light field datasets demonstrate the advantages over the current state of the art in terms of traditional evaluation metrics. Additionally, the LFSP self-similarity evaluations under different light field refocus levels show the refocus-invariance of the proposed algorithm. The full-sliced property of the proposed LFSP algorithm is verified by comparing it with the classical supervoxel algorithms. Finally, an LFSP-based application is demonstrated to show the effectiveness of LFSP in light field editing. Hao Zhu 0005, Qi Zhang 0029, Qing Wang 0006, Hongdong Li |
IEEE Trans. Image Process. | 4 |
| 2020 | Ground-Plane-Based Absolute Scale Estimation for Monocular Visual OdometryabstractRecovering an absolute metric scale from a monocular camera is a challenging but highly desirable problem for monocular camera-based systems. By using different kinds of cues, various approaches have been proposed for scale estimation, such as camera height and object size. In this paper, first, we summarize different kinds of scale estimation approaches. Then, we propose a robust divide-and-conquer absolute scale estimation method based on the ground plane and camera height by analyzing the advantages and disadvantages of different approaches. By using the estimated scale, an effective scale correction strategy has been proposed to reduce the scale drift during the monocular visual odometry estimation process. Finally, the effectiveness and robustness of the proposed method have been verified on both public and self-collected image sequences. Dingfu Zhou, Yuchao Dai, Hongdong Li |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Cousin Network Guided Sketch Recognition via Latent Attribute WarehouseabstractWe study the problem of sketch image recognition. This problem is plagued with two major challenges: 1) sketch images are often scarce in contrast to the abundance of natural images, rendering the training task difficult, and 2) the significant domain gap between sketch image and its natural image counterpart makes the task of bridging the two domains challenging. In order to overcome these challenges, in this paper we propose to transfer the knowledge of a network learned from natural images to a sketch network - a new deep net architecture which we term as cousin network. This network guides a sketch-recognition network to extract more relevant features that are close to those of natural images, via adversarial training. Moreover, to enhance the transfer ability of the classification model, a sketch-to-image attribute warehouse is constructed to approximate the transformation between the sketch domain and the real image domain. Extensive experiments conducted on the TU-Berlin dataset show that the proposed model is able to efficiently distill knowledge from natural images and achieves superior performance than the current state of the art. Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Hongdong Li |
AAAI | 4 |
| 2019 | Lending Orientation to Neural Networks for Cross-View Geo-LocalizationabstractThis paper studies image-based geo-localization (IBL) problem using ground-to-aerial cross-view matching. The goal is to predict the spatial location of a ground-level query image by matching it to a large geotagged aerial image database (e.g., satellite imagery). This is a challenging task due to the drastic differences in their viewpoints and visual appearances. Existing deep learning methods for this problem have been focused on maximizing feature similarity between spatially close-by image pairs, while minimizing other images pairs which are far apart. They do so by deep feature embedding based on visual appearance in those ground-and-aerial images. However, in everyday life, humans commonly use orientation information as an important cue for the task of spatial localization. Inspired by this insight, this paper proposes a novel method which endows deep neural networks with the `commonsense' of orientation. Given a ground-level spherical panoramic image as query input (and a large georeferenced satellite image database), we design a Siamese network which explicitly encodes the orientation (i.e., spherical directions) of each pixel of the images. Our method significantly boosts the discriminative power of the learned deep features, leading to a much higher recall and precision outperforming all previous methods. Our network is also more compact using only 1/5th number of parameters than a previously best-performing network. To evaluate the generalization of our method, we also created a large-scale cross-view localization benchmark containing 100K geotagged ground-aerial pairs covering a city. Our codes and datasets are available at https://github.com/Liumouliu/OriCNN. Liu Liu 0009, Hongdong Li |
CVPR | 2 |
| 2019 | The Alignment of the Spheres: Globally-Optimal Spherical Mixture Alignment for Camera Pose EstimationabstractDetermining the position and orientation of a calibrated camera from a single image with respect to a 3D model is an essential task for many applications. When 2D-3D correspondences can be obtained reliably, perspective-n-point solvers can be used to recover the camera pose. However, without the pose it is non-trivial to find cross-modality correspondences between 2D images and 3D models, particularly when the latter only contains geometric information. Consequently, the problem becomes one of estimating pose and correspondences jointly. Since outliers and local optima are so prevalent, robust objective functions and global search strategies are desirable. Hence, we cast the problem as a 2D-3D mixture model alignment task and propose the first globally-optimal solution to this formulation under the robust L2 distance between mixture distributions. We derive novel bounds on this objective function and employ branch-and-bound to search the 6D space of camera poses, guaranteeing global optimality without requiring a pose estimate. To accelerate convergence, we integrate local optimization, implement GPU bound computations, and provide an intuitive way to incorporate side information such as semantic labels. The algorithm is evaluated on challenging synthetic and real datasets, outperforming existing approaches and reliably converging to the global optimum. Dylan Campbell, Lars Petersson, Laurent Kneip, Hongdong Li, Stephen Gould |
CVPR | 4 |
| 2019 | Noise-Aware Unsupervised Deep Lidar-Stereo FusionabstractIn this paper, we present LidarStereoNet, the first unsupervised Lidar-stereo fusion network, which can be trained in an end-to-end manner without the need of ground truth depth maps. By introducing a novel ``Feedback Loop'' to connect the network input with output, LidarStereoNet could tackle both noisy Lidar points and misalignment between sensors that have been ignored in existing Lidar-stereo fusion work. Besides, we propose to incorporate the piecewise planar model into the network learning to further constrain depths to conform to the underlying 3D geometry. Extensive quantitative and qualitative evaluations on both real and synthetic datasets demonstrate the superiority of our method, which outperforms state-of-the-art stereo matching, depth completion and Lidar-Stereo fusion approaches significantly. Xuelian Cheng, Yiran Zhong, Yuchao Dai, Pan Ji, Hongdong Li |
CVPR | 5 |
| 2019 | ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous DrivingabstractAutonomous driving has attracted remarkable attention from both industry and academia. An important task is to estimate 3D properties (e.g. translation, rotation and shape) of a moving or parked vehicle on the road. This task, while critical, is still under-researched in the computer vision community – partially owing to the lack of large scale and fully-annotated 3D car database suitable for autonomous driving research. In this paper, we contribute the first large scale database suitable for 3D car instance understanding – ApolloCar3D. The dataset contains 5,277 driving images and over 60K car instances, where each car is fitted with an industry-grade 3D CAD model with absolute model size and semantically labelled keypoints. This dataset is above 20× larger than PASCAL3D+ and KITTI, the current state-of-the-art. To enable efficient labelling in 3D, we build a pipeline by considering 2D-3D keypoint correspondences for a single instance and 3D relationship among multiple instances. Equipped with such dataset, we build various baseline algorithms with the state-of-the-art deep convolutional neural networks. Specifically, we first segment each car with a pre-trained Mask R-CNN, and then regress towards its 3D pose and shape based on a deformable 3D car model with or without using semantic keypoints. We show that using keypoints significantly improves fitting performance. Finally, we develop a new 3D metric jointly considering 3D pose and 3D shape, allowing for comprehensive evaluation and ablation study. Xibin Song, Peng Wang 0001, Dingfu Zhou, Chenye Guan, Yuchao Dai, Hongdong Li, Ruigang Yang |
CVPR | 8 |
| 2019 | Deep Stacked Hierarchical Multi-Patch Network for Image DeblurringabstractDespite deep end-to-end learning methods have shown their superiority in removing non-uniform motion blur, there still exist major challenges with the current multi-scale and scale-recurrent models: 1) Deconvolution/upsampling operations in the coarse-to-fine scheme result in expensive runtime; 2) Simply increasing the model depth with finer-scale levels cannot improve the quality of deblurring. To tackle the above problems, we present a deep {hierarchical multi-patch network} inspired by Spatial Pyramid Matching to deal with blurry images via a fine-to-coarse hierarchical representation. To deal with the performance saturation w.r.t. depth, we propose a stacked version of our multi-patch model. Our proposed basic multi-patch model achieves the state-of-the-art performance on the GoPro dataset while enjoying a 40$\times$ faster runtime compared to current multi-scale methods. With 30ms to process an image at 1280$\times$720 resolution, it is the first real-time deep motion deblurring model for 720p images at 30fps. For stacked networks, significant improvements (over 1.2dB) are achieved on the GoPro dataset by increasing the network depth. Moreover, by varying the depth of the stacked model, one can adapt the performance and runtime of the same network for different application scenarios. Yuchao Dai, Hongdong Li, Piotr Koniusz |
CVPR | 3 |
| 2019 | Learning Joint Gait Representation via Quintuplet Loss MinimizationabstractGait recognition is an important biometric method popularly used in video surveillance, where the task is to identify people at a distance by their walking patterns from video sequences. Most of the current successful approaches for gait recognition either use a pair of gait images to form a cross-gait representation or rely on a single gait image for unique-gait representation. These two types of representations emperically complement one another. In this paper, we propose a new Joint Unique-gait and Cross-gait Network (JUCNet), to combine the advantages of unique-gait representation with that of cross-gait representation, leading to an significantly improved performance. Another key contribution of this paper is a novel quintuplet loss function, which simultaneously increases the inter-class differences by pushing representations extracted from different subjects apart and decreases the intra-class variations by pulling representations extracted from the same subject together. Experiments show that our method achieves the state-of-the-art performance tested on standard benchmark datasets, demonstrating its superiority over existing methods. Kaihao Zhang, Wenhan Luo, Lin Ma 0002, Wei Liu 0005, Hongdong Li |
CVPR | 5 |
| 2019 | Unsupervised Deep Epipolar Flow for Stationary or Dynamic ScenesabstractUnsupervised deep learning for optical flow computation has achieved promising results. Most existing deep-net based methods rely on image brightness consistency and local smoothness constraint to train the networks. Their performance degrades at regions where repetitive textures or occlusions occur. In this paper, we propose Deep Epipolar Flow, an unsupervised optical flow method which incorporates global geometric constraints into network learning. In particular, we investigate multiple ways of enforcing the epipolar constraint in flow estimation. To alleviate a ``chicken-and-egg'' type of problem encountered in dynamic scenes where multiple motions may be present, we propose a low-rank constraint as well as a union-of-subspaces constraint for training. Experimental results on various benchmarking datasets show that our method achieves competitive performance compared with supervised methods and outperforms state-of-the-art unsupervised deep-learning methods. Yiran Zhong, Pan Ji, Yuchao Dai, Hongdong Li |
CVPR | 5 |
| 2019 | Stochastic Attraction-Repulsion Embedding for Large Scale Image LocalizationabstractThis paper tackles the problem of large-scale image-based localization (IBL) where the spatial location of a query image is determined by finding out the most similar reference images in a large database. For solving this problem, a critical task is to learn discriminative image representation that captures informative information relevant for localization. We propose a novel representation learning method having higher location-discriminating power. It provides the following contributions: 1) we represent a place (location) as a set of exemplar images depicting the same landmarks and aim to maximize similarities among intra-place images while minimizing similarities among inter-place images; 2) we model a similarity measure as a probability distribution on L2-metric distances between intra-place and inter-place image representations; 3) we propose a new Stochastic Attraction and Repulsion Embedding (SARE) loss function minimizing the KL divergence between the learned and the actual probability distributions; 4) we give theoretical comparisons between SARE, triplet ranking and contrastive losses. It provides insights into why SARE is better by analyzing gradients. Our SARE loss is easy to implement and pluggable to any CNN. Experiments show that our proposed method improves the localization performance on standard benchmarks by a large margin. Demonstrating the broad applicability of our method, we obtained the third place out of 209 teams in the 2018 Google Landmark Retrieval Challenge. Our code and model are available at https://github.com/Liumouliu/deepIBL. Liu Liu 0009, Hongdong Li, Yuchao Dai |
ICCV | 2 |
| 2019 | Learning Trajectory Dependencies for Human Motion PredictionabstractHuman motion prediction, i.e., forecasting future body poses given observed pose sequence, has typically been tackled with recurrent neural networks (RNNs). However, as evidenced by prior work, the resulted RNN models suffer from prediction errors accumulation, leading to undesired discontinuities in motion prediction. In this paper, we propose a simple feed-forward deep network for motion prediction, which takes into account both temporal smoothness and spatial dependencies among human body joints. In this context, we then propose to encode temporal information by working in trajectory space, instead of the traditionally-used pose space. This alleviates us from manually defining the range of temporal dependencies (or temporal convolutional filter size, as done in previous work). Moreover, spatial dependency of human pose is encoded by treating a human pose as a generic graph (rather than a human skeletal kinematic tree) formed by links between every pair of body joints. Instead of using a pre-defined graph structure, we design a new graph convolutional network to learn graph connectivity automatically. This allows the network to capture long range dependencies beyond that of human kinematic tree. We evaluate our approach on several standard benchmark datasets for motion prediction, including Human3.6M, the CMU motion capture dataset and 3DPW. Our experiments clearly demonstrate that the proposed approach achieves state of the art performance, and is applicable to both angle-based and position-based pose representations. The code is available at https://github.com/wei-mao-2019/LearnTrajDep. Wei Mao 0001, Miaomiao Liu 0001, Mathieu Salzmann, Hongdong Li |
ICCV | 4 |
| 2019 | Neural Collaborative Subspace ClusteringabstractWe introduce the Neural Collaborative Subspace Clustering, a neural model that discovers clusters of data points drawn from a union of low-dimensional subspaces. In contrast to previous attempts, our model runs without the aid of spectral clustering. This makes our algorithm one of the kinds that can gracefully scale to large datasets. At its heart, our neural model benefits from a classifier which determines whether a pair of points lies on the same subspace or not. Essential to our model is the construction of two affinity matrices, one from the classifier and the other from a notion of subspace self-expressiveness, to supervise training in a collaborative scheme. We thoroughly assess and contrast the performance of our model against various state-of-the-art clustering algorithms including deep subspace-based ones. Tong Zhang 0023, Pan Ji, Mehrtash Harandi, Wenbing Huang 0001, Hongdong Li |
ICML | 5 |
| 2019 | Spatial-Aware Feature Aggregation for Image based Cross-View Geo-LocalizationabstractIn this paper, we develop a new deep network to explicitly address these inherent differences between ground and aerial views. We observe there exist some approximate domain correspondences between ground and aerial images. Specifically, pixels lying on the same azimuth direction in an aerial image approximately correspond to a vertical image column in the ground view image. Thus, we propose a two-step approach to exploit this prior knowledge. The first step is to apply a regular polar transform to warp an aerial image such that its domain is closer to that of a ground-view panorama. Note that polar transform as a pure geometric transformation is agnostic to scene content, hence cannot bring the two domains into full alignment. Then, we add a subsequent spatial-attention mechanism which further brings corresponding deep features closer in the embedding space. To improve the robustness of feature representation, we introduce a feature aggregation strategy via learning multiple spatial embeddings. By the above two-step approach, we achieve more discriminative deep representations, facilitating cross-view Geo-localization more accurate. Our experiments on standard benchmark datasets show significant performance boosting, achieving more than doubled recall rate compared with the previous state of the art. Yujiao Shi 0002, Liu Liu 0009, Xin Yu 0002, Hongdong Li |
NeurIPS | 4 |
| 2019 | Adversarial Spatio-Temporal Learning for Video DeblurringabstractCamera shake or target movement often leads to undesired blur effects in videos captured by a hand-held camera. Despite significant efforts having been devoted to video-deblur research, two major challenges remain: 1) how to model the spatio-temporal characteristics across both the spatial domain (i.e., image plane) and the temporal domain (i.e., neighboring frames) and 2) how to restore sharp image details with respect to the conventionally adopted metric of pixel-wise errors. In this paper, to address the first challenge, we propose a deblurring network (DBLRNet) for spatial-temporal learning by applying a 3D convolution to both the spatial and temporal domains. Our DBLRNet is able to capture jointly spatial and temporal information encoded in neighboring frames, which directly contributes to the improved video deblur performance. To tackle the second challenge, we leverage the developed DBLRNet as a generator in the generative adversarial network (GAN) architecture and employ a content loss in addition to an adversarial loss for efficient adversarial training. The developed network, which we name as deblurring GAN, is tested on two standard benchmarks and achieves the state-of-the-art performance. Kaihao Zhang, Wenhan Luo, Yiran Zhong, Lin Ma 0002, Wei Liu 0005, Hongdong Li |
IEEE Trans. Image Process. | 6 |
| 2019 | Canny-VO: Visual Odometry With RGB-D Cameras Based on Geometric 3-D-2-D Edge AlignmentabstractThis paper reviews the classical problem of free-form curve registration and applies it to an efficient RGB-D visual odometry system called Canny-VO, as it efficiently tracks all Canny edge features extracted from the images. Two replacements for the distance transformation commonly used in edge registration are proposed: approximate nearest neighbor fields and oriented nearest neighbor fields. 3-D-2-D edge alignment benefits from these alternative formulations in terms of both efficiency and accuracy. It removes the need for the more computationally demanding paradigms of data-to-model registration, bilinear interpolation, and subgradient computation. To ensure robustness of the system in the presence of outliers and sensor noise, the registration is formulated as a maximum a posteriori problem and the resulting weighted least-squares objective is solved by the iteratively reweighted least-squares method. A variety of robust weight functions are investigated and the optimal choice is made based on the statistics of the residual errors. Efficiency is furthermore boosted by an adaptively sampled definition of the nearest neighbor fields. Extensive evaluations on public SLAM benchmark sequences demonstrate state-of-the-art performance and an advantage over classical Euclidean distance fields. Yi Zhou 0010, Hongdong Li, Laurent Kneip |
IEEE Trans. Robotics | 2 |
| 2018 | Scalable Dense Non-Rigid Structure-From-Motion: A Grassmannian PerspectiveabstractThis paper addresses the task of dense non-rigid structure-front-motion (NRSfM) using multiple images. State-of-the-art methods to this problem are often hurdled by scalability, expensive computations, and noisy measurements. Further, recent methods to NRSfM usually either assume a small number of sparse feature points or ignore local non-linearities of shape deformations, and thus cannot reliably model complex non-rigid deformations. To address these issues, in this paper, we propose a new approach for dense NRSfM by modeling the problem on a Grassmann manifold. Specifically, we assume the complex non-rigid deformations lie on a union of local linear subspaces both spatially and temporally. This naturally allows for a compact representation of the complex non-rigid deformation over frames. We provide experimental results on several synthetic and real benchmark datasets. The procured results clearly demonstrate that our method, apart from being scalable and more accurate than state-of-the-art methods, is also more robust to noise and generalizes to highly nonlinear deformations. Suryansh Kumar 0001, Anoop Cherian, Yuchao Dai, Hongdong Li |
CVPR | 4 |
| 2018 | Structure From Recurrent Motion: From Rigidity to RecurrencyabstractThis paper proposes a new method for Non-Rigid Structure-from-Motion (NRSfM) from a long monocular video sequence observing a non-rigid object performing recurrent and possibly repetitive dynamic action. Departing from the traditional idea of using linear low-order or low-rank shape model for the task of NRSfM, our method exploits the property of shape recurrency (i.e., many deforming shapes tend to repeat themselves in time). We show that recurrency is in fact a generalized rigidity. Based on this, we reduce NRSfM problems to rigid ones provided that certain recurrency condition is satisfied. Given such a reduction, standard rigid-SfM techniques are directly applicable (without any change) to the reconstruction of non-rigid dynamic shapes. To implement this idea as a practical approach, this paper develops efficient algorithms for automatic recurrency detection, as well as camera view clustering via a rigidity-check. Experiments on both simulated sequences and real data demonstrate the effectiveness of the method. Since this paper offers a novel perspective on rethinking structure-from-motion, we hope it will inspire other new problems in the field. Xiu Li 0003, Hongdong Li, Hanbyul Joo, Yebin Liu, Yaser Sheikh |
CVPR | 2 |
| 2018 | Stereo Computation for a Single Mixture Image
Yiran Zhong, Yuchao Dai, Hongdong Li |
ECCV (9) | 3 |
| 2018 | Open-World Stereo Video Matching with Deep RNN
Yiran Zhong, Hongdong Li, Yuchao Dai |
ECCV (2) | 2 |
| 2018 | Semi-dense 3D Reconstruction with a Stereo Event Camera
Yi Zhou 0010, Guillermo Gallego 0002, Henri Rebecq, Laurent Kneip, Hongdong Li, Davide Scaramuzza 0001 |
ECCV (1) | 5 |
| 2018 | 3D Geometry-Aware Semantic Labeling of Outdoor Street ScenesabstractThis paper is concerned with the problem of how to better exploit 3D geometric information for dense semantic image labeling. Existing methods often treat the available 3D geometry information (e.g., 3D depth-map) simply as an additional image channel besides the R-G-B color channels, and apply the same technique for RGB image labeling. In this paper, we demonstrate that directly performing 3D convolution in the framework of a residual connected 3D voxel top-down modulation network can lead to superior results. Specifically, we propose a 3D semantic labeling method to label outdoor street scenes whenever a dense depth map is available. Experiments on the “Synthia” and “Cityscape” datasets show our method outperforms the state-of-the-art methods, suggesting such a simple 3D representation is effective in incorporating 3D geometric information. Yiran Zhong, Yuchao Dai, Hongdong Li |
ICPR | 3 |
| 2018 | Fully Convolutional Neural Networks for Road Detection with Multiple Cues IntegrationabstractRoad detection from images is a key task in autonomous driving. The recent advent of deep learning (and in particular, CNN or convolutional neural networks) has greatly improved the performance of road detection algorithms. In this paper, we show how to fuse multiple different cues under the same convolutional network framework. Specifically, we adopt a pre-trained Resnet-lOl to extract feature maps from RGB images; we then connect it with three extra deconvolution layers. These deconvolution layers is trained conditioning on appropriate image cues, and in our case they are a height image (i.e. elevation map obtained by e.g. Lidar scanner), image gradient, and position map. We also design two skip layers to speed up the convergence. Experiments on KITTI benchmark show competitive performance of our new networks. Jianfeng Lu 0003, Chunxia Zhao, Hongdong Li |
ICRA | 4 |
| 2018 | A simple way to detect disease-associated cellular molecular alterations from mixed-cell blood samplesabstractBlood is a promising surrogate for solid tissue to investigate disease-associated molecular biomarkers. However, proportion changes of the constituent cells in the often-used peripheral whole blood (PWB) or peripheral blood mononuclear cell (PBMC) samples may influence the detection of cell-specific alterations under disease states. We propose a simple method, Ref-REO, to detect molecular alterations in leukocytes using the mixed-cell blood samples. The method is based on the predetermined within-sample relative expression orderings (REOs) of genes in purified leukocytes of healthy people. Both the simulated and real mixed-cell blood gene expression profiles were used to evaluate the method. Approximately 99% of the differentially expressed genes (DEGs) detected by Ref-REO in the simulated mixed-cell data are owing to the transcriptional alterations in leukocytes rather than the proportion changes of leukocytes. For the real mixed-cell data, the DEGs detected by Ref-REO in the PBMCs expression data for systemic lupus erythematosus (SLE) patients overlap significantly with the DEGs detected in the expression data of SLE CD4 + T cells and B cells and they are mainly enriched with mRNA editing and interferon-associated genes. The detected DEGs in the PWB data for lung carcinoma patients are significantly enriched with coagulation-associated functional categories that are closely associated with cancer progression. In conclusion, the proposed method is capable of detecting the disease-associated leukocyte-specific molecular alterations, using mixed-cell blood samples, which provides simple, transferable and easy-to-use candidates for disease biomarkers. Guini Hong, Hongdong Li, Weicheng Zheng, Jing Li 0115, Meirong Chi, Zheng Guo 0002 |
Briefings Bioinform. | 2 |
| 2018 | Semisupervised and Weakly Supervised Road Detection Based on Generative Adversarial NetworksabstractRoad detection is a key component of autonomous driving; however, most fully supervised learning road detection methods suffer from either insufficient training data or high costs of manual annotation. To overcome these problems, we propose a semisupervised learning (SSL) road detection method based on generative adversarial networks (GANs) and a weakly supervised learning (WSL) method based on conditional GANs. Specifically, in our SSL method, the generator generates the road detection results of labeled and unlabeled images, and then they are fed into the discriminator, which assigns a label on each input to judge whether it is labeled. Additionally, in WSL method we add another network to predict road shapes of input images and use them in both generator and discriminator to constrain the learning progress. By training under these frameworks, the discriminators can guide a latent annotation process on the unlabeled data; therefore, the networks can learn better representations of road areas and leverage the feature distributions on both labeled and unlabeled data. The experiments are carried out on KITTI ROAD benchmark, and the results show our methods achieve the state-of-the-art performances. Jianfeng Lu 0003, Chunxia Zhao, Shaodi You, Hongdong Li |
IEEE Signal Process. Lett. | 5 |
| 2018 | Not All Negatives Are Equal: Learning to Track With Multiple Background ClustersabstractConventional tracking-by-detection approaches for visual object tracking often assume that the task at hand is a binary foreground-versus-background classification problem, in which the background is a single, generic, and all-inclusive class. In contrast, here we argue that the background appearance, for the most part, possesses a more complicated structure that would benefit from further partitioning into multiple contextual clusters. Our observation is that, although the background class is contemplated to contain a vast intra-class variation, during the tracking process, only a small portion of this diversity is present at the current frame around the foreground object. This observation motivates us to build multiple fine-grained foreground-versus-contextual-cluster models that provide more discriminative classifications, and consequently more robust and accurate foreground object tracking. For each cluster, we employ a structured output support vector machine (SSVM), and in an online manner, we combine the responses of multiple classifiers. To this end, we apply a top-level SSVM that models the tracked foreground object. We show that our refined modeling of the background is better than naïvely growing the complexity of a single foreground-background classifier, i.e., increasing the number of support vectors that existing approaches rely on, which cause overfitting issues. Our extensive evaluations on large benchmark data sets demonstrate that our tracker consistently outperforms the current state-of-the-art while having comparable computational requirements. Gao Zhu, Fatih Porikli, Hongdong Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Robust and Efficient Relative Pose With a Multi-Camera System for Autonomous Driving in Highly Dynamic EnvironmentsabstractThis paper studies the relative pose problem for autonomous vehicles driving in highly dynamic and possibly cluttered environments. This is a challenging scenario due to the existence of multiple, large, and independently moving objects in the environment, which often leads to an excessive portion of outliers and results in erroneous motion estimation. Existing algorithms cannot cope with such situations well. This paper proposes a new algorithm for relative pose estimation using a multi-camera system with multiple non-overlapping cameras. The method works robustly even when the number of outliers is overwhelming. By exploiting specific prior knowledge of the autonomous driving scene, we have developed an efficient 4-point algorithm for multi-camera relative pose estimation, which admits analytic solutions by solving a polynomial root finding equation, and runs extremely fast (at about 0.5 μs per root). When the solver is used in combination with a new random sample consensus sampling scheme by exploiting the conjugate motion constraint, we are able to quickly prune unpromising hypotheses and significantly improve the chance of finding inliers. Experiments on synthetic data have validated the performance of the proposed algorithm. Tests on real data further confirm the method's practical relevance. Liu Liu 0009, Hongdong Li, Yuchao Dai, Quan Pan 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Neural Aggregation Network for Video Face RecognitionabstractThis paper presents a Neural Aggregation Network (NAN) for video face recognition. The network takes a face video or face image set of a person with a variable number of face images as its input, and produces a compact, fixed-dimension feature representation for recognition. The whole network is composed of two modules. The feature embedding module is a deep Convolutional Neural Network (CNN) which maps each face image to a feature vector. The aggregation module consists of two attention blocks which adaptively aggregate the feature vectors to form a single feature inside the convex hull spanned by them. Due to the attention mechanism, the aggregation is invariant to the image order. Our NAN is trained with a standard classification or verification loss without any extra supervision signal, and we found that it automatically learns to advocate high-quality face images while repelling low-quality ones such as blurred, occluded and improperly exposed faces. The experiments on IJB-A, YouTube Face, Celebrity-1000 video face recognition benchmarks show that it consistently outperforms naive aggregation methods and achieves the state-of-the-art accuracy. Jiaolong Yang, Peiran Ren, Dongqing Zhang, Dong Chen 0003, Fang Wen 0001, Hongdong Li, Gang Hua 0001 |
CVPR | 6 |
| 2017 | Globally-Optimal Inlier Set Maximisation for Simultaneous Camera Pose and Feature CorrespondenceabstractEstimating the 6-DoF pose of a camera from a single image relative to a pre-computed 3D point-set is an important task for many computer vision applications. Perspective-n-Point (PnP) solvers are routinely used for camera pose estimation, provided that a good quality set of 2D-3D feature correspondences are known beforehand. However, finding optimal correspondences between 2D key-points and a 3D point-set is non-trivial, especially when only geometric (position) information is known. Existing approaches to the simultaneous pose and correspondence problem use local optimisation, and are therefore unlikely to find the optimal solution without a good pose initialisation, or introduce restrictive assumptions. Since a large proportion of outliers are common for this problem, we instead propose a globally-optimal inlier set cardinality maximisation approach which jointly estimates optimal camera pose and optimal correspondences. Our approach employs branch-and-bound to search the 6D space of camera poses, guaranteeing global optimality without requiring a pose prior. The geometry of SE(3) is used to find novel upper and lower bounds for the number of inliers and local optimisation is integrated to accelerate convergence. The evaluation empirically supports the optimality proof and shows that the method performs much more robustly than existing approaches, including on a large-scale outdoor data-set. Dylan Campbell, Lars Petersson, Laurent Kneip, Hongdong Li |
ICCV | 4 |
| 2017 | "Maximizing Rigidity" Revisited: A Convex Programming Approach for Generic 3D Shape Reconstruction from Multiple Perspective ViewsabstractRigid structure-from-motion (RSfM) and non-rigid structure-from-motion (NRSfM) have long been treated in the literature as separate (different) problems. Inspired by a previous work which solved directly for 3D scene structure by factoring the relative camera poses out, we revisit the principle of “maximizing rigidity” in structure-from-motion literature, and develop a unified theory which is applicable to both rigid and non-rigid structure reconstruction in a rigidity-agnostic way. We formulate these problems as a convex semi-definite program, imposing constraints that seek to apply the principle of minimizing non-rigidity. Our results demonstrate the efficacy of the approach, with stateof- the-art accuracy on various 3D reconstruction problems. Pan Ji, Hongdong Li, Yuchao Dai, Ian D. Reid 0001 |
ICCV | 2 |
| 2017 | Monocular Dense 3D Reconstruction of a Complex Dynamic Scene from Two Perspective FramesabstractThis paper proposes a new approach for monocular dense 3D reconstruction of a complex dynamic scene from two perspective frames. By applying superpixel over-segmentation to the image, we model a generically dynamic (hence non-rigid) scene with a piecewise planar and rigid approximation. In this way, we reduce the dynamic reconstruction problem to a “3D jigsaw puzzle ” problem which takes pieces from an unorganized “soup of superpixels". We show that our method provides an effective solution to the inherent relative scale ambiguity in structure-from-motion. Since our method does not assume a template prior, or per-object segmentation, or knowledge about the rigidity of the dynamic scene, it is applicable to a wide range of scenarios. Extensive experiments on both synthetic and real monocular sequences demonstrate the superiority of our method compared with the state-of-the-art methods. Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
ICCV | 3 |
| 2017 | Efficient Global 2D-3D Matching for Camera Localization in a Large-Scale 3D MapabstractGiven an image of a street scene in a city, this paper develops a new method that can quickly and precisely pinpoint at which location (as well as viewing direction) the image was taken, against a pre-stored large-scale 3D point-cloud map of the city. We adopt the recently developed 2D-3D direct feature matching framework for this task [23,31,32,42-44]. This is a challenging task especially for large-scale problems. As the map size grows bigger, many 3D points in the wider geographical area can be visually very similar-or even identical-causing severe ambiguities in 2D-3D feature matching. The key is to quickly and unambiguously find the correct matches between a query image and the large 3D map. Existing methods solve this problem mainly via comparing individual features' visual similarities in a local and per feature manner, thus only local solutions can be found, inadequate for large-scale applications. In this paper, we introduce a global method which harnesses global contextual information exhibited both within the query image and among all the 3D points in the map. This is achieved by a novel global ranking algorithm, applied to a Markov network built upon the 3D map, which takes account of not only visual similarities between individual 2D-3D matches, but also their global compatibilities (as measured by co-visibility) among all matching pairs found in the scene. Tests on standard benchmark datasets show that our method achieved both higher precision and comparable recall, compared with the state-of-the-art. Liu Liu 0009, Hongdong Li, Yuchao Dai |
ICCV | 2 |
| 2017 | Part-based fine-grained bird image retrieval respecting species correlationabstractMost of the existing works on fine-grained bird image categorization and retrieval focus on finding similar images from the same species and often give little importance to inter-species similarity. In this paper, we devise a new fine-grained retrieval task that searches similar instances from different species. To this end, we propose a two-step strategy. In the first step, we search for visually similar parts to a query image using a deep convolutional neural network (CNN). To improve the quality of the retrieved candidates, we incorporate structural cues into the CNN using a novel part-pooling layer. In the second step, we re-rank the retrieved candidates improving the species diversity. We achieve this by formulating a novel ranking function that balances between the similarity of the candidates to the queried parts, while decreasing the similarity to the query species. We provide experiments on the benchmark CUB200 dataset and demonstrate clear benefits of our schemes. Hongdong Li, Anoop Cherian, Hongxun Yao |
ICIP | 2 |
| 2017 | Semi-dense visual odometry for RGB-D cameras using approximate nearest neighbour fieldsabstractThis paper presents a robust and efficient semidense visual odometry solution for RGB-D cameras. The core of our method is a 2D-3D ICP pipeline which estimates the pose of the sensor by registering the projection of a 3D semidense map of a reference frame with the 2D semi-dense region extracted in the current frame. The processing is speeded up by efficiently implemented approximate nearest neighbour fields under the Euclidean distance criterion, which permits the use of compact Gauss-Newton updates in the optimization. The registration is formulated as a maximum a posterior problem to deal with outliers and sensor noise, and the equivalent weighted least squares problem is consequently solved by iteratively reweighted least squares method. A variety of robust weight functions are tested and the optimum is determined based on the probabilistic characteristics of the sensor model. Extensive evaluation on publicly available RGB-D datasets shows that the proposed method predominantly outperforms existing state-of-the-art methods. Yi Zhou 0010, Laurent Kneip, Hongdong Li |
ICRA | 3 |
| 2017 | Accurate extrinsic calibration between monocular camera and sparse 3D Lidar points without markersabstractIt is of practical interest to automatically calibrate the multiple sensors in autonomous vehicles. In this paper, we deal with an interesting case when used low-resolution Lidar and present a practical approach to extrinsic calibration between monocular camera and Lidar with sparse 3D measurements. We formulate the problem as directly minimizing the feature error evaluated between frames following the way of image warping. To overcome the difficulties in the optimization problem, we propose to use the distance transform and further projection error model to obtain the key approximated edge points that are sensitive to the loss function. Finally, the loss minimization is solved by an efficient random selection algorithm. Experimental results on KITTI dataset show that our proposed method can achieve competitive results and an improvement in translation estimation particularly. Zhipeng Xiao, Hongdong Li, Dingfu Zhou, Yuchao Dai, Bin Dai 0001 |
Intelligent Vehicles Symposium | 2 |
| 2017 | Deep Subspace Clustering NetworksabstractWe present a novel deep neural network architecture for unsupervised subspace clustering. This architecture is built upon deep auto-encoders, which non-linearly map the input data into a latent space. Our key idea is to introduce a novel self-expressive layer between the encoder and the decoder to mimic the "self-expressiveness" property that has proven effective in traditional subspace clustering. Being differentiable, our new self-expressive layer provides a simple but effective way to learn pairwise affinities between all data points through a standard back-propagation procedure. Being nonlinear, our neural-network based method is able to cluster data points having complex (often nonlinear) structures. We further propose pre-training and fine-tuning strategies that let us effectively learn the parameters of our subspace clustering networks. Our experiments show that the proposed method significantly outperforms the state-of-the-art unsupervised subspace clustering methods. Pan Ji, Tong Zhang 0023, Hongdong Li, Mathieu Salzmann, Ian D. Reid 0001 |
NIPS | 3 |
| 2017 | Automated detection and tracking of slalom paddlers from broadcast image sequences using cascade classifiers and discriminative correlation filters
Ami Drory, Gao Zhu, Hongdong Li, Richard I. Hartley |
Comput. Vis. Image Underst. | 3 |
| 2017 | Moving object detection and segmentation in urban environments from a moving platform
Dingfu Zhou, Vincent Frémont, Benjamin Quost, Yuchao Dai, Hongdong Li |
Image Vis. Comput. | 5 |
| 2017 | Spatio-temporal union of subspaces for multi-body non-rigid structure-from-motion
Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
Pattern Recognit. | 3 |
| 2016 | Multi-Body Non-Rigid Structure-from-MotionabstractIn this paper, we present the first multi-body non-rigid structure-from-motion (SFM) method, which simultaneously reconstructs and segments multiple objects that are undergoing non-rigid deformation over time. Under our formulation, 3D trajectories for each non-rigid object can be well approximated with a sparse affine combination of other 3D trajectories from the same object. The resultant optimization is solved by the alternating direction method of multipliers (ADMM). We demonstrate the efficacy of the proposed method through extensive experiments on both synthetic and real data sequences. Our method outperforms other alternative methods, such as first clustering the 2D feature tracks to groups and then doing non-rigid reconstruction in each group or first conducting 3D reconstruction by using single subspace assumption and then clustering the 3D trajectories into groups. Suryansh Kumar 0001, Yuchao Dai, Hongdong Li |
3DV | 3 |
| 2016 | Divide and Conquer: Efficient Density-Based Tracking of 3D Sensors in Manhattan Worlds
Yi Zhou 0010, Laurent Kneip, Cristian Rodriguez Opazo, Hongdong Li |
ACCV (5) | 4 |
| 2016 | Model-Free Multiple Object Tracking with Shared Proposals
Gao Zhu, Fatih Porikli, Hongdong Li |
ACCV (2) | 3 |
| 2016 | Rolling Shutter Camera Relative Pose: Generalized Epipolar GeometryabstractThe vast majority of modern consumer-grade cameras employ a rolling shutter mechanism. In dynamic geometric computer vision applications such as visual SLAM, the so-called rolling shutter effect therefore needs to be properly taken into account. A dedicated relative pose solver appears to be the first problem to solve, as it is of eminent importance to bootstrap any derivation of multi-view geometry. However, despite its significance, it has received inadequate attention to date. This paper presents a detailed investigation of the geometry of the rolling shutter relative pose problem. We introduce the rolling shutter essential matrix, and establish its link to existing models such as the push-broom cameras, summarized in a clean hierarchy of multi-perspective cameras. The generalization of well-established concepts from epipolar geometry is completed by a definition of the Sampson distance in the rolling shutter case. The work is concluded with a careful investigation of the introduced epipolar geometry for rolling shutter cameras on several dedicated benchmarks. Yuchao Dai, Hongdong Li, Laurent Kneip |
CVPR | 2 |
| 2016 | Robust Multi-Body Feature Tracker: A Segmentation-Free ApproachabstractFeature tracking is a fundamental problem in computer vision, with applications in many computer vision tasks, such as visual SLAM and action recognition. This paper introduces a novel multi-body feature tracker that exploits a multi-body rigidity assumption to improve tracking robustness under a general perspective camera model. A conventional approach to addressing this problem would consist of alternating between solving two subtasks: motion segmentation and feature tracking under rigidity constraints for each segment. This approach, however, requires knowing the number of motions, as well as assigning points to motion groups, which is typically sensitive to the motion estimates. By contrast, here, we introduce a segmentationfree solution to multi-body feature tracking that bypasses the motion assignment step and reduces to solving a series of subproblems with closed-form solutions. Our experiments demonstrate the benefits of our approach in terms of tracking accuracy and robustness to noise. Pan Ji, Hongdong Li, Mathieu Salzmann, Yiran Zhong |
CVPR | 2 |
| 2016 | Robust Optical Flow Estimation of Double-Layer Images under Transparency or ReflectionabstractThis paper deals with a challenging, frequently encountered, yet not properly investigated problem in two-frame optical flow estimation. That is, the input frames are compounds of two imaging layers - one desired background layer of the scene, and one distracting, possibly moving layer due to transparency or reflection. In this situation, the conventional brightness constancy constraint - the cornerstone of most existing optical flow methods - will no longer be valid. In this paper, we propose a robust solution to this problem. The proposed method performs both optical flow estimation, and image layer separation. It exploits a generalized double-layer brightness consistency constraint connecting these two tasks, and utilizes the priors for both of them. Experiments on both synthetic data and real images have confirmed the efficacy of the proposed method. To the best of our knowledge, this is the first attempt towards handling generic optical flow fields of two-frame images containing transparency or reflection. Jiaolong Yang, Hongdong Li, Yuchao Dai, Robby T. Tan |
CVPR | 2 |
| 2016 | Beyond Local Search: Tracking Objects Everywhere with Instance-Specific ProposalsabstractMost tracking-by-detection methods employ a local search window around the predicted object location in the current frame assuming the previous location is accurate, the trajectory is smooth, and the computational capacity permits a search radius that can accommodate the maximum speed yet small enough to reduce mismatches. These, however, may not be valid always, in particular for fast and irregularly moving objects. Here, we present an object tracker that is not limited to a local search window and has ability to probe efficiently the entire frame. Our method generates a small number of "high-quality" proposals by a novel instance-specific objectness measure and evaluates them against the object model that can be adopted from an existing tracking-by-detection approach as a core tracker. During the tracking process, we update the object model concentrating on hard false-positives supplied by the proposals, which help suppressing distractors caused by difficult background clutters, and learn how to re-rank proposals according to the object model. Since we reduce significantly the number of hypotheses the core tracker evaluates, we can use richer object descriptors and stronger detector. Our method outperforms most recent state-of-the-art trackers on popular tracking benchmarks, and provides improved robustness for fast moving objects as well as for ultra lowframerate videos. Gao Zhu, Fatih Porikli, Hongdong Li |
CVPR | 3 |
| 2016 | Learning Image Matching by Simply Watching Video
Gucan Long, Laurent Kneip, José M. Álvarez 0004, Hongdong Li |
ECCV (6) | 4 |
| 2016 | Non-iterative, fast SE(3) path smoothingabstractIn this paper, we present a fast, non-iterative approach to smooth a noisy input on the Special Euclidean Group, SE(3) manifold. The translational part can be smoothed by a simple Gaussian convolution. We then proposed a novel approach to rotation smoothing. Unlike existing rotation smoothing methods using either iterative optimization methods or stochastic filtering methods, our method allows direct computation of the smoothing result and allows parallelization of the computation. Furthermore, we have done a comparative study on Jia and Evans's method published in 2014, and shown that our method can better smooth an input rotation sequence, with shorter computational time. The smoothed camera path is then used for video stabilisation, which shows fluid and smooth camera motion. Yonhon Ng, Bomin Jiang, Changbin Yu, Hongdong Li |
IROS | 4 |
| 2016 | Real-time rotation estimation for dense depth sensors in piece-wise planar environmentsabstractLow-drift rotation estimation is a crucial part of any accurate odometry system. In this paper, we focus on the problem of 3D rotation estimation with dense depth sensors in environments that consist of piece-wise planar structures, such as corridors and office rooms. An efficient mean-shift paradigm is developed to extract and track planar modes in the surface normal vector distribution on the unit sphere. Robust and piece-wise drift-free behavior is achieved by registering the bundle of planar modes from the current frame with respect to a reference frame using a general ℓ1-norm regression scheme. We furthermore add a memory scheme to the regular birth and death of modes, which further compensates accumulated rotational drift when previously discovered modes are revisited. We discuss the robustness issue and evaluate our algorithm on both custom synthetic as well as real publicly available datasets. Our experimental results demonstrate high robustness and effectiveness of the proposed algorithm. Yi Zhou 0010, Laurent Kneip, Hongdong Li |
IROS | 3 |
| 2016 | Reliable scale estimation and correction for monocular Visual OdometryabstractRecovering absolute scale (i.e. metric information) from monocular vision system is a very challenging problem yet is highly desirable for vision-based autonomous driving. This paper proposes a new method for scale recovery, based on the idea of knowing camera height (relative to ground-plane). While this idea of using known camera height is not new in this context, existing implementations of this idea suffer significantly from severe numerical instability arisen in the ground plane homography decomposition stage. Our novel contribution of this work is to alleviate this issue by a divide and conquer approach, i.e. decomposing the motion parameters in the homography from the structure parameters of the ground plane. We also describe a robust procedure to correct scale drift in the monocular visual odometry system. Experimental results on KITTI standard benchmark dataset [1] and our self-collected driving dataset both show significant improvements. Dingfu Zhou, Yuchao Dai, Hongdong Li |
Intelligent Vehicles Symposium | 3 |
| 2016 | A proteogenomic approach to understand splice isoform functions through sequence and expression-based computational modelingabstractThe products of multi-exon genes are a mixture of alternatively spliced isoforms, from which the translated proteins can have similar, different or even opposing functions. It is therefore essential to differentiate and annotate functions for individual isoforms. Computational approaches provide an efficient complement to expensive and time-consuming experimental studies. The input data of these methods range from DNA sequence, to RNA selection pressure, to expressed sequence tags, to full-length complementary DNA, to exon array, to RNA-seq expression, to proteomic data. Notably, RNA-seq technology generates quantitative profiling of transcript expression at the genome scale, with an unprecedented amount of expression data available for developing isoform function prediction methods. Integrative analysis of these data at different molecular levels enables a proteogenomic approach to systematically interrogate isoform functions. Here, we briefly review the state-of-the-art methods according to their input data sources, discuss their advantages and limitations and point out potential ways to improve prediction accuracies. Hongdong Li, Gilbert S. Omenn, Yuanfang Guan |
Briefings Bioinform. | 1 |
| 2016 | Fast Rotation Search with Stereographic Projections for 3D RegistrationabstractRegistering two 3D point clouds involves estimating the rigid transform that brings the two point clouds into alignment. Recently there has been a surge of interest in using branch-and-bound (BnB) optimisation for point cloud registration. While BnB guarantees globally optimal solutions, it is usually too slow to be practical. A fundamental source of difficulty lies in the search for the rotational parameters. In this work, first by assuming that the translation is known, we focus on constructing a fast rotation search algorithm. With respect to an inherently robust geometric matching criterion, we propose a novel bounding function for BnB that is provably tighter than previously proposed bounds. Further, we also propose a fast algorithm to evaluate our bounding function. Our idea is based on using stereographic projections to precompute and index all possible point matches in spatial R-trees for rapid evaluations. The result is a fast and globally optimal rotation search algorithm. To conduct full 3D registration, we co-optimise the translation by embedding our rotation search kernel in a nested BnB algorithm. Since the inner rotation search is very efficient, the overall 6DOF optimisation is speeded up significantly without losing global optimality. On various challenging point clouds, including those taken out of lab settings, our approach demonstrates superior efficiency. Álvaro Parra Bustos, Tat-Jun Chin, Anders P. Eriksson, Hongdong Li, David Suter |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2016 | Go-ICP: A Globally Optimal Solution to 3D ICP Point-Set RegistrationabstractThe Iterative Closest Point (ICP) algorithm is one of the most widely used methods for point-set registration. However, being based on local iterative optimization, ICP is known to be susceptible to local minima. Its performance critically relies on the quality of the initialization and only local optimality is guaranteed. This paper presents the first globally optimal algorithm, named Go-ICP, for Euclidean (rigid) registration of two 3D point-sets under the$L_2$error metric defined in ICP. The Go-ICP method is based on a branch-and-bound scheme that searches the entire 3D motion space$SE(3)$. By exploiting the special structure of$SE(3)$geometry, we derive novel upper and lower bounds for the registration error function. Local ICP is integrated into the BnB scheme, which speeds up the new method while guaranteeing global optimality. We also discuss extensions, addressing the issue of outlier robustness. The evaluation demonstrates that the proposed method is able to produce reliable registration results regardless of the initialization. Go-ICP can be applied in scenarios where an optimal solution is desirable or where a good initialization is not always available. Jiaolong Yang, Hongdong Li, Dylan Campbell, Yunde Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Contour Completion Without Region SegmentationabstractContour completion plays an important role in visual perception, where the goal is to group fragmented low-level edge elements into perceptually coherent and salient contours. Most existing methods for contour completion have focused on pixelwise detection accuracy. In contrast, fewer methods have addressed the global contour closure effect, despite psychological evidences for its importance. This paper proposes a purely contour-based higher order CRF model to achieve contour closure, through local connectedness approximation. This leads to a simplified problem structure, where our higher order inference problem can be transformed into an integer linear program and be solved efficiently. Compared with the methods based on the same bottom-up edge detector, our method achieves a superior contour grouping ability (measured by Rand index), a comparable precision-recall performance, and more visually pleasing results. Our results suggest that contour closure can be effectively achieved in contour domain, in contrast to a popular view that segmentation is essential for this purpose. Yansheng Ming, Hongdong Li, Xuming He 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | SDICP: Semi-Dense Tracking based on Iterative Closest PointsabstractThe paper addresses the problem of camera tracking, which denotes the continuous image-based computation of a camera’s position and orientation with respect to a reference frame. The method aims at regular cameras, which means that 3D-3D registration methods applicable to RGB-D cameras are not an option. The tracked frame contains only 2D information, thus requiring a solution to the absolute pose or 2D-3D registration problem. While traditional solutions to camera tracking [3] rely on sparse feature correspondences, the community has recently seen a number of direct photometric registration methods such as Newcombe et al. [8] and Engel et al. [1]. [1] is conceptually similar to [8], however gains computational efficiency by reducing the computation from dense to semi-dense regions that correspond to a thresholded edge-map of the image. Photometric methods have the more general advantage of compensating for appearance variations caused by perspective view-point changes, whereas classical sparse methods often rely on static feature descriptors only (providing at most rotation and scale invariant properties [5, 6]). However, photometric registration techniques inherently suffer from the disability to overcome large disparities, where large sometimes means even just a couple of pixels [7]. Many photometric registration techniques therefore depend on pyramidal subsampling schemes in order to alleviate this problem. The goal of the present paper is a novel 2D-3D registration paradigm for semi-dense depth maps that relies on the Iterative Closest Point (ICP) technique, and thus a reintroduction of geometric error minimization as a valid alternative for real-time monocular camera tracking in the case of semi-dense features. An example semi-dense depth map is indicated in Figure 1. In comparison to photometric registration techniques, our ICP technique has the conceptual advantage of requiring neither isotropic enlargement of the employed semi-dense regions, nor pyramidal subsampling. The work is in line with Feldmar et al. [2], Tomono [9], and Klein and Murray [4], which already attempt curve or edge registration in 2D using ICP. Based on a hypothesized relative pose, the basic idea consists of warping a reference curve into the tracked image based on a prior 3D model or depth (in our case semi-dense) inside a reference frame. From a mathematical point of view, our idea may be formulated as follows. Let Laurent Kneip, Zhou Yi, Hongdong Li |
BMVC | 3 |
| 2015 | Iteratively reweighted graph cut for multi-label MRFs with non-convex priorsabstractWhile widely acknowledged as highly effective in computer vision, multi-label MRFs with non-convex priors are difficult to optimize. To tackle this, we introduce an algorithm that iteratively approximates the original energy with an appropriately weighted surrogate energy that is easier to minimize. Our algorithm guarantees that the original energy decreases at each iteration. In particular, we consider the scenario where the global minimizer of the weighted surrogate energy can be obtained by a multi-label graph cut algorithm, and show that our algorithm then lets us handle of large variety of non-convex priors. We demonstrate the benefits of our method over state-of-the-art MRF energy minimization techniques on stereo and inpainting problems. Thalaiyasingam Ajanthan, Richard I. Hartley, Mathieu Salzmann, Hongdong Li |
CVPR | 4 |
| 2015 | Dense, accurate optical flow estimation with piecewise parametric modelabstractThis paper proposes a simple method for estimating dense and accurate optical flow field. It revitalizes an early idea of piecewise parametric flow model. A key innovation is that, we fit a flow field piecewise to a variety of parametric models, where the domain of each piece (i.e., each piece's shape, position and size) is determined adaptively, while at the same time maintaining a global inter-piece flow continuity constraint. We achieve this by a multi-model fitting scheme via energy minimization. Our energy takes into account both the piecewise constant model assumption and the flow field continuity constraint, enabling the proposed method to effectively handle both homogeneous motions and complex motions. The experiments on three public optical flow benchmarks (KITTI, MPI Sintel, and Middlebury) show the superiority of our method compared with the state of the art: it achieves top-tier performances on all the three benchmarks. Jiaolong Yang, Hongdong Li |
CVPR | 2 |
| 2015 | Shape Interaction Matrix Revisited and Robustified: Efficient Subspace Clustering with Corrupted and Incomplete DataabstractThe Shape Interaction Matrix (SIM) is one of the earliest approaches to performing subspace clustering (i.e., separating points drawn from a union of subspaces). In this paper, we revisit the SIM and reveal its connections to several recent subspace clustering methods. Our analysis lets us derive a simple, yet effective algorithm to robustify the SIM and make it applicable to realistic scenarios where the data is corrupted by noise. We justify our method by intuitive examples and the matrix perturbation theory. We then show how this approach can be extended to handle missing data, thus yielding an efficient and general subspace clustering algorithm. We demonstrate the benefits of our approach over state-of-the-art subspace clustering methods on several challenging motion segmentation and face clustering problems, where the data includes corruptions and missing measurements. Pan Ji, Mathieu Salzmann, Hongdong Li |
ICCV | 3 |
| 2015 | Lie-Struck: Affine Tracking on Lie Groups Using Structured SVMabstractThis paper presents a novel and reliable tracking-by detection method for image regions that undergo affine transformations such as translation, rotation, scale, dilatation and shear deformations, which span the six degrees of freedom of motion. Our method takes advantage of the intrinsic Lie group structure of the 2D affine motion matrices and imposes this motion structure on a kernelized structured output SVM classifier that provides an appearance based prediction function to directly estimate the object transformation between frames using geodesic distances on manifolds unlike the existing methods proceeding by linearizing the motion. We demonstrate that these combined motion and appearance model structures greatly improve the tracking performance while an incorporated particle filter on the motion hypothesis space keeps the computational load feasible. Experimentally, we show that our algorithm is able to outperform state-of-the-art affine trackers in various scenarios. Gao Zhu, Fatih Porikli, Yansheng Ming, Hongdong Li |
WACV | 4 |
| 2015 | Kernel Methods on Riemannian Manifolds with Gaussian RBF KernelsabstractIn this paper, we develop an approach to exploiting kernel methods with manifold-valued data. In many computer vision problems, the data can be naturally represented as points on a Riemannian manifold. Due to the non-Euclidean geometry of Riemannian manifolds, usual Euclidean computer vision and machine learning algorithms yield inferior results on such data. In this paper, we define Gaussian radial basis function (RBF)-based positive definite kernels on manifolds that permit us to embed a given manifold with a corresponding metric in a high dimensional reproducing kernel Hilbert space. These kernels make it possible to utilize algorithms developed for linear spaces on nonlinear manifold-valued data. Since the Gaussian RBF defined with any given metric is not always positive definite, we present a unified framework for analyzing the positive definiteness of the Gaussian RBF on a generic metric space. We then use the proposed framework to identify positive definite kernels on two specific manifolds commonly encountered in computer vision: the Riemannian manifold of symmetric positive definite matrices and the Grassmann manifold, i.e., the Riemannian manifold of linear subspaces of a Euclidean space. We show that many popular algorithms designed for Euclidean spaces, such as support vector machines, discriminant analysis and principal component analysis can be generalized to Riemannian manifolds with the help of such positive definite Gaussian kernels. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Winding Number Constrained Contour DetectionabstractSalient contour detection can benefit from the integration of both contour cues and region cues. However, this task is difficult due to different nature of region representations and contour representations. To solve this problem, this paper proposes an energy minimization framework based on winding number constraints. In this framework, both region cues, such as color/texture homogeneity, and contour cues, such as local contrast and continuity, are represented in a joint objective function, which has both region and contour labels. The key problem is how to design constraints that ensure the topological consistency of the two kinds of labels. Our technique is based on the topological concept of winding number. Using a fast method for winding number computation, a small number of linear constraints are derived to ensure label consistency. Our method is instantiated by ratio-based energy functions. By successfully integrating both region and contour cues, our method shows advantages over competitive methods. Our method is extended to incorporate user interaction, which leads to further improvements. Yansheng Ming, Hongdong Li, Xuming He 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Optimizing over Radial Kernels on Compact ManifoldsabstractWe tackle the problem of optimizing over all possible positive definite radial kernels on Riemannian manifolds for classification. Kernel methods on Riemannian manifolds have recently become increasingly popular in computer vision. However, the number of known positive definite kernels on manifolds remain very limited. Furthermore, most kernels typically depend on at least one parameter that needs to be tuned for the problem at hand. A poor choice of kernel, or of parameter value, may yield significant performance drop-off. Here, we show that positive definite radial kernels on the unit n-sphere, the Grassmann manifold and Kendall's shape manifold can be expressed in a simple form whose parameters can be automatically optimized within a support vector machine framework. We demonstrate the benefits of our kernel learning algorithm on object, face, action and shape recognition. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
CVPR | 4 |
| 2014 | Efficient Computation of Relative Pose for Multi-camera SystemsabstractWe present a novel solution to compute the relative pose of a generalized camera. Existing solutions are either not general, have too high computational complexity, or require too many correspondences, which impedes an efficient or accurate usage within Ransac schemes. We factorize the problem as a low-dimensional, iterative optimization over relative rotation only, directly derived from well-known epipolar constraints. Common generalized cameras often consist of camera clusters, and give rise to omni-directional landmark observations. We prove that our iterative scheme performs well in such practically relevant situations, eventually resulting in computational efficiency similar to linear solvers, and accuracy close to bundle adjustment, while using less correspondences. Experiments on both virtual and real multi-camera systems prove superior overall performance for robust, real-time multi-camera motion-estimation. Laurent Kneip, Hongdong Li |
CVPR | 2 |
| 2014 | Expanding the Family of Grassmannian Kernels: An Embedding Perspective
Mehrtash Harandi, Mathieu Salzmann, Sadeep Jayasumana, Richard I. Hartley, Hongdong Li |
ECCV (7) | 5 |
| 2014 | Robust Motion Segmentation with Unknown Correspondences
Pan Ji, Hongdong Li, Mathieu Salzmann, Yuchao Dai |
ECCV (6) | 2 |
| 2014 | UPnP: An Optimal O(n) Solution to the Absolute Pose Problem with Universal Applicability
Laurent Kneip, Hongdong Li, Yongduek Seo |
ECCV (1) | 2 |
| 2014 | Optimal Essential Matrix Estimation via Inlier-Set Maximization
Jiaolong Yang, Hongdong Li, Yunde Jia |
ECCV (1) | 2 |
| 2014 | Null space clustering with applications to motion segmentation and face clusteringabstractThe problems of motion segmentation and face clustering can be addressed in a framework of subspace clustering methods. In this paper, we tackle the more general problem of clustering data points lying in a union of low-dimensional linear(or affine) subspaces, which can be naturally applied in motion segmentation and face clustering. For data points drawn from linear (or affine) subspaces, we propose a novel algorithm called Null Space Clustering (NSC), utilizing the null space of the data matrix to construct the affinity matrix. To better deal with noise and outliers, it is converted to an equivalent problem with Frobenius norm minimization, which can be solved efficiently. We demonstrate that the proposed NSC leads to improved performance in terms of clustering accuracy and efficiency when compared to state-of-the-art algorithms on two well-known datasets, i.e., Hopkins 155 and Extended Yale B. Pan Ji, Yiran Zhong, Hongdong Li, Mathieu Salzmann |
ICIP | 3 |
| 2014 | Object category detection by incorporating mid-level grouping cuesabstractMany state-of-the-art semantic object detection methods locate category-level objects by finding optimal bounding boxes. However, the accuracy of localization is compromised, when the shape of an object does not conform to rectangular bounding boxes. As a remedy, some recent work locates an object based on superpixel classification. However, the increased flexibility in shape modeling also means less control, and methods which mostly rely on high-level semantic (category-level) classification cue have difficulty in producing “regular” segments which align well with objects. To solve this problem, we propose a novel energy-minimization method which explicitly models the “objectness” of a segment by incorporating mid-level grouping cues. The highlevel classification cue is integrated with mid-level grouping features in a principled ratio energy function whose global optimal solution can be obtained efficiently. Our method compares favorably with state-of-the-art methods on public datasets. Gao Zhu, Yansheng Ming, Hongdong Li |
ICIP | 3 |
| 2014 | Efficient dense subspace clusteringabstractIn this paper, we tackle the problem of clustering data points drawn from a union of linear (or affine) subspaces. To this end, we introduce an efficient subspace clustering algorithm that estimates dense connections between the points lying in the same subspace. In particular, instead of following the standard compressive sensing approach, we formulate subspace clustering as a Frobenius norm minimization problem, which inherently yields denser con- nections between the data points. While in the noise-free case we rely on the self-expressiveness of the observations, in the presence of noise we simultaneously learn a clean dictionary to represent the data. Our formulation lets us address the subspace clustering problem efficiently. More specifically, the solution can be obtained in closed-form for outlier-free observations, and by performing a series of linear operations in the presence of outliers. Interestingly, we show that our Frobenius norm formulation shares the same solution as the popular nuclear norm minimization approach when the data is free of any noise, or, in the case of corrupted data, when a clean dictionary is learned. Our experimental evaluation on motion segmentation and face clustering demonstrates the benefits of our algorithm in terms of clustering accuracy and efficiency. Pan Ji, Mathieu Salzmann, Hongdong Li |
WACV | 3 |
| 2014 | Modeling dynamic functional relationship networks and application to ex vivo human erythroid differentiationabstractMOTIVATION: Functional relationship networks, which summarize the probability of co-functionality between any two genes in the genome, could complement the reductionist focus of modern biology for understanding diverse biological processes in an organism. One major limitation of the current networks is that they are static, while one might expect functional relationships to consistently reprogram during the differentiation of a cell lineage. To address this potential limitation, we developed a novel algorithm that leverages both differentiation stage-specific expression data and large-scale heterogeneous functional genomic data to model such dynamic changes. We then applied this algorithm to the time-course RNA-Seq data we collected for ex vivo human erythroid cell differentiation. RESULTS: Through computational cross-validation and literature validation, we show that the resulting networks correctly predict the (de)-activated functional connections between genes during erythropoiesis. We identified known critical genes, such as HBD and GATA1, and functional connections during erythropoiesis using these dynamic networks, while the traditional static network was not able to provide such information. Furthermore, by comparing the static and the dynamic networks, we identified novel genes (such as OSBP2 and PDZK1IP1) that are potential drivers of erythroid cell differentiation. This novel method of modeling dynamic networks is applicable to other differentiation processes where time-course genome-scale expression data are available, and should assist in generating greater understanding of the functional dynamics at play across the genome during development. AVAILABILITY AND IMPLEMENTATION: The network described in this article is available at http://guanlab.ccmb.med.umich.edu/stageSpecificNetwork. Lihong Shi, Hongdong Li, Ridvan Eksi, James Douglas Engel, Yuanfang Guan |
Bioinform. | 3 |
| 2014 | A Simple Prior-Free Method for Non-rigid Structure-from-Motion Factorization
Yuchao Dai, Hongdong Li, Mingyi He |
Int. J. Comput. Vis. | 2 |
| 2014 | Recognizing Gaits Across Views Through Correlated Motion Co-ClusteringabstractHuman gait is an important biometric feature, which can be used to identify a person remotely. However, view change can cause significant difficulties for gait recognition because it will alter available visual features for matching substantially. Moreover, it is observed that different parts of gait will be affected differently by view change. By exploring relations between two gaits from two different views, it is also observed that a part of gait in one view is more related to a typical part than any other parts of gait in another view. A new method proposed in this paper considers such variance of correlations between gaits across views that is not explicitly analyzed in the other existing methods. In our method, a novel motion co-clustering is carried out to partition the most related parts of gaits from different views into the same group. In this way, relationships between gaits from different views will be more precisely described based on multiple groups of the motion co-clustering instead of a single correlation descriptor. Inside each group, a linear correlation between gait information across views is further maximized through canonical correlation analysis (CCA). Consequently, gait information in one view can be projected onto another view through a linear approximation under the trained CCA subspaces. In the end, a similarity between gaits originally recorded from different views can be measured under the approximately same view. Comprehensive experiments based on widely adopted gait databases have shown that our method outperforms the state-of-the-art. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li, Liang Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2013 | Kernel Methods on the Riemannian Manifold of Symmetric Positive Definite MatricesabstractSymmetric Positive Definite (SPD) matrices have become popular to encode image information. Accounting for the geometry of the Riemannian manifold of SPD matrices has proven key to the success of many algorithms. However, most existing methods only approximate the true shape of the manifold locally by its tangent plane. In this paper, inspired by kernel methods, we propose to map SPD matrices to a high dimensional Hilbert space where Euclidean geometry applies. To encode the geometry of the manifold in the mapping, we introduce a family of provably positive definite kernels on the Riemannian manifold of SPD matrices. These kernels are derived from the Gaussian kernel, but exploit different metrics on the manifold. This lets us extend kernel-based algorithms developed for Euclidean spaces, such as SVM and kernel PCA, to the Riemannian manifold of SPD matrices. We demonstrate the benefits of our approach on the problems of pedestrian detection, object categorization, texture analysis, 2D motion segmentation and Diffusion Tensor Imaging (DTI) segmentation. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
CVPR | 4 |
| 2013 | Winding Number for Region-Boundary Consistent Salient Contour ExtractionabstractThis paper aims to extract salient closed contours froman image. For this vision task, both region segmentation cues (e.g. color/texture homogeneity) and boundary detection cues (e.g. local contrast, edge continuity and contour closure) play important and complementary roles. In this paper we show how to combine both cues in a unified framework. The main focus is given to how to maintain the consistency (compatibility) between the region cues and the boundary cues. To this ends, we introduce the use of winding number-a well-known concept in topology-as a powerful mathematical device. By this device, the region-boundary consistency is represented as aset of simple linear relationships. Our method is applied to the figure-ground segmentation problem. The experiments show clearly improved results. Yansheng Ming, Hongdong Li, Xuming He 0001 |
CVPR | 2 |
| 2013 | A Framework for Shape Analysis via Hilbert Space EmbeddingabstractWe propose a framework for 2D shape analysis using positive definite kernels defined on Kendall's shape manifold. Different representations of 2D shapes are known to generate different nonlinear spaces. Due to the nonlinearity of these spaces, most existing shape classification algorithms resort to nearest neighbor methods and to learning distances on shape spaces. Here, we propose to map shapes on Kendall's shape manifold to a high dimensional Hilbert space where Euclidean geometry applies. To this end, we introduce a kernel on this manifold that permits such a mapping, and prove its positive definiteness. This kernel lets us extend kernel-based algorithms developed for Euclidean spaces, such as SVM, MKL and kernel PCA, to the shape manifold. We demonstrate the benefits of our approach over the state-of-the-art methods on shape classification, clustering and retrieval. Sadeep Jayasumana, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
ICCV | 3 |
| 2013 | Multi-view 3D Reconstruction from Uncalibrated Radially-Symmetric CamerasabstractWe present a new multi-view 3D Euclidean reconstruction method for arbitrary uncalibrated radially-symmetric cameras, which needs no calibration or any camera model parameters other than radial symmetry. It is built on the radial 1D camera model [25], a unified mathematical abstraction to different types of radially-symmetric cameras. We formulate the problem of multi-view reconstruction for radial 1D cameras as a matrix rank minimization problem. Efficient implementation based on alternating direction continuation is proposed to handle scalability issue for real-world applications. Our method applies to a wide range of omni directional cameras including both dioptric and catadioptric (central and non-central) cameras. Additionally, our method deals with complete and incomplete measurements under a unified framework elegantly. Experiments on both synthetic and real images from various types of cameras validate the superior performance of our new method, in terms of numerical accuracy and robustness. Jae-Hak Kim, Yuchao Dai, Hongdong Li, Jonghyuk Kim |
ICCV | 3 |
| 2013 | Go-ICP: Solving 3D Registration Efficiently and Globally OptimallyabstractRegistration is a fundamental task in computer vision. The Iterative Closest Point (ICP) algorithm is one of the widely-used methods for solving the registration problem. Based on local iteration, ICP is however well-known to suffer from local minima. Its performance critically relies on the quality of initialization, and only local optimality is guaranteed. This paper provides the very first globally optimal solution to Euclidean registration of two 3D point sets or two 3D surfaces under the L2 error. Our method is built upon ICP, but combines it with a branch-and-bound (BnB) scheme which searches the 3D motion space SE(3) efficiently. By exploiting the special structure of the underlying geometry, we derive novel upper and lower bounds for the ICP error function. The integration of local ICP and global BnB enables the new method to run efficiently in practice, and its optimality is exactly guaranteed. We also discuss extensions, addressing the issue of outlier robustness. Jiaolong Yang, Hongdong Li, Yunde Jia |
ICCV | 2 |
| 2013 | Symmetry detection via contour groupingabstractThis paper presents a simple but effective model for detecting the symmetric axes of bilaterally symmetric objects in unsegmented natural scene images. Our model constructs a directed graph of symmetry interaction. Every node in the graph represents a matched pair of features, and every directed edge represents the interaction between nodes. The bilateral symmetry detection problem is then formulated as finding the star subgraph with maximal weight. The star structure ensures the consistency between grouped nodes while the optimal star subgraph can be found in polynomial time. Our model makes prediction based on contour cue: each node in the graph represents a pair of edge segments. Compared with the Loy and Eklundh's method which used SIFT feature, our model can often produce better results for the images containing limited texture. This advantage is demonstrated on two natural scene image sets. Yansheng Ming, Hongdong Li, Xuming He 0001 |
ICIP | 2 |
| 2013 | Single-shot extrinsic calibration of a generically configured RGB-D camera rig from scene constraintsabstractWith the increasing use of commodity RGB-D cameras for computer vision, robotics, mixed and augmented reality and other areas, it is of significant practical interest to calibrate the relative pose between a depth (D) camera and an RGB camera in these types of setups. In this paper, we propose a new single-shot, correspondence-free method to extrinsically calibrate a generically configured RGB-D camera rig. We formulate the extrinsic calibration problem as one of geometric 2D-3D registration which exploits scene constraints to achieve single-shot extrinsic calibration. Our method first reconstructs sparse point clouds from a single-view 2D image. These sparse point clouds are then registered with dense point clouds from the depth camera. Finally, we directly optimize the warping quality by evaluating scene constraints in 3D point clouds. Our single-shot extrinsic calibration method does not require correspondences across multiple color images or across different modalities and it is more flexible than existing methods. The scene constraints can be very simple and we demonstrate that a scene containing three sheets of paper is sufficient to obtain reliable calibration and with a lower geometric error than existing methods. Jiaolong Yang, Yuchao Dai, Hongdong Li, Henry J. Gardner, Yunde Jia |
ISMAR | 3 |
| 2013 | Rotation Averaging
Richard I. Hartley, Jochen Trumpf, Yuchao Dai, Hongdong Li |
Int. J. Comput. Vis. | 4 |
| 2013 | A Branch-and-Bound Approach to Correspondence and Grouping ProblemsabstractData correspondence/grouping under an unknown parametric model is a fundamental topic in computer vision. Finding feature correspondences between two images is probably the most popular application of this research field, and is the main motivation of our work. It is a key ingredient for a wide range of vision tasks, including three-dimensional reconstruction and object recognition. Existing feature correspondence methods are based on either local appearance similarity or global geometric consistency or a combination of both in some heuristic manner. None of these methods is fully satisfactory, especially in the presence of repetitive image textures or mismatches. In this paper, we present a new algorithm that combines the benefits of both appearance-based and geometry-based methods and mathematically guarantees a global optimization. Our algorithm accepts the two sets of features extracted from two images as input, and outputs the feature correspondences with the largest number of inliers, which verify both the appearance similarity and geometric constraints. Specifically, we formulate the problem as a mixed integer program and solve it efficiently by a series of linear programs via a branch-and-bound procedure. We subsequently generalize our framework in the context of data correspondence/grouping under an unknown parametric model and show it can be applied to certain classes of computer vision problems. Our algorithm has been validated successfully on synthesized data and challenging real images. Jean-Charles Bazin, Hongdong Li, In-So Kweon, Cédric Demonceaux, Pascal Vasseur, Katsushi Ikeuchi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Projective Multiview Structure and Motion from Element-Wise FactorizationabstractThe Sturm-Triggs type iteration is a classic approach for solving the projective structure-from-motion (SfM) factorization problem, which iteratively solves the projective depths, scene structure, and camera motions in an alternated fashion. Like many other iterative algorithms, the Sturm-Triggs iteration suffers from common drawbacks, such as requiring a good initialization, the iteration may not converge or may only converge to a local minimum, and so on. In this paper, we formulate the projective SfM problem as a novel and original element-wise factorization (i.e., Hadamard factorization) problem, as opposed to the conventional matrix factorization. Thanks to this formulation, we are able to solve the projective depths, structure, and camera motions simultaneously by convex optimization. To address the scalability issue, we adopt a continuation-based algorithm. Our method is a global method, in the sense that it is guaranteed to obtain a globally optimal solution up to relaxation gap. Another advantage is that our method can handle challenging real-world situations such as missing data and outliers quite easily, and all in a natural and unified manner. Extensive experiments on both synthetic and real images show comparable results compared with the state-of-the-art methods. Yuchao Dai, Hongdong Li, Mingyi He |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Systematically Differentiating Functions for Alternatively Spliced Isoforms through Integrating RNA-seq DataabstractIntegrating large-scale functional genomic data has significantly accelerated our understanding of gene functions. However, no algorithm has been developed to differentiate functions for isoforms of the same gene using high-throughput genomic data. This is because standard supervised learning requires 'ground-truth' functional annotations, which are lacking at the isoform level. To address this challenge, we developed a generic framework that interrogates public RNA-seq data at the transcript level to differentiate functions for alternatively spliced isoforms. For a specific function, our algorithm identifies the 'responsible' isoform(s) of a gene and generates classifying models at the isoform level instead of at the gene level. Through cross-validation, we demonstrated that our algorithm is effective in assigning functions to genes, especially the ones with multiple isoforms, and robust to gene expression levels and removal of homologous gene pairs. We identified genes in the mouse whose isoforms are predicted to have disparate functionalities and experimentally validated the 'responsible' isoforms using data from mammary tissue. With protein structure modeling and experimental evidence, we further validated the predicted isoform functional differences for the genes Cdkn2a and Anxa6. Our generic framework is the first to predict and differentiate functions for alternatively spliced isoforms, instead of genes, using genomic data. It is extendable to any base machine learner and other species with alternatively spliced isoforms, and shifts the current gene-centered function prediction to isoform-level predictions. Ridvan Eksi, Hongdong Li, Rajasree Menon, Yuchen Wen, Gilbert S. Omenn, Matthias Kretzler, Yuanfang Guan |
PLoS Comput. Biol. | 2 |
| 2013 | A New View-Invariant Feature for Cross-View Gait RecognitionabstractHuman gait is an important biometric feature which is able to identify a person remotely. However, change of view causes significant difficulties for recognizing gaits. This paper proposes a new framework to construct a new view-invariant feature for cross-view gait recognition. Our view-normalization process is performed in the input layer (i.e., on gait silhouettes) to normalize gaits from arbitrary views. That is, each sequence of gait silhouettes recorded from a certain view is transformed onto the common canonical view by using corresponding domain transformation obtained through invariant low-rank textures (TILTs). Then, an improved scheme of procrustes shape analysis (PSA) is proposed and applied on a sequence of the normalized gait silhouettes to extract a novel view-invariant gait feature based on procrustes mean shape (PMS) and consecutively measure a gait similarity based on procrustes distance (PD). Comprehensive experiments were carried out on widely adopted gait databases. It has been shown that the performance of the proposed method is promising when compared with other existing methods in the literature. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2012 | A simple prior-free method for non-rigid structure-from-motion factorizationabstractThis paper proposes a simple “prior-free” method for solving non-rigid structure-from-motion factorization problems. Other than using the basic low-rank condition, our method does not assume any extra prior knowledge about the nonrigid scene or about the camera motions. Yet, it runs reliably, produces optimal result, and does not suffer from the inherent basis-ambiguity issue which plagued many conventional nonrigid factorization techniques. Our method is easy to implement, which involves solving no more than an SDP (semi-definite programming) of small and fixed size, a linear Least-Squares or trace-norm minimization. Extensive experiments have demonstrated that it outperforms most of the existing linear methods of nonrigid factorization. This paper offers not only new theoretical insight, but also a practical, everyday solution, to non-rigid structure-from-motion. Yuchao Dai, Hongdong Li, Mingyi He |
CVPR | 2 |
| 2012 | Connected contours: A new contour completion model that respects the closure effectabstractContour Completion plays an important role in visual perception, where the goal is to group fragmented low-level edge elements into perceptually coherent and salient contours. This process is often considered as guided by some middle-level Gestalt principles. Most existing methods for contour completion have focused on utilizing rather local Gestalt laws such as good-continuity and proximity. In contrast, much fewer methods have addressed the global contour closure effect, despite that many psychological evidences have shown the usefulness of closure in perceptual grouping. This paper proposes a novel higher-order CRF model to address the contour closure effect, through local connectedness approximation. This leads to a simplified problem structure, where the higher-order inference can be formulated as an integer linear program (ILP) and solved by an efficient cutting-plane variant. Tested on the BSDS benchmark, our method achieves a comparable precision-recall performance, a superior contour grouping ability (measured by Rand index), and more visually pleasing results, compared with existing methods. Yansheng Ming, Hongdong Li, Xuming He 0001 |
CVPR | 2 |
| 2012 | An Efficient Hidden Variable Approach to Minimal-Case Camera Motion EstimationabstractIn this paper, we present an efficient new approach for solving two-view minimal-case problems in camera motion estimation, most notably the so-called five-point relative orientation problem and the six-point focal-length problem. Our approach is based on the hidden variable technique used in solving multivariate polynomial systems. The resulting algorithm is conceptually simple, which involves a relaxation which replaces monomials in all but one of the variables to reduce the problem to the solution of sets of linear equations, as well as solving a polynomial eigenvalue problem (polyeig). To efficiently find the polynomial eigenvalues, we make novel use of several numeric techniques, which include quotient-free Gaussian elimination, Levinson-Durbin iteration, and also a dedicated root-polishing procedure. We have tested the approach on different minimal cases and extensions, with satisfactory results obtained. Both the executables and source codes of the proposed algorithms are made freely downloadable. Richard I. Hartley, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Cross-view and multi-view gait recognitions based on view transformation model using multi-layer perceptron
Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
Pattern Recognit. Lett. | 4 |
| 2012 | Gait Recognition Under Various Viewing Angles Based on Correlated Motion RegressionabstractIt is well recognized that gait is an important biometric feature to identify a person at a distance, e.g., in video surveillance application. However, in reality, change of viewing angle causes significant challenge for gait recognition. A novel approach using regression-based view transformation model (VTM) is proposed to address this challenge. Gait features from across views can be normalized into a common view using learned VTM(s). In principle, a VTM is used to transform gait feature from one viewing angle (source) into another viewing angle (target). It consists of multiple regression processes to explore correlated walking motions, which are encoded in gait features, between source and target views. In the learning processes, sparse regression based on the elastic net is adopted as the regression function, which is free from the problem of overfitting and results in more stable regression models for VTM construction. Based on widely adopted gait database, experimental results show that the proposed method significantly improves upon existing VTM-based methods and outperforms most other baseline methods reported in the literature. Several practical scenarios of applying the proposed method for gait recognition under various views are also discussed in this paper. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2012 | Gait Recognition Across Various Walking Speeds Using Higher Order Shape Configuration Based on a Differential Composition ModelabstractGait has been known as an effective biometric feature to identify a person at a distance. However, variation of walking speeds may lead to significant changes to human walking patterns. It causes many difficulties for gait recognition. A comprehensive analysis has been carried out in this paper to identify such effects. Based on the analysis, Procrustes shape analysis is adopted for gait signature description and relevant similarity measurement. To tackle the challenges raised by speed change, this paper proposes a higher order shape configuration for gait shape description, which deliberately conserves discriminative information in the gait signatures and is still able to tolerate the varying walking speed. Instead of simply measuring the similarity between two gaits by treating them as two unified objects, a differential composition model (DCM) is constructed. The DCM differentiates the different effects caused by walking speed changes on various human body parts. In the meantime, it also balances well the different discriminabilities of each body part on the overall gait similarity measurements. In this model, the Fisher discriminant ratio is adopted to calculate weights for each body part. Comprehensive experiments based on widely adopted gait databases demonstrate that our proposed method is efficient for cross-speed gait recognition and outperforms other state-of-the-art methods. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2011 | Pairwise Shape configuration-based PSA for gait recognition under small viewing angle changeabstractTwo main components of Procrustes Shape Analysis (PSA) are adopted and adapted specifically to address gait recognition under small viewing angle change: 1) Procrustes Mean Shape (PMS) for gait signature description; 2) Procrustes Distance (PD) for similarity measurement. Pairwise Shape Configuration (PSC) is proposed as a shape descriptor in place of existing Centroid Shape Configuration (CSC) in conventional PSA. PSC can better tolerate shape change caused by viewing angle change than CSC. Small variation of viewing angle makes large impact only on global gait appearance. Without major impact on local spatio-temporal motion, PSC which effectively embeds local shape information can generate robust view-invariant gait feature. To enhance gait recognition performance, a novel boundary re-sampling process is proposed. It provides only necessary re-sampled points to PSC description. In the meantime, it efficiently solves problems of boundary point correspondence, boundary normalization and boundary smoothness. This re-sampling process adopts prior knowledge of body pose structure. Comprehensive experiment is carried out on the CASIA gait database. The proposed method is shown to significantly improve performance of gait recognition under small viewing angle change without additional requirements of supervised learning, known viewing angle and multi-camera system, when compared with other methods in literatures. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
AVSS | 4 |
| 2011 | Efficient Image Denoising by MRF Approximation with Uniform-Sampled Multi-spanning-treeabstractTraditionally, image processing based on Markov Random Field (MRF) is often addressed on a 4-connected grid graph defined on the image. This structure is not computationally efficient. In our work, we develop a multiple-trees structure to approximate the 4-connected grid. A set of spanning trees are generated by a new algorithm: re-weighted random walk (RWRW). This structure effectively covers the original grid and guarantees uniformly distributed occurrence of each edge. Exact maximum a posterior (MAP) inference is performed on each tree structure by dynamic programming and a median filter is chosen to merge the results together. As an important application, image denoising is used to validate our method. Experimentally, our algorithm provides better performance and higher computational efficiency than traditional methods (such as Loopy Belief Propagation) on a 4-connected MRF. Hongdong Li, Xuming He 0001 |
ICIG | 2 |
| 2011 | Speed-invariant gait recognition based on Procrustes Shape Analysis using higher-order shape configurationabstractWalking speed change is considered a typical challenge hindering reliable human gait recognition. This paper proposes a novel method to extract speed-invariant gait feature based on Procrustes Shape Analysis (PSA). Two major components of PSA, i.e., Procrustes Mean Shape (PMS) and Procrustes Distance (PD), are adopted and adapted specifically for the purpose of speed-invariant gait recognition. One of our major contributions in this work is that, instead of using conventional Centroid Shape Configuration (CSC) which is not suitable to describe individual gait when body shape changes particularly due to change of walking speed, we propose a new descriptor named Higher-order derivative Shape Configuration (HSC) which can generate robust speed-invariant gait feature. From the first order to the higher order, derivative shape configuration contains gait shape information of different levels. Intuitively, the higher order of derivative is able to describe gait with shape change caused by the larger change of walking speed. Encouraging experimental results show that our proposed method is efficient for speed-invariant gait recognition and evidently outperforms other existing methods in the literatures. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
ICIP | 4 |
| 2010 | Gradual Sampling and Mutual Information Maximisation for Markerless Motion Capture
Lei Wang 0001, Richard I. Hartley, Hongdong Li, Dan Xu 0001 |
ACCV (2) | 4 |
| 2010 | Compressive Evaluation in Human Motion Tracking
Lei Wang 0001, Richard I. Hartley, Hongdong Li, Dan Xu 0001 |
ACCV (4) | 4 |
| 2010 | Support vector regression for multi-view gait recognition based on local motion feature selectionabstractGait is a well recognized biometric feature that is used to identify a human at a distance. However, in real environment, appearance changes of individuals due to viewing angle changes cause many difficulties for gait recognition. This paper re-formulates this problem as a regression problem. A novel solution is proposed to create a View Transformation Model (VTM) from the different point of view using Support Vector Regression (SVR). To facilitate the process of regression, a new method is proposed to seek local Region of Interest (ROI) under one viewing angle for predicting the corresponding motion information under another viewing angle. Thus, the well constructed VTM is able to transfer gait information under one viewing angle into another viewing angle. This proposal can achieve view-independent gait recognition. It normalizes gait features under various viewing angles into a common viewing angle before similarity measurement is carried out. The extensive experimental results based on widely adopted benchmark dataset demonstrate that the proposed algorithm can achieve significantly better performance than the existing methods in literature. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
CVPR | 4 |
| 2010 | Multi-view structure computation without explicitly estimating motionabstractMost existing structure-from-motion methods follow a common two-step scheme, where relative camera motions are estimated in the first step and 3D structure is computed afterward in the second step. This paper presents a novel scheme which bypasses the motion-estimation step, and goes directly to structure computation step. By introducing graph rigidity theory to Sfm problems, we demonstrate that such a scheme is not only theoretically possible, but also technically feasible and effective. We also derive a new convex relaxation technique (based on semi-definite programming) which implements the above scheme very efficiently. Our new method provides other benefits as well, such as that it offers a new way to looking at Sfm, and that it is naturally suited for handling sparse large-scale Sfm problems. Hongdong Li |
CVPR | 1 |
| 2010 | Element-Wise Factorization for N-View Projective Reconstruction
Yuchao Dai, Hongdong Li, Mingyi He |
ECCV (4) | 2 |
| 2010 | Multi-view Gait Recognition Based on Motion Regression Using Multilayer PerceptronabstractIt has been shown that gait is an efficient biometric feature for identifying a person at a distance. However, it is a challenging problem to obtain reliable gait feature when viewing angle changes because the body appearance can be different under the various viewing angles. In this paper, the problem above is formulated as a regression problem where a novel View Transformation Model (VTM) is constructed by adopting Multilayer Perceptron (MLP) as regression tool. It smoothly estimates gait feature under an unknown viewing angle based on motion information in a well selected Region of Interest (ROI) under other existing viewing angles. Thus, this proposal can normalize gait features under various viewing angles into a common viewing angle before gait similarity measurement is carried out. Encouraging experimental results have been obtained based on widely adopted benchmark database. Worapan Kusakunniran, Qiang Wu 0001, Jian Zhang 0002, Hongdong Li |
ICPR | 4 |
| 2010 | Interactive color image segmentation with linear programming
Hongdong Li, Chunhua Shen |
Mach. Vis. Appl. | 1 |
| 2010 | Motion Estimation for Nonoverlapping Multicamera Rigs: Linear Algebraic and {\rm L}_\infty Geometric SolutionsabstractWe investigate the problem of estimating the ego-motion of a multicamera rig from two positions of the rig. We describe and compare two new algorithms for finding the 6 degrees of freedom (3 for rotation and 3 for translation) of the motion. One algorithm gives a linear solution and the other is a geometric algorithm that minimizes the maximum measurement error-the optimal L{infinity} solution. They are described in the context of the General Camera Model (GCM), and we pay particular attention to multicamera systems in which the cameras have nonoverlapping or minimally overlapping field of view. Many nonlinear algorithms have been developed to solve the multicamera motion estimation problem. However, no linear solution or guaranteed optimal geometric solution has previously been proposed. We made two contributions: 1) a fast linear algebraic method using the GCM and 2) a guaranteed globally optimal algorithm based on the L{infinity} geometric error using the branch-and-bound technique. In deriving the linear method using the GCM, we give a detailed analysis of degeneracy of camera configurations. In finding the globally optimal solution, we apply a rotation space search technique recently proposed by Hartley and Kahl. Our experiments conducted on both synthetic and real data have shown excellent results. Jae-Hak Kim, Hongdong Li, Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2010 | Estimating Relative Camera Motion from the Antipodal-Epipolar ConstraintabstractThis paper introduces a novel antipodal-epipolar constraint on relative camera motion. By using antipodal points, which are available in large Field-of-View cameras, the translational and rotational motions of a camera are geometrically decoupled, allowing them to be separately estimated as two problems in smaller dimensions. We present a new formulation based on discrete camera motions, which works over a larger range of motions compared to previous differential techniques using antipodal points. The use of our constraints is demonstrated with two robust and practical algorithms, one based on RANSAC and the other based on Hough-like voting. As an application of the motion decoupling property, we also present a new structure-from-motion algorithm that does not require explicitly estimating rotation (it uses only the translation found with our methods). Finally, experiments involving simulations and real image sequences will demonstrate that our algorithms perform accurately and robustly, with some advantages over the state-of-the-art. John Lim, Nick Barnes, Hongdong Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Rotation Averaging with Application to Camera-Rig Calibration
Yuchao Dai, Jochen Trumpf, Hongdong Li, Nick Barnes, Richard I. Hartley |
ACCV (2) | 3 |
| 2009 | Automatic Gait Recognition Using Weighted Binary Pattern on VideoabstractHuman identification by recognizing the spontaneous gait recorded in real-world setting is a tough and not yet fully resolved problem in biometrics research. Several issues have contributed to the difficulties of this task. They include various poses, different clothes, moderate to large changes of normal walking manner due to carrying diverse goods when walking, and the uncertainty of the environments where the people are walking. In order to achieve a better gait recognition, this paper proposes a new method based on Weighted Binary Pattern (WBP). WBP first constructs binary pattern from a sequence of aligned silhouettes. Then, adaptive weighting technique is applied to discriminate significances of the bits in gait signatures. Being compared with most of existing methods in the literatures, this method can better deal with gait frequency, local spatial-temporal human pose features, and global body shape statistics. The proposed method is validated on several well known benchmark databases. The extensive and encouraging experimental results show that the proposed algorithm achieves high accuracy, but with low complexity and computational time. Worapan Kusakunniran, Qiang Wu 0001, Hongdong Li, Jian Zhang 0002 |
AVSS | 3 |
| 2009 | Efficient reduction of L-infinity geometry problemsabstractThis paper presents a new method for computing optimal L∞solutions for vision geometry problems, particularly for those problems of fixed-dimension and of large-scale. Our strategy for solving a large L∞problem is to reduce it to a finite set of smallest possible subproblems. By using the fact that many of the problems in question are pseudoconvex, we prove that such a reduction is possible. To actually solve these small subproblems efficiently, we propose a direct approach which makes no use of any convex optimizer (e.g. SOCP or LP), but is based on a simple local Newton method. We give both theoretic justification and experimental validation to the new method. Potentially, our new method can be made extremely fast. Hongdong Li |
CVPR | 1 |
| 2009 | Consensus set maximization with guaranteed global optimality for robust geometry estimationabstractFinding the largest consensus set is one of the key ideas used by the original RANSAC for removing outliers in robust-estimation. However, because of its random and non-deterministic nature, RANSAC does not fulfill the goal of consensus set maximization exactly and optimally. Based on global optimization, this paper presents a new algorithm that solves the problem exactly. We reformulate the problem as a mixed integer programming (MIP), and solve it via a tailored branch-and-bound method, where the bounds are computed from the MIP's convex under-estimators. By exploiting the special structure of linear robust-estimation, the new algorithm is also made efficient from a computational point of view. Hongdong Li |
ICCV | 1 |
| 2009 | Two Efficient Algorithms for Outlier Removal in Multi-view Geometry Using L-Infinity NormabstractL∞ norm has been recently introduced to multi-view geometry computation to achieve globally optimal computation. It however suffers from a serious sensitivity to outliers. A few remedies have been proposed but with high computational complexity. This paper presents two efficient algorithms to overcome these problems. Our first algorithm is based on a cheap and effective local descent method (as opposed to the conventional but expensive SOCP(Second Order Cone Programming)). The second algorithm further improves the first one by using a Depth-first search heuristics. Both algorithms retain the nice property of global optimality of the L∞ scheme, while at cost only a small fraction of the original computation. Experiments on both synthetic data and real images have validated the proposed algorithms. Yuchao Dai, Mingyi He, Hongdong Li |
ICIG | 3 |
| 2009 | A probabilistic Demons algorithm for texture-rich image registrationabstractDemons algorithm has attracted considerable attention from the image processing community for registering (i.e., matching/aligning) deformable objects or images. It is observed that this algorithm is particularly successful when it applies to nonrigid object having homogenous region, but often fails when the object of interest is rich in texture. This is mainly because the Demons algorithm tends to overfit the many spurious edges inside the texture-rich region, consequently leading to erroneous thermodynamic `forces'. In this paper, we describe a probabilistic Demons algorithm that overcomes this problem. Our key idea is to re-formulate the deformable registration problem in Bayesian statistics framework. The result is a new and more robust Demons algorithm able to capture the essence (e.g., the mass) of a deformable image/object even it is rich in texture. This will significantly expand the applicable scopes of the traditional Demons algorithm. We give encouraging experimental results on real test images. Hongdong Li |
ICIP | 2 |
| 2008 | Motion estimation for multi-camera systems using global optimizationabstractWe present a motion estimation algorithm for multi-camera systems consisting of more than one calibrated camera securely attached on a moving object. So, they move all together, but do not require to have overlapping views across the cameras. The geometrically optimal solution of the motion for the multi-camera systems under Linfinnorm is provided in this paper using a global optimization technique which has been introduced recently in the computer vision research field. Taking advantage of an optimal estimate of the essential matrix through searching rotation space, we provide the optimal solution for translation by using linear programming and branch & bound algorithm. Synthetic and real data experiments are conducted, and they show more robust and improved performance than the previous methods. Jae-Hak Kim, Hongdong Li, Richard I. Hartley |
CVPR | 2 |
| 2008 | A linear approach to motion estimation using generalized camera modelsabstractA well-known theoretical result for motion estimation using the generalized camera model is that 17 corresponding image rays can be used to solve linearly for the motion of a generalized camera. However, this paper shows that for many common configurations of the generalized camera models (e.g., multi-camera rig, catadioptric camera etc.), such a simple 17-point algorithm does not exist, due to some previously overlooked ambiguities. We further discover that, despite the above ambiguities, we are still able to solve the motion estimation problem effectively by a new algorithm proposed in this paper. Our algorithm is essentially linear, easy to implement, and the computational efficiency is very high. Experiments on both real and simulated data show that the new algorithm achieves reasonably high accuracy as well. Hongdong Li, Richard I. Hartley, Jae-Hak Kim |
CVPR | 1 |
| 2008 | Supervised dimensionality reduction via sequential semidefinite programming
Chunhua Shen, Hongdong Li, Michael J. Brooks |
Pattern Recognit. | 2 |
| 2007 | A Convex Programming Approach to the Trace Quotient Problem
Chunhua Shen, Hongdong Li, Michael J. Brooks |
ACCV (2) | 2 |
| 2007 | Where's the Weet-Bix?
Yuhang Zhang 0001, Lei Wang 0001, Richard I. Hartley, Hongdong Li |
ACCV (1) | 4 |
| 2007 | Two-View Motion Segmentation from Linear Programming RelaxationabstractThis paper studies the problem of multibody motion segmentation, which is an important, but challenging problem due to its well-known chicken-and-egg-type recursive character. We propose a new mixture-of-fundamental-matrices model to describe the multibody motions from two views. Based on the maximum likelihood estimation, in conjunction with a random sampling scheme, we show that the problem can be naturally formulated as a linear programming (LP) problem. Consequently, the motion segmentation problem can be solved efficiently by linear program relaxation. Experiments demonstrate that: without assuming the actual number of motions our method produces accurate segmentation result. This LP formulation has also other advantages, such as easy to handle outliers and easy to enforce prior knowledge etc. Hongdong Li |
CVPR | 1 |
| 2007 | A practical algorithm for L triangulation with outliersabstractThis paper addresses the problem of robust optimal multi-view triangulation. We propose an abstract framework, as well as a practical algorithm, which finds the best 3D reconstruction with guaranteed global optimality even in the presence of outliers. Our algorithm is founded on the theory of LP-type problem. We have recognized that the Linfintriangulation is a concrete example of the LP-type problems. We propose a set of non-trivial basis operation subroutines that actually implement the idea. Experiments have validated the effectiveness and efficiency of the proposed algorithm. Hongdong Li |
CVPR | 1 |
| 2007 | The 3D-3D Registration Problem RevisitedabstractWe describe a new framework for globally solving the 3D-3D registration problem with unknown point correspondences. This problem is significant as it is frequently encountered in many applications. Existing methods are not fully satisfactory, mainly due to the risk of local minima. Our framework is grounded on the Lipschitz global optimization theory. It achieves a guaranteed global optimality without any initialization. By exploiting the special structure of the problem itself and of the 3D rotation space SO(3), we propose a box-and-ball algorithm, which solves the problem efficiently. The main idea of the work can be applied to many other problems as well. Hongdong Li, Richard I. Hartley |
ICCV | 1 |
| 2007 | Object-Respecting Color Image SegmentationabstractThe problem of foreground/background segmentation is of great importance in image processing and computer vision. We present a novel Linear-Programming (LP)-based algorithm for color image segmentation. This algorithm segments an image into a conceptually-meaningful foreground region (usually corresponding to the object of interest) and background regions. From a few user specified strokes we learn two Gaussian Mixture models corresponding to the foreground and background region respectively. The algorithm performs well even when the object region consists of several different colors and textures. Due to the global optimality of LP, our algorithm is free from the drawback of getting into local minima. Hongdong Li, Chunhua Shen |
ICIP (2) | 1 |
| 2007 | KLDA - An Iterative Approach to Fisher Discriminant AnalysisabstractIn this paper, we present an iterative approach to Fisher discriminant analysis called Kullback-Leibler discriminant analysis (KLDA) for both linear and nonlinear feature extraction. We pose the conventional problem of discriminative feature extraction into the setting of function optimization and recover the feature transformation matrix via maximization of the objective function. The proposed objective function is defined by pairwise distances between all pairs of classes and the Kullback-Leibler divergence is adopted to measure the disparity between the distributions of each pair of classes. Our proposed algorithm can be naturally extended to handle nonlinear data by exploiting the kernel trick. Experimental results on the real world databases demonstrate the effectiveness of both the linear and kernel versions of our algorithm. Fangfang Lu, Hongdong Li |
ICIP (2) | 2 |
| 2007 | Reconstruction of Underwater Image by BispectrumabstractReconstruction of an underwater object from a sequence of images distorted by moving water waves is a challenging task. A new approach is presented in this paper. We make use of the bispectrum technique to analyze the raw image sequences and recover the phase information of the true object. We test our approach on both simulated and real-world data, separately. Results show that our algorithm is very promising. Such technique has wide applications to areas such as ocean study and submarine observation. Zhiying Wen, Donald Fraser, Andrew J. Lambert, Hongdong Li |
ICIP (3) | 4 |
| 2007 | Conformal spherical representation of 3D genus-zero meshes
Hongdong Li, Richard I. Hartley |
Pattern Recognit. | 1 |
| 2006 | Plane-Based Calibration and Auto-calibration of a Fish-Eye Camera
Hongdong Li, Richard I. Hartley |
ACCV (1) | 1 |
| 2006 | New 3D Fourier Descriptors for Genus-Zero Mesh Objects
Hongdong Li, Richard I. Hartley |
ACCV (1) | 1 |
| 2006 | An LMI Approach for Reliable PTZ Camera Self-CalibrationabstractPTZ (Pan-Tilt-Zoom) cameras are widely used for large-area video surveillance. For many visual tracking and video analysis tasks, an accurate camera calibration is very important. Traditional off-line camera calibration algorithms are often not satisfactory because some of the PTZ camera intrinsic parameters (e.g., focal length) may change during working. Despite of their theoretical elegance, existing on-line camera self-calibration algorithms are not satisfactory either, because of the lack of numerical stability. This paper proposes a new method based on LMI (linear matrix inequality) optimization. This method automatically incorporates the required positive-definiteness constraint into the computation, thus delivers more reliable and more stable results. Experiments on both synthetic data and real images have validated the advantages of our method. Hongdong Li, Chunhua Shen |
AVSS | 1 |
| 2006 | Classification-Based Likelihood Functions for Bayesian TrackingabstractThe success of any Bayesian particle filtering based tracker relies heavily on the ability of the likelihood function to discriminate between the state that fits the image well and those that do not. This paper describes a general framework for learning probabilistic models of objects for exploiting these models for tracking objects in image sequences. We use a discriminative classifier to learn models of how they appear in images. In particular, we use a support vector machine (SVM) for training, which is able to extract useful non-linear information, and thus represent more complex characteristics of the tracked object and background. This is a particular advantage when tracking deformable objects and where appearance changes due to the unstable illumination and pose occur. A by-product of the SVM training procedure is the classification function, with which the tracking problem is cast into a binary classification problem. An object detector directly using the classification function is then available. To make the tracker robust, an object detector that directly uses the classification function is combined into the tracker for object verification. This provides the capability for automatic initialisation and recovery from momentary tracking failures. We demonstrate improved robustness in image sequences. Chunhua Shen, Hongdong Li, Michael J. Brooks |
AVSS | 2 |
| 2006 | A Simple Solution to the Six-Point Two-View Focal-Length Problem
Hongdong Li |
ECCV (4) | 1 |
| 2006 | Inverse tensor transfer with applications to novel view synthesis and multi-baseline stereo
Hongdong Li, Richard I. Hartley |
Signal Process. Image Commun. | 1 |
| 2005 | Inverse tensor transfer for novel view synthesisabstractThis paper provides a new transfer based novel view synthesis method. This method does not need a pre-computed dense depth map, therefore overcomes most common problems associated with conventional dense correspondence algorithms, yet still produce very photo-realistic novel images. The power of the method comes from the introducing and using of a novel inverse tensor transfer technique, which offers a simple mechanism to exploit both photometric constraints and geometric constraints across multiple input images. Our method works equally well for both calibrated images and un-calibrated images. Experiments on real sequences show promising results. Hongdong Li, Richard I. Hartley |
ICIP (2) | 1 |
| 2004 | A new and compact algorithm for simultaneously matching and estimationabstractFeature matching and transformation estimation are two fundamental problems in computer vision research. These two problems are often related and even interlocked; solving one is solving the other's precondition. This makes them hard to solve. In order to overcome this difficulty, the paper presents a new compact algorithm requiring less than 10 lines of Matlab code. We show that the solutions of correspondence and transformation are merely two factors of two grammian matrices, and can be worked out by a factorization method. A Newton-Schulz numerical iteration algorithm is used for the factorization. The two interlocked problems are solved in an alternate (flip-flop) way. The effectiveness and efficiency are illustrated by experiments on both synthetic and real images. Global and fast convergence is attained, even starting from randomly chosen initial guesses. Hongdong Li, Richard I. Hartley |
ICASSP (3) | 1 |
| 2002 | A fast and robust face location and feature extraction systemabstractAutomatic face location and feature detection in front view images has a wide range of usage. To fulfill both a fast and robust algorithm is still a challenge. We have proposed such a solution. In order to detect faces fast and accurately in various images, we have adopted a Gabor-like filtering scheme to locate all the possible features, followed by a verification procedure to check the "faceness" of all possible feature-blob combinations. Then the different features are segmented using integral projection. To detect the contours of different features precisely, we propose a heuristic knowledge-based contour tracking algorithm, using an adaptive oriented edge filtering and tracking scheme. We also have proposed a novel blob detector, which can detect blob-shaped dark image patterns (e.g. the iris, the nostril) efficiently. Our algorithm is fast, robust and accurate, which is proved by experiments on a pretty large database. Even in case of strange lighting, low-resolution, or strong distraction from other facial structures (such as facial hair), our algorithm can also work without serious deterioration in performance. Tianxiang Yao, Hongdong Li, Guangyao Liu, XiuQing Ye, WeiKang Gu, Yiqing Jin |
ICIP (1) | 2 |
| 2000 | A new and fast approach for DPIV using an incompressible affine flow model
Hongdong Li, Jilin Liu, WeiKang Gu |
Mach. Vis. Appl. | 1 |