Yihong Wu 0002

dblp:24/2219-2 · DBLP profile ↗
← Back
80ranked-venue papers
13as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 3 first-author · 17 since 2021Artificial intelligence and machine learning · 42 · 12 first-author · 19 since 2021Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 3 since 2021Theory of computation · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Sparse3DPR: Training-Free 3D Hierarchical Scene Parsing and Task-Adaptive Subgraph Reasoning from Sparse RGB Views
abstract
Recently, large language models (LLMs) have been explored widely for 3D scene understanding. Among them, training-free approaches are gaining attention for their flexibility and generalization over training-based methods. However, they typically struggle with accuracy and efficiency in practical deployment. To address the problems, we propose Sparse3DPR, a novel training-free framework for open-ended scene understanding, which leverages the reasoning capabilities of pre-trained LLMs and requires only sparse-view RGB inputs. Specifically, we introduce a hierarchical plane-enhanced scene graph that supports open vocabulary and adopts dominant planar structures as spatial anchors, which enables clearer reasoning chains and more reliable high-level inferences. Furthermore, we design a task-adaptive subgraph extraction method to filter query-irrelevant information dynamically, reducing contextual noise and improving 3D scene reasoning efficiency and accuracy. Experimental results demonstrate the superiority of Sparse3DPR, which achieves a 28.7% EM@1 improvement and a 78.2% speedup compared with ConceptGraphs on the Space3D-Bench. Moreover, Sparse3DPR obtains comparable performance to training-based methods on ScanQA, with additional real-world experiments confirming its robustness and generalization capability.
Haida Feng, Hao Wei 0008, Zewen Xu, Haolin Wang 0005, Chade Li, Yihong Wu 0002
AAAI6
2026 MR-COSMO: Visual-Text Memory Recall and Direct CrOSs-MOdal Alignment Method for Query-Driven 3D Segmentation
abstract
The rapid advancement of vision-language models (VLMs) in 3D domains has accelerated research in text-query-guided point cloud processing, though existing methods underperform in point-level segmentation due to inadequate 3D-text alignment that limits local feature-text context linking. To address this limitation, we propose MR-COSMO, a Visual-Text Memory Recall and Direct CrOSs-MOdal Alignment Method for Query-Driven 3D Segmentation, establishing explicit alignment between 3D point clouds and text/2D image data through a dedicated direct cross-modal alignment module while implementing a visual-text memory module with specialized feature banks. This direct alignment mechanism enables precise fusion of geometric and semantic features, while the memory module employs specialized banks storing text features, visual features, and their correspondence mappings to dynamically enhance scene-specific representations via attention-based knowledge recall. Comprehensive experiments across 3D instruction, reference, and semantic segmentation benchmarks confirm state-of-the-art performance.
Chade Li, Yihong Wu 0002
AAAI3
2026 SLNMapping: Super Lightweight Neural Mapping in Large-Scale Scenes
Chenhui Shi 0001, Fulin Tang, Hao Wei 0008, Yihong Wu 0002
Int. J. Comput. Vis.4
2026 Density-aware global-local attention network for point cloud segmentation
Chade Li, Yihong Wu 0002
Image Vis. Comput.4
2026 Temporal-camera-based semantic scene completion network with query-mask constraint and 3D layout initialization
Chade Li, Yihong Wu 0002
Pattern Recognit.4
2026 CGFMamba-PCR: Color-Geometric Fusion Mamba-Based Color Point Cloud Registration
abstract
Recently, color point cloud registration has begun to receive attention. Unlike geometric-only point clouds, color point clouds incorporate additional color information. Therefore, color point cloud registration can achieve higher accuracy than geometric-only point cloud registration. Despite the success, existing methods are computationally intensive due to the high resource demands of Transformer. In this paper, we propose a Mamba architecture based registration algorithm, CGFMamba-PCR, with color-geometric fusion. Specifically, we propose CGFMamba, a novel color-geometric fusion Mamba network for color point cloud registration, which enhances feature representation through color-geometric guided point ordering and positional encoding. For the input color point cloud pair, they are passed through CGFMamba based feature extraction module to obtain their corresponding features. Then, these features are passed through feature matching and outlier rejection modules to obtain final registration result. Furthermore, an ordering method for the Mamba architecture is proposed that clusters color hue and sorts spatial coordinates. Experiments on Color3DMatch and Color3DLoMatch datasets demonstrate that the proposed algorithm outperforms the state-of-the-art (SOTA) methods. The code of the proposed algorithm will be open-sourced upon acceptance of this paper.
Shiyi Guo, Tong Jia 0001, Bi Yang, Yihong Wu 0002, Hao Wei 0008, Nannan Liu, Ning An 0002
IEEE Trans. Circuits Syst. Video Technol.4
2026 Accurate Absolute Scale With Pseudo Depth Constraint for 3-D Sparse Construction of Monocular RGB Image
abstract
3D Sparse reconstruction and camera pose estimation from monocular sequences are inherently plagued by scale ambiguity, which prevents the recovery of the scene’s absolute scale. Existing solutions are primarily limited in two aspects: (1) traditional geometric methods often depend on additional sensors or restrictive pre-calibration, lacking general applicability; (2) learning-based approaches frequently suffer from scale biases when faced with domain shifts. To overcome these challenges, this paper introduces a unified framework that deeply integrates monocular depth estimation into a traditional reconstruction pipeline. Our method leverages the robustness of modern depth prediction networks to eliminate cumbersome prior assumptions while employing iterative geometric refinement to ensure precise scale accuracy. The proposed framework begins by estimating depth maps via RGB images. First, an absolute scale factor is initialized from the depth priors to calibrate the initial two-view reconstruction, effectively aligning the model into a metric space. Second, during incremental reconstruction, new points and poses are registered via depth-constrained triangulation, with an iterative weighting strategy to enhance stability. Third, we propose a joint bundle adjustment that harmonizes reprojection and geometric errors using an adaptive Mahalanobis distance, which implicitly optimizes the global scale and refines the scene geometry concurrently. Comprehensive evaluations on the ETH3D, TUM-RGBD, and ICL-NUIM benchmarks demonstrate that our framework achieves state-of-the-art performance in absolute scale recovery. The results confirm that our method successfully maintains high geometric fidelity across diverse scenarios, proving its effectiveness and robustness.
Erjie Jiao, Hengyou Wang, Jun Wan 0001, Yihong Wu 0002
IEEE Trans. Circuits Syst. Video Technol.5
2026 A Rotation-Translation Decoupled Solution for Visual-Inertial Initialization and Online Spatial-Temporal Calibration
abstract
We propose a novel initialization and online spatial-temporal calibration method for visual-inertial odometry (VIO), which decouples rotation and translation estimation to achieve higher accuracy and better robustness. Existing initialization methods suffer from limited accuracy or robustness (e.g., in scenarios with small translational motion) and rarely integrate simultaneous spatial-temporal calibration during initialization, despite its considerable practical value. Our proposed method leverages rotation-translation decoupling constraints to enable simultaneous estimation of gyroscope bias, extrinsic rotation, and camera-IMU time offset-even under pure rotational motion. Moreover, we are the first to conduct observability analysis on rotational constraints in rotation-translation decoupling methods, experimentally identifying the unobservable state-space directions under three degenerate motions within our approach. We also perform extensive experiments to delineate practical parameter solution boundaries for our method, with both efforts substantially enhancing the overall practical applicability of decoupling-based methods. Extensive experiments on simulated and real-world datasets demonstrate that our method outperforms state-of-the-art approaches in accuracy and robustness while maintaining computational efficiency. Furthermore, experiments verify that it significantly improves convergence in VIO systems.
Bo Xu 0022, Zewen Xu, Yijia He, Zhanpeng Ouyang, Hao Wei 0008, Yihong Wu 0002, Jiancheng Li, Hongdong Li
IEEE Trans. Robotics6
2026 Review of extrinsic parameter calibration of LiDAR and camera
abstract
LiDAR and camera are two of the most common sensors used in the fields of robot perception, autonomous driving, augmented reality, and virtual reality, where these sensors are widely used to perform various tasks such as odometry estimation and 3D reconstruction. Fusing the information from these two sensors can significantly increase the robustness and accuracy of these perception tasks. The extrinsic calibration between cameras and LiDAR is a fundamental prerequisite for multimodal systems. Recently, extensive studies have been conducted on the calibration of extrinsic parameters. Although several calibration methods facilitate sensor fusion, a comprehensive summary for researchers and, especially, non-expert users is lacking. Thus, we present an overview of extrinsic calibration and discuss diverse calibration methods from the perspective of calibration system design. Based on the calibration information sources, this study classifies these methods as target-based or targetless. For each type of calibration method, further classification was performed according to the diverse types of features or constraints used in the calibration process, and their detailed implementations and key characteristics were introduced. Thereafter, calibration-accuracy evaluation methods are presented. Finally, we comprehensively compare the advantages and disadvantages of each calibration method and suggest directions for practical applications and future research.
Shuo Wang 0018, Yihong Wu 0002
Virtual Real. Intell. Hardw.3
2025 FEAST-Mamba: FEAture and SpaTial Aware Mamba Network with Bidirectional Orthogonal Fusion for Cross-Modal Point Cloud Segmentation
abstract
Point cloud segmentation has a wide range of applications in autonomous driving, augmented reality and virtual reality. Multi-modal fusion strategies have received increasing attention in point cloud segmentation recently. Despite the success, existing methods usually generate unnecessary information loss or redundancy. In this paper, we propose FEAST-Mamba, a novel FEAture and SpaTial aware Mamba network to tackle multi-modal point cloud segmentation. To exploit the complementarity between different modals, we propose a bidirectional orthogonal attention module, where features are first bidirectionally interacted with each other through cross-modal attention, and then orthogonal fusion is used to reduce feature redundancy. Furthermore, a reordering strategy is proposed for the Mamba architecture that takes into account both spatial and semantic information during cross-modal feature ordering. Experiments on indoor datasets, S3DIS and ScanNet, and outdoor datasets, nuScenes and SemanticKITTI, show that the proposed method achieves state-of-the-art performances.
Chade Li, Hao Wei 0008, Yihong Wu 0002
AAAI5
2025 3D-SLNR: A Super Lightweight Neural Representation for Large-scale 3D Mapping
abstract
We propose 3D-SLNR, a new and ultra-lightweight neural representation with outstanding performance for large-scale 3D mapping. The representation defines a global signed distance function (SDF) in near-surface space based on a set of band-limited local SDFs anchored at support points sampled from point clouds. These SDFs are parameterized only by a tiny multi-layer perceptron (MLP) with no latent features, and the state of each SDF is modulated by three learnable geometric properties: position, rotation, and scaling, which make the representation adapt to complex geometries. Then, we develop a novel parallel algorithm tailored for this unordered representation to efficiently detect local SDFs where each sampled point is located, allowing for real-time updates of local SDF states during training. Additionally, a prune-and-expand strategy is introduced to enhance adaptability further. The synergy of our low-parameter model and its adaptive capabilities results in an extremely compact representation with excellent expressiveness. Extensive experiments demonstrate that our method achieves state-of-the-art reconstruction performance with less than 1/5 of the memory footprint compared with previous advanced methods.
Chenhui Shi 0001, Fulin Tang, Ning An 0002, Yihong Wu 0002
CVPR4
2025 Hi-Gaussian: Hierarchical Gaussians Under Normalized Spherical Projection for Single-View 3D Reconstruction
Binjian Xie, Hao Wei 0008, Yihong Wu 0002
ICCV4
2025 DOGE: An Extrinsic Orientation and Gyroscope Bias Estimation for Visual-Inertial Odometry Initialization
abstract
Most existing visual-inertial odometry (VIO) initialization methods rely on accurate pre-calibrated extrinsic parameters. However, during long-term use, irreversible structural deformation caused by temperature changes, mechanical squeezing, etc. will cause changes in extrinsic parameters, especially in the rotational part. Existing initialization methods that simultaneously estimate extrinsic parameters suffer from poor robustness, low precision, and long initialization latency due to the need for sufficient translational motion. To address these problems, we propose a novel VIO initialization method, which jointly considers extrinsic orientation and gyroscope bias within the normal epipolar constraints, achieving higher precision and better robustness without delayed rotational calibration. First, a rotation-only constraint is designed for extrinsic orientation and gyroscope bias estimation, which tightly couples gyroscope measurements and visual observations and can be solved in pure-rotation cases. Second, we propose a weighting strategy together with a failure detection strategy to enhance the precision and robustness of the estimator. Finally, we leverage Maximum A Posteriori to refine the results before enough translation parallax comes. Extensive experiments have demonstrated that our method outperforms the state-of-the-art methods in both accuracy and robustness while maintaining competitive efficiency.
Zewen Xu, Yijia He, Hao Wei 0008, Yihong Wu 0002
ICRA4
2025 Floorplan-SLAM: A Real-Time, High-Accuracy, and Long-Term Multi-Session Point-Plane SLAM for Efficient Floorplan Reconstruction
abstract
Floorplan reconstruction provides structural priors essential for reliable indoor robot navigation and high-level scene understanding. However, existing approaches either require time-consuming offline processing with a complete map, or rely on expensive sensors and substantial computational resources. To address the problems, we propose FloorplanSLAM, which incorporates floorplan reconstruction tightly into a multi-session SLAM system by seamlessly interacting with plane extraction, pose estimation, back-end optimization, and loop & map merging, achieving real-time, high-accuracy, and long-term floorplan reconstruction using only a stereo camera. Specifically, we present a robust plane extraction algorithm that operates in a compact plane parameter space and leverages spatially complementary features to accurately detect planar structures, even in weakly textured scenes. Furthermore, we propose a floorplan reconstruction module tightly coupled with the SLAM system, which uses continuously optimized plane landmarks and poses to formulate and solve a novel optimization problem, thereby enabling real-time and high-accuracy floorplan reconstruction. Note that by leveraging the map merging capability of multi-session SLAM, our method supports long-term floorplan reconstruction across multiple sessions without redundant data collection. Experiments on the VECtor and the self-collected datasets indicate that Floorplan-SLAM significantly outperforms state-of-the-art methods in terms of plane extraction robustness, pose estimation accuracy, and floorplan reconstruction fidelity and speed, achieving real-time performance at 25–45 FPS without GPU acceleration, which reduces the floorplan reconstruction time for a 1000 m2scene from 16 hours and 44 minutes to just 9.4 minutes.
Haolin Wang 0005, Zeren Lv, Hao Wei 0008, Haijiang Zhu, Yihong Wu 0002
IROS5
2025 Maximum Clique-Based Floorplan Association for Robust Multi-Session Stereo SLAM in Challenging Indoor Environments
abstract
Existing multi-session visual simultaneous localization and mapping (SLAM) systems struggle severely to achieve robust localization and map merging under extreme viewpoint and illumination variations, particularly when handling completely opposite viewpoints and drastic day-night lighting changes. These challenges stem largely from the limited viewpoint/illumination invariance of conventional low-level visual features and their inability to capture a global structural context. In this paper, we make the critical observation that a life-long floorplan not only encodes rich geometric and semantic information—serving as a robust high-level structural representation—but is also inherently more robust to severe viewpoint and illumination variations than purely visual data. Building on this insight, we propose a novel hierarchical framework for multi-session SLAM that integrates a floorplan-based map as a global feature to achieve robust indoor localization and map merging under drastic viewpoint and illumination shifts. In particular, we innovatively formulate floorplan association as a maximum clique problem augmented with trajectory data to achieve robust floorplan-level global localization. We further introduce a novel coarse-to-fine localization and map merging strategy that seamlessly integrates floorplan alignment, multistage point cloud registration, and feature matching, fully leveraging the macro-level stability of global features and the micro-level precision of local features to achieve keyframe-level fine localization. Extensive experiments on both public and self-collected datasets demonstrate that our method consistently outperforms state-of-the-art (SOTA) approaches reliant solely on low-level visual or geometric features. Crucially, it delivers superior accuracy and robustness even in the face of completely opposite viewpoints and extreme day–night illumination changes. This work underscores the promise of fusing macro-level floorplan representations with conventional SLAM frameworks to advance long-term, robust indoor localization and map merging under the most challenging conditions.
Haolin Wang 0005, Hao Wei 0008, Zeren Lv, Haijiang Zhu, Yihong Wu 0002
IROS5
2025 DGOcc: Depth-aware global query-based network for monocular 3D occupancy prediction
Yihong Wu 0002
Neurocomputing4
2025 Low-Overlap Point Cloud Registration by Semiglobal Block Matching
abstract
In recent years, the local feature-based point cloud registration methods have attracted more attention for their high robustness. The state-of-the-art local feature-based pipeline consists of local feature extraction, feature matching, and outlier rejection. However, the accuracy of this pipeline decreases significantly between low-overlap point clouds. In this article, we propose a novel hand-crafted framework to realize efficient low-overlap point cloud registration through point cloud partitioning, semiglobal point cloud block matching, and best transformation selection with refinement. Experimental results on benchmark datasets show that the proposed algorithm presents the superior performance over the previous methods on low-overlap scenes under different features and achieves very competitive accuracy on non-low-overlap data.
Shiyi Guo, Yihong Wu 0002, Binjian Xie, Bingxi Liu 0001, Tong Jia 0001
IEEE Trans. Ind. Informatics2
2025 CornerVINS: Accurate Localization and Layout Mapping for Structural Environments Leveraging Hierarchical Geometric Representations
abstract
A compact and consistent map of surroundings is critical for intelligent robots to understand their situations and realize robust navigation. Most existing techniques rely on infinite planes, which are sensitive to pose drift and may lead to confusing maps. Towards high-level perception in indoor environments, we propose CornerVINS, an innovative RGB-D inertial localization and layout mapping method leveraging hierarchical geometric features, i.e., points, planes, and box corners. Specifically, points are enhanced by fusing depth information, and planes are modeled as bounded patches using convex hulls to increase their discriminability. More importantly, box corners, lying at the intersection of three orthogonal planes, are parameterized with a 6-dimensional vector and integrated into the extended Kalman filter for the first time. We introduce a hierarchical mechanism to effectively extract and associate planes and corners, which are considered as layout components of scenes and serve as long-term landmarks to correct camera poses. Extensive experiments prove that the proposed box corners bring significant improvements, enabling accurate localization and consistent layout mapping at low computational cost. Overall, the proposed CornerVINS outperforms state-of-the-art systems in both accuracy and efficiency. The core implementations are available athttps://github.com/zydddd/CornerVINS.
Fulin Tang, Yihong Wu 0002
IEEE Trans. Robotics3
2024 Novel camera self-calibration method with clustering prior and nonlinear optimization from an image sequence
Xiaohui Jiang, Haijiang Zhu, Ning An 0002, Binjian Xie, Hao Wei 0008, Fulin Tang, Yihong Wu 0002
Multim. Tools Appl.7
2024 Covariant Peak Constraint for Accurate Keypoint Detection and Keypoint-Specific Descriptor Learning
abstract
Local feature extraction consists of keypoint detection and local descriptor extraction. Firstly, in keypoint detector learning, existing covariance constraint loss functions cannot constrain the probability distribution shapes in local probability maps that surround keypoints. And existing auxiliary peak loss functions, which are used to alleviate the problem, impair the performance of local feature methods. To solve this problem, we propose a novel Covariant Peak constraint Loss (CP Loss) which is defined as the expectations of local probability maps' position errors. Minimizing our CP Loss can make local probability maps accurately peak at reliable keypoints. Secondly, in descriptor learning, the Neural Reprojection Error (NRE) aims at constraining dense descriptor maps of images. But we argue that only those descriptors of keypoints need to be constrained. Thus, we propose a novel Conditional Neural Reprojection Error (CNRE) that is only conditioned on keypoints. Compared with the NRE, our CNRE can achieve much higher efficiency and produce more keypoint-specific descriptors with better matching performance. We use our CP Loss and CNRE to train a local feature network named as CPCN-Feat. Experimental results show that our CPCN-Feat achieves state-of-the-art performance on four challenging benchmarks.
Yujie Fu, Fulin Tang, Yihong Wu 0002
IEEE Trans. Multim.4
2023 PLPL-VIO: A Novel Probabilistic Line Measurement Model for Point-Line-Based Visual-Inertial Odometry
abstract
Point and line features are complementary in Visual-Inertial Odometry (VIO) or Visual-Inertial Simultaneous Localization And Mapping (VI-SLAM) systems. The advantage of combining these two types of features relies on their proper weighting in the cost function, usually set by their uncertainty. Compared with point features, setting line segment endpoints' uncertainty with isotropic distribution is unreasonable. But the uncertainty of line feature observation, especially for the endpoints' uncertainty along the line, is difficult to set due to occlusion and fragmentation problems. In this article, we use infinite lines as the line feature observations and prove that the uncertainty of these observations is only related to the vertical uncertainty of the endpoints, thus avoiding setting the parallel uncertainty of the endpoints. Besides, we introduce a novel consistent measurement model for line features. Furthermore, for long-time constraints, we add 3D line segments into the state vector and derive how to update them properly. Finally, we construct a point-line-based VIO system that takes into account the uncertainty of line feature observations and the consistency of line feature measurements. The proposed VIO system is validated on two public datasets. The results show that the proposed method obtains the best accuracy compared with the state-of-the-art point-based VIO systems (OpenVINS, VINS-Mono), a point-line-based VIO system (PL-VINS), and a structural line-based system (StructVIO).
Zewen Xu, Hao Wei 0008, Fulin Tang, Yihong Wu 0002, Gang Ma 0007, Shuzhe Wu, Xin Jin 0004
IROS5
2023 Efficient 6-DoF camera pose tracking with circular edges
Fulin Tang, Shaohuan Wu, Zhengda Qian, Yihong Wu 0002
Comput. Vis. Image Underst.4
2023 Learning to Reduce Scale Differences for Large-Scale Invariant Image Matching
abstract
Most image matching methods perform poorly when encountering large scale changes in images. To solve this problem, we propose a Scale-Difference-Aware Image Matching method (SDAIM) that reduces image scale differences before local feature extraction, via resizing both images of an image pair according to an estimated scale ratio. In order to accurately estimate the scale ratio for the proposed SDAIM, we propose a Covisibility-Attention-Reinforced Matching module (CVARM) and then design a novel neural network, termed as Scale-Net, based on CVARM. The proposed CVARM can lay more stress on covisible areas within the image pair and suppress the distraction from those areas visible in only one image. Quantitative and qualitative experiments confirm that the proposed Scale-Net has higher scale ratio estimation accuracy and much better generalization ability compared with all the existing scale ratio estimation methods. Further experiments on image matching and relative pose estimation tasks demonstrate that our SDAIM and Scale-Net are able to greatly boost the performance of representative local features and state-of-the-art local feature matching methods.
Yujie Fu, Bingxi Liu 0001, Zheng Rong, Yihong Wu 0002
IEEE Trans. Circuits Syst. Video Technol.5
2022 SRK-Net: Learning to Detect Repeatable Keypoints with Local Saliency Knowledge
abstract
The dominant approach for learning keypoint detectors relies on the covariance constraint. However, existing learned detectors sometimes extract unstable keypoints from edges. To solve this problem, we propose a novel method that exploits local saliency knowledge to train a keypoint detector, and obtain a keypoint detector, called as SRK-Net, which can extract stable and repeatable keypoints. Firstly, given an image, we propose a General Local Saliency Measure method (GLSM) to assess the local saliency value for each pixel and generate a local saliency map for this image. Then we propose a Local Salient Structure Maintaining loss (LSSM) and a two-stage progressive training manner tailored for leveraging the supervision of the covariance constraint and the local saliency maps provided by our GLSM. Experimental results show that the proposed SRK-Net performs better than all the existing keypoint detectors on HPatches dataset.
Yujie Fu, Zheng Rong, Yihong Wu 0002
ICIP3
2022 Cross-Attention-Based Feature Extraction Network for 3D Point Cloud Registration
abstract
In recent years, the feature-based point cloud registration methods have attracted more attention. However, most existing methods focus on extracting features with strong antiinterference ability from a single point cloud while neglecting the differences within point cloud pairs. In this paper, unlike these methods treating each point cloud independently, we instead consider the information between point cloud pairs when extracting features. Specifically, we propose a cross-attention-based network for modeling the correlation between a pair of point clouds, where a 3D cross-attention mechanism is proposed and combined with 3D convolution elegantly for feature extraction. The extracted features achieve better robustness under various conditions, such as rotation and translation changes. Then accurate point cloud registration is achieved by matching these features. Experimental results on 3DMatch dataset show that the proposed method achieves state-of-the-art performance on feature matching and point cloud registration tasks compared with the previous feature-based methods.
Shiyi Guo, Yujie Fu, Zhengda Qian, Zheng Rong, Yihong Wu 0002
ICME5
2022 Automatic Detection and Fitting of Ellipse Markers Using EllipseNet
abstract
Ellipses are important elements in projective geometry. Accurate extraction of ellipse information is the first step in many computer vision applications, such as ellipse-based camera calibration and camera pose estimation. At present, most ellipse detection algorithms rely on the edge features extracted by Canny, which leads to a number of wrong detection results since non-ellipse edges features are also involved. To address this problem, this paper proposes a novel ellipse marker detection neural network, called EllipseNet. Notably, we propose a new loss function to enhance the rotation perception ability of EllipseNet. Furthermore, a novel ellipse marker data enhancement method is proposed for saving the time cost of labelling ellipse parameters. Experiments show that EllipseNet can improve the detection precision of ellipse regions by more than 3% improvement compared with other SOTA general object detectors.
Zhengda Qian, Fulin Tang, Bingxi Liu 0001, Yujie Fu, Shaohuan Wu, Xiaohong Jia 0001, Yihong Wu 0002
ICPR7
2022 High-quality indoor scene 3D reconstruction with RGB-D cameras: A brief review
abstract
High-quality 3D reconstruction is an important topic in computer graphics and computer vision with many applications, such as robotics and augmented reality. The advent of consumer RGB-D cameras has made a profound advance in indoor scene reconstruction. For the past few years, researchers have spent significant effort to develop algorithms to capture 3D models with RGB-D cameras. As depth images produced by consumer RGB-D cameras are noisy and incomplete when surfaces are shiny, bright, transparent, or far from the camera, obtaining high-quality 3D scene models is still a challenge for existing systems. We here review high-quality 3D indoor scene reconstruction methods using consumer RGB-D cameras. In this paper, we make comparisons and analyses from the following aspects: (i) depth processing methods in 3D reconstruction are reviewed in terms of enhancement and completion, (ii) ICP-based, feature-based, and hybrid methods of camera pose estimation methods are reviewed, and (iii) surface reconstruction methods are reviewed in terms of surface fusion, optimization, and completion. The performance of state-of-the-art methods is also compared and analyzed. This survey will be useful for researchers who want to follow best practices in designing new high-quality 3D reconstruction methods.
Wei Gao 0014, Yihong Wu 0002, Yangdong Liu, Yanfei Shen
Comput. Vis. Media3
2022 Semi-global shape-aware attention network for image segmentation and retrieval
Jiagang Zhu, Chaofan Zhang, Zheng Rong, Yihong Wu 0002
Neurocomputing5
2022 Leveraging local and global descriptors in parallel to search correspondences for visual localization
Chaofan Zhang, Bingxi Liu 0001, Yihong Wu 0002
Pattern Recognit.4
2021 A Flexible and Efficient Loop Closure Detection Based on Motion Knowledge
abstract
Loop closure detection (LCD) is an essential module for simultaneous localization and mapping (SLAM), which can correct accumulated errors after long-term explorations. The widely used bag-of-words (BoW) model can not satisfy well the requirements of both low time consumption and high accuracy for a mobile platform. In this paper, we propose a novel LCD algorithm based on motion knowledge. We give a flexible and efficient detection strategy and also give flexible and efficient combinations of a global binary feature extracted by convolutional neural network (CNN) and a hand-crafted local binary feature. We take a continuous motion model, grid-based motion statistics (GMS) and motion states as motion knowledge. Furthermore, we fuse the proposed LCD with a visual-inertial odometry (VIO) system to correct localization errors by a pose graph optimization. Comparative experiments with state-of-the-art LCD algorithms on typical datasets have been carried out, and the results demonstrate that our proposed method achieves quite high recall rates and quite high speed at 100% precision. Moreover, experimental results from VIO further validate the effectiveness of the proposed method.
Bingxi Liu 0001, Fulin Tang, Yujie Fu, Yanqun Yang, Yihong Wu 0002
ICRA5
2021 Highly Efficient Line Segment Tracking with an IMU-KLT Prediction and a Convex Geometric Distance Minimization
abstract
Line segment features become popular in SLAM community. Usually, line-based SLAM systems utilize local appearance descriptors for line segment tracking. However, traditional descriptor-based line segment tracking algorithms suffer from the problem that accuracy and speed cannot be possessed simultaneously, which affects the performance of line-based SLAM systems negatively. We propose a novel line segment tracking method with an IMU-KLT line segment prediction and a convex geometric distance minimization to boost line segment tracking performance in both accuracy and speed. Particularly, the proposed convex geometric distance minimization uses a ℓ1-norm model to minimize geometric constraints between predicted line segments and extracted line segments efficiently. Furthermore, the line segment tracking is embedded into a VIO system and we adapt it to obtain more reliable point tracking. Experimental results on public datasets show that the proposed line segment tracking method achieves much higher accuracy and much less time cost than state-of-the-art level, where not only the number of correct matches increases but also the inlier ratios are increased by at least 35.1% along with a 3 times faster speed. Besides, the VIO system combining the proposed line segment tracking is improved in terms of accuracy.
Hao Wei 0008, Fulin Tang, Chaofan Zhang, Yihong Wu 0002
ICRA4
2021 Learning to Decompose and Restore Low-light Images with Wavelet Transform
abstract
Low-light images often suffer from low visibility and various noise. Most existing low-light image enhancement methods often amplify noise when enhancing low-light images, due to the neglect of separating valuable image information and noise. In this paper, we propose a novel wavelet-based attention network, where wavelet transform is integrated into attention learning for joint low-light enhancement and noise suppression. Particularly, the proposed wavelet-based attention network includes a Decomposition-Net, an Enhancement-Net and a Restoration-Net. In Decomposition-Net, to benefit denoising, wavelet transform layers are designed for separating noise and global content information into different frequency features. Furthermore, an attention-based strategy is introduced to progressively select suitable frequency features for accurately restoring illumination and reflectance according to Retinex theory. In addition, Enhancement-Net is introduced for further removing degradations in reflectance and adjusting illumination, while Restoration-Net employs conditional adversarial learning to adversarially improve the visual quality of final restored results based on enhanced illumination and reflectance. Extensive experiments on several public datasets demonstrate that the proposed method achieves more pleasing results than state-of-the-art methods.
Chaofan Zhang, Zheng Rong, Yihong Wu 0002
MMAsia4
2021 High Power-Efficient and Performance-Density FPGA Accelerator for CNN-Based Object Detection
Chaofan Zhang, Fulin Tang, Yihong Wu 0002, Xuezhi Yang
PRCV (1)5
2020 Scene-Unified Image Translation For Visual Localization
abstract
Visual localization is a key technology in the field of 3D robot vision. One of its major difficulties lies in how to deal with the challenges brought by the appearance changes of query images and database images caused by large time spans. Many methods focus on extracting more robust features from images to deal with the impact of complex scenes. In this paper, we explore the impact of image translation on visual localization tasks in complex scenes. We propose UniGAN - a modified image translation model, fusing semantic label constraints and finer reconstruction losses, to unify images captured under different environmental conditions to a standard scene more suitable for localization tasks. To estimate the 6-DOF camera pose, a two-stage localization framework composed of image retrieval and local matching is utilized. Experiments show that our method outperforms the state-of-the-art in terms of both accuracy and robustness to environmentally sensitive scenes.
Wei Gao 0014, Yiming Wan, Yihong Wu 0002
ICIP4
2020 Boosting Image-Based Localization Via Randomly Geometric Data Augmentation
abstract
Visual localization is a fundamental problem in computer vision and robotics. Recently, deep learning has shown to be effective for robust monocular localization. Most deep learning-based methods utilize convolution neural network (CNN) to regress global 6 degree-of-freedom (Dof) pose. However, these methods suffer from pose sparsity, leading to over-fitting during training and poor localization performance on unseen data. In this paper, we try to alleviate this issue by implementing randomly geometric augmentation (RGA) during training. Specifically, we firstly estimate the depth map using a depth estimation network for the initial training image. Combing the estimated depth, RGB image and its corresponding pose, we can randomly synthesize new images of different views. The synthesized and initial images are used to train the pose regression network. Experiment results show our geometric augmentation strategy can significantly improve the localization accuracy.
Yiming Wan, Wei Gao 0014, Yihong Wu 0002
ICIP4
2020 Dynamic Object-Aware Monocular Visual Odometry With Local And Global Information Aggregation
abstract
In this paper, we present a deep learning-based approach to monocular visual odometry. We propose a LCGR(Local Convolution and Global RNN) module which utilizes several independent 3D convolution layers to filter noise from features extracted by FlowNet, as well as to model local information, and a Bi-ConvLSTM layer to model time series and capture global information. In addition, our network jointly predicts optical flow as an auxiliary task by measuring photometric consistency in a self-supervised way to help the encoder for better motion feature extraction. In order to alleviate the effects of non-Lambertian surfaces and dynamical objects in the scene, a confidence mask layer is estimated and epipolar constraint is added to the training process. Experiment results indicate competitive performance of the proposed framework to the state-of-art methods.
Yiming Wan, Wei Gao 0014, Yihong Wu 0002
ICIP4
2020 3D Mapping and 6D Pose Computation for Real Time Augmented Reality on Cylindrical Objects
abstract
Visual Augmented Reality (AR) typically overlays virtual computer graphics or other virtual contents on the real world videos, attracting much interest from both academic and industrial communities. Although AR techniques on planes are well studied, cylindrical objects are seldom used for augmented reality. In this paper, we propose a new method for 3D reconstruction and 6D pose computation for augmented reality on a cylindrical object. The 6D pose is the relative pose between the camera and the cylindrical object, which is very convenient to make augmented reality. First, we capture some images with a cylindrical object and then reconstruct its 3D model with textures offline by using projective invariance and image contours. Second, according to the 3D model, we track the 6D relative pose between the camera and the cylindrical object online, where we propose a linear P3P RANSAC to remove outliers. Finally, the virtual images are exactly aligned with the cylindrical object in the real world. Experimental results show that the proposed method outperforms the state of the arts in terms of 3D mapping and 6D pose computation on cylindrical objects.
Fulin Tang, Yihong Wu 0002, Xiaohui Hou, Haibin Ling
IEEE Trans. Circuits Syst. Video Technol.2
2019 Classification Assisted Segmentation Network for Human Parsing
abstract
In human parsing task, it is important to fully exploit global and local structure information and get accurate and coherent results. In this paper, we propose a classification assisted segmentation network, in which a multi-label classification task can obtain the probability of each class in an image that used to learn better weights for parsing task. Our method takes advantages of both the global information from classification and the detail information from segmentation. Experiments demonstrate that our method could efficiently avoid the confusion between similar categories and get more reasonable results. Particularly, it significantly boosts performances of rare categories such as scarf, belt and sunglasses with mean IoU increased by 6.29%.
Zikun Liu 0001, Yinglu Liu, Zifeng Lian, Yihong Wu 0002
ICIP5
2019 Automatic Motion-Blurred Hand Matting for Human Soft Segmentation in Videos
abstract
Accurate hand segmentation is important for human segmentation. However, in videos, hand regions usually have serious motion blur, which reduces segmentation performance obviously. To solve this problem, we propose an automatic matting network to deal with motion-blurred hands. Then we combine the hand alpha mattes provided by matting network and the human segmentation results provided by segmentation network to generate our final human soft segmentation results. In addition, to train the matting network, we need a huge amount of motion-blurred hand images and their groundtruth alpha mattes. However these images are very difficult to obtain. To solve this problem, we propose an efficient semi-automatic synthetic data generation method and generate 36186 synthetic motion-blurred hand images and their alpha mattes. Experiments on synthetic images and real videos show that our method achieves state-of-art matting performance and successfully solve the problem of bad hand segmentation caused by serious motion blur.
Xiaomei Zhao, Yihong Wu 0002
ICIP2
2019 FMD Stereo SLAM: Fusing MVG and Direct Formulation Towards Accurate and Fast Stereo SLAM
abstract
We propose a novel stereo visual SLAM framework considering both accuracy and speed at the same time. The framework makes full use of the advantages of key-feature-based multiple view geometry (MVG) and direct-based formulation. At the front-end, our system performs direct formulation and constant motion model to predict a robust initial pose, reprojects local map to find 3D-2D correspondence and finally refines pose by the reprojection error minimization. This frontend process makes our system faster. At the back-end, MVG is used to estimate 3D structure. When a new keyframe is inserted, new mappoints are generated by triangulating. In order to improve the accuracy of the proposed system, bad mappoints are removed and a global map is kept by bundle adjustment. Especially, the stereo constraint is performed to optimize the map. This back-end process makes our system more accurate. Experimental evaluation on EuRoC dataset shows that the proposed algorithm can run at more than 100 frames per second on a consumer computer while achieving highly competitive accuracy.
Fulin Tang, Heping Li, Yihong Wu 0002
ICRA3
2019 Efficient conic fitting with an analytical Polar-N-Direction geometric distance
Yihong Wu 0002, Haoren Wang, Fulin Tang, Zhiheng Wang 0001
Pattern Recognit.1
2019 Automatically Extract Semi-Transparent Motion-Blurred Hand From a Single Image
abstract
When we use video chat, video game, or other video applications, motion-blurred hands often appear. Accurately extracting these hands is very useful for video editing and behavior analysis. However, existing motion-blurred object extraction methods either need user interactions, such as user supplied trimaps and scribbles, or need additional information, such as background images. In this letter, a novel method which can automatically extract the semi-transparent motion-blurred hand just according to the original RGB image is proposed. The proposed method separates the extraction task into two subtasks: alpha matte prediction and foreground prediction. These two subtasks are implemented by Xception based encoder-decoder networks. The images of extracted motion-blurred hands are calculated by multiplying the predicted alpha mattes and foreground images. Experiments on synthetic and real datasets show that the proposed method has promising performance.
Xiaomei Zhao, Yihong Wu 0002
IEEE Signal Process. Lett.2
2019 Real-time human segmentation by BowtieNet and a SLAM-based human AR system
abstract
Generally, it is difficult to obtain accurate pose and depth for a non-rigid moving object from a single RGB camera to create augmented reality (AR). In this study, we build an augmented reality system from a single RGB camera for a non-rigid moving human by accurately computing pose and depth, for which two key tasks are segmentation and monocular Simultaneous Localization and Mapping (SLAM). Most existing monocular SLAM systems are designed for static scenes, while in this AR system, the human body is always moving and non-rigid. In order to make the SLAM system suitable for a moving human, we first segment the rigid part of the human in each frame. A segmented moving body part can be regarded as a static object, and the relative motions between each moving body part and the camera can be considered the motion of the camera. Typical SLAM systems designed for static scenes can then be applied. In the segmentation step of this AR system, we first employ the proposed BowtieNet, which adds the atrous spatial pyramid pooling (ASPP) of DeepLab between the encoder and decoder of SegNet to segment the human in the original frame, and then we use color information to extract the face from the segmented human area. Based on the human segmentation results and a monocular SLAM, this system can change the video background and add a virtual object to humans. The experiments on the human image segmentation datasets show that BowtieNet obtains state-of-the-art human image segmentation performance and enough speed for real-time segmentation. The experiments on videos show that the proposed AR system can robustly add a virtual object to humans and can accurately change the video background.
Xiaomei Zhao, Fulin Tang, Yihong Wu 0002
Virtual Real. Intell. Hardw.3
2018 A Geometry-Based Point Cloud Reduction Method for Mobile Augmented Reality System
Hao-Ren Wang, Juan Lei, Yihong Wu 0002
J. Comput. Sci. Technol.4
2018 A deep learning model integrating FCNNs and CRFs for brain tumor segmentation
Xiaomei Zhao, Yihong Wu 0002, Guidong Song, Zhenye Li 0003, Yazhuo Zhang, Yong Fan 0001
Medical Image Anal.2
2017 Real-time SLAM relocalization with online learning of binary feature indexing
Youji Feng, Yihong Wu 0002, Lixin Fan
Mach. Vis. Appl.2
2016 Fast Localization in Large-Scale Environments Using Supervised Indexing of Binary Features
abstract
The essence of image-based localization lies in matching 2D key points in the query image and 3D points in the database. State-of-the-art methods mostly employ sophisticated key point detectors and feature descriptors, e.g., Difference of Gaussian (DoG) and Scale Invariant Feature Transform (SIFT), to ensure robust matching. While a high registration rate is attained, the registration speed is impeded by the expensive key point detection and the descriptor extraction. In this paper, we propose to use efficient key point detectors along with binary feature descriptors, since the extraction of such binary features is extremely fast. The naive usage of binary features, however, does not lend itself to significant speedup of localization, since existing indexing approaches, such as hierarchical clustering trees and locality sensitive hashing, are not efficient enough in indexing binary features and matching binary features turns out to be much slower than matching SIFT features. To overcome this, we propose a much more efficient indexing approach for approximate nearest neighbor search of binary features. This approach resorts to randomized trees that are constructed in a supervised training process by exploiting the label information derived from that multiple features correspond to a common 3D point. In the tree construction process, node tests are selected in a way such that trees have uniform leaf sizes and low error rates, which are two desired properties for efficient approximate nearest neighbor search. To further improve the search efficiency, a probabilistic priority search strategy is adopted. Apart from the label information, this strategy also uses non-binary pixel intensity differences available in descriptor extraction. By using the proposed indexing approach, matching binary features is no longer much slower but slightly faster than matching SIFT features. Consequently, the overall localization speed is significantly improved due to the much faster key point detection and descriptor extraction. It is empirically demonstrated that the localization speed is improved by an order of magnitude as compared with state-of-the-art methods, while comparable registration rate and localization accuracy are still maintained.
Youji Feng, Lixin Fan, Yihong Wu 0002
IEEE Trans. Image Process.3
2015 An efficient approach for 2D to 3D video conversion based on structure from motion
Wei Liu 0023, Yihong Wu 0002, Fusheng Guo, Zhanyi Hu
Vis. Comput.2
2014 Efficient pose tracking on mobile phones with 3D points grouping
abstract
With the rapid growth of computational capability and popularity of mobile phones, Mobile Augmented Reality (MAR) in large scale 3D scenes becomes an emerging field in recent years. The core of MAR is to continuously compute a precise 6 Degree-of-Freedom (DOF) camera pose for each frame, i.e. localization. However, as a crucial part of localization, the 2D-3D points matching is usually inefficient due to the usage of traditional features, e.g. SIFT, and the large number of 3D points candidates. This paper aims to tackle this problem by designing an efficient 6DOF pose tracking system on mobile phone. In this system, binary features are used in both offline sparse reconstruction and online tracking, while a PCA-based 3D points partition method is proposed to reduce the searching space of 2D-3D points matching, making it capable to achieve a low computational cost. Experiments on a NOKIA N900 smartphone show that our system could efficiently and robustly estimate the 6DOF camera pose.
Juan Lei, Yihong Wu 0002, Lixin Fan
ICME3
2014 Radial distortion invariants and lens evaluation under a single-optical-axis omnidirectional camera
Yihong Wu 0002, Zhanyi Hu, Youfu Li 0001
Comput. Vis. Image Underst.1
2012 A Novel Fast Method for L ∞ Problems in Multiview Geometry
Zhijun Dai, Yihong Wu 0002, Fengjun Zhang, Hongan Wang
ECCV (5)2
2012 On-line Object Reconstruction and Tracking for 3D Interaction
abstract
This paper presents a flexible and easy-to-use tracking method for 3D interaction. The method reconstructs points of a user-specified object from a video sequence, and recovers the 6 degrees of freedom (DOF) camera pose and position relative to the reconstructed points in each video frame. As opposed to most existing 3D object tracking methods, the proposed method does not need any off-line modeling or training process. Instead, it first segments the object from the background, then reconstructs and tracks the object using the Visual Simultaneous Localization And Mapping (VSLAM) techniques. To our knowledge, there are no existing works investigating this kind of on line reconstruction and tracking of moving objects. The proposed method employs the adapted pyramidal Lucas-Kanade tracker to increase the stability and the robustness of the tracking when dealing with a lightly textured or fast moving object. Experiments show that fast, accurate, stable and robust tracking can be achieved in everyday environment. Moreover, a simple stereo initialization approach is adopted to minimize user intervention. All these attributes conspire to make the method an adequate tool for some interaction applications. As a concrete example, an interactive 3D scene displaying system is demonstrated.
Youji Feng, Yihong Wu 0002, Lixin Fan
ICME2
2012 Self-calibration of hybrid central catadioptric and perspective cameras
Xiaoming Deng 0001, Fuchao Wu, Yihong Wu 0002, Fuqing Duan, Liang Chang 0001, Hongan Wang
Comput. Vis. Image Underst.3
2012 A calibration method for paracatadioptric camera from sphere images
Huixian Duan, Yihong Wu 0002
Pattern Recognit. Lett.2
2011 Graph Partition Based Bundle Adjustment for Structured Dataset
abstract
Bundle adjustment has been considered as one of the most important components in many visual tasks such as 3D reconstruction, photo grammetry, visual SLAM, etc. Unfortunately, both time and space complexity of this adjustment prevent it from being directly applied to large scale datasets. This paper presents a sub mapping method, which partitions a large scale dataset into disjointed subsets and adjusts them one by one or in parallel. Pair-wise sub maps are then "stitched" together by applying a similarity transformation. Both simulations and real applications show that our method scales well. Also some basic questions of this sub mapping method including map size, map fusion and global consistency are discussed.
Yuanfan Xie, Lixin Fan, Yihong Wu 0002
ICIG3
2011 Calibration of central catadioptric camera with one-dimensional object undertaking general motions
abstract
AID object is a segment with several known-distance markers, and calibration methods with ID objects are more flexible than those with 2D/3D objects. Under the pinhole camera model, it is proved that the calibration with free-moving ID objects is not possible. For a central catadioptric camera setup, can the camera be calibrated by a ID object under general motions? In this paper, we prove that a central catadioptric camera can indeed be calibrated, and propose a catadioptric camera calibration method using ID objects undertaking general motions. The proposed method consists of two steps. Firstly, the principal point is calculated with geometric invariants under catadioptric camera model; Secondly, we use images of ID object to calibrate the focal lengths, skew factor and mirror parameter. The method needs neither prior knowledge of catadioptric parameters nor conic fitting, and it is linear, which makes it easy to implement. Experiments demonstrate its usefulness and stability.
Xiaoming Deng 0001, Fuchao Wu, Yihong Wu 0002, Liang Chang 0001, Wei Liu 0023, Hongan Wang
ICIP3
2011 Paracatadioptric camera calibration using sphere images
abstract
The problem of calibrating paracatadioptric camera from sphere images is still open. In this paper, we propose a calibration method for paracatadioptric camera based on spheres. We notice that, under central catadioptric camera, a sphere is projected to two conies on the image plane, which can also be seen as the projections of two parallel circles on the viewing sphere by a virtue camera. These two conies are called a pair of antipodal sphere images. Firstly, we study properties of K(K >; 3) pairs of antipodal sphere images under paracatadioptric camera. Then, they can be estimated using these properties. Finally, paracatadioptric camera can be calibrated by three pairs of antipodal sphere images or more. The method only requires the projected contour of parabolic mirror is visible on the image plane in one view. Experimental results on both simulated and real image data have demonstrated the effectiveness of our method.
Huixian Duan, Yihong Wu 0002
ICIP2
2010 PCA-based structure refinement for reconstruction of urban scene
abstract
There is plenty of structured information (such as lines and planes) in urban scenes. Considering this, we propose a new method for making use of this information to enhance the reconstruction of urban outdoor scenes. Structured information (collinearity and coplanarity) is extracted from images by performing line detection and color image segmentation, which is used as hypothetic constraints of the 3D structure. In refining stage, we first build PCA subspaces for each structured components (collinear and coplanar point sets), during which the former hypothetic structure information is further inspected by the initial 3D structure. Then we iteratively update the structure through EM estimation. Experiments show that this method effectively improves the accuracy and robustness of reconstruction of urban scenes.
Qinxun Bai, Yihong Wu 0002, Lixin Fan
ICIP2
2010 Degeneracy from Twisted Cubic Under Two Views
Yihong Wu 0002, Zhanyi Hu
J. Comput. Sci. Technol.1
2009 Twisted Cubic: Degeneracy Degree and Relationship with General Degeneracy
Yihong Wu 0002, Zhanyi Hu
ACCV (2)2
2009 A Model Based Method for Overall Well Focused Catadioptric Image Acquisition with Multi-focal Images
Youfu Li 0001, Yihong Wu 0002
CAIP3
2009 Pointwise Motion Image (PMI): A Novel Motion Representation and Its Applications to Abnormality Detection and Behavior Recognition
abstract
In this paper, we propose a novel motion representation and apply it to abnormality detection and behavior recognition. At first, pointwise correspondences for the foreground in two consecutive video frames are established by performing a salient-region-based pointwise matching algorithm. Then, based on the established pointwise correspondences, a pointwise motion image (PMI) for each frame is built up to represent the motion status of the foreground. The PMI is more suitable for video analysis as it encapsulates a variety of motion information such as pointwise motion speed, pointwise motion orientation, pointwise motion duration, as well as the global shape of the foreground. In addition, it represents all of these pieces of information by a color image in the HSV space, by which many popular techniques in the image processing field can be straightforwardly adopted. By combining the PMI and AdaBoost, a method for abnormality detection and behavior recognition is proposed. The proposed method is shown to possess a high discriminative ability and is capable of dealing with local motion, global motion, and similar motions with different speeds. Experiments including a comparison with two existing methods demonstrate the effectiveness of the proposed representation in abnormality detection and behavior recognition.
Qiulei Dong, Yihong Wu 0002, Zhanyi Hu
IEEE Trans. Circuits Syst. Video Technol.2
2008 Visual metrology with uncalibrated radial distorted images
abstract
Visual metrology methods with radial distorted images usually require a radial distortion model and a pre-calibration. In this paper, we propose a novel 3D metrology algorithm with at least three uncalibrated radial distorted images, and also derive a 2D metrology algorithm with at least two uncalibrated radial distorted images. The algorithm does not require a radial distortion model or calibrating camera intrinsic parameters except for radial distortion center, which can be usually known as a prior or computed easily, and correspondences of control points with known coordinates. The algorithm is of high accuracy, and robust to noise due to no requirement for specific radial distortion models. Experimental results show the feasibility and accuracy of the algorithm.
Xiaoming Deng 0001, Fuchao Wu, Yihong Wu 0002, Fuqing Duan
ICPR3
2008 Detecting and Handling Unreliable Points for Camera Parameter Estimation
Yihong Wu 0002, Youfu Li 0001, Zhanyi Hu
Int. J. Comput. Vis.1
2008 A new linear algorithm for calibrating central catadioptric cameras
Fuchao Wu, Fuqing Duan, Zhanyi Hu, Yihong Wu 0002
Pattern Recognit.4
2007 MAPACo-Training: A Novel Online Learning Algorithm of Behavior Models
Heping Li, Zhanyi Hu, Yihong Wu 0002, Fuchao Wu
ACCV (1)3
2006 Gesture Recognition Using Quadratic Curves
Qiulei Dong, Yihong Wu 0002, Zhanyi Hu
ACCV (1)2
2006 Detecting Critical Configuration of Six Points
Yihong Wu 0002, Zhanyi Hu
ACCV (2)1
2006 Euclidean Structure from N geq 2 Parallel Circles: Theory and Algorithms
Pierre Gurdjos, Peter F. Sturm, Yihong Wu 0002
ECCV (1)3
2006 Easy Calibration for Para-catadioptric-like Camera
abstract
For omnidirectional cameras, most of the previous calibration methods from lines use conic fitting. This paper presents a calibration method for para-catadioptric-like cameras from lines without conic fitting under a single view. We establish equations on the five camera intrinsic parameters. These equations are linear for the focal lengths and skew factor once the principal point is known. The principal point can be approximated well by the center of the imaged mirror contour in practice or can be accurately estimated by quadric equations. After obtaining the principal point, we propose an algorithm to calibrate the focal lengths and skew factor. The algorithm needs neither prior structure knowledge nor conic fitting and is linear, which make it easy to implement. Other omnidirectional cameras can also use this presented work if high accuracy is not required. Experiments demonstrate the efficiency of the proposed algorithm.
Yihong Wu 0002, Youfu Li 0001, Zhanyi Hu
IROS1
2006 A robust method to recognize critical configuration for camera calibration
Yihong Wu 0002, Zhanyi Hu
Image Vis. Comput.1
2006 Coplanar circles, quasi-affine invariance and calibration
Yihong Wu 0002, Xinju Li, Fuchao Wu, Zhanyi Hu
Image Vis. Comput.1
2006 Euclidean reconstruction of a circular truncated cone only from its uncalibrated contours
Yihong Wu 0002, Guanghui Wang 0001, Fuchao Wu, Zhanyi Hu
Image Vis. Comput.1
2006 The Number of Independent Kruppa Constraints from N Images
Zhanyi Hu, Yihong Wu 0002, Fuchao Wu, Songde Ma
J. Comput. Sci. Technol.2
2005 Geometric Invariants and Applications under Catadioptric Camera Model
abstract
This paper presents geometric invariants of points and their applications under central catadioptric camera model. Although the image has severe distortions under the model, we establish some accurate projective geometric invariants of scene points and their image points. These invariants, being functions of principal point, are useful, from which a method for calibrating the camera principal point and a method for recovering planar scene structures are proposed. The main advantage of using these in variants for plane reconstruction is that neither camera motion nor the intrinsic parameters, except for the principal point, is needed. The theoretical correctness of the established invariants and robustness of the proposed methods are demonstrated by experiments. In addition, our results are found to be applicable to some more general camera models other than the catadioptric one
Yihong Wu 0002, Zhanyi Hu
ICCV1
2005 A new constraint on the imaged absolute conic from aspect ratio and its application
Yihong Wu 0002, Zhanyi Hu
Pattern Recognit. Lett.1
2004 Camera Calibration from the Quasi-affine Invariance of Two Parallel Circles
Yihong Wu 0002, Haijiang Zhu, Zhanyi Hu, Fuchao Wu
ECCV (1)1
2003 Automated short proof generation for projective geometric theorems with Cayley and bracket algebras: I. Incidence geometry
Hongbo Li 0012, Yihong Wu 0002
J. Symb. Comput.2
2003 Automated short proof generation for projective geometric theorems with Cayley and bracket algebras: II. Conic geometry
Hongbo Li 0012, Yihong Wu 0002
J. Symb. Comput.2
2003 The Invariant Representations of a Quadric Cone and a Twisted Cubic
abstract
Up to now, the shortest invariant representation of a quadric has 138 summands and there has been no invariant representation of a twisted cubic in 3D projective space, which limit to some extent the applications of invariants in 3D space. We give a very short invariant representation of a quadric cone, a special quadric, which has only two summands similar to the invariant representation of a planar conic, and give a short invariant representation of a twisted cubic. Then, a completely linear algorithm for generating the parametric equations of a twisted cubic is provided also. Finally, we exemplify some applications of our proposed invariant representations in the fields of computer vision and automated geometric theorem proving.
Yihong Wu 0002, Zhanyi Hu
IEEE Trans. Pattern Anal. Mach. Intell.1