Yifu Wang

dblp:00/3780 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 7 since 2021Systems, architecture and hardware · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 An intelligent retrievable object-tracking system with real-time edge inference capability
abstract
Abstract An intelligent retrievable object‐tracking system assists users in quickly and accurately locating lost objects. However, challenges such as real‐time processing on edge devices, low image resolution, and small‐object detection significantly impact the accuracy and efficiency of video‐stream‐based systems, especially in indoor home environments. To overcome these limitations, a novel real‐time intelligent retrievable object‐tracking system is designed. The system incorporates a retrievable object‐tracking algorithm that combines DeepSORT and sliding window techniques to enhance tracking capabilities. Additionally, the YOLOv7‐small‐scale model is proposed for small‐object detection, integrating a specialized detection layer and the convolutional batch normalization LeakyReLU spatial‐depth convolution module to enhance feature capture for small objects. TensorRT and INT8 quantization are used for inference acceleration on edge devices, doubling the frames per second. Experiments on a Jetson Nano (4 GB) using YOLOv7‐small‐scale show an 8.9% improvement in recognition accuracy over YOLOv7‐tiny in video stream processing. This advancement significantly boosts the system's performance in efficiently and accurately locating lost objects in indoor home settings.
Yujie Li 0002, Yifu Wang, Zihang Ma, Xinghe Wang, Benying Tan, Shuxue Ding
IET Image Process.2
2025 Weighted Squared Volume Minimization (WSVM) for Generating Uniform Tetrahedral Meshes
abstract
This paper presents a new algorithm, Weighted Squared Volume Minimization (WSVM), for generating high-quality tetrahedral meshes from closed triangle meshes. Drawing inspiration from the principle of minimal surfaces that minimize squared surface area, WSVM employs a new energy function integrating weighted squared volumes for tetrahedral elements. When minimized with constant weights, this energy promotes uniform volumes among the tetrahedra. Adjusting the weights to account for local geometry further achieves uniform dihedral angles within the mesh. The algorithm begins with an initial tetrahedral mesh generated via Delaunay tetrahedralization and proceeds by sequentially minimizing volume-oriented and then dihedral angle-oriented energies. At each stage, it alternates between optimizing vertex positions and refining mesh connectivity through the iterative process. The algorithm operates fully automatically and requires no parameter tuning. Evaluations on a variety of 3D models demonstrate that WSVM consistently produces tetrahedral meshes of higher quality, with fewer slivers and enhanced uniformity compared to existing methods.
Kaixin Yu, Yifu Wang, Peng Song 0001, Xiangqiao Meng, Ying He 0001, Jianjun Chen 0002
IEEE Trans. Vis. Comput. Graph.2
2024 Vision Transformer with 2D Explicit Position Encoding
abstract
Recently, the Vision Transformer (ViT) has achieved outstanding performance in various computer vision tasks. Positional encoding is an indispensable component of ViT for handling the inherent structural information of images. However, attaching position encodings manually is a time-consuming process that slows down the training speed of ViT. To address this issue, we propose an explicit approach for positional encoding, distinct from the original ViT’s implicit design. Our new implementation uses a 2D-based explicit positional encoding method that accelerates convergence and improves training efficiency. The proposed approach yields a remarkable improvement, especially in the initial stages of training, where the 2D explicit positional encoding offers improved compatibility with various input lengths and enhanced interpretability. The experimental results on the ImageNet dataset confirm the effectiveness of our proposed 2D explicit positional encoding approach. The proposed explicit 2D coordinate position encoding can achieve a maximum improvement of up to 437%.
Zihang Ma, Xinghe Wang, Yifu Wang, Benying Tan
ICASSP4
2024 Sod-Uav: Small Object Detection For Unmanned Aerial Vehicle Images Via Improved Yolov7
abstract
Detecting small objects in Unmanned Aerial Vehicle (UAV) images is pivotal for a multitude of applications. Given the high-altitude perspective of UAVs, the images they capture often feature intricate backgrounds, pronounced object heterogeneity, and a plethora of sparsely situated small targets. These characteristics pose significant challenges to conventional detection algorithms. In response, we introduce an enhanced YOLOv7-based technique specifically tailored for small object detection in UAV images. Our approach adds more layers dedicated to small object detection and leverages the Bi-directional Feature Pyramid Network (BiFPN) to extract features across diverse scales. Furthermore, we enhance the detection heads by standardizing channel configurations and incorporate attention mechanisms through the dynamic head framework (DyHead). This allows the model to adaptively modify the detection head structure, catering to varying scales, tasks, and features. Preliminary results on the VisDrone2019 dataset indicate that our method surpasses existing state-of-the-art algorithms in UAV-based small object detection.
Yifu Wang, Zihang Ma, Xinghe Wang, Yutao Tang
ICASSP2
2024 MAVIS: Multi-Camera Augmented Visual-Inertial SLAM using SE2(3) Based Exact IMU Pre-integration
abstract
We present a novel optimization-based Visual-Inertial SLAM system designed for multiple partially over-lapped camera systems, named MAVIS. Our framework fully exploits the benefits of wide field-of-view from multi-camera systems, and the metric scale measurements provided by an inertial measurement unit (IMU). We introduce an improved IMU pre-integration formulation based on the exponential function of an automorphism of SE2(3), which can effectively enhance tracking performance under fast rotational motion and extended integration time. Furthermore, we extend conventional front-end tracking and back-end optimization module designed for monocular or stereo setup towards multi-camera systems, and introduce implementation details that contribute to the performance of our system in challenging scenarios. The practical validity of our approach is supported by our experiments on public datasets. Our MAVIS won the first place in all the vision-IMU tracks (single and multi-session SLAM) on Hilti SLAM Challenge 2023 with 1.7 times the score compared to the second place1.
Yifu Wang, Yonhon Ng, Inkyu Sa, Álvaro Parra Bustos, Cristian Rodriguez Opazo, Hongdong Li
ICRA1
2024 Cross-Modal Semidense 6-DOF Tracking of an Event Camera in Challenging Conditions
abstract
Vision-based localization is a cost-effective and thus attractive solution for many intelligent mobile platforms. However, its accuracy and especially robustness still suffer from low illumination conditions, illumination changes, and aggressive motion. Event-based cameras are bio-inspired visual sensors that perform well in HDR conditions and have high temporal resolution, and thus provide an interesting alternative in such challenging scenarios. While purely event-based solutions currently do not yet produce satisfying mapping results, the present work demonstrates the feasibility of purely event-based tracking if an alternative sensor is permitted for mapping. The method relies on geometric 3D-2D registration of semi-dense maps and events, and achieves highly reliable and accurate cross-modal tracking results. Practically relevant scenarios are given by depth camera-supported tracking or map-based localization with a semi-dense map prior created by a regular image-based visual SLAM or structure-from-motion system. Conventional edge-based 3D-2D alignment is extended by a novel polarity-aware registration that makes use of signed time-surface maps (STSM) obtained from event streams. We furthermore introduce a novel culling strategy for occluded points. Both modifications increase the speed of the tracker and its robustness against occlusions or large view-point variations. The approach is validated on many real datasets covering the above-mentioned challenging conditions, and compared against similar solutions realised with regular cameras.
Wanting Xu, Xia Wang 0002, Yifu Wang, Laurent Kneip
IEEE Trans. Robotics4
2023 Revisiting Event-Based Video Frame Interpolation
abstract
Dynamic vision sensors or event cameras provide rich complementary information for video frame interpolation. Existing state-of-the-art methods follow the paradigm of combining both synthesis-based and warping networks. However, few of those methods fully respect the intrinsic characteristics of events streams. Given that event cameras only encode intensity changes and polarity rather than color intensities, estimating optical flow from events is arguably more difficult than from RGB information. We therefore propose to incorporate RGB information in an event-guided optical flow refinement strategy. Moreover, in light of the quasi-continuous nature of the time signals provided by event cameras, we propose a divide-and-conquer strategy in which event-based intermediate frame synthesis happens incrementally in multiple simplified stages rather than in a single, long stage. Extensive experiments on both synthetic and real-world datasets show that these modifications lead to more reliable and realistic intermediate frame results than previous video frame interpolation methods. Our findings underline that a careful consideration of event characteristics such as high temporal density and elevated noise benefits interpolation accuracy.
Jiaben Chen, Dongze Lian, Yifu Wang, Renrui Zhang, Xinhang Liu, Shenhan Qian, Laurent Kneip, Shenghua Gao
IROS5
2022 Accurate Calibration of Multi-Perspective Cameras from a Generalization of the Hand-Eye Constraint
abstract
Multi-perspective cameras are quickly gaining importance in many applications such as smart vehicles and virtual or augmented reality. However, a large system size or absence of overlap in neighbouring fields-of-view often complicate their calibration. We present a novel solution which relies on the availability of an external motion capture system. Our core contribution consists of an extension to the hand-eye calibration problem which jointly solves multi-eye-to-base problems in closed form. We furthermore demonstrate its equivalence to the multi-eye-in-hand problem. The practical validity of our approach is supported by our experiments, indicating that the method is highly efficient and accurate, and outperforms existing closed-form alternatives.
Yifu Wang, Wenqing Jiang, Sören Schwertfeger, Laurent Kneip
ICRA1
2022 DEVO: Depth-Event Camera Visual Odometry in Challenging Conditions
abstract
We present a novel real-time visual odometry framework for a stereo setup of a depth and high-resolution event camera. Our framework balances accuracy and robustness against computational efficiency towards strong performance in challenging scenarios. We extend conventional edge-based semi-dense visual odometry towards time-surface maps obtained from event streams. Semi-dense depth maps are generated by warping the corresponding depth values of the extrinsically calibrated depth camera. The tracking module updates the camera pose through efficient, geometric semi-dense 3D-2D edge alignment. Our approach is validated on both public and self-collected datasets captured under various conditions. We show that the proposed method performs comparable to state-of-the-art RGB-D camera-based alternatives in regular conditions, and eventually outperforms in challenging conditions such as high dynamics or low illumination.
Jiaben Chen, Xia Wang 0002, Yifu Wang, Laurent Kneip
ICRA5
2022 Globally-Optimal Contrast Maximisation for Event Cameras
abstract
Event cameras are bio-inspired sensors that perform well in challenging illumination conditions and have high temporal resolution. However, their concept is fundamentally different from traditional frame-based cameras. The pixels of an event camera operate independently and asynchronously. They measure changes of the logarithmic brightness and return them in the highly discretised form of time-stamped events indicating a relative change of a certain quantity since the last event. New models and algorithms are needed to process this kind of measurements. The present work looks at several motion estimation problems with event cameras. The flow of the events is modelled by a general homographic warping in a space-time volume, and the objective is formulated as a maximisation of contrast within the image of warped events. Our core contribution consists of deriving globally optimal solutions to these generally non-convex problems, which removes the dependency on a good initial guess plaguing existing methods. Our methods rely on branch-and-bound optimisation and employ novel and efficient, recursive upper and lower bounds derived for six different contrast estimation functions. The practical validity of our approach is demonstrated by a successful application to three different event camera motion estimation problems.
Xin Peng 0005, Ling Gao 0001, Yifu Wang, Laurent Kneip
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 B-splines for Purely Vision-based Localization and Mapping on Non-holonomic Ground Vehicles
abstract
Purely vision-based localization and mapping is a cost-effective and thus attractive solution to localization and mapping on smart ground vehicles. However, the accuracy and especially robustness of vision-only solutions remain rivalled by more expensive, lidar-based multi-sensor alternatives. We show that a significant increase in robustness can be achieved if taking non-holonomic kinematic constraints on the vehicle motion into account. Rather than using approximate planar motion models or simple, pair-wise regularization terms, we demonstrate the use of B-splines for an exact imposition of smooth, non-holonomic trajectories inside the 6 DoF bundle adjustment. We introduce both hard and soft formulations and compare their computational efficiency and accuracy against traditional solutions. Through results on both simulated and real data, we demonstrate a significant improvement in robustness and accuracy in degrading visual conditions.
Yifu Wang, Laurent Kneip
ICRA2
2021 Dynamic Event Camera Calibration
abstract
Camera calibration is an important prerequisite towards the solution of 3D computer vision problems. Traditional methods rely on static images of a calibration pattern. This raises interesting challenges towards the practical usage of event cameras, which notably require image change to produce sufficient measurements. The current standard for event camera calibration therefore consists of using flashing patterns. They have the advantage of simultaneously triggering events in all reprojected pattern feature locations, but it is difficult to construct or use such patterns in the field. We present the first dynamic event camera calibration algorithm. It calibrates directly from events captured during relative motion between camera and calibration pattern. The method is propelled by a novel feature extraction mechanism for calibration patterns, and leverages existing calibration tools before optimizing all parameters through a multi-segment continuous-time formulation. As demonstrated through our results on real data, the obtained calibration method is highly convenient and reliably calibrates from data sequences spanning less than 10 seconds.
Yifu Wang, Laurent Kneip
IROS2
2020 Globally-Optimal Event Camera Motion Estimation
Xin Peng 0005, Yifu Wang, Ling Gao 0001, Laurent Kneip
ECCV (26)2
2020 Reliable frame-to-frame motion estimation for vehicle-mounted surround-view camera systems
abstract
Modern vehicles are often equipped with a surround-view multi-camera system. The current interest in autonomous driving invites the investigation of how to use such systems for a reliable estimation of relative vehicle displacement. Existing camera pose algorithms either work for a single camera, make overly simplified assumptions, are computationally expensive, or simply become degenerate under non-holonomic vehicle motion. In this paper, we introduce a new, reliable solution able to handle all kinds of relative displacements in the plane despite the possibly non-holonomic characteristics. We furthermore introduce a novel two-view optimization scheme which minimizes a geometrically relevant error without relying on 3D point related optimization variables. Our method leads to highly reliable and accurate frame-to-frame visual odometry with a full-size, vehicle-mounted surround-view camera system.
Yifu Wang, Xin Peng 0005, Hongdong Li, Laurent Kneip
ICRA1
2019 Motion Estimation of Non-Holonomic Ground Vehicles From a Single Feature Correspondence Measured Over N Views
abstract
The planar motion of ground vehicles is often non-holonomic, which enables a solution of the two-view relative pose problem from a single point feature correspondence. Man-made environments such as underground parking lots are however dominated by line features. Inspired by the planar tri-focal tensor and its ability to handle lines, we establish an n-linear constraint on the locally circular motion of non-holonomic vehicles able to handle an arbitrarily large and dense window of views. We prove that this stays a uni-variate problem under the assumption of locally constant vehicle speed, and it can transparently handle both point and vertical line correspondences. In particular, we prove that an application of Viète's formulas for extrapolating trigonometric functions of angle multiples and the Weierstrass substitution casts the problem as one that merely seeks the roots of a uni-variate polynomial. We present the complete theory of this novel solver, and test it on both simulated and real data. Our results prove that it successfully handles a variety of relevant scenarios, eventually outperforming the 1-point two-view solver.
Yifu Wang, Laurent Kneip
CVPR2
2017 On Scale Initialization in Non-overlapping Multi-perspective Visual Odometry
Yifu Wang, Laurent Kneip
ICVS1
2013 Research on Task Allocation Strategy and Scheduling Algorithm of Multi-core Load Balance
abstract
Based on the research of multi-core load balancing's task scheduling and allocation, we proposed the static task graphs stratification algorithm, the static task group scheduling algorithm, and the minimum dynamic link algorithm, aiming at the characteristics of multi-core processors. When these algorithms allocate tasks, they are expected to complete multi-core load balancing. Firstly, the task allocation is divided into two stages: It needs to break dependencies among tasks and relatively independent tasks will be in the same group at the first stage. It conducts static allocation for the principle of load balancing and it allocates initial tasks which have almost the same time for the system hardware threads in the second stage. It allocates tasks which come from system's running for each hard ware thread with processor's speed as a standard in the third stage. From the verification of simulation experiment, the algorithms can achieve better load balancing and minimum completion time.
Yifu Wang, Aoyang Zhao, Tie Qiu 0001
CISIS2