EDBT 2026 Demo / reviewers in the wild / expert
Feihu Zhang
dblp:120/0587
· DBLP profile ↗
40ranked-venue papers
15as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 12 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorSystems, architecture and hardware · 3 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Video Depth Anything: Consistent Depth Estimation for Super-Long VideosabstractDepth Anything has achieved remarkable success in monocular depth estimation with strong generalization ability. However, it suffers from temporal inconsistency in videos, hindering its practical applications. Various methods have been proposed to alleviate this issue by leveraging video generation models or introducing priors from optical flow and camera poses. Nonetheless, these methods are only applicable to short videos (< 10 seconds) and require a trade-off between quality and computational efficiency. We propose Video Depth Anything for high-quality, consistent depth estimation in super-long videos (over several minutes) without sacrificing efficiency. We base our model on Depth Anything V2 and replace its head with an efficient spatial-temporal head. We design a straightforward yet effective temporal consistency loss by constraining the temporal depth gradient, eliminating the need for additional geometric priors. The model is trained on a joint dataset of video depth and unlabeled images, similar to Depth Anything V2. Moreover, a novel key-frame-based strategy is developed for long video inference. Experiments show that our model can be applied to arbitrarily long videos without compromising quality, consistency, or generalization ability. Comprehensive evaluations on multiple video benchmarks demonstrate that our approach sets a new state-of-the-art in zero-shot video depth estimation. We offer models of different scales to support a range of scenarios, with our smallest model capable of real-time performance at 30 FPS. Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Jiashi Feng, Bingyi Kang |
CVPR | 4 |
| 2025 | High-quality Text-to-3D Character Generation with SparseCubes and Sparse TransformersabstractCurrent state-of-the-art text-to-3D generation methods struggle to produce 3D models with fine details and delicate structures due to limitations in differentiable mesh representation techniques. This limitation is particularly pronounced in anime character generation, where intricate features such as fingers, hair, and facial details are crucial for capturing the essence of the characters.
In this paper, we introduce a novel, efficient, sparse differentiable mesh representation method, termed SparseCubes, alongside a sparse transformer network designed to generate high-quality 3D models. Our method significantly reduces computational requirements by over 95% and storage memory by 50%, enabling the creation of higher resolution meshes with enhanced details and delicate structures. We validate the effectiveness of our approach through its application to text-to-3D anime character generation, demonstrating its capability to accurately render subtle details and thin structures (e.g. individual fingers) in both meshes and textures. Jiachen Qian, Hongye Yang, Jingxi Xu 0001, Feihu Zhang |
ICLR | 5 |
| 2025 | Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse AttentionabstractGenerating high-resolution 3D shapes using volumetric representations such as Signed Distance Functions (SDFs) presents substantial computational and memory challenges. We introduce Direct3D-S2, a scalable 3D generation framework based on sparse volumes that achieves superior output quality with dramatically reduced training costs.
Our key innovation is the Spatial Sparse Attention (SSA) mechanism, which greatly enhances the efficiency of Diffusion Transformer (DiT) computations on sparse volumetric data. SSA allows the model to effectively process large token sets within sparse volumes, significantly reducing computational overhead and achieving a 3.9$\times$ speedup in the forward pass and a 9.6$\times$ speedup in the backward pass.
Our framework also includes a variational autoencoder (VAE) that maintains a consistent sparse volumetric format across input, latent, and output stages. Compared to previous methods with heterogeneous representations in 3D VAE, this unified design significantly improves training efficiency and stability.
Our model is trained on public datasets, and experiments demonstrate that Direct3D-S2 not only surpasses state-of-the-art methods in generation quality and efficiency, but also enables training at 1024³ resolution using only 8 GPUs—a task typically requiring at least 32 GPUs for volumetric representations at $256^3$ resolution, thus making gigascale 3D generation both practical and accessible. Project page: https://www.neural4d.com/research-page/direct3d-s2. Youtian Lin, Feihu Zhang, Yifei Zeng, Yajie Bao, Jiachen Qian, Siyu Zhu 0001, Xun Cao, Philip Torr 0001, Yao Yao 0008 |
NeurIPS | 3 |
| 2025 | A framework for super-resolution of side-scan sonar images: Combination of variational Bayes and regional feature selection
Chensheng Cheng, Feihu Zhang, Guang Pan |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | P2GCN: Pixel-patch mutual enhancement graph convolutional network for sonar image super-resolution
Xuanfeng Li, Lichuan Zhang, Qiaoqiao Zhao, Zixiao Zhu, Feihu Zhang |
Expert Syst. Appl. | 5 |
| 2025 | S-NeRF++: Autonomous Driving Simulation via Neural Reconstruction and GenerationabstractAutonomous driving simulation system plays a crucial role in enhancing self-driving data and simulating complex and rare traffic scenarios, ensuring navigation safety. However, traditional simulation systems, which often heavily rely on manual modeling and 2D image editing, struggled with scaling to extensive scenes and generating realistic simulation data. In this study, we present S-NeRF++, an innovative autonomous driving simulation system based on neural reconstruction. Trained on widely-used self-driving datasets, such as nuScenes and Waymo, S-NeRF++ can generate a large number of realistic street scenes and foreground objects with high rendering quality as well as offering considerable flexibility in manipulation and simulation. Specifically, S-NeRF++ is an enhanced neural radiance field for synthesizing large-scale scenes and moving vehicles, with improved scene parameterization and camera pose learning. The system effectively utilizes noisy and sparse LiDAR data to refine training and address depth outliers, ensuring high-quality reconstruction and novel-view rendering. It also provides a diverse foreground asset bank by reconstructing and generating different foreground vehicles to support comprehensive scenario creation. Moreover, we have developed an advanced foreground-background fusion pipeline that skillfully integrates illumination and shadow effects, further enhancing the realism of our simulations. With the high-quality simulated data provided by our S-NeRF++, we found the perception methods enjoy performance boosts on several autonomous driving downstream tasks, further demonstrating our proposed simulator's effectiveness. Yurui Chen, Junge Zhang, Ziyang Xie, Wenye Li 0002, Feihu Zhang, Li Zhang 0040 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | NeRF-LiDAR: Generating Realistic LiDAR Point Clouds with Neural Radiance FieldsabstractLabelling LiDAR point clouds for training autonomous driving is extremely expensive and difficult. LiDAR simulation aims at generating realistic LiDAR data with labels for training and verifying self-driving algorithms more efficiently. Recently, Neural Radiance Fields (NeRF) have been proposed for novel view synthesis using implicit reconstruction of 3D scenes. Inspired by this, we present NeRF-LIDAR, a novel LiDAR simulation method that leverages real-world information to generate realistic LIDAR point clouds. Different from existing LiDAR simulators, we use real images and point cloud data collected by self-driving cars to learn the 3D scene representation, point cloud generation and label rendering. We verify the effectiveness of our NeRF-LiDAR by training different 3D segmentation models on the generated LiDAR point clouds. It reveals that the trained models are able to achieve similar accuracy when compared with the same model trained on the real LiDAR data. Besides, the generated data is capable of boosting the accuracy through pre-training which helps reduce the requirements of the real labeled data. Code is available at https://github.com/fudan-zvg/NeRF-LiDAR Junge Zhang, Feihu Zhang, Shaochen Kuang |
AAAI | 2 |
| 2024 | Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise AttentionabstractIn this paper, we introduce **Era3D**, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resulting in poor-quality multiview images. Specifically, these methods assume that the input images should comply with a predefined camera type, e.g. a perspective camera with a fixed focal length, leading to distorted shapes when the assumption fails. Moreover, the full-image or dense multiview attention they employ leads to a dramatic explosion of computational complexity as image resolution increases, resulting in prohibitively expensive training costs. To bridge the gap between assumption and reality, Era3D first proposes a diffusion-based camera prediction module to estimate the focal length and elevation of the input image, which allows our method to generate images without shape distortions. Furthermore, a simple but efficient attention layer, named row-wise attention, is used to enforce epipolar priors in the multiview diffusion, facilitating efficient cross-view information fusion. Consequently, compared with state-of-the-art methods, Era3D generates high-quality multiview images with up to a 512×512 resolution while reducing computation complexity of multiview attention by 12x times. Comprehensive experiments demonstrate the superior generation power of Era3D- it can reconstruct high-quality and detailed 3D meshes from diverse single-view input images, significantly outperforming baseline multiview diffusion methods. Yuan Liu 0025, Xiaoxiao Long, Feihu Zhang, Cheng Lin 0001, Xingqun Qi, Shanghang Zhang, Wei Xue 0002, Wenhan Luo, Ping Tan 0002, Wenping Wang 0001, Yike Guo |
NeurIPS | 4 |
| 2024 | Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerabstractGenerating high-quality 3D assets from text and images has long been challenging, primarily due to the absence of scalable 3D representations capable of capturing intricate geometry distributions. In this work, we introduce Direct3D, a native 3D generative model scalable to in-the-wild input images, without requiring a multi-view diffusion model or SDS optimization. Our approach comprises two primary components: a Direct 3D Variational Auto-Encoder (D3D-VAE) and a Direct 3D Diffusion Transformer (D3D-DiT). D3D-VAE efficiently encodes high-resolution 3D shapes into a compact and continuous latent triplane space. Notably, our method directly supervises the decoded geometry using a semi-continuous surface sampling strategy, diverging from previous methods relying on rendered images as supervision signals. D3D-DiT models the distribution of encoded 3D latents and is specifically designed to fuse positional information from the three feature maps of the triplane latent, enabling a native 3D generative model scalable to large-scale 3D datasets. Additionally, we introduce an innovative image-to-3D generation pipeline incorporating semantic and pixel-level image conditions, allowing the model to produce 3D shapes consistent with the provided conditional image input. Extensive experiments demonstrate the superiority of our large-scale pre-trained Direct3D over previous image-to-3D approaches, achieving significantly better generation quality and generalization ability, thus establishing a new state-of-the-art for 3D content creation. Project page: https://www.neural4d.com/research/direct3d. Youtian Lin, Yifei Zeng, Feihu Zhang, Jingxi Xu 0001, Philip Torr 0001, Xun Cao, Yao Yao 0008 |
NeurIPS | 4 |
| 2024 | Transformation Decoupling Strategy Based on Screw Theory for Deterministic Point Cloud Registration With Gravity PriorabstractPoint cloud registration is challenging in the presence of heavy outlier correspondences. This paper focuses on addressing the robust correspondence-based registration problem with gravity prior that often arises in practice. The gravity directions are typically obtained by inertial measurement units (IMUs) and can reduce the degree of freedom (DOF) of rotation from 3 to 1. We propose a novel transformation decoupling strategy by leveraging the screw theory. This strategy decomposes the original 4-DOF problem into three sub-problems with 1-DOF, 2-DOF, and 1-DOF, respectively, enhancing computation efficiency. Specifically, the first 1-DOF represents the translation along the rotation axis, and we propose an interval stabbing-based method to solve it. The second 2-DOF represents the pole which is an auxiliary variable in screw theory, and we utilize a branch-and-bound method to solve it. The last 1-DOF represents the rotation angle, and we propose a global voting method for its estimation. The proposed method solves three consensus maximization sub-problems sequentially, leading to efficient and deterministic registration. In particular, it can even handle the correspondence-free registration problem due to its significant robustness. Extensive experiments on both synthetic and real-world datasets demonstrate that our method is more efficient and robust than state-of-the-art methods, even when dealing with outlier rates exceeding 99%. Zijian Ma, Yinlong Liu, Walter Zimmer, Hu Cao, Feihu Zhang, Alois C. Knoll |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | FDBANet: A Fusion Frequency-Domain Denoising and Multiscale Boundary Attention Network for Sonar Image Semantic SegmentationabstractAccurate semantic segmentation of sonar terrain images can identify and classify various structures on the seabed, which plays a vital role in marine exploration and construction. However, challenges such as high noise interference, low image resolution, and blurry target boundaries in sonar images limit the performance of semantic segmentation models that are typically designed for optical images. To address these issues, we propose an encoder-decoder network named FDBANet that integrates frequency-domain denoising and multiscale boundary feature extraction modules. We solve the problem of large noise in sonar images from the perspective of the frequency domain. In addition, we combine traditional boundary detection methods with deep convolutional networks to construct a multiscale boundary feature extraction module, which enhances the ability to reconstruct object boundaries in sonar images. To verify the effectiveness of the proposed model, we construct a sonar terrain image segmentation dataset and conduct comparative experiments. The results show that FDBANet achieves effective multiobject segmentation in sonar images while maintaining low computational complexity. In addition, ablation experiments were conducted on the proposed network to further verify the importance and effectiveness of each module. Qiaoqiao Zhao, Lichuan Zhang, Feihu Zhang, Xuanfeng Li, Guang Pan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | S-NeRF: Neural Radiance Fields for Street Views
Ziyang Xie, Junge Zhang, Wenye Li 0002, Feihu Zhang, Li Zhang 0001 |
ICLR | 4 |
| 2023 | Globally Optimal Robust Radar Calibration in Intelligent Transportation SystemsabstractRadar is among the most popular sensors in modern Intelligent Transportation Systems (ITSs), enabling weather-robust perception. The orientation and position of the traffic radar relative to the ITS coordinate system are necessary for the perception fusion in ITSs. However, due to the unknown target association, sparseness and noisiness of traffic radar measurements, the robust and accurate extrinsic calibration of traffic radar is challenging. In this paper, we propose a targetless traffic radar calibration method based on GPS to overcome the inconvenience during ITS operation, because the installation of a dedicated calibration target on the highway is impractical and dangerous. On the other hand, the high-precision GPS device installed on the moving vehicle can provide traffic radar with accurate positioning information of the detection target. Furthermore, during the optimization process of extrinsic calibration, we propose a globally optimal registration method, which is robust to noise and outliers in radar measurements, and is called Gaussian Mixture Robust Branch and Bound (GMRBnB). Specifically, we first construct the robust objective function by utilizing the Gaussian Mixture Model (GMM). Then, we derive novel relaxation bounds and present the GMRBnB algorithm that overcomes the susceptibility to local minima and the dependence on initialization of traditional optimization methods. Compared with existing methods, extensive experiments in synthetic and real-world data demonstrate that our method is not only globally optimal, but also more accurate and robust. Yinlong Liu, Venkatnarayanan Lakshminarasimhan, Hu Cao, Feihu Zhang, Alois C. Knoll |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | Separable Flow: Learning Motion Cost Volumes for Optical Flow EstimationabstractFull-motion cost volumes play a central role in current state-of-the-art optical flow methods. However, constructed using simple feature correlations, they lack the ability to encapsulate prior, or even non-local knowledge. This creates artifacts in poorly constrained ambiguous regions, such as occluded and textureless areas. We propose a separable cost volume module, a drop-in replacement to correlation cost volumes, that uses non-local aggregation layers to exploit global context cues and prior knowledge, in order to disambiguate motions in these regions. Our method leads both the now standard Sintel and KITTI optical flow benchmarks in terms of accuracy, and is also shown to generalize better from synthetic to real data. Feihu Zhang, Oliver J. Woodford, Victor Adrian Prisacariu, Philip Torr 0001 |
ICCV | 1 |
| 2021 | Looking Beyond Single Images for Contrastive Semantic Segmentation LearningabstractWe present an approach to contrastive representation learning for semantic segmentation. Our approach leverages the representational power of existing feature extractors to find corresponding regions across images. These cross-image correspondences are used as auxiliary labels to guide the pixel-level selection of positive and negative samples for more effective contrastive learning in semantic segmentation. We show that auxiliary labels can be generated from a variety of feature extractors, ranging from image classification networks that have been trained using unsupervised contrastive learning to segmentation models that have been trained on a small amount of labeled data. We additionally introduce a novel metric for rapidly judging the quality of a given auxiliary-labeling strategy, and empirically analyze various factors that influence the performance of contrastive learning for semantic segmentation. We demonstrate the effectiveness of our method both in the low-data as well as the high-data regime on various datasets. Our experiments show that contrastive learning with our auxiliary-labeling approach consistently boosts semantic segmentation accuracy when compared to standard ImageNet pretraining and outperforms existing approaches of contrastive and semi-supervised semantic segmentation. Feihu Zhang, Philip Torr 0001, René Ranftl, Stephan R. Richter |
NeurIPS | 1 |
| 2021 | Hypergraph convolution and hypergraph attention
Song Bai 0001, Feihu Zhang, Philip Torr 0001 |
Pattern Recognit. | 2 |
| 2020 | Deep FusionNet for Point Cloud Semantic Segmentation
Feihu Zhang, Benjamin W. Wah, Philip Torr 0001 |
ECCV (24) | 1 |
| 2020 | Domain-Invariant Stereo Matching Networks
Feihu Zhang, Xiaojuan Qi 0001, Ruigang Yang, Victor Adrian Prisacariu, Benjamin W. Wah, Philip Torr 0001 |
ECCV (2) | 1 |
| 2020 | Instance Segmentation of LiDAR Point CloudsabstractWe propose a robust baseline method for instance segmentation which are specially designed for large-scale outdoor LiDAR point clouds. Our method includes a novel dense feature encoding technique, allowing the localization and segmentation of small, far-away objects, a simple but effective solution for single-shot instance prediction and effective strategies for handling severe class imbalances. Since there is no public dataset for the study of LiDAR instance segmentation, we also build a new publicly available LiDAR point cloud dataset to include both precise 3D bounding box and point-wise labels for instance segmentation, while still being about 3~20 times as large as other existing LiDAR datasets. The dataset will be published at https://github.com/feihuzhang/LiDARSeg. Feihu Zhang, Chenye Guan, Song Bai 0001, Ruigang Yang, Philip Torr 0001, Victor Adrian Prisacariu |
ICRA | 1 |
| 2020 | High accuracy correspondence field estimation via MST based patch matching
Feihu Zhang, Shibiao Xu, Xiaopeng Zhang 0001 |
Multim. Tools Appl. | 1 |
| 2019 | GA-Net: Guided Aggregation Net for End-To-End Stereo MatchingabstractIn the stereo matching task, matching cost aggregation is crucial in both traditional methods and deep neural network models in order to accurately estimate disparities. We propose two novel neural net layers, aimed at capturing local and the whole-image cost dependencies respectively. The first is a semi-global aggregation layer which is a differentiable approximation of the semi-global matching, the second is the local guided aggregation layer which follows a traditional cost filtering strategy to refine thin structures. These two layers can be used to replace the widely used 3D convolutional layer which is computationally costly and memory-consuming as it has cubic computational/memory complexity. In the experiments, we show that nets with a two-layer guided aggregation block easily outperform the state-of-the-art GC-Net which has nineteen 3D convolutional layers. We also train a deep guided aggregation network (GA-Net) which gets better accuracies than state-of-the-art methods on both Scene Flow dataset and KITTI benchmarks. Feihu Zhang, Victor Adrian Prisacariu, Ruigang Yang, Philip Torr 0001 |
CVPR | 1 |
| 2019 | High-Speed Scene Flow on Embedded Commercial Off-the-Shelf SystemsabstractScene flow is an essential part of a stereo-based perception system for autonomous driving and mobile robotics. As in most of these platforms, the computing resource is limited but the computing requirement is high, embedded and parallelized algorithms are of vital importance for real-time tasks. This paper develops a cross-platform embedded scene flow algorithm by using an OpenCL (Open Computing Language) programming. Meanwhile, we propose a method to achieve a good performance by using a novel coarse-grained software pipeline for the embedded stream application. Experimental results show that the proposed algorithm can boost the average processing speed to 50 fps for different commercial off-the-shelf (COTS) hardware, including desktop graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and mobile phone platforms. For certain GPUs, the peak frame rates can also reach 1000 fps. By comparing the efficiency among the serial platform, we illustrate that with the help of OpenCL programming, COTS platforms can provide enough computing resources for the stereo-based perception algorithm. Long Chen 0005, Mingyue Cui, Feihu Zhang, Biao Hu 0001, Kai Huang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2018 | Color Image Segmentation Based on Evidence Theory and Two-Dimensional HistogramabstractImage segmentation is one of the most important tasks in image processing and recognition. Image segmentation based on two-dimensional histogram considers not only the target pixel information but also its neighborhood information. It segments the image according to the calculated threshold, which is a hard decision method actually. However, there is uncertainty when labeling the pixels around the threshold. In this paper, we propose a new binary segmentation method for color image based on information fusion. We use two thresholds to model the uncertainty and use Cautious OWA with evidential reasoning (COWA-ER) to implement the fusion-based color image segmentation. Experimental results show that the proposed method achieves better performance compared with the traditional two-dimensional histogram method. Deqiang Han, Zhe Zhang 0031, Weifeng Liu 0004, Feihu Zhang |
FUSION | 5 |
| 2018 | Fundamental Principles on Learning New Features for Effective Dense MatchingabstractIn dense matching (including stereo matching and optical flow), nearly all existing approaches are based on simple features, such as gray or RGB color, gradient or simple transformations like census, to calculate matching costs. These features do not perform well in complex scenes that may involve radiometric changes, noises, overexposure and/or textureless regions. Various problems may appear, such as wrong matching at the pixel or region level, flattening/breaking of edges and/or even entire structural collapse. In this paper, we propose two fundamental principles based on the consistency and the distinctiveness of features. We show that almost all existing problems in dense matching are caused by features that violate one or both of these principles. To systematically learn good features for dense matching, we develop a general multi-objective optimization based on these two principles and apply convolutional neural networks to find new features that lie on the Pareto frontier. By using two-frame optical flow and stereo matching as applications, our experimental results show that the features learned can significantly improve the performance of state-of-the-art approaches. Based on the KITTI benchmarks, our method ranks first on the two stereo benchmarks and is the best among existing two-frame optical-flow algorithms on flow benchmarks. Feihu Zhang, Benjamin W. Wah |
IEEE Trans. Image Process. | 1 |
| 2017 | Slope angle estimation based on multi-sensor fusion for a snake-like robotabstractIn this paper, we report on a body state and ground profile estimator for a snake-like robot executing a rolling gait to travel from flat ground to a slope. With the help of the estimator, the snake-like robot can adaptively adjust the body shape and locomotion speed by changing the gait parameters for the purpose of tackling a steep slope. Specifically, we propose a repeating sequence of continuous time dynamical models to fuse kinematic encoder data with on-board Inertial Measurement Unit (IMU) measurements based on extended Kalman filter (EKF). All the sensors are mounted inside each module of the snake-like robot, which measure the joint position, the three-axis acceleration, and the three-axis angular velocity. Further, the robot changes its moving pattern under our policy, judging by the estimated angle of the ground profile. We implement this estimation procedure off-line, using data extracted from repeated runs of the snake-like robot by simulation and evaluate its performance compared to the ground truth. Zhenshan Bing, Long Cheng 0007, Alois C. Knoll, Anyang Zhong, Kai Huang 0001, Feihu Zhang |
FUSION | 6 |
| 2017 | Graph based vehicle infrastructure cooperative localizationabstractThis paper presents a novel and an improved approach for estimating the position of a vehicle using vehicle-infrastructure cooperative localization. In our previous work we presented a Factor Graph based solution which added the topology (inter-vehicle distance) as a constraint while localizing the vehicle using data from sensors from both inside and outside the vehicle. This paper extends the work by reducing the error in calculating the precision of the position by almost 27% in the best case and lowering the computational time by at least 50% over our previously proposed solution. This is achieved by modifying current topology constraints to be also dependent on the previous state estimate. The proposed solution remains scalable for many vehicles without increasing the execution complexity. Finally, simulations indicate that incorporating the new topology information via Factor Graphs can improve performance over the traditional, state of the art, Kalman Filter approach. Dhiraj Gulati, Feihu Zhang, Daniel Malovetz, Daniel Clarke 0001, Gereon Hinz, Alois C. Knoll |
FUSION | 2 |
| 2017 | Robust cooperative localization in a dynamic environment using factor graphs and probability data association filterabstractAutonomous vehicles operating in dynamic environments rely on precise localization. In this paper we present a novel approach for cooperative localization of vehicular systems and an infrastructure RADAR which is resilient against outliers generated from the RADAR. The problem of cooperative localization is represented as a factor graph, where interrelated topologies (including that of outliers) are added as constraint factor between vehicle states. Corresponding probabilities for multiple topologies between states of the two vehicles are calculated using the Probability Data Association Filter and assigned to the respective edges in the graph. Simulation results indicate that this technique has significant benefits in the context of improving the resilience against outliers while optimizing joint state estimates. The methodology presented in this paper has the potential to provide a robust and flexible framework for cooperative localization in the presence of clutter, obscuration and targets entering and leaving the field of view. Dhiraj Gulati, Feihu Zhang, Daniel Malovetz, Daniel Clarke 0001, Alois C. Knoll |
FUSION | 2 |
| 2017 | Feature uncertainty estimation in sensor fusion applied to autonomous vehicle locationabstractWithin the complex driving environment, progress in autonomous vehicles is supported by advances in sensing and data fusion. Safe and robust autonomous driving can only be guaranteed provided that vehicles and infrastructure are fully aware of the driving scenario. This paper proposes a methodology for feature uncertainty prediction for sensor fusion by generating neural network surrogate models directly from data. This technique is particularly applied to vehicle location through odometry measurements, vehicle speed and orientation, to estimate the location uncertainty at any point along the trajectory. Neural networks are shown to be a suitable modeling technique, presenting good generalization capability and robust results. Clara Marina Martinez, Feihu Zhang, Daniel Clarke 0001, Gereon Hinz, Dongpu Cao |
FUSION | 2 |
| 2017 | Supplementary Meta-Learning: Towards a Dynamic Model for Deep Neural NetworksabstractData diversity in terms of types, styles, as well as radiometric, exposure and texture conditions widely exists in training and test data of vision applications. However, learning in traditional neural networks (NNs) only tries to find a model with fixed parameters that optimize the average behavior over all inputs, without using data-specific properties. In this paper, we develop a meta-level NN (MLNN) model that learns meta-knowledge on data-specific properties of images during learning and that dynamically adapts its weights during application according to the properties of the images input. MLNN consists of two parts: the dynamic supplementary NN (SNN) that learns meta-information on each type of inputs, and the fixed base-level NN (BLNN) that incorporates the meta-information from SNN into its weights at run time to realize the generalization for each type of inputs. We verify our approach using over ten network architectures under various application scenarios and loss functions. In low-level vision applications on image super-resolution and demising, MLNN has 0.1~0.3 dB improvements on PSNR, whereas for high-level image classification, MLNN has accuracy improvement of 0.4~0.6% for Cifar10 and 1.2~2.1% for ImageNet when compared to convolutional NNs (CNNs). Improvements are more pronounced as the scale or diversity of data is increased. Feihu Zhang, Benjamin W. Wah |
ICCV | 1 |
| 2016 | Vehicle infrastructure cooperative localization using Factor GraphsabstractHighly assisted and Autonomous Driving is dependent on the accurate localization of both the vehicle and other targets within the environment. With increasing traffic on roads and wider proliferation of low cost sensors, a vehicle-infrastructure cooperative localization scenario can provide improved performance over traditional mono-platform localization. The paper highlights the various challenges in the process and proposes a solution based on Factor Graphs which utilizes the concept of topology of vehicles. A Factor Graph represents probabilistic graphical model as a bipartite graph. It is used to add the inter-vehicle distance as constraints while localizing the vehicle. The proposed solution is easily scalable for many vehicles without increasing the execution complexity. Finally simulation indicates that incorporating the topology information as a state estimate can improve performance over the traditional Kalman Filter approach. Dhiraj Gulati, Feihu Zhang, Daniel Clarke 0001, Alois C. Knoll |
Intelligent Vehicles Symposium | 2 |
| 2016 | Cooperative vehicle-infrastructure localization based on the symmetric measurement equation filter
Feihu Zhang, Gereon Hinz, Dhiraj Gulati, Daniel Clarke 0001, Alois C. Knoll |
GeoInformatica | 1 |
| 2015 | Fully Connected Guided Image FilteringabstractThis paper presents a linear time fully connected guided filter by introducing the minimum spanning tree (MST) to the guided filter (GF). Since the intensity based filtering kernel of GF is apt to overly smooth edges and the fixed-shape local box support region adopted by GF is not geometric-adaptive, our filter introduces an extra spatial term, the tree similarity, to the filtering kernel of GF and substitutes the box window with the implicit support region by establishing all-pairs-connections among pixels in the image and assigning the spatial-intensity-aware similarity to these connections. The adaptive implicit support region composed by the pixels with large kernel weights in the entire image domain has a big advantage over the predefined local box window in presenting the structure of an image for the reason that: 1, MST can efficiently present the structure of an image, 2, the kernel weight of our filter considers the tree distance defined on the MST. Due to these reasons, our filter achieves better edge-preserving results. We demonstrate the strength of the proposed filter in several applications. Experimental results show that our method produces better results than state-of-the-art methods. Longquan Dai, Mengke Yuan, Feihu Zhang, Xiaopeng Zhang 0001 |
ICCV | 3 |
| 2015 | Segment Graph Based Image Filtering: Fast Structure-Preserving SmoothingabstractIn this paper, we design a new edge-aware structure, named segment graph, to represent the image and we further develop a novel double weighted average image filter (SGF) based on the segment graph. In our SGF, we use the tree distance on the segment graph to define the internal weight function of the filtering kernel, which enables the filter to smooth out high-contrast details and textures while preserving major image structures very well. While for the external weight function, we introduce a user specified smoothing window to balance the smoothing effects from each node of the segment graph. Moreover, we also set a threshold to adjust the edge-preserving performance. These advantages make the SGF more flexible in various applications and overcome the "halo" and "leak" problems appearing in most of the state-of-the-art approaches. Finally and importantly, we develop a linear algorithm for the implementation of our SGF, which has an O(N) time complexity for both gray-scale and high dimensional images, regardless of the kernel size and the intensity range. Typically, as one of the fastest edge-preserving filters, our CPU implementation achieves 0.15s per megapixel when performing filtering for 3-channel color images. The strength of the proposed filter is demonstrated by various applications, including stereo matching, optical flow, joint depth map upsampling, edge-preserving smoothing, edges detection, image abstraction and texture editing. Feihu Zhang, Longquan Dai, Shiming Xiang, Xiaopeng Zhang 0001 |
ICCV | 1 |
| 2015 | Fast Minimax Path-Based Joint Depth InterpolationabstractWe propose a fast minimax path-based depth interpolation method. The algorithm computes for each target pixel varying contributions from reliable depth seeds, and weighted averaging is used to interpolate missing depths. Compared with state-of-the-art joint geodesic upsampling method which selects the K nearest seeds to interpolate missing depths with O(Kn) complexity, our method does not need to limit the number of seeds to K and reduces the computational complexity to O(n). In addition, the minimax path chooses a path with the smallest maximum immediate pairwise pixel difference on it, so it tends to preserve sharp depth discontinuities better. In contrast to the results of previous depth upsampling algorithms, our approach can provide accurate depths with fewer artifacts. Longquan Dai, Feihu Zhang, Xing Mei, Xiaopeng Zhang 0001 |
IEEE Signal Process. Lett. | 2 |
| 2015 | PM-PM: PatchMatch With Potts Model for Object Segmentation and Stereo MatchingabstractThis paper presents a unified variational formulation for joint object segmentation and stereo matching, which takes both accuracy and efficiency into account. In our approach, depth-map consists of compact objects, each object is represented through three different aspects: 1) the perimeter in image space; 2) the slanted object depth plane; and 3) the planar bias, which is to add an additional level of detail on top of each object plane in order to model depth variations within an object. Compared with traditional high quality solving methods in low level, we use a convex formulation of the multilabel Potts Model with PatchMatch stereo techniques to generate depth-map at each image in object level and show that accurate multiple view reconstruction can be achieved with our formulation by means of induced homography without discretization or staircasing artifacts. Our model is formulated as an energy minimization that is optimized via a fast primal-dual algorithm, which can handle several hundred object depth segments efficiently. Performance evaluations in the Middlebury benchmark data sets show that our method outperforms the traditional integer-valued disparity strategy as well as the original PatchMatch algorithm and its variants in subpixel accurate disparity estimation. The proposed algorithm is also evaluated and shown to produce consistently good results for various real-world data sets (KITTI benchmark data sets and multiview benchmark data sets). Shibiao Xu, Feihu Zhang, Xiaofei He 0001, Xukun Shen, Xiaopeng Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2013 | Learning to Track Multi-target Online by Boosting and Scene LayoutabstractWe address two principal difficulties of multi-target tracking in a real traffic scenario. Firstly, fast moving traffic scenarios lead to large displacements and complex interactions with occlusions and ambiguities. Secondly, the tracking application for real traffic scenarios has the online requirement. To surmount these difficulties, we propose an approach to track the multi-target online by Boosting and scene context reasoning. To this end, we use a two-stage system, where the first stage learns a non-linear classifier which is capable of generating the observation similarities. In the second stage, we demonstrate a novel relationship between observations and the scene layout parameters. Using a probabilistic formulation and the above relationship, our method has the unique ability to handle exceptions. To evaluate our method, we create three real traffic data sets, covering urban, rural, and highway conditions. We hope that these datasets will push forward the performance of tracking systems when being moved outside the laboratory to the real world. Guang Chen 0001, Feihu Zhang, Daniel Clarke 0001, Alois C. Knoll |
ICMLA (1) | 2 |
| 2013 | Multiple vehicle cooperative localization under random finite set frameworkabstractThis paper presents a new multiple vehicle cooperative localization approach based on Random Finite Set (RFS) theory. Assuming vehicles are equipped with proprioceptive and exteroceptive sensors to localize the positions, a solution based on RFS statistics is therefore proposed to consider the whole group behavior instead of each vehicle. For this, we rely on Probability Hypothesis Density (PHD) filtering. Compared to other methods, our approach presents a recursive filtering algorithm that provides dynamic estimation of multiple vehicle states. The proposed method addresses the current challenges in multiple vehicle cooperative localization domain such as communication bandwidth issue, data association uncertainty and the over-convergence problem. A comparative study based on simulations demonstrates the reliability and the feasibility of the proposed approach in large scale environments. Feihu Zhang, Hauke Stähle, Guang Chen 0001, Christian Buckl, Alois C. Knoll |
IROS | 1 |
| 2013 | A lane marking extraction approach based on Random Finite Set StatisticsabstractWithin the past few years, lane detection technology has become of high interest in the field of intelligent vehicles; however, robustness is still an issue. The challenge is to extract the lane markings effectively from the complex urban environment. In this paper, we present a novel approach based on Random Finite Set Statistics for estimating the position of lane markings. We rely on Probability Hypothesis Density (PHD) filtering and apply this technique to lane marking extraction in urban environment. Our method is based on two phases: an image preprocessing phase to extract pixels that potentially represent lanes and a tracking phase to identify lane markings. Compared to other approaches, our method presents a recursive filtering algorithm which extracts lane markings in the presence of clutter and non-lane markings. The experimental results exhibit the high performance of the proposed approach under various scenarios. Feihu Zhang, Hauke Stähle, Chao Chen 0021, Christian Buckl, Alois C. Knoll |
Intelligent Vehicles Symposium | 1 |
| 2012 | Single camera visual odometry based on Random Finite Set StatisticsabstractThis paper presents a novel approach based on Random Finite Set (RFS) Statistics for estimating a vehicle's trajectory in complex urban environments by using a fixed single camera. For this, we extend our earlier works which used Probability Hypothesis Density (PHD) filtering under sensor fusion framework and are among the first to apply this technique to visual odometry in real traffic scenes. We consider features acquired from the camera as a group targets, use the PHD filter to update the overall group state and then estimate the ego-motion vector of the camera. Compared to other approaches, our approach presents a recursive filtering algorithm that provides dynamic estimation of multiple-targets states in the presence of clutter and avoids the association problem. Experimental results show that this method provides good robustness under real traffic scenarios. Feihu Zhang, Hauke Stähle, Andre Gaschler, Christian Buckl, Alois C. Knoll |
IROS | 1 |
| 2012 | Visual odometry based on Random Finite Set Statistics in urban environmentabstractThis paper presents a novel approach for estimating the vehicle's trajectory in complex urban environments. In previous work, we presented a visual odometry solution that estimates frame-to-frame motion from a single camera based on Random Finite Set (RFS) Statistics. This paper extends that work by combining the stereo cameras and gyroscope sensor. We are among the first to apply RFS statistics to visual odometry in real traffic scenes. The method is based on two phases: a preprocessing phase to extract features from the image and transform the coordinates from the image space to vehicle coordinates; a tracking phase to estimate the egomotion vector of the camera. We consider features as a group target and use the Probability Hypothesis Density (PHD) filter to update the overall group state as the motion vector. Compared to other approaches, our method presents a recursive filtering algorithm that provides dynamic estimation of multiple-targets states in the presence of clutter and high association uncertainty. The experimental results show that this method exhibits good robustness under various scenarios. Feihu Zhang, Guang Chen 0001, Hauke Stähle, Christian Buckl, Alois C. Knoll |
Intelligent Vehicles Symposium | 1 |