EDBT 2026 Demo / reviewers in the wild / expert
Xiaozhi Chen
dblp:150/3655
· DBLP profile ↗
29ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Deep Representation Learning for Event-Enhanced Visual Autonomous Perception: The eAP DatasetabstractRecent visual autonomous perception systems achieve remarkable performances with deep representation learning. However, they fail in scenarios with challenging illumination. While event cameras can mitigate this problem, there is a lack of a large-scale dataset to develop event-enhanced deep visual perception models in autonomous driving scenes. To address the gap, we present theeAP(event-enhancedAutonomousPerception) dataset, the largest dataset with event cameras for autonomous perception. We demonstrate howeAPcan facilitate the study of different autonomous perception tasks, including 3D vehicle detection and object time-to-contact (TTC) estimation, through deep representation learning. Based oneAP, we demonstrate the first successful use of events to improve a popular 3D vehicle detection network in challenging illumination scenarios.eAPalso enables a devoted study of the representation learning problem of object TTC estimation. We show how a geometry-aware representation learning framework leads to the best event-based object TTC estimation network that operates at 200 FPS. The dataset, code, and pre-trained models will be made publicly available for future research. Shichao Li 0002, Qing Lian, Peiliang Li 0001, Xiaozhi Chen, Yi Zhou 0010 |
IEEE Trans. Robotics | 5 |
| 2025 | Learning Better Representations for Crowded Pedestrians in Offboard LiDAR-Camera 3D Tracking-by-detectionabstractPerceiving pedestrians in highly crowded urban environments is a difficult long-tail problem for learning-based autonomous perception. Speeding up 3D ground truth generation for such challenging scenes is performance-critical yet very challenging. The difficulties include the sparsity of the captured pedestrian point cloud and a lack of suitable benchmarks for a specific system design study. To tackle the challenges, we first collect a new multi-view LiDAR-camera 3D multiple-object-tracking benchmark of highly crowded pedestrians for in-depth analysis. We then build an offboard auto-labeling system that reconstructs pedestrian trajectories from LiDAR point cloud and multi-view images. To improve the generalization power for crowded scenes and the performance for small objects, we propose to learn high-resolution representations that are density-aware and relationship-aware. Extensive experiments validate that our approach significantly improves the 3D pedestrian tracking performance towards higher auto-labeling efficiency. The code will be publicly available at this HTTP URL11https://github.com/Nicholasli1995/PCP-MV. Shichao Li 0002, Peiliang Li 0004, Qing Lian, Peng Yun, Xiaozhi Chen |
ICRA | 5 |
| 2024 | Adaptive Fusion of Single-View and Multi-View Depth for Autonomous DrivingabstractMulti- view depth estimation has achieved impressive performance over various benchmarks. However, almost all current multi-view systems rely on given ideal camera poses, which are unavailable in many real-world scenarios, such as autonomous driving. In this work, we propose a new robustness benchmark to evaluate the depth estimation system under various noisy pose settings. Surprisingly, we find current multi-view depth estimation methods or single-view and multi-view fusion methods will fail when given noisy pose settings. To address this challenge, we propose a single-view and multi-view fused depth estimation system, which adaptively integrates high-confident multi-view and single-view results for both robust and accurate depth es-timations. The adaptive fusion module performs fusion by dynamically selecting high-confidence regions between two branches based on a wrapping confidence map. Thus, the system tends to choose the more reliable branch when facing textureless scenes, inaccurate calibration, dynamic ob-jects, and other degradation or challenging conditions. Our method outperforms state-of-the-art multi-view and fusion methods under robustness testing. Furthermore, we achieve state-of-the-art performance on challenging benchmarks (KITTI and DDAD) when given accurate pose estimations. Project website: https://github.com/Junda24/Afnet/. Junda Cheng, Wei Yin 0006, Xiaozhi Chen, Xin Yang 0008 |
CVPR | 4 |
| 2024 | UC-NERF: Neural Radiance Field for Under-Calibrated Multi-View Cameras in Autonomous DrivingabstractMulti-camera setups find widespread use across various applications, such as autonomous driving, as they greatly expand sensing capabilities.
Despite the fast development of Neural radiance field (NeRF) techniques and their wide applications in both indoor and outdoor scenes, applying NeRF to multi-camera systems remains very challenging. This is primarily due to the inherent under-calibration issues in multi-camera setup, including inconsistent imaging effects stemming from separately calibrated image signal processing units in diverse cameras, and system errors arising from mechanical vibrations during driving that affect relative camera poses.
In this paper, we present UC-NeRF, a novel method tailored for novel view synthesis in under-calibrated multi-view camera systems.
Firstly, we propose a layer-based color correction to rectify the color inconsistency in different image regions. Second, we propose virtual warping to generate more viewpoint-diverse but color-consistent virtual views for color correction and 3D recovery. Finally, a spatiotemporally constrained pose refinement is designed for more robust and accurate pose calibration in multi-camera systems.
Our method not only achieves state-of-the-art performance of novel view synthesis in multi-camera setups, but also effectively facilitates depth estimation in large-scale outdoor scenes with the synthesized novel views. Xiaoxiao Long, Wei Yin 0006, Jin Wang 0001, Zhiqiang Wu 0001, Yuexin Ma, Xiaozhi Chen, Xuejin Chen |
ICLR | 8 |
| 2024 | GIM: Learning Generalizable Image Matcher From Internet VideosabstractImage matching is a fundamental computer vision problem. While learning-based methods achieve state-of-the-art performance on existing benchmarks, they generalize poorly to in-the-wild images. Such methods typically need to train separate models for different scene types (e.g., indoor vs. outdoor) and are impractical when the scene type is unknown in advance. One of the underlying problems is the limited scalability of existing data construction pipelines, which limits the diversity of standard image matching datasets. To address this problem, we propose GIM, a self-training framework for learning a single generalizable model based on any image matching architecture using internet videos, an abundant and diverse data source. Given an architecture, GIM first trains it on standard domain-specific datasets and then combines it with complementary matching methods to create dense labels on nearby frames of novel videos. These labels are filtered by robust fitting, and then enhanced by propagating them to distant frames. The final model is trained on propagated data with strong augmentations. Not relying on complex 3D reconstruction makes GIM much more efficient and less likely to fail than standard SfM-and-MVS based frameworks. We also propose ZEB, the first zero-shot evaluation benchmark for image matching. By mixing data from diverse domains, ZEB can thoroughly assess the cross-domain generalization performance of different methods. Experiments demonstrate the effectiveness and generality of GIM. Applying GIM consistently improves the zero-shot performance of 3 state-of-the-art image matching architectures as the number of downloaded videos increases (Fig. 1 (a)); with 50 hours of YouTube videos, the relative zero-shot performance improves by 6.9% − 18.1%. GIM also enables generalization to extreme cross-domain data such as Bird Eye View (BEV) images of projected 3D point clouds (Fig. 1 (c)). More importantly, our single zero-shot model consistently outperforms domain-specific baselines when evaluated on downstream tasks inherent to their respective domains. The code will be released upon acceptance. Xuelun Shen, Zhipeng Cai 0003, Wei Yin 0006, Matthias Müller 0011, Zijun Li 0006, Xiaozhi Chen, Cheng Wang 0003 |
ICLR | 7 |
| 2023 | Learning to Fuse Monocular and Multi-view Cues for Multi-frame Depth Estimation in Dynamic ScenesabstractMulti-frame depth estimation generally achieves high accuracy relying on the multi-view geometric consistency. When applied in dynamic scenes, e.g., autonomous driving, this consistency is usually violated in the dynamic areas, leading to corrupted estimations. Many multi-frame methods handle dynamic areas by identifying them with explicit masks and compensating the multi-view cues with monocular cues represented as local monocular depth or features. The improvements are limited due to the uncontrolled quality of the masks and the underutilized benefits of the fusion of the two types of cues. In this paper, we propose a novel method to learn to fuse the multi-view and monocular cues encoded as volumes without needing the heuristically crafted masks. As unveiled in our analyses, the multiview cues capture more accurate geometric information in static areas, and the monocular cues capture more useful contexts in dynamic areas. To let the geometric perception learned from multi-view cues in static areas propagate to the monocular representation in dynamic areas and let monocular cues enhance the representation of multi-view cost volume, we propose a cross-cue fusion (CCF) module, which includes the cross-cue attention (CCA) to encode the spatially non-local relative intra-relations from each source to enhance the representation of the other. Experiments on real-world datasets prove the significant effectiveness and generalization ability of the proposed method. Rui Li 0013, Dong Gong, Wei Yin 0006, Hao Chen 0041, Yu Zhu 0004, Xiaozhi Chen, Jinqiu Sun, Yanning Zhang 0001 |
CVPR | 7 |
| 2023 | Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageabstractReconstructing accurate 3D scenes from images is a long-standing vision task. Due to the ill-posedness of the single-image reconstruction problem, most well-established methods are built upon multi-view geometry. State-of-the-art (SOTA) monocular metric depth estimation methods can only handle a single camera model and are unable to perform mixed-data training due to metric ambiguity. Meanwhile, SOTA monocular methods trained on large mixed datasets achieve zero-shot generalization by learning affine-invariant depths, which cannot recover real-world metrics. In this work, we show that the key to a zero-shot single-view metric depth model lies in the combination of large-scale data training and resolving the metric ambiguity from various camera models. We propose a canonical camera space transformation module, which explicitly addresses the ambiguity problems and can be effortlessly plugged into existing monocular models. Equipped with our module, monocular models can be stably trained over 8 millions of images with thousands of camera models, resulting in zero-shot generalization to in-the-wild images with unseen camera settings. Experiments demonstrate SOTA performance of our method on 7 zero-shot benchmarks. Notably, our method won the championship in the 2nd Monocular Depth Estimation Challenge. Our method enables the accurate recovery of metric 3D structures on randomly collected internet images, paving the way for plausible single-image metrology. The potential benefits extend to downstream tasks, which can be significantly improved by simply plugging in our model. For example, our model relieves the scale drift issues of monocular-SLAM (Fig. 1), leading to high-quality metric scale dense mapping. The code is available at https://github.com/YvanYin/Metric3D. Wei Yin 0006, Chi Zhang 0007, Hao Chen 0041, Zhipeng Cai 0003, Gang Yu 0002, Xiaozhi Chen, Chunhua Shen |
ICCV | 7 |
| 2023 | UniFusion: Unified Multi-view Fusion Transformer for Spatial-Temporal Representation in Bird's-Eye-ViewabstractBird’s eye view (BEV) representation is a new perception formulation for autonomous driving, which is based on spatial fusion. Further, temporal fusion is also introduced in BEV representation and gains great success. In this work, we propose a new method that unifies both spatial and temporal fusion and merges them into a unified mathematical formulation. The unified fusion could not only provide a new perspective on BEV fusion but also brings new capabilities. With the proposed unified spatial-temporal fusion, our method could support long-range fusion, which is hard to achieve in conventional BEV methods. Moreover, the BEV fusion in our work is temporal-adaptive and the weights of temporal fusion are learnable. In contrast, conventional methods mainly use fixed and equal weights for temporal fusion. Besides, the proposed unified fusion could avoid information lost in conventional BEV fusion methods and make full use of features. Extensive experiments and ablation studies on the NuScenes dataset show the effectiveness of the proposed method and our method gains the state-of-the-art performance in the map and vehicle segmentation task. Zequn Qin, Xiaozhi Chen, Xi Li 0001 |
ICCV | 4 |
| 2023 | Are All Point Clouds Suitable for Completion? Weakly Supervised Quality Evaluation Network for Point Cloud CompletionabstractIn the practical application of point cloud completion tasks, real data quality is usually much worse than the CAD datasets used for training. A small amount of noisy data will usually significantly impact the overall system's accuracy. In this paper, we propose a quality evaluation network to score the point clouds and help judge the quality of the point cloud before applying the completion model. We believe our scoring method can help researchers select more appropriate point clouds for subsequent completion and reconstruction and avoid manual parameter adjustment. Moreover, our evaluation model is fast and straightforward and can be directly inserted into any model's training or use process to facilitate the automatic selection and post-processing of point clouds. We propose a complete dataset construction and model evaluation method based on ShapeNet. We verify our network using detection and flow estimation tasks on KITTI, a real-world dataset for autonomous driving. The experimental results show that our model can effectively distinguish the quality of point clouds and help in practical tasks. Jieqi Shi, Peiliang Li 0001, Xiaozhi Chen, Shaojie Shen |
ICRA | 3 |
| 2022 | MonoJSG: Joint Semantic and Geometric Cost Volume for Monocular 3D Object DetectionabstractDue to the inherent ill-posed nature of 2D-3D projection, monocular 3D object detection lacks accurate depth recovery ability. Although the deep neural network (DNN) enables monocular depth-sensing from high-level learned features, the pixel-level cues are usually omitted due to the deep convolution mechanism. To benefit from both the pow-erful feature representation in DNN and pixel-level geomet-ric constraints, we reformulate the monocular object depth estimation as a progressive refinement problem and propose a joint semantic and geometric cost volume to model the depth error. Specifically, we first leverage neural networks to learn the object position, dimension, and dense normal-ized 3D object coordinates. Based on the object depth, the dense coordinates patch together with the corresponding object features is reprojected to the image space to build a cost volume in a joint semantic and geometric error man-ner. The final depth is obtained by feeding the cost volume to a refinement network, where the distribution of semantic and geometric error is regularized by direct depth supervision. Through effectively mitigating depth error by the re-finement framework, we achieve state-of-the-art results on both the KITTI and Waymo datasets.11Code available at https://github.com/lianqingll/MonoJSG Qing Lian, Peiliang Li 0001, Xiaozhi Chen |
CVPR | 3 |
| 2022 | PUA-MOS: End-to-End Point-wise Uncertainty Weighted Aggregation for Moving Object SegmentationabstractSegmenting moving objects in the 3D LiDAR point cloud can provide important guidance to localization, mapping and decision-making for self-driving vehicles. As for the conventional approaches to point cloud segmentation, they rely on semantic-level information, which makes it inevitable for long-tail problems to arise as there are always unseen types of objects on the road. To achieve moving segmentation while avoiding the reliance on the object category, the point motion is identified in this paper by fully exploring and aggregating the point-level geometric consistency in sequential point clouds. More specifically, an end-to-end point-wise uncertainty weighted aggregation approach known as PUA-MOS is proposed to segment the moving points in 3D LiDAR Data. Our method is applicable to estimate point-wise moving mask, scene flow and rigid-body transformation simultaneously in a coarse- to-fine network, where the relations between each prediction are implicitly learned. To explicitly model the inner and inter relations across these predictions among all points, the point- wise estimation and the average value of the same motion points are aggregated according to a predicted uncertainty. Then, the aggregated estimation is fed again into the next-level fusion, where the points will be re-segmented using the aggregated mask from the last level. Through iterative joint aggregation, our PUA-MOS outperforms the previous methods significantly on both KITTI [4] and Waymo [26] datasets. The code will be provided to generate the moving segmentation labels on both datasets for reproduction. Peiliang Li 0001, Xiaozhi Chen, Xin Yang 0008 |
IROS | 3 |
| 2022 | Multi-Camera Collaborative Depth Prediction via Consistent Structure EstimationabstractDepth map estimation from images is an important task in robotic systems. Existing methods can be categorized into two groups including multi-view stereo and monocular depth estimation. The former requires cameras to have large overlapping areas and sufficient baseline between cameras, while the latter that processes each image independently can hardly guarantee the structure consistency between cameras. In this paper, we propose a novel multi-camera collaborative depth prediction method that does not require large overlapping areas while maintaining structure consistency between cameras. Specifically, we formulate the depth estimation as a weighted combination of depth basis, in which the weights are updated iteratively by a refinement network driven by the proposed consistency loss. During the iterative update, the results of depth estimation are compared across cameras and the information of overlapping areas is propagated to the whole depth maps with the help of basis formulation. Experimental results on DDAD and NuScenes datasets demonstrate the superior performance of our method. Jialei Xu, Xianming Liu 0005, Yuanchao Bai, Junjun Jiang, Xiaozhi Chen, Xiangyang Ji |
ACM Multimedia | 6 |
| 2021 | Geometry-based Distance Decomposition for Monocular 3D Object DetectionabstractMonocular 3D object detection is of great significance for autonomous driving but remains challenging. The core challenge is to predict the distance of objects in the absence of explicit depth information. Unlike regressing the distance as a single variable in most existing methods, we propose a novel geometry-based distance decomposition to recover the distance by its factors. The decomposition factors the distance of objects into the most representative and stable variables, i.e. the physical height and the projected visual height in the image plane. Moreover, the decomposition maintains the self-consistency between the two heights, leading to robust distance prediction when both predicted heights are inaccurate. The decomposition also enables us to trace the causes of the distance uncertainty for different scenarios. Such decomposition makes the distance prediction interpretable, accurate, and robust. Our method directly predicts 3D bounding boxes from RGB images with a compact architecture, making the training and inference simple and efficient. The experimental results show that our method achieves the state-of-the-art performance on the monocular 3D Object Detection and Bird’s Eye View tasks of the KITTI dataset, and can generalize to images with different camera intrinsics1. Xuepeng Shi, Qi Ye 0001, Xiaozhi Chen, Chuangrong Chen, Zhixiang Chen 0003, Tae-Kyun Kim 0001 |
ICCV | 3 |
| 2019 | Stereo R-CNN Based 3D Object Detection for Autonomous DrivingabstractWe propose a 3D object detection method for autonomous driving by fully exploiting the sparse and dense, semantic and geometry information in stereo imagery. Our method, called Stereo R-CNN, extends Faster R-CNN for stereo inputs to simultaneously detect and associate object in left and right images. We add extra branches after stereo Region Proposal Network (RPN) to predict sparse keypoints, viewpoints, and object dimensions, which are combined with 2D left-right boxes to calculate a coarse 3D object bounding box. We then recover the accurate 3D bounding box by a region-based photometric alignment using left and right RoIs. Our method does not require depth input and 3D position supervision, however, outperforms all existing fully supervised image-based methods. Experiments on the challenging KITTI dataset show that our method outperforms the state-of-the-art stereo-based method by around 30% AP on both 3D detection and 3D localization tasks. Code will be made publicly available. Peiliang Li 0001, Xiaozhi Chen, Shaojie Shen |
CVPR | 2 |
| 2019 | On the Over-Smoothing Problem of CNN Based Disparity EstimationabstractCurrently, most deep learning based disparity estimation methods have the problem of over-smoothing at boundaries, which is unfavorable for some applications such as point cloud segmentation, mapping, etc. To address this problem, we first analyze the potential causes and observe that the estimated disparity at edge boundary pixels usually follows multimodal distributions, causing over-smoothing estimation. Based on this observation, we propose a single-modal weighted average operation on the probability distribution during inference, which can alleviate the problem effectively. To integrate the constraint of this inference method into training stage, we further analyze the characteristics of different loss functions and found that using cross entropy with gaussian distribution consistently further improves the performance. For quantitative evaluation, we propose a novel metric that measures the disparity error in the local structure of edge boundaries. Experiments on various datasets using various networks show our method's effectiveness and general applicability. Code will be available at https://github.com/chenchr/otosp. Chuangrong Chen, Xiaozhi Chen |
ICCV | 2 |
| 2018 | 3D Object Proposals Using Stereo Imagery for Accurate Object Class DetectionabstractThe goal of this paper is to perform 3D object detection in the context of autonomous driving. Our method aims at generating a set of high-quality 3D object proposals by exploiting stereo imagery. We formulate the problem as minimizing an energy function that encodes object size priors, placement of objects on the ground plane as well as several depth informed features that reason about free space, point cloud densities and distance to the ground. We then exploit a CNN on top of these proposals to perform object detection. In particular, we employ a convolutional neural net (CNN) that exploits context and depth information to jointly regress to 3D bounding box coordinates and object pose. Our experiments show significant performance gains over existing RGB and RGB-D object proposal methods on the challenging KITTI benchmark. When combined with the CNN, our approach outperforms all existing results in object detection and orientation estimation tasks for all three KITTI object classes. Furthermore, we experiment also with the setting where LIDAR information is available, and show that using both LIDAR and stereo leads to the best result. Xiaozhi Chen, Kaustav Kundu, Yukun Zhu, Huimin Ma 0001, Sanja Fidler, Raquel Urtasun |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Edge Preserving and Multi-Scale Contextual Neural Network for Salient Object DetectionabstractIn this paper, we propose a novel edge preserving and multi-scale contextual neural network for salient object detection. The proposed framework is aiming to address two limits of the existing CNN based methods. First, region-based CNN methods lack sufficient context to accurately locate salient object since they deal with each region independently. Second, pixel-based CNN methods suffer from blurry boundaries due to the presence of convolutional and pooling layers. Motivated by these, we first propose an end-to-end edge-preserved neural network based on Fast R-CNN framework (named RegionNet) to efficiently generate saliency map with sharp object boundaries. Later, to further improve it, multi-scale spatial context is attached to RegionNet to consider the relationship between regions and the global scenes. Furthermore, our method can be generally applied to RGB-D saliency detection by depth refinement. The proposed framework achieves both clear detection boundary and multi-scale contextual robustness simultaneously for the first time, and thus achieves an optimized performance. Experiments on six RGB and two RGB-D benchmark datasets demonstrate that the proposed method achieves state-of-the-art performance. Xiang Wang 0003, Huimin Ma 0001, Xiaozhi Chen, Shaodi You |
IEEE Trans. Image Process. | 3 |
| 2017 | Multi-view 3D Object Detection Network for Autonomous DrivingabstractThis paper aims at high-accuracy 3D object detection in autonomous driving scenario. We propose Multi-View 3D networks (MV3D), a sensory-fusion framework that takes both LIDAR point cloud and RGB images as input and predicts oriented 3D bounding boxes. We encode the sparse 3D point cloud with a compact multi-view representation. The network is composed of two subnetworks: one for 3D object proposal generation and another for multi-view feature fusion. The proposal network generates 3D candidate boxes efficiently from the birds eye view representation of 3D point cloud. We design a deep fusion scheme to combine region-wise features from multiple views and enable interactions between intermediate layers of different paths. Experiments on the challenging KITTI benchmark show that our approach outperforms the state-of-the-art by around 25% and 30% AP on the tasks of 3D localization and 3D detection. In addition, for 2D detection, our approach obtains 14.9% higher AP than the state-of-the-art on the hard data among the LIDAR-based methods. Xiaozhi Chen, Huimin Ma 0001, Ji Wan, Bo Li 0018 |
CVPR | 1 |
| 2017 | Boundary-aware box refinement for object proposal generation
Xiaozhi Chen, Huimin Ma 0001, Chenzhuo Zhu, Xiang Wang 0003, Zhichen Zhao |
Neurocomputing | 1 |
| 2017 | Generalized symmetric pair model for action classification in still images
Zhichen Zhao, Huimin Ma 0001, Xiaozhi Chen |
Pattern Recognit. | 3 |
| 2016 | Monocular 3D Object Detection for Autonomous DrivingabstractThe goal of this paper is to perform 3D object detection from a single monocular image in the domain of autonomous driving. Our method first aims to generate a set of candidate class-specific object proposals, which are then run through a standard CNN pipeline to obtain high-quality object detections. The focus of this paper is on proposal generation. In particular, we propose an energy minimization approach that places object candidates in 3D using the fact that objects should be on the ground-plane. We then score each candidate box projected to the image plane via several intuitive potentials encoding semantic segmentation, contextual information, size and location priors and typical object shape. Our experimental evaluation demonstrates that our object proposal generation approach significantly outperforms all monocular approaches, and achieves the best detection performance on the challenging KITTI benchmark, among published monocular competitors. Xiaozhi Chen, Kaustav Kundu, Huimin Ma 0001, Sanja Fidler, Raquel Urtasun |
CVPR | 1 |
| 2016 | Salient object detection via fast R-CNN and low-level cuesabstractRecent advances in salient object detection have exploited the deep Convolutional Neural Network (CNN) to represent high-level semantic, however, due to the presence of convolutional and pooling layers, it is difficult for CNN to generate saliency map with sharp boundaries. In this paper, we propose multi-scale mask-based Fast R-CNN framework which generate saliency score of each region. Since the regions are segmented using edge-preserved methods, the results are naturally with sharp boundaries. To consider context information, we also propose low-level contrast and backgroundness prior which are complementary with high-level semantic. Finally, an edge-based propagation method which takes advantages of edge information is proposed to refine the saliency map. Experiments on three benchmark datasets demonstrate that the proposed method outperforms previous methods and achieves state-of-the-art performance. Xiang Wang 0003, Huimin Ma 0001, Xiaozhi Chen |
ICIP | 3 |
| 2016 | Multi-scale region candidate combination for action recognitionabstractIn still images, multi-scale regions contain rich information of different granularity. However, only semantically meaningful regions provide auxiliary cues for action recognition. Moreover, regions at different scales contribute differently. Motivated by the two observations, we propose an approach that is composed of three components: 1) detecting semantic region candidates at multiple scales, 2) training networks at each scale, 3) extracting features and learning to fuse them. The proposed approach captures multi-scale cues and highlights the optimal scale for each action, Experimental results show that our approach reaches the state-of-the-art performance on two challenging benchmarks: 1) PASCAL VOC 2012 and 2) Stanford-40. Zhichen Zhao, Huimin Ma 0001, Xiaozhi Chen |
ICIP | 3 |
| 2016 | Geodesic weighted Bayesian model for saliency optimization
Xiang Wang 0003, Huimin Ma 0001, Xiaozhi Chen |
Pattern Recognit. Lett. | 3 |
| 2016 | Semantic parts based top-down pyramid for action recognition
Zhichen Zhao, Huimin Ma 0001, Xiaozhi Chen |
Pattern Recognit. Lett. | 3 |
| 2015 | Improving object proposals with multi-thresholding straddling expansionabstractRecent advances in object detection have exploited object proposals to speed up object searching. However, many of existing object proposal generators have strong localization bias or require computationally expensive diversification strategies. In this paper, we present an effective approach to address these issues. We first propose a simple and useful localization bias measure, called superpixel tightness. Based on the characteristics of superpixel tightness distribution, we propose an effective method, namely multi-thresholding straddling expansion (MTSE) to reduce localization bias via fast diversification. Our method is essentially a box refinement process, which is intuitive and beneficial, but seldom exploited before. The greatest benefit of our method is that it can be integrated into any existing model to achieve consistently high recall across various intersection over union thresholds. Experiments on PASCAL VOC dataset demonstrates that our approach improves numerous existing models significantly with little computational overhead. Xiaozhi Chen, Huimin Ma 0001, Xiang Wang 0003, Zhichen Zhao |
CVPR | 1 |
| 2015 | Geodesic weighted Bayesian model for salient object detectionabstractIn recent years, a variety of salient object detection methods under Bayesian framework have been proposed and many achieved state of the art. However, those ignore spatial relationships and thus background regions similar to the objects are also highlighted. In this paper, we propose a novel geodesic weighted Bayesian model to address this issue. We consider spatial relationships by attaching more importance to regions which are more likely to be parts of a salient object, thus suppressing background regions. First, we learn a combined similarity via multiple features to measure similarity of adjacent regions. Then, we apply the combined similarity as edge weight to construct an undirected weighted graph and compute geodesic distance. Last, we utilize the geodesic distance to weight the observation likelihood to infer a more precise saliency map. Experiments on several benchmark datasets demonstrate the effectiveness of our model. Xiang Wang 0003, Huimin Ma 0001, Xiaozhi Chen |
ICIP | 3 |
| 2015 | 3D Object Proposals for Accurate Object Class DetectionabstractThe goal of this paper is to generate high-quality 3D object proposals in the context of autonomous driving. Our method exploits stereo imagery to place proposals in the form of 3D bounding boxes. We formulate the problem as minimizing an energy function encoding object size priors, ground plane as well as several depth informed features that reason about free space, point cloud densities and distance to the ground. Our experiments show significant performance gains over existing RGB and RGB-D object proposal methods on the challenging KITTI benchmark. Combined with convolutional neural net (CNN) scoring, our approach outperforms all existing results on all three KITTI object classes. Xiaozhi Chen, Kaustav Kundu, Yukun Zhu, Andrew G. Berneshawi, Huimin Ma 0001, Sanja Fidler, Raquel Urtasun |
NIPS | 1 |
| 2014 | Learning a compact latent representation of the Bag-of-Parts modelabstractThe Bag-of-Parts (BoP) model, which employs distinctive parts to represent images, has shown superior performance in vision recognition tasks. Our work is motivated by the need of reducing redundancy in tens of thousands parts. We propose a novel method to learn a compact latent representation from redundant part responses. We address this problem by employing spectral clustering and a multi-column coding scheme. The BoP model is viewed as a multi-scale convolutional model and additional sparse autoencoders are used to infer the latent patterns embedded in high-dimensional part-based representations. Spatial and semantic information is preserved by sparse learning on multiple spatial regions individually. Experiments demonstrate that the learnt representation achieves competitive performance with state-of-the-art methods on PASCAL VOC 2007 dataset. Xiaozhi Chen, Huimin Ma 0001 |
ICIP | 1 |