EDBT 2026 Demo / reviewers in the wild / expert
Chen Feng 0002
dblp:01/161-2
· DBLP profile ↗
66ranked-venue papers
5as first author
42since 2021 · last 2025
0000-0003-3211-1576ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 1 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 2 first-author · 15 since 2021Systems, architecture and hardware · 23 · 1 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CityWalker: Learning Embodied Urban Navigation from Web-Scale VideosabstractNavigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progress, existing visual navigation methods struggle in map-free or off-street settings, limiting the deployment of autonomous agents like last-mile delivery robots. To overcome these obstacles, we propose a scalable, data-driven approach for human-like urban navigation by training agents on thousands of hours of in-the-wild city walking and driving videos sourced from the web. We introduce a simple and scalable data processing pipeline that extracts action supervision from these videos, enabling large-scale imitation learning without costly annotations. Our model learns sophisticated navigation policies to handle diverse challenges and critical scenarios. Experimental results show that training on large-scale, diverse datasets significantly enhances navigation performance, surpassing current methods. This work shows the potential of using abundant online video data to develop robust navigation policies for embodied agents in dynamic urban settings. Xinhao Liu 0003, Jintong Li, Niranjan Sujay, Juexiao Zhang, John Abanes, Chen Feng 0002 |
CVPR | 9 |
| 2025 | Extrapolated Urban View Synthesis BenchmarkabstractPhotorealistic simulators are essential for the training and evaluation of vision-centric autonomous vehicles (AVs). At their core is Novel View Synthesis (NVS), a crucial capability that generates diverse unseen viewpoints to accommodate the broad and continuous pose distribution of AVs. Recent advances in radiance fields, such as 3D Gaussian Splatting, achieve photorealistic rendering at real-time speeds and have been widely used in modeling large-scale driving scenes. However, their performance is commonly evaluated using an interpolated setup with highly correlated training and test views. In contrast, extrapolation, where test views largely deviate from training views, remains underexplored, limiting progress in generalizable simulation technology. To address this gap, we leverage publicly available AV datasets with multiple traversals, multiple vehicles, and multiple cameras to build the first Extrapolated Urban View Synthesis (EUVS) benchmark. Meanwhile, we conduct both quantitative and qualitative evaluations of state-of-the-art NVS methods across different evaluation settings. Our results show that current NVS methods are prone to overfitting to training views. Besides, incorporating diffusion priors and improving geometry cannot fundamentally improve NVS under large view changes, highlighting the need for more robust approaches and large-scale training. We will release the data to help advance self-driving and urban robotics simulation technology. Xiangyu Han, Boyi Li 0001, Yan Wang 0051, Boris Ivanovic, Yurong You, Lingjie Liu, Yue Wang 0041, Marco Pavone 0001, Chen Feng 0002, Yiming Li 0003 |
ICCV | 10 |
| 2025 | GARF: Learning Generalizable 3D Reassembly for Real-World Fracturesabstract3D reassembly is a challenging spatial intelligence task with broad applications across scientific domains. While large-scale synthetic datasets have fueled promising learning-based approaches, their generalizability to different domains is limited. Critically, it remains uncertain whether models trained on synthetic datasets can generalize to real-world fractures where breakage patterns are more complex. To bridge this gap, we propose GARF, a generalizable 3D reassembly framework for real-world fractures. GARF leverages fracture-aware pretraining to learn fracture features from individual fragments, with flow matching enabling precise 6-DoF alignments. At inference time, we introduce one-step preassembly, improving robustness to unseen objects and varying numbers of fractures. In collaboration with archaeologists, paleoanthropologists, and ornithologists, we curate Fractura, a diverse dataset for vision and learning communities, featuring real-world fracture types across ceramics, bones, eggshells, and lithics. Comprehensive experiments have shown our approach consistently outperforms state-of-the-art methods on both synthetic and real-world datasets, achieving 82.87\% lower rotation error and 25.15\% higher part accuracy. This sheds light on training on synthetic data to advance real-world 3D puzzle solving, demonstrating its strong generalization across unseen object shapes and diverse fracture types. GARF's code, data and demo are available at https://ai4ce.github.io/GARF/. Sihang Li 0001, Grace Chen, Siqi Tan, Irving Fang, Kristof Zyskowski, Shannon P. McPherron, Radu Iovita, Chen Feng 0002 |
ICCV | 11 |
| 2025 | Adversarial Exploitation of Data Diversity Improves Visual Localization
Sihang Li 0001, Siqi Tan, Bowen Chang, Chen Feng 0002, Yiming Li 0003 |
ICCV | 5 |
| 2025 | Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View ReconstructionabstractHumans effortlessly integrate common-sense knowledge with sensory input from vision and touch to understand their surroundings. Emulating this capability, we introduce FusionSense, a novel 3D reconstruction framework that enables robots to fuse priors from foundation models with highly sparse observations from vision and tactile sensors. FusionSense addresses three key challenges: (i) How can robots efficiently acquire robust global shape information about the surrounding scene and objects? (ii) How can robots strategically select touch points on the object using geometric and commonsense priors? (iii) How can partial observations such as tactile signals improve the overall representation of the object? Our framework employs 3D Gaussian Splatting as a core representation and incorporates a hierarchical optimization strategy involving global structure construction, object visual hull pruning and local geometric constraints. This advancement results in fast and robust perception in environments with traditionally challenging objects that are transparent, reflective, or dark, enabling more downstream manipulation or navigation tasks. Experiments on real-world data suggest that our framework outperforms previously state-of-the-art sparse-view methods. All code and data are open-sourced on the project website. Irving Fang, Kairui Shi, Xujin He, Siqi Tan, Hanwen Zhao, Hung-Jui Huang, Wenzhen Yuan 0001, Chen Feng 0002 |
ICRA | 9 |
| 2025 | NYC-Event-VPR: A Large-Scale High-Resolution Event-Based Visual Place Recognition Dataset in Dense Urban EnvironmentsabstractVisual place recognition (VPR) enables autonomous robots to identify previously visited locations, which contributes to tasks like simultaneous localization and mapping (SLAM). VPR faces challenges such as accurate image neighbor retrieval and appearance change in scenery. Event cameras, also known as dynamic vision sensors, are a new sensor modality for VPR and offer a promising solution to the challenges with their unique attributes: high temporal resolution (1MHz clock), ultra-low latency (in μs), and high dynamic range (>120dB). These attributes make event cameras less susceptible to motion blur and more robust in variable lighting conditions, making them suitable for addressing VPR challenges. However, the scarcity of event-based VPR datasets, partly due to the novelty and cost of event cameras, hampers their adoption. To fill this data gap, our paper introduces the NYC-Event-VPR dataset to the robotics and computer vision communities, featuring the Prophesee IMX636 HD event sensor (1280x720 resolution), combined with RGB camera and GPS module. It encompasses over 13 hours of geotagged event data, spanning 260 kilometers across New York City, covering diverse lighting and weather conditions, day/night scenarios, and multiple visits to various locations. Furthermore, our paper employs three frameworks to conduct generalization performance assessments, promoting innovation in event-based VPR and its integration into robotics applications. Taiyi Pan, Junyang He, Yiming Li 0003, Chen Feng 0002 |
ICRA | 5 |
| 2025 | VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language ModelabstractLarge Vision Language Models (VLMs) have been adopted in robotics for their strong common sense understanding and generalization capabilities. Existing works leverage VLMs for task and motion planning based on language instructions and robot observations. In this work, we explore using VLM to interpret long-horizon human demonstration videos to generate a sequence of robot task plans in natural language. To achieve this, we propose SeeDo, an agent that integrates keyframe selection module, visual prompting module, and a VLM interpreter into a pipeline that enables the VLM to "see" human demonstrations and generate step-by-step plans for robots to "do" them. To evaluate, we curate a benchmark of long-horizon human demonstration videos of pick-and-place tasks in three diverse categories and designed comprehensive evaluation metrics. The experiments demonstrate SeeDo’s superior performance in generating subtask planning in natural language from long-horizon human demo videos. Experiments show SeeDo outperforms state-of-the-art video VLMs in generating subtask plans. By further integrating SeeDo with low-level action primitive functions and language model programs, we validated SeeDo in both simulated and real-world deployments. The code, demos, prompts and data can be found at ai4ce.github.io/SeeDo. Beichen Wang, Juexiao Zhang, Shuwen Dong, Irving Fang, Chen Feng 0002 |
IROS | 5 |
| 2025 | One-Shot Cross-Domain Instance Detection With Universal Representation
Chen Feng 0002, Jian Cheng 0001, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | CODNet: Infrared Small-Target Detection by Mitigating the Curse of DimensionalityabstractInfrared small-target detection (ISTD) presents significant challenges due to the blurred edges and low signal-to-noise ratio (SNR) of the targets, which often results in the loss of critical target features during convolution and downsampling. While researchers have explored complex architectures to preserve these features, such designs often come at the cost of increased computational cost. Our analysis reveals that the curse of dimensionality (COD) is a fundamental issue in ISTD. Due to the sparse information content of infrared small targets, they struggle to effectively utilize high-dimensional feature spaces, leading to a significant feature loss phenomenon. To address this, we propose the CODNet, which integrates spatial–channel shuffle (SCS) attention, dynamic gated channel attention (DGCA), and DCT-based frequency attention (DFA) in the UNet framework. SCS enhances feature representation through spatial attention and channel shuffling, making the feature space more compact and informative. DGCA performs selective channel compression to improve the feature SNR while reducing redundancy and sparsity. DFA incorporates frequency-domain features to supplement spatial representations, alleviating feature sparsity and improving information utilization. Experimental results demonstrate that the CODNet achieves state-of-the-art performance on four public datasets while maintaining relatively high operation efficiency. Shuaiyuan Du, Chen Feng 0002, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Multiagent Multitraversal Multimodal Self-Driving: Open MARS DatasetabstractLarge-scale datasets have fueled recent advancements in AI-based autonomous vehicle research. However, these datasets are usually collected from a single vehicle's one-time pass of a certain location, lacking multiagent inter-actions or repeated traversals of the same place. Such information could lead to transformative enhancements in autonomous vehicles' perception, prediction, and planning capabilities. To bridge this gap, in collaboration with the self-driving company May Mobility, we present MARS dataset which unifies scenarios that enable MultiAgent, multitraveRSal, and multimodal autonomous vehicle re-search. More specifically, MARS is collected with a fleet of autonomous vehicles driving within a certain geographical area. Each vehicle has its own route and different vehicles may appear at nearby locations. Each vehicle is equipped with a LiDAR and surround-view RGB cameras. We curate two subsets in MARS: one facilitates collaborative driving with multiple vehicles simultaneously present at the same location, and the other enables memory retrospection through asynchronous traversals of the same location by multiple vehicles. We conduct experiments in place recog-nition and neural reconstruction. More importantly, MARS introduces new research opportunities and challenges such as multitraversal 3D reconstruction, multiagent perception, and unsupervised object discovery. Our data and codes can be found at https://ai4ce.github.io/MARS/. Yiming Li 0003, Nuo Chen 0003, Moonjun Gong, Zonglin Lyu, Zehong Wang, Peili Jiang, Chen Feng 0002 |
CVPR | 8 |
| 2024 | LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic ImagesabstractLithic Use-Wear Analysis (LUWA) using microscopic images is an underexplored vision-for-science research area. It seeks to distinguish the worked material, which is critical for understanding archaeological artifacts, material interactions, tool functionalities, and dental records. However, this challenging task goes beyond the well-studied image classification problem for common objects. It is affected by many confounders owing to the complex wear mechanism and microscopic imaging, which makes it difficult even for human experts to identify the worked material successfully. In this paper, we investigate the following three questions on this unique vision task for the first time:(i) How well can state-of-the-art pre-trained models (like DINOv2) generalize to the rarely seen domain? (ii) How can few-shot learning be exploited for scarce microscopic images? (iii) How do the ambiguous magnification and sensing modality influence the classification accuracy? To study these, we collaborated with archaeologists and built the first open-source and the largest LUWA dataset containing 23,130 microscopic images with different magnifications and sensing modalities. Extensive experiments show that existing pretrained models notably outperform human experts but still leave a large gap for improvements. Most importantly, the LUWA dataset provides an underexplored opportunity for vision and learning communities and complements existing image classification problems on common objects. Irving Fang, Akshat Kaushik, Alice Rodriguez, Hanwen Zhao, Juexiao Zhang, Zhuo Zheng, Radu Iovita, Chen Feng 0002 |
CVPR | 10 |
| 2024 | EgoPAT3Dv2: Predicting 3D Action Target from 2D Egocentric Vision for Human-Robot InteractionabstractA robot’s ability to anticipate the 3D action target location of a hand’s movement from egocentric videos can greatly improve safety and efficiency in human-robot interaction (HRI). While previous research predominantly focused on semantic action classification or 2D target region prediction, we argue that predicting the action target’s 3D coordinate could pave the way for more versatile downstream robotics tasks, especially given the increasing prevalence of headset devices. This study expands EgoPAT3D, the sole dataset dedicated to egocentric 3D action target prediction. We augment both its size and diversity, enhancing its potential for generalization. Moreover, we substantially enhance the baseline algorithm by introducing a large pre-trained model and human prior knowledge. Remarkably, our novel algorithm can now achieve superior prediction outcomes using solely RGB images, eliminating the previous need for 3D point clouds and IMU input. Furthermore, we deploy our enhanced baseline algorithm on a real-world robotic platform to illustrate its practical utility in straightforward HRI tasks. The demonstrations showcase the real-world applicability of our advancements and may inspire more HRI use cases involving egocentric vision. All code and data are open-sourced and can be found on the project website. Irving Fang, Jianghan Zhang, Xibo He, Weibo Gao, Yiming Li 0003, Chen Feng 0002 |
ICRA | 11 |
| 2024 | ActFormer: Scalable Collaborative Perception via Active QueriesabstractCollaborative perception leverages rich visual observations from multiple robots to extend a single robot’s perception ability beyond its field of view. Many prior works receive messages broadcast from all collaborators, leading to a scalability challenge when dealing with a large number of robots and sensors. In this work, we aim to address scalable camera-based collaborative perception with a Transformer-based architecture. Our key idea is to enable a single robot to intelligently discern the relevance of the collaborators and their associated cameras according to a learned spatial prior. This proactive understanding of the visual features’ relevance does not require the transmission of the features themselves, enhancing both communication and computation efficiency. Specifically, we present ActFormer, a Transformer that learns bird’s eye view (BEV) representations by using predefined BEV queries to interact with multi-robot multi-camera inputs. Each BEV query can actively select relevant cameras for information aggregation based on pose information, instead of interacting with all cameras indiscriminately. Experiments on the V2X-Sim dataset demonstrate that ActFormer improves the detection performance from 29.89% to 45.15% in terms of [email protected] with about 50% fewer queries, showcasing the effectiveness of ActFormer in multi-agent collaborative 3D object detection. Suozhi Huang, Juexiao Zhang, Yiming Li 0003, Chen Feng 0002 |
ICRA | 4 |
| 2024 | Robust Collaborative Perception without External Localization and Clock DevicesabstractA consistent spatial-temporal coordination across multiple agents is fundamental for collaborative perception, which seeks to improve perception abilities through information exchange among agents. To achieve this spatial-temporal alignment, traditional methods depend on external devices to provide localization and clock signals. However, hardware-generated signals could be vulnerable to noise and potentially malicious attack, jeopardizing the precision of spatial-temporal alignment. Rather than relying on external hardwares, this work proposes a novel approach: aligning by recognizing the inherent geometric patterns within the perceptual data of various agents. Following this spirit, we propose a robust collaborative perception system that operates independently of external localization and clock devices. The key module of our system, FreeAlign, constructs a salient object graph for each agent based on its detected boxes and uses a graph neural network to identify common subgraphs between agents, leading to accurate relative pose and time. We validate FreeAlign on both real-world and simulated datasets. The results show that, the FreeAlign empowered robust collaborative perception system perform comparably to systems relying on precise localization and clock devices. ${\mathbf{Code}}$ will be released. Zixing Lei, Zhenyang Ni, Rui-Ze Han, Chen Feng 0002, Siheng Chen, Yanfeng Wang 0001 |
ICRA | 5 |
| 2024 | NYC-Indoor-VPR: A Long-Term Indoor Visual Place Recognition Dataset with Semi-Automatic AnnotationabstractVisual Place Recognition (VPR) in indoor environments is beneficial to humans and robots for better localization and navigation. It is challenging due to appearance changes at various frequencies, and difficulties of obtaining ground truth metric trajectories for training and evaluation. This paper introduces the NYC-Indoor-VPR dataset, a unique and rich collection of over 36,000 images compiled from 13 distinct crowded scenes in New York City taken under varying lighting conditions with appearance changes. Each scene has multiple revisits across a year. To establish the ground truth for VPR, we propose a semiautomatic annotation approach that computes the positional information of each image. Our method specifically takes pairs of videos as input and yields matched pairs of images along with their estimated relative locations. The accuracy of this matching is refined by human annotators, who utilize our annotation software to correlate the selected keyframes. Finally, we present a benchmark evaluation of several state-of-the-art VPR algorithms using our annotated dataset, revealing its challenge and thus value for VPR research. Diwei Sheng, Anbang Yang, John-Ross Rizzo, Chen Feng 0002 |
ICRA | 4 |
| 2024 | LiDAR-based 4D Occupancy Completion and ForecastingabstractScene completion and forecasting are two popular perception problems in research for mobile agents like autonomous vehicles. Existing approaches treat the two problems in isolation, resulting in a separate perception of the two aspects. In this paper, we introduce a novel LiDAR perception task of Occupancy Completion and Forecasting (OCF) in the context of autonomous driving to unify these aspects into a cohesive framework. This task requires new algorithms to address three challenges altogether: (1) sparse-to-dense reconstruction, (2) partial-to-complete hallucination, and (3) 3D-to-4D prediction. To enable supervision and evaluation, we curate a large-scale dataset termed OCFBench from public autonomous driving datasets. We analyze the performance of closely related existing baselines and variants on our dataset. We envision that this research will inspire and call for further investigation in this evolving and crucial area of 4D perception. Our code for data curation and baseline implementation is available at https://github.com/ai4ce/Occ4cast. Xinhao Liu 0003, Moonjun Gong, Yiming Li 0003, Hang Zhao 0021, Chen Feng 0002 |
IROS | 7 |
| 2024 | SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous DrivingabstractMonocular scene understanding is a foundational component of autonomous systems. Within the spectrum of monocular perception topics, one crucial and useful task for holistic 3D scene understanding is semantic scene completion (SSC), which jointly completes semantic information and geometric details from RGB input. However, progress in SSC, particularly in large-scale street views, is hindered by the scarcity of high-quality datasets. To address this issue, we introduce SSCBench, a comprehensive benchmark that integrates scenes from widely used automotive datasets (e.g., KITTI-360, nuScenes, and Waymo). SSCBench follows an established setup and format in the community, facilitating the easy exploration of SSC methods in various street views. We benchmark models using monocular, trinocular, and point cloud input to assess the performance gap resulting from sensor coverage and modality. Moreover, we have unified semantic labels across diverse datasets to simplify cross-domain generalization testing. We commit to including more datasets and SSC models to drive further advancements in this field. Our data and code are available at https://github.com/ai4ce/SSCBench. Yiming Li 0003, Sihang Li 0001, Xinhao Liu 0003, Moonjun Gong, Kenan Li, Nuo Chen 0003, Fisher Yu 0001, Yue Wang 0041, Hang Zhao 0021, Zhiding Yu, Chen Feng 0002 |
IROS | 14 |
| 2024 | Memorize What Matters: Emergent Scene Decomposition from MultitraverseabstractHumans naturally retain memories of permanent elements, while ephemeral moments often slip through the cracks of memory. This selective retention is crucial for robotic perception, localization, and mapping. To endow robots with this capability, we introduce 3D Gaussian Mapping (3DGM), a self-supervised, camera-only offline mapping framework grounded in 3D Gaussian Splatting. 3DGM converts multitraverse RGB videos from the same region into a Gaussian-based environmental map while concurrently performing 2D ephemeral object segmentation. Our key observation is that the environment remains consistent across traversals, while objects frequently change. This allows us to exploit self-supervision from repeated traversals to achieve environment-object decomposition. More specifically, 3DGM formulates multitraverse environmental mapping as a robust 3D representation learning problem, treating pixels of the environment and objects as inliers and outliers, respectively. Using robust feature distillation, feature residual mining, and robust optimization, 3DGM simultaneously performs 2D segmentation and 3D mapping without human intervention. We build the Mapverse benchmark, sourced from the Ithaca365 and nuPlan datasets, to evaluate our method in unsupervised 2D segmentation, 3D reconstruction, and neural rendering. Extensive results verify the effectiveness and potential of our method for self-driving and robotics. Yiming Li 0003, Zehong Wang, Yue Wang 0041, Zhiding Yu, Zan Gojcic, Marco Pavone 0001, Chen Feng 0002, José M. Álvarez 0004 |
NeurIPS | 7 |
| 2024 | Multiview Scene GraphabstractA proper scene representation is central to the pursuit of spatial intelligence where agents can robustly reconstruct and efficiently understand 3D scenes. A scene representation is either metric, such as landmark maps in 3D reconstruction, 3D bounding boxes in object detection, or voxel grids in occupancy prediction, or topological, such as pose graphs with loop closures in SLAM or visibility graphs in SfM.
In this work, we propose to build Multiview Scene Graphs (MSG) from unposed images, representing a scene topologically with interconnected place and object nodes.
The task of building MSG is challenging for existing representation learning methods since it needs to jointly address both visual place recognition, object detection, and object association from images with limited fields of view and potentially large viewpoint changes.
To evaluate any method tackling this task, we developed an MSG dataset and annotation based on a public 3D dataset.
We also propose an evaluation metric based on the intersection-over-union score of MSG edges.
Moreover, we develop a novel baseline method built on mainstream pretrained vision models, combining visual place recognition and object association into one Transformer decoder architecture.
Experiments demonstrate that our method has superior performance compared to existing relevant baselines. Juexiao Zhang, Gao Zhu, Sihang Li 0001, Xinhao Liu 0003, Haorui Song, Xinran Tang, Chen Feng 0002 |
NeurIPS | 7 |
| 2024 | Late better than early: A decision-level information fusion approach for RGB-Thermal crowd counting with illumination awareness
Jian Cheng 0001, Chen Feng 0002, Yang Xiao 0007, Zhiguo Cao 0001 |
Neurocomputing | 2 |
| 2024 | Advance One-Shot Multispectral Instance Detection With Text's SupervisionabstractOne key issue within one-shot multispectral instance detection (OMID) is to extract features of strong instance discriminative power, domain adaptation capability, and instance-wise generality. Existing methods generally only rely on visual clues. Comparatively, text is advantageous due to its structured information, high semantics, and low noise. Inspired by recent emergence of large image-text datasets and breakthrough visual-language models, we propose to advance OMID with text's supervision for the first time. To this end, our key idea is to establish the relationship between one-shot multispectral instance with ImageNet class labels via the CLIP model. Particularly, we retrieve, rank, and ensemble the text features of ImageNet labels via instance image feature as query. Then the resulting instance image and text features are realigned and fused to obtain a multimodal feature. Meanwhile, a multispectral contrastive learning approach is proposed to drive multimodal feature learning for OMID. Note that all the procedures are end-to-end trained in a unified network. In this way, the instance discriminative power and domain adaptation capability are facilitated simultaneously. Experiments on two tailored multispectral instance detection datasets verify the effectiveness of our method. Chen Feng 0002, Jian Cheng 0001, Yang Xiao 0007, Zhiguo Cao 0001 |
IEEE Signal Process. Lett. | 1 |
| 2023 | DeepMapping2: Self-Supervised Large-Scale LiDAR Map OptimizationabstractLiDAR mapping is important yet challenging in selfdriving and mobile robotics. To tackle such a global point cloud registration problem, DeepMapping [1] converts the complex map estimation into a self-supervised training of simple deep networks. Despite its broad convergence range on small datasets, DeepMapping still cannot produce satisfactory results on large-scale datasets with thousands of frames. This is due to the lack of loop closures and exact cross-frame point correspondences, and the slow convergence of its global localization network. We propose DeepMapping2 by adding two novel techniques to address these issues: (1) organization of training batch based on map topology from loop closing, and (2) self-supervised local-to-global point consistency loss leveraging pairwise registration. Our experiments and ablation studies on public datasets such as KITTI, NCLT, and Nebula demonstrate the effectiveness of our method. Xinhao Liu 0003, Yiming Li 0003, Li Ding 0009, Chen Feng 0002 |
CVPR | 5 |
| 2023 | VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene CompletionabstractHumans can easily imagine the complete 3D geometry of occluded objects and scenes. This appealing ability is vital for recognition and understanding. To enable such capability in AI systems, we propose VoxFormer, a Transformer-based semantic scene completion framework that can output complete 3D volumetric semantics from only 2D images. Our framework adopts a two-stage design where we start from a sparse set of visible and occupied voxel queries from depth estimation, followed by a densification stage that generates dense 3D voxels from the sparse ones. A key idea of this design is that the visual features on 2D images correspond only to the visible scene structures rather than the occluded or empty spaces. Therefore, starting with the fea-turization and prediction of the visible structures is more reliable. Once we obtain the set of sparse queries, we apply a masked autoencoder design to propagate the information to all the voxels by self-attention. Experiments on SemanticKITTI show that VoxFormer outperforms the state of the art with a relative improvement of 20.0% in geometry and 18.1% in semantics and reduces GPU memory during training to less than 16GB. Our code is available on https://github.com/NV1abs/VoxFormer. Yiming Li 0003, Zhiding Yu, Christopher B. Choy, Chaowei Xiao, José M. Álvarez 0004, Sanja Fidler, Chen Feng 0002, Anima Anandkumar |
CVPR | 7 |
| 2023 | Among Us: Adversarially Robust Collaborative Perception by ConsensusabstractMultiple robots could perceive a scene (e.g., detect objects) collaboratively better than individuals, although easily suffer from adversarial attacks when using deep learning. This could be addressed by the adversarial defense, but its training requires the often-unknown attacking mechanism. Differently, we propose ROBOSAC, a novel sampling-based defense strategy generalizable to unseen attackers. Our key idea is that collaborative perception should lead to consensus rather than dissensus in results compared to individual perception. This leads to our hypothesize-and-verify framework: perception results with and without collaboration from a random subset of teammates are compared until reaching a consensus. In such a framework, more teammates in the sampled subset often entail better perception performance but require longer sampling time to reject potential attackers. Thus, we derive how many sampling trials are needed to ensure the desired size of an attacker-free subset, or equivalently, the maximum size of such a subset that we can successfully sample within a given number of trials. We validate our method on the task of collaborative 3D object detection in autonomous driving scenarios. Yiming Li 0003, Jiamu Bai, Siheng Chen, Felix Juefei-Xu, Chen Feng 0002 |
ICCV | 6 |
| 2023 | Learning Simultaneous Navigation and Construction in Grid Worlds
Wenyu Han, Eisuke Hirota, Alexander Gao, Lerrel Pinto, Ludovic Righetti, Chen Feng 0002 |
ICLR | 7 |
| 2023 | Robust Collaborative 3D Object Detection in Presence of Pose ErrorsabstractCollaborative 3D object detection exploits information exchange among multiple agents to enhance accuracy of object detection in presence of sensor impairments such as occlusion. However, in practice, pose estimation errors due to imperfect localization would cause spatial message misalignment and significantly reduce the performance of collaboration. To alleviate adverse impacts of pose errors, we propose CoAlign, a novel hybrid collaboration framework that is robust to unknown pose errors. The proposed solution relies on a novel agent-object pose graph modeling to enhance pose consistency among collaborating agents. Furthermore, we adopt a multiscale data fusion strategy to aggregate intermediate features at multiple spatial resolutions. Comparing with previous works, which require ground-truth pose for training supervision, our proposed CoAlign is more practical since it doesn't require any ground-truth pose supervision in the training and makes no specific assumptions on pose errors. Extensive evaluation of the proposed method is carried out on multiple datasets, certifying that CoAlign significantly reduce relative localization error and achieving the state of art detection performance when pose errors exist. Code are made available for the use of the research community at https://github.com/yifanlu0227/CoAlign. Quanhao Li, Baoan Liu, Mehrdad Dianati, Chen Feng 0002, Siheng Chen, Yanfeng Wang 0001 |
ICRA | 5 |
| 2023 | Uncertainty Quantification of Collaborative Detection for Self-DrivingabstractSharing information between connected and autonomous vehicles (CAVs) fundamentally improves the performance of collaborative object detection for self-driving. However, CAVs still have uncertainties on object detection due to practical challenges, which will affect the later modules in self-driving such as planning and control. Hence, uncertainty quantification is crucial for safety-critical systems such as CAVs. Our work is the first to estimate the uncertainty of collaborative object detection. We propose a novel uncertainty quantification method, called Double- M Quantification, which tailors a moving block bootstrap (MBB) algorithm with direct modeling of the multivariant Gaussian distribution of each corner of the bounding box. Our method captures both the epistemic uncertainty and aleatoric uncertainty with one inference pass based on the offline Double- M training process. And it can be used with different collaborative object detectors. Through experiments on the comprehensive collaborative perception dataset, we show that our Double-M method achieves more than 4× improvement on uncertainty score and more than 3% accuracy improvement, compared with the state-of-the-art uncertainty quantification methods. Our code is public on https://coperception.github.io/double-m-quantification/. Sanbao Su, Yiming Li 0003, Sihong He, Songyang Han, Chen Feng 0002, Caiwen Ding, Fei Miao |
ICRA | 5 |
| 2023 | Toward Zero-Shot Sim-to-Real Transfer Learning for Pneumatic Soft Robot 3D Proprioceptive SensingabstractPneumatic soft robots present many advantages in manipulation tasks. Notably, their inherent compliance makes them safe and reliable in unstructured and fragile environments. However, full-body shape sensing for pneumatic soft robots is challenging because of their high degrees of freedom and complex deformation behaviors. Vision-based proprioception sensing methods relying on embedded cameras and deep learning provide a good solution to proprioception sensing by extracting the full-body shape information from the high-dimensional sensing data. But the current training data collection process makes it difficult for many applications. To address this challenge, we propose and demonstrate a robust sim-to-real pipeline that allows the collection of the soft robot's shape information in high-fidelity point cloud representation. The model trained on simulated data was evaluated with real internal camera images. The results show that the model performed with averaged Chamfer distance of 8.85 mm and tip position error of 10.12 mm even with external perturbation for a pneumatic soft robot with a length of 100.0 mm. We also demonstrated the sim-to-real pipeline's potential for exploring different configurations of visual patterns to improve vision-based reconstruction results. The code and dataset are available at https://github.com/DeepSoRo/DeepSoRoSim2Real. Uksang Yoo, Hanwen Zhao, Alvaro Altamirano, Wenzhen Yuan 0001, Chen Feng 0002 |
ICRA | 5 |
| 2023 | Multi-spectral template matching based object detection in a few-shot learning manner
Chen Feng 0002, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang, Joey Tianyi Zhou |
Inf. Sci. | 1 |
| 2022 | Egocentric Prediction of Action Target in 3DabstractWe are interested in anticipating as early as possible the target location of a person's object manipulation action in a 3D workspace from egocentric vision. It is important in fields like human-robot collaboration, but has not yet received enough attention from vision and learning communities. To stimulate more research on this challenging egocentric vision task, we propose a large multimodality dataset of more than 1 million frames of RGB-D and IMU streams, and provide evaluation metrics based on our high-quality 2D and 3D labels from semi-automatic annotation. Meanwhile, we design baseline methods using recurrent neural networks and conduct various ablation studies to validate their effectiveness. Our results demonstrate that this new task is worthy of further study by researchers in robotics, vision, and learning communities. Yiming Li 0003, Ziang Cao, Andrew Liang, Benjamin Liang, Luoyao Chen, Hang Zhao 0021, Chen Feng 0002 |
CVPR | 7 |
| 2022 | Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic CountingabstractClass-agnostic counting (CAC) aims to count all instances in a query image given few exemplars. A standard pipeline is to extract visual features from exemplars and match them with query images to infer object counts. Two essential components in this pipeline are feature representation and similarity metric. Existing methods either adopt a pretrained network to represent features or learn a new one, while applying a naive similarity metric with fixed inner product. We find this paradigm leads to noisy similarity matching and hence harms counting performance. In this work, we propose a similarity-aware CAC framework that jointly learns representation and similarity metric. We first instantiate our framework with a naive baseline called Bilinear Matching Network (BMNet), whose key component is a learnable bilinear similarity metric. To further embody the core of our framework, we extend BMNet to BMNet+ that models similarity from three aspects: 1) representing the instances via their self-similarity to enhance feature robustness against intra-class variations; 2) comparing the similarity dynamically to focus on the key patterns of each exemplar; 3) learning from a supervision signal to impose explicit constraints on matching results. Extensive experiments on a recent CAC dataset FSC147 show that our models significantly outperform state-of-the-art CAC approaches. In addition, we also validate the cross-dataset generality of BMNet and BMNet+ on a car counting dataset CARPK. Code is at tiny.one/BMNet Min Shi 0004, Hao Lu 0003, Chen Feng 0002, Zhiguo Cao 0001 |
CVPR | 3 |
| 2022 | Self-supervised Spatial Reasoning on Multi-View Line DrawingsabstractSpatial reasoning on multi-view line drawings by state-of-the-art supervised deep networks is recently shown with puzzling low performances on the SPARE3D dataset [14]. Based on the fact that self-supervised learning is helpful when a large number of data are available, we propose two self-supervised learning approaches to improve the baseline performance for view consistency reasoning and camera pose reasoning tasks on the SPARE3D dataset. For the first task, we use a self-supervised binary classification network to contrast the line drawing differences between various views of any two similar 3D objects, enabling the trained networks to effectively learn detail-sensitive yet view-invariant line drawing representations of 3D objects. For the second type of task, we propose a self-supervised multi-class classification framework to train a model to select the correct corresponding view from which a line drawing is rendered. Our method is even helpful for the downstream tasks with unseen camera poses. Experiments show that our method could significantly increase the baseline performance in SPARE3D, while some popular self-supervised learning methods cannot. Siyuan Xiang, Anbang Yang, Yanfei Xue, Yaoqing Yang 0002, Chen Feng 0002 |
CVPR | 5 |
| 2022 | A Deep Reinforcement Learning Environment for Particle Robot Navigation and Object ManipulationabstractParticle robots are novel biologically-inspired robotic systems where locomotion can be achieved collectively and robustly, but not independently. While its control is currently limited to a hand-crafted policy for basic locomotion tasks, such a multi-robot system could be potentially controlled via Deep Reinforcement Learning (DRL) for different tasks more efficiently. However, the particle robot system presents a new set of challenges for DRL differing from existing swarm robotics systems: the low degrees of freedom of each robot and the increased necessity of coordination between robots. We present a 2D particle robot simulator using the OpenAI Gym interface and Pymunk as the physics engine, and introduce new tasks and challenges to research the underexplored applications of DRL in the particle robot system. Moreover, we use Stable-baselines3 to provide a set of benchmarks for the tasks. Current baseline DRL algorithms show signs of achieving the tasks but are yet unable to reach the performance of the hand-crafted policy. Further development of DRL algorithms is necessary in order to accomplish the proposed tasks. Jeremy Shen, Erdong Xiao, Chen Feng 0002 |
ICRA | 4 |
| 2022 | Onboard Real-Time Aerial Tracking With Efficient Siamese Anchor Proposal NetworkabstractObject tracking approaches based on the Siamese network have demonstrated their huge potential in the remote sensing field recently. Nevertheless, due to the limited computing resource of aerial platforms and special challenges in aerial tracking, most existing Siamese-based methods can hardly meet the real-time and state-of-the-art performance simultaneously. Consequently, a novel Siamese-based method is proposed in this work for onboard real-time aerial tracking, i.e., SiamAPN. The proposed method is a no-prior two-stage method, i.e., Stage-1 for proposing adaptive anchors to enhance the ability of object perception and Stage-2 for fine-tuning the proposed anchors to obtain accurate results. Distinct from the traditional predefined anchors, the proposed anchors can adapt automatically to the tracking object. Besides, the internal information of adaptive anchors is utilized to feedback SiamAPN for enhancing the object perception. Attributing to the feature fusion network, different semantic information is integrated, enriching the information flow that is significant for robust aerial tracking. In the end, the regression and multiclassification operation refine the proposed anchors meticulously. Comprehensive evaluations on three well-known aerial tracking benchmarks have proven the superior performance of the presented approach. Moreover, to verify the practicability of the proposed method, SiamAPN is implemented onboard a typical embedded aerial tracking platform to conduct the real-world evaluations on specific aerial tracking scenarios, e.g., fast motion, long-term tracking, and low resolution. The results have demonstrated the efficiency and accuracy of the proposed approach, with a processing speed of over 30 frames/s. In addition, the image sequences in the real-world evaluations are collected and annotated as a new aerial tracking benchmark, i.e., UAVTrack112. Changhong Fu 0001, Ziang Cao, Yiming Li 0003, Junjie Ye 0004, Chen Feng 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | A Dataset-Dispersion Perspective on Reconstruction Versus Recognition in Single-View 3D Reconstruction NetworksabstractNeural networks (NN) for single-view 3D reconstruction (SVR) have gained in popularity. Recent work points out that for SVR, most cutting-edge NNs have limited performance on reconstructing unseen objects because they rely primarily on recognition (i.e., classification-based methods) rather than shape reconstruction. To understand this issue in depth, we provide a systematic study on when and why NNs prefer recognition to reconstruction and vice versa. Our finding shows that a leading factor in determining recognition versus reconstruction is how “dispersed” the training data is. Thus, we introduce the dispersion score, a new data-driven metric, to quantify this leading factor and study its effect on NNs. We hypothesize that NNs are biased toward recognition when training images are more dispersed and training shapes are less dispersed. Our hypothesis is supported and the dispersion score is proved effective through our experiments on synthetic and benchmark datasets. We show that the proposed metric is a principal way to analyze reconstruction quality and provides novel information in addition to the conventional reconstruction score. We have open-sourced our code.1 Yefan Zhou, Yiru Shen, Yujun Yan, Chen Feng 0002, Yaoqing Yang 0002 |
3DV | 4 |
| 2021 | Fooling LiDAR Perception via Adversarial Trajectory PerturbationabstractLiDAR point clouds collected from a moving vehicle are functions of its trajectories, because the sensor motion needs to be compensated to avoid distortions. When autonomous vehicles are sending LiDAR point clouds to deep networks for perception and planning, could the motion compensation consequently become a wide-open backdoor in those networks, due to both the adversarial vulnerability of deep learning and GPS-based vehicle trajectory estimation that is susceptible to wireless spoofing? We demonstrate such possibilities for the first time: instead of directly attacking point cloud coordinates which requires tampering with the raw LiDAR readings, only adversarial spoofing of a self-driving car’s trajectory with small perturbations is enough to make safety-critical objects undetectable or detected with incorrect positions. Moreover, polynomial trajectory perturbation is developed to achieve a temporally-smooth and highly-imperceptible attack. Extensive experiments on 3D object detection have shown that such attacks not only lower the performance of the state-of-the-art detectors effectively, but also transfer to other detectors, raising a red flag for the community. The code is available on https://ai4ce.github.io/FLAT/. Yiming Li 0003, Congcong Wen, Felix Juefei-Xu, Chen Feng 0002 |
ICCV | 4 |
| 2021 | Siamese Anchor Proposal Network for High-Speed Aerial TrackingabstractIn the domain of visual tracking, most deep learning-based trackers highlight the accuracy but casting aside efficiency. Therefore, their real-world deployment on mobile platforms like the unmanned aerial vehicle (UAV) is impeded. In this work, a novel two-stage Siamese network-based method is proposed for aerial tracking, i.e., stage-1 for high-quality anchor proposal generation, stage-2 for refining the anchor proposal. Different from anchor-based methods with numerous pre-defined fixed-sized anchors, our no-prior method can 1) increase the robustness and generalization to different objects with various sizes, especially to small, occluded, and fast-moving objects, under complex scenarios in light of the adaptive anchor generation, 2) make calculation feasible due to the substantial decrease of anchor numbers. In addition, compared to anchor-free methods, our framework has better performance owing to refinement at stage-2. Comprehensive experiments on three benchmarks have proven the superior performance of our approach, with a speed of ∼200 frames/s. Changhong Fu 0001, Ziang Cao, Yiming Li 0003, Junjie Ye 0004, Chen Feng 0002 |
ICRA | 5 |
| 2021 | Projector-Guided Non-Holonomic Mobile 3D PrintingabstractFused deposition modeling (FDM) using mobile robots instead of the gantry-based 3D printer enables additive manufacturing at a larger scale with higher speed. This introduces challenges including accurate localization, control of the printhead, and design of a stable mobile manipulator with low vibrations and proper degrees of freedom. We proposed and developed a low-cost non-holonomic mobile 3D printing system guided by a projector via learning-based visual servoing. It requires almost no manual calibration of the system parameters. Using a regular top-down projector without any expensive external localization device for pose feedback, this system enabled mobile robots to accurately follow pre-designed millimeter-level printing trajectories with speed control. We evaluate the system in terms of its trajectory accuracy and printing quality compared with original 3D designs. We further demonstrated the potential of this system using two such mobile robots to collaboratively print a 3D object with dimensions of 80 cm × 30 cm size, which exceeds the limitation of common desktop FDM 3D printers. Xuchu Xu, Chen Feng 0002 |
ICRA | 3 |
| 2021 | Mobile 3D Printing Robot Simulation with Viscoelastic FluidsabstractThe system design and algorithm development of mobile 3D printing robots need a realistic simulation. They require a mobile robot simulation platform to interoperate with a physics-based material simulation for handling interactions between the time-variant deformable 3D printing materials and other simulated rigid bodies in the environment, which is not available for roboticists yet. To bridge this gap and enable the real-time simulation of mobile 3D printing processes, we develop a simulation framework that includes particle-based viscoelastic fluid simulation and particle-to-mesh conversion in the widely adopted Gazebo robotics simulator, avoiding the bottlenecks of traditional additive manufacturing simulation approaches. This framework is the first of its kind that enables the simulation of robot arms or mobile manipulators together with viscoelastic fluids. The method is tested using various material properties and multiple collaborating robots to demonstrate its simulation ability for the robots to plan and control the printhead trajectories and to visually sense at the same time the printed fluid materials as a free-form mesh. The scalability as a function of available material particles in the simulation was also studied. A simulation with an average of 5 FPS was achieved on a regular desktop computer. Uljad Berdica, Yuewei Fu, Emmanouil Angelidis, Chen Feng 0002 |
IROS | 5 |
| 2021 | NYU-VPR: Long-Term Visual Place Recognition Benchmark with View Direction and Data Anonymization InfluencesabstractVisual place recognition (VPR) is critical in not only localization and mapping for autonomous driving vehicles, but also assistive navigation for the visually impaired population. To enable a long-term VPR system on a large scale, several challenges need to be addressed. First, different applications could require different image view directions, such as front views for self-driving cars while side views for the low vision people. Second, VPR in metropolitan scenes can often cause privacy concerns due to the imaging of pedestrian and vehicle identity information, calling for the need for data anonymization before VPR queries and database construction. Both factors could lead to VPR performance variations that are not well understood yet. To study their influences, we present the NYU-VPR dataset that contains more than 200,000 images over a 2km×2km area near the New York University campus, taken within the whole year of 2016. We present benchmark results on several popular VPR algorithms showing that side views are significantly more challenging for current VPR methods while the influence of data anonymization is almost negligible, together with our hypothetical explanations and in-depth analysis. Diwei Sheng, Yuxiang Chai, Chen Feng 0002, Jianzhe Lin, Cláudio T. Silva, John-Ross Rizzo |
IROS | 4 |
| 2021 | Learning Distilled Collaboration Graph for Multi-Agent PerceptionabstractTo promote better performance-bandwidth trade-off for multi-agent perception, we propose a novel distilled collaboration graph (DiscoGraph) to model trainable, pose-aware, and adaptive collaboration among agents. Our key novelties lie in two aspects. First, we propose a teacher-student framework to train DiscoGraph via knowledge distillation. The teacher model employs an early collaboration with holistic-view inputs; the student model is based on intermediate collaboration with single-view inputs. Our framework trains DiscoGraph by constraining post-collaboration feature maps in the student model to match the correspondences in the teacher model. Second, we propose a matrix-valued edge weight in DiscoGraph. In such a matrix, each element reflects the inter-agent attention at a specific spatial region, allowing an agent to adaptively highlight the informative regions. During inference, we only need to use the student model named as the distilled collaboration network (DiscoNet). Attributed to the teacher-student framework, multiple agents with the shared DiscoNet could collaboratively approach the performance of a hypothetical teacher model with a holistic view. Our approach is validated on V2X-Sim 1.0, a large-scale multi-agent perception dataset that we synthesized using CARLA and SUMO co-simulation. Our quantitative and qualitative experiments in multi-agent 3D object detection show that DiscoNet could not only achieve a better performance-bandwidth trade-off than the state-of-the-art collaborative perception methods, but also bring more straightforward design rationale. Our code is available on https://github.com/ai4ce/DiscoNet. Yiming Li 0003, Shunli Ren, Pengxiang Wu, Siheng Chen, Chen Feng 0002 |
NeurIPS | 5 |
| 2021 | Learning dynamic regression with automatic distractor repression for real-time UAV tracking
Changhong Fu 0001, Fangqiang Ding, Yiming Li 0003, Chen Feng 0002 |
Eng. Appl. Artif. Intell. | 5 |
| 2020 | LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility LikelihoodabstractModern face alignment methods have become quite accurate at predicting the locations of facial landmarks, but they do not typically estimate the uncertainty of their predicted locations nor predict whether landmarks are visible. In this paper, we present a novel framework for jointly predicting landmark locations, associated uncertainties of these predicted locations, and landmark visibilities. We model these as mixed random variables and estimate them using a deep network trained using our proposed Location, Uncertainty, and Visibility Likelihood (LUVLi) loss. In addition, we release an entirely new labeling of a large face alignment dataset with over 19,000 face images in a full range of head poses. Each face is manually labeled with the ground-truth locations of 68 landmarks, with the additional information of whether each landmarks is visible, self-occluded (due to extreme head poses), or externally occluded. Not only does our joint estimation yield accurate estimates of the uncertainty of predicted landmark locations, but it also yields state-of-the-art estimates for the landmark locations themselves on mulitple standard face alignment datasets. Our method's estimates of the uncertainty of predicted landmark locations could be used to automatically identify input images on which face alignment fails, which can be critical for downstream tasks. Abhinav Kumar 0004, Tim K. Marks, Wenxuan Mou, Ye Wang 0001, Michael J. Jones 0001, Anoop Cherian, Toshiaki Koike-Akino, Xiaoming Liu 0002, Chen Feng 0002 |
CVPR | 9 |
| 2020 | SPARE3D: A Dataset for SPAtial REasoning on Three-View Line DrawingsabstractSpatial reasoning is an important component of human intelligence. We can imagine the shapes of 3D objects and reason about their spatial relations by merely looking at their three-view line drawings in 2D, with different levels of competence. Can deep networks be trained to perform spatial reasoning tasks? How can we measure their "spatial intelligence"? To answer these questions, we present the SPARE3D dataset. Based on cognitive science and psychometrics, SPARE3D contains three types of 2D-3D reasoning tasks on view consistency, camera pose, and shape generation, with increasing difficulty. We then design a method to automatically generate a large number of challenging questions with ground truth answers for each task. They are used to provide supervision for training our baseline models using state-of-the-art architectures like ResNet. Our experiments show that although convolutional networks have achieved superhuman performance in many visual learning tasks, their spatial reasoning performance in SPARE3D is almost equal to random guesses. We hope SPARE3D can stimulate new problem formulations and network designs for spatial reasoning to empower intelligent robots to operate effectively in the 3D world via 2D sensors. Wenyu Han, Siyuan Xiang, Ruoyu Wang 0012, Chen Feng 0002 |
CVPR | 5 |
| 2020 | Regularizing Neural Networks via Minimizing Hyperspherical EnergyabstractInspired by the Thomson problem in physics where the distribution of multiple propelling electrons on a unit sphere can be modeled via minimizing some potential energy, hyperspherical energy minimization has demonstrated its potential in regularizing neural networks and improving their generalization power. In this paper, we first study the important role that hyperspherical energy plays in neural network training by analyzing its training dynamics. Then we show that naively minimizing hyperspherical energy suffers from some difficulties due to highly non-linear and non-convex optimization as the space dimensionality becomes higher, therefore limiting the potential to further improve the generalization. To address these problems, we propose the compressive minimum hyperspherical energy (CoMHE) as a more effective regularization for neural networks. Specifically, CoMHE utilizes projection mappings to reduce the dimensionality of neurons and minimizes their hyperspherical energy. According to different designs for the projection mapping, we propose several distinct yet well-performing variants and provide some theoretical guarantees to justify their effectiveness. Our experiments show that CoMHE consistently outperforms existing regularization methods, and can be easily applied to different neural networks. Rongmei Lin, Weiyang Liu, Zhen Liu 0019, Chen Feng 0002, Zhiding Yu, James M. Rehg, Li Xiong 0001 |
CVPR | 4 |
| 2020 | P2B: Point-to-Box Network for 3D Object Tracking in Point CloudsabstractTowards 3D object tracking in point clouds, a novel point-to-box network termed P2B is proposed in an end-to-end learning manner. Our main idea is to first localize potential target centers in 3D search area embedded with target information. Then point-driven 3D target proposal and verification are executed jointly. In this way, the time-consuming 3D exhaustive search can be avoided. Specifically, we first sample seeds from the point clouds in template and search area respectively. Then, we execute permutation-invariant feature augmentation to embed target clues from template into search area seeds and represent them with target-specific features. Consequently, the augmented search area seeds regress the potential target centers via Hough voting. The centers are further strengthened with seed-wise targetness scores. Finally, each center clusters its neighbors to leverage the ensemble power for joint 3D target proposal and verification. We apply PointNet++ as our backbone and experiments on KITTI tracking dataset demonstrate P2B's superiority (~10%'s improvement over state-of-the-art). Note that P2B can run with 40FPS on a single NVIDIA 1080Ti GPU. Our code and model are available at https://github.com/HaozheQi/P2B. Haozhe Qi, Chen Feng 0002, Zhiguo Cao 0001, Yang Xiao 0007 |
CVPR | 2 |
| 2020 | Parallel Network to Learn Novelty from the Known
Shuaiyuan Du, Chaoyi Hong, Chen Feng 0002, Zhiguo Cao 0001 |
ICPR | 4 |
| 2020 | DR2Track: Towards Real-Time Visual Tracking for UAV via Distractor Repressed Dynamic RegressionabstractVisual tracking has yielded promising applications with unmanned aerial vehicle (UAV). In literature, the advanced discriminative correlation filter (DCF) type trackers generally distinguish the foreground from the background with a learned regressor which regresses the implicit circulated samples into a fixed target label. However, the predefined and unchanged regression target results in low robustness and adaptivity to uncertain aerial tracking scenarios. In this work, we exploit the local maximum points of the response map generated in the detection phase to automatically locate current distractors1. By repressing the response of distractors in the regressor learning, we can dynamically and adaptively alter our regression target to leverage the tracking robustness as well as adaptivity. Substantial experiments conducted on three challenging UAV benchmarks demonstrate both excellent performance and extraordinary speed (~50fps on a cheap CPU) of our tracker. Changhong Fu 0001, Fangqiang Ding, Yiming Li 0003, Chen Feng 0002 |
IROS | 5 |
| 2020 | Automatic Failure Recovery and Re-Initialization for Online UAV Tracking with Joint Scale and Aspect Ratio OptimizationabstractCurrent unmanned aerial vehicle (UAV) visual tracking algorithms are primarily limited with respect to: (i) the kind of size variation they can deal with, (ii) the implementation speed which hardly meets the real-time requirement. In this work, a real-time UAV tracking algorithm with powerful size estimation ability is proposed. Specifically, the overall tracking task is allocated to two 2D filters: (i) translation filter for location prediction in the space domain, (ii) size filter for scale and aspect ratio optimization in the size domain. Besides, an efficient two-stage re-detection strategy is introduced for long-term UAV tracking tasks. Large-scale experiments on four UAV benchmarks demonstrate the superiority of the presented method which has computation feasibility on a low-cost CPU. Fangqiang Ding, Changhong Fu 0001, Yiming Li 0003, Chen Feng 0002 |
IROS | 5 |
| 2020 | A Low-Vision Navigation Platform for Economies in Transition CountriesabstractAn ability to move freely, when wanted, is an essential activity for healthy living. Visually impaired and completely blinded persons encounter many disadvantages in their day-to-day activities, including performing work-related tasks. They are at risk of mobility losses, illness, debility, social isolation, and premature mortality. A novel wearable device and computing platform called VIS4ION is reducing the disadvantage gaps and raising living standards for the visually challenged. It provides personal mobility navigational services that serves as a customizable, human-in-the-loop, sensing-to-feedback platform to deliver functional assistance. The platform is configured as a wearable that provides on-board microcomputers, human-machine interfaces, and sensory augmentation. Mobile edge computing enhances functionality as more services are unleashed with the computational gains. The meta-level goal is to support spatial cognition, personal freedom, and activities, and to promoting health and wellbeing. VIS4ION can be conceptualized as the dovetailing of two thrusts: an on-person navigational and computing device and a multimodal functional aid providing microservices through the cloud. The device has on-board wireless capabilities connected through Wi-Fi or 4/5G. The cloud-based microservices reduce hardware and power requirements while allowing existing and new services to be enhanced and added such as loading new map and real-time communication via haptic or audio signals. This technology can be made available and affordable in the economies of transition countries. John-Ross Rizzo, Chen Feng 0002, Wachara Riewpaiboon, Pattanasak Mongkolwat |
SERVICES | 2 |
| 2020 | Deep Unsupervised Learning of 3D Point Clouds via Graph Topology Inference and FilteringabstractWe propose a deep autoencoder with graph topology inference and filtering to achieve compact representations of unorganized 3D point clouds in an unsupervised manner. Many previous works discretize 3D points to voxels and then use lattice-based methods to process and learn 3D spatial information; however, this leads to inevitable discretization errors. In this work, we try to handle raw 3D points without such compromise. The proposed networks follow the autoencoder framework with a focus on designing the decoder. The encoder of the proposed networks adopts similar architectures as in PointNet, which is a well-acknowledged method for supervised learning of 3D point clouds. The decoder of the proposed networks involves three novel modules: the folding module, the graph-topology-inference module, and the graph-filtering module. The folding module folds a canonical 2D lattice to the underlying surface of a 3D point cloud, achieving coarse reconstruction; the graph-topology-inference module learns a graph topology to represent pairwise relationships between 3D points, pushing the latent code to preserve both coordinates and pairwise relationships of points in 3D point clouds; and the graph-filtering module couples the above two modules, refining the coarse reconstruction through a learnt graph topology to obtain the final reconstruction. The proposed decoder leverages a learnable graph topology to push the codeword to preserve representative features and further improve the unsupervised-learning performance. We further provide theoretical analyses of the proposed architecture. We provide an upper bound for the reconstruction loss and further show the superiority of graph smoothness over spatial smoothness as a prior to model 3D point clouds. In the experiments, we validate the proposed networks in three tasks, including 3D point cloud reconstruction, visualization, and transfer classification. The experimental results show that (1) the proposed networks outperform the state-of-the-art methods in various tasks, including reconstruction and transfer classification; (2) a graph topology can be inferred as auxiliary information without specific supervision on graph topology inference; (3) graph filtering refines the reconstruction, leading to better performances; and (4) designing a powerful decoder could improve the unsupervised-learning performance, just like a powerful encoder. Siheng Chen, Chaojing Duan, Yaoqing Yang 0002, Duanshun Li, Chen Feng 0002, Dong Tian |
IEEE Trans. Image Process. | 5 |
| 2019 | DeepMapping: Unsupervised Map Estimation From Multiple Point CloudsabstractWe propose DeepMapping, a novel registration framework using deep neural networks (DNNs) as auxiliary functions to align multiple point clouds from scratch to a globally consistent frame. We use DNNs to model the highly non-convex mapping process that traditionally involves hand-crafted data association, sensor pose initialization, and global refinement. Our key novelty is that “training” these DNNs with properly defined unsupervised losses is equivalent to solving the underlying registration problem, but less sensitive to good initialization than ICP. Our framework contains two DNNs: a localization network that estimates the poses for input point clouds, and a map network that models the scene structure by estimating the occupancy status of global coordinates. This allows us to convert the registration problem to a binary occupancy classification, which can be solved efficiently using gradient-based optimization. We further show that DeepMapping can be readily extended to address the problem of Lidar SLAM by imposing geometric constraints between consecutive point clouds. Experiments are conducted on both simulated and real datasets. Qualitative and quantitative comparisons demonstrate that DeepMapping often enables more robust and accurate global registration of multiple point clouds than existing techniques. Our code is available at https://ai4ce.github.io/DeepMapping/. Li Ding 0009, Chen Feng 0002 |
CVPR | 2 |
| 2019 | An assistive low-vision platform that augments spatial cognition through proprioceptive guidance: Point-to-Tell-and-TouchabstractSpatial cognition, as gained through the sense of vision, is one of the most important capabilities of human beings. However, for the visually impaired (VI), lack of this perceptual capability poses great challenges in their life. Therefore, we have designed Point-to-Tell-and-Touch, a wearable system with an ergonomic human-machine interface, for assisting the VI with active environmental exploration, with a particular focus on spatial intelligence and navigation to objects of interest in an alien environment. Our key idea is to link visual signals, as decoded synthetically, to the VI's proprioception for more intelligible guidance, in addition to vision-to-audio assistance, i.e., finger pose, as indicated by pointing, is used as “proprioceptive laser pointer” to target an object in that line of sight. The whole system consists of two features, Point-to-Tell and Point-to-Touch, both of which can work independently or cooperatively. The Point-to-Tell feature contains a camera with a novel one-stage neural network tailored for blind-centered object detection and recognition, and a headphone telling the VI the semantic label and distance from the pointed object. the Point-to-Touch, the second feature, leverages a vibrating wrist band to create a haptic feedback tool that supplements the initial vectorial guidance provided by the first stage (hand pose being direction and the distance being the extent, offered through audio cues). Both platform features utilize proprioception or joint position sense. Through hand pose, the VI end user knows where he or she is pointing relative to their egocentric coordinate system and we are able to use this foundation to build spatial intelligence. Our successful indoor experiments demonstrate the proposed system to be effective and reliable in helping the VI gain spatial cognition and explore the world in a more intuitive way. Wenjun Gui, Shuaihang Yuan, John-Ross Rizzo, Lakshay Sharma, Chen Feng 0002, Anthony Tzes, Yi Fang 0006 |
IROS | 6 |
| 2018 | Mining Point Cloud Local Structures by Kernel Correlation and Graph PoolingabstractUnlike on images, semantic learning on 3D point clouds using a deep network is challenging due to the naturally unordered data structure. Among existing works, PointNet has achieved promising results by directly learning on point sets. However, it does not take full advantage of a point's local neighborhood that contains fine-grained structural information which turns out to be helpful towards better semantic learning. In this regard, we present two new operations to improve PointNet with a more efficient exploitation of local structures. The first one focuses on local 3D geometric structures. In analogy to a convolution kernel for images, we define a point-set kernel as a set of learnable 3D points that jointly respond to a set of neighboring data points according to their geometric affinities measured by kernel correlation, adapted from a similar technique for point cloud registration. The second one exploits local high-dimensional feature structures by recursive feature aggregation on a nearest-neighbor-graph computed from 3D positions. Experiments show that our network can efficiently capture local information and robustly achieve better performances on major datasets. Our code is available at http://www.merl.com/research/license#KCNet. Yiru Shen, Chen Feng 0002, Yaoqing Yang 0002, Dong Tian |
CVPR | 2 |
| 2018 | FoldingNet: Point Cloud Auto-Encoder via Deep Grid DeformationabstractRecent deep networks that directly handle points in a point set, e.g., PointNet, have been state-of-the-art for supervised learning tasks on point clouds such as classification and segmentation. In this work, a novel end-to-end deep auto-encoder is proposed to address unsupervised learning challenges on point clouds. On the encoder side, a graph-based enhancement is enforced to promote local structures on top of PointNet. Then, a novel folding-based decoder deforms a canonical 2D grid onto the underlying 3D object surface of a point cloud, achieving low reconstruction errors even for objects with delicate structures. The proposed decoder only uses about 7% parameters of a decoder with fully-connected neural networks, yet leads to a more discriminative representation that achieves higher linear SVM classification accuracy than the benchmark. In addition, the proposed decoder structure is shown, in theory, to be a generic architecture that is able to reconstruct an arbitrary point cloud from a 2D grid. Our code is available at http://www.merl.com/research/license#FoldingNet. Yaoqing Yang 0002, Chen Feng 0002, Yiru Shen, Dong Tian |
CVPR | 2 |
| 2018 | Simultaneous Edge Alignment and Learning
Zhiding Yu, Weiyang Liu, Yang Zou 0003, Chen Feng 0002, Srikumar Ramalingam, B. V. K. Vijaya Kumar, Jan Kautz |
ECCV (3) | 4 |
| 2018 | Robust Attentional Pooling via Feature SelectionabstractIn this paper we propose a novel network module, namely Robust Attentional Pooling (RAP), that potentially can be applied in an arbitrary network for generating single vector representations for classification. By taking a feature matrix for each data sample as the input, our RAP learns data-dependent weights that are used to generate a vector through linear transformations of the feature matrix. We utilize feature selection to control the sparsity in weights for compressing the data matrices as well as enhancing the robustness of attentional pooling. As exemplary applications, we plug RAP into PointNet and ResNet for point cloud and image recognition, respectively. We demonstrate that our RAP significantly improves the recognition performance for both networks whenever sparsity is high. For instance, in extreme cases where only one feature per matrix is selected for recognition, RAP achieves more than 60% improvement over PointNet in terms of accuracy on the ModelNet40 dataset. Jianbo Zheng, Teng-Yok Lee, Chen Feng 0002, Xiaohua Li 0003 |
ICPR | 3 |
| 2018 | VLASE: Vehicle Localization by Aggregating Semantic EdgesabstractWe propose VLASE, a framework to use semantic edge features from images to achieve on-road localization. Semantic edge features denote edge contours that separate pairs of distinct objects such as building-sky, road-sidewalk, and building-ground. While prior work has shown promising results by utilizing the boundary between prominent classes such as sky and building using skylines, we generalize this to consider 19 semantic classes. We extract semantic edge features using CASENet architecture and utilize VLAD framework to perform image retrieval. We achieve improvement over state-of-the-art localization algorithms such as SIFT-VLAD and its deep variant NetVLAD. Ablation study shows the importance of different semantic classes, and our unified approach achieves better performance compared to individual prominent features such as skylines. We also introduce SLC Marathon dataset, a challenging dataset covering most of Salt Lake City with sufficient lighting variations. Xin Yu 0003, Sagar Chaturvedi, Chen Feng 0002, Yuichi Taguchi, Teng-Yok Lee, Clinton Fernandes, Srikumar Ramalingam |
IROS | 3 |
| 2017 | CASENet: Deep Category-Aware Semantic Edge DetectionabstractBoundary and edge cues are highly beneficial in improving a wide variety of vision tasks such as semantic segmentation, object recognition, stereo, and object proposal generation. Recently, the problem of edge detection has been revisited and significant progress has been made with deep learning. While classical edge detection is a challenging binary problem in itself, the category-aware semantic edge detection by nature is an even more challenging multi-label problem. We model the problem such that each edge pixel can be associated with more than one class as they appear in contours or junctions belonging to two or more semantic classes. To this end, we propose a novel end-to-end deep semantic edge learning architecture based on ResNet and a new skip-layer architecture where category-wise edge activations at the top convolution layer share and are fused with the same set of bottom layer features. We then propose a multi-label loss function to supervise the fused activations. We show that our proposed architecture benefits this problem with better performance, and we outperform the current state-of-the-art semantic edge detection methods by a large margin on standard data sets such as SBD and Cityscapes. Zhiding Yu, Chen Feng 0002, Ming-Yu Liu 0001, Srikumar Ramalingam |
CVPR | 2 |
| 2017 | Contour-enhanced resampling of 3D point clouds via graphsabstractTo reduce storage and computational cost for processing and visualizing large-scale 3D point clouds, an efficient resampling strategy is needed to select a representative subset of 3D points that can preserve contours in the original 3D point cloud. We tackle this problem by using graph-based techniques as graphs can represent underlying surfaces and lend themselves well to efficient computation. We first construct a general graph for a 3D point cloud and then propose a graph-based metric to quantify the contour information via high-pass graph filtering. Finally, we obtain an optimal resampling distribution that preserves the contour information by solving an optimization problem. When browsing, the proposed graph-based resampling performs better than uniform resampling both for toy point clouds as well as real large-scale point clouds. Furthermore, as neither mesh construction nor surface normal calculation is involved, the proposed graph-based method is computationally more efficient than the mesh-based methods. Siheng Chen, Dong Tian, Chen Feng 0002, Anthony Vetro, Jelena Kovacevic |
ICASSP | 3 |
| 2017 | Compression of 3-D point clouds using hierarchical patch fittingabstractFor applications such as virtual reality and mobile mapping, point clouds are an effective means for representing 3-D environments. The need for compressing such data is rapidly increasing, given the widespread use and precision of these systems. This paper presents a method for compressing organized point clouds. 3-D point cloud data is mapped to a 2-D organizational grid, where each element on the grid is associated with a point in 3-D space and its corresponding attributes. The data on the 2-D grid is hierarchically partitioned, and a Bezier patch is fit to the 3-D coordinates associated with each partition. Residual values are quantized and signaled along with data necessary to reconstruct the patch hierarchy in the decoder. We show how this method can be used to process point clouds captured by a mobile-mapping system, in which laser-scanned point locations are organized and compressed. The performance of the patch-fitting codec exceeds or is comparable to that of an octree-based codec. Robert A. Cohen, Maja Krivokuca, Chen Feng 0002, Yuichi Taguchi, Hideaki Ochimizu, Dong Tian, Anthony Vetro |
ICIP | 3 |
| 2017 | Geometric distortion metrics for point cloud compressionabstractIt is challenging to measure the geometry distortion of point cloud introduced by point cloud compression. Conventionally, the errors between point clouds are measured in terms of point-to-point or point-to-surface distances, that either ignores the surface structures or heavily tends to rely on specific surface reconstructions. To overcome these drawbacks, we propose using point-to-plane distances as a measure of geometric distortions on point cloud compression. The intrinsic resolution of the point clouds is proposed as a normalizer to convert the mean square errors to PSNR numbers. In addition, the perceived local planes are investigated at different scales of the point cloud. Finally, the proposed metric is independent of the size of the point cloud and rather reveals the geometric fidelity of the point cloud. From experiments, we demonstrate that our method could better track the perceived quality than the point-to-point approach while requires limited computations. Dong Tian, Hideaki Ochimizu, Chen Feng 0002, Robert A. Cohen, Anthony Vetro |
ICIP | 3 |
| 2017 | FasTFit: A fast T-spline fitting algorithm
Chen Feng 0002, Yuichi Taguchi |
Comput. Aided Des. | 1 |
| 2014 | Fast plane extraction in organized point clouds using agglomerative hierarchical clusteringabstractReal-time plane extraction in 3D point clouds is crucial to many robotics applications. We present a novel algorithm for reliably detecting multiple planes in real time in organized point clouds obtained from devices such as Kinect sensors. By uniformly dividing such a point cloud into non-overlapping groups of points in the image space, we first construct a graph whose node and edge represent a group of points and their neighborhood respectively. We then perform an agglomerative hierarchical clustering on this graph to systematically merge nodes belonging to the same plane until the plane fitting mean squared error exceeds a threshold. Finally we refine the extracted planes using pixel-wise region growing. Our experiments demonstrate that the proposed algorithm can reliably detect all major planes in the scene at a frame rate of more than 35Hz for 640×480 point clouds, which to the best of our knowledge is much faster than state-of-the-art algorithms. Chen Feng 0002, Yuichi Taguchi, Vineet R. Kamat |
ICRA | 1 |
| 2013 | Point-plane SLAM for hand-held 3D sensorsabstractWe present a simultaneous localization and mapping (SLAM) algorithm for a hand-held 3D sensor that uses both points and planes as primitives. We show that it is possible to register 3D data in two different coordinate systems using any combination of three point/plane primitives (3 planes, 2 planes and 1 point, 1 plane and 2 points, and 3 points). Our algorithm uses the minimal set of primitives in a RANSAC framework to robustly compute correspondences and estimate the sensor pose. As the number of planes is significantly smaller than the number of points in typical 3D data, our RANSAC algorithm prefers primitive combinations involving more planes than points. In contrast to existing approaches that mainly use points for registration, our algorithm has the following advantages: (1) it enables faster correspondence search and registration due to the smaller number of plane primitives; (2) it produces plane-based 3D models that are more compact than point-based ones; and (3) being a global registration algorithm, our approach does not suffer from local minima or any initialization problems. Our experiments demonstrate real-time, interactive 3D reconstruction of indoor spaces using a hand-held Kinect sensor. Yuichi Taguchi, Yong-Dian Jian, Srikumar Ramalingam, Chen Feng 0002 |
ICRA | 4 |
| 2012 | SLAM using both points and planes for hand-held 3D sensorsabstractWe present a simultaneous localization and mapping (SLAM) algorithm for a hand-held 3D sensor that uses both points and planes as primitives. Our algorithm uses any combination of three point/plane primitives (3 planes, 2 planes and 1 point, 1 plane and 2 points, and 3 points) in a RANSAC framework to efficiently compute the sensor pose. As the number of planes is significantly smaller than the number of points in typical 3D scenes, our RANSAC algorithm prefers primitive combinations involving more planes than points. In contrast to existing approaches that mainly use points for registration, our algorithm has the following advantages: (1) it enables faster correspondence search and registration due to the smaller number of plane primitives; (2) it produces plane-based 3D models that are more compact than point-based ones; and (3) being a global registration algorithm, our approach does not suffer from local minima or any initialization problems. Our experiments demonstrate real-time, interactive 3D reconstruction of office spaces using a hand-held Kinect sensor. Yuichi Taguchi, Yong-Dian Jian, Srikumar Ramalingam, Chen Feng 0002 |
ISMAR | 4 |