VLDB 2026 Research / reviewers in the wild / expert
Junjie Hu 0003
dblp:123/0773-3
· DBLP profile ↗
32ranked-venue papers
9as first author
28since 2021 · last 2026
0000-0002-1911-4361ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 6 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 9 since 2021Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RoSe: Robust Self-Supervised Stereo Matching Under Adverse Weather ConditionsabstractRecent self-supervised stereo matching methods have made significant progress, but their performance significantly degrades under adverse weather conditions such as night, rain, and fog. We identify two primary weaknesses contributing to this performance degradation. First, adverse weather introduces noise and reduces visibility, making CNN-based feature extractors struggle with degraded regions like reflective and textureless areas. Second, these degraded regions can disrupt accurate pixel correspondences, leading to ineffective supervision based on the photometric consistency assumption. To address these challenges, we propose injecting robust priors derived from the visual foundation model into the CNN-based feature extractor to improve feature representation under adverse weather conditions. We then introduce scene correspondence priors to construct robust supervisory signals rather than relying solely on the photometric consistency assumption. Specifically, we create synthetic stereo datasets with realistic weather degradations. These datasets feature clear and adverse image pairs that maintain the same semantic context and disparity, preserving the scene correspondence property. With this knowledge, we propose a robust self-supervised training paradigm, consisting of two key steps: robust self-supervised scene correspondence learning and adverse weather distillation. Both steps aim to align underlying scene results from clean and adverse image pairs, thus improving model disparity estimation under adverse weather effects. Extensive experiments demonstrate the effectiveness and versatility of our proposed solution, which outperforms existing state-of-the-art self-supervised methods. Codes are available at https://github.com/cocowy1/RoSe-Robust-Self-supervised-Stereo-Matching-under-Adverse-Weather-Conditions. Yun Wang 0053, Junjie Hu 0003, Junhui Hou, Chenghao Zhang 0003, Renwei Yang, Dapeng Oliver Wu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Peer Learning Approach to Unbiased Scene Graph Generation for Traffic Scene UnderstandingabstractThe biased scene graph generation problem arises from the inherent long-tailed distributions of predicates, which are challenging to handle effectively with a single network. In this paper, we introduce a novel framework called peer learning, designed to address the issue of unbiased scene graph generation (USGG) through a divide-and-vote approach. To address the long-tailed problem, our framework operates in three steps. Firstly, we partition the heavily long-tailed distribution into subsets of more balanced sub-distribution groups, including head, body, and tail classes with a predicate sampling module. Next, we establish a peer network consisting of multiple peers, where each peer receives a combination of sub-distributions. This division enables peers to focus on different aspects of the scene graph generation task. Then, a novel peer learning loss function is introduced to cultivate the learning process among peer networks. Lastly, we employ the voting strategies for making final predictions within the peer network, boosting the influence of the majority’s opinion while downplaying the minority’s perspective. To illustrate the applicability of the proposed framework in intelligent transportation systems (ITSs), we further conduct qualitative evaluations on traffic scene understanding tasks. The results demonstrate that peer learning markedly enhances the reliability of interpreting complex traffic scenarios. Experimental results on the Visual Genome and Open Images V6 datasets further verify the effectiveness of our proposed model. These results highlight that the peer learning framework is well-suited for addressing the challenges of unbiased scene graph generation, offering practical benefits for ITS applications such as traffic analysis and monitoring. The code is available at: PL. Liguang Zhou, Junjie Hu 0003, Yuhongze Zhou, Tin Lun Lam, Yangsheng Xu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | DualNet: Robust Self-Supervised Stereo Matching with Pseudo-Label SupervisionabstractSelf-supervised stereo matching has drawn attention due to its ability to estimate disparity without needing ground-truth data. However, existing self-supervised stereo matching methods heavily rely on the photo-metric consistency assumption, which is vulnerable to natural disturbances, resulting in ambiguous supervision and inferior performance compared to the supervised ones. To relax the limitation of the photo-metric consistency assumption and even bypass this assumption, we propose a novel self-supervised framework named DualNet, which consists of two key steps: robust self-supervised teacher learning and pseudo-label supervised student training. Specifically, the teacher model is first trained in a self-supervised manner with a focus on feature-metric consistency and data augmentation consistency. Then, the output of the teacher model is geometrically constrained to obtain high-quality pseudo labels. Benefiting from these high-quality pseudo labels, the student model can outperform its teacher model by a large margin. With the two well-designed steps, the proposed framework DualNet ranks 1st among all self-supervised methods on multiple benchmarks, surprisingly even outperforming several supervised counterparts. Yun Wang 0053, Jiahao Zheng 0001, Chenghao Zhang 0003, Zhanjie Zhang, Kunhong Li 0001, Junjie Hu 0003 |
AAAI | 7 |
| 2025 | SGFormer: Satellite-Ground Fusion for 3D Semantic Scene CompletionabstractRecently, camera-based solutions have been extensively explored for scene semantic completion (SSC). Despite their success in visible areas, existing methods struggle to capture complete scene semantics due to frequent visual occlusions. To address this limitation, this paper presents the first satellite-ground cooperative SSC framework, i.e., SGFormer, exploring the potential of satellite-ground image pairs in the SSC task. Specifically, we propose a dual-branch architecture that encodes orthogonal satellite and ground views in parallel, unifying them into a common domain. Additionally, we design a ground-view guidance strategy that corrects satellite image biases during feature encoding, addressing misalignment between satellite and ground views. Moreover, we develop an adaptive weighting strategy that balances contributions from satellite and ground views. Experiments demonstrate that SG-Former outperforms the state of the art on SemanticKITTI and SSCBench-KITTI-360 datasets. Our code is available on https://github.com/gxytcrc/SGFormer. Xiyue Guo, Jiarui Hu 0004, Junjie Hu 0003, Hujun Bao, Guofeng Zhang 0001 |
CVPR | 3 |
| 2025 | Learning Robust Stereo Matching in the Wild with Selective Mixture-of-ExpertsabstractRecently, learning-based stereo matching networks have advanced significantly. However, they often lack robustness and struggle to achieve impressive cross-domain performance due to domain shifts and imbalanced disparity distributions among diverse datasets. Leveraging Vision Foundation Models (VFMs) can intuitively enhance the model's robustness, but integrating such a model into stereo matching cost-effectively to fully realize their robustness remains a key challenge. To address this, we propose SMoEStereo, a novel framework that adapts VFMs for stereo matching through a tailored, scene-specific fusion of Low-Rank Adaptation (LoRA) and Mixture-of-Experts (MoE) modules. SMoEStereo introduces MoE-LoRA with adaptive ranks and MoE-Adapter with adaptive kernel sizes. The former dynamically selects optimal experts within MoE to adapt varying scenes across domains, while the latter injects inductive bias into frozen VFMs to improve geometric feature extraction. Importantly, to mitigate computational overhead, we further propose a lightweight decision network that selectively activates MoE modules based on input complexity, balancing efficiency with accuracy. Extensive experiments demonstrate that our method exhibits state-of-the-art cross-domain and joint generalization across multiple benchmarks without dataset-specific adaptation. The code is available at \textcolor{red}{https://github.com/cocowy1/SMoE-Stereo}. Yun Wang 0053, Longguang Wang, Chenghao Zhang 0003, Zhanjie Zhang, Ao Ma 0005, Chenyou Fan, Tin Lun Lam, Junjie Hu 0003 |
ICCV | 9 |
| 2025 | Topology-Based Visual Active Room SegmentationabstractRoom segmentation plays a significant role in scene understanding, semantic mapping, and scene coverage for robots navigating in real-world indoor environments. However, most previous works take a passive segmentation that requires a complete and uncluttered grid map as input, often resulting in lower segmentation accuracy and cannot be deployed in unknown environments. In this paper, we propose an active room segmentation framework that can enable a robot to incrementally and autonomously perform room segmentation in cluttered indoor environments. Our framework consists of three key components: i) a door extraction module where a visual semantic feature, specifically, door, is extracted to better identify rooms in cluttered environments, ii) a within-room exploration module that detects frontiers within the currently exploring room, and iii) a topological module that represents connectivity between rooms and determines next room for exploration. We show through experiments that the proposed method depicts two distinct advantages against existing methods in segmentation accuracy and autonomy. The code is available at https://github.com/FreeformRobotics/Active_room_segmentation. Chenyu Bao, Junjie Hu 0003, Qiu Zheng, Tin Lun Lam |
ICRA | 2 |
| 2025 | Transferring Visual Knowledge: Semi-Supervised Instance Segmentation for Object Navigation Across Varying Height ViewpointsabstractThe object navigation task requires robots to understand the semantic regularities in their environments. However, existing modular object navigation frameworks rely on instance segmentation models trained at fixed camera height viewpoints, limiting generalization performance and increasing labeling costs for new height viewpoints. To tackle this issue, we propose a semi-supervised method that transfers knowledge from a source height to a target height, minimizing the need for additional labels. Our approach introduces three key innovations: i) a projection policy to enhance the teacher model's detection capabilities at the target height, ii) a dynamic weight mechanism that emphasizes high-confidence pseudo-labels to reduce overfitting, and iii) a prototype contrast transferring method to transfer knowl-edge effectively. Experiments on the Habitat- Matterport 3D (HM3D) dataset show our method outperforms state-of-the-art semi-supervised techniques, improving both segmentation accuracy and navigation performance. The code is available at: https://github.com/FreeformRobotics/TransferKnowledge. Qiu Zheng, Junjie Hu 0003, Zengfeng Zeng, Tin Lun Lam |
ICRA | 2 |
| 2025 | PPMStereo: Pick-and-Play Memory Construction for Consistent Dynamic Stereo MatchingabstractTemporally consistent depth estimation from stereo video is critical for real-world applications such as augmented reality, where inconsistent depth estimation disrupts the immersion of users.
Despite its importance, this task remains challenging due to the difficulty in modeling long-term temporal consistency in a computationally efficient manner.
Previous methods attempt to address this by aggregating spatio-temporal information but face a fundamental trade-off: limited temporal modeling provides only modest gains, whereas capturing long-range dependencies significantly increases computational cost.
To address this limitation, we introduce a memory buffer for modeling long-range spatio-temporal consistency while achieving efficient dynamic stereo matching.
Inspired by the two-stage decision-making process in humans, we propose a Pick-and-Play Memory (PPM) construction module for dynamic Stereo matching, dubbed as PPMStereo. PPM consists of a pick process that identifies the most relevant frames and a play process that weights the selected frames adaptively for spatio-temporal aggregation.
This two-stage collaborative process maintains a compact yet highly informative memory buffer while achieving temporally consistent information aggregation.
Extensive experiments validate the effectiveness of PPMStereo, demonstrating state-of-the-art performance in both accuracy and temporal consistency.Codes are available at \textcolor{blue}{https://github.com/cocowy1/PPMStereo}. Yun Wang 0053, Junjie Hu 0003, Qiaole Dong, Yanwei Fu 0001, Tin Lun Lam, Dapeng Oliver Wu |
NeurIPS | 2 |
| 2025 | Enhancing Human Trajectory Prediction with Reinforcement Learning from Quantified Human Preferences
Chenyou Fan, Kehui Tan, Yanzhao Chen, Tianqi Pang, Haiqi Jiang 0003, Junjie Hu 0003 |
PRCV (7) | 6 |
| 2025 | Robust Depth Estimation Under Sensor Degradations: A Multi-Sensor Fusion PerspectiveabstractThe significance of depth estimation has spurred recent endeavors to enhance it through Multi-Sensor Fusion (MSF). However, prevailing MSF methods exhibit limitations concerning accuracy and resilience when confronted with sensor degradations. While certain forms of degradation, such as suboptimal lighting and adverse weather conditions, can be mitigated by collecting pertinent data in data-driven learning, this approach proves ineffective for Out-of-Distribution (OOD) sensor degradations. In this paper, we propose a novel approach termed Combinable and Separable Multi-Sensor Fusion (CSMSF) designed to bolster depth estimation robustness against multiple sensor degradations. CSMSF hinges on four core principles: i) improved performance is achieved with an increased number of valid sensors, ii) a single valid sensor can independently enable its own depth estimation, iii) maintaining a judicious equilibrium between accuracy and model complexity, and iv) autonomous diagnosis of sensor observation failure. Leveraging these advantages, CSMSF identifies and rejects degraded sensors, allowing autonomous selection of valid sensors for scene depth estimation. The experimental results demonstrate the superior robustness of the proposed CSMSF, underscoring its efficacy in addressing challenges associated with sensor degradations across diverse environmental conditions. Junjie Hu 0003, Chenyou Fan, Mete Ozay, Qing Gao 0002, Yulan Guo, Tin Lun Lam |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Domain adaptive depth completion via spatial-error consistency
Lingyu Xiao, Junjie Hu 0003, Wankou Yang |
Pattern Recognit. | 3 |
| 2025 | ADStereo: Efficient Stereo Matching With Adaptive Downsampling and Disparity AlignmentabstractThe balance between accuracy and computational efficiency is crucial for the applications of deep learning-based stereo matching algorithms in real-world scenarios. Since matching cost aggregation is usually the most computationally expensive component, a common practice is to construct cost volumes at a low resolution for aggregation and then directly regress a high-resolution disparity map. However, current solutions often suffer from limitations such as the loss of discriminative features caused by downsampling operations that treat all pixels equally, and spatial misalignment resulting from repeated downsampling and upsampling. To overcome these challenges, this paper presents two sampling strategies: the Adaptive Downsampling Module (ADM) and the Disparity Alignment Module (DAM), to prioritize real-time inference while ensuring accuracy. The ADM leverages local features to learn adaptive weights, enabling more effective downsampling while preserving crucial structure information. On the other hand, the DAM employs a learnable interpolation strategy to predict transformation offsets of pixels, thereby mitigating the spatial misalignment issue. Building upon these modules, we introduce ADStereo, a real-time yet accurate network that achieves highly competitive performance on multiple public benchmarks. Specifically, our ADStereo runs over faster than the current state-of-the-art CREStereo (0.054s vs. ) under the same hardware while achieving comparable accuracy (1.82% vs. 1.69%) on the KITTI stereo 2015 benchmark. The codes are available at: https://github.com/cocowy1/ADStereo. Yun Wang 0053, Kunhong Li 0001, Longguang Wang, Junjie Hu 0003, Dapeng Oliver Wu, Yulan Guo |
IEEE Trans. Image Process. | 4 |
| 2025 | Unlocking Drone Perception in Low AGL Heights: Progressive Semi-Supervised Learning for Ground-to-Aerial Perception Knowledge TransferabstractWe explore the novel challenge of drone perception across varying low AGL (above ground level) heights, a task essential for dynamic tasks, unlike the fixed ground viewpoint in autonomous driving. Supervised learning for this incurs high annotation costs, and current semi-supervised methods struggle with viewpoint differences. In this paper, we introduce ground-to-aerial perception knowledge transfer and propose a progressive semi-supervised learning framework for drone perception using only labeled data from the ground viewpoint and unlabeled data from flying viewpoints. The framework hinges on four key components: 1) a dense viewpoint sampling strategy, segmenting the vertical flight height range into evenly distributed intervals; 2) nearest neighbor pseudo-labeling, inferring labels of the nearest neighbor viewpoint using a model learned on the preceding viewpoint; 3) MixView, generating augmented images among different viewpoints to mitigate viewpoint differences; and 4) a progressive distillation strategy, gradually learning until reaching the maximum flying height. To validate our approach, we create both synthesized and real-world datasets. Extensive experimental analyses reveal a remarkable relative accuracy improvement of 25.7% and 16.9% for the synthesized dataset and the real world, respectively. Code and datasets are available on https://github.com/FreeformRobotics/Progressive-Self-Distillation-for-Ground-to-Aerial-Perception-Knowledge-Transfer. Junjie Hu 0003, Chenyou Fan, Mete Ozay, Yuan Gao 0024, Tin Lun Lam |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Lifelong-MonoDepth: Lifelong Learning for Multidomain Monocular Metric Depth EstimationabstractWith the rapid advancements in autonomous driving and robot navigation, there is a growing demand for lifelong learning (LL) models capable of estimating metric (absolute) depth. LL approaches potentially offer significant cost savings in terms of model training, data storage, and collection. However, the quality of RGB images and depth maps is sensor-dependent, and depth maps in the real world exhibit domain-specific characteristics, leading to variations in depth ranges. These challenges limit existing methods to LL scenarios with small domain gaps and relative depth map estimation. To facilitate lifelong metric depth learning, we identify three crucial technical challenges that require attention: 1) developing a model capable of addressing the depth scale variation through scale-aware depth learning; 2) devising an effective learning strategy to handle significant domain gaps; and 3) creating an automated solution for domain-aware depth inference in practical applications. Based on the aforementioned considerations, in this article, we present 1) a lightweight multihead framework that effectively tackles the depth scale imbalance; 2) an uncertainty-aware LL solution that adeptly handles significant domain gaps; and 3) an online domain-specific predictor selection method for real-time inference. Through extensive numerical studies, we show that the proposed method can achieve good efficiency, stability, and plasticity, leading the benchmarks by 8%-15%. The code is available at https://github.com/FreeformRobotics/Lifelong-MonoDepth. Junjie Hu 0003, Chenyou Fan, Liguang Zhou, Qing Gao 0002, Honghai Liu 0001, Tin Lun Lam |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Meta-Reinforcement Learning Based Cooperative Surface Inspection of 3D Uncertain Structures using Multi-robot SystemsabstractThis paper presents a decentralized cooperative motion planning approach for surface inspection of 3D structures which includes uncertainties like size, number, shape, position, using multi-robot systems (MRS). Given that most of existing works mainly focus on surface inspection of single and fully known 3D structures, our motivation is two-fold: first, 3D structures separately distributed in 3D environments are complex, therefore the use of MRS intuitively can facilitate an inspection by fully taking advantage of sensors with different capabilities. Second, performing the aforementioned tasks when considering uncertainties is a complicated and time-consuming process because we need to explore, figure out the size and shape of 3D structures and then plan surface-inspection path. To overcome these challenges, we present a meta-learning approach that provides a decentralized planner for each robot to improve the exploration and surface inspection capabilities. The experimental results demonstrate our method can outperform other methods by approximately 10.5%-27% on success rate and 70%-75% on inspection speed. Yuan Gao 0024, Junjie Hu 0003, Fuqin Deng, Tin Lun Lam |
ICRA | 3 |
| 2024 | From Satellite to Ground: Satellite Assisted Visual Localization with Cross-view Semantic MatchingabstractOne of the key challenges of visual Simultaneous Localization and Mapping (SLAM) in large-scale environments is how to effectively use global localization to correct the cumulative errors from long-term tracking. This challenge presents itself in two main aspects: first, the difficulty for robots in revisiting previous locations to perform loop closure, and second, the considerable memory resources required to maintain point-cloud-based global maps. Recent solutions have resorted into neural networks, using satellite images as the references for ground-level localization. However, most of these methods merely provide cross-view patch-matching results, which leads to unfeasible in integration with the SLAM system. To address these issues, we present a semantic-based cross-view localization method. This approach combines semantic information with a reward and penalty mechanism, enabling us to obtain a global probability map and achieve precise 3-degree-of-freedom (3-DoF) localization. Based on that, we develop a SLAM system that capitalizes on satellite imagery for global localization. This strategy effectively bridges the gap between SLAM and real-world coordinates while also substantially reducing accumulated errors. Our experimental results demonstrate that our global localization method significantly outperforms existing satellite-based systems. Moreover, in scenarios where the robot struggles to find loop closures, employing our localization method improves the SLAM accuracy. Xiyue Guo, Haocheng Peng, Junjie Hu 0003, Hujun Bao, Guofeng Zhang 0001 |
ICRA | 3 |
| 2024 | Streamlining Forest Wildfire Surveillance: AI-Enhanced UAVs Utilizing the FLAME Aerial Video Dataset for Lightweight and Efficient MonitoringabstractIn recent years, unmanned aerial vehicles (UAVs) have played an increasingly crucial role in supporting disaster emergency response efforts by analyzing aerial images. While current deep-learning models focus on improving accuracy, they often overlook the limited computing resources of UAVs. This study recognizes the imperative for real-time data processing in disaster response scenarios and introduces a lightweight and efficient approach for aerial video understanding. Our methodology identifies redundant portions within the video through policy networks and eliminates this excess information using frame compression techniques. Additionally, we introduced the concept of a station point, which leverages future information in the sequential policy network, thereby enhancing accuracy. To validate our method, we employed the wildfire FLAME dataset. Compared to the baseline, our approach reduces computation costs by more than 10 times while improving accuracy by 3%. Moreover, our method can intelligently select salient frames from the video, refining the dataset. This feature enables sophisticated models to be effectively trained on a smaller dataset, significantly reducing the time spent during the training process. Lemeng Zhao, Junjie Hu 0003, Jianchao Bi, Yanbing Bai, Erick Mas, Shunichi Koshimura |
IROS | 2 |
| 2024 | Dense depth distillation with out-of-distribution simulated images
Junjie Hu 0003, Chenyou Fan, Mete Ozay, Hualie Jiang, Tin Lun Lam |
Knowl. Based Syst. | 1 |
| 2023 | Trajectory Prediction with Contrastive Pre-training and Social Rank Fine-Tuning
Chenyou Fan, Haiqi Jiang 0003, Aimin Huang, Junjie Hu 0003 |
ICONIP (10) | 4 |
| 2023 | Descriptor Distillation for Efficient Multi-Robot SLAMabstractPerforming accurate localization while maintaining the low-level communication bandwidth is an essential challenge of multi-robot simultaneous localization and mapping (MR-SLAM). In this paper, we tackle this problem by generating a compact yet discriminative feature descriptor with minimum inference time. We propose descriptor distillation that formulates the descriptor generation into a learning problem under the teacher-student framework. To achieve real-time descriptor generation, we design a compact student network and learn it by transferring the knowledge from a pre-trained large teacher model. To reduce the descriptor dimensions from the teacher to the student, we propose a novel loss function that enables the knowledge transfer between two different dimensional descriptors. The experimental results demonstrate that our model is 30% lighter than the state-of-the-art model and produces better descriptors in patch matching. Moreover, we build a MR-SLAM system based on the proposed method and show that our descriptor distillation can achieve higher localization performance for MR-SLAM with lower bandwidth. Xiyue Guo, Junjie Hu 0003, Hujun Bao, Guofeng Zhang 0001 |
ICRA | 2 |
| 2023 | Boosting LightWeight Depth Estimation via Knowledge Distillation
Junjie Hu 0003, Chenyou Fan, Hualie Jiang, Xiyue Guo, Yuan Gao 0024, Xiangyong Lu, Tin Lun Lam |
KSEM (1) | 1 |
| 2023 | Few-Shot Multi-Agent Perception With Ranking-Based Feature LearningabstractIn this article, we focus on performing few-shot learning (FSL) under multi-agent scenarios in which participating agents only have scarce labeled data and need to collaborate to predict labels of query observations. We aim at designing a coordination and learning framework in which multiple agents, such as drones and robots, can collectively perceive the environment accurately and efficiently under limited communication and computation conditions. We propose a metric-based multi-agent FSL framework which has three main components: an efficient communication mechanism that propagates compact and fine-grained query feature maps from query agents to support agents; an asymmetric attention mechanism that computes region-level attention weights between query and support feature maps; and a metric-learning module which calculates the image-level relevance between query and support data fast and accurately. Furthermore, we propose a specially designed ranking-based feature learning module, which can fully utilize the order information of training data by maximizing the inter-class distance, while minimizing the intra-class distance explicitly. We perform extensive numerical studies and demonstrate that our approach can achieve significantly improved accuracy in visual and acoustic perception tasks such as face identification, semantic segmentation, and sound genre recognition, consistently outperforming the state-of-the-art baselines by 5%-20%. Chenyou Fan, Junjie Hu 0003, Jianwei Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Deep Depth Completion From Extremely Sparse Data: A SurveyabstractDepth completion aims at predicting dense pixel-wise depth from an extremely sparse map captured from a depth sensor, e.g., LiDARs. It plays an essential role in various applications such as autonomous driving, 3D reconstruction, augmented reality, and robot navigation. Recent successes on the task have been demonstrated and dominated by deep learning based solutions. In this article, for the first time, we provide a comprehensive literature review that helps readers better grasp the research trends and clearly understand the current advances. We investigate the related studies from the design aspects of network architectures, loss functions, benchmark datasets, and learning strategies with a proposal of a novel taxonomy that categorizes existing methods. Besides, we present a quantitative comparison of model performance on three widely used benchmarks, including indoor and outdoor datasets. Finally, we discuss the challenges of prior works and provide readers with some insights for future research directions. Junjie Hu 0003, Chenyu Bao, Mete Ozay, Chenyou Fan, Qing Gao 0002, Honghai Liu 0001, Tin Lun Lam |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Asymmetric Self-Play-Enabled Intelligent Heterogeneous Multirobot Catching System Using Deep Multiagent Reinforcement LearningabstractAiming to develop a more robust and intelligent heterogeneous system for adversarial catching in security and rescue tasks, in this article, we discuss the specialities of applying asymmetric self-play and curriculum learning techniques to deal with the increasing heterogeneity and number of different robots in modern heterogeneous multirobot systems (HMRS). Our method, based on actor-critic multiagent reinforcement learning, provides a framework that can enable cooperative behaviors among heterogeneous multirobot teams. This leads to the development of an HMRS for complex catching scenarios that involve several robot teams and real-world constraints. We conduct simulated experiments to evaluate different mechanisms' influence on our method's performance, and real-world experiments to assess our system's performance in complex real-world catching problems. In addition, a bridging study is conducted to compare our method with a state-of-the-art method called S2M2 in heterogeneous catching problems, and our method performs better in adversarial settings. As a result, we show that the proposed framework, through fusing asymmetric self-play and curriculum learning during training, is able to successfully complete the HMRS catching task under realistic constraints in both simulation and the real world, thus providing a direction for future large-scale intelligent security & rescue HMRS. Yuan Gao 0024, Xi Chen 0051, Junjie Hu 0003, Fuqin Deng, Tin Lun Lam |
IEEE Trans. Robotics | 5 |
| 2022 | Private Semi-Supervised Federated LearningabstractWe study a federated learning (FL) framework to effectively train models from scarce and skewly distributed labeled data. We consider a challenging yet practical scenario: a few data sources own a small amount of labeled data, while the rest mass sources own purely unlabeled data. Classical FL requires each client to have enough labeled data for local training, thus is not applicable in this scenario. In this work, we design an effective federated semi-supervised learning framework (FedSSL) to fully leverage both labeled and unlabeled data sources. We establish a unified data space across all participating agents, so that each agent can generate mixed data samples to boost semi-supervised learning (SSL), while keeping data locality. We further show that FedSSL can integrate differential privacy protection techniques to prevent labeled data leakage at the cost of minimum performance degradation. On SSL tasks with as small as 0.17% and 1% of MNIST and CIFAR-10 datasets as labeled data, respectively, our approach can achieve 5-20% performance boost over the state-of-the-art methods. Chenyou Fan, Junjie Hu 0003, Jianwei Huang 0001 |
IJCAI | 2 |
| 2021 | PLNet: Plane and Line Priors for Unsupervised Indoor Depth EstimationabstractUnsupervised learning of depth from indoor monocular videos is challenging as the artificial environment contains many textureless regions. Fortunately, the indoor scenes are full of specific structures, such as planes and lines, which should help guide unsupervised depth learning. This paper proposes PLNet that leverages the plane and line priors to enhance the depth estimation. We first represent the scene geometry using local planar coefficients and impose the smoothness constraint on the representation. Moreover, we enforce the planar and linear consistency by randomly selecting some sets of points that are probably coplanar or collinear to construct simple and effective consistency losses. To verify the proposed method’s effectiveness, we further propose to evaluate the flatness and straightness of the predicted point cloud on the reliable planar and linear regions. The regularity of these regions indicates quality indoor reconstruction. Experiments on NYU Depth V2 and ScanNet show that PLNet outperforms existing methods. The code is available at https://github.com/HalleyJiang/PLNet. Hualie Jiang, Laiyan Ding, Junjie Hu 0003, Rui Huang 0001 |
3DV | 3 |
| 2021 | FEANet: Feature-Enhanced Attention Network for RGB-Thermal Real-time Semantic SegmentationabstractThe RGB-Thermal (RGB-T) information for semantic segmentation has been extensively explored in recent years. However, most existing RGB-T semantic segmentation usually compromises spatial resolution to achieve real-time inference speed, which leads to poor performance. To better extract detail spatial information, we propose a two-stage Feature-Enhanced Attention Network (FEANet) for the RGB-T semantic segmentation task. Specifically, we introduce a Feature-Enhanced Attention Module (FEAM) to excavate and enhance multi-level features from both the channel and spatial views. Benefited from the proposed FEAM module, our FEANet can preserve the spatial information and shift more attention to high-resolution features from the fused RGB-T images. Extensive experiments on the urban scene dataset demonstrate that our FEANet outperforms other state-of-the-art (SOTA) RGB-T methods in terms of objective metrics and subjective visual comparison (+2.6% in global mAcc and +0.8% in global mIoU). For the 480 × 640 RGB-T test images, our FEANet can run with a real-time speed on an NVIDIA GeForce RTX 2080 Ti card. Fuqin Deng, Mingjian Liang, Hongmin Wang, Yuan Gao 0024, Junjie Hu 0003, Xiyue Guo, Tin Lun Lam |
IROS | 8 |
| 2021 | Few-Shot Multi-Agent PerceptionabstractWe study few-shot learning (FSL) under multi-agent scenarios, in which participating agents only have local scarce labeled data and need to collaborate to predict query data labels. Though each of the agents, such as drones and robots, has minimal communication and computation capability, we aim at designing coordination schemes such that they can collectively perceive the environment accurately and efficiently. We propose a novel metric-based multi-agent FSL framework which has three main components: an efficient communication mechanism that propagates compact and fine-grained query feature maps from query agents to support agents; an asymmetric attention mechanism that computes region-level attention weights between query and support feature maps; and a metric-learning module which calculates the image-level relevance between query and support data fast and accurately. Through analysis and extensive numerical studies, we demonstrate that our approach can save communication and computation costs and significantly improve performance in both visual and acoustic perception tasks such as face identification, semantic segmentation, and sound genre recognition. Chenyou Fan, Junjie Hu 0003, Jianwei Huang 0001 |
ACM Multimedia | 2 |
| 2020 | Extending information maximization from a rate-distortion perspective
Yan Zhang 0055, Junjie Hu 0003, Takayuki Okatani |
Neurocomputing | 2 |
| 2019 | Visualization of Convolutional Neural Networks for Monocular Depth EstimationabstractRecently, convolutional neural networks (CNNs) have shown great success on the task of monocular depth estimation. A fundamental yet unanswered question is: how CNNs can infer depth from a single image. Toward answering this question, we consider visualization of inference of a CNN by identifying relevant pixels of an input image to depth estimation. We formulate it as an optimization problem of identifying the smallest number of image pixels from which the CNN can estimate a depth map with the minimum difference from the estimate from the entire image. To cope with a difficulty with optimization through a deep CNN, we propose to use another network to predict those relevant image pixels in a forward computation. In our experiments, we first show the effectiveness of this approach, and then apply it to different depth estimation networks on indoor and outdoor scene datasets. The results provide several findings that help exploration of the above question. Junjie Hu 0003, Yan Zhang 0055, Takayuki Okatani |
ICCV | 1 |
| 2019 | Revisiting Single Image Depth Estimation: Toward Higher Resolution Maps With Accurate Object BoundariesabstractThis paper considers the problem of single image depth estimation. The employment of convolutional neural networks (CNNs) has recently brought about significant advancements in the research of this problem. However, most existing methods suffer from loss of spatial resolution in the estimated depth maps; a typical symptom is distorted and blurry reconstruction of object boundaries. In this paper, toward more accurate estimation with a focus on depth maps with higher spatial resolution, we propose two improvements to existing approaches. One is about the strategy of fusing features extracted at different scales, for which we propose an improved network architecture consisting of four modules: an encoder, decoder, multi-scale feature fusion module, and refinement module. The other is about loss functions for measuring inference errors used in training. We show that three loss terms, which measure errors in depth, gradients and surface normals, respectively, contribute to improvement of accuracy in an complementary fashion. Experimental results show that these two improvements enable to attain higher accuracy than the current state-of-the-arts, which is given by finer resolution reconstruction, for example, with small objects and object boundaries. Junjie Hu 0003, Mete Ozay, Yan Zhang 0055, Takayuki Okatani |
WACV | 1 |
| 2017 | NON-rigid structure from motion via sparse self-expressive representationabstractTo simultaneously recover 3D shapes of non-rigid object and camera motions from 2D corresponding points is a difficult task in computer vision. This task is called Non-rigid Structure from motion(NRSfM). To solve this ill-posed problem, many existing methods rely on low rank assumption. However, the value of rank has to be accurately predefined because incorrect value can largely degrade the reconstruction performance. Unfortunately, these is no automatic solution to determine this value. In this paper, we present a self-expressive method that models 3D shapes with a sparse combination of other 3D shapes from the same structure. One of the biggest advantages is that it doesn't need the rank to be predefined. Also, unlike other learning-based methods, our method doesn't need learning step. Experimental results validate the efficiency of our method. Junjie Hu 0003, Terumasa Aoki |
ICIP | 1 |