EDBT 2026 Demo / reviewers in the wild / expert
Haonan Luo 0002
dblp:175/7532-2
· DBLP profile ↗
22ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0002-9121-2687ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SparseSurf: Sparse-View 3D Gaussian Splatting for Surface ReconstructionabstractRecent advances in optimizing Gaussian Splatting for scene geometry have enabled efficient reconstruction of detailed surfaces from images. However, when input views are sparse, such optimization is prone to overfitting, leading to suboptimal reconstruction quality. Existing approaches address this challenge by employing flattened Gaussian primitives to better fit surface geometry, combined with depth regularization to alleviate geometric ambiguities under limited viewpoints. Nevertheless, the increased anisotropy inherent in flattened Gaussians exacerbates overfitting in sparse-view scenarios, hindering accurate surface fitting and degrading novel view synthesis performance. In this paper, we propose SparseSurf, a method that reconstructs more accurate and detailed surfaces while preserving high-quality novel view rendering. Our key insight is to introduce Stereo Geometry-Texture Alignment, which bridges rendering quality and geometry estimation, thereby jointly enhancing both surface reconstruction and view synthesis. In addition, we present a Pseudo-Feature Enhanced Geometry Consistency that enforces multi-view geometric consistency by incorporating both training and unseen views, effectively mitigating overfitting caused by sparse supervision. Extensive experiments on the DTU, BlendedMVS, and Mip-NeRF360 datasets demonstrate that our method achieves the state-of-the-art performance. Meiying Gu, Jiahe Li 0007, Xiaohan Yu 0001, Haonan Luo 0002, Xiao Bai 0001 |
AAAI | 5 |
| 2026 | Bidirectional chain-of-thought for zero-shot object navigation
Haonan Luo 0002, Yijie Zeng, Zihang Wang 0002, Botao Jiang, Xiruo Jiang |
Frontiers Comput. Sci. | 1 |
| 2026 | CurvLoc: Surface Curvature Prompted Gaussian Splatting for Visual Localization
Jiahe Li 0007, Botao Jiang, Zihang Wang 0002, Xiaohan Yu 0001, Xiao Bai 0001, Haonan Luo 0002 |
Int. J. Comput. Vis. | 9 |
| 2026 | Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization
Muhammad Kashif Ali, Eun Woo Im, Dongjin Kim 0004, Tae Hyun Kim 0006, Haonan Luo 0002, Tianrui Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | EIK-Nav: Boosting zero-shot object navigation with explicit and implicit knowledge
Botao Jiang, Haonan Luo 0002, Zihang Wang 0002, Jiahe Li 0007, Xiao Bai 0001 |
Pattern Recognit. | 2 |
| 2025 | Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score CollaborationabstractUnderstanding how humans cooperatively utilize semantic knowledge to explore unfamiliar environments and decide on navigation directions is critical for house service multi-robot systems. Previous methods primarily focused on single-robot centralized planning strategies, which severely limited exploration efficiency. Recent research has considered decentralized planning strategies for multiple robots, assigning separate planning models to each robot, but these approaches often overlook communication costs. In this work, we propose Multimodal Chain-of-Thought Co-Navigation (MCoCoNav), a modular approach that utilizes multimodal Chain-of-Thought to plan collaborative semantic navigation for multiple robots. MCoCoNav combines visual perception with Vision Language Models (VLMs) to evaluate exploration value through probabilistic scoring, thus reducing time costs and achieving stable outputs. Additionally, a global semantic map is used as a communication bridge, minimizing communication overhead while integrating observational results. Guided by scores that reflect exploration trends, robots utilize this map to assess whether to explore new frontier points or revisit history nodes. Experiments on HM3D_v0.2 and MP3D demonstrate the effectiveness of our approach. Zhixuan Shen, Haonan Luo 0002, Kexun Chen, Fengmao Lv, Tianrui Li 0001 |
AAAI | 2 |
| 2025 | A Continual Learning Approach for Embodied Question Answering with Generative Adversarial Imitation LearningabstractEmbodied Question Answering (EQA) is a task in artificial intelligence where an intelligent agent is required to answer questions about its environment. For example, to answer a question such as "Is the TV on or off?", the agent must navigate to the room with the TV and answer with either "On." or "Off." after recognizing the status. Unlike traditional question-answering systems that rely solely on text or static images, EQA involves agents that can move through a physical or simulated space, interact with the environment, and gather information to respond accurately. The agent must interpret both visual and linguistic inputs, navigate the environment, and complete tasks or locate objects based on the user’s questions. However, in the real world, the agent always faces unseen environments (i.e. different people’s houses), which makes the pre-trained model fail. Meanwhile, re-training in an unseen environment can cause high costs. Therefore, it is significant for the agent to learn continually by itself to cope with the challenges of unseen environments. In this work, we proposed a continual learning method based on generative adversarial imitation learning and self-supervision to support the agent when facing unseen environments. Besides, we designed a policy generator and policy quality discriminator to generate action policy sequences and evaluate the quality of the policy, respectively. Extensive experiments on the MP3D-EQA dataset demonstrate that our method reaches state-of-the-art performance. Haonan Luo 0002, Zihang Wang 0002, Zhixuan Shen, Tianrui Li 0001 |
ICASSP | 2 |
| 2025 | KLFormer: Karhunen-Loève Transform for Robust 3D Human Pose EstimationabstractIn the scope of 3D human pose estimation, the task encompasses estimating the 3D positions of key skeletal points (i.e., wrists, elbows, and knees) from a 2D image or video sequence. This technology demonstrates widespread applicability across diverse domains, encompassing domains such as kinematic analysis, virtual reality, augmented reality, and medical imaging analysis. The common approach is divided into two stages: i) 2D Keypoint Detection: Detecting 2D keypoints from images. ii) 2D-to-3D Lifting: Converting 2D keypoints into 3D coordinates. Present research predominantly concentrates on Stage 2 and leverages sophisticated deep learning architectures, notably Transformers, yet two key challenges exist: i) suboptimal performance in predicting local details; ii) susceptibility to noise interference. In this work, we propose a time-principal component fusion model with limb segment property tracking is proposed to address the challenges above. By integrating timedomain features with principal components extracted through the Karhunen-Loève Transform (KLT), the model aims to address challenges related to feature extraction and noise reduction. Furthermore, to address the issue of suboptimal performance in predicting local details, we devise a property transformer to track the lengths of limb segments and predict the fixed property. Extensive experiments demonstrate that KLFormer showcases state-of-the-art performance on the standard benchmark dataset, Human3.6M. Haonan Luo 0002, Zihang Wang 0002, Leyu Zhang, Tianrui Li 0001 |
ICASSP | 2 |
| 2025 | MoPE: Mixture of Policy Experts and Verification with Multimodal Information for Instance ImageGoal NavigationabstractInstance ImageGoal Navigation (IIN) entails an agent autonomously seeking out a specific object instance depicted by a goal image in an unknown environment. While Large Language Models (LLMs) have shown promise in navigation tasks similar to IIN, their application to IIN remains unexplored. Furthermore, existing LLM-based exploration faces challenges such as inaccurate reasoning due to informative environmental information available to the agent, especially in the early episode stages, and the inability of reinforcement learning(RL) exploration to fully leverage gathered information. Moreover, previous IIN methods did not productively verify potentially distant goal objects discovered during exploration. This work proposes MoPE–Mixture of Policy Experts for exploration and potential goal verification with multimodal information when exploring. Specifically, the hybrid exploration policy comprises an LLM and an RL-based Policy Network (RLPN) to generate an exploration goal to explore efficiently. Our MoPE model surpasses prior approaches on the HM3D datasets significantly. Yijie Zeng, Kexun Chen, Zhixuan Shen, Haonan Luo 0002, Tianrui Li 0001 |
ICME | 5 |
| 2025 | Safety-constrained Reinforcement Learning with Interaction-aware for Decision-making of Autonomous DrivingabstractReinforcement learning(RL) has made significant advancements in autonomous driving(AD). However, the stochastic nature of dynamic traffic scenario and the diversity of road type make it challenging for autonomous vehicles to make safe and efficient decisions. To tackle these problems, this paper proposes a novel RL framework that incorporates the motion prediction model to enhance the agent’s decision-making capability. We first utilize Transformer to model driving scenarios and capture interaction-aware relationships between the ego vehicle and scenarios, then design a safety-constraint and integrate it into the Proximal Policy Optimization (PPO) algorithm so as to guarantee the safety and feasibility of the policy. To improve data efficiency and filter noisy samples, we construct a dual network to communicate and guide each other. Experimental results show that compared with popular RL algorithms, our method demonstrates superior performance in success rate, completion time, safety, and data efficiency. Haonan Luo 0002, Honglin Dong, Jianfeng Lu 0003 |
ICME | 2 |
| 2025 | 3D human avatar reconstruction with neural fields: A recent survey
Meiying Gu, Jiahe Li 0007, Haonan Luo 0002, Xiao Bai 0001 |
Image Vis. Comput. | 4 |
| 2025 | Investigating Synthetic-to-Real Transfer Robustness for Stereo Matching and Optical Flow EstimationabstractWith advancements in robust stereo matching and optical flow estimation networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, their robustness can be seriously degraded when fine-tuning them in real-world scenarios. This paper investigates fine-tuning stereo matching and optical flow estimation networks without compromising their robustness to unseen domains. Specifically, we divide the pixels into consistent and inconsistent regions by comparing Ground Truth (GT) with Pseudo Label (PL) and demonstrate that the imbalance learning of consistent and inconsistent regions in GT causes robustness degradation. Based on our analysis, we propose the DKT framework, which utilizes PL to balance the learning of different regions in GT. The core idea is to utilize an exponential moving average (EMA) teacher to measure what the student network has learned and dynamically adjust the learning regions. We further propose the DKT++ framework, which improves target-domain performances and network robustness by applying slow-fast update teachers to generate more accurate PL, introducing the unlabeled data and synthetic data. We integrate our frameworks with state-of-the-art networks and evaluate their effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the robustness of stereo matching and optical flow networks during fine-tuning. Jiahe Li 0007, Lei Huang 0015, Haonan Luo 0002, Xiaohan Yu 0001, Lin Gu 0003, Xiao Bai 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Multi-Scale CNN-Transformer Hybrid Network for Rail Fastener Defect DetectionabstractDefect detection in rail fasteners is crucial for train safety, as defective fasteners can cause derailments and severe safety incidents. However, Existing algorithms often struggle in various real-world scenarios due to challenges such as obscured fasteners, motion blur in images, varying camera angles, and fasteners submerged in water. To address these challenges, we propose a Multi-scale CNN-Transformer Hybrid Network for Rail Fastener Defect Detection (MCHNet-RF2D), specifically designed to identify fastener defects in complex environments. Our approach constructs an efficient CNN block and a multi-scale Vision Transformer block to alternately extract local detail features and global semantic features of the fasteners. These features are seamlessly integrated through multi-scale fusion to enhance defect recognition robustness. By combining comprehensive global recognition with detailed local defect detection, MCHNet-RF2D outperforms existing CNN-Transformer hybrid networks by 2.8% and surpasses current fastener defect detection algorithms by 2.9%. In practical deployment on over 40 trains, our model successfully detected more than 2,000 fastener defects, demonstrating its effectiveness in diverse and challenging conditions. Wei Wang 0278, Fengmao Lv, Haonan Luo 0002, Gexiang Zhang, Zhenghua Chen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Adversarial Training with OCR modality Perturbation for Scene-Text Visual Question AnsweringabstractScene-Text Visual Question Answering (ST-VQA) aims to understand scene text in images and answer questions related to the text content. Most existing methods heavily rely on the accuracy of Optical Character Recognition (OCR) systems, and aggressive fine-tuning based on limited spatial location information and erroneous OCR text information often leads to inevitable overfitting. In this paper, we propose a multimodal adversarial training architecture with spatial awareness capabilities. Specifically, we introduce an Adversarial OCR Enhancement (AOE) module, which leverages adversarial training in the embedding space of OCR modality to enhance fault-tolerant representation of OCR texts, thereby reducing noise caused by OCR errors. Simultaneously, We add a Spatial-Aware Self-Attention (SASA) mechanism to help the model better capture the spatial relationships among OCR tokens. Various experiments demonstrate that our method achieves significant performance improvements on both the ST-VQA and TextVQA datasets and provides a novel paradigm for multimodal adversarial training. Zhixuan Shen, Haonan Luo 0002, Tianrui Li 0001 |
ICME | 2 |
| 2024 | Enhancing Robustness in Learning with Noisy Labels: An Asymmetric Co-Training ApproachabstractLabel noise, an inevitable issue in various real-world datasets, tends to impair the performance of deep neural networks. A large body of literature focuses on symmetric co-training, aiming to enhance model robustness by exploiting interactions between models with distinct capabilities. However, the symmetric training processes employed in existing methods often culminate in model consensus, diminishing their efficacy in handling noisy labels. To this end, we propose an Asymmetric Co-Training (ACT) method to mitigate the detrimental effects of label noise. Specifically, we introduce an asymmetric training framework in which one model (i.e., RTM) is robustly trained with a selected subset of clean samples while the other (i.e., NTM) is conventionally trained using the entire training set. We propose two novel criteria based on agreement and discrepancy between models, establishing asymmetric sample selection and mining. Moreover, a metric, derived from the divergence between models, is devised to quantify label memorization, guiding our method in determining the optimal stopping point for sample mining. Finally, we propose to dynamically re-weight identified clean samples according to their reliability inferred from historical information. We additionally employ consistency regularization to achieve further performance improvement. Extensive experimental results on synthetic and real-world datasets demonstrate the effectiveness and superiority of our method. Mengmeng Sheng, Zeren Sun, Gensheng Pei, Tao Chen 0012, Haonan Luo 0002, Yazhou Yao |
ACM Multimedia | 5 |
| 2024 | VLAI: Exploration and Exploitation based on Visual-Language Aligned Information for Robotic Object Goal Navigation
Haonan Luo 0002, Yijie Zeng, Kexun Chen, Zhixuan Shen, Fengmao Lv |
Image Vis. Comput. | 1 |
| 2024 | Transferring Multi-Modal Domain Knowledge to Uni-Modal Domain for Urban Scene SegmentationabstractSynthetic data (i.e., source domain) have been widely adopted to improve the semantic segmentation performance for real-world images (i.e., target domain), since obtaining pixel-level annotations is fairly easy in the synthetic environment. Traditional domain adaptation methods normally focus on learning in the RGB modality only. We notice that the synthetic environment can generate depth information of semantic objects at almost no cost, while it is nontrivial to collect such information in the real-world scenario. In this case, we employ the depth information of synthetic data in this work to further boost the segmentation performance, and then transform the uni-modal problem into a multi-modal one. In this work, we focus on urban scene understanding and make a pioneer attempt on learning uni-modal feature representations for real-world images by mining from multi-modal knowledge of synthetic images with additional depth information. To this end, we propose a novel method called Multi-modal Domain Knowledge Transfer (MDKT), which transfers the multi-modal knowledge of the source domain to the uni-modal target domain through domain adaptation. In MDKT, we first employ the Cross-Modal Correlation (CMC) module to enhance the source features by fusing the RGB and depth information. Then, the uni-modal target domain feature and multi-modal source domain feature are aligned through the Modal-Imbalanced Adversarial Training (MIAT) strategy, which transfers the multi-modal knowledge to the uni-modal network in the target domain. We conduct extensive experiments on several benchmark settings for urban scene understanding. The promising results clearly show the effectiveness of our proposed MDKT approach. Peng Liu 0049, Yanqi Ge, Lixin Duan, Wen Li 0001, Haonan Luo 0002, Fengmao Lv |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Robust-EQA: Robust Learning for Embodied Question Answering With Noisy LabelsabstractEmbodied question answering (EQA) is a recently emerged research field in which an agent is asked to answer the user's questions by exploring the environment and collecting visual information. Plenty of researchers turn their attention to the EQA field due to its broad potential application areas, such as in-home robots, self-driven mobile, and personal assistants. High-level visual tasks, such as EQA, are susceptible to noisy inputs, because they have complex reasoning processes. Before the profits of the EQA field can be applied to practical applications, good robustness against label noise needs to be equipped. To tackle this problem, we propose a novel label noise-robust learning algorithm for the EQA task. First, a joint training co-regularization noise-robust learning method is proposed for noisy filtering of the visual question answering (VQA) module, which trains two parallel network branches by one loss function. Then, a two-stage hierarchical robust learning algorithm is proposed to filter out noisy navigation labels in both trajectory level and action level. Finally, by taking purified labels as inputs, a joint robust learning mechanism is given to coordinate the work of the whole EQA system. Empirical results demonstrate that, under extremely noisy environments (45% of noisy labels) and low-level noisy environments (20% of noisy labels), the robustness of deep learning models trained by our algorithm is superior to the existing EQA models in noisy environments. Haonan Luo 0002, Guosheng Lin, Fumin Shen, Xingguo Huang, Yazhou Yao, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | ABF-FNN: A new fuzzy neural network for predicting coal mine gas concentration hazardabstractSerious accidents in coal mines are often caused by high concentration of gas. Therefore, accurate prediction of mine gas concentration is vital for ensuring the safety of coal mine production. Hence, a new method for predicting the risk of gas content in underground coal mines (ABFFNN, adaptive bandwidth feedback fuzzy neural network) is proposed in this paper. Namely, 1) Fuzzy theory and neural networks are combined in the method, and feedback mechanisms are incorporated to enhance network learning capability and robustness. 2) This method designs a novel approach to calculate adaptive bandwidths and dynamically adapts to changes in univariate and multivariate time series data. Weight training is performed using particle swarm optimization algorithm. 3) The evaluation of the developed prediction model is conducted using two publicly available datasets as well as actual industrial data collected at a mining site. The public datasets used include the Data of Qinggangping coal mine monitoring system in Shaanxi Province and the Box-Jenkins gas stove dataset. It is shown by the results that good prediction results are achieved by the developed prediction model on two public data sets and one private data set. In addition, it can be seen that the gas concentration hazard in underground coal mines can be better predicted by the method proposed in this paper when compared with other models. Yimin Sun, Yunyang Wu, Haihao Tang, Haonan Luo 0002 |
IEEE Big Data | 6 |
| 2023 | Depth and Video Segmentation Based Visual Attention for Embodied Question AnsweringabstractEmbodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real-world environment. It has attracted increasing research interests due to its broad applications in personal assistants and in-home robots. Most of the existing methods perform poorly in terms of answering and navigation accuracy due to the absence of fine-level semantic information, stability to the ambiguity, and 3D spatial information of the virtual environment. To tackle these problems, we propose a depth and segmentation based visual attention mechanism for Embodied Question Answering. First, we extract local semantic features by introducing a novel high-speed video segmentation framework. Then guided by the extracted semantic features, a depth and segmentation based visual attention mechanism is proposed for the Visual Question Answering (VQA) sub-task. Further, a feature fusion strategy is designed to guide the navigator's training process without much additional computational cost. The ablation experiments show that our method effectively boosts the performance of the VQA module and navigation module, leading to 4.9 % and 5.6 % overall improvement in EQA accuracy on House3D and Matterport3D datasets respectively. Haonan Luo 0002, Guosheng Lin, Yazhou Yao, Fayao Liu, Zichuan Liu, Zhenmin Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Dense Semantics-Assisted Networks for Video Action RecognitionabstractMost existing action recognition approaches directly leverage the video-level features to recognize human actions from videos. Although these methods have made remarkable progress, the accuracy is still unsatisfied. When the test video involves complex backgrounds and activities, existing methods usually suffer from a significant drop in accuracy. Human action is inherently a high-level concept. Merely applying a video classification model without a detailed semantic understanding of the video content, e.g., objects, scene context, object motions, object interactions, is inadequate to tackle the challenges for action recognition. Fine-level semantic understanding of videos generates elementary semantic concepts from the raw video data, such as the semantics of objects and background regions. It can be employed to bridge the gap between the raw video data and the high-level concept of human actions. In this work, we leverage dense semantic segmentation masks, which encode rich semantic details, provide extra information for the network training, and improve the performance of action recognition. We propose a novel deep architecture which is named as Dense Semantics-Assisted Convolutional Neural Networks (DSA-CNNs) to effectively utilize dense semantic information of video by a bottom-up attention way in the spatial stream, while by the way of branch fusion in the temporal stream. To verify the effectiveness of our approach, we conduct extensive experiments on publicly available datasets – UCF101, HMDB51, and Kinetics. The experimental results demonstrate that our approach substantially improves existing methods and achieves very competitive performance. It also shows that our approach is superior to other related methods that utilize extra information for action recognition. Haonan Luo 0002, Guosheng Lin, Yazhou Yao, Zhenmin Tang, Qingyao Wu, Xian-Sheng Hua 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | SegEQA: Video Segmentation Based Visual Attention for Embodied Question AnsweringabstractEmbodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real world environment. It has attracted increasing research interests due to its broad applications in automatic driving system, in-home robots, and personal assistants. Most of the existing methods perform poorly in terms of answering and navigation accuracy due to the absence of local details and vulnerability to the ambiguity caused by complicated vision conditions. To tackle these problems, we propose a segmentation based visual attention mechanism for Embodied Question Answering. Firstly, We extract the local semantic features by introducing a novel high-speed video segmentation framework. Then by the guide of extracted semantic features, a bottom-up visual attention mechanism is proposed for the Visual Question Answering (VQA) sub-task. Further, a feature fusion strategy is proposed to guide the training of the navigator without much additional computational cost. The ablation experiments show that our method boosts the performance of VQA module by 4.2% (68.99% vs 64.73%) and leads to 3.6% (48.59% vs 44.98%) overall improvement in EQA accuracy. Haonan Luo 0002, Guosheng Lin, Zichuan Liu, Fayao Liu, Zhenmin Tang, Yazhou Yao |
ICCV | 1 |