VLDB 2026 Research / reviewers in the wild / expert
Hongwei Fan
dblp:21/10174
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation ModelabstractExisting vision-and-language navigation models often deviate from the correct trajectory when executing instructions. However, these models lack effective error correction capability, hindering their recovery from errors. To address this challenge, we propose Self-correction Flywheel, a novel post-training paradigm. Instead of considering the model’s error trajectories on the training set as a drawback, our paradigm emphasizes their significance as a valuable data source. We have developed a method to identify deviations in these error trajectories and devised innovative techniques to automatically generate self-correction data for perception and action. These self-correction data serve as fuel to power the model’s continued training. The brilliance of our paradigm is revealed when we re-evaluate the model on the training set, uncovering new error trajectories. At this time, the self-correction flywheel begins to spin. Through multiple flywheel iterations, we progressively enhance our monocular RGB-based VLA navigation model CorrectNav. Experiments on R2R-CE and RxR-CE benchmarks show CorrectNav achieves new state-of-the-art success rates of 65.1% and 69.3%, surpassing prior best VLA navigation models by 8.2% and 16.4%. Real robot tests in various indoor and outdoor environments demonstrate \method's superior capability of error correction, dynamic obstacle avoidance, and long instruction following. Yuxing Long, Chengyan Zeng, Hongwei Fan, Jiyao Zhang, Hao Dong 0003 |
AAAI | 5 |
| 2026 | A multi-source domain-invariant acoustic feature extraction network for rotating machinery fault diagnosis under unknown cross-working conditions
Xiangang Cao, Hongwei Fan, Fuyuan Zhao |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Interacted Object Grounding in Spatio-Temporal Human-Object InteractionsabstractSpatio-temporal Human-Object Interaction (ST-HOI) understanding aims at detecting HOIs from videos, which is crucial for activity understanding. However, existing whole-body-object interaction video benchmarks overlook the truth that open-world objects are diverse, that is, they usually provide limited and predefined object classes. Therefore, we introduce a new open-world benchmark: Grounding Interacted Objects (GIO) including 1,098 interacted objects class and 290K interacted object boxes annotation. Accordingly, an object grounding task is proposed expecting vision systems to discover interacted objects. Even though today’s detectors and grounding methods have succeeded greatly, they perform unsatisfactorily in localizing diverse and rare objects in GIO. This profoundly reveals the limitations of current vision systems and poses a great challenge. Thus, we explore leveraging spatio-temporal cues to address object grounding and propose a 4D question-answering framework (4D-QA) to discover interacted objects from diverse videos. Our method demonstrates significant superiority in extensive experiments compared to current baselines. Xiaoyang Liu 0014, Boran Wen, Xinpeng Liu 0002, Zizheng Zhou, Hongwei Fan, Cewu Lu, Lizhuang Ma, Yong-Lu Li 0001 |
AAAI | 5 |
| 2025 | BiAssemble: Learning Collaborative Affordance for Bimanual Geometric AssemblyabstractShape assembly, the process of combining parts into a complete whole, is a crucial skill for robots with broad real-world applications. Among the various assembly tasks, geometric assembly—where broken parts are reassembled into their original form (e.g., reconstructing a shattered bowl)—is particularly challenging. This requires the robot to recognize geometric cues for grasping, assembly, and subsequent bimanual collaborative manipulation on varied fragments. In this paper, we exploit the geometric generalization of point-level affordance, learning affordance aware of bimanual collaboration in geometric assembly with long-horizon action sequences. To address the evaluation ambiguity caused by geometry diversity of broken parts, we introduce a real-world benchmark featuring geometric variety and global reproducibility. Extensive experiments demonstrate the superiority of our approach over both previous affordance-based and imitation-based methods. Yan Shen 0035, Ruihai Wu, Yubin Ke, Xiaoqi Li 0020, Hongwei Fan, Hao Dong 0003 |
ICML | 7 |
| 2025 | SimLauncher: Launching Sample-Efficient Real-World Robotic Reinforcement Learning via Simulation Pre-TrainingabstractAutonomous learning of dexterous, long-horizon robotic skills has been a longstanding pursuit of embodied AI. Recent advances in robotic reinforcement learning (RL) have demonstrated remarkable performance and robustness in real-world visuomotor control tasks. However, applying RL in the real world faces challenges such as low sample efficiency, slow exploration, and significant reliance on human intervention. In contrast, simulators offer a safe and efficient environment for extensive exploration and data collection, while the visual sim-to-real gap, often a limiting factor, can be mitigated using real-to-sim techniques. Building on these, we propose SimLauncher, a novel framework that combines the strengths of real-world RL and real-to-sim-to-real approaches to overcome these challenges. Specifically, we first pre-train a visuomotor policy in the digital twin simulation environment, which then benefits real-world RL in two ways: (1) bootstrapping target values using extensive simulated demonstrations and real-world demonstrations derived from pre-trained policy rollouts, and (2) Incorporating action proposals from the pre-trained policy for better exploration. We conduct comprehensive experiments across multi-stage, contact-rich, and dexterous hand manipulation tasks. Compared to prior real-world RL approaches, SimLauncher significantly improves sample efficiency and achieves near-perfect success rates. We hope this work serves as a proof of concept and inspires further research on leveraging large-scale simulation pre-training to benefit real-world robotic RL. Mingdong Wu, Lehong Wu, Yizhuo Wu, Weiyao Huang, Hongwei Fan, Zheyuan Hu 0003, Jinzhou Li, Jiahe Ying, Yuanpei Chen, Hao Dong 0003 |
IROS | 5 |
| 2025 | A prototype-guided federated learning based fault diagnosis method of mechanical transmission system under label distribution skew
Hongwei Fan, Shenglin Liu, Xiangang Cao |
Neurocomputing | 1 |
| 2024 | Vehicle-Based Machine Vision Approaches in Intelligent Connected SystemabstractThe application of machine vision techniques in Vehicle-to-Everything (V2X) scenarios within Intelligent Connected Systems (ICS) has gained increasing importance with advancements in 6G communication technology. However, the stringent latency and bandwidth requirements of most machine vision applications pose significant challenges to the existing infrastructure. Hence, there is a dearth of prior research examining whether the latency of real applications in ICS aligns with the needs of machine vision scenarios, let alone any performance evaluations conducted in this regard. In this paper, we conduct a comprehensive literature review and proposed a novel machine vision architecture that can analyze traffic data in real-time in the V2X scenario within ICS. Furthermore, based on the end-to-end latency assessment of the system, we outline a plan to optimize the latency as per the requirements of the machine vision application. Our findings show that with appropriate algorithms and architecture, the ICS system can meet the stringent needs of machine vision applications. Our research can provide valuable insights as a guideline on ICS with high latency requirements and therefore pave the way for future explorations in this field. Chendong Ma, Hongwei Fan, Xing Wu 0001, Tuo Sun |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Research on Dynamic Characteristics of Fuel Cell Stack with a Hybrid Quasi-2D Equivalent Circuit ModelabstractThe degradation issues significantly restrict the automotive application of proton exchange membrane fuel cell (PEMFC), while the dynamic characteristics of PEMFC stack with multiple cells play a critical role in fuel cell's lifetime and reliability. But it is challenging to efficiently and accurately predict the multi-cell voltages and performance uniformity under dynamic fuel cell operations due to the complex structure and multi-physical processes inside the stack. In this paper, we develop a hybrid quasi-2D equivalent circuit model to study the nonuniform dynamic characteristics of PEMFC stack. The 140-cell PEMFC stack with active area of 406 cm2are experimentally studied under various load changing processes to validate the proposed model. The proposed multi-cell equivalent circuit network could reproduce the dynamic cell voltage distributions with the error less than 15 mV (~2.1%) under varied operating conditions. The efficient prediction of dynamic cell voltage distributions of the stack is beneficial for control strategy optimization of fuel cell loading operations and improved stack lifetime. Cong Yin, Yuqin Luo, Guangyou Xie, Renkang Wang, Hongwei Fan |
IECON | 7 |
| 2023 | BioMassters: A Benchmark Dataset for Forest Biomass Estimation using Multi-modal Satellite Time-seriesabstractAbove Ground Biomass is an important variable as forests play a crucial role in mitigating climate change as they act as an efficient, natural and cost-effective carbon sink. Traditional field and airborne LiDAR measurements have been proven to provide reliable estimations of forest biomass. Nevertheless, the use of these techniques at a large scale can be challenging and expensive. Satellite data have been widely used as a valuable tool in estimating biomass on a global scale. However, the full potential of dense multi-modal satellite time series data, in combination with modern deep learning approaches, has yet to be fully explored. The aim of the "BioMassters" data challenge and benchmark dataset is to investigate the potential of multi-modal satellite data (Sentinel-1 SAR and Sentinel-2 MSI) to estimate forest biomass at a large scale using the Finnish Forest Centre's open forest and nature airborne LiDAR data as a reference. The performance of the top three baseline models shows the potential of deep learning to produce accurate and higher-resolution biomass maps. Our benchmark dataset is publically available at https://huggingface.co/datasets/nascetti-a/BioMassters (doi:10.57967/hf/1009) and the implementation of the top three winning models are available at https://github.com/drivendataorg/the-biomassters. Andrea Nascetti, Ritu Yadav, Kirill Brodt, Qixun Qu, Hongwei Fan, Iurii Shendryk, Isha Shah, Christine Chung 0002 |
NeurIPS | 5 |
| 2023 | An Intelligent Diagnosis Approach Combining Resampling and CWGAN-GP of Single-to-Mixed Faults of Rolling Bearings Under Unbalanced Small SamplesabstractRolling bearing is a key component with the high fault rate in the rotary machines, and its fault diagnosis is important for the safe and healthy operation of the entire machine. In recent years, the deep learning has been widely used for the mechanical fault diagnosis. However, in the process of equipment operation, its state data always presents unbalanced. Number of effective data in different states is different and usually the gap is large, which makes it difficult to directly conduct deep learning. This paper proposes a new data enhancement method combining the resampling and Conditional Wasserstein Generative Adversarial Networks-Gradient Penalty (CWGAN-GP), and uses the gray images-based Convolutional Neural Network (CNN) to realize the intelligent fault diagnosis of rolling bearings. First, the resampling is used to expand the small number of samples to a large level. Second, the conditional label in Conditional Generative Adversarial Networks (CGAN) is combined with WGAN-GP to control the generated samples. Meanwhile, the Maximum Mean Discrepancy (MMD) is used to filter the samples to obtain the high-quality expanded data set. Finally, CNN is used to train the obtained dataset and carry out the fault classification. In the experiment, a single, compound and mixed fault cases of rolling bearings are successively simulated. For each case, the different sets considering the imbalance ratio of data are constructed, respectively. The results show that the method proposed significantly improves the fault diagnosis accuracy of rolling bearings, which provides a feasible way for the intelligent diagnosis of mechanical component with the complex fault modes and unbalanced small data. Hongwei Fan, Jiateng Ma, Xiangang Cao, Qinghua Mao |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2023 | Reducing environment exposure to COVID-19 by IoT sensing and computing with deep learning
Chendong Ma, Hongwei Fan, Xing Wu 0001, Tuo Sun, Jiemin Xie |
Neural Comput. Appl. | 4 |
| 2023 | ST-Bikes: Predicting Travel-Behaviors of Sharing-Bikes Exploiting Urban Big DataabstractWith the development of the modern smart city, sharing-bikes require behaviors prediction for grid-level areas which is essential for intelligent transportation systems. A model which can predict bike sharing demand behaviours accurately can allocate sharing-bikes in advance to satisfy travel demands alongside saving energy, reducing traffic, cutting down waste for those sharing-bikes companies putting excessive sharing-bikes in unsaturated demand areas. In this paper, we abandon the traditional time series prediction method and use a more efficient deep learning method to solve the traffic forecasting problem. Moreover, instead of considering spatial relation and temporal relation relatively, we produced a deep multi-view spatial-temporal network to combine them into one prediction model framework. In the experimental section, we investigate in the experiment on enormous amount of real sharing-bikes application use data in the core region of Beijing to test the performance of the model framework with a 1 km$\times $1 km grid-level scale and compare it with other existing machine learning approaches and prediction models. And the 4G/5G/6G communication technology facilitate the real-time control of the space-time locations of sharing bikes dynamically. Thus, it provides the basis for high-frequency analysis of space-time patterns, especially supported by the 6G large-scale application in the future. Jun Chai, Hongwei Fan, Le Zhang 0004, Bing Guo 0003, Yawen Xu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Intelligent Wear Debris Identification of Gearbox Based on Virtual Ferrographic Images and Two-Level Transfer LearningabstractFerrography analysis is one of main means to identify wear state of mechanical equipment, and its key is the intelligent recognition of wear debris ferrographic images. Ferrographic image acquisition is a complex and time-consuming work, so the direct deep learning cannot been carried out for the small tested samples. A virtual ferrographic image dataset is prepared firstly and then two-level transfer learning scheme is proposed to improve the identification rate of the tested samples based on the deep learning model trained by the virtual samples. A combined network of YOLOv3 and DarkNet53 is constructed, and the application effect of model is improved by two-level transfer learning of virtual dataset to open dataset and then open dataset to tested dataset, and the model errors before and after twice transfer learning are analyzed. The average identification accuracy of the model in the validation dataset is 86.1%, which is 44.5% higher than that without two-level transfer learning, and the average recall reaches 95.8%. The experimental results prove the proposed method have a high identification rate for the tested ferrographic images of an actual gearbox. Hongwei Fan, Shuoqi Gao, Ningge Ma, Xiangang Cao |
Int. J. Pattern Recognit. Artif. Intell. | 1 |