VLDB 2026 Research / reviewers in the wild / expert
Kevin W. Tong
dblp:11/4740-2 · also Wei Tong 0002
· DBLP profile ↗
18ranked-venue papers
7as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 7 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neural Rendering and Flow-Assisted Unsupervised Multi-View Stereo for Real-Time Monocular Tracking and Scene PerceptionabstractThe existing camera tracking and perception methods mainly rely on sparse SLAM, which limits the dense perception ability of the scene and affects the reliability of auxiliary decision-making. Different from this, this work proposes a real-time tracking and unsupervised dense sensing framework. Firstly, the dense depth value of the scene is predicted by unsupervised multi-view stereo to remove the dependence on labeled data. Then, the quality of synthetic pseudo-reference image is quantified according to the predicted depth map and used as a weighted guidance to train the unsupervised model, thus reducing the ambiguity of feature matching in areas such as specular reflection. Moreover, the sparse optical flow of the keyframes is solved by real-time and robust ORB feature matching operator, which assists the high-precision training of unsupervised depth inference model. To increase the prediction accuracy of occluded area, a novel rendering consistency loss via neural radiance fields is designed to constrain the geometric characteristics of object surface. Finally, dense direct image alignment is performed from a global model to improve the tracking robustness, which is incrementally constructed from dense depth prediction. Extensive experiments on synthetic datasets and real datasets validate the effectiveness and practicability of the proposed work, which is an effective supplement to the existing SLAM work. Kevin W. Tong, Yandong Cai, Yu-Wen Jie, Ya Duan, Yuhong Hou, Qi Wu 0003 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Inferring Pilot Cognitive States Using a Multimodal Physiology-Based Graph Fusion ModelabstractAccurate recognition of pilots' cognitive states is essential for intelligent cockpit systems operating in complex and high-load flight environments. Existing methods either rely on single-modal physiological signals or fail to fully exploit the intrinsic temporal, spatial, and cross-modal structures of multimodal data. To address these limitations, we propose a multimodal physiology-based graph fusion (MPGFusion) network, an end-to-end deep framework that explicitly models the hierarchical graph structure and spatiotemporal dependencies of multimodal physiological signals. Unlike conventional approaches that depend on hand-crafted features or independent modality-specific encoders, MPGFusion directly takes raw multimodal physiological time series as input and jointly captures first, multiscale temporal dynamics via parallel temporal kernels, second, intramodality spatial relationships through single-modal graphs, and third, intermodality heterogeneous interactions using a multimodal graph integrated with graph transformation network, graph convolutional network, and gated recurrent unit. This hierarchical graph-based design enables structured cross-modal feature interaction while preserving modality-specific characteristics. Experiments conducted on the CogPilot dataset demonstrate that MPGFusion significantly outperforms representative baselines, achieving 97.49% classification accuracy. Extensive ablation studies validate the contribution of each architectural component and highlight the complementary nature of multimodal physiological signals. The proposed framework provides a generalizable and effective solution for multimodal cognitive state recognition in aviation scenarios. Shuai Wu 0004, Xuefeng Men, Kevin W. Tong |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2026 | Multimodal Feature Interaction and High-Quality Pseudolabel Generation With Self-Training for Cognitive State DetectionabstractCognitive state detection holds significant research value in the field of human–computer interaction and neural engineering. However, existing works are insufficient in modeling the temporal dynamics of multimodal physiological signals, which leads to heterogeneous distribution differences in cross-modal feature interactions. In addition, domain shift issues under cross-subject and few-sample conditions restrict the model generalization performance. To cope with these problems, this work proposes a cognitive state detection framework that integrates Transformer-based multimodal feature interaction and self-training of pseudolabel optimization. First, the multihead attention mechanism is introduced to model the temporal evolution patterns across modalities, dynamically harmonizing cross-modal contributions to extract cognitive state-related shared features. Then, a dual-model cross-validation strategy is designed to filter high-quality pseudolabeled samples from the target domain for subsequent self-training, effectively avoiding the dependency on auxiliary modules in domain adaptation. Finally, Extensive experiments show that the proposed work significantly improves the recognition accuracy, and the designed pseudolabel optimization mechanism can be transferred to related tasks without increasing model complexity. Kevin W. Tong, Xuefeng Men, Haoran Duan 0001, Shaojun Cai, Changyu Li, Ping Li 0044, Guangyu Zhu 0001, Qi Wu 0003, Limin Zhu 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | FMDConv: Fast multi-attention dynamic convolution via speed-accuracy trade-offabstractSpatial convolution is fundamental in constructing deep Convolutional Neural Networks (CNNs) for visual recognition. While dynamic convolution enhances model accuracy by adaptively combining static kernels, it incurs significant computational overhead, limiting its deployment in resource-constrained environments such as federated edge computing. To address this, we propose Fast Multi-Attention Dynamic Convolution (FMDConv), which integrates input attention, temperature-degraded kernel attention, and output attention to optimize the speed-accuracy trade-off. FMDConv achieves a better balance between accuracy and efficiency by selectively enhancing feature extraction with lower complexity. Furthermore, we introduce two novel quantitative metrics, the Inverse Efficiency Score and Rate-Correct Score, to systematically evaluate this trade-off. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet demonstrate that FMDConv reduces the computational cost by up to 49.8% on ResNet-18 and 42.2% on ResNet-50 compared to prior multi-attention dynamic convolution methods while maintaining competitive accuracy. These advantages make FMDConv highly suitable for real-world, resource-constrained applications. This figure presents an overview of the proposed Fast Multi-Attention Dynamic Convolution (FMDConv) framework, which integrates Input Attention, Temperature-Degraded Kernel Attention, and Output Attention to optimize the speed-accuracy trade-off in convolutional neural networks. The diagram illustrates how these mechanisms enhance feature selection at different stages, significantly reducing computational cost while maintaining competitive accuracy, making FMDConv suitable for resource-constrained applications. • Introduces IES and RCS to quantify speed-accuracy trade-off in CNNs. • Evaluates channel, kernel, and filter attention for effective structures. • Develops a new CNN with kernel attention to enhance efficiency and accuracy. • Demonstrates FMDConv’s advantages via standard benchmark testing. Fan Wan, Haoran Duan 0001, Kevin W. Tong, Jingjing Deng 0001, Yang Long 0001 |
Knowl. Based Syst. | 4 |
| 2025 | Semi-Supervised Image Domain Adaption for Aerial Refueling Drogue Detection on Embedded Chip Under Foggy ConditionsabstractThe application of aerial refueling technology to UAVs can reduce the dependence on the pilot’s operation, which has unique advantages in carrying out battlefield reconnaissance, monitoring suspicious targets and collecting intelligence through all-weather work. The existing vision-based drogue detection methods are assumed to be carried out under daily lighting conditions, but special weather, such as fog, makes it difficult to identify the characteristics of the drogue, which will greatly degrade the model performance or even fail. Moreover, the traditional computing architecture is difficult to be directly applied to real airborne equipment, so it is necessary to adopt AI processor module with faster and better computing power and supporting parallel computing to meet the requirements of low delay and high security in aerial refueling. Therefore, this work proposes a robust detection network based on image domain adaption. Firstly, an end-to-end image defogging module is designed to deal with foggy image enhancement under weak supervision. Then, knowledge distillation with the teacher-student network is applied to guide the student model to obtain the instance-level features of the unlabeled target domain. In addition, the detection model is compiled and transplanted on the system-on-chip chip of JFMQL100TAI. The comprehensive experimental results on public datasets and real refueling datasets validate the effectiveness and feasibility of the proposed work, which effectively complements the drogue detection of special autonomous aerial refueling tasks. Note to Practitioners—As a widely used refueling technology in the field of national defense, probe-and-drogue refueling has developed from manual control docking to monitoring auxiliary docking. However, it is difficult for pilots to accurately and quickly obtain the relative position of the refueling drogue through visual perception. In this work, a semi-supervised drogue detection network for special aerial refueling task is designed. The proposed work has good application potential in refueling scenes, which can provide fast and accurate drogue positioning under foggy conditions. Kevin W. Tong, Ai Gu, Xiangyang Deng, Yandong Cai, Ya Duan, Yuhong Hou |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Edge-Assisted Epipolar Transformer for Industrial Scene ReconstructionabstractGiven a set of calibrated images, Multiple View Stereo (MVS) applies end-to-end depth inference network to recover scene structure. However, previous methods designed pixel-visibility modules to aggregate cross-view cost, ignoring the consistency assumption of 2D contextual features in the 3D depth direction. The current multi-stage depth inference model also relies on intensive depth samples, which requires high memory consumption. To alleviate these problems, this work exploits edge-assisted epipolar Transformer for multi-view depth inference. The improvements of this work are summarized as follows: 1) The epipolar Transformer block is developed for reliable cross-view cost aggregation, and the edge detection branch is designed to constrain the consistency of epipolar geometry and edge features. 2) The dynamic depth range sampling mechanism based on probability volume is adopted to improve the accuracy of uncertain areas. Comprehensive comparisons with the state-of-the-art works indicate that our work can reconstruct dense scene representations with limited memory bottleblockNote to Practitioners—Learning-based MVS can obtain dense point clouds with accurate depth map estimation, which are widely applied in the fields of unmanned driving, battlefield environment perception and robot navigation. MVS-based scene reconstruction technology is the premise of the subsequent planning, decision-making and control of the human-machine system. To obtain dense scene representation with limited memory and runtime, this work proposes a multi-view stereo network with edge-assisted epipolar Transformer. Experiments on public benchmarks verify the feasibility and effectiveness of our model, which has good potential in battlefield environment reconstruction and human-computer interaction fields, and can provide intuitive and dense scene representation for decision-making assistance. Kevin W. Tong, Xiaorong Guan, Miaomiao Zhang 0001, Ping Li 0044, Qi Wu 0003, Limin Zhu 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Robust Drogue Positioning System Based on Detection and Tracking for Autonomous Aerial Refueling of UAVsabstractIn modern war, endurance mileage and combat radius are important factors for the combat effectiveness of military aircraft. However, the existing aerial refueling schemes mainly use manual docking operations and complex conditions require pilots to have high operating skills. Therefore, this work proposes a visual positioning system for autonomous aerial refueling without adding cooperation marks, which mainly includes a dynamic graph convolution module for drogue detection and a kernel correlation filter for drogue tracking. Firstly, a dynamic GCN module is designed to generate the correlation matrix to aggregate adjacent high-order features and fuse them with global features extracted from the CNN stream to achieve accurate drogue detection. Then, the multi-scale features extracted from the drogue detector network and the HOG features are input to the filter learning module together, and weighted response maps are fused to alleviate the occlusion and scale change problems in the tracking process. In addition, a visual positioning scheme combining a drogue detector and tracker is introduced to output an accurate drogue ROI area. The effectiveness and robustness of the proposed work are verified by comparison with the mainstream methods on COCO detection datasets and real aerial refueling datasetsNote to Practitioners—For probe-and-drogue refueling, the combination of autonomous aerial refueling technology and machine vision technology can improve the real-time and robustness of autonomous refueling docking system in complex aerial visual scenes. Therefore, this work studies the accurate recognition of drogue detection and real-time drogue tracking. In addition, a drogue ROI positioning system is also designed. Experiments on public datasets and real refueling scenarios validate the feasibility and effectiveness of the proposed work, which has good application potential in refueling scenes and can provide decision-making for autonomous refueling and manual docking refueling of UAVs. Kevin W. Tong, Yu-Hong Hou |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | "Jumpingly" Perceive Time Series: Image Generation Approach to Modeling Functional Brain ActivationabstractThis paper presents a novel Linear Mapping Field (LMF) to map time series into two-dimensional images. The LMF extracts deeper features of fNIRS signals, which makes fNIRS less reliant on some prior. The developed convolution neural networks detect more prominent features than the state-of-the-art methods. The experimental results indicate that different from RNNs which can only perceive the time series in a “sequential” manner, LMF’s characteristic of “jumpingly” perception is the key to achieving excellent results.Note to Practitioners—As an optical and non-invasive technique to obtain the changes of oxyhemoglobin (O2Hb) and deoxyhemoglobin (HHb), fNIRS can be used to measure the changes of cerebral hemodynamics related to brain activities. This work proposes a linear mapping field and other mapping fields to map fNIRS signals to two-dimensional images, and the deep features of these generated images can be further extracted by convolutional neural networks, establishing an end-to-end bridge to detect the activation of brain functions under different tasks. Compared with mainstream methods, this work can be used as an effective mapping for fNIRS with low computational complexity and good performance. Kevin W. Tong, Miaomiao Zhang 0001, Zhiyi Shi, Yuhong Hou |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Adaptive Guidance in Dynamic Environments: A Deep Reinforcement Learning Approach for Highly Maneuvering TargetsabstractIn future battlefields, missiles are expected to become highly precise and efficient strike weapons, with missile intelligence emerging as a critical development trend. To address the problem of optimizing 3-D missile interception guidance laws, this article introduces the deep Q-network (DQN) algorithm on the foundation of proportional navigation guidance (PNG) and proposes an adaptive proportional guidance algorithm based on deep reinforcement learning (DRL). The proposed algorithm uses air combat situational information as the state space and incorporates parameters such as the missile-target relative distance and line-of-sight (LOS) angle into the reward function design. The optimal proportional navigation coefficient$K^{*}$for low-overload maneuvering targets is determined through network search, and the longitudinal and lateral control commands of the missile are decoupled by designing the proportional coefficient increment$\Delta K$, constructing a discretized action space. Simulation results show that, compared to the PNG with a constant$K^{*}$, the proposed method significantly improves the hit probability of high-overload maneuvering targets while maintaining the hit rate for low-overload maneuvering targets. As an exploration of future intelligent combat scenarios, this guidance law design method holds both theoretical significance and practical application value. Longjun Zhu, Yandong Cai, Kevin W. Tong, Shuai Wu 0004, Fengtao Xiang, Ya Duan, Yuhong Hou, Guangyu Zhu 0001, Qi Wu 0003 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | Concept-Aware Entity Alignment Network for Industrial Knowledge GraphabstractThe industrial knowledge graph (IKG) can improve the cognitive intelligence of the manufacturing system and is recognized as one of the cores of the next-generation industrial management information system. Due to the multisource heterogeneous nature of industrial data, aligning entities with the same semantics (entity alignment) is the core technology for building large-scale, high-coverage IKGs. Existing approaches show that embedded learning of IKGs performs well for this task. However, most advanced methods ignore concept information when learning topological information about IKGs. Inspired by the ontology matching theory, in this article, we realize the importance of entity concepts in alignment. The conceptual semantics of entities can usually be obtained through the is–a relation. However, the IKG is usually constructed by triples (entity, relation, entity) automatically extracted from a large text corpus. This will lead to entities in the IKG having problems such as lacking conceptual information, belonging to multiple concepts, or having different concept granularities. To solve the two problems of lacking conceptual information and different concept granularity, we propose the concept-aware entity alignment network (CAEA), aggregating bidirectional relations and attributes to get the entity concept semantics by a novel concept-aware graph attention mechanism. The excellent performance of the CAEA can better support the construction of large and complete IKGs and support downstream applications such as industrial knowledge recommendation and assisted decision-making. To verify the performance of the CAEA on the IKG, we construct a new entity alignment benchmark using industrial control network security data and verify the effectiveness of the CAEA on the new benchmark and several mainstream datasets. Experimental results show that our method outperforms other state-of-the-art (SOTA) methods and promotes the development of IKGs. Shuai Wu 0004, Kevin W. Tong, Yuhong Hou, Ping Li 0044, Weidong Yang 0001, Qi Wu 0003 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | A parallel neural networks for emotion recognition based on EEG signals
Yuwen Jie, Kevin W. Tong, Miaomiao Zhang 0001, Guangyu Zhu 0001, Qi Wu 0003 |
Neurocomputing | 3 |
| 2024 | Robust Depth Estimation Based on Parallax Attention for Aerial Scene PerceptionabstractGiven the precalibrated image pairs, stereo matching aims to infer the scene depth information in real-time, which has important research value in the fields of high-precision 3-D reconstruction of the Earth’s surface, automatic driving and unmanned aerial vehicle (UAV) navigation. The cost volume-based stereo matching method adopts a coarse-to-fine manner to construct cascaded cost volume, and applies 3-D convolution to capture the correspondence of feature matching to infer the disparity map, which achieves comparable performance. However, the existing method has difficulty dealing with jitter regions with disparity change, and direct disparity regression easily leads to overfitting of cost volume regularization. To alleviate the above two problems, this work proposes an end-to-end disparity estimation network based on Transformer. Its specific improvements are as follows. 1) The cross-view feature interaction module based on Transformer is introduced to realize the feature interaction of global context information. 2) A parallax attention mechanism is designed to impose global geometric constraints on the epipolar line to improve the reliability of feature matching. 3) Focal loss is applied for the training of the disparity classification model to emphasize one-hot supervision in ambiguous regions. Comprehensive experiments on public datasets Sceneflow, KITTI2015, ETH3D, and aerial WHU datasets validate that the proposed work can effectively enhance the performance of disparity estimation. Kevin W. Tong, Miaomiao Zhang 0001, Guangyu Zhu 0001, Xin Xu 0001, Qi Wu 0003 |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Cognitive State Detection in Task Context Based on Graph Attention Network During FlightabstractThis work provides a graph network solution for pilot brain fatigue state inference based on electroencephalography (EEG) fatigue indicators. Two graph methods are built as follows. The first one uses a single EEG signal sample as a node, and fatigue detection as a node classification task in a graph network. The developed graph network is then utilized to extract the correlation among different samples to achieve multisample joint decision making. The second method uses a single EEG signal sample as a graph structure, and EEG fatigue prediction as a graph classification task. Electrode position correlation is used to construct a graph. The feature fusion of adjacent electrodes is obtained through the connection relationship among nodes in a graph structure to improve network learning accuracy. In addition, a Bayesian optimization method is proposed to model the randomness of attention weights, and a Bayesian graph attention network is built. This work constructs a based-graph deep learning structures to achieve a pilot fatigue detection model with high accuracy, good generalization, and strong adaptability. Experimental results demonstrate the effectiveness of the proposed model. Qi Wu 0003, Yubing Gao, Kevin W. Tong, Yuhong Hou, Rob Law 0001, Guangyu Zhu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2024 | SQIX: QMIX Algorithm Activated by General Softmax Operator for Cooperative Multiagent Reinforcement LearningabstractMultiagent cooperative systems can be used to conceptualize many real-world problems. Reinforcement learning is a particularly effective tool. The issue of bias in$Q$-function value estimation in single-agent reinforcement learning has garnered a lot of interest and substantial study. Indeed, this challenge endures in multiagent reinforcement learning, primarily owing to the inclusion of maximization operations. The crux of the matter lies in the inability to seamlessly extrapolate single-agent reinforcement learning algorithms to their multiagent counterparts. In this article, we introduce a more encompassing and straightforward principle: the notion of appropriate value correction. We suggest replacing the maximization operation with a monotonically nondecreasing function to obtain more accurate value estimates. We theoretically demonstrate that this operation effectively reduces the potential overestimation bias in the QMIX algorithm. Ultimately, our methodology, dubbed the SMIX algorithm—a fusion of the QMIX algorithm empowered by the Softmax operator, attains state-of-the-art outcomes across diverse multiagent cooperative tasks. This success extends to challenging domains such as StarCraft II, marking it as one of the most formidable games to date. Miaomiao Zhang 0001, Kevin W. Tong, Guangyu Zhu 0001, Xin Xu 0001, Qi Wu 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2023 | Multistage Pixel-Visibility Learning With Cost Regularization for Multiview StereoabstractMultiple-view stereo has potential applications in robotic operations and autonomous driving (unstructured environment construction, visual servo). With assisted depth information, inertial navigation systems can achieve precise navigation. It is, especially suitable for GPS failures in complex environments. Accurate depth estimation is a challenge in low-textured or occluded regions. To alleviate the inference of incorrect depth, a multi-stage pixel-visibility learning-based stereo network is presented in this paper. Its improvements are as follows: 1) a new content-adaptive cost volume aggregation mechanism based on neighboring pixel-wise visibility is designed to effectively produce more accurate and smoother depth map predictions in the object boundary. 2) global convolution block and boundary refinement block are developed to regularize its cost volume, they can learn the inherent constraints of feature matching correspondence and effectively mitigate the depth estimation uncertainty in low-textured regions. 3) a new loss function is designed to measure the uncertainty of predicted probability distribution and enhance the reliability of depth map inference. Experimental results on the indoor DTU datasets and the outdoor Tanks & Temples datasets indicate that our method can achieve superior performance and has a powerful generalization ability, which is comparable to state-of-the-art works. Note to Practitioners—Multiple-view stereo (MVS) can estimate dense 3D representations of scenes, which is widely used in autonomous driving, robotic navigation, virtual reality (VR), and augmented reality (AR). Aiming at the problem of incorrect depth inference in low-textured or occluded regions, this work proposes a novel multi-stage depth prediction method based on neighboring pixel-wise visibility. Our method cannot only achieve accurate depth estimation for robot perception but also make no concession to real-time performance. It is clear that the proposed method has good potential in 3D reconstruction, robotic navigation, and VR/AR fields to provide accurate depth estimation in real-time with limited memory consumption. Xiaorong Guan, Kevin W. Tong, Shan Jiang 0022, Zhao-Hui Sun, Qi Wu 0003, Guimin Chen |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2023 | Anti-Disturbance Path-Following Control for Snake Robots With Spiral MotionabstractThree-dimensional spiral gait enables a snake robot to climb over obstacles, cross caves, and adapt to complex environments. This article reports an antidisturbance path-following control method for a snake robot with a spiral gait. This method reduces the deviation of the robot's position in following the ideal path by estimating the time-varying parameters, the external disturbances, and the viscous friction coefficients. The estimations are used to compensate for the control inputs of the system, which can improve the adaptability of the robot to the environment. Then, the attitude and position errors can rapidly converge to the origin. An appropriate Lyapunov function is adopted to explore the stability of following errors. Experimental results show that the proposed method can accelerate the convergence rate of errors, reduce the fluctuation peak, and improve the following stability of snake robots. Dongfang Li 0001, Kevin W. Tong, Ping Li 0044, Rob Law 0001, Xin Xu 0001, Limin Zhu 0001, Qi Wu 0003 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Robust Neural Dynamics Method for Redundant Robot Manipulator Control With Physical ConstraintsabstractRedundant robot manipulators play a significant role in modern industry. In this article, we propose a solution scheme to the trajectory tracking problem of the redundant robot manipulator with physical constraints through the Zhang neural dynamics method. Such problem is integrated into a time-varying system consisting of time-varying nonlinear equation (TVNE) and time-varying linear inequality (TVLI) and solved online by the varying-parameter Zhang neural dynamics (VPZND) model. It is ensured that the redundant robot manipulator can still perform the tracking task perfectly under the coexistence of time-varying bounded noise and physical constraints. Theoretical analysis proves that this VPZND model also has an explicit fixed convergence time. Numerical experiments confirm the feasibility of our VPZND model for TVLI. The trajectory tracking problem of the redundant robot manipulator with six or three degrees of freedom under the dual influence of physical constraints and noise is perfectly solved by the VPZND model, which is enough to verify its practical value. Miaomiao Zhang 0001, Kevin W. Tong, Ping Li 0044, Yuhong Hou, Xin Xu 0001, Limin Zhu 0001, Qi Wu 0003 |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Normal Assisted Pixel-Visibility Learning With Cost Aggregation for Multiview StereoabstractMultiple-View Stereo (MVS) aims to reconstruct the dense 3D representations of scenes. MVS has potential applications in the fields of autonomous driving (unstructured environment construction) and robotic navigation (visual-inertial navigation). To mitigate the error of depth estimation in low-textured or occluded regions, this work proposes a two-stage multi-view stereo network for fast and accurate depth estimation. The improvements of this work over the state of the art are as follows: 1) Sparse costs are constructed to jointly predict the initial depth map and surface normal by cost regularization, which proves that the surface normals can be estimated in this way with low memory consumption. 2) A new edge refinement block is developed to refine the coarse surface normal to obtain a fine-grained surface normal map. 3) Instead of using the general variance-based metric to equally aggregate cost, a new content-adaptive cost aggregation mechanism based on the similarity of the neighboring surface normal is designed for reliable cost aggregation. To the best of our knowledge, the proposed work is the first trainable network that leverages surface normal as guidance to capture neighboring pixel-visibility, which is an effective supplement to existing depth/normal estimation frameworks. Experimental results indicate that our method can not only achieve accurate depth estimation for scene perception but also make no concession to the real-time performance and limited memory bottleblock. Multiple-view stereo (MVS) aims to reconstruct the dense 3D representations of scenes. It is widely used in the fields of industrial measurement, autonomous driving, and robotic navigation. To mitigate the error of depth estimation in challenging scenarios, this work proposes a two-stage multi-view stereo network for fast and accurate depth estimation. Our method is the first trainable network that leverages surface normal as pixel-visibility guidance to aggregate reliable cost, which could achieve accurate depth estimation and provide the perception ability for the robot. The proposed method has great potential in the fields of 3D reconstruction, industrial measurement, and robotic navigation to estimate real-time and accurate depth with limited memory consumption. Kevin W. Tong, Xiaorong Guan, Jian Kang 0005, Zhao-Hui Sun, Rob Law 0001, Pedram Ghamisi, Qi Wu 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |