Weichen Dai 0001

dblp:230/3555-1 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-5019-2755ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Dense structure distillation towards recurrent scene flow estimation for sparse point cloud
abstract
Abstract Scene flow, a key representation of motion information in 3D space, plays a critical role in numerous downstream tasks. However, existing point cloud scene flow estimation methods often experience significant performance degradation due to insufficient feature expressiveness or mismatching under sparse point cloud conditions. To alleviate these problems, we propose a novel dense structure distillation towards recurrent scene flow estimation method for sparse point clouds, significantly enhancing performance under low-density point cloud conditions through a dense structure distillation strategy. This module addresses the information loss caused by point cloud sparsity by leveraging high-quality point cloud features to effectively guide the learning of sparse point cloud feature extractors. To further refine the estimation results, a recurrent update strategy is adopted to gradually improve the accuracy and stability of scene flow estimation. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on both the public FlyingThings3D and KITTI datasets, particularly under sparse point cloud conditions, and outperforms existing methods.
Weichen Dai 0001, Ziyue Meng, Xiaoyang Weng, Wenhan Su, Wanzeng Kong
Comput. J.1
2026 Brain-Eye Collaborative Camouflaged Target Detection
abstract
Brain-computer interface (BCI) based on rapid serial visual presentation (RSVP) are widely applied in target detection but suffer from low decoding accuracy. While some methods integrate eye movements to improve localization, they often underutilize EEG’s coarse spatial cues, limiting overall detection effectiveness. We propose a brain-eye collaborative method for camouflaged target detection that integrates EEG and eye movement signals in a coarse-to-fine framework. In the coarse stage, target presence and its image quadrant are determined using brain-eye fusion and contrastive learning on bimodal features. This leverages complementary spatiotemporal information from both modalities to generate a coarse localization region. In the fine stage, eye movement data further refines the target location within this region by identifying high-interest areas. This collaborative approach significantly improves both recognition and localization performance in complex visual search tasks. Our method achieves an F1 score of 83.16% and a balanced accuracy of 85.38%, outperforming existing state-of-the-art methods.
Longjie Ma, Weichen Dai 0001, Ziyue Yang 0007, Chenyi Hong, Jianting Cao, Wanzeng Kong
Int. J. Hum. Comput. Interact.2
2026 Brain-Machine Enhanced Intelligence for Semi-Supervised Facial Emotion Recognition
abstract
Machine learning, particularly deep learning, typically achieves high facial emotion image recognition accuracy benefiting from multiple labeled data. However, the datasets usually contain insufficient labeled samples and numerous unlabeled data since human labeling is a costly endeavor. For semi-supervised learning of these datasets, self-training procedure solely based on the visual features of images fails to comprehensively understand the intricate high-level semantic features. Since EEG signals contain not only visual information related to the visual stimulus but also emotional information related to brain activity, they are highly suitable as supervisory signals for labeling unlabeled facial emotion images. In this study, we specifically employ EEG signals evoked by visual image stimuli in conjunction with EEGNet3D to learn a discriminative EEG class representation manifold of brain activity. The one-hot class label is replaced with the EEG class representation as the supervisory to train the base model. Then, better pseudo-labeling is achieved using the base model in the EEG class representation manifold. Based on pseudo-labeling results, the utilization of unlabeled data is further improved. Interestingly, our findings reveal that when utilizing EEG class representations as supervisory information for the base model, the base model demonstrates a learning pattern that involves focusing more on the eye area when making judgments about emotions. This behavior closely resembles how the human brain decodes emotions. Experiments show that the performance of the proposed method can be effectively enhanced by combining labeled and pseudo-labeled images. Further experiments demonstrate that our method exhibits strong generalization abilities when applied to new image datasets and other visual networks.
Dongjun Liu, Weichen Dai 0001, Hangjie Yi, Honggang Liu, Jianting Cao, Qibin Zhao, Fabio Babiloni, Wanzeng Kong
IEEE Trans. Affect. Comput.2
2025 Multi-Modal Synergistic Implicit Image Enhancement for Efficient Optical Flow Estimation
abstract
As a fundamental visual task, optical flow estimation has widespread applications in computer vision. However, it faces significant challenges under adverse lighting conditions, where low texture and noise make accurate optical flow estimation particularly difficult. In this paper, we propose an optical flow method that employs implicit image enhancement through multi-modal synergistic training. To supplement the scene information missing in the original low-quality image, we utilize a high-low frequency feature enhancement network. The enhancement network is implicitly guided by multi-modal data and the specific subsequent tasks, enabling the model to learn multi-modal knowledge that enhances feature information suitable for optical flow estimation during inference. By using RGBD multi-modal data, the proposed method avoids the reliance on the images captured from the same view, a common limitation in traditional image enhancement methods. During training, the encoded features extracted from the enhanced images are synergistically supervised by features from the RGBD fusion as well as by the optical flow task. Experiments conducted on both synthetic and real datasets demonstrate that the proposed method significantly improves performance on public datasets.
Weichen Dai 0001, Hexing Wu, Xiaoyang Weng, Yuhang Ming 0001, Wanzeng Kong
CVPR1
2025 MPFDAN: Multi-Perspective Feature Dynamic Adaptation Network for Domain Adaptive Object Detection
abstract
Based on adversarial training and hierarchical alignment structure, domain adaptive object detection methods have made impressive progress. However, current adaptation methods treats each feature equally, and exerts unchanged alignment strength during alignment process, without sufficiently considering the transferability inconsistency of different features. Such static alignment can only achieve approximate alignment and will bring about negative transfer eventually. To address these issues, we propose the Multi-Perspective Feature Dynamic Adaptation Network (MPFDAN). In this network, the transferability of features is thoroughly considered and utilized from three different perspectives, allowing the alignment process to be dynamically adjusted in different ways. Firstly, regarding local transferability, Shannon entropy is used to adjust the weights of features in different local regions to focus more on regions with higher transferability. Next, from the perspective of global alignment, we dynamically adjust the alignment strength applied during the image-level adaptation process to avoid overfitting. Finally, category information is introduced to achieve category-aware instance-level adaptation, dynamically adjusted based on the differences in category transferability. Experiments on various domain transfer scenarios demonstrate that our MPFDAN outperforms all compared methods, thereby proving the effectiveness of our proposed approach.
Wenchao Weng, Weichen Dai 0001, Andrzej Cichocki, Wanzeng Kong
IJCNN4
2025 Brain-Machine Cross-Modal Alignment via Sample Relational Learning for Visual Classification
abstract
Recent works on visual classification tasks have leveraged EEG signals to provide additional supervisory information, further improving the performance of the models on natural images. However, previous methods often force machine models to directly match EEG signals, which involves the transfer of modal-specific representations, leading to potentially distorted alignment of modal-shared representations. Moreover, focusing solely on aligning individual sample features neglects the alignment of relationships between samples, making it difficult to capture the potential relational reasoning capabilities in EEG signals. This relational reasoning ability is key to the human brain’s outstanding performance in visual classification tasks. Similarly, for a machine model, the complex relationships between instances are more critical than individual instances. Inspired by this, our idea is to enhance machine visual classification capabilities by imparting human-like relational reasoning, encouraging machine models to focus on the relational structure within EEG signals. To this end, we propose a brain-machine relation alignment method that constructs a cognitive model and a visual model to process EEG signals and visual images, respectively. Instead of forcing the visual model to mimic the output of an individual EEG data sample represented by the cognitive model, we encourage it to learn the mutual relations of EEG data samples. By penalizing the difference in relational structures between EEG signals and visual images, we facilitate the transfer of relational knowledge. Experiments demonstrate that the proposed method significantly improves the classification performance of the visual model. This highlights the potential of relational alignment as a robust mechanism for integrating human relational reasoning into machine learning models.
Dongjun Liu, Weichen Dai 0001, Honggang Liu, Hangjie Yi, Wanzeng Kong
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Implicit guidance for enhancing low-light optical flow estimation via channel attention networks
Weichen Dai 0001, Hexing Wu, Xiaoyang Weng, Wanzeng Kong
Vis. Comput.1
2024 AEGIS-Net: Attention-Guided Multi-Level Feature Aggregation for Indoor Place Recognition
abstract
We present AEGIS-Net, a novel indoor place recognition model that takes in RGB point clouds and generates global place descriptors by aggregating lower-level color, geometry features and higher-level implicit semantic features. However, rather than simple feature concatenation, self-attention modules are employed to select the most important local features that best describe an indoor place. Our AEGIS-Net is made of a semantic encoder, a semantic decoder and an attention-guided feature embedding. The model is trained in a 2-stage process with the first stage focusing on an auxiliary semantic segmentation task and the second one on the place recognition task. We evaluate our AEGIS-Net on the ScanNetPR dataset and compare its performance with a pre-deep-learning feature-based method and five state-of-the-art deep-learning-based methods. Our AEGIS-Net achieves exceptional performance and outperforms all six methods.
Yuhang Ming 0001, Jian Ma 0001, Xingrui Yang 0001, Weichen Dai 0001, Yong Peng 0001, Wanzeng Kong
ICASSP4
2023 Adaptive Positional Encoding for Bundle-Adjusting Neural Radiance Fields
abstract
Neural Radiance Fields have shown great potential to synthesize novel views with only a few discrete image observations of the world. However, the requirement of accurate camera parameters to learn scene representations limits its further application. In this paper, we present adaptive positional encoding (APE) for bundle-adjusting neural radiance fields to reconstruct the neural radiance fields from unknown camera poses (or even intrinsics). Inspired by Fourier series regression, we investigate its relationship with the positional encoding method and therefore propose APE where all frequency bands are trainable. Furthermore, we introduce period-activated multilayer perceptrons (PMLPs) to construct the implicit network for the high-order scene representations and fine-grained gradients during backpropagation. Experimental results on public datasets demonstrate that the proposed method with APE and PMLPs can outperform the state-of-the-art methods in accurate camera poses and high-fidelity view synthesis.
Zelin Gao, Weichen Dai 0001, Yu Zhang 0018
ICCV2
2023 Continuous-Time LiDAR-Inertial-Vehicle Odometry Method with Lateral Acceleration Constraint
abstract
In this paper, we propose a continuous-time-based LiDAR-inertial-vehicle odometry method, which can tightly fuse the data from Light Detection And Ranging (LiDAR), inertial measurement units (IMU), and vehicle measurements. The lateral acceleration constraint is further added to trajectory estimation to make the estimated trajectory follow the motion characteristics of vehicles. In addition, since vehicle model parameters vary with different motion conditions and tyre pressure, we estimate vehicle correction factors that rectify changes in vehicle model parameters online, and also analyze the observability of these vehicle correction factors. In experiments, the proposed method is evaluated and compared with state-of-the-art methods in the public dataset. The experimental results show that the proposed method achieves more accurate results in all sequences since we add additional sensor measurements and utilize the characteristic of vehicle motion to restrict the trajectory estimation. The ablation study also proved the effectiveness of continuous-time representation, online correction factor estimation, and incorporation of lateral acceleration constraint.
Weichen Dai 0001, Zeyu Wan, Yu Zhang 0018
ICRA2
2023 Brain-Machine Coupled Learning Method for Facial Emotion Recognition
abstract
Neural network models of machine learning have shown promising prospects for visual tasks, such as facial emotion recognition (FER). However, the generalization of the model trained from a dataset with a few samples is limited. Unlike the machine, the human brain can effectively realize the required information from a few samples to complete the visual tasks. To learn the generalization ability of the brain, in this article, we propose a novel brain-machine coupled learning method for facial emotion recognition to let the neural network learn the visual knowledge of the machine and cognitive knowledge of the brain simultaneously. The proposed method utilizes visual images and electroencephalogram (EEG) signals to couple training the models in the visual and cognitive domains. Each domain model consists of two types of interactive channels, common and private. Since the EEG signals can reflect brain activity, the cognitive process of the brain is decoded by a model following reverse engineering. Decoding the EEG signals induced by the facial emotion images, the common channel in the visual domain can approach the cognitive process in the cognitive domain. Moreover, the knowledge specific to each domain is found in each private channel using an adversarial strategy. After learning, without the participation of the EEG signals, only the concatenation of both channels in the visual domain is used to classify facial emotion images based on the visual knowledge of the machine and the cognitive knowledge learned from the brain. Experiments demonstrate that the proposed method can produce excellent performance on several public datasets. Further experiments show that the proposed method trained from the EEG signals has good generalization ability on new datasets and can be applied to other network models, illustrating the potential for practical applications.
Dongjun Liu, Weichen Dai 0001, Hangkui Zhang, Xuanyu Jin, Jianting Cao, Wanzeng Kong
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Dynamically Adjust Word Representations Using Unaligned Multimodal Information
abstract
Multimodal Sentiment Analysis is a promising research area for modeling multiple heterogeneous modalities. Two major challenges that exist in this area are a) multimodal data is unaligned in nature due to the different sampling rates of each modality, and b) long-range dependencies between elements across modalities. These challenges increase the difficulty of conducting efficient multimodal fusion. In this work, we propose a novel end-to-end network named Cross Hyper-modality Fusion Network (CHFN). The CHFN is an interpretable Transformer-based neural model that provides an efficient framework for fusing unaligned multimodal sequences. The heart of our model is to dynamically adjust word representations in different non-verbal contexts using unaligned multimodal sequences. It is concerned with the influence of non-verbal behavioral information at the scale of the entire utterances and then integrates this influence into verbal expression. We conducted experiments on both publicly available multimodal sentiment analysis datasets CMU-MOSI and CMU-MOSEI. The experiment results demonstrate that our model surpasses state-of-the-art models. In addition, we visualize the learned interactions between language modality and non-verbal behavior information and explore the underlying dynamics of multimodal language data.
Jiwei Guo, Jiajia Tang, Weichen Dai 0001, Yu Ding 0001, Wanzeng Kong
ACM Multimedia3
2022 RGB-D SLAM in Dynamic Environments Using Point Correlations
abstract
In this paper, a simultaneous localization and mapping (SLAM) method that eliminates the influence of moving objects in dynamic environments is proposed. This method utilizes the correlation between map points to separate points that are part of the static scene and points that are part of different moving objects into different groups. A sparse graph is first created using Delaunay triangulation from all map points. In this graph, the vertices represent map points, and each edge represents the correlation between adjacent points. If the relative position between two points remains consistent over time, there is correlation between them, and they are considered to be moving together rigidly. If not, they are considered to have no correlation and to be in separate groups. After the edges between the uncorrelated points are removed during point-correlation optimization, the remaining graph separates the map points of the moving objects from the map points of the static scene. The largest group is assumed to be the group of reliable static map points. Finally, motion estimation is performed using only these points. The proposed method was implemented for RGB-D sensors, evaluated with a public RGB-D benchmark, and tested in several additional challenging environments. The experimental results demonstrate that robust and accurate performance can be achieved by the proposed SLAM method in both slightly and highly dynamic environments. Compared with other state-of-the-art methods, the proposed method can provide competitive accuracy with good real-time performance.
Weichen Dai 0001, Yu Zhang 0018, Ping Li 0017, Zheng Fang 0001, Sebastian A. Scherer
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 A Multi-spectral Dataset for Evaluating Motion Estimation Systems
abstract
Visible images have been widely used for motion estimation. Thermal images, in contrast, are more challenging to be used in motion estimation since they typically have lower resolution, less texture, and more noise. In this paper, a novel dataset for evaluating the performance of multi-spectral motion estimation systems is presented. All the sequences are recorded from a handheld multi-spectral device. It consists of a standard visible-light camera, a long-wave infrared camera, an RGB-D camera, and an inertial measurement unit (IMU). The multi-spectral images, including both color and thermal images in full sensor resolution (640 × 480), are obtained from a standard and a long-wave infrared camera at 32Hz with hardware-synchronization. The depth images are captured by a Microsoft Kinect2 and can have benefits for learning cross-modalities stereo matching. For trajectory evaluation, accurate ground-truth camera poses obtained from a motion capture system are provided. In addition to the sequences with bright illumination, the dataset also contains dim, varying, and complex illumination scenes. The full dataset, including raw data and calibration data with detailed data format specifications, is publicly available.
Weichen Dai 0001, Yu Zhang 0018, Shenzhou Chen, Donglei Sun, Da Kong
ICRA1
2019 Multi-Spectral Visual Odometry without Explicit Stereo Matching
abstract
Multi-spectral sensors consisting of a standard (visible-light) camera and a long-wave infrared camera can simultaneously provide both visible and thermal images. Since thermal images are independent from environmental illumination, they can help to overcome certain limitations of standard cameras under complicated illumination conditions. However, due to the difference in the information source of the two types of cameras, their images usually share very low texture similarity. Hence, traditional texture-based feature matching methods cannot be directly applied to obtain stereo correspondences. To tackle this problem, a multi-spectral visual odometry method without explicit stereo matching is proposed in this paper. Bundle adjustment of multi-view stereo is performed on the visible and the thermal images using direct image alignment. Scale drift can be avoided by additional temporal observations of map points with the fixed-baseline stereo. Experimental results indicate that the proposed method can provide accurate visual odometry results with recovered metric scale. Moreover, the proposed method can also provide a metric 3D reconstruction in semi-dense density with multi-spectral information, which is not available from existing multi-spectral methods.
Weichen Dai 0001, Yu Zhang 0018, Donglei Sun, Naira Hovakimyan, Ping Li 0017
3DV1
2018 Feature Regions Segmentation Based RGB-D Visual Odometry in Dynamic Environment
abstract
A novel RGB-D visual odometry method for dynamic environment is proposed. Majority of visual odometry systems can only work in static environments, which limits their applications in real world. In order to improve the accuracy and robustness of visual odometry in dynamic environment, a Feature Regions Segmentation algorithm is proposed to resist the disturbance caused by the moving objects. The matched features are divided into different regions to separate the moving objects from the static background. The features in the largest region which belong to the static background are used to estimate the camera pose finally. The effectiveness of our visual odometry method is verified in a dynamic environment of our lab. Furthermore, an exhaustive experimental evaluation is conducted on benchmark datasets including static environments and dynamic environments compared with the state-of-art visual odometry systems. The accuracy comparison results show that the proposed algorithm outperforms those systems in large scale dynamic environments. Our method tracks the camera movement correctly while others failed. In addition, our method can give the same good performances in static environment. Experiments demonstrate that the proposed RGB-D visual odometry can obtain accurate and robust estimation results in dynamic environments.
Yu Zhang 0018, Weichen Dai 0001, Ping Li 0017, Zheng Fang 0001
IECON2