Nan Dong

dblp:40/9048 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Physics-based and data-driven adaptive compressor optimization in gas pipeline systems under dynamic and non-isothermal conditions
Nan Dong, Xinmin Wang, Ling Jian
Eng. Appl. Artif. Intell.2
2025 AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting
Yongtao Wang, Yufei Wei, Nan Dong, Ming-Hsuan Yang 0001
ICCV5
2025 LLM-MANUF: An integrated framework of Fine-Tuning large language models for intelligent Decision-Making in manufacturing
Kaze Du, Bo Yang 0042, Keqiang Xie, Nan Dong, Zhengping Zhang, Shilong Wang 0001
Adv. Eng. Informatics4
2025 A manufacturing knowledge graph completion method based on a lightweight dual encoding model
Keqiang Xie, Nan Dong
Appl. Intell.6
2024 BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios
abstract
Existing LiDAR-based 3D object detection methods for autonomous driving scenarios mainly adopt the training-from-scratch paradigm. Unfortunately, this paradigm heavily relies on large-scale labeled data, whose collection can be expensive and time-consuming. Self-supervised pre-training is an effective and desirable way to alleviate this dependence on extensive annotated data. In this work, we present BEV-MAE, an efficient masked autoencoder pre-training framework for LiDAR-based 3D object detection in autonomous driving. Specifically, we propose a bird's eye view (BEV) guided masking strategy to guide the 3D encoder learning feature representation in a BEV perspective and avoid complex decoder design during pre-training. Furthermore, we introduce a learnable point token to maintain a consistent receptive field size of the 3D encoder with fine-tuning for masked point cloud inputs. Based on the property of outdoor point clouds in autonomous driving scenarios, i.e., the point clouds of distant objects are more sparse, we propose point density prediction to enable the 3D encoder to learn location information, which is essential for object detection. Experimental results show that BEV-MAE surpasses prior state-of-the-art self-supervised methods and achieves a favorably pre-training efficiency. Furthermore, based on TransFusion-L, BEV-MAE achieves new state-of-the-art LiDAR-based 3D object detection results, with 73.6 NDS and 69.6 mAP on the nuScenes benchmark. The source code will be released at https://github.com/VDIGPKU/BEV-MAE.
Yongtao Wang, Shengxiang Qi, Nan Dong, Ming-Hsuan Yang 0001
AAAI4
2024 RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object Detection
abstract
Three-dimensional object detection is one of the key tasks in autonomous driving. To reduce costs in practice, low-cost multi-view cameras for 3D object detection are proposed to replace the expansive LiDAR sensors. However, relying solely on cameras is difficult to achieve highly accurate and robust 3D object detection. An effective solution to this issue is combining multi-view cameras with the economical millimeter-wave radar sensor to achieve more reliable multi-modal 3D object detection. In this paper, we introduce RCBEVDet, a radar-camera fusion 3D object detection method in the bird's eye view (BEV). Specifically, we first design RadarBEVNet for radar BEV feature extraction. RadarBEVNet consists of a dual-stream radar backbone and a Radar Cross-Section (RCS) aware BEV encoder. In the dual-stream radar backbone, a point-based encoder and a transformer-based encoder are proposed to extract radar features, with an injection and extraction module to facilitate communication between the two encoders. The RCS-aware BEV encoder takes RCS as the object size prior to scattering the point feature in BEV. Besides, we present the Cross-Attention Multi-layer Fusion module to automatically align the multi-modal BEV feature from radar and camera with the deformable attention mechanism, and then fuse the feature with channel and spatial fusion layers. Experimental results show that RCBEVDet achieves new state-of-the-art radar-camera fusion results on nuScenes and view-of-delft (VoD) 3D object detection benchmarks. Furthermore, RCBEVDet achieves better 3D detection results than all real-time camera-only and radar-camera 3D object detectors with a faster inference speed at 21∼28 FPS. The source code will be released at https://github.com/VDIGPKU/RCBEVDet.
Zhongyu Xia, Yongtao Wang, Shengxiang Qi, Nan Dong, Ce Zhu
CVPR8
2024 TEOcc: Radar-Camera Multi-Modal Occupancy Prediction via Temporal Enhancement
abstract
As a novel 3D scene representation, semantic occupancy has gained much attention in autonomous driving. However, existing occupancy prediction methods mainly focus on designing better occupancy representations, such as tri-perspective view or neural radiance fields, while ignoring the advantages of using long-temporal information. In this paper, we propose a radar-camera multi-modal temporal enhanced occupancy prediction network, dubbed TEOcc. Our method is inspired by the success of utilizing temporal information in 3D object detection. Specifically, we introduce a temporal enhancement branch to learn temporal occupancy prediction. In this branch, we randomly discard the t−k input frame of the multi-view camera and predict its 3D occupancy by long-term and short-term temporal decoders separately with the information from other adjacent frames and multi-modal inputs. Besides, to reduce computational costs and incorporate multi-modal inputs, we specially designed 3D convolutional layers for long-term and short-term temporal decoders. Furthermore, since the lightweight occupancy prediction head is a dense classification head, we propose to use a shared occupancy prediction head for the temporal enhancement and main branches. It is worth noting that the temporal enhancement branch is only performed during training and is discarded during inference. Experiment results demonstrate that TEOcc achieves state-of-the-art occupancy prediction on nuScenes benchmarks. In addition, the proposed temporal enhancement branch is a plug-and-play module that can be easily integrated into existing occupancy prediction methods to improve the performance of occupancy prediction. The source code and models will be released at https://github.com/VDIGPKU/TEOcc.
Hongbo Jin, Yongtao Wang, Yufei Wei, Nan Dong
ECAI5
2024 HENet: Hybrid Encoding for End-to-End Multi-task 3D Perception from Multi-view Cameras
Zhongyu Xia, Yongtao Wang, Shengxiang Qi, Nan Dong, Ming-Hsuan Yang 0001
ECCV (50)7
2024 Multidomain neural process model based on source attention for industrial robot anomaly detection
Bo Yang 0042, Keqiang Xie, Nan Dong
Adv. Eng. Informatics7
2022 Semi-Blind Multi-cell Interference Detection and Cancellation in 5G Uplink OFDM Systems
abstract
As interference becomes one of the key factors restricting the performance of wireless network, many interference cancellation schemes are studied. However, most of these schemes have disadvantages in one way or another, such as high complexity and cost of the accurate feedback. In this case, blind interference cancellation schemes are proposed, which can eliminate the interference according to the received signal without any prior information, but with a very high searching complexity. To solve above issues, we propose a semi-blind interference parameter detection (Semi-BIPD) and signal restoration scheme in this paper. Firstly, a semi-blind interference detection module is investigated to detect the parameters related to strong inter-ference with the help of the received demodulation reference signal (DM-RS) sequence. Then, the channel parameters of both target user and interfering users are estimated. Finally, the signal restoration based on Semi-BIPD is conducted in the data sequence to eliminate the interference. Simulation results demonstrate that the proposed scheme can achieve better mean square error (MSE) with a low searching complexity in 5G uplink orthogonal frequency division multiplexing (OFDM) systems.
Yanzan Sun, Jiaqi Kang, Wenshu Sui, Shunqing Zhang, Xiaojing Chen 0001, Nan Dong
IWCMC6
2018 Incremental generalized multiple maximum scatter difference with applications to feature extraction
Ning Zheng 0003, Xin Guo 0005, Tie Yun, Nan Dong, Lin Qi 0001, Ling Guan
J. Vis. Commun. Image Represent.4
2016 The design and implementation of the privacy protection system of a Regional Health Information Platform
abstract
Objective The Regional Health Information Platform (RHIP) is an integrated information sharing system for the collection, transmission, storage, sharing, and application of health related data in a city. This paper introduced the design and the implementation of the privacy protection system of the RHIP. Methods Both technical and administrative methods are applied for privacy protection. The specific measures are derived from the analysis of the business needs, system architecture, possibilities of privacy leakage, and the political rules of China. Results Five principles are established to protect the privacy issues in the RHIP. A series of key techniques and administrative strategies for protecting privacy issues of the RHIP are developed according to these five principles. A whole system of preventing information leakage and responding to the emergency was created. Conclusion It is an effective exploration of protecting privacy in sharing medical information to provide some beneficial experience of developing new information safety systems.
Nan Dong, Yi Zhou 0005, Zhaosheng Gao
BIBM3
2015 A Fast 4D Facial Expression Recognition Method for Low-resolution Videos
Nan Dong
ICPRAM (1)2
2015 An Approach to Ballet Dance Training through MS Kinect and Visualization in a CAVE Virtual Reality Environment
abstract
This article proposes a novel framework for the real-time capture, assessment, and visualization of ballet dance movements as performed by a student in an instructional, virtual reality (VR) setting. The acquisition of human movement data is facilitated by skeletal joint tracking captured using the popular Microsoft (MS) Kinect camera system, while instruction and performance evaluation are provided in the form of 3D visualizations and feedback through a CAVE virtual environment, in which the student is fully immersed. The proposed framework is based on the unsupervised parsing of ballet dance movement into a structured posture space using the spherical self-organizing map (SSOM). A unique feature descriptor is proposed to more appropriately reflect the subtleties of ballet dance movements, which are represented as gesture trajectories through posture space on the SSOM. This recognition subsystem is used to identify the category of movement the student is attempting when prompted (by a virtual instructor) to perform a particular dance sequence. The dance sequence is then segmented and cross-referenced against a library of gestural components performed by the teacher. This facilitates alignment and score-based assessment of individual movements within the context of the dance sequence. An immersive interface enables the student to review his or her performance from a number of vantage points, each providing a unique perspective and spatial context suggestive of how the student might make improvements in training. An evaluation of the recognition and virtual feedback systems is presented.
Matthew J. Kyan, Guoyu Sun, Paisarn Muneesawang, Nan Dong, Bruce Elder, Ling Guan
ACM Trans. Intell. Syst. Technol.6
2014 An Advanced Computational Intelligence System for Training of Ballet Dance in a Cave Virtual Reality Environment
abstract
This paper presents a computer-based system for assessment and training of ballet dance in a CAVE virtual reality environment. The system utilizes Kinect sensor to capture student's dance and extracts features from skeleton joints. This system depends on a structured posture space, which comprises a set of dance elements that represent key moments -- "postures", that typically will be so briefly held as to experience as a fleeting moment in a flux -- in the dance movements whose performance we are attempting to assess. The recording captured from the Kinect allows the parsing of dance movement into a structured posture space using the spherical self-organizing map (SSOM). From this, a unique descriptor can be obtained by following gesture trajectories through posture space on the SSOM, which appropriately reflects the subtleties of ballet dance movements. Consequently, the system can recognize the category of movement the student is attempting, and this allows us make a quantitative assessment of individual movements. Based on the experimental results, the proposed system appears to be very effective for recognition and offering generalization across instances of movement. Thus, it is possible for the construction of assessment and visualization of ballet dance movements performed by the student in an instructional, virtual reality setting.
Guoyu Sun, Paisarn Muneesawang, Matthew J. Kyan, Nan Dong, Bruce Elder, Ling Guan
ISM6
2013 Multi-part sparse representation in random crowded scenes tracking
Jie Shao 0010, Nan Dong, Minglei Tong
Pattern Recognit. Lett.2
2010 Traffic Abnormality Detection through Directional Motion Behavior Map
abstract
Automatic traffic abnormality detection through visual surveillance is one of the critical requirements for Intelligent Transportation Systems (ITS). In this paper, we present a novel algorithm to detect abnormal traffic events in crowded scenes. Our algorithm can be deployed with few setup steps to automatically monitor traffic status. Different from other approaches, we don't need to define region of interests (ROI) or tripwires nor to configure object detection and tracking parameters. A novel object behavior descriptor directional motion behavior descriptors are proposed. The directional motion behavior descriptors collect foreground objects' direction and speed information from a video sequence with normal traffic events, and then these descriptors are accumulated to generate a directional motion behavior map which models the normal traffic status. During detection steps, we first extract the directional motion behavior map from the newly observed video and then measure the differences between the normal behavior map and the new map. If new direction motion behaviors are very different from the descriptors in the normal behavior map, then the corresponding regions in the observed video contain traffic abnormalities. Our proposed algorithm has been tested using both synthesized and real surveillance videos. Experimental results demonstrated that our algorithm is effective and efficient for practical real-time traffic surveillance applications.
Nan Dong, Jie Shao 0010, Ziyou Xiong, Pei-Yuan Peng
AVSS1
2009 Content-free image orientation detection using Local Binary Pattern
abstract
Image orientation detection is a basic, but important subject in image processing. It is also a difficult task because of the uncertainty of the context information of the image and the illumination. In this paper, we present a theoretically and computationally simple yet efficient content-free image orientation detection approach. The proposed approach is based on Local Binary Patterns which is recognized as gray scale and rotation invariant. By using the binary mode and the histogram computed over the overlapping region of two images, the experiments show that it can get excellent results in both artificial and real images. Furthermore, the method can be used in different images with similar contents. Through optimization, the run time of the algorithm reduces by a large amount and makes it applicable in real applications.
Nan Dong, Ling Guan
MMSP1