EDBT 2026 Demo / reviewers in the wild / expert
Haitao Meng
dblp:53/4309
· DBLP profile ↗
17ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MTRT: Multitask Reconstruction Transformer for Sleep Stage Classification Based on Internet-of-Things-Enabled Wearable DevicesabstractSleep stage classification is an effective approach for diagnosing sleep disorders and holds significant importance in sleep research and biomedical studies. In recent years, sleep stage classification models based on machine learning and deep learning have achieved remarkable progress. However, several challenges remain. First, conventional deep learning-based sleep stage classification models struggle to effectively handle noise and artifacts, leading to inaccurate feature extraction. Second, the scarcity of minority class data, such as deep sleep and REM stages, prevents models from learning crucial features of these stages. Additionally, computational resource limitations hinder the deployment of previous models on wearable devices. To address these issues, we propose a Multi-Task Reconstruction Transformer (MTRT) for sleep stage classification based on internet of things and wearable devices. MTRT integrates multiple advanced techniques, including masked autoencoders, Transformer networks, Focal Loss, and model pruning, to enhance classification performance and applicability. By introducing a masked autoencoder within a multi-task learning framework, the model effectively removes noise and artifacts from EEG signals, enabling the extraction of high-quality signal representations, thereby improving classification robustness and accuracy. The combination of signal reconstruction and sleep stage classification tasks facilitates feature sharing between the encoder and decoder, enhancing the model’s ability to learn from minority class samples and improving generalization. Moreover, model pruning reduces complexity, minimizing computational and memory requirements, making MTRT more suitable for real-time applications on wearable devices. Extensive experiments on real-world datasets demonstrate the superior performance of the proposed MTRT model. Haitao Meng, Xing Shao |
IEEE Internet Things J. | 1 |
| 2026 | Rotation-invariant representation learning by sector convolution neural networks
Wenwei Lin, Xunpei Sun, Chonghao Zhong, Haitao Meng, Gang Chen 0023, Bingxian Zhang, Zonghua Gu 0001 |
Pattern Recognit. | 4 |
| 2026 | CS-YOLO:A small object detection model based on YOLO for UAV aerial photography
Renhao Jiao, Weigui Nan, Haitao Meng, Abin Jiang, Xiaojia Yang, Jin Dang, Yanshan Tian, Baiying Dong, Xiaoli Luo |
Signal Process. Image Commun. | 4 |
| 2025 | OmniStereo: Real-time Omnidireactional Depth Estimation with Multiview Fisheye CamerasabstractFast and reliable omnidirectional 3D sensing is essential to many applications such as autonomous driving, robotics and drone navigation. While many well-recognized methods have been developed to produce high-quality omnidirectional 3D information, they are too slow for real-time computation, limiting their feasibility in practical applications. Motivated by these shortcomings, we propose an efficient omnidirectional depth sensing framework, called OmniStereo, which generates high-quality 3D information in real-time. Unlike prior works, OmniStereo employs Cassini projection to simplify the photometric matching and introduces a lightweight stereo matching network to minimize computational overhead. Additionally, OmniStereo proposes a novel fusion method to handle depth discontinuities and invalid pixels complemented by a refinement module to reduce mapping-introduced errors and recover fine details. As a result, OmniStereo achieves state-of-the-art (SOTA) accuracy, surpassing the second-best method over 32% in MAE, while maintaining real-time efficiency. It operates more than 16.5× faster than the second-best method in accuracy on TITAN RTX, achieving 12.3 FPS on embedded device Jetson AGX Orin, underscoring its suitability for real-world deployment. The code is available at https://github.com/DengJiaxi1/OmniStereo. Jiaxi Deng, Yushen Wang, Haitao Meng, Zuoxun Hou, Yi Chang 0002, Gang Chen 0023 |
CVPR | 3 |
| 2025 | FastPoint: Super Lightweight Keypoint Detection, Description and Depth Estimation FrameworkabstractKeypoint detection, description and depth estimation are fundamental parts for many computer vision tasks, like image matching and simultaneous localization and mapping (SLAM) systems. Early methods are based on human heuristics, extracted handcrafted feature from image pixel, which may lead to unstable and confusable results. Recent years learned-based methods have achieved remarkable results in this field. However, most of these methods cannot ensure real-time application, and the lightweight deployment of network models remains challenging. Additionally, the black-box nature of neural network operations lacks interpretability, which constrains further application expansion. This paper proposed a super lightweight framework, FastPoint, used shallow binary neural network combine with handcrafted differentiable module FastLayer, which outputs key-points location, descriptor and depth estimation simultaneously. Our FastPoint is trained on data synthetically created form KITTI Raw dataset and evaluated on HPatches. The results demonstrate that our method achieves competitive performance compared to existing approaches while significantly reducing computational complexity. Chonghao Zhong, Haitao Meng, Wenwei Lin, Gang Chen 0023, Alois C. Knoll |
ICTAI | 2 |
| 2023 | ER3D: An Efficient Real-time 3D Object Detection Framework for Autonomous Drivingabstract3D object detection is a vital computer vision task in mobile robotics and autonomous driving. However, most existing methods have exclusively focused on achieving high accuracy, leading to complex and bulky systems that can not be deployed in a real-time manner. In this paper, we propose the ER3D (Efficient and Real-time 3D) object detection framework, which takes stereo images as input and predicts 3D bounding boxes. Instead of using the complex network architecture, we leverage a fast-but-inaccurate method of semi-global matching (SGM) for depth estimation. To eliminate the accuracy degradation in 3D detection caused by inaccurate depth estimation, we introduce decoupled regression head and 3D distance-consistency loU loss to boost the accuracy performance of the 3D detector with a small computing overhead. ER3D achieves both high-precision and real-time performance to enable practical applications of 3D object detection systems on robotic systems. Extensive experiments with the comparison of the state of the arts demonstrate the superior practicability of ER3D, which achieves comparable detection accuracy with significant leadership on inference efficiency. Haitao Meng, Changcai Li, Gang Chen 0023, Zonghua Gu 0001, Alois C. Knoll |
ICPADS | 1 |
| 2023 | FastOmniMVS: Real-time Omnidirectional Depth Estimation from Multiview Fisheye ImagesabstractOmnidirectional depth sensing, which can build 3-D scene structure of the objects of interest in all directions without any blind regions, is essential for ensuring the safety of robotic systems in complex environments. Recent researches attempt to use end-to-end deep neural network (DNN) models for omnidirectional depth estimation from multiview fisheye images. However, most of these existing DNN-based researches cannot achieve real-time performance. In this paper, we propose, FastOmniMVS, a lightweight end-to-end deep neural network (DNN) to achieve real-time omnidirectional depth estimation. FastOmniMVS consists of three parts: a feature extraction module that extracts features from the input fisheye images, spherical sweeping that maps the features to the spherical space to form a 4D cost volume, and finally, the depth is computed by a 3D encoder-decoder. Additionally, we conduct Quantization-Aware Training (QAT) for this network and deploy the quantized model on Jetson Orin. Through extensive experiments, we demonstrate that FastOmniMVS achieves high computational efficiency while maintaining performance close to the state-of-the-art (SOTA) accuracy. It can achieve a real-time inference speed of 35 frames per second (FPS) on the Jetson Orin device. Yushen Wang, Jiaxi Deng, Haitao Meng, Gang Chen 0023 |
ICPADS | 4 |
| 2023 | Rotation-Invariant Descriptors Learned with Circulant Convolution Neural NetworksabstractExtracting local features for accurate correspondences between image pairs is an essential basis for various computer vision tasks. Recent works have shown that deep neural networks (DNNs) have demonstrated promising performance in challenging environments. However, these state-of-the-art DNN-based approaches are not well suited for the scenario with geometry rotations due to their intrinsic deficiencies of square kernel structure. That is, square kernel structures in standard DNNs cannot fully identify the essentials of the rotations in geometry. To address this problem, we present RICNN, a novel deep learning framework that encodes invariance against the rotations in geometry explicitly into convolutional neural networks. Rather than using the square-shaped kernel structure, RICNN adopts sectorshaped convolutional kernels to achieve encoding invariance in all rotations. With the explicitness of such rotation encoding, RICNN enables the transfer of perspective DNN models to obtain rotation-invariant descriptions. Furthermore, we propose a novel multi-level hinge triplet loss function to strengthen the matching constraints against geometry rotations. Comprehensive experiments demonstrate the strong generalization ability of the RICNN descriptor on the HPatches dataset. Toward the rotation invariance evaluation, our method shows state-of-the-art results. Wenwei Lin, Chonghao Zhong, Xunpei Sun, Haitao Meng, Gang Chen 0023, Biao Hu 0001, Zonghua Gu 0001 |
ICTAI | 4 |
| 2023 | A robust and real-time DNN-based multi-baseline stereo accelerator in FPGAs
Yehua Ling, Haitao Meng, Gang Chen 0023 |
J. Syst. Archit. | 4 |
| 2022 | Lite-Stereo: A Resource-Efficient Hardware Accelerator for Real-Time High-Quality Stereo Estimation Using Binary Neural NetworkabstractStereo estimation plays a key role in many autonomous systems, such as robotics and self-driving cars. Recent work on StereoEngine, an FPGA-based accelerator for deep neural network (DNN)-based stereo estimation, has been demonstrated as a promising solution to achieve both real-time and high accuracy performance for depth sensing. However, this solution still suffers from over-utilizing the hardware resource of FPGAs. In this article, we present Lite-Stereo, a resource-efficient DNN-based stereo vision accelerator to improve the hardware efficiency for StereoEngine running on a resource-constrained FPGA. To achieve this, we design a set of optimized hardware architectures for resource-demanding bottleneck modules. In order to balance the gap between the processing speed and resource efficiency, the process elements in binary neural network modules are shared within and across modules. In addition, we provide reusing strategies on path aggregation and neighbor calculation to improve the resource efficiency of the semi-global matching module. Evaluation results demonstrate that Lite-Stereo reduces the hardware cost of ALUTs and RAM bits by 60% and 29%, respectively, without compromising the accuracy and energy efficiency compared with StereoEngine. Yehua Ling, Haitao Meng, Kai Huang 0001, Gang Chen 0023 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | A GPU -accelerated Deep Stereo- LiDAR Fusion for Real-time High-precision Dense Depth SensingabstractActive LiDAR and stereo vision are the most commonly used depth sensing techniques in autonomous vehicles. Each of them alone has weaknesses in terms of density and reliability and thus cannot perform well on all practical scenarios. Recent works use deep neural networks (DNNs) to exploit their complementary properties, achieving a superior depth-sensing. However, these state-of-the-art solutions are not satisfactory on real-time responsiveness due to the high computational complexities of DNNs. In this paper, we present FastFusion, a fast deep stereo-LiDAR fusion framework for real-time high-precision depth estimation. FastFusion provides an efficient two-stage fusion strategy that leverages binary neural network to integrate stereo-LiDAR information as input and use cross-based LiDAR trust aggregation to further fuse the sparse LiDAR measurements in the back-end of stereo matching. More importantly, we present a GPU-based acceleration framework for providing a low latency implementation of FastFusion, gaining both accuracy improvement and real-time responsiveness. In the experiments, we demonstrate the effectiveness and practicability of FastFusion, which obtains a significant speedup over state-of-the-art baselines while achieving comparable accuracy on depth sensing. Haitao Meng, Chonghao Zhong, Jianfeng Gu 0001, Gang Chen 0023 |
DATE | 1 |
| 2021 | An efficient GPU-accelerated inference engine for binary neural network on mobile phones
Shengyu He, Haitao Meng, Zhaoheng Zhou, Kai Huang 0001, Gang Chen 0023 |
J. Syst. Archit. | 2 |
| 2021 | Hardware accelerator for an accurate local stereo matching algorithm using binary neural network
Yehua Ling, Haitao Meng, Gang Chen 0023 |
J. Syst. Archit. | 3 |
| 2020 | PhoneBit: Efficient GPU-Accelerated Binary Neural Network Inference Engine for Mobile PhonesabstractOver the last years, a great success of deep neural networks (DNNs) has been witnessed in computer vision and other fields. However, performance and power constraints make it still challenging to deploy DNNs on mobile devices due to their high computational complexity. Binary neural networks (BNNs) have been demonstrated as a promising solution to achieve this goal by using bit-wise operations to replace most arithmetic operations. Currently, existing GPU-accelerated implementations of BNNs are only tailored for desktop platforms. Due to architecture differences, mere porting of such implementations to mobile devices yields suboptimal performance or is impossible in some cases. In this paper, we propose PhoneBit, a GPU-accelerated BNN inference engine for Android-based mobile devices that fully exploits the computing power of BNNs on mobile GPUs. PhoneBit provides a set of operator-level optimizations including locality-friendly data layout, bit packing with vectorization and layers integration for efficient binary convolution. We also provide a detailed implementation and parallelization optimization for PhoneBit to optimally utilize the memory bandwidth and computing power of mobile GPUs. We evaluate PhoneBit with AlexNet, YOLOv2 Tiny and VGG16 with their binary version. Our experiment results show that PhoneBit can achieve significant speedup and energy efficiency compared with state-of-the-art frameworks for mobile devices. Gang Chen 0023, Shengyu He, Haitao Meng, Kai Huang 0001 |
DATE | 3 |
| 2020 | StereoEngine: An FPGA-Based Accelerator for Real-Time High-Quality Stereo Estimation With Binary Neural NetworkabstractStereo estimation is essential to many applications such as mobile autonomous robots, most of which ask for real-time response, high energy, and storage efficiency. Deep neural networks (DNNs) have shown to yield significant gains in improving accuracy. However, these DNN-based algorithms are challenging to be deployed on energy and resource-constrained devices due to the high computational complexities of DNNs. In this article, we present StereoEngine, a fully pipelined end-to-end stereo vision accelerator that computes accurate dense depth in a real-time and energy-efficient manner. An efficient stereo algorithm is developed and optimized for a high-quality hardware-friendly implementation, that leverages binary neural network (BNN) to learn discriminative binary descriptors to improve the disparity. The design of StereoEngine is a standalone DNN-based stereo vision system where all processing procedures are implemented on a hardware platform. The effectiveness of StereoEngine is evaluated by comprehensive experiments. Compared with software-based implementations on the highend and embedded Nvidia GPUs, StereoEngine achieves up to 3×, 13×, and 50× speedups, as well as up to 211×, 58×, and 73× energy efficiency improvement, respectively. Furthermore, StereoEngine achieves leading accuracy when compared to state-of-the-art hardware implementations on the challenging KITTI dataset. Gang Chen 0023, Yehua Ling, Haitao Meng, Shengyu He, Kai Huang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | GPU-Accelerated Real-Time Stereo Estimation With Binary Neural NetworkabstractDepth estimation from stereo images is essential to many applications such as robotics and autonomous vehicles, most of which ask for the real-time response, high energy and storage efficiency. Recent work has shown deep neural networks (DNN) perform extremely well for stereo estimation. However, these state-of-the-art DNN based algorithms are challenging to be deployed into real-world applications due to the high computational complexities of DNNs. Most of them are too slow for real-time inference and require several seconds of GPU computation to process image frames. In this article, we address the problem of fast stereo estimation and propose an efficient and light-weighted stereo matching system, called StereoBit, to produce a disparity map in a real-time manner while achieving close to state-of-the-art accuracy. To achieve this goal, we propose a binary neural network to generate weighted Hamming distance for an efficient similarity join in stereo estimation. In addition, we propose a novel approximation approach to derive StereoBit network directly from the well-trained network with the cosine similarity. Our approximation strategies enable a significant speedup while maintaining almost the same accuracy compared to the network with the cosine similarity. Furthermore, we present an optimization framework for fully exploiting the computing power of StereoBit. The framework provides a significant speedup of stereo estimation routines, and at the same time, reduces the memory usage for storing parameters. The effectiveness of StereoBit is evaluated by comprehensive experiments. StereoBit can achieve 60 frames per second on an NVIDIA TITAN Xp GPU on KITTI 2012 benchmark while achieving 3-pixel non-occluded stereo error 3.56 percent. Gang Chen 0023, Haitao Meng, Yucheng Liang, Kai Huang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2006 | A Flexible Algorithm for Extracting Periodic Signals
Zhi-Lin Zhang, Haitao Meng |
ISNN (2) | 2 |