EDBT 2026 Demo / reviewers in the wild / expert
Mingwei Cao
dblp:34/10413
· DBLP profile ↗
18ranked-venue papers
13as first author
13since 2021 · last 2026
0000-0002-8161-0382ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 9 first-author · 9 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | APMVS: Learning Multi-View Stereo Based on Adjacent Stage and Pair-Wise Stage Uncertainty EstimationabstractMany multi-view stereo (MVS) networks with a cascaded structure can effectively estimate depth while saving memory. However, the accuracy of the depth map in the fine stage depends on the depth map estimated in the coarse stage. Additionally, the multi-stage depth maps generated by the cascaded structure are used to compute losses but are not reused, resulting in a loss of inter-stage differentiation information. To address these issues, we propose a dual-uncertainty estimation MVS method that learns an MVS network based on adjacent stage and pair-wise stage uncertainty estimation, named APMVS. The core of the proposed APMVS is to employ dual-uncertainty estimation to mitigate the adverse effects of the cascaded structure. Specifically, it involves two estimation modules: adjacent stage uncertainty (ASU) and pair-wise stage uncertainty (PSU). The ASU estimation module dynamically adjusts the depth-hypothesis range by leveraging uncertainty from the previous stage, thereby improving the accuracy of depth-map prediction in the current stage. The PSU estimation module estimates the uncertainty between each pair of stages. Thus, regions with high uncertainty have minimal impact. We evaluate the proposed APMVS on the DTU, Tanks and Temples, and BlendedMVS datasets. Experimental results show that our method achieves superior reconstruction quality compared with other state-of-the-art methods. Mingwei Cao, Siqi Nian, Haifeng Zhao 0001, Feng Xue 0002, Zhihan Lyu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2026 | SpatioGS: spatiotemporal-aware density control for dynamic scene rendering with Gaussian splatting
Mingwei Cao |
Vis. Comput. | 1 |
| 2025 | DOMVS: Unsupervised Multi-view Stereo for Dealing With Occlusion Scenes
Mingwei Cao, Qiuju Wang, Haifeng Zhao 0001 |
CGI (2) | 1 |
| 2025 | StarGS: Towards Real-Time Dynamic Scene Rendering via Gaussian Splatting with Spatiotemporal-Aware Density ControlabstractThe 3D Gaussian Splatting (3DGS) technique has yielded significant achievements in scene rendering. However, it encounters difficulties when representing dynamic scenes marked by complexity, high degrees of freedom, and limited viewpoints. To overcome these challenges, we introduce StarGS, an innovative method for real-time dynamic scene rendering that employs spatiotemporal-aware density control in Gaussian splatting. StarGS features an adaptive density mechanism that dynamically modifies the Gaussian distribution based on spatiotemporal characteristics, thereby precisely capturing structural variations and intricate texture details in areas with complex motion. Firstly, we devise a motion-aware keypoints mechanism, guided by spatio-temporal features, which isolates notable moving objects within the scene and promotes local density augmentation around them, enhancing their structural representations. Secondly, we incorporate a pixel-wise error-guided pruning module to eliminate low-contribution Gaussians and noise, effectively minimizing computational redundancy and enhancing rendering efficiency. Thirdly, we suggest a structural feature-anchor rigidity constraint that bolsters the local consistency of Gaussians, reducing motion artifacts in dynamic scenes. We assess StarGS using the HyperNeRF and Neu3D datasets and compare it with state-of-the-art methods. Experimental results reveal that StarGS achieves 86 FPS at a$1352 \times 1014$resolution, while outperforming state-of-the-art methods in image quality. Source code available at: https://github.com/caomw/stargs Mingwei Cao, Haifeng Zhao 0001 |
CW | 1 |
| 2025 | DAMR: Multi-scale graph contrastive learning with dynamic adjustment and mutual rectification
Dengdi Sun, Mingwei Cao, Zhifu Tao, Zhuanlian Ding |
Knowl. Based Syst. | 3 |
| 2025 | A Chiplet Platform for Intelligent Radar/Sonar Leveraging Domain-Specific Reusable Active InterposerabstractThrough chiplet reuse, chiplet-based system designs have emerged as a cost-effective solution for system-on-chips (SoCs), yet considerable silicon interposer costs often negate the benefits. Though general reusable interposers (GRIs) can lower the cost, they often compromise on performance and energy efficiency. In this article, a domain-specific reusable active interposer (active DSRI) approach is proposed for a better cost-efficiency tradeoff. Moreover, a chiplet platform based on an active DSRI designed for the intelligent radar/sonar (IRS) domain is introduced to facilitate rapid and customized SoC development. This platform offers flexible and energy-efficient interconnections tailored for IRS, platform infrastructure functions, and peripherals to simplify the chiplets. Furthermore, it integrates lightweight, composable standard 3-D interfaces across the chiplets and interposer, delivering up to 96-Gb/s bandwidth, 11.1-ns latency, and 0.62-pJ/bit energy efficiency, well controlling the cost and power penalties of SoC partition. Demonstrated with a customized hand gesture recognition sonar system (HGRSS) baseband SoC implemented on the proposed platform, it achieves similar performance to a monolithic SoC, with a recognition frame rate of 6286 frames/s, where overhead of the 3-D interface is only 6.86% in area and 4.84% in power. Our approach proves cost-effective, energy efficient, and customizable, moving system volume breakeven point forward by$3.22\sim 3.36$times, and reducing the cost by 58.5%~59.8%. This represents a pioneering demonstration of reusable chiplets in HGRSS, showcasing the potential of our approach for broader domains. Chaoqin Zhang, Yunlai Zhang, Mingwei Cao, Shouyi Yin |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | DUE-MVSNet: Learning Multi-view Stereo Based on Dual Uncertainty Estimation
Mingwei Cao, Siqi Nian |
CGI (2) | 1 |
| 2024 | DI-MVS: Learning Efficient Multi-View Stereo With Depth-Aware IterationsabstractLearning-based Multi-View Stereo (MVS) methods aim to reconstruct 3D scenes from a set of 2D calibrated images. However, existing learning-based MVS methods often overlook depth maps that include the geometric shapes of the scene when constructing the cost volume. This can result in suboptimal reconstructions, particularly in low-texture or repetitive-texture regions where valuable geometric information is absent. To address this issue, we develop DI-MVS, a coarse-to-fine framework that effectively incorporates context-guided depth geometry into the cost volume using a depth-aware iterator. First, we employ the proposed depth-aware cost completion module to update the cost volume, followed by 2D ConvGRUs to iteratively optimize depth maps efficiently. Second, we propose a hybrid loss strategy that combines two loss functions’ strengths to improve depth estimation’s robustness. Extensive experiments demonstrate that DI-MVS outperforms state-of-the-art methods on the DTU dataset and the Tanks & Temples benchmark. The source code is available at: https://github.com/JianfeiJ/DI-MVS. Jianfei Jiang 0005, Mingwei Cao, Chenglong Li 0002 |
ICASSP | 2 |
| 2024 | BCS-NeRF: Bundle Cross-Sensing Neural Radiance Fields
Mingwei Cao, Fengna Wang, Dengdi Sun, Haifeng Zhao 0001 |
MMAsia | 1 |
| 2023 | G2PL: Lexicon Enhanced Chinese Polyphone Disambiguation Using Bert Adapter with a New DatasetabstractPolyphone disambiguation is the core of grapheme-to-phoneme(G2P) module for the Chinese speech synthesis system. However, there is a lack of datasets and only one public for polyphone disambiguation. Moreover, due to the double long-tail distribution of polyphones, the ratio of pronunciation data for most polyphones is extremely unbalanced after sampling. To solve these problems, we propose a new dataset with 57,000 sentences from various domains by a new strategy for sampling. In addition, we propose the G2PL, which integrates word features into the bottom of BERT to assist in predicting the correct pronunciation of polyphone. In the experiment, we train the G2PL model to outperform other methods on our and public datasets. Our dataset, codes and user-friendly package are freely available. Haifeng Zhao 0001, Hongzhi Wan, Lili Huang 0006, Mingwei Cao |
ICASSP | 4 |
| 2023 | LCSNet: End-to-end Lipreading with Channel-aware Feature SelectionabstractLipreading is a task of decoding the movement of the speaker’s lip region into text. In recent years, lipreading methods based on deep neural network have attracted widespread attention, and the accuracy has far surpassed that of experienced human lipreaders. The visual differences in some phonemes are extremely subtle and pose a great challenge to lipreading. Most of the lipreading existing methods do not process the extracted visual features, which mainly suffer from two problems. First, the extracted features contain lot of useless information such as noise caused by differences in speech speed and lip shape, for example. In addition, the extracted features are not abstract enough to distinguish phonemes with similar pronunciation. These problems have a bad effect on the performance of lipreading. To extract features from the lip regions that are more distinguishable and more relevant to the speech content, this article proposes an end-to-end deep neural network-based lipreading model (LCSNet). The proposed model extracts the short-term spatio-temporal features and the motion trajectory features from the lip region in the video clips. The extracted features are filtered by the channel attention module to eliminate the useless features and then used as input to the proposed Selective Feature Fusion Module (SFFM) to extract the high-level abstract features. Afterwards, these features are used as input to the bidirectional GRU network in time order for temporal modeling to obtain the long-term spatio-temporal features. Finally, a Connectionist Temporal Classification (CTC) decoder is used to generate the output text. The experimental results show that the proposed model achieves a 1.0% CER and 2.3% WER on the GRID corpus database, which, respectively, represents an improvement of 52% and 47% compared to LipNet. Feng Xue 0002, Kang Liu 0024, Zikun Hong, Mingwei Cao, Dan Guo 0001, Richang Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2021 | Accurate 3-D Reconstruction Under IoT Environments and Its Applications to Augmented RealityabstractWith the remarkable development of sensor devices and the Internet of Things (IoT), today's researchers can easily know what changes have taken place in the real world by acquiring a 3-D model. Conversely, a large amount of image data promotes the development of perceptual computing technology. In this article, we focus on modeling 3-D scenes from the multisource image data obtained from the IoT with cameras. Although great progress has been made in 3-D reconstruction, it is still challenging to recover the 3-D model from IoT data because the captured images are usually noisy, incomplete, varying scale, and with repetitive structures or features. In this article, we propose an accurate 3-D reconstruction method under IoT environments for perceptual computing of the scene. This method consists of sparse, dense, and surface reconstruction processes, which can gradually recover high-quality geometric models from the image data and efficiently deal with various repetitive structures. By analyzing the reconstructed model, we can detect the changes of scenes. We evaluate the proposed method on the benchmark data sets (i.e., tanks and temples) and publicly available data sets(in which samples usually contain repeated structures, lighting change, and different scales). Experimental results show that the proposed method outperforms the state-of-the-art methods according to the standard evaluation metric. We also use our method to enhance the real scenes with virtual objects, thus producing promising results. Mingwei Cao, Liping Zheng, Wei Jia 0001, Huimin Lu 0001, Xiaoping Liu 0003 |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Joint 3D Reconstruction and Object Tracking for Traffic Video Analysis Under IoV EnvironmentabstractBenefits from artificial intelligence and the Internet of Vehicles (IoV), Management of modern transportation have great progress, especially in urban areas. However, traditional traffic video analysis and visualization are usually conducted in offsite and textural environments, i.e., text and number, which do not promote user's sensorial perception and interaction. Thus, the problem that how to use modern novel techniques to analyze traffic video for improving intelligent transportation is so emergency. In this paper, we introduce a joint 3D reconstruction and object tracking approach to traffic video analysis under the IoV environment, which is an integrative framework and consists of 3D reconstruction, object detection, and visual tracking. The 3D reconstruction system is connected to the Internet of Vehicles and integrated into the system to retrieve image data for recovering the 3D model of vehicles, and then, visualizing vehicle trajectories in real-time by augmented reality. And the system can also locate the vehicle's position in real-time. The experiments in both laboratory and practice show great feedback, which will effectively contribute to intelligent transportation. Mingwei Cao, Liping Zheng, Wei Jia 0001, Xiaoping Liu 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Real-time video stabilization via camera path correction and its applications to augmented reality on edge devices
Mingwei Cao, Liping Zheng, Wei Jia 0001, Xiaoping Liu 0003 |
Comput. Commun. | 1 |
| 2020 | Constructing big panorama from video sequence based on deep local feature
Mingwei Cao, Liping Zheng, Wei Jia 0001, Xiaoping Liu 0003 |
Image Vis. Comput. | 1 |
| 2018 | Evaluation of Local Features for Structure from Motion
Mingwei Cao, Wei Jia 0001, Yujie Li 0001, Zhihan Lyu, Liping Zheng, Xiaoping Liu 0003 |
Multim. Tools Appl. | 1 |
| 2017 | Robust bundle adjustment for large-scale structure from motion
Mingwei Cao, Wei Jia 0001, Shanglin Li, Xiaoping Liu 0003 |
Multim. Tools Appl. | 1 |
| 2008 | Bandwidth Efficient Cooperative Diversity Scheme Based on Constellation RotationabstractA bandwidth efficient three-user cooperative diversity scheme is proposed. Each user has two partners who receive its symbols and rotate the phases of these symbols. Then, two partners retransmit the real parts and imaginary parts of these rotated symbols, respectively. This scheme can offer diversity order of three, and the bandwidth efficiency is only decreased by 1/2 compared to a noncooperative diversity scheme. The optimal rotation angles are derived when QPSK and 8PSK modulations are employed, respectively. Then, a new symbol mapping for 8PSK modulation is proposed, and it is a simple way to improve the system performance. Mingwei Cao, Guangguo Bi |
IEEE Signal Process. Lett. | 1 |