Sheng Miao

dblp:59/3196 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 FreeFix: Boosting 3D Gaussian Splatting via Fine-Tuning-Free Diffusion Models
abstract
Neural Radiance Fields and 3D Gaussian Splatting have advanced novel view synthesis, yet still rely on dense inputs and often degrade at extrapolated views. Recent approaches leverage generative models, such as diffusion models, to provide additional supervision, but face a tradeoff between generalization and fidelity: fine-tuning diffusion models for artifact removal improves fidelity but risks overfitting, while fine-tuning-free methods preserve generalization but often yield lower fidelity. We introduce FreeFix, a fine-tuning-free approach that pushes the boundary of this trade-off by enhancing extrapolated rendering with pretrained image diffusion models. We present an interleaved 2D-3D refinement strategy, showing that image diffusion models can be leveraged for consistent refinement without relying on costly video diffusion models. Furthermore, we take a closer look at the guidance signal for 2D refinement and propose a per-pixel confidence mask to identify uncertain regions for targeted improvement. Experiments across multiple datasets show that FreeFix improves multiframe consistency and achieves performance comparable to or surpassing fine-tuning-based methods, while retaining strong generalization ability. Our project page is at https://xdimlab.github.io/freefix.
Zisen Shao, Sheng Miao, Dongfeng Bai, Yiyi Liao
3DV3
2026 Towards Depth Foundation Models: Recent Trends in Vision-Based Depth Estimation
abstract
Depth estimation is a fundamental task in 3D computer vision, crucial for applications such as 3D reconstruction, free-viewpoint rendering, robotics, autonomous driving, and AR/VR technologies. Traditional methods relying on hardware sensors like LiDAR are often limited by their high costs, low resolution, and sensitivity to the environment, limiting their applicability to real-world scenarios. Recent advances in vision-based methods offer a promising alternative, yet they face challenges in generalization and stability due to either the low capacity of model architectures or reliance on domain-specific and small-scale datasets. The emergence of scaling laws and foundation models in other domains has inspired the development of “depth foundation models”: deep neural networks trained on large datasets with strong zeroshot generalization capabilities. This paper surveys the evolution of deep learning architectures and paradigms for depth estimation across monocular, stereo, multiview, and monocular video settings. We explore the potential of these models to address existing challenges and we also provide a comprehensive overview of large-scale datasets that can facilitate their development. By identifying key architectures and training strategies, we aim to highlight the path towards robust depth foundation models, offering insights for future research and applications.
Zhen Xu 0008, Sida Peng, Haotong Lin, Jiahao Shao, Peishan Yang, Qinglin Yang, Sheng Miao, Yifan Wang 0026, Ruizhen Hu, Yiyi Liao, Xiaowei Zhou 0001, Hujun Bao
Comput. Vis. Media9
2025 EVolSplat: Efficient Volume-based Gaussian Splatting for Urban View Synthesis
abstract
Novel view synthesis of urban scenes is essential for autonomous driving-related applications. Existing NeRF and 3DGS-based methods show promising results in achieving photorealistic renderings but require slow, per-scene optimization. We introduce EVolSplat, an efficient 3D Gaussian Splatting model for urban scenes that works in a feed-forward manner. Unlike existing feed-forward, pixelaligned 3DGS methods, which often suffer from issues like multi-view inconsistencies and duplicated content, our approach predicts 3D Gaussians across multiple frames within a unified volume using a 3D convolutional network. This is achieved by initializing 3D Gaussians with noisy depth predictions, and then refining their geometric properties in 3D space and predicting color based on 2D textures. Our model also handles distant views and the sky with a flexible hemisphere background model. This enables us to perform fast, feed-forward reconstruction while achieving real-time rendering. Experimental evaluations on the KITTI-360 and Waymo datasets show that our method achieves state-of-the-art quality compared to existing feedforward 3DGS- and NeRF-based methods.
Sheng Miao, Jiaxin Huang 0012, Dongfeng Bai, Xu Yan 0005, Yue Wang 0020, Andreas Geiger 0001, Yiyi Liao
CVPR1
2025 Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting
Jiaxin Huang 0012, Sheng Miao, Bangbang Yang, Yuewen Ma, Yiyi Liao
ICCV2
2025 Application of visual attribute transfer technology in analysing changes in emotional expression in picture books
abstract
Abstract In picture books, readers can obtain different emotional perceptions according to different image style attributes. Artists often use different combinations of colours, textures, materials, and other style elements in images to convey different emotions in their creations. Especially in picture books for children, there is a strong correlation between the perceived effect of the work and the accuracy and degree of emotional expression. In the process of creating picture books, various factors will affect the efficiency of artists trying to transfer styles to meet their creative needs. With the development of image style transfer technology based on a deep convolutional neural network, artists can use this technology to create works with different styles of emotional changes efficiently. In this paper, we select illustrations of picture books and use deep convolutional neural networks to transfer image styles from three aspects: colour style transfer, texture style, and material style transfer. Through sampling survey experiments, we discuss the changes in image attributes, emotional expression, and emotional perception in picture books for children. The survey results found that the most direct and evident influence on the emotional changes of picture book images is the transfer of colour style attributes, material style attributes, and texture style attributes. The results of this study can provide a valuable reference for improving the accuracy of emotional expression, the depth of meaning extension, and the height of artistic value in picture books for children during the process of an artist's creation. This research stands out by systematically analysing the distinct impact of each style attribute transfer, offering a comprehensive framework that can be utilized by artists and technologists alike to enhance the emotional and artistic quality of children's picture books.
Yue Wang 0138, Yansu Qi, Sheng Miao, Weijun Gao 0002
Expert Syst. J. Knowl. Eng.4
2024 Efficient Depth-Guided Urban View Synthesis
Sheng Miao, Jiaxin Huang 0012, Dongfeng Bai, Weichao Qiu, Andreas Geiger 0001, Yiyi Liao
ECCV (30)1
2023 Augmented Digital Twins for Predictive Automatic Regulation and Fault Alarm in Sewage Plan
abstract
In this paper, Digital Twins(DT) is combined with the sewage plant. Through Digital Twins, the actual needs are analyzed to solve the problems existing in the sewage plant. Combined with Augmented Reality(AR), Machine Learning(ML) and automatic control algorithms, various functions of sewage plant can be achieved. The system uses Long Short Term Memory(LSTM), Gate Recurrent Unit(GRU) and Fuzzy Neural Network(FNN) to predict the Chemical Oxygen Demand(COD) concentration in water quality. By using these algorithms, the Digital Twins Sewage Plant(DTSP) can be better interacted with workers. Through remote control, fault alarm, automatic regulation and prediction, Digital Twins can improve the efficiency of sewage treatment.
Zhihan Lyu, Sheng Miao
ACM Multimedia4
2023 Application of hybrid artificial bee colony algorithm based on load balancing in aerospace composite material manufacturing
Yufang Wang, Jiarong Ge, Sheng Miao, Tianhua Jiang, Xiaoning Shen
Expert Syst. Appl.3
2022 Monitoring of Green Tide Under Cloud Based on GOCI-II Data
abstract
The outbreak of green tides has caused various degrees of damage to coastal areas, and the extraction of green tides based on optical data is vulnerable to cloud. Based on GOCI$-\Vert$(Geostationary Ocean Color Imager -$\Vert)$data, this paper proposes a multi-band threshold combination method for the extraction of green tide under clouds by analyzing the green tide sensitive wavebands. The monitoring results of this algorithm are compared with the monitoring results two existing algorithms (NDVI (normalized difference vegetation index) threshold method and IGAG algorithm), and it is found that the multi-band threshold combination method has better results in extracting green tide under clouds, and can obtain more detailed and comprehensive green tide information. Compared with visual interpretation results, the extraction accuracy of thin cloud and thick cloud is 80% and 50%, respectively. The extraction accuracy is both due to NDVI and IGAG methods, indicating that the multi-band threshold combination method has higher accuracy.
Xiangrong Xin, Xirong Liu, Sheng Miao
IGARSS3
2022 Open-set iris recognition based on deep learning
abstract
Abstract The existing iris recognition methods offer excellent recognition performance for known classes, but they do not consider the rejection of unknown classes. It is important to reject an unknown object class for a reliable iris recognition system. This study proposes open‐set iris recognition based on deep learning. In the method, by training the deep network, the extracted iris features are clustered near the feature centre of each kind of iris image. Then, the authors build an open‐class features outlier network (OCFON) containing distance features, which maps the features extracted by the deep network to a new feature space and classifies them. Finally, the unknown class samples are determined by a SoftMax probability threshold. The authors conducted experiments on the open iris dataset constructed using the iris datasets CASIA‐Iris‐Twins and CASIA‐Iris‐Lamp. The experiment shows that the proposed method has good open‐set iris recognition performance, can effectively distinguish iris samples of unknown classes, and has little impact on the recognition ability of known classes of iris samples.
Shipeng Zhao, Sheng Miao
IET Image Process.3
2022 A Multi-Target Passive Location Method Based on GDOP Value and Beam Resolution
abstract
In passive location systems on the ground, the judgment and location of multi-target is more challenging compared with the case of single target. In this paper, we propose a method for multi-target identification and location in an arbitrary structure with three base stations (BSs). First of all, we discuss the scene of multi-targets judgment based on geometric dilution of precision (GDOP) value. Secondly, we propose an algorithm that calculates the system coverage radius based on arbitrary three BS structures. The algorithm helps to identify the number of targets for unsupervised learning. Finally, we locate each target individually located again based on the linear constrained minimum variance (LCMV) beam former and time difference of arrival (TDOA) algorithm. In the simulations, we analyzed the location dispersion under different signal-to-noise ratio (SNR), then calculated the termination threshold of the k-means algorithm under different SNR. The simulation results show that, compared to the probability hypothesis density (PHD) filter and TDOA-angle-of-arrival (AOA) joint algorithm, the proposed method can increase more than 12.5% and 15.6% points. With the increase of the number of targets, the running time of our algorithm is controllable with better stability.
Sheng Miao
Int. J. Pattern Recognit. Artif. Intell.1
2019 Probabilistic guided polycystic ovary syndrome recognition using learned quality kernel
Dongyun He, Sheng Miao, Xiaoli Tong, Minjia Sheng
J. Vis. Commun. Image Represent.3
2017 An Exploratory Case Study to Support Young Children with Spinal Muscular Atrophy (SMA)
abstract
In this paper, we describe a preliminary case study that examines the challenges faced by very young children with Type I Spinal Muscular Atrophy (SMA) and how technology may help these children live a more independent life. Several input solutions were examined to support interaction between a young patient and computer-based systems. We started working with the patient when he was two and a half years old. The challenges observed and lessons learned regarding both working with very young children with severe disabilities and the use of specific technical solutions are discussed.
Sheng Miao, Ziying Tang, Jinjuan Feng, Amanda Jozkowski, Molly Lichtenwalner
ASSETS1
2017 Utilizing human processing for fuzzy-based military situation awareness based on social media
abstract
With the increasing development of social media, we now face large amounts of up-to-date data, new information resources, and fast and transparent information propagating methods. They have changed the way we understand the world, which means we can have a different perspective when computing situation awareness. In this paper, we give a detailed explanation on how social media data affect situation awareness and our focus is set on the military environment. We discuss why subject matter expert based human computation is necessary and essential for this procedure and how to involve it in the situation awareness architecture. The goal of our paper is to give readers some suggestions on how to use social media and human processing power within the military situational awareness domain.
Sheng Miao, Ziying Tang
FUZZ-IEEE1
2015 Integrating complementary/contradictory information into fuzzy-based VoI determinations
abstract
In today's military environment vast amounts of disparate information are available. To aid situational awareness it is vital to have some way to judge information importance. Recent research has developed a fuzzy-based system to assign a Value of Information (VoI) determination for individual pieces of information. This paper presents an investigation of the effect of integrating subsequent complementary and/or contradictory information into the VoI process. Specifically, the idea of using complementary and/or contradictory new information to impact the previously used fuzzy membership values for the information content characteristic applied in the VoI calculations is shown to be a particularly suitable approach.
Sheng Miao, Robert J. Hammell II, Ziying Tang, Tim Hanratty, John Dumer, John T. Richardson
CISDA1
2006 Artificial Neural Network Methodology for Soil Liquefaction Evaluation Using CPT Values
Ben-yu Liu, Liao-yuan Ye, Mei-ling Xiao, Sheng Miao
ICIC (1)4
2006 Peak Ground Velocity Evaluation by Artificial Neural Network for West America Region
Ben-yu Liu, Liao-yuan Ye, Mei-ling Xiao, Sheng Miao
ICONIP (2)4
2006 Artificial Neural Network Methodology for Three-Dimensional Seismic Parameters Attenuation Analysis
Ben-yu Liu, Liao-yuan Ye, Mei-ling Xiao, Sheng Miao, Jing-yu Su
ISNN (2)4