Cheng Shang

dblp:121/1048 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Consistent 3D Human Reconstruction From Monocular Video: Learning Correctable Appearance and Temporal Motion Priors
abstract
Recent advancements in rendering dynamic humans using NeRF and 3D Gaussian splatting have made significant progress, leveraging implicit geometry learning and image appearance rendering to create digital humans. However, in monocular video rendering, there are still challenges in rendering subtle and complex motion from different viewpoints and states, primarily due to the imbalance of viewpoints. Additionally, ensuring continuity between adjacent frames when rendering from novel and free viewpoints remains a difficult task. To address these challenges, we first propose a pixel-level motion correction module that adjusts the errors in the learned representation between different viewpoints. We also introduce a temporal information-based model to improve motion continuity by leveraging adjacent frames. Experimental results on dynamic human rendering, using the NeuMan, ZJU-Mocap, and People-Snapshot datasets, demonstrate that our method outperforms state-of-the-art techniques both quantitatively and qualitatively.
Cheng Shang, Liang An 0001, Jiajun Zhang 0012, Yuxiang Zhang 0006, Jidong Tian, Yebin Liu, Xubo Yang
IEEE Trans. Vis. Comput. Graph.1
2026 DanceAgent: Dance Movement Refinement With LLM Agent
abstract
Recent research on motion generation and text-to-motion synthesis focus on coarse-grained motion descriptions, neglecting fine-grained motion details and motion quality refinement. Additionally, current text-to-motion models, such as MotionGPT, lack multi-turn interaction capabilities, relying on single-turn and single-modality transformations, which limit their ability to integrate information from different modalities across interaction stages. These gaps leave critical questions, such as "How well is the motion performed" and "How can it be refined?" largely unaddressed. To address these issues, first, we introduce two fine-grained dance datasets-one focusing on jazz dance and the other on folk dance, which we have independently collected. Second, considering that dance motions are inherently complex and consist of long sequential actions, we introduce both global and local optimization during the motion encoding phase and employ Hidden Markov Model (HMM) temporal modeling to capture differential features between correct and incorrect movements, thereby optimizing the training process. Finally, we propose a multi-turn historical dialogue framework that enables three stages generation-motion assess, text instructions, and motion refinement-for input videos. This framework assists dance beginners by providing feedback on their movements, offering textual instructions, and delivering motion-based refinement. Experimental results on the jazz dance and folk dance datasets demonstrate that our method surpasses existing approaches in both quantitative and qualitative metrics, establishing a new benchmark for motion-text generation in the field of dance training.
Cheng Shang, Liang An 0001, Jiajun Zhang 0012, Yuxiang Zhang 0006, Yebin Liu, Xubo Yang
IEEE Trans. Vis. Comput. Graph.1
2025 FaceCapGes: Real-Time Frame-by-Frame Gesture Generation from Audio, Facial Capture, and Head Pose
Jun Hanaizumi, Cheng Shang, Xubo Yang
CGI (3)2
2025 Dynamic Quadruple Optimization Based Transfer Learning for Animal Biometric Identification
abstract
With the progress of computer vision and machine learning, the research of object detection and pedestrian recognition has demonstrated significant performance. However, the identification studies in domestic animals, especially in the same species of domestic animals, remains a significant challenge. His study focuses on distinguishing cashmere and dairy goats, which share similar traits. Our contributions are: (1) Proposing a dynamic quadruple optimization algorithm to optimize goat images from local and global dimensions, enhancing network representation with a multi-branch structure; (2) Introducing a novel transfer learning algorithm based on goat granularity to preview dataset knowledge; (3) Validating our approach on our goat dataset and a public bird dataset. We achieved recognition accuracies of 95% for cashmere goats, 94.04% for dairy goats, and 82.48% on the public dataset, demonstrating the effectiveness of our methods for animal biometric identification.
Cheng Shang, Chong He, Xubo Yang, Yongliang Qiao, Meili Wang 0001
CSCWD1
2025 A universal sampling method based on feature and structural comprehensive proximity measure
Jinhui Pang, Cheng Shang, Ziyu Jia, Peng Hao 0003, Xiaoshuai Hao
Neurocomputing2
2024 Free-view Rendering of Dynamic Human from Monocular Video Via Modeling Temporal Information Globally and Locally among Adjacent Frames
abstract
Recent research developments on rendering dynamic humans using neural radiance fields are remarkable. These methods often utilize learning implicit geometry and image appearance rendering for digital humans. However, keeping the complex and fast motions in detail, such as fingers, clothes, and faces, remains a challenge. Inspired by temporal information from human motion, we propose an architecture among adjacent frames by constructing a model on global and local levels. For the global level, we propose a hidden Markov model (HMM)based method to capture the global similarity among adjacent frames. At the local level, we introduce a module composed of a multi-head attention mechanism on a triplet canonical space structure for patch-level local temporal information. Experiments on two public datasets of dynamic human rendering (ZJU-MoCap and the People-Snapshot dataset) demonstrate that the proposed method outperforms advanced methods quantitatively and qualitatively.
Cheng Shang, Jidong Tian, Jiannan Ye, Xubo Yang
ICME1
2024 Monitoring Color for Small Water Bodyies by Fusing Sentinel-2 Multispectral Data and Gaofen -1 Panchromatic Imagery
abstract
Water color is an intuitive parameter for indicating the condition of natural waters. It can be captured by satellite images on a large scale and over a long period of time. However, it can be challenging to capture this information for small bodies of water due to the trade-off between spatial and spectral resolution of imaging sensors. We propose a deep learning-based method, called Spectral Angle (SA) Consistent Unsupervised Generative Adversarial Network (SA-UGAN), to fuse Sentinel-2 multispectral data and China's Gaofen-1 panchromatic images. The loss function is optimized to emphasize spectral preservation through an SA-based loss. SA-UGAN outperforms three traditional approaches in terms of quality assessment. The hue angle, which is used to describe water color, was monitored based on the fused image. The results showed that the spatial distribution of hue angle obtained from the SA-UGAN fused images was consistent with the original results.
Xueer Geng, Baoyin He, Feng Ling 0003, Cheng Shang
IGARSS7
2024 Neural foveated super-resolution for real-time VR rendering
abstract
Abstract As virtual reality display technologies advance, resolutions and refresh rates continue to approach human perceptual limits, presenting a challenge for real‐time rendering algorithms. Neural super‐resolution is promising in reducing the computation cost and boosting the visual experience by scaling up low‐resolution renderings. However, the added workload of running neural networks cannot be neglected. In this article, we try to alleviate the burden by exploiting the foveated nature of the human visual system, in a way that we upscale the coarse input in a heterogeneous manner instead of uniform super‐resolution according to the visual acuity decreasing rapidly from the focal point to the periphery. With the help of dynamic and geometric information (i.e., pixel‐wise motion vectors, depth, and camera transformation) available inherently in the real‐time rendering content, we propose a neural accumulator to effectively aggregate the amortizedly rendered low‐resolution visual information from frame to frame recurrently. By leveraging a partition‐assemble scheme, we use a neural super‐resolution module to upsample the low‐resolution image tiles to different qualities according to their perceptual importance and reconstruct the final output adaptively. Perceptually high‐fidelity foveated high‐resolution frames are generated in real‐time, surpassing the quality of other foveated super‐resolution methods.
Jiannan Ye, Xiaoxu Meng, Daiyun Guo, Cheng Shang, Haotian Mao, Xubo Yang
Comput. Animat. Virtual Worlds4
2023 Individual identification of cashmere goats via method of fusion of multiple optimization
abstract
Abstract Facial recognition technology and related research have matured over time, but research in the field of individual animal recognition is still very limited. Therefore, this article focuses on the identification of cashmere goats with similar characteristics. First, the single shot multibox detector network was used to process the dataset. Next, transfer learning was applied to learn the characteristics of the goats, as well as the loss function is composed of Triplet Loss and Label Smoothing CrossEntropy Loss function. The result of Label Smoothing CrossEntropy Loss function is fused by multiple different branches, which is convenient for classification. We added a small number of images of 24 different breeds of sheep to each cashmere goat dataset with different ID to promote the distance between training individuals, and then used the trained model to find the number of goats with the lowest recognition accuracy. The Cycle‐Consistent Adversarial Network (Cycle‐GAN) learned the goat dataset with a high error rate in individual identification. Unlike previous studies using the Cycle‐GAN, we took the novel approach of using this network to learn and combine the features seen in photos of cashmere goats. Since the learned features were all observed in the same goats, this method achieved better results in learning the features of the goats. Finally, we found that recognition can be performed on this data with an accuracy of 93.75%. These results suggest that identification based on deep learning has a high accuracy rate, as well as great value in identifying individual cashmere goats.
Cheng Shang, Hongke Zhao, Meili Wang 0001, Qiang Gao 0015
Comput. Animat. Virtual Worlds1
2022 Cattle behavior recognition based on feature fusion under a dual attention mechanism
Cheng Shang, Meili Wang 0001, Qiang Gao 0015
J. Vis. Commun. Image Represent.1
2022 Superresolution Land Cover Mapping Using a Generative Adversarial Network
abstract
Superresolution mapping (SRM) is a commonly used method to cope with the problem of mixed pixels when predicting the spatial distribution within low-resolution pixels. Central to the popular SRM method is the spatial pattern model, which is utilized to represent the land cover spatial distribution within mixed pixels. The use of an inappropriate spatial pattern model limits such SRM analyses. Alternative approaches, such as deep-learning-based algorithms, which learn the spatial pattern from training data through a convolutional neural network, have been shown to have considerable potential. Deep learning methods, however, are limited by issues such as the way the fraction images are utilized. Here, a novel SRM model based on a generative adversarial network (GAN), GAN-SRM, is proposed that uses an end-to-end network to address the main limitations of existing SRM methods. The potential of the proposed GAN-SRM model was assessed using four land cover subsets and compared to hard classification and several popular SRM methods. The experimental results show that of the set of methods explored, the GAN-SRM model was able to generate the most accurate high-resolution land cover maps.
Cheng Shang, Xiaodong Li 0006, Giles M. Foody, Feng Ling 0003
IEEE Geosci. Remote. Sens. Lett.1
2022 Spatiotemporal Reflectance Fusion Using a Generative Adversarial Network
abstract
The spatiotemporal reflectance fusion method is used to blend high-temporal and low-spatial resolution images with their low-temporal and high-spatial resolution counterparts that were previously acquired by various satellite sensors. Recently, a wide variety of learning-based solutions have been developed, but challenges remain. These solutions usually require two sets of data acquired before and after the prediction time, making them unsuitable for near-real-time predicting. The solutions are always trained band by band and thus do not consider the spectral correlation. High-resolution temporal changes are difficult to reconstruct accurately with the network structure used, which lowers the accuracy of the fusion result. To address these problems, this study proposes a novel spatiotemporal adaptive reflectance fusion model using a generative adversarial network (GASTFN). In GASTFN, an end-to-end network, including a generative and discriminative network, is simultaneously trained for all spectral bands. The proposed model can be applied to the one-pair case, consider the spectral correlation of each band, and improve the process of producing super-resolution imagery by adopting the discriminative network for image reflectance values rather than temporal changes in reflectance. The proposed model has been verified with two actual satellite data sets acquired in heterogeneous landscapes and areas with abrupt changes, with a comparison of the state-of-art methods. The results show that GASTFN can generate the most accurate fusion images with more detailed textures, more realistic spatial shapes, and higher accuracy, demonstrating that the GASTFN is effective for predicting near-real-time changes in image reflectance and preserves the most valuable spatial information.
Cheng Shang, Xiaodong Li 0006, Yihang Zhang 0001, Feng Ling 0003
IEEE Trans. Geosci. Remote. Sens.1
2020 Bas-reliefs modelling based on learning deformable 3D models
abstract
Bas-relief is a special art-form, which can be used as decorations. Deformable 3D models are usually used in image segmentation and 3D reconstruction, aiming at the reconstruction of certain category especially. To quickly reconstruct different objects of the same category and take the psychological process and visual characteristics of human observation images into account, in this paper, we propose a bas-reliefs generation method for flowers based on deformable 3D models which is a type of parametric model that can deform shapes by changing related parameters. First, we create a dataset by labelling images with keypoints. Then, we use Non Rigid Structure from Motion (NRSfM) algorithm to estimate viewpoints. A deformable 3D model is reconstructed by Visual Hull. Next, we deform the 3D model to a variant of the object according to an input 2D image. After that, we enhance and smoothen the image by applying gamma correction on the gray scale image with gradient. Finally, we calculate height values as details using illumination model and add them to the variant. The experimental results demonstrate that our method can reconstruct bas-reliefs for flowers in different shapes with complex background, which is also proven to be able to extend to other categories, such as birds and fishes.
Siyuan Zhu, Cheng Shang
IJCNN2