Tianyu Shen

dblp:238/0027 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The Structure-Equivalent Prior: Unifying Temporal Dynamics and 3D Evolution in 4D Latent Space
abstract
Recent advances in deep learning-based 3D representation have achieved remarkable success, particularly in modeling static high-fidelity geometries. However, the extension of these techniques to dynamic 3D scenes introduces a critical challenge of effectively representing spatio-temporal dependencies, i.e., jointly modeling detailed spatial structures within frames and temporal dynamics across frames. To address this challenge, this paper proposes that the temporal evolution observed in dynamic 3D scenes is fundamentally attributable to the deformation of underlying spatial structures. To capture this relationship, we introduce a unified continuous 4D latent space representation incorporating a structure-equivalence prior, named SEP-4D. The core of SEP-4D is an efficient 4D tensor decomposition-fusion approach. This method fuses decomposed learnable 2D feature planes via a plane-wise spatio-temporal fusion mechanism of planar distributions, explicitly enforcing the principle that temporal evolution originates from geometric deformations of the 3D structure. To mitigate the associated computational demands, we sample the 3D probability volumes generated by VAE-based fusion into a spatio-temporally consistent 4D latent representation. The efficacy of our approach is validated through experiments on the fundamental task of 4D occupancy reconstruction. Extensive results demonstrate that, by leveraging the inherent equivalence of temporal dynamics and structural deformation, our method achieves high-quality reconstruction across various sequence lengths. Notably, for 4-frame scenes, we attain an impressive 91.68% mIoU, significantly outperforming state-of-the-art baselines on standard benchmarks.
Jingyuan Gao, Tianyu Shen, Ruosen Hao, Te Guo 0003, Zhiwei Li 0011, Kunfeng Wang
AAAI2
2026 PSR: Proactive soft-orthogonal regulation for long-tailed class-incremental learning
Zhihan Fu, Shipeng Liao, Zerun Chen, Tianyu Shen
Pattern Recognit.6
2026 Self-Supervised Depth Completion Guided by 3D Perception and Geometry Consistency
abstract
Depth completion which aims at predicting dense depth maps from sparse depth measurements, plays a crucial role in many computer graphics and computer vision applications. Previous supervised learning based approaches have demonstrated overwhelming success in this task, while unsupervised high-precision depth completion without relying on the ground-truth data still remains challenging. One main drawback of most previous unsupervised solutions comes from the ignorance of 3D structural information, which often leads to inaccurate spatial propagation and mixed-depth problems. To alleviate the above challenges, this paper explores the utilization of 3D perceptual features and multi-view geometry consistency to devise a high-precision self-supervised depth completion method. Our key contribution is a 3D perceptual spatial propagation constructed with a point cloud representation and an attention weighting mechanism, to capture more reasonable and favorable neighbors during the depth propagation process. Based on the 3D perceptual spatial propagation, we also introduce multi-view geometric constraints between adjacent views to guide the optimization of the whole depth completion model, which achieves geometry consistent depth completion in a self-supervised manner. Extensive experiments on benchmark datasets of NYU-Depth-v2, VOID and KITTI Depth Completion demonstrate that the proposed model achieves the state-of-the-art depth completion performance compared with other unsupervised methods, and even competitive performance compared with previous supervised methods.
Tianyu Shen, Shi-Sheng Huang, Hua Huang 0001
IEEE Trans. Vis. Comput. Graph.2
2025 Computer-aided diagnosis of pituitary microadenoma on dynamic contrast-enhanced MRI based on spatio-temporal features
abstract
Computer-aided diagnosis (CAD) of pituitary microadenoma (PM) can assist doctors in decision-making, leading to improved lesion detection rates and diagnostic accuracy. However, the performance of existing CAD methods for PM detection has been hindered by the difficulty in obtaining high-quality segmentation results. This is primarily due to the small size of PM lesions and the relatively low resolution of Magnetic Resonance Imaging (MRI) images. To address these challenges, this paper proposes a new medical image detection and segmentation model based on spatio-temporal information. The proposed model aims to addresses the disease classification of PM by designing a network module based on multi-scale feature fusion. This module ensures comprehensive extraction of target semantic information while retaining clear spatial information, achieving classification from dynamic contrast-enhanced MRI(DCE-MRI) to identify positive PM samples. For the lesion segmentation of PM, after ROI Align alignment, the model further adds a semantic segmentation module named Dual-path Semantic Segmentation Module (DSSM) behind the mask head and classification head. This module captures more precise spatio-temporal semantic information, reducing accuracy loss and achieving pituitary segmentation. Finally, leveraging the results of pituitary detection, a feature pyramid network (FPN) layer is redesigned named Reuse Underlying Information Module (RUIM) to reuse low-level information, enhancing the detection capability for PM and thus achieving precise object detection and segmentation. The proposed model achieves an accuracy of 97.10% for PM, mAP of 50.24%, which is superior to multiple representative deep models for medical data. The code is available at https://github.com/BUCT-IUSRC/Research__PM-CAD .
Te Guo 0003, Jixin Luan, Jingyuan Gao, Tianyu Shen, Guolin Ma, Kunfeng Wang
Expert Syst. Appl.5
2025 MIPD: A Multi-Sensory Interactive Perception Dataset for Embodied Intelligent Driving
abstract
During the process of driving, humans usually rely on multiple senses to gather information and make decisions. Analogously, in order to achieve embodied intelligence in autonomous driving, it is essential to integrate multidimensional sensory information in order to facilitate interaction with the environment. However, the current multi-modal fusion sensing schemes often neglect these additional sensory inputs, hindering the realization of fully autonomous driving. This paper considers multi-sensory information and proposes a multi-modal interactive perception dataset named MIPD, enabling expanding the current autonomous driving algorithm framework, for supporting the research on embodied intelligent driving. In addition to the conventional camera, lidar, and 4D radar data, our dataset incorporates multiple sensor inputs including sound, light intensity, vibration intensity and vehicle speed to enrich the dataset comprehensiveness. Comprising 126 consecutive sequences, many exceeding twenty seconds, MIPD features over 8,500 meticulously synchronized and annotated frames. Moreover, it encompasses many challenging scenarios, covering various road and lighting conditions. The dataset has undergone thorough experimental validation, producing valuable insights for the exploration of next-generation autonomous driving frameworks. Data, development kit and more details will be available athttps://github.com/BUCT-IUSRC/Dataset__MIPD
Zhiwei Li 0011, Tingzhen Zhang, Meihua Zhou, Dandan Tang, Wenzhuo Liu, Qiaoning Yang, Tianyu Shen, Kunfeng Wang, Huaping Liu 0001
IEEE Trans. Intell. Transp. Syst.8
2024 SMFuse: Two-Stage Structural Map Aware Network for Multi-focus Image Fusion
Tianyu Shen, Hui Li 0037, Chunyang Cheng, Xiaoning Song
ICPR (22)1
2024 Fixed-Time Composite Learning Fuzzy Control With Disturbance Rejection for Uncertain Engineering Systems Toward Industry 5.0
abstract
Intelligent control is a crucial technology for realizing Industry 5.0, which makes industrial engineering systems more efficient, robust, and resilient. It is noteworthy that uncertainties and disturbances will inevitably be detrimental to the control performances of Industry 5.0 engineering applications. To deal with these issues, we propose a novel super-twisting-like continuous extended state observer-based fixed-time composite learning fuzzy control scheme and apply it to a typical engineering system. Unlike conventional fixed-time adaptive fuzzy control methods that update parameters merely by closed-loop stability conditions, the proposed fixed-time control scheme utilizes both tracking errors and prediction errors to update parameters compositely, which achieves better-tracking performance and fuzzy approximation precision. First, fuzzy logic systems (FLSs) are developed to identify the unknown model functions in the Industry 5.0 engineering system. Second, to deal with the remaining approximation errors of the FLSs, parameter uncertainties, and external disturbances, the novel super-twisting-like continuous extended state observers are designed to estimate these lumped disturbances. Third, the prediction errors that indicate the fuzzy approximation precision are constructed by developing fixed-time parallel estimators. Moreover, rigorous Lyapunov stability analysis is carried out to illustrate the fixed-time convergence of the entire closed-loop control system. Finally, the proposed control scheme is applied to a practical buck converter engineering system toward Industry 5.0, and comparative hardware experiments verified the advantages of the control scheme.
Jinlin Sun, Yafei Chang, Tianyu Shen, Shihong Ding
IEEE Trans. Syst. Man Cybern. Syst.4
2023 Dynamic View Synthesis with Spatio-Temporal Feature Warping from Sparse Views
abstract
Significant progress has been made in realizing novel view synthesis of dynamic scenes from sparse input views. However, achieving spatio-temporal consistency in dynamic view synthesis remains to be challenging for previous approaches, since the spatio-temporal correlation for view synthesis has not been fully explored. In this paper, we propose a spatio-temporal feature warping (STFW) mechanism, which can be embedded into a deep model to produce high-quality and spatio-temporally consistent view synthesis results. The two core components of STFW are: (1) a spatial feature warping (SFW) module, which enables adaptive perception of multi-view context-consistent geometric information with a compact point cloud representation, and (2) a temporal feature warping (TFW) module that implicitly models the dynamic geometry by approaching the pixel shift in image coordinate. In the optimization process of view synthesis, the SFW and TFW are integrated to exploit the spatio-temporal correlation cues across sparse input views and novel views. Leveraging the STFW, we further build an end-to-end dynamic view synthesis model with sparse input views. Qualitative and quantitative evaluation on public multi-view datasets demonstrate that our view synthesis pipeline achieves better performance compared to previous methods in terms of visual quality.
Deqi Li, Shi-Sheng Huang, Tianyu Shen, Hua Huang 0001
ACM Multimedia3
2023 A domain density peak clustering algorithm based on natural neighbor
abstract
Density peaks clustering (DPC) is as an efficient algorithm due for the cluster centers can be found quickly. However, this approach has some disadvantages. Firstly, it is sensitive to the cutoff distance; secondly, the neighborhood information of the data is not considered when calculating the local density; thirdly, during allocation, one assignment error may cause more errors. Considering these problems, this study proposes a domain density peak clustering algorithm based on natural neighbor (NDDC). At first, natural neighbor is introduced innovatively to obtain the neighborhood of each point. Then, based on the natural neighbors, several new methods are proposed to calculate corresponding metrics of the points to identify the centers. At last, this study proposes a new two-step assignment strategy to reduce the probability of data misclassification. A series of experiments are conducted that the NDDC offers higher accuracy and robustness than other methods.
Tao Du 0002, Jin Zhou 0003, Tianyu Shen
Intell. Data Anal.4
2023 Depth-Aware Multi-Person 3D Pose Estimation With Multi-Scale Waterfall Representations
abstract
Estimating absolute 3D poses of multiple people from monocular image is challenging due to the presence of occlusions and the scale variation among different persons. Among the existing methods, the top-down paradigms are highly dependent on human detection which is prone to the influence from inter-person occlusions, while the bottom-up paradigms suffer from the difficulties in keypoint feature extraction caused by scale variation and unreliable joint grouping caused by occlusions. To address these challenges, we introduce a novel multi-person 3D pose estimation framework, aided by multi-scale feature representations and human depth perceiving. Firstly, a waterfall-based architecture is incorporated for multi-scale feature representations to achieve a more accurate estimation of occluded joints with a better detection of human shapes. Then the global and local representations are fused for handling the effects of inter-person occlusion and scale variation in depth perceiving and keypoint feature extraction. Finally, with the guidance of the fused multi-scale representations, a depth-aware model is exploited for better 2D joint grouping and 3D pose recovering. Quantitative and qualitative evaluations on benchmark datasets of MuCo-3DHP and MuPoTS-3D prove the effectiveness of our proposed method. Furthermore, we produce an occluded MuPoTS-3D dataset and the experiments on it validate the superiority of our method for overcoming the occlusions.
Tianyu Shen, Deqi Li, Fei-Yue Wang 0001, Hua Huang 0001
IEEE Trans. Multim.1
2023 VirtualClassroom: A Lecturer-Centered Consumer-Grade Immersive Teaching System in Cyber-Physical-Social Space
abstract
Lecturers, as the guidance of the classroom, play a significant role in the teaching process. However, the lecturers’ sense of space immersion has been ignored in current virtual teaching systems. In this article, we explore the cyber–physical–social intelligence for Edu-Metaverse in cyber–physical–social space and specially design a lecturer-centered immersive teaching system, taking the social and lecturers’ factors into consideration. We call this system VirtualClassroom (V-Classroom). Specifically, we first introduce the cyber–physical–social system (CPSS) paradigm of V-Classroom so that the workflow is standardized and significantly simplified, and the systems can be constructed with off-the-shelf hardware. The key component of V-Classroom is a cyber-world representation of a physical-world classroom instrumented with sparse consumer-grade RGBD cameras for capturing the 3-D geometry and texture of the classrooms. We provide each V-Classroom lecturer with a physical device for sending 6DoF view-change messages and showing view-dependent content of the remote classroom. Following the above paradigm, we develop the V-Classroom algorithms, including V-Classroom depth algorithm (V-DA) and V-Classroom view algorithm (V-VA), to achieve the real-time rendering of remote classrooms. V-DA is dedicated to recovering accurate depth information of the classrooms while V-VA is devoted to real-time novel view synthesis. Finally, we illustrate our implemented CPSS-driven V-Classroom prototype, based on real-world classroom scenarios we collected, and discuss the main challenges and future direction.
Tianyu Shen, Shi-Sheng Huang, Deqi Li, Fei-Yue Wang 0001, Hua Huang 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2021 Binary thresholding defense against adversarial attacks
Yutong Wang 0001, Tianyu Shen, Hui Yu 0001, Fei-Yue Wang 0001
Neurocomputing3
2020 Quantized generalized maximum correntropy criterion based kernel recursive least squares for online time series prediction
Tianyu Shen, Min Han 0001
Eng. Appl. Artif. Intell.1
2020 A Systematic Review of the Personality of Robot: Mapping Its Conceptualization, Operationalization, Contextualization and Effects
abstract
Robots are becoming prevalent as they could socially interact with humans and provide service or companionship. As people attribute personality traits to machines, the personality of robot (POR) has attracted considerable scholarly attention from researchers of human-robot interaction. However, due to the complexity of personality, the ways to design personality into robotics vary on a wide range. This systematic review attempts to map the approaches to designing the personality of robot and understand its effects on human-robot interaction. Following the guidelines of the Preferred Reporting Items for Systematic Reviews and Meta-Analyses, a review of 40 peer-reviewed publications was conducted. The conceptualization, operationalization, contextualization and effects of POR were summarized in the review. In general, positive POR was preferred and associated with desirable social responses. Suggestions on future design of robotics were discussed. Specifically, it is recommended that the design of POR should match users’ expectations in different social contexts. Social cues such as eye gaze, gestures, and voice should be applied at a self-explanatory level to help users efficiently predict and engage with the behaviors of social robots.
Yi Mou, Changqian Shi, Tianyu Shen, Kun Xu 0006
Int. J. Hum. Comput. Interact.3
2020 Simultaneous Segmentation and Classification of Mass Region From Mammograms Using a Mixed-Supervision Guided Deep Model
abstract
Automatic diagnosis based on medical imaging necessitates both lesion segmentation and disease classification. Lesion segmentation requires pixel-level annotations while disease classification only requires image-level annotations. The two tasks are usually studied separately despite the latter problem relies on the former. Motivated by the close correlation between them, we propose a mixed-supervision guided method and a residual-aided classification U-Net model (ResCU-Net) for joint segmentation and benign-malignant classification. By coupling the strong supervision in the form of segmentation mask and weak supervision in the form of benign-malignant label through a simple annotation procedure, our method efficiently segments tumor regions while simultaneously predicting a discriminative map for identifying the benign-malignant types of tumors. Our network, ResCU-Net, extends U-Net by incorporating the residual module and the SegNet architecture to exploit multilevel information for achieving improved tissue identification. With experiments on a public mammogram database of INbreast, we validate the effectiveness of our method and achieve consistent improvements over state-of-the-art models.
Tianyu Shen, Chao Gou, Jiangong Wang, Fei-Yue Wang 0001
IEEE Signal Process. Lett.1
2020 Hierarchical Fused Model With Deep Learning and Type-2 Fuzzy Learning for Breast Cancer Diagnosis
abstract
Breast cancer diagnosis based on medical imaging necessitates both fine-grained lesion segmentation and disease grading. Although deep learning (DL) offers an emerging and powerful paradigm of feature learning for these two tasks, it is hampered from popularizing in practical application due to the lack of interpretability, generalization ability, and large labeled training sets. In this article, we propose a hierarchical fused model based on DL and fuzzy learning to overcome the drawbacks for pixelwise segmentation and disease grading of mammography breast images. The proposed system consists of a segmentation model (ResU-segNet) and a hierarchical fuzzy classifier (HFC) that is a fusion of interval type-2 possibilistic fuzzy c-means and fuzzy neural network. The ResU-segNet segments the masks of mass regions from the images through convolutional neural networks, while the HFC encodes the features from mass images and masks to obtain the disease grading through fuzzy representation and rule-based learning. Through the integration of feature extraction aided by domain knowledge and fuzzy learning, the system achieves favorable performance in a few-shot learning manner, and the deterioration of cross-dataset generalization ability is alleviated. In addition, the interpretability is further enhanced. The effectiveness of the proposed system is analyzed on the publicly available mammogram database of INbreast and a private database through cross-validation. Thorough comparative experiments are also conducted and demonstrated.
Tianyu Shen, Jiangong Wang, Chao Gou, Fei-Yue Wang 0001
IEEE Trans. Fuzzy Syst.1