EDBT 2026 Demo / reviewers in the wild / expert
Yuanhong Zhong
dblp:11/9048
· DBLP profile ↗
24ranked-venue papers
16as first author
22since 2021 · last 2026
0000-0001-5689-1146ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 12 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards personalized long-term learning modeling in knowledge tracing
Shanshan Wang 0008, Jianqi Qiu, Jiaxin Pang, Xun Yang 0001, Ke Xu 0011, Zhangling Duan, Yuanhong Zhong, Xingyi Zhang 0001 |
Expert Syst. Appl. | 7 |
| 2026 | BADD: Sample-Wise Reallocation of KL Divergence Supervision for Online Mutual Learning
Yuanhong Zhong |
IEEE Signal Process. Lett. | 2 |
| 2026 | Local-Global Feature Fusion for Enhancing 3D Human Pose EstimationabstractBased on its excellent capability to extract temporal features, transformer has been widely used in monocular 3D human pose estimation. However, due to its global perspective, it performs inadequately in extracting spatial features, which hinders breakthroughs in performance. In this paper, we propose a local-global feature fusion method based on GCN and transformer for 3D human pose estimation. Our method integrates GCN with multiscale transformer to extract local spatiotemporal features of poses. These are then integrated with the global spatiotemporal features extracted by vanilla transformer to reconstruct 3D human poses accurately. In addition, we introduce a hierarchical feature fusion method to better capturing the underlying 3D pose structure. It blends deep abstract features with shallow raw features. We evaluate our model on the Human3.6M and MPI-INF-3DHP datasets, and experimental results demonstrate that our approach outperforms existing state-of-the-art methods. We achieve advanced performance on both datasets with errors of 37.7mm and 16.4mm under MPJPE, respectively. The code and model are available at https://github.com/ygx7/LG3DPose. Yuanhong Zhong, Guangxia Yang, Daidi Zhong, Xun Yang 0001, Shanshan Wang 0008, Zhangling Duan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | ASCD: An Adaptive Framework for Long-Tailed Cognitive DiagnosisabstractCognitive Diagnosis Modeling is a fundamental task in intelligent education, intending to assess students’ mastery levels on knowledge concepts through interactions. Previous methodologies prioritized enhancing average diagnostic accuracy. However, they often neglected the Long-Tailed issue in interactions. To address this issue, we propose an Adaptive Self-Supervised Graph Learning for Cognitive Diagnosis (ASCD) framework that leverages relatively balanced sparse views to forcing the graph network to focus on long-tailed nodes, aiming to tackle the long-tailed problem in graph-based cognitive diagnosis. Additionally, we employ self-supervised manners to mitigate the impact of dropped information on other nodes. Our approach leverages the adaptive graph confusion method to create sparse views of the original student-exercise interaction graph. In these sparse views, both long-tailed and head students carry similar weights, pushing the graph network to allocate attention impartially across all students. We integrate two different graph confusion techniques into adaptive graph confusion to accommodate varying degrees of data sparsity: edge dropout and feature masking. ASCD can serve as a plug-and-play module integrated into any graph-based cognitive diagnosis model, enhancing its performance regarding long-tailed scenarios. Extensive experiments on real-world datasets show the effectiveness of our approach, especially on the students with much sparser interaction records. Shanshan Wang 0008, Xun Yang 0001, Yuanhong Zhong, Xingyi Zhang 0001, Meng Wang 0001 |
ACM Trans. Inf. Syst. | 4 |
| 2025 | Learning multiscale residual prototypes and global-local correspondence for video anomaly detection
Yongting Hu, Yuanhong Zhong |
Comput. Vis. Image Underst. | 2 |
| 2025 | A Two-Stage Framework With Memory for Anomaly Detection via Video Decomposition and Bidirectional ConsistencyabstractExisting unsupervised video anomaly detection methods based on prediction typically employ a memory module to limit the generalization ability of the network so that normal frames can be accurately reconstructed or predicted while abnormal frames cannot. These memory-based methods usually utilize memory to record the fusion prototypes of appearance and motion. However, the motion part of the fusion prototypes is obtained from implicit motion representation, which is incomplete and constrain the ability for abnormal detection. To tackle the above issue, we proposed a Two-Stage framework with Memory via video Decomposition and bidirectional Consistency (TSMDC), which employs explicit motion data to learn comprehensive motion prototypes and use video decomposition and bidirectional consistency to learn fine granularity and advanced prototypes. We first decompose the video clip into three components: motion, scene, and object. In the first stage, the features of the motion are extracted to obtain a comprehensive motion representation, and its prototypes are stored in the motion memory. For further use of the bidirectional information of video, we present a cascaded frame prediction network that is utilized to learn and record the advanced spatio-temporal prototypes in the second stage. Specifically, the fine-granularity features of the scene and object are extracted and fused to predict the future frame. Then the initial frame of the video clip is predicted based on bidirectional consistency and motion prototypes enhancement. And the advanced spatio-temporal prototypes of video are recorded in this process. Anomalies are evaluated using the combination anomaly score of the predicted future and initial frame.Extensive experimental results on three public datasets indicate the effectiveness of the proposed method. Code will be available at https://github.com/yangugu/TSMDC. Yuanhong Zhong, Yongting Hu, Ruyue Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Inter-Clip Feature Similarity Based Weakly Supervised Video Anomaly Detection via Multi-Scale Temporal MLPabstractThe major paradigm of weakly supervised video anomaly detection (WSVAD) is treating it as a multiple instance learning (MIL) problem, with only video-level labels available for training. Due to the rarity and ambiguity of anomaly, the selection of potential abnormal training sample is the prime challenge for WSVAD. Considering the temporal relevance and length variation of anomaly events, how to integrate the temporal information is also a controversial topic in WSVAD area. To address forementioned problems, we propose a novel method named Inter-clip Feature Similarity based Video Anomaly Detection (IFS-VAD). In the proposed IFS-VAD, to make use of both the global and local temporal relation, a Multi-scale Temporal MLP (MT-MLP) is leveraged. To better capture the ambiguous abnormal instances in positive bags, we introduce a novel anomaly criterion based on the Inter-clip Feature Similarity (IFS). The proposed IFS criterion can assist in discerning anomaly, as an additional anomaly score in the prediction process of anomaly classifier. Extensive experiments show that IFS-VAD demonstrates state-of-the-art performance on ShanghaiTech with an AUC of 97.95%, UCF-Crime with an AUC of 86.57% and XD-Violence with an AP of 83.14%. Our code implementation is accessible athttps://github.com/Ria5331/IFS-VAD. Yuanhong Zhong, Ruyue Zhu, Ping Gan, Xuerui Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Exploring Invariance Matters for Domain GeneralizationabstractDomain generalization (DG) aims to solve the problem of significant performance degradation when target domain data collected from the Out-Of-Distribution (O.O.D). Previous efforts try to exploit invariant features in the source domain through CNN networks. However, inspired by causal mechanisms, we find that the complex spurious-invariant information is still hidden in this view invariant features, and the impact of domain and class discrepancies on extracting invariance has not been effectively mitigated. To alleviate these issues, we propose a self-weighted multi-view mining invariance domain generalization framework (SMIDG). On the one hand, to make up for the insufficiency of traditional single-view convolutional feature extraction networks, we propose to mine features from another frequency view and use the self-adaptive adversarial masks to eliminate some spurious correlations, ensuring causal invariance in the coarse-grained generalization. However, due to inconsistencies in discriminative information between inter-domain and intra-domain samples, as well as inter-class and intra-class samples, the coarse-grained elimination of spurious associations does not fully resolve this issue. On the other hand, we also consider the fine-grained generalization from two aspects. Firstly, to better tackle the domain discrepancies, we propose a novel progressive contrastive learning strategy that learns the underlying specific features of samples while gradually mitigating domain discrepancies, thereby ensuring domain invariance in fine-grained generalization. Secondly, due to the issue of feature inconsistency, we adopt a self-adaptive hard sample mining method with information gain to ensure that the model pays more attention on hard disentangled samples, thus maintaining feature invariance. Extensive experiments on five benchmark datasets demonstrate that our method outperforms state-of-the-art approaches. Our code is available at https://github.com/bihhm/SMIDG. Shanshan Wang 0008, Houmeng He, Xun Yang 0001, Zhipu Liu, Yuanhong Zhong, Xingyi Zhang 0001, Meng Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Early Traffic Accident Anticipation via Feature Consistency Representation and Soft Label RegressionabstractEarly traffic accident prediction using dashcam videos plays a crucial role in enhancing the safety of intelligent vehicles. Accurately predicting accidents in advance can significantly reduce traffic accidents and improve overall road safety. However, despite extensive research efforts to capture more visual information by employing different feature extraction methods within the same frame, the consistency between features within the same frame and the discrepancy between features across different frames have not been sufficiently emphasized. To address this critical issue, we introduce contrastive learning into the field of accident prediction and propose a novel feature fusion module for the deep integration of diverse features. Our method treats features from the same frame as positive pairs, neighboring frames as sub-positive pairs due to their high correlation, and features from temporally distant frames as negative pairs. This approach effectively strengthens the representation capability of the model, thereby improving overall predictive performance. Additionally, we redefine the accident prediction task by converting it into an anomaly score regression problem using soft labels. This redefinition allows the model to better quantify the likelihood of an accident, offering a more nuanced and accurate prediction. We evaluate our method comprehensively on two publicly available Dashcam Accident Dataset (DAD) and Car Crash Dataset (CCD) datasets to assess its performance. The results demonstrate that our method outperforms state-of-the-art accident prediction approaches, highlighting its potential for practical applications. Code will be available at https://github.com/yangugu/TAP_CC . Yuanhong Zhong, Ruyue Zhu, Ping Gan, Xuerui Shen |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Wavelet-guided network with fine-grained feature extraction for vessel segmentation
Yuanhong Zhong, Daidi Zhong |
Vis. Comput. | 1 |
| 2024 | FaSRnet: a feature and semantics refinement network for human pose estimationabstractDue to factors such as motion blur, video out-of-focus, and occlusion, multi-frame human pose estimation is a challenging task. Exploiting temporal consistency between consecutive frames is an efficient approach for addressing this issue. Currently, most methods explore temporal consistency through refinements of the final heatmaps. The heatmaps contain the semantics information of key points, and can improve the detection quality to a certain extent. However, they are generated by features, and feature-level refinements are rarely considered. In this paper, we propose a human pose estimation framework with refinements at the feature and semantics levels. We align auxiliary features with the features of the current frame to reduce the loss caused by different feature distributions. An attention mechanism is then used to fuse auxiliary features with current features. In terms of semantics, we use the difference information between adjacent heatmaps as auxiliary features to refine the current heatmaps. The method is validated on the large-scale benchmark datasets PoseTrack2017 and PoseTrack2018, and the results demonstrate the effectiveness of our method. Yuanhong Zhong, Qianfeng Xu, Daidi Zhong, Xun Yang 0001, Shanshan Wang 0008 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2024 | Frame-Padded Multiscale Transformer for Monocular 3D Human Pose EstimationabstractMonocular 3D human pose estimation is an ill-posed problem in computer vision due to its depth ambiguity. Most existing works supplement the depth information by extracting temporal pose features from video frames, and they have made notable progress. However, these approaches divide a long sequence of video frames into multiple short sequences for separate processing, which leads to the loss of complementary information between sequences. Furthermore, the short-term temporal correlation among frames in a sequence is often not fully exploited. To model temporal dependencies efficiently, we propose the frame-padded multiscale transformer approach, which includes a frame-padded video sequence preprocessing step and a multiscale temporal transformer backbone. Our approach addresses the omission of the temporal features of edge frames in existing approaches by padding video frames in the shallow layer. In addition, we extract the temporal information of 3D human poses using a multiscale transformer to enhance the short-term correlation of human pose skeleton keypoints. Extensive experiments validate the effectiveness of our approach on two popular datasets: Human3.6M and MPI-INF-3DHP. The results show that our approach achieves state-of-the-art performance. Yuanhong Zhong, Guangxia Yang, Daidi Zhong, Xun Yang 0001, Shanshan Wang 0008 |
IEEE Trans. Multim. | 1 |
| 2024 | Video Compressed Sensing Reconstruction via an Untrained Network with Low-Rank RegularizationabstractDeep image prior (DIP) is an emerging technology that indicates that the structure of an untrained network can serve as an excellent prior for image restoration. It bridges the gap between training-based and training-free methods and exhibits considerable potential in image compressed sensing (CS) reconstruction. In this article, we extend DIP and propose a novel Low-Rank Regularization Video Compressed Sensing Network for CS video reconstruction (dubbed LRR-VCSNet). We explore the application of a low-rank latent tensor with an untrained network for global low-rank regularization on video reconstruction, and the interframe low-rank approximation for framewise nonlocal low-rank regularization in the data space is also exploited. In addition, we design the structure of the untrained network based on the encoder-decoder architecture to improve the performance. Extensive experiments on six standard CIF video sequences show that LLR-VCSNet significantly outperforms traditional video CS methods and achieves competitive results when compared with the state-of-the-art training-based video CS method. Yuanhong Zhong, Chenxu Zhang 0001, Xun Yang 0001, Shanshan Wang 0008 |
IEEE Trans. Multim. | 1 |
| 2023 | Associative Memory With Spatio-Temporal Enhancement for Video Anomaly DetectionabstractMemory network has been extensively used to record prototypical normal patterns to prevent overgeneralization of the network to reconstruct anomalies for video anomaly detection. However, existing memory-based methods only record the lossy representation of normal item prototypes, without recording the rich relationships between them. In this work, we propose an Associative Memory with Spatio-Temporal Enhancement (AMSTE) which introduces the global context information constraint of motion to enhance the appearance features and learn the normal item prototypes and their relationship. Specifically, we utilize two encoders to extract spatio-temporal features with the Spatio-Temporal Enhancement Module (STEM) to enhance appearance features with global motion constraints. Then, the prototypical patterns of normal data and their relationships are recorded in the item memory and relational memory, respectively. Finally, we retrieve features from the memory pools and reconstruct the video frame through the decoder. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our approach. The code will be released athttps://github.com/HuYongting/AMSTE. Yuanhong Zhong, Yongting Hu, Panliang Tang, Heng Wang 0003 |
IEEE Signal Process. Lett. | 1 |
| 2023 | Image Compressed Sensing Reconstruction via Deep Image Prior With Structure-Texture DecompositionabstractDeep image prior has been successfully applied to image compressed sensing, allowing capture implicit prior using only the network architecture without training data. However, existing methods fail to take full advantage of the characteristics of the different components of the image signal, resulting in loss of details, and the network architecture is designed in a homogeneous way, which limits the performance. We propose a novel network architecture to capture distinct implicit priors for different image components under the guidance of a designed loss function. In addition, we design a novel module to extract and fuse the local and global features to facilitate the interaction between the two components and boost the performance. Sufficient experiments demonstrate the competitive performance and effectiveness of our method. Yuanhong Zhong, Chenxu Zhang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2022 | MBMR-Net: multi-branches multi-resolution cross-projection network for single image super-resolution
Binglian Zhu, Yuanhong Zhong |
Appl. Intell. | 3 |
| 2022 | Reverse erasure guided spatio-temporal autoencoder with compact feature representation for video anomaly detection
Yuanhong Zhong, Jinyang Jiang 0002 |
Sci. China Inf. Sci. | 1 |
| 2022 | PRPN: Progressive region prediction network for natural scene text detection
Yuanhong Zhong, Tao Chen 0004, Jing Zhang 0037, Zhaokun Zhou |
Knowl. Based Syst. | 1 |
| 2022 | Joining features by global guidance with bi-relevance trihard loss for person re-identification
Wen-cheng Qin, Zhiyong Huang 0004, Lamia Tahsin, Daming Sun, Yuanhong Zhong |
Neural Comput. Appl. | 6 |
| 2022 | A cascade reconstruction model with generalization ability evaluation for anomaly detection in videos
Yuanhong Zhong, Jinyang Jiang 0002 |
Pattern Recognit. | 1 |
| 2022 | Bidirectional Spatio-Temporal Feature Learning With Multiscale Evaluation for Video Anomaly DetectionabstractVideo anomaly detection aims to detect the segments containing abnormal events from video sequence, which is a current research hotspot due to the importance in maintaining social security. Recent detection methods tend to build frame reconstruction or frame prediction model based on deep learning to learn features of events. The reconstruction-based methods reproduce the input frame one-to-one, inevitably losing some temporal features. The prediction-based methods predict frames according to the natural time order, but ignore the reverse time information, causing the deviation in information learning. Besides, anomaly evaluation methods based on patch-level error neglect the diversity of object sizes in complex scenes, and it is difficult to determine the optimal size of the error patch accurately. For these issues, we propose a bidirectional spatio-temporal feature learning framework with multi-scale anomaly evaluation strategy. A video sequence is input to a double-encoder double-decoder network, and bidirectional spatio-temporal features are obtained for bidirectional prediction by fusing forward and backward features extracted from the two encoders. The multi-scale anomaly evaluation method is implemented based on error pyramid and mean pooling, which effectively detects target objects with different sizes. Experiments on several publicly video datasets show that our method outperforms most of existing methods. Yuanhong Zhong, Yongting Hu, Panliang Tang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Recovery of image and video based on compressive sensing via tensor approximation and Spatio-temporal correlation
Yuanhong Zhong, Jing Zhang 0037, Zhaokun Zhou |
Multim. Tools Appl. | 1 |
| 2020 | Specific category region proposal network for text detection in natural sceneabstractNatural scene text usually carries considerable abstract semantic information, which is closely related to the surrounding environment. Thus, natural scene text detection plays a vital role in image content retrieval and understanding. In this study, the authors propose a novel specific category region proposal network (SCRPN) based on maximally stable extremal regions (MSER) and fully convolutional network (FCN) for natural scene text detection. First, FCN for pixel‐level recognition is utilised to obtain the text saliency map and MSER is used to obtain oversegmented regions. Then, the multiple features of oversegmented regions and text saliency map are used for region aggregation. Next, single‐linkage clustering method is adopted to cluster the segmentation regions to obtain a hierarchical structure of text region proposals. Finally, for the top‐ranking region proposals, SCRPN built an end‐to‐end pipeline for scene text detection directly. Experiments on street view text and international conference on document analysis and recognition (ICDAR) 2013 have demonstrated the effectiveness of SCRPN for generating the text proposals. SCRPN could work with various two‐stage text detection networks; thus, faster region convolutional neural network was used as the text detection framework to evaluate the performance of SCRPN in the ICDAR 2015 and MSRA‐TD500 benchmarks. The experimental results confirmed that SCRPN makes text detection more robust in complex scenarios. Yuanhong Zhong, Zhaokun Zhou, Jing Zhang 0037 |
IET Image Process. | 1 |
| 2010 | Single Relay Selection With Feedback and Power Allocation in Multisource Multidestination Cooperative NetworksabstractSingle relay selection has been shown to be an attractive strategy for cooperative communications. In this work, we extend the distributed selection cooperation protocol with feedback to multisource multidestination cooperative networks and investigate the best relay confliction problem existing in the scenario. In the high signal-to-noise ratio (SNR) regime, due to the fact that sharing the same best relay by multiple source nodes has no influence on the diversity-multiplexing tradeoff (DMT) performance for each communication pair, the single relay sharing method is a simple and effective solution. For some practical systems with low or medium SNR, we propose two power allocation algorithms for the shared best relay and show that the method based on maximizing the number of successful relay flows is efficient in distributed scenarios with low complexity. Heng Wang 0003, Shizhong Yang, Jinzhao Lin, Yuanhong Zhong |
IEEE Signal Process. Lett. | 4 |