EDBT 2026 Demo / reviewers in the wild / expert
Pengyi Hao
dblp:08/7950
· DBLP profile ↗
40ranked-venue papers
17as first author
24since 2021 · last 2026
0000-0002-1515-3937ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 12 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A High-Order Semantic Dual-Stream Synergistic Network for 3D Facial Landmark Detection
Fuli Wu, Chaoran Hu, Tao Qiu, Pengyi Hao, Xiangtao Liu, Mengfei Yu |
ICIC (18) | 4 |
| 2026 | Graph-Structured Sparse Transformer with Semantic Priors for Multimodal Video Summarization
Qingyuan Zhu, Zhengqi Zhao, Pengyi Hao |
ICIC (19) | 3 |
| 2026 | Feature-based decoupled distillation
Shuang Wang 0016, Qingyuan Zhu, Fuli Wu, Pengyi Hao |
Knowl. Based Syst. | 5 |
| 2026 | GraphVSum:graph guided multimodal video summarization
Zhengqi Zhao, Cong Bai, Pengyi Hao |
Multim. Syst. | 3 |
| 2025 | Deformable Blur Sensing and Regression Analysis ReID Feature Fusion for Multitarget Multicamera Tracking Systems in Highway ScenariosabstractIn highway scenarios, the rapid motion of vehicles can cause deformation and blur in camera footage, significantly affecting the accuracy of vehicle detection and re-identification (ReID) in multitarget multicamera tracking (MTMCT) systems. To address this issue, this article develops the deformable and blur sensing and regression analysis ReID feature fusion MTMCT system (DSRF). First, a deformable and blur sensing detection module (DFB) in DSRF is designed to overcome the limitations of cameras in capturing fast-moving objects, thereby accurately detecting vehicles moving at high speeds on highways. Then, a regression-based ReID feature fusion algorithm (RARF) in DSRF is proposed, which enhances ReID features by modeling the relationship between vehicle motion and its features, thereby better associating the detected vehicles in consecutive frames into trajectories and establishing intertrajectory relationships. Finally, extensive experiments are conducted on the highway surveillance traffic (HST) dataset developed by our team and the public dataset (CityFlow). Promising results are achieved, validating the effectiveness of our proposed method. Sixian Chan 0001, Shenghao Ni, Jie Hu 0041, Tinglong Tang, Xiaolong Zhou 0001, Pengyi Hao |
IEEE Trans. Comput. Soc. Syst. | 7 |
| 2024 | Live on the Hump: Self Knowledge Distillation via Virtual Teacher-Students Mutual LearningabstractFor solving the limitations of the current self knowledge distillation including never fully utilizing the knowledge of shallow exits and neglecting the impact of auxiliary exits' structure on the performance of network, a novel self knowledge distillation framework via virtual teacher-students mutual learning named LOTH is proposed in this paper. A knowledgeable virtual teacher is constructed from the rich feature maps of each exit to help the learning of each exit. Meanwhile, the logit knowledges of each exit are incorporated to guide the learning of the virtual teacher. They learn mutually through the well-designed loss in LOTH. Moreover, two kinds of auxiliary building blocks are designed to balance the efficiency and effectiveness of network. Extensive experiments with diverse backbones on CIFAR-100 and Tiny-ImageNet validate the effectiveness of LOTH, which realizes superior performance with less resource by the comparison with the state-of-the-art distillation methods. The code of LOTH is available on Github https://github.com/cloak-s/LOTH. Shuang Wang 0016, Pengyi Hao, Fuli Wu, Cong Bai |
ACM Multimedia | 2 |
| 2024 | Intent-Aware Graph-Level Embedding Learning Based Recommendation
Pengyi Hao, Si-Hao Liu, Cong Bai |
J. Comput. Sci. Technol. | 1 |
| 2023 | Multi-scale Hybrid Transformer Network with Grouped Convolutional Embedding for Automatic Cephalometric Landmark Detection
Fuli Wu, Lijie Chen 0009, Pengyi Hao |
CAD/Graphics | 4 |
| 2023 | A Structure-Fusion Network for Medical Image ClassificationabstractThe Convolutional Neural Networks and Transformer cannot provide satisfactory performance in medical image classification due to insufficient data, high resolution, and a lot of redundancy. To achieve better performance, this paper proposes a structure-fusion network that combines the architecture of convolution and transformer. To reduce the computational overhead incurred by the transformer structure, we optimize it by aggregating adjacent features. We further modify the Multilayer Perceptron using convolution to increase the network capacity. The network is verified on the grading of the Anterior Cruciate Ligament from knee MRI and the screening of pneumonia and COVID-19 from Chest X-ray. Compared with current advanced methods, the proposed network not only reduces FLOPs but also achieves improvement on AUC and F1score. Fuli Wu, Pengyi Hao, Shu-yuan Tian |
ICIP | 3 |
| 2023 | MOOCs Dropout Prediction via Classmates Augmented Time-Flow Hybrid Network
Guanbao Liang, Zhaojie Qian, Shuang Wang 0016, Pengyi Hao |
ICONIP (15) | 4 |
| 2023 | Uncertainty-aware iterative learning for noisy-labeled medical image segmentationabstractAbstract Medical image segmentation from noisy labels is an important task since obtaining high‐quality annotations is extremely difficult and expensive. There are a lot of approaches proposed for such task. However, some issues like the overfitting on noisy annotations, the limited learning of boundary features, and no consideration of the corrupted local pixels are still not solved. Therefore, a novel approach named uncertainty‐aware iterative learning (UaIL) is proposed for medical image segmentation with noisy labels. UaIL iteratively and jointly trains two deep networks using the original images and their argumented ones through a joint loss function including softened label loss, hard label loss and consistency loss, which encourages UaIL to produce segmentations that are robust to the perturbations in arbitrary semantic space. The uncertainty of labels is estimated based on the predictions in iterative learning, then the original labels are refined, which improves the learning of boundary features in segmentation. To avoid overfitting, a stopping strategy is designed based on the dice coefficient in iterative learning. Experiments on two public datasets verify the effectiveness of UaIL under different levels of annotation noise. Especially, when there are serious noises in the labels, the dice achieved by UaIL is 1.43% to 15.03% higher than the competing approaches on the two public datasets. The UaIL is further verified on a private dataset, which shows its ability of applying in the real application with noisy labels. Pengyi Hao, Kangjian Shi, Shu-yuan Tian, Fuli Wu |
IET Image Process. | 1 |
| 2023 | Meta-relationship for course recommendation in MOOCs
Pengyi Hao, Cong Bai |
Multim. Syst. | 1 |
| 2023 | Community aware graph embedding learning for item recommendation
Pengyi Hao, Zhaojie Qian, Shuang Wang 0016, Cong Bai |
World Wide Web (WWW) | 1 |
| 2022 | Learning Tucker Compression for Deep CNNabstractRecently, tensor decomposition approaches are used to compress deep convolutional neural networks (CNN) for getting a faster CNN with fewer parameters. However, there are two problems of tensor decomposition based CNN compression approaches, one is that they usually decompose CNN layer by layer, ignoring the correlation between layers, the other is that training and compressing a CNN is separated, easily leading to local optimum of ranks. In this paper, Learning Tucker Compression (LTC) is proposed. It gets the best tucker ranks by jointly optimizing of CNN's loss function and Tucker's cost function, which means that training and compressing is carried out at the same time. It can directly optimize the CNN without decomposing the whole network layer by layer and can directly fine-tune the whole network without using fixed parameters. LTC is verified on two public datasets. Experiments show that LTC can make a network like ResNet, VGG faster with nearly the same classification accuracy, which surpasses current tensor decomposition approaches. Pengyi Hao, Fuli Wu |
DCC | 1 |
| 2022 | Deep Guided Context-aware Network for Anomaly Detection in Musculoskeletal RadiographsabstractAutomated anomaly detection and localization in musculoskeletal radiographs are essential for large-scale screening in the radiograph workflow. However, the anomaly is a localized pattern that may be affected by the extra irrelevant areas, and the ambiguous and subtle features of abnormal regions are difficult to detect. To tackle these two problems, we propose a deep guided context-aware network (DR-Net) for anomaly detection in musculoskeletal X-rays. Specifically, we design a positional guide module, which embeds the spatial positional information as prior knowledge to guide the network to enhance the feature representation of a specific region. Then, to detect subtle anomalies, we construct a contextual relation module. It can obtain context-aware features by capturing the spatial dependence of any two positions from the entire X-ray image. It combines context appearance information and selects more distinguishable features from space and channel, producing a detailed visualization of the anomaly region. Note that only image-level labels are required. The extensive experiments on the two radiograph datasets show that DR-Net has a promising performance in anomaly detection and localization. Kangjian Shi, Fuli Wu, Pengyi Hao |
ICPR | 4 |
| 2022 | Structural and Temporal Learning for Dropout Prediction in MOOCs
Tianxing Han, Pengyi Hao, Cong Bai |
KSEM (2) | 2 |
| 2022 | Deep Correlation based Concept Recommendation for MOOCsabstractThe current course recommendation in massive open online courses (MOOCs) usually ignores students' interests in some certain type of knowledge concepts, resulting in low completion of most courses.Therefore, it requires a concept recommendation to help students accurately choose courses in MOOCs.In this paper, we propose Deep Correlation based Concept Recommendation (DCCR) for MOOCs.It gathers the interactive information obtained by different entities through meta-paths in MOOCs and extracts the semantic information of concepts.To deeply capture the correlation information among users, a multi-relation graph is built to generate the correlation features which aggregates the abundant information under different meta-paths.Then through the graph convolutional neural networks, entity embeddings of users and knowledge concepts are generated.Additionally, a concatenation-based fusion function is designed to get the final joint representations reasonably.By verifying on two public datasets, experiments show that DCCR outperforms the state-of-the-art methods. Shengyu Mao, Pengyi Hao, Cong Bai |
SEKE | 2 |
| 2022 | DINs: Deep Interactive Networks for Neurofibroma Segmentation in Neurofibromatosis Type 1 on Whole-Body MRIabstractNeurofibromatosis type 1 (NF1) is an autosomal dominant tumor predisposition syndrome that involves the central and peripheral nervous systems. Accurate detection and segmentation of neurofibromas are essential for assessing tumor burden and longitudinal tumor size changes. Automatic convolutional neural networks (CNNs) are sensitive and vulnerable as tumors' variable anatomical location and heterogeneous appearance on MRI. In this study, wepropose deep interactive networks (DINs) to address the above limitations. User interactions guide the model to recognize complicated tumors and quickly adapt to heterogeneous tumors. We introduce a simple but effective Exponential Distance Transform (ExpDT) that converts user interactions into guide maps regarded as the spatial and appearance prior. Comparing with popular Euclidean and geodesic distances, ExpDT is more robust to various image sizes, which reserves the distribution of interactive inputs. Furthermore, to enhance the tumor-related features, we design a deep interactive module to propagate the guides into deeper layers. We train and evaluate DINs on three MRI data sets from NF1 patients. The experiment results yield significant improvements of 44% and 14% in DSC comparing with automated and other interactive methods, respectively. We also experimentally demonstrate the efficiency of DINs in reducing user burden when comparing with conventional interactive methods. Jianwei Zhang 0015, Wei Chen 0001, K. Ina Ly, Xubin Zhang, Fan Yan, Justin Jordan, Gordon J. Harris, Scott Plotkin, Pengyi Hao, Wenli Cai |
IEEE J. Biomed. Health Informatics | 9 |
| 2021 | A Novel Feature Fusion Network for Myocardial Infarction Screening Based on ECG Images
Pengyi Hao, Fuli Wu, Fan Zhang 0056 |
ICIG (2) | 1 |
| 2021 | Classmates Enhanced Diversity-Self-Attention Network for Dropout Prediction in MOOCs
Dongen Wu, Pengyi Hao, Tianxing Han, Cong Bai |
ICONIP (4) | 2 |
| 2021 | Community Enhanced Course Concept Recommendation in MOOCs with Multiple Entities
Binglong Ye, Shengyu Mao, Pengyi Hao, Wei Chen 0001, Cong Bai |
KSEM | 3 |
| 2021 | Radiographs and texts fusion learning based deep networks for skeletal bone age assessment
Pengyi Hao, Taotao Ye, Xuhang Xie, Fuli Wu, Wuheng Zuo, Wei Chen 0001, Jian Wu 0001 |
Multim. Tools Appl. | 1 |
| 2021 | Multi-modality learning for human action recognition
Ziliang Ren, Qieshi Zhang, Xiangyang Gao, Pengyi Hao, Jun Cheng 0002 |
Multim. Tools Appl. | 4 |
| 2021 | Self-supervised deep subspace clustering network for faces in videos
Yunhao Qiu, Pengyi Hao |
Vis. Comput. | 2 |
| 2020 | Overlap classification mechanism for skeletal bone age assessmentabstractThe bone development is a continuous process, however, discrete labels are usually used to represent bone ages. This inevitably causes a semantic gap between actual situation and label representation scope. In this paper, we present a novel method named as overlap classification network to narrow the semantic gap in bone age assessment. In the proposed network, discrete bone age labels (such as 0-228 month) are considered as a sequence that is used to generate a series of subsequences. Then the proposed network makes use of the overlapping information between adjacent subsequences and output several bone age ranges at the same time for one case. The overlapping part of these age ranges is considered as the final predicted bone age. The proposed method without any preprocessing can achieve a much smaller mean absolute error compared with state-of-the-art methods on a public dataset. Pengyi Hao, Xuhang Xie, Tianxing Han, Cong Bai |
MMAsia | 1 |
| 2020 | Lung adenocarcinoma diagnosis in one stage
Pengyi Hao, Kun You, Haozhe Feng, Xinnan Xu, Fan Zhang 0056, Fuli Wu, Peng Zhang 0043, Wei Chen 0001 |
Neurocomputing | 1 |
| 2020 | Cross-domain representation learning by domain-migration generative adversarial network for sketch based image retrieval
Cong Bai, Jian Chen 0009, Pengyi Hao, Shengyong Chen |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | Texture branch network for chronic kidney disease screening based on ultrasound imagesabstractChronic kidney disease (CKD) is a widespread renal disease throughout the world. Once it develops to the advanced stage, serious complications and high risk of death will follow. Hence, early screening is crucial for the treatment of CKD. Since ultrasonography has no side effects and enables radiologists to dynamically observe the morphology and pathological features of the kidney, it is commonly used for kidney examination. In this study, we propose a novel convolutional neural network (CNN) framework named the texture branch network to screen CKD based on ultrasound images. This introduces a texture branch into a typical CNN to extract and optimize texture features. The model can automatically generate texture features and deep features from input images, and use the fused information as the basis of classification. Furthermore, we train the base part of the network by means of transfer learning, and conduct experiments on a dataset with 226 ultrasound images. Experimental results demonstrate the effectiveness of the proposed approach, achieving an accuracy of 96.01% and a sensitivity of 99.44%. Pengyi Hao, Shu-yuan Tian, Fuli Wu, Wei Chen 0001, Jian Wu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2019 | End-to-End Panoptic Segmentation with Pixel-Level Non-Overlapping EmbeddingabstractRecent panoptic segmentation even instance segmentation methods usually rely on the region-based method or highly-specialized combination with heuristics module, followed by post-processing techniques. While most of the recent methods neglect low-fill rate linear objects and cannot recognize pixels located in bounding box margins. We propose a branched, end-to-end trainable multi-task architecture focusing on pixel-level grouping problems for panoptic segmentation. The embedding branch regress pixels into an embedding space, so that pixels from the same group are at close range while those from different groups have a specified margin. Every pixel can be considered in an image without overlapping. And semantic branch produces best seed scores with labels as clustering center. The further-embedding branch disentangles each pixel in pixel embedding space. Thus, we are able to segment both thing and stuff classes, and explain all the pixels in the image. We obtain state-of-the-art results on Pascal VOC2012 and Cityscapes. Qieshi Zhang, Jun Cheng 0002, Cong Bai, Pengyi Hao |
ICME | 5 |
| 2019 | Video Summarization based on Sparse Subspace Clustering with Automatically Estimated Number of ClustersabstractAdvancements in technology resulted in a sharp growth in the number of digital cameras at people's disposal all across the world. Consequently, the huge storage space consumed by the videos from these devices on video repositories make the job of video processing and analysis to be time-consuming. Furthermore, this also slows down the video browsing and retrieval. Video summarization plays a very crucial role in solving these issues. Despite the number of video summarization approaches proposed up to the present time, the goal is to take a long video and generate a video summary in form of a short video skim without losing the meaning or the message transmitted by the original lengthy video. This is done by selecting the important frames called key-frames. The approach proposed by this work performs automatic summarization of digital videos based on detected objects' deep features. To this end, we apply sparse subspace clustering with an automatically estimated number of clusters to the objects' deep features. The summary generated from our scheme will store the meta-data for each short video inferred from the clustering results. In this paper, we also suggest a new video dataset for video summarization. We evaluate the performance of our work using the TVSum dataset and our video summarization dataset. Pengyi Hao, Edwin Manhando, Taotao Ye, Cong Bai |
MMAsia | 1 |
| 2018 | Co-consistent Regularization with Discriminative Feature for Zero-Shot Learning
Yanling Tian, Qieshi Zhang, Jun Cheng 0002, Pengyi Hao |
ICONIP (1) | 5 |
| 2017 | Coexistence and Local Exponential Stability of Multiple Equilibria in Memristive Neural Networks with a Class of General Nonmonotonic Activation Functions
Jie Xiao 0003, Pengyi Hao |
ISNN (1) | 4 |
| 2017 | Adaptive Weberfaces for occlusion-robust face representation and recognitionabstractIn order to deal with facial occlusion effectively, the authors propose a powerful but simple face representation method, called adaptive Weberfaces (AdapWeber), based on human visual perception change model and the Weber ratio R implied in Weber's law. Specifically, human perception is naturally highly selective and robust to occlusions, and the Weber ratio R is very important to enhance feature redundancy. As feature redundancy and locality are two guiding principles against facial occlusion, they further develop eight variants of AdapWeber, collectively referred to as single‐scale and single‐orientation (SSSO) AdapWeber, by shrinking the kernel locality and varying the kernel orientation of the original AdapWeber, and integrate them to formulate a multi‐scale and multi‐orientation (MSMO) AdapWeber. A natural by‐product of MSMO AdapWeber is MSMO Weberfaces. Experiments on four benchmark databases, including Extended Yale B, AR, UMB‐DB, and LFW, showed that MSMO AdapWeber/Weberfaces, rather than any variant of SSSO AdapWeber/Weberfaces, outperformed several popular feature extraction approaches in many scenarios, especially when the occlusion level is very high or the image dimension is very low. This result demonstrates that several occlusion‐weak features can be combined together to construct an occlusion‐robust feature. Lin He 0001, Pengyi Hao |
IET Image Process. | 3 |
| 2017 | Discriminative Histogram Intersection Metric Learning and Its Applications
Pengyi Hao, Shengyong Chen |
J. Comput. Sci. Technol. | 1 |
| 2013 | An efficient video retrieval scheme based on facial signaturesabstractThe topic of retrieving videos containing a desired person by just using facial content has many applications like video surveillance, social network, etc. In this paper, we propose a compact, discriminative and low-dimensional signature to describe an person with a set of high-dimensional features. The signature is generated by linear discriminant analysis with maximum correntropy criterion that is robust to outliers and noises. Based on the proposed signatures, a new video retrieval scheme is given for fast finding the desired videos by measuring the similarities between the signature of a query and the ones in the dataset. Evaluations on a large dataset of videos show that the proposed video retrieval scheme has the potential to substantially reduce the response time and slightly increase the mean average precision of retrieval. Pengyi Hao |
ICIP | 1 |
| 2013 | Maximum correntropy criterion for discriminative dictionary learningabstractIn this paper, a novel discriminative dictionary learning with pairwise constraints by maximum correntropy criterion is proposed for pair matching problem. Comparing with the conventional dictionary learning approaches, the proposed method has several advantages: (i) It can deal with the outliers and noises problem more efficiently during the reconstruction step. (ii) It can be effectively solved by half-quadratic optimization algorithm, and in each iteration step, the complex optimization problem can be reduced to a general problem that can be efficiently solved by feature-sign search optimization. (iii) The proposed method is capable of analyzing non-Gaussian noise to reduce the influence of large outliers substantially, resulting in a robust and discriminative dictionary. We test the performance of the proposed method on two applications: face verification on the challenging restricted protocol of Labeled Faces in the Wild (LFW) benchmark and face-track identification on a dataset with more than 7,000 face-tracks. Compared with the recent state-of-the-art approaches, the outstanding performance of the proposed method validates its robustness and discriminability. Pengyi Hao |
ICIP | 1 |
| 2013 | Facial signatures for fast individual retrieval from video datasetabstractThe topic of retrieving videos containing a desired person by using the content of faces without any help of textual information has many interesting applications like video surveillance, social network, video mining, etc. However, face-by-face matching leads to an unacceptable response time for a video dataset with a large number of detected faces and may also reduce the accuracy of searching. Therefore, in this paper we propose a scheme to generate facial signatures for fast retrieving videos containing the same person with a query. First, we summarize each video as a set of person-oriented individuals based on detected faces, which are represented as high dimensional vectors in a feature space. Then, each person with a collection of high dimensional vectors is projected to a compact and reduced dimensionality representation that is called facial signature for this person. The projection is realized by constructing a matcher using linear discriminant analysis with maximum correntropy criterion optimization. In this research, two kinds of signatures are provided, which are called 1D facial signature and 2D facial signature. The proposed searching scheme can support two types of queries: face image and video clip. Evaluations on a large dataset of videos show reliable measurement of similarities by using facial signature to represent each person generated from videos and also demonstrate that the proposed searching scheme has the potential to substantially reduce the response time and slightly increase the mean average precision of retrieval. Pengyi Hao |
ICME | 1 |
| 2012 | Unsupervised people organization and its application on individual retrieval from videos
Pengyi Hao |
ICPR | 1 |
| 2009 | A KFCM and SIFT Based Matching Approach to Similarity Retrieval of ImagesabstractRecently, keypoint descriptors such as Scale Invariant Feature Transform (SIFT) have been proved promising in similarity retrieval of images, which adopts matching score as similarity. However, the matching score is easy to be decreased once there are little variances between image details, and hence lead to low retrieval performance. In this paper, we propose a novel retrieval approach that improves the matching score with reduced time of matching by Kernel-based Fuzzy C-Means clustering (KFCM), which proves to be a better trade-off between matching and retrieval precision. Experiments conducted on three representative image databases show that our retrieval approach is surprisingly effective, outperforming the SIFT based method, not only in object-based image retrieval but also for searching scenes with similar semantic. Pengyi Hao, Youdong Ding, Yuchun Fang, Shuhan Wei |
ICIG | 1 |
| 2009 | Blotch Detection Based on Texture Matching and Adaptive Multi-thresholdabstractBlotch is a typical artifact in old films and the detection of them is an important step in film restoration. The existing simplified rank-ordered difference detector achieves higher detection rates by reducing the value of threshold. However, the corresponding higher number of false alarms is undesirable. To maximize the ratio between correct detections and false alarms, this paper proposes an improved blotch detector based on adaptive multi- threshold. According to different objects of blotches, the proposed detector can achieve the most appropriate threshold by convergence confinement. Meanwhile, texture matching is introduced to avoid the possible deviation caused by motion vector estimation in the regions with blotches. Performance evaluation is taken to the image sequences with both real blotches and artificially corrupted ones. The experimental results indicate higher correct detection rates and fewer false alarms simultaneously. Shuhan Wei, Pengyi Hao, Youdong Ding |
ICIG | 3 |