Shengye Yan

dblp:85/3795 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
5since 2021 · last 2024
0000-0003-3583-0146ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 The Attempt on Combining Three Talents by KD with Enhanced Boundary in Co-Salient Object Detection
Ziyi Cao, Shengye Yan
BMVC2
2024 Lightweight Human Pose Estimation with Enhanced Knowledge Review
Shengye Yan
BMVC2
2024 Internal Location Assistance for Temporal Action Proposal Generation
abstract
Temporal action proposal generation (TAPG) aims to locate action instances in untrimmed videos for video analysis tasks. In this paper, we propose a novel approach called Internal Location Assistance Net (ILAN) to take advantage of the internal action point instead of only utilizing the start and end points themselves. Specifically, instead of solely predicting the action’s start and end positions, we also predict extra internal points which could be the one-eighth, one-fourth or center points etc. Then the pair of left and right internal positions of the action are matched to generate center region proposals. Then the predicted center region is constrained to lie in the predicted overall action region. And both confidences of the overall action region and the center region are combined to get the final action proposal. Besides, we incorporate the window transformer to enhance the feature extraction for capturing more precise action boundaries. Extensive experiments are conducted on two popular benchmark datasets: THUMOS14 and ActivityNet-v1.3. The experimental results demonstrate that the proposed method outperforms the state-of-the-arts.
Songsong Feng, Shengye Yan
ICASSP2
2024 Bounding Box-Guided Pseudo Point Clouds Early-Fusion and Density Optimize for 3D Object Detection
abstract
3D point clouds are used for 3D object detection. The sparsity and noises of point clouds poses a significant challenge primarily due to difficulties in capturing LiDAR data and LiDAR sensor performance. Typically, the number of captured point cloud data is lower than the real point density of real scenes. To address this issue, we adopted point clouds early-fusion, which integrates multi-modal data into a comprehensive 3D point clouds. We propose a straightforward and efficient pseudo-point cloud generation method that can be easily integrated with existing point cloud data. This involves generating pseudo point clouds and then fusing them with the original point cloud data. We also propose vertical density adaptation method to address the problem of density inhomogeneity in point cloud data following early-fusion. We conduct experiments on the well-known KITTI dataset using three state-of-the-art LiDAR 3D object detectors. Our results show that our fusion strategy can enhance the performance of existing detectors, which could potentially inspire future detector designs.
Shengye Yan
ICASSP2
2024 Self Knowledge Distillation Based On Layer-Wise Weighted Feature Imitation For Efficient Object Detection
abstract
Knowledge Distillation (KD)[1] is a widely-used technology to inherit information from cumbersome teacher models to compact student models, consequently realizing model compression and acceleration. Compared with image classification, object detection is a more complex task, and designing specific KD methods for object detection is non-trivial. In this paper, we propose Layer-wise Weighted Feature Imitation (LWFI), the model loaded with pre-trained parameters acts as a teacher to guide students who have not been pretrained. We use the feature maps of multiple intermediate positions in the teacher network to guide the corresponding positions in the student network, and allocate corresponding weights according to the magnitude of distillation loss. We have carried out a series of experiments on the VOC and KITTI datasets. Specifically, YOLOv6N on VOC improved from 66.0% to 67.3%, YOLOv6S improved from 70.8% to 71.9%, and on KITTI, YOLOv6N increased from 56.4% to 58.1%, outperforming similar-sized networks.
Liangqi Zhong, Shengye Yan
ICASSP2
2017 Sparse multiple instance learning as document classification
Shengye Yan, Jianxin Wu 0001
Multim. Tools Appl.1
2015 Inferring occluded features for fast object detection
Shengye Yan, Qingshan Liu 0001
Signal Process.1
2015 Image Classification With Densely Sampled Image Windows and Generalized Adaptive Multiple Kernel Learning
abstract
We present a framework for image classification that extends beyond the window sampling of fixed spatial pyramids and is supported by a new learning algorithm. Based on the observation that fixed spatial pyramids sample a rather limited subset of the possible image windows, we propose a method that accounts for a comprehensive set of windows densely sampled over location, size, and aspect ratio. A concise high-level image feature is derived to effectively deal with this large set of windows, and this higher level of abstraction offers both efficient handling of the dense samples and reduced sensitivity to misalignment. In addition to dense window sampling, we introduce generalized adaptive l(p)-norm multiple kernel learning (GA-MKL) to learn a robust classifier based on multiple base kernels constructed from the new image features and multiple sets of prelearned classifiers from other classes. With GA-MKL, multiple levels of image features are effectively fused, and information is shared among different classifiers. Extensive evaluation on benchmark datasets for object recognition (Caltech256 and Caltech101) and scene recognition (15Scenes) demonstrate that the proposed method outperforms the state-of-the-art under a broad range of settings.
Shengye Yan, Xinxing Xu, Dong Xu 0001, Stephen Lin 0001, Xuelong Li 0001
IEEE Trans. Cybern.1
2014 Learning the object location, scale and view for image categorization with adapted classifier
Shengye Yan, Xinxing Xu, Qingshan Liu 0001
Inf. Sci.1
2013 Region-Based Spatial Sampling for Image Classification
abstract
Local descriptors with Bag-of-Words representation were widely used for image classification. Especially, local descriptors of dense spatial sampling were demonstrated to be able to further improve performances of image classification. However, denser spatial sampling is impractical due to huge computation cost. To handle this issue, we propose a new region-based sampling strategy in this paper. We first perform an over-segmentation to get image regions, and then we extract local descriptors around the region boundaries and inside the regions respectively with a popular sampling strategy. Thus, almost at the same computation cost, the proposed method can capture more salient points and obtain much better classification performance. Extensive experiments are conducted on two widely-used datasets (UIUC Sports and Caltech-101). The experimental results demonstrate the effectiveness of the proposed method. Specifically, the proposed method updates the state-of-the-art on UIUC Sports dataset with a classification accuracy of 89.38%.
Pengpeng Ji, Shengye Yan
ICIG2
2012 Beyond Spatial Pyramids: A New Feature Extraction Framework with Dense Spatial Sampling for Image Classification
Shengye Yan, Xinxing Xu, Dong Xu 0001, Stephen Lin 0001, Xuelong Li 0001
ECCV (4)1
2009 Fea-Accu cascade for face detection
abstract
Aiming at unloading the high training time burden of the popular cascaded classifier, in this paper, a novel cascade structure called Fea-Accu cascade is proposed. In Fea-Accu cascade training, the times of feature selection are largely reduced by enhancing the correlation among different stage classifiers of the cascaded classifier. In detail, for each stage classifier, before selecting new features out, the features selected out by previous stage classifiers are reused through creating new corresponding weak classifiers. To verify the efficiency and effectiveness of the proposed method, experiment is designed on frontal face detection problem. The experimental results show that it can largely reduce the training time. A frontal face detector with state-of-the-art classification performance can be learned in less than 10 hours.
Shengye Yan, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001
ICIP1
2008 Locally Assembled Binary (LAB) feature with feature-centric cascade for fast and accurate face detection
abstract
In this paper, we describe a novel type of feature for fast and accurate face detection. The feature is called Locally Assembled Binary (LAB) Haar feature. LAB feature is basically inspired by the success of Haar feature and Local Binary Pattern (LBP) for face detection, but it is far beyond a simple combination. In our method, Haar features are modified to keep only the ordinal relationship (named by binary Haar feature) rather than the difference between the accumulated intensities. Several neighboring binary Haar features are then assembled to capture their co-occurrence with similar idea to LBP. We show that the feature is more efficient than Haar feature and LBP both in discriminating power and computational cost. Furthermore, a novel efficient detection method called feature-centric cascade is proposed to build an efficient detector, which is developed from the feature-centric method. Experimental results on the CMU+MIT frontal face test set and CMU profile test set show that the proposed method can achieve very good results and amazing detection speed.
Shengye Yan, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001
CVPR1
2007 Matrix-Structural Learning (MSL) of Cascaded Classifier from Enormous Training Set
abstract
Aiming at the problem when both positive and negative training set are enormous, this paper proposes a novel matrix-structural learning (MSL) method, as an extension to Viola and Jones' cascade learning method for object detection. Briefly speaking, unlike Viola and Jones' method that learn linearly by bootstrapping only negative samples, the proposed MSL method bootstraps both positive and negative samples in a matrix-like structure. Moreover, an accumulative way is further presented to improve the training efficiency of MSL by inheriting features learned previously during training procedure. The proposed method is evaluated on face detection problem. On a positive set containing 230000 face samples, only 12 hours are needed on a common PC with a 3.20 GHz Pentium IV processor to learn a classifier with false alarm rate less than 1/1000000. What's more, the accuracy of the learned detector exceeds the state-of-the-art results on the CMU+MIT frontal face test set.
Shengye Yan, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001, Jie Chen 0001
CVPR1
2007 Enhancing Human Face Detection by Resampling Examples Through Manifolds
abstract
As a large-scale database of hundreds of thousands of face images collected from the Internet and digital cameras becomes available, how to utilize it to train a well-performed face detector is a quite challenging problem. In this paper, we propose a method to resample a representative training set from a collected large-scale database to train a robust human face detector. First, in a high-dimensional space, we estimate geodesic distances between pairs of face samples/examples inside the collected face set by isometric feature mapping (Isomap) and then subsample the face set. After that, we embed the face set to a low-dimensional manifold space and obtain the low-dimensional embedding. Subsequently, in the embedding, we interweave the face set based on the weights computed by locally linear embedding (LLE). Furthermore, we resample nonfaces by Isomap and LLE likewise. Using the resulting face and nonface samples, we train an AdaBoost-based face detector and run it on a large database to collect false alarms. We then use the false detections to train a one-class support vector machine (SVM). Combining the AdaBoost and one-class SVM-based face detector, we obtain a stronger detector. The experimental results on the MIT + CMU frontal face test set demonstrated that the proposed method significantly outperforms the other state-of-the-art methods.
Jie Chen 0001, Ruiping Wang 0001, Shengye Yan, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001
IEEE Trans. Syst. Man Cybern. Part A3
2003 Rotated face detection in color images using radial template (RT)
abstract
We propose a face detection algorithm to locate faces rotated in any orientation. Detecting rotated faces is important for a face detection system. First, we present a novel model named radial template (RT) to detect rotated faces. This template is designed to find stable features of center-rotated objects in edge maps. Based on skin detection and edge extraction, our method searches for face-like areas and gets their orientations by RT searching. Then the candidates are rotated upright and a frontal face detector is used to determine the existence of faces. A system integrating these techniques is presented. Experimental results show that our algorithm is effective for detecting human faces rotated at any angle with different sizes, lighting conditions and backgrounds.
Shengye Yan, Xilin Chen 0001, Wen Gao 0001
ICASSP (3)2
2003 Rotated face detection in color images using radial template (RT)
abstract
In this paper, we propose a face detection algorithm to locate faces rotated in any orientation. Detecting rotated faces is important for a face detection system. First, we present a novel model named radial template (RT) to detect rotated faces. This template is designed to find stable features of center-rotated objects in edge maps. Based on skin detection and edge extraction, our method searches for face-like areas and gets their orientations by RT searching. Then the candidates are rotated upright and a frontal face detector is used to determine existence of faces. A system integrating these techniques is presented. Experimental results show that our algorithm is effective to detect human faces rotated in any angle with different sizes, lighting conditions and backgrounds.
Shengye Yan, Xilin Chen 0001, Wen Gao 0001
ICME2