Songhao Zhu

dblp:91/4729 · DBLP profile ↗
← Back
48ranked-venue papers
25as first author
20since 2021 · last 2026
0000-0002-9891-5692ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 16 first-author · 12 since 2021Artificial intelligence and machine learning · 21 · 9 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Few-shot object detection via dynamic feature enhancement and attention template matching
Ruqi Su, Songhao Zhu
Appl. Intell.3
2026 Enhanced image restoration via multi-scale cross-domain feature fusion
Songhao Zhu
Neural Comput. Appl.3
2025 Unsupervised cross-domain pedestrian re-identification via squeeze-excitation attention and latent feature mining
Songhao Zhu
Multim. Syst.3
2025 Multi-branch blind-spot network with multi-class replacement refinement for self-supervised image denoising
Songhao Zhu, Kangkang Ji
Multim. Tools Appl.1
2025 Lightweight multi-scale vehicle object detection algorithm based on optimized YOLOv8
Hanqing Duan, Songhao Zhu
Pattern Anal. Appl.3
2024 Future Does Matter: Boosting 3D Object Detection with Temporal Motion Estimation in Point Cloud Sequences
Runkai Zhao, Cong Nie, Songhao Zhu
BMVC6
2024 Subdomain alignment based open-set domain adaptation image classification
Kangkang Ji, Qingliang Zhang, Songhao Zhu
J. Vis. Commun. Image Represent.3
2024 CTFCD: Channel transformer based on full convolutional decoder for single image deraining
Shaohan Tan, Songhao Zhu
J. Vis. Commun. Image Represent.3
2024 Multi-level similarity transfer and adaptive fusion data augmentation for few-shot object detection
Songhao Zhu
J. Vis. Commun. Image Represent.1
2024 Occluded pedestrian re-identification via Res-ViT double-branch hybrid network
Yunbin Zhao, Songhao Zhu
Multim. Syst.2
2024 Domain-adaptive person re-identification via domain alignment and mutual pseudo-label refinement
Songhao Zhu
Multim. Syst.1
2024 Few-shot object detection via data augmentation and distribution calibration
Songhao Zhu
Mach. Vis. Appl.1
2024 SMFRNet: Complex Scene Lane Detection With Start Point-Guided Multi-Dimensional Feature Refinement
abstract
Lane lines play a crucial role in the traffic system. However, due to the diversity of lane categories, road conditions and weather environments, as well as the different aspect ratios of lane lines, lane detection algorithms face many challenges. This paper proposes a multi-dimensional feature refinement method for complex scene lane detection based on start point guidance. Due to the fact that lane lines often traverse the entire image, capturing sufficient context is crucial, and lane lines also have specific local patterns that require detailed low-level features for accurate localization. We propose global feature refinement (GFR) and lane aware gather (LAG) to refine the features from the following two dimensions: global enhancement and local refinement. To generate high-quality anchors, we predict the start point coordinate of lane instances through start point coordinate prediction (SPCP). To better fit 2D lane detection, we adopt the more general penalty LaneIoU (PLIoU) as the loss function to evaluate the predicted results. Experimental results demonstrate that the proposed method performs well in in lane detection tasks in complex scenes and has strong competitiveness among existing methods.
Shaohan Tan, Songhao Zhu
IEEE Trans. Circuits Syst. Video Technol.3
2023 Infrared-visible person re-identification via Dual-Channel attention mechanism
Zhihan Lyu, Songhao Zhu, Dongsheng Wang 0010
Multim. Tools Appl.2
2022 Attention-calibration based double-branch cross-domain person re-identification
Kangkang Ji, Songhao Zhu
Knowl. Based Syst.3
2022 Cross-modality person re-identication with triple-attentive feature aggregation
Songhao Zhu, Dongsheng Wang 0010
Multim. Tools Appl.2
2022 Whole constraint and partial triplet-center loss for infrared-visible re-identification
Zhihan Lyu, Songhao Zhu, Dongsheng Wang 0010
Neural Comput. Appl.2
2022 Short range correlation transformer for occluded person re-identification
Yunbin Zhao, Songhao Zhu, Dongsheng Wang 0010
Neural Comput. Appl.2
2021 Pitch contours curve frequency domain fitting with vocabulary matching based music generation
Runnan Lang, Songhao Zhu, Dongsheng Wang 0010
Multim. Tools Appl.2
2021 An approach to improve SSD through mask prediction of multi-scale feature maps
Yaqin Zhao, Songhao Zhu
Pattern Anal. Appl.3
2019 Joint Learning of Self-Representation and Indicator for Multi-View Image Clustering
abstract
Multi-view subspace clustering aims to divide a set of multisource data into several groups according to their underlying subspace structure. Although the spectral clustering based methods achieve promotion in multi-view clustering, their utility is limited by the separate learning manner in which affinity matrix construction and cluster indicator estimation are isolated. In this paper, we propose to jointly learn the self-representation, continue and discrete cluster indicators in an unified model. Our model can explore the subspace structure of each view and fusion them to facilitate clustering simultaneously. Experimental results on two benchmark datasets demonstrate that our method outperforms other existing competitive multi-view clustering methods.
Songsong Wu, Zhiqiang Lu, Hao Tang 0005, Yan Yan 0002, Songhao Zhu, Xiaoyuan Jing
ICIP5
2019 Target Tracking via Two-Branch Spatio-Temporal Regularized Correlation Filter Network
Wenbo Guo 0006, Songhao Zhu
PRCV (1)3
2018 Sparse Representation and Weighted Clustering Based Abnormal Behavior Detection
abstract
Abnormal activity detection is a challenging problem divided into global abnormal activity detection and local abnormal activity detection. First, the hybrid histogram of optical flow feature is extracted; then, the double sparse representation is proposed to tackle the issue of global abnormal activity detection; finally, for the issue of local abnormal activity detection, the foreground of region of interest within the current frame is first detected, and then the method of online weighted clustering is utilized to detect local abnormal activity. Experiments results conducted on UMN datasets and UCSD datasets validate the advantages of the proposed method.
Dongliang Jin, Songhao Zhu, Songsong Wu, Xiaoyuan Jing
ICPR2
2018 Weak Supervised Learning Based Abnormal Behavior Detection
abstract
Artificial features are adopted in most of the existing abnormal behavior detection. However, it is difficult to choose and design an effective behavior feature for the reason of highly computational complexity and complex scenarios. To solve this problem, temporal consistency based weak supervised abnormal behavior detection method is proposed in this paper. First, temporal gram matrices are constructed for a given video sequence. Then, a pair of behavior units (Candidate action fragment) are formed by exploiting the temporal consistency and the smoothness of human behavior, which helps to locate the start frame and end frame of the related class of abnormal behavior in the video sequence to train the corresponding classifier. Finally, sparse reconstruction is here utilized to detect abnormal behavior. Experiments conducted on common database including CAVIAR and Crossing demonstrate the effectiveness of the proposed method.
Xian Sun 0002, Songhao Zhu, Songsong Wu, Xiaoyuan Jing
ICPR2
2017 Integration of semantic and visual hashing for image retrieval
Songhao Zhu, Dongliang Jin, Yajie Sun, Guozheng Xu
J. Vis. Commun. Image Represent.1
2016 Multi-view semi-supervised learning for image classification
Songhao Zhu, Xian Sun 0002, Dongliang Jin
Neurocomputing1
2016 Multi-target tracking via hierarchical association learning
Songhao Zhu, Chengjian Sun
Neurocomputing1
2016 Local abnormal behavior detection based on optical flow and spatio-temporal gradient
Songhao Zhu, Juanjuan Hu
Multim. Tools Appl.1
2016 Tracklet association based multi-target tracking
Songhao Zhu, Chengjian Sun
Multim. Tools Appl.1
2016 Statistical background model-based target detection
Songhao Zhu, Lingling Chen
Pattern Anal. Appl.2
2015 Image Annotation Based on Multi-view Learning
Songhao Zhu, Chengjian Sun
ICIG (2)2
2015 Deep neural network based image annotation
Songhao Zhu, Chengjian Sun, Shuhan Shen
Pattern Recognit. Lett.1
2014 Content based image retrieval via a transductive model
Songhao Zhu, Liming Zou, Baojie Fang
J. Intell. Inf. Syst.1
2013 Image annotation using high order statistics in non-Euclidean spaces
Songhao Zhu, Juanjuan Hu, Baoyun Wang, Shuhan Shen
J. Vis. Commun. Image Represent.1
2013 Transductive multi-distance learning for video search
Songhao Zhu, Yuncai Liu
Pattern Anal. Appl.1
2012 Spreading activation theory based image annotation
abstract
The overwhelming amounts of digital images on the Web and personal computers have triggered the requirement of an effective tool to retrieve images of interest using semantic concepts. Due to the semantic gap between low level content features and its high level semantic features of an image, however, the performances of many existing automatic image annotation algorithms are not so satisfactory. In this paper, a novel approach based on the cognitive science is proposed to improve the quality of annotations. The main idea is that the tags of an image are considered as nodes in a semantic network, and the relevance between the tags and image contents is regulated using the spreading activation theory. After the spreading activation process finishes, each tag will be assigned an appropriate value with respect to its relation to other tags. Experimental results conducted on the 50,000 Flickr images demonstrate that the proposed scheme can effectively improve the performance in automatic image annotation.
Songhao Zhu, Baoyun Wang, Yuncai Liu
ICASSP1
2012 Using non-parametric quantum theory to rank images
abstract
Recently learning to rank has become one of the popular means to create a ranking model for social image search. However, the results of existing approaches are not as satisfactory for the large gap between low-level visual features and high-level semantic concepts, and these sophisticated approaches require a significant amount of parameters tuning to be effective and efficient. In this paper, we propose a novel framework for social image re-ranking based on a non-parametric quantum technique, which reranks top retrieved images by considering the interrelationship between images through the quantum estimation and requires no explicit parameter tuning. The basic idea of the proposed framework is inspired by the photon polarization experiment supporting the theory of quantum measurement. Experimental results conducted on the Flickr dataset demonstrate the effectiveness and efficiency of the proposed framework.
Songhao Zhu, Baoyun Wang, Yuncai Liu
ICASSP1
2012 Movie abstraction via the progress of the storyline
abstract
An appropriate movie abstraction is helpful for movie producers to promote the progress of the storyline as well as for the audiences to capture the theme of the movie before watching the full-length movie. Most existing movie abstraction schemes rely heavily on video content only, which may not deliver ideal results because of the semantic gap between computer calculated low-level audiovisual features and human used high-level perceptual understanding. In this study, the authors incorporate script into movie content understanding and present a new movie abstraction approach via the progress of the storyline, which is the soul of a film that actually catches the audiences' attention. The authors first segment the movie scenes by analysis of the movie script. Then the authors conduct storyline analysis using the attention analysis and audiovisual features. Given the transition intensity values, the authors calculate the storyline progress score and adopt this as the criterion to generate movie abstraction. The promising experimental results demonstrate that the analysis of storyline evolution is an effective approach for the abstraction and understanding of movie content.
Songhao Zhu, Yaqin Zhao, Xiaoyuan Jing
IET Signal Process.1
2012 Face feature extraction and recognition based on discriminant subclass-center manifold preserving projection
Xiaoyuan Jing, Chao Lan, David Zhang 0001, Jing-Yu Yang 0001, Sheng Li 0001, Songhao Zhu
Pattern Recognit. Lett.7
2011 Video Retrieval via Learning Collaborative Semantic Distance
abstract
Graph-based semi-supervised learning approaches have been proven effective and efficient in solving the problem of the inefficiency of labeled data in many real-world application areas, such as video annotation. However, the pairwise similarity metric, a significant factor of existing approaches, has not been fully investigated. That is, these graph-based semi-supervised approaches estimate the pairwise similarity between samples mainly according to the spatial property of video data. On the other hand, temporal property, an essential characteristic of video data, is not embedded into the pairwise similarity measure. Accordingly, a novel framework for video annotation, called Joint Spatio-Temporal Correlation Learning (JSTCL), is proposed in this paper. This framework is characterized by simultaneously taking into account the spatial and temporal property of video data to achieve more accurate pairwise similarity values. We apply the proposed framework to video annotation and report superior performance compared to key existing approaches over the benchmark TRECVID data set.
Songhao Zhu, Xiaoyuan Jing
Int. J. Pattern Recognit. Artif. Intell.1
2009 Improving Semantic Scene Categorization by Exploiting Audio-Visual Features
abstract
We address the issue of categorizing scenes from feature films into semantic classifications based on the audio-visual cues. Specifically, we first exploit the grammar of film production to specify the semantic content of scenes. Then, each scene is classified into one of the following categories: conversation, action and suspense. Finally, to achieve more specific scene and consist with human perception, conversation scene is further categorizes into emotional conversation and common one, and action scene is further categorizes into gunfight, beating and chasing scene. This work is a step toward browsing and retrieval content of feature films in limited bandwidth, video repository, and rating of feature films of interest effectively and efficiently.
Songhao Zhu, Junchi Yan, Yuncai Liu
ICIG1
2009 A novel semantic model for video concept detection
abstract
Graph-based semi-supervised learning approaches have been proven effective and efficient in solving the problem of the inefficiency of training samples in many real-world application areas, such as video annotation. As a significant factor of these algorithms, however, pairwise similarity metric has not been fully investigated. On the one hand, for general video annotation methods, the estimation of pairwise similarity between two samples relies on the spatial property of video data. On the other hand, temporal property, which is an essential character of video data, is not embedded into the pairwise similarity metric. Therefore, in this paper, a novel method, called Joint Spatio-Temporal Correlation Learning (JSTCL), is proposed to improve the accuracy of video annotation. This method is characterized by simultaneously taking into account both the spatial and temporal property of video data to well represent the pairwise similarity. Experiments conducted on the TRECVID demonstrate the efficiency of the proposed method.
Songhao Zhu, Yuncai Liu
ICIP1
2009 Automatic scene detection for advanced story retrieval
Songhao Zhu, Yuncai Liu
Expert Syst. Appl.1
2009 Video scene segmentation and semantic representation using a novel scheme
Songhao Zhu, Yuncai Liu
Multim. Tools Appl.1
2009 Semi-Supervised Learning Model Based Efficient Image Annotation
abstract
Automatic image annotation is a promising way to achieve more effective image management and retrieval. However, system performances of the existing state-of-the-art keyword annotation schemes are often not so satisfactory. Therefore, image annotation refinement is crucial to improve the imprecise annotation results. In this paper, a novel approach is developed to automatically annotate image content by a semi-supervised learning model. With perceptual visual characteristics, the candidate annotations of unlabelled images are first obtained based on a progressive model. Then, a transducitive model, random walk with restart algorithm is used to refine these candidate annotations and the top ones are reserved as the final annotations. Experiments conducted on the typical Corel dataset show the effectiveness of the proposed approach.
Songhao Zhu, Yuncai Liu
IEEE Signal Process. Lett.1
2008 A novel scheme for video scenes segmentation and semantic representation
abstract
Grouping video contents into semantic segments is the crucial pass to content-based video summarization and retrieval. In this paper, we present a novel scene segmentation and semantic representation scheme for various video types. We first detect video shot using a coarse-to-fine algorithm. The key frames without useful information are detected and removed using template matching. Spatio-temporal coherent shots are then grouped into the same scene based on the temporal constraint of video content and visual similarity of shot activity. With general editing technique used in the continuously recorded video, semantic representation of scene content is specified to satisfy human demand on video retrieval. The proposed algorithm has been performed on various types of videos containing movie and TV program. Promising experimental results shows that the proposed method makes sense to efficient retrieval of video contents of interest.
Songhao Zhu, Yuncai Liu
ICME1
2008 Image annotation refinement using semantic similarity correlation
abstract
Automatic image annotation is a promising way to achieve more effective image management and retrieval by using keywords. However, system performances of the existing state-of-the-art keyword annotation schemes are often not so satisfactory. Therefore, image annotation refinement is crucial to improve the imprecise annotation results. In this paper, a novel approach is developed to automatically refine the initial annotation of images. First, for a query image, the candidate annotations are obtained by a step-up model-based algorithm using perceptual visual characteristic. Then, a refine algorithm, fast random walk with restart is used to re-rank the candidate annotations and the top ones are reserved as the final annotations. Experiments conducted on the typical Corel dataset shows that the proposed scheme can effectively improve the automatic annotation performance.
Songhao Zhu, Yuncai Liu
ICPR1
2008 Scene Segmentation and Semantic Representation for High-Level Retrieval
abstract
In this letter, a novel framework to segment video scene and represent scene content is proposed. Firstly, video shots are detected using a rough-to-fine algorithm. Secondly, key frames are selected adaptively, and redundant key frames are removed using template matching. Then, spatio-temporal coherent shots are clustered into the same scene. Finally, under the full analysis of typical characters on continuously recorded videos, video scene content is semantically represented to satisfy human demand on video retrieval. Experimental results show the proposed method makes sense to efficient retrieval of video content of interest.
Songhao Zhu, Yuncai Liu
IEEE Signal Process. Lett.1