EDBT 2026 Demo / reviewers in the wild / expert
Zhanjie Song
dblp:31/5015
· DBLP profile ↗
26ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-6654-767XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 2 since 2021Theory of computation · 2 · 2 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Point Evolution Hierarchy Network for Weak Single-Point Human ParsingabstractSparsely single-point human parsing aims at segmenting the human body into fine-grained categories via weak point-level labels (e.g., point-level, scribble-level, or image-level, etc). The point-level label, especially single-point supervision, can simultaneously preserve spatial positions as well as take light annotation time, which is particularly advantageous in alleviating the human labeling burden. However, how to obtain satisfactory parsing performance under limited sparse point annotations is challenging, which requires further investigation. In this paper, we propose a novel end-to-end Point Evolution Hierarchy human parsing Network (PEHNet) for fine-grained human parsing task that just leverages single-point supervision. Motivated by the concept of a divide-and-conquer strategy, we partition all pixels into three distinct groups, i.e., single-point labels, pseudo-region labels, and unlabeled pixels, then optimize each group with suitable mechanisms. To expand the coverage of single-point labels, we introduce a point dissemination module that generates high-quality pseudo-region labels. Furthermore, the point-level spatial position information inherently preserves the structural characteristics of the human body. Inspired by this hierarchical property, we devise a point-level human hierarchy-wise constraint that guides the prediction probabilities to align with the inherent hierarchy of the human body. Experimental results demonstrate that the proposed PEHNet outperforms state-of-the-art parsing methods on two popular human parsing benchmark datasets (LIP and ATR) and one semantic segmentation dataset (Pascal VOC 2012). Sanyi Zhang, Xiaochun Cao, Long Ye, Zhanjie Song, Guo-Jun Qi, Jie Zhou 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Frame-part-activated deep reinforcement learning for Action Prediction
Lei Chen 0069, Zhanjie Song |
Pattern Recognit. Lett. | 2 |
| 2024 | Toward Generalizable Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection has achieved great success in past years, which can be used in autonomous driving for intelligent transportation system. Most existing multispectral pedestrian detection approaches are developed on the assumption that training and test data belong to an identical distribution, which does not guarantee a good generalization to cross-domain (unseen) data. In this paper, we aim to develop a generalizable multispectral pedestrian detector, which achieves a favorable performance on both intra-dataset evaluation and cross-dataset evaluation. To achieve this goal, we conduct intra-dataset and cross-dataset experiments using single-modal and multi-modal data. By deep analysis, we find that, compared to visible or multi-modal data, thermal data not only has a best cross-dataset generalization, but also generates high-quality proposals on intra-dataset and cross-dataset evaluations. Inspired by this, we propose a novel thermal-first and fusion-second network (called TFNet) for multispectral pedestrian detection. In our TFNet, we first employ a thermal-based proposal network to extract candidate pedestrian proposals. After that, we design a transformer fusion based head network to further classify/regress these proposals. Experiments are performed on three public datasets. The comprehensive results demonstrate the effectiveness of our proposed TFNet on both intra-dataset and cross-dataset evaluations. We hope that our simple design can promote the future study on generalizable multispectral pedestrian detection. Fuchen Chu, Jiale Cao, Zhanjie Song, Yanwei Pang, Xuelong Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Learning cross-task relations for panoptic driving perception
Zhanjie Song, Linqing Zhao |
Pattern Recognit. Lett. | 1 |
| 2022 | Ambiguousness-Aware State Evolution for Action PredictionabstractIn this paper, we propose an ambiguousness-aware state evolution (AASE) method which represents the uncertainty of the input sequence and evolves the subsequent skeletons to generate a reasonable full-length sequence for action prediction. Unlike most existing methods that enforce partial sequences with the labels of full-length videos and ignore the semantic information of the subsequent action, we develop an evolution method by predicting the instructional actions and generating the reasonable candidate subsequent actions, so that the ambiguity of the full sequence’s label supervising for the partial actions can be effectively alleviated. Our method generates the rational subsequent actions under the instructional action class to complement the partially observed action sequence. We design two criteria for a rational generation: 1) the instruction of subsequent action keeps the semantic consistency with the observed sequence; 2) the generation sequence is satisfied with the distribution of the sequence of real data. Moreover, we design an uncertainty module to decide the instructional action class for the generation network. AASE predicts instructional actions with uncertainty learning and evolves different instructional actions by generating the subsequent skeletons, which find the most probable action to represent the partially observed action by learning the way of perceiving the tendency of the ongoing action. We conduct experiments on seven widely used action datasets: NTU-60, NTU-120, UCF101, UT-Interaction, BIT, PKU-MMD and HMDB51, and our experimental results clearly demonstrate that our method achieves very competitive performance with state-of-the-art. Lei Chen 0069, Jiwen Lu, Zhanjie Song, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Learning Hybrid Semantic Affinity for Point Cloud SegmentationabstractIn this paper, we present a hybrid semantic affinity learning method (HSA) to capture and leverage the dependencies of categories for 3D semantic segmentation. Unlike existing methods that only use the cross-entropy loss to perform one-to-one supervision and ignore the semantic relations between points, our approach aims to learn the label dependencies between 3D points from a hybrid perspective. From a global view, we introduce the structural correlations among different classes to provide global priors for point features. Specifically, we fuse word embeddings of labels and scene-level features as category nodes, which are processed via a graph convolutional network (GCN) to produce the sample-adapted global priors. These priors are then combined with point features to enhance the rationality of semantic predictions. From a local view, we propose the concept of local affinity to effectively model the intra-class and inter-class semantic similarities for adjacent neighborhoods, making the predictions more discriminative. Experimental results show that our method consistently improves the performance of state-of-the-art models across indoor (S3DIS, ScanNet), outdoor (SemanticKITTI), and synthetic (ShapeNet) datasets. Zhanjie Song, Linqing Zhao, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | AIParsing: Anchor-Free Instance-Level Human ParsingabstractMost state-of-the-art instance-level human parsing models adopt two-stage anchor-based detectors and, therefore, cannot avoid the heuristic anchor box design and the lack of analysis on a pixel level. To address these two issues, we have designed an instance-level human parsing network which is anchor-free and solvable on a pixel level. It consists of two simple sub-networks: an anchor-free detection head for bounding box predictions and an edge-guided parsing head for human segmentation. The anchor-free detector head inherits the pixel-like merits and effectively avoids the sensitivity of hyper-parameters as proved in object detection applications. By introducing the part-aware boundary clue, the edge-guided parsing head is capable to distinguish adjacent human parts from among each other up to 58 parts in a single human instance, even overlapping instances. Meanwhile, a refinement head integrating box-level score and part-level parsing quality is exploited to improve the quality of the parsing results. Experiments on two multiple human parsing datasets (i.e., CIHP and LV-MHP-v2.0) and one video instance-level human parsing dataset (i.e., VIP) show that our method achieves the best global-level and instance-level performance over state-of-the-art one-stage top-down alternatives. Sanyi Zhang, Xiaochun Cao, Guo-Jun Qi, Zhanjie Song, Jie Zhou 0001 |
IEEE Trans. Image Process. | 4 |
| 2021 | Efficient iterative thresholding algorithms with functional feedbacks and null space tuning
Ningning Han, Shidong Li, Zhanjie Song |
Signal Process. | 3 |
| 2021 | Recurrent Semantic Preserving Generation for Action PredictionabstractIn this paper, we propose a recurrent semantic preserving generation (RSPG) method for action prediction. Unlike most existing methods which don't make full use of information from partially observed sequences, we develop a generation architecture to complement the sequence of skeletons for predicting the action, which can exploit more potential information of the movement tendency. Our method learns to capture the tendency of observed sequences and complement the subsequent action with adversarial learning under some constrains, which preserves the consistency between the generation sequence and the observed sequence. By generating the subsequent action, our method can predict the action with the most probability. Moreover, the redundant generation introduces the noise and disturbs the prediction. The insufficient generation cannot exploit the potential information for improving the effect of predicting the action. Our RSPG controls the generation step in a recurrent manner for maximizing the discriminative information of actions, which can adapt to the variable length of different actions. We evaluate our method on four popular action datasets: NTU, UCF101, BIT, and UT-Interaction, and experimental results show that our method achieves very competitive performance with the state-of-the-art. Lei Chen 0069, Jiwen Lu, Zhanjie Song, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Human Parsing With Pyramidical Gather-Excite ContextabstractHuman parsing, especially in the wild, has attracted a lot of attention due to its great potential in many real-world applications. The Pyramid Spatial Parsing (PSP) module has shown superior performances in scene and human parsing tasks. However, the basic AvgPool operation in PSP equally aggregates spatial clues of a local region, and thus mixes up influences of different human parts presented in this region. It results in failures in capturing useful contexts relevant to parsing different parts. To address this problem, a suitable mechanism to collect spatial clues aligning with different human parts is proposed in this paper. We employ a Gather-Excite (GE) operation, a replacement of the AvgPool-Upsample operation in a pyramidical structure, to accurately reflect relevant human parts of various scales. The GE operation contains two steps: the gather operation that adaptively aggregates spatial clues to relevant human parts, and the excite operation that generates new feature maps with the gathered contextual information. This results in a novel Pyramidical Gather-Excite Context (PGEC) module to solve the multi-scale problem and parse person at various scales. The PGEC module is composed of multiple GE operations with different spatial extents and aggregates local and global spatial clues for better modeling multi-scale contextual information in parallel. Moreover, we integrate the PGEC module with fine-grained details, edge preserving module and deep supervision to formulate a novel PGEC Network (PGECNet) for human parsing. The proposed PGECNet has achieved state-of-the-art performance on four single-person human parsing datasets (i.e., LIP, PPSS, ATR and Fashion Clothing) and two multi-person human parsing datasets (i.e., PASCAL-Person-Part and CIHP). The experimental results show that the proposed PGEC is superior to the PSP and ASPP modules especially in single-human parsing task. The source code is publicly available at https://github.com/31sy/PGECNet. Sanyi Zhang, Guo-Jun Qi, Xiaochun Cao, Zhanjie Song, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Task-Aware Attention Model for Clothing Attribute PredictionabstractClothing attribute recognition, especially in unconstrained street images, is a challenging task for multimedia. Existing methods for multi-task clothing attribute prediction often ignore the relation between specific attributes and positions. However, the attribute response is always location-sensitive, i.e., different spatial locations have various contributions to attributes. Inspired by the locality of clothing attributes, in this paper, we introduce the attention mechanism to incorporate the impact of positions for clothing attribute prediction with only image-level annotations. However, the performance improvement is limited if we directly use the traditional spatial attention model for each task since it does not take the influence from other tasks into account. Instead, we propose a novel task-aware attention mechanism, which estimates the importance of each position across different tasks. We first evaluate a task attention network with an end-to-end multi-task clothing attribute learning architecture on the shop domain. And then, we employ curriculum learning strategy, which transfers the well-trained shop domain attribute knowledge to the street domain attribute prediction. Experiments are conducted on three clothing benchmarks, i.e., cross-domain clothing attribute dataset, woman clothing dataset, and man clothing dataset. The performance of attribute prediction demonstrates the superiority of the proposed task-aware attention mechanism over several state-of-the-art methods both in shop and street domains. Sanyi Zhang, Zhanjie Song, Xiaochun Cao, Hua Zhang 0008, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Learning principal orientations and residual descriptor for action recognition
Lei Chen 0069, Zhanjie Song, Jiwen Lu, Jie Zhou 0001 |
Pattern Recognit. | 2 |
| 2018 | Part-Activated Deep Reinforcement Learning for Action Prediction
Lei Chen 0069, Jiwen Lu, Zhanjie Song, Jie Zhou 0001 |
ECCV (3) | 3 |
| 2018 | Watch fashion shows to tell clothing attributes
Sanyi Zhang, Si Liu 0001, Xiaochun Cao, Zhanjie Song, Jie Zhou 0001 |
Neurocomputing | 4 |
| 2017 | Learning Pooling for Convolutional Neural Network
Manli Sun, Zhanjie Song, Xiaoheng Jiang, Yanwei Pang |
Neurocomputing | 2 |
| 2017 | Parameterization of LSB in Self-Recovery Speech Watermarking Framework in Big Data MiningabstractThe privacy is a major concern in big data mining approach. In this paper, we propose a novel self-recovery speech watermarking framework with consideration of trustable communication in big data mining. In the framework, the watermark is the compressed version of the original speech. The watermark is embedded into the least significant bit (LSB) layers. At the receiver end, the watermark is used to detect the tampered area and recover the tampered speech. To fit the complexity of the scenes in big data infrastructures, the LSB is treated as a parameter. This work discusses the relationship between LSB and other parameters in terms of explicit mathematical formulations. Once the LSB layer has been chosen, the best choices of other parameters are then deduced using the exclusive method. Additionally, we observed that six LSB layers are the limit for watermark embedding when the total bit layers equaled sixteen. Experimental results indicated that when the LSB layers changed from six to three, the imperceptibility of watermark increased, while the quality of the recovered signal decreased accordingly. This result was a trade-off and different LSB layers should be chosen according to different application conditions in big data infrastructures. Zhanjie Song, Wenhuan Lu, Daniel Sun 0004, Jianguo Wei |
Secur. Commun. Networks | 2 |
| 2016 | Cluster-based image super-resolution via jointly low-rank and sparse representation
Ningning Han, Zhanjie Song |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Salient object detection using color spatial distribution and minimum spanning tree weight
Chang Tang, Chunping Hou, Pichao Wang, Zhanjie Song |
Multim. Tools Appl. | 4 |
| 2016 | Stereoscopic image quality assessment method based on binocular combination saliency model
Qinggang Meng, Zhihan Lyu, Zhanjie Song, Zhiqun Gao |
Signal Process. | 5 |
| 2015 | Error evaluation of free-form surface based on distance function of measured point to surface
Gaiyun He, Zhanjie Song |
Comput. Aided Des. | 3 |
| 2015 | A perceptual stereoscopic image quality assessment model accounting for binocular combination behavior
Zhiqun Gao, Rongrong Chu, Zhanjie Song |
J. Vis. Commun. Image Represent. | 5 |
| 2015 | Truncation Error Analysis on Reconstruction of Signal From Unsymmetrical Local Average SamplingabstractThe classical Shannon sampling theorem is suitable for reconstructing a band-limited signal from its sampled values taken at regular instances with equal step by using the well-known sinc function. However, due to the inertia of the measurement apparatus, it is impossible to measure the value of a signal precisely at such discrete time. In practice, only unsymmetrically local averages of signal near the regular instances can be measured and used as the inputs for a signal reconstruction method. In addition, when implemented in hardware, the traditional sinc function cannot be directly used for signal reconstruction. We propose using the Taylor expansion of sinc function to reconstruct signal sampled from unsymmetrically local averages and give the upper bound of the reconstruction error (i.e., truncation error). The convergency of the reconstruction method is also presented. Yanwei Pang, Zhanjie Song, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2014 | Beautifying Fisheye Images using Orientation and Shape CuesabstractFisheye images, due to their wide range of vision, become more and more popular in our daily life. However, the fisheye images usually suffer from misalignment that reduces their visual pleasure. In this paper, we develop a computational method for enhancing the aesthetics of such images by exploiting the orientation and shape cues. More specifically, the orientation cue is based on the observation that cameras are often oriented when taking photos, so that their upvectors are parallel to vertical linear structures in the scene. While the shape one refers to that after repositing the fisheye image, the circular shape should be preserved. By employing these two rules as our basic aesthetic guidelines, our method can correct the rotation angle between the camera coordinate and the world coordinate to make the virtual camera oriented, and complete the missing part. Experimental results on a number of challenging indoor and outdoor fisheye images show the effectiveness of our approach, and demonstrate the superior aesthetics of the proposed method compared to the state-of-the-arts. Xiaobo Wang 0001, Xiaochun Cao, Xiaojie Guo 0001, Zhanjie Song |
ACM Multimedia | 4 |
| 2013 | Balance between object and background: Object-enhanced features for scene image classification
Zhong Ji, Yuting Su 0001, Zhanjie Song, Shikai Xing |
Neurocomputing | 4 |
| 2012 | An Improved Nyquist-Shannon Irregular Sampling Theorem From Local AveragesabstractThe Nyquist–Shannon sampling theorem is on the reconstruction of a band-limited signal from its uniformly sampled samples. The higher the signal bandwidth gets, the more challenging the uniform sampling may become. To deal with this problem, signal reconstruction from local averages has been studied in the literature. In this paper, we obtain an improved Nyquist–Shannon sampling theorem from general local averages. In practice, the measurement apparatus gives a weighted average over an asymmetrical interval. As a special case, for local averages from symmetrical interval, we show that the sampling rate is much lower than that of a result by Gröchenig. Moreover, we obtain two exact dual frames from local averages, one of which improves a result by Sun and Zhou. At the end of this paper, as an example application of local average sampling, we consider a reconstruction algorithm: the piecewise linear approximations. Zhanjie Song, Yanwei Pang, Chunping Hou, Xuelong Li 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2007 | An Average Sampling Theorem for Bandlimited Stochastic ProcessesabstractIn this correspondence, we give an average sampling theorem for bandlimited stochastic processes which shows that many average sampling theorems have their counterparts for stochastic signals. Zhanjie Song, Wenchang Sun, Xingwei Zhou, Hengxin Hou |
IEEE Trans. Inf. Theory | 1 |