Songzhi Su

dblp:36/7815 · also Song-Zhi Su · DBLP profile ↗
← Back
57ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0001-8961-9405ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 25 · 13 since 2021Systems, architecture and hardware · 4 · 3 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ARDiff: Anisotropic Residual Diffusion for Heterogeneous Graph Learning
abstract
Learning representations on graphs is foundational for many downstream tasks, and its synergy with diffusion models has emerged as a promising direction. However, diffusion-based methods for heterogeneous graphs remain underexplored, confronting two principal challenges: (1) The presence of noise and structural heterogeneity in graphs makes it challenging to accurately capture semantic transitions among diverse relation types. (2) The isotropic Gaussian noise used in forward diffusion fails to reflect graphs' inherent semantics and structural anisotropy. To address these, we propose ARDiff, a novel framework that integrates residual diffusion with anisotropic noise for heterogeneous graph learning. Specifically, we propose a semantic residual diffusion mechanism that progressively refines node embeddings by orchestrating transitions from low-semantic (high-noise) to high-semantic (low-noise) relational contexts, thus enabling step-wise distillation of task-relevant information. In addition, to address the limitations of conventional diffusion, we introduce an anisotropic diffusion strategy: in the forward process, noise injection is oriented by structural and semantic priors; in the denoising step, a conditional diffusion mechanism is guided by a random walk encoding, enhancing both topological consistency and semantic alignment. Extensive evaluation on heterogeneous graph datasets demonstrates that ARDiff significantly surpasses current leading methods in link prediction and node classification, setting a new paradigm and benchmark in heterogeneous graph representation learning.
Li Li 0122, Nannan Zong, Songzhi Su
AAAI5
2026 MTSCL-Net: Multi-level temporal spatial contrastive learning for robust breast tumor segmentation in DCE-MRI
Jiezhou He, Zhiming Luo, Songzhi Su, Shaozi Li
Pattern Recognit.5
2026 MAPLE-VAE: MAP-based LaplacE VAE for robust density estimation
Nannan Zong, Li Li 0122, Wenfang Xiang, Changle Zhou, Songzhi Su
Pattern Recognit.5
2025 PFDiff: Training-Free Acceleration of Diffusion Models Combining Past and Future Scores
abstract
Diffusion Probabilistic Models (DPMs) have shown remarkable potential in image generation, but their sampling efficiency is hindered by the need for numerous denoising steps. Most existing solutions accelerate the sampling process by proposing fast ODE solvers. However, the inevitable discretization errors of the ODE solvers are significantly magnified when the number of function evaluations (NFE) is fewer. In this work, we propose PFDiff, a novel training-free and orthogonal timestep-skipping strategy, which enables existing fast ODE solvers to operate with fewer NFE. Specifically, PFDiff initially utilizes score replacement from past time steps to predict a springboard. Subsequently, it employs this ``springboard" along with foresight updates inspired by Nesterov momentum to rapidly update current intermediate states. This approach effectively reduces unnecessary NFE while correcting for discretization errors inherent in first-order ODE solvers. Experimental results demonstrate that PFDiff exhibits flexible applicability across various pre-trained DPMs, particularly excelling in conditional DPMs and surpassing previous state-of-the-art training-free methods. For instance, using DDIM as a baseline, we achieved 16.46 FID (4 NFE) compared to 138.81 FID with DDIM on ImageNet 64x64 with classifier guidance, and 13.06 FID (10 NFE) on Stable Diffusion with 7.5 guidance scale. Code is available at https://github.com/onefly123/PFDiff.
Guangyi Wang, Yuren Cai, Lijiang Li, Wei Peng 0009, Songzhi Su
ICLR5
2025 Diffusion Sampling Correction via Approximately 10 Parameters
abstract
While powerful for generation, Diffusion Probabilistic Models (DPMs) face slow sampling challenges, for which various distillation-based methods have been proposed. However, they typically require significant additional training costs and model parameter storage, limiting their practicality. In this work, we propose **P**CA-based **A**daptive **S**earch (PAS), which optimizes existing solvers for DPMs with minimal additional costs. Specifically, we first employ PCA to obtain a few basis vectors to span the high-dimensional sampling space, which enables us to learn just a set of coordinates to correct the sampling direction; furthermore, based on the observation that the cumulative truncation error exhibits an ``S"-shape, we design an adaptive search strategy that further enhances the sampling efficiency and reduces the number of stored parameters to approximately 10. Extensive experiments demonstrate that PAS can significantly enhance existing fast solvers in a plug-and-play manner with negligible costs. E.g., on CIFAR10, PAS optimizes DDIM's FID from 15.69 to 4.37 (NFE=10) using only **12 parameters and sub-minute training** on a single A100 GPU. Code is available at https://github.com/onefly123/PAS.
Guangyi Wang, Wei Peng 0009, Lijiang Li, Yuren Cai, Songzhi Su
ICML6
2025 Multi-Scale Loss Components in 3DGS-based RGB-D SLAM Systems
abstract
We introduce a novel multi-scale loss components method designed to enhance camera pose tracking in 3D Gaussian Splatting (3DGS) based RGB-D SLAM systems. Although existing 3DGS SLAM methods show promising results, they often falter under conditions of rapid camera motion and tend to rely on simple, fine-grained L1 loss, which typically provides only single-scale optimization gradients. To overcome these limitations, our method integrates multi-scale loss components that combines Scene Loss for global consistency, Object Loss and Keypoint Loss for semi-coarse alignment, and pixel-level L1 loss for fine-grained optimization. Extensive experimental evaluations on the TUM-RGBD and Replica datasets have demonstrated that our approach markedly improves tracking accuracy and stability, especially in challenging scenarios characterized by rapid camera movements. Quantitative results indicate that our method achieves up to a statistically significant reduction in Absolute Trajectory Error compared to state-of-the-art baselines, while maintaining real-time performance. Our proposed framework thus offers a more robust and precise solution for camera pose tracking in neural implicit SLAM systems.
Songzhi Su
IJCNN2
2025 Dynamic Similarity Weighted Regression
abstract
Data imbalance poses a significant challenge in deep learning for regression tasks, where continuous labels display intrinsic ordering and vary in density across the label space. Traditional methods, mainly developed for classification problems, do not adequately address the continuous characteristics of regression. To overcome this, we introduce Dynamic Similarity Weighted Regression (DSWR), an innovative approach designed to effectively bridge the gap between label space and feature space in regression scenarios. DSWR leverages information from the label space to inform and constrain feature space representation, ensuring that samples with similar labels also exhibit similarity in their feature representations. This strategy enhances the model’s ability to learn from imbalanced data by emphasizing the continuity and structured order of the label space. It also extends the model’s applicability to complex, high-dimensional label spaces, such as those found in depth map regression tasks. Unlike traditional methods that may focus solely on local distributions or employ smoothing techniques, DSWR considers both proximate and distant label relationships. This promotes a feature space arrangement that accurately reflects the nuanced structure of the label space. Validated on three distinct datasets, DSWR has demonstrated superior performance, offering a new perspective on tackling data imbalance in regression models.
Wenfang Xiang, Nannan Zong, Songzhi Su
IJCNN3
2025 VoRec: Enhancing Recommendation with Voronoi Diagram in Hyperbolic Space
abstract
The sparse user-item interactions in recommender systems hinder the quality of embedding representations and degraded recommendation performance. Existing methods attempt to alleviate this sparsity issue by incorporating auxiliary information via item tags, but often neglect structured characteristics of embedding space, such as semantic distribution and logical relations. To this end, we propose VoRec, a novel framework that explores the spatial distribution of items and their associated tags to achieve accurate recommendations in hyperbolic space. Specifically, we employ the Voronoi diagram to partition hyperbolic space into logically related subspaces, based on tag distributions and relationships derived from existing tag taxonomies. In addition, we combine the Voronoi diagram with the Hyperbolic Graph Convolutional Network (HGCN) and exploit the respective advantages of the Poincaré and Lorentz models in hyperbolic space. Finally, we develop two types of Voronoi site update strategies, namely active and passive ones, to optimize the Voronoi diagram for recommendation tasks. The active strategy employs contrastive learning to guide updates to the Voronoi diagram, while the passive strategy adaptively freezes parameters based on information gain to regulate the update rate. Extensive experiments on four real-world benchmark datasets demonstrate that our proposed VoRec framework delivers substantial performance improvements, achieving an average 16.35% enhancement in Recall and NDCG metrics compared to state-of-the-art baselines. The model implementation is publicly available at: https://github.com/s35lay/VoRec.
Li Li 0122, Wei Peng 0009, Songzhi Su
SIGIR4
2025 HRCUNet: Hierarchical Region Contrastive Learning for Segmentation of Breast Tumors in DCE-MRI
abstract
ABSTRACT Segmenting breast tumors from dynamic contrast‐enhanced magnetic resonance images is a critical step in the early detection and diagnosis of breast cancer. However, this task becomes significantly more challenging due to the diverse shapes and sizes of tumors, which make it difficult to establish a unified perception field for modeling them. Moreover, tumor regions are often subtle or imperceptible during early detection, exacerbating the issue of extreme class imbalance. This imbalance can lead to biased training and challenge accurately segmenting tumor regions from the predominant normal tissues. To address these issues, we propose a hierarchical region contrastive learning approach for breast tumor segmentation. Our approach introduces a novel hierarchical region contrastive learning loss function that addresses the class imbalance problem. This loss function encourages the model to create a clear separation between feature embeddings by maximizing the inter‐class margin and minimizing the intra‐class distance across different levels of the feature space. In addition, we design a novel Attention‐based 3D Multi‐scale Feature Fusion Residual Module to explore more granular multi‐scale representations to improve the feature learning ability of tumors. Extensive experiments on two breast DCE‐MRI datasets demonstrate that the proposed algorithm is more competitive against several state‐of‐the‐art approaches under different segmentation metrics.
Jiezhou He, Zhiming Luo, Wei Peng 0009, Songzhi Su, Shaozi Li
Concurr. Comput. Pract. Exp.4
2024 CC-DA: Cross-Domain Consistency Data Augmentation for 3D Tumor Segmentation
abstract
Deep learning-based tumor segmentation in 3D medical images faces the challenges of limited annotated data and class imbalance. In this paper, we proposed a novel Cross-domain Consistency Data Augmentation (CC-DA) for 3D tumor segmentation. Specifically, we copy the tumor from source data and apply random transformations to enhance its diversity. Then, we paste the enhanced tumor into the organ area of target data to generate a new sample. This process can alleviate class imbalance by regulating the merged tumor pixel ratio. To further enhance the generated data credibility, we proposed a domain consistency constraint that aligns the source data distribution with the target data distribution. We conduct extensive experiments on KiTS19 and LiTS17 datasets. The promising results clearly show that our CC-DA method can effectively improve the existing state-of-the-art 3D tumor segmentation performance.
Jiezhou He, Zhiming Luo, Wei Peng 0009, Songzhi Su, Shaozi Li
ICASSP4
2024 MeshStyle: Text-driven Efficient and High-Quality 3D Mesh Stylization via Hypergraph Convolution
abstract
Text-driven 3D mesh stylization aims to transform unstylized meshes into vivid stylized 3D mesh according to the provided text prompts. Existing works have achieved impressive results in this task, but they lack a specific design for processing mesh features, such as internal interactions within mesh data, resulting in unsatisfactory stylization, and slower convergence rates. To overcome these limitations, we propose a text-driven 3D stylization framework called MeshStyle, including a novel mesh feature processing module named Mesh HyperGraph Neural Network (MHGNN). Hypergraph neural network is exploited to aggregate mesh spatial features, addressing the issue of the lack of inner connectivity in mesh data. Furthermore, we incorporate depth prior and an extra diffusion prior to enhance geometry and appearance optimization, respectively. We also constructed a new dataset collected from various public 3D datasets, along with the evaluation protocol. Through both qualitative and quantitative experiments, we validate the capability of our MeshStyle.
Shihao Gao, Songzhi Su, Xizhi Chen
ICME3
2024 TSESNet: Temporal-Spatial Enhanced Breast Tumor Segmentation in DCE-MRI Using Feature Perception and Separability
Jiezhou He, Zhiming Luo, Songzhi Su, Shaozi Li
IJCAI4
2024 Patch Trajectories for Visual Odometry in Dynamic scenes
abstract
Accurately differentiating dynamic elements within dynamic scenes is crucial for enhancing the robustness of visual odometry. Previous studies segment dynamic object regions using semantic information or dense optical flow between two frames. However, information from multiple frames is necessary for more accurate motion estimation, as the motion in the real-world is highly time-dependent. In addition, dense motion estimation requires a substantial computational cost, rendering it impractical for most real cases. To address these issues, we introduce Patch Trajectories for Visual Odometry (PTVO), a method based on sparse patch trajectories for pose optimization to make visual odometry more robust in dynamic scenes. Specifically, by analyzing the trajectories of inter-frame patches, PTVO can predict the motion status of patches. Then, based on the result, PTVO applies weight reduction to dynamic patches, which can significantly enhance the performance of visual odometry in dynamic scenes. Experiments show that, without re-training, PTVO surpasses existing methods in terms of accuracy and computational efficiency in dynamic scenes.
Lijian Zhuang, Songzhi Su
IJCNN2
2024 Boosting semi-supervised learning under imbalanced regression via pseudo-labeling
abstract
Summary Imbalanced samples are widespread, which impairs the generalization and fairness of models. Semi‐supervised learning can overcome the deficiency of rare labeled samples, but it is challenging to select high‐quality pseudo‐label data. Unlike discrete labels that can be matched one‐to‐one with points on a numerical axis, labels in regression tasks are consecutive and cannot be directly chosen. Besides, the distribution of unlabeled data is imbalanced, which easily leads to an imbalanced distribution of pseudo‐label data, exacerbating the imbalance in the semi‐supervised dataset. To solve this problem, this article proposes a semi‐supervised imbalanced regression network (SIRN), which consists of two components: A, designed to learn the relationship between features and labels (targets), and B, dedicated to learning the relationship between features and target deviations. To measure target deviations under imbalanced distribution, the target deviation function is introduced. To select continuous pseudo‐labels, the deviation matching strategy is designed. Furthermore, an adaptive selection function is developed to mitigate the risk of skewed distributions due to imbalanced pseudo‐label data. Finally, the effectiveness of the proposed method is validated through evaluations of two regression tasks. The results show a great reduction in predicted value error, particularly in few‐shot regions. This empirical evidence confirms the efficacy of our method in addressing the issue of imbalanced samples in regression tasks.
Nannan Zong, Songzhi Su, Changle Zhou
Concurr. Comput. Pract. Exp.2
2023 EventPoint: Self-Supervised Interest Point Detection and Description for Event-based Camera
abstract
This paper proposes a self-supervised learned local detector and descriptor, called EventPoint, for event stream/camera tracking and registration. Event-based cameras have grown in popularity because of their biological inspiration and low power consumption. Despite this, applying local features directly to the event stream is difficult due to its peculiar data structure. We propose a new time-surface-like event stream representation method called Ten-code. The event stream data processed by Tencode can obtain the pixel-level positioning of interest points while also simultaneously extracting descriptors through a neural network. Instead of using costly and unreliable manual annotation, our network leverages the prior knowledge of local feature extraction on color images and conducts self-supervised learning via homographic and spatio-temporal adaptation. To the best of our knowledge, our proposed method is the first research on event-based local features learning using a deep neural network. We provide comprehensive experiments of feature point detection and matching, and three public datasets are used for evaluation (i.e. DSEC, N-Caltech101, and HVGA ATIS Corner Dataset). The experimental findings demonstrate that our method outperforms SOTA in terms of feature point detection and description.
Ze Huang, Li Sun 0005, Cheng Zhao 0002, Songzhi Su
WACV5
2023 Long short-distance topology modelling of 3D point cloud segmentation with a graph convolution neural network
abstract
Abstract 3D point cloud segmentation is a non‐trivial problem due to its irregular, sparse, and unordered data structure. Existing methods only consider structural relationships of a 3D point and its spatial neighbours. However, the inner‐point interactions and long‐distance context of a 3D point cloud have been less investigated. In this study, we propose an effective plug‐and‐play module called the Long Short‐Distance Topologically Modelled (LSDTM) Graph Convolutional Neural Network (GCNN) to learn the underlying structure of 3D point clouds. Specifically, we introduce the concept of subgraph to model the contextual‐point relationships within a short distance. Then the proposed topology can be reconstructed by recursive aggregation of subgraphs, and importantly, to propagate the contextual scope to a long range. The proposed LSDTM can parse the point cloud data with maximisation of preserving the geometric structure and contextual structure, and the topological graph can be trained end‐to‐end through a seamlessly integrated GCNN. We provide a case study of triple‐layer ternary topology and experimental results on ShapeNetPart, Stanford 3D Indoor Semantics and ScanNet datasets, indicating a significant improvement on the task of 3D point cloud segmentation and validating the effectiveness of our research.
Songzhi Su, Qingqi Hong, Beizhan Wang, Li Sun 0005
IET Comput. Vis.2
2023 DistVAE: Distributed Variational Autoencoder for sequential recommendation
Li Li 0122, Jianbing Xiahou, Fan Lin, Songzhi Su
Knowl. Based Syst.4
2022 VEFNet: an Event-RGB Cross Modality Fusion Network for Visual Place Recognition
abstract
Visual Place Recognition (VPR) on natural image is challenging due to the illumination variance and seasonal changes. In terms of long-term localization, the emerging event stream cameras are naturally resilient to appearance changes. In this paper, we propose a novel multi-modal network, e.g. VEFNet for VPR by learning location-specific cross RGB-event modality feature representations. Specifically, we firstly extract dense visual features via shared Convolutional Neural Network (CNN) backbone from RGB and event frames separately. Then, two branch features are fed to the cross-modality attention module to establish correspondences between the dual-modality. We also employ a self-attention module to enhance the contextual integration within densely encoded features. Finally, the learned global descriptor is used as the place representation of the dual-modality inputs for VPR. Experimental results demonstrate the state-of-the-art (SOTA) performance on the public datasets
Ze Huang, Li Sun 0005, Cheng Zhao 0002, Min Huang 0004, Songzhi Su
ICIP6
2021 NDT-Transformer: Large-Scale 3D Point Cloud Localisation using the Normal Distribution Transform Representation
abstract
3D point cloud-based place recognition is highly demanded by autonomous driving in GPS-challenged environments and serves as an essential component (i.e. loop-closure detection) in lidar-based SLAM systems. This paper proposes a novel approach, named NDT-Transformer, for real-time and large-scale place recognition using 3D point clouds. Specifically, a 3D Normal Distribution Transform (NDT) representation is employed to condense the raw, dense 3D point cloud as probabilistic distributions (NDT cells) to provide the geometrical shape description. Then a novel NDT-Transformer network learns a global descriptor from a set of 3D NDT cell representations. Benefiting from the NDT representation and NDT-Transformer network, the learned global descriptors are enriched with both geometrical and contextual information. Finally, descriptor retrieval is achieved using a query-database for place recognition. Compared to the state-of-the-art methods, the proposed approach achieves an improvement of 7.52% on average top 1 recall and 2.73% on average top 1% recall on the Oxford Robotcar benchmark.
Cheng Zhao 0002, Daniel Adolfsson, Songzhi Su, Yang Gao 0002, Tom Duckett, Li Sun 0005
ICRA4
2021 Deep 3D caricature face generation with identity and structure consistency
Songzhi Su, Juncong Lin, Guo-Rong Cai, Li Sun 0005
Neurocomputing2
2020 Local k-NNs pattern in Omni-Direction graph convolution neural network for 3D point clouds
Songzhi Su, Beizhan Wang, Qingqi Hong, Li Sun 0005
Neurocomputing2
2019 Reading Digital Numbers of Water Meter with Deep Learning Based Object Detector
Shirong Liao, Lianglin Wang, Songzhi Su
PRCV (1)4
2019 Cover patches: A general feature extraction strategy for spoofing detection
abstract
Summary Face anti‐spoofing has attracted many attentions in security applications, such as mobile payment and entrance guard. Until now, face anti‐spoofing technique is still a challenging task. Mainstream image‐based spoofing algorithms usually use global motion or texture information to distinguish whether an input face is live or fake. However, the performance of these methods are sensitive in light changes, or images acquired from different sensors. The main reason is that spoofed face image always has slight different texture in local areas, such as landmark or salient region of face. To this end, this paper proposes a novel multi‐patches feature extraction strategy to detect spoofing. First, a set of patches with specific combination scheme is selected to cover the face image. Second, features such as hand‐crafted Gray Level Co‐occurrence Matrix (GLCM), Local Binary Patterns (LBP), or deep features are extracted from these patches. Third, all features are combined as the global descriptor of the face image, then fed into an SVM classifier to verify the anti‐spoofing detection. Experimental results show that the proposed strategy can effectively enhance the performance, concerning with the accuracy of spoofed face detection in four widely used anti‐spoofing databases.
Guo-Rong Cai, Songzhi Su, Chengcai Leng, Jipeng Wu, Yun-Dong Wu, Shaozi Li
Concurr. Comput. Pract. Exp.2
2018 Combining 2D and 3D features to improve road detection based on stereo cameras
abstract
Road detection is a fundamental component of autonomous driving systems since it provides validspace and candidate regions of objects for driving decision. The core of roaddetection methods is extracting effective and discriminative features. Sincetwo‐dimensional (2D) and 3D features are complementary, the authors propose arobust multi‐feature combination and optimisation framework for stereo imagepairs, called Feature++. First, several 2D and 3D features such as Gabor andplane are, respectively, extracted after the generation of 2D super‐pixel and a3D depth image from stereo matching. Second, the combined features are fed intoa three‐layer shallow neural network classifier to decide whether a super‐pixelis road region or not. Finally, the classified results are further refined usingfully connected conditional random field (CRF), taking the content informationinto consideration. We extensively evaluate the performance of four 2D features,four 3D features, and their combinations. Experiments conducted on the KITTIROAD benchmark show that (i) the combinations of 2D and 3D features greatlyimprove the road detection performance and (ii) using CRF as a refinement stepis necessary. Overall, their proposed ‘Feature + +’ method outperforms mostmanually designed features, and is comparable with state‐of‐the‐art methods thatare based on deep learning methods.
Guo-Rong Cai, Songzhi Su, Wenli He, Yun-Dong Wu, Shaozi Li
IET Comput. Vis.2
2018 Discriminative parts learning for 3D human action recognition
Min Huang 0004, Guo-Rong Cai, Hongbo Zhang 0002, Sheng Yu 0007, Dong-Ying Gong, Donglin Cao, Shaozi Li, Songzhi Su
Neurocomputing8
2018 Attention guided U-Net for accurate iris segmentation
Sheng Lian, Zhiming Luo, Zhun Zhong, Songzhi Su, Shaozi Li
J. Vis. Commun. Image Represent.5
2018 Multi-label learning with label-specific features by resolving label correlations
Jia Zhang 0019, Candong Li, Donglin Cao, Yaojin Lin, Songzhi Su, Shaozi Li
Knowl. Based Syst.5
2018 Traffic Analytics With Low-Frame-Rate Videos
abstract
In this paper, we investigate the possibility of monitoring highway traffic based on videos whose frame rate is too low to accurately estimate motion features. The goal of the proposed method is to recognize traffic conditions instead of measuring them, as is usually the case. The main advantage of our approach comes from its ability to process low-frame-rate videos for which motion features cannot be estimated. Our method takes advantage of the highly redundant nature of traffic scenes that are pictured from a top-down perspective showing vehicles on a predominant asphalted road surrounded by background objects. Due to the limited variety of objects pictured in traffic scenes, our method gets to learn features that are specific to such images. With these features, our method is able to segment traffic images, classify traffic scenes, and estimate traffic density without requiring motion features. Different convolutional neural network models are proposed to segment traffic images in three different classes (Road, Car, and Background), classify traffic images into different categories (Empty, Fluid, Heavy, and Jam), and predict traffic density. We also propose a procedure to perform transfer learning of any of these models to new traffic scenes.
Zhiming Luo, Pierre-Marc Jodoin, Songzhi Su, Shaozi Li, Hugo Larochelle
IEEE Trans. Circuits Syst. Video Technol.3
2018 Multifeature Selection for 3D Human Action Recognition
abstract
In mainstream approaches for 3D human action recognition, depth and skeleton features are combined to improve recognition accuracy. However, this strategy results in high feature dimensions and low discrimination due to redundant feature vectors. To solve this drawback, a multi-feature selection approach for 3D human action recognition is proposed in this paper. First, three novel single-modal features are proposed to describe depth appearance, depth motion, and skeleton motion. Second, a classification entropy of random forest is used to evaluate the discrimination of the depth appearance based features. Finally, one of the three features is selected to recognize the sample according to the discrimination evaluation. Experimental results show that the proposed multi-feature selection approach significantly outperforms other approaches based on single-modal feature and feature fusion.
Min Huang 0004, Songzhi Su, Hongbo Zhang 0002, Guo-Rong Cai, Dong-Ying Gong, Donglin Cao, Shaozi Li
ACM Trans. Multim. Comput. Commun. Appl.2
2017 Feature++: Cross dimension feature fusion for road detection
abstract
Road detection is a key component of Advanced Driving Assistance Systems, which provides valid space and candidate regions of objects for vehicles. Mainstream road detection methods have focused on extracting discriminative features. In this paper, we propose a robust feature fusion framework, called “Feature++”, which is combined with superpixel feature and 3D feature extracted from stereo images. Then a neural network classifier is been trained to decide whether a superpixel is road region or not. Finally, the classified results are further refined by conditional random field. Experiments conducted on the KITTI ROAD benchmark show that the proposed “Feature++” method outperforms most manually designed features, and are comparable with state-of-the-art methods that based on deep learning architecture.
Wenli He, Guo-Rong Cai, Zhun Zhong, Songzhi Su
ICASSP4
2017 Meta-action descriptor for action recognition in RGBD video
abstract
Action recognition is one of the hottest research topics in computer vision. Recent methods represent actions based on global or local video features. These approaches, however, lack semantic structure and may not provide a deep insight into the essence of an action. In this work, the authors argue that semantic clues, such as joint positions and part‐level motion clustering, help verify actions. To this end, a meta‐action descriptor for action recognition in RGBD video is proposed in this study. Specifically, two discrimination‐based strategies – dynamic and discriminative part clustering – are introduced to improve accuracy. Experiments conducted on the MSR Action 3D dataset show that the proposed method significantly outperforms the methods without joint position semantic.
Min Huang 0004, Songzhi Su, Guo-Rong Cai, Hongbo Zhang 0002, Donglin Cao, Shaozi Li
IET Comput. Vis.2
2017 Learning rich features from objectness estimation for human lying-pose detection
Daoxun Xia, Songzhi Su, Li-Chuan Geng, Guoxi Wu, Shaozi Li
Multim. Syst.2
2017 Stratified pooling based deep convolutional neural networks for human action recognition
Sheng Yu 0007, Songzhi Su, Guo-Rong Cai, Shaozi Li
Multim. Tools Appl.3
2017 Detecting ground control points via convolutional neural network for stereo matching
Zhun Zhong, Songzhi Su, Donglin Cao, Shaozi Li, Zhihan Lyu
Multim. Tools Appl.2
2017 Improving pedestrian detection using motion-guided filtering
Yi Wang 0025, Sébastien Piérard, Songzhi Su, Pierre-Marc Jodoin
Pattern Recognit. Lett.3
2016 Probability-based method for boosting human action recognition using scene context
abstract
In this study, the authors investigate the possibility of boosting action recognition performance by exploiting the associated scene context. Towards this end, the authors model a scene as a mid‐level ‘middle layer’ in order to bridge action descriptors and action categories. This is achieved via a scene topic model, in which hybrid visual descriptors, including spatial–temporal action features and scene descriptors, are first extracted from a video sequence. Then, the authors learn a joint probability distribution between scene and action using a naive Bayes nearest neighbour algorithm, which is adopted to jointly infer the action categories online by combining off‐the‐shelf action recognition algorithms. The authors demonstrate the advantages of their approach by comparing it with state‐of‐the‐art approaches using several action recognition benchmarks.
Hongbo Zhang 0002, Duansheng Chen, Bineng Zhong 0001, Jialin Peng, Jixiang Du, Songzhi Su
IET Comput. Vis.7
2016 CBDF: Compressed Binary Discriminative Feature
Li-Chuan Geng, Pierre-Marc Jodoin, Songzhi Su, Shaozi Li
Neurocomputing3
2016 Detection based object labeling of 3D point cloud for indoor scenes
Wei Liu 0005, Shaozi Li, Donglin Cao, Songzhi Su, Rongrong Ji
Neurocomputing4
2016 Decomposed human localization from social photo album
Shaozi Li, Songzhi Su, Bing Shuai, Rongrong Ji
Multim. Syst.3
2016 Fast verification via statistical geometric for mobile visual search
Shaozi Li, Xianming Lin, Songzhi Su, Rongrong Ji
Multim. Syst.4
2016 Multi-view fall detection based on spatio-temporal interest points
Songzhi Su, Sin-Sian Wu, Shu-Yuan Chen, Der-Jyh Duh, Shaozi Li
Multim. Tools Appl.1
2015 Traffic analysis without motion features
abstract
In this paper, we investigate the possibility of monitoring traffic without using any motion features. The goal of our system is to process videos with ultra-low frame rate, i.e. videos for which reliable motion features cannot be computed. In this work, we investigate how 2D spatial features combined with a machine learning method can assess traffic conditions such as fluid traffic, dense traffic, and traffic jam. The underlying hypothesis that we ought to validate is that traffic images are heavily characterized by their 2D spatial textures. In that perspective, we tested different 2D texture features and machine learning methods to see how accurate such an approach can be. We also performed a regression on the image descriptor in order to estimate traffic density. Experimental results obtained on the UCSD traffic dataset reveal that our approach generalizes well to various weather and lighting conditions. It even outperforms state-of-the-art traffic analysis methods relying on spatio-temporal features.
Zhiming Luo, Pierre-Marc Jodoin, Shaozi Li, Songzhi Su
ICIP4
2015 Feature learning based on SAE-PCA network for human gesture recognition in RGBD images
Shaozi Li, Wei Wu 0072, Songzhi Su, Rongrong Ji
Neurocomputing4
2015 Sparse auto-encoder based feature learning for human body detection in depth image
Songzhi Su, Zhi-Hui Liu, Suping Xu, Shaozi Li, Rongrong Ji
Signal Process.1
2014 Lying-pose detection with training dataset expansion
abstract
We propose a rotation and scale invariant method to locate people lying on the ground. Unlike conventional human-shape detection methods which assume that all human shapes are in upright position, a person lying on the ground can have arbitrary orientation and pose. Accounting for every possible body configuration would thus require a huge training dataset that would be challenging to gather. In this paper, we propose a method which increases the size of a small training dataset and allows to detect multiple body poses. To do so, our method increases the size of the dataset with a geometric distortion method followed by a rejection sampling method. Then, it automatically identifies K body configurations in the training set, realign it in upright position and trains K SVM classifiers, one for each body configuration. Lying pose detection is then performed by considering a max pooling strategy across all K SVM classifiers.
Daoxun Xia, Songzhi Su, Shaozi Li, Pierre-Marc Jodoin
ICIP2
2014 Pursuing Detector Efficiency for Simple Scene Pedestrian Detection
De-Dong Yuan, Songzhi Su, Shaozi Li, Rongrong Ji
MMM (2)3
2014 Kinship classification based on discriminative facial patches
abstract
Recently there has been a large explosive growth of image data on social networks and how to use computer vision and machine learning technology to verify people relationships on these huge amount of human-centered image data remains a challenging issue. Remarkably, there have been few research attempts to analyze the possible human relationships on images, especially kin relationships. In this paper, we tackle a challenging and relatively new issue in kinship classification: determining the family that a query face image belongs to. To address this challenge, we propose a kinship classification method in three steps: (l)Discriminative patches are detected automatically in the facial landmark regions. (2) Appearance features, Histogram of Gradient (HOG), Scale-Invariant Feature Transform (SIFT) and Four-Patch Local Binary Pattern (FPLBP) are extracted from these patches respectively, and then we concatenate the features to create a high-dimensional feature vector. (3) Linear Support Vector Machine (SVM) with polynomial kernel is adopted to accomplish kinship classification task. Experimental evaluation results on Cornell Family 101 dataset demonstrate that our proposed method significantly outperforms the state-of-the-art kinship classification approaches.
Songzhi Su, Shaozi Li
VCIP3
2014 Online MIL tracking with instance-level semi-supervised learning
Si Chen 0002, Shaozi Li, Songzhi Su, Qi Tian 0001, Rongrong Ji
Neurocomputing3
2014 Perspective-Invariant Image Matching Framework with Binary Feature Descriptor and APSO
abstract
A novel perspective invariant image matching framework is proposed in this paper, noted as Perspective-Invariant Binary Robust Independent Elementary Features (PBRIEF). First, we use the homographic transformation to simulate the distortion between two corresponding patches around the feature points. Then, binary descriptors are constructed by comparing the intensity of sample points surrounding the feature location. We transform the location of the sample points with simulated homographic matrices. This operation is to ensure that the intensities which we compared are the realistic corresponding pixels between two image patches. Since the exact perspective transform matrix is unknown, an Adaptive Particle Swarm Optimization (APSO) algorithm-based iterative procedure is proposed to estimate the real transformation angles. Experimental results obtained on five different datasets show that PBRIEF outperforms significantly the existing methods on images with large viewpoint difference. Moreover, the efficiency of our framework is also improved comparing with Affine-Scale Invariant Feature Transform (ASIFT).
Li-Chuan Geng, Songzhi Su, Donglin Cao, Shaozi Li
Int. J. Pattern Recognit. Artif. Intell.2
2014 Online semi-supervised compressive coding for robust visual tracking
Si Chen 0002, Shaozi Li, Songzhi Su, Donglin Cao, Rongrong Ji
J. Vis. Commun. Image Represent.3
2014 Logo detection with extendibility and discrimination
Kuo-Wei Li, Shu-Yuan Chen, Songzhi Su, Der-Jyh Duh, Hongbo Zhang 0002, Shaozi Li
Multim. Tools Appl.3
2014 Adaptive photograph retrieval method
Hongbo Zhang 0002, Shang-An Li, Shu-Yuan Chen, Songzhi Su, Der-Jyh Duh, Shaozi Li
Multim. Tools Appl.4
2013 Saliency detection by adaptive clustering
abstract
Saliency detection plays an important role in image segmentation, content-aware resizing and object recognition. Most approaches obtain promising performance recently, which is useful for the postprocessing. We propose a clustering-based method to detect refined regions with comparative performance. For coarse-grained classification with unknown clusters number, an adaptive algorithm called f-means is developed in this paper. Pixels are clustered by f-means based on color and spatial features, and then the centroids are used to compute their saliency values. Experiments show that our algorithm generates more fine maps, which outperform the state-of-the-art approaches on MSRA dataset. Relying on the saliency map, we also get superior results in foreground extracting, image resizing and thumbnails generation.
Hai Cao, Shaozi Li, Songzhi Su, Rongrong Ji
VCIP3
2013 A new camera self-calibration method based on CSA
abstract
A large number of computer vision applications rely on camera calibration. Camera self-calibration which only depends on the relationship between corresponding points of a pair of images draws much attention for its simplicity. Almost all the camera self-calibration methods rely on the solution of Kruppa equations which are difficult to be directly solved. The state-of-the-art self-calibration algorithms usually convert the solution of these equations to non-linear optimization problem, traditional optimization methods usually have the drawback of convergent to local extreme. Artificial immune system (AIS) has the ability to fast convergent to global extreme. To address this problem, we proposed an artificial immune system based method which can fast convergent to the global optimization solutions. We demonstrate the performance of the proposed method with synthetic and real data.
Li-Chuan Geng, Shaozi Li, Songzhi Su, Donglin Cao, Rongrong Ji
VCIP3
2013 Decomposed human localization in personal photo albums
abstract
Recent years have seen tremendous progress in human detection, whereas only upright poses are usually considered. In this paper, we relax this constraint to localizing highly deformable persons, as commonly exhibited in personal photo albums. Human localization based on arbitrary pose is extremely challenging, due to the large pose variances, disabling the traditional part based template detectors. To tackle this issue, we propose a decomposition-based human localization model dealing with this issue in three-step: a stable upper-body is firstly detected, then a set of bigger bounding boxes are extended, from which the most appropriate instance is distinguished by a discriminative Whole Person Model. The experiment results demonstrated that our decomposition-based model worked very well at localizing deformable persons, which boosted the average precision by 10% compared to state-of-the-art person detectors. On the other hand, Similar Pose Feature(SPF) provides the feasibility of projecting persons with similar poses into same clusters, facilitating a novel pose-based photo album browsing functionality.
Bing Shuai, Songzhi Su, Shaozi Li, Rongrong Ji
VCIP2
2013 Seeing actions through scene context
abstract
Recognizing human actions is not alone, as hinted by the scene herein. In this paper, we investigate the possibility to boost the action recognition performance by exploiting their scene context associated. To this end, we model the scene as a mid-level “hidden layer” to bridge action descriptors and action categories. This is achieved via a scene topic model, in which hybrid visual descriptors including spatiotemporal action features and scene descriptors are first extracted from the video sequence. Then, we learn a joint probability distribution between scene and action by a Naive Bayesian N-earest Neighbor algorithm, which is adopted to jointly infer the action categories online by combining off-the-shelf action recognition algorithms. We demonstrate our merits by comparing to state-of-the-arts in several action recognition benchmarks.
Hongbo Zhang 0002, Songzhi Su, Shaozi Li, Duansheng Chen, Bineng Zhong 0001, Rongrong Ji
VCIP2
2013 Perspective-SIFT: An efficient tool for low-altitude remote sensing image registration
Guo-Rong Cai, Pierre-Marc Jodoin, Shaozi Li, Yun-Dong Wu, Songzhi Su, Zhenkun Huang
Signal Process.5