VLDB 2026 Research / reviewers in the wild / expert
Hongbo Zhang 0002
dblp:24/3333-2 · also Hong-Bo Zhang 0002
· DBLP profile ↗
51ranked-venue papers
9as first author
36since 2021 · last 2026
0000-0001-5536-5224ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 5 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 12 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Skeleton-Guided Spatio-Temporal Video Representation for Long-Term Action Quality AssessmentabstractLong-term Action Quality Assessment (AQA) is of significant importance in applications such as sports analysis and medical rehabilitation. Compared with short-term AQA, long-term AQA involves longer temporal spans and more complex motion structures, posing greater challenges to spatio-temporal modeling. Existing methods typically rely on a single modality, either video or skeleton data. Video-based approaches are easily affected by background noise, while skeleton-based methods lack visual appearance information, making it difficult to comprehensively represent action quality. To address these limitations, we propose a multimodal framework termed Skeleton-Guided Spatio-Temporal Video Representation (SG-STVR). By introducing skeletal priors, we develop a Spatial Prior Injection (SPI) strategy that utilizes human foreground masks generated from skeleton data to guide the model to focus on action-relevant regions and suppress background interference. In addition, we propose a Temporal Prior Guidance (TPG) strategy, which enhances the temporal modeling capability of video features based on the motion sensitivity of skeleton sequences. Experimental results on the RG and Fis-V datasets demonstrate that SG-STVR outperforms other approaches, reflecting its competitiveness in long-term AQA task. Zhiwei Hong, Hongbo Zhang 0002, Jixiang Du |
ICMR | 3 |
| 2026 | Cascaded-parallel decoders and anchor-guided query generator for human-object interaction
Hongbo Zhang 0002, Jia-Yin Luo, Zhen-Zhen Sun, Jixiang Du |
Comput. Vis. Image Underst. | 2 |
| 2026 | Partial label feature selection with dynamic streaming labels
Tianlang Li, Hongbo Zhang 0002, Zhenzhen Sun, Jia Zhong, Jin Gou |
Pattern Recognit. | 3 |
| 2026 | Learnable token for visual tracking
Yan Chen 0017, Zhongkang Jiang, Jixiang Du, Hongbo Zhang 0002 |
Signal Process. Image Commun. | 4 |
| 2026 | Consensus Labeling: Prompt-Guided Clustering Refinement for Weakly Supervised Text-Based Person Re-IdentificationabstractWeakly supervised text-based person re-identification aims to retrieve specific pedestrians based on textual descriptions without identity labels available during training. This task remains challenging due to the inherent cross-modal heterogeneity and lack of identity annotations. There is a common issue of modality gap in vision language models, which in turn affects the performance of downstream tasks such as cross-modal retrieval and multimodal clustering. Specifically, in our research and experiments, we found that there is a problem of inter-modal misalignment between image and text modalities. However, existing methods rely on mutual enhancement strategies between image and text clustering, leading to the accumulation of clustering noise and affecting the final retrieval performance. To address this issue, we propose a Consensus Labelling: Prompt-guided Clustering refinement (CLPC) framework for weakly supervised text-based person re-identification. Specifically, we introduce a textual inversion network to learn a pseudo token that captures visual context, which is then integrated into natural language sentences as personalized textual prompt. To further improve clustering quality, we introduce a Nearest Neighbor-Guided Pseudo Label Mining (NGPM) method, which uses the clusters derived from personalized textual prompts to refine the clustering of image features. Additionally, we design a Dynamic Margin Triplet (DMT) loss, where the margin is adaptively adjusted using a sigmoid-based function to enhance the model’s ability to distinguish hard negative samples. We have also introduce a Normalized Distribution Matching (NDM) loss to minimize the KL divergence between the image-text matching scores and the normalized soft matching scores. The extensive experimental results on three public datasets have demonstrated the superiority of our method. Our code is available at https://github.com/LeviWeiZhi/CLPC. Chengji Wang, Weizhi Nie, Hongbo Zhang 0002, Hao Sun 0014, Mang Ye |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | LSPEL: Label-Specific Feature-Based Partial Label Learning for Emerging New LabelsabstractIn partial label learning (PLL) tasks, each training instance is assigned a candidate label set, with only one label being correct. Previous studies on PLL have focused on scenarios where the class label set remains fixed, i.e., the label set for test data is the same as that used during training. However, in many real-world applications, the environment is dynamic, and new labels may emerge, requiring methods that can detect and classify these new labels. Moreover, previous methods typically learn from partial label data by manipulating the same feature set, which may be suboptimal as it overlooks the semantic relationships between instances and labels. To this end, we develop a novel PLL approach called Label-Specific feature-based Partial label learning with Emerging new Labels (LSPEL), which works by iteratively learning label-specific features during the label disambiguation process to support new label detection and model update. It consists of three key components: (1) model training based on label-specific feature learning, (2) construction of a new label detector that works in conjunction with the classifier to predict known labels, and (3) model updating and induction to further enhance the prediction results for known labels. Extensive experiments on synthetic and real-world PL datasets demonstrate that LSPEL is effective in handling emerging new labels. Hongbo Zhang 0002, Jin Gou, Yaojin Lin |
ACM Trans. Knowl. Discov. Data | 3 |
| 2025 | Contrastive Single-Stream Spatio-Temporal Joint Modeling for Few-Shot Action RecognitionabstractPrior work on few-shot action recognition predominantly adopts two strategies: spatio-temporal separated frame matching and multi-stream multi-modal networks. However, each suffering from either incomplete spatio-temporal modeling or an over-reliance on additional annotation data. To address these limitations, we propose a Contrastive Single-Stream Spatio-Temporal joint modeling Few-Shot Action Recognition (CS3T-FSAR) model. In terms of spatio-temporal modeling, our approach directly constructs high-quality three-dimensional spatio-temporal representations to fully capture the global associations among video frames. Regarding the loss function design, we integrate a triplet loss to achieve precise matching while reducing both inference cost and computational complexity. Ultimately, our method achieves significant performance improvements across four benchmark datasets, demonstrating its competitiveness in few-shot action recognition. Xingyang Xu, Jixiang Du, Jing Wang 0049, Hongbo Zhang 0002, Lijing Ye, Jiayu Xiong |
ICMR | 4 |
| 2025 | IFFN: Irrelevant Frames Filtering Network for Online Action DetectionabstractOnline Action Detection (OAD) is a fundamental task in video action understanding, aiming to predict ongoing actions in real-time from untrimmed video streams. However, existing models face difficulties in distinguishing feature differences between action frames and neutral transition frames. To address this challenge, we propose the Irrelevant Frames Filtering Network (IFFN), a dual-branch architecture comprising an action branch and an irrelevant frames filtering branch. The IFFN employs an Irrelevant Frame Filtering Unit (IFFU) to extract informative foreground frames from streaming videos. To further enhance the model’s discriminative ability, we introduce a novel training strategy that utilizes opposing training objectives and weight-sharing mechanisms to better distinguish neutral transition frames from foreground frames. We evaluate IFFN on three benchmark datasets: the THUMOS’14, TVSeries, and MultiTHUMOS. Experimental results demonstrate that IFFN outperforms all models used for comparison. Furthermore, we integrate IFFN into OadTR, LSTR, and MiniROAD, which are representative models based on Transformer and RNN architectures. Experiments demonstrate that IFFN can serve as a plug-and-play module to consistently enhance the performance of OAD models. Ming-Xuan Lin, Hongbo Zhang 0002, Bo-Sheng Zheng, Zhenzhen Sun |
MMAsia | 2 |
| 2025 | Interaction Confidence Attention for Human-Object Interaction Detection
Hongbo Zhang 0002, Wang-Kai Lin, Jixiang Du |
Int. J. Comput. Vis. | 1 |
| 2025 | Learning referee evaluation and assessing action quality from coarse to fine in diving sport
Hong-Ming Qiu, Hongbo Zhang 0002, Jixiang Du |
Neurocomputing | 2 |
| 2025 | Object tracking based on temporal and spatial context information
Yan Chen 0017, Jixiang Du, Hongbo Zhang 0002 |
Image Vis. Comput. | 4 |
| 2025 | Skeletal spatio-temporal decoupling transformer for long-duration action quality assessment
Long Yao, Hongbo Zhang 0002, Jixiang Du |
Knowl. Based Syst. | 3 |
| 2025 | Multi-stage query-based feature generating and encoding for robust early action recognition
Wei-Xiang Pan, Hongbo Zhang 0002, Ming-Xuan Lin |
Vis. Comput. | 3 |
| 2024 | Spatial and temporal consistency learning for monocular 6D pose estimation
Hongbo Zhang 0002, Jia-Yu Liang, Jia-Xin Hong, Jixiang Du |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Semantics feature sampling for point-based 3D object detection
Jing-Dong Huang, Jixiang Du, Hongbo Zhang 0002, Huai-Jin Liu |
Image Vis. Comput. | 3 |
| 2024 | PVConvNet: Pixel-Voxel Sparse Convolution for multimodal 3D object detection
Huaijin Liu, Jixiang Du, Yong Zhang 0066, Hongbo Zhang 0002, Jiandian Zeng |
Pattern Recognit. | 4 |
| 2024 | Learning implicit labeling-importance and label correlation for multi-label feature selection with streaming labels
Yaojin Lin, Lijie Yang 0001, Hongbo Zhang 0002 |
Pattern Recognit. | 5 |
| 2024 | MSSA: Multi-Representation Semantics-Augmented Set Abstraction for 3D Object DetectionabstractAccurate recognition and localization of 3D objects is a fundamental research problem in 3D computer vision. Benefiting from transformation-free point cloud processing and flexible receptive fields, point-based methods have become accurate in 3D point cloud modeling, but still fall behind voxel-based competitors in 3D detection. We observe that the set abstraction module, commonly utilized by point-based methods for downsampling points, tends to retain excessive irrelevant background information, thus hindering the effective learning of features for object detection tasks. To address this issue, we propose MSSA, a Multi-representation Semantics-augmented Set Abstraction for 3D object detection. Specifically, we first design a backbone network to encode different representation features of point clouds, which extracts point-wise features through PointNet to preserve fine-grained geometric structure features, and adopts VoxelNet to extract voxel features and BEV features to enhance the semantic features of key points. Second, to efficiently fuse different representation features of keypoints, we propose a Point feature-guided Voxel feature and BEV feature fusion (PVB-Fusion) module to adaptively fuse multi-representation features and remove noise. At last, a novel Multi-representation Semantic-guided Farthest Point Sampling (MS-FPS) algorithm is designed to help set abstraction modules progressively downsample point clouds, thereby improving instance recall and detection performance with more important foreground points. We evaluate MSSA on the widely used KITTI dataset and the more challenging nuScenes dataset. Experimental results show that compared to PointRCNN, our method improves the AP of “moderate” level for three classes of objects by 7.02%, 6.76%, and 5.44%, respectively. Compared to the advanced point-voxel-based method PV-RCNN, our method improves the AP of “moderate” level by 1.23%, 2.84%, and 0.55% for the three classes, respectively. Huaijin Liu, Jixiang Du, Yong Zhang 0066, Hongbo Zhang 0002, Jiandian Zeng |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Multi-skeleton structures graph convolutional network for action quality assessment in long videos
Hongbo Zhang 0002, Jixiang Du, Shangce Gao |
Appl. Intell. | 3 |
| 2023 | ASFS: A novel streaming feature selection for multi-label data based on neighborhood rough set
Yaojin Lin, Jixiang Du, Hongbo Zhang 0002, Ziyi Chen 0001, Jia Zhang 0019 |
Appl. Intell. | 4 |
| 2023 | Label-reconstruction-based pseudo-subscore learning for action quality assessment in sporting events
Hongbo Zhang 0002, Li-Jia Dong, Lijie Yang 0001, Jixiang Du |
Appl. Intell. | 1 |
| 2023 | Extracting geometric and semantic point cloud features with gateway attention for accurate 3D object detection
Huaijin Liu, Jixiang Du, Yong Zhang 0066, Hongbo Zhang 0002 |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Multi-label feature selection based on label distribution and neighborhood rough set
Yaojin Lin, Weiping Ding 0001, Hongbo Zhang 0002, Cheng Wang 0003, Jixiang Du |
Neurocomputing | 4 |
| 2023 | Fuzzy Mutual Information-Based Multilabel Feature Selection With Label Dependency and Streaming LabelsabstractMultilabel feature selection (MFS) has received widespread attention in various big data applications. However, most of the existing methods either explicitly or implicitly assume that all labels are given in advance before feature selection starts; or that all labels are independent. In fact, in many practical applications, the available labels usually arrive dynamically, and they may be interdependent with each other. Moreover, labels may be generated dynamically in a minibatch manner, which makes it more difficult to explore label dependency. In this article, we propose a novel fuzzy mutual information-based multilabel feature selection approach MSDS, which is able to solve single streaming label, minibatch streaming labels, and exploit label dependency simultaneously. In specific, we first promote fuzzy mutual information to be suitable for multilabel learning. This model can effectively consider the relationship between two labels, and has good applicability for measuring the relationship between multiple labels. Then, we analyze feature relevance and feature redundancy based on the combination of label dependency and streaming labels, which helps to facilitate the selection of high-quality feature subsets. Finally, a feature conversion is designed to fuse the representative features of new arrival streaming labels. Comprehensive experiments on twelve multilabel datasets clearly reveal the superiority of the proposed method against two streaming labels based algorithms and five state-of-the-art static label space based algorithms. Yaojin Lin, Weiping Ding 0001, Hongbo Zhang 0002, Jixiang Du |
IEEE Trans. Fuzzy Syst. | 4 |
| 2023 | Point-Based Learnable Query Generator for Human-Object Interaction DetectionabstractTransformer-based and interaction point-based methods have demonstrated promising performance and potential in human-object interaction detection. However, due to differences in structure and properties, direct integration of these two types of models is not feasible. Recent Transformer-based methods divide the decoder into two branches: an instance decoder for human-object pair detection and a classification decoder for interaction recognition. While the attention mechanism within the Transformer enhances the connection between localization and classification, this paper focuses on further improving HOI detection performance by increasing the intrinsic correlation between instance and action features. To address these challenges, this paper proposes a novel Transformer-based HOI Detection framework. In the proposed method, the decoder contains three parts: learnable query generator, instance decoder, and interaction classifier. The learnable query generator aims to build an effective query to guide the instance decoder and interaction classifier to learn more accurate instance and interaction features. These features are then applied to update the query generator for the next layer. Especially, inspired by the interaction point-based HOI and object detection methods, this paper introduces the prior bounding boxes, keypoints detection and spatial relation feature to build the novel learnable query generator. Finally, the proposed method is verified on HICO-DET and V-COCO datasets. The experimental results show that the proposed method has the better performance compared with the state-of-the-art methods. Wang-Kai Lin, Hongbo Zhang 0002, Zongwen Fan, Lijie Yang 0001, Jixiang Du |
IEEE Trans. Image Process. | 2 |
| 2023 | Effective skeleton topology and semantics-guided adaptive graph convolution network for action recognition
Zhong-Xiang Qiu, Hongbo Zhang 0002, Wei-Mo Deng, Jixiang Du |
Vis. Comput. | 2 |
| 2022 | Pairwise Contrastive Learning Network for Action Quality Assessment
Hongbo Zhang 0002, Zongwen Fan, Jixiang Du |
ECCV (4) | 2 |
| 2022 | Cross-scale feature fusion connection for a YOLO detectorabstractAbstract Multi‐scale feature fusion is often used to address the issue of scale variations in object detection. However, most of the proposed network architectures only combine the features of two adjacent levels sequentially, so the first fusion nodes in both top‐down and bottom‐up pathways must be blank nodes that only have one input with no feature fusion. In this work, cross‐scale feature fusion connection (CFFC) is proposed which aims to enhance the entire feature hierarchy by propagating the features of each level more efficiently. The proposed method reuses and aggregates all the features of other scales to the blank nodes in both top‐down and bottom‐up pathways. Furthermore, the authors remove the 1 × 1 convolutional layer and replace the shortcut with concatenation before fusing multiple features. These concatenated feature maps are then supervised by the channel attention block at the fusion nodes. This modification allows the network to learn the important degree of each level in concatenated feature maps along the channel dimension. It is also observed that the proposed method alleviates the inconsistency in feature pyramids with fewer parameters. The performance of a YOLO object detector equipped with the proposed method on the COCO test‐dev 2017 is evaluated. The results show that the proposed method outperforms other architectures presented in the literature. Zhongling Ruan, Jianzhong Cao, Hongbo Zhang 0002 |
IET Comput. Vis. | 4 |
| 2022 | Late feature supplement network for early action prediction
Hongbo Zhang 0002, Miao-Hui Zhang, Jixiang Du |
Image Vis. Comput. | 2 |
| 2022 | Multi-feature fusion refine network for video captioningabstractDescribing video content using natural language is an important part of video understanding. It needs to not only understand the spatial information on video, but also capture the motion information. Meanwhile, video captioning is a cross-modal problem between vision and language. Traditional video captioning methods follow the encoder-decoder framework that transfers the video to sentence. But the semantic alignment from sentence to video is ignored. Hence, finding a discriminative visual representation as well as narrowing the semantic gap between video and text has great influence on generating accurate sentences. In this paper, we propose an approach based on multi-feature fusion refine network (MFRN), which can not only capture the spatial information and motion information by exploiting multi-feature fusion, but also can get better semantic aligning of different models by designing a refiner to explore the sentence to video stream. The main novelties and advantages of our method are: (1) multi-feature fusion: Both two-dimension convolutional neural networks and three-dimension convolutional neural networks pre-trained on ImageNet and Kinetic respectively are used to construct spatial information and motion information, and then fused to get better visual representation. (2) Sematic alignment refiner: the refiner is designed to restrain the decoder and reproduce the video features to narrow semantic gap between different modal. Experiments on two widely used datasets demonstrate our approach achieves state-of-the-art performance in terms of BLEU@4, METEOR, ROUGE and CIDEr metrics. Guan-Hong Wang, Jixiang Du, Hongbo Zhang 0002 |
J. Exp. Theor. Artif. Intell. | 3 |
| 2022 | Improved human-object interaction detection through skeleton-object relationsabstractCurrent methods for human-object interaction detection often use the spatial relation between a human and an object as an interaction pattern. However, this strategy is relatively simple and has low discrimination in similar interactions. To solve this drawback, the spatial relation between skeletons and objects is proposed to model the interaction pattern and improve the detection accuracy. First, the skeleton-object interaction pattern image is extracted for each interaction proposal. Second, a deep neural network is applied to learn the interaction features from these images. Finally, the interaction feature is added to the human-object interaction detection network by a multistream structure. In the experiments, we evaluate the proposed method on the HICO-DET and V-COCO datasets. Experimental results show that the proposed method can achieve the best performance compared with state-of-art methods. Hongbo Zhang 0002, Yi-Zhong Zhou, Jixiang Du, Jin-Long Huang, Lijie Yang 0001 |
J. Exp. Theor. Artif. Intell. | 1 |
| 2022 | Skeleton-based deep pose feature learning for action quality assessment on figure skating videos
Hongbo Zhang 0002, Jixiang Du, Shangce Gao |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | Pose attention and object semantic representation-based human-object interaction detection network
Wei-Mo Deng, Hongbo Zhang 0002, Jixiang Du, Min Huang 0004 |
Multim. Tools Appl. | 2 |
| 2022 | Shuffle-invariant Network for Action Recognition in VideosabstractThe local key features in video are important for improving the accuracy of human action recognition. However, most end-to-end methods focus on global feature learning from videos, while few works consider the enhancement of the local information in a feature. In this article, we discuss how to automatically enhance the ability to discriminate the local information in an action feature and improve the accuracy of action recognition. To address these problems, we assume that the critical level of each region for the action recognition task is different and will not change with the region location shuffle. We therefore propose a novel action recognition method called the shuffle-invariant network. In the proposed method, the shuffled video is generated by regular region cutting and random confusion to enhance the input data. The proposed network adopts the multitask framework, which includes one feature backbone network and three task branches: local critical feature shuffle-invariant learning, adversarial learning, and an action classification network. To enhance the local features, the feature response of each region is predicted by a local critical feature learning network. To train this network, an L 1-based critical feature shuffle-invariant loss is defined to ensure that the ordered feature response list of these regions remains unchanged after region location shuffle. Then, the adversarial learning is applied to eliminate the noise caused by the region shuffle. Finally, the action classification network combines these two tasks to jointly guide the training of the feature backbone network and obtain more effective action features. In the testing phase, only the action classification network is applied to identify the action category of the input video. We verify the proposed method on the HMDB51 and UCF101 action datasets. Several ablation experiments are constructed to verify the effectiveness of each module. The experimental results show that our approach achieves the state-of-the-art performance. Qinghongya Shi, Hongbo Zhang 0002, Jixiang Du |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Tile selection method based on error minimization for photomosaic image creation
Hongbo Zhang 0002, Jixiang Du, Lijie Yang 0001 |
Frontiers Comput. Sci. | 1 |
| 2021 | Learning and fusing multiple hidden substages for action quality assessment
Li-Jia Dong, Hongbo Zhang 0002, Qinghongya Shi, Jixiang Du, Shangce Gao |
Knowl. Based Syst. | 2 |
| 2020 | PON: Proposal Optimization Network for Temporal Action Proposal Generation
Xiao-Xiao Peng, Jixiang Du, Hongbo Zhang 0002 |
ICIC (3) | 3 |
| 2020 | Brushwork master: Chinese ink painting synthesis for animating brushwork processabstractAbstract Generally, it is regarded as challenge work to grasp the drawing style of an ancient masterpiece in Chinese painting learning. This paper presents a novel approach to the generation of a Chinese ink painting in a certain style and animating its brushwork process with expert skills. In order to demonstrate the techniques of brush and ink inside a stroke, a serials of geometric properties of a brush stroke, are first extracted, then through rational deformation calculation, the best stroke source is mapped onto the stroking path, which is sketched by the user, and finally a new Chinese painting can be synthesized by style migration and natural stroke composition. So with the generated strokes, the lifelike brushwork process of the new painting can be represented dramatically. Actually, by showing the authentic painting process, the tool we implemented helps the learners, who have no profound skills and knowledge in domain of Chinese painting, master the essence of a great painting style, and also provides an easy way to art creation and the comprehension of mysterious Chinese traditional art. Lijie Yang 0001, Tianchen Xu, Jixiang Du, Hongbo Zhang 0002, Enhua Wu |
Comput. Animat. Virtual Worlds | 4 |
| 2019 | Periodic Action Temporal Localization Method Based on Two-Path Architecture for Product Counting in Sewing Video
Jin-Long Huang, Hongbo Zhang 0002, Jixiang Du, Xiao-Xiao Peng |
ICIC (3) | 2 |
| 2018 | Fine-Grained Recognition of Vegetable Images Based on Multi-scale Convolution Neural Network
Xiu-Hong Yang, Jixiang Du, Hongbo Zhang 0002, Wentao Fan 0001 |
ICIC (2) | 3 |
| 2018 | Discriminative parts learning for 3D human action recognition
Min Huang 0004, Guo-Rong Cai, Hongbo Zhang 0002, Sheng Yu 0007, Dong-Ying Gong, Donglin Cao, Shaozi Li, Songzhi Su |
Neurocomputing | 3 |
| 2018 | A hierarchical representation for human action recognition in realistic scenes
Hongbo Zhang 0002, Minghai Xin, Yiqiao Cai |
Multim. Tools Appl. | 2 |
| 2018 | Multifeature Selection for 3D Human Action RecognitionabstractIn mainstream approaches for 3D human action recognition, depth and skeleton features are combined to improve recognition accuracy. However, this strategy results in high feature dimensions and low discrimination due to redundant feature vectors. To solve this drawback, a multi-feature selection approach for 3D human action recognition is proposed in this paper. First, three novel single-modal features are proposed to describe depth appearance, depth motion, and skeleton motion. Second, a classification entropy of random forest is used to evaluate the discrimination of the depth appearance based features. Finally, one of the three features is selected to recognize the sample according to the discrimination evaluation. Experimental results show that the proposed multi-feature selection approach significantly outperforms other approaches based on single-modal feature and feature fusion. Min Huang 0004, Songzhi Su, Hongbo Zhang 0002, Guo-Rong Cai, Dong-Ying Gong, Donglin Cao, Shaozi Li |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2017 | Meta-action descriptor for action recognition in RGBD videoabstractAction recognition is one of the hottest research topics in computer vision. Recent methods represent actions based on global or local video features. These approaches, however, lack semantic structure and may not provide a deep insight into the essence of an action. In this work, the authors argue that semantic clues, such as joint positions and part‐level motion clustering, help verify actions. To this end, a meta‐action descriptor for action recognition in RGBD video is proposed in this study. Specifically, two discrimination‐based strategies – dynamic and discriminative part clustering – are introduced to improve accuracy. Experiments conducted on the MSR Action 3D dataset show that the proposed method significantly outperforms the methods without joint position semantic. Min Huang 0004, Songzhi Su, Guo-Rong Cai, Hongbo Zhang 0002, Donglin Cao, Shaozi Li |
IET Comput. Vis. | 4 |
| 2017 | Sparse Representation-Based Semi-Supervised Regression for People CountingabstractLabel imbalance and the insufficiency of labeled training samples are major obstacles in most methods for counting people in images or videos. In this work, a sparse representation-based semi-supervised regression method is proposed to count people in images with limited data. The basic idea is to predict the unlabeled training data, select reliable samples to expand the labeled training set, and retrain the regression model. In the algorithm, the initial regression model, which is learned from the labeled training data, is used to predict the number of people in the unlabeled training dataset. Then, the unlabeled training samples are regarded as an over-complete dictionary. Each feature of the labeled training data can be expressed as a sparse linear approximation of the unlabeled data. In turn, the labels of the labeled training data can be estimated based on a sparse reconstruction in feature space. The label confidence in labeling an unlabeled sample is estimated by calculating the reconstruction error. The training set is updated by selecting unlabeled samples with minimal reconstruction errors, and the regression model is retrained on the new training set. A co-training style method is applied during the training process. The experimental results demonstrate that the proposed method has a low mean square error and mean absolute error compared with those of state-of-the-art people-counting benchmarks. Hongbo Zhang 0002, Bineng Zhong 0001, Jixiang Du, Jialin Peng, Duansheng Chen, Xiao Ke |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2016 | Probability-based method for boosting human action recognition using scene contextabstractIn this study, the authors investigate the possibility of boosting action recognition performance by exploiting the associated scene context. Towards this end, the authors model a scene as a mid‐level ‘middle layer’ in order to bridge action descriptors and action categories. This is achieved via a scene topic model, in which hybrid visual descriptors, including spatial–temporal action features and scene descriptors, are first extracted from a video sequence. Then, the authors learn a joint probability distribution between scene and action using a naive Bayes nearest neighbour algorithm, which is adopted to jointly infer the action categories online by combining off‐the‐shelf action recognition algorithms. The authors demonstrate the advantages of their approach by comparing it with state‐of‐the‐art approaches using several action recognition benchmarks. Hongbo Zhang 0002, Duansheng Chen, Bineng Zhong 0001, Jialin Peng, Jixiang Du, Songzhi Su |
IET Comput. Vis. | 1 |
| 2016 | Higher order partial least squares for object tracking: A 4D-tracking method
Bineng Zhong 0001, Xiangnan Yang, Yingju Shen, Cheng Wang 0020, Tian Wang 0001, Zhen Cui 0001, Hongbo Zhang 0002, Xiaopeng Hong, Duansheng Chen |
Neurocomputing | 7 |
| 2015 | Online learning 3D context for robust visual tracking
Bineng Zhong 0001, Yingju Shen, Yan Chen 0017, Weibo Xie, Zhen Cui 0001, Hongbo Zhang 0002, Duansheng Chen, Tian Wang 0001, Xin Liu 0011, Shu-Juan Peng, Jin Gou, Jixiang Du, Jing Wang 0049, Wenming Zheng |
Neurocomputing | 6 |
| 2014 | Logo detection with extendibility and discrimination
Kuo-Wei Li, Shu-Yuan Chen, Songzhi Su, Der-Jyh Duh, Hongbo Zhang 0002, Shaozi Li |
Multim. Tools Appl. | 5 |
| 2014 | Adaptive photograph retrieval method
Hongbo Zhang 0002, Shang-An Li, Shu-Yuan Chen, Songzhi Su, Der-Jyh Duh, Shaozi Li |
Multim. Tools Appl. | 1 |
| 2013 | Seeing actions through scene contextabstractRecognizing human actions is not alone, as hinted by the scene herein. In this paper, we investigate the possibility to boost the action recognition performance by exploiting their scene context associated. To this end, we model the scene as a mid-level “hidden layer” to bridge action descriptors and action categories. This is achieved via a scene topic model, in which hybrid visual descriptors including spatiotemporal action features and scene descriptors are first extracted from the video sequence. Then, we learn a joint probability distribution between scene and action by a Naive Bayesian N-earest Neighbor algorithm, which is adopted to jointly infer the action categories online by combining off-the-shelf action recognition algorithms. We demonstrate our merits by comparing to state-of-the-arts in several action recognition benchmarks. Hongbo Zhang 0002, Songzhi Su, Shaozi Li, Duansheng Chen, Bineng Zhong 0001, Rongrong Ji |
VCIP | 1 |