Hua Zhang 0003

dblp:69/2745-3 · DBLP profile ↗
← Back
49ranked-venue papers
0as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 6 since 2021Artificial intelligence and machine learning · 14 · 1 since 2021Computer networks · 6Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 A Snippets Relation and Hard-Snippets Mask Network for Weakly-Supervised Temporal Action Localization
abstract
Weakly-supervised temporal action localization (WTAL) is a problem learning an action localization model with only video-level labels available. In recent years, many WTAL methods have developed. However, hard-to-predict snippets near action boundaries are often not considered in these existing approaches, causing action incompleteness and action over-complete issues. To solve these issues, in this work, an end-to-end snippets relation and hard-snippets mask network (SRHN) is proposed. Specifically, a hard-snippets mask module is applied to mask the hard-to-predict snippets adaptively, and in this way, the trained model focuses more on those snippets with low uncertainty. Then, a snippets relation module is designed to capture the relationship among snippets and can make hard-to-predict snippets easy to predict by aggregating the information of multiple temporal receptive fields. Finally, a snippet enhancement loss is further developed to reduce the action probabilities that are not present in videos for hard-to-predict snippets and other snippets, enlarging the action probabilities that exist in videos. Extensive experiments on THUMOS14, ActivityNet1.2, and ActivityNet1.3 datasets demonstrate the effectiveness of the SRHN method.
Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Weili Guan, Meng Wang 0001, Shengyong Chen
IEEE Trans. Circuits Syst. Video Technol.2
2023 A Novel Action Saliency and Context-Aware Network for Weakly-Supervised Temporal Action Localization
abstract
Temporal action localization is a challenging task in computer vision, and it tries to find the start time and the end time of the actions and predict their categories. However, compared to temporal action localization, weakly supervised temporal action localization (WTAL) is a more challenging task due to its poor annotations. With only video-level annotation, some background frames, similar to actions, would be classified as actions and produce inaccurate results. In addition, the two-stream fusion problem, ignored previously, also needs to be further considered. To resolve these issues, we propose a novel action saliency and context-aware network (ASCN) for weakly supervised temporal action localization tasks. Specifically, the temporal saliency and context module is designed to enhance the global saliency and context information of the RGB and the flow features to suppress the backgrounds and enhance the actions. In addition, a hybrid attention mechanism using frame differences and two-stream attention is designed to model the local action context information and further enlarge the scores of the potential action regions and suppress the background regions. Finally, to obtain two-stream consistency and solve the fusion problem, we use the similarity loss and a channel self-attention module to adaptively fuse the enhanced RGB and flow features. Extensive experiments demonstrate that ASCN can outperform all of the SOTA WTAL methods on the THUMOS14 dataset and the ActivityNet1.3 dataset with an average mAP that can reach 37.2% on the THUMOS14 dataset and attains an average mAP of 26.3% on the ActivityNet1.3 dataset. On the ActivityNet1.2 dataset, ASCN can also obtain comparable results. Compared with AdapNet (TNNLS20), MMSD (TIP22), and FTCL (CVPR22) on the THUMOS14 dataset, ASCN can outperform them by 13.5%, 2.9%, and 2.8%, respectively.
Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Wen Gao 0001, Meng Wang 0001, Shengyong Chen
IEEE Trans. Multim.2
2022 Multi-Level View Associative Convolution Network for View-Based 3D Model Retrieval
abstract
With the continuous improvement of image processing capabilities, a three-dimensional (3D) model that can contain rich information is becoming the fourth type of multimedia data (in addition to sound, image, and video). Moreover, since there is a wide range of applications of 3D models, how to quickly and effectively obtain the correct target model from the massive data has become a key issue. To date, 3D model retrieval approaches have been proposed, and in these approaches, view-based 3D model retrieval methods can achieve satisfactory performance. In the 3D model retrieval task, the latent relationship mining of all images in a 3D model, the adaptive fusion of different images, and the discriminative feature extraction are the main challenges, but in most existing solutions, these issues are separately performed and they are not explored in an end-to-end network architecture. To solve these issues, in this work, we propose a novel and effective multi-level view associative convolution network (MLVACN) to realize view-based 3D model retrieval, where the relationship exploration of multiple-view images, the fusion of different images, and the feature discrimination learning are realized in a unified end-to-end framework. Specifically, we design the group association layer and the block association layer to study the latent relationships among different views from the view-level and the block-level, respectively. Moreover, the weight fusion layer is further designed to adaptively fuse different views in a 3D model. In addition, these three layers are embedded into theMLVACN. Finally, the pairwise discrimination loss function is proposed to learn the discriminative features of the 3D model. Extensive experimental results on three 3D model retrieval datasets including ModelNet40, ModelNet10, and ShapeNetCore55 demonstrate thatMLVACNcan outperform state-of-the-art methods in term of mAP. When the ModelNet40 dataset is used, the mAP ofMLVACNis improved by 13.25%, 7.75%, 3.95%, and 0.61% as compared to those of the MVCNN, GVCNN, PVNet, and MLVCNN methods, respectively.
Zan Gao 0001, Yan Zhang 0154, Hua Zhang 0003, Weili Guan, Shengyong Chen
IEEE Trans. Circuits Syst. Video Technol.3
2022 A Novel Multiple-View Adversarial Learning Network for Unsupervised Domain Adaptation Action Recognition
abstract
Abstract-domain adaptation action recognition is a hot research topic in machine learning and some effective approaches have been proposed. However, samples in the target domain with label information are often required by these approaches. Moreover, domain-invariant discriminative feature learning, feature fusion, and classifier module learning have not been explored in an end-to-end framework. Thus, in this study, we propose a novel end-to-end multiple-view adversarial learning network (MAN) for unsupervised domain adaptation action recognition in which the fusion of RGB and optical-flow features, domain-invariant discrimination feature learning, and action recognition is conducted in a unified framework. Specifically, a robust spatiotemporal feature extraction network, including a spatial transform network and an adaptive intrachannel weight network, is proposed to improve the scale invariance and robustness of the method. Then, a self-attention mechanism fusion module is designed to adaptively fuse the RGB and optical-flow features. Moreover, a multiview adversarial learning loss is developed to obtain domain-invariant discriminative features. In addition, three benchmark datasets are constructed for unsupervised domain adaptation action recognition, for which all actions and samples are carefully collected from public action datasets, and their action categories are hierarchically augmented, which can guide how to extend existing action datasets. We conduct extensive experiments on four benchmark datasets, and the experimental results demonstrate that our proposed MAN can outperform several state-of-the-art unsupervised domain adaptation action recognition approaches. When the SDAI Action II-6 and SDAI Action II-11 datasets are used, MAN can achieve 3.7% ( H → U ) and 6.1% ( H → U ) improvements over the temporal attentive adversarial adaptation network (published in ICCV 2019) module, respectively. As an added contribution, the SDAI Action II-6, SDAI Action II-11, and SDAI Action II-16 datasets will be released to facilitate future research on domain adaptation action recognition.
Zan Gao 0001, Yibo Zhao 0001, Hua Zhang 0003, Da Chen 0002, Anan Liu, Shengyong Chen
IEEE Trans. Cybern.3
2022 A Temporal-Aware Relation and Attention Network for Temporal Action Localization
abstract
Temporal action localization is currently an active research topic in computer vision and machine learning due to its usage in smart surveillance. It is a challenging problem since the categories of the actions must be classified in untrimmed videos and the start and end of the actions need to be accurately found. Although many temporal action localization methods have been proposed, they require substantial amounts of computational resources for the training and inference processes. To solve these issues, in this work, a novel temporal-aware relation and attention network (abbreviated as TRA) is proposed for the temporal action localization task. TRA has an anchor-free and end-to-end architecture that fully uses temporal-aware information. Specifically, a temporal self-attention module is first designed to determine the relationship between different temporal positions, and more weight is given to features within the actions. Then, a multiple temporal aggregation module is constructed to aggregate the temporal domain information. Finally, a graph relation module is designed to obtain the aggregated graph features, which are used to refine the boundaries and classification results. Most importantly, these three modules are jointly explored in a unified framework, and temporal awareness is always fully used. Extensive experiments demonstrate that the proposed method can outperform all state-of-the-art methods on the THUMOS14 dataset with an average mAP that reaches 67.6% and obtain a comparable result on the ActivityNet1.3 dataset with an average mAP that reaches 34.4%. Compared with A2Net (TIP20), PCG-TAL (TIP21), and AFSD (CVPR21) TRA can achieve improvements of 11.7%, 4.4%, and 1.8%, respectively on the THUMOS14 dataset.
Yibo Zhao 0001, Hua Zhang 0003, Zan Gao 0001, Weili Guan, Jie Nie, Anan Liu, Meng Wang 0001, Shengyong Chen
IEEE Trans. Image Process.2
2021 A bus passenger re-identification dataset and a deep learning baseline using triplet embedding
Junliang Guo, Yanbing Xue, Zan Gao 0002, Guangping Xu, Hua Zhang 0003
Multim. Tools Appl.6
2021 DCR: A Unified Framework for Holistic/Partial Person ReID
abstract
Person reidentification (ReID) is a very popular research topic in machine learning and computer vision. According to the occlusions, it can be divided into holistic person ReID and partial person ReID tasks. Occlusions commonly exist in the partial person ReID task but pose or observation perspective changes often occur in the holistic person ReID task; thus, many different algorithms or different network architectures have been designed for each task. However, this approach increases the cost in practice and hinders the development of ReID techniques. To solve this problem, in this work, a unified framework is proposed for holistic/partial person ReID, which can effectively and efficiently address changes in pose or observation perspective and the occlusions in both tasks. In detail, we first employ a fully convolutional network (FCN) to generate feature maps for an arbitrarily sized image and then use spatial pyramid pooling (SPP) to obtain its spatial pyramid feature. Thereafter, to efficiently solve the matching problem between the query image and gallery images, we build a deep spatial pyramid feature collaborative reconstruction model (DCR). In DCR, the reconstruction errors mainly come from similar blocks (uncovered parts), and the influence of the reconstruction errors of dissimilar blocks (covered parts or changed parts) is minimized. In addition, we also use the deep mutual learning approach to jointly learn the features in the training process and promote model training. Experimental results on two partial person ReID datasets and three holistic person ReID datasets demonstrate that the DCR outperforms the state-of-the-art approaches on both tasks and all datasets. Specifically, it outperforms all competitors with a large margin and achieves an improvement of 9.07% and 5.95% over the DSR method (published in CVPR18) on the Partial ReID and Partial-iLIDS datasets with Rank-1, respectively. Similarly, it also achieves an improvement of 5.08% over the VPM method (published in CVPR19) on the DukeMTMC-ReID dataset with Rank-1. Additionally, the running time of our method for each query is more than 7 faster than that of the DSR or DuATM methods.
Zan Gao 0001, Li-Shuai Gao, Hua Zhang 0003, Zhiyong Cheng 0001, Richang Hong, Shengyong Chen
IEEE Trans. Multim.3
2020 Attention Stereo Matching Network
abstract
Despite great progress, previous stereo matching algorithms still lack the ability to match textureless regions and slender structure areas. To tackle this problem, we propose ASM-Net, an attention stereo matching network. Attention module and disparity refinement module are constructed in the ASMNet. The attention module can improve correlation information between two images by channels and spatial attention. The feature-guided disparity refinement module learns more geometry information in different feature levels to refine the coarse prediction resolution constantly. The proposed approach was evaluated on several benchmark datasets. Experiments show that the proposed method achieves competitive results on KITTI and Scene-Flow datasets while running in real-time at 14ms.
Doudou Zhang, Yanbing Xue, Hua Zhang 0003
ICPR5
2020 Texture Semantically Aligned with Visibility-aware for Partial Person Re-identification
abstract
In real person re-identification (ReID) tasks, pedestrians are often obscured by other pedestrians or objects; moreover, changes in poses or observation perspectives also commonly exist in partial-person ReID. To the best of our knowledge, few works simultaneously focus on these two issues. In this work, we propose a novel texture semantic alignment (TSA) approach with the visibility-aware for partial person ReID task where the occlusion issue and changes in poses are simultaneously explored in an end-to-end unified framework. Specifically, we first employ a texture alignment scheme with the semantic visibility of a person's image to solve the issue of changes in poses that can enhance the alignment and generalization capability of the models. Second, we design a human pose-based partial region alignment scheme to solve the occlusion problem that makes TSA method emphasize the shared body parts. Finally, these two networks jointly learn these aspects. Extensive experimental results demonstrate that our proposed TSA method is very effective and robust for simultaneously handling occlusion and changes in pose, and it can outperform state-of-the-art approaches by a large margin and achieves an improvement of 5% and 6.4% on the rank-1 accuracy over the visibility-aware part model (VPM) method (published in CVPR 2019) on the Partial ReID and Partial-iLIDS datasets, respectively.
Li-Shuai Gao, Hua Zhang 0003, Zan Gao 0001, Weili Guan, Zhiyong Cheng 0001, Meng Wang 0001
ACM Multimedia2
2019 Deep Spatial Pyramid Features Collaborative Reconstruction for Partial Person ReID
abstract
Partial person re-identification (ReID) is a hot research problem in computer vision. Accurate partial ReID is very challenging due to the common occlusion problem. To address this problem, in this paper, we propose a novel D eep spatial pyramid feature C ollaborative R econstruction approach (DCR ) for partial person ReID, which can effectively and efficiently tackle the occlusion in arbitrary sizes. Specifically, a fully convolutional network (FCN) is first leveraged to extract feature maps of an arbitrary-size image, and then the spatial pyramid pooling (SPP) is adopted to obtain spatial pyramid features. Thereafter, our DCR method is designed to efficiently solve the matching problem between the partial person and the holistic person in the partial person ReID task where the occlusion problem often occurs. Experiments on two partial person ReID datasets demonstrate the efficiency and efficacy of the proposed method by comparing to several state-of-the-art partial person ReID approaches. Our method outperforms all the competitors with a large margin and can achieve an improvement of 9.07% and 5.95% over the DSR method on the Partial REID and Partial-iLIDS Person ReID datasets in terms of the Rank-1 accuracy, respectively.
Zan Gao 0001, Li-Shuai Gao, Hua Zhang 0003, Zhiyong Cheng 0001, Richang Hong
ACM Multimedia3
2019 Cognitive-inspired class-statistic matching with triple-constrain for camera free 3D object retrieval
Zan Gao 0002, Shaohua Wan 0001, Hua Zhang 0003, Yinglong Wang 0001
Future Gener. Comput. Syst.4
2019 As-global-as-possible stereo matching with adaptive smoothness prior
abstract
More global matching (MGM) overcomes the limitation of one‐dimensional scanline optimisation in semi‐global matching (SGM). Nevertheless, the possible weaknesses of the MGM algorithm are as follows: (i) only two directions are considered for each image traversal direction, which may lead to massive mismatches; (ii) disparity estimation around the object boundaries usually performs terrible since the smoothness term is designed independent of the image prior. In this research, the authors consider all of the four directions for each image traversal direction through a novel model. Besides utilising the prior of neighboured pixels' correlation, adaptive smoothness terms are modelled and augmented into the energy function. These contributions encourage ‘as‐global‐as‐possible (AGAP)’. More importantly, different from the recent works in which the aggregated algorithms have been conducted as the data term of an energy function, conversely, the authors make the energy function as a part of cost aggregation framework. Performance evaluations on Middlebury v.2 and v.3 stereo data sets demonstrate that the proposed AGAP outperforms other four most challenging stereo matching algorithms, and also performs better on Microsoft i2i stereo videos. In addition, under various strategies of parallelisation, the presented AGAP shows a near real‐time execution time.
Hua Zhang 0003, Yanbing Xue, Shengyong Chen
IET Image Process.2
2019 Adaptive Fusion and Category-Level Dictionary Learning Model for Multiview Human Action Recognition
abstract
Human actions are often captured by multiple cameras (or sensors) to overcome the significant variations in viewpoints, background clutter, object speed, and motion patterns in video surveillance, and action recognition systems often benefit from fusing multiple types of cameras (sensors). Therefore, adaptive fusion of the information from multiple domains is mandatory for multiview human action recognition. Two widely applied fusion schemes are feature-level fusion and score-level fusion. We point out that limitations still exist and there is tremendous room for improvement, including the separate computation of feature fusion and action recognition, or the fixed weights for each action and each camera. However, previous fusion methods cannot accomplish them. In this paper, inspired by nature, the above limitations are addressed for multiview action recognition by developing a novel adaptive fusion and category-level dictionary learning model (abbreviated to AFCDL). It can jointly learn the adaptive weight for each camera and optimize the reconstruction of samples toward the action recognition task. To induce the dictionary learning and the reconstruction of query set (or test samples), the induced set for each category is built, and the corresponding induced regularization term is designed for the objective function. Extensive experiments on four public multiview action benchmarks show that AFCDL can significantly outperforms the state-of-the-art methods with 3% to 10% improvement in recognition accuracy.
Hai-Zhen Xuan, Hua Zhang 0003, Shaohua Wan 0001, Kim-Kwang Raymond Choo
IEEE Internet Things J.3
2019 Multi-view and multivariate gaussian descriptor for 3D object retrieval
Kaixin Xue, Hua Zhang 0003
Multim. Tools Appl.3
2018 Group-Pair Convolutional Neural Networks for Multi-View Based 3D Object Retrieval
abstract
In recent years, research interest in object retrieval has shifted from 2D towards 3D data. Despite many well-designed approaches, we point out that limitations still exist and there is tremendous room for improvement, including the heavy reliance on hand-crafted features, the separated optimization of feature extraction and object retrieval, and the lack of sufficient training samples. In this work, we address the above limitations for 3D object retrieval by developing a novel end-to-end solution named Group Pair Convolutional Neural Network (GPCNN). It can jointly learn the visual features from multiple views of a 3D model and optimize towards the object retrieval task. To tackle the insufficient training data issue, we innovatively employ a pair-wise learning scheme, which learns model parameters from the similarity of each sample pair, rather than the traditional way of learning from sparse label–sample matching. Extensive experiments on three public benchmarks show that our GPCNN solution significantly outperforms the state-of-the-art methods with 3% to 42% improvement in retrieval accuracy.
Xiangnan He 0001, Hua Zhang 0003
AAAI4
2018 AGO: Accelerating Global Optimization for Accurate Stereo Matching
Hua Zhang 0003, Yanbing Xue, Shengyong Chen
MMM (1)2
2018 MSCS: MeshStereo with Cross-Scale Cost Filtering for fast stereo matching
abstract
MeshStereo (MS) and cross‐scale cost filtering (CSCF) are two most recently celebrated models for stereo matching. On one hand, MS model enlightens for fast solving the dense stereo correspondence problem according to a region‐based opinion. On the other hand, CSCF model could generate more robust matching cost volumes than single scale. In this study, the authors weave these two models together for attaining greater and faster disparity estimation. With CSCF, more powerful initial volumes of matching cost are computed and they are conducted as the data term of MS energy function model. More importantly, the novel‐fused stereo model also draws a closer connection between multi‐scale aggregated and global algorithms. Integrating the advantages of both stereo models, they name the presented one as MS with cross‐scale (MSCS). Performance evaluations on Middlebury v.2 and v.3 stereo data sets demonstrate that the proposed MSCS outperforms other four most challenging stereo matching algorithms; and also performs better on Microsoft i2i stereo videos. In addition, thanks to this novel‐fused model, MSCS requires fewer iteration times for optimising and makes it surprisingly possesses a much faster execution time.
Hua Zhang 0003, Yanbing Xue, Shengyong Chen
IET Comput. Vis.2
2018 3D object recognition based on pairwise Multi-view Convolutional Neural Networks
Zan Gao 0002, Yanbin Xue, Guangping Xu, Hua Zhang 0003, Yinglong Wang 0001
J. Vis. Commun. Image Represent.5
2018 MMA: a multi-view and multi-modality benchmark dataset for human action recognition
Zan Gao 0002, Tao-tao Han, Hua Zhang 0003, Yanbing Xue, Guangping Xu
Multim. Tools Appl.3
2018 Semantic segmentation based on fusion of features and classifiers
Yanbing Xue, Huiqiang Geng, Hua Zhang 0003, Zhenshan Xue, Guangping Xu
Multim. Tools Appl.3
2017 Segment-tree based cost aggregation for stereo matching with enhanced segmentation advantage
abstract
Segment-tree (ST) based cost aggregation algorithm for stereo matching successfully integrates the information of segmentation with non-local cost aggregation framework. The tree structure which is generated by the segmentation strategy directly determines the final results for this kind of algorithms. However, the original strategy performs unreasonable due to its coarse performance and ignores to meet the disparity consistency assumption. To improve these weaknesses we propose a novel segmentation algorithm for constructing a more faithful ST with enhanced segmentation advantage according to a robust initial over-segmentation. Then we implement non-local cost aggregation framework on this new ST structure and obtain improved disparity maps. Performance evaluations on all 31 Middlebury stereo pairs show that the proposed algorithm outperforms than other five state-of-the-art aggregated based algorithms and also keeps time efficiency.
Hua Zhang 0003, Yanbing Xue, Mian Zhou, Guangping Xu, Zan Gao 0002, Shengyong Chen
ICASSP2
2017 SPMVP: Spatial PatchMatch Stereo with Virtual Pixel Aggregation
Hua Zhang 0003, Yanbing Xue, Shengyong Chen
ICONIP (3)2
2017 3D human action recognition model based on image set and regularized multi-task leaning
Zan Gao 0002, Guotai Zhang, Hua Zhang 0003, Yanbin Xue, Guangping Xu
Neurocomputing3
2017 Collaborative sparse representation leaning model for RGBD action recognition
Su-hua Li, Yujia Zhu, Hua Zhang 0003
J. Vis. Commun. Image Represent.5
2017 Evaluation of regularized multi-task leaning algorithms for single/multi-view human action recognition
Su-hua Li, Guotai Zhang, Yujia Zhu, Hua Zhang 0003
Multim. Tools Appl.6
2016 Extremal graphic model in optimizing fractional repetition codes for efficient storage repair
abstract
Consider that a set of balls of n different colors are thrown into m bins with the assumption that the ball number of each color is constant and the number of balls in each bin is also constant. Our optimal goal is to find a feasible placement such that the distinct colors of remaining balls should be at least c after removing any k bins (k ≤ m) with the minimum number of balls. We present that the optimal colored bins in bins is equivalent to the optimization of Fractional Repetition (FR) codes in distributed storage systems. Here balls correspond to coded packets and bins correspond to storage nodes. This problem can be represented as biregualr graph and then deduced to the Zarankiewicz problem, which is a well-known extremal graph theoretic problem. We present the problem with the relation to combinatorial design theory, especially t-designs and propose the explicit construction algorithm for the optimization problem from t-designs. Some constructions of the optimized FR codes by 2-designs are analyzed to tolerate the desired k fault-tolerance with c = n - 1.
Guangping Xu, Qunfang Mao, Sheng Lin 0002, Kai Shi 0002, Hua Zhang 0003
ICC5
2016 Iterative color-depth MST cost aggregation for stereo matching
abstract
The minimum spanning tree (MST) based non-local cost aggregation algorithm performs well in accuracy and time efficiency. However, it can still be improved in two aspects. First, we propose a logarithmic transformation on matching cost function to improve the matching efficiency in texture less regions. The textureless neighbors can provide effective contributions in cost aggregation by the proposed monotone increasing function. Hence the algorithm can distinguish different pixels in textureless regions. Second, MST algorithm only utilizes color information in weight function while aggregating, which leads 3D cues missing. We introduce depth weight computed from the original MST algorithm into an edge weight function. With the proposed color-depth weight, we further iteratively rebuild the tree and obtain enhanced disparity map. Performance evaluations on 19 Middlebury stereo pairs and Microsoft stereo videos show that the proposed algorithm outperforms than other five state-of-the-art cost aggregation algorithms.
Hua Zhang 0003, Yanbing Xue, Mian Zhou, Guangping Xu, Zan Gao 0002
ICME2
2016 A Fast 3D Retrieval Algorithm via Class-Statistic and Pair-Constraint Model
abstract
With the development of 3D technologies and devices, 3D model retrieval becomes a hot research topic where multi-view matching algorithms have demonstrated satisfying performance. However, exciting works overlook the common factors among objects in a single class, and they are time consuming in retrieval processing. In this paper, a class-statistics and pair-constraint model (CSPC) method is originally proposed for 3D model retrieval, which is composed of supervised class-based statistics model and pair-constraint object retrieval model. In our CSPC model, we firstly convert view-based distance measure into object-based distance measure without falling in performance, which will advance 3D model retrieval speed. Secondly, the generality of the distribution of each feature dimension in each class is computed to judge category information, and then we further adopt this distribution information to build class models. Finally, an object-based pairwise constraint is introduced on the base of the class-statistic measure, which can remove a lot of false alarm samples in retrieval. Experimental results on ETH, NTU-60, MVRED and PSB 3D datasets show that our method is fast, and its performance is also comparable with the-state-of-the-art algorithms.
Zan Gao 0002, Hua Zhang 0003, Yanbing Xue, Guangping Xu
ACM Multimedia3
2016 Reverse Testing Image Set Model Based Multi-view Human Action Recognition
Yan Zhang 0154, Hua Zhang 0003, Guangping Xu, Yanbing Xue
MMM (1)3
2016 Evaluation of local spatial-temporal features for cross-view action recognition
Zan Gao 0002, Weizhi Nie, Anan Liu, Hua Zhang 0003
Neurocomputing4
2016 Multi-dimensional human action recognition model based on image set and group sparisty
Yan Zhang 0154, Hua Zhang 0003, Yanbin Xue, Guangping Xu
Neurocomputing3
2016 Human action recognition on depth dataset
Zan Gao 0002, Hua Zhang 0003, Anan Liu, Guangping Xu, Yanbing Xue
Neural Comput. Appl.2
2015 Performance Optimization and Evaluation of Space Management in Cloud Storage Systems
Guangping Xu, Qunfang Mao, Sheng Lin 0002, Hua Zhang 0003
ICA3PP (4)5
2015 Single Face Image Super-Resolution via Multi-dictionary Bayesian Non-parametric Learning
Hua Zhang 0003, Yanbing Xue, Mian Zhou, Guangping Xu, Zan Gao 0002
ICONIP (1)2
2015 Multi-perspective and multi-modality joint representation and recognition model for 3D action recognition
Zan Gao 0002, Hua Zhang 0003, Guangping Xu, Yanbin Xue
Neurocomputing2
2015 Multi-view discriminative and structured dictionary learning with group sparsity for human action recognition
Zan Gao 0002, Hua Zhang 0003, Guangping Xu, Yanbin Xue, Alex Hauptmann 0001
Signal Process.2
2014 Enhanced and hierarchical structure algorithm for data imbalance problem in semantic extraction under massive video dataset
Zan Gao 0002, Ming-yu Chen 0001, Alex Hauptmann 0001, Hua Zhang 0003, Anni Cai
Multim. Tools Appl.5
2013 Optimization for reliable erasure-coded storage allocation under multiple constraints
abstract
To maximize the reliability of erasure coded objects in cloud storage, the optimal allocation is a challenging problem constrained with the node heterogeneities and erasure coding budgets. We model the optimal problem from a combinatorial view under multiple constraints and propose an efficient search algorithm to find the most reliable allocation. The experiments are evaluated and analyzed including the optimal reliability, redundancy reduction and allocation pattern.
Guangping Xu, Hua Zhang 0003, Sheng Lin 0002, Chunxia Yang
IPCCC3
2013 Expander code: A scalable erasure-resilient code to keep up with data growth in distributed storage
abstract
To ensure high reliability and storage efficiency, erasure codes are preferred in storage systems. With the prevalent of distributed storage systems such as clouds storage, how to design a scalable and efficient erasure-resilient code is challenging. We propose a scalable binary linear code to keep up with data growth which has the following properties. Given the group size k and the code block length n, the proposed code corrects any two bit erasures among the n bits. The redundancy overhead of the code is 2/(k + 2), and each data bit affects exactly 2 parity bits. As results of these properties, if a data bit is changed or added, only two parity bits need to be updated; and the recovery of an erasured bit requires accessing at most k other bits and the recovery of two erasured bits requires at most 2k other bits. We give the construction algorithm by the order expansion of regular graphs; moreover, we optimize the failure resilience during the construction procedure. Compared with existing codes, our proposed code has notable benefits in storage scalability, redundancy overhead and I/O bandwidth. The deployment of the proposed code in distributed storage systems can be simple and practical.
Guangping Xu, Sheng Lin 0002, Hua Zhang 0003, Kai Shi 0002
IPCCC3
2013 Online Boosting Tracking with Fragmented Model
Dingcheng Shen, Hua Zhang 0003, Yanbing Xue, Guangping Xu, Zan Gao 0002
MMM (2)2
2012 Human action recognition based on sparse representation induced by L1/L2 regulations
Zan Gao 0002, Anan Liu, Hua Zhang 0003, Guangping Xu, Yanbing Xue
ICPR3
2012 HERO: Heterogeneity-aware erasure coded redundancy optimal allocation for reliable storage in distributed networks
abstract
Heterogeneity is the natural feature in distributed networks. Different from the traditional disk array, the amount of data allocated on heterogenous peers may be not the same. To maximize the reliability of stored data objects in heterogeneous networks, the optimal allocation of erasure-coded fragments is a challenging problem constrained with heterogeneous peer availabilities and redundancy overhead. This paper examines this optimal problem considered MDS erasure codes applied into distributed storage networks. First, we model the reliability of an allocation with the weighted-k-out-of-s model and extend its properties to efficiently calculate the reliability of an allocation; then we reduce the reliability computation of a given allocation to linear computation cost based on the weighted k-out-of-s model. Then, we deduce the problem to integer partition problem and propose two order-based search algorithms. Our experiments show that our proposed algorithms can be applied to find the optimal allocations efficiently in various practical coding cases. Furthermore, we evaluate the performance of our proposed search algorithms with some practical storage settings, and then present experimental results including the reliability, redundancy overheads and allocation pattern for the optimal allocation driven by practical network traces.
Guangping Xu, Sheng Lin 0002, Gang Wang 0001, Xiaoguang Liu 0001, Kai Shi 0002, Hua Zhang 0003
IPCCC6
2010 Performance Comparison of Erasure Codes for Different Churn Models in P2P Storage Systems
Jingxing Li, Guangping Xu, Hua Zhang 0003
ICIC (2)3
2010 Vector-valued Chan-Vese model driven by local histogram for texture segmentation
abstract
The Chan-Vese model is one of the most popular region-based active contours, and its vector-valued extension is also powerful for multichannel images. Very recently, the histogram is introduced into the Chan-Vese model due to the effectiveness of histogram to model region information. Motivated by the fact that the histogram is also a powerful tool to characterize texture, it is introduced into the vector-valued Chan-Vese model for texture segmentation in this work. In order to determine an optimal number of bins in the histogram, a Bayesian method is adopted. Experiments are conducted and the results show that the proposed strategy is effective for texture segmentation.
Yuanquan Wang 0001, Yue Xiong, Liping Lv, Hua Zhang 0003, Zuoliang Cao
ICIP4
2010 Image Segmentation Using Active Contours With Normally Biased GVF External Force
abstract
Gradient vector flow (GVF) is an effective external force for active contours, but its isotropic nature handicaps its performance. The recently proposed NGVF model is anisotropic since it only keeps the diffusion along the normal direction of the isophotes; however, it is sensitive to noise and could erase weak boundaries. In this letter, the normally biased GVF (NBGVF) external force is proposed for snake models, which keeps the diffusion along the tangential direction of the isophotes and biases that along the normal direction. The biasing weight approaches zero at boundaries and is 1 in homogeneous regions. Consequently, the NBGVF snake can preserve weak edges and smooth out noise while maintaining other desirable properties of GVF and NGVF snakes such as enlarged capture range, insensitivity to initialization and convergence to u-shape concavity. These properties are evaluated on synthetic and real images.
Yuanquan Wang 0001, Lixiong Liu, Hua Zhang 0003, Zuoliang Cao, Shaopei Lu
IEEE Signal Process. Lett.3
2009 Network measurement based redundancy model and maintenance in dynamic P2P storage systems
abstract
Peer-to-peer distributed storage systems aggregate the storage space of many peers spread over the Internet. Due to the dynamic and scalable nature of these systems, it is a challenging issue to access data in an available and reliable way through redundancy. Following the modeling methodology presented, we present the stochastic model to analyze redundancy evolution of these systems under churn. Different from the previous work based on the average peer availability, the stochastic model can be applied into the practice based on both conditional probabilities (alpha, theta) which can be obtained from network probing easily. First, we apply the model to characterize the redundancy evolution of a fragment system with temporary churn. Then, we use an empirical trace and a synthetic trace to validate the model. Second, based on the characteristics of different churn from the both probabilities, we propose the redundancy maintenance strategy assisted by network sampling. Our simulations evaluate the performance of the strategy driven by empirical and synthetic traces.
Guangping Xu, Hua Zhang 0003, Jing Liu 0010, Gang Wang 0001, Xiaoguang Liu 0001
IPCCC2
2009 A new watermarking approach based on probabilistic neural network in wavelet domain
Xian-Bin Wen, Hua Zhang 0003, Xue-Quan Xu, Jin-Juan Quan
Soft Comput.2
2008 SHTM: A Semantic Hierarchy Transaction Model for Web Services Transactions
abstract
Web Services (WS) have quickly evolved as an approach to integrate processes and applications at an inter-enterprise level. Such WS-based integrated applications should guarantee consistent data manipulation. Thus, WS should be extended to equip with transaction processing functionalities, i.e. WS transactions. WS transactions differ from traditional transactions in that they execute over long periods and cross multiple loosely-coupled organizations, so traditional transaction models and processing mechanism have been inappropriate for them. In this paper, we first propose a semantic hierarchy transaction model (called SHTM) for WS transactions. On the basis of this, the isolation property of WS transactions is relaxed by allowing uncommitted transactions to lend their data to other executing transaction. Further, we give the commit processing mechanism which can guarantee consistent data manipulation and correct outcomes of WS transactions.
Yingyuan Xiao, Hua Zhang 0003
APSCC2
2006 A Multiscale Self-growing Probabilistic Decision-Based Neural Network for Segmentation of SAR Imagery
Xian-Bin Wen, Hua Zhang 0003
PRICAI2