VLDB 2026 Research / reviewers in the wild / expert
Yongzhao Zhan 0001
dblp:67/6654-1 · also Yong-zhao Zhan 0001
· DBLP profile ↗
74ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0001-7475-2895ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 19 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Systems, architecture and hardware · 2Security and privacy · 2 · 1 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Subdomain adaptive feature enhancement via confidence-adjudicated dual-decision pseudo-labeling for cross-subject and cross-session EEG emotion recognition
Wenwen He, Qinghua Ren, Yongzhao Zhan 0001 |
Multim. Syst. | 5 |
| 2025 | Effective Crowdsourcing Ranking Solution: Performance-Aware Working Group Selection and Efficient Result AggregationabstractCrowdsourcing can harness grassroots intelligence to effectively solve complex ranking tasks. Two key issues for these tasks are how to recruit high-performance working groups and how to efficiently aggregate their submissions. In past work, many attempts have been made to maximize social benefits or motivate worker participation, without considering the performance and efficiency of task completion. Therefore, we investigate an effective crowdsourcing ranking solution which is concerned with the two problems of optimal working group selection and efficient result aggregation. A performance-aware working group selection method is proposed to solve the high time complexity of combinatorial optimization in working group selection with a fixed number of requester recruits. To improve the performance of aggregated results more efficiently, an efficient aggregation method based on differential evolution algorithm and Top-k pruning is proposed. Finally, the effectiveness of the proposed methods is verified by extensive experiments compared with recent related methods. Yuping Xing, Yongzhao Zhan 0001 |
CSCWD | 2 |
| 2025 | Inter-image Token Relation Learning for weakly supervised semantic segmentation
Jingfeng Tang, Keyang Cheng, Liutao Wei, Yongzhao Zhan 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | Dynamic prompting class distribution optimization for semi-supervised sound event detectionabstractSemi-supervised sound event detection (SSED) tasks typically leverage a large amount of unlabeled and synthetic data to facilitate model generalization during training, reducing overfitting on a limited set of labeled data. However, the generalization training process often encounters challenges from noisy interference introduced by pseudo-labels or domain knowledge gaps. To alleviate noisy interference in class distribution learning, we propose an efficient semi-supervised class distribution learning method through dynamic prompt tuning, named prompting class distribution optimization (PADO). Specifically, when modeling real labeled data, PADO dynamically incorporates independent learnable prompt tokens to explore prior knowledge about the true distribution. Then, the prior knowledge serves as prompt information, dynamically interacting with the posterior noisy-class distribution information. In this case, PADO achieves class distribution optimization while maintaining model generalization, leading to a significant improvement in the efficiency of class distribution learning. Compared with state-of-the-art methods on the SSED datasets from DCASE 2019, 2020, and 2021 challenges, PADO achieves significant performance improvements. Furthermore, it is readily extendable to other benchmark models. Lijian Gao, Qing Zhu 0002, Yaxin Shen, Qirong Mao, Yongzhao Zhan 0001 |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2025 | Constraint embedding for prompt tuning in vision-language pre-trained model
Keyang Cheng, Liutao Wei, Jingfeng Tang, Yongzhao Zhan 0001 |
Multim. Syst. | 4 |
| 2025 | Unsupervised subdomain adaptation framework guided by pseudo label for cross-subject and cross-session EEG emotion recognition
Wenwen He, Yalan Ye, Qinghua Ren, Yongzhao Zhan 0001 |
Multim. Syst. | 6 |
| 2025 | Infrared-Visible Image Fusion Using Dual-Branch Auto-Encoder With Invertible High-Frequency EncodingabstractIn the field of Infrared-Visible Image Fusion (IVIF), the preservation of details, edges, and texture is crucial for generating high-quality fused images. However, a major challenge arises due to the inevitable loss of high-frequency information during feature extraction, resulting in fused images that lack significant details. In this paper, we propose a dual-branch auto-encoder by exploiting an invertible high-frequency branch for detailed feature preservation and a transformer-based low-frequency branch for global dependencies modeling. First, the high-frequency branch employs the wavelet transforms and an Invertible Neural Networks (INN)-based encoder to model high-frequency features through an invertible transformation, including a forward process for image fusion and an inverse process for original image reconstruction. Additionally, a high-frequency loss is designed to enhance the high-frequency feature representation for high-quality image fusion. Second, a low-frequency branch based on a transformer encoder and an adaptive fusion module is introduced to capture the global contextual features of the infrared and visible images. Finally, the decoder integrates the low- and high-frequency features from both branches to generate the final fused image. Image fusion, object detection, and semantic segmentation experiments conducted on public datasets such as TNO, MFNet, and M3FD, show that our method outperforms the state-of-the-art (SOTA) image fusion methods. Qirong Mao, Ming Dong 0001, Yongzhao Zhan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | A post-processing framework for class-imbalanced learning in a transductive setting
Yu Lu 0017, Yongzhao Zhan 0001, Qirong Mao |
Expert Syst. Appl. | 4 |
| 2024 | Person re-identification via deep compound eye network and pose repair moduleabstractAbstract Person re‐identification is aimed at searching for specific target pedestrians from non‐intersecting cameras. However, in real complex scenes, pedestrians are easily obscured, which makes the target pedestrian search task time‐consuming and challenging. To address the problem of pedestrians' susceptibility to occlusion, a person re‐identification via deep compound eye network (CEN) and pose repair module is proposed, which includes (1) A deep CEN based on multi‐camera logical topology is proposed, which adopts graph convolution and a Gated Recurrent Unit to capture the temporal and spatial information of pedestrian walking and finally carries out pedestrian global matching through the Siamese network; (2) An integrated spatial‐temporal information aggregation network is designed to facilitate pose repair. The target pedestrian features under the multi‐level logic topology camera are utilised as auxiliary information to repair the occluded target pedestrian image, so as to reduce the impact of pedestrian mismatch due to pose changes; (3) A joint optimisation mechanism of CEN and pose repair network is introduced, where multi‐camera logical topology inference provides auxiliary information and retrieval order for the pose repair network. The authors conducted experiments on multiple datasets, including Occluded‐DukeMTMC, CUHK‐SYSU, PRW, SLP, and UJS‐reID. The results indicate that the authors’ method achieved significant performance across these datasets. Specifically, on the CUHK‐SYSU dataset, the authors’ model achieved a top‐1 accuracy of 89.1% and a mean Average Precision accuracy of 83.1% in the recognition of occluded individuals. Hongjian Gu, Wenxuan Zou, Keyang Cheng, Humaira abdul Ghafoor, Yongzhao Zhan 0001 |
IET Comput. Vis. | 6 |
| 2024 | Enhancing action discrimination via category-specific frame clustering for weakly-supervised temporal action localizationabstractTemporal action localization (TAL) is a task of detecting the start and end timestamps of action instances and classifying them in an untrimmed video. As the number of action categories per video increases, existing weakly-supervised TAL (W-TAL) methods with only video-level labels cannot provide sufficient supervision. Single-frame supervision has attracted the interest of researchers. Existing paradigms model single-frame annotations from the perspective of video snippet sequences, neglect action discrimination of annotated frames, and do not pay sufficient attention to their correlations in the same category. Considering a category, the annotated frames exhibit distinctive appearance characteristics or clear action patterns. Thus, a novel method to enhance action discrimination via category-specific frame clustering for W-TAL is proposed. Specifically, the K -means clustering algorithm is employed to aggregate the annotated discriminative frames of the same category, which are regarded as exemplars to exhibit the characteristics of the action category. Then, the class activation scores are obtained by calculating the similarities between a frame and exemplars of various categories. Category-specific representation modeling can provide complimentary guidance to snippet sequence modeling in the mainline. As a result, a convex combination fusion mechanism is presented for annotated frames and snippet sequences to enhance the consistency properties of action discrimination, which can generate a robust class activation sequence for precise action classification and localization. Due to the supplementary guidance of action discriminative enhancement for video snippet sequences, our method outperforms existing single-frame annotation based methods. Experiments conducted on three datasets (THUMOS14, GTEA, and BEOID) show that our method achieves high localization performance compared with state-of-the-art methods. Huifen Xia 0001, Yongzhao Zhan 0001, Xiaopeng Ren |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2024 | Object-based video anomaly detection using multi-attention and adaptive velocity attribute representation learning
Xiaopeng Ren, Huifen Xia 0001, Yongzhao Zhan 0001 |
Multim. Syst. | 3 |
| 2024 | Generalized Welsch penalty for edge-aware image decomposition
Yang Yang 0046, Shunli Ji, Lanling Zeng, Yongzhao Zhan 0001 |
Multim. Syst. | 5 |
| 2024 | Self-expressive induced clustered attention for video-text retrieval
Jingxuan Zhu, Xiangjun Shen, Sumet Mehta, Timothy Apasiba Abeo, Yongzhao Zhan 0001 |
Multim. Syst. | 5 |
| 2024 | β-CLVAE: a semantic disentangled generative model
Keyang Cheng, Chunyun Meng, Guojian Ma, Yongzhao Zhan 0001 |
Multim. Tools Appl. | 4 |
| 2024 | Tiny Object Detection via Regional Cross Self-Attention NetworkabstractAs vision sensor technology continues to evolve, the requirements for detecting targets of interest in the images captured by the sensors are increasing. Considering fast detection and high accuracy, the industry favors geometric key point-based solutions. However, there are a large number of small and fuzzy objects in the real world. Geometric key point detectors do not effectively utilize the contextual features of the region of interest, leading to excessive false positive and false negative results. In this work, a simple, effective, and interpretable tiny object detection method called Regional Cross Self-Attention Object Detection Network (RCSANet) is proposed. It adopts Region Proposal Networks and transformers to capture regional background relations and uses regional background relations to generate key point sequences. The regional cross self-attention mechanism is introduced to curtail computation redundancy and minimize the interference of redundant information to the target region. Additionally, a position coding called dynamic implicit position coding is proposed to cooperate with regional cross self-attentiveness. Dynamic implicit location coding can encode arbitrarily long input sequences. The computational cost of RCSANet is significantly lower than that of state-of-the-art object detection solutions. Moreover, RCSANet improves the performance on the four benchmark datasets, of MSCOCO, Tinyperson, DOTA, and AI-TOD, by about 3.0%AP. Keyang Cheng, Honggang Cui, Humaira abdul Ghafoor, Qirong Mao, Yongzhao Zhan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | A semi-supervised resampling method for class-imbalanced learning
Yu Lu 0017, Yongzhao Zhan 0001, Qirong Mao |
Expert Syst. Appl. | 4 |
| 2023 | Fast bilateral filter with spatial subsampling
Yang Yang 0046, Yiwen Xiong, Yanqing Cao, Lanling Zeng, Yan Zhao 0038, Yongzhao Zhan 0001 |
Multim. Syst. | 6 |
| 2023 | Deep cascaded action attention network for weakly-supervised temporal action localization
Huifen Xia 0001, Yongzhao Zhan 0001 |
Multim. Tools Appl. | 2 |
| 2023 | Semi-Supervised Clustering Under a "Compact-Cluster" AssumptionabstractSemi-supervised clustering (SSC) aims to improve clustering performance with the support of prior knowledge (i.e., side information). Compared with pairwise constraints, the partial labeling information is more natural to characterize the data distribution in a high level. However, the natural gap between the class information and the clustering is not adequately taken into account in exiting SSC methods when utilizing partial labeling information to guide the clustering procedure. In order to address this problem, we present a compact-cluster assumption for SSC to utilize the partial labeling information via a cluster-splitting technique. Based on this assumption, a general framework, CSSC, is proposed to supervise the traditional clustering with an objective function which is defined by incorporating an item to measure the compact degree of clusters. Furthermore, we provide two effective solutions for Kmeans and spectral clustering within the CSSC framework and derive the corresponding algorithms to seek the optimum number of clusters and their centroids. Corresponding theoretical analyses demonstrate the feasibility and effectivity of the proposed method. Finally, the extensive experiments on eight real-world datasets demonstrate the superiority of our method over other state-of-the-art SSC methods. Yongzhao Zhan 0001, Qirong Mao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Deep Robust Low Rank Correlation With Unifying Clustering Structure for Cross Domain AdaptationabstractCross domain adaptation aims to improve the performance of the target domain model by making full use of information rich source domain samples. However, as information becomes richer, the noise also increases. In order to improve the reliability of cross domain adaptation, we propose a novel method based on deep robust low rank correlation. Borrowed from the traditional idea of Canonical Correlation Analysis (CCA), we developed a robust correlation model to maximize the correlation between source and target domains. Also, the low-rank characteristics of cross domain data can effectively reduce the negative influence of noisy data. Furthermore, in order that the cross-domain data can share a unifying clustering structure, we introduced a common Laplacian affinity structure. Then the learned features can be smoothed and aligned to the unifying structure. In this way, we obtain a deep robust low rank correlation model with the help of the unifying clustering structure, which can effectively reduce the influence of noise and improve the performance of cross domain adaptation. Experimental results on three datasets including Office-31, ImageCLEF-DA and Office-Home show that our model significantly outperforms state-of-the-art cross domain adaptation methods. Xiangjun Shen, Yanan Cai, Stanley Ebhohimhen Abhadiomhen, Yongzhao Zhan 0001, Jianping Fan 0007 |
IEEE Trans. Multim. | 5 |
| 2023 | $L_{1}$-Regularized Reconstruction Model for Edge-Preserving FilteringabstractSmoothing images while preserving salient edges is a crucial task in computational photography. Existing edge-preserving filters suffer from various artifacts, such as halos, gradient reversals, and intensity shifts. Observing that various artifacts are strongly related to salient edges with large gradients, we propose a continuous mapping function to process the gradients. The proposed function is literally edge-preserving, i.e., it keeps large gradients intact while attenuating small gradients. We propose an L1-regularized reconstruction model based on the processed gradients for edge-preserving image filtering. The L1-regularization facilitates the edge-preserving property in the reconstructed results. To solve the proposed L1-regularized model, we implement an efficient algorithm based on the alternating direction method of multipliers (ADMM) and Fourier domain optimization. We have conducted qualitative and quantitative experiments to evaluate the proposed filter. The results demonstrate that our filter better handles various artifacts and delivers superior image quality on various applications. The proposed filter is highly efficient, our GPU implementation takes 70ms to process a color image with 1 megapixel on an NVIDIA GTX 1070 GPU. Yang Yang 0046, Lanling Zeng, Xiangjun Shen, Yongzhao Zhan 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Deep Weighted Guided Upsampling Network for Depth of Field Image UpsamplingabstractDepth-of-field (DoF) rendering is an important technique in computational photography that simulates the human visual attention system. Existing DoF rendering methods usually suffer from a high computational cost. The task of DoF rendering can be accelerated by guided upsampling methods. However, the state-of-the-art guided upsampling methods fail to distinguish the focus and defocus areas, resulting in unsatisfying DoF effects. In this paper, we propose a novel deep weighted guided upsampling network (DWGUN) based on a encoder and decoder framework to jointly upsample the low-resolution DoF image under the guidance of the corresponding high-resolution all-in-focus image. Due to the intuitive weight design, the traditional weighted image upsampling is not tailored to DoF image upsampling. We propose a deep refocus-defocus edge-aware module (DREAM) to learn the spatially-varying weights and embed them in the deep weighted guided upsampling block (DWGUB). We have conducted comprehensive experiments to evaluate the proposed method. Rigorous ablation studies are also conducted to validate the rationality of the proposed components. Lanling Zeng, Lianxiong Wu, Yang Yang 0046, Xiangjun Shen, Yongzhao Zhan 0001 |
MMAsia | 5 |
| 2022 | A representation coefficient-based k-nearest centroid neighbor classifier
Jianping Gou, Liyuan Sun 0004, Lan Du 0002, Hongxing Ma, Taisong Xiong, Weihua Ou, Yongzhao Zhan 0001 |
Expert Syst. Appl. | 7 |
| 2022 | Robust low-rank representation via residual projection for image classification
Kaifa Hui, Xiangjun Shen, Stanley Ebhohimhen Abhadiomhen, Yongzhao Zhan 0001 |
Knowl. Based Syst. | 4 |
| 2022 | Spatial-temporal correlations learning and action-background jointed attention for weakly-supervised temporal action localization
Huifen Xia 0001, Yongzhao Zhan 0001, Keyang Cheng |
Multim. Syst. | 2 |
| 2022 | Edge-Preserving Image Filtering Based on Soft ClusteringabstractEdge-preserving image filtering is an essential task in computational photography and imaging. In this paper, we propose a simple yet effective global edge-preserving filter based on soft clustering, and we propose a novel soft clustering algorithm based on a restricted Gaussian mixture model. Given specified parameters, the soft clustering process is firstly performed on the image to derive the partition matrix, from which the affinity matrix is then constructed for filtering. The filtering output is calculated as the weighted average of the pixels in the local window, so the proposed filter could suppress the intensity shift artifacts that impede most global filters. Besides, the weights in the proposed filter are derived by clustering, which properly separates dissimilar pixels, so the proposed filter could handle the halo artifacts that haunt many local filters. Moreover, our filter provides flexible control over the amount of smoothing that is deficient in the deep learning-based filters. Besides the efficacy in smoothing, the proposed filter naturally has low computational complexity. Qualitative and quantitative results suggest that the proposed filter benefits various applications, including edge-preserving smoothing, image enhancing, flash/non-flash fusion, HDR tone mapping, and dehazing. Yang Yang 0046, Hongjun Hui, Lanling Zeng, Yan Zhao 0038, Yongzhao Zhan 0001, Tao Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Class mean-weighted discriminative collaborative representation for classificationabstractRepresentation-based classification (RBC) has been attracting a great deal of attention in pattern recognition. As a typical extension to RBC, collaborative representation-based classification (CRC) has demonstrated its superior performance in various image classification tasks. Ideally, we expect that the learned class-specific representations for a testing sample are discriminative, and the representation computed for the true class dominates the final representation of the testing sample. Most existing CRC-based methods can learn pattern discrimination, but cannot differentiate the contribution of class-specific representations to the classification of each testing sample. It is challenging for a representation-based classifier to retain both properties. To address this challenge and further improve CRC's classification performance, we propose a novel CRC-based method, class mean-weighted discriminative collaborative representation-based classifier (CMW-DCRC). Its objective function penalises the standard l 2 -norm residuals with two discriminative regularisation terms. A decorrelating term makes the class-specific representations more discriminative, and a newly designed class mean-weighted term that promotes the training samples from individual classes to competitively reconstruct the testing sample while boosting the contribution of the true class. To further enhance the robustness of CRC, we extend CMW-DCRC by replacing the l2-norm coding residual with a l1-norm coding residual, and solve the optimisation problem with an iteratively reweighted least square algorithm. Extensive experimental results on nine image data sets have shown that our methods outperform the state-of-the-art RBC-based methods. Jianping Gou, Lan Du 0002, Shaoning Zeng, Yongzhao Zhan 0001, Zhang Yi 0001 |
Int. J. Intell. Syst. | 5 |
| 2021 | Robust deflated canonical correlation analysis via feature factoring for multi-view image classification
Kaifa Hui, Ernest Domanaanmwi Ganaa, Yongzhao Zhan 0001, Xiangjun Shen |
Multim. Tools Appl. | 3 |
| 2020 | Discriminative globality and locality preserving graph embedding for dimensionality reduction
Jianping Gou, Zhang Yi 0001, Jiancheng Lv 0001, Qirong Mao, Yongzhao Zhan 0001 |
Expert Syst. Appl. | 6 |
| 2020 | A multi-kernel method of measuring adaptive similarity for spectral clustering
Augustine Monney, Yongzhao Zhan 0001, Ben-Bright Benuwa |
Expert Syst. Appl. | 2 |
| 2020 | Hierarchical attributes learning for pedestrian re-identification via parallel stochastic gradient descent combined with momentum correction and adaptive learning rate
Keyang Cheng, Yongzhao Zhan 0001, Maozhen Li 0001, Kenli Li 0001 |
Neural Comput. Appl. | 3 |
| 2020 | Double graphs-based discriminant projections for dimensionality reduction
Jianping Gou, Ya Xue, Hongxing Ma, Yongzhao Zhan 0001, Jia Ke |
Neural Comput. Appl. | 5 |
| 2019 | Video semantic analysis based kernel locality-sensitive discriminative sparse representation
Ben-Bright Benuwa, Yongzhao Zhan 0001, Augustine Monney, Benjamin Ghansah, Ernest K. Ansah |
Expert Syst. Appl. | 2 |
| 2019 | Locality constrained representation-based K-nearest neighbor classification
Jianping Gou, Wenmo Qiu, Zhang Yi 0001, Xiangjun Shen, Yongzhao Zhan 0001, Weihua Ou |
Knowl. Based Syst. | 5 |
| 2019 | Multimodal shared features learning for emotion recognition by enhanced sparse local discriminative canonical correlation analysis
Jiamin Fu, Qirong Mao, Juanjuan Tu, Yongzhao Zhan 0001 |
Multim. Syst. | 4 |
| 2019 | Group sparse based locality - sensitive dictionary learning for video semantic analysis
Ben-Bright Benuwa, Yongzhao Zhan 0001, Jianping Gou, Benjamin Ghansah, Ernest K. Ansah |
Multim. Tools Appl. | 2 |
| 2019 | Face detection based on multilayer feed-forward neural network and Haar featuresabstractSummary Fast and accurate detection of a facial data is crucial for both face and facial expression recognition systems. These systems include internet protocol video surveillance systems, crime scene photographs systems, and criminals' databases. The aim for this study is both improvement of accuracy and speed. The salient facial features are extracted through Haar techniques. The sizes of the images are reduced by Bessel down‐sampling algorithm. This method preserved the details and perceptual quality of the original image. Then, image normalization was done by anisotropic smoothing. Multilayer feed‐forward neural network with a back‐propagation algorithm was used as classifier. A detection accuracy of 98.5% with acceptable false positives was registered with test sets from FDDB, CMU‐MIT, and Champions databases. The speed of execution was also promising. An evaluation of the proposed method with other popular detectors on the FDDB set shows great improvement. Ebenezer Owusu, Jamal-Deen Abdulai, Yongzhao Zhan 0001 |
Softw. Pract. Exp. | 3 |
| 2019 | A Local Mean Representation-based K-Nearest Neighbor ClassifierabstractK -nearest neighbor classification method (KNN), as one of the top 10 algorithms in data mining, is a very simple and yet effective nonparametric technique for pattern recognition. However, due to the selective sensitiveness of the neighborhood size k , the simple majority vote, and the conventional metric measure, the KNN-based classification performance can be easily degraded, especially in the small training sample size cases. In this article, to further improve the classification performance and overcome the main issues in the KNN-based classification, we propose a local mean representation-based k -nearest neighbor classifier (LMRKNN). In the LMRKNN, the categorical k -nearest neighbors of a query sample are first chosen to calculate the corresponding categorical k -local mean vectors, and then the query sample is represented by the linear combination of the categorical k -local mean vectors; finally, the class-specific representation-based distances between the query sample and the categorical k -local mean vectors are adopted to determine the class of the query sample. Extensive experiments on many UCI and KEEL datasets and three popular face databases are carried out by comparing LMRKNN to the state-of-art KNN-based methods. The experimental results demonstrate that the proposed LMRKNN outperforms the related competitive KNN-based methods with more robustness and effectiveness. Jianping Gou, Wenmo Qiu, Zhang Yi 0001, Yong Xu 0001, Qirong Mao, Yongzhao Zhan 0001 |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2019 | An emotion-based responding model for natural language conversation
Qirong Mao, Liangjun Wang, Nelson Ruwa, Jianping Gou, Yongzhao Zhan 0001 |
World Wide Web | 6 |
| 2018 | ALSTM: Adaptive LSTM for Durative Sequential DataabstractLong short-term memory (LSTM) network is an effective model architecture for deep learning approaches to sequence modeling tasks. However, the current LSTMs can't use the property of sequential data when dealing with the sequence components, which last for a certain period of time. This may make the model unable to benefit from the inherent characteristics of time series and result in poor performance as well as lower efficiency. In this paper, we present a novel adaptive LSTM for durative sequential data which exploits the temporal continuance of the input data in designing a new LSTM unit. By adding a new mask gate and maintaining span, the cell's memory update is not only determined by the input data but also affected by its duration. An adaptive memory update method is proposed according to the change of the sequence input at each time step. This breaks the limitation that the cells calculate the cell state and hidden output for each input always in a unified manner, making the model more suitable for processing the sequences with continuous data. The experimental results on various sequence training tasks show that under the same iteration epochs, the proposed method can achieve higher accuracy, but need relatively less training time compared with the standard LSTM architecture. DeJiao Niu, Zheng Xia, Tao Cai 0003, Tianquan Liu, Yongzhao Zhan 0001 |
ICTAI | 6 |
| 2018 | Least squares kernel ensemble regression in Reproducing Kernel Hilbert Space
Xiangjun Shen, Yong Dong, Jianping Gou, Yongzhao Zhan 0001, Jianping Fan 0001 |
Neurocomputing | 4 |
| 2018 | Two-phase linear reconstruction measure-based classification for face recognition
Jianping Gou, Yong Xu 0001, David Zhang 0001, Qirong Mao, Lan Du 0002, Yongzhao Zhan 0001 |
Inf. Sci. | 6 |
| 2018 | Discriminative self-adapted locality-sensitive sparse representation for video semantic analysis
Jianping Gou, Yongzhao Zhan 0001, Qirong Mao |
Multim. Tools Appl. | 3 |
| 2018 | Spatially Coherent Feature Learning for Pose-Invariant Facial Expression RecognitionabstractFeature learning has enjoyed much attention and achieved good performance in recent studies of image processing. Unlike the required training conditions often assumed there, far less labeled data is available for training emotion classification systems. In addition, current feature learning is typically performed on an entire face image without considering the dependency between features. These approaches ignore the fact that faces are structured and the neighboring features are dependent. Thus, the learned features lack the power to describe visually coherent facial images. Our method is therefore designed with the goal of simplifying the problem domain by removing expression-irrelevant factors from the input images, with a key region-based mechanism, which is an effort to reduce the amount of data required to effectively train the feature-learning methods. Meanwhile, we can construct geometric constraints between the key regions and its detected positions. To this end, we introduce a Spatially Coherent featurelearning method for Pose-invariant Facial Expression Recognition (SC-PFER). In our model, we first perform face frontalization through a 3D pose-normalization technique, which could normalize poses while preserving the identity information through synthesizing frontal faces for facial images with arbitrary views. Subsequently, we select a sequence of key regions around 51 key points in the synthetic frontal face images for efficient unsupervised feature learning. Finally, we introduce a linkage structure over the learning-based features and the corresponding geometry information of each key region to encode the dependencies of the regions. Our method, on the whole, does not require training multiple models for each specific pose and avoids separating training and parameter tuning for each pose. The proposed framework has been evaluated on two benchmark databases, BU-3DFE and SFEW, for pose-invariant Facial Expression Recognition (FER). The experimental results demonstrate that our algorithm outperforms current state-of-the-art FER methods. Specifically, our model achieves an improvement of 1.72% and 1.11% FER accuracy, on average, on BU-3DFE and SFEW, respectively. Feifei Zhang 0001, Qirong Mao, Xiangjun Shen, Yongzhao Zhan 0001, Ming Dong 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | A Multi-local Means Based Nearest Neighbor ClassifierabstractIn this paper, we propose a multi-local means based nearest neighbor classifier (MLMNN). In the MLMNN, k categorical nearest neighbors of a query sample are first found and used to calculate the corresponding k categorical multi-local mean vectors which can represent different local class-specific sample distributions. Then, the query sample is represented by a linear combination of k categorical local mean vectors and the representation coefficient of each local mean vector as the contribution to representing and classifying the query sample is obtained. Finally, the class-specific representation-based distance (i.e. reconstruction residual) between the query sample and k categorical multi-local mean vectors is adopted to determine the class label of the query sample. The experimental results on three popular face databases show that the proposed MLMNN method outperforms the related competitive KNN-based methods. Jianping Gou, Wenmo Qiu, Qirong Mao, Yongzhao Zhan 0001, Xiangjun Shen, Yunbo Rao |
ICTAI | 4 |
| 2017 | AL-DDCNN: a distributed crossing semantic gap learning for person re-identificationabstractSummary By the reason of the variability of light and pedestrians' appearance, it is hard for a camera to obtain a clear human figure. Person re‐identification with different cameras is a difficult visual recognition task. In this paper, a novel approach called attribute learning based on distributed deep convolutional neural network model is proposed to address person re‐identification task. It shows how attributes, namely the mid‐level medium between classes and features, are obtained automatically and how they are employed to re‐identify person with semantics when an author‐topic model is used to mapping category. Besides, considering the ability to operate on raw pixel input without the need to design special features, deep convolutional neural network is employed to generate features without supervision for attributes learning model. To overcome the model's weakness in computing speed, parallelized implementations such as distributed parameter manipulation and attributes learning are employed in attribute learning based on distributed deep convolutional neural network model. Experiments show that the proposed approach achieves state‐of‐the‐art recognition performance in the VIPeR data set and is with a good semantic explanation. Copyright © 2016 John Wiley & Sons, Ltd. Keyang Cheng, Yongzhao Zhan 0001, Man Qi |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Sparse representations based distributed attribute learning for person re-identification
Keyang Cheng, Kaifa Hui, Yongzhao Zhan 0001, Maozhen Li 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Unsupervised domain adaptation for speech emotion recognition using PCANet
Zhengwei Huang, Wentao Xue, Qirong Mao, Yongzhao Zhan 0001 |
Multim. Tools Appl. | 4 |
| 2017 | Learning emotion-discriminative and domain-invariant features for domain adaptation in speech emotion recognition
Qirong Mao, Guopeng Xu, Wentao Xue, Jianping Gou, Yongzhao Zhan 0001 |
Speech Commun. | 5 |
| 2016 | Domain adaptation for speech emotion recognition by sharing priors between related source and target classesabstractIn speech emotion recognition (SER), speech data is usually captured from different scenarios, which often leads to significant performance degradation due to the inherent mismatch between training and test set. To cope with this problem, we propose a domain adaptation method called Sharing Priors between Related Source and Target classes (SPRST) based on a two-layer neural network. The classifier parameters, namely the weights of the second layer, are imposed the common priors between the related classes, so that the classes with few labeled data in target domain can borrow knowledge from the related classes in source domain. The method is evaluated on the INTERSPEECH 2009 Emotion Challenge two-class task. Experimental results show that our approach significantly improves the performance when only a small number of target labeled instances are available. Qirong Mao, Wentao Xue, Qiyu Rao, Feifei Zhang 0001, Yongzhao Zhan 0001 |
ICASSP | 5 |
| 2016 | Multi-pose Facial Expression Recognition Using Transformed Dirichlet ProcessabstractDriven by recent advances in human-centered computing, Facial Expression Recognition (FER) has attracted significant attention in many applications. In this paper, we propose a novel graphical model, multi-level Transformed Dirichlet Process (ml-TDP), for multi-pose FER. In our approach, pose is explicitly introduced into ml-TDP so that separate training and parameter tuning for each pose is not required. In addition, ml-TDP can learn an intermediate facial expression representation subject to geometric constraints. By sharing the pool of spatially-coherent features over expressions and poses, we provide a scalable solution for multi-pose FER. Extensive experimental result on benchmark facial expression databases shows the superior performance of ml-TDP. Feifei Zhang 0001, Qirong Mao, Ming Dong 0001, Yongzhao Zhan 0001 |
ACM Multimedia | 4 |
| 2016 | Pose-robust feature learning for facial expression recognition
Feifei Zhang 0001, Qirong Mao, Jianping Gou, Yongzhao Zhan 0001 |
Frontiers Comput. Sci. | 5 |
| 2016 | A video semantic detection method based on locality-sensitive discriminant sparse representation and weighted KNN
Yongzhao Zhan 0001, Jianping Gou, Minchao Wang |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | A Socioecological Model for Advanced Service Discovery in Machine-to-Machine Communication NetworksabstractThe new development of embedded systems has the potential to revolutionize our lives and will have a significant impact on future Internet of Thing (IoT) systems if required services can be automatically discovered and accessed at runtime in Machine-to-Machine (M2M) communication networks. It is a crucial task for devices to perform timely service discovery in a dynamic environment of IoTs. In this article, we propose a Socioecological Service Discovery (SESD) model for advanced service discovery in M2M communication networks. In the SESD network, each device can perform advanced service search to dynamically resolve complex enquires and autonomously support and co-operate with each other to quickly discover and self-configure any services available in M2M communication networks to deliver a real-time capability. The proposed model has been systematically evaluated and simulated in a dynamic M2M environment. The experiment results show that SESD can self-adapt and self-organize themselves in real time to generate higher flexibility and adaptability and achieve a better performance than the existing methods in terms of the number of discovered service and a better efficiency in terms of the number of discovered services per message. Lu Liu 0001, Nick Antonopoulos, Minghui Zheng, Yongzhao Zhan 0001, Zhijun Ding |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2015 | Multi-pose facial expression recognition based on SURF boostingabstractToday Human Computer Interaction (HCI) is one of the most important topics in machine vision and image processing fields. The ability to handle multi-pose facial expressions is important for computers to understand affective behavior under less constrained environment. In this paper, we propose a SURF (Speeded-Up Robust Features) boosting framework to address challenging issues in multi-pose facial expression recognition (FER). Local SURF features from different overlapping patches are selected by boosting in our model to focus on more discriminable representations of facial expression. And this paper proposes a novel training step during boosting. The experiments using the proposed method demonstrate favorable results on RaFD and KDEF databases. Qiyu Rao, Xing Qu, Qirong Mao, Yongzhao Zhan 0001 |
ACII | 4 |
| 2015 | A Video Semantic Analysis Method Based on Kernel Discriminative Sparse Representation and Weighted KNNabstractTo improve the video semantic analysis for video surveillance, a new video semantic analysis method based on the kernel discriminative sparse representation (KSVD) and weighted K nearest neighbors (KNN) is proposed in this paper. A discriminative model is built by introducing a kernel discriminative function to the KSVD dictionary optimization algorithm, mapping the sparse representation features into a high-dimensional space. The optimal dictionary is then generated and applied to compute the sparse representations of video features. For video semantic analysis, a weighted KNN algorithm based on the optimal sparse representation is proposed. In the algorithm, a kernel function is introduced to establish discrimination about sparse representation features and the classification vote result is weighted, the purpose of which is to improve the accuracy and rationality for video semantic analysis. The experimental results show that the proposed method significantly improves the discrimination of sparse representation features when compared with the traditional KSVD-based support vector machine method. The method can effectively detect the concept and event, which can be potentially useful for improving the video surveillance. Yongzhao Zhan 0001, Shan Dai, Qirong Mao, Lu Liu 0001, Wei Sheng |
Comput. J. | 1 |
| 2015 | Using Kinect for real-time emotion recognition via facial expressionsabstractEmotion recognition via facial expressions (ERFE) has attracted a great deal of interest with recent advances in artificial intelligence and pattern recognition. Most studies are based on 2D images, and their performance is usually computationally expensive. In this paper, we propose a real-time emotion recognition approach based on both 2D and 3D facial expression features captured by Kinect sensors. To capture the deformation of the 3D mesh during facial expression, we combine the features of animation units (AUs) and feature point positions (FPPs) tracked by Kinect. A fusion algorithm based on improved emotional profiles (IEPs) and maximum confidence is proposed to recognize emotions with these real-time facial expression features. Experiments on both an emotion dataset and a real-time video show the superior performance of our method. Qirong Mao, Yongzhao Zhan 0001, Xiangjun Shen |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2015 | Performance evaluation and simulation of peer-to-peer protocols for Massively Multiplayer Online Games
Lu Liu 0001, Nick Antonopoulos, Zhijun Ding, Yongzhao Zhan 0001 |
Multim. Tools Appl. | 5 |
| 2015 | A semi-supervised incremental learning method based on adaptive probabilistic hypergraph for video semantic detection
Yongzhao Zhan 0001, Jiayao Sun, DeJiao Niu, Qirong Mao, Jianping Fan 0001 |
Multim. Tools Appl. | 1 |
| 2014 | Speech Emotion Recognition Using CNNabstractDeep learning systems, such as Convolutional Neural Networks (CNNs), can infer a hierarchical representation of input data that facilitates categorization. In this paper, we propose to learn affect-salient features for Speech Emotion Recognition (SER) using semi-CNN. The training of semi-CNN has two stages. In the first stage, unlabeled samples are used to learn candidate features by contractive convolutional neural network with reconstruction penalization. The candidate features, in the second step, are used as the input to semi-CNN to learn affect-salient, discriminative features using a novel objective function that encourages the feature saliency, orthogonality and discrimination. Our experiment results on benchmark datasets show that our approach leads to stable and robust recognition performance in complex scenes (e.g., with speaker and environment distortion), and outperforms several well-established SER features. Zhengwei Huang, Ming Dong 0001, Qirong Mao, Yongzhao Zhan 0001 |
ACM Multimedia | 4 |
| 2014 | An SVM-AdaBoost facial expression recognition system
Ebenezer Owusu, Yongzhao Zhan 0001, Qirong Mao |
Appl. Intell. | 2 |
| 2014 | A neural-AdaBoost based facial expression recognition systemabstractThis study improves the recognition accuracy and execution time of facial expression recognition system. Various techniques were utilized to achieve this. The face detection component is implemented by the adoption of Viola–Jones descriptor. The detected face is down-sampled by Bessel transform to reduce the feature extraction space to improve processing time then. Gabor feature extraction techniques were employed to extract thousands of facial features which represent various facial deformation patterns. An AdaBoost-based hypothesis is formulated to select a few hundreds of the numerous extracted features to speed up classification. The selected features were fed into a well designed 3-layer neural network classifier that is trained by a back-propagation algorithm. The system is trained and tested with datasets from JAFFE and Yale facial expression databases. An average recognition rate of 96.83% and 92.22% are registered in JAFFE and Yale databases, respectively. The execution time for a 100 × 100 pixel size is 14.5 ms. The general results of the proposed techniques are very encouraging when compared with others. Ebenezer Owusu, Yongzhao Zhan 0001, Qirong Mao |
Expert Syst. Appl. | 2 |
| 2014 | An SVM-AdaBoost-based face detection systemabstractFace detection is the first significant step in face recognition and many computer vision applications. The goal of this work was to improve detection accuracy as well as reducing the execution time. Images are pre-processed, scaled and normalised with the discrete cosine transform. Gabor feature extraction techniques were employed to extract thousands of facial vectors. An AdaBoost-based feature selection tool was formulated to select a few hundreds of the Gabor wavelets. These vectors representing significant salient local features are used as input vectors to a support vector machine classifier. The classifier is trained and becomes capable of detecting faces. A detection rate of 97.6% with acceptable false positives was registered with a test set of 507 faces. The execution time of a pixel of size 320 × 240 is 0.0285 s, which is very promising. A comparative evaluation of receiver operating characteristic (ROC) curves of different detectors on FDDB set shows that the proposed method is very effective. Ebenezer Owusu, Yongzhao Zhan 0001, Qirong Mao |
J. Exp. Theor. Artif. Intell. | 2 |
| 2014 | Improved pseudo nearest neighbor classification
Jianping Gou, Yongzhao Zhan 0001, Yunbo Rao, Xiangjun Shen, Wu He |
Knowl. Based Syst. | 2 |
| 2014 | Learning Salient Features for Speech Emotion Recognition Using Convolutional Neural NetworksabstractAs an essential way of human emotional behavior understanding, speech emotion recognition (SER) has attracted a great deal of attention in human-centered signal processing. Accuracy in SER heavily depends on finding good affect- related , discriminative features. In this paper, we propose to learn affect-salient features for SER using convolutional neural networks (CNN). The training of CNN involves two stages. In the first stage, unlabeled samples are used to learn local invariant features (LIF) using a variant of sparse auto-encoder (SAE) with reconstruction penalization. In the second step, LIF is used as the input to a feature extractor, salient discriminative feature analysis (SDFA), to learn affect-salient, discriminative features using a novel objective function that encourages feature saliency, orthogonality, and discrimination for SER. Our experimental results on benchmark datasets show that our approach leads to stable and robust recognition performance in complex scenes (e.g., with speaker and language variation, and environment distortion) and outperforms several well-established SER features. Qirong Mao, Ming Dong 0001, Zhengwei Huang, Yongzhao Zhan 0001 |
IEEE Trans. Multim. | 4 |
| 2013 | Speaker-independent speech emotion recognition by fusion of functional and accompanying paralanguage featuresabstractFunctional paralanguage includes considerable emotion information, and it is insensitive to speaker changes. To improve the emotion recognition accuracy under the condition of speaker-independence, a fusion method combining the functional paralanguage features with the accompanying paralanguage features is proposed for the speaker-independent speech emotion recognition. Using this method, the functional paralanguages, such as laughter, cry, and sigh, are used to assist speech emotion recognition. The contributions of our work are threefold. First, one emotional speech database including six kinds of functional paralanguage and six typical emotions were recorded by our research group. Second, the functional paralanguage is put forward to recognize the speech emotions combined with the accompanying paralanguage features. Third, a fusion algorithm based on confidences and probabilities is proposed to combine the functional paralanguage features with the accompanying paralanguage features for speech emotion recognition. We evaluate the usefulness of the functional paralanguage features and the fusion algorithm in terms of precision, recall, and F1-measurement on the emotional speech database recorded by our research group. The overall recognition accuracy achieved for six emotions is over 67% in the speaker-independent condition using the functional paralanguage features. Qirong Mao, Xiao-lei Zhao, Zhengwei Huang, Yongzhao Zhan 0001 |
J. Zhejiang Univ. Sci. C | 4 |
| 2012 | An Investigation into the Evolution of Security Usage in Home Wireless NetworksabstractWireless networks are an integral part of life for many residential properties. The use of laptops and smartphones has lead to a large increase in the number of wireless networks. The research within this project revealed how residential users were securing their networks, this research mirrored a previous investigation into wireless security from 2009, as well as a comparison with other authors work dating back to 2006 and 2007. These findings were then analyzed and reasons behind the positive trend in encryption utilization were examined. Thomas Stimpson, Lu Liu 0001, Yongzhao Zhan 0001 |
TrustCom | 3 |
| 2012 | A Secure Node Localization Method Based on the Congruity of Time in Wireless Sensor NetworksabstractWith the development of the theory and technology of wireless sensor networks (WSN), location-based applications, such as location-based access control, pose new challenges. In order to improve the accuracy of node localization, the energy consumption must be considered in conjunction with security. Improving the accuracy of node localization as much as possible under the premise of ensuring the security of the localization process forms the basis of location-based applications in wireless sensor networks. In this paper, a secure localization method of nodes based on the congruity of time is proposed. This method does not need to meet time synchronization conditions between user nodes and base stations. It calculates the congruity of time according to the communication delay between nodes, and then estimates the location of the user node. It can ensure the accuracy of localization and the security of localization processes. Yongzhao Zhan 0001, Lu Liu 0001, Hussain Al-Aqrabi |
TrustCom | 1 |
| 2010 | A New Classifier for Facial Expression Recognition: Fuzzy Buried Markov Model
Yongzhao Zhan 0001, Keyang Cheng, Ya-Bi Chen, Chuan-Jun Wen |
J. Comput. Sci. Technol. | 1 |
| 2008 | Application Research of Ontology in E-Learning EnvironmentabstractIn this paper, awareness and situation ontology model is presented to process awareness and learning situations in E-Learning environment. By this model, knowledge domain, awareness information and learning situations can be described reasonably and efficiently. In addition, reasoning rules for awareness and situation are given in this model. According to these reasoning rules, the learnerpsilas awareness information and learning situations can be concluded. We design and realize an E-Learning system based on this ontology model. The experiment results show that learner's learning efficiency is raised greatly after awareness and situation ontology model is adopted. Qirong Mao, Yongzhao Zhan 0001 |
CW | 3 |
| 2005 | The shared knowledge space model in Web-based cooperative learning coalitionabstractIn order to make learners share learning data in multi-Web sites, in this paper, based on nested knowledge space model, Web-based collaborative learning coalition is defined, and the shared knowledge space model is put forward, by which the learning data in multi-Web sites can be organized together and formed the shared knowledge space. Each Web site in the coalition is able to customize the shared knowledge space according to its needs, then the nested customized knowledge space is formed in the local site integrated with the local learning data. Furthermore, the management and consistency maintenance strategy for the shared knowledge space and the nested customized knowledge space is proposed. In the end, the Web-based collaborative learning coalition is implemented according to the model and strategy mentioned above, and the application effect shows that the data in multi-Web sites in the coalition is shared and their different needs for learning data are met. Qirong Mao, Yongzhao Zhan 0001, Shunlin Song |
CSCWD (2) | 2 |
| 2004 | Design and Simulation of Multicast Routing Protocol for Mobile Internet
Guangsheng Li, Qirong Mao, Yongzhao Zhan 0001, Yibin Hou |
APWeb | 4 |
| 2004 | Facial expression recognition based on Gabor wavelet transformation and elastic templates matchingabstractIn order to achieve subject-independent facial expression recognition and obtain robustness against illumination variety and image deformation, facial expression recognition methods based on Gabor wavelet transformation and elastic templates matching are presented in this paper. Firstly, given a still image containing facial expression information, preprocessors are executed. Secondly, Gabor wavelets are adopted to extract expression features. Then the elastic graph for expression features is constructed. Finally, elastic templates matching algorithm is used to recognize facial expression. Experiments show that the expression features can be extracted effectively by Gabor wavelet transformation and high recognition rate can be obtained using elastic templates matching algorithm. Yongzhao Zhan 0001, Jingfu Ye, DeJiao Niu |
ICIG | 1 |
| 2004 | DRMR: Dynamic-Ring-Based Multicast Routing Protocol for Ad Hoc Networks
Guangsheng Li, Yongzhao Zhan 0001, Qirong Mao, Yibin Hou |
J. Comput. Sci. Technol. | 3 |