VLDB 2026 Research / reviewers in the wild / expert
Zhijing Yang
dblp:57/7614
· DBLP profile ↗
82ranked-venue papers
15as first author
61since 2021 · last 2026
0000-0001-8336-5109ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 6 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 21 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 5 · 2 first-author · 1 since 2021Computer networks · 4 · 4 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal feature disentangle-fusion network for detecting and grounding multi-modal media manipulation
Jinghui Qin, Tianshui Chen, Zhijing Yang |
Expert Syst. Appl. | 4 |
| 2026 | Multimodal progressive fusion and enhancement network for multimodal sentiment analysis
Jinghui Qin, Qite Zhou, Lihuang Fang, Zhijing Yang |
Expert Syst. Appl. | 5 |
| 2026 | Learning semantic-aware threshold for multi-label image recognition with partial labels
Haoxian Ruan, Zhihua Xu, Zhijing Yang, Guang Ma, Jieming Xie, Changxiang Fan, Tianshui Chen |
Expert Syst. Appl. | 3 |
| 2026 | Neural clothing tryer: Customized virtual try-on via semantic enhancement and controlling diffusion model
Zhijing Yang, Yukai Shi, Junpeng Tan, Tianshui Chen, Liruo Zhong |
Expert Syst. Appl. | 1 |
| 2026 | Webly supervised multi-label recognition: Evaluation benchmark and Dual-Branch Multi-Label Contrastive Learning
Zhihua Xu, Zhijing Yang, Tianshui Chen |
Image Vis. Comput. | 2 |
| 2026 | Exploring label co-occurrence metric and graph contrastive learning method for multi-label image recognition with partial labels
Zhijing Yang, Yu Cheng 0010, Jing Ling, Haoxian Ruan, Yongyi Lu |
Knowl. Based Syst. | 2 |
| 2026 | Revisiting DIRE: towards universal AI-generated image detection
Huanqi Lin, Jinghui Qin, Xiaoqi Wu, Tianshui Chen, Zhijing Yang |
Neural Networks | 5 |
| 2026 | Adversarial incomplete multi-view clustering with adaptive contrastive learning
Shuzhao Xu, Zhijing Yang, Feiping Nie 0001 |
Neural Networks | 3 |
| 2026 | Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression ManipulationabstractSpeech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated with spoken content. Current works depend on inaccessible paired training samples for the person, where two aligned frames exhibit the same speech content yet differ in emotional expression, limiting the SPFEM applications in real-world scenarios. In this work, we discover that speakers who convey the same content with different emotions exhibit highly correlated local facial animations in both spatial and temporal spaces, providing valuable supervision for SPFEM. To capitalize on this insight, we propose a novel spatial-temporal coherent correlation learning (STCCL) algorithm, which models the aforementioned correlations as explicit metrics and integrates the metrics to supervise manipulating facial expression and meanwhile better preserving the facial animation of spoken content. To this end, it first learns a spatial coherent correlation metric, ensuring that the visual correlations of adjacent local regions within an image linked to a specific emotion closely resemble those of corresponding regions in an image linked to a different emotion. Simultaneously, it develops a temporal coherent correlation metric, ensuring that the visual correlations of specific regions across adjacent image frames associated with one emotion are similar to those in the corresponding regions of frames associated with another emotion. Recognizing that visual correlations are not uniform across all regions, we have also crafted a correlation-aware adaptive strategy that prioritizes regions that present greater challenges. During SPFEM model training, we construct the spatial-temporal coherent correlation metric between corresponding local regions of the input and output image frames as additional loss to supervise the generation process. We conduct extensive experiments on various datasets, and the results demonstrate the effectiveness of the proposed STCCL algorithm. Tianshui Chen, Jianman Lin, Zhijing Yang, Chunmei Qing, Guangrun Wang, Liang Lin 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head Animation
Zhihua Xu, Tianshui Chen, Zhijing Yang, Chunmei Qing, Keze Wang, Liang Lin 0004 |
IEEE Trans. Multim. | 3 |
| 2026 | Exploring Talking Head Models with Adjacent Frame Prior for Speech-Preserving Facial Expression ManipulationabstractSpeech-Preserving Facial Expression Manipulation (SPFEM) is an innovative technique aimed at altering facial expressions in images and videos while retaining the original mouth movements. Despite advancements, SPFEM still struggles with accurate lip synchronization due to the complex interplay between facial expressions and mouth shapes. Capitalizing on the advanced capabilities of Audio-Driven Talking Head Generation (AD-THG) models in synthesizing precise lip movements, our research introduces a novel integration of these models with SPFEM. We present a new framework, Talking Head Facial Expression Manipulation (THFEM), which utilizes AD-THG models to generate frames with accurately synchronized lip movements from audio inputs and SPFEM-altered images. However, increasing the number of frames generated by AD-THG models tends to compromise the realism and expression fidelity of the images. To counter this, we develop an adjacent frame learning strategy that finetunes AD-THG models to predict sequences of consecutive frames. This strategy enables the models to incorporate information from neighboring frames, significantly improving image quality during testing. Our extensive experimental evaluations demonstrate that this framework effectively preserves mouth shapes during expression manipulations, highlighting the substantial benefits of integrating AD-THG with SPFEM. Zhenxuan Lu, Zhihua Xu, Zhijing Yang, Feng Gao 0014, Yongyi Lu, Keze Wang, Tianshui Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Adaptive Few-shot Prompting for Machine Translation with Pre-trained Language ModelsabstractRecently, Large Language Models (LLMs) with in-context learning have demonstrated remarkable potential in handling neural machine translation. However, existing evidence shows that LLMs are prompt-sensitive and it is sub-optimal to apply the fixed prompt to any input for downstream machine translation tasks. To address this issue, we propose an adaptive few-shot prompting (AFSP) framework to automatically select suitable translation demonstrations for various source input sentences to further elicit the translation capability of an LLM for better machine translation. First, we build a translation demonstration retrieval module based on LLM's embedding to retrieve top-k semantic-similar translation demonstrations from aligned parallel translation corpus. Rather than using other embedding models for semantic demonstration retrieval, we build a hybrid demonstration retrieval module based on the embedding layer of the deployed LLM to build better input representation for retrieving more semantic-related translation demonstrations. Then, to ensure better semantic consistency between source inputs and target outputs, we force the deployed LLM itself to generate multiple output candidates in the target language with the help of translation demonstrations and rerank these candidates. Besides, to better evaluate the effectiveness of our AFSP framework on the latest language and extend the research boundary of neural machine translation, we construct a high-quality diplomatic Chinese-English parallel dataset that consists of 5,528 parallel Chinese-English sentences. Finally, extensive experiments on the proposed diplomatic Chinese-English parallel dataset and the United Nations Parallel Corpus (Chinese-English part) show the effectiveness and superiority of our proposed AFSP. Jinghui Qin, Wenxuan Ye, Hao Tan 0007, Zhijing Yang |
AAAI | 5 |
| 2025 | Enhancing Fairness in Gaussian Mixture Clustering through Impact FactorabstractClustering is a common method used in machine learning to group sample points in a dataset. Gaussian Mixture Clustering (GMC) is a clustering method based on maximum likelihood estimation and expectation maximisation (EM) algorithms. Traditional GMC does not consider the fairness between different sensitive groups, leading to biased clustering results. In this paper, we propose a novel Fair Gaussian Mixture Clustering (FGMC) to solve the fairness problem in clustering tasks. We incorporate a fairness optimiser, the impact factor, into the GMC to ensure that the fairness of the clustering is gradually optimised over the iterations. Using FGMC ensures that the clustering results are not overly biased towards a particular group defined by sensitive attributes such as age or race. We evaluated FGMC on a real-world dataset and found that it significantly improved clustering fairness. FGMC is a promising direction for clustering that requires ethical considerations. Zhijing Yang, Chuan Qian, Yiding Tang, Boyang Yan, Hui Zhang 0055 |
ICASSP | 1 |
| 2025 | Individual Fairness for Fuzzy C-Means ClusteringabstractIn the field of clustering algorithms, the Fuzzy CMeans algorithm stands out for its ability to deal with uncertainty by assigning membership degrees to data points. However, research on the fairness of Fuzzy C-Means algorithms has mainly focused on group fairness, with limited attention to individual fairness. To fill this gap, this paper proposes an Individual fair Fuzzy C-Means algorithm. By establishing a relaxed Lipschitz condition as the theoretical foundation and incorporating the random walk Laplacian matrix constructed from the similarity matrix into the clustering process, the Individual Fair Fuzzy CMeans algorithm optimizes the individual fairness of clustering. Experimental results show that individual fairness is significantly improved while maintaining clustering quality comparable to traditional Fuzzy C-Means algorithms. Zhijing Yang, Boyang Yan, Yiding Tang, Chuan Qian, Hui Zhang 0055 |
ICASSP | 1 |
| 2025 | Game-Theoretic Optimization for Scale Fair Spectral Clustering
Zhijing Yang, Hui Zhang 0055 |
ICIC (12) | 2 |
| 2025 | DeltaDiff: Reality-Driven Diffusion with Anchor Residuals for Faithful SR
Zhijing Yang |
ICIC (2) | 5 |
| 2025 | Improving Fairness in Density Peak Clustering through Fair Constraints and Multi-Objective Optimization AllocationabstractClustering is a fundamental technique in data analysis and machine learning, essential for uncovering hidden patterns. Density Peak Clustering (DPC) is known for its ability to intuitively and rapidly discover hidden clusters in datasets. However, it lacks fairness considerations, leading to biased results. This issue is increasingly critical as fairness becomes a widely discussed and emphasized aspect of clustering algorithms. Notably, this paper introduces Fair Density Peak Clustering (FDPC). Inspired by the cut-off distance in DPC, FDPC introduces a novel distance metric called fair cut-off distance (fdc), derived using the Lagrange multiplier method to ensure minimal fairness loss. To enhance overall fairness, the fdc is used in both the selection of cluster centers and the point assignment strategy. Furthermore, the assignment strategy considers the constraint of a fairness deviation upper bound. Additionally, this paper integrates Pareto multi-objective optimization theory into the assignment strategy, considering both the distances of unassigned points to clusters and their fairness deviations. Comparative experiments demonstrate that FDPC significantly improves fairness while maintaining comparable clustering performance. The code for our work is available at https://github.com/HurryUp1234/FDPC. Zhijing Yang, Yiding Tang, Boyang Yan, Chuan Qian, Hui Zhang 0055 |
IJCNN | 1 |
| 2025 | Contrastive Decoupled Representation Learning and Regularization for Speech-Preserving Facial Expression Manipulation
Tianshui Chen, Jianman Lin, Zhijing Yang, Chumei Qing, Yukai Shi, Liang Lin 0004 |
Int. J. Comput. Vis. | 3 |
| 2025 | Individual fair fuzzy C-means clustering via density-adaptive spectral regularization
Boyang Yan, Zhijing Yang, Yiding Tang, Hui Zhang 0055 |
Neurocomputing | 2 |
| 2025 | Fair Laplace: A unified framework for fair spectral clustering
Zhijing Yang, Hui Zhang 0055, Chunming Yang, Bo Li 0065, Xujian Zhao, Yin Long |
Inf. Process. Manag. | 1 |
| 2025 | Spectral clustering with scale fairness constraints
Zhijing Yang, Hui Zhang 0055, Chunming Yang, Bo Li 0065, Xujian Zhao, Yin Long |
Knowl. Inf. Syst. | 1 |
| 2025 | Dual constraint based semi-supervised nonnegative matrix factorization for multi-view clustering
Zimeng Huangfu, Wenyun Xie, Zhijing Yang, Feiping Nie 0001 |
Knowl. Based Syst. | 4 |
| 2025 | Robust Semi-Supervised Deep Nonnegative Matrix Factorization With Constraint Propagation for Data RepresentationabstractDeep nonnegative matrix factorization (DNMF) technique has attached much great attention in recent year, since it can effectively discover the underlying hierarchical structure of complex data. However, most existing unsupervised and semi-supervised DNMF approaches not only suffer from the noisy data seriously, but also fail to enhance the decomposition quality of DNMF obviously by using the obtained supervisory information. To overcome these drawbacks, a robust semi-supervised DNMF method, called the correntropy based semi-supervised DNMF with constraint propagation (CSDCP), is proposed in this paper for learning a compact and meaningful data representation from the original data. Particularly, instead of adopting the traditional Frobenius norm, CSDCP employs the nonlinear and local similarity measure (e.g., correntropy) as the loss function in DNMF to enhance the robustness of DNMF for the noisy data. In addition, the hypergraph based constraint propagation (HCP) algorithm is adopted in CSDCP to exploit the limited supervisory information fully for capturing good data representation. Moreover, algorithm analysis of CSDCP is presented in this paper, including convergence analysis, robustness analysis supervised information analysis, and computational complexity. Extensive experimental results have illustrated that, in comparison to the most related DNMF approaches, CSDCP usually has better clustering results on six nonnegative datasets in clustering tasks. Jingxing Yin, Zhijing Yang, Feiping Nie 0001, Badong Chen |
IEEE Trans. Big Data | 3 |
| 2025 | Neural Scene Designer: Self-Styled Semantic Image ManipulationabstractMaintaining stylistic consistency is crucial for the cohesion and aesthetic appeal of images, a fundamental requirement in effective image editing and inpainting. However, existing methods primarily focus on the semantic control of generated content, often neglecting the critical task of preserving this consistency. In this work, we introduce the Neural Scene Designer (NSD), a novel framework that enables photo-realistic manipulation of user-specified scene regions while ensuring both semantic alignment with user intent and stylistic consistency with the surrounding environment. NSD leverages an advanced diffusion model, incorporating two parallel cross-attention mechanisms that separately process text and style information to achieve the dual objectives of semantic control and style consistency. To capture fine-grained style representations, we propose the Progressive Self-style Representational Learning (PSRL) module. This module is predicated on the intuitive premise that different regions within a single image share a consistent style, whereas regions from different images exhibit distinct styles. The PSRL module employs a style contrastive loss that encourages high similarity between representations from the same image while enforcing dissimilarity between those from different images. Furthermore, to address the lack of standardized evaluation protocols for this task, we establish a comprehensive benchmark. This benchmark includes competing algorithms, dedicated style-related metrics, and diverse datasets and settings to facilitate fair comparisons. Extensive experiments conducted on our benchmark demonstrate the effectiveness of the proposed framework. Jianman Lin, Tianshui Chen, Chunmei Qing, Zhijing Yang, Shuangping Huang, Yuheng Ren, Liang Lin 0004 |
IEEE Trans. Image Process. | 4 |
| 2025 | Learning Semantic-aware Representation in Visual-Language Models for Multi-label Recognition with Partial LabelsabstractMulti-label recognition with partial labels (MLR-PL), in which only some labels are known while others are unknown for each image, is a practical task in computer vision, since collecting large-scale and complete multi-label datasets is difficult in real application scenarios. Recently, vision language models (e.g., CLIP) have demonstrated impressive transferability to downstream tasks in data limited or label limited settings. However, current CLIP-based methods suffer from semantic confusion in MLR task due to the lack of fine-grained information in the single global visual and textual representation for all categories. In this work, we address this problem by introducing a semantic decoupling module and a category-specific prompt optimization method in CLIP-based framework. Specifically, the semantic decoupling module following the visual encoder learns category-specific feature maps by utilizing the semantic-guided spatial attention mechanism. Moreover, the category-specific prompt optimization method is introduced to learn text representations aligned with category semantics. Therefore, the prediction of each category is independent, which alleviate the semantic confusion problem. Extensive experiments on Microsoft COCO 2014 and Pascal VOC 2007 datasets demonstrate that the proposed framework significantly outperforms current state-of-art methods with a simpler model structure. Additionally, visual analysis shows that our method effectively separates information from different categories and achieves better performance compared to CLIP-based baseline method. Haoxian Ruan, Zhihua Xu, Zhijing Yang, Yongyi Lu, Jinghui Qin, Tianshui Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | NegVSR: Augmenting Negatives for Generalized Noise Modeling in Real-world Video Super-ResolutionabstractThe capability of video super-resolution (VSR) to synthesize high-resolution (HR) video from ideal datasets has been demonstrated in many works. However, applying the VSR model to real-world video with unknown and complex degradation remains a challenging task. First, existing degradation metrics in most VSR methods are not able to effectively simulate real-world noise and blur. On the contrary, simple combinations of classical degradation are used for real-world noise modeling, which led to the VSR model often being violated by out-of-distribution noise. Second, many SR models focus on noise simulation and transfer. Nevertheless, the sampled noise is monotonous and limited. To address the aforementioned problems, we propose a Negatives augmentation strategy for generalized noise modeling in Video Super-Resolution (NegVSR) task. Specifically, we first propose sequential noise generation toward real-world data to extract practical noise sequences. Then, the degeneration domain is widely expanded by negative augmentation to build up various yet challenging real-world noise sets. We further propose the augmented negative guidance loss to learn robust features among augmented negatives effectively. Extensive experiments on real-world datasets (e.g., VideoLQ and FLIR) show that our method outperforms state-of-the-art methods with clear margins, especially in visual quality. Project page is available at: https://negvsr.github.io/. Yexing Song, Meilin Wang, Zhijing Yang, Xiaoyu Xian, Yukai Shi |
AAAI | 3 |
| 2024 | Learning Adaptive Spatial Coherent Correlations for Speech-Preserving Facial Expression ManipulationabstractSpeech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated with spoken content. Current works depend on inaccessible paired training samples for the person, where two aligned frames exhibit the same speech content yet differ in emotional expression, limiting the SPFEM applications in real-world scenarios. In this work, we discover that speak-ers who convey the same content with different emotions exhibit highly correlated local facial animations, providing valuable supervision for SPFEM. To capitalize on this insight, we propose a novel adaptive spatial coherent correlation learning (ASCCL) algorithm, which models the aforementioned correlation as an explicit metric and integrates the metric to supervise manipulating facial expression and meanwhile better preserving the facial animation of spoken contents. To this end, it first learns a spatial coherent correlation metric, ensuring the visual disparities of adjacent local regions of the image belonging to one emotion are similar to those of the corresponding counterpart of the image belonging to another emotion. Recognizing that visual disparities are not uniform across all regions, we have also crafted a disparity-aware adaptive strategy that prioritizes regions that present greater challenges. During SPFEM model training, we construct the adaptive spatial coherent correlation metric between corresponding local regions of the input and output images as addition loss to supervise the generation process. We conduct extensive experiments on variant datasets, and the results demonstrate the effectiveness of the proposed ASCCL algorithm. Code is publicly available at https://githiub.com/jianmanlincjx/ASCCL Tianshui Chen, Jianman Lin, Zhijing Yang, Chunmei Qing, Liang Lin 0004 |
CVPR | 3 |
| 2024 | Multi-similarity clustering algorithm by ordered pair of normalized real numbersabstractThis paper extends the Fuzzy c-means (FCM) algorithm and proposes the Ordered pair of normalized real numbers clustering (OPNC) algorithm. The OPNC algorithm adopts the paradigm of learning in parallel universes and simultaneously uses multiple similarity measures to convert ordinary data into ordered pairs of normalized real numbers (OPNs). Clustering is performed with OPNs, and OPNs contain different similarity information, so the OPNC algorithm can further improve the clustering performance by combining different similarity measures. Experiments on multiple real datasets and comparisons with other clustering algorithms verified that the OPNC algorithm has excellent performance. Hui Zhang 0055, Zhijing Yang, Chunming Yang, Bo Li 0065 |
IJCNN | 3 |
| 2024 | Self-Supervised Emotion Representation Disentanglement for Speech-Preserving Facial Expression ManipulationabstractSpeech-preserving Facial Expression Manipulation (SPFEM) aims to alter facial emotions in video content while preserving the facial movements associated with speech. Current works often fall short due to the inadequate representation of emotion as well as the absence of time-aligned paired data-two corresponding frames from the same speaker that showcase the same speech content but differ in emotional expression. In this work, we introduce a novel framework, Self-Supervised Emotion Representation Disentanglement (SSERD), to disentangle emotion representation for accurate emotion transfer while implementing a paired data construction module to facilitate automated, photorealistic facial animations. Specifically, We developed a module for learning emotion latent codes using StyleGAN's latent space, employing a cross-attention mechanism to extract and predict emotion editing codes, with contrastive learning to differentiate emotions. To overcome the lack of strictly paired data in the SPFEM task, we exploit pretrained StyleGAN to generate paired data, focusing on expression vectors unrelated to mouth shape. Additionally, we employed a hybrid training strategy using both synthetic paired and real unpaired data to enhance the realism of SPFEM model's generated images. Extensive experiments conducted on benchmark datasets, including MEAD and RAVDESS, have validated the effectiveness of our framework, demonstrating its superior capability in generating photorealistic and expressive facial animations. Zhihua Xu, Tianshui Chen, Zhijing Yang, Chunmei Qing, Yukai Shi, Liang Lin 0004 |
ACM Multimedia | 3 |
| 2024 | Scale Fairness on Spectral ClusteringabstractThe fairness and bias of spectral clustering algorithms have attracted considerable research interest in recent years. Currently fair spectral clustering algorithms are based on the notions of group fairness and individual fairness, which effectively reduce decision bias for similar individuals and sensitive groups. Existing fair spectral clustering algorithms achieve a certain degree of resource redistribution during the clustering process for a particular individual or part of a group, but there is still a situation where the final decision is unfair to the oversized or undersized result clusters. To this end, we present the first principled study of Scale Fairness on Spectral Clustering and propose the SFSC algorithm, which aims to effectively reduce the possibility of the results being oversized or undersized clusters by introducing entropy computation into the spectral clustering process. We measure the scale fairness of clusters by two statistical metrics, and demonstrate on eight classical and real-world datasets that SFSC has better fairness performance compared to spectral clustering while having comparable clustering effect. To the best of our knowledge, this paper is the first study to propose scale fairness for spectral clustering. Zhijing Yang, Hui Zhang 0055, Chunming Yang, Bo Li 0065, Xujian Zhao, Yin Long |
SSDBM | 1 |
| 2024 | Individual Fair Density-Peaks Clustering Based on Local Similar Center Graph and Similar Decision MatrixabstractClustering, as a core technique in data mining, plays a crucial role in uncovering latent patterns in data. Among the many clustering algorithms, Density-Peaks Clustering (DPC) has garnered significant attention due to its ability to efficiently form high-density clusters. Recent research has primarily focused on improving the accuracy and speed of DPC. As concerns about fairness in data science continue to grow, clustering algorithms have gradually started incorporating fairness constraints. Nevertheless, DPC and its variants have remained largely unexplored from the perspective of fairness. Consequently, this paper proposes a novel algorithm, Individual Fair Density-Peaks Clustering (IFDPC), which enhancing individual fairness by Local Similar Center Graph (LSCG), dynamically assigning rest data based on updating Similar Decision Matrix. Experimental results demonstrate that, compared to DPC and its variants, IFDPC not only achieves better fairness but also delivers comparable clustering performance. This work is the first attempt to introduce individual fairness in DPC even in density-based clustering. Code is available on https://github.com/HurryUp1234/IFDPC. Yiding Tang, Zhijing Yang, Yufan Peng, Hui Zhang 0055 |
TrustCom | 2 |
| 2024 | Asymmetric low-rank double-level cooperation for scalable discrete cross-modal hashing
Junpeng Tan, Yinghong Zhou, Zhijing Yang, Feiping Nie 0001, Tianshui Chen |
Expert Syst. Appl. | 4 |
| 2024 | Dual-perspective semantic-aware representation blending for multi-label image recognition with partial labels
Tao Pu 0002, Tianshui Chen, Hefeng Wu, Yukai Shi, Zhijing Yang, Liang Lin 0004 |
Expert Syst. Appl. | 5 |
| 2024 | Heterogeneous Semantic Transfer for Multi-label Recognition with Partial Labels
Tianshui Chen, Tao Pu 0002, Lingbo Liu, Yukai Shi, Zhijing Yang, Liang Lin 0004 |
Int. J. Comput. Vis. | 5 |
| 2024 | Unsupervised multi-perspective fusing semantic alignment for cross-modal hashing retrieval
Yongfeng Chen, Junpeng Tan, Zhijing Yang, Yukai Shi, Jinghui Qin |
Multim. Tools Appl. | 3 |
| 2024 | Discriminative latent semantics-preserving similarity embedding hashing for cross-modal retrieval
Yongfeng Chen, Junpeng Tan, Zhijing Yang, Yongqiang Cheng 0001 |
Neural Comput. Appl. | 3 |
| 2024 | A weighted prior tensor train decomposition method for community detection in multi-layer networks
Zhijing Yang, Tianshui Chen, Jieming Xie, Guang Ma |
Neural Networks | 3 |
| 2024 | Graph Representation and Prototype Learning for webly supervised fine-grained image recognition
Jiantao Lin, Tianshui Chen, Ying-Cong Chen, Zhijing Yang, Yuefang Gao |
Pattern Recognit. Lett. | 4 |
| 2024 | Dynamic Correlation Learning and Regularization for Multi-Label Confidence CalibrationabstractModern visual recognition models often display overconfidence due to their reliance on complex deep neural networks and one-hot target supervision, resulting in unreliable confidence scores that necessitate calibration. While current confidence calibration techniques primarily address single-label scenarios, there is a lack of focus on more practical and generalizable multi-label contexts. This paper introduces the Multi-Label Confidence Calibration (MLCC) task, aiming to provide well-calibrated confidence scores in multi-label scenarios. Unlike single-label images, multi-label images contain multiple objects, leading to semantic confusion and further unreliability in confidence scores. Existing single-label calibration methods, based on label smoothing, fail to account for category correlations, which are crucial for addressing semantic confusion, thereby yielding sub-optimal performance. To overcome these limitations, we propose the Dynamic Correlation Learning and Regularization (DCLR) algorithm, which leverages multi-grained semantic correlations to better model semantic confusion for adaptive regularization. DCLR learns dynamic instance-level and prototype-level similarities specific to each category, using these to measure semantic correlations across different categories. With this understanding, we construct adaptive label vectors that assign higher values to categories with strong correlations, thereby facilitating more effective regularization. We establish an evaluation benchmark, re-implementing several advanced confidence calibration algorithms and applying them to leading multi-label recognition (MLR) models for fair comparison. Through extensive experiments, we demonstrate the superior performance of DCLR over existing methods in providing reliable confidence scores in multi-label scenarios. Tianshui Chen, Weihang Wang 0008, Tao Pu 0002, Jinghui Qin, Zhijing Yang, Jie Liu 0022, Liang Lin 0004 |
IEEE Trans. Image Process. | 5 |
| 2024 | DPHANet: Discriminative Parallel and Hierarchical Attention Network for Natural Language Video LocalizationabstractNatural Language Video Localization (NLVL) has recently attracted much attention because of its practical significance. However, the existing methods still face the following challenges: 1) When the models learn intra-modal semantic association, the temporal causal interaction information and contextual semantic discriminative information are ignored, resulting in the lack of intra-modal semantic context connection; 2) When learning fusion representations, existing cross-modal interaction modules lack hierarchical attention function to extract inter-modal similarity information and intra-modal self-correlation information, resulting in insufficient cross-modal information interaction; and 3) When the loss function is optimized, the existing models ignore the correlation of causal inference between the start and end boundaries, resulting in inaccurate start and end boundary calibrations. To conquer the above challenges, we proposed a novel NLVL model, called Discriminative Parallel and Hierarchical Attention Network (DPHANet). Specifically, we emphasized the importance of temporal causal interaction information and contextual semantic discriminative information and correspondingly proposed a Discriminative Parallel Attention Encoder (DPAE) module to infer and encode the above critical information. Besides, to overcome the shortcomings of the existing cross-modal interaction modules, we designed a Video-Query Hierarchical Attention (VQHA) module, which can perform cross-modal interaction and intra-modal self-correlation modeling in a hierarchical manner. Furthermore, a novel deviation loss function was proposed to capture the correlation of causal inference between the start and end boundaries and force the model to focus on the continuity and temporal causality in the video. Finally, extensive experiments on three benchmark datasets demonstrated the superiority of our proposed DPHANet model, which has achieved about 1.5% and 3.5% average performance improvement and about 2.5% and 7.5% maximum performance improvement on the Charades-STA and TACoS datasets respectively. Junpeng Tan, Zhijing Yang, Yongqiang Cheng 0001, Liang Lin 0004 |
IEEE Trans. Multim. | 3 |
| 2024 | Extensible Max-Min Collaborative Retention for Online Mini-Batch Learning Hash RetrievalabstractAlong with the concern of similarity measures in linear space, supervised online hash methods have been applied to the retrieval task. However, they ignored multi-dimensional space semantic mining and association characteristics will cause quantization errors of information hash code: 1) The similarity relation of discretized data needs to be considered in different spaces; 2) Latent semantic features need to be continuously embedded into hash code learning; 3) The correlation between the structure similarity and discrete hash matrices needs to be continuously optimized. To tackle these challenges, this paper proposes a novel Extensible Max-min Collaborative Retention Online Hash retrieval method based on mini-batch training data (EMCROH). It mainly includes the Max-min Bayesian Similarity Sparse Latent Hash module (MBSSLH), and the Repetition Collaborative Projection Learning module (RCPL). Specifically, MBSSLH is a max-min optimization model. Firstly, to explore the semantic similarity of multi-dimensional space, we propose a novel liner and nonlinear semantic similarity discrimination mechanism based on the log maximum likelihood similarity estimation with Euclidean space and minimize the input batch data features with a common projection matrix. Moreover, to further mine the potential semantic information of the discretization, we also propose a robust sparse discrete latent semantic information extraction submodule based on double latent factors. RCPL can extend the data externally using the repetition collaborative projection matrix with robustness regularization constraint. Finally, a novel max-min embedding iterative step is proposed to solve the batch discrete optimization problem based on Augmented Lagrange Multipliers (ALM) with Alternating Direction Minimization (ADM). Extensive experiments on several well-known large databases demonstrate that EMCROH outperforms the state-of-the-art hash methods. Code and datasets have been publicly available athttps://github.com/Tjeep-Tan/EMCROH. Junpeng Tan, Zhijing Yang, Yongyi Lu, Liang Lin 0004 |
IEEE Trans. Multim. | 3 |
| 2023 | Cost Guarantee for Individual Fairness on Spectral ClusteringabstractThe graph mining algorithm has been widely used in various fields in recent years, among which the spectral clustering algorithm is based on spectral graph theory, which has the ability to cluster on an arbitrarily shaped sample space and converge to the global optimal solution compared with the traditional clustering algorithm. As algorithmic fairness has become a research hot-spot recently, more and more fairness constraints have been introduced into spectral clustering. Most studies focus on group fairness, with only a small number providing individual-level fairness constraints. Existing fair spectral clustering algorithms focus only on whether the clustering results are fair or the decreased rate of fairness loss. To this end, we propose an Individual Fair Spectral Clustering with Cost constraints (IFSCC), which ensures the clustering effect while also improving a certain degree of individual fairness. The experimental results show that IFSCC has the lowest COST value while having a comparable clustering effect compared to other individual fair spectral clustering algorithms. To the best of our knowledge, this paper is the first study to make a trade-off between the clustering effect and fairness performance. Zhijing Yang, Hui Zhang 0055, Chunming Yang, Bo Li 0065 |
ICPADS | 1 |
| 2023 | Reference-free low-light image enhancement by associating hierarchical wavelet representations
Junhong Gong, Lianpei Wu, Zhijing Yang, Yukai Shi, Feiping Nie 0001 |
Expert Syst. Appl. | 4 |
| 2023 | ESA: Excitation-Switchable Attention for convolutional neural networks
Shanshan Zhong, Zhongzhan Huang, Wushao Wen, Zhijing Yang, Jinghui Qin |
Neurocomputing | 4 |
| 2023 | Cross-modal hash retrieval based on semantic multiple similarity learning and interactive projection matrix learning
Junpeng Tan, Zhijing Yang, Jielin Ye, Yongqiang Cheng 0001, Jinghui Qin, Yongfeng Chen |
Inf. Sci. | 2 |
| 2023 | Hypergraph based semi-supervised symmetric nonnegative matrix factorization for image clustering
Jingxing Yin, Zhijing Yang, Badong Chen, Zhiping Lin 0001 |
Pattern Recognit. | 3 |
| 2023 | Adaptive graph regularization method based on least square regression for clusteringabstractThe low-rank representation with adaptive graph regularization (LRAGR) method has been successfully proposed for clustering applications, since it can solve the problems of missing local information and unclear graph structure in data clustering tasks. However, LRAGR utilizes the low-rank representation technique to mine data information, which makes the obtained coefficient matrix too dense and is not conducive to cluster division. Moreover, due to the singular value decomposition, the constraint of the coefficient matrix requires high computational complexity and is difficult to be applied in practical tasks. To address the above issues, in this paper, a new adaptive graph regularization method, called the adaptive graph regularization method based on least squares regression (LSAGR), is proposed for data clustering applications. Specifically, the proposed LSAGR method adopts the Frobenius norm instead of kernel norm to approximate the rank function to satisfy the clustering effect, that is, the coefficients of the cluster-related data are approximately equal. Compared with the traditional LRAGR method, in clustering tasks, the LSAGR method can better reveal the real subspace membership, improve the clustering performance and reduce the computational complexity. Finally, extensive experimental results demonstrate that, compared with several related state-of-the-art methods, the proposed LSAGR method usually has better clustering performance on six real-world image datasets in clustering applications. Jiang-Zhong Cao, Qiaomei Peng, Zhijing Yang |
Signal Process. Image Commun. | 5 |
| 2023 | Multiview Clustering via Hypergraph Induced Semi-Supervised Symmetric Nonnegative Matrix FactorizationabstractNonnegative matrix factorization (NMF) based multiview technique has been commonly used in multiview data clustering tasks. However, previous NMF based multiview clustering approaches fail to take advantage of a small amount of supervisory information to effectively improve the clustering performance, and are easily affected by the additional post-processing method in clustering tasks. To cope with these issues, a novel framework named multiview clustering via hypergraph induced semi-supervised symmetric NMF (MVCHSS) is proposed in this paper for multiview data clustering applications. Specifically, the proposed method has the following features: 1) a new multiview based hypergraph pairwise constraints propagation (MHPCP) algorithm is developed in MVCHSS to construct a set of informative similarity matrices, revealing the high-order relationships effectively and fully utilizing the limited pairwise constraint supervisory information among samples of each view data; 2) the obtained similarity matrices with much supervisory information are not only enforced into the symmetric NMF (SNMF) model, but also incorporated into the graph regularization for each view data; 3) the optimization problem of MVCHSS is formulated for multiview data clustering tasks to acquire a more discriminative clustering indicator matrix (or called consensus assignment matrix) without additional post-processing method. Moreover, the proof of convergence and the computational complexity for MVCHSS are presented. Extensive experiments on five multiview datasets demonstrate that the proposed MVCHSS framework outperforms several state-of-the-art multiview clustering methods. Jingxing Yin, Zhijing Yang, Badong Chen, Zhiping Lin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Real-World Image Super-Resolution by Exclusionary Dual-LearningabstractReal-world image super-resolution is a practical image restoration problem that aims to obtain high-quality images from in-the-wild input, has recently received considerable attention with regard to its tremendous application potentials. Although deep learning-based methods have achieved promising restoration quality on real-world image super-resolution datasets, they ignore the relationship between L1- and perceptual- minimization and roughly adopt auxiliary large-scale datasets for pre-training. In this paper, we discuss the image types within a corrupted image and the property of perceptual- and Euclidean- based evaluation protocols. Then we propose a method, Real-World image Super-Resolution by Exclusionary Dual-Learning (RWSR-EDL) to address the feature diversity in perceptual- and L1- based cooperative learning. Moreover, a noise-guidance data collection strategy is developed to address the training time consumption in multiple datasets optimization. When an auxiliary dataset is incorporated, RWSR-EDL achieves promising results and repulses any training time increment by adopting the noise-guidance data collection strategy. Extensive experiments show that RWSR-EDL achieves competitive performance over state-of-the-art methods on four in-the-wild image super-resolution datasets. Hao Li 0058, Jinghui Qin, Zhijing Yang, Pengxu Wei, Jinshan Pan, Liang Lin 0004, Yukai Shi |
IEEE Trans. Multim. | 3 |
| 2023 | OccluMix: Towards De-Occlusion Virtual Try-on by Semantically-Guided MixupabstractImage Virtual try-on aims at replacing the cloth on a personal image with a garment image (in-shop clothes), which has attracted increasing attention from the multimedia and computer vision communities. Prior methods successfully preserve the character of clothing images, however, occlusion remains a pernicious effect for realistic virtual try-on. In this work, we first present a comprehensive analysis of the occlusions and categorize them into two aspects: i) Inherent-Occlusion: the ghost of the former cloth still exists in the try-on image; ii) Acquired-Occlusion: the target cloth warps to the unreasonable body part. Based on the in-depth analysis, we find that the occlusions can be simulated by a novel semantically-guided mixup module, which can generate semantic-specific occluded images that work together with the try-on images to facilitate training a de-occlusion try-on (DOC-VTON) framework. Specifically, DOC-VTON first conducts a sharpened semantic parsing on the try-on person. Aided by semantics guidance and pose prior, various complexities of texture are selectively blending with human parts in a copy-and-paste manner. Then, the Generative Module (GM) is utilized to take charge of synthesizing the final try-on image and learning to de-occlusion jointly. In comparison to the state-of-the-art methods, DOC-VTON achieves better perceptual quality by reducing occlusion effects. Zhijing Yang, Junyang Chen 0002, Yukai Shi, Hao Li 0058, Tianshui Chen, Liang Lin 0004 |
IEEE Trans. Multim. | 1 |
| 2023 | Progressive Transformer Machine for Natural Character ReenactmentabstractCharacter reenactment aims to control a target person’s full-head movement by a driving monocular sequence that is made up of the driving character video. Current algorithms utilize convolution neural networks in generative adversarial networks, which extract historical and geometric information to iteratively generate video frames. However, convolution neural networks can merely capture local information with limited receptive fields and ignore global dependencies that play a crucial role in face synthesis, leading to generating unnatural video frames. In this work, we design a progressive transformer module that introduces multi-head self-attention with convolution refinement to simultaneously capture global-local dependencies. Specifically, we utilize the non-lapping window-based multi-head self-attention mechanism with hierarchical architecture to obtain the larger receptive fields at low-resolution feature map and thus extract global information. To better model local dependencies, we introduce the convolution operation to further refine the attentional weight in the multi-head self-attention mechanism. Finally, we use several stacked progressive transformer modules with the down-sampling operation to encode information of appearance information of previously generated frames and parameterized 3D face information of the current frame. Similarly, we use several stacked progressive transformer modules with the up-sampling operation to iteratively generate video frames. In this way, it can capture global-local information to facilitate generating video frames that are globally natural while preserving sharp outlines and rich detail information. Extensive experiments on several standard benchmarks suggest that the proposed method outperforms current leading algorithms. Yongzong Xu, Zhijing Yang, Tianshui Chen, Chunmei Qing |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | DFAEN: Double-order knowledge fusion and attentional encoding network for texture recognition
Zhijing Yang, Shujian Lai, Yukai Shi, Yongqiang Cheng 0001, Chunmei Qing |
Expert Syst. Appl. | 1 |
| 2022 | Progressive representation recalibration for lightweight super-resolution
Ruimian Wen, Zhijing Yang, Tianshui Chen, Hao Li 0058 |
Neurocomputing | 2 |
| 2022 | Dual semi-supervised convex nonnegative matrix factorization for data representation
Zhijing Yang, Bingo Wing-Kuen Ling, Badong Chen, Zhiping Lin 0001 |
Inf. Sci. | 2 |
| 2022 | DnSwin: Toward real-world denoising via a continuous Wavelet Sliding Transformer
Hao Li 0058, Zhijing Yang, Ziying Zhao, Junyang Chen 0001, Yukai Shi, Jinshan Pan |
Knowl. Based Syst. | 2 |
| 2022 | Correntropy based semi-supervised concept factorization with adaptive neighbors for clustering
Zhijing Yang, Feiping Nie 0001, Badong Chen, Zhiping Lin 0001 |
Neural Networks | 2 |
| 2022 | A Novel Robust Low-rank Multi-view Diversity Optimization Model with Adaptive-Weighting Based Manifold Learning
Junpeng Tan, Zhijing Yang, Jinchang Ren, Yongqiang Cheng 0001, Bingo Wing-Kuen Ling |
Pattern Recognit. | 2 |
| 2022 | Criteria Comparative Learning for Real-Scene Image Super-ResolutionabstractReal-scene image super-resolution aims to restore real-world low-resolution images into their high-quality versions. A typical RealSR framework usually includes the optimization of multiple criteria which are designed for different image properties, by making the implicit assumption that the ground-truth images can provide a good trade-off between different criteria. However, this assumption could be easily violated in practice due to the inherent contrastive relationship between different image properties. Contrastive learning (CL) provides a promising recipe to relieve this problem by learning discriminative features using the triplet contrastive losses. Though CL has achieved significant success in many computer vision tasks, it is non-trivial to introduce CL to RealSR due to the difficulty in defining valid positive image pairs in this case. Inspired by the observation that the contrastive relationship could also exist between the criteria, in this work, we propose a novel training paradigm for RealSR, named Criteria Comparative Learning (Cria-CL), by developing contrastive losses defined on criteria instead of image patches. In addition, a spatial projector is proposed to obtain a good view for Cria-CL in RealSR. Our experiments demonstrate that compared with the typical weighted regression strategy, our method achieves a significant improvement under similar parameter settings. Yukai Shi, Hao Li 0058, Sen Zhang 0006, Zhijing Yang, Xiao Wang 0014 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Dynamic User Activity and Data Detection for Grant-Free NOMA via Weighted ℓ2, 1 MinimizationabstractGrant-free non-orthogonal multiple access (NOMA) has recently received wide attention for reducing signaling overhead and transmission latency in massive machine-type communications (mMTC). In grant-free NOMA systems, user activity and data (UAD) has to be detected, which is challenging in practice. As an emerging technique, compressive sensing (CS) shows great promise in solving this problem by exploiting the inherent sparsity nature of user activity. This paper proposes to use the weighted$\ell _{2, 1}$minimization (WL21M) to jointly detect UAD in realistic dynamic scenarios. At first, the average recoverability of the WL21M is analyzed. This analysis reveals the fact that the WL21M can improve the detection performance by means of an appropriate weighting and the incorporation of intrinsic temporal correlation. Motivated by the analysis, a collaborative hierarchical match pursuit (C-HiMP) algorithm is proposed for dynamic UAD detection. In the C-HiMP, a sequence of WL21M problems are solved in the subspaces spanned by all of the components in the hierarchical estimated support sets, where the weights are collaboratively updated by the solutions in previous time slots so that an attractive self-correction capacity is obtained. Simulation results demonstrate that the proposed C-HiMP can obtain significant performance improvements, in terms of detection accuracy and computational complexity, compared with several state-of-the-art CS-based detection algorithms. Jun Zhang 0026, Zhijing Yang, Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2021 | SRAGL-AWCL: A two-step multi-view clustering via sparse representation and adaptive weighted cooperative learning
Junpeng Tan, Zhijing Yang, Yongqiang Cheng 0001, Jielin Ye |
Pattern Recognit. | 2 |
| 2021 | Unsupervised Multi-View Clustering by Squeezing Hybrid Knowledge From Cross View and Each ViewabstractMulti-view clustering methods have been a focus in recent years because of their superiority in clustering performance. However, typical traditional multi-view clustering algorithms still have shortcomings in some aspects, such as removal of redundant information, utilization of various views and fusion of multi-view features. In view of these problems, this paper proposes a new multi-view clustering method, low-rank subspace multi-view clustering based on adaptive graph regularization. We construct two new data matrix decomposition models into a unified optimization model. In this framework, we address the significance of the common knowledge shared by the cross view and the unique knowledge of each view by presenting new low-rank and sparse constraints on the sparse subspace matrix. To ensure that we achieve effective sparse representation and clustering performance on the original data matrix, adaptive graph regularization and unsupervised clustering constraints are also incorporated in the proposed model to preserve the internal structural features of the data. Finally, the proposed method is compared with several state-of-the-art algorithms. Experimental results for five widely used multi-view benchmarks show that our proposed algorithm surpasses other state-of-the-art methods by a clear margin. Junpeng Tan, Yukai Shi, Zhijing Yang, Caizhen Wen, Liang Lin 0004 |
IEEE Trans. Multim. | 3 |
| 2020 | MIMN-DPP: Maximum-information and minimum-noise determinantal point processes for unsupervised hyperspectral band selection
Weizhao Chen, Zhijing Yang, Jinchang Ren, Jiang-Zhong Cao, Nian Cai, Huimin Zhao 0001, Peter W. T. Yuen |
Pattern Recognit. | 2 |
| 2020 | DDet: Dual-Path Dynamic Enhancement Network for Real-World Image Super-ResolutionabstractDifferent from traditional image super-resolution task, real image super-resolution(Real-SR) focus on the relationship between real-world high-resolution(HR) and low-resolution(LR) image. Most of the traditional image SR obtains the LR sample by applying a fixed down-sampling operator. Real-SR obtains the LR and HR image pair by incorporating different quality optical sensors. Generally, Real-SR has more challenges as well as broader application scenarios. Previous image SR methods fail to exhibit similar performance on Real-SR as the image data is not aligned inherently. In this article, we propose a Dual-path Dynamic Enhancement Network(DDet) for Real-SR, which addresses the cross-camera image mapping by realizing a dual-way dynamic sub-pixel weighted aggregation and refinement. Unlike conventional methods which stack up massive convolutional blocks for feature representation, we introduce a content-aware framework to study non-inherently aligned image pair in image SR issue. First, we use a content-adaptive component to exhibit the Multi-scale Dynamic Attention(MDA). Second, we incorporate a long-term skip connection with a Coupled Detail Manipulation(CDM) to perform collaborative compensation and manipulation. The above dual-path model is joint into a unified model and works collaboratively. Extensive experiments on the challenging benchmarks demonstrate the superiority of our model. Yukai Shi, Haoyu Zhong, Zhijing Yang, Liang Lin 0004 |
IEEE Signal Process. Lett. | 3 |
| 2020 | Locality Regularized Robust-PCRC: A Novel Simultaneous Feature Extraction and Classification Framework for Hyperspectral ImagesabstractDespite the successful applications of probabilistic collaborative representation classification (PCRC) in pattern classification, it still suffers from two challenges when being applied on hyperspectral images (HSIs) classification: 1) ineffective feature extraction in HSIs under noisy situation; and 2) lack of prior information for HSIs classification. To tackle the first problem existed in PCRC, we impose the sparse representation to PCRC, i.e., to replace the 2-norm with 1-norm for effective feature extraction under noisy condition. In order to utilize the prior information in HSIs, we first introduce the Euclidean distance (ED) between the training samples and the testing samples for the PCRC to improve the performance of PCRC. Then, we bring the coordinate information (CI) of the HSIs into the proposed model, which finally leads to the proposed locality regularized robust PCRC (LRR-PCRC). Experimental results show the proposed LRR-PCRC outperformed PCRC and other state-of-the-art pattern recognition and machine learning algorithms. Zhijing Yang, Faxian Cao, Yongqiang Cheng 0001, Bingo Wing-Kuen Ling, Ruo Hu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Dimensionality reduction based on determinantal point process and singular spectrum analysis for hyperspectral imagesabstractDimensionality reduction is of high importance in hyperspectral data processing, which can effectively reduce the data redundancy and computation time for improved classification accuracy. Band selection and feature extraction methods are two widely used dimensionality reduction techniques. By integrating the advantages of the band selection and feature extraction, the authors propose a new method for reducing the dimension of hyperspectral image data. First, a new and fast band selection algorithm is proposed for hyperspectral images based on an improved determinantal point process (DPP). To reduce the amount of calculation, the dual‐DPP is used for fast sampling representative pixels, followed by k‐nearest neighbour‐based local processing to explore more spatial information. These representative pixel points are used to construct multiple adjacency matrices to describe the correlation between bands based on mutual information. To further improve the classification accuracy, two‐dimensional singular spectrum analysis is used for feature extraction from the selected bands. Experiments show that the proposed method can select a low‐redundancy and representative band subset, where both data dimension and computation time can be reduced. Furthermore, it also shows that the proposed dimensionality reduction algorithm outperforms a number of state‐of‐the‐art methods in terms of classification accuracy. Weizhao Chen, Zhijing Yang, Faxian Cao, Yijun Yan, Meilin Wang, Chunmei Qing, Yongqiang Cheng 0001 |
IET Image Process. | 2 |
| 2019 | Local Block Multilayer Sparse Extreme Learning Machine for Effective Feature Extraction and Classification of Hyperspectral ImagesabstractAlthough extreme learning machines (ELM) have been successfully applied for the classification of hyperspectral images (HSIs), they still suffer from three main drawbacks. These include: 1) ineffective feature extraction (FE) in HSIs due to a single hidden layer neuron network used; 2) ill-posed problems caused by the random input weights and biases; and 3) lack of spatial information for HSIs classification. To tackle the first problem, we construct a multilayer ELM for effective FE from HSIs. The sparse representation is adopted with the multilayer ELM to tackle the ill-posed problem of ELM, which can be solved by the alternative direction method of multipliers. This has resulted in the proposed multilayer sparse ELM (MSELM) model. Considering that the neighboring pixels are more likely from the same class, a local block extension is introduced for MSELM to extract the local spatial information, leading to the local block MSELM (LBMSELM). The loopy belief propagation is also applied to the proposed MSELM and LBMSELM approaches to further utilize the rich spectral and spatial information for improving the classification. Experimental results show that the proposed methods have outperformed the ELM and other state-of-the-art approaches. Faxian Cao, Zhijing Yang, Jinchang Ren, Weizhao Chen, Guojun Han, Yuzhen Shen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Optimal design of orders of DFrFTs for sparse representationsabstractThis study proposes an optimal design of the orders of the discrete fractional Fourier transforms (DFrFTs) and construct an overcomplete transform using the DFrFTs with these orders for performing the sparse representations. The design problem is formulated as an optimisation problem with an ‐norm non‐convex objective function. To avoid all the orders of the DFrFTs to be the same, the exclusive OR of two constraints are imposed. The constrained optimisation problem is further reformulated to an optimal frequency sampling problem. A method based on solving the roots of a set of harmonic functions is employed for finding the optimal sampling frequencies. As the designed overcomplete transform can exploit the physical meanings of the signals in terms of representing the signals as the sums of the components in the time–frequency plane, the designed overcomplete transform can be applied to many applications. Xiao-Zhi Zhang, Bingo Wing-Kuen Ling, Ran Tao 0003, Zhijing Yang, Wai Lok Woo, Saeid Sanei, Kok Lay Teo |
IET Signal Process. | 4 |
| 2018 | Camera identification based on very low bit rate videos with overall noise pattern having time varying statistics
Nili Tian, Bingo Wing-Kuen Ling, Chunmei Qing, Zhijing Yang |
Multim. Tools Appl. | 4 |
| 2018 | Joint bilateral filtering and spectral similarity-based sparse representation: A generic framework for effective feature extraction and data classification in hyperspectral imaging
Zhijing Yang, Jinchang Ren, Peter W. T. Yuen, Huimin Zhao 0001, Genyun Sun, Stephen Marshall, Jón Atli Benediktsson |
Pattern Recognit. | 2 |
| 2018 | Sparse Representation-Based Augmented Multinomial Logistic Extreme Learning Machine With Weighted Composite Features for Spectral-Spatial Classification of Hyperspectral ImagesabstractAlthough extreme learning machine (ELM) has successfully been applied to a number of pattern recognition problems, only with the original ELM it can hardly yield high accuracy for the classification of hyperspectral images (HSIs) due to two main drawbacks. The first is due to the randomly generated initial weights and bias, which cannot guarantee optimal output of ELM. The second is the lack of spatial information in the classifier as the conventional ELM only utilizes spectral information for classification of HSI. To tackle these two problems, a new framework for ELM-based spectral-spatial classification of HSI is proposed, where probabilistic modeling with sparse representation and weighted composite features (WCFs) is employed to derive the optimized output weights and extract spatial features. First, ELM is represented as a concave logarithmic-likelihood function under statistical modeling using the maximum a posteriori estimator. Second, sparse representation is applied to the Laplacian prior to efficiently determine a logarithmic posterior with a unique maximum in order to solve the ill-posed problem of ELM. The variable splitting and the augmented Lagrangian are subsequently used to further reduce the computation complexity of the proposed algorithm. Third, the spatial information is extracted using the WCFs to construct the spectral-spatial classification framework. In addition, the lower bound of the proposed method is derived by a rigorous mathematical proof. Experimental results on three publicly available HSI data sets demonstrate that the proposed methodology outperforms ELM and also a number of state-of-the-art approaches. Faxian Cao, Zhijing Yang, Jinchang Ren, Bingo Wing-Kuen Ling, Huimin Zhao 0001, Meijun Sun, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Blind inpainting using the fully convolutional neural network
Nian Cai, Zhenghang Su, Zhineng Lin, Han Wang 0017, Zhijing Yang, Bingo Wing-Kuen Ling |
Vis. Comput. | 5 |
| 2016 | Stabilization of single bit high order interpolative sigma delta modulators for analog-to-digital conversion in wireless mobile handset based electromyogram acquisition systemabstractWireless mobile handsets are widely used for the electromyogram acquisitions because of their portability property. However, as the electromyograms are with low amplitudes and they are very sensitive to the noises, a good analog-to-digital conversion system plays a very important role for the processing of the electromyograms. Among all the analog-to-digital converters, the sigma delta modulator is the most common analog-to-digital convertor employed in the mobile handsets. This is because the oversampling mechanism for bandlimited signals can be implemented using the existing hardware. However, the sigma delta modulator may suffer from the stability issue. This paper considers the stabilization of a single bit high order interpolative sigma delta modulator via flipping some values of the quantizer output. The values of the quantizer output are determined in such that certain frequency contents of the input signal are cancelled by those of the quantizer output. A stability condition for the proposed control strategy is derived. Since only some frequency detectors and a simple inverter are required for the implementation of the proposed control strategy, the implementation cost is low. Yu-Fan Zeng, Bingo Wing-Kuen Ling, Yuping Gui, Zhijing Yang |
INDIN | 4 |
| 2016 | The graphics of the solutions by learning a RBM: Discussion on a special caseabstractThe study of Restricted Boltzmann Machine(RBM) attracts considerable attentions in recent years. RBM training algorithm is an unsupervised learning method with many applications, moreover, it is the basic module in deep learning. Maximizing the log-likelihood by gradient ascent method, RBM training algorithm can approximate the probability distribution underlying the observing data. For a simple RBM system, we give the closed-form representation for the solutions of the training algorithm, which forms a manifold. Based on the result, it is understood that the solution of the maximization of the log-likelihood function calculated by the gradient ascent method give only one point on the manifold. To illustrate the phenomenon more clearly, a family of new parameters are introduced to express the solution manifold. Qian Zhang 0044, Zhijing Yang, Chao Huang 0004, Zhihua Yang, Lihua Yang 0001 |
INDIN | 2 |
| 2016 | Instantaneous magnitudes and instantaneous frequencies of signals with their positivity constraints via non-smooth non-convex functional constrained optimisationabstractThis study proposes an iterative method to approximate an N‐dimensional optimisation problem with a weighted Lp norm and L2 norm objective function by a sequence of N independent one‐dimensional optimisation problems. This iterative method is inspired by the existing weighted L1 norm and L2 norm separable surrogate functional (SSF) iterative shrinkage algorithm. However, as these independent one‐dimensional optimisation problems consist of weighted Lp norm and L2 norm objective functions, these optimisation problems are non‐convex and they may have more than one locally optimal solutions. In general, it is very difficult to find their globally optimal solutions. To address this difficulty, this study proposes to partition the feasible set of each approximated problem into various regions such that the sign of the convexity of the objective function in each region remains unchanged. In this case, there is no more than one stationary point in each region. By finding the stationary point in each region, the globally optimal solution of each approximated optimisation problem can be found. Besides, this study also shows that the sequence of the globally optimal solutions of the approximated problems converge to the globally optimal solution of the original optimisation problem. Zhijing Yang, Wei-Chao Kuang, Bingo Wing-Kuen Ling |
IET Signal Process. | 1 |
| 2016 | Novel segmented stacked autoencoder for effective dimensionality reduction and feature extraction in hyperspectral imaging
Jaime Zabalza, Jinchang Ren, Jiangbin Zheng 0001, Huimin Zhao 0001, Chunmei Qing, Zhijing Yang, Peijun Du, Stephen Marshall |
Neurocomputing | 6 |
| 2015 | Detecting Parkinson's diseases via the characteristics of the intrinsic mode functions of filtered electromyogramsabstractThis paper proposes a novel method for detecting the Parkinson's diseases via applying the empirical mode decomposition to filtered electromyograms. First, the electromyograms are processed by different linear phase finite impulse response bandpass filters with different pairs of cutoff frequencies. Second, each filtered electromyogram is decomposed into several intrinsic mode functions. Third, both the entropies and the total numbers of the extrema of the intrinsic mode functions of each filtered electromyogram are computed and they are used as the features for detecting the Parkinson's diseases. Computer numerical simulation results show that the features are linearly separable. Hence, a simple perceptron can be employed for the detection of the Parkinson's diseases. Finally, the algorithm is implemented via a mobile application. Compared to conventional empirical mode decomposition approaches in which a predefined number of features is employed for detecting the Parkinson's diseases, our proposed method allows to use a flexible number of features for detecting the Parkinson's diseases. This is because the total number of filters to be employed is very flexible. As a result, our proposed method is more flexible than the existing methods. Yizhong Dai, Wei-Chao Kuang, Bingo Wing-Kuen Ling, Zhijing Yang, Kim Fung Tsang, Hao Ran Chi, Chung Kit Wu, Henry S. H. Chung, Gerhard P. Hancke 0001 |
INDIN | 4 |
| 2015 | Content removal via both thresholding averaging and two dimensional discrete fractional Fourier transformabstractThis paper proposes a novel content removal technique for enhancing the camera identification performance. Here, very low bit rate videos with the overall noise patterns having time varying statistics are considered. First, different two dimensional discrete fractional Fourier transforms with different rotational angles are applied to the overall noise pattern of each frame of each video. Second, the modulus of each element of each transformed matrix is normalized to one if the rotational angles of the transforms are not equal to the integer multiples of π. Third, the corresponding two dimensional inverse discrete fractional Fourier transform is applied to each normalized matrix and the corresponding real part is taken out for the further processing. Fourth, the absolute values of the elements in each normalized real valued matrix are bounded by a certain threshold value. Finally, the processed matrices are averaged over all the rotational angles and all the frames of the videos corresponding to the same camera. Extensive computer numerical simulation results on the correlation performances are presented. It is found that the proposed method outperforms the existing method for a wide range of rotational angles. Bingo Wing-Kuen Ling, Zhijing Yang, Nian Cai |
PCS | 3 |
| 2015 | Approximate affine linear relationship between L1 norm objective functional values and L2 norm constraint boundsabstractFor an optimisation problem with an L 1 norm objective function subject to an L 2 norm inequality constraint, this study shows that there is an approximately linear relationship between the L 1 norm objective functional values and the L 2 norm specifications. This relationship is verified through the use of random and real world industrial data. The obtained results can be employed for (i) estimating the L 1 norm objective functional value without solving the optimisation problem numerically; (ii) providing an insight for defining the L 2 norm specification in which a simple method is proposed in this study; and (iii) testing whether the obtained solutions are the globally optimal solutions or not. These advantages are demonstrated via the use of random data. Zhijing Yang, Bingo Wing-Kuen Ling, Chris Bingham |
IET Signal Process. | 1 |
| 2013 | Extracting underlying trend and predicting power usage via joint SSA and sparse binary programmingabstractThis paper proposes a novel methodology for extracting the underlying trend and predicting the power usage through a joint singular spectrum analysis (SSA) and sparse binary programming approach. The underlying trend is approximated by the sum of a part of SSA components, in which the total number of the SSA components in the sum is minimized subject to a specification on the maximum absolute difference between the original signal and the approximated underlying trend. As the selection of the SSA components is binary, this selection problem is to minimize the L0norm of the selection vector subject to the L∞norm constraint on the difference between the original signal and the approximated underlying trend as well as the binary valued constraint on the elements of the selection vector. This problem is actually a sparse binary programming problem. To solve this problem, first the corresponding continuous valued sparse optimization problem is solved. That is, to solve the same problem without the consideration of the binary valued constraint. This problem can be approximated by a linear programming problem when the isometry condition is satisfied, and the solution of the linear programming problem can be obtained via existing simplex methods or interior point methods. By applying the binary quantization to the obtained solution of the linear programming problem, the approximated solution of the original sparse binary programming problem is obtained. Unlike previously reported techniques that require a pre-cursor model or parameter specifications, the proposed method is completely adaptive. Experiment results show that our proposed method is very effective and efficient for extracting the underlying trend and predicting the power usage. Zhijing Yang, Bingo Wing-Kuen Ling, Chris Bingham |
ISCAS | 1 |
| 2012 | An Evolutionary Based Clustering Algorithm Applied to Dada Compression for Industrial Systems
Jun Chen 0009, Mahdi Mahfouf, Chris Bingham, Yu Zhang 0001, Zhijing Yang, Michael Gallimore |
IDA | 5 |
| 2012 | Unit Operational Pattern Analysis and Forecasting Using EMD and SSA for Industrial Systems
Zhijing Yang, Chris Bingham, Bingo Wing-Kuen Ling, Yu Zhang 0001, Michael Gallimore, Jill Stewart |
IDA | 1 |
| 2010 | Normalized Co-Occurrence Mutual Information for Facial Pose Detection Inside VideosabstractHuman faces captured inside videos are often presented with variable poses, making it difficult to recognize and thus pose detection becomes crucial for such face recognition under non-controlled environment. While existing mutual in formation (MI) primarily considers the relationship between corresponding individual pixels, we propose a normalized co occurrence mutual information in this letter to capture the information embedded not only in corresponding pixel values but also in their geographical locations. In comparison with the existing Mis, the proposed presents an essential advantage that both marginal entropy and joint entropy can be optimally exploited in measuring the similarity between two given images. When developed into a facial pose detection algorithm inside video sequences, we show, through extensive experiments, that such design is capable of achieving the best performances among all the representative existing techniques compared. Chunmei Qing, Jianmin Jiang, Zhijing Yang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |