Zhihui Li 0001

dblp:95/5287-1 · DBLP profile ↗
← Back
58ranked-venue papers
12as first author
29since 2021 · last 2026
0000-0001-9642-8009ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 11 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
abstract
Text-guided image inpainting aims to inpaint masked image regions based on a textual prompt while preserving the background. Although diffusion-based methods have become dominant, their property of modeling the entire image in latent space makes it challenging for the results to align well with prompt details and maintain a consistent background. To address these issues, we explore Mask AutoRegressive (MAR) models for this task. MAR naturally supports image inpainting by generating latent tokens corresponding to mask regions, enabling better local controllability without altering the background. However, directly applying MAR to this task makes the inpainting content either ignore the prompts or be disharmonious with the background context. Through analysis of the attention maps from the inpainting images, we identify the impact of background tokens on text tokens during the MAR generation, and leverage this to designToken Painter, a training-free text-guided image inpainting method based on MAR. Our approach introduces two key components: (1) Dual-Stream Encoder Information Fusion (DEIF), which fuses the semantic and context information from text and background in frequency domain to produce novel guidance tokens, allowing MAR to generate text-faithful inpainting content while keeping harmonious with background context. (2) Adaptive Decoder Attention Score Enhancing (ADAE), which adaptively enhances attention scores on guidance tokens and inpainting tokens to further enhance the alignment of prompt details and the content visual quality. Extensive experiments demonstrate that our training-free method outperforms prior state-of-the-art methods across almost all metrics.
Longtao Jiang, Jie Huang 0017, Mingfei Han 0002, Yongqiang Yu, Feng Zhao 0004, Xiaojun Chang, Zhihui Li 0001
AAAI8
2026 Entropy-Guided Condensing for Vision Transformer
Sihao Lin, Pumeng Lyu, Dongrui Liu, Zhihui Li 0001, Wenguan Wang, Xiaojun Chang, Yuhui Zheng
Int. J. Comput. Vis.4
2026 Advancing In-Context Learning for Efficient and Stable Medical Report Generation
abstract
Vision-language models (VLMs) have shown strong generalization across multimodal tasks, but adapting them to medical report generation (MRG) often demands extensive paired image-text data that are limited due to data privacy and annotation cost. In-context learning (ICL) offers a promising training-free alternative, yet standard ICL approaches rely on long demonstration prompts that are computationally inefficient and often yield inconsistent or clinically inaccurate descriptions. To address these challenges, we propose Principal In-Context Vectors (PCVs), a compact latent-guidance framework that distills multimodal demonstrations into stable semantic representations. By extracting hidden states from auto-regressive VLMs and applying principal component analysis (PCA), we identify robust semantic directions that remain stable under input perturbations. These PCVs are then injected into new queries to steer generation toward accurate and clinically meaningful outputs without any model tuning. Extensive experiments on four MRG benchmark datasets show that our approach can enhance both zero-shot and fully supervised generation quality across diverse settings, including cross-center, cross-disease, and longitudinal scenarios. This work provides a lightweight and scalable approach to adapt pre-trained VLMs for practical clinical deployment.
Mingjie Li 0006, Zeyi Shi, Mingfei Han 0002, Lina Yao 0001, Zhihui Li 0001, Xiaojun Chang, Kilian M. Pohl, Md Tauhidul Islam, Lei Xing 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Distilling Object Detectors via Monte Carlo Dropout
abstract
Knowledge distillation (KD) has become a fundamental technique for model compression in object detection tasks. The data noise and training randomness may cause the knowledge of the teacher model to be unreliable, referred to as knowledge uncertainty. Existing methods neglect this uncertainty, potentially hindering the student's capacity to capture and understand latent "dark knowledge". In this work, we introduce a novel strategy that explicitly incorporates knowledge uncertainty, named Uncertainty-Driven Knowledge Extraction and Transfer (UET). Given the unknown, high-dimensional nature of the knowledge distribution, we employ Monte Carlo dropout to effectively estimate the teacher's uncertainty. Leveraging information theory, we combine uncertainty with deterministic knowledge, enabling the student to benefit from both precision and diversity. UET is a plug-and-play method that integrates seamlessly with existing distillation techniques. We validate our approach through comprehensive experiments across various distillation strategies, detectors, and backbones. Specifically, UET achieves state-of-the-art results, with a ResNet50-based GFL detector obtaining 44.1% mAP on the COCO dataset-surpassing baseline performance by 3.9%.
Junfei Yi, Hui Zhang 0023, Jianxu Mao, Tengfei Liu 0005, Mingjie Li 0006, Sihao Lin, Hanyu Gu, Zhihui Li 0001, Xiaojun Chang, Yaonan Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2026 SELongVLM: Empowering Long Video Language Models With Self-Corrective Clip Selection
abstract
Recent advances in multimodal large language models (MLLMs) have enabled impressive progress in visual-language reasoning, yet long-video understanding remains a formidable challenge due to the need for coherent reasoning over ultra-long spatiotemporal dependencies. Existing methods struggle with the vast candidate space for relevant information in long videos, often failing to distinguish meaningful events from redundant content. We identify two critical and previously under-explored issues: absolute redundancy, where static visual content inflates token counts without adding narrative value, and relative redundancy, where task-irrelevant segments introduce noise that impairs reasoning. Compounding these issues is the weak spatiotemporal modeling in current MLLMs, which limits their ability to capture complex event dynamics. To address these multifaceted challenges, we introduce SELongVLM, a dynamically lenient-to-stringent selection long video language model. SELongVLM integrates two coordinated branches: a Residual Token Pruner (RTP) that removes repetitive background tokens via inter-frame residual modeling thus mitigating absolute redundancy while preserving motion cues, and a Semantic-aware Self-Correction Selector (SCSelector) that progressively refines query-relevant clip selection without frame-level annotations to reduce relative redundancy, guided by a stringent-to-lenient self-correcting mechanism during optimization. To ensure causal continuity and bolster spatiotemporal reasoning across disjoint clips, the framework further incorporates an action-aware operation for intra-clip dynamics and a temporal memory for cross-clip context, enabling robust spatiotemporal inference on long videos. Extensive experiments across eight benchmarks demonstrate that SELongVLM markedly outperforms existing models on both general and specialized long-video tasks. Specifically, it achieves 65.5% on VideoMME and 69.8% on MLVU for general benchmarks, and delivers strong performance on four specialized benchmarks - for example, 39.2% on TOMATO for fine-grained temporal reasoning and 69.2% on EventBench for event-level understanding.
Kecheng Zhang, Zongxin Yang, Mingfei Han 0002, Yunzhi Zhuge, Haihong Hao, Zhihui Li 0001, Xiaojun Chang
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation
abstract
Mainstream research in audio-visual learning has focused on designing task-specific expert models, primarily implemented through sophisticated multimodal fusion approaches. Recently, a few efforts have aimed to develop more task-independent or universal audiovisual embedding networks, encoding advanced representations for use in various audiovisual downstream tasks. This is typically achieved by fine-tuning large pretrained transformers, such as Swin-V2-L and HTS-AT, in a parameter-efficient manner through techniques such as tuning only a few adapter layers inserted into the pretrained transformer backbone. Although these methods are parameter-efficient, they suffer from significant training memory consumption due to gradient backpropagation through the deep transformer backbones, which limits accessibility for researchers with constrained computational resources. In this paper, we present Meta-Token Learning (Mettle), a simple and memory-efficient method for adapting large-scale pretrained transformer models to downstream audio-visual tasks. Instead of sequentially modifying the output feature distribution of the transformer backbone, Mettle utilizes a lightweight Layer-Centric Distillation (LCD) module to distill in parallel the intact audio or visual features embedded by each transformer layer into compact meta-tokens. This distillation process considers both pretrained knowledge preservation and task-specific adaptation. The obtained meta-tokens can be directly applied to classification tasks, such as audio-visual event localization and audio-visual video parsing. To further support fine-grained segmentation tasks, such as audio-visual segmentation, we introduce a Meta-Token Injection (MTI) module, which utilizes the audio and visual meta-tokens distilled from the top transformer layer to guide feature adaptation in earlier layers. Extensive experiments on multiple audiovisual benchmarks demonstrate that our method significantly reduces memory usage and training time while maintaining parameter efficiency and competitive accuracy.
Jinxing Zhou, Zhihui Li 0001, Yongqiang Yu, Yanghao Zhou, Ruohao Guo, Guangyao Li 0001, Yuxin Mao, Mingfei Han 0002, Xiaojun Chang, Meng Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Top-Push Polynomial Ranking Embedded Dictionary Learning for Enhanced Re-Id
Ying Chen 0041, De Cheng, Zhihui Li 0001, Andy Song
ICAART (3)3
2025 Variational methods with application to medical image segmentation: A survey
Xiu Shu, Zhihui Li 0001, Xiaojun Chang, Di Yuan 0002
Neurocomputing2
2025 An active learning model based on image similarity for skin lesion segmentation
Xiu Shu, Zhihui Li 0001, Chunwei Tian, Xiaojun Chang, Di Yuan 0002
Neurocomputing2
2025 TeST: Temporal-spatial separated transformer for temporal action localization
Herun Wan, Minnan Luo, Zhihui Li 0001
Neurocomputing3
2025 Dual Prototypes-Based Personalized Federated Adversarial Cross-Modal Hashing
abstract
With the rapid advances in wireless communication and IoT platforms, it is increasingly difficult to analyze relevant multi-modal data distributed across geographically diverse and heterogeneous platforms. One promising approach is to rely on federated learning to build compact cross-modal hash codes. However, existing federated learning methods easily exhibit degenerative performance in the global model due to the distributed data being derived from diverse domains. In addition, directly forcing each client to adopt the same global parameters as local parameters, without effective local training, significantly reduces the performance of each client. To overcome these challenges, we propose a novel federated adversarial cross-modal hashing, called Dual Prototypes-based personalized Federated Adversarial (DP-FeAd), which provides iterated training of shared dual prototypes. Specifically, aiming to expand local hashing models beyond their knowledge realms, DP-FeAd enables participating clients to engage in cooperative learning through two constructions: cluster prototypes and unbiased prototypes, instead of the traditional global prototypes, ensuring both generalization and stability. Specifically, the cluster prototypes are derived from local class-level prototypes and adversarially trained with local approximate hash codes to align their distributions. The unbiased prototypes are averaged from cluster prototypes and integrated into the training of local hashing models to maintain consistency across different local class-level prototypes further. The experiments conducted on two benchmark datasets demonstrate that our proposed method significantly enhances the performance of deep cross-modal hashing models in both IID (Independent and Identically Distributed) and non-IID scenarios.
Lingchen Gu, Xiaojuan Shen, Jiande Sun 0001, Jing Li 0046, Zhihui Li 0001, Sen-Ching S. Cheung, Wenbo Wan
IEEE Trans. Circuits Syst. Video Technol.6
2025 Fine-Grained Feature and Template Reconstruction for TIR Object Tracking
abstract
Thermal infrared (TIR) object tracking is a significant subject within the field of computer vision. Currently, TIR object tracking faces challenges such as insufficient representation of object texture information and underutilization of temporal information, which severely affects the tracking accuracy of TIR tracking methods. To address these issues, we propose a TIR object tracking method (called: FFTR) based on fine-grained feature and template reconstruction. Specifically, aiming at the fine-grained information of the TIR object, we employ a frequency channel attention mechanism that transforms TIR images into the frequency domain using discrete cosine transform features. By capturing the fine-grained feature of TIR images from the frequency domain, we enhance the model’s ability to comprehend these images. To better leverage temporal information, we utilize a template region reconstruction method. This method reconstructs the template from the previous frame based on the search area of the current frame, which is then incorporated into the attention computation for the subsequent frame, thereby improving the tracking capability of TIR objects. Extensive quantitative and qualitative experiments show that our method achieves competitive tracking performance on the TIR benchmarks.
Donghai Liao, Xiu Shu, Zhihui Li 0001, Qiao Liu 0001, Di Yuan 0002, Xiaojun Chang, Zhenyu He 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Masked Distillation Advances Self-Supervised Transformer Architecture Search
abstract
Transformer architecture search (TAS) has achieved remarkable progress in automating the neural architecture design process of vision transformers. Recent TAS advancements have discovered outstanding transformer architectures while saving tremendous labor from human experts. However, it is still cumbersome to deploy these methods in real-world applications due to the expensive costs of data labeling under the supervised learning paradigm. To this end, this paper proposes a masked image modelling (MIM) based self-supervised neural architecture search method specifically designed for vision transformers, termed as MaskTAS, which completely avoids the expensive costs of data labeling inherited from supervised learning. Based on the one-shot NAS framework, MaskTAS requires to train various weight-sharing subnets, which can easily diverged without strong supervision in MIM-based self-supervised learning. For this issue, we design the search space of MaskTAS as a siamesed teacher-student architecture to distill knowledge from pre-trained networks, allowing for efficient training of the transformer supernet. To achieve self-supervised transformer architecture search, we further design a novel unsupervised evaluation metric for the evolutionary search algorithm, where each candidate of the student branch is rated by measuring its consistency with the larger teacher network. Extensive experiments demonstrate that the searched architectures can achieve state-of-the-art accuracy on CIFAR-10, CIFAR-100, and ImageNet datasets even without using manual labels. Moreover, the proposed MaskTAS can generalize well to various data domains and tasks by searching specialized transformer architectures in self-supervised manner.
Caixia Yan, Xiaojun Chang, Zhihui Li 0001, Lina Yao 0001, Minnan Luo
ICLR3
2023 HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object Segmentation
abstract
Referring Video Object Segmentation (RVOS) is to segment the object instance from a given video, according to the textual description of this object. However, in the open world, the object descriptions are often diversified in contents and flexible in lengths. This leads to the key difficulty in RVOS, i.e., various descriptions of different objects are corresponding to different temporal scales in the video, which is ignored by most existing approaches with single stride of frame sampling. To tackle this problem, we propose a concise Hybrid Temporal-scale Multimodal Learning (HTML) framework, which can effectively align lingual and visual features to discover core object semantics in the video, by learning multimodal interaction hierarchically from different temporal scales. More specifically, we introduce a novel inter-scale multimodal perception module, where the language queries dynamically interact with visual features across temporal scales. It can effectively reduce complex object confusion by passing video context among different scales. Finally, we conduct extensive experiments on the widely used benchmarks, including Ref-Youtube-VOS, Ref-DAVIS17, A2D-Sentences and JHMDB-Sentences, where our HTML achieves state-of-the-art performance on all these datasets.
Mingfei Han 0002, Yali Wang 0001, Zhihui Li 0001, Lina Yao 0001, Xiaojun Chang, Yu Qiao 0001
ICCV3
2023 Unsupervised Visible-Infrared Person ReID by Collaborative Learning with Neighbor-Guided Label Refinement
abstract
Unsupervised learning visible-infrared person re-identification (USL-VI-ReID) aims at learning modality-invariant features from unlabeled cross-modality dataset, which is crucial for practical applications in video surveillance systems. The key to essentially address the USL-VI-ReID task is to solve the cross-modality data association problem for further heterogeneous joint learning. To address this issue, we propose a Dual Optimal Transport Label Assignment (DOTLA) framework to simultaneously assign the generated labels from one modality to its counterpart modality. The proposed DOTLA mechanism formulates a mutual reinforcement and efficient solution to cross-modality data association, which could effectively reduce the side-effects of some insufficient and noisy label associations. Besides, we further propose a cross-modality neighbor consistency guided label refinement and regularization module, to eliminate the negative effects brought by the inaccurate supervised signals, under the assumption that the prediction or label distribution of each example should be similar to its nearest neighbors'. Extensive experimental results on the public SYSU-MM01 and RegDB datasets demonstrate the effectiveness of the proposed method, surpassing existing state-of-the-art approach by a large margin of 7.76% mAP on average, which even surpasses some supervised VI-ReID methods.
De Cheng, Xiaojian Huang, Nannan Wang 0001, Zhihui Li 0001, Xinbo Gao 0001
ACM Multimedia5
2023 MARLlib: A Scalable and Efficient Multi-agent Reinforcement Learning Library
abstract
A significant challenge facing researchers in the area of multi-agent reinforcement learning (MARL) pertains to the identification of a library that can offer fast and compatible development for multi-agent tasks and algorithm combinations, while obviating the need to consider compatibility issues. In this paper, we present MARLlib, a library designed to address the aforementioned challenge by leveraging three key mechanisms: 1) a standardized multi-agent environment wrapper, 2) an agent-level algorithm implementation, and 3) a flexible policy mapping strategy. By utilizing these mechanisms, MARLlib can effectively disentangle the intertwined nature of the multi-agent task and the learning process of the algorithm, with the ability to automatically alter the training strategy based on the current task's attributes. The MARLlib library's source code is publicly accessible on GitHub: https://github.com/Replicable-MARL/MARLlib.
Siyi Hu 0001, Yifan Zhong, Minquan Gao, Weixun Wang, Hao Dong 0003, Xiaodan Liang, Zhihui Li 0001, Xiaojun Chang, Yaodong Yang 0001
J. Mach. Learn. Res.7
2023 Adaptive Manifold Graph representation for Two-Dimensional Discriminant Projection
Jinlong Qu, Xiaowei Zhao 0002, Xiaojun Chang, Zhihui Li 0001, Xuanhong Wang
Knowl. Based Syst.5
2023 A Comprehensive Survey of Scene Graphs: Generation and Application
abstract
Scene graph is a structured representation of a scene that can clearly express the objects, attributes, and relationships between objects in the scene. As computer vision technology continues to develop, people are no longer satisfied with simply detecting and recognizing objects in images; instead, people look forward to a higher level of understanding and reasoning about visual scenes. For example, given an image, we want to not only detect and recognize objects in the image, but also understand the relationship between objects (visual relationship detection), and generate a text description (image captioning) based on the image content. Alternatively, we might want the machine to tell us what the little girl in the image is doing (Visual Question Answering (VQA)), or even remove the dog from the image and find similar images (image editing and retrieval), etc. These tasks require a higher level of understanding and reasoning for image vision tasks. The scene graph is just such a powerful tool for scene understanding. Therefore, scene graphs have attracted the attention of a large number of researchers, and related research is often cross-modal, complex, and rapidly developing. However, no relatively systematic survey of scene graphs exists at present. To this end, this survey conducts a comprehensive investigation of the current scene graph research. More specifically, we first summarize the general definition of the scene graph, then conducte a comprehensive and systematic discussion on the generation method of the scene graph (SGG) and the SGG with the aid of prior knowledge. We then investigate the main applications of scene graphs and summarize the most commonly used datasets. Finally, we provide some insights into the future development of scene graphs.
Xiaojun Chang, Pengzhen Ren, Pengfei Xu 0003, Zhihui Li 0001, Xiaojiang Chen, Alex Hauptmann 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Vision Transformers
abstract
Dynamic networks have shown their promising capability in reducing theoretical computation complexity by adapting their architectures to the input during inference. However, their practical runtime usually lags behind the theoretical acceleration due to inefficient sparsity. In this paper, we explore a hardware-efficient dynamic inference regime, named dynamic weight slicing, that can generalized well on multiple dimensions in both CNNs and transformers (e.g. kernel size, embedding dimension, number of heads, etc.). Instead of adaptively selecting important weight elements in a sparse way, we pre-define dense weight slices with different importance level by nested residual learning. During inference, weights are progressively sliced beginning with the most important elements to less important ones to achieve different model capacity for inputs with diverse difficulty levels. Based on this conception, we present DS-CNN++ and DS-ViT++, by carefully designing the double headed dynamic gate and the overall network architecture. We further propose dynamic idle slicing to address the drastic reduction of embedding dimension in DS-ViT++. To ensure sub-network generality and routing fairness, we propose a disentangled two-stage optimization scheme. In Stage I, in-place bootstrapping (IB) and multi-view consistency (MvCo) are proposed to stablize and improve the training of DS-CNN++ and DS-ViT++ supernet, respectively. In Stage II, sandwich gate sparsification (SGS) is proposed to assist the gate training. Extensive experiments on 4 datasets and 3 different network architectures demonstrate our methods consistently outperform the state-of-the-art static and dynamic model compression methods by a large margin (up to 6.6%). Typically, we achieves 2-4× computation reduction and up to 61.5% real-world acceleration on MobileNet, ResNet-50 and Vision Transformer, with minimal accuracy drops on ImageNet. Code release: https://github.com/changlin31/DS-Net.
Guangrun Wang, Bing Wang 0013, Xiaodan Liang, Zhihui Li 0001, Xiaojun Chang
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 When Object Detection Meets Knowledge Distillation: A Survey
abstract
Object detection (OD) is a crucial computer vision task that has seen the development of many algorithms and models over the years. While the performance of current OD models has improved, they have also become more complex, making them impractical for industry applications due to their large parameter size. To tackle this problem, knowledge distillation (KD) technology was proposed in 2015 for image classification and subsequently extended to other visual tasks due to its ability to transfer knowledge learned by complex teacher models to lightweight student models. This paper presents a comprehensive survey of KD-based OD models developed in recent years, with the aim of providing researchers with an overview of recent progress in the field. We conduct an in-depth analysis of existing works, highlighting their advantages and limitations, and explore future research directions to inspire the design of models for related tasks. We summarize the basic principles of designing KD-based OD models, describe related KD-based OD tasks, including performance improvements for lightweight models, catastrophic forgetting in incremental OD, small object detection, and weakly/semi-supervised OD. We also analyze novel distillation techniques, i.e. different types of distillation loss, feature interaction between teacher and student models, etc. Additionally, we provide an overview of the extended applications of KD-based OD models on specific datasets, such as remote sensing images and 3D point cloud datasets. We compare and analyze the performance of different models on several common datasets and discuss promising directions for solving specific OD problems.
Zhihui Li 0001, Pengfei Xu 0003, Xiaojun Chang, Luyao Yang, Lina Yao 0001, Xiaojiang Chen
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 TN-ZSTAD: Transferable Network for Zero-Shot Temporal Activity Detection
abstract
An integral part of video analysis and surveillance is temporal activity detection, which means to simultaneously recognize and localize activities in long untrimmed videos. Currently, the most effective methods of temporal activity detection are based on deep learning, and they typically perform very well with large scale annotated videos for training. However, these methods are limited in real applications due to the unavailable videos about certain activity classes and the time-consuming data annotation. To solve this challenging problem, we propose a novel task setting called zero-shot temporal activity detection (ZSTAD), where activities that have never been seen in training still need to be detected. We design an end-to-end deep transferable network TN-ZSTAD as the architecture for this solution. On the one hand, this network utilizes an activity graph transformer to predict a set of activity instances that appear in the video, rather than produces many activity proposals in advance. On the other hand, this network captures the common semantics of seen and unseen activities from their corresponding label embeddings, and it is optimized with an innovative loss function that considers the classification property on seen activities and the transfer property on unseen activities together. Experiments on the THUMOS'14, Charades, and ActivityNet datasets show promising performance in terms of detecting unseen activities.
Lingling Zhang 0005, Xiaojun Chang, Jun Liu 0002, Minnan Luo, Zhihui Li 0001, Lina Yao 0001, Alex Hauptmann 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Trace ratio criterion for multi-view discriminant analysis
Mei Shi, Zhihui Li 0001, Xiaowei Zhao 0002, Pengfei Xu 0003, Baoying Liu, Jun Guo 0020
Appl. Intell.2
2022 ZeroNAS: Differentiable Generative Adversarial Networks Search for Zero-Shot Learning
abstract
In recent years, remarkable progress in zero-shot learning (ZSL) has been achieved by generative adversarial networks (GAN). To compensate for the lack of training samples in ZSL, a surge of GAN architectures have been developed by human experts through trial-and-error testing. Despite their efficacy, however, there is still no guarantee that these hand-crafted models can consistently achieve good performance across diversified datasets or scenarios. Accordingly, in this paper, we turn to neural architecture search (NAS) and make the first attempt to bring NAS techniques into the ZSL realm. Specifically, we propose a differentiable GAN architecture search method over a specifically designed search space for zero-shot learning, referred to as ZeroNAS. Considering the relevance and balance of the generator and discriminator, ZeroNAS jointly searches their architectures in a min-max player game via adversarial training. Extensive experiments conducted on four widely used benchmark datasets demonstrate that ZeroNAS is capable of discovering desirable architectures that perform favorably against state-of-the-art ZSL and generalized zero-shot learning (GZSL) approaches. Source code is at https://github.com/caixiay/ZeroNAS.
Caixia Yan, Xiaojun Chang, Zhihui Li 0001, Weili Guan, ZongYuan Ge, Lei Zhu 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Fuzzy K-Means Clustering With Discriminative Embedding
abstract
Fuzzy K-Means (FKM) clustering is of great importance for analyzing unlabeled data. FKM algorithms assign each data point to multiple clusters with some degree of certainty measured by the membership function. In these methods, the fuzzy membership degree matrix is obtained based on the calculation of the distance between data points in the original space. However, this operation may lead to suboptimal results because of the influence of noises and redundant features. Besides, some FKM clustering methods ignore the importance of the weighting exponent. In this paper, we propose a novel FKM method called Fuzzy K-Means Clustering With Discriminative Embedding. Within this method, we simultaneously conduct dimensionality reduction along with fuzzy membership degree learning. To retain most information in the embedding subspace and improve the robustness of this method, principal component analysis is incorporated into our framework. An iterative optimization algorithm is proposed to solve the model. To validate the efficacy of the proposed method, we perform comprehensive analyses, including convergence behavior, parameter determination and computational complexity. Moreover, we also match a appropriate weighting exponent for each data set. Experimental results on benchmark data sets show that the proposed method is more discriminative and effective for clustering tasks.
Feiping Nie 0001, Xiaowei Zhao 0002, Rong Wang 0001, Xuelong Li 0001, Zhihui Li 0001
IEEE Trans. Knowl. Data Eng.5
2022 Learning Adaptive Spatial-Temporal Context-Aware Correlation Filters for UAV Tracking
abstract
Tracking in the unmanned aerial vehicle (UAV) scenarios is one of the main components of target-tracking tasks. Different from the target-tracking task in the general scenarios, the target-tracking task in the UAV scenarios is very challenging because of factors such as small scale and aerial view. Although the discriminative correlation filter (DCF)-based tracker has achieved good results in tracking tasks in general scenarios, the boundary effect caused by the dense sampling method will reduce the tracking accuracy, especially in UAV-tracking scenarios. In this work, we propose learning an adaptive spatial-temporal context-aware (ASTCA) model in the DCF-based tracking framework to improve the tracking accuracy and reduce the influence of boundary effect, thereby enabling our tracker to more appropriately handle UAV-tracking tasks. Specifically, our ASTCA model can learn a spatial-temporal context weight, which can precisely distinguish the target and background in the UAV-tracking scenarios. Besides, considering the small target scale and the aerial view in UAV-tracking scenarios, our ASTCA model incorporates spatial context information within the DCF-based tracker, which could effectively alleviate background interference. Extensive experiments demonstrate that our ASTCA method performs favorably against state-of-the-art tracking methods on some standard UAV datasets.
Di Yuan 0002, Xiaojun Chang, Zhihui Li 0001, Zhenyu He 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Three Birds with One Stone: Multi-Task Temporal Action Detection via Recycling Temporal Annotations
abstract
Temporal action detection on unconstrained videos has seen significant research progress in recent years. Deep learning has achieved enormous success in this direction. However, collecting large-scale temporal detection datasets to ensuring promising performance in the real-world is a laborious, impractical and time consuming process. Accordingly, we present a novel improved temporal action localization model that is better able to take advantage of limited labeled data available. Specifically, we design two auxiliary tasks by reconstructing the available label information and then facilitate the learning of the temporal action detection model. Each task generates their supervision signal by recycling the original annotations, and are jointly trained with the temporal action detection model in a multi-task learning fashion. Note that the proposed approach can be pluggable to any region proposal based temporal action detection models. We conduct extensive experiments on three benchmark datasets, namely THUMOS’14 [15], Charades [35] and ActivityNet [14]. Our experimental results confirm the effectiveness of the proposed model.
Zhihui Li 0001, Lina Yao 0001
CVPR1
2021 Dynamic Slimmable Network
abstract
Current dynamic networks and dynamic pruning methods have shown their promising capability in reducing theoretical computation complexity. However, dynamic sparse patterns on convolutional filters fail to achieve actual acceleration in real-world implementation, due to the extra burden of indexing, weight-copying, or zero-masking. Here, we explore a dynamic network slimming regime, named Dynamic Slimmable Network (DS-Net), which aims to achieve good hardware-efficiency via dynamically adjusting filter numbers of networks at test time with respect to different inputs, while keeping filters stored statically and contiguously in hardware to prevent the extra burden. Our DS-Net is empowered with the ability of dynamic inference by the proposed double-headed dynamic gate that comprises an attention head and a slimming head to predictively adjust network width with negligible extra computation cost. To ensure generality of each candidate architecture and the fairness of gate, we propose a disentangled two-stage training scheme inspired by one-shot NAS. In the first stage, a novel training technique for weight-sharing networks named In-place Ensemble Bootstrapping is proposed to improve the supernet training efficacy. In the second stage, Sandwich Gate Sparsification is proposed to assist the gate training by identifying easy and hard samples in an online way. Extensive experiments demonstrate our DS-Net consistently outperforms its static counterparts as well as state-of-the-art static and dynamic model compression methods by a large margin (up to 5.9%). Typically, DS-Net achieves 2-4× computation reduction and 1.62× real-world acceleration over ResNet-50 and MobileNet with minimal accuracy drops on ImageNet.1
Guangrun Wang, Bing Wang 0013, Xiaodan Liang, Zhihui Li 0001, Xiaojun Chang
CVPR5
2021 Weighted multi-view common subspace learning method
Mei Shi, Jun Guo 0020, Xiaoqing Gong, Zhihui Li 0001
Pattern Recognit. Lett.6
2021 Self-weighted Robust LDA for Multiclass Classification with Edge Classes
abstract
Linear discriminant analysis (LDA) is a popular technique to learn the most discriminative features for multi-class classification. A vast majority of existing LDA algorithms are prone to be dominated by the class with very large deviation from the others, i.e., edge class, which occurs frequently in multi-class classification. First, the existence of edge classes often makes the total mean biased in the calculation of between-class scatter matrix. Second, the exploitation of ℓ2-norm based between-class distance criterion magnifies the extremely large distance corresponding to edge class. In this regard, a novel self-weighted robust LDA with ℓ2,1-norm based pairwise between-class distance criterion, called SWRLDA, is proposed for multi-class classification especially with edge classes. SWRLDA can automatically avoid the optimal mean calculation and simultaneously learn adaptive weights for each class pair without setting any additional parameter. An efficient re-weighted algorithm is exploited to derive the global optimum of the challenging ℓ2,1-norm maximization problem. The proposed SWRLDA is easy to implement and converges fast in practice. Extensive experiments demonstrate that SWRLDA performs favorably against other compared methods on both synthetic and real-world datasets while presenting superior computational efficiency in comparison with other techniques.
Caixia Yan, Xiaojun Chang, Minnan Luo, Xiaoqin Zhang 0002, Zhihui Li 0001, Feiping Nie 0001
ACM Trans. Intell. Syst. Technol.6
2020 Adaptive Two-Dimensional Embedded Image Clustering
abstract
With the rapid development of mobile devices, people are generating huge volumes of images data every day for sharing on social media, which draws much research attention to understanding the contents of images. Image clustering plays an important role in image understanding systems. Often, most of the existing image clustering algorithms flatten digital images that are originally represented by matrices into 1D vectors as the image representation for the subsequent learning. The drawbacks of vector-based algorithms include limited consideration of spatial relationship between pixels and computational complexity, both of which blame to the simple vectorized representation. To overcome the drawbacks, we propose a novel image clustering framework that can work directly on matrices of images instead of flattened vectors. Specifically, the proposed algorithm simultaneously learn the clustering results and preserve the original correlation information within the image matrix. To solve the challenging objective function, we propose a fast iterative solution. Extensive experiments have been conducted on various benchmark datasets. The experimental results confirm the superiority of the proposed algorithm.
Zhihui Li 0001, Lina Yao 0001, Sen Wang 0001, Salil S. Kanhere, Xue Li 0001, Huaxiang Zhang 0001
AAAI1
2020 Robust Self-Weighted Multi-View Projection Clustering
abstract
Many real-world applications involve data collected from different views and with high data dimensionality. Furthermore, multi-view data always has unavoidable noise. Clustering on this kind of high-dimensional and noisy multi-view data remains a challenge due to the curse of dimensionality and ineffective de-noising and integration of multiple views. Aiming at this problem, in this paper, we propose a Robust Self-weighted Multi-view Projection Clustering (RSwMPC) based on ℓ2,1-norm, which can simultaneously reduce dimensionality, suppress noise and learn local structure graph. Then the obtained optimal graph can be directly used for clustering while no further processing is required. In addition, a new method is introduced to automatically learn the optimal weight of each view with no need to generate additional parameters to adjust the weight. Extensive experimental results on different synthetic datasets and real-world datasets demonstrate that the proposed algorithm outperforms other state-of-the-art methods on clustering performance and robustness.
Beilei Wang, Zhihui Li 0001, Xuanhong Wang, Xiaojiang Chen, Dingyi Fang
AAAI3
2020 Grounding Visual Concepts for Zero-Shot Event Detection and Event Captioning
abstract
The flourishing of social media platforms requires techniques for understanding the content of media on a large scale. However, state-of-the art video event understanding approaches remain very limited in terms of their ability to deal with data sparsity, semantically unrepresentative event names, and lack of coherence between visual and textual concepts. Accordingly, in this paper, we propose a method of grounding visual concepts for large-scale Multimedia Event Detection (MED) and Multimedia Event Captioning (MEC) in zero-shot setting. More specifically, our framework composes the following: (1) deriving the novel semantic representations of events from their textual descriptions, rather than event names; (2) aggregating the ranks of grounded concepts for MED tasks. A statistical mean-shift outlier rejection model is proposed to remove the outlying concepts which are incorrectly grounded; and (3) defining MEC tasks and augmenting the MEC training set by the videos detected in MED in a zero-shot setting. To the best of our knowledge, this work is the first time to define and solve the MEC task, which is a further step towards understanding video events. We conduct extensive experiments and achieve state-of-the-art performance on the TRECVID MEDTest dataset, as well as our newly proposed TRECVID-MEC dataset.
Zhihui Li 0001, Xiaojun Chang, Lina Yao 0001, Shirui Pan, ZongYuan Ge, Huaxiang Zhang 0001
KDD1
2020 Visual saliency guided complex image retrieval
Haoxiang Wang 0001, Zhihui Li 0001, Brij B. Gupta, Chang Choi
Pattern Recognit. Lett.2
2020 Diverse fuzzy c-means for image clustering
Lingling Zhang 0005, Minnan Luo, Jun Liu 0002, Zhihui Li 0001
Pattern Recognit. Lett.4
2020 Fusion of Multiple Person Re-id Methods With Model and Data-Aware Abilities
abstract
Person re-identification (person re-id) has attracted rapidly increasing attention in computer vision and pattern recognition research community in recent years. With the goal of providing match ranking results between each query person image and the gallery ones, the person re-id technique has been widely explored and a large number of person re-id methods have been developed. As these algorithms leverage different kinds of prior assumptions, image features, distance matching functions, et al., each of them has its own strengths and weaknesses. Inspired by these facts, this paper proposes a novel person re-id method based on the idea of inferring superior fusion results from a variety of previous base person re-id algorithms using different methodologies or features. To achieve this goal, we propose a novel framework which mainly consists of two steps: 1) a number of existing person re-id methods are implemented, and the ranking results are obtained in the test datasets. and 2) the robust fusion strategy is applied to obtain better re-ranked matching results by simultaneously considering the recognition abilities of various base re-id methods and the difficulties of different gallery person images to be correctly recognized under the generative model of labels, abilities, and difficulties framework. Comprehensive experiments show the effectiveness of our proposed method, and we have received state-of-the-art results on recent popular person re-id datasets.
De Cheng, Zhihui Li 0001, Yihong Gong, Dingwen Zhang
IEEE Trans. Cybern.2
2020 Joint Principal Component and Discriminant Analysis for Dimensionality Reduction
abstract
via principal component analysis (PCA), the LDA algorithm can avoid the small sample size problem. Most existing supervised dimensionality reduction methods extract the principal component of data first, and then conduct LDA on it. However, "most variance" is very often the most important, but not always in PCA. Thus, this two-step strategy may not be able to obtain the most discriminant information for classification tasks. Different from traditional approaches which conduct PCA and LDA in sequence, we propose a novel method referred to as joint principal component and discriminant analysis (JPCDA) for dimensionality reduction. Using this method, we are able to not only avoid the small sample size problem but also extract discriminant information for classification tasks. An iterative optimization algorithm is proposed to solve the method. To validate the efficacy of the proposed method, we perform extensive experiments on several benchmark data sets in comparison with some state-of-the-art dimensionality reduction methods. A large number of experimental results illustrate that the proposed method has quite promising classification performance.
Xiaowei Zhao 0002, Jun Guo 0020, Feiping Nie 0001, Ling Chen 0006, Zhihui Li 0001, Huaxiang Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2019 Zero-Shot Object Detection with Textual Descriptions
abstract
Object detection is important in real-world applications. Existing methods mainly focus on object detection with sufficient labelled training data or zero-shot object detection with only concept names. In this paper, we address the challenging problem of zero-shot object detection with natural language description, which aims to simultaneously detect and recognize novel concept instances with textual descriptions. We propose a novel deep learning framework to jointly learn visual units, visual-unit attention and word-level attention, which are combined to achieve word-proposal affinity by an element-wise multiplication. To the best of our knowledge, this is the first work on zero-shot object detection with textual descriptions. Since there is no directly related work in the literature, we investigate plausible solutions based on existing zero-shot object detection for a fair comparison. We conduct extensive experiments on three challenging benchmark datasets. The extensive experimental results confirm the superiority of the proposed model.
Zhihui Li 0001, Lina Yao 0001, Xiaoqin Zhang 0002, Xianzhi Wang 0001, Salil S. Kanhere, Huaxiang Zhang 0001
AAAI1
2019 RS3CIS: Robust Single-Step Spectral Clustering with Intrinsic Subspace
Pengzhen Ren, Zhihui Li 0001, Xiaojiang Chen, Xin Wang 0004, Dingyi Fang
AAAI3
2019 Domain-Aware Unsupervised Cross-dataset Person Re-identification
Zhihui Li 0001, Wenhe Liu, Xiaojun Chang, Lina Yao 0001, Mahesh Prakash, Huaxiang Zhang 0001
ADMA1
2019 Knowledge driven temporal activity localization
Zhihui Li 0001, ZongYuan Ge, Mingjie Li 0006
J. Vis. Commun. Image Represent.2
2019 Zero-shot event detection via event-adaptive concept relevance mining
Zhihui Li 0001, Lina Yao 0001, Xiaojun Chang, Kun Zhan, Jiande Sun 0001, Huaxiang Zhang 0001
Pattern Recognit.1
2019 Spectral Clustering of Customer Transaction Data With a Two-Level Subspace Weighting Method
abstract
Finding customer groups from transaction data is very important for retail and e-commerce companies. Recently, a "Purchase Tree" data structure is proposed to compress the customer transaction data and a local PurTree spectral clustering method is proposed to cluster the customer transaction data. However, in the PurTree distance, the node weights for the children nodes of a parent node are set as equal and the differences between different nodes are not distinguished. In this paper, we propose a two-level subspace weighting spectral clustering (TSW) algorithm for customer transaction data. In the new method, a PurTree subspace metric is proposed to measure the dissimilarity between two customers represented by two purchase trees, in which a set of level weights are introduced to distinguish the importance of different tree levels and a set of sparse node weights are introduced to distinguish the importance of different tree nodes in a purchase tree. TSW learns an adaptive similarity matrix from the local distances in order to better uncover the cluster structure buried in the customer transaction data. Simultaneously, it learns a set of level weights and a set of sparse node weights in the PurTree subspace distance. An iterative optimization algorithm is proposed to optimize the proposed model. We also present an efficient method to compute a regularization parameter in TSW. TSW was compared with six clustering algorithms on ten benchmark data sets and the experimental results show the superiority of the new method.
Xiaojun Chen 0006, Wenya Sun, Zhihui Li 0001, Xizhao Wang, Yunming Ye
IEEE Trans. Cybern.4
2019 Large-Scale Robust Semisupervised Classification
abstract
Semisupervised learning aims to leverage both labeled and unlabeled data to improve performance, where most of them are graph-based methods. However, the graph-based semisupervised methods are not capable for large-scale data since the computational consumption on the construction of graph Laplacian matrix is huge. On the other hand, the substantial unlabeled data in training stage of semisupervised learning could cause large uncertainties and potential threats. Therefore, it is crucial to enhance the robustness of semisupervised classification. In this paper, a novel large-scale robust semisupervised learning method is proposed in the framework of capped ℓ2,p-norm. This strategy is superior not only in computational cost because it makes the graph Laplacian matrix unnecessary, but also in robustness to outliers since the capped ℓ2,p-norm used for loss measurement. An efficient optimization algorithm is exploited to solve the nonconvex and nonsmooth challenging problem. The complexity of the proposed algorithm is analyzed and discussed in theory detailedly. Finally, extensive experiments are conducted over six benchmark data sets to demonstrate the effectiveness and superiority of the proposed method.
Lingling Zhang 0005, Minnan Luo, Zhihui Li 0001, Feiping Nie 0001, Huaxiang Zhang 0001, Jun Liu 0002
IEEE Trans. Cybern.3
2018 Trace Ratio Optimization With Feature Correlation Mining for Multiclass Discriminant Analysis
abstract
Fisher's linear discriminant analysis is a widely accepted dimensionality reduction method, which aims to find a transformation matrix to convert feature space to a smaller space by maximising the between-class scatter matrix while minimising the within-class scatter matrix. Although the fast and easy process of finding the transformation matrix has made this method attractive, overemphasizing the large class distances makes the criterion of this method suboptimal. In this case, the close class pairs tend to overlap in the subspace. Despite different weighting methods having been developed to overcome this problem, there is still a room to improve this issue. In this work, we study a weighted trace ratio by maximising the harmonic mean of the multiple objective reciprocals. To further improve the performance, we enforce the l2,1-norm to the developed objective function. Additionally, we propose an iterative algorithm to optimise this objective function. The proposed method avoids the domination problem of the largest objective, and guarantees that no objectives will be too small. This method can be more beneficial if the number of classes is large. The extensive experiments on different datasets show the effectiveness of our proposed method when compared with four state-of-the-art methods.
Forough Rezaei Boroujeni, Sen Wang 0001, Zhihui Li 0001, Nicholas West, Bela Stantic, Lina Yao 0001, Guodong Long
AAAI3
2018 Balanced Clustering via Exclusive Lasso: A Pragmatic Approach
abstract
Clustering is an effective technique in data mining to generate groups that are the matter of interest.Among various clustering approaches, the family of k-means algorithms and min-cut algorithms gain most popularity due to their simplicity and efficacy. The classical k-means algorithm partitions a number of data points into several subsets by iteratively updating the clustering centers and the associated data points. By contrast, a weighted undirected graph is constructed in min-cut algorithms which partition the vertices of the graph into two sets. However, existing clustering algorithms tend to cluster minority of data points into a subset, which shall be avoided when the target dataset is balanced. To achieve more accurate clustering for balanced dataset, we propose to leverage exclusive lasso on k-means and min-cut to regulate the balance degree of the clustering results. By optimizing our objective functions that build atop the exclusive lasso, we can make the clustering result as much balanced as possible. Extensive experiments on several large-scale datasets validate the advantage of the proposed algorithms compared to the state-of-the-art clustering algorithms.
Zhihui Li 0001, Feiping Nie 0001, Xiaojun Chang, Zhigang Ma, Yi Yang 0001
AAAI1
2018 Multi-Rate Gated Recurrent Convolutional Networks for Video-Based Pedestrian Re-Identification
abstract
Matching pedestrians across multiple camera views has attracted lots of recent research attention due to its apparent importance in surveillance and security applications.While most existing works address this problem in a still-image setting, we consider the more informative and challenging video-based person re-identification problem, where a video of a pedestrian as seen in one camera needs to be matched to a gallery of videos captured by other non-overlapping cameras. We employ a convolutional network to extract the appearance and motion features from raw video sequences, and then feed them into a multi-rate recurrent network to exploit the temporal correlations, and more importantly, to take into account the fact that pedestrians, sometimes even the same pedestrian, move in different speeds across different camera views. The combined network is trained in an end-to-end fashion, and we further propose an initialization strategy via context reconstruction to largely improve the performance. We conduct extensive experiments on the iLIDS-VID and PRID-2011 datasets, and our experimental results confirm the effectiveness and the generalization ability of our model.
Zhihui Li 0001, Lina Yao 0001, Feiping Nie 0001, Dingwen Zhang
AAAI1
2018 Top-k multi-class SVM using multiple features
Caixia Yan, Minnan Luo, Huan Liu 0012, Zhihui Li 0001
Inf. Sci.4
2018 Unsupervised multi-view feature extraction with dynamic graph learning
Dan Shi 0003, Lei Zhu 0002, Zhiyong Cheng 0001, Zhihui Li 0001, Huaxiang Zhang 0001
J. Vis. Commun. Image Represent.4
2018 An efficient multi-feature SVM solver for complex event detection
Huan Liu 0012, Zhihui Li 0001, Tao Qin 0002, Lei Zhu 0002
Multim. Tools Appl.3
2018 A Multiview Learning Framework With a Linear Computational Cost
abstract
Learning features from multiple views has attracted much research attention in different machine learning tasks, such as multiclass and multilabel classification problems. In this paper, we propose a multiclass multilabel multiview learning framework with a linear computational cost where an example is associated with at least one label and represented by multiple information sources. We simultaneously analyze all features by learning an integrated projection matrix. We can also automatically select more important views for subsequent classifier to predict each class. As the proposed objective function is nonsmooth and difficult to solve, we apply a novel optimization method that converts the multiview learning problem to a set of linear single-view learning problems by bridging our problem to an easily solvable approach. Compared to the conventional methods which learn the entire projection matrix, our algorithm independently optimizes each column of the projection matrix for each class, which can be easily parallelized. In each column optimization, the most computationally intensive step is pure and simple matrix-by-vector multiplication. As a result, our algorithm is much more applicable to large-scale problems than the multiview learning methods with a nonlinear computational cost. Moreover, rigorous convergence proof of the proposed algorithm is also provided. To evaluate the effectiveness of the proposed approach, experimental comparisons are made with state-of-the-art algorithms in multiclass and multilabel classification tasks on many multiview benchmarks. We also report the efficiency comparison results on different numbers of data samples. The experimental results demonstrate that our algorithm can achieve superior performance to all the compared algorithms.
Xiaowei Xue, Feiping Nie 0001, Zhihui Li 0001, Sen Wang 0001, Xue Li 0001
IEEE Trans. Cybern.3
2018 Two-Stream Multirate Recurrent Neural Network for Video-Based Pedestrian Reidentification
abstract
Video-based pedestrian reidentification is an emerging task in video surveillance and is closely related to several real-world applications. Its goal is to match pedestrians across multiple nonoverlapping network cameras. Despite the recent effort, the performance of pedestrian reidentification needs further improvement. Hence, we propose a novel two-stream multirate recurrent neural network for video-based pedestrian reidentification with two inherent advantages: First, capturing the static spatial and temporal information; Second,Author: Figure II is not cited in the text. Please cite it at the appropriate place. dealing with motion speed variance. Given video sequences of pedestrians, we start with extracting spatial and motion features using two different deep neural networks. Then, we explore the feature correlation which results in a regularized fusion network integrating the two aforementioned networks. Considering that pedestrians, sometimes even the same pedestrian, move in different speeds across different camera views, we extend our approach by feeding the two networks into a multirate recurrent network to exploit the temporal correlations. Extensive experiments have been conducted on two real-world video-based pedestrian reidentification benchmarks: iLIDS-VID and PRID 2011 datasets. The experimental results confirm the efficacy of the proposed method. Our code will be released upon acceptance.
Zhihui Li 0001, De Cheng, Huaxiang Zhang 0001, Kun Zhan, Yi Yang 0001
IEEE Trans. Ind. Informatics2
2018 Multi-Modal Joint Clustering With Application for Unsupervised Attribute Discovery
abstract
Utilizing multiple descriptions/views of an object is often useful in image clustering tasks. Despite many works that have been proposed to effectively cluster multi-view data, there are still unaddressed problems such as the errors introduced by the traditional spectral-based clustering methods due to the two disjoint stages: 1) eigendecomposition and 2) the discretization of new representations. In this paper, we propose a unified clustering framework which jointly learns the two stages together as well as utilizing multiple descriptions of the data. More specifically, two learning methods from this framework are proposed: 1) through a graph construction from different views and 2) through combining multiple graphs. Furthermore, benefiting from the separability and local graph preserving properties of the proposed methods, a novel unsupervised automatic attribute discovery method is proposed. We validate the efficacy of our methods on five data sets, showing that the proposed joint learning clustering methods outperform the recent state-of-the-art methods. We also show that it is possible to derive a novel method to address the unsupervised automatic attribute discovery tasks.
Feiping Nie 0001, Arnold Wiliem, Zhihui Li 0001, Teng Zhang 0004, Brian C. Lovell
IEEE Trans. Image Process.4
2018 Rank-Constrained Spectral Clustering With Flexible Embedding
abstract
Spectral clustering (SC) has been proven to be effective in various applications. However, the learning scheme of SC is suboptimal in that it learns the cluster indicator from a fixed graph structure, which usually requires a rounding procedure to further partition the data. Also, the obtained cluster number cannot reflect the ground truth number of connected components in the graph. To alleviate these drawbacks, we propose a rank-constrained SC with flexible embedding framework. Specifically, an adaptive probabilistic neighborhood learning process is employed to recover the block-diagonal affinity matrix of an ideal graph. Meanwhile, a flexible embedding scheme is learned to unravel the intrinsic cluster structure in low-dimensional subspace, where the irrelevant information and noise in high-dimensional data have been effectively suppressed. The proposed method is superior to previous SC methods in that: 1) the block-diagonal affinity matrix learned simultaneously with the adaptive graph construction process, more explicitly induces the cluster membership without further discretization; 2) the number of clusters is guaranteed to converge to the ground truth via a rank constraint on the Laplacian matrix; and 3) the mismatch between the embedded feature and the projected feature allows more freedom for finding the proper cluster structure in the low-dimensional subspace as well as learning the corresponding projection matrix. Experimental results on both synthetic and real-world data sets demonstrate the promising performance of the proposed algorithm.
Zhihui Li 0001, Feiping Nie 0001, Xiaojun Chang, Liqiang Nie, Huaxiang Zhang 0001, Yi Yang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Dynamic Affinity Graph Construction for Spectral Clustering Using Multiple Features
abstract
Spectral clustering (SC) has been widely applied to various computer vision tasks, where the key is to construct a robust affinity matrix for data partitioning. With the increase in visual features, conventional SC methods are facing two challenges: 1) how to effectively generate an affinity matrix based on multiple features? and 2) how to deal with high-dimensional visual features which could be redundant? To address these issues mentioned earlier, we present a new approach to: 1) learn a robust affinity matrix using multiple features, allowing us to simultaneously determine optimal weights for each feature; and 2) decide a set of optimal projection matrixes, one for each feature, that decide the lower dimensional space, as well as the optimal affinity weight of each data pair in the lower dimensional space. There are two major advantages of our new approach over the existing clustering techniques. First, our approach assigns affinity weights for data points on a per-data-pair basis. The learning procedure avoids the explicit specification of the size of the neighborhood in the affinity matrix, and the bandwidth parameter required to compute the Gaussian kernel, both of which are sensitive and yet difficult to determine beforehand. Second, the affinity weights are based on the distances in a lower dimensional space, while the low-dimensional space is inferred according to the optimized affinity weights. Both variables are jointly optimized so as to leverage mutual benefits. The experimental results outperform the compared alternatives, which indicate that the proposed method is effective in simultaneously learning the affinity graph and feature fusion, resulting in better clustering results.
Zhihui Li 0001, Feiping Nie 0001, Xiaojun Chang, Yi Yang 0001, Chengqi Zhang, Nicu Sebe
IEEE Trans. Neural Networks Learn. Syst.1
2018 Exploring Auxiliary Context: Discrete Semantic Transfer Hashing for Scalable Image Retrieval
abstract
Unsupervised hashing can desirably support scalable content-based image retrieval for its appealing advantages of semantic label independence, memory, and search efficiency. However, the learned hash codes are embedded with limited discriminative semantics due to the intrinsic limitation of image representation. To address the problem, in this paper, we propose a novel hashing approach, dubbed as discrete semantic transfer hashing (DSTH). The key idea is to directly augment the semantics of discrete image hash codes by exploring auxiliary contextual modalities. To this end, a unified hashing framework is formulated to simultaneously preserve visual similarities of images and perform semantic transfer from contextual modalities. Furthermore, to guarantee direct semantic transfer and avoid information loss, we explicitly impose the discrete constraint, bit-uncorrelation constraint, and bit-balance constraint on hash codes. A novel and effective discrete optimization method based on augmented Lagrangian multiplier is developed to iteratively solve the optimization problem. The whole learning process has linear computation complexity and desirable scalability. Experiments on three benchmark data sets demonstrate the superiority of DSTH compared with several state-of-the-art approaches.
Lei Zhu 0002, Zi Huang, Zhihui Li 0001, Liang Xie 0001, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.3
2017 Simple to complex cross-modal learning to rank
Minnan Luo, Xiaojun Chang, Zhihui Li 0001, Liqiang Nie, Alex Hauptmann 0001
Comput. Vis. Image Underst.3
2017 Refined Spectral Clustering via Embedded Label Propagation
abstract
Spectral clustering is a key research topic in the field of machine learning and data mining. Most of the existing spectral clustering algorithms are built on gaussian Laplacian matrices, which is sensitive to parameters. We propose a novel parameter-free distance-consistent locally linear embedding. The proposed distance-consistent LLE can promise that edges between closer data points are heavier. We also propose a novel improved spectral clustering via embedded label propagation. Our algorithm is built on two advancements of the state of the art. First is label propagation, which propagates a node's labels to neighboring nodes according to their proximity. We perform standard spectral clustering on original data and assign each cluster with [Formula: see text]-nearest data points and then we propagate labels through dense unlabeled data regions. Second is manifold learning, which has been widely used for its capacity to leverage the manifold structure of data points. Extensive experiments on various data sets validate the superiority of the proposed algorithm compared to state-of-the-art spectral algorithms.
Yan-shuo Chang, Feiping Nie 0001, Zhihui Li 0001, Xiaojun Chang, Heng Huang 0001
Neural Comput.3
2017 Beyond Trace Ratio: Weighted Harmonic Mean of Trace Ratios for Multiclass Discriminant Analysis
abstract
Linear discriminant analysis (LDA) is one of the most important supervised linear dimensional reduction techniques which seeks to learn low-dimensional representation from the original high-dimensional feature space through a transformation matrix, while preserving the discriminative information via maximizing the between-class scatter matrix and minimizing the within class scatter matrix. However, the conventional LDA is formulated to maximize the arithmetic mean of trace ratios which suffers from the domination of the largest objectives and might deteriorate the recognition accuracy in practical applications with a large number of classes. In this paper, we propose a new criterion to maximize the weighted harmonic mean of trace ratios, which effectively avoid the domination problem while did not raise any difficulties in the formulation. An efficient algorithm is exploited to solve the proposed challenging problems with fast convergence, which might always find the globally optimal solution just using eigenvalue decomposition in each iteration. Finally, we conduct extensive experiments to illustrate the effectiveness and superiority of our method over both of synthetic datasets and real-life datasets for various tasks, including face recognition, human motion recognition and head pose recognition. The experimental results indicate that our algorithm consistently outperforms other compared methods on all of the datasets.
Zhihui Li 0001, Feiping Nie 0001, Xiaojun Chang, Yi Yang 0001
IEEE Trans. Knowl. Data Eng.1