VLDB 2026 Research / reviewers in the wild / expert
Lei Wang 0001
dblp:w/LeiWang1
· DBLP profile ↗
194ranked-venue papers
20as first author
72since 2021 · last 2026
0000-0002-0961-0441ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 132 · 13 first-author · 49 since 2021Graphics, computer vision, multimedia, augmented reality and games · 105 · 10 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 7 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorSystems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report GenerationabstractAutomated radiology report generation (R2Gen) has advanced significantly, yet evaluation remains challenging due to the complexity of assessing report quality. Traditional metrics often misalign with human judgments, failing to identify specific deficiencies. To address this, we introduce ReFINE, a framework for training an Evaluation Model using a novel margin-based reward enforcement loss. This approach decomposes report quality into fine-grained sub-scores across user-defined criteria, improving interpretability. Leveraging GPT-4, we generate diverse training data with paired accepted and rejected reports to train our model under a reward-based system. The trained ReFINE Score provides both granular sub-scores and an aggregated quality assessment, enabling criterion-specific evaluation. Experimental results demonstrate ReFINE's superior alignment with human judgments, outperforming traditional metrics in model selection. Its robustness is validated across three expert-annotated datasets—including chest X-rays and multimodal reports covering 9 imaging modalities—and under two distinct scoring systems. Yunyi Liu, Yingshu Li 0001, Zhanyu Wang, Lingqiao Liu, Lei Wang 0001, Luping Zhou |
AAAI | 6 |
| 2026 | Global-and-local guidance with synthesized view for unpaired multi-view clustering
Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014 |
Inf. Process. Manag. | 3 |
| 2026 | A continual learning framework with long-term and multiple short-term memory networks
Shangge Liu, Lei Wang 0001, Rui Yan 0005, Jing Huo, Wenbin Li 0006, Yang Gao 0001 |
Neural Networks | 2 |
| 2026 | AMPL: An adaptive meta-prompt learner for few-shot image classification
Zhiping Wu, Lian Huai, Zeyu Shangguan, Lei Wang 0001, Jing Huo, Wenbin Li 0006, Yang Gao 0001, Xingqun Jiang |
Neural Networks | 5 |
| 2025 | Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-TuningabstractRecent advancements in Multimodal Large Language Models (MLLMs) have generated significant interest in their ability to autonomously interact with and interpret Graphical User Interfaces (GUIs). A major challenge in these systems is grounding—accurately identifying critical GUI components such as text or icons based on a GUI image and a corresponding text query. Traditionally, this task has relied on fine-tuning MLLMs with specialized training data to predict component locations directly. However, in this paper, we propose a novel Tuning-free Attention-driven Grounding (TAG) method that leverages the inherent attention patterns in pretrained MLLMs to accomplish this task without the need for additional fine-tuning. Our method involves identifying and aggregating attention maps from specific tokens within a carefully constructed query prompt. Applied to MiniCPM-Llama3-V 2.5, a state-of-the-art MLLM, our tuning-free approach achieves performance comparable to tuning-based methods, with notable success in text localization. Additionally, we demonstrate that our attention map-based grounding technique significantly outperforms direct localization predictions from MiniCPM-Llama3-V 2.5, highlighting the potential of using attention maps from pretrained MLLMs and paving the way for future innovations in this domain. Qi Chen 0014, Lei Wang 0001, Lingqiao Liu |
AAAI | 3 |
| 2025 | Enhancing Few-Shot Class-Incremental Learning via Training-Free Bi-Level Modality CalibrationabstractFew-shot Class-Incremental Learning (FSCIL) challenges models to adapt to new classes with limited samples, presenting greater difficulties than traditional class-incremental learning. While existing approaches rely heavily on visual models and require additional training during base or incremental phases, we propose a training-free framework that leverages pre-trained visual-language models like CLIP. At the core of our approach is a novel Bi-level Modality Calibration (BiMC) strategy. Our framework initially performs intra-modal calibration, combining LLM-generated fine-grained category descriptions with visual prototypes from the base session to achieve precise classifier estimation. This is further complemented by inter-modal calibration that fuses pre-trained linguistic knowledge with task-specific visual priors to mitigate modality-specific biases. To enhance prediction robustness, we introduce additional metrics and strategies that maximize the utilization of limited data. Extensive experimental results demonstrate that our approach significantly outperforms existing methods. Code is available at: https://github.com/yychen016/BiMC. Tianyu Ding, Lei Wang 0001, Jing Huo, Yang Gao 0001, Wenbin Li 0006 |
CVPR | 3 |
| 2025 | Dynamic Multi-Layer Null Space Projection for Vision-Language Continual Learning
Borui Kang, Lei Wang 0001, Zhiping Wu, Yawen Li 0001, Yang Gao 0001, Wenbin Li 0006 |
ICCV | 2 |
| 2025 | Optimizing Efficiency and Visual-Textual Alignment for LLM-Based Radiology Report GenerationabstractLLM-based radiology report generation (R2Gen) systems have demonstrated promising performance but face significant challenges in bridging the gap between the visual encoder and the LLM. Specifically, two issues hinder progress: (1) parameter-heavy visual projector that increases complexity and degrades performance, and (2) insufficient alignment between visual and textual modalities, limiting system efficacy. To address these, we propose R2Gen-EVA, a novel framework emphasizing Efficiency and Visual-Textual Alignment (VTA), which introduces two key innovations: (1) a parameter-free visual projector that enhances model efficiency while improving performance, and (2) an LLM-adapted VTA module that enhances the alignment of visual features with LLM’s textual embeddings. Our design significantly improves model efficacy without adding extra parameters, achieving both streamlined complexity and higher computational efficiency during inference. Extensive experiments demonstrate that R2Gen-EVA enhances the fluency and clinical accuracy of generated reports, establishing it as a more effective and efficient solution for LLM-based R2Gen. The code is available at https://github.com/zailongchen/R2Gen-EVA. Zailong Chen, Yujian Lee, Johan Barthelemy, Luping Zhou, Lei Wang 0001 |
ICME | 6 |
| 2025 | Adaptive Gradient Learning for Spiking Neural Networks by Exploiting Membrane Potential DynamicsabstractRecent advancements have focused on directly training high-performance spiking neural networks (SNNs) by estimating the approximate gradients of spiking activity through a continuous function with constant sharpness, known as surrogate gradient (SG) learning. However, as spikes propagate within neurons and among layers, the distribution of membrane potential dynamics (MPD) will deviate from the gradient-available interval of fixed SG, hindering SNNs from searching the optimal solution space. To maintain the stability of gradient flows, SG needs to align with evolving MPD. Here, we propose a novel adaptive gradient learning for SNNs by exploiting MPD, namely MPD-AGL. It fully accounts for the underlying factors contributing to membrane potential shifts and establishes a dynamic association between SG and MPD at different timesteps to relax gradient estimation, which provides a new degree of freedom for SG learning. Experimental results demonstrate that our method achieves excellent performance at low latency. Moreover, it increases the proportion of neurons that fall into the gradient-available interval compared to fixed SG, effectively mitigating the gradient vanishing problem. Code is available at https://github.com/jqjiang1999/MPD-AGL. Jiaqiang Jiang, Lei Wang 0001, Runhao Jiang, Rui Yan 0005 |
IJCAI | 2 |
| 2025 | Collaborative 3D object detection by smart vehicles considering semantic information and agent heterogeneity
Yongcai Wang, Deying Li 0001, Yunjun Han, Lei Wang 0001 |
Adv. Eng. Informatics | 5 |
| 2025 | ONNXPruner: ONNX-Based General Model Pruning AdapterabstractRecent advancements in model pruning have focused on developing new algorithms and improving upon benchmarks. However, the practical application of these algorithms across various models and platforms remains a significant challenge. To address this challenge, we propose ONNXPruner, a versatile pruning adapter designed for the ONNX format models. ONNXPruner streamlines the adaptation process across diverse deep learning frameworks and hardware platforms. A novel aspect of ONNXPruner is its use of node association trees, which automatically adapt to various model architectures. These trees clarify the structural relationships between nodes, guiding the pruning process, particularly highlighting the impact on interconnected nodes. Furthermore, we introduce a tree-level evaluation method. By leveraging node association trees, this method allows for a comprehensive analysis beyond traditional single-node evaluations, enhancing pruning performance without the need for extra operations. Experiments across multiple models and datasets confirm ONNXPruner's strong adaptability and increased efficacy. Our work aims to advance the practical application of model pruning. Dongdong Ren, Wenbin Li 0006, Tianyu Ding, Lei Wang 0001, Jing Huo, Hongbing Pan, Yang Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Diagnostic Captioning by Cooperative Task Interactions and Sample-Graph ConsistencyabstractRadiographic images are similar to each other, making it challenging for diagnostic captioning to narrate fine-grained visual differences of clinical importance. In this paper, we propose a self-boosting framework integrating two novel strategies to learn tightly correlated image and text features for diagnostic captioning. The first strategy explicitly aligns image and text features through training an auxiliary task of image-text matching (ITM) jointly with the main task of report generation (RG) as two branches of a network model. The ITM branch explicitly learns image-text alignment and provides highly correlated visual and textual features for the RG branch to generate high-quality reports. The high-quality reports generated by RG branch, in turn, are utilized as additional harder negative samples to push the ITM branch to evolve towards better image-text alignment. These two branches help improve each other progressively, so that the whole model is self-boosted without requiring external resources. The second strategy aligns image-sample space and report-sample space to achieve consistent image and text feature embeddings. To achieve this, the sample graph of the embedded ground-truth reports is built and used as the target to train the sample graph of the embedded images so that the fine discrepancy in the ground-truth reports could be captured by the learned visual feature embeddings. Our proposed framework demonstrates its superiority on two medical report generation benchmarks, including the largest dataset MIMIC-CXR. Zhanyu Wang, Lei Wang 0001, Xiu Li 0001, Luping Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Leveraging Frequency Analysis for Image Denoising Network PruningabstractAs a common model compression technique, network pruning is widely used to reduce storage and computational cost of deep models in the resource-constrained regime. However, most current pruning methods are designed for high-level vision tasks, with few developed for low-level vision tasks. We observed that the norm-based pruning criterion, originally designed for high-level vision tasks, is highly unsuitable for low-level image denoising networks. This difference arises because image denoising networks pursue distinct feature granularities and goals compared to typical high-level vision tasks. To address this issue, we propose a novel filter evaluation method, termed High-Frequency Components Pruning (HFCP), specifically tailored for image denoising network pruning. HFCP assesses filter importance based on high-frequency components. To the best of our knowledge, this is the first pruning method designed specifically for image denoising tasks, straightforward and applicable to various types of noise. Furthermore, HFCP enhances the pruned model's high-frequency information content with high reliability and interpretability. This facilitates the network's ability to distinguish high-frequency signals from noise. We comprehensively analyzed multiple image denoising networks and validated HFCP's effectiveness across four mainstream networks. Dongdong Ren, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Hongbing Pan, Yang Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | E2MPL: An Enduring and Efficient Meta Prompt Learning Framework for Few-Shot Unsupervised Domain AdaptationabstractFew-shot unsupervised domain adaptation (FS-UDA) leverages a limited amount of labeled data from a source domain to enable accurate classification in an unlabeled target domain. Despite recent advancements, current approaches of FS-UDA continue to confront a major challenge: models often demonstrate instability when adapted to new FS-UDA tasks and necessitate considerable time investment. To address these challenges, we put forward a novel framework called Enduring and Efficient Meta-Prompt Learning (E2MPL) for FS-UDA. Within this framework, we utilize the pre-trained CLIP model as the backbone of feature learning. Firstly, we design domain-shared prompts, consisting of virtual tokens, which primarily capture meta-knowledge from a wide range of meta-tasks to mitigate the domain gaps. Secondly, we develop a task prompt learning network that adaptively learns task-specific prompts with the goal of achieving fast and stable task generalization. Thirdly, we formulate the meta-prompt learning process as a bilevel optimization problem, consisting of (outer) meta-prompt learner and (inner) task-specific classifier and domain adapter. Also, the inner objective of each meta-task has the closed-form solution, which enables efficient prompt learning and adaptation to new tasks in a single step. Extensive experimental studies demonstrate the promising performance of our framework in a domain adaptation benchmark dataset DomainNet. Compared with state-of-the-art methods, our approach has improved the average accuracy by at least 15 percentage points and reduces the average time by 64.67% in the 5-way 1-shot task; in the 5-way 5-shot task, it achieves at least a 9-percentage-point improvement in average accuracy and reduces the average time by 63.18%. Moreover, our method exhibits more enduring and stable performance than the other methods, i.e., reducing the average IQR value by over 40.80% and 25.35% in the 5-way 1-shot and 5-shot task, respectively. Wanqi Yang, Lei Wang 0001, Ming Yang 0014, Yang Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Enhancing Radiology Report Generation via Multi-Phased SupervisionabstractRadiology report generation using large language models has recently produced reports with more realistic styles and better language fluency. However, their clinical accuracy remains inadequate. Considering the significant imbalance between clinical phrases and general descriptions in a report, we argue that using an entire report for supervision is problematic as it fails to emphasize the crucial clinical phrases, which require focused learning. To address this issue, we propose a multi-phased supervision method, inspired by the spirit of curriculum learning where models are trained by gradually increasing task complexity. Our approach organizes the learning process into structured phases at different levels of semantical granularity, each building on the previous one to enhance the model. During the first phase, disease labels are used to supervise the model, equipping it with the ability to identify underlying diseases. The second phase progresses to use entity-relation triples to guide the model to describe associated clinical findings. Finally, in the third phase, we introduce conventional whole-report-based supervision to quickly adapt the model for report generation. Throughout the phased training, the model remains the same and consistently operates in the generation mode. As experimentally demonstrated, this proposed change in the way of supervision enhances report generation, achieving state-of-the-art performance in both language fluency and clinical accuracy. Our work underscores the importance of training process design in radiology report generation. Our code is available on https://github.com/zailongchen/MultiP-R2Gen. Zailong Chen, Yingshu Li 0002, Zhanyu Wang, Johan Barthelemy, Luping Zhou, Lei Wang 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Asynchronous Functional Brain Network Construction With Spatiotemporal Transformer for MCI ClassificationabstractConstruction and analysis of functional brain networks (FBNs) with resting-state functional magnetic resonance imaging (rs-fMRI) is a promising method to diagnose functional brain diseases. Nevertheless, the existing methods suffer from several limitations. First, the functional connectivities (FCs) of the FBN are usually measured by the temporal co-activation level between rs-fMRI time series from regions of interest (ROIs). While enjoying simplicity, the existing approach implicitly assumes simultaneous co-activation of all the ROIs, and models only their synchronous dependencies. However, the FCs are not necessarily always synchronous due to the time lag of information flow and cross-time interactions between ROIs. Therefore, it is desirable to model asynchronous FCs. Second, the traditional methods usually construct FBNs at individual level, leading to large variability and degraded diagnosis accuracy when modeling asynchronous FBN. Third, the FBN construction and analysis are conducted in two independent steps without joint alignment for the target diagnosis task. To address the first limitation, this paper proposes an effective sliding-window-based method to model spatiotemporal FCs in Transformer. Regarding the second limitation, we propose to learn common and individual FBNs adaptively with the common FBN as prior knowledge, thus alleviating the variability and enabling the network to focus on the individual disease-specific asynchronous FCs. To address the third limitation, the common and individual asynchronous FBNs are built and analyzed by an integrated network, enabling end-to-end training and improving the flexibility and discriminability. The effectiveness of the proposed method is consistently demonstrated on three data sets for mild cognitive impairment (MCI) diagnosis. Jianjia Zhang, Xiaotong Wu, Xiang Tang, Luping Zhou, Lei Wang 0001, Weiwen Wu, Dinggang Shen |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Selective Contrastive Learning for Unpaired Multi-View ClusteringabstractIn this article, we investigate a novel but insufficiently studied issue, unpaired multi-view clustering (UMC), where no paired observed samples exist in multi-view data, and the goal is to leverage the unpaired observed samples in all views for effective joint clustering. Existing methods in incomplete multi-view clustering usually utilize the sample pairing relationship between views to connect the views for joint clustering, but unfortunately, it is invalid for the UMC case. Therefore, we strive to mine a consistent cluster structure between views and propose an effective method, namely selective contrastive learning for UMC (scl-UMC), which needs to solve the following two challenging issues: 1) uncertain clustering structure under no supervision information and 2) uncertain pairing relationship between the clusters of views. Specifically, for the first one, we design an inner-view (IV) selective contrastive learning module to enhance the clustering structures and alleviate the uncertainty, which selects confident samples near the cluster centroids to perform contrastive learning in each view. For the second one, we design a cross-view (CV) selective contrastive learning module to first iteratively match the clusters between views and then tighten the matched clusters. Also, we utilize mutual information to further enhance the correlation of the matched clusters between views. Extensive experiments show the efficiency of our methods for UMC, compared with the state-of-the-art methods. Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Unpaired Multiview Clustering via Reliable View GuidanceabstractThis article focuses on unpaired multiview clustering (UMC), a challenging problem, where paired observed samples are unavailable across multiple views. The goal is to perform effective joint clustering using the unpaired observed samples in all views. In incomplete multiview clustering (IMC), existing methods typically rely on sample pairing between views to capture their complementary. However, this is not applicable in the case of UMC. Hence, we aim to extract the consistent cluster structure across views. In UMC, two challenging issues arise: the uncertain cluster structure due to the lack of labels and the uncertain pairing relationship due to the absence of paired samples. We assume that the view with a good cluster structure is the reliable view, which acts as a supervisor to guide the clustering of the other views. With the guidance of reliable views, a more certain cluster structure of these views is obtained while achieving alignment between the reliable views and the other views. Then, we propose reliable view guided UMC with one reliable view (RG-UMC) and reliable view guided UMC with multiple reliable views (RGs-UMC). Specifically, we design alignment modules with one reliable view and multiple reliable views, respectively, to adaptively guide the optimization process. Also, we utilize the compactness module to enhance the relationship of samples within the same cluster. Meanwhile, an orthogonal constraint is applied to the latent representation to obtain discriminate features. Extensive experiments show that both RG-UMC and RGs-UMC outperform the best state-of-the-art method by an average of 24.14% and 29.42% in normalized mutual information (NMI), respectively. Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Multilevel Reliable Guidance for Unpaired Multiview ClusteringabstractIn this article, we address the challenging problem of unpaired multiview clustering (UMC), which aims to achieve effective joint clustering using unpaired samples observed across multiple views. Traditional incomplete multiview clustering (IMC) methods typically rely on paired samples to capture complementary information between views. However, such strategies become impractical in the UMC due to the absence of paired samples. Although some researchers have attempted to address this issue by preserving consistent cluster structures across views, effectively mining such consistency remains challenging when the cluster structures with low confidence. Therefore, we propose a novel method, multilevel reliable guidance for UMC (MRG-UMC), which integrates multilevel clustering and reliable view guidance to learn consistent and confident cluster structures from three perspectives. Specifically, inner view multilevel clustering exploits high-confidence sample pairs across different levels to reduce the impact of boundary samples, resulting in more confident cluster structures. Synthesized-view alignment leverages a synthesized view to mitigate cross-view discrepancies and promote consistency. Cross-view guidance employs a reliable view guidance strategy to enhance the clustering confidence of poorly clustered views. These three modules are jointly optimized across multiple levels to achieve consistent and confident cluster structures. Furthermore, theoretical analyses verify the effectiveness of MRG-UMC in enhancing clustering confidence. Extensive experimental results show that MRG-UMC outperforms state-of-the-art UMC methods, achieving an average NMI improvement of 12.95% on multiview datasets. The source code is available at https://anonymous.4open.science/r/MRG-UMC-5E20. Like Xin, Wanqi Yang, Lei Wang 0001, Ming Yang 0014 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Roll with the Punches: Expansion and Shrinkage of Soft Label Selection for Semi-supervised Fine-Grained LearningabstractWhile semi-supervised learning (SSL) has yielded promising results, the more realistic SSL scenario remains to be explored, in which the unlabeled data exhibits extremely high recognition difficulty, e.g., fine-grained visual classification in the context of SSL (SS-FGVC). The increased recognition difficulty on fine-grained unlabeled data spells disaster for pseudo-labeling accuracy, resulting in poor performance of the SSL model. To tackle this challenge, we propose Soft Label Selection with Confidence-Aware Clustering based on Class Transition Tracking (SoC) by reconstructing the pseudo-label selection process by jointly optimizing Expansion Objective and Shrinkage Objective, which is based on a soft label manner. Respectively, the former objective encourages soft labels to absorb more candidate classes to ensure the attendance of ground-truth class, while the latter encourages soft labels to reject more noisy classes, which is theoretically proved to be equivalent to entropy minimization. In comparisons with various state-of-the-art methods, our approach demonstrates its superior performance in SS-FGVC. Checkpoints and source code are available at https://github.com/NJUyued/SoC4SS-FGVC. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
AAAI | 5 |
| 2024 | Exploiting Inter-sample and Inter-feature Relations in Dataset DistillationabstractDataset distillation has emerged as a promising approach in deep learning, enabling efficient training with small synthetic datasets derived from larger real ones. Particularly, distribution matching-based distillation methods attract attention thanks to its effectiveness and low computational cost. However, these methods face two primary limitations: the dispersed feature distribution within the same class in synthetic datasets, reducing class discrim-ination, and an exclusive focus on mean feature consistency, lacking precision and comprehensiveness. To address these challenges, we introduce two novel constraints: a class centralization constraint and a covariance matching constraint. The class centralization constraint aims to enhance class discrimination by more closely clustering samples within classes. The covariance matching constraint seeks to achieve more accurate feature distribution matching between real and synthetic datasets through local feature covariance matrices, particularly beneficial when sample sizes are much smaller than the number of features. Experiments demonstrate notable improvements with these constraints, yielding performance boosts of up to 6.6% on CIFAR10, 2.9% on SVHN, 2.5% on CIFAR100, and 2.5% on TinyImageNet, compared to the state-of-the-art relevant methods. In addition, our method maintains robust performance in cross-architecture settings, with a maximum performance drop of 1.7% on four architectures. Code is avail-able at https://github.com/VincenDen/IID. Wenxiao Deng, Wenbin Li 0006, Tianyu Ding, Lei Wang 0001, Kuihua Huang, Jing Huo, Yang Gao 0001 |
CVPR | 4 |
| 2024 | STDP-based Associative Memory Model on Spiking Neural NetworksabstractIn the cognitive function of the brain, memory plays a crucial role, which involves a complex process of encoding, storing, and retrieving information. The key mechanism in this process is synaptic plasticity, which allows excitatory and inhibitory neural circuits in the brain to adjust the strength of synapses based on experience and learning, facilitating the storage of information. In this study, we designed an associative memory model using spiking neural networks (SNNs) with spike-timing-dependent plasticity (STDP), aiming to simulate the excitatory and inhibitory neural circuits in the brain responsible for encoding, storing, and retrieving memories. In this model, two groups of excitatory neurons and a group of inhibitory neurons in the memory layer receive inputs and activate. Synaptic states between these neurons are modified through STDP during this process, including strengthening, weakening, or forming new synapses related to memory. Subsequently, through the generated synapses, cues guide the activation of one group of excitatory neurons in the memory layer and then trigger the response of another group of excitatory neurons, thereby achieving memory retrieval, namely recall. The results show that our model successfully achieved memory storage and retrieval in auto-associative and hetero-associative tasks. We also discussed the relationship between synaptic utilization of the model and the number of input patterns. Finally, we verified the role of inhibitory neurons in maintaining the stability of the memory network. Wenwu Jiang, Lei Wang 0001, Rui Yan 0005 |
IJCNN | 2 |
| 2024 | KARGEN: Knowledge-Enhanced Automated Radiology Report Generation Using Large Language Models
Yingshu Li 0002, Zhanyu Wang, Yunyi Liu, Lei Wang 0001, Lingqiao Liu, Luping Zhou |
MICCAI (5) | 4 |
| 2024 | MRScore: Evaluating Medical Report with LLM-Based Reward System
Yunyi Liu, Zhanyu Wang, Yingshu Li 0002, Lingqiao Liu, Lei Wang 0001, Luping Zhou |
MICCAI (3) | 6 |
| 2024 | RoCo: Robust Cooperative Perception By Iterative Object Matching and Pose AdjustmentabstractCollaborative autonomous driving with multiple vehicles usually requires the data fusion from multiple modalities. To ensure effective fusion, the data from each individual modality shall maintain a reasonably high quality. However, in collaborative perception, the quality of object detection based on a modality is highly sensitive to the relative pose errors among the agents. It leads to feature misalignment and significantly reduces collaborative performance. To address this issue, we propose RoCo, a novel unsupervised framework to conduct iterative object matching and agent pose adjustment. To the best of our knowledge, our work is the first to model the pose correction problem in collaborative perception as an object matching task, which reliably associates common objects detected by different agents. On top of this, we propose a graph optimization process to adjust the agent poses by minimizing the alignment errors of the associated objects, and the object matching is re-done based on the adjusted agent poses. This process is carried out iteratively until convergence. Experimental study on both simulated and real-world datasets demonstrates that the proposed framework RoCo consistently outperforms existing relevant methods in terms of the collaborative object detection performance, and exhibits highly desired robustness when the pose information of agents is with high-level noise. Ablation studies are also provided to show the impact of its key parameters and components. The code is released at https://github.com/HuangZhe885/RoCo. Shuo Wang 0015, Yongcai Wang, Deying Li 0001, Lei Wang 0001 |
ACM Multimedia | 6 |
| 2024 | Guest Editorial: Special Issue on ACCV 2022
Lei Wang 0001, Juergen Gall, Tat-Jun Chin, Imari Sato, Rama Chellappa |
Int. J. Comput. Vis. | 1 |
| 2024 | Constructing hierarchical attentive functional brain networks for early AD diagnosis
Jianjia Zhang, Yunan Guo, Luping Zhou, Lei Wang 0001, Weiwen Wu, Dinggang Shen |
Medical Image Anal. | 4 |
| 2024 | Question-Aware Global-Local Video Understanding Network for Audio-Visual Question AnsweringabstractAs a newly emerging task, audio-visual question answering (AVQA) has attracted research attention. Compared with traditional single-modality (e.g., audio or visual) QA tasks, it poses new challenges due to the higher complexity of feature extraction and fusion brought by the multimodal inputs. First, AVQA requires more comprehensive understanding of the scene which involves both audio and visual information; Second, in the presence of more information, feature extraction has to be better connected with a given question; Third, features from different modalities need to be sufficiently correlated and fused. To address this situation, this work proposes a novel framework for multimodal question answering task. It characterises an audiovisual scene at both global and local levels, and within each level, the features from different modalities are well fused. Furthermore, the given question is utilised to guide not only the feature extraction at the local level but also the final fusion of global and local features to predict the answer. Our framework provides a new perspective for audio-visual scene understanding through focusing on both general and specific representations as well as aggregating multimodalities by prioritizing question-related information. As experimentally demonstrated, our method significantly improves the existing audio-visual question answering performance, with the averaged absolute gain of 3.3% and 3.1% on MUSIC-AVQA and AVQA datasets, respectively. Moreover, the ablation study verifies the necessity and effectiveness of our design. Our code will be publicly released. Zailong Chen, Lei Wang 0001, Peng Wang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | MutexMatch: Semi-Supervised Learning With Mutex-Based Consistency RegularizationabstractThe core issue in semi-supervised learning (SSL) lies in how to effectively leverage unlabeled data, whereas most existing methods tend to put a great emphasis on the utilization of high-confidence samples yet seldom fully explore the usage of low-confidence samples. In this article, we aim to utilize low-confidence samples in a novel way with our proposed mutex-based consistency regularization, namely MutexMatch. Specifically, the high-confidence samples are required to exactly predict "what it is" by the conventional true-positive classifier (TPC), while low-confidence samples are employed to achieve a simpler goal-to predict with ease "what it is not" by the true-negative classifier (TNC). In this sense, we not only mitigate the pseudo-labeling errors but also make full use of the low-confidence unlabeled data by the consistency of dissimilarity degree. MutexMatch achieves superior performance on multiple benchmark datasets, i.e., Canadian Institute for Advanced Research (CIFAR)-10, CIFAR-100, street view house numbers (SVHN), self-taught learning 10 (STL-10), and mini-ImageNet. More importantly, our method further shows superiority when the amount of labeled data is scarce, e.g., 92.23% accuracy with only 20 labeled data on CIFAR-10. Code has been released at https://github.com/NJUyued/MutexMatch4SSL. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Lei Wang 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Iterative Multiview Subspace Learning for Unpaired Multiview ClusteringabstractIn real applications, several unpredictable or uncertain factors could result in unpaired multiview data, i.e., the observed samples between views cannot be matched. Since joint clustering among views is more effective than individual clustering in each view, we investigate unpaired multiview clustering (UMC), which is a valuable but insufficiently studied problem. Due to lack of matched samples between views, we could fail to build the connection between views. Therefore, we aim to learn the latent subspace shared by views. However, existing multiview subspace learning methods usually rely on the matched samples between views. To address this issue, we propose an iterative multiview subspace learning strategy [iterative unpaired multiview clustering (IUMC)], aiming to learn a complete and consistent subspace representation among views for UMC. Moreover, based on IUMC, we design two effective UMC methods: 1) Iterative unpaired multiview clustering via covariance matrix alignment (IUMC-CA) that further aligns the covariance matrix of subspace representations and then performs clustering on the subspace and 2) iterative unpaired multiview clustering via one-stage clustering assignments (IUMC-CY) that performs one-stage multiview clustering (MVC) by replacing the subspace representations with clustering assignments. Extensive experiments show the excellent performance of our methods for UMC, compared with the state-of-the-art methods. Also, the clustering performance of observed samples in each view can be considerably improved by those observed samples from the other views. In addition, our methods have good applicability in incomplete MVC. Wanqi Yang, Like Xin, Lei Wang 0001, Ming Yang 0014, Wenzhu Yan, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | High-Level Semantic Feature Matters Few-Shot Unsupervised Domain AdaptationabstractIn few-shot unsupervised domain adaptation (FS-UDA), most existing methods followed the few-shot learning (FSL) methods to leverage the low-level local features (learned from conventional convolutional models, e.g., ResNet) for classification. However, the goal of FS-UDA and FSL are relevant yet distinct, since FS-UDA aims to classify the samples in target domain rather than source domain. We found that the local features are insufficient to FS-UDA, which could introduce noise or bias against classification, and not be used to effectively align the domains. To address the above issues, we aim to refine the local features to be more discriminative and relevant to classification. Thus, we propose a novel task-specific semantic feature learning method (TSECS) for FS-UDA. TSECS learns high-level semantic features for image-to-class similarity measurement. Based on the high-level features, we design a cross-domain self-training strategy to leverage the few labeled samples in source domain to build the classifier in target domain. In addition, we minimize the KL divergence of the high-level feature distributions between source and target domains to shorten the distance of the samples between the two domains. Extensive experiments on DomainNet show that the proposed method significantly outperforms SOTA methods in FS-UDA by a large margin (i.e., ~10%). Wanqi Yang, Shengqi Huang, Lei Wang 0001, Ming Yang 0014 |
AAAI | 4 |
| 2023 | Learning Partial Correlation based Deep Visual Representation for Image ClassificationabstractVisual representation based on covariance matrix has demonstrates its efficacy for image classification by characterising the pairwise correlation of different channels in convolutional feature maps. However, pairwise correlation will become misleading once there is another channel correlating with both channels of interest, resulting in the “confounding” effect. For this case, “partial correlation” which removes the confounding effect shall be estimated instead. Nevertheless, reliably estimating partial correlation requires to solve a symmetric positive definite matrix optimisation, known as sparse inverse covariance estimation (SICE). How to incorporate this process into CNN remains an open issue. In this work, we formulate SICE as a novel structured layer of CNN. To ensure end-to-end trainability, we develop an iterative method to solve the above matrix optimisation during forward and backward propagation steps. Our work obtains a partial correlation based deep visual representation and mitigates the small sample problem often encountered by covariance matrix estimation in CNN. Computationally, our model can be effectively trained with GPU and works well with a large number of channels of advanced CNNs. Experiments show the efficacy and superior classification performance of our deep visual representation compared to covariance matrix based counterparts. Saimunur Rahman, Piotr Koniusz, Lei Wang 0001, Luping Zhou, Peyman Moghadam, Changming Sun |
CVPR | 3 |
| 2023 | METransformer: Radiology Report Generation by Transformer with Multiple Learnable Expert TokensabstractIn clinical scenarios, multi-specialist consultation could significantly benefit the diagnosis, especially for intricate cases. This inspires us to explore a “multi-expert joint diagnosis” mechanism to upgrade the existing “single expert” framework commonly seen in the current literature. To this end, we propose METransformer, a method to realize this idea with a transformer-based backbone. The key design of our method is the introduction of multiple learnable “expert” tokens into both the transformer encoder and decoder. In the encoder, each expert token interacts with both vision tokens and other expert tokens to learn to attend different image regions for image representation. These expert tokens are encouraged to capture complementary information by an orthogonal loss that minimizes their over-lap. In the decoder, each attended expert token guides the cross-attention between input words and visual tokens, thus influencing the generated report. A metrics-based expert voting strategy is further developed to generate the final report. By the multi-experts concept, our model enjoys the merits of an ensemble-based approach but through a manner that is computationally more efficient and supports more sophisticated interactions among experts. Experimental results demonstrate the promising performance of our proposed model on two widely used benchmarks. Last but not least, the framework-level innovation makes our work ready to incorporate advances on existing “single-expert” models to further improve its performance. Zhanyu Wang, Lingqiao Liu, Lei Wang 0001, Luping Zhou |
CVPR | 3 |
| 2023 | Automatic Radiology Report Generation by Learning with Increasingly Hard NegativesabstractAutomatic radiology report generation is challenging as medical images or reports are usually similar to each other due to the common content of anatomy. This makes a model hard to capture the uniqueness of individual images and is prone to producing undesired generic or mismatched reports. This situation calls for learning more discriminative features that could capture even fine-grained mismatches between images and reports. To achieve this, this paper proposes a novel framework to learn discriminative image and report features by distinguishing them from their closest peers, i.e., hard negatives. Especially, to attain more discriminative features, we gradually raise the difficulty of such a learning task by creating increasingly hard negative reports for each image in the feature space during training, respectively. By treating the increasingly hard negatives as auxiliary variables, we formulate this process as a min-max alternating optimisation problem. At each iteration, conditioned on a given set of hard negative reports, image and report features are learned as usual by minimising the loss functions related to report generation. After that, a new set of harder negative reports will be created by maximising a loss reflecting image-report alignment. By solving this optimisation, we attain a model that can generate more specific and accurate reports. It is noteworthy that our framework enhances discriminative feature learning without introducing extra network weights. Also, in contrast to the existing way of generating hard negatives, our framework extends beyond the granularity of the dataset by generating harder samples out of the training set. Experimental study on benchmark datasets verifies the efficacy of our framework and shows that it can serve as a plug-in to readily improve existing medical report generation models. The code is publicly available at https://github.com/Bhanu068/ITHN. Bhanu Prakash Voutharoja, Lei Wang 0001, Luping Zhou |
ECAI | 2 |
| 2023 | Towards Semi-supervised Learning with Non-random Missing LabelsabstractSemi-supervised learning (SSL) tackles the label missing problem by enabling the effective usage of unlabeled data. While existing SSL methods focus on the traditional setting, a practical and challenging scenario called label Missing Not At Random (MNAR) is usually ignored. In MNAR, the labeled and unlabeled data fall into different class distributions resulting in biased label imputation, which deteriorates the performance of SSL models. In this work, class transition tracking based Pseudo-Rectifying Guidance (PRG) is devised for MNAR. We explore the class-level guidance information obtained by the Markov random walk, which is modeled on a dynamically created graph built over the class tracking matrix. PRG unifies the historical information of class distribution and class transitions caused by the pseudo-rectifying procedure to maintain the model’s unbiased enthusiasm towards assigning pseudo-labels to all classes, so as the quality of pseudo-labels on both popular classes and rare classes in MNAR could be improved. Finally, we show the superior performance of PRG across a variety of MNAR scenarios, outperforming the latest SSL approaches combining bias removal solutions by a large margin. Code and model weights are available at https://github.com/NJUyued/PRG4SSL-MNAR. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
ICCV | 5 |
| 2023 | Enhancing Sample Utilization through Sample Adaptive Augmentation in Semi-Supervised LearningabstractIn semi-supervised learning, unlabeled samples can be utilized through augmentation and consistency regularization. However, we observed certain samples, even undergoing strong augmentation, are still correctly classified with high confidence, resulting in a loss close to zero. It indicates that these samples have been already learned well and do not provide any additional optimization benefits to the model. We refer to these samples as "naive samples". Unfortunately, existing SSL models overlook the characteristics of naive samples, and they just apply the same learning strategy to all samples. To further optimize the SSL model, we emphasize the importance of giving attention to naive samples and augmenting them in a more diverse manner. Sample adaptive augmentation (SAA) is proposed for this stated purpose and consists of two modules: 1) sample selection module; 2) sample augmentation module. Specifically, the sample selection module picks out naive samples based on historical training information at each epoch, then the naive samples will be augmented in a more diverse manner in the sample augmentation module. Thanks to the extreme ease of implementation of the above modules, SAA is advantageous for being simple and lightweight. We add SAA on top of FixMatch and FlexMatch respectively, and experiments demonstrate SAA can significantly improve the models. For example, SAA helped improve the accuracy of FixMatch from 92.50% to 94.76% and that of FlexMatch from 95.01% to 95.31% on CIFAR-10 with 40 labels. The code is available at https://github.com/GuanGui-nju/SAA. Guan Gui 0002, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
ICCV | 5 |
| 2023 | Learning Spatial-context-aware Global Visual Feature Representation for Instance Image RetrievalabstractIn instance image retrieval, considering local spatial information within an image has proven effective to boost retrieval performance, as demonstrated by local visual descriptor based geometric verification. Nevertheless, it will be highly valuable to make ordinary global image representations spatial-context-aware because global representation based image retrieval is appealing thanks to its algorithmic simplicity, low memory cost, and being friendly to sophisticated data structures. To this end, we propose a novel feature learning framework for instance image retrieval, which embeds local spatial context information into the learned global feature representations. Specifically, in parallel to the visual feature branch in a CNN backbone, we design a spatial context branch that consists of two modules called online token learning and distance encoding. For each local descriptor learned in CNN, the former module is used to indicate the types of its surrounding descriptors, while their spatial distribution information is captured by the latter module. After that, the visual feature branch and the spatial context branch are fused to produce a single global feature representation per image. As experimentally demonstrated, with the spatial-context-aware characteristic, we can well improve the performance of global representation based image retrieval while maintaining all of its appealing properties. Our code is available at https://github.com/Zy-Zhang/SpCa. Zhongyan Zhang, Lei Wang 0001, Luping Zhou, Piotr Koniusz |
ICCV | 2 |
| 2023 | LibFewShot: A Comprehensive Library for Few-Shot LearningabstractFew-shot learning, especially few-shot image classification, has received increasing attention and witnessed significant advances in recent years. Some recent studies implicitly show that many generic techniques or "tricks", such as data augmentation, pre-training, knowledge distillation, and self-supervision, may greatly boost the performance of a few-shot learning method. Moreover, different works may employ different software platforms, backbone architectures and input image sizes, making fair comparisons difficult and practitioners struggle with reproducibility. To address these situations, we propose a comprehensive library for few-shot learning (LibFewShot) by re-implementing eighteen state-of-the-art few-shot learning methods in a unified framework with the same single codebase in PyTorch. Furthermore, based on LibFewShot, we provide comprehensive evaluations on multiple benchmarks with various backbone architectures to evaluate common pitfalls and effects of different training tricks. In addition, with respect to the recent doubts on the necessity of meta- or episodic-training mechanism, our evaluation results confirm that such a mechanism is still necessary especially when combined with pre-training. We hope our work can not only lower the barriers for beginners to enter the area of few-shot learning but also elucidate the effects of nontrivial tricks to facilitate intrinsic research on few-shot learning. Wenbin Li 0006, Xuesong Yang, Chuanqi Dong, Pinzhuo Tian, Tiexin Qin, Jing Huo, Yinghuan Shi, Lei Wang 0001, Yang Gao 0001, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2023 | Defensive Few-Shot LearningabstractThis article investigates a new challenging problem called defensive few-shot learning in order to learn a robust few-shot model against adversarial attacks. Simply applying the existing adversarial defense methods to few-shot learning cannot effectively solve this problem. This is because the commonly assumed sample-level distribution consistency between the training and test sets can no longer be met in the few-shot setting. To address this situation, we develop a general defensive few-shot learning (DFSL) framework to answer the following two key questions: (1) how to transfer adversarial defense knowledge from one sample distribution to another? (2) how to narrow the distribution gap between clean and adversarial examples under the few-shot setting? To answer the first question, we propose an episode-based adversarial training mechanism by assuming a task-level distribution consistency to better transfer the adversarial defense knowledge. As for the second question, within each few-shot task, we design two kinds of distribution consistency criteria to narrow the distribution gap between clean and adversarial examples from the feature-wise and prediction-wise perspectives, respectively. Extensive experiments demonstrate that the proposed framework can effectively make the existing few-shot models robust against adversarial attacks. Code is available at https://github.com/WenbinLee/DefensiveFSL.git. Wenbin Li 0006, Lei Wang 0001, Xingxing Zhang 0001, Lei Qi 0001, Jing Huo, Yang Gao 0001, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Dataset-Driven Unsupervised Object Discovery for Region-Based Instance Image RetrievalabstractInstance image retrieval could greatly benefit from discovering objects in the image dataset. This not only helps produce more reliable feature representation but also better informs users by delineating query-matched object regions. However, object classes are usually not predefined in a retrieval dataset and class label information is generally unavailable in image retrieval. This situation makes object discovery a challenging task. To address this, we propose a novel dataset-driven unsupervised object discovery framework. By utilizing deep feature representation and weakly-supervised object detection, we explore supervisory information from within an image dataset, construct class-wise object detectors, and assign multiple detectors to each image for detection. To efficiently construct object detectors for large image datasets, we propose a novel "base-detector repository" and derive a fast way to generate the base detectors. In addition, the whole framework is designed to work in a self-boosting manner to iteratively refine object discovery. Compared with existing unsupervised object detection methods, our framework produces more accurate object discovery results. Different from supervised detection, we need neither manual annotation nor auxiliary datasets to train object detectors. Experimental study demonstrates the effectiveness of the proposed framework and the improved performance for region-based instance image retrieval. Zhongyan Zhang, Lei Wang 0001, Yang Wang 0002, Luping Zhou, Jianjia Zhang, Fang Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Unsupervised generalizable multi-source person re-identification: A Domain-specific adaptive framework
Lei Qi 0001, Lei Wang 0001, Yinghuan Shi, Xin Geng 0001 |
Pattern Recognit. | 3 |
| 2023 | Global- and local-aware feature augmentation with semantic orthogonality for few-shot image classification
Boyao Shi, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Yang Gao 0001 |
Pattern Recognit. | 5 |
| 2023 | Kernel-based feature aggregation framework in point cloud networks
Jianjia Zhang, Lei Wang 0001, Luping Zhou, Xiaocai Zhang, Weiwen Wu |
Pattern Recognit. | 3 |
| 2023 | A Novel Mix-Normalization Method for Generalizable Multi-Source Person Re-IdentificationabstractPerson re-identification (Re-ID) has achieved great success in the supervised scenario. However, it is difficult to directly transfer the supervised model to arbitrary unseen domains due to the model overfitting to the seen source domains. In this paper, we aim to tackle the generalizable multi-source person Re-ID task (i.e., there are multiple available source domains, and the testing domain is unseen during training) from the data augmentation perspective, thus we put forward a novel method, termed MixNorm. It consists of domain-aware mix-normalization (DMN) and domain-aware center regularization (DCR). Different from the conventional data augmentation, the proposed domain-aware mix-normalization enhances the diversity of features during training from the normalization perspective of the neural network, which can effectively alleviate the model overfitting to the source domains, so as to boost the generalization capability of the model in the unseen domain. To further promote the efficacy of the proposed DMN, we exploit the domain-aware center regularization to better map the diversely generated features into the same space. Extensive experiments on multiple benchmark datasets validate the effectiveness of the proposed method and show that the proposed method can outperform the state-of-the-art methods. Besides, further analysis also reveals the superiority of the proposed method. Lei Qi 0001, Lei Wang 0001, Yinghuan Shi, Xin Geng 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | A Multilayer Framework for Online Metric LearningabstractOnline metric learning (OML) has been widely applied in classification and retrieval. It can automatically learn a suitable metric from data by restricting similar instances to be separated from dissimilar instances with a given margin. However, the existing OML algorithms have limited performance in real-world classifications, especially, when data distributions are complex. To this end, this article proposes a multilayer framework for OML to capture the nonlinear similarities among instances. Different from the traditional OML, which can only learn one metric space, the proposed multilayer OML (MLOML) takes an OML algorithm as a metric layer and learns multiple hierarchical metric spaces, where each metric layer follows a nonlinear layer for the complicated data distribution. Moreover, the forward propagation (FP) strategy and backward propagation (BP) strategy are employed to train the hierarchical metric layers. To build a metric layer of the proposed MLOML, a new Mahalanobis-based OML (MOML) algorithm is presented based on the passive-aggressive strategy and one-pass triplet construction strategy. Furthermore, in a progressively and nonlinearly learning way, MLOML has a stronger learning ability than traditional OML in the case of limited available training data. To make the learning process more explainable and theoretically guaranteed, theoretical analysis is provided. The proposed MLOML enjoys several nice properties, indeed learns a metric progressively, and performs better on the benchmark datasets. Extensive experiments with different settings have been conducted to verify these properties of the proposed MLOML. Wenbin Li 0006, Jing Huo, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Jiebo Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | LaSSL: Label-Guided Self-Training for Semi-supervised LearningabstractThe key to semi-supervised learning (SSL) is to explore adequate information to leverage the unlabeled data. Current dominant approaches aim to generate pseudo-labels on weakly augmented instances and train models on their corresponding strongly augmented variants with high-confidence results. However, such methods are limited in excluding samples with low-confidence pseudo-labels and under-utilization of the label information. In this paper, we emphasize the cruciality of the label information and propose a Label-guided Self-training approach to Semi-supervised Learning (LaSSL), which improves pseudo-label generations from two mutually boosted strategies. First, with the ground-truth labels and iteratively-polished pseudo-labels, we explore instance relations among all samples and then minimize a class-aware contrastive loss to learn discriminative feature representations that make same-class samples gathered and different-class samples scattered. Second, on top of improved feature representations, we propagate the label information to the unlabeled samples across the potential data manifold at the feature-embedding level, which can further improve the labelling of samples with reference to their neighbours. These two strategies are seamlessly integrated and mutually promoted across the whole training process. We evaluate LaSSL on several classification benchmarks under partially labeled settings and demonstrate its superiority over the state-of-the-art approaches. Zhen Zhao 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001 |
AAAI | 3 |
| 2022 | Kernelized Few-shot Object Detection with Efficient Integral AggregationabstractWe design a Kernelized Few-shot Object Detector by leveraging kernelized matrices computed over multiple proposal regions, which yield expressive non-linear representations whose model complexity is learned on the fly. Our pipeline contains several modules. An Encoding Network encodes support and query images. Our Kernelized Autocorrelation unit forms the linear, polynomial and RBF kernelized representations from features extracted within support regions of support images. These features are then cross-correlated against features of a query image to obtain attention weights, and generate query proposal regions via an Attention Region Proposal Net. As the query proposal regions are many, each described by the linear, polynomial and RBF kernelized matrices, their formation is costly but that cost is reduced by our proposed Integral Region-of-Interest Aggregation unit. Finally, the Multi-head Relation Net combines all kernelized (second-order) representations with the first-order feature maps to learn support-query class relations and locations. We outperform the state of the art on novel classes by 3.8%, 5.4% and 5.7% mAP on PASCAL VOC 2007, FSOD, and COCO. Shan Zhang 0002, Lei Wang 0001, Naila Murray, Piotr Koniusz |
CVPR | 2 |
| 2022 | DC-SSL: Addressing Mismatched Class Distribution in Semi-supervised LearningabstractConsistency-based Semi-supervised learning (SSL) has achieved promising performance recently. However, the success largely depends on the assumption that the labeled and unlabeled data share an identical class distribution, which is hard to meet in real practice. The distribution mismatch between the labeled and unlabeled sets can cause severe bias in the pseudo-labels of SSL, resulting in significant performance degradation. To bridge this gap, we put forward a new SSL learning framework, named Distribution Consistency SSL (DC-SSL), which rectifies the pseudolabels from a distribution perspective. The basic idea is to directly estimate a reference class distribution (RCD), which is regarded as a surrogate of the ground truth class distribution about the unlabeled data, and then improve the pseudo-labels by encouraging the predicted class distribution (PCD) of the unlabeled data to approach RCD gradually. To this end, this paper revisits the Exponentially Moving Average (EMA) model and utilizes it to estimate RCD in an iteratively improved manner, which is achieved with a momentum-update scheme throughout the training procedure. On top of this, two strategies are proposed for RCD to rectify the pseudo-label prediction, respectively. They correspond to an efficient training-free scheme and a training-based alternative that generates more accurate and reliable predictions. DC-SSL is evaluated on multiple SSL benchmarks and demonstrates remarkable performance improvement over competitive methods under matched- and mismatched-distribution scenarios. Zhen Zhao 0001, Luping Zhou, Yue Duan, Lei Wang 0001, Lei Qi 0001, Yinghuan Shi |
CVPR | 4 |
| 2022 | RDA: Reciprocal Distribution Alignment for Robust Semi-supervised Learning
Yue Duan, Lei Qi 0001, Lei Wang 0001, Luping Zhou, Yinghuan Shi |
ECCV (30) | 3 |
| 2022 | Time-rEversed DiffusioN tEnsor Transformer: A New TENET of Few-Shot Object Detection
Shan Zhang 0002, Naila Murray, Lei Wang 0001, Piotr Koniusz |
ECCV (20) | 3 |
| 2022 | Few-Shot Unsupervised Domain Adaptation via Meta LearningabstractUnsupervised domain adaptation (UDA) has raised a lot of interests in recent years. However, current UDA methods are still not capable enough in dealing with two issues: 1) the scarcity of labeled data in source domain and 2) the need of a general model that can quickly adapt to solve new UDA tasks. To address this situation, we investigate available but rarely-studied setting called few-shot unsupervised domain adaptation (FS-UDA), in which the data of source domain is few-shot per category and the data of target domain remains unlabeled. To realize effective adaptation for FS-UDA tasks in the same source and target domains, we propose a novel meta learning method namely meta-FUDA, which leverages meta learning to perform task-level transfer and domain-level transfer jointly. Extensive experiments demonstrate the promising performance of our method on multiple benchmark data sets. Wanqi Yang, Chengmei Yang, Shengqi Huang, Lei Wang 0001, Ming Yang 0014 |
ICME | 4 |
| 2022 | A Medical Semantic-Assisted Transformer for Radiographic Report Generation
Zhanyu Wang, Mingkang Tang, Lei Wang 0001, Xiu Li 0001, Luping Zhou |
MICCAI (3) | 3 |
| 2022 | Domain Generalization by Learning and Removing Domain-specific FeaturesabstractDeep Neural Networks (DNNs) suffer from domain shift when the test dataset follows a distribution different from the training dataset. Domain generalization aims to tackle this issue by learning a model that can generalize to unseen domains. In this paper, we propose a new approach that aims to explicitly remove domain-specific features for domain generalization. Following this approach, we propose a novel framework called Learning and Removing Domain-specific features for Generalization (LRDG) that learns a domain-invariant model by tactically removing domain-specific features from the input images. Specifically, we design a classifier to effectively learn the domain-specific features for each source domain, respectively. We then develop an encoder-decoder network to map each input image into a new image space where the learned domain-specific features are removed. With the images output by the encoder-decoder network, another classifier is designed to learn the domain-invariant features to conduct image classification. Extensive experiments demonstrate that our framework achieves superior performance compared with state-of-the-art methods. Lei Wang 0001, Bin Liang 0003, Shuming Liang, Yang Wang 0002, Fang Chen 0001 |
NeurIPS | 2 |
| 2022 | Improving Barely Supervised Learning by Discriminating Unlabeled Samples with Super-ClassabstractIn semi-supervised learning (SSL), a common practice is to learn consistent information from unlabeled data and discriminative information from labeled data to ensure both the immutability and the separability of the classification model. Existing SSL methods suffer from failures in barely-supervised learning (BSL), where only one or two labels per class are available, as the insufficient labels cause the discriminative information being difficult or even infeasible to learn. To bridge this gap, we investigate a simple yet effective way to leverage unlabeled samples for discriminative learning, and propose a novel discriminative information learning module to benefit model training. Specifically, we formulate the learning objective of discriminative information at the super-class level and dynamically assign different classes into different super-classes based on model performance improvement. On top of this on-the-fly process, we further propose a distribution-based loss to learn discriminative information by utilizing the similarity relationship between samples and super-classes. It encourages the unlabeled samples to stay closer to the distribution of their corresponding super-class than those of others. Such a constraint is softer than the direct assignment of pseudo labels, while the latter could be very noisy in BSL. We compare our method with state-of-the-art SSL and BSL methods through extensive experiments on standard SSL benchmarks. Our method can achieve superior results, \eg, an average accuracy of 76.76\% on CIFAR-10 with merely 1 label per class. Guan Gui 0002, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
NeurIPS | 5 |
| 2022 | Dual-scale correlation analysis for robust multi-label classification
Kaixiang Wang 0001, Ming Yang 0014, Wanqi Yang, Lei Wang 0001 |
Appl. Intell. | 4 |
| 2022 | ASMFS: Adaptive-similarity-based multi-modality feature selection for classification of Alzheimer's disease
Yuang Shi, Chen Zu, Luping Zhou, Lei Wang 0001, Xi Wu 0004, Jiliu Zhou, Daoqiang Zhang, Yan Wang 0015 |
Pattern Recognit. | 5 |
| 2022 | From Regional to Global Brain: A Novel Hierarchical Spatial-Temporal Neural Network Model for EEG Emotion RecognitionabstractIn this paper, we propose a novel Electroencephalograph (EEG) emotion recognition method inspired by neuroscience with respect to the brain response to different emotions. The proposed method, denoted by R2G-STNN, consists of spatial and temporal neural network models with regional to global hierarchical feature learning process to learn discriminative spatial-temporal EEG features. To learn the spatial features, a bidirectional long short term memory (BiLSTM) network is adopted to capture the intrinsic spatial relationships of EEG electrodes within brain region and between brain regions, respectively. Considering that different brain regions play different roles in the EEG emotion recognition, a region-attention layer into the R2G-STNN model is also introduced to learn a set of weights to strengthen or weaken the contributions of brain regions. Based on the spatial feature sequences, BiLSTM is adopted to learn both regional and global spatial-temporal features and the features are fitted into a classifier layer for learning emotion-discriminative features, in which a domain discriminator working corporately with the classifier is used to decrease the domain shift between training and testing data. Finally, to evaluate the proposed method, we conduct both subject-dependent and subject-independent EEG emotion recognition experiments on SEED database, and the experimental results show that the proposed method achieves state-of-the-art performance. Yang Li 0019, Wenming Zheng, Lei Wang 0001, Yuan Zong, Zhen Cui 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | Adversarial Camera Alignment Network for Unsupervised Cross-Camera Person Re-IdentificationabstractIn person re-identification (Re-ID), supervised methods usually need a large amount of expensive label information, while unsupervised ones are still unable to deliver satisfactory identification performance. In this paper, we introduce a novel person Re-ID task called unsupervised cross-camera person Re-ID, which only needs the within-camera (intra-camera) label information but not cross-camera (inter-camera) labels which are more expensive to obtain. In real-world applications, the intra-camera label information can be easily captured by tracking algorithms and few manual annotations. In this situation, the main challenge becomes the distribution discrepancy across different camera views, caused by the various body pose, occlusion, image resolution, illumination conditions, and background noises in different cameras. To address this situation, we propose a novel Adversarial Camera Alignment Network (ACAN) for unsupervised cross-camera person Re-ID. It consists of the camera-alignment task and the supervised within-camera learning task. To achieve the camera alignment, we develop a Multi-Camera Adversarial Learning (MCAL) to map images of different cameras into a shared subspace. Particularly, we investigate two different schemes, including the existing GRL (i.e., gradient reversal layer) scheme and the proposed scheme called “other camera equiprobability” (OCE), to conduct the multi-camera adversarial task. Based on this shared subspace, we then leverage the within-camera labels to train the network. Extensive experiments on five large-scale datasets demonstrate the superiority of ACAN over multiple state-of-the-art unsupervised methods that take advantage of labeled source domains and generated images by GAN-based models. In particular, we verify that the proposed multi-camera adversarial task does contribute to the significant improvement. Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Xin Geng 0001, Yang Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Automated Radiographic Report Generation Purely on Transformer: A Multicriteria Supervised ApproachabstractAutomated radiographic report generation is challenging in at least two aspects. First, medical images are very similar to each other and the visual differences of clinic importance are often fine-grained. Second, the disease-related words may be submerged by many similar sentences describing the common content of the images, causing the abnormal to be misinterpreted as the normal in the worst case. To tackle these challenges, this paper proposes a pure transformer-based framework to jointly enforce better visual-textual alignment, multi-label diagnostic classification, and word importance weighting, to facilitate report generation. To the best of our knowledge, this is the first pure transformer-based framework for medical report generation, which enjoys the capacity of transformer in learning long range dependencies for both image regions and sentence words. Specifically, for the first challenge, we design a novel mechanism to embed an auxiliary image-text matching objective into the transformer's encoder-decoder structure, so that better correlated image and text features could be learned to help a report to discriminate similar images. For the second challenge, we integrate an additional multi-label classification task into our framework to guide the model in making correct diagnostic predictions. Also, a term-weighting scheme is proposed to reflect the importance of words for training so that our model would not miss key discriminative information. Our work achieves promising performance over the state-of-the-arts on two benchmark datasets, including the largest dataset MIMIC-CXR. Zhanyu Wang, Lei Wang 0001, Xiu Li 0001, Luping Zhou |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Diffusion Kernel Attention Network for Brain Disorder ClassificationabstractConstructing and analyzing functional brain networks (FBN) has become a promising approach to brain disorder classification. However, the conventional successive construct-and-analyze process would limit the performance due to the lack of interactions and adaptivity among the subtasks in the process. Recently, Transformer has demonstrated remarkable performance in various tasks, attributing to its effective attention mechanism in modeling complex feature relationships. In this paper, for the first time, we develop Transformer for integrated FBN modeling, analysis and brain disorder classification with rs-fMRI data by proposing a Diffusion Kernel Attention Network to address the specific challenges. Specifically, directly applying Transformer does not necessarily admit optimal performance in this task due to its extensive parameters in the attention module against the limited training samples usually available. Looking into this issue, we propose to use kernel attention to replace the original dot-product attention module in Transformer. This significantly reduces the number of parameters to train and thus alleviates the issue of small sample while introducing a non-linear attention mechanism to model complex functional connections. Another limit of Transformer for FBN applications is that it only considers pair-wise interactions between directly connected brain regions but ignores the important indirect connections. Therefore, we further explore diffusion process over the kernel attention to incorporate wider interactions among indirectly connected brain regions. Extensive experimental study is conducted on ADHD-200 data set for ADHD classification and on ADNI data set for Alzheimer's disease classification, and the results demonstrate the superior performance of the proposed method over the competing methods. Jianjia Zhang, Luping Zhou, Lei Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Contrastive Learning Based Hybrid Networks for Long-Tailed Image ClassificationabstractLearning discriminative image representations plays a vital role in long-tailed image classification because it can ease the classifier learning in imbalanced cases. Given the promising performance contrastive learning has shown recently in representation learning, in this work, we explore effective supervised contrastive learning strategies and tailor them to learn better image representations from imbalanced data in order to boost the classification accuracy thereon. Specifically, we propose a novel hybrid network structure being composed of a supervised contrastive loss to learn image representations and a cross-entropy loss to learn classifiers, where the learning is progressively transited from feature learning to the classifier learning to embody the idea that better features make better classifiers. We explore two variants of contrastive loss for feature learning, which vary in the forms but share a common idea of pulling the samples from the same class together in the normalized embedding space and pushing the samples from different classes apart. One of them is the recently proposed supervised contrastive (SC) loss, which is designed on top of the state-of-the-art unsupervised contrastive loss by incorporating positive samples from the same class. The other is a prototypical supervised contrastive (PSC) learning strategy which addresses the intensive memory consumption in standard SC loss and thus shows more promise under limited memory budget. Extensive experiments on three long-tailed classification datasets demonstrate the advantage of the proposed contrastive learning based hybrid networks in long-tailed classification. Peng Wang 0023, Kai Han 0001, Xiu-Shen Wei, Lei Zhang 0054, Lei Wang 0001 |
CVPR | 5 |
| 2021 | A Self-Boosting Framework for Automated Radiographic Report GenerationabstractAutomated radiographic report generation is a challenging task since it requires to generate paragraphs describing fine-grained visual differences of cases, especially for those between the diseased and the healthy. Existing image captioning methods commonly target at generic images, and lack mechanism to meet this requirement. To bridge this gap, in this paper, we propose a self-boosting framework that improves radiographic report generation based on the cooperation of the main task of report generation and an auxiliary task of image-text matching. The two tasks are built as the two branches of a network model and influence each other in a cooperative way. On one hand, the image-text matching branch helps to learn highly text-correlated visual features for the report generation branch to output high quality reports. On the other hand, the improved reports produced by the report generation branch provide additional harder samples for the image-text matching branch and enforce the latter to improve itself by learning better visual and text feature representations. This, in turn, helps improve the report generation branch again. These two branches are jointly trained to help improve each other iteratively and progressively, so that the whole model is self-boosted without requiring external resources. Experimental results demonstrate the effectiveness of our method on two public datasets, showing its superior performance over multiple state-of-the-art image captioning and medical report generation methods. Zhanyu Wang, Luping Zhou, Lei Wang 0001, Xiu Li 0001 |
CVPR | 3 |
| 2021 | LoFGAN: Fusing Local Representations for Few-shot Image GenerationabstractGiven only a few available images for a novel unseen category, few-shot image generation aims to generate more data for this category. Previous works attempt to globally fuse these images by using adjustable weighted coefficients. However, there is a serious semantic misalignment between different images from a global perspective, making these works suffer from poor generation quality and diversity. To tackle this problem, we propose a novel Local-Fusion Generative Adversarial Network (LoFGAN) for fewshot image generation. Instead of using these available images as a whole, we first randomly divide them into a base image and several reference images. Next, LoFGAN matches local representations between the base and reference images based on semantic similarities, and replaces the local features with the closest related local features. In this way, LoFGAN can produce more realistic and diverse images at a more fine-grained level, and simultaneously enjoy the characteristic of semantic alignment. Furthermore, a local reconstruction loss is also proposed, which can provide better training stability and generation quality. We conduct extensive experiments on three datasets, which successfully demonstrates the effectiveness of our proposed method for few-shot image generation and downstream visual applications with limited data. Code is available at https://github.com/edward3862/LoFGAN-pytorch. Zheng Gu 0001, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Yang Gao 0001 |
ICCV | 4 |
| 2021 | BV-Person: A Large-scale Dataset for Bird-view Person Re-identificationabstractPerson Re-IDentification (ReID) aims at re-identifying persons from non-overlapping cameras. Existing person ReID studies focus on horizontal-view ReID tasks, in which the person images are captured by the cameras from a (nearly) horizontal view. In this work we introduce a new ReID task, bird-view person ReID, which aims at searching for a person in a gallery of horizontal-view images with the query images taken from a bird's-eye view, i.e., an elevated view of an object from above. The task is important because there are a large number of video surveillance cameras capturing persons from such an elevated view at public places. However, it is a challenging task in that the images from the bird view (i) provide limited person appearance information and (ii) have a large discrepancy compared to the persons in the horizontal view. We aim to facilitate the development of person ReID from this line by introducing a large-scale real-world dataset for this task. The proposed dataset, named BV-Person, contains 114k images of 18k identities in which nearly 20k images of 7.4k identities are taken from the bird's-eye view. We further introduce a novel model for this new ReID task. Large-scale experiments are performed to evaluate our model and 11 current state-of-the-art ReID models on BV-Person to establish performance benchmarks from multiple perspectives. The empirical results show that our model consistently and substantially outperforms the state-of-the-art models on all five datasets derived from BV-Person. Our model also achieves state-of-the-art performance on two general ReID datasets. The BV-Person dataset is available at: https://git.io/BVPerson Guansong Pang, Lei Wang 0001, Jile Jiao, Xuetao Feng, Chunhua Shen |
ICCV | 3 |
| 2021 | Inter-Modality Fusion Based Attention for Zero-Shot Cross-Modal RetrievalabstractZero-shot cross-modal retrieval (ZS-CMR) performs the task of cross-modal retrieval where the classes of test categories have a different scope than the training categories. It borrows the intuition from zero-shot learning which targets to transfer the knowledge inferred during the training phase for seen classes to the testing phase for unseen classes. It mimics the real-world scenario where new object categories are continuously populating the multi-media data corpus. Unlike existing ZS-CMR approaches which use generative adversarial networks (GANs) to generate more data, we propose Inter-Modality Fusion based Attention (IMFA) and a framework ZS_INN_FUSE (Zero-Shot cross-modal retrieval using INNer product with image-text FUSEd). It exploits the rich semantics of textual data as guidance to infer additional knowledge during the training phase. This is achieved by generating attention weights through the fusion of image and text modalities to focus on the important regions in an image. We carefully create a zero-shot split based on the large-scale MS-COCO and Flickr30k datasets to perform experiments. The results show that our method achieves improvement over the ZS-CMR baseline and self-attention mechanism, demonstrating the effectiveness of inter-modality fusion in a zero-shot scenario. Bela Chakraborty, Peng Wang 0023, Lei Wang 0001 |
ICIP | 3 |
| 2021 | Few-shot Unsupervised Domain Adaptation with Image-to-Class Sparse Similarity EncodingabstractThis paper investigates a valuable setting called few-shot unsupervised domain adaptation (FS-UDA), which has not been sufficiently studied in the literature. In this setting, the source domain data are labelled, but with few-shot per category, while the target domain data are unlabelled. To address the FS-UDA setting, we develop a general UDA model to solve the following two key issues: the few-shot labeled data per category and the domain adaptation between support and query sets. Our model is general in that once trained it will be able to be applied to various FS-UDA tasks from the same source and target domains. Inspired by the recent local descriptor based few-shot learning (FSL), our general UDA model is fully built upon local descriptors (LDs) for image classification and domain adaptation. By proposing a novel concept called similarity patterns (SPs), our model not only effectively considers the spatial relationship of LDs that was ignored in previous FSL methods, but also makes the learned image similarity better serve the required domain alignment. Specifically, we propose a novel IMage-to-class sparse Similarity Encoding (IMSE) method. It learns SPs to extract the local discriminative information for classification and meanwhile aligns the covariance matrix of the SPs for domain adaptation. Also, domain adversarial training and multi-scale local feature matching are performed upon LDs. Extensive experiments conducted on a multi-domain benchmark dataset DomainNet demonstrates the state-of-the-art performance of our IMSE for the novel setting of FS-UDA. In addition, for FSL, our IMSE can also show better performance than most of recent FSL methods on miniImageNet. Shengqi Huang, Wanqi Yang, Lei Wang 0001, Luping Zhou, Ming Yang 0014 |
ACM Multimedia | 3 |
| 2021 | Interactive medical image segmentation via a point-based interaction
Jian Zhang 0090, Yinghuan Shi, Jinquan Sun, Lei Wang 0001, Luping Zhou, Yang Gao 0001, Dinggang Shen |
Artif. Intell. Medicine | 4 |
| 2021 | Systematic evaluation of machine learning methods for identifying human-pathogen protein-protein interactionsabstractIn recent years, high-throughput experimental techniques have significantly enhanced the accuracy and coverage of protein-protein interaction identification, including human-pathogen protein-protein interactions (HP-PPIs). Despite this progress, experimental methods are, in general, expensive in terms of both time and labour costs, especially considering that there are enormous amounts of potential protein-interacting partners. Developing computational methods to predict interactions between human and bacteria pathogen has thus become critical and meaningful, in both facilitating the detection of interactions and mining incomplete interaction maps. In this paper, we present a systematic evaluation of machine learning-based computational methods for human-bacterium protein-protein interactions (HB-PPIs). We first reviewed a vast number of publicly available databases of HP-PPIs and then critically evaluate the availability of these databases. Benefitting from its well-structured nature, we subsequently preprocess the data and identified six bacterium pathogens that could be used to study bacterium subjects in which a human was the host. Additionally, we thoroughly reviewed the literature on 'host-pathogen interactions' whereby existing models were summarized that we used to jointly study the impact of different feature representation algorithms and evaluate the performance of existing machine learning computational models. Owing to the abundance of sequence information and the limited scale of other protein-related information, we adopted the primary protocol from the literature and dedicated our analysis to a comprehensive assessment of sequence information and machine learning models. A systematic evaluation of machine learning models and a wide range of feature representation algorithms based on sequence information are presented as a comparison survey towards the prediction performance evaluation of HB-PPIs. Huaming Chen, Fuyi Li, Lei Wang 0001, Yaochu Jin, Chihung Chi, Lukasz A. Kurgan, Jiangning Song, Jun Shen 0001 |
Briefings Bioinform. | 3 |
| 2021 | Beyond Covariance: SICE and Kernel Based Visual Feature Representation
Jianjia Zhang, Lei Wang 0001, Luping Zhou, Wanqing Li 0001 |
Int. J. Comput. Vis. | 2 |
| 2021 | Visual place recognition: A survey from deep learning perspective
Xiwu Zhang, Lei Wang 0001, Yan Su 0002 |
Pattern Recognit. | 2 |
| 2021 | SA-LuT-Nets: Learning Sample-Adaptive Intensity Lookup Tables for Brain Tumor SegmentationabstractIn clinics, the information about the appearance and location of brain tumors is essential to assist doctors in diagnosis and treatment. Automatic brain tumor segmentation on the images acquired by magnetic resonance imaging (MRI) is a common way to attain this information. However, MR images are not quantitative and can exhibit significant variation in signal depending on a range of factors, which increases the difficulty of training an automatic segmentation network and applying it to new MR images. To deal with this issue, this paper proposes to learn a sample-adaptive intensity lookup table (LuT) that dynamically transforms the intensity contrast of each input MR image to adapt to the following segmentation task. Specifically, the proposed deep SA-LuT-Net framework consists of a LuT module and a segmentation module, trained in an end-to-end manner: the LuT module learns a sample-specific nonlinear intensity mapping function through communication with the segmentation module, aiming at improving the final segmentation performance. In order to make the LuT learning sample-adaptive, we parameterize the intensity mapping function by exploring two families of non-linear functions (i.e., piece-wise linear and power functions) and predict the function parameters for each given sample. These sample-specific parameters make the intensity mapping adaptive to samples. We develop our SA-LuT-Nets separately based on two backbone networks for segmentation, i.e., DMFNet and the modified 3D Unet, and validate them on BRATS2018 and BRATS2019 datasets for brain tumor segmentation. Our experimental results clearly demonstrate the superior performance of the proposed SA-LuT-Nets using either single or multiple MR modalities. It not only significantly improves the two baselines (DMFNet and the modified 3D Unet), but also wins a set of state-of-the-art segmentation methods. Moreover, we show that, the LuTs learnt using one segmentation model could also be applied to improving the performance of another segmentation model, indicating the general segmentation information captured by LuTs. Biting Yu, Luping Zhou, Lei Wang 0001, Wanqi Yang, Ming Yang 0014, Pierrick Bourgeat, Jurgen Fripp |
IEEE Trans. Medical Imaging | 3 |
| 2021 | GreyReID: A Novel Two-stream Deep Framework with RGB-grey Information for Person Re-identificationabstractIn this article, we observe that most false positive images (i.e., different identities with query images) in the top ranking list usually have the similar color information with the query image in person re-identification (Re-ID). Meanwhile, when we use the greyscale images generated from RGB images to conduct the person Re-ID task, some hard query images can obtain better performance compared with using RGB images. Therefore, RGB and greyscale images seem to be complementary to each other for person Re-ID. In this article, we aim to utilize both RGB and greyscale images to improve the person Re-ID performance. To this end, we propose a novel two-stream deep neural network with RGB-grey information, which can effectively fuse RGB and greyscale feature representations to enhance the generalization ability of Re-ID. First, we convert RGB images to greyscale images in each training batch. Based on these RGB and greyscale images, we train the RGB and greyscale branches, respectively. Second, to build up connections between RGB and greyscale branches, we merge the RGB and greyscale branches into a new joint branch. Finally, we concatenate the features of all three branches as the final feature representation for Re-ID. Moreover, in the training process, we adopt the joint learning scheme to simultaneously train each branch by the independent loss function, which can enhance the generalization ability of each branch. Besides, a global loss function is utilized to further fine-tune the final concatenated feature. The extensive experiments on multiple benchmark datasets fully show that the proposed method can outperform the state-of-the-art person Re-ID methods. Furthermore, using greyscale images can indeed improve the person Re-ID performance in the proposed deep framework. Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Differentiable Meta-Learning Model for Few-Shot Semantic SegmentationabstractTo address the annotation scarcity issue in some cases of semantic segmentation, there have been a few attempts to develop the segmentation model in the few-shot learning paradigm. However, most existing methods only focus on the traditional 1-way segmentation setting (i.e., one image only contains a single object). This is far away from practical semantic segmentation tasks where the K-way setting (K > 1) is usually required by performing the accurate multi-object segmentation. To deal with this issue, we formulate the few-shot semantic segmentation task as a learning-based pixel classification problem, and propose a novel framework called MetaSegNet based on meta-learning. In MetaSegNet, an architecture of embedding module consisting of the global and local feature branches is developed to extract the appropriate meta-knowledge for the few-shot segmentation. Moreover, we incorporate a linear model into MetaSegNet as a base learner to directly predict the label of each pixel for the multi-object segmentation. Furthermore, our MetaSegNet can be trained by the episodic training mechanism in an end-to-end manner from scratch. Experiments on two popular semantic segmentation datasets, i.e., PASCAL VOC and COCO, reveal the effectiveness of the proposed MetaSegNet in the K-way few-shot semantic segmentation task. Pinzhuo Tian, Zhangkai Wu, Lei Qi 0001, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001 |
AAAI | 4 |
| 2020 | Few-Shot Object Detection by Second-Order Pooling
Shan Zhang 0002, Dawei Luo, Lei Wang 0001, Piotr Koniusz |
ACCV (4) | 3 |
| 2020 | ReDro: Efficiently Learning Large-Sized SPD Visual Representation
Saimunur Rahman, Lei Wang 0001, Changming Sun, Luping Zhou |
ECCV (15) | 2 |
| 2020 | Asymmetric Distribution Measure for Few-shot LearningabstractThe core idea of metric-based few-shot image classification is to directly measure the relations between query images and support classes to learn transferable feature embeddings. Previous work mainly focuses on image-level feature representations, which actually cannot effectively estimate a class's distribution due to the scarcity of samples. Some recent work shows that local descriptor based representations can achieve richer representations than image-level based representations. However, such works are still based on a less effective instance-level metric, especially a symmetric metric, to measure the relation between a query image and a support class. Given the natural asymmetric relation between a query image and a support class, we argue that an asymmetric measure is more suitable for metric-based few-shot learning. To that end, we propose a novel Asymmetric Distribution Measure (ADM) network for few-shot learning by calculating a joint local and global asymmetric measure between two multivariate local distributions of a query and a class. Moreover, a task-aware Contrastive Measure Strategy (CMS) is proposed to further enhance the measure function. On popular miniImageNet and tieredImageNet, ADM can achieve the state-of-the-art results, validating our innovative design of asymmetric distribution measures for few-shot learning. The source code can be downloaded from https://github.com/WenbinLee/ADM.git. Wenbin Li 0006, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001, Jiebo Luo 0001 |
IJCAI | 2 |
| 2020 | HIME: Mining and Ensembling Heterogeneous Information for Protein-Protein Interactions PredictionabstractResearch on protein-protein interactions (PPIs) data paves the way towards understanding the mechanisms of infectious diseases, however improving the prediction performance of PPIs of inter-species remains a challenge. Since one single type of sequence data such as amino acid composition may be deficient for high-quality prediction of protein interactions, we have investigated a broader range of heterogeneous information of sequences data. This paper proposes a novel framework for PPIs prediction based on Heterogeneous Information Mining and Ensembling (HIME) process to effectively learn from the interaction data. In particular, the proposed approach introduces an ensemble process together with substantial features that generate better performance of PPIs prediction task. The performance of the proposed framework is validated on real protein interaction datasets. The extensive experiments show that HIME achieves higher performance over all existing methods reported in literature so far. Huaming Chen, Yaochu Jin, Lei Wang 0001, Chihung Chi, Jun Shen 0001 |
IJCNN | 3 |
| 2020 | Learning Sample-Adaptive Intensity Lookup Table for Brain Tumor Segmentation
Biting Yu, Luping Zhou, Lei Wang 0001, Wanqi Yang, Ming Yang 0014, Pierrick Bourgeat, Jurgen Fripp |
MICCAI (4) | 3 |
| 2020 | APEX2S: A two-layer machine learning model for discovery of host-pathogen protein-protein interactions on cloud-based multiomics dataabstractSummary Presented by the avalanche of biological interactions data, computational biology is now facing greater challenges on big data analysis and solicits more studies to mine and integrate cloud‐based multiomics data, especially when the data are related to infectious diseases. Meanwhile, machine learning techniques have recently succeeded in different computational biology tasks. In this article, we have calibrated the focus for host‐pathogen protein‐protein interactions study, aiming to apply the machine learning techniques for learning the interactions data and making predictions. A comprehensive and practical workflow to harness different cloud‐based multiomics data is discussed. In particular, a novel two‐layer machine learning model, namely APEX2S, is proposed for discovery of the protein‐protein interactions data. The results show that our model can better learn and predict from the accumulated host‐pathogen protein‐protein interactions. Huaming Chen, Jun Shen 0001, Lei Wang 0001, Chihung Chi |
Concurr. Comput. Pract. Exp. | 3 |
| 2020 | Deep learning based HEp-2 image classification: A comprehensive review
Saimunur Rahman, Lei Wang 0001, Changming Sun, Luping Zhou |
Medical Image Anal. | 2 |
| 2020 | Absent Multiple Kernel Learning AlgorithmsabstractMultiple kernel learning (MKL) has been intensively studied during the past decade. It optimally combines the multiple channels of each sample to improve classification performance. However, existing MKL algorithms cannot effectively handle the situation where some channels of the samples are missing, which is not uncommon in practical applications. This paper proposes three absent MKL (AMKL) algorithms to address this issue. Different from existing approaches where missing channels are first imputed and then a standard MKL algorithm is deployed on the imputed data, our algorithms directly classify each sample based on its observed channels, without performing imputation. Specifically, we define a margin for each sample in its own relevant space, a space corresponding to the observed channels of that sample. The proposed AMKL algorithms then maximize the minimum of all sample-based margins, and this leads to a difficult optimization problem. We first provide two two-step iterative algorithms to approximately solve this problem. After that, we show that this problem can be reformulated as a convex one by applying the representer theorem. This makes it readily be solved via existing convex optimization packages. In addition, we provide a generalization error bound to justify the proposed AMKL algorithms from a theoretical perspective. Extensive experiments are conducted on nine UCI and six MKL benchmark datasets to compare the proposed algorithms with existing imputation-based methods. As demonstrated, our algorithms achieve superior performance and the improvement is more significant with the increase of missing ratio. Xinwang Liu 0002, Lei Wang 0001, Xinzhong Zhu, Miaomiao Li 0001, En Zhu, Tongliang Liu, Li Liu 0002, Yong Dou, Jianping Yin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Multiple Kernel $k$k-Means with Incomplete KernelsabstractMultiple kernel clustering (MKC) algorithms optimally combine a group of pre-specified base kernel matrices to improve clustering performance. However, existing MKC algorithms cannot efficiently address the situation where some rows and columns of base kernel matrices are absent. This paper proposes two simple yet effective algorithms to address this issue. Different from existing approaches where incomplete kernel matrices are first imputed and a standard MKC algorithm is applied to the imputed kernel matrices, our first algorithm integrates imputation and clustering into a unified learning procedure. Specifically, we perform multiple kernel clustering directly with the presence of incomplete kernel matrices, which are treated as auxiliary variables to be jointly optimized. Our algorithm does not require that there be at least one complete base kernel matrix over all the samples. Also, it adaptively imputes incomplete kernel matrices and combines them to best serve clustering. Moreover, we further improve this algorithm by encouraging these incomplete kernel matrices to mutually complete each other. The three-step iterative algorithm is designed to solve the resultant optimization problems. After that, we theoretically study the generalization bound of the proposed algorithms. Extensive experiments are conducted on 13 benchmark data sets to compare the proposed algorithms with existing imputation-based methods. Our algorithms consistently achieve superior performance and the improvement becomes more significant with increasing missing ratio, verifying the effectiveness and advantages of the proposed joint imputation and clustering. Xinwang Liu 0002, Xinzhong Zhu, Miaomiao Li 0001, Lei Wang 0001, En Zhu, Tongliang Liu, Marius Kloft, Dinggang Shen, Jianping Yin, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Progressive Cross-Camera Soft-Label Learning for Semi-Supervised Person Re-IdentificationabstractIn this paper, we focus on the semi-supervised person re-identification (Re-ID) case, which only has the intra-camera (within-camera) labels but not inter-camera (cross-camera) labels. In real-world applications, these intra-camera labels can be readily captured by tracking algorithms or few manual annotations, when compared with cross-camera labels. In this case, it is very difficult to explore the relationships between cross-camera persons in the training stage due to the lack of cross-camera label information. To deal with this issue, we propose a novel Progressive Cross-camera Soft-label Learning (PCSL) framework for the semi-supervised person Re-ID task, which can generate cross-camera soft-labels and utilize them to optimize the network. Concretely, we calculate an affinity matrix based on person-level features and adapt them to produce the similarities between cross-camera persons (i.e., cross-camera soft-labels). To exploit these soft-labels to train the network, we investigate the weighted cross-entropy loss and the weighted triplet loss from the classification and discrimination perspectives, respectively. Particularly, the proposed framework alternately generates progressive cross-camera soft-labels and gradually improves feature representations in the whole learning course. Extensive experiments on five large-scale benchmark datasets show that PCSL significantly outperforms the state-of-the-art unsupervised methods that employ labeled source domains or the images generated by the GANs-based models. Furthermore, the proposed method even has a competitive performance with respect to deep supervised Re-ID methods. Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Attention-Diffusion-Bilinear Neural Network for Brain Network AnalysisabstractBrain network provides essential insights in diagnosing many brain disorders. Integrative analysis of multiple types of connectivity, e.g, functional connectivity (FC) and structural connectivity (SC), can take advantage of their complementary information and therefore may help to identify patients. However, traditional brain network methods usually focus on either FC or SC for describing node interactions and only consider the interaction between paired network nodes. To tackle this problem, in this paper, we propose an Attention-Diffusion-Bilinear Neural Network (ADB-NN) framework for brain network analysis, which is trained in an end-to-end manner. The proposed network seamlessly couples FC and SC to learn wider node interactions and generates a joint representation of FC and SC for diagnosis. Specifically, a brain network (graph) is first defined, where each node corresponding to a brain region is governed by the features of brain activities (i.e., FC) extracted from functional magnetic resonance imaging (fMRI), and the presence of edges is determined by neural fiber physical connections (i.e., SC) extracted from Diffusion Tensor Imaging (DTI). Based on this graph, we train two Attention-Diffusion-Bilinear (ADB) modules jointly. In each module, an attention model is utilized to automatically learn the strength of node interactions. This information further guides a diffusion process that generates new node representations by considering the influence from other nodes as well. After that, the second-order statistics of these node representations are extracted by bilinear pooling to form connectivity-based features for disease prediction. The two ADB modules correspond to the one-step and two-step diffusion, respectively. Experiments on a real epilepsy dataset demonstrate the effectiveness and advantages of our proposed method. Jiashuang Huang, Luping Zhou, Lei Wang 0001, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Sample-Adaptive GANs: Linking Global and Local Mappings for Cross-Modality MR Image SynthesisabstractGenerative adversarial network (GAN) has been widely explored for cross-modality medical image synthesis. The existing GAN models usually adversarially learn a global sample space mapping from the source-modality to the target-modality and then indiscriminately apply this mapping to all samples in the whole space for prediction. However, due to the scarcity of training samples in contrast to the complicated nature of medical image synthesis, learning a single global sample space mapping that is "optimal" to all samples is very challenging, if not intractable. To address this issue, this paper proposes sample-adaptive GAN models, which not only cater for the global sample space mapping between the source- and the target-modalities but also explore the local space around each given sample to extract its unique characteristic. Specifically, the proposed sample-adaptive GANs decompose the entire learning model into two cooperative paths. The baseline path learns a common GAN model by fitting all the training samples as usual for the global sample space mapping. The new sample-adaptive path additionally models each sample by learning its relationship with its neighboring training samples and using the target-modality features of these training samples as auxiliary information for synthesis. Enhanced by this sample-adaptive path, the proposed sample-adaptive GANs are able to flexibly adjust themselves to different samples, and therefore optimize the synthesis performance. Our models have been verified on three cross-modality MR image synthesis tasks from two public datasets, and they significantly outperform the state-of-the-art methods in comparison. Moreover, the experiment also indicates that our sample-adaptive strategy could be utilized to improve various backbone GAN models. It complements the existing GANs models and can be readily integrated when needed. Biting Yu, Luping Zhou, Lei Wang 0001, Yinghuan Shi, Jurgen Fripp, Pierrick Bourgeat |
IEEE Trans. Medical Imaging | 3 |
| 2019 | Distribution Consistency Based Covariance Metric Networks for Few-Shot LearningabstractFew-shot learning aims to recognize new concepts from very few examples. However, most of the existing few-shot learning methods mainly concentrate on the first-order statistic of concept representation or a fixed metric on the relation between a sample and a concept. In this work, we propose a novel end-to-end deep architecture, named Covariance Metric Networks (CovaMNet). The CovaMNet is designed to exploit both the covariance representation and covariance metric based on the distribution consistency for the few-shot classification tasks. Specifically, we construct an embedded local covariance representation to extract the second-order statistic information of each concept and describe the underlying distribution of this concept. Upon the covariance representation, we further define a new deep covariance metric to measure the consistency of distributions between query samples and new concepts. Furthermore, we employ the episodic training mechanism to train the entire network in an end-to-end manner from scratch. Extensive experiments in two tasks, generic few-shot image classification and fine-grained fewshot image classification, demonstrate the superiority of the proposed CovaMNet. The source code can be available from https://github.com/WenbinLee/CovaMNet.git. Wenbin Li 0006, Jinglin Xu, Jing Huo, Lei Wang 0001, Yang Gao 0001, Jiebo Luo 0001 |
AAAI | 4 |
| 2019 | Hyperparameter Estimation in SVM with GPU Acceleration for Prediction of Protein-Protein InteractionsabstractFor classification tasks, such as protein-protein interactions (PPI), support vector machines (SVMs) have been continually utilised as a standard machine learning model. However, most practices in PPIs classifications are limited to common circumstances with small datasets and low feature dimensions, due to the big computation burden of kernel functions and quadratic optimization of SVM. Alternatively, these practical experiences might tend to employ a linear model once the dataset becomes larger, which may have exclusively lost the kernel function's potential. Since there are different defined kernels and various groups of hyperparameter, the time costs in estimating a best set of hyperparameter by traditional grid search are subsequently tremendous for PPI classification. To address this challenge, in this paper, we present a more efficient solution of hyperparameter estimation by gaining acceleration with GPU, which trains SVM efficiently and accurately with kernel functions calculation accelerated on various PPI datasets. The experiments are firstly conducted on PPI classification task, and we have exclusively evaluated the effectiveness on five public classification datasets. Our solution demonstrates a faster and more accurate performance comparing with the state-of-the-art. Huaming Chen, Lei Wang 0001, Yaochu Jin, Chihung Chi, Fucun Li, Huaiyuan Chu, Jun Shen 0001 |
IEEE BigData | 2 |
| 2019 | Revisiting Local Descriptor Based Image-To-Class Measure for Few-Shot LearningabstractFew-shot learning in image classification aims to learn a classifier to classify images when only few training examples are available for each class. Recent work has achieved promising classification performance, where an image-level feature based measure is usually used. In this paper, we argue that a measure at such a level may not be effective enough in light of the scarcity of examples in few-shot learning. Instead, we think a local descriptor based image-to-class measure should be taken, inspired by its surprising success in the heydays of local invariant features. Specifically, building upon the recent episodic training mechanism, we propose a Deep Nearest Neighbor Neural Network (DN4 in short) and train it in an end-to-end manner. Its key difference from the literature is the replacement of the image-level feature based measure in the final layer by a local descriptor based image-to-class measure. This measure is conducted online via a k-nearest neighbor search over the deep local descriptors of convolutional feature maps. The proposed DN4 not only learns the optimal deep local descriptors for the image-to-class measure, but also utilizes the higher efficiency of such a measure in the case of example scarcity, thanks to the exchangeability of visual patterns across the images in the same class. Our work leads to a simple, effective, and computationally efficient framework for few-shot learning. Experimental study on benchmark datasets consistently shows its superiority over the related state-of-the-art, with the largest absolute improvement of 17% over the next best. The source code can be available from https://github.com/WenbinLee/DN4.git. Wenbin Li 0006, Lei Wang 0001, Jinglin Xu, Jing Huo, Yang Gao 0001, Jiebo Luo 0001 |
CVPR | 2 |
| 2019 | A Novel Unsupervised Camera-Aware Domain Adaptation Framework for Person Re-IdentificationabstractUnsupervised cross-domain person re-identification (Re-ID) faces two key issues. One is the data distribution discrepancy between source and target domains, and the other is the lack of discriminative information in target domain. From the perspective of representation learning, this paper proposes a novel end-to-end deep domain adaptation framework to address them. For the first issue, we highlight the presence of camera-level sub-domains as a unique characteristic in person Re-ID, and develop a “camera-aware” domain adaptation method via adversarial learning. With this method, the learned representation reduces distribution discrepancy not only between source and target domains but also across all cameras. For the second issue, we exploit the temporal continuity in each camera of target domain to create discriminative information. This is implemented by dynamically generating online triplets within each batch, in order to maximally take advantage of the steadily improved representation in training process. Together, the above two methods give rise to a new unsupervised domain adaptation framework for person Re-ID. Extensive experiments and ablation studies conducted on benchmark datasets demonstrate its superiority and interesting properties. Lei Qi 0001, Lei Wang 0001, Jing Huo, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
ICCV | 2 |
| 2019 | Miss Detection vs. False Alarm: Adversarial Learning for Small Object Segmentation in Infrared ImagesabstractA key challenge of infrared small object segmentation (ISOS) is to balance miss detection (MD) and false alarm (FA). This usually needs ``opposite'' strategies to suppress the two terms, and has not been well resolved in the literature. In this paper, we propose a deep adversarial learning framework to improve this situation. Departing from the tradition of jointly reducing MD and FA via a single objective, we decompose this difficult task into two sub-tasks handled by two models trained adversarially, with each focusing on reducing either MD or FA. Such a new design brings forth at least three advantages. First, as each model focuses on a relatively simpler sub-task, the overall difficulty of ISOS is somehow decreased. Second, the adversarial training of the two models naturally produces a delicate balance of MD and FA, and low rates for both MD and FA could be achieved at Nash equilibrium. Third, this MD-FA detachment gives us more flexibility to develop specific models dedicated to each sub-task. To realize the above design, we propose a conditional Generative Adversarial Network comprising of two generators and one discriminator. Each generator strives for one sub-task, while the discriminator differentiates the three segmentation results from the two generators and the ground truth. Moreover, in order to better serve the sub-tasks, the two generators, based on context aggregation networks, utilzse different size of receptive fields, providing both local and global views of objects for segmentation. As verified on multiple infrared image data sets, our method consistently achieves better segmentation than many state-of-the-art ISOS methods. Luping Zhou, Lei Wang 0001 |
ICCV | 3 |
| 2019 | A Mask Based Deep Ranking Neural Network for Person RetrievalabstractPerson retrieval faces many challenges including cluttered background, appearance variations (e.g., illumination, pose, occlusion) among different camera views and the similarity among different person's images. To address these issues, we put forward a novel mask based deep ranking neural network with a skipped fusing layer. Firstly, to alleviate the problem of cluttered background, masked images with only the foreground regions are incorporated as input in the proposed neural network. Secondly, to reduce the impact of the appearance variations, the multi-layer fusion scheme is developed to obtain more discriminative fine-grained information. Lastly, considering person retrieval is a special image retrieval task, we propose a novel ranking loss to optimize the whole network. The proposed ranking loss can further mitigate the interference problem of similar negative samples when producing ranking results. The extensive experiments validate the superiority of the proposed method compared with the state-of-the-art methods on many benchmark datasets. Lei Qi 0001, Jing Huo, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001 |
ICME | 3 |
| 2019 | Integrating Functional and Structural Connectivities via Diffusion-Convolution-Bilinear Neural Network
Jiashuang Huang, Luping Zhou, Lei Wang 0001, Daoqiang Zhang |
MICCAI (3) | 3 |
| 2019 | Calibrated Multi-label Classification with Label Correlations
Zhifen He, Ming Yang 0014, Hui-Dong Liu, Lei Wang 0001 |
Neural Process. Lett. | 4 |
| 2019 | Late Fusion Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering optimally integrates a group of pre-specified incomplete views to improve clustering performance. Among various excellent solutions, multiple kernel $k$k-means with incomplete kernels forms a benchmark, which redefines the incomplete multi-view clustering as a joint optimization problem where the imputation and clustering are alternatively performed until convergence. However, the comparatively intensive computational and storage complexities preclude it from practical applications. To address these issues, we propose Late Fusion Incomplete Multi-view Clustering (LF-IMVC) which effectively and efficiently integrates the incomplete clustering matrices generated by incomplete views. Specifically, our algorithm jointly learns a consensus clustering matrix, imputes each incomplete base matrix, and optimizes the corresponding permutation matrices. We develop a three-step iterative algorithm to solve the resultant optimization problem with linear computational complexity and theoretically prove its convergence. Further, we conduct comprehensive experiments to study the proposed LF-IMVC in terms of clustering accuracy, running time, advantages of late fusion multi-view clustering, evolution of the learned consensus clustering matrix, parameter sensitivity and convergence. As indicated, our algorithm significantly and consistently outperforms some state-of-the-art algorithms with much less running time and memory. Xinwang Liu 0002, Xinzhong Zhu, Miaomiao Li 0001, Lei Wang 0001, Chang Tang, Jianping Yin, Dinggang Shen, Huaimin Wang 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | A Probabilistic Approach to Cross-Region Matching-Based Image RetrievalabstractWith deep convolutional features, cross-region matching (CRM) has recently shown superior performance on image retrieval. It evaluates image similarity by comparing image regions at different locations and scales, and is, therefore, more robust to geometric variance of objects. This paper first scrutinizes CRM-based image retrieval to provide a rigorous probabilistic interpretation by following the probability ranking principle. In addition to manifesting the assumptions implicitly taken by CRM, our interpretation highlights a fundamental issue hindering the performance of CRM-when comparing two image regions, CRM ignores modeling the distribution of the visual concept class associated with an image region, making the similarity comparison less precise. Taking advantage of the unprecedented representation capability of deep convolutional features, this paper proposes one approach to tackle that issue. It treats locally clustered image regions as a pseudo-labeled class sharing the same visual concept and utilizes them to model the distribution of the visual concept class associated with an image region. Both non-parametric and parametric methods are developed for this purpose, with careful probabilistic justification. Extensive experimental study on multiple benchmark data sets demonstrates the superior performance of the proposed pseudo-label approach to CRM and other comparable methods, with the maximum improvement of more than 10 percentage points over CRM. Zhimin Gao, Lei Wang 0001, Luping Zhou |
IEEE Trans. Image Process. | 2 |
| 2019 | 3D Auto-Context-Based Locality Adaptive Multi-Modality GANs for PET SynthesisabstractPositron emission tomography (PET) has been substantially used recently. To minimize the potential health risk caused by the tracer radiation inherent to PET scans, it is of great interest to synthesize the high-quality PET image from the low-dose one to reduce the radiation exposure. In this paper, we propose a 3D auto-context-based locality adaptive multi-modality generative adversarial networks model (LA-GANs) to synthesize the high-quality FDG PET image from the low-dose one with the accompanying MRI images that provide anatomical information. Our work has four contributions. First, different from the traditional methods that treat each image modality as an input channel and apply the same kernel to convolve the whole image, we argue that the contributions of different modalities could vary at different image locations, and therefore a unified kernel for a whole image is not optimal. To address this issue, we propose a locality adaptive strategy for multi-modality fusion. Second, we utilize 1 ×1 ×1 kernel to learn this locality adaptive fusion so that the number of additional parameters incurred by our method is kept minimum. Third, the proposed locality adaptive fusion mechanism is learned jointly with the PET image synthesis in a 3D conditional GANs model, which generates high-quality PET images by employing large-sized image patches and hierarchical features. Fourth, we apply the auto-context strategy to our scheme and propose an auto-context LA-GANs model to further refine the quality of synthesized images. Experimental results show that our method outperforms the traditional multi-modality fusion methods used in deep networks, as well as the state-of-the-art PET estimation approaches. Yan Wang 0015, Luping Zhou, Biting Yu, Lei Wang 0001, Chen Zu, David S. Lalush, Weili Lin, Xi Wu 0004, Jiliu Zhou, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2019 | Ea-GANs: Edge-Aware Generative Adversarial Networks for Cross-Modality MR Image SynthesisabstractMagnetic resonance (MR) imaging is a widely used medical imaging protocol that can be configured to provide different contrasts between the tissues in human body. By setting different scanning parameters, each MR imaging modality reflects the unique visual characteristic of scanned body part, benefiting the subsequent analysis from multiple perspectives. To utilize the complementary information from multiple imaging modalities, cross-modality MR image synthesis has aroused increasing research interest recently. However, most existing methods only focus on minimizing pixel/voxel-wise intensity difference but ignore the textural details of image content structure, which affects the quality of synthesized images. In this paper, we propose edge-aware generative adversarial networks (Ea-GANs) for cross-modality MR image synthesis. Specifically, we integrate edge information, which reflects the textural structure of image content and depicts the boundaries of different objects in images, to reduce this gap. Corresponding to different learning strategies, two frameworks are proposed, i.e., a generator-induced Ea-GAN (gEa-GAN) and a discriminator-induced Ea-GAN (dEa-GAN). The gEa-GAN incorporates the edge information via its generator, while the dEa-GAN further does this from both the generator and the discriminator so that the edge similarity is also adversarially learned. In addition, the proposed Ea-GANs are 3D-based and utilize hierarchical features to capture contextual information. The experimental results demonstrate that the proposed Ea-GANs, especially the dEa-GAN, outperform multiple state-of-the-art methods for cross-modality MR image synthesis in both qualitative and quantitative measures. Moreover, the dEa-GAN also shows excellent generality to generic image synthesis tasks on benchmark datasets about facades, maps, and cityscapes. Biting Yu, Luping Zhou, Lei Wang 0001, Yinghuan Shi, Jurgen Fripp, Pierrick Bourgeat |
IEEE Trans. Medical Imaging | 3 |
| 2018 | A Joint Local and Global Deep Metric Learning Method for Caricature Recognition
Wenbin Li 0006, Jing Huo, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Jiebo Luo 0001 |
ACCV (4) | 5 |
| 2018 | Towards Biological Sequence Data Service with InsightsabstractTestable prediction outcomes generated by computational models based on available databases are the primary sources helping to design biological experiments. Although numerous databases have been designed by collecting data either only from literature manually or together with prediction outcomes from computational models, there is currently not a comprehensive data service framework delivering better insights for these results. In this paper, we introduce a biological sequence data service towards delivering deeper insights and helping better biological experiments design. The service includes following major components: a comprehensive database for storing biological data, data analytics tools for analysing biological data, and computational models for delivering testable prediction outcomes. Specifically, we present this service in a framework for studies on host-pathogen interactions. The design of this framework aims to improve the understanding of host-pathogen interactions. The relationships of hierarchical databases and their working mechanism, specifically between PPIs and DDIs, are also presented in this framework. Finally, the preliminary and practical experiences of building computational model for prediction is discussed. Huaming Chen, Jun Shen 0001, Lei Wang 0001, Chihung Chi |
IEEE BigData | 3 |
| 2018 | Modelling Diffusion Process by Deep Neural Networks for Image Retrieval
Yan Zhao 0019, Lei Wang 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
BMVC | 2 |
| 2018 | DeepKSPD: Learning Kernel-Matrix-Based SPD Representation For Fine-Grained Image Recognition
Melih Engin, Lei Wang 0001, Luping Zhou, Xinwang Liu 0002 |
ECCV (2) | 2 |
| 2018 | Locality Adaptive Multi-modality GANs for High-Quality PET Image Synthesis
Yan Wang 0015, Luping Zhou, Lei Wang 0001, Biting Yu, Chen Zu, David S. Lalush, Weili Lin, Xi Wu 0004, Jiliu Zhou, Dinggang Shen |
MICCAI (1) | 3 |
| 2018 | Instance Image Retrieval by Aggregating Sample-based Discriminative CharacteristicsabstractIdentifying the discriminative characteristic of a query is important for image retrieval. For retrieval without human interaction, such characteristic is usually obtained by average query expansion (AQE) or its discriminative variant (DQE) learned from pseudo-examples online, among others. In this paper, we propose a new query expansion method to further improve the above ones. The key idea is to learn a "unique'' discriminative characteristic for each database image, in an offline manner. During retrieval, the characteristic of a query is obtained by aggregating the unique characteristics of the query-relevant images collected from an initial retrieval result. Compared with AQE which works in the original feature space, our method works in the space of the unique characteristics of database images, significantly enhancing the discriminative power of the characteristic identified for a query. Compared with DQE, our method needs neither pseudo-labeled negatives nor the online learning process, leading to more efficient retrieval and even better performance. The experimental study conducted on seven benchmark datasets verifies the considerable improvement achieved by the proposed method, and also demonstrates its application to the state-of-the-art diffusion-based image retrieval. Zhongyan Zhang, Lei Wang 0001, Yang Wang 0002, Luping Zhou, Jianjia Zhang, Fang Chen 0001 |
ICMR | 2 |
| 2018 | OPML: A one-pass closed-form solution for online metric learning
Wenbin Li 0006, Yang Gao 0001, Lei Wang 0001, Luping Zhou, Jing Huo, Yinghuan Shi |
Pattern Recognit. | 3 |
| 2018 | Detection and Separation of Smoke From Single Image FramesabstractThis paper proposes novel methods for detecting and separating smoke from a single image frame. Specifically, an image formation model is derived based on the atmospheric scattering models. The separation of a frame into quasi-smoke and quasi-background components is formulated as convex optimization that solves a sparse representation problem using dual dictionaries for the smoke and background components, respectively. A novel feature is constructed as a concatenation of the respective sparse coefficients for detection. In addition, a method based on the concept of image matting is developed to separate the true smoke and background components from the smoke detection results. Extensive experiments on detection were conducted and the results showed that the proposed feature significantly outperforms existing features for smoke detection. In particular, the proposed method is able to differentiate smoke from other challenging objects (e.g. fog/haze, cloud, and so on) with similar visual appearance in a gray-scale frame. Experiments on smoke separation also demonstrated that the proposed separation method can effectively estimate/separate the true smoke and background components. Hongda Tian, Wanqing Li 0001, Philip Ogunbona, Lei Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Incomplete-Data Oriented Multiview Dimension Reduction via Sparse Low-Rank RepresentationabstractFor dimension reduction on multiview data, most of the previous studies implicitly take an assumption that all samples are completed in all views. Nevertheless, this assumption could often be violated in real applications due to the presence of noise, limited access to data, equipment malfunction, and so on. Most of the previous methods will cease to work when missing values in one or multiple views occur, thus an incomplete-data oriented dimension reduction becomes an important issue. To this end, we mathematically formulate the above-mentioned issue as sparse low-rank representation through multiview subspace (SRRS) learning to impute missing values, by jointly measuring intraview relations (via sparse low-rank representation) and interview relations (through common subspace representation). Moreover, by exploiting various subspace priors in the proposed SRRS formulation, we develop three novel dimension reduction methods for incomplete multiview data: 1) multiview subspace learning via graph embedding; 2) multiview subspace learning via structured sparsity; and 3) sparse multiview feature selection via rank minimization. For each of them, the objective function and the algorithm to solve the resulting optimization problem are elaborated, respectively. We perform extensive experiments to investigate their performance on three types of tasks including data recovery, clustering, and classification. Both two toy examples (i.e., Swiss roll and -curve) and four real-world data sets (i.e., face images, multisource news, multicamera activity, and multimodality neuroimaging data) are systematically tested. As demonstrated, our methods achieve the performance superior to that of the state-of-the-art comparable methods. Also, the results clearly show the advantage of integrating the sparsity and low-rankness over using each of them separately. Wanqi Yang, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Ming Yang 0014 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Multiple Kernel k-Means with Incomplete KernelsabstractMultiple kernel clustering (MKC) algorithms optimally combine a group of pre-specified base kernels to improve clustering performance. However, existing MKC algorithms cannot efficiently address the situation where some rows and columns of base kernels are absent. This paper proposes a simple while effective algorithm to address this issue. Different from existing approaches where incomplete kernels are firstly imputed and a standard MKC algorithm is applied to the imputed kernels, our algorithm integrates imputation and clustering into a unified learning procedure. Specifically, we perform multiple kernel clustering directly with the presence of incomplete kernels, which are treated as auxiliary variables to be jointly optimized. Our algorithm does not require that there be at least one complete base kernel over all the samples. Also, it adaptively imputes incomplete kernels and combines them to best serve clustering. A three-step iterative algorithm with proved convergence is designed to solve the resultant optimization problem. Extensive experiments are conducted on four benchmark data sets to compare the proposed algorithm with existing imputation-based methods. Our algorithm consistently achieves superior performance and the improvement becomes more significant with increasing missing ratio, verifying the effectiveness and advantages of the proposed joint imputation and clustering. Xinwang Liu 0002, Miaomiao Li 0001, Lei Wang 0001, Yong Dou, Jianping Yin, En Zhu |
AAAI | 3 |
| 2017 | Collaborative data analytics towards prediction on pathogen-host protein-protein interactionsabstractNowadays more and more data are being sequenced and accumulated in system biology, which brings the data analytics researchers to a brand new era, namely `big data', to extract the inner relationship and knowledge from the huge amount of data. Bridging the gap between computational methodology and biology to accelerate the development of biology analytics has been a hot area. In this paper, we focus on these enormous amounts of data generated with the speedy development of high throughput technologies during the past decades, especially for protein-protein interactions, which are the critical molecular process in biology. Since pathogen-host protein-protein interactions are the major and basic problems for not only infectious diseases but also drug design, molecular level interactions between pathogen and host play very critical role for the study of infection mechanisms. In this paper, we built a basic framework for analyzing the specific problems about pathogen-host protein-protein interactions (PHPPI), meanwhile, we also presented the state-of-art deep learning method results on prediction of PHPPI comparing with other machine learning methods. Utilizing the evaluation methods, specifically by considering the high skewed imbalanced ratio and huge amount of data, we detailed the pipeline solution on both storing and learning for PHPPI. This work contributes as a basis for a further investigation of protein and protein-protein interactions, with the collaboration of data analytics results from the vast amount of data dispersedly available in biology literature. Huaming Chen, Jun Shen 0001, Lei Wang 0001, Jiangning Song |
CSCWD | 3 |
| 2017 | Revisiting Metric Learning for SPD Matrix Based Visual RepresentationabstractThe success of many visual recognition tasks largely depends on a good similarity measure, and distance metric learning plays an important role in this regard. Meanwhile, Symmetric Positive Definite (SPD) matrix is receiving increased attention for feature representation in multiple computer vision applications. However, distance metric learning on SPD matrices has not been sufficiently researched. A few existing works approached this by learning either d2× p or d × k transformation matrix for d× d SPD matrices. Different from these methods, this paper proposes a new member to the family of distance metric learning for SPD matrices. It learns only d parameters to adjust the eigenvalues of the SPD matrices through an efficient optimisation scheme. Also, it is shown that the proposed method can be interpreted as learning a sample-specific transformation matrix, instead of the fixed transformation matrix learned for all the samples in the existing works. The optimised d parameters can be used to massage the SPD matrices for better discrimination while still keeping them in the original space. From this perspective, the proposed method complements, rather than competes with, the existing linear-transformation-based methods, as the latter can always be applied to the output of the former to perform distance metric learning in further. The proposed method has been tested on multiple SPD-based visual representation data sets used in the literature, and the results demonstrate its interesting properties and attractive performance. Luping Zhou, Lei Wang 0001, Jianjia Zhang, Yinghuan Shi, Yang Gao 0001 |
CVPR | 2 |
| 2017 | Infomax principle based pooling of deep convolutional activations for image retrievalabstractNeural activations produced by deep convolutional networks have recently become state-of-the-art representation for image retrieval. To obtain a global image representation, sum-pooling has been frequently used to aggregate activations of convolutional feature maps. This work first presents an understanding on the effectiveness of sum-pooling via probabilistic interpretation, by proving that sum-pooling is an upper bound of the probability that a visual pattern is present in an image. To further answer the optimality of sum-pooling, a quantitative analysis based on the Infomax principle in neural networks is provided. It shows that sum-pooling aligns well with the leading eigenvector of principal component analysis (PCA) applied to the activations of a feature map. Moreover, considering the 2D matrix structure of feature maps, a two-directional 2DPCA-based pooling scheme is proposed to aggregate the convolutional activations. Experiments on multiple benchmark image retrieval datasets demonstrate the above analysis and the superiority of the proposed pooling scheme. Zhimin Gao, Lei Wang 0001, Luping Zhou, Ming Yang 0014 |
ICME | 2 |
| 2017 | Compositional Model Based Fisher Vector Coding for Image ClassificationabstractDeriving from the gradient vector of a generative model of local features, Fisher vector coding (FVC) has been identified as an effective coding method for image classification. Most, if not all, FVC implementations employ the Gaussian mixture model (GMM) as the generative model for local features. However, the representative power of a GMM can be limited because it essentially assumes that local features can be characterized by a fixed number of feature prototypes, and the number of prototypes is usually small in FVC. To alleviate this limitation, in this work, we break the convention which assumes that a local feature is drawn from one of a few Gaussian distributions. Instead, we adopt a compositional mechanism which assumes that a local feature is drawn from a Gaussian distribution whose mean vector is composed as a linear combination of multiple key components, and the combination weight is a latent random variable. In doing so we greatly enhance the representative power of the generative model underlying FVC. To implement our idea, we design two particular generative models following this compositional approach. In our first model, the mean vector is sampled from the subspace spanned by a set of bases and the combination weight is drawn from a Laplace distribution. In our second model, we further assume that a local feature is composed of a discriminative part and a residual part. As a result, a local feature is generated by the linear combination of discriminative part bases and residual part bases. The decomposition of the discriminative and residual parts is achieved via the guidance of a pre-trained supervised coding method. By calculating the gradient vector of the proposed models, we derive two new Fisher vector coding strategies. The first is termed Sparse Coding-based Fisher Vector Coding (SCFVC) and can be used as the substitute of traditional GMM based FVC. The second is termed Hybrid Sparse Coding-based Fisher vector coding (HSCFVC) since it combines the merits of both pre-trained supervised coding methods and FVC. Using pre-trained Convolutional Neural Network (CNN) activations as local features, we experimentally demonstrate that the proposed methods are superior to traditional GMM based FVC and achieve state-of-the-art performance in various image classification tasks. Lingqiao Liu, Peng Wang 0023, Chunhua Shen, Lei Wang 0001, Anton van den Hengel, Chao Wang 0063, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Subject-adaptive Integration of Multiple SICE Brain Networks with Different Sparsity
Jianjia Zhang, Luping Zhou, Lei Wang 0001 |
Pattern Recognit. | 3 |
| 2017 | HEp-2 Cell Image Classification With Deep Convolutional Neural NetworksabstractEfficient Human Epithelial-2 cell image classification can facilitate the diagnosis of many autoimmune diseases. This paper proposes an automatic framework for this classification task, by utilizing the deep convolutional neural networks (CNNs) which have recently attracted intensive attention in visual recognition. In addition to describing the proposed classification framework, this paper elaborates several interesting observations and findings obtained by our investigation. They include the important factors that impact network design and training, the role of rotation-based data augmentation for cell images, the effectiveness of cell image masks for classification, and the adaptability of the CNN-based classification system across different datasets. Extensive experimental study is conducted to verify the above findings and compares the proposed framework with the well-established image classification models in the literature. The results on benchmark datasets demonstrate that 1) the proposed framework can effectively outperform existing models by properly applying data augmentation, 2) our CNN-based framework has excellent adaptability across different datasets, which is highly desirable for cell image classification under varying laboratory settings. Our system is ranked high in the cell image classification competition hosted by ICPR 2014. Zhimin Gao, Lei Wang 0001, Luping Zhou, Jianjia Zhang |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Two-Stage Friend Recommendation Based on Network Alignment and Series Expansion of Probabilistic Topic ModelabstractPrecise friend recommendation is an important problem in social media. Although most social websites provide some kinds of auto friend searching functions, their accuracies are not satisfactory. In this paper, we propose a more precise auto friend recommendation method with two stages. In the first stage, by utilizing the information of the relationship between texts and users, as well as the friendship information between users, we align different social networks and choose some “possible friends.” In the second stage, with the relationship between image features and users, we build a topic model to further refine the recommendation results. Because some traditional methods, such as variational inference and Gibbs sampling, have their limitations in dealing with our problem, we develop a novel method to find out the solution of the topic model based on series expansion. We conduct experiments on the Flickr dataset to show that the proposed algorithm recommends friends more precisely and faster than traditional methods. Shangrong Huang, Jian Zhang 0002, Dan Schonfeld, Lei Wang 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | A Graph-Embedding Approach to Hierarchical Visual Word MergenceabstractAppropriately merging visual words are an effective dimension reduction method for the bag-of-visual-words model in image classification. The approach of hierarchically merging visual words has been extensively employed, because it gives a fully determined merging hierarchy. Existing supervised hierarchical merging methods take different approaches and realize the merging process with various formulations. In this paper, we propose a unified hierarchical merging approach built upon the graph-embedding framework. Our approach is able to merge visual words for any scenario, where a preferred structure and an undesired structure are defined, and, therefore, can effectively attend to all kinds of requirements for the word-merging process. In terms of computational efficiency, we show that our algorithm can seamlessly integrate a fast search strategy developed in our previous work and, thus, well maintain the state-of-the-art merging speed. To the best of our survey, the proposed approach is the first one that addresses the hierarchical visual word mergence in such a flexible and unified manner. As demonstrated, it can maintain excellent image classification performance even after a significant dimension reduction, and outperform all the existing comparable visual word-merging methods. In a broad sense, our work provides an open platform for applying, evaluating, and developing new criteria for hierarchical word-merging tasks. Lei Wang 0001, Lingqiao Liu, Luping Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Multiple Kernel k-Means Clustering with Matrix-Induced RegularizationabstractMultiple kernel k-means (MKKM) clustering aims to optimally combine a group of pre-specified kernels to improve clustering performance. However, we observe that existing MKKM algorithms do not sufficiently consider the correlation among these kernels. This could result in selecting mutually redundant kernels and affect the diversity of information sources utilized for clustering, which finally hurts the clustering performance. To address this issue, this paper proposes an MKKM clustering with a novel, effective matrix-induced regularization to reduce such redundancy and enhance the diversity of the selected kernels. We theoretically justify this matrix-induced regularization by revealing its connection with the commonly used kernel alignment criterion. Furthermore, this justification shows that maximizing the kernel alignment for clustering can be viewed as a special case of our approach and indicates the extendability of the proposed matrix-induced regularization for designing better clustering algorithms. As experimentally demonstrated on five challenging MKL benchmark data sets, our algorithm significantly improves existing MKKM and consistently outperforms the state-of-the-art ones in the literature, verifying the effectiveness and advantages of incorporating the proposed matrix-induced regularization. Xinwang Liu 0002, Yong Dou, Jianping Yin, Lei Wang 0001, En Zhu |
AAAI | 4 |
| 2016 | Multiple Kernel Clustering with Local Kernel Alignment Maximization
Miaomiao Li 0001, Xinwang Liu 0002, Lei Wang 0001, Yong Dou, Jianping Yin, En Zhu |
IJCAI | 3 |
| 2016 | A Generalized Probabilistic Framework for Compact Codebook CreationabstractCompact and discriminative visual codebooks are preferred in many visual recognition tasks. In the literature, a number of works have taken the approach of hierarchically merging visual words of an initial large-sized codebook, but implemented this approach with different merging criteria. In this work, we propose a single probabilistic framework to unify these merging criteria, by identifying two key factors: the function used to model the class-conditional distribution and the method used to estimate the distribution parameters. More importantly, by adopting new distribution functions and/or parameter estimation methods, our framework can readily produce a spectrum of novel merging criteria. Three of them are specifically discussed in this paper. For the first criterion, we adopt the multinomial distribution with the Bayesian method; For the second criterion, we integrate the Gaussian distribution with maximum likelihood parameter estimation. For the third criterion, which shows the best merging performance, we propose a max-margin-based parameter estimation method and apply it with the multinomial distribution. Extensive experimental study is conducted to systematically analyze the performance of the above three criteria and compare them with existing ones. As demonstrated, the best criterion within our framework achieves the overall best merging performance among the compared merging criteria developed in the literature. Lingqiao Liu, Lei Wang 0001, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Learning Discriminative Bayesian Networks from High-Dimensional Continuous Neuroimaging DataabstractDue to its causal semantics, Bayesian networks (BN) have been widely employed to discover the underlying data relationship in exploratory studies, such as brain research. Despite its success in modeling the probability distribution of variables, BN is naturally a generative model, which is not necessarily discriminative. This may cause the ignorance of subtle but critical network changes that are of investigation values across populations. In this paper, we propose to improve the discriminative power of BN models for continuous variables from two different perspectives. This brings two general discriminative learning frameworks for Gaussian Bayesian networks (GBN). In the first framework, we employ Fisher kernel to bridge the generative models of GBN and the discriminative classifiers of SVMs, and convert the GBN parameter learning to Fisher kernel learning via minimizing a generalization error bound of SVMs. In the second framework, we employ the max-margin criterion and build it directly upon GBN models to explicitly optimize the classification performance of the GBNs. The advantages and disadvantages of the two frameworks are discussed and experimentally compared. Both of them demonstrate strong power in learning discriminative parameters of GBNs for neuroimaging based brain network analysis, as well as maintaining reasonable representation capacity. The contributions of this paper also include a new Directed Acyclic Graph (DAG) constraint with theoretical guarantee to ensure the graph validity of GBN. Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Guest Editors' Introduction: Special issue on deep learning with applications to visual representation and analysis
Lei Wang 0001, Ce Zhu, Jieping Ye, Juergen Gall |
Signal Process. Image Commun. | 1 |
| 2016 | Social Friend Recommendation Based on Multiple Network CorrelationabstractFriend recommendation is an important recommender application in social media. Major social websites such as Twitter and Facebook are all capable of recommending friends to individuals. However, most of these websites use simple friend recommendation algorithms such as similarity, popularity, or “friend's friends are friends,” which are intuitive but consider few of the characteristics of the social network. In this paper we investigate the structure of social networks and develop an algorithm for network correlation-based social friend recommendation (NC-based SFR). To accomplish this goal, we correlate different “social role” networks, find their relationships and make friend recommendations. NC-based SFR is characterized by two key components: 1) related networks are aligned by selecting important features from each network, and 2) the network structure should be maximally preserved before and after network alignment. After important feature selection has been made, we recommend friends based on these features. We conduct experiments on the Flickr network, which contains more than ten thousand nodes and over 30 thousand tags covering half a million photos, to show that the proposed algorithm recommends friends more precisely than reference methods. Shangrong Huang, Jian Zhang 0002, Lei Wang 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | Learning Discriminative Stein Kernel for SPD Matrices and Its ApplicationsabstractStein kernel (SK) has recently shown promising performance on classifying images represented by symmetric positive definite (SPD) matrices. It evaluates the similarity between two SPD matrices through their eigenvalues. In this paper, we argue that directly using the original eigenvalues may be problematic because: 1) eigenvalue estimation becomes biased when the number of samples is inadequate, which may lead to unreliable kernel evaluation, and 2) more importantly, eigenvalues reflect only the property of an individual SPD matrix. They are not necessarily optimal for computing SK when the goal is to discriminate different classes of SPD matrices. To address the two issues, we propose a discriminative SK (DSK), in which an extra parameter vector is defined to adjust the eigenvalues of input SPD matrices. The optimal parameter values are sought by optimizing a proxy of classification performance. To show the generality of the proposed method, three kernel learning criteria that are commonly used in the literature are employed as a proxy. A comprehensive experimental study is conducted on a variety of image classification tasks to compare the proposed DSK with the original SK and other methods for evaluating the similarity between SPD matrices. The results demonstrate that the DSK can attain greater discrimination and better align with classification tasks by altering the eigenvalues. This makes it produce higher classification performance than the original SK and other commonly used methods. Jianjia Zhang, Lei Wang 0001, Luping Zhou, Wanqing Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Absent Multiple Kernel LearningabstractMultiple kernel learning (MKL) optimally combines the multiple channels of each sample to improve classification performance. However, existing MKL algorithms cannot effectively handle the situation where some channels are missing, which is common in practical applications. This paper proposes an absent MKL (AMKL) algorithm to address this issue. Different from existing approaches where missing channels are firstly imputed and then a standard MKL algorithm is deployed on the imputed data, our algorithm directly classifies each sample with its observed channels. In specific, we define a margin for each sample in its own relevant space, which corresponds to the observed channels of that sample. The proposed AMKL algorithm then maximizes the minimum of all sample-based margins, and this leads to a difficult optimization problem. We show that this problem can be reformulated as a convex one by applying the representer theorem. This makes it readily be solved via existing convex optimization packages. Extensive experiments are conducted on five MKL benchmark data sets to compare the proposed algorithm with existing imputation-based methods. As observed, our algorithm achieves superior performance and the improvement is more significant with the increasing missing ratio. Xinwang Liu 0002, Lei Wang 0001, Jianping Yin, Yong Dou, Jian Zhang 0002 |
AAAI | 2 |
| 2015 | Beyond Covariance: Feature Representation with Nonlinear Kernel MatricesabstractCovariance matrix has recently received increasing attention in computer vision by leveraging Riemannian geometry of symmetric positive-definite (SPD) matrices. Originally proposed as a region descriptor, it has now been used as a generic representation in various recognition tasks. However, covariance matrix has shortcomings such as being prone to be singular, limited capability in modeling complicated feature relationship, and having a fixed form of representation. This paper argues that more appropriate SPD-matrix-based representations shall be explored to achieve better recognition. It proposes an open framework to use the kernel matrix over feature dimensions as a generic representation and discusses its properties and advantages. The proposed framework significantly elevates covariance representation to the unlimited opportunities provided by this new representation. Experimental study shows that this representation consistently outperforms its covariance counterpart on various visual recognition tasks. In particular, it achieves significant improvement on skeleton-based human action recognition, demonstrating the state-of-the-art performance over both the covariance and the existing non-covariance representations. Lei Wang 0001, Jianjia Zhang, Luping Zhou, Chang Tang, Wanqing Li 0001 |
ICCV | 1 |
| 2015 | Multiple kernel extreme learning machine
Xinwang Liu 0002, Lei Wang 0001, Guang-Bin Huang, Jian Zhang 0002, Jianping Yin |
Neurocomputing | 2 |
| 2015 | An efficient radius-incorporated MKL algorithm for Alzheimer's disease prediction
Xinwang Liu 0002, Luping Zhou, Lei Wang 0001, Jian Zhang 0002, Jianping Yin, Dinggang Shen |
Pattern Recognit. | 3 |
| 2015 | Density Maximization for Improving Graph Matching With Its ApplicationsabstractGraph matching has been widely used in both image processing and computer vision domain due to its powerful performance for structural pattern representation. However, it poses three challenges to image sparse feature matching: 1) the combinatorial nature limits the size of the possible matches; 2) it is sensitive to outliers because its objective function prefers more matches; and 3) it works poorly when handling many-to-many object correspondences, due to its assumption of one single cluster of true matches. In this paper, we address these challenges with a unified framework called density maximization (DM), which maximizes the values of a proposed graph density estimator both locally and globally. DM leads to the integration of feature matching, outlier elimination, and cluster detection. Experimental evaluation demonstrates that it significantly boosts the true matches and enables graph matching to handle both outliers and many-to-many object correspondences. We also extend it to dense correspondence estimation and obtain large improvement over the state-of-the-art methods. We further demonstrate the usefulness of our methods using three applications: 1) instance-level image retrieval; 2) mask transfer; and 3) image enhancement. Chao Wang 0063, Lei Wang 0001, Lingqiao Liu |
IEEE Trans. Image Process. | 2 |
| 2014 | Sample-Adaptive Multiple Kernel LearningabstractExisting multiple kernel learning (MKL) algorithms \textit{indiscriminately} apply a same set of kernel combination weights to all samples. However, the utility of base kernels could vary across samples and a base kernel useful for one sample could become noisy for another. In this case, rigidly applying a same set of kernel combination weights could adversely affect the learning performance. To improve this situation, we propose a sample-adaptive MKL algorithm, in which base kernels are allowed to be adaptively switched on/off with respect to each sample. We achieve this goal by assigning a latent binary variable to each base kernel when it is applied to a sample. The kernel combination weights and the latent variables are jointly optimized via margin maximization principle. As demonstrated on five benchmark data sets, the proposed algorithm consistently outperforms the comparable ones in the literature. Xinwang Liu 0002, Lei Wang 0001, Jian Zhang 0002, Jianping Yin |
AAAI | 2 |
| 2014 | Single Image Smoke Detection
Hongda Tian, Wanqing Li 0001, Philip Ogunbona, Lei Wang 0001 |
ACCV (2) | 4 |
| 2014 | Discriminative Sparse Inverse Covariance Matrix: Application in Brain Functional Network ClassificationabstractRecent studies show that mental disorders change the functional organization of the brain, which could be investigated via various imaging techniques. Analyzing such changes is becoming critical as it could provide new biomarkers for diagnosing and monitoring the progression of the diseases. Functional connectivity analysis studies the covary activity of neuronal populations in different brain regions. The sparse inverse covariance estimation (SICE), also known as graphical LASSO, is one of the most important tools for functional connectivity analysis, which estimates the interregional partial correlations of the brain. Although being increasingly used for predicting mental disorders, SICE is basically a generative method that may not necessarily perform well on classifying neuroimaging data. In this paper, we propose a learning framework to effectively improve the discriminative power of SICEs by taking advantage of the samples in the opposite class. We formulate our objective as convex optimization problems for both one-class and two-class classifications. By analyzing these optimization problems, we not only solve them efficiently in their dual form, but also gain insights into this new learning framework. The proposed framework is applied to analyzing the brain metabolic covariant networks built upon FDG-PET images for the prediction of the Alzheimer's disease, and shows significant improvement of classification performance for both one-class and two-class scenarios. Moreover, as SICE is a general method for learning undirected Gaussian graphical models, this paper has broader meanings beyond the scope of brain research. Luping Zhou, Lei Wang 0001, Philip Ogunbona |
CVPR | 2 |
| 2014 | Progressive Mode-Seeking on Graphs for Sparse Feature Matching
Chao Wang 0063, Lei Wang 0001, Lingqiao Liu |
ECCV (2) | 2 |
| 2014 | A Method of Discriminative Information Preservation and In-Dimension Distance Minimization Method for Feature SelectionabstractPreserving sample's pair wise similarity is essential for feature selection. In supervised learning, labels can be used as a direct measure to check whether two samples are similar with each other. In unsupervised learning, however, such similarity information is usually unavailable. In this paper, we propose a new feature selection method through spectral clustering based on discriminative information as an underlying data structure. Laplacian matrix is used to obtain more partitioning information than other previously proposed structures such as the Eigen space of original data. The high dimension of sample data is projected into a low dimensional space. The in-dimension distance is also considered to get a better compact clustering result. The proposed method can be solved efficiently by updating the projection matrix and its inverse normalized diagonal matrix. A comprehensive experimental study has demonstrated that the proposed method outperforms many state-of-the-art feature selection algorithms with different criterion including the accuracy of clustering/classification and Jaccard score. Shangrong Huang, Jian Zhang 0002, Xinwang Liu 0002, Lei Wang 0001 |
ICPR | 4 |
| 2014 | Max-Margin Based Learning for Discriminative Bayesian Network from Neuroimaging Data
Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen |
MICCAI (3) | 2 |
| 2014 | Encoding High Dimensional Local Features by Sparse Coding Based Fisher Vectors
Lingqiao Liu, Chunhua Shen, Lei Wang 0001, Anton van den Hengel, Chao Wang 0063 |
NIPS | 3 |
| 2014 | Smoke Detection in Video: An Image Separation Approach
Hongda Tian, Wanqing Li 0001, Lei Wang 0001, Philip Ogunbona |
Int. J. Comput. Vis. | 3 |
| 2014 | A Hierarchical Word-Merging Algorithm with Class Separability MeasureabstractIn image recognition with the bag-of-features model, a small-sized visual codebook is usually preferred to obtain a low-dimensional histogram representation and high computational efficiency. Such a visual codebook has to be discriminative enough to achieve excellent recognition performance. To create a compact and discriminative codebook, in this paper we propose to merge the visual words in a large-sized initial codebook by maximally preserving class separability. We first show that this results in a difficult optimization problem. To deal with this situation, we devise a suboptimal but very efficient hierarchical word-merging algorithm, which optimally merges two words at each level of the hierarchy. By exploiting the characteristics of the class separability measure and designing a novel indexing structure, the proposed algorithm can hierarchically merge 10,000 visual words down to two words in merely 90 seconds. Also, to show the properties of the proposed algorithm and reveal its advantages, we conduct detailed theoretical analysis to compare it with another hierarchical word-merging algorithm that maximally preserves mutual information, obtaining interesting findings. Experimental studies are conducted to verify the effectiveness of the proposed algorithm on multiple benchmark data sets. As shown, it can efficiently produce more compact and discriminative codebooks than the state-of-the-art hierarchical word-merging algorithms, especially when the size of the codebook is significantly reduced. Lei Wang 0001, Luping Zhou, Chunhua Shen, Lingqiao Liu, Huan Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | HEp-2 cell image classification with multiple linear descriptors
Lingqiao Liu, Lei Wang 0001 |
Pattern Recognit. | 2 |
| 2014 | Global and Local Structure Preservation for Feature SelectionabstractThe recent literature indicates that preserving global pairwise sample similarity is of great importance for feature selection and that many existing selection criteria essentially work in this way. In this paper, we argue that besides global pairwise sample similarity, the local geometric structure of data is also critical and that these two factors play different roles in different learning scenarios. In order to show this, we propose a global and local structure preservation framework for feature selection (GLSPFS) which integrates both global pairwise sample similarity and local geometric data structure to conduct feature selection. To demonstrate the generality of our framework, we employ methods that are well known in the literature to model the local geometric data structure and develop three specific GLSPFS-based feature selection algorithms. Also, we develop an efficient optimization algorithm with proven global convergence to solve the resulting feature selection problem. A comprehensive experimental study is then conducted in order to compare our feature selection algorithms with many state-of-the-art ones in supervised, unsupervised, and semisupervised learning scenarios. The result indicates that: 1) our framework consistently achieves statistically significant improvement in selection performance when compared with the currently used algorithms; 2) in supervised and semisupervised learning scenarios, preserving global pairwise similarity is more important than preserving local geometric data structure; 3) in the unsupervised scenario, preserving local geometric data structure becomes clearly more important; and 4) the best feature selection performance is always obtained when the two factors are appropriately integrated. In summary, this paper not only validates the advantages of the proposed GLSPFS framework but also gains more insight into the information to be preserved in different feature selection tasks. Xinwang Liu 0002, Lei Wang 0001, Jian Zhang 0002, Jianping Yin, Huan Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | Efficient Dual Approach to Distance Metric LearningabstractDistance metric learning is of fundamental interest in machine learning because the employed distance metric can significantly affect the performance of many learning methods. Quadratic Mahalanobis metric learning is a popular approach to the problem, but typically requires solving a semidefinite programming (SDP) problem, which is computationally expensive. The worst case complexity of solving an SDP problem involving a matrix variable of size D×D with O(D) linear constraints is about O(D(6.5)) using interior-point methods, where D is the dimension of the input data. Thus, the interior-point methods only practically solve problems exhibiting less than a few thousand variables. Because the number of variables is D(D+1)/2, this implies a limit upon the size of problem that can practically be solved around a few hundred dimensions. The complexity of the popular quadratic Mahalanobis metric learning approach thus limits the size of problem to which metric learning can be applied. Here, we propose a significantly more efficient and scalable approach to the metric learning problem based on the Lagrange dual formulation of the problem. The proposed formulation is much simpler to implement, and therefore allows much larger Mahalanobis metric learning problems to be solved. The time complexity of the proposed method is roughly O(D(3)), which is significantly lower than that of the SDP approach. Experiments on a variety of data sets demonstrate that the proposed method achieves an accuracy comparable with the state of the art, but is applicable to significantly larger problems. We also show that the proposed method can be applied to solve more general Frobenius norm regularized SDP problems approximately. Chunhua Shen, Junae Kim, Fayao Liu, Lei Wang 0001, Anton van den Hengel |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2013 | A Fast Approximate AIB Algorithm for Distributional Word ClusteringabstractDistributional word clustering merges the words having similar probability distributions to attain reliable parameter estimation, compact classification models and even better classification performance. Agglomerative Information Bottleneck (AIB) is one of the typical word clustering algorithms and has been applied to both traditional text classification and recent image recognition. Although enjoying theoretical elegance, AIB has one main issue on its computational efficiency, especially when clustering a large number of words. Different from existing solutions to this issue, we analyze the characteristics of its objective function-the loss of mutual information, and show that by merely using the ratio of word-class joint probabilities of each word, good candidate word pairs for merging can be easily identified. Based on this finding, we propose a fast approximate AIB algorithm and show that it can significantly improve the computational efficiency of AIB while well maintaining or even slightly increasing its classification performance. Experimental study on both text and image classification benchmark data sets shows that our algorithm can achieve more than 100 times speedup on large real data sets over the state-of-the-art method. Lei Wang 0001, Jianjia Zhang, Luping Zhou, Wanqing Li 0001 |
CVPR | 1 |
| 2013 | Discriminative Brain Effective Connectivity Analysis for Alzheimer's Disease: A Kernel Learning Approach upon Sparse Gaussian Bayesian NetworkabstractAnalyzing brain networks from neuroimages is becoming a promising approach in identifying novel connectivity-based biomarkers for the Alzheimer's disease (AD). In this regard, brain ``effective connectivity" analysis, which studies the causal relationship among brain regions, is highly challenging and of many research opportunities. Most of the existing works in this field use generative methods. Despite their success in data representation and other important merits, generative methods are not necessarily discriminative, which may cause the ignorance of subtle but critical disease-induced changes. In this paper, we propose a learning-based approach that integrates the benefits of generative and discriminative methods to recover effective connectivity. In particular, we employ Fisher kernel to bridge the generative models of sparse Bayesian networks (SBN) and the discriminative classifiers of SVMs, and convert the SBN parameter learning to Fisher kernel learning via minimizing a generalization error bound of SVMs. Our method is able to simultaneously boost the discriminative power of both the generative SBN models and the SBN-induced SVM classifiers via Fisher kernel. The proposed method is tested on analyzing brain effective connectivity for AD from ADNI data, and demonstrates significant improvements over the state-of-the-art work. Luping Zhou, Lei Wang 0001, Lingqiao Liu, Philip Ogunbona, Dinggang Shen |
CVPR | 2 |
| 2013 | A Scalable Unsupervised Feature Merging Approach to Efficient Dimensionality Reduction of High-Dimensional Visual DataabstractTo achieve a good trade-off between recognition accuracy and computational efficiency, it is often needed to reduce high-dimensional visual data to medium-dimensional ones. For this task, even applying a simple full-matrix-based linear projection causes significant computation and memory use. When the number of visual data is large, how to efficiently learn such a projection could even become a problem. The recent feature merging approach offers an efficient way to reduce the dimensionality, which only requires a single scan of features to perform reduction. However, existing merging algorithms do not scale well with high-dimensional data, especially in the unsupervised case. To address this problem, we formulate unsupervised feature merging as a PCA problem imposed with a special structure constraint. By exploiting its connection with k-means, we transform this constrained PCA problem into a feature clustering problem. Moreover, we employ the hashing technique to improve its scalability. These produce a scalable feature merging algorithm for our dimensionality reduction task. In addition, we develop an extension of this method by leveraging the neighborhood structure in the data to further improve dimensionality reduction performance. In further, we explore the incorporation of bipolar merging - a variant of merging function which allows the subtraction operation - into our algorithms. Through three applications in visual recognition, we demonstrate that our methods can not only achieve good dimensionality reduction performance with little computational cost but also help to create more powerful representation at both image level and local feature level. Lingqiao Liu, Lei Wang 0001 |
ICCV | 2 |
| 2013 | Improving Graph Matching via Density MaximizationabstractGraph matching has been widely used in various applications in computer vision due to its powerful performance. However, it poses three challenges to image sparse feature matching: (1) The combinatorial nature limits the size of the possible matches, (2) It is sensitive to outliers because the objective function prefers more matches, (3) It works poorly when handling many-to-many object correspondences, due to its assumption of one single cluster for each graph. In this paper, we address these problems with a unified framework-Density Maximization. We propose a graph density local estimator (DLE) to measure the quality of matches. Density Maximization aims to maximize the DLE values both locally and globally. The local maximization of DLE finds the clusters of nodes as well as eliminates the outliers. The global maximization of DLE efficiently refines the matches by exploring a much larger matching space. Our Density Maximization is orthogonal to specific graph matching algorithms. Experimental evaluation demonstrates that it significantly boosts the true matches and enables graph matching to handle both outliers and many-to-many object correspondences. Chao Wang 0063, Lei Wang 0001, Lingqiao Liu |
ICCV | 2 |
| 2013 | Learning with multi-resolution overlapping communities
Xufei Wang, Lei Tang 0001, Huan Liu 0001, Lei Wang 0001 |
Knowl. Inf. Syst. | 4 |
| 2013 | An Efficient Approach to Integrating Radius Information into Multiple Kernel LearningabstractIntegrating radius information has been demonstrated by recent work on multiple kernel learning (MKL) as a promising way to improve kernel learning performance. Directly integrating the radius of the minimum enclosing ball (MEB) into MKL as it is, however, not only incurs significant computational overhead but also possibly adversely affects the kernel learning performance due to the notorious sensitivity of this radius to outliers. Inspired by the relationship between the radius of the MEB and the trace of total data scattering matrix, this paper proposes to incorporate the latter into MKL to improve the situation. In particular, in order to well justify the incorporation of radius information, we strictly comply with the radius-margin bound of support vector machines (SVMs) and thus focus on the l2-norm soft-margin SVM classifier. Detailed theoretical analysis is conducted to show how the proposed approach effectively preserves the merits of incorporating the radius of the MEB and how the resulting optimization is efficiently solved. Moreover, the proposed approach achieves the following advantages over its counterparts: 1) more robust in the presence of outliers or noisy training samples; 2) more computationally efficient by avoiding the quadratic optimization for computing the radius at each iteration; and 3) readily solvable by the existing off-the-shelf MKL packages. Comprehensive experiments are conducted on University of California, Irvine, protein subcellular localization, and Caltech-101 data sets, and the results well demonstrate the effectiveness and efficiency of our approach. Xinwang Liu 0002, Lei Wang 0001, Jianping Yin, En Zhu, Jian Zhang 0002 |
IEEE Trans. Cybern. | 2 |
| 2013 | An Adaptive Approach to Learning Optimal Neighborhood KernelsabstractLearning an optimal kernel plays a pivotal role in kernel-based methods. Recently, an approach called optimal neighborhood kernel learning (ONKL) has been proposed, showing promising classification performance. It assumes that the optimal kernel will reside in the neighborhood of a "pre-specified" kernel. Nevertheless, how to specify such a kernel in a principled way remains unclear. To solve this issue, this paper treats the pre-specified kernel as an extra variable and jointly learns it with the optimal neighborhood kernel and the structure parameters of support vector machines. To avoid trivial solutions, we constrain the pre-specified kernel with a parameterized model. We first discuss the characteristics of our approach and in particular highlight its adaptivity. After that, two instantiations are demonstrated by modeling the pre-specified kernel as a common Gaussian radial basis function kernel and a linear combination of a set of base kernels in the way of multiple kernel learning (MKL), respectively. We show that the optimization in our approach is a min-max problem and can be efficiently solved by employing the extended level method and Nesterov's method. Also, we give the probabilistic interpretation for our approach and apply it to explain the existing kernel learning methods, providing another perspective for their commonness and differences. Comprehensive experimental results on 13 UCI data sets and another two real-world data sets show that via the joint learning process, our approach not only adaptively identifies the pre-specified kernel, but also achieves superior classification performance to the original ONKL and the related MKL algorithms. Xinwang Liu 0002, Jianping Yin, Lei Wang 0001, Lingqiao Liu, Jun Liu 0003, Chenping Hou, Jian Zhang 0002 |
IEEE Trans. Cybern. | 3 |
| 2013 | On Similarity Preserving Feature SelectionabstractIn the literature of feature selection, different criteria have been proposed to evaluate the goodness of features. In our investigation, we notice that a number of existing selection criteria implicitly select features that preserve sample similarity, and can be unified under a common framework. We further point out that any feature selection criteria covered by this framework cannot handle redundant features, a common drawback of these criteria. Motivated by these observations, we propose a new "Similarity Preserving Feature Selection” framework in an explicit and rigorous way. We show, through theoretical analysis, that the proposed framework not only encompasses many widely used feature selection criteria, but also naturally overcomes their common weakness in handling feature redundancy. In developing this new framework, we begin with a conventional combinatorial optimization formulation for similarity preserving feature selection, then extend it with a sparse multiple-output regression formulation to improve its efficiency and effectiveness. A set of three algorithms are devised to efficiently solve the proposed formulations, each of which has its own advantages in terms of computational complexity and selection performance. As exhibited by our extensive experimental study, the proposed framework achieves superior feature selection performance and attractive properties. Zheng Zhao 0002, Lei Wang 0001, Huan Liu 0001, Jieping Ye |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | What has my classifier learned? Visualizing the classification rules of bag-of-feature model by support region detectionabstractIn the past decade, the bag-of-feature model has established itself as the state-of-the-art method in various visual classification tasks. Despite its simplicity and high performance, it normally works as a black box and the classification rule is not transparent to users. However, to better understand the classification process, it is favorable to look into the black box to see how an image is recognized. To fill this gap, we developed a tool called Restricted Support Region Set (RSRS) Detection which can be utilized to visualize the image regions that are critical to the classification decision. More specifically, we define the Restricted Support Region Set for a given image as such a set of size-restricted and non-overlapped regions that if any one of them is removed the image will be wrongly classified. Focusing on the state-of-the-art bag-of-feature classification system, we developed an efficient RSRS detection algorithm and discussed its applications. We showed that it can be used to identify the limitation of a classifier, predict its failure mode, discover the classification rules and reveal the database bias. Moreover, as experimentally demonstrated, this tool also enables common users to efficiently tune the classifier by removing the inappropriate support regions, which can lead to a better generalization performance. Lingqiao Liu, Lei Wang 0001 |
CVPR | 2 |
| 2012 | A Novel Video-Based Smoke Detection Method Using Image SeparationabstractIn the state-of-the-art video-based smoke detection methods, the representation of smoke mainly depends on the visual information in the current image frame. In the case of light smoke, the original background can be still seen and may deteriorate the characterization of smoke. The core idea of this paper is to demonstrate the superiority of using smoke component for smoke detection. In order to obtain smoke component, a blended image model is constructed, which basically is a linear combination of background and smoke components. Smoke opacity which represents a weighting of the smoke component is also defined. Based on this model, an optimization problem is posed. An algorithm is devised to solve for smoke opacity and smoke component, given an input image and the background. The resulting smoke opacity and smoke component are then used to perform the smoke detection task. The experimental results on both synthesized and real image data verify the effectiveness of the proposed method. Hongda Tian, Wanqing Li 0001, Lei Wang 0001, Philip Ogunbona |
ICME | 3 |
| 2012 | Incorporation of radius-info can be simple with SimpleMKL
Xinwang Liu 0002, Lei Wang 0001, Jianping Yin, Lingqiao Liu |
Neurocomputing | 2 |
| 2012 | Positive Semidefinite Metric Learning Using Boosting-like Algorithms
Chunhua Shen, Junae Kim, Lei Wang 0001, Anton van den Hengel |
J. Mach. Learn. Res. | 3 |
| 2011 | A generalized probabilistic framework for compact codebook creationabstractCompact and discriminative visual codebooks are preferred in many visual recognition tasks. In the literature, a few researchers have taken the approach of hierarchically merging visual words of a initial large-size code-book, but implemented this idea with different merging criteria. In this work, we show that by defining different class-conditional distribution function and parameter estimation method, these merging criteria can be unified under a single probabilistic framework. More importantly, by adopting new distribution functions and/or parameter estimation methods, we can generalize this framework to produce a spectrum of novel merging criteria. Two of them are particularly focused in this work. For one criterion, we adopt the multinomial distribution to model each object class, and for the other criterion we propose a max-margin-based parameter estimation method. Both theoretical analysis and experimental study demonstrate the superior performance of the two new merging criteria and the general applicability of our probabilistic framework. Lingqiao Liu, Lei Wang 0001, Chunhua Shen |
CVPR | 2 |
| 2011 | A scalable dual approach to semidefinite metric learningabstractDistance metric learning plays an important role in many vision problems. Previous work of quadratic Mahalanobis metric learning usually needs to solve a semidefinite programming (SDP) problem. A standard interior-point SDP solver has a complexity of O(D6.5) (with D the dimension of input data), and can only solve problems up to a few thousand variables. Since the number of variables is D(D + l)/2, this corresponds to a limit around D3) with t ≈ 20 ~ 30 for most problems in our experiments. The proposed algorithm is scalable and easy to implement. Experiments on various datasets show its similar accuracy compared with state-of-the-art. We also demonstrate that this idea may also be able to be applied to other SDP problems such as maximum variance unfolding. Chunhua Shen, Junae Kim, Lei Wang 0001 |
CVPR | 3 |
| 2011 | In defense of soft-assignment codingabstractIn object recognition, soft-assignment coding enjoys computational efficiency and conceptual simplicity. However, its classification performance is inferior to the newly developed sparse or local coding schemes. It would be highly desirable if its classification performance could become comparable to the state-of-the-art, leading to a coding scheme which perfectly combines computational efficiency and classification performance. To achieve this, we revisit soft-assignment coding from two key aspects: classification performance and probabilistic interpretation. For the first aspect, we argue that the inferiority of soft-assignment coding is due to its neglect of the underlying manifold structure of local features. To remedy this, we propose a simple modification to localize the soft-assignment coding, which surprisingly achieves comparable or even better performance than existing sparse or local coding schemes while maintaining its computational advantage. For the second aspect, based on our probabilistic interpretation of the soft-assignment coding, we give a probabilistic explanation to the magic max-pooling operation, which has successfully been used by sparse or local coding schemes but still poorly understood. This probability explanation motivates us to develop a new mix-order max-pooling operation which further improves the classification performance of the proposed coding scheme. As experimentally demonstrated, the localized soft-assignment coding achieves the state-of-the-art classification performance with the highest computational efficiency among the existing coding schemes. Lingqiao Liu, Lei Wang 0001, Xinwang Liu 0002 |
ICCV | 2 |
| 2011 | The Effect of the Characteristics of the Dataset on the Selection StabilityabstractFeature selection is an effective technique to reduce the dimensionality of a data set and to select relevant features for the domain problem. Recently, stability of feature selection methods has gained increasing attention. In fact, it has become a crucial factor in determining the goodness of a feature selection algorithm besides the learning performance. In this work, we conduct an extensive experimental study using verity of data sets and different well-known feature selection algorithms in order to study the behavior of these algorithms in terms of the stability. Salem Alelyani, Huan Liu 0001, Lei Wang 0001 |
ICTAI | 3 |
| 2011 | Exploring latent class information for image retrieval using the bag-of-feature modelabstractRecently, the Bag-of-Feature (BoF) model has shown promising performance in object and generic image retrieval. The similarity between two images is typically measured by the distance between the two histograms. Due to the imperfection of local descriptor and quantization error, visually similar image patches can be wrongly quantized into different visual words, making this distance-based measure less accurate. To address this issue, this paper explores the information of latent class, which is formed by all the database images that share the same visual concept with the one being compared to a given query. We then cast image similarity as the probability of the query and a database image belonging to a same latent class. Considering that a class of images together can better depict a visual concept, the shift from image-to-image to image-to-class comparison is expected to bring a more robust similarity measure. Because the ground truth of the latent class is not accessible in image retrieval, we define a latent class prior in our probabilistic model and derive its marginal distribution. This gives rise to a novel and efficient image similarity measure. It can significantly improve retrieval performance without prolonging retrieval process. Experimental study on multiple benchmark data sets demonstrates its advantages. Lingqiao Liu, Lei Wang 0001 |
ACM Multimedia | 2 |
| 2010 | Efficient Spectral Feature Selection with Minimum RedundancyabstractSpectral feature selection identifies relevant features by measuring their capability of preserving sample similarity. It provides a powerful framework for both supervised and unsupervised feature selection, and has been proven to be effective in many real-world applications. One common drawback associated with most existing spectral feature selection algorithms is that they evaluate features individually and cannot identify redundant features. Since redundant features can have significant adverse effect on learning performance, it is necessary to address this limitation for spectral feature selection. To this end, we propose a novel spectral feature selection algorithm to handle feature redundancy, adopting an embedded model. The algorithm is derived from a formulation based on a sparse multi-output regression with a L2,1-norm constraint. We conduct theoretical analysis on the properties of its optimal solutions, paving the way for designing an efficient path-following solver. Extensive experiments show that the proposed algorithm can do well in both selecting relevant features and removing redundancy. Zheng Zhao 0002, Lei Wang 0001, Huan Liu 0001 |
AAAI | 2 |
| 2010 | Efficient Structured Support Vector Regression
Ke Jia, Lei Wang 0001, Nianjun Liu |
ACCV (3) | 2 |
| 2010 | Gradual Sampling and Mutual Information Maximisation for Markerless Motion Capture
Lei Wang 0001, Richard I. Hartley, Hongdong Li, Dan Xu 0001 |
ACCV (2) | 2 |
| 2010 | Compressive Evaluation in Human Motion Tracking
Lei Wang 0001, Richard I. Hartley, Hongdong Li, Dan Xu 0001 |
ACCV (4) | 2 |
| 2010 | Fast Multi-labelling for Stereo Matching
Yuhang Zhang 0001, Richard I. Hartley, Lei Wang 0001 |
ECCV (3) | 3 |
| 2010 | Efficient Learning to Label ImagesabstractConditional random field methods (CRFs) have gained popularity for image labeling tasks in recent years. In this paper, we describe an alternative discriminative approach, by extending the large margin principle to incorporate spatial correlations among neighboring pixels. In particular, by explicitly enforcing the sub modular condition, graph-cuts is conveniently integrated as the inference engine to attain the optimal label assignment efficiently. Our approach allows learning a model with thousands of parameters, and is shown to be capable of readily incorporating higher-order scene context. Empirical studies on a variety of image datasets suggest that our approach performs competitively compared to the state-of-the-art scene labeling methods. Ke Jia, Li Cheng 0001, Nianjun Liu, Lei Wang 0001 |
ICPR | 4 |
| 2010 | Inter-Domain Routing Validator Based Spoofing Defence SystemabstractIP spoofing remains a problem today in the Internet. In this paper, a new system called Inter-Domain Routing Validator Based Spoofing Defence System (SDS) for filtering spoofed IP packets is proposed. SDS uses efficient symmetric key message authentication code (UMAC) as its tag to verify that a source IP address is valid. Different ASes border routers obtain a shared key via the Inter-Domain Routing Validator (IRV) servers which will manage the secret keys and exchange keys among different ASes via security communication channel. SDS is efficient, secure and easy to cooperate with other defence mechanisms. Lei Wang 0001, Tianbing Xia, Jennifer Seberry |
ISI | 1 |
| 2010 | Hippocampal Shape Classification Using Redundancy Constrained Feature Selection
Luping Zhou, Lei Wang 0001, Chunhua Shen, Nick Barnes |
MICCAI (2) | 2 |
| 2010 | Scalable large-margin Mahalanobis distance metric learningabstractFor many machine learning algorithms such as k-nearest neighbor ( k-NN) classifiers and k-means clustering, often their success heavily depends on the metric used to calculate distances between different data points. An effective solution for defining such a metric is to learn it from a set of labeled training samples. In this work, we propose a fast and scalable algorithm to learn a Mahalanobis distance metric. The Mahalanobis metric can be viewed as the Euclidean distance metric on the input data that have been linearly transformed. By employing the principle of margin maximization to achieve better generalization performances, this algorithm formulates the metric learning as a convex optimization problem and a positive semidefinite (p.s.d.) matrix is the unknown variable. Based on an important theorem that a p.s.d. trace-one matrix can always be represented as a convex combination of multiple rank-one matrices, our algorithm accommodates any differentiable loss function and solves the resulting optimization problem using a specialized gradient descent procedure. During the course of optimization, the proposed algorithm maintains the positive semidefiniteness of the matrix variable that is essential for a Mahalanobis metric. Compared with conventional methods like standard interior-point algorithms or the special solver used in large margin nearest neighbor , our algorithm is much more efficient and has a better performance in scalability. Experiments on benchmark data sets suggest that, compared with state-of-the-art metric learning algorithms, our algorithm can achieve a comparable classification accuracy with reduced computational complexity. Chunhua Shen, Junae Kim, Lei Wang 0001 |
IEEE Trans. Neural Networks | 3 |
| 2010 | Feature selection with redundancy-constrained class separabilityabstractScatter-matrix-based class separability is a simple and efficient feature selection criterion in the literature. However, the conventional trace-based formulation does not take feature redundancy into account and is prone to selecting a set of discriminative but mutually redundant features. In this brief, we first theoretically prove that in the context of this trace-based criterion the existence of sufficiently correlated features can always prevent selecting the optimal feature set. Then, on top of this criterion, we propose the redundancy-constrained feature selection (RCFS). To ensure the algorithm's efficiency and scalability, we study the characteristic of the constraints with which the resulted constrained 0-1 optimization can be efficiently and globally solved. By using the totally unimodular (TUM) concept in integer programming, a necessary condition for such constraints is derived. This condition reveals an interesting special case in which qualified redundancy constraints can be conveniently generated via a clustering of features. We study this special case and develop an efficient feature selection approach based on Dinkelbach's algorithm. Experiments on benchmark data sets demonstrate the superior performance of our approach to those without redundancy constraints. Luping Zhou, Lei Wang 0001, Chunhua Shen |
IEEE Trans. Neural Networks | 2 |
| 2009 | A Scalable Algorithm for Learning a Mahalanobis Distance Metric
Junae Kim, Chunhua Shen, Lei Wang 0001 |
ACCV (3) | 3 |
| 2009 | Positive Semidefinite Metric Learning with BoostingabstractThe learning of appropriate distance metrics is a critical problem in classification. In this work, we propose a boosting-based technique, termed BoostMetric, for learning a Mahalanobis distance metric. One of the primary difficulties in learning such a metric is to ensure that the Mahalanobis matrix remains positive semidefinite. Semidefinite programming is sometimes used to enforce this constraint, but does not scale well. BoostMetric is instead based on a key observation that any positive semidefinite matrix can be decomposed into a linear positive combination of trace-one rank-one matrices. BoostMetric thus uses rank-one positive semidefinite matrices as weak learners within an efficient and scalable boosting-based learning process. The resulting method is easy to implement, does not require tuning, and can accommodate various types of constraints. Experiments on various datasets show that the proposed algorithm compares favorably to those state-of-the-art methods in terms of classification accuracy and running time. Chunhua Shen, Junae Kim, Lei Wang 0001, Anton van den Hengel |
NIPS | 3 |
| 2009 | Identifying Anatomical Shape Difference by Regularized Discriminative DirectionabstractIdentifying the shape difference between two groups of anatomical objects is important for medical image analysis and computer-aided diagnosis. A method called "discriminative direction" in the literature has been proposed to solve this problem. In that method, the shape difference between groups is identified by deforming a shape along the discriminative direction. This paper conducts a thorough study about inferring this discriminative direction in an efficient and accurate way. First, finding the discriminative direction is reformulated as a preimage problem in kernel-based learning. This provides a complementary but conceptually simpler solution than the previous method. More importantly, we find that a shape deforming along the original discriminative direction cannot faithfully maintain its anatomical correctness. This unnecessarily introduces spurious shape differences and leads to inaccurate analysis. To overcome this problem, this paper further proposes a regularized discriminative direction by requiring a shape to conform to its underlying distribution when it deforms. Two different approaches are developed to impose the regularization, one from the perspective of probability distributions and the other from a geometric point of view, and their relationship is discussed. After verifying their superior performance through controlled experiments, we apply the proposed methods to detecting and localizing the hippocampal shape difference between sexes. We get results consistent with other independent research, providing a more compact representation of the shape difference compared with the established discriminative direction method. Luping Zhou, Richard I. Hartley, Lei Wang 0001, Paulette Lieby, Nick Barnes |
IEEE Trans. Medical Imaging | 3 |
| 2008 | A Fast Algorithm for Creating a Compact and Discriminative Visual Codebook
Lei Wang 0001, Luping Zhou, Chunhua Shen |
ECCV (4) | 1 |
| 2008 | Regularized Discriminative Direction for Shape Difference Analysis
Luping Zhou, Richard I. Hartley, Lei Wang 0001, Paulette Lieby, Nick Barnes |
MICCAI (1) | 3 |
| 2008 | PSDBoost: Matrix-Generation Linear Programming for Positive Semidefinite Matrices LearningabstractIn this work, we consider the problem of learning a positive semidefinite matrix. The critical issue is how to preserve positive semidefiniteness during the course of learning. Our algorithm is mainly inspired by LPBoost [1] and the general greedy convex optimization framework of Zhang [2]. We demonstrate the essence of the algorithm, termed PSDBoost (positive semidefinite Boosting), by focusing on a few different applications in machine learning. The proposed PSDBoost algorithm extends traditional Boosting algorithms in that its parameter is a positive semidefinite matrix with trace being one instead of a classifier. PSDBoost is based on the observation that any trace-one positive semidefinitematrix can be decomposed into linear convex combinations of trace-one rank-one matrices, which serve as base learners of PSDBoost. Numerical experiments are presented. Chunhua Shen, Alan H. Welsh, Lei Wang 0001 |
NIPS | 3 |
| 2008 | AdaBoost with SVM-based component classifiers
Xuchun Li, Lei Wang 0001, Eric Sung |
Eng. Appl. Artif. Intell. | 2 |
| 2008 | Feature Selection with Kernel Class SeparabilityabstractClassification can often benefit from efficient feature selection. However, the presence of linearly nonseparable data, quick response requirement, small sample problem and noisy features makes the feature selection quite challenging. In this work, a class separability criterion is developed in a high-dimensional kernel space, and feature selection is performed by the maximization of this criterion. To make this feature selection approach work, the issues of automatic kernel parameter tuning, the numerical stability, and the regularization for multi-parameter optimization are addressed. Theoretical analysis uncovers the relationship of this criterion to the radius-margin bound of the SVMs, the KFDA, and the kernel alignment criterion, providing more insight on using this criterion for feature selection. This criterion is applied to a variety of selection modes with different search strategies. Extensive experimental study demonstrates its efficiency in delivering fast and robust feature selection. Lei Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | A Kernel-Induced Space Selection Approach to Model Selection in KLDAabstractModel selection in kernel linear discriminant analysis (KLDA) refers to the selection of appropriate parameters of a kernel function and the regularizer. By following the principle of maximum information preservation, this paper formulates the model selection problem as a problem of selecting an optimal kernel-induced space in which different classes are maximally separated from each other. A scatter-matrix-based criterion is developed to measure the "goodness" of a kernel-induced space, and the kernel parameters are tuned by maximizing this criterion. This criterion is computationally efficient and is differentiable with respect to the kernel parameters. Compared with the leave-one-out (LOO) or k-fold cross validation (CV), the proposed approach can achieve a faster model selection, especially when the number of training samples is large or when many kernel parameters need to be tuned. To tune the regularization parameter in the KLDA, our criterion is used together with the method proposed by Saadi (2004). Experiments on benchmark data sets verify the effectiveness of this model selection approach. Lei Wang 0001, Kap Luk Chan, Ping Xue 0001, Luping Zhou |
IEEE Trans. Neural Networks | 1 |
| 2008 | Two Criteria for Model Selection in Multiclass Support Vector MachinesabstractPractical applications call for efficient model selection criteria for multiclass support vector machine (SVM) classification. To solve this problem, this paper develops two model selection criteria by combining or redefining the radius-margin bound used in binary SVMs. The combination is justified by linking the test error rate of a multiclass SVM with that of a set of binary SVMs. The redefinition, which is relatively heuristic, is inspired by the conceptual relationship between the radius-margin bound and the class separability measure. Hence, the two criteria are developed from the perspective of model selection rather than a generalization of the radius-margin bound for multiclass SVMs. As demonstrated by extensive experimental study, the minimization of these two criteria achieves good model selection on most data sets. Compared with the k-fold cross validation which is often regarded as a benchmark, these two criteria give rise to comparable performance with much less computational overhead, particularly when a large number of model parameters are to be optimized. Lei Wang 0001, Ping Xue 0001, Kap Luk Chan |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2007 | Feature Subset Selection for Multi-class SVM Based Image Classification
Lei Wang 0001 |
ACCV (2) | 1 |
| 2007 | Where's the Weet-Bix?
Yuhang Zhang 0001, Lei Wang 0001, Richard I. Hartley, Hongdong Li |
ACCV (1) | 2 |
| 2007 | Toward A Discriminative Codebook: Codeword Selection across Multi-resolutionabstractIn patch-based object recognition, there are two important issues on the codebook generation: (I) resolution: a coarse codebook lacks sufficient discriminative power, and an over-fine one is sensitive to noise; (2) codeword selection: non-discriminative codewords not only increase the codebook size, but also can hurt the recognition performance. To achieve a discriminative codebook for better recognition, this paper argues that these two issues are strongly related and should be solved as a whole. In this paper, a multi-resolution codebook is first designed via hierarchical clustering. With a reasonable size, it includes all of the codewords which cross a large number of resolution levels. More importantly, it forms a diverse candidate codeword set that is critical to codeword selection. A Boosting feature selection approach is modified to select the discriminative codewords from this multi-resolution code-book. By doing so, the obtained codebook is composed of the most discriminative codewords culled from different levels of resolution. Experimental study demonstrates the better recognition performance attained by this codebook. Lei Wang 0001 |
CVPR | 1 |
| 2006 | A novel multi-resolution video representation scheme based on kernel PCA
Xiao-Dong Yu, Lei Wang 0001, Qi Tian 0002, Ping Xue 0001 |
Vis. Comput. | 2 |
| 2005 | A Framework of 2D Fisher Discriminant Analysis: Application to Face Recognition with Small Number of Training SamplesabstractA novel framework called 2D Fisher discriminant analysis (2D-FDA) is proposed to deal with the small sample size (SSS) problem in conventional one-dimensional linear discriminant analysis (1D-LDA). Different from the 1D-LDA based approaches, 2D-FDA is based on 2D image matrices rather than column vectors so the image matrix does not need to be transformed into a long vector before feature extraction. The advantage arising in this way is that the SSS problem does not exist any more because the between-class and within-class scatter matrices constructed in 2D-FDA are both of full-rank. This framework contains unilateral and bilateral 2D-FDA. It is applied to face recognition where only few training images exist for each subject. Both the unilateral and bilateral 2D-FDA achieve excellent performance on two public databases: ORL database and Yale face database B. Hui Kong 0001, Lei Wang 0001, Eam Khwang Teoh, Jian-Gang Wang 0001, Ronda Venkateswarlu |
CVPR (2) | 2 |
| 2005 | Retrieval with Knowledge-driven Kernel Design: An Approach to Improving SVM-Based CBIR with Relevance FeedbackabstractThe performance of SVM-based image retrieval is often constrained by the scarcity of training samples. The total number of image samples labeled by users in a retrieval session is very limited, and this small number of labeled samples cannot effectively represent the true distributions of positive and negative image classes, especially for the negative image class. This paper proposes a novel approach to deal with this problem. Instead of treating it as a problem, the mere existence of the small number of labeled images and their desired distribution in the kernel space is considered as prior knowledge from image retrieval to aid the design of the kernel used by SVMs. This is achieved by maximizing a criterion, such as one based on scatter matrices, through gradient-based search methods, incurring very little computational overhead to real-time retrieval process. Experimental results on two benchmark image databases demonstrate the improved retrieval performance by the dynamically designed kernel and hence the effectiveness of the proposed approach for SVM based image retrieval Lei Wang 0001, Kap Luk Chan, Ping Xue 0001, Weiyun Yau |
ICCV | 1 |
| 2005 | Generalized 2D principal component analysisabstractA two-dimensional principal component analysis (2DPCA) by J. Yang et al. (2004) was proposed and the authors have demonstrated its superiority over the conventional principal component analysis (PCA) in face recognition. But the theoretical proof why 2DPCA is better than PCA has not been given until now. In this paper, the essence of 2DPCA is analyzed and a framework of generalized 2D principal component analysis (G2DPCA) is proposed to extend the original 2DPCA in two perspectives: a bilateral-projection-based 2DPCA (B2DPCA) and a kernel-based 2DPCA (K2DPCA) schemes are introduced. Experimental results in face recognition show its excellent performance. Hui Kong 0001, Xuchun Li, Lei Wang 0001, Earn Khwang Teoh, Ronda Venkateswarlu |
IJCNN | 3 |
| 2005 | A study of AdaBoost with SVM based weak learnersabstractIn this article, we focus on designing an algorithm, named AdaBoostSVM, using SVM as weak learners for AdaBoost. To obtain a set of effective SVM weak learners, this algorithm adaptively adjusts the kernel parameter in SVM instead of using a fixed one. Compared with the existing AdaBoost methods, the AdaBoostSVM has advantages of easier model selection and better generalization performance. It also provides a possible way to handle the over-fitting problem in AdaBoost. An improved version called Diverse AdaBoostSVM is further developed to deal with the accuracy/diversity dilemma in Boosting methods. By implementing some parameter adjusting strategies, the distributions of accuracy and diversity over these SVM weak learners are tuned to achieve a good balance. To the best of our knowledge, such a mechanism that can conveniently and explicitly balances this dilemma has not been seen in the literature. Experimental results demonstrated that both proposed algorithms achieve better generalization performance than AdaBoost using other kinds of weak learners. Benefiting from the balance between accuracy and diversity, the Diverse AdaBoostSVM achieves the best performance. In addition, the experiments on unbalanced data sets showed that the AdaBoostSVM performed much better than SVM. Xuchun Li, Lei Wang 0001, Eric Sung |
IJCNN | 2 |
| 2005 | A novel framework for SVM-based image retrieval on large databasesabstractIn this paper, a novel framework is proposed to deliver a fast, robust, and generally applicable SVM-based image retrieval for large databases. A quick test scheme is developed, and on-line kernel learning is employed to realize it after analyzing the relationship between them. Then an upper bound on maximum test scope is derived to speed up testing further. Also, the general applicability is well maintained because this framework does not need a kernel function and index structure to be pre-defined. Taking the advantages of this framework, more sophisticated SVM can be used to improve retrieval performance while keeping short response time. Experimental results on large image databases verify the effectiveness and efficiency of the proposed framework. Lei Wang 0001, Xuchun Li, Ping Xue 0001, Kap Luk Chan |
ACM Multimedia | 1 |
| 2005 | Generalized 2D principal component analysis for face image representation and recognition
Hui Kong 0001, Lei Wang 0001, Eam Khwang Teoh, Xuchun Li, Jian-Gang Wang 0001, Ronda Venkateswarlu |
Neural Networks | 2 |
| 2005 | A criterion for optimizing kernel parameters in KBDA for image retrievalabstractA criterion is proposed to optimize the kernel parameters in Kernel-based Biased Discriminant Analysis (KBDA) for image retrieval. Kernel parameter optimization is performed by optimizing the kernel space such that the positive images are well clustered while the negative ones are pushed far away from the positives. The proposed criterion measures the goodness of a kernel space, and the optimal kernel parameter set is obtained by maximizing this criterion. Retrieval experiments on two benchmark image databases demonstrate the effectiveness of proposed criterion for KBDA to achieve the best possible performance at the cost of a small fractional computational overhead. Lei Wang 0001, Kap Luk Chan, Ping Xue 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2004 | Multi-label SVM active learning for image classificationabstractImage classification is an important task in computer vision. However, how to assign suitable labels to images is a subjective matter, especially when some images can be categorized into multiple classes simultaneously. Multilabel image classification focuses on the problem that each image can have one or multiple labels. It is known that manually labelling images is time-consuming and expensive. In order to reduce the human effort of labelling images, especially multilabel images, we proposed a multilabel SVM active learning method. We also proposed two selection strategies: Max Loss strategy and Mean Max Loss strategy. Experimental results on both artificial data and real-world images demonstrated the advantage of proposed method. Xuchun Li, Lei Wang 0001, Eric Sung |
ICIP | 2 |
| 2004 | Robust multi-level video representation using mean shift analysisabstractA robust method for multi-level video representation based on the mean shift analysis (MSA) of low-level visual features is proposed in this paper. By tuning the bandwidth of MSA, video representation from the coarse level to the fine level can be achieved. This representation form provides a flexible scheme for content-based video analysis such as summarization, classification, and retrieval. Compared with the conventional k-means or fuzzy c-means algorithms, our method can adjust the resolution of representation in a more straightforward way, and is more robust since it does not need to initialize the cluster centers Hai Gao, Xiao-Dong Yu, Lei Wang 0001, Ping Xue 0001, Qi Tian 0002 |
ICME | 3 |
| 2004 | Multi-Level Video Representation with Application to Keyframe ExtractionabstractContent-based video analysis calls for efficient video representation. In this paper, a novel multi-level representation of video is proposed based on the principle components derived from low-level visual features. It can characterize the video content from the coarse level to the fine level according to its intrinsic structure. This representation form provides a flexible scheme for video content analysis such as summarization, classification, and retrieval. A newly proposed subspace method, kernel based PCA, is explored to achieve this conveniently. The application in keyframe extraction is investigated to demonstrate the benefits of this representation. Xiao-Dong Yu, Lei Wang 0001, Qi Tian 0002, Ping Xue 0001 |
MMM | 2 |
| 2003 | Bootstrapping SVM Active Learning by Incorporating Unlabelled Images for Image RetrievalabstractThe performance of image retrieval with SVM active learning is known to be poor when started with few labeled images only. In this paper, the problem is solved by incorporating the unlabelled images into the bootstrapping of the learning process. In this work, the initial SVM classifier is trained with the few labeled images and the unlabelled images randomly selected from the image database. Both theoretical analysis and experimental results show that by incorporating unlabelled images in the bootstrapping, the efficiency of SVM active learning can be improved, and thus improves the overall retrieval performance. Lei Wang 0001, Kap Luk Chan |
CVPR (1) | 1 |
| 2003 | Image retrieval with SVM active learning embedding Euclidean searchabstractImage retrieval with relevance feedback suffers from the small sample problem. Recently, SVM active learning has been proposed to tackle this problem, showing promising results. However, a small but sufficient number of initially labelled samples are still required to ensure subsequent efficient active learning and good retrieval performance. In the existing method, the user is asked to label more images before active learning starts. In this paper, a method of embedding Euclidean search into SVM active learning is proposed. With the help of Euclidean search, the adverse effect on retrieval performance due to lack of initially labelled samples can be reduced. Experimental results demonstrate the improvement by the proposed method, especially when the number of initially labelled samples is small. Lei Wang 0001, Kap Luk Chan, Yap-Peng Tan |
ICIP (1) | 1 |
| 2003 | A Dynamic Sub-vector Weighting Scheme for Image Retrieval with Relevance Feedback
Lei Wang 0001, Kap Luk Chan |
Pattern Anal. Appl. | 1 |
| 2000 | Bayesian Learning for Image Retrieval Using Multiple Features
Lei Wang 0001, Kap Luk Chan |
IDEAL | 1 |