Lei Huang 0015

dblp:18/1763-15 · DBLP profile ↗
← Back
56ranked-venue papers
17as first author
29since 2021 · last 2026
0000-0003-0502-168XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 14 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 11 first-author · 11 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
Zonghao Ying, Aishan Liu, Siyuan Liang 0004, Lei Huang 0015, Jinyang Guo 0002, Wenbo Zhou 0004, Xianglong Liu 0001, Dacheng Tao
Int. J. Comput. Vis.4
2025 Clustering Properties of Self-Supervised Learning
abstract
Self-supervised learning (SSL) methods via joint embedding architectures have proven remarkably effective at capturing semantically rich representations with strong clustering properties, magically in the absence of label supervision. Despite this, few of them have explored leveraging these untapped properties to improve themselves. In this paper, we provide an evidence through various metrics that the encoder's output *encoding* exhibits superior and more stable clustering properties compared to other components. Building on this insight, we propose a novel positive-feedback SSL method, termed **Re**presentation **S**elf-**A**ssignment (ReSA), which leverages the model's clustering properties to promote learning in a self-guided manner. Extensive experiments on standard SSL benchmarks reveal that models pretrained with ReSA outperform other state-of-the-art SSL methods by a significant margin. Finally, we analyze how ReSA facilitates better clustering properties, demonstrating that it effectively enhances clustering performance at both fine-grained and coarse-grained levels, shaping representations that are inherently more structured and semantically meaningful.
Xi Weng, Jianing An, Xudong Ma, Binhang Qi, Jie Luo 0004, Jin Song Dong 0001, Lei Huang 0015
ICML8
2025 Verification of Bit-Flip Attacks against Quantized Neural Networks
abstract
In the rapidly evolving landscape of neural network security, the resilience of neural networks against bit-flip attacks (i.e., an attacker maliciously flips an extremely small amount of bits within its parameter storage memory system to induce harmful behavior), has emerged as a relevant area of research. Existing studies suggest that quantization may serve as a viable defense against such attacks. Recognizing the documented susceptibility of real-valued neural networks to such attacks and the comparative robustness of quantized neural networks (QNNs), in this work, we introduce BFAVerifier, the first verification framework designed to formally verify the absence of bit-flip attacks against QNNs or to identify all vulnerable parameters in a sound and rigorous manner. BFAVerifier comprises two integral components: an abstraction-based method and an MILP-based method. Specifically, we first conduct a reachability analysis with respect to symbolic parameters that represent the potential bit-flip attacks, based on a novel abstract domain with a sound guarantee. If the reachability analysis fails to prove the resilience of such attacks, then we encode this verification problem into an equivalent MILP problem which can be solved by off-the-shelf solvers. Therefore, BFAVerifier is sound, complete, and reasonably efficient. We conduct extensive experiments, which demonstrate its effectiveness and efficiency across various activation functions, quantization bit-widths, and adversary capabilities.
Yedi Zhang, Lei Huang 0015, Fu Song, Jun Sun 0001, Jin Song Dong 0001
Proc. ACM Program. Lang.2
2025 Investigating Synthetic-to-Real Transfer Robustness for Stereo Matching and Optical Flow Estimation
abstract
With advancements in robust stereo matching and optical flow estimation networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, their robustness can be seriously degraded when fine-tuning them in real-world scenarios. This paper investigates fine-tuning stereo matching and optical flow estimation networks without compromising their robustness to unseen domains. Specifically, we divide the pixels into consistent and inconsistent regions by comparing Ground Truth (GT) with Pseudo Label (PL) and demonstrate that the imbalance learning of consistent and inconsistent regions in GT causes robustness degradation. Based on our analysis, we propose the DKT framework, which utilizes PL to balance the learning of different regions in GT. The core idea is to utilize an exponential moving average (EMA) teacher to measure what the student network has learned and dynamically adjust the learning regions. We further propose the DKT++ framework, which improves target-domain performances and network robustness by applying slow-fast update teachers to generate more accurate PL, introducing the unlabeled data and synthetic data. We integrate our frameworks with state-of-the-art networks and evaluate their effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the robustness of stereo matching and optical flow networks during fine-tuning.
Jiahe Li 0007, Lei Huang 0015, Haonan Luo 0002, Xiaohan Yu 0001, Lin Gu 0003, Xiao Bai 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Dual-mask: Progressively sparse multi-task architecture learning
abstract
Multi-task architecture learning has achieved significant success by learning optimal sharing architectures for different tasks. However, previous works to learn branched architectures for different tasks can sometimes lead to unsatisfying multi-task performance, as not all detailed branches are relevant to a specific task. Task-relevant architectures can be sparse, including only partial channels or layers in the entire architecture (i.e., a sub-network). In addition, most previous works rely on a heuristic architecture selection procedure that could not support continuous architecture optimization. To this end, in this paper, we propose dual-mask , a progressively sparse multi-task architecture learning method. Starting with a task-free architecture, it identifies the informative features along two-level, channels and layers, for each task, while suppressing conflicting or noisy parts in a differentiable manner, so that better task-specific sub-networks are captured. Specifically, the channel and layer selection modules produce respective hybrid binary and real value masks, designed to pick salient channels and layers for each task, respectively. To jointly optimize masks with model parameters, we propose an importance-guided relaxation method for solving the stochastic binary optimization problem , after which the interference or noise parts can be pruned by masks. Additionally, a progressive training strategy with continuation is provided that gradually sparsity the task-specific sub-networks. Experiments show that dual-mask achieves superior performance than SOTA multi-task methods.
Jiejie Zhao, Tongyu Zhu, Leilei Sun, Bowen Du 0001, Lei Huang 0015
Pattern Recognit.6
2025 Understanding the Dimensional Need of Noncontrastive Learning
abstract
Noncontrastive self-supervised learning methods offer an effective alternative to contrastive approaches by avoiding the need for negative samples to avoid representation collapse. Noncontrastive learning methods explicitly or implicitly optimize the representation space, yet they often require large representation dimensions, leading to dimensional inefficiency. To provide negative samples, contrastive learning methods often require large batch sizes, thus regarded as sample inefficient, while noncontrastive learning methods require large representation dimensions, thus regarded as dimension inefficient. Although we have some understanding of the noncontrastive learning method, theoretical analysis of such phenomenon still remains largely unexplored. We present a theoretical analysis of the dimensional need for noncontrastive learning. We investigate the transfer between upstream representation learning and downstream tasks' performance, demonstrating how noncontrastive methods implicitly increase interclass distances within the representation space and how the distance affects the model performance of evaluation performance. We prove that the performance of noncontrastive methods is affected by the output dimension and the number of latent classes, and illustrate why performance degrades significantly when the output dimension is substantially smaller than the number of latent classes. We demonstrate our findings through experiments on image classification experiments, and enrich the verification in audio, graph and text modalities. We also perform empirical evaluation for image models on extensive detection and segmentation tasks beyond classification that show satisfactory correspondence to our theorem.
Zhexiao Cao, Lei Huang 0015, Tian Wang 0002, Yinquan Wang, Jingang Shi, Aichun Zhu, Tianyun Shi, Hichem Snoussi
IEEE Trans. Cybern.2
2024 Robust Synthetic-to-Real Transfer for Stereo Matching
abstract
With advancements in domain generalized stereo matching networks, models pre-trained on synthetic data demonstrate strong robustness to unseen domains. However, few studies have investigated the robustness after fine-tuning them in real-world scenarios, during which the domain generalization ability can be seriously degraded. In this paper, we explore fine-tuning stereo matching networks without compromising their robustness to unseen domains. Our motivation stems from comparing Ground Truth (GT) versus Pseudo Label (PL) for fine-tuning: GT degrades, but PL preserves the domain generalization ability. Empirically, we find the difference between GT and PL implies valuable information that can regularize networks during fine-tuning. We also propose a framework to utilize this difference for fine-tuning, consisting of a frozen Teacher, an exponential moving average (EMA) Teacher, and a Student network. The core idea is to utilize the EMA Teacher to measure what the Student has learned and dynamically improve GT and PL for fine-tuning. We integrate our framework with state-of-the-art networks and evaluate its effectiveness on several real-world datasets. Extensive experiments show that our method effectively preserves the domain generalization ability during fine-tuning. Code is available at: https://github.com/jiaw-z/DKT-Stereo.
Jiahe Li 0007, Lei Huang 0015, Xiaohan Yu 0001, Lin Gu 0003, Xiao Bai 0001
CVPR3
2024 CoR-GS: Sparse-View 3D Gaussian Splatting via Co-regularization
Jiahe Li 0007, Xiaohan Yu 0001, Lei Huang 0015, Lin Gu 0003, Xiao Bai 0001
ECCV (1)4
2024 Modulate Your Spectrum in Self-Supervised Learning
abstract
Whitening loss offers a theoretical guarantee against feature collapse in self-supervised learning (SSL) with joint embedding architectures. Typically, it involves a hard whitening approach, transforming the embedding and applying loss to the whitened output. In this work, we introduce Spectral Transformation (ST), a framework to modulate the spectrum of embedding and to seek for functions beyond whitening that can avoid dimensional collapse. We show that whitening is a special instance of ST by definition, and our empirical investigations unveil other ST instances capable of preventing collapse. Additionally, we propose a novel ST instance named IterNorm with trace loss (INTL). Theoretical analysis confirms INTL's efficacy in preventing collapse and modulating the spectrum of embedding toward equal-eigenvalues during optimization. Our experiments on ImageNet classification and COCO object detection demonstrate INTL's potential in learning superior representations. The code is available at https://github.com/winci-ai/INTL.
Xi Weng, Yunhao Ni, Tengwei Song, Jie Luo 0004, Rao Muhammad Anwer, Salman Khan 0001, Fahad Shahbaz Khan, Lei Huang 0015
ICLR8
2024 On the Nonlinearity of Layer Normalization
abstract
Layer normalization (LN) is a ubiquitous technique in deep learning but our theoretical understanding to it remains elusive. This paper investigates a new theoretical direction for LN, regarding to its nonlinearity and representation capacity. We investigate the representation capacity of a network with layerwise composition of linear and LN transformations, referred to as LN-Net. We theoretically show that, given $m$ samples with any label assignment, an LN-Net with only 3 neurons in each layer and $O(m)$ LN layers can correctly classify them. We further show the lower bound of the VC dimension of an LN-Net. The nonlinearity of LN can be amplified by group partition, which is also theoretically demonstrated with mild assumption and empirically supported by our experiments. Based on our analyses, we consider to design neural architecture by exploiting and amplifying the nonlinearity of LN, and the effectiveness is supported by our experiments.
Yunhao Ni, Junlong Jia, Lei Huang 0015
ICML4
2024 Towards Defending Multiple ℓ p-Norm Bounded Adversarial Perturbations via Gated Batch Normalization
Aishan Liu, Shiyu Tang, Lei Huang 0015, Haotong Qin, Xianglong Liu 0001, Dacheng Tao
Int. J. Comput. Vis.4
2024 Exploring the Usage of Pre-trained Features for Stereo Matching
Lei Huang 0015, Xiao Bai 0001, Lin Gu 0003, Edwin R. Hancock
Int. J. Comput. Vis.2
2024 Understanding Whitening Loss in Self-Supervised Learning
abstract
A desirable objective in self-supervised learning (SSL) is to avoid feature collapse. Whitening loss guarantees collapse avoidance by minimizing the distance between embeddings of positive pairs under the conditioning that the embeddings from different views are whitened. In this paper, we propose a framework with an informative indicator to analyze whitening loss, which provides a clue to demystify several interesting phenomena and a pivoting point connecting to other SSL methods. We show that batch whitening (BW) based methods do not impose whitening constraints on the embedding but only require the embedding to be full-rank. This full-rank constraint is also sufficient to avoid dimensional collapse. We further demonstrate that the stable rank of the embedding is invariant during training by gradient descent, given the assumption that embedding is updated with an infinitely small learning rate. Based on our analysis, we propose channel whitening with random group partition (CW-RGP), which exploits the advantages of BW-based methods in preventing collapse and avoids their disadvantages requiring large batch size. Experimental results on ImageNet classification and COCO object detection reveal that the proposed CW-RGP possesses a promising potential for learning good representations.
Lei Huang 0015, Yunhao Ni, Xi Weng, Rao Muhammad Anwer, Salman Khan 0001, Ming-Hsuan Yang 0001, Fahad Shahbaz Khan
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 On the Number of Linear Regions of Convolutional Neural Networks With Piecewise Linear Activations
abstract
One fundamental problem in deep learning is understanding the excellent performance of deep Neural Networks (NNs) in practice. An explanation for the superiority of NNs is that they can realize a large family of complicated functions, i.e., they have powerful expressivity. The expressivity of a Neural Network with Piecewise Linear activations (PLNN) can be quantified by the maximal number of linear regions it can separate its input space into. In this paper, we provide several mathematical results needed for studying the linear regions of Convolutional Neural Networks with Piecewise Linear activations (PLCNNs), and use them to derive the maximal and average numbers of linear regions for one-layer PLCNNs. Furthermore, we obtain upper and lower bounds for the number of linear regions of multi-layer PLCNNs. Our results suggest that deeper PLCNNs have more powerful expressivity than shallow PLCNNs, while PLCNNs have more expressivity than fully-connected PLNNs per parameter, in terms of the number of linear regions.
Huan Xiong, Lei Huang 0015, Wenston J. T. Zang, Xiantong Zhen, Guosen Xie, Bin Gu 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Combating medical noisy labels by disentangled distribution learning and consistency regularization
Yi Zhou 0007, Lei Huang 0015, Tao Zhou 0002, Hanshi Sun
Future Gener. Comput. Syst.2
2023 Normalization Techniques in Training DNNs: Methodology, Analysis and Application
abstract
Normalization techniques are essential for accelerating the training and improving the generalization of deep neural networks (DNNs), and have successfully been used in various applications. This paper reviews and comments on the past, present and future of normalization methods in the context of DNN training. We provide a unified picture of the main motivation behind different approaches from the perspective of optimization, and present a taxonomy for understanding the similarities and differences between them. Specifically, we decompose the pipeline of the most representative normalizing activation methods into three components: the normalization area partitioning, normalization operation and normalization representation recovery. In doing so, we provide insight for designing new normalization technique. Finally, we discuss the current progress in understanding normalization methods, and provide a comprehensive review of the applications of normalization for particular tasks, in which it can effectively solve the key issues.
Lei Huang 0015, Jie Qin 0004, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Delving into the Estimation Shift of Batch Normalization in a Network
abstract
Batch normalization (BN) is a milestone technique in deep learning. It normalizes the activation using mini-batch statistics during training but the estimated population statistics during inference. This paper focuses on investigating the estimation of population statistics. We define the estimation shift magnitude of BN to quantitatively measure the difference between its estimated population statistics and expected ones. Our primary observation is that the estimation shift can be accumulated due to the stack of BN in a network, which has detriment effects for the test performance. We further find a batch-free normalization (BFN) can block such an accumulation of estimation shift. These observations motivate our design of XBNBlock that replace one BN with BFN in the bottleneck block of residual-style networks. Experiments on the ImageNet and COCO benchmarks show that XBNBlock consistently improves the performance of different architectures, including ResNet and ResNeXt, by a significant margin and seems to be more robust to distribution shift.
Lei Huang 0015, Yi Zhou 0007, Tian Wang 0002, Jie Luo 0004, Xianglong Liu 0001
CVPR1
2022 Bi-level Doubly Variational Learning for Energy-based Latent Variable Models
abstract
Energy-based latent variable models (EBLVMs) are more expressive than conventional energy-based models. However, its potential on visual tasks are limited by its training process based on maximum likelihood estimate that requires sampling from two intractable distributions. In this paper, we propose Bi-level doubly variational learning (BiDVL), which is based on a new bi-level optimization framework and two tractable variational distributions to facilitate learning EBLVMs. Particularly, we lead a decoupled EBLVM consisting of a marginal energy-based distribution and a structural posterior to handle the difficulties when learning deep EBLVMs on images. By choosing a symmetric KL divergence in the lower level of our framework, a compact BiDVL for visual tasks can be obtained. Our model achieves impressive image generation performance over related works. It also demonstrates the significant capacity of testing image reconstruction and out-of-distribution detection.
Ge Kan, Jinhu Lü 0001, Tian Wang 0002, Baochang Zhang 0001, Aichun Zhu, Lei Huang 0015, Guodong Guo, Hichem Snoussi
CVPR6
2022 Revisiting Domain Generalized Stereo Matching Networks from a Feature Consistency Perspective
abstract
Despite recent stereo matching networks achieving impressive performance given sufficient training data, they suffer from domain shifts and generalize poorly to unseen domains. We argue that maintaining feature consistency between matching pixels is a vital factor for promoting the generalization capability of stereo matching networks, which has not been adequately considered. Here we address this issue by proposing a simple pixel-wise contrastive learning across the viewpoints. The stereo contrastive feature loss function explicitly constrains the consistency between learned features of matching pixel pairs which are observations of the same 3D points. A stereo selective whitening loss is further introduced to better preserve the stereo feature consistency across domains, which decorrelates stereo features from stereo viewpoint-specific style information. Counter-intuitively, the generalization of feature consistency between two viewpoints in the same scene translates to the generalization of stereo matching performance to unseen domains. Our method is generic in nature as it can be easily embedded into existing stereo networks and does not require access to the samples in the target domain. When trained on synthetic data and generalized to four real-world testing sets, our method achieves superior performance over several state-of-the-art networks. The code is available online11https://github.com/jiaw-z/FCStereo.
Xiang Wang 0014, Xiao Bai 0001, Chen Wang 0026, Lei Huang 0015, Lin Gu 0003, Jun Zhou 0001, Tatsuya Harada, Edwin R. Hancock
CVPR5
2022 Understanding the Failure of Batch Normalization for Transformers in NLP
abstract
Batch Normalization (BN) is a core and prevalent technique in accelerating the training of deep neural networks and improving the generalization on Computer Vision (CV) tasks. However, it fails to defend its position in Natural Language Processing (NLP), which is dominated by Layer Normalization (LN). In this paper, we are trying to answer why BN usually performs worse than LN in NLP tasks with Transformer models. We find that the inconsistency between training and inference of BN is the leading cause that results in the failure of BN in NLP. We define Training Inference Discrepancy (TID) to quantitatively measure this inconsistency and reveal that TID can indicate BN's performance, supported by extensive experiments, including image classification, neural machine translation, language modeling, sequence labeling, and text classification tasks. We find that BN can obtain much better test performance than LN when TID keeps small through training. To suppress the explosion of TID, we propose Regularized BN (RBN) that adds a simple regularization term to narrow the gap between batch statistics and population statistics of BN. RBN improves the performance of BN consistently and outperforms or is on par with LN on 17 out of 20 settings, including ten datasets and two common variants of Transformer.
Ji Wu 0002, Lei Huang 0015
NeurIPS3
2022 An Investigation into Whitening Loss for Self-supervised Learning
abstract
A desirable objective in self-supervised learning (SSL) is to avoid feature collapse. Whitening loss guarantees collapse avoidance by minimizing the distance between embeddings of positive pairs under the conditioning that the embeddings from different views are whitened. In this paper, we propose a framework with an informative indicator to analyze whitening loss, which provides a clue to demystify several interesting phenomena as well as a pivoting point connecting to other SSL methods. We reveal that batch whitening (BW) based methods do not impose whitening constraints on the embedding, but they only require the embedding to be full-rank. This full-rank constraint is also sufficient to avoid dimensional collapse. Based on our analysis, we propose channel whitening with random group partition (CW-RGP), which exploits the advantages of BW-based methods in preventing collapse and avoids their disadvantages requiring large batch size. Experimental results on ImageNet classification and COCO object detection reveal that the proposed CW-RGP possesses a promising potential for learning good representations. The code is available at https://github.com/winci-ai/CW-RGP.
Xi Weng, Lei Huang 0015, Rao Muhammad Anwer, Salman Khan 0001, Fahad Shahbaz Khan
NeurIPS2
2021 Slimmable Generative Adversarial Networks
abstract
Generative adversarial networks (GANs) have achieved remarkable progress in recent years, but the continuously growing scale of models make them challenging to deploy widely in practical applications. In particular, for real-time generation tasks, different devices require generators of different sizes due to varying computing power. In this paper, we introduce slimmable GANs (SlimGANs), which can flexibly switch the width of the generator to accommodate various quality-efficiency trade-offs at runtime. Specifically, we leverage multiple discriminators that share partial parameters to train the slimmable generator. To facilitate the consistency between generators of different widths, we present a stepwise inplace distillation technique that encourages narrow generators to learn from wide ones. As for class-conditional generation, we propose a sliceable conditional batch normalization that incorporates the label information into different widths. Our methods are validated, both quantitatively and qualitatively, by extensive experiments and a detailed ablation study.
Zehuan Yuan, Lei Huang 0015, Huawei Shen, Xueqi Cheng 0001, Changhu Wang
AAAI3
2021 Many-to-One Distribution Learning and K-Nearest Neighbor Smoothing for Thoracic Disease Identification
abstract
Chest X-rays are an important and accessible clinical imaging tool for the detection of many thoracic diseases. Over the past decade, deep learning, with a focus on the convolutional neural network (CNN), has become the most powerful computer-aided diagnosis technology for improving disease identification performance. However, training an effective and robust deep CNN usually requires a large amount of data with high annotation quality. For chest X-ray imaging, annotating large-scale data requires professional domain knowledge and is time-consuming. Thus, existing public chest X-ray datasets usually adopt language pattern based methods to automatically mine labels from reports. However, this results in label uncertainty and inconsistency. In this paper, we propose many-to-one distribution learning (MODL) and K-nearest neighbor smoothing (KNNS) methods from two perspectives to improve a single model's disease identification performance, rather than focusing on an ensemble of models. MODL integrates multiple models to obtain a soft label distribution for optimizing the single target model, which can reduce the effects of original label uncertainty. Moreover, KNNS aims to enhance the robustness of the target model to provide consistent predictions on images with similar medical findings. Extensive experiments on the public NIH Chest X-ray and CheXpert datasets show that our model achieves consistent improvements over the state-of-the-art methods.
Yi Zhou 0007, Lei Huang 0015, Tianfei Zhou, Ling Shao 0001
AAAI2
2021 Group Whitening: Balancing Learning Efficiency and Representational Capacity
abstract
Batch normalization (BN) is an important technique commonly incorporated into deep learning models to perform standardization within mini-batches. The merits of BN in improving a model’s learning efficiency can be further amplified by applying whitening, while its drawbacks in estimating population statistics for inference can be avoided through group normalization (GN). This paper proposes group whitening (GW), which exploits the advantages of the whitening operation and avoids the disadvantages of normalization within mini-batches. In addition, we analyze the constraints imposed on features by normalization, and show how the batch size (group number) affects the performance of batch (group) normalized networks, from the perspective of model’s representational capacity. This analysis provides theoretical guidance for applying GW in practice. Finally, we apply the proposed GW to ResNet and ResNeXt architectures and conduct experiments on the ImageNet and COCO benchmarks. Results show that GW consistently improves the performance of different architectures, with absolute gains of 1.02% ∼ 1.49% in top-1 accuracy on ImageNet and 1.82% ∼ 3.21% in bounding box AP on COCO.
Lei Huang 0015, Yi Zhou 0007, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001
CVPR1
2021 S2-BNN: Bridging the Gap Between Self-Supervised Real and 1-Bit Neural Networks via Guided Distribution Calibration
abstract
Previous studies dominantly target at self-supervised learning on real-valued networks and have achieved many promising results. However, on the more challenging binary neural networks (BNNs), this task has not yet been fully explored in the community. In this paper, we focus on this more difficult scenario: learning networks where both weights and activations are binary, meanwhile, without any human annotated labels. We observe that the commonly used contrastive objective is not satisfying on BNNs for competitive accuracy, since the backbone network contains relatively limited capacity and representation ability. Hence instead of directly applying existing self-supervised methods, which cause a severe decline in performance, we present a novel guided learning paradigm from real-valued to distill binary networks on the final prediction distribution, to minimize the loss and obtain desirable accuracy. Our proposed method can boost the simple contrastive learning baseline by an absolute gain of 5.5∼15% on BNNs. We further reveal that it is difficult for BNNs to recover the similar predictive distributions as real-valued models when training without labels. Thus, how to calibrate them is key to address the degradation in performance. Extensive experiments are conducted on the large-scale ImageNet and downstream datasets. Our method achieves substantial improvement over the simple contrastive learning baseline, and is even comparable to many mainstream supervised BNN methods. Code is available at https://github.com/szq0214/S2-BNN.
Zechun Liu, Jie Qin 0004, Lei Huang 0015, Kwang-Ting Cheng, Marios Savvides
CVPR4
2021 CCT-Net: Category-Invariant Cross-Domain Transfer for Medical Single-to-Multiple Disease Diagnosis
abstract
A medical imaging model is usually explored for the diagnosis of a single disease. However, with the expanding demand for multi-disease diagnosis in clinical applications, multi-function solutions need to be investigated. Previous works proposed to either exploit different disease labels to conduct transfer learning through fine-tuning, or transfer knowledge across different domains with similar diseases. However, these methods still cannot address the real clinical challenge - a multi-disease model is required but annotations for each disease are not always available. In this paper, we introduce the task of transferring knowledge from single-disease diagnosis (source domain) to enhance multi-disease diagnosis (target domain). A category-invariant cross-domain transfer (CCT) method is proposed to address this single-to-multiple extension. First, for domain-specific task learning, we present a confidence weighted pooling (CWP) to obtain coarse heatmaps for different disease categories. Then, conditioned on these heatmaps, category-invariant feature refinement (CIFR) blocks are proposed to better localize discriminative semantic regions related to the corresponding diseases. The category-invariant characteristic enables transferability from the source domain to the target domain. We validate our method in two popular areas: extending diabetic retinopathy to identifying multiple ocular diseases, and extending glioma identification to the diagnosis of other brain tumors.
Yi Zhou 0007, Lei Huang 0015, Tao Zhou 0002, Ling Shao 0001
ICCV2
2021 Visual-Textual Attentive Semantic Consistency for Medical Report Generation
abstract
Automatic report generation on medical radiographs have recently gained interest. However, identifying diseases as well as correctly predicting their corresponding sizes, locations and other medical description patterns, which is essential for generating high-quality reports, is challenging. Although previous methods focused on producing readable reports, how to accurately detect and describe findings that match with the query X-Ray has not been successfully addressed. In this paper, we propose a multi-modality semantic attention model to integrate visual features, predicted key finding embeddings, as well as clinical features, and progressively decode reports with visual-textual semantic consistency. First, multi-modality features are extracted and attended with the hidden states from the sentence de-coder, to encode enriched context vectors for better decoding a report. These modalities include regional visual features of scans, semantic word embeddings of the top-K findings predicted with high probabilities, and clinical features of indications. Second, the progressive report decoder consists of a sentence decoder and a word decoder, where we propose image-sentence matching and description accuracy losses to constrain the visual-textual semantic consistency. Extensive experiments on the public MIMIC-CXR and IU X-Ray datasets show that our model achieves consistent improvements over the state-of-the-art methods.
Yi Zhou 0007, Lei Huang 0015, Tao Zhou 0002, Huazhu Fu, Ling Shao 0001
ICCV2
2021 Rot-Pro: Modeling Transitivity by Projection in Knowledge Graph Embedding
abstract
Knowledge graph embedding models learn the representations of entities and relations in the knowledge graphs for predicting missing links (relations) between entities. Their effectiveness are deeply affected by the ability of modeling and inferring different relation patterns such as symmetry, asymmetry, inversion, composition and transitivity. Although existing models are already able to model many of these relations patterns, transitivity, a very common relation pattern, is still not been fully supported. In this paper, we first theoretically show that the transitive relations can be modeled with projections. We then propose the Rot-Pro model which combines the projection and relational rotation together. We prove that Rot-Pro can infer all the above relation patterns. Experimental results show that the proposed Rot-Pro model effectively learns the transitivity pattern and achieves the state-of-the-art results on the link prediction task in the datasets containing transitive relations.
Tengwei Song, Jie Luo 0004, Lei Huang 0015
NeurIPS3
2021 A Benchmark for Studying Diabetic Retinopathy: Segmentation, Grading, and Transferability
abstract
People with diabetes are at risk of developing an eye disease called diabetic retinopathy (DR). This disease occurs when high blood glucose levels cause damage to blood vessels in the retina. Computer-aided DR diagnosis has become a promising tool for the early detection and severity grading of DR, due to the great success of deep learning. However, most current DR diagnosis systems do not achieve satisfactory performance or interpretability for ophthalmologists, due to the lack of training data with consistent and fine-grained annotations. To address this problem, we construct a large fine-grained annotated DR dataset containing 2,842 images (FGADR). Specifically, this dataset has 1,842 images with pixel-level DR-related lesion annotations, and 1,000 images with image-level labels graded by six board-certified ophthalmologists with intra-rater consistency. The proposed dataset will enable extensive studies on DR diagnosis. Further, we establish three benchmark tasks for evaluation: 1. DR lesion segmentation; 2. DR grading by joint classification and segmentation; 3. Transfer learning for ocular multi-disease identification. Moreover, a novel inductive transfer learning method is introduced for the third task. Extensive experiments using different state-of-the-art methods are conducted on our FGADR dataset, which can serve as baselines for future research. Our dataset will be released in https://csyizhou.github.io/FGADR/.
Yi Zhou 0007, Lei Huang 0015, Shanshan Cui, Ling Shao 0001
IEEE Trans. Medical Imaging3
2020 Controllable Orthogonalization in Training DNNs
abstract
Orthogonality is widely used for training deep neural networks (DNNs) due to its ability to maintain all singular values of the Jacobian close to 1 and reduce redundancy in representation. This paper proposes a computationally efficient and numerically stable orthogonalization method using Newton's iteration (ONI), to learn a layer-wise orthogonal weight matrix in DNNs. ONI works by iteratively stretching the singular values of a weight matrix towards 1. This property enables it to control the orthogonality of a weight matrix by its number of iterations. We show that our method improves the performance of image classification networks by effectively controlling the orthogonality to provide an optimal tradeoff between optimization benefits and representational capacity reduction. We also show that ONI stabilizes the training of generative adversarial networks (GANs) by maintaining the Lipschitz continuity of a network, similar to spectral normalization (SN), and further outperforms SN by providing controllable orthogonality.
Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Diwen Wan, Zehuan Yuan, Bo Li 0026, Ling Shao 0001
CVPR1
2020 An Investigation Into the Stochasticity of Batch Whitening
abstract
Batch Normalization (BN) is extensively employed in various network architectures by performing standardization within mini-batches. A full understanding of the process has been a central target in the deep learning communities. Unlike existing works, which usually only analyze the standardization operation, this paper investigates the more general Batch Whitening (BW). Our work originates from the observation that while various whitening transformations equivalently improve the conditioning, they show significantly different behaviors in discriminative scenarios and training Generative Adversarial Networks (GANs). We attribute this phenomenon to the stochasticity that BW introduces. We quantitatively investigate the stochasticity of different whitening transformations and show that it correlates well with the optimization behaviors during training. We also investigate how stochasticity relates to the estimation of population statistics during inference. Based on our analysis, we provide a framework for designing and comparing BW algorithms in different scenarios. Our proposed BW algorithm improves the residual networks by a significant margin on ImageNet classification. Besides, we show that the stochasticity of BW can improve the GAN's performance with, however, the sacrifice of the training stability.
Lei Huang 0015, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001
CVPR1
2020 An Efficient Agreement Mechanism in CapsNets by Pairwise Product
Lei Huang 0015
ECAI3
2020 Layer-Wise Conditioning Analysis in Exploring the Learning Dynamics of DNNs
Lei Huang 0015, Jie Qin 0004, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001
ECCV (2)1
2020 Invertible Zero-Shot Recognition Flows
Yuming Shen, Jie Qin 0004, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001
ECCV (16)3
2020 On the Number of Linear Regions of Convolutional Neural Networks
abstract
One fundamental problem in deep learning is understanding the outstanding performance of deep Neural Networks (NNs) in practice. One explanation for the superiority of NNs is that they can realize a large class of complicated functions, i.e., they have powerful expressivity. The expressivity of a ReLU NN can be quantified by the maximal number of linear regions it can separate its input space into. In this paper, we provide several mathematical results needed for studying the linear regions of CNNs, and use them to derive the maximal and average numbers of linear regions for one-layer ReLU CNNs. Furthermore, we obtain upper and lower bounds for the number of linear regions of multi-layer ReLU CNNs. Our results suggest that deeper CNNs have more powerful expressivity than their shallow counterparts, while CNNs have more expressivity than fully-connected NNs per parameter.
Huan Xiong, Lei Huang 0015, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001
ICML2
2020 Deep Local Binary Coding for Person Re-Identification by Delving into the Details
abstract
Person re-identification (ReID) has recently received extensive research interests due to its diverse applications in multimedia analysis and computer vision. However, the majority of existing works focus on improving matching accuracy, while ignoring matching efficiency. In this work, we present a novel binary representation learning framework for efficient person ReID, namely Deep Local Binary Coding (DLBC). Different from existing deep binary ReID approaches, DLBC attempts to learn discriminative binary codes by explicitly interacting with local visual details. Specifically, DLBC first extracts a set of local features from spatially salient regions of pedestrian images. Subsequently, DLBC formulates a new binary-local semantic mutual information (BSMI) maximization term, based on which a self-lifting (SL) block is built to further exploit the semantic importance of local features. The BSMI term together with the SL block simultaneously enhances the dependency of binary codes on selected local features as well as their robustness to cross-view visual inconsistency. In addition, an efficient optimizing method is developed to train the proposed deep models with orthogonal and binary constraints. Extensive experiments reveal that DLBC significantly minimizes the accuracy gap between binary ReID methods and the state-of-the-art real-valued ones, whilst remarkably reducing query time and memory cost.
Jiaxin Chen 0002, Jie Qin 0004, Yichao Yan, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001
ACM Multimedia4
2020 Projection based weight normalization: Efficient method for optimization on oblique manifold in DNNs
Lei Huang 0015, Xianglong Liu 0001, Jie Qin 0004, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001
Pattern Recognit.1
2020 Deep quantization generative networks
Diwen Wan, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Lei Huang 0015, Mengyang Yu, Heng Tao Shen, Ling Shao 0001
Pattern Recognit.5
2020 A general non-parametric active learning framework for classification on multiple manifolds
Lei Huang 0015, Yuqing Ma, Xianglong Liu 0001
Pattern Recognit. Lett.1
2019 Iterative Normalization: Beyond Standardization Towards Efficient Whitening
abstract
Batch Normalization (BN) is ubiquitously employed for accelerating neural network training and improving the generalization capability by performing standardization within mini-batches. Decorrelated Batch Normalization (DBN) further boosts the above effectiveness by whitening. However, DBN relies heavily on either a large batch size, or eigen-decomposition that suffers from poor efficiency on GPUs. We propose Iterative Normalization (IterNorm), which employs Newton’s iterations for much more efficient whitening, while simultaneously avoiding the eigen-decomposition. Furthermore, we develop a comprehensive study to show IterNorm has better trade-off between optimization and generalization, with theoretical and experimental support. To this end, we exclusively introduce Stochastic Normalization Disturbance (SND), which measures the inherent stochastic uncertainty of samples when applied to normalization operations. With the support of SND, we provide natural explanations to several phenomena from the perspective of optimization, e.g., why group-wise whitening of DBN generally outperforms full-whitening and why the accuracy of BN degenerates with reduced batch sizes. We demonstrate the consistently improved performance of IterNorm with extensive experiments on CIFAR-10 and ImageNet over BN and DBN.
Lei Huang 0015, Yi Zhou 0007, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001
CVPR1
2019 Collaborative Learning of Semi-Supervised Segmentation and Classification for Medical Images
abstract
Medical image analysis has two important research areas: disease grading and fine-grained lesion segmentation. Although the former problem often relies on the latter, the two are usually studied separately. Disease severity grading can be treated as a classification problem, which only requires image-level annotations, while the lesion segmentation requires stronger pixel-level annotations. However, pixel-wise data annotation for medical images is highly time-consuming and requires domain experts. In this paper, we propose a collaborative learning method to jointly improve the performance of disease grading and lesion segmentation by semi-supervised learning with an attention mechanism. Given a small set of pixel-level annotated data, a multi-lesion mask generation model first performs the traditional semantic segmentation task. Then, based on initially predicted lesion maps for large quantities of image-level annotated data, a lesion attentive disease grading model is designed to improve the severity classification accuracy. Meanwhile, the lesion attention model can refine the lesion maps using class-specific information to fine-tune the segmentation model in a semi-supervised manner. An adversarial architecture is also integrated for training. With extensive experiments on a representative medical problem called diabetic retinopathy (DR), we validate the effectiveness of our method and achieve consistent improvements over state-of-the-art methods on three public datasets.
Yi Zhou 0007, Xiaodong He 0004, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Shanshan Cui, Ling Shao 0001
CVPR3
2018 Orthogonal Weight Normalization: Solution to Optimization Over Multiple Dependent Stiefel Manifolds in Deep Neural Networks
abstract
Orthogonal matrix has shown advantages in training Recurrent Neural Networks (RNNs), but such matrix is limited to be square for the hidden-to-hidden transformation in RNNs. In this paper, we generalize such square orthogonal matrix to orthogonal rectangular matrix and formulating this problem in feed-forward Neural Networks (FNNs) as Optimization over Multiple Dependent Stiefel Manifolds (OMDSM). We show that the orthogonal rectangular matrix can stabilize the distribution of network activations and regularize FNNs. We propose a novel orthogonal weight normalization method to solve OMDSM. Particularly, it constructs orthogonal transformation over proxy parameters to ensure the weight matrix is orthogonal. To guarantee stability, we minimize the distortions between proxy parameters and canonical weights over all tractable orthogonal transformations. In addition, we design orthogonal linear module (OLM) to learn orthogonal filter banks in practice, which can be used as an alternative to standard linear module. Extensive experiments demonstrate that by simply substituting OLM for standard linear module without revising any experimental protocols, our method improves the performance of the state-of-the-art networks, including Inception and residual networks on CIFAR and ImageNet datasets.
Lei Huang 0015, Xianglong Liu 0001, Bo Lang, Adams Wei Yu, Bo Li 0026
AAAI1
2018 Decorrelated Batch Normalization
abstract
Batch Normalization (BN) is capable of accelerating the training of deep models by centering and scaling activations within mini-batches. In this work, we propose Decorrelated Batch Normalization (DBN), which not just centers and scales activations but whitens them. We explore multiple whitening techniques, and find that PCA whitening causes a problem we call stochastic axis swapping, which is detrimental to learning. We show that ZCA whitening does not suffer from this problem, permitting successful learning. DBN retains the desirable qualities of BN and further improves BN's optimization efficiency and generalization ability. We design comprehensive experiments to show that DBN can improve the performance of BN on multilayer perceptrons and convolutional neural networks. Furthermore, we consistently improve the accuracy of residual networks on CIFAR-10, CIFAR-100, and ImageNet.
Lei Huang 0015, Bo Lang, Jia Deng 0001
CVPR1
2018 DGCNN: Disordered graph convolutional neural network based on the Gaussian mixture model
Bo Wu 0021, Yang Liu 0088, Bo Lang, Lei Huang 0015
Neurocomputing4
2017 Centered Weight Normalization in Accelerating Training of Deep Neural Networks
abstract
Training deep neural networks is difficult for the pathological curvature problem. Re-parameterization is an effective way to relieve the problem by learning the curvature approximately or constraining the solutions of weights with good properties for optimization. This paper proposes to reparameterize the input weight of each neuron in deep neural networks by normalizing it with zero-mean and unit-norm, followed by a learnable scalar parameter to adjust the norm of the weight. This technique effectively stabilizes the distribution implicitly. Besides, it improves the conditioning of the optimization problem and thus accelerates the training of deep neural networks. It can be wrapped as a linear module in practice and plugged in any architecture to replace the standard linear module. We highlight the benefits of our method on both multi-layer perceptrons and convolutional neural networks, and demonstrate its scalability and efficiency on SVHN, CIFAR-10, CIFAR-100 and ImageNet datasets.
Lei Huang 0015, Xianglong Liu 0001, Yang Liu 0088, Bo Lang, Dacheng Tao
ICCV1
2017 Learning Joint Multimodal Representation Based on Multi-fusion Deep Neural Networks
Zepeng Gu, Bo Lang, Tongyu Yue, Lei Huang 0015
ICONIP (2)4
2016 Efficient segmentation for Region-based Image Retrieval using Edge Integrated Minimum Spanning Tree
abstract
Region-based Image Retrieval (RBIR), which bases itself on image segmentation rather than global features or key-point-based local features, is a branch of Content-based Image Retrieval. This paper proposes a novel RBIR-oriented image segmentation algorithm named Edge Integrated Minimum Spanning Tree (EI-MST). The difference between EI-MST and the traditional MST-based methods is that EI-MST generates MSTs over edge-maps rather than the original images, which achieved high retrieval performance cooperating with state-of-the-art matching strategies. In addition, by limiting the nodes in every MST with adaptive scale selection, EI-MST is efficient especially when processing high resolution images. The experiments on four popular public datasets proved that, EI-MST is capable of achieving higher retrieval accuracy over four widely used segmentation methods while only consuming moderate amount of time in both online and offline parts of RBIR systems.
Yang Liu 0088, Lei Huang 0015, Xianglong Liu 0001, Bo Lang
ICPR2
2016 Full-duplex based successive interference cancellation in heterogeneous networks
abstract
This paper studies the mitigation of cross-tier inter-cell interference (ICI) generated by a macro base station to a small-cell user equipment (SUE) in heterogeneous networks. A full-duplex (FD) based successive ICI cancellation (SIC) scheme, called fSICIC, is devised by applying FD technique at the small-cell base station (SBS). The basic idea of the fSICIC is to let the SBS send the desired signal and forward the overheard cross-tier ICI simultaneously to the SUE, where the forwarded ICI is controlled to enhance the ICI at the SUE to facilitate SIC. We first investigate the feasibility of the fSICIC, and then optimize the fSICIC to maximize the data rate of the SUE. Simulation results demonstrate the advantages of the fSICIC on mitigating cross-tier ICI, especially for strong ICI.
Lei Huang 0015, Shengqian Han, Chenyang Yang 0001, Gang Wang 0009
PIMRC1
2016 A novel rotation adaptive object detection method based on pair Hough model
Yang Liu 0088, Lei Huang 0015, Xianglong Liu 0001, Bo Lang
Neurocomputing2
2016 Query-Adaptive Hash Code Ranking for Large-Scale Multi-View Visual Search
abstract
Hash-based nearest neighbor search has become attractive in many applications. However, the quantization in hashing usually degenerates the discriminative power when using Hamming distance ranking. Besides, for large-scale visual search, existing hashing methods cannot directly support the efficient search over the data with multiple sources, and while the literature has shown that adaptively incorporating complementary information from diverse sources or views can significantly boost the search performance. To address the problems, this paper proposes a novel and generic approach to building multiple hash tables with multiple views and generating fine-grained ranking results at bitwise and tablewise levels. For each hash table, a query-adaptive bitwise weighting is introduced to alleviate the quantization loss by simultaneously exploiting the quality of hash functions and their complement for nearest neighbor search. From the tablewise aspect, multiple hash tables are built for different data views as a joint index, over which a query-specific rank fusion is proposed to rerank all results from the bitwise ranking by diffusing in a graph. Comprehensive experiments on image search over three well-known benchmarks show that the proposed method achieves up to 17.11% and 20.28% performance gains on single and multiple table search over the state-of-the-art methods.
Xianglong Liu 0001, Lei Huang 0015, Cheng Deng 0002, Bo Lang, Dacheng Tao
IEEE Trans. Image Process.2
2015 Multi-View Complementary Hash Tables for Nearest Neighbor Search
abstract
Recent years have witnessed the success of hashing techniques in fast nearest neighbor search. In practice many applications (eg., visual search, object detection, image matching, etc.) have enjoyed the benefits of complementary hash tables and information fusion over multiple views. However, most of prior research mainly focused on compact hash code cleaning, and rare work studies how to build multiple complementary hash tables, much less to adaptively integrate information stemming from multiple views. In this paper we first present a novel multi-view complementary hash table method that learns complementarity hash tables from the data with multiple views. For single multi-view table, using exemplar based feature fusion, we approximate the inherent data similarities with a low-rank matrix, and learn discriminative hash functions in an efficient way. To build complementary tables and meanwhile maintain scalable training and fast out-of-sample extension, an exemplar reweighting scheme is introduced to update the induced low-rank similarity in the sequential table construction framework, which indeed brings mutual benefits between tables by placing greater importance on exemplars shared by mis-separated neighbors. Extensive experiments on three large-scale image datasets demonstrate that the proposed method significantly outperforms various naive solutions and state-of-the-art multi-table methods.
Xianglong Liu 0001, Lei Huang 0015, Cheng Deng 0002, Jiwen Lu, Bo Lang
ICCV2
2015 Online semi-supervised annotation via proxy-based local consistency propagation
Lei Huang 0015, Xianglong Liu 0001, Binqiang Ma, Bo Lang
Neurocomputing1
2014 Ontology-based Concept Similarity Integrating Image Semantic and Visual Information
abstract
In recent years, the concept similarity measure has received wide attention in many applications, such as ontology construction, text analysis, image retrieval, etc.Currently, the concept similarity measure depends on the information mining in various knowledge bases, like dictionaries, ontologies, image annotation labels, and search engines.However, these knowledge bases usually only contain semantic information.With the development of the Internet and the popularity of the digital imaging devices, a lot of images and related texts have appeared, which help us to further mine the concept similarity relationships.The concept similarity is the outcome of human subjective perception.In addition to analysis of semantic information, the content of image itself precisely provides the visual perception information, which also plays an important role in the access of concept similarity relationships.To integrate both image semantic and visual information, in this paper we propose an ontology concept similarity measure that simultaneously utilizes the image semantic annotations and visual features to optimize the ontology-based metrics.The experiment result on the Corel dataset demonstrates the effectiveness of our proposed method.
Mengyun Wang, Xianglong Liu 0001, Lei Huang 0015, Bo Lang, Hailiang Yu
FedCSIS3
2014 Graph-based active semi-supervised learning: A new perspective for relieving multi-class annotation labor
abstract
Semi-supervised learning and active learning are important techniques to build more accurate model while labeled data are scarce. The objective of this paper is combining both to effectively relieve user labor for multi-class annotation. We propose a novel graph-based active semi-supervised learning framework which aim at efficiently learning a multi-class model with minimal human labor. In particular, we propose Minimize Expected Global Uncertainty algorithm to actively select examples (for labels), which naturally integrates with the probabilistic results of graph-based semi-supervised learning. Meanwhile, we update the model incrementally by decomposed formulation while the new example are incorporated for training, which only has the time complexity of O(n), compared to the original re-training of O(n3). Extensive evaluations over three real-world datasets demonstrate that our proposed method has the superior performance comparing with the baselines and the capability to efficiently build more accurate model with fractional human labor.
Lei Huang 0015, Yang Liu 0088, Xianglong Liu 0001, Xindong Wang, Bo Lang
ICME1
2014 Query-Adaptive Hash Code Ranking for Fast Nearest Neighbor Search
abstract
Recently hash-based nearest neighbor search has become attractive in many applications due to its compressed storage and fast query speed. However, the quantization in the hashing process usually degenerates its discriminative power when using Hamming distance ranking. To enable fine-grained ranking, hash bit weighting has been proved as a promising solution. Though achieving satisfying performance improvement, state-of-the-art weighting methods usually heavily rely on the projection's distribution assumption, and thus can hardly be directly applied to more general types of hashing algorithms. In this paper, we propose a new ranking method named QRank with query-adaptive bitwise weights by exploiting both the discriminative power of each hash function and their complement for nearest neighbor search. QRank is a general weighting method for all kinds of hashing algorithms without any strict assumptions. Experimental results on two well-known benchmarks MNIST and NUS-WIDE show that the proposed method can achieve up to 17.11\% performance gains over state-of-the-art methods.
Tianxu Ji, Xianglong Liu 0001, Cheng Deng 0002, Lei Huang 0015, Bo Lang
ACM Multimedia4
2013 Efficient semi-supervised annotation with Proxy-based Local Consistency Propagation
abstract
Semi-supervised learning methods can largely leverage the image annotation problem using both labeled and unlabeled data, especially when the labeled information is quite limited. However, most of them suffer the expensive computation stemming from the batch learning on large training dataset. In this paper we proposed a highly efficient semi-supervised annotation approach with the partial label propagation based on the graph representation. Specifically, the label information is first propagated from labeled samples to the unlabeled ones, and then spreads only among unlabeled ones like a spreading activation network. Our approach takes advantage of the decomposed formulation to achieve a fast incremental learning instead of the expensive batch one without accuracy loss. Extensive evaluations over two large datasets demonstrate the superior performance of the proposed method and its significant efficiency.
Lei Huang 0015, Xianglong Liu 0001, Bo Lang
ICME1