Liansheng Zhuang

dblp:94/5815 · DBLP profile ↗
← Back
60ranked-venue papers
13as first author
32since 2021 · last 2025
0000-0002-4345-856XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 43 · 10 first-author · 22 since 2021Artificial intelligence and machine learning · 26 · 5 first-author · 16 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 EPERM: An Evidence Path Enhanced Reasoning Model for Knowledge Graph Question and Answering
abstract
Due to the remarkable reasoning ability, Large language models (LLMs) have demonstrated impressive performance in knowledge graph question answering (KGQA) tasks, which find answers to natural language questions over knowledge graphs (KGs). To alleviate the hallucinations and lack of knowledge issues of LLMs, existing methods often retrieve the question-related information from KGs to enrich the input context. However, most methods focus on retrieving the relevant information while ignoring the importance of different types of knowledge in reasoning, which degrades their performance. To this end, this paper reformulates the KGQA problem as a graphical model and proposes a three-stage framework named the Evidence Path Enhanced Reasoning Model (EPERM) for KGQA. In the first stage, EPERM uses the fine-tuned LLM to retrieve a subgraph related to the question from the original knowledge graph. In the second stage, EPERM filters out the evidence paths that faithfully support the reasoning of the questions, and score their importance in reasoning. Finally, EPERM uses the weighted evidence paths to reason the final answer. Since considering the importance of different structural information in KGs for reasoning, EPERM can improve the reasoning ability of LLMs in KGQA tasks. Extensive experiments on benchmark datasets demonstrate that EPERM achieves superior performances in KGQA tasks.
Liansheng Zhuang, Aodi Li, Minghong Yao, Shafei Wang
AAAI2
2025 Seeking Consistent Flat Minima for Better Domain Generalization via Refining Loss Landscapes
abstract
Domain generalization aims to learn a model from multiple training domains and generalize it to unseen test domains. Recent theory has shown that seeking the deep models, whose parameters lie in the flat minima of the loss landscape, can significantly reduce the out-of-domain generalization error. However, existing methods often neglect the consistency of loss landscapes in different domains, resulting in models that are not simultaneously in the optimal flat minima in all domains, which limits their generalization ability. To address this issue, this paper proposes an iterative Self-Feedback Training (SFT) framework to seek consistent flat minima that are shared across different domains by progressively refining loss landscapes during training. It alternatively generates a feedback signal by measuring the inconsistency of loss landscapes in different domains and refines these loss landscapes for greater consistency using this feedback signal. Benefiting from the consistency of the flat minima within these refined loss landscapes, our SFT helps achieve better out-of-domain generalization. Extensive experiments on DomainBed demonstrate superior performances of SFT when compared to state-of-the-art sharpness-aware methods and other prevalent DG baselines. On average across five DG benchmarks, SFT surpasses the sharpness-aware minimization by 2.6% with ResNet-50 and 1.5% with ViT-B/16, respectively.
Aodi Li, Liansheng Zhuang, Minghong Yao, Shafei Wang
CVPR2
2025 Progressive Neural Architecture Generation with Weaker Predictors
Zhengzhuo Zhang, Liansheng Zhuang
MMM (3)2
2025 Purified Policy Space Response Oracles for Symmetric Zero-Sum Games
abstract
Policy space response oracles (PSRO) is a promising tool to find an approximate Nash equilibrium (NE) in a two-player zero-sum game. It solves the equilibrium by iteratively expanding a small-scale meta-game formed by a restricted strategy population consisting of historical approximate best responses of the meta-games. However, since these best responses have a strong correlation with each other, existing PSRO and its variants often have the slow diversity growth of the strategy population, and thus suffer from poor exploration efficiency and slow convergence rate. To address this problem, this article proposes Purified PSRO, which deliberately maintains a pure strategy population formed by pure strategy bases of approximate best responses. A novel module namely non-best response suppression (NBRS) is introduced to calculate a pure strategy base with better orthogonality to expand the strategy population at each epoch. In this way, Purified PSRO can quickly increase the diversity of the strategy population, thus greatly enhance the efficiency of exploration. Theoretically, we prove the convergence of Purified PSRO. Moreover, we introduce an early stop module to reduce computation cost, and give the upper bound of the exploitability when the algorithm stops early. Extensive experiments on random games of skill (RGoS) and real-world meta-games show that Purified PSRO can consistently outperform existing SOTA methods, sometimes with a large margin.
Zhengdao Shao, Liansheng Zhuang, Houqiang Li, Shafei Wang
IEEE Trans. Neural Networks Learn. Syst.2
2025 COPSRO: An Offline Empirical Game Theoretic Method With Conservative Critic
abstract
This article studies how to learn approximate Nash equilibrium (NE) from static historical datasets by empirical game-theoretic analysis (EGTA), which provides a simulation-based framework to model complex multiagent interactions. Generally, EGTA requires plentiful interactions with the environment or simulator to estimate a cogent and tractable game model approximating the underlying game. However, these exploratory interactions often suffer from low data utilization efficiency and may not be feasible in risk-sensitive applications. To address these problems, this article investigates a new EGTA paradigm for offline settings and introduces a novel algorithm called conservative offline policy space response oracle (COPSRO) to identify NE from fixed datasets without active data collection. COPSRO initiates by extracting a set of strategies from the offline dataset to construct an overcomplete strategy population, achieving an approximation to the policy space of the original game. Then, COPSRO integrates the conservative critic (CC) to tackle the challenge of overestimation inherent in offline learning scenarios. Additionally, it devises the offline NE solver to iteratively compute approximate NE. Consequently, COPSRO can ascertain equilibrium strategies without real-world interaction, markedly enhancing its utility in risk-averse settings. This article provides both theoretical analysis and empirical evaluation to demonstrate the effectiveness and superiority of COPSRO across various real-world tasks in the offline setting. Our method surpasses existing approaches in terms of convergence and exploitability, especially when the coverage ration of dataset is low (20% or 10%).
Zhengdao Shao, Liansheng Zhuang, Houqiang Li, Shafei Wang
IEEE Trans. Neural Networks Learn. Syst.2
2024 KGDM: A Diffusion Model to Capture Multiple Relation Semantics for Knowledge Graph Embedding
abstract
Knowledge graph embedding (KGE) is an efficient and scalable method for knowledge graph completion. However, most existing KGE methods suffer from the challenge of multiple relation semantics, which often degrades their performance. This is because most KGE methods learn fixed continuous vectors for entities (relations) and make deterministic entity predictions to complete the knowledge graph, which hardly captures multiple relation semantics. To tackle this issue, previous works try to learn complex probabilistic embeddings instead of fixed embeddings but suffer from heavy computational complexity. In contrast, this paper proposes a simple yet efficient framework namely the Knowledge Graph Diffusion Model (KGDM) to capture the multiple relation semantics in prediction. Its key idea is to cast the problem of entity prediction into conditional entity generation. Specifically, KGDM estimates the probabilistic distribution of target entities in prediction through Denoising Diffusion Probabilistic Models (DDPM). To bridge the gap between continuous diffusion models and discrete KGs, two learnable embedding functions are defined to map entities and relation to continuous vectors. To consider connectivity patterns of KGs, a Conditional Entity Denoiser model is introduced to generate target entities conditioned on given entities and relations. Extensive experiments demonstrate that KGDM significantly outperforms existing state-of-the-art methods in three benchmark datasets.
Liansheng Zhuang, Aodi Li, Jiuchang Wei, Houqiang Li, Shafei Wang
AAAI2
2024 Learning Label Dependencies for Visual Information Extraction
Minghong Yao, Liansheng Zhuang, Houqiang Li, Jiuchang Wei
IJCAI2
2024 Conservative In-Distribution Q-Learning for Offline Reinforcement Learning
abstract
Offline Reinforcement Learning (RL) aims to learn policies from pre-collected datasets without any additional interaction. In order to perform well and robustly in dynamic environments with noise or disturbances, the learned value function and derived policy should generalize well within and near the dataset distribution, rather than ‘over-fitting’ to training samples. To meet this requirement, we propose a new approach called Conservative In-Distribution Q-learning (CIDQL) that takes a step towards in-distribution offline RL. CIDQL is designed to learn in-distribution with respect to the dataset, using a perturbation-based interpolation technique and a quantile method for value regularization. It prohibits bootstrapping during value iteration, ensuring stable Q-value learning that is separated from policy improvement. The approach has theoretical guarantees for both Q-value underestimation and non-underestimation, and outperforms most SOTA algorithms on D4RL gym-MuJoCo benchmarks.
Zhengdao Shao, Liansheng Zhuang, Liting Chen
IJCNN2
2024 Private Gradient Estimation is Useful for Generative Modeling
abstract
While generative models have proved successful in many domains, they may pose a privacy leakage risk in practical deployment. To address this issue, differentially private generative model learning has emerged as a solution to train private generative models for different downstream tasks. However, existing private generative modeling approaches face significant challenges in generating high-dimensional data due to the inherent complexity involved in modeling such data. In this work, we present a new private generative modeling approach where samples are generated via Hamiltonian dynamics with gradients of the private dataset estimated by a well-trained network. In the approach, we achieve differential privacy by perturbing the projection vectors in the estimation of gradients with sliced score matching. In addition, we enhance the reconstruction ability of the model by incorporating a residual enhancement module during the score matching. For sampling, we perform Hamiltonian dynamics with gradients estimated by the well-trained network, allowing the sampled data close to the private dataset's manifold step by step. In this way, our model is able to generate data with a resolution of 256×256. Extensive experiments and analysis clearly demonstrate the effectiveness and rationality of the proposed approach.
Bochao Liu, Weijia Guo, Liansheng Zhuang, Weiping Wang 0005, Shiming Ge
ACM Multimedia5
2024 Fact Embedding through Diffusion Model for Knowledge Graph Completion
abstract
Knowledge graph embedding (KGE) is an efficient and scalable method for knowledge graph completion tasks. Existing KGE models typically map entities and relations into a unified continuous vector space and define a score function to capture the connectivity patterns among the elements (entities and relations) of facts. The score on a fact measures its plausibility in a knowledge graph (KG). However, since the connectivity patterns are very complex in a real knowledge graph, it is difficult to define an explicit and efficient score function to capture them, which also limits their performance. This paper argues that plausible facts in a knowledge graph come from a distribution in the low-dimensional fact space. Inspired by this insight, this paper proposes a novel framework called Fact Embedding through Diffusion Model (FDM) to address the knowledge graph completion task. Instead of defining a score function to measure the plausibility of facts in a knowledge graph, this framework directly learns the distribution of plausible facts from the known knowledge graph and casts the entity prediction task into the conditional fact generation task. Specifically, we concatenate the elements embedding in a fact as a whole and take it as input. Then, we introduce a Conditional Fact Denoiser to learn the reverse denoising diffusion process and generate the target fact embedding from noised data. Extensive experiments demonstrate that FDM significantly outperforms existing state-of-the-art methods in three benchmark datasets.
Liansheng Zhuang, Aodi Li, Houqiang Li, Shafei Wang
WWW2
2024 Hierarchical Augmentation and Distillation for Class Incremental Audio-Visual Video Recognition
abstract
Audio-visual video recognition (AVVR) integrates audio and visual cues to accurately categorize videos. While current methods using provided datasets achieve satisfactory results, they face challenges in retaining historical class knowledge when new classes appear in real-world situations. There are no dedicated methods to address this issue, prompting this paper to explore Class Incremental Audio-Visual Video Recognition (CIAVVR). CIAVVR aims to preserve historical knowledge contained in stored data and learned models to prevent catastrophic forgetting. Audio-visual data and models inherently have hierarchical structures, where the model contains both low-level and high-level semantic information, and data includes snippet-level, video-level, and distribution-level spatial information. It is crucial to fully exploit these hierarchical structures for data knowledge preservation and model knowledge preservation. However, existing image class incremental learning methods do not explicitly consider these hierarchical structures. Therefore, we introduce Hierarchical Augmentation and Distillation (HAD), which includes the Hierarchical Augmentation Module (HAM) and Hierarchical Distillation Module (HDM). These modules efficiently utilize the hierarchical structure of data and models. Specifically, HAM uses a novel augmentation strategy, segmental feature augmentation, to preserve hierarchical model knowledge. Simultaneously, HDM employs newly designed hierarchical logical distillation (video-distribution) and hierarchical correlative distillation (snippet-video) to maintain intra-sample and inter-sample hierarchical knowledge. Evaluations on four benchmarks (AVE, AVK-100, AVK-200, and AVK-400) show that HAD effectively captures hierarchical information, enhancing the preservation of historical class knowledge and performance. We also provide a theoretical analysis to support the segmental feature augmentation strategy.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 A Robust Framework for One-Shot Key Information Extraction via Deep Partial Graph Matching
abstract
Text field labelling plays a key role in Key Information Extraction (KIE) from structured document images. However, existing methods ignore the field drift and outlier problems, which limit their performance and make them less robust. This paper casts the text field labelling problem into a partial graph matching problem and proposes an end-to-end trainable framework called Deep Partial Graph Matching (dPGM) for the one-shot KIE task. It represents each document as a graph and estimates the correspondence between text fields from different documents by maximizing the graph similarity of different documents. Our framework obtains a strict one-to-one correspondence by adopting a combinatorial solver module with an extra one-to-(at most)-one mapping constraint to do the exact graph matching, which leads to the robustness of the field drift problem and the outlier problem. Finally, a large one-shot KIE dataset named DKIE is collected and annotated to promote research of the KIE task. This dataset will be released to the research and industry communities. Extensive experiments on both the public and our new DKIE datasets show that our method can achieve state-of-the-art performance and is more robust than existing methods.
Minghong Yao, Liansheng Zhuang, Liangwei Wang 0004, Houqiang Li
IEEE Trans. Image Process.3
2023 LayoutDM: Transformer-based Diffusion Model for Layout Generation
abstract
Automatic layout generation that can synthesize high-quality layouts is an important tool for graphic design in many applications. Though existing methods based on generative models such as Generative Adversarial Networks (GANs) and Variational Auto-Encoders (VAEs) have progressed, they still leave much room for improving the quality and diversity of the results. Inspired by the recent success of diffusion models in generating high-quality images, this paper explores their potential for conditional layout generation and proposes Transformer-based Layout Diffusion Model (LayoutDM) by instantiating the conditional denoising diffusion probabilistic model (DDPM) with a purely transformer-based architecture. Instead of using convolutional neural networks, a transformer-based conditional Layout Denoiser is proposed to learn the reverse diffusion process to generate samples from noised layout data. Benefitting from both transformer and DDPM, our LayoutDM is of desired properties such as high-quality generation, strong sample diversity, faithful distribution coverage, and stationary training in comparison to GANs and VAEs. Quantitative and qualitative experimental results show that our method outperforms state-of-the-art generative models in terms of quality and diversity.
Shang Chai, Liansheng Zhuang, Fengying Yan
CVPR2
2023 Bayesian Sharpness-Aware Prompt Tuning for Cross-Domain Few-shot Learning
abstract
Few-shot learning aims to learn a classifier to recognize novel classes with only few labeled images in each class. Fine-tuning is a promising tool to solve the few-shot learning problem, which pre-trains a large-scale model on source domains and then adapts it to target domains. However, existing methods have poor generalization when encountering the domain-shifting problem in the cross-domain scenario. Inspired by recent advances on domain generalization and prompt-based tuning methods, this paper proposes Bayesian Sharpness-Aware Prompt Tuning (BSAPT) for the cross-domain few-shot learning task. Instead of learning deterministic prompts like existing methods, our BSAPT learns a weight distribution over prompts to model the uncertainty caused by limited training data and resist overfitting. To improve the generalization ability, our BSAPT seeks the prompts which lie in neighborhoods having uniformly low loss by simultaneously minimizing the training loss value and loss sharpness. Benefiting from deterministic pre-trained training and Bayesian inference, our BSAPT has better generalization ability and less overfitting than existing fine-tuning methods. Extensive experiments on public datasets show that our BSAPT outperforms state-of-the-art methods and achieves new state-of-the-art performance in the cross-domain few-shot learning task.
Shuo Fan, Liansheng Zhuang, Aodi Li
IJCNN2
2023 Curriculum Learning Based Multi-Agent Path Finding for Complex Environments
abstract
Multi-agent reinforcement learning (MARL) is a promising tool to solve the Multi-Agent Path Finding (MAPF) task, which aims to find conflict-free paths for multiple agents, one for each agent, from a start position to a goal position. It uses global information to learn a mechanism for cooperation among agents by maximising the cumulative team rewards, which are often very sparse. However, the sparsity of rewards implies that agents have to blindly explore all possible paths, which makes MARL methods difficult to converge in complex environments. To address this issue, this paper proposes a novel Curriculum based Path-finding Learning (CPL) under the framework of curriculum learning, which allows agents to start with simple skills and to learn cooperative strategies stage-by-stage for more efficient training. Specifically, CPL divides the training process into three stages and speeds up the learning by changing the difficulty of the tasks from easy to hard. Experiments on random obstacle grid worlds show that our proposed method performs significantly better in terms of success rate and makespan than state-of-the-art learning-based methods.
Liansheng Zhuang
IJCNN2
2023 Two-stage Content-Aware Layout Generation for Poster Designs
abstract
Automatic layout generation models can generate numerous design layouts in a few seconds, which significantly reduces the amount of repetitive work for designers. However, most of these models consider the layout generation task as arranging layout elements with different attributes on a blank canvas, thus struggle to handle the case when an image is used as the layout background. Additionally, existing layout generation models often fail to incorporate explicit aesthetic principles such as alignment and non-overlap, and neglect implicit aesthetic principles which are hard to model. To address these issues, this paper proposes a two-stage content-aware layout generation framework for poster layout generation. Our framework consists of an aesthetics-conditioned layout generation module and a layout ranking module. The diffusion model based layout generation module utilizes an aesthetics-guided layout denoising process to sample layout proposals that meet explicit aesthetic constraints. The Auto-Encoder based layout ranking module then measures the distance between those proposals and real designs to determine the layout that best meets implicit aesthetic principles. Quantitative and qualitative experiments demonstrate that our method outperforms state-of-the-art content-aware layout generation models.
Shang Chai, Liansheng Zhuang, Fengying Yan, Zihan Zhou 0001
ACM Multimedia2
2023 Learning Generalized Representations for Open-Set Temporal Action Localization
abstract
Open-set Temporal Action Localization (OSTAL) is a critical and challenging task that aims to recognize and temporally localize human actions in untrimmed videos in open word scenarios. The main challenge in this task is the knowledge transfer from known actions to unknown actions. However, existing methods utilize limited training data and overparameterized deep neural network, which have poor generalization. This paper proposes a novel Generalized OSTAL model (namely GOTAL) to learn generalized representations of actions. GOTAL utilizes a Transformer network to model actions and a open-set detection head to perform action localization and recognition. Benefitting from Transformer's temporal modeling capabilities, GOTAL facilitates the extraction of human motion information from videos to mitigate the effects of irrelevant background data. Furthermore, a sharpness minimization algorithm is used to learn the network parameters of GOTAL, which facilitates the convergence of network parameters towards flatter minima by simultaneously minimizing the training loss value and sharpness of the loss plane. The collaboration of the above components significantly enhances the generalization of the representation. Experimental results demonstrate that GOTAL achieves the state-of-the-art performance on THUMOS14 and ActivityNet1.3 benchmarks, confirming the effectiveness of our proposed method.
Junshan Hu, Liansheng Zhuang, Weisong Dong, Shiming Ge, Shafei Wang
ACM Multimedia2
2023 Learning a Robust Model with Pseudo Boundaries for Noisy Temporal Action Localization
abstract
Temporal Action Localization (TAL) aims to locate starting and ending times of actions and recognize categories in untrimmed videos. Significant progress has been made in developing deep models for TAL. The success of previous methods relies on large-scale training data with precise boundary annotations. However, fully accurate annotations are unpractical to be obtained due to the ambiguities of the action boundaries and the crowd-sourcing labeling process, leading to a degradation in performance. In this work, we take the first step into learning with inaccurate boundaries in TAL tasks. Motivated by the fact that inaccurate boundary annotations harm localization precision more than classification accuracy, we propose to use classification as a guidance signal to improve localization precision. Specifically, we introduce a pseudo-boundary generation and refinement method (PbGaR). PbGaR first treats each action segment as a bag of instances to select the instances with more accurate boundaries for training. Then these boundaries are refined via two strategies for higher quality. The proposed method significantly alleviates the degraded performance of TAL models under inaccurate boundaries. Extensive experiments on two popular datasets demonstrate the effectiveness of our method.
Liansheng Zhuang
MMAsia2
2023 Dual Structural Knowledge Interaction for Domain Adaptation
abstract
Domain adaptation aims to transfer knowledge from a label-rich source domain to an unlabeled target domain. A common strategy is to assign pseudo-labels to unlabeled target samples for performing representation learning. However, most existing methods only apply the source-guided classifier to generate the source-biased pseudo-labels for self-training, leading to biased target representations. Moreover, the generated pseudo-labels ignore the manifold assumption that neighboring samples are likely to have the same labels. To address the above problem, we formulate a novel structural knowledge to assign target-oriented and manifold-guided pseudo-labels for unlabeled target samples. The structural knowledge consists of cluster-based knowledge and locality-based knowledge. The cluster-based knowledge denotes the label consistency between the target samples and the non-parametric target cluster centers, making the pseudo-labels target-oriented. The locality-based knowledge constrains the target sample and its neighbors to satisfy the manifold assumption. As the neighbors contain the source and target samples, the source and target locality-based knowledge are utilized to boost the descriptions. With the structural knowledge, we propose a novel Dual Structural Knowledge Interaction (DSKI) framework for domain adaptation. For generating aligned and discriminative features, knowledge consistency constraint and instance mutual constraint are proposed in DSKI. Evaluations on three benchmarks demonstrate the effectiveness of the Dual Structural Knowledge Interaction,e.g.,74.9%, 87.7%, and 90.8% for Office-Home, VisDa-2017, and Office-31, respectively.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Multim.3
2022 Learning Common and Specific Visual Prompts for Domain Generalization
Aodi Li, Liansheng Zhuang, Shuo Fan, Shafei Wang
ACCV (6)2
2022 HaViT: Hybrid-Attention Based Vision Transformer for Video Classification
Liansheng Zhuang, Shenghua Gao, Shafei Wang
ACCV (4)2
2022 Neural-based Mixture Probabilistic Query Embedding for Answering FOL queries on Knowledge Graphs
abstract
Query embedding (QE)-which aims to embed entities and first-order logical (FOL) queries in a vector space, has shown great power in answering FOL queries on knowledge graphs (KGs).Existing QE methods divide a complex query into a sequence of mini-queries according to its computation graph and perform logical operations on the answer sets of mini-queries to get answers.However, most of them assume that answer sets satisfy an individual distribution (e.g., Uniform, Beta, or Gaussian), which is often violated in real applications and limit their performance.In this paper, we propose a Neural-based Mixture Probabilistic Query Embedding Model (NMP-QEM) that encodes the answer set of each mini-query as a mixed Gaussian distribution with multiple means and covariance parameters, which can approximate any random distribution arbitrarily well in real KGs.Additionally, to overcome the difficulty in defining the closed solution of negation operation, we introduce neural-based logical operators of projection, intersection and negation for a mixed Gaussian distribution to answer all the FOL queries.Extensive experiments demonstrate that NMP-QEM significantly outperforms existing stateof-the-art methods on benchmark datasets.In NELL995, NMP-QEM achieves a 31% relative improvement over the state-of-the-art.
Liansheng Zhuang, Aodi Li, Shafei Wang, Houqiang Li
EMNLP2
2022 Efficient Double Oracle for Extensive-Form Two-Player Zero-Sum Games
Liansheng Zhuang
ICONIP (2)2
2022 VAAC: V-value Attention Actor-Critic for Cooperative Multi-agent Reinforcement Learning
Liansheng Zhuang
ICONIP (1)2
2022 Estimation of Reliable Proposal Quality for Temporal Action Detection
abstract
Temporal action detection (TAD) aims to locate and recognize the actions in an untrimmed video. Anchor-free methods have made remarkable progress which mainly formulate TAD into two tasks: classification and localization using two separate branches. This paper reveals the temporal misalignment between the two tasks hindering further progress. To address this, we propose a new method that gives insights into moment and region perspectives simultaneously to align the two tasks by acquiring reliable proposal quality. For the moment perspective, Boundary Evaluate Module (BEM) is designed which focuses on local appearance and motion evolvement to estimate boundary quality and adopts a multi-scale manner to deal with varied action durations. For the region perspective, we introduce Region Evaluate Module (REM) which uses a new and efficient sampling method for proposal feature representation containing more contextual information compared with point feature to refine category score and proposal boundary. The proposed B oundary Evaluate Module and R egion E valuate M odule (BREM) are generic, and they can be easily integrated with other anchor-free TAD methods to achieve superior performance. In our experiments, BREM is combined with two different frameworks and improves the performance on THUMOS14 by 3.6% and 1.0% respectively, reaching a new state-of-the-art (63.6% average mAP). Meanwhile, a competitive result of 36.2% average mAP is achieved on ActivityNet-1.3 with the consistent improvement of BREM. The codes are released at \urlhttps://github.com/Junshan233/BREM.
Junshan Hu, Chaoxu Guo, Liansheng Zhuang, Tiezheng Ge, Yuning Jiang 0001, Houqiang Li
ACM Multimedia3
2022 MEViT: Motion Enhanced Video Transformer for Video Classification
Liansheng Zhuang
MMM (2)2
2022 Question-relationship guided graph attention network for visual question answer
Liansheng Zhuang, Zhou Yu 0001, Tian Bai 0005
Multim. Syst.2
2022 PMIVec: a word embedding model guided by point-wise mutual information criterion
Minghong Yao, Liansheng Zhuang, Shafei Wang, Houqiang Li
Multim. Syst.2
2022 Margin-Based Adversarial Joint Alignment Domain Adaptation
abstract
Domain adaptation aims to transfer the knowledge learned from a labeled source domain to an unlabeled target domain, which has different data distribution with the source domain. Most of the existing methods focus on aligning the data distribution between the source and target domains but ignore the discrimination of the feature space among categories, leading the samples close to the decision boundary to be misclassified easily. To address the above issue, we propose a Margin-based Adversarial Joint Alignment (MAJA) to constrain the feature spaces of source and target domains to be aligned and discriminative. The proposed MAJA consists of two components: joint alignment module and margin-based generative module. The joint alignment module is proposed to align the source and target feature spaces by considering the joint distribution of features and labels. Therefore, the embedding features and the corresponding labels treated as pair data are applied for domain alignment. Furthermore, the margin-based generative module is proposed to boost the discrimination of the feature space,i.e.,make all samples as far away from the decision boundary as possible. The margin-based generative module first employs the Generative Adversarial Networks (GAN) to generate a lot of fake images for each category, then applies the adversarial learning to enlarge and reduce the category margin for the true images and generated fake images, respectively. The evaluations on three benchmarks,e.g.,small image datasets, VisDA-2017, and Office-31, verify the effectiveness of the proposed method.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.3
2022 Seek Common Ground While Reserving Differences: A Model-Agnostic Module for Noisy Domain Adaptation
abstract
Noisy domain adaptation aims to solve the problem that the source dataset contains noisy labels in domain adaptation. Previous methods handle noisy labels by selecting the small-loss samples with inconsistent predictions between two models and discarding the consistent samples, resulting in many noises contained in the selected samples. By jointly considering the consistent and inconsistent samples, we propose a model-agnostic module, named Seek Common Ground While Reserving Differences (SCGWRD), to reduce the impact of noisy samples. The proposed SCGWRD module consists of Seek Common Ground (SCG) component and Reserve Differences (RD) component by utilizing the outputs of two symmetrical domain adaptation models. As the common samples with consistent predictions between two models are more likely to be clean samples, the SCG component applies the small-loss strategy to select the reliable samples with consistent predictions. Unlike SCG, the RD component maintains the divergences between two models with mutual learning and reduces the effect of noisy data using the samples with different predictions and small losses. Evaluations on three benchmarks demonstrate the effectiveness and robustness of the proposed SCGWRD module for noisy domain adaptation.
Yukun Zuo, Hantao Yao, Liansheng Zhuang, Changsheng Xu
IEEE Trans. Multim.3
2021 Path Ranking Model for Entity Prediction
abstract
Knowledge graphs (KGs) often encounter knowledge incompleteness, necessitating a demand for KG completion. Path-based methods are one of the most important approaches to this task. However, since the number of entities is much larger than that of relations in a knowledge graph, existing path-based methods are only used to predict the relations between entity pairs, and are rarely applied to solve the entity prediction task. To address the issue, this paper proposes a new framework called Path Ranking Model (PRM) for the knowledge graph completion task. Our key idea is to exploit both the observable patterns and latent semantic information in relation paths to predict the entities. Extensive experiments on public popular datasets demonstrate the effectiveness of our proposed framework in the entity prediction task.
Minghong Yao, Liansheng Zhuang, Houqiang Li, Shafei Wang
ICME3
2021 Learn the Approximation Distribution of Sparse Coding with Mixture Sparsity Network
Liansheng Zhuang, Shafei Wang
PRCV (4)3
2020 Time-Sensitive Collaborative Interest Aware Model for Session-Based Recommendation
abstract
Session-based recommendation, which aims at predicting user's action based on anonymous sessions, is a challenging problem due to the uncertainty of user's behaviors and limited clicked information. Existing methods model users' interests to relieve the uncertainty of user's behavior prediction. However, most methods mainly focus on the current session, ignoring collaborative information (i.e., collaborative interest) in neighborhood sessions with similar interests. We argue that relying on limited implicit feedbacks within a session is insufficient to precisely infer user's interest, especially in the absence of user's profiles and historical behaviors. This paper proposes a novel model called Time-Sensitive Collaborative Interest Aware (TSCIA) to tackle this problem. It explicitly aggregates similar interests from neighborhood sessions to model the general collaborative interest, and simultaneously takes users' interest drifts into account. Finally, both current session and collaborative information are used for next-item prediction. Extensive experiments on public datasets demonstrate the effectiveness of our model.
Liansheng Zhuang, Pengyu Luo, Houqiang Li, Zhengjun Zha
ICME2
2020 Phrase-Level Global-Local Hybrid Model For Sentence Embedding
abstract
Latent structure models have drawn much attention due to the ability to learn an optimal latent hierarchical structures without explicit structure annotations. However, most existing models suffer from high computation complexity and hard training. To this end, this paper proposes a novel phrase-level global-local hybrid model, which inherits the advantages of existing latent structure models while requires less time complexity. Our model splits a sentence into multiple phrases by a category-selection module. Then, it encodes the context dependency by a phrase-level global encoding module, and encodes the task-specific information by a phrase-level local encoding module. Finally, sentence embedding is obtained by integrating the global encoding and task-specific encoding. Experiments on public benchmarks show that, our model achieves state-of-the-art performance on the tasks of sentence classification and natural language inference. Meanwhile, our model is at least 10 times faster than existing state-of-the-art method at the training stage.
Mingyu Tang, Liansheng Zhuang, Houqiang Li, Yanqun Guo
ICME2
2020 Quantile Regression Hindsight Experience Replay
Qiwei He, Liansheng Zhuang, Houqiang Li
ICONIP (4)2
2019 Continuous Sign Language Recognition via Reinforcement Learning
abstract
In this paper, we propose an approach to apply the Transformer with reinforcement learning (RL) for continuous sign language recognition (CSLR) task. The Transformer has an encoder-decoder structure, where the encoder network encodes the sign video into the context vector representation, while the decoder network generates the target sentence word by word based on the context vector. To avoid the intrinsic defects of supervised learning (SL) in our task, e.g., the exposure bias and non-differentiable task metrics issues, we propose to train the Transformer directly on non-differentiable metrics, i.e., word error rate (WER), through RL. Moreover, a policy gradient algorithm with baseline, which we call Self-critic REINFORCE, is employed to reduce variance while training. Experimental results on RWTH-PHOENIX-Weather benchmark verify the effectiveness of our method and demonstrate that our method achieves the comparable performance.
Junfu Pu, Liansheng Zhuang, Wengang Zhou 0001, Houqiang Li
ICIP3
2019 Learning Motion-Aware Policies for Robust Visual Tracking
abstract
Visual object tracking aims to locate a moving target specified at the initial frame. Although this task is closely related to the temporal motion information, the motion model typically draws limited attention. In this paper, we propose a motion-aware multi-domain network for robust visual tracking. In our approach, a motion-aware agent is trained via reinforcement learning, which can infer the parameters of the particle filter in a continuous action space. Different from existing tracking-by-detection frameworks that the particle filter merely relies on the previous target state, our motion-aware agent, after receiving the current state, can adaptively change the parameters of the particle filter (e.g., particle location and scale range). As a result, our approach samples high-quality candidates for further classification/tracking, thus can better handle challenges such as fast motion and scale variation. Extensive experiments on large-scale benchmarks verify the effectiveness of our method.
Liansheng Zhuang, Ning Wang 0020, Wengang Zhou 0001, Houqiang Li
ICME2
2019 Dynamic Cascaded Regression Network with Reinforcement Learning for Robust Face Alignment
abstract
Regression-based methods for facial landmark detection usually learn a series of regressors to update the landmark positions from an initial shape with a fixed number of iterations. Their accuracy is sensitive to the initial shape, and the fixed number of iterations always leads to massive unnecessary computation. In this paper, we propose a Dynamic Cascaded Regression Network (DCRN) with a two-stage architecture to address these issues. In the first stage, we introduce a Global Estimation Network (GEN) to provide a coarse landmark estimation. In the second stage, we propose a Local Regression Network (LRN) to iteratively refine the coarse estimation in a reinforcement learning (RL) paradigm. Our DCRN takes the face image as input, and adaptively learns facial landmarks. Extensive experiments on 300W, COFW, and AFLW datasets show the effectiveness of our proposed method and demonstrate that DCRN consistently achieves the state-of-the-art performance.
Liansheng Zhuang, Wengang Zhou 0001, Houqiang Li
ICME2
2018 CCNet: Cluster-Coordinated Net for Learning Multi-agent Communication Protocols with Reinforcement Learning
abstract
Multi-agent system is crucial for many practical applications. Recent years have witnessed numerous research on multi-agent task with reinforcement learning (RL) algorithms. Traditional reinforcement learning algorithms often fail to learn the cooperation between different agents, which is vital for multi-agent problems. A promising solution is to establish a communication protocol among agents. However, existing approaches often suffer from generalization challenges especially in tasks with partial observation and dynamic variation of agent amount. In this paper, we develop a Cluster-Coordinated Network (CCNet) to address the “Learning-to-communicate” problem in multi-agent system by utilizing the combination of a trainable Vector of Locally Aggregated Descriptor (VLAD) algorithm and reinforcement learning. Embedding with a VLAD based end-to-end trainable communication information processing module (called VLAD Processing Core), CCNet can learn efficient communication protocols even from scratch under partially observable environments and possesses robustness to the dynamic changes of agent number as well. Moreover, with the help of communication, CCNet is with less non-stationarity when training the network by common RL algorithms. We evaluated the proposed CCNet on two multi-agent partially observable tasks, \emph{i.e.}, Traffic Junction and Combat Task. The experimental results have demonstrated that CCNet is effective and improves the performance by a large margin over the state-of-the-art methods.
Zhengjun Zha, Zilei Wang, Liansheng Zhuang, Houqiang Li
ACML4
2018 A Feature-Adaptive Semi-Supervised Framework for Co-saliency Detection
abstract
Co-saliency detection, which refers to the discovery of common salient foreground regions in a group of relevant images, has attracted increasing attention due to its widespread applications in many vision tasks. Existing methods assemble features from multiple views toward a comprehensive representation, however overlook the efficacy disparity among various features in detecting co-saliency. This paper proposes a novel feature-adaptive semi-supervised (FASS) framework for co-saliency detection, which seamlessly integrates multi-view feature learning, graph structure optimization and co-saliency prediction in a unified solution. In particular, the FASS exploits the efficacy disparity of multi-view features at both view and element levels by a joint formulation of view-wise feature weighting and element-wise feature selection, leading to an effective representation robust to feature noise and redundancy as well as adaptive to the task at hand. It predicts co-saliency map by optimizing co-saliency label prorogation over a graph of both labeled and unlabeled image regions. The graph structure is optimized jointly with feature learning and co-saliency prediction to precisely characterize underlying correlation among regions. The FASS is thus able to produce satisfactory co-saliency map based on the effective exploration of multi-view features as well as inter-region correlation. Extensive experiments on three benchmark datasets, i.e., iCoseg, Cosal2015 and MSRC, have demonstrated that the proposed FASS outperforms the state-of-the-art methods.
Xiaoju Zheng, Zhengjun Zha, Liansheng Zhuang
ACM Multimedia3
2017 Label Information Guided Graph Construction for Semi-Supervised Learning
abstract
In the literature, most existing graph-based semi-supervised learning methods only use the label information of observed samples in the label propagation stage, while ignoring such valuable information when learning the graph. In this paper, we argue that it is beneficial to consider the label information in the graph learning stage. Specifically, by enforcing the weight of edges between labeled samples of different classes to be zero, we explicitly incorporate the label information into the state-of-the-art graph learning methods, such as the low-rank representation (LRR), and propose a novel semi-supervised graph learning method called semi-supervised low-rank representation. This results in a convex optimization problem with linear constraints, which can be solved by the linearized alternating direction method. Though we take LRR as an example, our proposed method is in fact very general and can be applied to any self-representation graph learning methods. Experiment results on both synthetic and real data sets demonstrate that the proposed graph learning method can better capture the global geometric structure of the data, and therefore is more effective for semi-supervised learning tasks.
Liansheng Zhuang, Zihan Zhou 0001, Shenghua Gao, Jingwen Yin, Zhouchen Lin, Yi Ma 0001
IEEE Trans. Image Process.1
2016 Part-based multi-graph ranking for visual tracking
abstract
Recently, graph ranking-based methods have been introduced to visual tracking and achieved promising results due to the local structure preserving property. However, existing graph ranking-based trackers use holistic templates to construct the graphs which makes the trackers sensitive to occlusions. In this paper, we propose a part-based multi-graph ranking algorithm for robust visual tracking. In our method, template samples are divided into local parts. Multiple graphs are constructed based on different part samples and different feature representations. Then, the multiple graphs are integrated into a regularization framework with each graph assigned a weight. Furthermore, by imposing the l2,1 norm on the weight matrix of graphs, the confident parts are selected to reduce the effects of occluded ones. An effective optimization scheme is proposed to learn the weight matrix and the rank scores jointly. Experimental results on various challenging video sequences demonstrate our proposed algorithm outperforms state-of-the-art trackers.
Jingjing Wang 0005, Chi Fei, Liansheng Zhuang, Nenghai Yu
ICIP3
2016 Multi-level visual tracking with hierarchical tree structural constraint
Jingjing Wang 0005, Nenghai Yu, Feng Zhu 0006, Liansheng Zhuang
Neurocomputing4
2016 Locality-preserving low-rank representation for graph construction from nonlinear manifolds
Liansheng Zhuang, Jingjing Wang 0005, Zhouchen Lin, Allen Y. Yang, Yi Ma 0001, Nenghai Yu
Neurocomputing1
2015 Neither Global Nor Local: Regularized Patch-Based Representation for Single Sample Per Person Face Recognition
Shenghua Gao, Kui Jia, Liansheng Zhuang, Yi Ma 0001
Int. J. Comput. Vis.3
2015 Sparse Illumination Learning and Transfer for Single-Sample Face Recognition with Image Corruption and Misalignment
Liansheng Zhuang, Tsung-Han Chan, Allen Y. Yang, S. Shankar Sastry, Yi Ma 0001
Int. J. Comput. Vis.1
2015 Constructing a Nonnegative Low-Rank and Sparse Graph With Data-Adaptive Features
abstract
This paper aims at constructing a good graph to discover the intrinsic data structures under a semisupervised learning setting. First, we propose to build a nonnegative low-rank and sparse (referred to as NNLRS) graph for the given data representation. In particular, the weights of edges in the graph are obtained by seeking a nonnegative low-rank and sparse reconstruction coefficients matrix that represents each data sample as a linear combination of others. The so-obtained NNLRS-graph captures both the global mixture of subspaces structure (by the low-rankness) and the locally linear structure (by the sparseness) of the data, hence it is both generative and discriminative. Second, as good features are extremely important for constructing a good graph, we propose to learn the data embedding matrix and construct the graph simultaneously within one framework, which is termed as NNLRS with embedded features (referred to as NNLRS-EF). Extensive NNLRS experiments on three publicly available data sets demonstrate that the proposed method outperforms the state-of-the-art graph construction method by a large margin for both semisupervised classification and discriminative analysis, which verifies the effectiveness of our proposed method.
Liansheng Zhuang, Shenghua Gao, Jinhui Tang 0001, Jingjing Wang 0005, Zhouchen Lin, Yi Ma 0001, Nenghai Yu
IEEE Trans. Image Process.1
2014 Unsupervised Feature Learning for RGB-D Image Classification
I-Hong Jhuo, Shenghua Gao, Liansheng Zhuang, D. T. Lee, Yi Ma 0001
ACCV (1)3
2013 Single-Sample Face Recognition with Image Corruption and Misalignment via Sparse Illumination Transfer
abstract
Single-sample face recognition is one of the most challenging problems in face recognition. We propose a novel face recognition algorithm to address this problem based on a sparse representation based classification (SRC) framework. The new algorithm is robust to image misalignment and pixel corruption, and is able to reduce required training images to one sample per class. To compensate the missing illumination information typically provided by multiple training images, a sparse illumination transfer (SIT) technique is introduced. The SIT algorithms seek additional illumination examples of face images from one or more additional subject classes, and form an illumination dictionary. By enforcing a sparse representation of the query image, the method can recover and transfer the pose and illumination information from the alignment stage to the recognition stage. Our extensive experiments have demonstrated that the new algorithms significantly outperform the existing algorithms in the single-sample regime and with less restrictions. In particular, the face alignment accuracy is comparable to that of the well-known Deformable SRC algorithm using multiple training images, and the face recognition accuracy exceeds those of the SRC and Extended SRC algorithms using hand labeled alignment initialization.
Liansheng Zhuang, Allen Y. Yang, Zihan Zhou 0001, S. Shankar Sastry, Yi Ma 0001
CVPR1
2013 Regularized Semi-Supervised Latent Dirichlet Allocation for visual concept learning
Liansheng Zhuang, Haoyuan Gao, Jiebo Luo 0001, Zhouchen Lin
Neurocomputing1
2012 Non-negative low rank and sparse graph for semi-supervised learning
abstract
Constructing a good graph to represent data structures is critical for many important machine learning tasks such as clustering and classification. This paper proposes a novel non-negative low-rank and sparse (NNLRS) graph for semi-supervised learning. The weights of edges in the graph are obtained by seeking a nonnegative low-rank and sparse matrix that represents each data sample as a linear combination of others. The so-obtained NNLRS-graph can capture both the global mixture of subspaces structure (by the low rankness) and the locally linear structure (by the sparseness) of the data, hence is both generative and discriminative. We demonstrate the effectiveness of NNLRS-graph in semi-supervised classification and discriminative analysis. Extensive experiments testify to the significant advantages of NNLRS-graph over graphs obtained through conventional means.
Liansheng Zhuang, Haoyuan Gao, Zhouchen Lin, Yi Ma 0001, Xin Zhang 0051, Nenghai Yu
CVPR1
2011 A New Graph Constructor for Semi-supervised Discriminant Analysis via Group Sparsity
abstract
Semi-supervised dimensionality reduction is very important in mining high-dimensional data due to the lack of costly labeled data. This paper studies the Semi-supervised Discriminant Analysis (SDA) algorithm, which aims at dimensionality reduction utilizing both limited labeled data and abundant unlabeled data. Different from other relative work, we pay our attention to graph construction, which plays a key role in graph based SSL methods. Inspired by the advances of compressive sensing, we propose a novel graph construction method via group sparsity, which means to constrain the reconstruct data to be sparse for each sample, and constrain the representation in each class to be quite similar. Experimental results show that our method can significantly improve the performance of SDA, and outperform state-of-the-art methods.
Haoyuan Gao, Liansheng Zhuang, Nenghai Yu
ICIG2
2011 Dynamic Background Subtraction Using Spatial-Color Binary Patterns
abstract
In this paper, an efficient approach for background modeling and subtraction is proposed. It's based on a novel spatial-color feature extraction operator named spatial-color binary patterns(SCBP). As the name implies, features extracted by this operator include spatial texture and color information. In addition, a refine module is designed to refine the contour of moving objects. Using the proposed method, we improve the accuracy of subtracting the background and detecting moving objects in dynamic scenes. A data-driven model is used in our method. For each pixel, first, a histogram of SCBP is extracted from the circular egion, and then a model consist of several histograms is built. For a new observed frame, each pixel is labeled either background or foreground according to the matching degree between its SCBP histogram and its model, then the label is refined and finally the model of this pixel is updated. The proposed approach is tested on challenging video sequences, which shows that the proposed method performs much better than several texture-based methods.
Weiming Zhang 0001, Liansheng Zhuang, Nenghai Yu
ICIG4
2011 Semi-supervised Classification via Low Rank Graph
abstract
Graph plays a very important role in graph based semi-supervised learning (SSL) methods. However, most current graph construction methods emphasize on local properties of the graph. In this paper, inspired by the advances of compressive sensing, we present a novel method to construct a so-called low-rank graph (LR-graph) for graph based SSL methods. Assuming that the graph is sparse and low rank, our proposed method uses both the local property and the global property of the graph, and thus is better at capturing the global structure of all data. Compared with current graphs, LR-graph is more informative and discriminative, and robust to outliers. Experiments on generic object recognition show that LR-graph achieves state-of-the-art performance for graph based SSL methods.
Liansheng Zhuang, Haoyuan Gao, Nenghai Yu
ICIG1
2011 Regularized Semi-supervised Latent Dirichlet Allocation for Visual Concept Learning
Liansheng Zhuang, Lanbo She, Jiebo Luo 0001, Nenghai Yu
MMM (1)1
2010 SVD based linear filtering in DCT domain
abstract
Efficient linear filtering in DCT domain is important in the area of processing and manipulation of image and video streams compressed in DCT-based method. In this paper, we proposed a novel method for linear filtering in DCT domain, regardless of filter type. We decompose any filter by SVD into weighted separable sub-filters which are well studied. Then we do fast linear filtering using these separable subfilters in DCT domain, and combine their results. To our best knowledge, it is the first method capable to do linear filtering with any type of filters directly in DCT domain. The scheme is demonstrated and discussed by doing Gabor filtering in DCT domain. Experiment results show that convolution result using the proposed solution is the same as that in spatial domain. Furthermore, our scheme is well suitable for distributed computing, which will improve computing speed greatly.
Liansheng Zhuang, Rui Zhao 0001, Nenghai Yu, Bin Liu 0016
ICIP1
2010 A robust part-based tracker
abstract
In this paper, we propose a new method for modeling appearance variances in generic object tracking task. Although object tracking has been studied by many researchers for a long time, there are still many challenging problems, which is mainly due to the complex variances of object's appearance. While most of traditional methods using a global or pixel-wise approach, we proposed a part-based tracking framework. We divide an object region into several non-overlapping parts (note they are not semantic as limbs and head of a human), and then a local classifier is updated on-line for each part. We gain a global confidence map by applying these local classifiers to the next frame, and find the new location of target object, i.e. the peak of confidence map, using mean-shift. Our tracker runs real-time, and is robust to some kinds of appearance variance (e.g. change of illumination, occlusion, change of pose, deformation of shape, object/camera movement and so on). Experiments show that our method outperforms the other states of the art approaches, especially on dealing with occlusion.
Liansheng Zhuang, Nenghai Yu
ICME2
2009 Image Classification via Semi-supervised pLSA
abstract
In this paper, we propose Semi-Supervised pLSA(SSpLSA) for image classification. Compared with the classic non-supervised pLSA, our method overcomes the shortcoming of poor classification performance when the features of two categories are quite similar. By introducing category label information into EM algorithm, the iteration process can be directed carefully to the desired result. SS-pLSA greatly prevents the inter-impact between different categories. The experiment results show that the proposed SS-pLSA significantly improves the performance of image classification, especially when two categories' features are similar and difficult to distinguish by classic pLSA. In contrast to these totally supervised algorithm, SS-pLSA almost has no loss in detection rate while sharply reduces the difficulty of collecting training samples. With highly flexibility, SS-pLSA enables users to explore the trade-off between labeled number and accuracy.
Liansheng Zhuang, Lanbo She, Yuning Jiang 0001, Ketan Tang, Nenghai Yu
ICIG1
2009 Low-Resolution Face Recognition via Sparse Representation of Patches
abstract
Images resolution plays an important role during face recognition. Low-resolution face images will reduce drastically the performance of face recognition algorithms. In this paper, we propose a novel approach for low-resolution face recognition. Our method first exacts patches with different size from the face images. Each patch is represented by its LBP feature. Then, we find the sparse representation of these patches based on corresponding LBP features of high-resolution face image patches. At last, we use AdaBoost to select the most discriminative patches, each of which is treated as a weak classifier, and make the ensemble of these patches weak classifiers for final decision. Experiments on Extended Yale B face database showed our method achieved high performance for low-resolution face recognition.
Liansheng Zhuang, Mengliao Wang, Nenghai Yu, Yangchun Qian
ICIG1
2008 Learning object from small and imbalanced dataset with Boost-BFKO
abstract
One of the main drawbacks of boosting is its overfitting and poor predictive accuracy when the training dataset is small and imbalanced. In this paper, we introduce a novel learning algorithm Boost-BFKO, which combines boosting and data generation. It is suitable for small and imbalanced training datasets. To enlarge training sets, Boost-BFKO uses the adaptive Balanced Feature Knockout procedure (BFKO) to generate new synthetic samples. To enrich the training sets, Boost-BFKO selects seed samples from the minority class, and rebalances the total weights of the different classes in the updated training dataset. Experiments on Caltech 101 database showed that our method achieves a desirable performance when only a few training samples are available for binary classification and multiple object classification.
Liansheng Zhuang, Qi Tian 0001, Nenghai Yu
ICME1