EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyang Tan
dblp:79/768
· DBLP profile ↗
88ranked-venue papers
12as first author
31since 2021 · last 2026
0000-0002-2683-8667ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 72 · 8 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSecurity and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Variational OOD State Correction for Offline Reinforcement LearningabstractThe performance of Offline reinforcement learning is significantly impacted by the issue of state distributional shift, and out-of-distribution (OOD) state correction is a popular approach to address this problem. However, previous methods correct the agent's transition distributions in a supervised way, which significantly degrades the flexibility and robustness. In this paper, we propose a novel method named Density-Aware Safety Perception (DASP) for OOD state correction. Specifically, our method encourages the agent to prioritize actions that lead to outcomes with higher data density, thereby promoting its operation within or the return to in-distribution (safe) regions. To achieve this, we optimize the objective within a variational framework that concurrently considers both the potential outcomes of decision-making and their density, thus providing crucial contextual information for safe decision-making. Finally, we validate the effectiveness and feasibility of our proposed method through extensive experimental evaluations on the offline MuJoCo and AntMaze suites. Ke Jiang 0002, Xiaoyang Tan |
AAAI | 3 |
| 2026 | Beyond non-expert demonstrations: Outcome-driven action constraint for offline reinforcement learning
Ke Jiang 0002, Xiaoyang Tan |
Pattern Recognit. | 4 |
| 2026 | Toward Reliable Offline Reinforcement Learning via Lyapunov Uncertainty ControlabstractLearning trustworthy and reliable offline policies presents significant challenges due to the inherent uncertainty in pre-collected datasets. In this article, we propose a novel offline reinforcement learning (RL) method to tackle this issue. Inspired by the concepts of Lyapunov stability and control-invariant sets from control theory, the central idea is to introduce a restricted state space for the agent to operate within, which allows the learned models to exhibit reduced Bellman uncertainty and make reliable decisions. To achieve this, we regulate the expected Bellman uncertainty associated with the new policy, ensuring that its growth trend in subsequent states remains within acceptable limits. The resulting method, termed Lyapunov uncertainty control (LUC), is shown to guarantee that the agent remains within a low-uncertainty state enclosure throughout its entire trajectory. Furthermore, we perform extensive theoretical and experimental analysis to showcase the effectiveness and feasibility of the proposed LUC. Ke Jiang 0002, Xiaoyang Tan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2026 | ORAL: Adaptive Gap Increasing for Advantage Learning via Occam's Razor PrincipleabstractBenefiting from the gap increasing between the optimal action and its competitors, the advantage learning (AL) operator is more robust to estimation errors in the approximated $Q$ -functions than the Bellman optimality operator in reinforcement learning (RL). However, our analysis reveals that its robustness and larger action gaps come at the cost of a worse performance loss bound, leading to slower convergence of value functions. To address this issue, we present a novel method, named Occam's Razor-based AL (ORAL), which follows Occam's Razor principle and takes the necessity into consideration when increasing the action gap. Specifically, our ORAL can adaptively increase the action gap for different state-action pairs, depending on the proximity of their $Q$ values to the optimal ones. We first propose a naive implementation of ORAL, employing a nonsmooth clipping function to realize the above idea, and then introduce a smooth version of ORAL aimed at achieving more stable learning. Furthermore, our methods can be easily plugged into other AL-based operators and extended to more complex continuous-control tasks. Theoretical analysis supports the feasibility of our approaches, demonstrating their ability to balance the gap increasing with fast convergence. Empirical results further validate its effectiveness, showing significant performance improvements across multiple benchmarks. Yongle Zhou, Yuyang Long, Jia Zhang 0019, Juanjuan Weng, Zhetao Li, Yaozhong Gan, Xiaoyang Tan |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | RoGA: Towards Generalizable Deepfake Detection through Robust Gradient AlignmentabstractRecent advancements in domain generalization for deepfake detection have attracted significant attention, with previous methods often incorporating additional modules to prevent overfitting to domain-specific patterns. However, such regularization can hinder the optimization of the empirical risk minimization (ERM) objective, ultimately degrading model performance. In this paper, we propose a novel learning objective that aligns generalization gradient updates with ERM gradient updates. The key innovation is the application of perturbations to model parameters, aligning the ascending points across domains, which specifically enhances the robustness of deepfake detection models to domain shifts. This approach effectively preserves domain-invariant features while managing domain-specific characteristics, without introducing additional regularization. Experimental results on multiple challenging deepfake detection datasets demonstrate that our gradient alignment strategy outperforms state-of-the-art domain generalization techniques, confirming the efficacy of our method. The code is available at https://github.com/Lynn0925/RoGA. Lingyu Qiu, Ke Jiang 0002, Xiaoyang Tan |
ICME | 3 |
| 2025 | Candidate ratio guided proximal policy optimization
Xiaoyang Tan |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Boundary-aware adversarial ensemble learning for multivariate time series anomaly detection
Xiaoyang Tan, Yuehua Cheng |
Knowl. Based Syst. | 2 |
| 2025 | Highly valued subgoal generation for efficient goal-conditioned reinforcement learning
Yuhui Wang 0004, Xiaoyang Tan |
Neural Networks | 3 |
| 2025 | RLGrid: Reinforcement Learning Controlled Grid Deformation for Coarse-to-Fine Point Cloud CompletionabstractMany point cloud completion methods typically rely on two steps: coarse generation and 2D Grid deformed fine output. However, in the fine generation, the expansion range (2D Grid Scale) required by each point cloud sample may be vastly different. For example, if the expansion range for a vessel shape is applied to a table shape, the final output may be blurry or sparse. To this end, we propose the RLGrid, Reinforcement Learning Controlled Grid Deformation. In detail, we firstly obtain two point cloud skeletons by two branches. One is to use an autoencoder, and the other is to convert the randomly generated normal distribution to coarse point cloud by GAN. We choose the one with smaller chamfer distance between coarse output and incomplete input as the input of the second stage. Then, a Reinforcement Learning (RL) agent is designed to select the appropriate expansion range based on the feature of each point cloud, and generate a 2D Grid. Finally, all the features are concatenated and sent into a Multilayer Perceptron to obtain the detailed complete point cloud. Experimental results show that RLGrid achieves state-of-the-art performance on various datasets. To the best of our knowledge, RL is not widely used in point cloud completion task due to lack of custom environment, and the proposed RLGrid provides an insight on how to formulate 2D Grid deformation as a sequential decision making problem. Further, it can also be plug-and-play on any 2D Grid features. Pan Gao 0001, Xiaoyang Tan, Wei Xiang 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | An Implicit Trust Region Approach to Behavior Regularized Offline Reinforcement LearningabstractWe revisit behavior regularization, a popular approach to mitigate the extrapolation error in offline reinforcement learning (RL), showing that current behavior regularization may suffer from unstable learning and hinder policy improvement. Motivated by this, a novel reward shaping-based behavior regularization method is proposed, where the log-probability ratio between the learned policy and the behavior policy is monitored during learning. We show that this is equivalent to an implicit but computationally lightweight trust region mechanism, which is beneficial to mitigate the influence of estimation errors of the value function, leading to more stable performance improvement. Empirical results on the popular D4RL benchmark verify the effectiveness of the presented method with promising performance compared with some state-of-the-art offline RL algorithms. Xiaoyang Tan |
AAAI | 2 |
| 2024 | A Temporal Consistency Learning Framework for Face Forgery Detection
Xiaoyang Tan |
ISNN | 4 |
| 2024 | Multi-level Distributional Discrepancy Enhancement for Cross Domain Face Forgery Detection
Lingyu Qiu, Ke Jiang 0002, Sinan Liu, Xiaoyang Tan |
PRCV (15) | 4 |
| 2023 | ProxyFormer: Proxy Alignment Assisted Point Cloud Completion with Missing Part Sensitive TransformerabstractProblems such as equipment defects or limited view-points will lead the captured point clouds to be incomplete. Therefore, recovering the complete point clouds from the partial ones plays an vital role in many practical tasks, and one of the keys lies in the prediction of the missing part. In this paper, we propose a novel point cloud completion approach namely ProxyFormer that divides point clouds into existing (input) and missing (to be predicted) parts and each part communicates information through its proxies. Specifically, we fuse information into point proxy via feature and position extractor, and generate features for missing point proxies from the features of existing point proxies. Then, in order to better perceive the position of missing points, we design a missing part sensitive transformer, which converts random normal distribution into reasonable position information, and uses proxy alignment to refine the missing proxies. It makes the predicted point proxies more sensitive to the features and positions of the missing part, and thus makes these proxies more suitable for subsequent coarse-to-fine processes. Experimental results show that our method outperforms state-of-the-art completion networks on several benchmark datasets and has the fastest inference speed. Pan Gao 0001, Xiaoyang Tan, Mingqiang Wei |
CVPR | 3 |
| 2023 | Adaptive Reward Shifting Based on Behavior Proximity for Offline Reinforcement LearningabstractOne of the major challenges of the current offline reinforcement learning research is to deal with the distribution shift problem due to the change in state-action visitations for the new policy. To address this issue, we present a novel reward shifting-based method. Specifically, to regularize the behavior of the new policy at each state, we modify the reward to be received by the new policy by shifting it adaptively according to its proximity to the behavior policy, and apply the reward shifting along opposite directions for in-distribution actions and the ones not. In this way we are able to guide the learning procedure of the new policy itself by influencing the consequence of its actions explicitly, helping it to achieve a better balance between behavior constraints and policy improvement. Empirical results on the popular D4RL benchmarks show that the proposed method obtains competitive performance compared to the state-of-art baselines. Xiaoyang Tan |
IJCAI | 2 |
| 2023 | Recovering from Out-of-sample States via Inverse Dynamics in Offline Reinforcement LearningabstractIn this paper we deal with the state distributional shift problem commonly encountered in offline reinforcement learning during test, where the agent tends to take unreliable actions at out-of-sample (unseen) states. Our idea is to encourage the agent to follow the so called state recovery principle when taking actions, i.e., besides long-term return, the immediate consequences of the current action should also be taken into account and those capable of recovering the state distribution of the behavior policy are preferred. For this purpose, an inverse dynamics model is learned and employed to guide the state recovery behavior of the new policy. Theoretically, we show that the proposed method helps aligning the transited state distribution of the new policy with the offline dataset at out-of-sample states, without the need of explicitly predicting the transited state distribution, which is usually difficult in high-dimensional and complicated environments. The effectiveness and feasibility of the proposed method is demonstrated with the state-of-the-art performance on the general offline RL benchmarks. Ke Jiang 0002, Jia-Yu Yao, Xiaoyang Tan |
NeurIPS | 3 |
| 2023 | Self-imitation guided goal-conditioned reinforcement learning
Xiaoyang Tan |
Pattern Recognit. | 3 |
| 2023 | Deep continual hashing with gradient-aware memory for cross-modal retrieval
Xiaoyang Tan, Ming Yang 0014 |
Pattern Recognit. | 2 |
| 2023 | Robust RGB-T Tracking via Graph Attention-Based Bilinear PoolingabstractRGB-T tracker possesses strong capability of fusing two different yet complementary target observations, thus providing a promising solution to fulfill all-weather tracking in intelligent transportation systems. Existing convolutional neural network (CNN)-based RGB-T tracking methods often consider the multisource-oriented deep feature fusion from global viewpoint, but fail to yield satisfactory performance when the target pair only contains partially useful information. To solve this problem, we propose a four-stream oriented Siamese network (FS-Siamese) for RGB-T tracking. The key innovation of our network structure lies in that we formulate multidomain multilayer feature map fusion as a multiple graph learning problem, based on which we develop a graph attention-based bilinear pooling module to explore the partial feature interaction between the RGB and the thermal targets. This can effectively avoid uninformed image blocks disturbing feature embedding fusion. To enhance the efficiency of the proposed Siamese network structure, we propose to adopt meta-learning to incorporate category information in the updating of bilinear pooling results, which can online enforce the exemplar and current target appearance obtaining similar sematic representation. Extensive experiments on grayscale-thermal object tracking (GTOT) and RGBT234 datasets demonstrate that the proposed method outperforms the state-of-the-art methods for the task of RGB-T tracking. Bin Kang, Dong Liang 0008, Junxi Mei, Xiaoyang Tan, Dengyin Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | SMIX(λ): Enhancing Centralized Value Functions for Cooperative Multiagent Reinforcement LearningabstractLearning a stable and generalizable centralized value function (CVF) is a crucial but challenging task in multiagent reinforcement learning (MARL), as it has to deal with the issue that the joint action space increases exponentially with the number of agents in such scenarios. This article proposes an approach, named SMIX( λ ), that uses an OFF-policy training to achieve this by avoiding the greedy assumption commonly made in CVF learning. As importance sampling for such OFF-policy training is both computationally costly and numerically unstable, we proposed to use the λ -return as a proxy to compute the temporal difference (TD) error. With this new loss function objective, we adopt a modified QMIX network structure as the base to train our model. By further connecting it with the Q(λ) approach from a unified expectation correction viewpoint, we show that the proposed SMIX( λ ) is equivalent to Q(λ) and hence shares its convergence properties, while without being suffered from the aforementioned curse of dimensionality problem inherent in MARL. Experiments on the StarCraft Multiagent Challenge (SMAC) benchmark demonstrate that our approach not only outperforms several state-of-the-art MARL methods by a large margin but also can be used as a general tool to improve the overall performance of other centralized training with decentralized execution (CTDE)-type algorithms by enhancing their CVFs. Xinghu Yao, Yuhui Wang 0004, Xiaoyang Tan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Smoothing Advantage LearningabstractAdvantage learning (AL) aims to improve the robustness of value-based reinforcement learning against estimation errors with action-gap-based regularization. Unfortunately, the method tends to be unstable in the case of function approximation. In this paper, we propose a simple variant of AL, named smoothing advantage learning (SAL), to alleviate this problem. The key to our method is to replace the original Bellman Optimal operator in AL with a smooth one so as to obtain more reliable estimation of the temporal difference target. We give a detailed account of the resulting action gap and the performance bound for approximate SAL. Further theoretical analysis reveals that the proposed value smoothing technique not only helps to stabilize the training procedure of AL by controlling the trade-off between convergence rate and the upper bound of the approximation errors, but is beneficial to increase the action gap between the optimal and sub-optimal action value as well. Yaozhong Gan, Xiaoyang Tan |
AAAI | 3 |
| 2022 | Robust Action Gap Increasing with Clipped Advantage LearningabstractAdvantage Learning (AL) seeks to increase the action gap between the optimal action and its competitors, so as to improve the robustness to estimation errors. However, the method becomes problematic when the optimal action induced by the approximated value function does not agree with the true optimal action. In this paper, we present a novel method, named clipped Advantage Learning (clipped AL), to address this issue. The method is inspired by our observation that increasing the action gap blindly for all given samples while not taking their necessities into account could accumulate more errors in the performance loss bound, leading to a slow value convergence, and to avoid that, we should adjust the advantage value adaptively. We show that our simple clipped AL operator not only enjoys fast convergence guarantee but also retains proper action gaps, hence achieving a good balance between the large action gap and the fast convergence. The feasibility and effectiveness of the proposed method are verified empirically on several RL benchmarks with promising performance. Yaozhong Gan, Xiaoyang Tan |
AAAI | 3 |
| 2022 | Multi-scale Intermediate Flow Estimation for Video Frame InterpolationabstractVideo frame interpolation is one of the most chal-lenging tasks in video processing, which aims to synthesize intermediate frames between consecutive frames. In this work, we propose a flow-based approach called Multi-scale Intermediate Flow Estimation (MIFE) to balance the fineness and estimation range of the flows. MIFE consists of two main modules. Specifically, (1) Refined Flow Estimation uses a shifted window to estimate low-resolution intermediate flows at three levels. The refined full-resolution flow of each level is a weighted combination of nearby low-resolution flows, where the weights are determined by the similarity scores of the input frames and the reliability scores of the flows. (2) Multi-scale Flow Fusion generates fusion masks based on the estimable flow range and the estimated flow size. It fuses three levels of flows and refines the results. Experimental results show that the proposed method achieves good performance on various datasets. The source code is available at https://github.com/fzh169/MIFE. Zehua Fan, Xiaoyang Tan |
ICTAI | 4 |
| 2022 | A Cooperative-Competitive Multi-Agent Framework for Auto-bidding in Online AdvertisingabstractIn online advertising, auto-bidding has become an essential tool for advertisers to optimize their preferred ad performance metrics by simply expressing high-level campaign objectives and constraints. Previous works designed auto-bidding tools from the view of single-agent, without modeling the mutual influence between agents. In this paper, we instead consider this problem from a distributed multi-agent perspective, and propose a general \underlineM ulti-\underlineA gent reinforcement learning framework for \underlineA uto-\underlineB idding, namely MAAB, to learn the auto-bidding strategies. First, we investigate the competition and cooperation relation among auto-bidding agents, and propose a temperature-regularized credit assignment to establish a mixed cooperative-competitive paradigm. By carefully making a competition and cooperation trade-off among agents, we can reach an equilibrium state that guarantees not only individual advertiser's utility but also the system performance (i.e., social welfare). Second, to avoid the potential collusion behaviors of bidding low prices underlying the cooperation, we further propose bar agents to set a personalized bidding bar for each agent, and then alleviate the revenue degradation due to the cooperation. Third, to deploy MAAB in the large-scale advertising system with millions of advertisers, we propose a mean-field approach. By grouping advertisers with the same objective as a mean auto-bidding agent, the interactions among the large-scale advertisers are greatly simplified, making it practical to train MAAB efficiently. Extensive experiments on the offline industrial dataset and Alibaba advertising platform demonstrate that our approach outperforms several baseline methods in terms of social welfare and revenue. Zhilin Zhang 0003, Zhenzhe Zheng 0001, Yuhui Wang 0004, Xiaoyang Tan, Chuan Yu 0002, Jian Xu 0015, Fan Wu 0006, Guihai Chen, Xiaoqiang Zhu, Bo Zheng 0007 |
WSDM | 9 |
| 2022 | Delving deep into spatial pooling for squeeze-and-excitation networks
Xin Jin 0023, Yanping Xie, Xiu-Shen Wei, Borui Zhao, Xiaoyang Tan |
Pattern Recognit. | 6 |
| 2022 | Alleviating the estimation bias of deep deterministic policy gradient via co-regularization
Yuhui Wang 0004, Yaozhong Gan, Xiaoyang Tan |
Pattern Recognit. | 4 |
| 2022 | A Lightweight Encoder-Decoder Path for Deep Residual NetworksabstractIn this article, we present a novel lightweight path for deep residual neural networks. The proposed method integrates a simple plug-and-play module, i.e., a convolutional encoder-decoder (ED), as an augmented path to the original residual building block. Due to the abstract design and ability of the encoding stage, the decoder part tends to generate feature maps where highly semantically relevant responses are activated, while irrelevant responses are restrained. By a simple elementwise addition operation, the learned representations derived from the identity shortcut and original transformation branch are enhanced by our ED path. Furthermore, we exploit lightweight counterparts by removing a portion of channels in the original transformation branch. Fortunately, our lightweight processing does not cause an obvious performance drop but brings a computational economy. By conducting comprehensive experiments on ImageNet, MS-COCO, CUB200-2011, and CIFAR, we demonstrate the consistent accuracy gain obtained by our ED path for various residual architectures, with comparable or even lower model complexity. Concretely, it decreases the top-1 error of ResNet-50 and ResNet-101 by 1.22% and 0.91% on the task of ImageNet classification and increases the mmAP of Faster R-CNN with ResNet-101 by 2.5% on the MS-COCO object detection task. The code is available at https://github.com/Megvii-Nanjing/ED-Net. Xin Jin 0023, Yanping Xie, Xiu-Shen Wei, Borui Zhao, Xiaoyang Tan, Yang Yu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | Stabilizing Q Learning Via Soft Mellowmax OperatorabstractLearning complicated value functions in high dimensional state space by function approximation is a challenging task, partially due to that the max-operator used in temporal difference updates can theoretically cause instability for most linear or non-linear approximation schemes. Mellowmax is a recently proposed differentiable and non-expansion softmax operator that allows a convergent behavior in learning and planning. Unfortunately, the performance bound for the fixed point it converges to remains unclear, and in practice, its parameter is sensitive to various domains and has to be tuned case by case. Finally, the Mellowmax operator may suffer from oversmoothing as it ignores the probability being taken for each action when aggregating them. In this paper we address all the above issues with an enhanced Mellowmax operator, named SM2 (Soft Mellowmax). Particularly, the proposed operator is reliable, easy to implement, and has provable performance guarantee, while preserving all the advantages of Mellowmax. Furthermore, we show that our SM2 operator can be applied to the challenging multi-agent reinforcement learning scenarios, leading to stable value function approximation and state of the art performance. Yaozhong Gan, Xiaoyang Tan |
AAAI | 3 |
| 2021 | Deep Recurrent Belief Propagation Network for POMDPsabstractIn many real-world sequential decision-making tasks, especially in continuous control like robotic control, it is rare that the observations are perfect, that is, the sensory data could be incomplete, noisy or even dynamically polluted due to the unexpected malfunctions or intrinsic low quality of the sensors. Previous methods handle these issues in the framework of POMDPs and are either deterministic by feature memorization or stochastic by belief inference. In this paper, we present a new method that lies somewhere in the middle of the spectrum of research methodology identified above and combines the strength of both approaches. In particular, the proposed method, named Deep Recurrent Belief Propagation Network (DRBPN), takes a hybrid style belief updating procedure − an RNN-type feature extraction step followed by an analytical belief inference, significantly reducing the computational cost while faithfully capturing the complex dynamics and maintaining the necessary uncertainty for generalization. The effectiveness of the proposed method is verified on a collection of benchmark tasks, showing that our approach outperforms several state-of-the-art methods under various challenging scenarios. Yuhui Wang 0004, Xiaoyang Tan |
AAAI | 2 |
| 2021 | Cross-scene foreground segmentation with supervised and unsupervised model communication
Dong Liang 0008, Bin Kang, Pan Gao 0001, Xiaoyang Tan, Shun'ichi Kaneko |
Pattern Recognit. | 5 |
| 2021 | Deep robust multilevel semantic hashing for multi-label cross-modal retrieval
Xiaoyang Tan, Jun Zhao 0007, Ming Yang 0014 |
Pattern Recognit. | 2 |
| 2021 | Real-world Cross-modal Retrieval via Sequential LearningabstractCross-modal retrieval is playing an increasingly important role in our daily life with the explosive growth of multimedia data. However, its learning paradigm under real-life environments is less studied, and most existing approaches are developed in the pre-desired settings (e.g., unchanging modalities and explicitly modal-aligned samples). Inspired by the recent achievement in the field of cognition mechanism on how the human brain acquires knowledge, we present a new sequential learning method for real-world cross-modal retrieval. In this method, a unified model is maintained to capture the common knowledge of various modalities but are learned in a sequential manner such that it behaves adaptively according to the evolving distribution of different modalities, and needs no laborious alignment operations among multimodal data before learning. Furthermore, we reformulate the objective of optimization-based meta-learning and propose a novel meta-learning method to overcome the catastrophic forgetting encountered in sequential learning. Extensive experiments are conducted on four popular image-text multimodal datasets and a five-modal dataset, showing that our method achieves state-of-the-art cross-modal retrieval performance without explicit modal-alignment. Xiaoyang Tan |
IEEE Trans. Multim. | 2 |
| 2020 | SMIX(λ): Enhancing Centralized Value Functions for Cooperative Multi-Agent Reinforcement LearningabstractThis work presents a sample efficient and effective value-based method, named SMIX(λ), for reinforcement learning in multi-agent environments (MARL) within the paradigm of centralized training with decentralized execution (CTDE), in which learning a stable and generalizable centralized value function (CVF) is crucial. To achieve this, our method carefully combines different elements, including 1) removing the unrealistic centralized greedy assumption during the learning phase, 2) using the λ-return to balance the trade-off between bias and variance and to deal with the environment's non-Markovian property, and 3) adopting an experience-replay style off-policy training. Interestingly, it is revealed that there exists inherent connection between SMIX(λ) and previous off-policy Q(λ) approach for single-agent learning. Experiments on the StarCraft Multi-Agent Challenge (SMAC) benchmark show that the proposed SMIX(λ) algorithm outperforms several state-of-the-art MARL methods by a large margin, and that it can be used as a general tool to improve the overall performance of a CTDE-type method by enhancing the evaluation quality of its CVF. We open-source our code at: https://github.com/chaovven/SMIX. Xinghu Yao, Yuhui Wang 0004, Xiaoyang Tan |
AAAI | 4 |
| 2020 | ACRM: Attention Cascade R-CNN with Mix-NMS for Metallic Surface Defect DetectionabstractMetallic surface defect detection is of great significance in quality control for production. However, this task is very challenging due to the noise disturbance, large appearance variation, and the ambiguous definition of the defect individual. Traditional image processing methods are unable to detect the damaged region effectively and efficiently. In this paper, we propose a new defect detection method, Attention Cascade R-CNN with Mix-NMS (ACRM), to classify and locate defects robustly. Three submodules are developed to achieve this goal: 1) a lightweight attention block is introduced, which can improve the ability in capture global and local feature both in the spatial and channel dimension; 2) we firstly apply the cascade R-CNN to our task, which exploits multiple detectors to sequentially refine the detection result robustly; 3) we introduce a new method named Mix Non-Maximum Suppression (Mix-NMS), which can significantly improve its ability in filtering the redundant detection result in our task. Extensive experiments on a real industrial dataset show that ACRM achieves state-of-the-art results compared to the existing methods, demonstrating the effectiveness and robustness of our detection method. Junting Fang, Xiaoyang Tan, Yuhui Wang 0004 |
ICPR | 2 |
| 2020 | Deep code operation network for multi-label image retrieval
Xiaoyang Tan |
Comput. Vis. Image Underst. | 2 |
| 2019 | Trust Region-Guided Proximal Policy OptimizationabstractProximal policy optimization (PPO) is one of the most popular deep reinforcement learning (RL) methods, achieving state-of-the-art performance across a wide range of challenging tasks. However, as a model-free RL method, the success of PPO relies heavily on the effectiveness of its exploratory policy search. In this paper, we give an in-depth analysis on the exploration behavior of PPO, and show that PPO is prone to suffer from the risk of lack of exploration especially under the case of bad initialization, which may lead to the failure of training or being trapped in bad local optima. To address these issues, we proposed a novel policy optimization method, named Trust Region-Guided PPO (TRGPPO), which adaptively adjusts the clipping range within the trust region. We formally show that this method not only improves the exploration ability within the trust region but enjoys a better performance bound compared to the original PPO as well. Extensive experiments verify the advantage of the proposed method. Yuhui Wang 0004, Xiaoyang Tan, Yaozhong Gan |
NeurIPS | 3 |
| 2019 | Truly Proximal Policy Optimization
Yuhui Wang 0004, Xiaoyang Tan |
UAI | 3 |
| 2019 | Robust Class-Specific Autoencoder for Data Cleaning and Classification in the Presence of Label Noise
Weining Zhang, Dong Wang 0015, Xiaoyang Tan |
Neural Process. Lett. | 3 |
| 2019 | Bayesian denoising hashing for robust image retrieval
Dong Wang 0015, Xiaoyang Tan |
Pattern Recognit. | 3 |
| 2019 | Pornographic Image Recognition via Weighted Multiple Instance LearningabstractIn the era of Internet, recognizing pornographic images is of great significance for protecting children's physical and mental health. However, this task is very challenging as the key pornographic contents (e.g., breast and private part) in an image often lie in local regions of small size. In this paper, we model each image as a bag of regions, and follow a multiple instance learning (MIL) approach to train a generic region-based recognition model. Specifically, we take into account the regions' degree of pornography, and make three main contributions. First, we show that based on very few annotations of the key pornographic contents in a training image, we can generate a bag of properly sized regions, among which the potential positive regions usually contain useful contexts that can aid recognition. Second, we present a simple quantitative measure of a region's degree of pornography, which can be used to weigh the importance of different regions in a positive image. Third, we formulate the recognition task as a weighted MIL problem under the convolutional neural network framework, with a bag probability function introduced to combine the importance of different regions. Experiments on our newly collected large scale dataset demonstrate the effectiveness of the proposed method, achieving an accuracy with 97.52% true positive rate at 1% false positive rate, tested on 100K pornographic images and 100K normal images. Yuhui Wang 0004, Xiaoyang Tan |
IEEE Trans. Cybern. | 3 |
| 2019 | Deep Memory Network for Cross-Modal RetrievalabstractWith the explosive growth of multimedia data on the Internet, cross-modal retrieval has attracted a great deal of attention in both computer vision and multimedia communities. However, this task is challenging due to the heterogeneity gap between different modalities. Current approaches typically involve a common representation learning process that maps data from different modalities into a common space by linear or nonlinear embedding. Yet, most of them only handle the dual-modal situation and generalize poorly to complex cases that involve multiple modalities. In addition, they often require expensive fine-grained alignment of training data among diverse modalities. In this paper, we address these with a novel cross-modal memory network (CMMN), in which memory contents across modalities are simultaneously learned from end to end without the need of exact alignment. We further account for the diversity across multiple modalities using the strategy of adversarial learning. Extensive experimental results on several large-scale datasets demonstrate that the proposed CMMN approach achieves state-of-the-art performance in the task of cross-modal retrieval. Dong Wang 0015, Xiaoyang Tan |
IEEE Trans. Multim. | 3 |
| 2018 | Data Cleaning and Classification in the Presence of Label Noise with Class-Specific Autoencoder
Weining Zhang, Dong Wang 0015, Xiaoyang Tan |
ISNN | 3 |
| 2018 | Learning Multilevel Semantic Similarity for Large-Scale Multi-Label Image RetrievalabstractWe present a novel Deep Supervised Hashing with code operation (DSOH) method for large-scale multi-label image retrieval. This approach is in contrast with existing methods in that we respect both the intention gap and the intrinsic multilevel similarity of multi-labels. Particularly, our method allows a user to simultaneously present multiple query images rather than a single one to better express her intention, and correspondingly a separate sub-network in our architecture is specifically designed to fuse the query intention represented by each single query. Furthermore, as in the training stage, each image is annotated with multiple labels to enrich its semantic representation, we propose a new margin-adaptive triplet loss to learn the fine-grained similarity structure of multi-labels, which is known to be hard to capture. The whole system is trained in an end-to-end manner, and our experimental results demonstrate that the proposed method is not only able to learn useful multilevel semantic similarity-preserving binary codes but also achieves state-of-the-art retrieval performance on three popular datasets. Xiaoyang Tan |
ICMR | 2 |
| 2018 | Robust Distance Metric Learning via Bayesian InferenceabstractDistance metric learning (DML) has achieved great success in many computer vision tasks. However, most existing DML algorithms are based on point estimation, and thus are sensitive to the choice of training examples and tend to be over-fitting in the presence of label noise. In this paper, we present a robust DML algorithm based on Bayesian inference. In particular, our method is essentially a Bayesian extension to a previous classic DML method-large margin nearest neighbor classification and we use stochastic variational inference to estimate the posterior distribution of the transformation matrix. Furthermore, we theoretically show that the proposed algorithm is robust against label noise in the sense that an arbitrary point with label noise has bounded influence on the learnt model. With some reasonable assumptions, we derive a generalization error bound of this method in the presence of label noise. We also show that the DML hypothesis class in which our model lies is probably approximately correct-learnable and give the sample complexity. The effectiveness of the proposed method1is demonstrated with state of the art performance on three popular data sets with different types of label noise.1A MATLAB implementation of this method is made available at http://parnec.nuaa.edu.cn/xtan/Publication.htm. Dong Wang 0015, Xiaoyang Tan |
IEEE Trans. Image Process. | 2 |
| 2018 | Bayesian Neighborhood Component AnalysisabstractLearning a distance metric in feature space potentially improves the performance of the nearest neighbor classifier and is useful in many real-world applications. Many metric learning (ML) algorithms are, however, based on the point estimation of a quadratic optimization problem, which is time-consuming, susceptible to overfitting, and lacks a natural mechanism to reason with parameter uncertainty-a property useful especially when the training set is small and/or noisy. To deal with these issues, we present a novel Bayesian ML (BML) method, called Bayesian neighborhood component analysis (NCA), based on the well-known NCA method, in which the metric posterior is characterized by the local label consistency constraints of observations, encoded with a similarity graph instead of independent pairwise constraints. For efficient Bayesian inference, we explore the variational lower bound over the log-likelihood of the original NCA objective. Experiments on several publicly available data sets demonstrate that the proposed method is able to learn robust metric measures from small size data set and/or from challenging training set with labels contaminated by errors. The proposed method is also shown to outperform a previous pairwise constrained BML method. Dong Wang 0015, Xiaoyang Tan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Cross-modal Retrieval via Memory Network
Xiaoyang Tan |
BMVC | 2 |
| 2017 | Face alignment in-the-wild: A Survey
Xiaoyang Tan |
Comput. Vis. Image Underst. | 2 |
| 2017 | Hierarchical deep hashing for image retrieval
Xiaoyang Tan |
Frontiers Comput. Sci. | 2 |
| 2017 | Selective Weakly Supervised Human Detection under Arbitrary Poses
Yawei Cai, Xiaosong Tan, Xiaoyang Tan |
Pattern Recognit. | 3 |
| 2016 | Weakly supervised human body detection under arbitrary posesabstractIn this work we study the problem of weakly supervised human body detection under difficult poses (e.g., multiview and/or arbitrary poses) within the framework of multi-instance learning (MIL). We first point out the existence of the so-called “vanishing gradient” problem in MIL with a noisy-or rule as its bagging model. This is mainly due to the independence assumption of the noisy-or rule, which significantly reducing the magnitude of gradient under a weak initial instance-level model. To address this issue, we propose an iterative selective MIL method in which 1) the noisy-or rule is replaced with the max rule and only a few instances are included for MIL learning for each bag and for each time, and 2) prior knowledge about the positive instances in terms of few fully supervised samples are employed to improve the robustness. The method is shown to outperform the previous state-of-the-art methods by over 20.0% in accuracy. Finally, we present a new large-scale data set called MPHB (Multiple Poses Human Body) for human body detection under arbitrary poses. Yawei Cai, Xiaoyang Tan |
ICIP | 2 |
| 2016 | Pornographic image recognition by strongly-supervised deep multiple instance learningabstractIn this paper, we propose a principled framework for pornographic image recognition. Specifically, we present our definition of pornographic images, which characterizes the pornographic contents in images as the exposure of private body parts. As the private body parts often lie in local image regions, we model each image as a bag of local image patches (instances), and assume that for each pornographic image at least one instance accounts for the pornographic content within it. This treatment allows us to cast the model training as a Multiple Instance Learning (MIL) problem. Furthermore, we propose a strongly-supervised setting for MIL by identifying the most likely pornographic instances in positive bags, which effectively prevents the algorithm from getting trapped in a bad local optima. Last but not least, we formulate our strongly-supervised MIL under the deep CNN framework to learn deep representations; hence we call it Strongly-supervised Deep MIL (SD-MIL). We demonstrate that our SD-MIL based system produces remarkable accuracy with 97.01% TPR at 1% FPR, testing on 117K pornographic images and 117K normal images from our newly-collected large scale dataset. Yuhui Wang 0004, Xiaoyang Tan |
ICIP | 3 |
| 2016 | Max-margin non-negative matrix factorization with flexible spatial constraints based on factor analysis
Dakun Liu, Xiaoyang Tan |
Frontiers Comput. Sci. | 2 |
| 2016 | Local subspace smoothness alignment for constrained local model fitting
Dakun Liu, Xiaoyang Tan |
Neurocomputing | 2 |
| 2016 | Mixed bi-subject kinship verification via multi-view multi-task learning
Xiaoqian Qin, Xiaoyang Tan, Songcan Chen |
Neurocomputing | 2 |
| 2016 | Face alignment by robust discriminative Hough voting
Xiaoyang Tan |
Pattern Recognit. | 2 |
| 2016 | Unsupervised feature learning with C-SVDDNet
Dong Wang 0015, Xiaoyang Tan |
Pattern Recognit. | 2 |
| 2015 | Tri-Subject Kinship Verification: Understanding the Core of A FamilyabstractOne major challenge in computer vision is to go beyond the modeling of individual objects and to investigate the bi- (one-versus-one) or tri- (one-versus-two) relationship among multiple visual entities, answering such questions as whether a child in a photo belongs to the given parents. The child-parents relationship plays a core role in a family, and understanding such kin relationship would have a fundamental impact on the behavior of an artificial intelligent agent working in the human world. In this work, we tackle the problem of one-versus-two (tri-subject) kinship verification and our contributions are threefold: 1) a novel relative symmetric bilinear model (RSBM) is introduced to model the similarity between the child and the parents, by incorporating the prior knowledge that a child may resemble one particular parent more than the other; 2) a spatially voted method for feature selection, which jointly selects the most discriminative features for the child-parents pair, while taking local spatial information into account; and 3) a large-scale tri-subject kinship database characterized by over 1,000 child-parents families. Extensive experiments on KinFaceW, Family101, and our newly released kinship database show that the proposed method outperforms several previous state of the art methods, while could also be used to significantly boost the performance of one-versus-one kinship verification when the information about both parents are available. Xiaoqian Qin, Xiaoyang Tan, Songcan Chen |
IEEE Trans. Multim. | 2 |
| 2014 | Robust Distance Metric Learning in the Presence of Label NoiseabstractMany distance learning algorithms have been developed in recent years. However, few of them consider the problem when the class labels of training data are noisy, and this may lead to serious performance deterioration. In this paper, we present a robust distance learning method in the presence of label noise, by extending a previous non-parametric discriminative distance learning algorithm, i.e., Neighbourhood Components Analysis (NCA). Particularly, we analyze the effect of label noise on the derivative of likelihood with respect to the transformation matrix, and propose to model the conditional probability of the true label of each point so as to reduce that effect. The model is then optimized within the EM framework, with additional regularization used to avoid overfitting. Our experiments on several UCI datasets and a real dataset with unknown noise patterns show that the proposed RNCA is more tolerant to class label noise compared to the original NCA method. Dong Wang 0015, Xiaoyang Tan |
AAAI | 2 |
| 2014 | Learning One-Shot Exemplar SVM from the Web for Face Verification
Fengyi Song, Xiaoyang Tan |
ACCV (3) | 2 |
| 2014 | Action Recognition from a Single Web Image Based on an Ensemble of Pose Experts
Peihao Zhang, Xiaoyang Tan |
ACCV (1) | 2 |
| 2014 | Label-Denoising Auto-encoder for Classification with Inaccurate Supervision InformationabstractLabel noise is not uncommon in machine learning applications nowadays and imposes great challenges for many existing classifiers. In this paper we propose a new type of auto-encoder coined label-denoising auto-encoder to learn a representation for robust classification under this situation. For this purpose, we include both the feature and the (noisy) label of a data point in the input layer of the auto-encoder network, and during each learning iteration, we disturb the label according to the posterior probability of the data estimated by a soft max regression classifier. The learnt representation is shown to be robust against label noise on three real-world data-sets. Dong Wang 0015, Xiaoyang Tan |
ICPR | 2 |
| 2014 | Exploiting relationship between attributes for improved face verification
Fengyi Song, Xiaoyang Tan, Songcan Chen |
Comput. Vis. Image Underst. | 2 |
| 2014 | Part-based pose estimation with local and non-local contextual informationabstractIn this study, the authors propose a new method for part‐based human pose estimation. The key idea of the authors method is to improve the accuracies for leaf parts localisations – an issue that was largely ignored by the previous study – by incorporating both local and non‐local contextual information into the model. In particular, they use the local contextual information to reduce or eliminate the influences of the noises, while the non‐local contextual information helps to improve the detection accuracies of the leaf parts. Since more accurate parts localisations usually mean a more reasonable active set of spatial constraints, this potentially enhances the effectiveness of the subsequent optimisation procedure. Furthermore, they keep the basic structure of the tree‐based model, hence taking advantage of its conceptual simplicity and computationally efficient inference. Their experiments on two challenging real‐world datasets demonstrate the feasibility and the effectiveness of the proposed method. Xiaoyang Tan |
IET Comput. Vis. | 2 |
| 2014 | Sparse representations based attribute learning for flower classification
Keyang Cheng, Xiaoyang Tan |
Neurocomputing | 2 |
| 2014 | Comparative study among three strategies of incorporating spatial structures to ordinal image regression
Qing Tian 0002, Songcan Chen, Xiaoyang Tan |
Neurocomputing | 3 |
| 2014 | Eyes closeness detection from still images with multi-scale histograms of principal oriented gradients
Fengyi Song, Xiaoyang Tan, Songcan Chen |
Pattern Recognit. | 2 |
| 2013 | Centering SVDD for Unsupervised Feature Representation in Object Classification
Dong Wang 0015, Xiaoyang Tan |
ICONIP (3) | 2 |
| 2013 | A literature survey on robust and efficient eye localization in real-life scenarios
Fengyi Song, Xiaoyang Tan, Songcan Chen, Zhi-Hua Zhou |
Pattern Recognit. | 2 |
| 2013 | Two-dimensional bar code out-of-focus deblurring via the Increment Constrained Least Squares filter
Ningzhong Liu, Xingming Zheng, Xiaoyang Tan |
Pattern Recognit. Lett. | 4 |
| 2012 | Exploiting relationship between attributes for improved face verificationabstractAbstract Recent work has shown the advantages of using high level representation such as attribute-based descriptors over low-level feature sets in face verification. However, in most work each attribute is coded with extremely short information length (e.g., “is Male”, “has Beard”) and all the attributes belonging to the same object are assumed to be independent of each other when using them for prediction. To address the above two problems, we propose a discriminative distributed-representation for attribute description; on the basis of this description, we present a novel method to model the relationship between attributes and exploit such relationship to improve the performance of face verification, in the meantime taking uncertainty in attribute responses into account. Specifically, inspired by the vector representation of words in the literature of text categorization, we first represent the meaning of each attribute as a high-dimensional vector in the subject space, then construct an attribute-relationship graph based on the distribution of attributes in that space. With this graph, we are able to explicitly constrain the searching space of parameter values of a discriminative classifier to avoid over-fitting. The effectiveness of the proposed method is verified on two challenging face databases (i.e., LFW and PubFig) and the a-Pascal object dataset. Furthermore, we extend the proposed method to the case with continuous attributes with promising results. Fengyi Song, Xiaoyang Tan, Songcan Chen |
BMVC | 2 |
| 2010 | Face Liveness Detection from a Single Image with Sparse Low Rank Bilinear Discriminative Model
Xiaoyang Tan, Jun Liu 0003 |
ECCV (6) | 1 |
| 2010 | Sparsity preserving projections with applications to face recognition
Lishan Qiao, Songcan Chen, Xiaoyang Tan |
Pattern Recognit. | 3 |
| 2010 | Sparsity preserving discriminant analysis for single training image face recognition
Lishan Qiao, Songcan Chen, Xiaoyang Tan |
Pattern Recognit. Lett. | 3 |
| 2010 | Enhanced Local Texture Feature Sets for Face Recognition Under Difficult Lighting ConditionsabstractMaking recognition more reliable under uncontrolled lighting conditions is one of the most important challenges for practical face recognition systems. We tackle this by combining the strengths of robust illumination normalization, local texture-based face representations, distance transform based matching, kernel-based feature extraction and multiple feature fusion. Specifically, we make three main contributions: 1) we present a simple and efficient preprocessing chain that eliminates most of the effects of changing illumination while still preserving the essential appearance details that are needed for recognition; 2) we introduce local ternary patterns (LTP), a generalization of the local binary pattern (LBP) local texture descriptor that is more discriminant and less sensitive to noise in uniform regions, and we show that replacing comparisons based on local spatial histograms with a distance transform based similarity metric further improves the performance of LBP/LTP based face recognition; and 3) we further improve robustness by adding Kernel principal component analysis (PCA) feature extraction and incorporating rich local appearance cues from two complementary sources--Gabor wavelets and LBP--showing that the combination is considerably more accurate than either feature set alone. The resulting method provides state-of-the-art performance on three data sets that are widely used for testing recognition under difficult illumination conditions: Extended Yale-B, CAS-PEAL-R1, and Face Recognition Grand Challenge version 2 experiment 4 (FRGC-204). For example, on the challenging FRGC-204 data set it halves the error rate relative to previously published methods, achieving a face verification rate of 88.1% at 0.1% false accept rate. Further experiments show that our preprocessing method outperforms several existing preprocessors for a range of feature sets, data sets and lighting conditions. Xiaoyang Tan, Bill Triggs |
IEEE Trans. Image Process. | 1 |
| 2010 | Generalized low-rank approximations of matrices revisitedabstractCompared to singular value decomposition (SVD), generalized low-rank approximations of matrices (GLRAM) can consume less computation time, obtain higher compression ratio, and yield competitive classification performance. GLRAM has been successfully applied to applications such as image compression and retrieval, and quite a few extensions have been successively proposed. However, in literature, some basic properties and crucial problems with regard to GLRAM have not been explored or solved yet. For this sake, we revisit GLRAM in this paper. First, we reveal such a close relationship between GLRAM and SVD that GLRAM's objective function is identical to SVD's objective function except the imposed constraints. Second, we derive a lower bound of GLRAM's objective function, and discuss when the lower bound can be touched. Moreover, from the viewpoint of minimizing the lower bound, we answer one open problem raised by Ye (Machine Learning, 2005), i.e., a theoretical justification of the experimental phenomenon that, under given number of reduced dimension, the lowest reconstruction error is obtained when the left and right transformations have equal number of columns. Third, we explore when and why GLRAM can perform well in terms of compression, which is a fundamental problem concerning the usability of GLRAM. Jun Liu 0003, Songcan Chen, Zhi-Hua Zhou, Xiaoyang Tan |
IEEE Trans. Neural Networks | 4 |
| 2009 | Enhanced Pictorial Structures for precise eye localization under incontrolled conditionsabstractIn this paper, we present an enhanced pictorial structure (PS) model for precise eye localization, a fundamental problem involved in many face processing tasks. PS is a computationally efficient framework for part-based object modelling. For face images taken under uncontrolled conditions, however, the traditional PS model is not flexible enough for handling the complicated appearance and structural variations. To extend PS, we 1) propose a discriminative PS model for a more accurate part localization when appearance changes seriously, 2) introduce a series of global constraints to improve the robustness against scale, rotation and translation, and 3) adopt a heuristic prediction method to address the difficulty of eye localization with partial occlusion. Experimental results on the challenging LFW (Labeled Face in the Wild) database show that our model can locate eyes accurately and efficiently under a broad range of uncontrolled variations involving poses, expressions, lightings, camera qualities, occlusions, etc. Xiaoyang Tan, Fengyi Song, Zhi-Hua Zhou, Songcan Chen |
CVPR | 1 |
| 2009 | Face recognition under occlusions and variant expressions with partial similarityabstractRecognition in uncontrolled situations is one of the most important bottlenecks for practical face recognition systems. In particular, few researchers have addressed the challenge to recognize noncooperative or even uncooperative subjects who try to cheat the recognition system by deliberately changing their facial appearance through such tricks as variant expressions or disguise (e.g., by partial occlusions). This paper addresses these problems within the framework of similarity matching. A novel perception-inspired nonmetric partial similarity measure is introduced, which is potentially useful in dealing with the concerned problems because it can help capture the prominent partial similarities that are dominant in human perception. Two methods, based on the general golden section rule and the maximum margin criterion, respectively, are proposed to automatically set the similarity threshold. The effectiveness of the proposed method in handling large expressions, partial occlusions, and other distortions is demonstrated on several well-known face databases. Xiaoyang Tan, Songcan Chen, Zhi-Hua Zhou, Jun Liu 0003 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2008 | A study on three linear discriminant analysis based methods in small sample size problem
Jun Liu 0003, Songcan Chen, Xiaoyang Tan |
Pattern Recognit. | 3 |
| 2008 | Fractional order singular value decomposition representation for face recognition
Jun Liu 0003, Songcan Chen, Xiaoyang Tan |
Pattern Recognit. | 3 |
| 2007 | Efficient Pseudoinverse Linear Discriminant Analysis and its Nonlinear Form for Face RecognitionabstractPseudoinverse Linear Discriminant Analysis (PLDA) is a classical and pioneer method that deals with the Small Sample Size (SSS) problem in LDA when applied to such applications as face recognition. However, it is expensive in computation and storage due to direct manipulation on extremely large d × d matrices, where d is the dimension of the sample image. As a result, although frequently cited in literature, PLDA is hardly compared in terms of classification performance with the newly proposed methods. In this paper, we propose a new feature extraction method named RSw + LDA, which is (1) much more efficient than PLDA in both computation and storage; and (2) theoretically equivalent to PLDA, meaning that it produces the same projection matrix as PLDA. Further, to make PLDA deal better with data of nonlinear distribution, we propose a Kernel PLDA (KPLDA) method with the well-known kernel trick. Finally, our experimental results on AR face dataset, a challenging dataset with variations in expression, lighting and occlusion, show that PLDA (or RSw + LDA) can achieve significantly higher classification accuracy than the recently proposed Linear Discriminant Analysis via QR decomposition and Discriminant Common Vectors, and KPLDA can yield better classification performance compared to PLDA and Kernel PCA. Jun Liu 0003, Songcan Chen, Xiaoyang Tan, Daoqiang Zhang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2007 | Comments on "Efficient and Robust Feature Extraction by Maximum Margin Criterion"abstractThe goal of this comment is to first point out two loopholes in the paper by Li (2006): (1) so-designed efficient maximal margin criterion (MMC) algorithm for small sample size (SSS) problem is problematic and (2) the discussion on the equivalence with the null-space-based methods in SSS problem does not hold. Then, we will present a really efficient MMC algorithm for SSS problem. Jun Liu 0003, Songcan Chen, Xiaoyang Tan, Daoqiang Zhang |
IEEE Trans. Neural Networks | 3 |
| 2006 | Learning Non-Metric Partial Similarity Based on Maximal Margin CriterionabstractThe performance of many computer vision and machine learning algorithms critically depends on the quality of the similarity measure defined over the feature space. Previous works usually utilize metric distances which are ofen epistemologically different from the perceptual distance of human beings. In this paper a novel non-metric partial similarity measure is introduced, which is born to automatically capture the prominent partial similarity between two images while ignoring the confusing unimportant dissimilarity. This measure is potentially useful in face recognition since it can help identify the inherent intra-personal similarity and thus reducing the influence caused by large variations such as expression and occlusions. Moreover; to make this method practical, this paper proposes an automatic and class-dependent similarity threshold setting mechanism based on the maximal margin criterion, and uses a Self- Organization Map-based embedding technique to alleviate the computational problem. Experimental results show the feasibility and effectiveness of the proposed method. Xiaoyang Tan, Songcan Chen, Zhi-Hua Zhou |
CVPR (1) | 1 |
| 2006 | Recognition from a Single Sample per Person with Multiple SOM Fusion
Xiaoyang Tan, Jun Liu 0003, Songcan Chen |
ISNN (2) | 1 |
| 2006 | Sub-intrapersonal space analysis for face recognition
Xiaoyang Tan, Jun Liu 0003, Songcan Chen |
Neurocomputing | 1 |
| 2006 | Face recognition from a single image per person: A survey
Xiaoyang Tan, Songcan Chen, Zhi-Hua Zhou, Fuyan Zhang |
Pattern Recognit. | 1 |
| 2005 | Weighted SOM-Face: Selecting Local Features for Recognition from Individual Face Image
Xiaoyang Tan, Jun Liu 0003, Songcan Chen, Fuyan Zhang |
IDEAL | 1 |
| 2005 | Feature Selection for High Dimensional Face Image Using Self-organizing Maps
Xiaoyang Tan, Songcan Chen, Zhi-Hua Zhou, Fuyan Zhang |
PAKDD | 1 |
| 2005 | Recognizing partially occluded, expression variant faces from single training image per person with SOM and soft k-NN ensembleabstractMost classical template-based frontal face recognition techniques assume that multiple images per person are available for training, while in many real-world applications only one training image per person is available and the test images may be partially occluded or may vary in expressions. This paper addresses those problems by extending a previous local probabilistic approach presented by Martinez, using the self-organizing map (SOM) instead of a mixture of Gaussians to learn the subspace that represented each individual. Based on the localization of the training images, two strategies of learning the SOM topological space are proposed, namely to train a single SOM map for all the samples and to train a separate SOM map for each class, respectively. A soft kappa nearest neighbor (soft kappa-NN) ensemble method, which can effectively exploit the outputs of the SOM topological space, is also proposed to identify the unlabeled subjects. Experiments show that the proposed method exhibits high robust performance against the partial occlusions and variant expressions. Xiaoyang Tan, Songcan Chen, Zhi-Hua Zhou, Fuyan Zhang |
IEEE Trans. Neural Networks | 1 |
| 2004 | Robust Face Recognition from a Single Training Image per Person with Kernel-Based SOM-Face
Xiaoyang Tan, Songcan Chen, Zhi-Hua Zhou, Fuyan Zhang |
ISNN (1) | 1 |