VLDB 2026 Research / reviewers in the wild / expert
Xinwen Hou
dblp:76/5119
· DBLP profile ↗
46ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-8468-001XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 7 since 2021Security and privacy · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ICaRe: hierarchical causality-inspired representation learning for task-specific disentanglement in multi-task scenarios
Wenhao Jing, Bo Fan 0008, Jilong Zhong, Lixia Xu, Xiaoyu Zhai, Xinwen Hou |
Mach. Vis. Appl. | 8 |
| 2025 | GaussMarker: Robust Dual-Domain Watermark for Diffusion ModelsabstractAs Diffusion Models (DM) generate increasingly realistic images, related issues such as copyright and misuse have become a growing concern. Watermarking is one of the promising solutions. Existing methods inject the watermark into the *single-domain* of initial Gaussian noise for generation, which suffers from unsatisfactory robustness.
This paper presents the first *dual-domain* DM watermarking approach using a pipelined injector to consistently embed watermarks in both the spatial and frequency domains. To further boost robustness against certain image manipulations and advanced attacks, we introduce a model-independent learnable Gaussian Noise Restorer (GNR) to refine Gaussian noise extracted from manipulated images and enhance detection robustness by integrating the detection scores of both watermarks.
GaussMarker efficiently achieves state-of-the-art performance under eight image distortions and four advanced attacks across three versions of Stable Diffusion with better recall and lower false positive rates, as preferred in real applications. Kecen Li, Xinwen Hou |
ICML | 3 |
| 2025 | From Easy to Hard: Building a Shortcut for Differentially Private Image SynthesisabstractDifferentially private (DP) image synthesis aims to generate synthetic images from a sensitive dataset, alleviating the privacy leakage concerns of organizations sharing and utilizing synthetic images. Although previous methods have significantly progressed, especially in training diffusion models on sensitive images with DP Stochastic Gradient Descent (DP-SGD), they still suffer from unsatisfactory performance. In this work, inspired by curriculum learning, we propose a two-stage DP image synthesis framework, where diffusion models learn to generate DP synthetic images from easy to hard. Unlike existing methods that directly use DP-SGD to train diffusion models, we propose an easy stage in the beginning, where diffusion models learn simple features of the sensitive images. To facilitate this easy stage, we propose to use ‘central images’, simply aggregations of random samples of the sensitive dataset. Intuitively, although those central images do not show details, they demonstrate useful characteristics of all images and only incur minimal privacy costs, thus helping early-phase model training. We conduct experiments to present that on the average of four investigated image datasets, the fidelity and utility metrics of our synthetic images are 33.1% and 2.1% better than the state-of-the-art method. The replication package and datasets can be accessed online11.https://github.comJSunnierLee/DP-FETA. Kecen Li, Chen Gong 0005, Yuzhong Zhao, Xinwen Hou, Tianhao Wang 0001 |
SP | 5 |
| 2024 | MEPE: A Minimalist Ensemble Policy Evaluation Operator for Deep Reinforcement LearningabstractEnsemble deep reinforcement learning (DRL) is a popular approach to mitigate the risk of overspecialization to a particular distribution and improve performance. However, it is commonly plagued by high computational resource requirements that arise from the introduction of multiple value and policy functions. To avoid the notorious resource consumption issue, we design a simple yet effective ensemble policy evaluation operator, termed Minimalist Ensemble Policy Evaluation (MEPE), which leverages a modified dropout operator combined with the Bellman equation. MEPE is equivalent to integrating multiple models into a single model. The MEPE operator holds ensemble property by keeping the dropout consistency of both sides of the Bellman equation and can be combined with any DRL algorithms if they have a policy evaluation phase. To verify the MEPE’s ability, we perform experiments on both low-dimensional and high-dimensional environments, which presents that the algorithms combined with the MEPE operator outperform or achieve a similar level of performance as the current state-of-the-art ensemble methods and model-free methods without increasing additional computational resource costs. To the best of our knowledge, our work is the first that applies the dropout operator both in low-dimensional control tasks and high-dimensional games. Our code is available at https://github.com/sweetice/MEPE. Xinwen Hou |
ICASSP | 2 |
| 2024 | GAN Inversion for Image Editing via Unsupervised Domain AdaptationabstractExisting GAN inversion methods work brilliantly in reconstructing high-quality (HQ) images while struggling with more common low-quality (LQ) inputs in practical application. To address this issue, we propose Unsupervised Domain Adaptation (UDA) in the inversion process, namely UDA-inversion, for effective inversion and editing of both HQ and LQ images. Regarding unpaired HQ images as the source domain and LQ images as the unlabeled target domain, we introduce a theoretical guarantee: loss value in the target domain is upper-bounded by loss in the source domain and a novel discrepancy function measuring the difference between two domains. Following that, we can only minimize this upper bound to obtain accurate latent codes for HQ and LQ images. Thus, constructive representations of HQ images can be spontaneously learned and transformed into LQ images without supervision. UDA-Inversion achieves a better PSNR of 22.14 on FFHQ dataset and performs comparably to supervised methods. Siyu Xing, Chen Gong 0005, Hewei Guo, Xinwen Hou, Yu Liu 0078 |
ICME | 5 |
| 2024 | Baffle: Hiding Backdoors in Offline Reinforcement Learning DatasetsabstractReinforcement learning (RL) makes an agent learn from trial-and-error experiences gathered during the interaction with the environment. Recently, offline RL has become a popular RL paradigm because it saves the interactions with environments. In offline RL, data providers share large pre-collected datasets, and others can train high-quality agents without interacting with the environments. This paradigm has demonstrated effectiveness in critical tasks like robot control, autonomous driving, etc. However, less attention is paid to investigating the security threats to the offline RL system. This paper focuses on backdoor attacks, where some perturbations are added to the data (observations) such that given normal observations, the agent takes high-rewards actions, and low-reward actions on observations injected with triggers. In this paper, we propose Baffle (Backdoor Attack for Offline Reinforcement Learning), an approach that automatically implants backdoors to RL agents by poisoning the offline RL dataset, and evaluate how different offline RL algorithms react to this attack. Our experiments conducted on four tasks and nine offline RL algorithms expose a disquieting fact: none of the existing offline RL algorithms has been immune to such a backdoor attack. More specifically, Baffle modifies 10% of the datasets for four tasks (3 robotic controls and 1 autonomous driving). Agents trained on the poisoned datasets perform well in normal settings. However, when triggers are presented, the agents’ performance decreases drastically by 63.2%, 53.9%, 64.7%, and 47.4% in the four tasks on average. The backdoor still persists after fine-tuning poisoned agents on clean datasets. We further show that the inserted backdoor is also hard to be detected by a popular defensive method. This paper calls attention to developing more effective protection for the open-source offline RL dataset. Chen Gong 0005, Zhou Yang 0003, Yunpeng Bai, Junda He, Jieke Shi, Kecen Li, Arunesh Sinha, Xinwen Hou, David Lo 0001, Tianhao Wang 0001 |
SP | 9 |
| 2024 | PrivImage: Differentially Private Synthetic Image Generation using Diffusion Models with Semantic-Aware Pretraining
Kecen Li, Chen Gong 0005, Yuzhong Zhao, Xinwen Hou, Tianhao Wang 0001 |
USENIX Security Symposium | 5 |
| 2023 | Frustratingly Easy Regularization on Representation Can Boost Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) gives the promise that an agent learns good policy from high-dimensional information, whereas representation learning removes irrelevant and redundant information and retains pertinent information. In this work, we demonstrate that the learned representation of the Q-network and its target Q-network should, in theory, satisfy a favorable distinguishable representation property. Specifically, there exists an upper bound on the representation similarity of the value functions of two adjacent time steps in a typical DRL setting. However, through illustrative experiments, we show that the learned DRL agent may violate this property and lead to a sub-optimal policy. Therefore, we propose a simple yet effective regularizer called P_olicy E_valuation with E_asy Regularization on Representation (PEER), which aims to maintain the distinguishable representation property via explicit regularization on internal representations. And we provide the convergence rate guarantee of PEER. Implementing PEER requires only one line of code. Our experiments demonstrate that incorporating PEER into DRL can significantly improve performance and sample efficiency. Comprehensive experiments show that PEER achieves state-of-the-art performance on all 4 environments on PyBullet, 9 out of 12 tasks on DM-Control, and 19 out of 26 games on Atari. To the best of our knowledge, PEER is the first work to study the inherent representation property of Q-network and its target. Our code is available at https://sites.google.com/view/peer-cvpr2023/. Huangyuan Su, Jieyu Zhang 0001, Xinwen Hou |
CVPR | 4 |
| 2023 | Template-guided Hierarchical Feature Restoration for Anomaly DetectionabstractTargeting for detecting anomalies of various sizes for complicated normal patterns, we propose a Template-guided Hierarchical Feature Restoration method, which introduces two key techniques, bottleneck compression and template-guided compensation, for anomaly-free feature restoration. Specially, our framework compresses hierarchical features of an image by bottleneck structure to preserve the most crucial features shared among normal samples. We design template-guided compensation to restore the distorted features towards anomaly-free features. Particularly, we choose the most similar normal sample as the template, and leverage hierarchical features from the template to compensate the distorted features. The bottleneck could partially filter out anomaly features, while the compensation further converts the reminding anomaly features towards normal with template guidance. Finally, anomalies are detected in terms of the cosine distance between the pre-trained features of an inference image and the corresponding restored anomaly-free features. Experimental results demonstrate the effectiveness of our approach, which achieves the state-of-the-art performance on the MVTec LOCO AD dataset. Hewei Guo, Liping Ren, Jingjing Fu, Yuwang Wang, Zhizheng Zhang 0004, Cuiling Lan, Haoqian Wang, Xinwen Hou |
ICCV | 8 |
| 2023 | Keep Various Trajectories: Promoting Exploration of Ensemble Policies in Continuous ControlabstractThe combination of deep reinforcement learning (DRL) with ensemble methods has been proved to be highly effective in addressing complex sequential decision-making problems. This success can be primarily attributed to the utilization of multiple models, which enhances both the robustness of the policy and the accuracy of value function estimation. However, there has been limited analysis of the empirical success of current ensemble RL methods thus far. Our new analysis reveals that the sample efficiency of previous ensemble DRL algorithms may be limited by sub-policies that are not as diverse as they could be. Motivated by these findings, our study introduces a new ensemble RL algorithm, termed \textbf{T}rajectories-awar\textbf{E} \textbf{E}nsemble exploratio\textbf{N} (TEEN). The primary goal of TEEN is to maximize the expected return while promoting more diverse trajectories. Through extensive experiments, we demonstrate that TEEN not only enhances the sample diversity of the ensemble policy compared to using sub-policies alone but also improves the performance over ensemble RL algorithms. On average, TEEN outperforms the baseline ensemble DRL algorithms by 41\% in performance on the tested representative environments. Chen Gong 0005, Xinwen Hou |
NeurIPS | 4 |
| 2022 | Curiosity-Driven and Victim-Aware Adversarial PoliciesabstractRecent years have witnessed great potential in applying Deep Reinforcement Learning (DRL) in various challenging applications, such as autonomous driving, nuclear fusion control, complex game playing, etc. However, recently researchers have revealed that deep reinforcement learning models are vulnerable to adversarial attacks: malicious attackers can train adversarial policies to tamper with the observations of a well-trained victim agent, the latter of which fails dramatically when faced with such an attack. Understanding and improving the adversarial robustness of deep reinforcement learning is of great importance in enhancing the quality and reliability of a wide range of DRL-enabled systems. Chen Gong 0005, Zhou Yang 0003, Yunpeng Bai, Jieke Shi, Arunesh Sinha, David Lo 0001, Xinwen Hou |
ACSAC | 8 |
| 2022 | POPO: Pessimistic Offline Policy OptimizationabstractOffline reinforcement learning (RL) aims to optimize policy from large pre-recorded datasets without interaction with the environment. This setting offers the promise of utilizing diverse and static datasets to obtain policies without costly, risky, active exploration. However, commonly used off-policy deep RL methods perform poorly when facing arbitrary off-policy datasets. In this work, we show that there exists an estimation gap of value-based deep RL algorithms in the offline setting. To eliminate the estimation gap, we propose a novel offline RL algorithm that we term Pessimistic Offline Policy Optimization (POPO), which learns a pessimistic value function. To demonstrate the effectiveness of POPO, we perform experiments on various quality datasets. And we find that POPO performs surprisingly well and scales to tasks with high-dimensional state and action space, comparing or outperforming tested state-of-the-art offline RL algorithms on benchmark tasks. Xinwen Hou, Yu Liu 0078 |
ICASSP | 2 |
| 2022 | Cooperative Multi-Agent Reinforcement Learning with Hypergraph ConvolutionabstractRecent years have witnessed the great success of multi-agent systems (MAS). Value decomposition, which decom-poses joint action values into individual action values, has been an important work in MAS. However, many value decomposition methods ignore the coordination among different agents, leading to the notorious “lazy agents” problem. To enhance the coordination in MAS, this paper proposes HyperGraph CoNvo-lution MIX (HGCN-MIX), a method that incorporates hyper-graph convolution with value decomposition. HGCN-MIX models agents as well as their relationships as a hypergraph, where agents are nodes and hyperedges among nodes indicate that the corresponding agents can coordinate to achieve larger rewards. Then, it trains a hypergraph that can capture the collaborative relationships among agents. Leveraging the learned hypergraph to consider how other agents' observations and actions affect their decisions, the agents in a MAS can better coordinate. We evaluate HGCN-MIX in the StarCraft II multi-agent challenge benchmark. The experimental results demonstrate that HGCN-MIX can train joint policies that outperform or achieve a similar level of performance as the current state-of-the-art techniques. We also observe that HGCN-MIX has an even more significant improvement of performance in the scenarios with a large amount of agents. Besides, we conduct additional analysis to emphasize that when the hypergraph learns more relationships, HGCN-MIX can train stronger joint policies. Yunpeng Bai, Chen Gong 0005, Bin Zhang 0052, Xinwen Hou, Yu Liu 0078 |
IJCNN | 5 |
| 2021 | Wide-Sense Stationary Policy Optimization with Bellman Residual on Video GamesabstractDeep Reinforcement Learning (DRL) has an increasing application in video games. However, it usually suffers from unstable training, low sampling efficiency, etc. Under the assumption that Bellman residual follows a stationary random process when the training process is convergent, we propose the Wide-sense Stationary Policy Optimization (WSPO) framework, which leverages the Wasserstein distance from the Bellman Residual Distribution (BRD) between two adjacent time steps, to stabilize the training stage and improve the sampling efficiency. We minimize the Wasserstein distance with Quantile Regression, where the specific form of BRD is not needed. Finally, we combine WSPO with Advantage Actor-Critic (A2C) algorithm and Deep Deterministic Policy Gradient (DDPG) algorithm. We evaluate WSPO on Atari 2600 video games and continuous control tasks, illustrating that WSPO compares or outperforms the state-of-the-art algorithms we tested. Chen Gong 0005, Yunpeng Bai, Xinwen Hou, Yu Liu 0078 |
ICME | 4 |
| 2021 | Minimizing Wasserstein-1 Distance by Quantile Regression for GANs Model
Xinwen Hou, Yu Liu 0078 |
PRCV (4) | 2 |
| 2020 | Highway Transformer: Self-Gating Enhanced Self-Attentive NetworksabstractSelf-attention mechanisms have made striking state-of-the-art (SOTA) progress in various sequence learning tasks, standing on the multiheaded dot product attention by attending to all the global contexts at different locations.Through a pseudo information highway, we introduce a gated component self-dependency units (SDU) that incorporates LSTM-styled gating units to replenish internal semantic importance within the multi-dimensional latent space of individual representations.The subsidiary content-based SDU gates allow for the information flow of modulated latent embeddings through skipped connections, leading to a clear margin of convergence speed with gradient descent algorithms.We may unveil the role of gating mechanism to aid in the contextbased Transformer modules, with hypothesizing that SDU gates, especially on shallow layers, could push it faster to step towards suboptimal points during the optimization process. Yekun Chai, Jin Shuo, Xinwen Hou |
ACL | 3 |
| 2020 | Stable Training of Bellman Error in Reinforcement Learning
Chen Gong 0005, Yunpeng Bai, Xinwen Hou, Xiaohui Ji |
ICONIP (5) | 3 |
| 2020 | WD3: Taming the Estimation Bias in Deep Reinforcement LearningabstractThe overestimation phenomenon caused by function approximation is a well-known issue in value-based reinforcement learning algorithms such as deep Q-networks and DDPG, which could lead to suboptimal policies. To address this issue, TD3 takes the minimum value between a pair of critics, which introduces underestimation bias. By unifying these two opposites, we propose a novel Weighted Delayed Deep Deterministic Policy Gradient algorithm, which can reduce the estimation error and further improve the performance by weighting a pair of critics. We compare the learning process of value function between DDPG, TD3, and our proposed algorithm, which verifies that our algorithm could indeed eliminate the estimation error of value function. We evaluate our algorithm in the OpenAI Gym continuous control tasks, outperforming the state-of-the-art algorithms on every environment tested. Xinwen Hou |
ICTAI | 2 |
| 2020 | An Improvement based on Wasserstein GAN for Alleviating Mode CollapsingabstractIn the past few years, Generative Adversarial Networks as a deep generative model has received more and more attention. Mode collapsing is one of the challenges in the study of Generative Adversarial Networks. In order to solve this problem, we deduce a new algorithm on the basis of Wasserstein GAN. We add a generated distribution entropy term to the objective function of generator net and maximize the entropy to increase the diversity of fake images. And then Stein Variational Gradient Descent algorithm is used for optimization. We named our method SW-GAN. In order to substantiate our theoretical analysis, we perform experiments on MNIST and CIFAR-10, and the results demonstrate superiority of our method. Xinwen Hou |
IJCNN | 2 |
| 2019 | Mixing Update Q-value for Deep Reinforcement LearningabstractThe value-based reinforcement learning methods are known to overestimate action values such as deep Q-learning, which could lead to suboptimal policies. This problem also persists in an actor-critic algorithm. In this paper, we propose a novel mechanism to minimize its effects on both the critic and the actor. Our mechanism builds on Double Q-learning, by mixing update action value based on the minimum and maximum between a pair of critics to limit the overestimation. We then propose a specific adaptation to the Twin Delayed Deep Deterministic policy gradient algorithm (TD3) and show that the resulting algorithm not only reduces the observed overestimations, as hypothesized, but that this also leads to much better performance on several tasks. Zhunan Li, Xinwen Hou |
IJCNN | 2 |
| 2019 | Exponential Moving Averaged Q-Network for DDPG
Xiangxiang Shen, Chuanhuan Yin, Yekun Chai, Xinwen Hou |
PRCV (1) | 4 |
| 2017 | Traffic Sign Detection Using a Cascade Method With Fast Feature Extraction and Saliency TestabstractAutomatic traffic sign detection is challenging due to the complexity of scene images, and fast detection is required in real applications such as driver assistance systems. In this paper, we propose a fast traffic sign detection method based on a cascade method with saliency test and neighboring scale awareness. In the cascade method, feature maps of several channels are extracted efficiently using approximation techniques. Sliding windows are pruned hierarchically using coarse-to-fine classifiers and the correlation between neighboring scales. The cascade system has only one free parameter, while the multiple thresholds are selected by a data-driven approach. To further increase speed, we also use a novel saliency test based on mid-level features to pre-prune background windows. Experiments on two public traffic sign data sets show that the proposed method achieves competing performance and runs 2~7 times as fast as most of the state-of-the-art methods. Xinwen Hou, Jiawei Xu 0004, Shigang Yue, Cheng-Lin Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2016 | Exploiting coarse-to-fine mechanism for fine-grained recognitionabstractFine-grained object recognition is more challenging than generic categorization due to the subtle difference between subcategories under the large intra-class pose change and appearance variations. The state-of-the-art fine-grained recognition methods usually utilize part detection or pose alignment to alleviate the pose variation, and then use convolutional neural networks (CNNs) to extract local discriminative features. Although the hierarchical structure of deep CNNs enables rich and discriminative visual feature extraction, the recognition methods so far mostly use the features of only the last convolutional layer for classification. In this paper, by exploiting the correlation of the convolutional features of within-layer and between-layer, we propose a method to integrate multi-layer convolutional features based on coarse-to-fine mechanism for improving the discrimination capability. Experiments on a number of public datasets show that the proposed method, without part annotation or pose alignment, yields superior or comparable performance to the state-of-the-art methods. Yongzhong Wang, Xu-Yao Zhang, Yan-Ming Zhang 0001, Xinwen Hou, Cheng-Lin Liu 0001 |
ICIP | 4 |
| 2015 | A saliency-based cascade method for fast traffic sign detectionabstractWe propose a cascade method for fast and accurate traffic sign detection. The main feature of the method is that mid-level saliency test is used to efficiently and reliably eliminate background windows. Fast feature extraction is adopted in the subsequent stages for rejecting more negatives. Combining with neighbor scales awareness in window search, the proposed method runs at 3~5 fps for high resolution (1360×800) images, 2~7 times as fast as most state-of-the-art methods. Compared with them, the proposed method yields competitive performance on prohibitory signs while sacrifices performance moderately on danger and mandatory signs. Shigang Yue, Jiawei Xu 0004, Xinwen Hou, Cheng-Lin Liu 0001 |
Intelligent Vehicles Symposium | 4 |
| 2014 | Learning Locality Preserving Graph from DataabstractMachine learning based on graph representation, or manifold learning, has attracted great interest in recent years. As the discrete approximation of data manifold, the graph plays a crucial role in these kinds of learning approaches. In this paper, we propose a novel learning method for graph construction, which is distinct from previous methods in that it solves an optimization problem with the aim of directly preserving the local information of the original data set. We show that the proposed objective has close connections with the popular Laplacian Eigenmap problem, and is hence well justified. The optimization turns out to be a quadratic programming problem with n(n-1)/2 variables (n is the number of data points). Exploiting the sparsity of the graph, we further propose a more efficient cutting plane algorithm to solve the problem, making the method better scalable in practice. In the context of clustering and semi-supervised learning, we demonstrated the advantages of our proposed method by experiments. Yan-Ming Zhang 0001, Kaizhu Huang, Xinwen Hou, Cheng-Lin Liu 0001 |
IEEE Trans. Cybern. | 3 |
| 2012 | Multiple Outlooks Learning with Support Vector Machines
Yinglu Liu, Xu-Yao Zhang, Kaizhu Huang, Xinwen Hou, Cheng-Lin Liu 0001 |
ICONIP (3) | 4 |
| 2012 | Insect species recognition using discriminative local soft coding
An Lu, Xinwen Hou, Cheng-Lin Liu 0001 |
ICPR | 2 |
| 2011 | A Hybrid Approach to Detect and Localize Texts in Natural Scene ImagesabstractText detection and localization in natural scene images is important for content-based image analysis. This problem is challenging due to the complex background, the non-uniform illumination, the variations of text font, size and line orientation. In this paper, we present a hybrid approach to robustly detect and localize texts in natural scene images. A text region detector is designed to estimate the text existing confidence and scale information in image pyramid, which help segment candidate text components by local binarization. To efficiently filter out the non-text components, a conditional random field (CRF) model considering unary component properties and binary contextual component relationships with supervised parameter learning is proposed. Finally, text components are grouped into text lines/words with a learning-based energy minimization method. Since all the three stages are learning-based, there are very few parameters requiring manual tuning. Experimental results evaluated on the ICDAR 2005 competition dataset show that our approach yields higher precision and recall performance compared with state-of-the-art methods. We also evaluated our approach on a multilingual image dataset with promising results. Yi-Feng Pan, Xinwen Hou, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2010 | Transductive Learning on Adaptive GraphsabstractGraph-based semi-supervised learning methods are based on some smoothness assumption about the data. As a discrete approximation of the data manifold, the graph plays a crucial role in the success of such graph-based methods. In most existing methods, graph construction makes use of a predefined weighting function without utilizing label information even when it is available. In this work, by incorporating label information, we seek to enhance the performance of graph-based semi-supervised learning by learning the graph and label inference simultaneously. In particular, we consider a particular setting of semi-supervised learning called transductive learning. Using the LogDet divergence to define the objective function, we propose an iterative algorithm to solve the optimization problem which has closed-form solution in each step. We perform experiments on both synthetic and real data to demonstrate improvement in the graph and in terms of classification accuracy. Yan-Ming Zhang 0001, Yu Zhang 0006, Dit-Yan Yeung, Cheng-Lin Liu 0001, Xinwen Hou |
AAAI | 5 |
| 2010 | Gaussian Process Latent Random FieldabstractIn this paper, we propose a novel supervised extension of GPLVM, called Gaussian process latent random field (GPLRF), by enforcing the latent variables to be a Gaussian Markov random field with respect to a graph constructed from the supervisory information. Guoqiang Zhong 0001, Wu-Jun Li, Dit-Yan Yeung, Xinwen Hou, Cheng-Lin Liu 0001 |
AAAI | 4 |
| 2010 | Fast scene text localization by learning-based filtering and verificationabstractThis paper proposes a new method for fast text localization in natural scene images by combining learning-based region filtering and verification in a coarse-to-fine strategy. In each pyramid layer, a boosted region filter is used to extract candidate text regions, which are segmented into candidate text lines by multi-orientation projection analysis. A polynomial classifier with combined features is used to verify patches of candidate text lines for removing non-texts. The remaining text patches over all pyramid layers are grouped into text lines based on their spatial relationships. The text lines are further refined and partitioned into words by connected component analysis. Experimental results show that the proposed method provides competitive localization performance at high speed. Yi-Feng Pan, Cheng-Lin Liu 0001, Xinwen Hou |
ICIP | 3 |
| 2010 | Multi-class AdaBoost with Hypothesis MarginabstractMost AdaBoost algorithms for multi-class problems have to decompose the multi-class classification into multiple binary problems, like the Adaboost.MH and the LogitBoost. This paper proposes a new multi-class AdaBoost algorithm based on hypothesis margin, called AdaBoost.HM, which directly combines multi-class weak classifiers. The hypothesis margin maximizes the output about the positive class meanwhile minimizes the maximal outputs about the negative classes. We discuss the upper bound of the training error about AdaBoost.HM and a previous multi-class learning algorithm AdaBoost.M1. Our experiments using feed forward neural networks as weak learners show that the proposed AdaBoost.HM yields higher classification accuracies than the AdaBoost.M1 and the AdaBoost.MH, and meanwhile, AdaBoost.HM is computationally efficient in training. Xiao-Bo Jin, Xinwen Hou, Cheng-Lin Liu 0001 |
ICPR | 2 |
| 2010 | Boosting Incremental Semi-supervised Discriminant Analysis for TrackingabstractTracking is recently formulated as a problem of discriminating the object from its nearby background, where the classifier is updated by new samples successively arriving during tracking. Depending on whether labeling the samples or not, the tracker can be designed in a supervised or semi-supervised manner. This paper proposes a novel semi-supervised algorithm for tracking by combining Semi-supervised Discriminant Analysis (SDA) with an online boosting framework. Using the local geometric structure information from the samples, the SDA-based weak classifier is made more robust to outliers. Meanwhile, we design an incremental updating mechanism for SDA so that it can adapt to appearance changes. We further propose an Extended SDA (ESDA) algorithm, which gives better discrimination ability. Results on several challenging video sequences demonstrate the effectiveness of the method. Xinwen Hou, Cheng-Lin Liu 0001 |
ICPR | 2 |
| 2010 | Regularized margin-based conditional log-likelihood loss for prototype learning
Xiao-Bo Jin, Cheng-Lin Liu 0001, Xinwen Hou |
Pattern Recognit. | 3 |
| 2009 | Text Localization in Natural Scene Images Based on Conditional Random FieldabstractThis paper proposes a novel hybrid method to robustly and accurately localize texts in natural scene images. A text region detector is designed to generate a text confidence map, based on which text components can be segmented by local binarization approach. A conditional random field (CRF) model, considering the unary component property as well as binary neighboring component relationship, is then presented to label components as "text" or "non-text". Last, text components are grouped into text lines with an energy minimization approach. Experimental results show that the proposed method gives promising performance comparing with the existing methods on ICDAR 2003 competition dataset. Yi-Feng Pan, Xinwen Hou, Cheng-Lin Liu 0001 |
ICDAR | 2 |
| 2009 | Object tracking by bidirectional learning with feature selectionabstractThis paper proposes a new tracking algorithm which combines object and background information, via building object and background appearance models simultaneously by non-parametric kernel density estimation. The major contribution is a novel bidirectional learning framework for discrimination between the object and background. It has the following advantages: 1) it embeds background information, unlike most other methods that focus on the object only, 2) it provides a mechanism to detect occlusion and distraction, which are two main causes of tracking failure, 3) it performs feature selection, making the tracker more robust to outliers. By this learning framework, we are able to embed discriminative information into the generative appearance model. Experimental results demonstrate that the tracker is able to model drastic appearance changes and robust to occlusion and distraction. Xinwen Hou, Cheng-Lin Liu 0001 |
ICIP | 2 |
| 2009 | Subspace Regularization: A New Semi-supervised Learning Method
Yan-Ming Zhang 0001, Xinwen Hou, Shiming Xiang, Cheng-Lin Liu 0001 |
ECML/PKDD (2) | 2 |
| 2008 | A Robust System to Detect and Localize Texts in Natural Scene ImagesabstractIn this paper, we present a robust system to accurately detect and localize texts in natural scene images. For text detection, a region-based method utilizing multiple features and cascade AdaBoost classifier is adopted. For text localization, a window grouping method integrating text line competition analysis is used to generate text lines. Then within each text line, local binarization is used to extract candidate connected components (CCs) and non-text CCs are filtered out by Markov Random Fields (MRF) model, through which text line can be localized accurately. Experiments on the public benchmark ICDAR 2003 Robust Reading and Text Locating Dataset show that our system is comparable to the best existing methods both in accuracy and speed. Yi-Feng Pan, Xinwen Hou, Cheng-Lin Liu 0001 |
Document Analysis Systems | 2 |
| 2008 | Prototype learning with margin-based conditional log-likelihood lossabstractThe classification performance of nearest prototype classifiers largely relies on the prototype learning algorithms, such as the learning vector quantization (LVQ) and the minimum classification error (MCE). This paper proposes a new prototype learning algorithm based on the minimization of a conditional log-likelihood loss (CLL), called log-likelihood of margin (LOGM). A regularization term is added to avoid over-fitting in training. The CLL loss in LOGM is a convex function of margin, and so, gives better convergence than the MCE algorithm. Our empirical study on a large suite of benchmark datasets demonstrates that the proposed algorithm yields higher accuracies than the MCE, the generalized LVQ (GLVQ), and the soft nearest prototype classifier (SNPC). Xiao-Bo Jin, Cheng-Lin Liu 0001, Xinwen Hou |
ICPR | 3 |
| 2008 | A pooled subspace mixture density model for pattern classification in high-dimensional spacesabstractDensity estimation in high-dimensional data spaces is a challenge due to the sparseness of data which is known as ldquothe curse of dimensionalityrdquo. Researchers often resort to low-dimensional subspaces for such tasks, while discard the distribution in the complementary subspace. In this paper, we propose a new mixture density model based on pooled subspace. In our method, the Gaussian components of each class share a subspace and the complementary subspace is incorporated in the density function. The subspace and Gaussian mixture density are estimated simultaneously in EM iteration steps. We apply the density model to pattern classification in experiments on UCI datasets and compare the proposed method with previous ones. The experimental results demonstrate the superiority of the proposed method. Xiao-Hua Liu, Cheng-Lin Liu 0001, Xinwen Hou |
IJCNN | 3 |
| 2006 | Recognize Multi-people Interaction Activity by PCA-HMMs
Ying Wang 0003, Xinwen Hou, Tieniu Tan |
ACCV (1) | 2 |
| 2006 | Learning Boosted Asymmetric Classifiers for Object DetectionabstractObject detection can be posted as those classification tasks where the rare positive patterns are to be distinguished from the enormous negative patterns. To avoid the danger of missing positive patterns, more attention should be payed on them. Therefore there should be different requirements for False Reject Rate (FRR) and False Accept Rate (FAR) , and learning a classifier should use an asymmetric factor to balance between FRR and FAR. In this paper, a normalized asymmetric classification error is proposed for the task of rejecting negative patterns. Minimizing it not only controls the ratio of FRR and FAR, but more importantly limits the upper-bound of FRR. The latter characteristic is advantageous for those tasks where there is a requirement for low FRR. Based on this normalized asymmetric classification error, we develop an asymmetric AdaBoost algorithm with variable asymmetric factor and apply it to the learning of cascade classifiers for face detection. Experiments demonstrate that the proposed method achieves less complex classifiers and better performance than some previous AdaBoost methods. Xinwen Hou, Cheng-Lin Liu 0001, Tieniu Tan |
CVPR (1) | 1 |
| 2005 | Learning multiview face subspaces and facial pose estimation using independent component analysisabstractAn independent component analysis (ICA) based approach is presented for learning view-specific subspace representations of the face object from multiview face examples. ICA, its variants, namely independent subspace analysis (ISA) and topographic independent component analysis (TICA), take into account higher order statistics needed for object view characterization. In contrast, principal component analysis (PCA), which de-correlates the second order moments, can hardly reveal good features for characterizing different views, when the training data comprises a mixture of multiview examples and the learning is done in an unsupervised way with view-unlabeled data. We demonstrate that ICA, TICA, and ISA are able to learn view-specific basis components unsupervisedly from the mixture data. We investigate results learned by ISA in an unsupervised way closely and reveal some surprising findings and thereby explain underlying reasons for the emergent formation of view subspaces. Extensive experimental results are presented. Stan Z. Li, Xiaoguang Lu, Xinwen Hou, Xianhua Peng |
IEEE Trans. Image Process. | 3 |
| 2001 | Direct Appearance ModelsabstractActive appearance model (AAM), which makes ingenious use of both shape and texture constraints, is a powerful tool for face modeling, alignment and facial feature extraction under shape deformations and texture variations. However, as we show through our analysis and experiments, there exist admissible appearances that are not modeled by AAM and hence cannot be reached by AAM search; also the mapping from the texture subspace to the shape subspace is many-to-one and therefore a shape should be determined entirely by the texture in it. We propose a new appearance model, called direct appearance model (DAM), without combining from shape and texture as in AAM. The DAM model uses texture information directly in the prediction of the shape and in the estimation of position and appearance (hence the name DAM). In addition, DAM predicts the new face position and appearance based on principal components of texture difference vectors, instead of the raw vectors themselves as in AAM. These lead to the following advantages over AAM: (1) DAM subspaces include admissible appearances previously unseen in AAM, (2) convergence and accuracy are improved, and (3) memory requirement is cut down to a large extent. The advantages are substantiated by comparative experimental results. Xinwen Hou, Stan Z. Li, HongJiang Zhang |
CVPR (1) | 1 |
| 2001 | Learning Spatially Localized, Parts-Based RepresentationabstractIn this paper, we propose a novel method, called local non-negative matrix factorization (LNMF), for learning spatially localized, parts-based subspace representation of visual patterns. An objective function is defined to impose a localization constraint, in addition to the non-negativity constraint in the standard NMF. This gives a set of bases which not only allows a non-subtractive (part-based) representation of images but also manifests localized features. An algorithm is presented for the learning of such basic components. Experimental results are presented to compare LNMF with the NMF and PCA methods for face representation and recognition, which demonstrates advantages of LNMF. Stan Z. Li, Xinwen Hou, HongJiang Zhang |
CVPR (1) | 2 |
| 2001 | Learning Low Dimensional Invariant Signature of 3-D Object under Varying View and Illumination from 2-D Appearances
Stan Z. Li, Xinwen Hou, HongJiang Zhang |
ICCV | 3 |