EDBT 2026 Demo / reviewers in the wild / expert
Hao Liu 0019
dblp:09/3214-19
· DBLP profile ↗
42ranked-venue papers
11as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 5 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion-Based Multi-Agent With Reinforcement Learning for Multimodal-Based RecommendationabstractMultimodal-based recommendation integrates visual, textual, and acoustic modalities of items to comprehensively capture user preferences, thereby playing a crucial role in modern multimedia online platforms. In this paper, we propose aDiffusion-based Multi-agent withReinforcement Learning forMultimodal-basedRecommendation (DRMRec). Specifically, our DRMRec first leverages a hierarchicalDiffusion-basedMulti-Agent (DMA) to reconstruct the user-item interaction graph. Subsequently, these reconstructed graphs are fused with multimodal information to form contrastive views. During the reconstruction, each agent is injected with multimodal information through aCross-ModalAligner (CMA), thereby bringing cross-modal information closer to interaction embeddings and facilitating alignment across modalities. Then, we introduce aMetric-AwareDiffusionReinforcer (MADR), a reinforcement learning framework that leverages validation-set recommendation metrics as reward signals to enable dynamic and individual fine-tuning of each agent, thereby actively aligning model optimization with recommendation tasks. Next, we apply cross-modal graph contrastive learning to contrastive views, alleviating data sparsity while further enhancing cross-modal alignment. Extensive experiments on three real-world multimedia platform datasets demonstrate that DRMRec consistently outperforms state-of-the-art approaches in multimodal-based recommendation. It is noteworthy that the performance improvements across all three metrics on both the TikTok and Sports datasets exceed 10%. Rui Tang 0020, Hao Liu 0019, Xian Mo |
IEEE Trans. Big Data | 3 |
| 2025 | Appearance Contrasts for Unconstrained Age EstimationabstractIn this paper, we propose a Dual-Constraint Diffusion Model (DCDM) to contrast aging appearance for facial age estimation, addressing the key issue of noisy labels. Existing methods for face age estimation are plagued by class imbalance and noisy supervision signals, which disrupt the ordinal relationships between age categories and hinder effective feature decoupling in existing models. To overcome these challenges, the proposed DCDM develops a label-independent Paired Comparison, ensuring accurate sample labeling and maintaining continuity in age estimation. Moreover, we incorporate a Dual-Constraint Diffusion Model to effectively separate and recombine age-related and unrelated features, thus facilitating the generation of high-fidelity and continuous age-progressed facial representations. Lastly, we optimize our model parameters by exploiting the age difference information via an active learning framework. Comparative evaluations on several in-the-wild datasets demonstrate that our DCDM significantly achieves superior results compared to existing state-of-the-art methods in facial age estimation. Jilong Wei, Yangyang Hu, Xiangjuan Wu, Yiqiang Wu, Hao Liu 0019 |
ACM Multimedia | 5 |
| 2025 | Low-Rank Approximation CLIP to Improve Cross-Modal Consistency in Language-Guided Age Estimation
Jilong Wei, Xiangjuan Wu, Jiao Feng, Hao Liu 0019 |
PRCV (15) | 5 |
| 2025 | Intelligible graph contrastive learning with attention-aware for recommendation
Xian Mo, Zihang Zhao, Xiaoru He, Hao Liu 0019 |
Neurocomputing | 5 |
| 2025 | Multi-relational graph contrastive learning with learnable graph augmentationabstractMulti-relational graph learning aims to embed entities and relations in knowledge graphs into low-dimensional representations, which has been successfully applied to various multi-relationship prediction tasks, such as information retrieval, question answering, and etc. Recently, contrastive learning has shown remarkable performance in multi-relational graph learning by data augmentation mechanisms to deal with highly sparse data. In this paper, we present a Multi-Relational Graph Contrastive Learning architecture (MRGCL) for multi-relational graph learning. More specifically, our MRGCL first proposes a Multi-relational Graph Hierarchical Attention Networks (MGHAN) to identify the importance between entities, which can learn the importance at different levels between entities for extracting the local graph dependency. Then, two graph augmented views with adaptive topology are automatically learned by the variant MGHAN, which can automatically adapt for different multi-relational graph datasets from diverse domains. Moreover, a subgraph contrastive loss is designed, which generates positives per anchor by calculating strongly connected subgraph embeddings of the anchor as the supervised signals. Comprehensive experiments on multi-relational datasets from three application domains indicate the superiority of our MRGCL over various state-of-the-art methods. Our datasets and source code are published at https://github.com/Legendary-L/MRGCL. Xian Mo, Jun Pang 0001, Binyuan Wan, Rui Tang 0020, Hao Liu 0019, Shuyu Jiang |
Neural Networks | 5 |
| 2025 | ISDAT: An image-semantic dual adversarial training framework for robust image classification
Chenhong Sui, Hao Liu 0019, Qingtao Gong, Jing Yao 0002, Danfeng Hong |
Pattern Recognit. | 4 |
| 2025 | Knowledge-Aware Diffusion-Enhanced Multimedia RecommendationabstractMultimedia recommendations aim to use rich multimedia content to enhance historical user-item interaction information, which can not only indicate the content relatedness among items but also reveal finer-grained preferences of users. In this paper, we propose aKnowledge-awareDiffusion-Enhanced architecture using contrastive learning paradigms (KDiffE) for multimedia recommendations. Specifically, we first utilize original user-item graphs to build an attention-aware matrix into graph neural networks, which can learn the importance between users and items for main view construction. The attention-aware matrix is constructed by adopting a random walk with a restart strategy, which can preserve the importance between users and items to generate aggregation of attention-aware node features. Then, we propose a guided diffusion model to generate strongly task-relevant knowledge graphs with less noise for constructing a knowledge-aware contrastive view, which utilizes user embeddings with an edge connected to an item to guide the generation of strongly task-relevant knowledge graphs for enhancing the item's semantic information. We perform comprehensive experiments on three multimedia datasets that reveal the effectiveness of our KDiffE and its components on various state-of-the-art methods. Our source codes are available. Xian Mo, Rui Tang 0020, Jin-Tao Gao, Hao Liu 0019 |
IEEE Trans. Multim. | 5 |
| 2024 | Toward Quantifiable Face Age Transformation Under Attribute UnbiasabstractPrevious works condition aging patterns utilizing one-hot or artificial predefined distributions. Nevertheless, different age groups show different intraclass variations. This property made it challenging to express differences in apparent age across all age groups discriminately. Adaptive aging feature distribution by learning the target age group in training data is a promising solution. Unfortunately, existing datasets commonly suffer from diverse degrees of semantic-level attribute imbalance, which leads to the tendency for previous approaches to generate paradoxical appearances. To address the aforementioned issues, we propose a novel framework containing three modules: the Causal Aging (CA) module, the Shapley Value Quantization (SVQ) module, and the Differentiated Age Embedding Transformation (DAT) module. Specifically, to eliminate the effect of attribute imbalance on the adaptive distribution of learning target age groups, we design the CA module, which controls the effect of momentum on aging features by De-confound training. Meanwhile, the influence of the aging-independent attribute, which appears abundantly in training data, on the target aging feature is eliminated by counterfactual inference subtraction. Subsequently, the SVQ module quantifies the contribution of different attributes to age based on the results of the CA module. This operation allows us to obtain adaptive age distributions for different age groups. Eventually, the DAT module takes a target age vector, sampled from the target age distribution quantized by SVQ, and modulates the age representation of the generated image. Extensive experimental results on four face aging datasets show that our model achieves convincing performance compared to the current state-of-the-art methods. Ling Lin 0002, Hao Liu 0019, Congcong Zhu, Jingrun Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Consensus-Agent Deep Reinforcement Learning for Face AgingabstractFace aging tasks aim to simulate changes in the appearance of faces over time. However, due to the lack of data on different ages under the same identity, existing models are commonly trained using mapping between age groups. This makes it difficult for most existing aging methods to accurately capture the correspondence between individual identities and aging features, leading to generating faces that do not match the real aging appearance. In this paper, we re-annotate the CACD2000 dataset and propose a consensus-agent deep reinforcement learning method to solve the aforementioned problem. Specifically, we define two agents, the aging process agent and the aging personalization agent, and model the task of matching aging features as a Markov decision process. The aging process agent simulates the aging process of an individual, while the aging personalization agent calculates the difference between the aging appearance of an individual and the average aging appearance. The two agents iteratively adjust the matching degree between the target aging feature and the current identity through a form of synergistic cooperation. Extensive experimental results on four face aging datasets show that our model achieves convincing performance compared to the current state-of-the-art methods. Ling Lin 0002, Hao Liu 0019, Jinqiao Liang, Jiao Feng, Hu Han 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | HeadDiff: Exploring Rotation Uncertainty With Diffusion Models for Head Pose EstimationabstractIn this paper, we propose a probabilistic regression diffusion model for head pose estimation, dubbed HeadDiff, which typically addresses the rotation uncertainty, especially when faces are captured in wild conditions. Unlike conventional image-to-pose methods which cannot explicitly establish the rotational manifold of head poses, our HeadDiff aims to ensure the pose rotation via the diffusion process and in parallel, refine the mapping process iteratively. Specifically, we initially formulate the head pose estimation problem as a reverse diffusion process, defining a paradigm for progressive denoising on the manifold, which explores the uncertainty by decomposing the large gap into intermediate steps. Moreover, our HeadDiff is equipped with an isotropic Gaussian distribution by encoding the incoherence information in our rotation representation. Finally, we learn the facial relationship of nearest neighbors with a cycle-consistent constraint for robust pose estimation versus diverse shape variations. Experimental results on multiple datasets demonstrate that our proposed method outperforms existing state-of-the-art techniques without auxiliary data. Yaoxing Wang, Hao Liu 0019, Yaowei Feng, Xiangjuan Wu, Congcong Zhu |
IEEE Trans. Image Process. | 2 |
| 2023 | Joint Statistical and Causal Feature Modulated Face Anti-SpoofingabstractIn this paper, we propose a hierarchical feature modulation (HFM) approach for stable face anti-spoofing in unseen domains and unseen attacks. The conventional multi-domain based generalizable approaches likely lead to local optima due to the complicated or heuristic learning paradigm. Inspired by the fact that high-level semantic disturbances and low-level miscellaneous bias jointly cause the distribution shift, HFM aims to modulate the fine-grained feature in a hierarchical manner. Specifically, we complement the structural feature with patch-wise learnable statistical information, i.e. local difference histogram, to relieve the overfitting on high-level semantics. We further introduce the structural causal model (SCM) with imaging color model to reveal that presenting mediums and capturing devices destroy the liveness-relevant information from the low level. Thus we model this hidden entanglement as a distribution mixture problem and propose the expectation-maximization (EM) based causal intervention to remove these miscellanies. Experimental results on public datasets demonstrate the effectiveness of HFM, especially in out-of-distribution settings. Hao Liu 0019 |
ICME | 4 |
| 2023 | Cross-Modality Fourier Feature for Medical Image SynthesisabstractIn this paper, we propose a cross-modality fourier feature (CMFF) method via frequency selection, which learns the rational anatomical structure for targeting medical modality images. Unlike existing works seeking pixel-wise intensity discrepancy likely misleading to bias anatomical structures, our approach strongly holds the medical prior that different modality MRI images should share the same anatomical structure. To achieve this, our method instead learns to convert MRI images to the auxiliary frequency domain. Moreover, we adopt the Shapley value to quantify the contribution of each frequency, with respect to the structure for MRI image pairs of different modalities. Thus our approach learns to refine the anatomical structure generated by the target modality iteratively. Extensive experimental results on the BraTS dataset show that our model surpasses the performance compared to SOTAs. Mei Ma, Ling Lin 0002, Hao Liu 0019 |
ICME | 5 |
| 2023 | Structural Equivariance Self-Supervised Learning for Facial Pose EstimationabstractIn this paper, we propose a self-supervised learning method for robust facial pose estimation. Conventional methods usually split the coherent head motion into discrete and finite outputs, likely leading to bias prediction because the performance of head pose estimation highly relies on structural facial appearance. To address this issue, our model achieves structural equivariance to poses through a self-supervised learning strategy from extrinsic attributes of face neighbors and underlying local associations. Specifically, we construct a complete neighbor graph to capture the extrinsic properties of face neighbors, where different latent semantic attributes are assigned to each subgraph. Accordingly, we design a set of proxy tasks based on different attribute subgraphs, where the model is encouraged to learn the underlying relation of local features under pose variation. Extensive experimental results on the challenging, widely evaluated datasets indicate the effectiveness of our model compared with the state of the arts. Yaoxing Wang, Xian Mo, Hao Liu 0019 |
ICME | 5 |
| 2023 | A relation-aware heterogeneous graph convolutional network for relationship prediction
Xian Mo, Rui Tang 0020, Hao Liu 0019 |
Inf. Sci. | 3 |
| 2023 | UCSL: Toward Unsupervised Common Subspace Learning for Cross-Modal Image ClassificationabstractThe emerging research line of cross-modal learning focuses on the issue of transferring feature representation manner learned from limited multimodal data with labelings to the testing phase with partial modalities. This is essentially common and practical in the remote sensing community when only modal-incomplete data are in users’ hands due to inevitable imaging or access restrictions under large-scale observation scenarios. However, most of the existing cross-modal learning methods have been designed with exclusive reliance on labeling, which can be either limited or noisy due to their costly production. To address this issue, we explore in this paper the possibility to learn cross-modal feature representation in an unsupervised fashion. By integrating the multimodal data into a fully recombined matrix form, we propose 1) the use of common subspace representation as the regression target instead of conventionally adopted binary labels, and 2) the orthogonality and manifold alignment regularization terms to shrink the solution space whilst preserving the pairwise manifold correlations. Through this manner, the modality-specific and mutual latent representations in this common subspace as well as their corresponding projections can be learned simultaneously and their optimums can be efficiently reached through a nearly one-step computation with the help of Eigen decomposition. Finally, we show the superiority of our method through extensive image classification experiments on three multimodal datasets with four remotely sensed modalities involved (i.e., hyperspectral, multispectral, synthetic aperture radar, and light detection and ranging data). The code and dataset will be made freely available at https://github.com/jingyao16/UCSL after a possible publication to encourage the reproduction of our method and further use. Jing Yao 0002, Danfeng Hong, Hao Liu 0019, Jocelyn Chanussot |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Exploiting Unfairness With Meta-Set Learning for Chronological Age EstimationabstractFacial age estimation aims to rank the face aging data by taking in the correlation among age categories. Conventional age estimation models are trained based on assumed high-quality training annotations in a totally-supervised manner. However, noisy data in a sparse distribution collected from unconstrained environment usually account for the corruption of produced gradients and ordinal relationships, which may fail to fairly describe the correlated face aging data. In this paper, we propose a meta-set learning (MSL) approach for exploiting the unfairness of face aging datasets, thus achieving unbiased age classification in unconstrained conditions. To address this, we elaborately create an unfairness filtration network under the meta-learning paradigm, which exceeds a reliable margin-reweighting initialization suffering from class variance, simultaneously exploiting the meta-reweighting intervention to minimize the training bias caused by class imbalance. Moreover, our proposed model leverages the learned instance-level margin between logits and develops a unimodal constrained logits loss, further surviving age regression models from unfairness. Experimental results on multiple in-the-wild datasets demonstrate that our proposed method achieves superior results compared to existing state-of-the-art methods. Chenyang Wang 0004, Xian Mo, Xiaofen Tang, Hao Liu 0019 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Siamese Graph Learning for Semi-Supervised Age EstimationabstractIn this paper, we propose a Siamese graph learning (SGL) approach to alleviate aging dataset bias. While numerous semi-supervised algorithms have been successfully applied to classification tasks, most of them assume that both the labeled and unlabeled samples are drawn from identical distributions. However, this assumption may not hold due to the heterogeneity of face aging data, which gives rise to a bias and unpromising prediction. Motivated by this, our SGL learns to align the sparse distribution with the dense one for dataset debias with preserving the real aging smoothness. To achieve this, we adopt a mixup strategy to plausibly generate hallucinatory samples, which leverages amounts of unlabeled data to enhance the diversity of unbalanced classes. Moreover, we develop a graph contrastive regularization to suppress the noise introduced by auxiliary unlabeled samples. Extensive experimental results show compelling performance by only utilizing the limited scalability of training annotations. Hao Liu 0019, Mei Ma, Zixian Gao, Zongyong Deng, Fengjun Li |
IEEE Trans. Multim. | 1 |
| 2022 | Meta Descent Learning for Class Imbalanced Age EstimationabstractIn this paper, we propose a meta descent learning method (MDL) for class imbalanced age estimation while preserving the relative ordinal information. The class imbalanced problem causes head classes with enough samples dominant the gradient descent process, thus suppressing the performance of tail classes. Due to its superiority for quick adaptive descent optimization, we utilize meta learning to adjust the learning gradient iteratively for balancing training. Specifically, we reweight the sample loss using meta descent learning to prevent the dominance of head classes in feature space. Furthermore, we propose an order-consistent loss to explicitly constrain the ordinal information on the output logits, helping the exploit of semantic correlation in aging images. Experimental results on several datasets including Morph II and ChaLearn LAP demonstrate the effectiveness of our method. Hao Liu 0019 |
ICME | 3 |
| 2021 | PML: Progressive Margin Loss for Long-Tailed Age ClassificationabstractIn this paper, we propose a progressive margin loss (PML) approach for unconstrained facial age classification. Conventional methods make strong assumption on that each class owns adequate instances to outline its data distribution, likely leading to bias prediction where the training samples are sparse across age classes. Instead, our PML aims to adaptively refine the age label pattern by enforcing a couple of margins, which fully takes in the in-between discrepancy of the intra-class variance, inter-class variance and class center. Our PML typically incorporates with the ordinal margin and the variational margin, simultaneously plugging in the globally-tuned deep neural network paradigm. More specifically, the ordinal margin learns to exploit the correlated relationship of the real-world age labels. Accordingly, the variational margin is leveraged to minimize the influence of head classes that misleads the prediction of tailed samples. Moreover, our optimization carefully seeks a series of indicator curricula to achieve robust and efficient model training. Extensive experimental results on three face aging datasets demonstrate that our PML achieves compelling performance compared to state of the art. Code will be made publicly. Zongyong Deng, Hao Liu 0019, Yaoxing Wang, Chenyang Wang 0004, Zekuan Yu, Xuehong Sun |
CVPR | 2 |
| 2021 | Open Set Face Anti-Spoofing in Unseen AttacksabstractIn this paper, we propose an end-to-end open set face anti-spoofing (OSFA) approach for unseen attack recognition. Previous domain generalization approaches aim to align multiple domains beyond one common subspace, leading to performance degradation due to the discrepancy of different domains. To address this issue, our approach formulates face anti-spoofing (FAS) in an open set recognition framework, which learns compact representation for each known class in parallel to recognizing unseen attack examples. To this end, we introduce the statistical extreme value theory incorporated in our objective under the multi-task framework. Moreover, we develop an identity-aware contrastive learning method, preventing us from confusion in unseen attack examples versus hard examples. Experimental results on four datasets demonstrate the robustness of our proposed OSFA, especially under diverse categories of unseen attacks. Hao Liu 0019, Pengyuan Lv, Zekuan Yu |
ACM Multimedia | 2 |
| 2021 | Exploiting Invariance of Mining Facial LandmarksabstractIn this paper, we propose an invariant learning method for facial landmark mining in a self-supervised manner. The conventional methods mostly train with raw data of paired facial appearances and landmarks, assuming that they are evenly distributed. However, assumptions like this tend to lead to failures in challenging cases even undergo costly training since they usually don't hold in real-world scenarios. To address this issue, our model achieves to be invariant to facial biases by learning through the landmark-anchored distributions. Specifically, we generate faces from these distributions, then group them based on the appearance sources and the probe facial landmarks into intra-identities and intra-landmarks classes, respectively. Thus, we construct intra-class invariance losses to disentangle the spatial structures from appearances. In addition, we adopt a reconstruction loss to produce more realistic faces with probe landmarks. Extensive experimental results on four standard facial landmark datasets demonstrate that our method achieves compelling performance compared with supervised and unsupervised methods. Jiangming Shi, Zixian Gao, Hao Liu 0019, Zekuan Yu, Fengjun Li |
ACM Multimedia | 3 |
| 2021 | Cross-modality Attention Method for Medical Image Enhancement
Zebin Hu, Hao Liu 0019, Zekuan Yu |
PRCV (3) | 2 |
| 2021 | Non-local Network Routing for Perceptual Image Super-Resolution
Zexin Ji, Zekuan Yu, Hao Liu 0019 |
PRCV (3) | 5 |
| 2021 | Geometry-attentive relational reasoning for robust facial landmark detection
Zongyong Deng, Hao Liu 0019 |
Neurocomputing | 2 |
| 2020 | Towards Omni-Supervised Face Alignment for Large Scale Unlabeled VideosabstractIn this paper, we propose a spatial-temporal relational reasoning networks (STRRN) approach to investigate the problem of omni-supervised face alignment in videos. Unlike existing fully supervised methods which rely on numerous annotations by hand, our learner exploits large scale unlabeled videos plus available labeled data to generate auxiliary plausible training annotations. Motivated by the fact that neighbouring facial landmarks are usually correlated and coherent across consecutive frames, our approach automatically reasons about discriminative spatial-temporal relationships among landmarks for stable face tracking. Specifically, we carefully develop an interpretable and efficient network module, which disentangles facial geometry relationship for every static frame and simultaneously enforces the bi-directional cycle-consistency across adjacent frames, thus allowing the modeling of intrinsic spatial-temporal relations from raw face sequences. Extensive experimental results demonstrate that our approach surpasses the performance of most fully supervised state-of-the-arts. Congcong Zhu, Hao Liu 0019, Zhenhua Yu 0002, Xuehong Sun |
AAAI | 2 |
| 2020 | Learning Neighborhood-Reasoning Label Distribution (NRLD) for Facial Age EstimationabstractIn this paper, we propose to learn a neighborhood-reasoning label distribution (NRLD) for facial age estimation. Unlike conventional label distribution methods with fixed-structural aging patterns, in this work, our NRLD aims to reason about more resilient and adaptive label distribution by disentangling the graph of face neighbors. In particular, our model holds the assumption on that the sample-specific age label distribution is principally influenced by a mixture of interpretable and meaningful factors, which typically cause plausible edges connected to the anchors. Under the scenario of each factor, we specifically collect the subset of graph edges and then convolute them with face samples to regress a mean-variance label distribution. During the training process, the mixture hyperparameters of our label distribution are iteratively optimized by following the Expectation-Maximization schema. Extensive experimental results on three challenging widely-evaluated datasets indicate the superiority in comparisons with most state of the arts. Zongyong Deng, Mo Zhao, Hao Liu 0019, Zhenhua Yu 0002 |
ICME | 3 |
| 2020 | Network Architecture Reasoning Via Deep Deterministic Policy GradientabstractIn this paper, we introduce global compression learning (GCL) for finding reduced network architecture from a pre-trained network by removing both intra-layer and inter-layer structural redundancy. To accomplish this, we first derive architecture features from a binary representation of the network structure that effectively characterize the relationships between different layers. We then leverage reinforcement learning to iteratively compress the network via deep deterministic policy gradient based on the learned architecture features. To void extensive exploration of the huge space of network architectures, we bound feasible solutions within a small subspace by following a strict accuracy loss tolerance. Benchmarking tests show GCL outperforms the state-of-the-art models. On CIFAR-10 dataset, our model reduces 60.5% FLOPs and 93.3% parameters on VGG-16 without hurting the network accuracy, and yields a significantly compressed architecture for ResNet-110 by reductions of 71.92% FLOPs and 79.62% parameters with the cost of only 0.11% accuracy loss. Huidong Liu, Fang Du, Xiaofen Tang, Hao Liu 0019, Zhenhua Yu 0002 |
ICME | 4 |
| 2020 | Learning Reasoning-Decision Networks for Robust Face AlignmentabstractIn this paper, we propose an end-to-end reasoning-decision networks (RDN) approach for robust face alignment via policy gradient. Unlike the conventional coarse-to-fine approaches which likely lead to bias prediction due to poor initialization, our approach aims to learn a policy by leveraging raw pixels to reason a subset of shape candidates, sequentially making plausible decisions to remove outliers for robust initialization. To achieve this, we formulate face alignment as a Markov decision process by defining an agent, which typically interacts with a trajectory of states, actions, state transitions and rewards. The agent seeks an optimal shape searching policy over the whole shape space by maximizing a discounted sum of the received values. To further improve the alignment performance, we develop an LSTM-based value function to evaluate the shape quality. During the training procedure, we adjust the gradient of our value function in directions of the policy gradient. This prevents our training goal from being trapped into local optima entangled by both the pose deformations and appearance variations especially in unconstrained environments. Experimental results show that our proposed RDN consistently outperforms most state-of-the-art approaches on four widely-evaluated challenging datasets. Hao Liu 0019, Jiwen Lu, Suping Wu, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Learning spatial-temporal deformable networks for unconstrained face alignment and tracking in videos
Hao Liu 0019, Congcong Zhu, Zongyong Deng, Xuehong Sun |
Pattern Recognit. | 2 |
| 2020 | Similarity-Aware and Variational Deep Adversarial Learning for Robust Facial Age EstimationabstractIn this paper, we propose a similarity-aware deep adversarial learning (SADAL) approach for facial age estimation. Instead of making full access to the limited training samples which likely leads to bias age prediction, our SADAL aims to seek batches of unobserved hard-negative samples based on existing training samples, which typically reinforces the discriminativeness of the learned feature representation for facial ages. Motivated by the fact that age labels are usually correlated in real-world scenarios, we carefully develop a similarity-aware function to well measure the distance of each face pair based on the age value gaps. Consequently, the age-difference information is exploited in the synthetic feature space for robust age estimation. During the learning process, we jointly optimize both procedures of generating hard negatives and learning discriminative age ranker via a sequence of adversarial-game iterations. Another major issue lies on that existing methods only enforce the indiscriminativeness within each class, which is probably trapped into model overfitting and thus the generation capacity is limited particularly on unseen age classes with many individuals. To circumvent this problem, we propose a variational deep adversarial learning (VDAL) paradigm, which learns to encode each face sample in two factorized parts, i.e., the intra-class variance distribution and the intra-class invariant class center. Moreover, our VDAL principally optimizes the variational confidence lower bound on the variational factorized feature representation. To better enhance the discriminativeness of the age representation, our VDAL further learns to encode the ordinal relationship among age labels in the reconstructed subspace. Experimental results on folds of widely-evaluated benchmarking datasets demonstrate that our approach achieves promising performance in contrast to most state-of-the-art age estimation methods. Hao Liu 0019, Penghui Sun, Suping Wu, Zhenhua Yu 0002, Xuehong Sun |
IEEE Trans. Multim. | 1 |
| 2019 | Learning Deformable Hourglass Networks (DHGN) for Unconstrained Face AlignmentabstractIn this paper, we propose a deformable hourglass networks (DHGN) approach to investigate the problem of face alignment, especially in such challenging cases when faces undergo large variations including severe poses, diverse expressions and partial occlusions in unconstrained environments. Unlike conventional feature extractions which cannot explicitly exploit irregular geometric structures for facial shapes, our DHGN learns a deformable mask to reduce the variances of facial deformation and extract attentional facial regions for robust feature representation. To achieve this, we carefully design a differential module, dubbed the deformable transformer, which typically incorporates with a regression sub-net to predict a set of offsets and a masking operator to filter the semantic facial parts for feature representation learning. To further reinforce the alignment performance, we integrate our designed modules in the paradigm of stacked hourglass networks and jointly optimize the network parameters in an end-to-end manner. Extensive experimental results demonstrate very compelling performance in comparisons to most state-of-the-art methods. Congcong Zhu, Suping Wu, Zhenhua Yu 0002, Xuehong Sun, Hao Liu 0019 |
ICIP | 6 |
| 2019 | Multi-Agent Deep Collaboration Learning for Face Alignment Under Different PerspectivesabstractIn this paper, we propose a multi-agent deep collaboration learning method (MADCL) for simultaneously detecting 2D facial landmarks and 3D facial landmarks projected from 3D to 2D, which aims at distinguishing the ambiguity caused by different perspectives. Above two facial annotations, there are a large number of public semantic areas and some very important private semantic areas. Our single agent captures and memorizes private features for iterations and multiple agents collaborate to learn public features. To achieve this, we design a collaboration learning mechanism to capture, memorize and share semantic information for enhancing the feature representation. Moreover, the input of traditional cascade regression methods is cropped directly from the raw facial image via the shape-indexed manner, which leads that the poor initial shapes likely bring about the predicted results getting worse and worse. We introduce the Markov decision process (MDP) to reason a better position of the initial shape by a reward function that reflects the shape quality. Authentic experimental results indicate that our MADCL consistently outperforms most state-of-the-art methods on two widely-evaluated challenging datasets. Congcong Zhu, Suping Wu, Zhenhua Yu 0002, Hao Liu 0019 |
ICIP | 5 |
| 2019 | Similarity-Aware Deep Adversarial Learning for Facial Age EstimationabstractIn this paper, we propose a similarity-aware deep adversarial learning (SADAL) approach for facial age estimation. Instead of making access to limited training samples which likely leads to sub-optima, our SADAL seeks sets of unobserved and plausible hard-examples based on existing training samples, which typically reinforces the discriminativeness of the learned feature descriptor for ages. Motivated by the fact that age labels are usually correlated in the real-world applications, we carefully develop a similarity-aware function in our approach, which dynamically measures each face pair with different weights based on different age value gaps. During the learning process, we jointly optimize both procedures of generating hard-examples and learning age estimator via a sequence of adversarial-game iterations. As a result, the smoothing aging pattern is exploited in the reconstructed hard-example space for robust age estimation. Experimental results on two standard benchmarking datasets show that our approach achieves superior performance compared with most state-of-the-art age estimation methods. Penghui Sun, Hao Liu 0019, Zhenhua Yu 0002, Suping Wu |
ICME | 2 |
| 2019 | Ordinal Deep Learning for Facial Age EstimationabstractIn this paper, we propose an ordinal deep learning approach for facial age estimation. Unlike conventional hand-crafted feature-based methods that require prior and expert knowledge, we propose an ordinal deep feature learning (ODFL) method to learn feature descriptors for face representation directly from raw pixels. Motivated by the fact that age labels are chronologically correlated and age estimation is an ordinal learning problem, our proposed ODFL enforces two criteria on the descriptors, which are learned at the top of the deep networks: 1) the topology-preserving ordinal relation is employed to exploit the order information in the learned feature space and 2) the age-difference cost information is leveraged to dynamically measure face pairs with different age value gaps. However, both the procedures of feature extraction and age estimation are learned independently in ODFL, which may lead to a sub-optimal problem. To address this, we further propose an end-to-end ordinal deep learning (ODL) framework, where the complementary information of both the procedures is exploited to reinforce our model. Extensive experimental results on five face aging datasets show that both our ODFL and ODL achieve superior performance in comparisons with most state-of-the-art methods. Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Two-Stream Transformer Networks for Video-Based Face AlignmentabstractIn this paper, we propose a two-stream transformer networks (TSTN) approach for video-based face alignment. Unlike conventional image-based face alignment approaches which cannot explicitly model the temporal dependency in videos and motivated by the fact that consistent movements of facial landmarks usually occur across consecutive frames, our TSTN aims to capture the complementary information of both the spatial appearance on still frames and the temporal consistency information across frames. To achieve this, we develop a two-stream architecture, which decomposes the video-based face alignment into spatial and temporal streams accordingly. Specifically, the spatial stream aims to transform the facial image to the landmark positions by preserving the holistic facial shape structure. Accordingly, the temporal stream encodes the video input as active appearance codes, where the temporal consistency information across frames is captured to help shape refinements. Experimental results on the benchmarking video-based face alignment datasets show very competitive performance of our method in comparisons to the state-of-the-arts. Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Label-Sensitive Deep Metric Learning for Facial Age EstimationabstractIn this paper, we present a label-sensitive deep metric learning (LSDML) approach for facial age estimation. Motivated by the fact that human age labels are chronologically correlated, our proposed LSDML aims to seek a series of hierarchical nonlinear transformations by deep residual network to project face samples to a latent common space, where the similarity of face pairs is equivalently isotonic to the age difference in a ranking-preserving manner. Since traversal access to total negative samples catastrophically costs and leads to suboptimal, our model learns to mine hard meaningful samples in parallel to learning feature similarity, so that the local manifold of face samples is preserved in the transformed subspace. To better improve the performance on the data set that contains few labeled samples, we further extend our LSDML to a multi-source LSDML method, which aims at maximizing the cross-population correlation of different face aging data sets. Extensive experimental results on four benchmarking data sets show the effectiveness of our proposed approach. Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | Ordinal Deep Feature Learning for Facial Age EstimationabstractIn this paper, we propose an ordinal deep feature learning (ODFL) approach for facial age estimation. Unlike conventional age estimation methods which utilize hand-crafted features, our ODFL develops deep convolutional neural networks to learn discriminative feature descriptors directly from image pixels for face representation. Motivated by the fact that age labels are chronologically correlated and age estimation is an ordinal learning computer vision problem, we enforce two criterions on the descriptors which are learned at the top of our network: 1) the topology-aware ordinal relation of face samples is preserved in the learned feature space, and 2) the age difference information of the embedded feature representation is exploited in a ranking-preserving manner. Extensive experimental results on four face aging datasets show that our approach achieves promising performance compared with the state-of-the-art methods. Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
FG | 1 |
| 2017 | Group-aware deep feature learning for facial age estimation
Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
Pattern Recognit. | 1 |
| 2017 | Learning Deep Sharable and Structural Detectors for Face AlignmentabstractFace alignment aims at localizing multiple facial landmarks for a given facial image, which usually suffers from large variances of diverse facial expressions, aspect ratios and partial occlusions, especially when face images were captured in wild conditions. Conventional face alignment methods extract local features and then directly concatenate these features for global shape regression. Unlike these methods which cannot explicitly model the correlation of neighbouring landmarks and motivated by the fact that individual landmarks are usually correlated, we propose a deep sharable and structural detectors (DSSD) method for face alignment. To achieve this, we firstly develop a structural feature learning method to explicitly exploit the correlation of neighbouring landmarks, which learns to cover semantic information to disambiguate the neighbouring landmarks. Moreover, our model selectively learns a subset of sharable latent tasks across neighbouring landmarks under the paradigm of the multi-task learning framework, so that the redundancy information of the overlapped patches can be efficiently removed. To better improve the performance, we extend our DSSD to a recurrent DSSD (R-DSSD) architecture by integrating with the complementary information from multi-scale perspectives. Experimental results on the widely used benchmark datasets show that our methods achieve very competitive performance compared to the state-of-the-arts. Hao Liu 0019, Jiwen Lu, Jianjiang Feng, Jie Zhou 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Set-label modeling and deep metric learning on person re-identification
Hao Liu 0019, Bingpeng Ma, Junbiao Pang, Chunjie Zhang 0001, Qingming Huang |
Neurocomputing | 1 |
| 2015 | Dense depth image synthesis via energy minimization for three-dimensional video
You Yang 0002, Qiong Liu 0001, Hao Liu 0019, Li Yu 0003, Fanglin Wang |
Signal Process. | 3 |
| 2013 | Set-based classification for person re-identification utilizing mutual-informationabstractIdentifying individuals in multi-view camera network, known as person re-identification, becomes an emerging topic for video surveillance. In this paper, we address person re-identification as a set-based classification problem and introduce mutual-information to fully utilize gallery information. Firstly, we define a set-based structure that contains pairwise features between query image and gallery images. Then these features are fed into a set-class model, which exploits the relationship between set and class label (person identity) using mutual-information. Finally, we estimate and rank the mutual-information scores, and the corresponding label of the highest score is assigned to the query image. Our method has gained a superior performance compared with the state-of-the-art in the benchmark datasets i-LIDS and ETHZ. Hao Liu 0019, Zhongwei Cheng, Qingming Huang |
ICIP | 1 |