EDBT 2026 Demo / reviewers in the wild / expert
Hai Liu 0004
dblp:46/2375-4
· DBLP profile ↗
53ranked-venue papers
21as first author
40since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 8 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 8 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DRKT: Learning differential relationships for efficient knowledge tracing with learner's knowledge internalization representation
Zhaoli Zhang, Hai Liu 0004, Erqi Zhang, Tingting Liu 0006, Minhong Wang 0001 |
Expert Syst. Appl. | 3 |
| 2026 | CR2Former: Contour-relation aware representation learning for fine-grained bird image classification via Transformers
Hai Liu 0004, Tingting Liu 0006, Xiaolan Yang, Zhaoli Zhang |
Neurocomputing | 1 |
| 2026 | Exploiting structure-semantic consistency for photorealistic SLAM with 3D Gaussian splatting
Qianang Zhou, Hai Liu 0004, Youfu Li 0001, Junlin Xiong |
Neurocomputing | 3 |
| 2026 | RetFER: Exploiting reticular relationship representation via textural architectonics modeling for facial expression recognition
Zhaoli Zhang, Qinglei Han, Liqian Deng, Tingting Liu 0006, Hai Liu 0004, Youfu Li |
Neurocomputing | 6 |
| 2026 | CoSimR: Similarity Cues-Aware Consanguinity Relationship Mining Framework for IoT-Based Bird Monitoring SystemsabstractInternet of Things (IoT)-based bird monitoring system faces challenges from occlusion, arbitrary postures, and similar species. To address these, we exploit two key properties of bird images: self-correlation within individual birds and appearance similarity across species, revealing self-correlation and consanguinity relationships. We propose a similarity cues-aware consanguinity relationship mining (CoSimR) framework, which leverages these relationships for robust classification. CoSimR comprises two modules: Consanguinity Relationship Mining (CRM) and Cross-species Avian Prediction (CAP). CRM can capture skeletal structures by generating the correlation tokens, while consanguinity tokens encode p information across five taxonomic levels (class, order, family, genus, species). CAP leverages a consanguinity-driven multi-loss function, including homogeneity loss for intra-species consistency and affinity loss for cross-species similarity, to guide discriminative feature learning. Experiments conducted on two fine-grained bird image classification datasets demonstrate that the CoSimR model achieves better performance compared with state-of-the-art methods. It highlights the effectiveness of integrating biological hierarchy with visual features, paving the way for similarity cues-aware approaches in fine-grained classification tasks. Tingting Liu 0006, Hai Liu 0004, Zhaoli Zhang, Naixue Xiong |
IEEE Internet Things J. | 4 |
| 2026 | TransSIL: A Silhouette Cue-Aware Image Classification Framework for Bird Ecological Monitoring SystemsabstractHow to automatically recognize the bird species has caused concerns in ecological systems and intelligent ecological monitoring system due to the increasing threat to bird ecological society and bird species diversity. However, traditional monitoring systems are susceptible to specific challenges when operating in complex environments, such as complex environments, multifarious postures and backlight scenarios. To effectively address these challenges, we present a novel bird ecological intelligent detection system (TransSIL) for fine-grained bird image classification (FBIC) in diverse ecological to learn discriminative features by explicitly incorporating silhouette structural information alongside critical visual cues. Specifically, the approach begins with a silhouette token construction module to estimate the bird silhouette and extract silhouette tokens. Then, a silhouette relationship mining module is developed to fuse visual and silhouette tokens and capture long-range dependencies between them. In addition, to learn bird distinctive features at multiple levels, a critical cues awareness module is embedded within TransSIL. The performance of TransSIL was evaluated on two bird datasets: CUB200-2011 and NABirds. The framework demonstrates significant improvements over existing ecological intelligent surveillance methods. By utilizing silhouette and visual dependencies, we anticipate that our approach will ultimately contribute to the conservation of avian ecological societies. Hai Liu 0004, Tingting Liu 0006, Lin Chen 0033, Zhaoli Zhang, Xiaolan Yang, Naixue Xiong |
IEEE Internet Things J. | 1 |
| 2026 | SC2R: similarity cues-aware evolutionary relationship mining for fine-grained bird image classification
Hai Liu 0004, Feifei Li 0003, Zhiyi Du, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
Pattern Recognit. | 1 |
| 2026 | PhyTrans: Learning Phylogenetic Relationships for FBIC via Hierarchical Taxonomy RepresentationabstractHow to accurately identify endangered bird species in complex natural environments has become an important research topic jointly concerned by the computer vision and biological conservation communities. However, they remain limited in systematically modeling cross-species semantic similarity and effectively exploiting structural stability under pose variations, making robust discrimination in highly similar species scenarios difficult. To address these challenges, we propose PhyTrans, a phylogeny-driven fine-grained bird recognition framework that achieves unified representation learning by jointly modeling inter-species phylogenetic relationships and intra-image skeletal invariance across different poses. Specifically, a phylogenetic token construction (PTC) module is designed to leverage hierarchical taxonomic information, ranging from class to species, and embed phylogenetic relationships into a hyperbolic space, which preserves hierarchical semantic distances while explicitly modeling appearance similarity induced by evolutionary relatedness. Building upon this, phylogenetic representations and intra-image skeletal structural cues are further integrated within a unified Transformer architecture through the proposed phylogenetic relationship mining (PRM) module, enabling collaborative modeling of cross-species similarity and structural invariance. Extensive experiments on the CUB-200-2011 and NABirds datasets demonstrate that PhyTrans outperforms state-of-the-art approaches, validating the critical role of phylogenetic relationships in advancing ecological visual recognition. Hai Liu 0004, Tingting Liu 0006, Dazhen Shen, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | MHPE: Learning Morphology Relationships for Robust Head Pose Estimation With Facial Rotation RepresentationabstractAlthough accurate head pose estimation is critical for natural human-computer interaction, it remains challenging due to occlusion, extreme poses, illumination conditions, and data ambiguity issues. To address these challenges, a novel morphology aware Transformer framework (MHPE) is proposed, which can learn morphological relationships during facial rotation. The methodology is based on two key findings: cross-region geometric dependencies and angle-specific morphodynamic representations. The proposed framework incorporates two key components: adversarial feature generation, which generates robust rotation representations by adaptive multi-scale feature interaction; and morphology relationship inference, which establishes long-range dependencies between facial features through a cross-modal attention mechanism that incorporates morphological priors. Extensive evaluations on three demanding benchmarks (BIWI, AFLW2000, and 300W-LP) demonstrate state-of-the-art performance, particularly in demanding scenarios. The Python implementation will be available on request to facilitate reproducibility. Tingting Liu 0006, Jianping Ju, Zhixiong Song, Shijia Qian, Ning Rao, Hai Liu 0004, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | HPCTrans: Heterogeneous Plumage Cues-Aware Texton Correlation Representation for FBIC via TransformersabstractFine-grained bird image classification (FBIC) for distinguishing bird subspecies is challenging because of several issues, including a camouflaged appearance, body occlusion, and an arbitrary bird posture. To address these challenges, we propose a novel heterogeneous plumage cues-aware texton correlation representation for FBIC, which leverages texton correlation in various functional plumage regions for effective learning. Two key findings are revealed: 1) texton structural discrepancies of heterogeneous plumage; and 2) abstract region information for specific birds. On this basis, this model introduces texton coherence extraction module (TCEM) and abstract representation selection (ARS). Specifically, considering bird characteristics, TCEM is introduced to exploit the spatial statistical properties of local textons in heterogeneous plumage. To the best of our knowledge, this study is the first to introduce heterogeneous plumage cues for mining texton correlation relationship representations in FBIC tasks. In addition, a Multiscale Information Cross-Attention Transformer (MICAformer) is proposed for better modeling texton correlation representation. The experimental results on the CUB-200-2011 dataset and NABirds show the effectiveness of the proposed HPCTrans model over the state-of-the-art methods. Hai Liu 0004, Shuang Zeng, Liqian Deng, Tingting Liu 0006, Xionghua Liu, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Texture Affinity Cue-Aware Relationship Representation via Transformers for Facial Expression Recognition in Affective RobotsabstractAutomatic facial expression recognition (FER) from facial videos is a key component in enabling machines to understand human emotional states, which is crucial for affective robots designed to be interactive companions and applied in smart healthcare. However, FER is susceptible to challenges such as occlusion, arbitrary orientations, and illumination, making it difficult to implement precise FER models in robots. To address these issues, we propose a texture affinity cues-aware relationship representation method (FTATrans), which learns to associate facial texture with facial expressions in videos. The research reveals two key findings: 1) interaction of facial textures, and 2) texture affinity effects. On this basis, FTATrans mainly consists of two key networks: semantic-information feature generation (SFG) and texture-affinity relationship mining (TAR). In particular, the semantic relationships between different facial regions can be learned through SFG. TAR is used to capture the texture affinity relationship and integrate them with the overall facial expression information. Additionally, a loss function focused on expression-specific texture variations is proposed to guide the model in learning discriminative expression information. Experiments conducted on five video-based FER datasets demonstrate that the FTATrans model achieves state-of-the-art performance. Hai Liu 0004, Feifei Li 0003, Tingting Liu 0006, Zhaoli Zhang, Naixue Xiong, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2026 | SkeFormer: Skeletal Cues-Aware Bone Point Relationship Learning for Efficient FBIC via TransformersabstractHow to identify endangered bird species in complex outdoor environments has attracted significant attention in the fields of computer vision and machine learning. Previous studies on fine-grained bird image classification (FBIC) face numerous challenges, such as environmental occlusions and arbitrary postures, which limit the accuracy and robustness of existing methods. To address these challenges and enable more reliable bird species identification in extreme outdoor conditions, we propose a novel skeletal cues-aware bone point relationship learning for efficient FBIC via Transformers (SkeFormer). To the best of our knowledge, this is the first time skeletal relationships have been introduced to the FBIC task. Our model introduces three key modules: the skeletal relationship mining (SRM) module, the multilevel feature generation (MFG) module, and the key feature selection (KFS) module. Specifically, in SRM, the model mines the skeletal relationships among different bird species. In MFG, multiscale information is aggregated by connecting features across multiple layers. The KFS module selects key immutable regions of birds based on the learned skeletal relationships. Extensive experiments on two benchmark datasets, CUB-200-2011 and NABirds, show that SkeFormer outperforms existing state-ofthe- art models. The code for SkeFormer will be publicly available. Hai Liu 0004, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Multim. | 1 |
| 2026 | HomLLM: Exploiting Semantic Homology Relationship for Fine-Grained Bird Image Classification via Large Language ModelsabstractHow to recognize endangered bird species in complex outdoor environments has attracted considerable attention in the fields of computer vision and machine learning. However, fine-grained bird image classification (FBIC) is susceptible to problems such as arbitrary postures, interclass discriminability, and occlusions. We propose a novel semantic homology relationship representation learning for fine-grained bird classification with large language models, namely HomLLM, to address these challenges in FBIC effectively. Our proposed model aims to learn homology relationship representations adaptively by identifying invariant structural correspondences between visual features and semantic descriptions, using limited bird data and base class labels. Our approach yields two key findings: 1) invariant homology in key regions of birds that maintain structural consistency across different postures and 2) homological relationship that establish essential taxonomic markers among similar bird classes. Based on these insights, we propose two new modules of the model: the semantic homology generation (SHG) module and homology relationship mining (HRM) module. Specifically, in SHG, bird features are described at multiple granularities through a large language model (LLM) to establish semantic homology. In HRM, feature adaptation is performed separately for textual and visual information, and cross-modal homological interaction is performed hierarchically. In addition, we propose a hierarchical homology interaction scheme to integrate multilevel homological features while preserving structural consistency. Experiments on the commonly used bird datasets CUB-200-2011 and NABirds demonstrate that HomLLM exhibits better performance than state-of-the-art (SOTA) methods. Hai Liu 0004, Tingting Liu 0006, Lin Chen 0033, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | MDT: A multiscale differencing transformer with sequence feature relationship mining for robust action recognition
Zengzhao Chen, Fumei Ma, Hai Liu 0004, Tingting Liu 0006 |
Appl. Intell. | 3 |
| 2025 | KRMNet: learning core representations for partial discharge pattern recognition via masked autoencoders and mixed position coding
Quan Xie, Hai Liu 0004 |
Appl. Intell. | 5 |
| 2025 | Pixel-Level Semantics Boosted Fine-Grained Bird Image Classification
Yongjian Deng, Bochen Xie, Hai Liu 0004, Youfu Li 0001, Zhen Yang 0004 |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Neuromorphic event-based recognition boosted by motion-aware learning
Yuhan Liu 0021, Yongjian Deng, Bochen Xie, Hai Liu 0004, Zhen Yang 0004, Youfu Li 0001 |
Neurocomputing | 4 |
| 2025 | HRHPE: FRoI guides heterogeneous relationship representation learning for precise head pose estimation
Hai Liu 0004, Shijia Qian, Tingting Liu 0006, Zelin Cao, Minhong Wang 0001, Jianping Ju, Zhaoli Zhang |
Neurocomputing | 1 |
| 2025 | ESERNet: Learning spectrogram structure relationship for effective speech emotion recognition with swin transformer in classroom discourse analysis
Tingting Liu 0006, Minghong Wang, Hai Liu 0004, Shaoxin Yi |
Neurocomputing | 4 |
| 2025 | Exploiting module evolution correlation relationship for fine-grained bird image classification with structural functional representation
Shuang Zeng, Hai Liu 0004, Tingting Liu 0006, Qiuxia Liu, Minhong Wang 0001, Zhaoli Zhang |
Neurocomputing | 2 |
| 2025 | MHRN: A multi-perspective hierarchical relation network for knowledge graph embedding
Zengcan Xue, Zhaoli Zhang, Hai Liu 0004, Zhifei Li 0009, Shuyun Han, Erqi Zhang |
Knowl. Based Syst. | 3 |
| 2025 | TransIFC: Invariant Cues-Aware Feature Concentration Learning for Efficient Fine-Grained Bird Image ClassificationabstractFine-grained bird image classification (FBIC) is not only meaningful for endangered bird observation and protection but also a prevalent task for image classification in multimedia processing and computer vision. However, FBIC suffers from several challenges, such as bird molting, complex background, and arbitrary bird posture. To effectively tackle these challenges, we present a novel invariant cues-aware feature concentration Transformer (TransIFC), which learns invariant and core information in bird images. To this end, two novel modules are proposed to leverage the characteristics of bird images, namely, the hierarchy stage feature aggregation (HSFA) module and the feature in feature abstraction (FFA) module. The HSFA module aggregates the multiscale information of bird images by concatenating multilayer features. The FFA module extracts the invariant cues of birds through feature selection based on discrimination scores. Transformer is employed as the backbone to reveal the long-dependent semantic relationships in bird images. Moreover, abundant visualizations are provided to prove the interpretability of the HSFA and FFA modules in TransIFC. Comprehensive experiments demonstrate that TransIFC can achieve state-of-the-art performance on the CUB-200-2011 dataset (91.0%) and the NABirds dataset (90.9%). Finally, extended experiments have been conducted on the Stanford Cars dataset to suggest the potential of generalizing our method on other fine-grained visual classification tasks. Hai Liu 0004, Cheng Zhang 0020, Yongjian Deng, Bochen Xie, Tingting Liu 0006, Youfu Li 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | LDCNet: Limb Direction Cues-Aware Network for Flexible HPE in Industrial Behavioral Biometrics Systemsabstract2D Human pose estimation (HPE) has been widely used in the many fields such as behavioral understanding, identity authentication, and industrial automatic manufacturing. Most of the previous studies have encountered many constraints, such as restricted scenarios and strict inputs. To solve this problem, we present a simple yet effective HPE network called limb direction cues-aware network (LDCNet) with limb direction cues and differentiated Cauchy labels, which can efficiently suppress uncertainties and prevent deep networks from over-fitting uncertain keypoint positions. In particular, LDCNet suppresses the uncertainties from two aspects. (1) A differentiated Cauchy coordinate encoding method is designed to reveal the limb direction information among adjacent keypoints. (2) Jeffreys divergence is introduced as loss function to measure the prediction heatmap and ground-truth one. Positions of keypoints are perceived at the limb direction based deep network in an end-to-end manner. An extensive study on two benchmark data sets (i.e., MS COCO and MPII) illustrates the superiority of the proposed LDCNet model over state-of-the-art approaches. Tingting Liu 0006, Hai Liu 0004, Zhaoli Zhang |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | MMATrans: Muscle Movement Aware Representation Learning for Facial Expression Recognition via TransformersabstractHow to automatically recognize facial expression has caused concerns in industrial human–robot interaction. However, facial expression recognition (FER) is susceptible to problems, such as occlusion, arbitrary orientations, and illumination. To effectively address these challenges in FER, we present a novel facial muscle movement aware representation learning that can learn the semantic relationships of facial muscle movements in facial expression images. Two key findings are revealed: 1) muscle movements from different facial regions often show semantic relationships; and 2) not all facial muscle regions have equal contributions for different facial expressions. On this basis, this model presents two novel modules, namely, discriminative feature generation (DFG) and muscle relationship mining (MRM). Specifically, in DFG, the memory of our model for mislabeling decreases. In MRM, muscle–motion interaction among diverse facial regions is learned through visual transformers (MMATrans). Experiments on three in-the-wild FER datasets (RAF-DB, FERPlus, and AffectNet) show that our MMATrans yields better performance compared with state-of-the-art methods. Hai Liu 0004, Qiyun Zhou, Cheng Zhang 0020, Junyan Zhu, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | EHPE: Skeleton Cues-Based Gaussian Coordinate Encoding for Efficient Human Pose EstimationabstractHuman pose estimation (HPE) has many wide applications such as multimedia processing, behavior understanding and human-computer interaction. Most previous studies have encountered many constraints, such as restricted scenarios and RGB inputs. To mitigate constraints to estimating the human poses in general scenarios, we present an efficient human pose estimation model (i.e., EHPE) with joint direction cues and Gaussian coordinate encoding. Specifically, we propose an anisotropic Gaussian coordinate coding method to describe the skeleton direction cues among adjacent keypoints. To the best of our knowledge, this is the first time that the skeleton direction cues is introduced to the heatmap encoding in HPE task. Then, a multi-loss function is proposed to constrain the output to prevent the overfitting. The Kullback-Leibler divergence is introduced to measure the predication label and its ground truth one. The performance of EHPE is evaluated on two HPE datasets: MS COCO and MPII. Experimental results demonstrate that EHPE can obtain robust results, and it significantly outperforms existing state-of-the-art HPE methods. Lastly, we extend the experiments on infrared images captured by our research group. The experiments achieved the impressive results regardless of insufficient color and texture information. Hai Liu 0004, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | TokenHPE: Learning Orientation Tokens for Efficient Head Pose Estimation via TransformersabstractHead pose estimation (HPE) has been widely used in the fields of human machine interaction, self-driving, and attention estimation. However, existing methods cannot deal with extreme head pose randomness and serious occlusions. To address these challenges, we identify three cues from head images, namely, neighborhood similarities, significant facial changes, and critical minority relationships. To leverage the observed findings, we propose a novel critical minority relationship-aware method based on the Transformer architecture in which the facial part relationships can be learned. Specifically, we design several orientation tokens to explicitly encode the basic orientation regions. Meanwhile, a novel token guide multiloss function is designed to guide the orientation tokens as they learn the desired regional similarities and relationships. We evaluate the proposed method on three challenging benchmark HPE datasets. Experiments show that our method achieves better performance compared with state-of-the-art methods. Our code is publicly available at https://github.com/zc2023/TokenHPE. Cheng Zhang 0020, Hai Liu 0004, Yongjian Deng, Bochen Xie, Youfu Li 0001 |
CVPR | 2 |
| 2023 | Learning knowledge graph embedding with multi-granularity relational augmentation network
Zengcan Xue, Zhaoli Zhang, Hai Liu 0004, Shuoqiu Yang, Shuyun Han |
Expert Syst. Appl. | 3 |
| 2023 | MSRANet: Learning discriminative embeddings for speaker verification via channel and spatial attention mechanism in alterable scenarios
Qiuyu Zheng, Zengzhao Chen, Hai Liu 0004, Tingting Liu 0006 |
Expert Syst. Appl. | 3 |
| 2023 | Orientation Cues-Aware Facial Relationship Representation for Head Pose Estimation via TransformerabstractHead pose estimation (HPE) is an indispensable upstream task in the fields of human-machine interaction, self-driving, and attention detection. However, practical head pose applications suffer from several challenges, such as severe occlusion, low illumination, and extreme orientations. To address these challenges, we identify three cues from head images, namely, critical minority relationships, neighborhood orientation relationships, and significant facial changes. On the basis of the three cues, two key insights on head poses are revealed: 1) intra-orientation relationship and 2) cross-orientation relationship. To leverage two key insights above, a novel relationship-driven method is proposed based on the Transformer architecture, in which facial and orientation relationships can be learned. Specifically, we design several orientation tokens to explicitly encode basic orientation regions. Besides, a novel token guide multi-loss function is accordingly designed to guide the orientation tokens as they learn the desired regional similarities and relationships. Experimental results on three challenging benchmark HPE datasets show that our proposed TokenHPE achieves state-of-the-art performance. Moreover, qualitative visualizations are provided to verify the effectiveness of the token-learning methodology. Hai Liu 0004, Cheng Zhang 0020, Yongjian Deng, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | A Voxel Graph CNN for Object Classification with Event CamerasabstractEvent cameras attract researchers' attention due to their low power consumption, high dynamic range, and extremely high temporal resolution. Learning models on event-based object classification have recently achieved massive success by accumulating sparse events into dense frames to apply traditional 2D learning methods. Yet, these approaches necessitate heavy-weight models and are with high computational complexity due to the redundant information introduced by the sparse-to-dense conversion, limiting the potential of event cameras on real-life applications. This study aims to address the core problem of balancing accuracy and model complexity for event-based classification models. To this end, we introduce a novel graph representation for event data to exploit their sparsity better and customize a lightweight voxel graph convolutional neural network (EV-VGCNN) for event-based classification. Specifically, (1) using voxel-wise vertices rather than previous point-wise inputs to explicitly exploit regional 2D semantics of event streams while keeping the sparsity; (2) proposing a multi-scale feature relational layer (MFRL) to extract spatial and motion cues from each vertex discriminatively concerning its distances to neighbors. Comprehensive experiments show that our model can advance state-of-the-art classification accuracy with extremely low model complexity (merely 0.84M parameters). Yongjian Deng, Hao Chen 0034, Hai Liu 0004, Youfu Li 0001 |
CVPR | 3 |
| 2022 | Multi-perspective social recommendation method with graph representation learning
Hai Liu 0004, Duantengchuan Li, Zhaoli Zhang, Ke Lin 0001, Xiaoxuan Shen, Naixue Xiong, Jiazhang Wang |
Neurocomputing | 1 |
| 2022 | ARHPE: Asymmetric Relation-Aware Representation Learning for Head Pose Estimation in Industrial Human-Computer InteractionabstractHead pose estimation (HPE) has wide industrial applications, such as online education, human–robot interaction, and automatic manufacturing. In this article, we address two key problems in HPE based on label learning and asymmetric relation cues: 1) how to bridge the gap between the better prediction performance of networks and incorrectly label pose images in the HPE datasets and 2) how to take full advantage of the adjacent poses information around the centered pose image. We reconstruct all the incorrect labels as a two-dimensional Lorentz distribution to tackle the first problem. Instead of directly adopting the angle values ashardlabels, we assign part of the probability values (softlabels) to adjacent labels for learning discriminative feature representations. To address the second problem, we reveal the asymmetric relation nature of HPE datasets. The yaw direction and pitch direction are assigned different weights by introducing the half at half-maximum of the Lorentz distribution. Compared with the traditional end-to-end frameworks, the proposed one can leverage the asymmetric relation cues for predicting the head pose angle in the incorrect label scenarios. Extensive experiments on two public datasets and our infrared dataset demonstrate that the proposed ARHPE network significantly outperforms other state-of-the-art approaches. Hai Liu 0004, Tingting Liu 0006, Zhaoli Zhang, Arun Kumar Sangaiah, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | EDMF: Efficient Deep Matrix Factorization With Review Feature Learning for Industrial Recommender SystemabstractRecommendation accuracy is a fundamental problem in the quality of the recommendation system. In this article, we propose an efficient deep matrix factorization (EDMF) with review feature learning for the industrial recommender system. Two characteristics in user’s review are revealed. First, interactivity between the user and the item, which can also be considered as the former’s scoring behavior on the latter, is exploited in a review. Second, the review is only a partial description of the user’s preferences for the item, which is revealed as the sparsity property. Specifically, in the first characteristic, EDMF extracts the interactive features of onefold review by convolutional neural networks with word-attention mechanism. Subsequently,${L}_{0}$norm is leveraged to constrain the review considering that the review information is a sparse feature, which is the second characteristic. Furthermore, the loss function is constructed by maximuma posterioriestimation theory, where the interactivity and sparsity property are converted as two prior probability functions. Finally, the alternative minimization algorithm is introduced to optimize the loss functions. Experimental results on several datasets demonstrate that the proposed methods, which show good industrial conversion application prospects, outperform the state-of-the-art methods in terms of effectiveness and efficiency. Hai Liu 0004, Duantengchuan Li, Xiaoxuan Shen, Ke Lin 0001, Jiazhang Wang, Zhaoli Zhang, Naixue Xiong |
IEEE Trans. Ind. Informatics | 1 |
| 2022 | Multi-Scale Dynamic Convolutional Network for Knowledge Graph EmbeddingabstractKnowledge graphs are large graph-structured knowledge bases with incomplete or partial information. Numerous studies have focused on knowledge graph embedding to identify the embedded representation of entities and relations, thereby predicting missing relations between entities. Previous embedding models primarily regard (subject entity, relation, and object entity) triplet as translational distance or semantic matching in vector space. However, these models only learn a few expressive features and hard to handle complex relations, i.e., 1-to-N, N-to-1, and N-to-N, in knowledge graphs. To overcome these issues, we introduce a multi-scale dynamic convolutional network (M-DCN) model for knowledge graph embedding. This model features topnotch performance and an ability to generate richer and more expressive feature embeddings than its counterparts. The subject entity and relation embeddings in M-DCN are composed in an alternating pattern in the input layer, which helps extract additional feature interactions and increase the expressiveness. Multi-scale filters are generated in the convolution layer to learn different characteristics among input embeddings. Specifically, the weights of these filters are dynamically related to each relation to model complex relations. The performance of M-DCN on the five benchmark datasets is tested via experiments. Results show that the model can effectively handle complex relations and achieve state-of-the-art link prediction results on most evaluation metrics. Zhaoli Zhang, Zhifei Li 0009, Hai Liu 0004, Naixue Xiong |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | MFDNet: Collaborative Poses Perception and Matrix Fisher Distribution for Head Pose EstimationabstractHead pose estimation suffers from several problems, including low pose tolerance under different disturbances and ambiguity arising from common head pose representation. In this study, a robust three-branch model with triplet module and matrix Fisher distribution module is proposed to address these problems. Based on metric learning, the triplet module employs triplet architecture and triplet loss. It is implemented to maximize the distance between embeddings with different pose pairs and minimize the distance between embeddings with same pose pairs. It can learn a highly discriminate and robust embedding related to head pose. Moreover, the rotation matrix instead of Euler angle and unit quaternion is utilized to represent head pose. An exponential probability density model based on the rotation matrix (referred to as the matrix Fisher distribution) is developed to model head rotation uncertainty. The matrix Fisher distribution can further analyze the head pose, and its maximum likelihood obtained using singular value decomposition provides enhanced accuracy. Extensive experiments executed over AFLW2000 and BIWI datasets demonstrate that the proposed model achieves state-of-the-art performance in comparison with traditional methods. Hai Liu 0004, Shuai Fang, Zhaoli Zhang, Duantengchuan Li, Ke Lin 0001, Jiazhang Wang |
IEEE Trans. Multim. | 1 |
| 2022 | Learning Knowledge Graph Embedding With Heterogeneous Relation Attention NetworksabstractKnowledge graph (KG) embedding aims to study the embedding representation to retain the inherent structure of KGs. Graph neural networks (GNNs), as an effective graph representation technique, have shown impressive performance in learning graph embedding. However, KGs have an intrinsic property of heterogeneity, which contains various types of entities and relations. How to address complex graph data and aggregate multiple types of semantic information simultaneously is a critical issue. In this article, a novel heterogeneous GNNs framework based on attention mechanism is proposed. Specifically, the neighbor features of an entity are first aggregated under each relation-path. Then the importance of different relation-paths is learned through the relation features. Finally, each relation-path-based features with the learned weight values are aggregated to generate the embedding representation. Thus, the proposed method not only aggregates entity features from different semantic aspects but also allocates appropriate weights to them. This method can capture various types of semantic information and selectively aggregate informative features. The experiment results on three real-world KGs demonstrate superior performance when compared with several state-of-the-art methods. Zhifei Li 0009, Hai Liu 0004, Zhaoli Zhang, Tingting Liu 0006, Naixue Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | CARM: Confidence-aware recommender model via review representation learning and historical rating behavior in the online platforms
Duantengchuan Li, Hai Liu 0004, Zhaoli Zhang, Ke Lin 0001, Shuai Fang, Zhifei Li 0009, Naixue Xiong |
Neurocomputing | 2 |
| 2021 | Recalibration convolutional networks for learning interaction knowledge graph embedding
Zhifei Li 0009, Hai Liu 0004, Zhaoli Zhang, Tingting Liu 0006, Jiangbo Shu |
Neurocomputing | 2 |
| 2021 | Anisotropic angle distribution learning for head pose estimation and attention understanding in human-computer interaction
Hai Liu 0004, Hanwen Nie, Zhaoli Zhang, Youfu Li 0001 |
Neurocomputing | 1 |
| 2021 | Deep Variational Matrix Factorization with Knowledge Embedding for Recommendation SystemabstractAutomatic recommendation has become an increasingly relevant problem to industries, which allows users to discover new items that match their tastes and enables the system to target items to the right users. In this article, we have proposed a deep learning based fully Bayesian treatment recommendation framework, DVMF, which has high-quality performance and ability to integrate any kinds of side information handily and efficiently. In DVMF, the variational inference technique and the reparameterization tricks are introduced to make DVMF possible to be optimized by the stochastic gradient-based methods, in addition, two novel deep neural networks have been constructed to infer the hyper-parameters of the distributions of latent factors from the knowledge of user and item, which are represented as low-dimensional real-valued vectors retaining primary features. Experimental results on five public databases indicate that the proposed method performs better than the state-of-the-art recommendation algorithms on prediction accuracy in terms of quantitative assessments. Xiaoxuan Shen, Baolin Yi, Hai Liu 0004, Wei Zhang 0139, Zhaoli Zhang, Sannyuya Liu, Naixue Xiong |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | 3D Gaze Estimation for Head-Mounted Devices based on Visual SaliencyabstractCompared with the maturity of 2D gaze tracking technology, 3D gaze tracking has gradually become a research hotspot in recent years. The head-mounted gaze tracker has shown great potential for gaze estimation in 3D space due to its appealing flexibility and portability. The general challenge for 3D gaze tracking algorithms is that calibration is necessary before the usage, and calibration targets cannot be easily applied in some situations or might be blocked by moving human and objects. Besides, the accuracy on depth direction has always come to be a crucial problem. Regarding the issues mentioned above, a 3D gaze estimation with auto-calibration method is proposed in this study. We use an RGBD camera as the scene camera to acquire the accurate 3D structure of the environment. The automatic calibration is achieved by uniting gaze vectors with saliency maps of the scene which aligned depth information. Finally, we determine the 3D gaze point through a point cloud generated from the RGBD camera. The experiment result demonstrates that our proposed method achieves 4.34° of average angle error in the field from 0.5m to 3m and the average depth error is 23.22mm, which is sufficient for 3D gaze estimation in the real scene. Meng Liu 0021, Youfu Li 0001, Hai Liu 0004 |
IROS | 3 |
| 2020 | Infrared head pose estimation with multi-scales feature fusion on the IRHP database for human attention recognition
Hai Liu 0004, Xiang Wang 0024, Wei Zhang 0139, Zhaoli Zhang, Youfu Li 0001 |
Neurocomputing | 1 |
| 2020 | Infrared facial expression recognition via Gaussian-based label distribution learning in the dark illumination environment for human emotion detection
Zhaoli Zhang, Chenghang Lai, Hai Liu 0004, Youfu Li 0001 |
Neurocomputing | 3 |
| 2020 | Flexible FTIR Spectral Imaging Enhancement for Industrial Robot Infrared Vision SensingabstractInfrared (IR) spectral imaging sensing is a powerful visual technique for industrial material recognition in robot vision systems. However, the imaging sensing data have issues of random noise and band overlap. Resolution enhancement is usually the first step in the preprocessing procedure of industrial robot vision sensing. In this article, we develop a resolution-enhancement algorithm with total variation (TV) constraints for the degraded Fourier transform IR (FTIR) spectrum due to overlap and noise degradation in the robot vision sensing. The kernel function is calculated using the spectrometer imaging systems and Fourier optical theory. The proposed model not only can remove noises effectively but also can estimate the kernel function because of the adaptive TV as constraint regularization. This model is examined by a set of simulated FTIR spectra with the Poisson noises and a series of real FTIR spectra. The proposed model is compared with the other state-of-the-art methods in terms of performance. Experimental results demonstrate that the proposed approach can split the overlap band effectively while the spectral structure details are retained satisfactorily. The enhanced high-resolution imaging spectrum data can raise the robot vision sensing accuracy in industrial intelligent systems. Tingting Liu 0006, Hai Liu 0004, Youfu Li 0001, Zengzhao Chen, Zhaoli Zhang, Sannyuya Liu |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | DISR: Deep Infrared Spectral Restoration Algorithm for Robot Sensing and Intelligent Visual Tracking SystemsabstractInfrared imaging spectrometer (IRIS) often suffers from overlapped bands and random noises, which limit the precision of subsequent processing in robot vision sensing. To address this problem, we propose a novel Gabor transform-based infrared spectrum restoration method by successfully exploring the intrinsic structure of the clean IR spectrum from the degraded one. At first, a total variation (TV) regularized Gabor coefficients adjustment descriptor is designed and incorporated into the spectrum restoration model. Then, the proposed model is inferred via an efficient optimization approach based on split Bregman iteration method. Comprehensive experiments illustrate the significant and consistent improvements of the developed model over state-of-the-art approaches. The restored high-resolution spectrum can be utilized for detecting the different materials in the robot visual tracking systems. Hai Liu 0004, Youfu Li 0001, Dan Su 0001, Zhaoli Zhang, Sannyuya Liu, Tingting Liu 0006 |
IROS | 1 |
| 2019 | Deep Matrix Factorization With Implicit Feedback Embedding for Recommendation SystemabstractAutomatic recommendation has become an increasingly relevant problem to industries, which allows users to discover new items that match their tastes and enables the system to target items to the right users. In this paper, we propose a deep learning (DL) based collaborative filtering framework, namely, deep matrix factorization (DMF), which can integrate any kind of side information effectively and handily. In DMF, two feature transforming functions are built to directly generate latent factors of users and items from various input information. As for the implicit feedback that is commonly used as input of recommendation algorithms, implicit feedback embedding (IFE) is proposed. IFE converts the high-dimensional and sparse implicit feedback information into a low-dimensional real-valued vector retaining primary features. Using IFE could reduce the scale of model parameters conspicuously and increase model training efficiency. Experimental results on five public databases indicate that the proposed method performs better than the state-of-the-art DL-based recommendation algorithms on both accuracy and training efficiency in terms of quantitative assessments. Baolin Yi, Xiaoxuan Shen, Hai Liu 0004, Zhaoli Zhang, Wei Zhang 0139, Sannyuya Liu, Naixue Xiong |
IEEE Trans. Ind. Informatics | 3 |
| 2018 | A content-based recommendation algorithm for learning resources
Jiangbo Shu, Xiaoxuan Shen, Hai Liu 0004, Baolin Yi, Zhaoli Zhang |
Multim. Syst. | 3 |
| 2018 | Fast Blind Instrument Function Estimation Method for Industrial Infrared SpectrometersabstractInfrared (IR) spectrometers, particularly the aging ones, often suffer from the band overlap and random noise. In this paper, a blind estimation method based on discrete cosine transform (DCT) regularization is proposed for IR spectrum measured from an aging spectrometer instrument. Motivated by the observation that the DCT coefficient distribution of the ground-truth spectrum is sparser than that of the observed spectrum, an IR spectral deconvolution model is formulated in our method to regularize the distribution of the observed spectrum by total variation regularization. Then, the split Bregman method is exploited to solve the resulting optimization problem. The experimental results demonstrate an encouraging performance of the proposed approach to suppress noise and preserve spectral details. The novelty of our method lies on its ability to estimate instrument function and latent spectrum in a joint framework; thus, mitigating the effects of instrument aging to a large extent. The recovered IR spectra can efficiently capture the spectral features and interpret the unknown chemical mixture in industrial applications. Tingting Liu 0006, Hai Liu 0004, Zengzhao Chen, Alan M. Lesgold |
IEEE Trans. Ind. Informatics | 2 |
| 2015 | Destriping algorithm with L0 sparsity prior for remote sensing imagesabstractRemote sensing image often suffers from the common problems of stripe noise and random noise. In this paper, we present a destriping method with unidirectional gradient L0 norm and L0 sparsity priori. The major novelty of the proposed method is that combining the unidirectional gradient L0 norm with the sparsity priori to address the destriping and denoising issues. Moreover, doubly augmented Lagrangian (DAL) method is adopted to solve the L0 regularized minimization problem. The proposed method is verified on heavily striped remote sensing images. Comparative results demonstrate that the proposed method outperforms the-state-of-art methods, which can suppress noise effectively as well as preserve image structures well. Hai Liu 0004, Zhaoli Zhang, Sanya Liu, Tingting Liu 0006, Yi Chang 0002 |
ICIP | 1 |
| 2014 | Simultaneous Destriping and Denoising for Remote Sensing Images With Unidirectional Total Variation and Sparse RepresentationabstractRemote sensing images destriping and denoising are both classical problems, which have attracted major research efforts separately. This letter shows that the two problems can be successfully solved together within a unified variational framework. To do this, we proposed a joint destriping and denoising method by integrating the unidirectional total variation and sparse representation regularizations. Experimental results on simulated and real data in terms of qualitative and quantitative assessments show significant improvements over conventional methods. Yi Chang 0002, Luxin Yan, Houzhang Fang, Hai Liu 0004 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2014 | A Gravity Gradient Differential Ratio Method for Underwater Object DetectionabstractThe smooth operation of autonomous underwater vehicles (AUVs) relies heavily on the accurate detection of surrounding objects. Toward this end, this letter presents a novel method for underwater object detection based on the gravity gradient differential and the gravity gradient differential ratio caused by the relative motion between the AUV and the object. Unlike the existing techniques, the proposed method works in a passive manner and achieves AUV invisibility without energy emission. In addition, for the proposed method, no gravity map or gravity gradient map is required, which improves its practicality. Experimental results demonstrate that the proposed method performs better than the existing methods. Zu Yan, Jie Ma 0003, Jinwen Tian, Hai Liu 0004, Jingang Yu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2013 | Joint blind deblurring and destriping for remote sensing imagesabstractDeblurring and destriping are both classical problems for remote sensing images, which are known to be difficult. Treating deblurring and destriping separately, such a straightforward approach, however, suffers greatly from the defective output. This paper shows that the two problems can be successfully solved together and benefit greatly from each other within a unified variational framework. To do this, we propose a joint deblurring and destriping method by combining the framelet regularization and unidirectional total variation. Extensive experiments on simulation and real remote sensing images are carried out and the results of our joint model show significant improvement over conventional methods of treating the two tasks separately. Yi Chang 0002, Houzhang Fang, Luxin Yan, Hai Liu 0004 |
ICIP | 4 |
| 2012 | Atmospheric-Turbulence-Degraded Astronomical Image Restoration by Minimizing Second-Order Central MomentabstractAtmospheric turbulence affects imaging systems by virtue of wave propagation through a medium with a nonuniform index of refraction. It can lead to blurring in images acquired from a long distance away. In this letter, it is observed that blurring increases the second-order central moment (SOCM) of images, and we introduce a new parametric blur identification method by minimizing SOCM. The method applies to finite-support images, in which the scene consists of a finite-extent object against a uniformly black, gray, or white background. The SOCM method has been validated by direct comparisons with other methods on simulated and real degraded images. Luxin Yan, Mingzhi Jin, Houzhang Fang, Hai Liu 0004, Tianxu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |