EDBT 2026 Demo / reviewers in the wild / expert
Tingting Liu 0006
dblp:25/3327-6
· DBLP profile ↗
32ranked-venue papers
8as first author
27since 2021 · last 2026
0000-0002-9347-5974ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DRKT: Learning differential relationships for efficient knowledge tracing with learner's knowledge internalization representation
Zhaoli Zhang, Hai Liu 0004, Erqi Zhang, Tingting Liu 0006, Minhong Wang 0001 |
Expert Syst. Appl. | 5 |
| 2026 | CR2Former: Contour-relation aware representation learning for fine-grained bird image classification via Transformers
Hai Liu 0004, Tingting Liu 0006, Xiaolan Yang, Zhaoli Zhang |
Neurocomputing | 3 |
| 2026 | RetFER: Exploiting reticular relationship representation via textural architectonics modeling for facial expression recognition
Zhaoli Zhang, Qinglei Han, Liqian Deng, Tingting Liu 0006, Hai Liu 0004, Youfu Li |
Neurocomputing | 5 |
| 2026 | CoSimR: Similarity Cues-Aware Consanguinity Relationship Mining Framework for IoT-Based Bird Monitoring SystemsabstractInternet of Things (IoT)-based bird monitoring system faces challenges from occlusion, arbitrary postures, and similar species. To address these, we exploit two key properties of bird images: self-correlation within individual birds and appearance similarity across species, revealing self-correlation and consanguinity relationships. We propose a similarity cues-aware consanguinity relationship mining (CoSimR) framework, which leverages these relationships for robust classification. CoSimR comprises two modules: Consanguinity Relationship Mining (CRM) and Cross-species Avian Prediction (CAP). CRM can capture skeletal structures by generating the correlation tokens, while consanguinity tokens encode p information across five taxonomic levels (class, order, family, genus, species). CAP leverages a consanguinity-driven multi-loss function, including homogeneity loss for intra-species consistency and affinity loss for cross-species similarity, to guide discriminative feature learning. Experiments conducted on two fine-grained bird image classification datasets demonstrate that the CoSimR model achieves better performance compared with state-of-the-art methods. It highlights the effectiveness of integrating biological hierarchy with visual features, paving the way for similarity cues-aware approaches in fine-grained classification tasks. Tingting Liu 0006, Hai Liu 0004, Zhaoli Zhang, Naixue Xiong |
IEEE Internet Things J. | 1 |
| 2026 | TransSIL: A Silhouette Cue-Aware Image Classification Framework for Bird Ecological Monitoring SystemsabstractHow to automatically recognize the bird species has caused concerns in ecological systems and intelligent ecological monitoring system due to the increasing threat to bird ecological society and bird species diversity. However, traditional monitoring systems are susceptible to specific challenges when operating in complex environments, such as complex environments, multifarious postures and backlight scenarios. To effectively address these challenges, we present a novel bird ecological intelligent detection system (TransSIL) for fine-grained bird image classification (FBIC) in diverse ecological to learn discriminative features by explicitly incorporating silhouette structural information alongside critical visual cues. Specifically, the approach begins with a silhouette token construction module to estimate the bird silhouette and extract silhouette tokens. Then, a silhouette relationship mining module is developed to fuse visual and silhouette tokens and capture long-range dependencies between them. In addition, to learn bird distinctive features at multiple levels, a critical cues awareness module is embedded within TransSIL. The performance of TransSIL was evaluated on two bird datasets: CUB200-2011 and NABirds. The framework demonstrates significant improvements over existing ecological intelligent surveillance methods. By utilizing silhouette and visual dependencies, we anticipate that our approach will ultimately contribute to the conservation of avian ecological societies. Hai Liu 0004, Tingting Liu 0006, Lin Chen 0033, Zhaoli Zhang, Xiaolan Yang, Naixue Xiong |
IEEE Internet Things J. | 3 |
| 2026 | SC2R: similarity cues-aware evolutionary relationship mining for fine-grained bird image classification
Hai Liu 0004, Feifei Li 0003, Zhiyi Du, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
Pattern Recognit. | 6 |
| 2026 | PhyTrans: Learning Phylogenetic Relationships for FBIC via Hierarchical Taxonomy RepresentationabstractHow to accurately identify endangered bird species in complex natural environments has become an important research topic jointly concerned by the computer vision and biological conservation communities. However, they remain limited in systematically modeling cross-species semantic similarity and effectively exploiting structural stability under pose variations, making robust discrimination in highly similar species scenarios difficult. To address these challenges, we propose PhyTrans, a phylogeny-driven fine-grained bird recognition framework that achieves unified representation learning by jointly modeling inter-species phylogenetic relationships and intra-image skeletal invariance across different poses. Specifically, a phylogenetic token construction (PTC) module is designed to leverage hierarchical taxonomic information, ranging from class to species, and embed phylogenetic relationships into a hyperbolic space, which preserves hierarchical semantic distances while explicitly modeling appearance similarity induced by evolutionary relatedness. Building upon this, phylogenetic representations and intra-image skeletal structural cues are further integrated within a unified Transformer architecture through the proposed phylogenetic relationship mining (PRM) module, enabling collaborative modeling of cross-species similarity and structural invariance. Extensive experiments on the CUB-200-2011 and NABirds datasets demonstrate that PhyTrans outperforms state-of-the-art approaches, validating the critical role of phylogenetic relationships in advancing ecological visual recognition. Hai Liu 0004, Tingting Liu 0006, Dazhen Shen, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | MHPE: Learning Morphology Relationships for Robust Head Pose Estimation With Facial Rotation RepresentationabstractAlthough accurate head pose estimation is critical for natural human-computer interaction, it remains challenging due to occlusion, extreme poses, illumination conditions, and data ambiguity issues. To address these challenges, a novel morphology aware Transformer framework (MHPE) is proposed, which can learn morphological relationships during facial rotation. The methodology is based on two key findings: cross-region geometric dependencies and angle-specific morphodynamic representations. The proposed framework incorporates two key components: adversarial feature generation, which generates robust rotation representations by adaptive multi-scale feature interaction; and morphology relationship inference, which establishes long-range dependencies between facial features through a cross-modal attention mechanism that incorporates morphological priors. Extensive evaluations on three demanding benchmarks (BIWI, AFLW2000, and 300W-LP) demonstrate state-of-the-art performance, particularly in demanding scenarios. The Python implementation will be available on request to facilitate reproducibility. Tingting Liu 0006, Jianping Ju, Zhixiong Song, Shijia Qian, Ning Rao, Hai Liu 0004, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | HPCTrans: Heterogeneous Plumage Cues-Aware Texton Correlation Representation for FBIC via TransformersabstractFine-grained bird image classification (FBIC) for distinguishing bird subspecies is challenging because of several issues, including a camouflaged appearance, body occlusion, and an arbitrary bird posture. To address these challenges, we propose a novel heterogeneous plumage cues-aware texton correlation representation for FBIC, which leverages texton correlation in various functional plumage regions for effective learning. Two key findings are revealed: 1) texton structural discrepancies of heterogeneous plumage; and 2) abstract region information for specific birds. On this basis, this model introduces texton coherence extraction module (TCEM) and abstract representation selection (ARS). Specifically, considering bird characteristics, TCEM is introduced to exploit the spatial statistical properties of local textons in heterogeneous plumage. To the best of our knowledge, this study is the first to introduce heterogeneous plumage cues for mining texton correlation relationship representations in FBIC tasks. In addition, a Multiscale Information Cross-Attention Transformer (MICAformer) is proposed for better modeling texton correlation representation. The experimental results on the CUB-200-2011 dataset and NABirds show the effectiveness of the proposed HPCTrans model over the state-of-the-art methods. Hai Liu 0004, Shuang Zeng, Liqian Deng, Tingting Liu 0006, Xionghua Liu, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Texture Affinity Cue-Aware Relationship Representation via Transformers for Facial Expression Recognition in Affective RobotsabstractAutomatic facial expression recognition (FER) from facial videos is a key component in enabling machines to understand human emotional states, which is crucial for affective robots designed to be interactive companions and applied in smart healthcare. However, FER is susceptible to challenges such as occlusion, arbitrary orientations, and illumination, making it difficult to implement precise FER models in robots. To address these issues, we propose a texture affinity cues-aware relationship representation method (FTATrans), which learns to associate facial texture with facial expressions in videos. The research reveals two key findings: 1) interaction of facial textures, and 2) texture affinity effects. On this basis, FTATrans mainly consists of two key networks: semantic-information feature generation (SFG) and texture-affinity relationship mining (TAR). In particular, the semantic relationships between different facial regions can be learned through SFG. TAR is used to capture the texture affinity relationship and integrate them with the overall facial expression information. Additionally, a loss function focused on expression-specific texture variations is proposed to guide the model in learning discriminative expression information. Experiments conducted on five video-based FER datasets demonstrate that the FTATrans model achieves state-of-the-art performance. Hai Liu 0004, Feifei Li 0003, Tingting Liu 0006, Zhaoli Zhang, Naixue Xiong, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | SkeFormer: Skeletal Cues-Aware Bone Point Relationship Learning for Efficient FBIC via TransformersabstractHow to identify endangered bird species in complex outdoor environments has attracted significant attention in the fields of computer vision and machine learning. Previous studies on fine-grained bird image classification (FBIC) face numerous challenges, such as environmental occlusions and arbitrary postures, which limit the accuracy and robustness of existing methods. To address these challenges and enable more reliable bird species identification in extreme outdoor conditions, we propose a novel skeletal cues-aware bone point relationship learning for efficient FBIC via Transformers (SkeFormer). To the best of our knowledge, this is the first time skeletal relationships have been introduced to the FBIC task. Our model introduces three key modules: the skeletal relationship mining (SRM) module, the multilevel feature generation (MFG) module, and the key feature selection (KFS) module. Specifically, in SRM, the model mines the skeletal relationships among different bird species. In MFG, multiscale information is aggregated by connecting features across multiple layers. The KFS module selects key immutable regions of birds based on the learned skeletal relationships. Extensive experiments on two benchmark datasets, CUB-200-2011 and NABirds, show that SkeFormer outperforms existing state-ofthe- art models. The code for SkeFormer will be publicly available. Hai Liu 0004, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | HomLLM: Exploiting Semantic Homology Relationship for Fine-Grained Bird Image Classification via Large Language ModelsabstractHow to recognize endangered bird species in complex outdoor environments has attracted considerable attention in the fields of computer vision and machine learning. However, fine-grained bird image classification (FBIC) is susceptible to problems such as arbitrary postures, interclass discriminability, and occlusions. We propose a novel semantic homology relationship representation learning for fine-grained bird classification with large language models, namely HomLLM, to address these challenges in FBIC effectively. Our proposed model aims to learn homology relationship representations adaptively by identifying invariant structural correspondences between visual features and semantic descriptions, using limited bird data and base class labels. Our approach yields two key findings: 1) invariant homology in key regions of birds that maintain structural consistency across different postures and 2) homological relationship that establish essential taxonomic markers among similar bird classes. Based on these insights, we propose two new modules of the model: the semantic homology generation (SHG) module and homology relationship mining (HRM) module. Specifically, in SHG, bird features are described at multiple granularities through a large language model (LLM) to establish semantic homology. In HRM, feature adaptation is performed separately for textual and visual information, and cross-modal homological interaction is performed hierarchically. In addition, we propose a hierarchical homology interaction scheme to integrate multilevel homological features while preserving structural consistency. Experiments on the commonly used bird datasets CUB-200-2011 and NABirds demonstrate that HomLLM exhibits better performance than state-of-the-art (SOTA) methods. Hai Liu 0004, Tingting Liu 0006, Lin Chen 0033, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | MDT: A multiscale differencing transformer with sequence feature relationship mining for robust action recognition
Zengzhao Chen, Fumei Ma, Hai Liu 0004, Tingting Liu 0006 |
Appl. Intell. | 5 |
| 2025 | HRHPE: FRoI guides heterogeneous relationship representation learning for precise head pose estimation
Hai Liu 0004, Shijia Qian, Tingting Liu 0006, Zelin Cao, Minhong Wang 0001, Jianping Ju, Zhaoli Zhang |
Neurocomputing | 3 |
| 2025 | ESERNet: Learning spectrogram structure relationship for effective speech emotion recognition with swin transformer in classroom discourse analysis
Tingting Liu 0006, Minghong Wang, Hai Liu 0004, Shaoxin Yi |
Neurocomputing | 1 |
| 2025 | Exploiting module evolution correlation relationship for fine-grained bird image classification with structural functional representation
Shuang Zeng, Hai Liu 0004, Tingting Liu 0006, Qiuxia Liu, Minhong Wang 0001, Zhaoli Zhang |
Neurocomputing | 3 |
| 2025 | TransIFC: Invariant Cues-Aware Feature Concentration Learning for Efficient Fine-Grained Bird Image ClassificationabstractFine-grained bird image classification (FBIC) is not only meaningful for endangered bird observation and protection but also a prevalent task for image classification in multimedia processing and computer vision. However, FBIC suffers from several challenges, such as bird molting, complex background, and arbitrary bird posture. To effectively tackle these challenges, we present a novel invariant cues-aware feature concentration Transformer (TransIFC), which learns invariant and core information in bird images. To this end, two novel modules are proposed to leverage the characteristics of bird images, namely, the hierarchy stage feature aggregation (HSFA) module and the feature in feature abstraction (FFA) module. The HSFA module aggregates the multiscale information of bird images by concatenating multilayer features. The FFA module extracts the invariant cues of birds through feature selection based on discrimination scores. Transformer is employed as the backbone to reveal the long-dependent semantic relationships in bird images. Moreover, abundant visualizations are provided to prove the interpretability of the HSFA and FFA modules in TransIFC. Comprehensive experiments demonstrate that TransIFC can achieve state-of-the-art performance on the CUB-200-2011 dataset (91.0%) and the NABirds dataset (90.9%). Finally, extended experiments have been conducted on the Stanford Cars dataset to suggest the potential of generalizing our method on other fine-grained visual classification tasks. Hai Liu 0004, Cheng Zhang 0020, Yongjian Deng, Bochen Xie, Tingting Liu 0006, Youfu Li 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | LDCNet: Limb Direction Cues-Aware Network for Flexible HPE in Industrial Behavioral Biometrics Systemsabstract2D Human pose estimation (HPE) has been widely used in the many fields such as behavioral understanding, identity authentication, and industrial automatic manufacturing. Most of the previous studies have encountered many constraints, such as restricted scenarios and strict inputs. To solve this problem, we present a simple yet effective HPE network called limb direction cues-aware network (LDCNet) with limb direction cues and differentiated Cauchy labels, which can efficiently suppress uncertainties and prevent deep networks from over-fitting uncertain keypoint positions. In particular, LDCNet suppresses the uncertainties from two aspects. (1) A differentiated Cauchy coordinate encoding method is designed to reveal the limb direction information among adjacent keypoints. (2) Jeffreys divergence is introduced as loss function to measure the prediction heatmap and ground-truth one. Positions of keypoints are perceived at the limb direction based deep network in an end-to-end manner. An extensive study on two benchmark data sets (i.e., MS COCO and MPII) illustrates the superiority of the proposed LDCNet model over state-of-the-art approaches. Tingting Liu 0006, Hai Liu 0004, Zhaoli Zhang |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | MMATrans: Muscle Movement Aware Representation Learning for Facial Expression Recognition via TransformersabstractHow to automatically recognize facial expression has caused concerns in industrial human–robot interaction. However, facial expression recognition (FER) is susceptible to problems, such as occlusion, arbitrary orientations, and illumination. To effectively address these challenges in FER, we present a novel facial muscle movement aware representation learning that can learn the semantic relationships of facial muscle movements in facial expression images. Two key findings are revealed: 1) muscle movements from different facial regions often show semantic relationships; and 2) not all facial muscle regions have equal contributions for different facial expressions. On this basis, this model presents two novel modules, namely, discriminative feature generation (DFG) and muscle relationship mining (MRM). Specifically, in DFG, the memory of our model for mislabeling decreases. In MRM, muscle–motion interaction among diverse facial regions is learned through visual transformers (MMATrans). Experiments on three in-the-wild FER datasets (RAF-DB, FERPlus, and AffectNet) show that our MMATrans yields better performance compared with state-of-the-art methods. Hai Liu 0004, Qiyun Zhou, Cheng Zhang 0020, Junyan Zhu, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | EHPE: Skeleton Cues-Based Gaussian Coordinate Encoding for Efficient Human Pose EstimationabstractHuman pose estimation (HPE) has many wide applications such as multimedia processing, behavior understanding and human-computer interaction. Most previous studies have encountered many constraints, such as restricted scenarios and RGB inputs. To mitigate constraints to estimating the human poses in general scenarios, we present an efficient human pose estimation model (i.e., EHPE) with joint direction cues and Gaussian coordinate encoding. Specifically, we propose an anisotropic Gaussian coordinate coding method to describe the skeleton direction cues among adjacent keypoints. To the best of our knowledge, this is the first time that the skeleton direction cues is introduced to the heatmap encoding in HPE task. Then, a multi-loss function is proposed to constrain the output to prevent the overfitting. The Kullback-Leibler divergence is introduced to measure the predication label and its ground truth one. The performance of EHPE is evaluated on two HPE datasets: MS COCO and MPII. Experimental results demonstrate that EHPE can obtain robust results, and it significantly outperforms existing state-of-the-art HPE methods. Lastly, we extend the experiments on infrared images captured by our research group. The experiments achieved the impressive results regardless of insufficient color and texture information. Hai Liu 0004, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | MSRANet: Learning discriminative embeddings for speaker verification via channel and spatial attention mechanism in alterable scenarios
Qiuyu Zheng, Zengzhao Chen, Hai Liu 0004, Tingting Liu 0006 |
Expert Syst. Appl. | 6 |
| 2023 | Orientation Cues-Aware Facial Relationship Representation for Head Pose Estimation via TransformerabstractHead pose estimation (HPE) is an indispensable upstream task in the fields of human-machine interaction, self-driving, and attention detection. However, practical head pose applications suffer from several challenges, such as severe occlusion, low illumination, and extreme orientations. To address these challenges, we identify three cues from head images, namely, critical minority relationships, neighborhood orientation relationships, and significant facial changes. On the basis of the three cues, two key insights on head poses are revealed: 1) intra-orientation relationship and 2) cross-orientation relationship. To leverage two key insights above, a novel relationship-driven method is proposed based on the Transformer architecture, in which facial and orientation relationships can be learned. Specifically, we design several orientation tokens to explicitly encode basic orientation regions. Besides, a novel token guide multi-loss function is accordingly designed to guide the orientation tokens as they learn the desired regional similarities and relationships. Experimental results on three challenging benchmark HPE datasets show that our proposed TokenHPE achieves state-of-the-art performance. Moreover, qualitative visualizations are provided to verify the effectiveness of the token-learning methodology. Hai Liu 0004, Cheng Zhang 0020, Yongjian Deng, Tingting Liu 0006, Zhaoli Zhang, Youfu Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | Understanding Learner Continuance Intention: A Comparison of Live Video Learning, Pre-Recorded Video Learning and Hybrid Video Learning in COVID-19 PandemicabstractIn response to the COVID-19 pandemic, the emergency policy for online learning at home has been launched in China. Live video, prerecorded video, and their combination are the three main video learning modes. Parental involvement emerges as a unique factor that may influence online learning effects. Students’ behaviors such as interaction, engagement may vary from traditional online learning context due to multiple reasons. Hence, the previous acceptance models are not applicable. In this paper, by modifying and extending the Expectation Confirmation Model, a unified acceptance and continuance intention model (UACIM) is proposed to explore the factors affecting students online learning continuance intention. Confirmation is omitted for the lack of a priori knowledge. Perceived ease of use is included as the factor assessing technology performance. Parental involvement, interaction, and students’ engagement are considered as external variables. The hypotheses are validated by the data collected from 306,139 students from Grade 4 to Grade 9 in Hubei Province, China. The results of structural equation modeling reveal that all variables are significant in explaining continuance intention, which indicates the vital roles of the external factors. Some negative correlations are found with suggestions that practitioners should be prudent when implementing interactions as well as offering individualized instructional designs. Furthermore, three video learning modes are found to be successful. The most preferable way is the hybrid mode with the most interactions, engagement, strongest satisfaction, perceived usefulness, and continuance intention. There are significant differences among three groups for most paths. Therefore, appropriate interventions can be devised to enhance learning effects and continuance intention when utilizing different video learning. Tingting Liu 0006 |
Int. J. Hum. Comput. Interact. | 2 |
| 2022 | ARHPE: Asymmetric Relation-Aware Representation Learning for Head Pose Estimation in Industrial Human-Computer InteractionabstractHead pose estimation (HPE) has wide industrial applications, such as online education, human–robot interaction, and automatic manufacturing. In this article, we address two key problems in HPE based on label learning and asymmetric relation cues: 1) how to bridge the gap between the better prediction performance of networks and incorrectly label pose images in the HPE datasets and 2) how to take full advantage of the adjacent poses information around the centered pose image. We reconstruct all the incorrect labels as a two-dimensional Lorentz distribution to tackle the first problem. Instead of directly adopting the angle values ashardlabels, we assign part of the probability values (softlabels) to adjacent labels for learning discriminative feature representations. To address the second problem, we reveal the asymmetric relation nature of HPE datasets. The yaw direction and pitch direction are assigned different weights by introducing the half at half-maximum of the Lorentz distribution. Compared with the traditional end-to-end frameworks, the proposed one can leverage the asymmetric relation cues for predicting the head pose angle in the incorrect label scenarios. Extensive experiments on two public datasets and our infrared dataset demonstrate that the proposed ARHPE network significantly outperforms other state-of-the-art approaches. Hai Liu 0004, Tingting Liu 0006, Zhaoli Zhang, Arun Kumar Sangaiah, Youfu Li 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | Learning Knowledge Graph Embedding With Heterogeneous Relation Attention NetworksabstractKnowledge graph (KG) embedding aims to study the embedding representation to retain the inherent structure of KGs. Graph neural networks (GNNs), as an effective graph representation technique, have shown impressive performance in learning graph embedding. However, KGs have an intrinsic property of heterogeneity, which contains various types of entities and relations. How to address complex graph data and aggregate multiple types of semantic information simultaneously is a critical issue. In this article, a novel heterogeneous GNNs framework based on attention mechanism is proposed. Specifically, the neighbor features of an entity are first aggregated under each relation-path. Then the importance of different relation-paths is learned through the relation features. Finally, each relation-path-based features with the learned weight values are aggregated to generate the embedding representation. Thus, the proposed method not only aggregates entity features from different semantic aspects but also allocates appropriate weights to them. This method can capture various types of semantic information and selectively aggregate informative features. The experiment results on three real-world KGs demonstrate superior performance when compared with several state-of-the-art methods. Zhifei Li 0009, Hai Liu 0004, Zhaoli Zhang, Tingting Liu 0006, Naixue Xiong |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Recalibration convolutional networks for learning interaction knowledge graph embedding
Zhifei Li 0009, Hai Liu 0004, Zhaoli Zhang, Tingting Liu 0006, Jiangbo Shu |
Neurocomputing | 4 |
| 2021 | NGDNet: Nonuniform Gaussian-label distribution learning for infrared head pose estimation and on-task behavior understanding in the classroom
Tingting Liu 0006 |
Neurocomputing | 1 |
| 2020 | Flexible FTIR Spectral Imaging Enhancement for Industrial Robot Infrared Vision SensingabstractInfrared (IR) spectral imaging sensing is a powerful visual technique for industrial material recognition in robot vision systems. However, the imaging sensing data have issues of random noise and band overlap. Resolution enhancement is usually the first step in the preprocessing procedure of industrial robot vision sensing. In this article, we develop a resolution-enhancement algorithm with total variation (TV) constraints for the degraded Fourier transform IR (FTIR) spectrum due to overlap and noise degradation in the robot vision sensing. The kernel function is calculated using the spectrometer imaging systems and Fourier optical theory. The proposed model not only can remove noises effectively but also can estimate the kernel function because of the adaptive TV as constraint regularization. This model is examined by a set of simulated FTIR spectra with the Poisson noises and a series of real FTIR spectra. The proposed model is compared with the other state-of-the-art methods in terms of performance. Experimental results demonstrate that the proposed approach can split the overlap band effectively while the spectral structure details are retained satisfactorily. The enhanced high-resolution imaging spectrum data can raise the robot vision sensing accuracy in industrial intelligent systems. Tingting Liu 0006, Hai Liu 0004, Youfu Li 0001, Zengzhao Chen, Zhaoli Zhang, Sannyuya Liu |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | DISR: Deep Infrared Spectral Restoration Algorithm for Robot Sensing and Intelligent Visual Tracking SystemsabstractInfrared imaging spectrometer (IRIS) often suffers from overlapped bands and random noises, which limit the precision of subsequent processing in robot vision sensing. To address this problem, we propose a novel Gabor transform-based infrared spectrum restoration method by successfully exploring the intrinsic structure of the clean IR spectrum from the degraded one. At first, a total variation (TV) regularized Gabor coefficients adjustment descriptor is designed and incorporated into the spectrum restoration model. Then, the proposed model is inferred via an efficient optimization approach based on split Bregman iteration method. Comprehensive experiments illustrate the significant and consistent improvements of the developed model over state-of-the-art approaches. The restored high-resolution spectrum can be utilized for detecting the different materials in the robot visual tracking systems. Hai Liu 0004, Youfu Li 0001, Dan Su 0001, Zhaoli Zhang, Sannyuya Liu, Tingting Liu 0006 |
IROS | 6 |
| 2018 | Fast Blind Instrument Function Estimation Method for Industrial Infrared SpectrometersabstractInfrared (IR) spectrometers, particularly the aging ones, often suffer from the band overlap and random noise. In this paper, a blind estimation method based on discrete cosine transform (DCT) regularization is proposed for IR spectrum measured from an aging spectrometer instrument. Motivated by the observation that the DCT coefficient distribution of the ground-truth spectrum is sparser than that of the observed spectrum, an IR spectral deconvolution model is formulated in our method to regularize the distribution of the observed spectrum by total variation regularization. Then, the split Bregman method is exploited to solve the resulting optimization problem. The experimental results demonstrate an encouraging performance of the proposed approach to suppress noise and preserve spectral details. The novelty of our method lies on its ability to estimate instrument function and latent spectrum in a joint framework; thus, mitigating the effects of instrument aging to a large extent. The recovered IR spectra can efficiently capture the spectral features and interpret the unknown chemical mixture in industrial applications. Tingting Liu 0006, Hai Liu 0004, Zengzhao Chen, Alan M. Lesgold |
IEEE Trans. Ind. Informatics | 1 |
| 2016 | Blind image restoration with sparse priori regularization for passive millimeter-wave images
Tingting Liu 0006, Zengzhao Chen, Sanya Liu, Zhaoli Zhang, Jiangbo Shu |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | Destriping algorithm with L0 sparsity prior for remote sensing imagesabstractRemote sensing image often suffers from the common problems of stripe noise and random noise. In this paper, we present a destriping method with unidirectional gradient L0 norm and L0 sparsity priori. The major novelty of the proposed method is that combining the unidirectional gradient L0 norm with the sparsity priori to address the destriping and denoising issues. Moreover, doubly augmented Lagrangian (DAL) method is adopted to solve the L0 regularized minimization problem. The proposed method is verified on heavily striped remote sensing images. Comparative results demonstrate that the proposed method outperforms the-state-of-art methods, which can suppress noise effectively as well as preserve image structures well. Hai Liu 0004, Zhaoli Zhang, Sanya Liu, Tingting Liu 0006, Yi Chang 0002 |
ICIP | 4 |