EDBT 2026 Demo / reviewers in the wild / expert
Lin Zhao 0003
dblp:72/2195-3
· DBLP profile ↗
45ranked-venue papers
8as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 4 first-author · 26 since 2021Artificial intelligence and machine learning · 21 · 3 first-author · 16 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Video Anomaly Detection for Edge Devices via Background Feature CachingabstractVideo anomaly detection (VAD) is essential for intelligent surveillance and public security. While modern VAD models have achieved high accuracy, their heavy computational demands make them difficult to deploy on resource-constrained edge devices, where real-time anomaly detection is most needed. In this paper, we propose an efficient VAD method (BC-VAD) designed for edge devices, which utilizes a cross-frame information fusion mechanism with background feature caching. Specifically, the proposed method utilizes the first frame as a static background prior to construct a Key-Value (KV) cache. By reusing these cached features, we eliminate the need to repeatedly encode background content, which is a major source of inefficiency in traditional VAD methods. Furthermore, we introduce a data augmentation strategy inspired by physical principles to enhance model robustness in real-world applications. We successfully deploy our method on typical edge devices, achieving superior inference speeds with only 1.7M parameters and 0.7 GFLOPs. Meanwhile, our method maintains competitive accuracy on the Avenue and ShanghaiTech benchmarks. Lin Zhao 0003, Wenyan Xing, Di Wang 0011, Chen Gong 0002 |
ICMR | 2 |
| 2026 | Efficient greedy optimization method for k-means
Yuan Yuan 0001, Lin Zhao 0003, Shenfei Pei, Feiping Nie 0001 |
Pattern Recognit. | 2 |
| 2026 | Efficient hierarchical multi-resolution k-means clustering
Lin Zhao 0003, Yuan Yuan 0001, Feiping Nie 0001 |
Pattern Recognit. | 1 |
| 2026 | Generating Imperceptible Perturbations to Attack Human Pose Estimation NetworksabstractAdversarial attacks on deep networks have received significant attention recently. However, most existing research focuses on classification tasks, with limited exploration of adversarial attacks on human pose estimation networks. To bridge this gap, we propose a novel attack framework specifically tailored for human pose estimation by exploiting unique HPE characteristics including heatmap-sensitive localization, joint-influence imbalance, and keypoint-focused perception. Our objective is to substantially diminish object keypoint similarity while introducing minimal perturbations to the image. We have devised a two-stage framework to implement the attack. The first stage involves a gradient attack framework that induces deviation in the adversarial heatmap from the original heatmap by exploiting heatmap-sensitive localization. In the second stage, the perturbations are optimized and restricted to the vicinity of keypoints to make the attacks imperceptible by exploiting joint-influence imbalance and keypoint-focused perception. To achieve this, we incorporate low-frequency constraints to limit the perturbations to high-frequency components and utilize a perceptual color distance metric to control the perturbation's magnitude. Extensive experimental results on the COCO and MPII datasets demonstrate that our attack can generate adversarial examples with high strength and low detectability. Junlong Mu, Lin Zhao 0003, Di Wang 0011, Chen Gong 0002, Nannan Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | FD2-Net: Frequency-Driven Feature Decomposition Network for Infrared-Visible Object DetectionabstractInfrared-visible object detection (IVOD) seeks to harness the complementary information in infrared and visible images, thereby enhancing the performance of detectors in complex environments. However, existing methods often neglect the frequency characteristics of complementary information, such as the abundant high-frequency details in visible images and the valuable low-frequency thermal information in infrared images, thus constraining detection performance. To solve this problem, we introduce a novel Frequency-Driven Feature Decomposition Network for IVOD, called FD2-Net, which effectively captures the unique frequency representations of complementary information across multimodal visual spaces. Specifically, we propose a feature decomposition encoder, wherein the high-frequency unit (HFU) utilizes discrete cosine transform to capture representative high-frequency features, while the low-frequency unit (LFU) employs dynamic receptive fields to model the multi-scale context of diverse objects. Next, we adopt a parameter-free complementary strengths strategy to enhance multimodal features through seamless inter-frequency recoupling. Furthermore, we innovatively design a multimodal reconstruction mechanism that recovers image details lost during feature extraction, further leveraging the complementary information from infrared and visible images to enhance overall representational capacity. Extensive experiments demonstrate that FD2-Net outperforms state-of-the-art (SoTA) models across various IVOD benchmarks, i.e. LLVIP (96.2% mAP), FLIR (82.9% mAP), and M3FD (83.5% mAP). Ke Li 0024, Di Wang 0011, Zhangyuan Hu, Weiping Ni, Lin Zhao 0003, Quan Wang 0006 |
AAAI | 6 |
| 2025 | Pre-training a Density-Aware Pose Transformer for Robust LiDAR-based 3D Human Pose EstimationabstractWith the rapid development of autonomous driving, LiDAR-based 3D Human Pose Estimation (3D HPE) is becoming a research focus. However, due to the noise and sparsity of LiDAR-captured point clouds, robust human pose estimation remains challenging. Most of the existing methods use temporal information, multi-modal fusion, or SMPL optimization to correct biased results. In this work, we try to obtain sufficient information for 3D HPE only by modeling the intrinsic properties of low-quality point clouds. Hence, a simple yet powerful method is proposed, which provides insights both on modeling and augmentation of point clouds. Specifically, we first propose a concise and effective density-aware pose transformer (DAPT) to get stable keypoint representations. By using a set of joint anchors and a carefully designed exchange module, valid information is extracted from point clouds with different densities. Then 1D heatmaps are utilized to represent the precise locations of the keypoints. Secondly, a comprehensive LiDAR human synthesis and augmentation method is proposed to pre-train the model, enabling it to acquire a better human body prior. We increase the diversity of point clouds by randomly sampling human positions and orientations and by simulating occlusions through the addition of laser-level masks. Extensive experiments have been conducted on multiple datasets, including IMU-annotated LidarHuman26M, SLOPER4D, and manually annotated Waymo Open Dataset v2.0 (Waymo), HumanM3. Our method demonstrates SOTA performance in all scenarios. In particular, compared with LPFormer on Waymo, we reduce the average MPJPE by 10.0mm. Compared with PRN on SLOPER4D, we notably reduce the average MPJPE by 20.7mm. Xiaoqi An, Lin Zhao 0003, Chen Gong 0002, Jun Li 0027, Jian Yang 0003 |
AAAI | 2 |
| 2025 | Provable Discriminative Hyperspherical Embedding for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection aims to identify the test examples that do not belong to the distribution of training data. The distance-based methods, which identify OOD examples based on their distances from the centroids of in-distribution (ID) examples, have demonstrated promising OOD detection performance. However, the objectives utilized in prior approaches are typically designed for classification and thus might not yield sufficient discriminative power to distinguish between ID and OOD examples. Therefore, this paper proposes a prototype-based contrastive learning framework for OOD detection, which is termed provable Discriminative Hyperspherical Embedding (DHE). The proposed framework provides a theoretical analysis of inter-class dispersion, which is proved to be fundamental in reducing the false positive rate (FPR) on OOD examples. Based on this, we devise an angular spread loss to achieve the maximal dispersion of the prototypes of different classes prior to training. Subsequently, a prototype-enhanced contrastive loss is introduced to align embeddings of ID examples closely with their corresponding prototypes. In our proposed DHE, the maximal prototype dispersion is theoretically proved, thereby avoiding the pitfalls of local optima commonly encountered by most existing methods. Experimental results demonstrate the effectiveness of our proposed DHE, which showcases a remarkable reduction in FPR95 (i.e., 5.37% on CIFAR-100) and more than doubling the computational efficiency when compared with the state-of-the-art methods. Zhipeng Zou, Sheng Wan, Bo Han 0003, Tongliang Liu, Lin Zhao 0003, Chen Gong 0002 |
AAAI | 6 |
| 2025 | DynPose: Largely Improving the Efficiency of Human Pose Estimation by a Simple Dynamic FrameworkabstractTop-down approaches for human pose estimation (HPE) have reached a high level of sophistication, exemplified by models such as HRNet and ViTPose. Nonetheless, the low efficiency of top-down methods is a recognized issue that has not been sufficiently explored in current research. Our analysis suggests that the primary cause of inefficiency stems from the substantial diversity found in pose samples. On one hand, simple poses can be accurately estimated without requiring the computational resources of larger models. On the other hand, a more prominent issue arises from the abundance of bounding boxes, which remain excessive even after NMS. In this paper, we present a straightforward yet effective dynamic framework called DynPose, designed to match diverse pose samples with the most appropriate models, thereby ensuring optimal performance and high efficiency. Specifically, the framework contains a lightweight router and two pre-trained HPE models: one small and one large. The router is optimized to classify samples and dynamically determine the appropriate inference paths. Extensive experiments demonstrate the effectiveness of the framework. For example, using ResNet-50 and HRNet-W32 as the pretrained models, our DynPose achieves an almost 50% increase in speed over HRNet-W32 while maintaining the same-level accuracy. More importantly, the framework can be generalized to other pre-trained models and datasets without re-training or fine-tuning. Code is available at https://github.com/Aritoria/DynPose. Yalong Xu, Lin Zhao 0003, Chen Gong 0002, Di Wang 0011, Nannan Wang 0001 |
CVPR | 2 |
| 2025 | Efficient Anchor Graph Clustering Through Enhanced Within-Cluster HomogeneityabstractAnchor-based clustering methods have gained attention for their efficiency in subspace, multi-view, and ensemble clustering tasks. Most existing methods focus on using anchors to reduce computational complexity in the original data space. However, clustering directly on anchors, followed by label propagation to the original data, can significantly improve computational efficiency. In this paper, we propose an Efficient Anchor Graph Clustering (EAGC) method that maximizes within-cluster homogeneity among anchors. Inspired by the relaxation and discretization model in spectral clustering, we propose two corresponding models, namely EAGC-R and EAGC-D. EAGC-R first obtains relaxed spectral embedding of anchors and then the embedding is discretized by K-Means. EAGC-D directly solves the discrete anchor membership matrix by coordinate descent method. Once anchor clustering results are obtained, original data labels can be obtained through anchor label transmission. Extensive experiments conducted on both synthetic and real datasets illustrate the effectiveness of proposed methods. Fangyuan Xie, Lin Zhao 0003, Feiping Nie 0001, Weizhong Yu, Xuelong Li 0001 |
ICASSP | 2 |
| 2025 | Prototype-Based Contrastive Learning with Stage-Wise Progressive Augmentation for Self-Supervised Fine-Grained Learning
Baofeng Tan, Xiu-Shen Wei, Lin Zhao 0003 |
ICCV | 3 |
| 2025 | MambaPose: Efficient 2D Human Pose Estimation with Pose-Prior Guided State Space ModelabstractTransformer-based methods have achieved remarkable accuracy in 2D human pose estimation (HPE), exemplified by ViTPose. However, their high computational cost remains a critical limitation. Recently, Mamba, known for its linear complexity, has demonstrated impressive efficiency across various vision tasks. Despite this, there has been limited exploration of Mamba to 2D HPE. We find that incorporating pose priors can notably enhance the modeling capacity of Mamba. Therefore, we propose MambaPose, an efficient and accurate method that introduces a Mamba-based Pose Information Fusion (PIF) module. Specifically, PIF integrates global information by leveraging semantic similarity between tokens, followed by bidirectional cyclic scanning that incorporates pose structural priors to capture geometry-aware local details. Our method achieves 75.0 AP on the COCO validation set at 5.2 GFLOPs, establishing a new benchmark under limited computational resources. Code is available at https://github.com/Aritoria/MambaPose. Yalong Xu, Mengting Jiang, Junlong Mu, Di Wang 0011, Lin Zhao 0003 |
ICME | 6 |
| 2025 | Learning Attribute-Aware Hash Codes for Fine-Grained Image Retrieval via Query OptimizationabstractFine-grained hashing has become a powerful solution for rapid and efficient image retrieval, particularly in scenarios requiring high discrimination between visually similar categories. To enable each hash bit to correspond to specific visual attributes, we propose a novel method that harnesses learnable queries for attribute-aware hash code learning. This method deploys a tailored set of queries to capture and represent nuanced attribute-level information within the hashing process, thereby enhancing both the interpretability and relevance of each hash bit. Building on this query-based optimization framework, we incorporate an auxiliary branch to help alleviate the challenges of complex landscape optimization often encountered with low-bit hash codes. This auxiliary branch models high-order attribute interactions, reinforcing the robustness and specificity of the generated hash codes. Experimental results on benchmark datasets demonstrate that our method generates attribute-aware hash codes and consistently outperforms state-of-the-art techniques in retrieval accuracy and robustness, especially for low-bit hash codes, underscoring its potential in fine-grained image hashing tasks. Peng Wang 0107, Yong Li 0032, Lin Zhao 0003, Xiu-Shen Wei |
ICML | 3 |
| 2025 | A Prior Representation-Guided Method for Low-Resolution Human Pose EstimationabstractHuman pose estimation has achieved significant progress on high-resolution (HR) images, but it experiences severe performance degradation on low-resolution (LR) images. One key reason is that LR images lack sufficient appearance details and fine-grained spatial information. In this paper, we propose a prior representation-guided method (PRG) for low-resolution human pose estimation. Our approach consists of two stages: in the first stage, we design a prior representation extraction network to obtain prior representation from HR images. Then we propose dynamic residual blocks that utilize the extracted prior representation to guide the pose estimation network in focusing on detailed features around joint areas. In the second stage, we use a compact diffusion model with fewer iterations to generate the consistent prior representation from LR images, eliminating the reliance on HR images. Extensive experiments demonstrate that our method achieves significant improvements across various resolutions and backbone networks. In particular, our method improves 16.4 AP compared to the SimCC-Res50 baseline at a resolution of 32×32. Mengting Jiang, Xiaoqi An, Yalong Xu, Di Wang 0011, Lin Zhao 0003 |
ICMR | 6 |
| 2025 | Text-Guided Attribute Enhancement Framework for Composed Image RetrievalabstractComposed image retrieval is a challenging multimodal task that refers to the process of retrieving target image by taking advantage of both complementary and synergistic image and text input. Existing efforts often focus on designing interaction models to fuse global query image and text features. However, these approaches struggle to capture fine-grained semantic association information between query image and text, especially when it comes to identifying specific objects or attributes in the query image that need to be modified in the text. In addition, these methods fail to adequately model cross-modal attention when dealing with composed query and target image, resulting in the model's inability to accurately map the semantic information from the composed query to the corresponding regions in the target image. To address these challenges, we propose a Text-Guided Attribute Enhancement Framework for Composed Image Retrieval (TAE-CIR). Our approach consists of three key modules: (a) Multi-granularity vision aggregation module, which extracts multi-granularity visual features and captures fine-grained object-level features related to the query text, refining object and attribute representations for more precise retrieval; (b) Multi-level fusion interaction module, which facilitates deep cross-modal interactions between the composed query and target image features, effectively capturing complex semantic relationships from the composed query to target image; (c) Composed feature alignment, which fuses multi-granularity visual features with the text using a text-guided Q-Former and contrastive learning to ensure accurate alignment between the composed query and the target image. Our extensive experiments on benchmark datasets FashionIQ and CIRR demonstrate the superiority of our proposed method. Yizi Huang, Di Wang 0011, Bo Wan 0002, Lin Zhao 0003, Quan Wang 0006 |
ICMR | 5 |
| 2025 | CheXPO: Preference Optimization for Chest X-ray VLMs with Counterfactual RationaleabstractVision-language models (VLMs) are prone to hallucinations that critically compromise reliability in medical applications. While preference optimization can mitigate these hallucinations through clinical feedback, its implementation faces challenges such as clinically irrelevant training samples, imbalanced data distributions, and prohibitive expert annotation costs. To address these challenges, we introduce CheXPO, a Chest X-ray Preference Optimization strategy that combines confidence-similarity joint mining with counterfactual rationale. Our approach begins by synthesizing a unified, fine-grained multi-task chest X-ray visual instruction dataset across different question types for supervised fine-tuning (SFT). We then identify hard examples through token-level confidence analysis of SFT failures and use similarity-based retrieval to expand hard examples for balancing preference sample distributions, while synthetic counterfactual rationales provide fine-grained clinical preferences, eliminating the need for additional expert input. Experiments show that CheXPO achieves 8.93% relative performance gain using only 5% of SFT samples, reaching state-of-the-art performance across diverse clinical tasks. 1Code: https://github.com/ResearchGroup-MedVLLM/CheX-Phi35V Di Wang 0011, Lin Zhao 0003, Ronghan Li, Bo Wan 0002, Quan Wang 0006 |
ACM Multimedia | 5 |
| 2025 | Cross-Domain Attribute Alignment with CLIP: A Rehearsal-Free Approach for Class-Incremental Unsupervised Domain AdaptationabstractClass-Incremental Unsupervised Domain Adaptation (CI-UDA) aims to adapt a model from a labeled source domain to an unlabeled target domain, where the sets of potential target classes appearing at different time steps are disjoint and are subsets of the source classes. The key to solving this problem lies in avoiding catastrophic forgetting of knowledge about previous target classes during continuously mitigating the domain shift. Most previous works cumbersomely combine two technical components. On one hand, they need to store and utilize rehearsal target sample from previous time steps to avoid catastrophic forgetting; on the other hand, they perform alignment only between classes shared across domains at each time step. Consequently, the memory will continuously increase and the asymmetric alignment may inevitably result in knowledge forgetting. In this paper, we propose to mine and preserve domain-invariant and class-agnostic knowledge to facilitate the CI-UDA task. Specifically, via using CLIP, we extract the class-agnostic properties which we name as ''attribute''. In our framework, we learn a ''key-value'' pair to represent an attribute, where the key corresponds to the visual prototype and the value is the textual prompt. We maintain two attribute dictionaries, each corresponding to a different domain. Then we perform attribute alignment across domains to mitigate the domain shift, via encouraging visual attention consistency and prediction consistency. Through attribute modeling and cross-domain alignment, we effectively reduce catastrophic knowledge forgetting while mitigating the domain shift, in a rehearsal-free way. Experiments on three CI-UDA benchmarks demonstrate that our method outperforms previous state-of-the-art methods and effectively alleviates catastrophic forgetting. Code is available at https://github.com/RyunMi/VisTA. Kerun Mi, Guoliang Kang, Lin Zhao 0003, Tao Zhou 0002, Chen Gong 0002 |
ACM Multimedia | 4 |
| 2025 | Scene-enhanced multi-scale temporal aware network for video moment retrieval
Di Wang 0011, Yousheng Yu, Haodi Zhong, Lin Zhao 0003 |
Pattern Recognit. | 6 |
| 2025 | A shoreline extraction method based on dual-loop network framework
Xuanpeng Li, Hengshuo Cao, Lin Zhao 0003 |
Vis. Comput. | 5 |
| 2024 | SHaRPose: Sparse High-Resolution Representation for Human Pose EstimationabstractHigh-resolution representation is essential for achieving good performance in human pose estimation models. To obtain such features, existing works utilize high-resolution input images or fine-grained image tokens. However, this dense high-resolution representation brings a significant computational burden. In this paper, we address the following question: "Only sparse human keypoint locations are detected for human pose estimation, is it really necessary to describe the whole image in a dense, high-resolution manner?" Based on dynamic transformer models, we propose a framework that only uses Sparse High-resolution Representations for human Pose estimation (SHaRPose). In detail, SHaRPose consists of two stages. At the coarse stage, the relations between image regions and keypoints are dynamically mined while a coarse estimation is generated. Then, a quality predictor is applied to decide whether the coarse estimation results should be refined. At the fine stage, SHaRPose builds sparse high-resolution representations only on the regions related to the keypoints and provides refined high-precision human pose estimations. Extensive experiments demonstrate the outstanding performance of the proposed method. Specifically, compared to the state-of-the-art method ViTPose, our model SHaRPose-Base achieves 77.4 AP (+0.5 AP) on the COCO validation set and 76.7 AP (+0.5 AP) on the COCO test-dev set, and infers at a speed of 1.4x faster than ViTPose-Base. Code is available at https://github.com/AnxQ/sharpose. Xiaoqi An, Lin Zhao 0003, Chen Gong 0002, Nannan Wang 0001, Di Wang 0011, Jian Yang 0003 |
AAAI | 2 |
| 2024 | Open-World Semi-Supervised Learning under Compound Distribution Shifts
Shijia Xu, Lin Zhao 0003, Jialiang Tang, Chen Gong 0002 |
BMVC | 2 |
| 2024 | TMFN: A Target-oriented Multi-grained Fusion Network for End-to-end Aspect-based Multimodal Sentiment AnalysisabstractEnd-to-end multimodal aspect-based sentiment analysis (MABSA) combines multimodal aspect terms extraction (MATE) with multimodal aspect sentiment classification (MASC), aiming to simultaneously extract aspect words and classify the sentiment polarity of each aspect. However, existing MABSA methods have overlooked two issues: (i) They only focus on fusing image regional information and textual words for two subtasks of MABSA. Whereas, MATE subtask relies more on global image information to assist in obtaining the quantity and attributes of aspects. Ignoring the integration with global information may affect the performance of MABSA methods. (ii) They fail to take advantage of target information. Nevertheless, the fine-grained details of targets are important for classifying sentiments of aspects. To solve these problems, we propose a Target-oriented Multi-grained Fusion Network(TMFN). It fuses text information with global coarse-grained image information for MATE subtask and with fine-grained image information for MASC subtask. In addition, a target-oriented feature alignment (TOFA) module is designed to enhance target-related information in image features with target details. In such a way, image features will contain more target emotional-related information which is beneficial to sentiment classification. Extensive experiments show that our method outperforms state-of-the-art methods on two benchmark datasets. Di Wang 0011, Yuzheng He, Yumin Tian, Lin Zhao 0003 |
LREC/COLING | 6 |
| 2024 | An Asymmetric Augmented Self-Supervised Learning Method for Unsupervised Fine-Grained Image HashingabstractUnsupervised fine-grained image hashing aims to learn compact binary hash codes in unsupervised settings, addressing challenges posed by large-scale datasets and dependence on supervision. In this paper, we first identify a granularity gap between generic and fine-grained datasets for unsupervised hashing methods, highlighting the inadequacy of conventional self-supervised learning for fine-grained visual objects. To bridge this gap, we propose the Asymmetric Augmented Self-Supervised Learning (A2-SSL) method, comprising three modules. The asymmetric augmented SSL module employs suitable augmentation strategies for positive/negative views, preventing fine-grained category confusion inherent in conventional SSL. Part-oriented dense contrastive learning utilizes the Fisher Vector framework to capture and model fine- grained object parts, enhancing unsupervised representations through part-level dense contrastive learning. Self-consistent hash code learning introduces a reconstruction task aligned with the self-consistency principle, guiding the model to emphasize comprehensive features, particularly fine-grained patterns. Experimental results on five benchmark datasets demonstrate the superiority of A2-SSL over existing methods, affirming its efficacy in unsupervised fine-grained image hashing. Feiran Hu, Chen-Lin Zhang, Jiangliang Guo, Xiu-Shen Wei, Lin Zhao 0003, Lingyan Gao |
CVPR | 5 |
| 2024 | A Comprehensive Framework for Occluded Human Pose EstimationabstractOcclusion presents a significant challenge in human pose estimation. The challenges posed by occlusion can be attributed to the following factors: 1) Data: The collection and annotation of occluded human pose samples are relatively challenging. 2) Feature: Occlusion can cause feature confusion due to the high similarity between the target person and interfering individuals. 3) Inference: Robust inference becomes challenging due to the loss of complete body structural information. The existing methods designed for occluded human pose estimation usually focus on addressing only one of these factors. In this paper, we propose a comprehensive framework DAG (Data, Attention, Graph) to address the performance degradation caused by occlusion. Specifically, we introduce the mask joints with instance paste data augmentation technique to simulate occlusion scenarios. Additionally, an Adaptive Discriminative Attention Module (ADAM) is proposed to effectively enhance the features of target individuals. Furthermore, we present the Feature-Guided Multi-Hop GCN (FGMP-GCN) to fully explore the prior knowledge of body structure and improve pose estimation results. Through extensive experiments conducted on three benchmark datasets for occluded human pose estimation, we demonstrate that the proposed method outperforms existing methods. Code and data will be publicly available. Linhao Xu, Lin Zhao 0003, Xinxin Sun, Di Wang 0011, Kedong Yan |
ICASSP | 2 |
| 2024 | Domain Adaptive Pose Estimation Via Multi-level AlignmentabstractDomain adaptive pose estimation aims to enable deep models trained on source domain (synthesized) datasets produce similar results on the target domain (real-world) datasets. The existing methods have made significant progress by conducting image-level or feature-level alignment. However, only aligning at a single level is not sufficient to fully bridge the domain gap and achieve excellent domain adaptive results. In this paper, we propose a multi-level domain adaptation approach, which aligns different domains at the image, feature, and pose levels. Specifically, we first utilize image style transfer to ensure that images from the source and target domains have a similar distribution. Subsequently, at the feature level, we employ adversarial training to make the features from the source and target domains preserve domain-invariant characteristics as much as possible. Finally, at the pose level, a self-supervised approach is utilized to enable the model to learn diverse knowledge, implicitly addressing the domain gap. Experimental results demonstrate that significant improvement can be achieved by the proposed multi-level alignment method in pose estimation, which outperforms previous state-of-the-art in human pose by up to 2.4% and animal pose estimation by up to 3.1% for dogs and 1.4% for sheep. The codes are available at the link. Yugan Chen, Lin Zhao 0003, Yalong Xu, Honglei Zu, Xiaoqi An |
ICME | 2 |
| 2024 | Alignment and Multimodal Reasoning for Remote Sensing Visual Question AnsweringabstractRecently, visual question answering for remote sensing data (RSVQA) has emerged as a prominent research area in the field of remote sensing. Transformer-based approaches have demonstrated impressive results, attributed to their superior performance in jointly modeling visual and textual modalities. However, existing Remote Sensing Visual Question Answering (RSVQA) methods often overlook the modality biases present in visual-language interactions, leading to in-accuracies in answers. To address this issue, we propose a novel Transformer-based approach aimed at mitigating modality biases in RSVQA. Specifically, we introduce a contrastive learning loss to align image and text representations before cross-modal fusion, facilitating foundational learning of visual and language representations. Subsequently, we design a cross-modal decoder to comprehensively understand the correlations between images and text. Notably, in addition to predicting answers to questions, we incorporate an extra head for regression prediction of question types. Experimental results demonstrate that our approach achieves higher accuracy in answer prediction compared to state-of-the-art (SoTA) methods, establishing a new record. Yumin Tian, Di Wang 0011, Ke Li 0024, Lin Zhao 0003 |
IGARSS | 5 |
| 2024 | Fine-grained Semantics-aware Representation Learning for Text-based Person RetrievalabstractText-based person retrieval aims to search for target persons based on a given text description query. However, existing methods often have the following problems: (1) Ignoring local attribute information between different persons in feature learning, which results in the low distinguishability of similar people's feature representations. (2) Lacking fine-grained semantics alignment between visual images and text descriptions, which leads to inconsistency in person details between query and target. To address these issues, we propose a Fine-grained Semantics-aware Representation Learning (FSRL) method that establishing intra-modal local attribute correlations and inter-modal fine-grained semantic correlations. Specifically, we first design an identity self-distillation module, which explores soft identity labels that reflect local attribute similarities among different people. The soft identity labels assist the model in learning discriminative features associated with fine-grained attributes of persons. Secondly, we propose a visual-language relationship modeling module that enforces the model to proofread "error words" randomly changed in text during the cross-modal interaction process to establish fine-grained image-text semantic correlations. Extensive experiments show that the proposed method achieves new state-of-the-art results on three benchmark datasets and also performs well on the domain generalization task. Our code is available at https://github.com/y416f/FSRL. Di Wang 0011, Yifeng Wang 0004, Lin Zhao 0003, Haodi Zhong |
ICMR | 4 |
| 2024 | Long-tailed Object Detection Pretraining: Dynamic Rebalancing Contrastive Learning with Dual ReconstructionabstractPre-training plays a vital role in various vision tasks, such as object recognition and detection. Commonly used pre-training methods, which typically rely on randomized approaches like uniform or Gaussian distributions to initialize model parameters, often fall short when confronted with long-tailed distributions, especially in detection tasks. This is largely due to extreme data imbalance and the issue of simplicity bias. In this paper, we introduce a novel pre-training framework for object detection, called Dynamic Rebalancing Contrastive Learning with Dual Reconstruction (2DRCL). Our method builds on a Holistic-Local Contrastive Learning mechanism, which aligns pre-training with object detection by capturing both global contextual semantics and detailed local patterns. To tackle the imbalance inherent in long-tailed data, we design a dynamic rebalancing strategy that adjusts the sampling of underrepresented instances throughout the pre-training process, ensuring better representation of tail classes. Moreover, Dual Reconstruction addresses simplicity bias by enforcing a reconstruction task aligned with the self-consistency principle, specifically benefiting underrepresented tail classes. Experiments on COCO and LVIS v1.0 datasets demonstrate the effectiveness of our method, particularly in improving the mAP/AP scores for tail classes. Chen-Long Duan, Yong Li 0032, Xiu-Shen Wei, Lin Zhao 0003 |
NeurIPS | 4 |
| 2024 | Global semantic enhancement network for video captioning
Xuemei Luo, Xiaotong Luo, Di Wang 0011, Bo Wan 0002, Lin Zhao 0003 |
Pattern Recognit. | 6 |
| 2024 | Dual-Perspective Fusion Network for Aspect-Based Multimodal Sentiment AnalysisabstractAspect-based multimodal sentiment analysis (ABMSA) is an important sentiment analysis task that analyses aspect-specific sentiment in data with different modalities (usually multimodal data with text and images). Previous works usually ignore the overall sentiment tendency when analyzing the sentiment of each aspect term. However, the overall sentiment tendency is highly correlated with aspect-specific sentiment. In addition, existing methods neglect to explore and make full use of the fine-grained multimodal information closely related to aspect terms. To address these limitations, we propose a dual-perspective fusion network (DPFN) that considers both global and local fine-grained sentiment information in multimodal data. From the global perspective, we use text-image caption pairs to obtain a global representation containing information about the overall sentiment tendencies. From the local fine-grained perspective, we construct two graph structures to explore the fine-grained information in texts and images. Finally, aspect-level sentiment polarities can be obtained by analyzing the combination of global and local fine-grained sentiment information. Experimental results on two multimodal Twitter datasets show that the proposed DPFN model outperforms state-of-the-art methods. Di Wang 0011, Changning Tian, Lin Zhao 0003, Lihuo He, Quan Wang 0006 |
IEEE Trans. Multim. | 4 |
| 2023 | Hierarchical Semantic Structure Preserving Hashing for Cross-Modal RetrievalabstractCross-modal hashing has become a vital technique in cross-modal retrieval due to its fast query speed and low storage cost in recent years. Generally, most of the priors supervised cross-modal hashing methods are flat methods which are designed for non-hierarchical labeled data. They treat different categories independently and ignore the inter-category correlations. In practical applications, many instances are labeled with hierarchical categories. The hierarchical label structure provides rich information among different categories. To rationally take use of category correlations, hierarchical cross-modal hashing is proposed. However, existing methods intend to preserve instance-pairwise or class-pairwise similarities, which cannot fully explore the semantic correlations among different categories and make the learned hash codes less discriminative. In this paper, we propose a deep cross-modal hashing method named hierarchical semantic structure preserving hashing (HSSPH), which directly exploits the label hierarchy information to learn discriminative hash codes. Specifically, HSSPH learns a set of class-wise hash codes for each layer. By augmenting class-wise codes with labels, it generates layer-wise prototype codes which reflect the semantic structure of each layer. In order to enhance the discriminative ability of hash codes, HSSPH supervises the hash codes learning with both labels and semantic structures to preserve the hierarchical semantics. Besides, efficient optimization algorithms are developed to directly learn the discrete hash codes for each instance and each class. Extensive experiments on two benchmark datasets show the superiority of HSSPH over several state-of-the-art methods. Di Wang 0011, Caiping Zhang, Quan Wang 0006, Yumin Tian, Lihuo He, Lin Zhao 0003 |
IEEE Trans. Multim. | 6 |
| 2022 | Towards High Performance One-Stage Human Pose EstimationabstractMaking top-down human pose estimation method present both good performance and high efficiency is appealing. Mask RCNN can largely improve the efficiency by conducting person detection and pose estimation in a single framework, as the features provided by the backbone are able to be shared by the two tasks. However, the performance is not as good as traditional two-stage methods. In this paper, we aim to largely advance the human pose estimation results of Mask-RCNN and still keep the efficiency. Specifically, we make improvements on the whole process of pose estimation, which contains feature extraction and keypoint detection. The part of feature extraction is ensured to get enough and valuable information of pose. Then, we introduce a Global Context Module into the keypoints detection branch to enlarge the receptive field, as it is crucial to successful human pose estimation. On the COCO val2017 set, our model using the ResNet-50 backbone achieves an AP of 68.1, which is 2.6 higher than Mask RCNN (AP of 65.5). Compared to the classic two-stage top-down method SimpleBaseline, our model largely narrows the performance gap (68.1 APkp vs. 68.9 APkp) with a much faster inference speed (77 ms vs. 168 ms), demonstrating the effectiveness of the proposed method. Code is available at: https://github.com/lingl_space/maskrcnn_keypoint_refined. Lin Zhao 0003, Linhao Xu, Jie Xu 0021 |
MMAsia | 2 |
| 2022 | Laplacian Welsch Regularization for Robust Semisupervised LearningabstractSemisupervised learning (SSL) has been widely used in numerous practical applications where the labeled training examples are inadequate while the unlabeled examples are abundant. Due to the scarcity of labeled examples, the performances of the existing SSL methods are often affected by the outliers in the labeled data, leading to the imperfect trained classifier. To enhance the robustness of SSL methods to the outliers, this article proposes a novel SSL algorithm called Laplacian Welsch regularization (LapWR). Specifically, apart from the conventional Laplacian regularizer, we also introduce a bounded, smooth, and nonconvex Welsch loss which can suppress the adverse effect brought by the labeled outliers. To handle the model nonconvexity caused by the Welsch loss, an iterative half-quadratic (HQ) optimization algorithm is adopted in which each subproblem has an ideal closed-form solution. To handle the large datasets, we further propose an accelerated model by utilizing the Nyström method to reduce the computational complexity of LapWR. Theoretically, the generalization bound of LapWR is derived based on analyzing its Rademacher complexity, which suggests that our proposed algorithm is guaranteed to obtain satisfactory performance. By comparing LapWR with the existing representative SSL algorithms on various benchmark and real-world datasets, we experimentally found that LapWR performs robustly to outliers and is able to consistently achieve the top-level results. Jingchen Ke, Chen Gong 0002, Tongliang Liu, Lin Zhao 0003, Jian Yang 0003, Dacheng Tao |
IEEE Trans. Cybern. | 4 |
| 2021 | Tiny Person Pose Estimation via Image and Feature Super Resolution
Jie Xu 0021, Yunan Liu 0001, Lin Zhao 0003, Shanshan Zhang 0001, Jian Yang 0003 |
ICIG (3) | 3 |
| 2021 | Learning to Acquire the Quality of Human Pose EstimationabstractMaking human poses serve high-level computer vision tasks such as action recognition, recognizing the quality of estimated poses is of critical importance. Conventionally, the mean confidence of each keypoint is used as pose quality in most human pose estimation frameworks. However, because different types of keypoint are not identical in visibility and size, they should not contribute equally, which produces biased quality scores. In the paper, we propose end-to-end human pose quality learning, which adds a quality prediction block alongside pose regression. The proposed block learns the object keypoint similarity (OKS) between the estimated pose and its corresponding ground truth by sharing the pose features with heatmap regression. The predicted OKS correlates well with pose quality, making the selection of reliable poses straightforward. Moreover, utilizing the learned quality as pose score improves pose estimation performance during COCO AP evaluation, because it ranks more accurate ones high among all pose detections. We conduct extensive experiments based on the three most popular human pose estimation frameworks, including Hourglass, SimpleBaseline and HRNet. Adding the proposed quality learning block is able to consistently bring nearly 1 percent AP improvement on all the frameworks. Lin Zhao 0003, Jie Xu 0021, Chen Gong 0002, Jian Yang 0003, Wangmeng Zuo, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Estimating Human Pose Efficiently by Parallel Pyramid NetworksabstractGood performance and high efficiency are both critical for estimating human pose in practice. Recent state-of-the-art methods have greatly boosted the pose detection accuracy through deep convolutional neural networks, however, the strong performance is typically achieved without high efficiency. In this paper, we design a novel network architecture for human pose estimation, which aims to strike a fine balance between speed and accuracy. Two essential tasks for successful pose estimation, preserving spatial location and extracting semantic information, are handled separately in the proposed architecture. Semantic knowledge of joint type is obtained through deep and wide sub-networks with low-resolution input, and high-resolution features indicating joint location are processed by shallow and narrow sub-networks. Because accurate semantic analysis mainly asks for adequate depth and width of the network and precise spatial information mostly requests preserving high-resolution features, good results can be produced by fusing the outputs of the sub-networks. Moreover, the computational cost can be considerably reduced comparing with existing networks, since the main part of the proposed network only deals with low-resolution features. We refer to the architecture as "parallel pyramid" network (PPNet), as features of different resolutions are processed at different levels of the hierarchical model. The superiority of our network is empirically demonstrated on two benchmark datasets: the MPII Human Pose dataset and the COCO keypoint detection dataset. PPNet outcompetes all recent methods by using less computation and memory to achieve better human pose estimation results. Lin Zhao 0003, Nannan Wang 0001, Chen Gong 0002, Jian Yang 0003, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Graph-based motion prediction for abnormal action detectionabstractAbnormal action detection is the most noteworthy part of anomaly detection, which tries to identify unusual human behaviors in videos. Previous methods typically utilize future frame prediction to detect frames deviating from the normal scenario. While this strategy enjoys success in the accuracy of anomaly detection, critical information such as the cause and location of the abnormality is unable to be acquired. This paper proposes human motion prediction for abnormal action detection. We employ sequence of human poses to represent human motion, and detect irregular behavior by comparing the predicted pose with the actual pose detected in the frame. Hence the proposed method is able to explain why the action is regarded as irregularity and locate where the anomaly happens. Moreover, pose sequence is robust to noise, complex background and small targets in videos. Since posture information is non-Euclidean data, graph convolutional network is adopted for future pose prediction, which not only leads to greater expressive power but also stronger generalization capability. Lin Zhao 0003, Zhaoliang Yao, Chen Gong 0002, Jian Yang 0003 |
MMAsia | 2 |
| 2020 | Perceiving heavily occluded human poses by assigning unbiased score
Lin Zhao 0003, Jie Xu 0021, Shanshan Zhang 0001, Chen Gong 0002, Jian Yang 0003, Xinbo Gao 0001 |
Inf. Sci. | 1 |
| 2020 | Integrating prediction and reconstruction for anomaly detection
Lin Zhao 0003, Shanshan Zhang 0001, Chen Gong 0002, Jian Yang 0003 |
Pattern Recognit. Lett. | 2 |
| 2020 | LPR-Net: Recognizing Chinese license plate in complex environments
Di Wang 0011, Yumin Tian, Wenhui Geng, Lin Zhao 0003, Chen Gong 0002 |
Pattern Recognit. Lett. | 4 |
| 2020 | Multi-task learning for object keypoints detection and classification
Jie Xu 0021, Lin Zhao 0003, Shanshan Zhang 0001, Chen Gong 0002, Jian Yang 0003 |
Pattern Recognit. Lett. | 2 |
| 2019 | Coarse-to-Fine 3D Human Pose Estimation
Yu Guo 0006, Lin Zhao 0003, Shanshan Zhang 0001, Jian Yang 0003 |
ICIG (3) | 2 |
| 2015 | A deep structure for human pose estimation
Lin Zhao 0003, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001 |
Signal Process. | 1 |
| 2015 | Tracking Human Pose Using Max-Margin Markov ModelsabstractWe present a new method for tracking human pose by employing max-margin Markov models. Representing a human body by part-based models, such as pictorial structure, the problem of pose tracking can be modeled by a discrete Markov random field. Considering max-margin Markov networks provide an efficient way to deal with both structured data and strong generalization guarantees, it is thus natural to learn the model parameters using the max-margin technique. Since tracking human pose needs to couple limbs in adjacent frames, the model will introduce loops and will be intractable for learning and inference. Previous work has resorted to pose estimation methods, which discard temporal information by parsing frames individually. Alternatively, approximate inference strategies have been used, which can overfit to statistics of a particular data set. Thus, the performance and generalization of these methods are limited. In this paper, we approximate the full model by introducing an ensemble of two tree-structured sub-models, Markov networks for spatial parsing and Markov chains for temporal parsing. Both models can be trained jointly using the max-margin technique, and an iterative parsing process is proposed to achieve the ensemble inference. We apply our model on three challengeable data sets, which contains highly varied and articulated poses. Comprehensive experimental results demonstrate the superior performance of our method over the state-of-the-art approaches. Lin Zhao 0003, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2015 | Learning a Tracking and Estimation Integrated Graphical Model for Human Pose TrackingabstractWe investigate the tracking of 2-D human poses in a video stream to determine the spatial configuration of body parts in each frame, but this is not a trivial task because people may wear different kinds of clothing and may move very quickly and unpredictably. The technology of pose estimation is typically applied, but it ignores the temporal context and cannot provide smooth, reliable tracking results. Therefore, we develop a tracking and estimation integrated model (TEIM) to fully exploit temporal information by integrating pose estimation with visual tracking. However, joint parsing of multiple articulated parts over time is difficult, because a full model with edges capturing all pairwise relationships within and between frames is loopy and intractable. In previous models, approximate inference was usually resorted to, but it cannot promise good results and the computational cost is large. We overcome these problems by exploring the idea of divide and conquer, which decomposes the full model into two much simpler tractable submodels. In addition, a novel two-step iteration strategy is proposed to efficiently conquer the joint parsing problem. Algorithmically, we design TEIM very carefully so that: 1) it enables pose estimation and visual tracking to compensate for each other to achieve desirable tracking results; 2) it is able to deal with the problem of tracking loss; and 3) it only needs past information and is capable of tracking online. Experiments are conducted on two public data sets in the wild with ground truth layout annotations, and the experimental results indicate the effectiveness of the proposed TEIM framework. Lin Zhao 0003, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Sparse frontal face image synthesis from an arbitrary profile image
Lin Zhao 0003, Xinbo Gao 0001, Yuan Yuan 0001, Dapeng Tao |
Neurocomputing | 1 |